Levelrail
Skip to content

Self-host vLLM with Docker ​

A fast, high-throughput inference server that exposes an OpenAI-compatible API for open-weight language models.

This page describes the Levelrail template for vLLM (AI). It deploys the services below as one Docker Compose app on your own server.

Project site: https://docs.vllm.ai.

Recommended memory: about 8192 MiB.

This template reserves an NVIDIA GPU, so it needs a node with a GPU and the NVIDIA container runtime.

Services, ports and volumes ​

ServiceImageContainer portsVolumes
vllmvllm/vllm-openai:v0.30.08000
vllm_hf_cache -> /root/.cache/huggingface

Environment variables ​

ServiceVariableValue
vllmHF_HOMEPreset in the template

Passwords and keys marked as generated are created for you when the app is deployed and stored as secrets. Values are not shown here.

Deploy vLLM with Levelrail ​

In the dashboard, open Apps, choose New app, then Browse templates, and select vLLM. Review the Compose body and deploy.

With the CLI:

levelrail-cli templates deploy vllm --name my-vllm

See Service template catalog for how templates work, and Getting started if you have not installed Levelrail yet.

More AI templates ​

All templates are listed in the self-host gallery.

Released under the Apache 2.0 License.