Self-host llama.cpp Server with Docker
The llama.cpp HTTP server: run quantized GGUF models on plain CPUs with an OpenAI-compatible API.
This page describes the Levelrail template for llama.cpp Server (AI). It deploys the services below as one Docker Compose app on your own server.
Project site: https://github.com/ggml-org/llama.cpp/tree/master/tools/server.Recommended memory: about 4096 MiB.
Services, ports and volumes
| Service | Image | Container ports | Volumes |
|---|---|---|---|
llama-cpp | ghcr.io/ggml-org/llama.cpp:server | 8080 | llama_cpp_models -> /models |
Environment variables
| Service | Variable | Value |
|---|---|---|
llama-cpp | LLAMA_CACHE | Preset in the template |
Passwords and keys marked as generated are created for you when the app is deployed and stored as secrets. Values are not shown here.
Deploy llama.cpp Server with Levelrail
In the dashboard, open Apps, choose New app, then Browse templates, and select llama.cpp Server. Review the Compose body and deploy.
With the CLI:
levelrail-cli templates deploy llama-cpp --name my-llama-cppSee Service template catalog for how templates work, and Getting started if you have not installed Levelrail yet.
More AI templates
All templates are listed in the self-host gallery.