Levelrail
Skip to content

Self-host llama.cpp Server with Docker ​

The llama.cpp HTTP server: run quantized GGUF models on plain CPUs with an OpenAI-compatible API.

This page describes the Levelrail template for llama.cpp Server (AI). It deploys the services below as one Docker Compose app on your own server.

Project site: https://github.com/ggml-org/llama.cpp/tree/master/tools/server.

Recommended memory: about 4096 MiB.

Services, ports and volumes ​

ServiceImageContainer portsVolumes
llama-cppghcr.io/ggml-org/llama.cpp:server8080
llama_cpp_models -> /models

Environment variables ​

ServiceVariableValue
llama-cppLLAMA_CACHEPreset in the template

Passwords and keys marked as generated are created for you when the app is deployed and stored as secrets. Values are not shown here.

Deploy llama.cpp Server with Levelrail ​

In the dashboard, open Apps, choose New app, then Browse templates, and select llama.cpp Server. Review the Compose body and deploy.

With the CLI:

levelrail-cli templates deploy llama-cpp --name my-llama-cpp

See Service template catalog for how templates work, and Getting started if you have not installed Levelrail yet.

More AI templates ​

All templates are listed in the self-host gallery.

Released under the Apache 2.0 License.