Self-host Text Generation Inference with Docker
Hugging Face's production inference server for transformer language models, with continuous batching and streaming.
This page describes the Levelrail template for Text Generation Inference (AI). It deploys the services below as one Docker Compose app on your own server.
Project site: https://huggingface.co/docs/text-generation-inference.Recommended memory: about 8192 MiB.
This template reserves an NVIDIA GPU, so it needs a node with a GPU and the NVIDIA container runtime.
Services, ports and volumes
| Service | Image | Container ports | Volumes |
|---|---|---|---|
tgi | ghcr.io/huggingface/text-generation-inference:3.3.6 | 80 | tgi_data -> /data |
Environment variables
This template sets no environment variables.
Passwords and keys marked as generated are created for you when the app is deployed and stored as secrets. Values are not shown here.
Deploy Text Generation Inference with Levelrail
In the dashboard, open Apps, choose New app, then Browse templates, and select Text Generation Inference. Review the Compose body and deploy.
With the CLI:
levelrail-cli templates deploy text-generation-inference --name my-text-generation-inferenceSee Service template catalog for how templates work, and Getting started if you have not installed Levelrail yet.
More AI templates
All templates are listed in the self-host gallery.