Levelrail
Skip to content

Docker access and the API guard ​

The Docker socket is root equivalent: anything that can talk to it can start a privileged container that mounts / from the host. The control plane and every node agent hold that socket, so a bug or a compromise in either one is, by default, root on the machine.

Levelrail narrows this with an in-process Docker API guard. Each process starts a private Unix socket proxy in front of the real daemon, and every Docker client inside the process dials the proxy instead of the daemon. The proxy forwards only the endpoints listed below, and for container create, update, exec and volume create it reads the request body and refuses configurations that would hand a container the host.

Modes ​

APP_DOCKER_GUARDBehavior
audit (default)Everything is forwarded. A request a rule would deny is logged and written to the audit log as docker_guard.would_deny.
enforceA denied request gets 403 with the rule id and never reaches Docker. Every denial is written to the audit log as docker_guard.denied and logged with the rule id.
offNo proxy. Clients dial the daemon directly, as before the guard existed.

audit is the default so an upgrade never breaks a running deploy. After a week of audit mode with no would-be denials, the attention feed offers the switch to enforce. Change the mode from Settings, Security, Docker API guard, or with the CLI:

levelrail-cli docker-guard status
levelrail-cli docker-guard set --mode enforce

The API is GET and PUT /api/v1/system/docker-guard (PUT needs the root ability). Switching between audit and enforce applies immediately. Switching to or from off takes effect at the next restart; until then off behaves like audit. When APP_DOCKER_GUARD is set on the server it wins and the API returns 409. An unrecognized value fails closed to enforce, and in enforce mode a guard that cannot start stops the process instead of running unguarded.

levelrail-cli doctor reports two checks: docker_guard (the mode and what it flagged in the summary window) and docker_privilege (whether the daemon is rootless or uses userns-remap, and whether the service user is root or in the docker group).

Settings ​

VariableDefaultMeaning
APP_DOCKER_GUARDauditoff, audit or enforce. Overrides the setting saved from the dashboard.
APP_DOCKER_GUARD_SOCKET<data dir>/guard/docker-guard.sockWhere the proxy listens. Its directory must be 0700; the socket is 0600.
APP_DOCKER_GUARD_ALLOW_HOST_NETWORKfalseLets a container declared with host networking through network_host.
APP_DOCKER_GUARD_MAX_BODY_BYTES4194304Largest create, update, exec or volume body the guard reads. Larger is denied.
APP_DOCKER_GUARD_AUDIT_DEDUP1hIn audit mode, one audit row per rule and endpoint per interval. Enforce mode records every denial.
APP_DOCKER_GUARD_SUMMARY_WINDOW168hWindow for the doctor, attention item and status summary.
APP_DOCKER_GUARD_HEADER_TIMEOUT30sTime a client has to send request headers to the proxy.
APP_DOCKER_GUARD_RECORD_QUEUE256Decisions waiting to be written. A full queue drops audit rows, never requests.

The mode chosen in the dashboard is saved to <data dir>/docker-guard.json. A node agent reads APP_DOCKER_GUARD only and writes its decisions to <data dir>/docker-guard-audit.jsonl on that node.

Rules ​

Every rule id appears in the 403 body, the log line and the audit row (ability column).

RuleDenies
endpoint_not_allowedAny method and path not in the allowlist below: plugins, swarm, services, secrets, attach, commit, export, classic /build, prune of containers, volumes or networks.
path_noncanonicalPaths with .. or . segments, encoded / or \ (%2F, %5C), or NUL. Repeated slashes are collapsed and the version prefix (/v1.47/) is kept, so the daemon receives exactly the path that was matched.
body_too_large, body_invalidA body over the limit, or one that does not decode as the endpoint's JSON type.
privilegedHostConfig.Privileged.
host_pid, host_ipc, host_uts, host_userns, host_cgroupnsPidMode, IpcMode, UTSMode, UsernsMode, CgroupnsMode set to host. container:<id> sharing stays allowed.
network_hostNetworkMode: host, unless the create was declared with host networking and APP_DOCKER_GUARD_ALLOW_HOST_NETWORK=true. app.yaml has no host networking field today, so in practice this is always denied.
cap_add_disallowedA CapAdd entry outside the hardening profile: the minimal set from Container hardening plus APP_CONTAINER_HARDENING_CAP_ADD. NET_ADMIN and NET_RAW pass only for a create that declared them (the egress sidecar).
devices_not_grantedDeviceRequests, or Devices other than /dev/nvidia* and /dev/dri/*, on a container whose app did not request a GPU. Also device changes through container update.
device_cgroup_rulesAny DeviceCgroupRules.
security_opt_weakenedseccomp= (any override, including unconfined), apparmor=unconfined, label=disable, label=type:spc_t, systempaths=unconfined, no-new-privileges=false.
masked_paths_overrideAny MaskedPaths or ReadonlyPaths, which replace the default masking of /proc and /sys.
bind_sensitive_pathA bind mount (legacy Binds or Mounts) of a protected host path, or of a path that resolves to one through a symlink, or of an ancestor of one: /, /etc, /root, /boot, /sys, /proc, /dev, /run, /var/run, /var/run/docker.sock, /var/lib/docker, /var/lib/containerd, /lib/modules and the Levelrail data directory. The app volume model refuses these paths too, so no declaration can open them.
bind_undeclaredA bind mount whose host path the app volume model did not declare for this container. Named volumes and tmpfs are always allowed.
mount_type_disallowedA mount type other than volume, bind or tmpfs.
volumes_fromVolumesFrom, which inherits another container's mounts.
volume_bind_driver_optsA volume, created directly or inline in a mount, with the local driver's o=bind (or rbind) option, the trick that turns a named volume into a host bind mount. NFS and CIFS options pass.
exec_privilegedAn exec with Privileged: true.
image_importPOST /images/create?fromSrc=, importing a root filesystem instead of pulling an image.

How declarations work ​

internal/docker's Create builds every app, database and sidecar container from the stored app spec. Right before the create call it registers a declaration for that container name: its bind host paths, whether it asked for a GPU, host networking, and extra capabilities. The guard looks the declaration up by the name query parameter and drops it when Create returns. A create request that did not come through Create, or names a container Create is not creating at that moment, carries no grants. In enforce mode the guard forwards a re-encoded copy of the body it validated, so unknown fields and case or duplicate key tricks never reach the daemon in a form the guard did not see.

The Engine API surface Levelrail uses ​

This is the allowlist. Paths are shown without the version prefix; * is one path segment and ** is one or more (image references contain slashes).

Method and pathUsed byWhy
GET, HEAD /_ping, GET /version, GET /infoclient setup, doctor, GPU and CDI detectionAPI version negotiation, daemon health, runtime, storage root, security options
GET /eventsreconcilerObserved state comes from the event stream, not polling
GET /system/dfstatus, orphaned volume sizesDisk usage
POST /authregistry credential testdocker login check
GET /containers/json, GET /containers/*/jsonreconciler, orphan sweep, browser and one-shot helpersList and inspect
POST /containers/createapp, database, sidecar, one-shot scanner, headless browser, volume chown helperBody validated
POST /containers/*/start, stop, kill, wait, DELETE /containers/*reconciler, helpersLifecycle
POST /containers/*/updateresource changes without recreateBody validated
GET /containers/*/logs, GET /containers/*/statslog store, live tail, metricsLong-lived streams
GET, HEAD, PUT /containers/*/archivevolume ownership check, one-shot file injectionCopy files in or out of a container
POST /containers/*/exec, POST /exec/*/start, POST /exec/*/resize, GET /exec/*/jsonterminal, database tools, backupsExec create body validated; start is an HTTP upgrade (hijacked stream)
GET /images/json, GET /images/**/json, DELETE /images/**, POST /images/**/tag, POST /images/prunerollback pinning, garbage collectionImage bookkeeping
POST /images/createpullPull only, fromSrc denied
POST /images/load, GET /images/get, GET /images/**/getbuilt image load, image moves between nodesStreamed tar archives
GET /distribution/**/jsondigest resolutionRegistry manifest lookup
POST /grpc, POST /session, POST /build/pruneBuildKit inside dockerdHTTP upgrade to h2c for the BuildKit API, and build cache pruning
GET /networks, GET /networks/*, POST /networks/create, POST /networks/*/connect, POST /networks/*/disconnect, DELETE /networks/*per-app networks, egress, bridge gatewayNetworking
GET /volumes, GET /volumes/*, POST /volumes/create, DELETE /volumes/*named volumes, NFS and CIFS sharesVolume create body validated

Streaming and hijacked endpoints (events, logs, stats, exec start, /grpc, /session, image load and save) pass through unbuffered: responses flush as they arrive and upgraded connections are spliced end to end.

What the guard does not protect against ​

  • A process that is already compromised at code execution level. The guard runs inside the same process and the service user can still open the real socket. It stops bugs, injected specs, and a control plane asking an agent for something dangerous; it does not stop an attacker who can run arbitrary code as the service user. For that, run Docker rootless or with userns-remap, which docker_privilege reports.
  • The node agent container. The default agent runs as root with the host Docker socket mounted. Its guard wraps its own Docker client, but anything running as root in that container can bypass it.
  • BuildKit build steps. The guard sees the /grpc upgrade, not the BuildKit requests inside it. Builds run in BuildKit's own sandbox; insecure entitlements stay off unless the daemon itself allows them.
  • Image contents and allowed operations. A pulled image runs with the hardened profile, but the guard does not judge what an image does. Exec into a running container, and copying files into it, stay allowed.
  • Symlinked bind sources on a containerized process. When the process runs in a container it cannot see the host filesystem, so a bind source is checked by path only, not by what it resolves to on the host.
  • Agent denials in the dashboard. An agent's denials are logged on the node and appended to its local audit file; the control plane sees them as the failed operation's error, which names the rule, not as audit rows.

Prior art ​

CandidateLicenceWhat it givesDecision
Tecnativa docker-socket-proxyApache-2.0HAProxy in a container, one env var per API section plus a POST switchNot used: section granularity cannot tell a hardened create from a privileged one, and it is a separate container to run
wollomatic socket-proxyMIT, parts Apache-2.0Go proxy with a per-method regex allowlist and bind mount source restrictionsNot used: same idea as ours, but a separate process configured by regex; we need per-container declarations from the app spec. No code copied
Docker authorization pluginsDocker docsDaemon-side allow or deny hook with the request bodyNot used: needs daemon configuration and a restart on every host, and skips gRPC calls; a good later addition for operators who can configure the daemon
gVisor, SysboxApache-2.0Stronger container sandboxesOut of scope here: they confine containers, not the socket holder. Compatible with the guard
Portainer, Coolify, DokployvariousMount the socket into their own container, or drive Docker over SSH, with full API accessShows the gap: none restrict what their own process can ask Docker to do

Released under the Apache 2.0 License.