Prove your backups work
A backup file is not a backup until a restore of it has worked. Levelrail separates the two: last backup says an object was uploaded, last verified restore says a copy of it was restored into a scratch resource and checked. This guide shows how to get the second number green and what to do when it is not.
For targets, schedules and retention basics, see Backups and storage.
What a drill does
A restore drill takes a stored backup and:
- Checks the object still exists and has the size that was uploaded (detects a deleted or truncated object without downloading it).
- Downloads it and compares its SHA-256 with the checksum recorded at backup time.
- Restores it into a scratch resource that is not attached to any app.
- Validates the result.
- Destroys the scratch resource, whether the drill passed or failed.
| Resource | Scratch resource | Validation |
|---|---|---|
| App volume | A new Docker volume named drill-<id> | The restored tree is read back and compared with the archive: file count, bytes, paths, modes, owners, symlink targets and content hashes. It is also compared with the manifest recorded at backup time. |
| Postgres, MySQL, MariaDB | A throwaway container from the database's own image | The dump is replayed with the normal restore command, then tables are counted. A dump with a failing statement fails the drill. |
| Other engines | None | Object check, checksum and the engine's dump trailer or header check. This proves the object is whole, not that it replays. |
| SQLite volume snapshot | None | Checksum and the SQLite file header. |
Drills run on the control plane's own Docker daemon. A volume that lives on another node is restored for the drill from the object store, so the drill still proves the backup itself, not that node.
Run one
Dashboard: open Backups and press Prove restore on a row. CLI:
levelrail-cli backups drill run --backup bkh_abc123
levelrail-cli backups drill list --app web --volume data
levelrail-cli backups drill show bkd_xyz789drill run waits for the result and exits non-zero when the drill fails, so it works in a script or a cron job.
Automatic drills
Every resource that has a succeeded backup is drilled on its newest backup once per APP_BACKUP_DRILL_INTERVAL_HOURS (default 168, one week; 0 turns scheduled drills off). The result history is on the Backups page and in levelrail-cli backups drill list.
A failed drill fires the platform-wide restore_drill_failed alert rule. Create it under Alerts like any other rule; it resolves when the next drill of that resource passes. A resource also shows Failing on the Backups page until then.
Read backup health
The Backups page groups every backed up database and app volume by app:
| Column | Meaning |
|---|---|
| State | Healthy, Needs attention, Failing, Not verified, or No backup. The reason is printed under the badge. |
| Last backup | When the newest succeeded backup started. |
| Last verified restore | When a drill last passed. "Never verified" means you only know a file exists. |
| Size | Total stored size of the resource's succeeded backups. |
| Next run | The next scheduled backup, from the cron schedule. |
Terminal view: levelrail-cli backups health.
Restore
Restores never overwrite by default. Pick Restore on a volume row for the wizard, or:
# A new volume next to the original, on the same node
levelrail-cli backups volumes restore web data --backup bkh_abc123
# A new volume that the app "web-staging" will mount when it deploys
levelrail-cli backups volumes restore web data --backup bkh_abc123 --target-app web-staging
# Onto another node
levelrail-cli backups volumes restore web data --backup bkh_abc123 --node node_2The restore refuses a volume name that already exists on the node. For an in-place restore over the live volume, use app-volume-backups restore, which asks you to type the volume name.
Encrypted backups are decrypted with this control plane's key. To restore somewhere with no control plane, decrypt by hand:
age -d -i backup-encryption.key object.tar.zst.age | zstd -d | tar -x -C ./restoreFor databases, the same Backups row links to the database page, where Restore as new database and the point-in-time restore picker live. Point-in-time restore needs PITR enabled on the database first; see Managing databases.
Encryption and the key
New volume backups are compressed with zstd and encrypted with age on the control plane before upload, so the bucket never sees plaintext. The key is generated on first use and written to backup-encryption.key in APP_DATA_DIR (override the path with APP_BACKUP_ENCRYPTION_KEY_FILE).
WARNING
Copy that file somewhere that is not this server. Lose it and encrypted backups cannot be restored, including after a disaster that destroys the control plane.
Set APP_BACKUP_ENCRYPTION=off to keep compression but skip encryption. Older backups taken before this feature stay readable; each backup records how it was encoded.
Consistency hooks
A running database's files are only consistent if it is quiet. Per volume you can set, from the Policy button or levelrail-cli backups volumes policy set:
- Command before backup: runs inside the app's container. A non-zero exit fails the backup, so an inconsistent volume is never archived.
- Command after backup: runs afterwards, even if the backup failed.
- Pause the app: freezes the container while the archive streams. It resumes when the upload ends.
Retention
Besides the schedule's keep-last and age limits, a volume can keep the newest backup of each of the last N days, weeks and months:
levelrail-cli backups volumes policy set web data --retain-daily 7 --retain-weekly 4 --retain-monthly 6Buckets count periods that actually have a backup, so a week with no backups does not shrink what is kept. The newest backup is never removed by retention. Pruning runs after each successful scheduled backup.
Budgets
| Variable | Default | Effect |
|---|---|---|
APP_BACKUP_MAX_SIZE_MB | unset | A volume archive larger than this fails the backup and is not retried. |
APP_BACKUP_TIMEOUT_MINUTES | unset | One backup, archive plus upload, is cancelled after this long. |
APP_BACKUP_UPLOAD_ATTEMPTS | 1 | How many times a failed upload is retried from a fresh archive. |
APP_BACKUP_DRILL_INTERVAL_HOURS | 168 | Scheduled drill period per resource. 0 disables. |
APP_BACKUP_STALE_RUNNING_MINUTES | 360 | A backup still "running" after this long is settled from the bucket. |
A live volume cannot be re-read from an offset, so a half-uploaded object is aborted and the volume is archived again rather than resumed. If the control plane stops mid-backup, a sweep finishes the row: an object that reached the bucket is checksummed and recorded as succeeded, a missing one is marked failed as interrupted.
Bucket protection
Check buckets (or levelrail-cli backups protection --refresh) probes each backup target for object lock and versioning, and writes and deletes one small canary object to learn whether the key can delete. The result is also refreshed daily.
| Level | Meaning |
|---|---|
| Locked | Object lock is on. A leaked key cannot erase backups inside the retention window. |
| Versioned | Versioning is on but not object lock. Old versions survive an overwrite. |
| Open | Neither. A leaked key or a bad script can delete or overwrite every backup. |
Retention needs delete permission, so the key must be able to delete; object lock is what makes that safe. Backup object names include the backup ID, so a new backup never overwrites an old one.
What was tested
These ran against real Docker volumes and a real S3 compatible server (SeaweedFS), no fakes:
| Check | Result |
|---|---|
| Back up a volume with a uid 1234:5678 directory (750), a 640 file with a 2020 mtime and a symlink, delete the volume, restore into a new volume | Content hashes, owners, modes, mtime and symlink target identical; manifest digest identical |
| Drill a healthy backup | Passes, scratch volume removed |
| Flip one byte of the stored object | Drill fails at restore, scratch volume removed |
| Truncate the stored object | Drill fails at the object check |
| Restore a truncated object over a live volume | Refused, the live volume is untouched |
| Pre-backup hook, post-backup hook and pause | Hooks ran inside the container, container not left paused, a failing pre-hook fails the backup |
| Postgres dump drilled into a scratch Postgres | Restores, two tables counted; a modified dump fails |
Limits: drills for engines other than Postgres, MySQL and MariaDB check the object, not a replay. Object lock and versioning detection depends on the provider answering those API calls; a provider that does not shows as Open.