Observability: metrics, logs, and alerts
This platform keeps metrics and logs on each node's local store, with a federated query layer and alert engine on top. Every app gets CPU, memory, and request metrics automatically, full-text log search with live tail, and alert rules you can route to email, Slack, Discord, or a webhook, with no separate observability stack to install.
For contributors: where this lives in the source
internal/telemetry- storage and queryinternal/alerting- rule evaluation and notificationinternal/api/metrics.go,node_metrics.go,database_metrics.go- metrics handlerslogs.go,live_logs.go,logs_download.go- log handlersapp_resource_usage.go,alerts.go,notification_channels.go- supporting handlers
Why node-local, not a central store
Every other platform in this category (and every hosted observability vendor) centralizes: ship every metric and log line to one place, index it there, query it there. That creates write amplification for every container on every node, plus a central index that keeps growing whether or not anyone queries it.
This platform keeps metrics and logs where they are collected: each node's own SQLite-backed store (the same modernc.org/sqlite used elsewhere in the control plane). The control plane queries agents on demand instead of ingesting continuously.
TIP
Right now there is exactly one node (control plane and agent share a process, see internal/agent's in-memory transport), so "federated query" fans out to a single source. But the query interface (TelemetryQuerier in internal/api/metrics.go) is already shaped for Phase 3's real multi-node federation. Nothing changes when a second node appears, only how many sources the querier asks.
How it actually works
This avoids write amplification to a central index. Each node keeps its own metrics and logs, answering queries on demand.
Metrics collection:
internal/telemetry.Collector samples every running container's Docker stats every 15 seconds (metricsCollectionInterval in cmd/levelrail/main.go). Samples are stored under resource IDs:
service:<name>for an appdatabase:<name>for a managed databasenode:<id>for per-node readings like disk usage
This resource-ID scheme is used across internal/alerting, the log store, and resource-usage ranking. An app's metrics, logs, and alert rules are always found under the one identifier its reconciler uses.
Log collection:
internal/telemetry's log collector reads each container's Docker log stream directly (not the json-file driver's raw files), chunks and compresses it, and indexes it for full-text search. A LogBroadcaster fans out each line to live SSE subscribers at the same time it's written to the store. This makes live tailing and historical search two views over the same pipe, not separate systems.
Retention policy:
Retention uses fixed sweeps, not query-time filtering. The functions runMetricsRetentionSweep and runLogsRetentionSweep (in cmd/levelrail/main.go) delete anything older than the retention window once per hour.
Both windows default to 15 days and are independently overridable:
APP_METRICS_RETENTION(Go duration string, e.g."360h")APP_LOGS_RETENTION(same format)
The sweep interval itself (one hour) is not configurable; only the retention window is.
What gets collected
Per app or database (service:<name> / database:<name>):
The collector currently writes 7 metrics, matching web/src/types/metrics.ts's MetricName:
cpu_percentmemory_usage_bytesmemory_limit_bytesnetwork_rx_bytesnetwork_tx_bytesdisk_read_bytesdisk_write_bytes
Per node (node:<id>):
The endpoint GET /api/v1/nodes/{id}/metrics returns two types of readings:
Sum across placed services for
cpu_percent,memory_usage_bytes,network_rx_bytes,network_tx_bytes,disk_read_bytes,disk_write_bytes. This is a sum of already-collected per-container samples, not true host utilization (sinceinternal/agenthas no host-level stats collection). Note:memory_limit_bytesis excluded from the sum to avoid multiplying the host's approximate total memory by N containers.Real per-node readings for disk usage and OS patch counts (
disk_used_bytes,disk_total_bytes,os_patches_available,os_security_patches_available). These are collected directly byHostDiskCollector/HostPatchCollector.
The response includes resource_count: how many placed services actually contributed a sample (for summed metrics), or 1/0 for whether host data exists (for real per-node metrics).
Request metrics (RED), measured at the ingress with no app changes: see Request metrics below.
Not collected today: container restart count, build duration and deploy frequency are recorded as discrete events or computed client-side from deploy-attempt history, not as continuous collectors.
Request metrics (RED)
Every proxy and static route in the embedded Caddy is wrapped by a small request_stats handler (internal/ingress/request_stats.go). It times the request and counts it into fixed in-memory counters per route host, so idle cost is near zero and nothing is written to log files. Paths, query strings and client addresses are never read, and cardinality is bounded by the number of configured hosts. The sampler drains the counters every 15 seconds, maps hosts to apps through the ingress route table (domain to app), and writes only non-zero values under the app's service:<name> resource id. Platform routes (dashboard, registry, models) are not attributed to an app. Requests rejected by the WAF or rate limiter count as 4xx. Load balancer and proxy failures (502, 503, 504 raised by the reverse proxy) also count under http_upstream_errors.
Metrics (per-tick counter deltas; a tick with no traffic writes nothing):
| Metric | Meaning |
|---|---|
http_requests | Requests finished |
http_responses_2xx, _3xx, _4xx, _5xx | Requests by status class |
http_upstream_errors | Proxy failures (502, 503, 504) returned by the reverse proxy |
http_bytes_in, http_bytes_out | Request content length and response body bytes |
http_latency_bucket_<le> | Latency histogram, upper bounds 5, 10, 25, 50, 100, 250, 500, 1000, 2500, 5000, 10000 ms plus inf |
Percentiles are estimated from the summed histogram buckets (linear interpolation inside a bucket), so they merge correctly across time buckets, rollup tiers and nodes.
Read it with GET /api/v1/apps/{name}/requests?from=&to=&step= (rate, 4xx and 5xx error rate, p50, p95, p99, bytes, upstream errors and a summary), levelrail-cli apps requests <name> (or apps metrics <name> --requests), the get_app_requests MCP tool, or the Metrics tab of an app. GET /api/v1/apps/{name} also carries a requests summary (rate, error rates, p95 over APP_REQUESTS_SUMMARY_WINDOW, default 5m) that later features such as auto-rollback and SLO alerts can consume. In Go, use telemetry.SummarizeRequests(ctx, querier, app, window, now).
Rollups and retention
Raw 15 second samples are rolled up into 1 minute and 1 hour buckets by a background job (telemetry.Maintainer). Each run computes at most APP_METRICS_ROLLUP_MAX_BUCKETS closed buckets per tier and commits them together with a per-tier watermark in one transaction, so an interrupted run resumes where it stopped and recomputing a bucket is idempotent. Counters (http_*) are summed, gauges are averaged. Queries pick a tier by range (raw up to APP_METRICS_RAW_QUERY_MAX_RANGE, 1 minute up to APP_METRICS_MINUTE_QUERY_MAX_RANGE, otherwise 1 hour), move to a coarser tier when the finer one no longer retains the range start, and fill the not yet rolled up tail from the finer tier.
| Variable | Default | Meaning |
|---|---|---|
APP_METRICS_RETENTION | 360h (15 days) | Raw sample retention |
APP_METRICS_RETENTION_1M | 720h (30 days) | 1 minute rollup retention |
APP_METRICS_RETENTION_1H | 8760h (365 days) | 1 hour rollup retention |
APP_METRICS_MAX_DB_BYTES | 1073741824 | Live size cap for telemetry.db, 0 disables |
APP_METRICS_RAW_QUERY_MAX_RANGE | 6h | Widest range served from raw samples |
APP_METRICS_MINUTE_QUERY_MAX_RANGE | 168h | Widest range served from 1 minute rollups |
APP_METRICS_ROLLUP_LAG | 30s | Delay before a bucket is closed |
APP_METRICS_ROLLUP_MAX_BUCKETS | 1440 | Buckets computed per tier per run |
When the database exceeds the size cap, the oldest tenth of the finest tier is deleted, but only rows already covered by the next tier, so raw samples that are not yet rolled up are never dropped for size. The cap measures the whole file, which also holds logs; only metric tiers are pruned, so a database dominated by logs can stay over the cap (a warning is logged).
Dashboard pages
Per-app and per-database metrics:
/apps/$name/metrics-MetricsDashboardshows charts for all 7 collected metrics, with deploy attempts overlaid as colored lines (green succeeded, red failed, gray running)./databases/$name/metrics-DatabaseMetricsDashboardshows the same charts scoped to a managed database.- Node metrics -
NodeMetricsDashboardshows the sum-across-placed-services view plus real disk/patch readings.
App health timeline:
/apps/$name/overview-AppHealthTimelineis a compact 24h/7d strip answering "what happened to this app, and when". It plots deploy markers (green succeeded, red failed, blue running), container restarts (amber, bucketed so a burst reads as one marker with a count), and shaded error windows (a failed deploy from start to finish, or a crashloop of 3 or more restarts with under 15 minutes between them). Markers are keyboard focusable with aria labels and show details on hover or focus; Enter or click opens the deploy's logs (deploys) or the app's logs (restarts). It reads only existing data (GET /api/v1/apps/{name}/deploy-attemptsand thecontainer_restart_countmetric), so there is no new API or CLI surface: the same data is available fromlevelrail apps deploys listand the metrics API. If telemetry is not configured, only deploys are shown.
Logs:
/apps/$name/logs- Two tabs:Live(LiveLogViewer, default) andSearch(LogSearchPanel for historical full-text search). Scoped to the app's running container(s).- Separate from
/apps/$name/deploys/$deployId/logs, which tails a specific deploy attempt's output. /databases/$name/logs- The same live/search pair (LiveDatabaseLogViewer, DatabaseLogSearchPanel) for managed databases./databases/$name/slow-queries-DatabaseSlowQueriesPanelshows the slow query log for Postgres/MySQL databases, sortable by duration or timestamp with a minimum-duration filter.
Overview and alerts:
- Dashboard home -
FleetUtilizationSummaryshows a compact fleet-wide CPU/memory/disk card (backed byGET /api/v1/nodes/resource-usage, polled every 30 seconds), aboveTopResourceConsumers(ranks every app by latest CPU/memory/network reading, backed byGET /api/v1/apps/resource-usage) andFleetResourceChart(a 30-minute rolling history of total CPU and memory usage across all apps, polled every 30 seconds from the same app resource-usage endpoint). All three render nothing when telemetry is unconfigured or no samples exist. - Nodes list (
/nodes) - CPU/Memory/Disk columns read the sameGET /api/v1/nodes/resource-usagesnapshot, keyed by node ID. /apps/$name/alerts-AlertRulesPanellists, creates, edits, and deletes alert rules. Shows each rule's current firing state.- Settings -> Notification channels -
NotificationChannelTableto connect, edit, delete, test, and view delivery history for channels.
Live tailing vs stored search vs download

These are three different reads over the same underlying log store, not separate systems:
Live tail (GET /api/v1/apps/{name}/logs/stream, SSE)
- Opens with a short backfill (last 5 minutes, capped at 200 lines, oldest first).
- Streams every new line as
LogCollectorreceives it from Docker. - The handler subscribes to the live broadcaster before running the backfill query to avoid gaps.
Stored search (GET /api/v1/apps/{name}/logs)
- A request/response query over persisted logs.
- Filtered by
from/to(RFC3339, default last hour) and optionalqfull-text phrase. - Optional
level(minimum level: trace, debug, info, warn, error, fatal) keeps lines at or above it, using the JSONlevel/severityfield or a level token near the start of a plain line; lines with no detectable level are dropped when it is set. Optionallimitkeeps only the newest N matches. - The response carries
total(matches beforelimit) and each entry alevelwhen one was detected. - This is what "why was this app slow at 3am last Tuesday" queries.
Compact query for agents (levelrail logs query <app>, MCP query_logs)
- Filters by level, time window (
--since/--until, a duration or RFC3339), deploy attempt (--deploy, the window from that attempt's start to the next attempt's start) and text. - Returns an excerpt of the newest matching lines, never a full dump: at most
--max-lineslines (default 100) and a byte cap (default 8 KB, set with--max-bytesorAPP_MCP_LOG_MAX_BYTES), plus counts and a notice such asshowing 40 of 1,812 matching lines (newest), use since/until or level to narrow. - Uses the stored search above with
levelandlimit;--jsonprints the excerpt as one object.
Download (GET /api/v1/apps/{name}/logs/download)
- Same
from/to/qfilters as stored search. - Returns plain-text file attachment (not JSON), capped at 5,000 lines.
- Use for pulling copies into support tickets or archives.
- No CLI command today; browser download or
curlwith query params and bearer token.
Database logs (/api/v1/databases/{name}/logs, /api/v1/databases/{name}/logs/stream)
- Mirror app endpoints exactly (same query params, same SSE shape).
- Use resource ID prefix
database:instead ofservice:. - No download endpoint for databases yet, only apps.
Reading lines in the log viewer (client side, no API change)
- Each row gets a subtle level tag (ERR, WRN, INF, DBG) detected from JSON
level/severitykeys,level=WARNstyle key/value pairs, or a leadingERROR/WARN/INFO/DEBUGprefix. - The level chips (All, Errors, Warnings, Info, Debug) filter by detected level. Errors also includes every stderr line, matching the old "Errors only" toggle.
- Click a row to expand it: JSON lines are pretty printed, other lines wrap in full, and Copy line copies the raw text.
- "Jump to first error" scrolls to the first error or stderr line in the current view.
Database slow query log
GET /api/v1/databases/{name}/slow-queries returns structured slow-query entries (timestamp, duration, query text, rows examined where the engine reports it) for Postgres and MySQL. The two engines get their data differently:
- Postgres:
log_min_duration_statementis always set on the container (default 1000ms), so any statement at or above the threshold gets logged asduration: N ms statement: .... This lands in the same container log stream/logsalready reads, so the handler reads it from the log store, no separate collector. No rows-examined count; Postgres doesn't log it here. - MySQL: the built-in slow query log is enabled (
slow-query-log=1,long-query-timefrom the same threshold), writing to a file inside the container's own data directory. MySQL 8's FILE log sink cannot reliably open/dev/stderrfrom inside a container (confirmed against a real container:Could not use /dev/stderr for logging), so there is no working Docker-log-stream source for this engine. The handler execs into the running container instead (the samedocker.Runtime.Execprimitive backups already use to runpg_dump/mysqldump) and reads the file directly, bounded to the last 4,000 lines. - Redis and every other engine (MongoDB, MariaDB, KeyDB, Dragonfly, ClickHouse) return 400. Redis's SLOWLOG is a live-server command against an in-memory ring buffer, not something that appears in log output, so it doesn't fit this shape and isn't planned for this endpoint.
Threshold is control-plane-wide, not per-database: APP_DATABASE_SLOW_QUERY_THRESHOLD_MS (default 1000). Changing it takes effect on the database's next reconcile (container recreate). The MySQL path requires exec support configured on the control plane (same requirement apps/{name}/exec has); without it, MySQL slow-query requests return 501.
External log drains
Use PUT /api/v1/apps/{name}/log-drain to forward an app's container log stream to an external HTTP endpoint or syslog target. The drain is in addition to (never instead of) the node-local store.
A drain taps the same LogBroadcaster a live SSE viewer subscribes to; it does not replace local storage. Configure it from the app's log-drain card in the dashboard or levelrail-cli apps log-drain set.
Clearing a drain (DELETE) stops the external forward only. Historical search and live tail on this control plane are unaffected either way.
Resource-usage ranking
GET /api/v1/apps/resource-usage answers "what is every app doing right now" in one call, avoiding the N+1 pattern for page load.
The response includes:
- Every app that exists, including ones with no telemetry samples yet (freshly deployed apps appear as zero-usage rows, not missing).
- Each field (
cpu_percent,memory_usage_bytes,memory_limit_bytes,network_rx_bytes,network_tx_bytes) present only when a sample has been recorded. - One
LatestByMetriccall per metric, not one query per app.
Batched apps overview
GET /api/v1/apps-metrics returns, in one response, every app the caller can read with its latest CPU and memory, a one hour request rate, 5xx error rate, p95 latency, a 12 point request-rate sparkline (5 minute buckets) and the last deploy time. Pass ?names=a,b to limit it. The dashboard apps list and home tiles read only this endpoint, so a page of apps costs one request instead of two per row. The number of apps included is capped by APP_APPS_METRICS_MAX (default 200). From the CLI: levelrail-cli apps overview [name ...].
Fleet utilization
GET /api/v1/nodes/resource-usage answers "how full are my servers" in one call: the node-scoped counterpart to apps/resource-usage above, read by both the node list's CPU/memory/disk columns and the dashboard's fleet summary card.
The response is { "nodes": [...], "fleet": {...} }:
- Each node row's
cpu_percent/memory_usage_bytesis the sum of every placed service's latest sample, the same "sum of containers, not a true host read" contractapps/{id}/metricsdocuments above.memory_total_bytes/disk_used_bytes/disk_total_bytesare real host reads, but (today) only ever populated for the node running the control plane itself, since no other node has a host-metrics collector yet (see "Missing node metrics" below). A field is absent, not zero, when nothing has reported it. - The
fleetrollup sums every node's numbers.total_cpu_percentis a raw sum, not a percentage of fleet capacity (no node reports its core count anywhere in this codebase).memory_used_percent/disk_used_percentare only computed from nodes that actually reported a capacity figure;nodes_with_memory_capacity/nodes_with_disk_capacitysay how many ofnode_countthat covers, so a 3-node fleet where only 1 node has host metrics reads as "1/3 nodes reporting," never as if the percentage covered the whole fleet.
Alert rules
There are ten rule kinds, all stored in one table (alert_rules). The evaluation loop (internal/alerting.Engine) runs every 30 seconds (alertEvaluationInterval, fixed, not env-configurable).
Each rule tracks its own pending/firing state and notifies only on transitions (firing or resolved), never on every tick a rule stays in the same state. This prevents channels from being trained to ignore repeated alerts.
Ten rule kinds and their configuration
| Kind | Scope | What it watches | Key fields |
|---|---|---|---|
threshold | one app's own metric | latest value of metric vs threshold, debounced by for_duration | metric, comparator (>, <, >=, <=), threshold, for_duration |
crashloop | one app | container restarts within a rolling window | restart_count_threshold, restart_window |
cert_expiry | platform-wide | every stored TLS certificate approaching or past expiry, or stuck mid-renewal | none required |
patch_status | platform-wide | every node's pending OS security patch count | none required (threshold is a control-plane default/env var, not a rule field) |
node_disk_space | platform-wide | every node's disk-used percentage | none required |
node_offline | platform-wide | any node whose status is offline (agent stopped heartbeating); resolves when all are back | none required |
node_cert_expiring | platform-wide | any node's agent certificate inside the warning window (renewal failing) or expired; revoked nodes are left out (see agent certificates) | for_duration (optional, overrides APP_NODE_CERT_EXPIRY_WARNING, default 21 days) |
node_resource_usage | platform-wide | every node's summed placed-container CPU and memory | none required |
scheduled_task_failure | one app's own scheduled task | consecutive failed runs of one task | scheduled_task_id, restart_count_threshold (reused as the failure-count threshold) |
domain_health | one app's own domains | a DNS check gone bad (not resolving, or resolving somewhere else) on any of the app's configured domains | for_duration (optional debounce) |
backup_missing | one database (platform-wide) or one app's own volume | last successful backup trailing its own cron schedule's expected interval by more than a grace period | backup_resource_kind (database or volume), backup_database_name or backup_service_name/backup_volume_name, for_duration (reused as the overdue grace period, default 6h) |
control_plane_backup_stale | platform-wide | the newest control plane self-backup snapshot (see control plane backup) being older than a maximum age; quiet when scheduled backups are disabled or no snapshot exists yet | for_duration (reused as the maximum age, default 3d) |
log_archive_stale | platform-wide | a log archive policy's last run failed, or it has not succeeded within a maximum age (see object storage) | for_duration (reused as the maximum age, default three intervals, at least 2h) |
Platform-wide rule kinds (cert_expiry, patch_status, node_disk_space, node_resource_usage, node_offline, node_cert_expiring, control_plane_backup_stale, log_archive_stale)
These are created through an app's /apps/{name}/alerts URL, but that URL only decides where the rule appears in that app's list. The rule evaluates every certificate, node, or disk across the entire control plane regardless of which app created it.
Thresholds default sensibly and are overridable per control plane (not per rule):
- Cert expiry: 14-day warning window
- Patch status: 1 pending security patch
- Disk space: 90% used
- Node CPU: 80%
- Node memory: 4 GiB
Override via env vars: APP_ALERT_PATCH_STATUS_THRESHOLD, APP_ALERT_NODE_DISK_SPACE_THRESHOLD_PERCENT, APP_ALERT_NODE_CPU_THRESHOLD_PERCENT, APP_ALERT_NODE_MEMORY_THRESHOLD_BYTES, APP_ALERT_DOMAIN_HEALTH_CHECK_INTERVAL.
Backup missing rule logic
The backup_missing rule doesn't invent its own cadence calculation. It reads the same backup_schedule cron expression and backup_history rows that internal/backup.Scheduler uses. It computes the expected interval from the cron expression (internal/cronexpr) and fires once the last succeeded attempt is older than that interval plus a grace period.
Edge cases:
- A target with attempts but no success in the lookback fires, anchored to its oldest attempt (a silently-failing backup reads the same as one that stopped).
- A target with no schedule or no history yet stays quiet on day one.
Grace period defaults to 6h (DefaultBackupMissingGracePeriod). Override control-plane-wide via APP_ALERT_BACKUP_MISSING_GRACE_PERIOD, or per-rule via for_duration for tighter/looser windows on specific databases or volumes.
Crashloop alert attachments
A firing crashloop rule attaches the last 200 lines of the crashlooping container's logs (from the last 15 minutes; see crashloopLogLines/crashloopLogLookback in internal/alerting/engine.go) to the notification webhook.
This is useful context for recipients, but there's no API endpoint that reconstructs those exact lines after the fact. AlertRulesPanel links to the app's live/historical log view instead of trying to retrieve them.
Auto-rollback on crashloop
Crashloop detection alerts you; it doesn't fix anything by default. A failed deploy keeps retrying the same bad image forever until you intervene, which is correct (level-triggered, never edge-triggered) but has no escape hatch on its own.
Auto-rollback is the opt-in escape hatch, per app, off by default:
levelrail-cli apps auto-rollback enable <app>
levelrail-cli apps auto-rollback status <app>
levelrail-cli apps auto-rollback disable <app>Or from the dashboard: the app's Deploys tab has an "Auto-rollback on crashloop" toggle above the deploy history list.
Once enabled, the first time a crashloop rule for that app transitions to firing, the control plane automatically redeploys the most recent successful deploy attempt's image that differs from the current (crashlooping) one, through the exact same trigger path the dashboard's "Rollback to this build" button and apps rollback/apps deploy already use (POST /api/v1/apps/{name}/deploys). The resulting deploy attempt shows up in deploy history with source auto_rollback.
Guardrails:
- Fires at most once per crashloop episode: it only runs on the rule's pending-to-firing transition, the same transition the notification itself fires on, so a rule that stays firing across several evaluation ticks doesn't trigger a second rollback.
- If there's no older successful image to fall back to (the app has never deployed before, or has already been rolled back to its oldest recorded image), auto-rollback does nothing and leaves the crashloop to the alert notification alone, rather than rolling back to nothing.
- Independent of the
crashloopalert rule's own notification, which still fires either way.
Auto-rollback on SLO burn
The same escape hatch as crashloop auto-rollback above, for a slo_burn rule instead, with three modes rather than a plain on/off:
levelrail-cli apps auto-rollback-slo-burn set <app> auto # roll back immediately, same as crashloop
levelrail-cli apps auto-rollback-slo-burn set <app> dry_run # log what would have rolled back, never deploy
levelrail-cli apps auto-rollback-slo-burn set <app> pause_for_human # open a pending deploy approval instead of deploying directly
levelrail-cli apps auto-rollback-slo-burn set <app> off # default
levelrail-cli apps auto-rollback-slo-burn status <app>Or from the dashboard: the app's Deploys tab has an "Auto-rollback on SLO burn" mode selector next to the crashloop toggle.
auto mode rolls back through the exact same TriggerImageDeploy path crashloop auto-rollback uses, on the rule's pending-to-firing transition (and again if a genuinely new bad deploy lands while the rule stays continuously firing). dry_run never deploys; it records a slo_burn_would_rollback event in alert history so you can see what would have happened before turning auto on. pause_for_human opens a pending entry in GET /api/v1/deploy-approvals (the same queue a deploy or promote into a protected environment uses) instead of deploying directly, so a teammate approves or rejects the rollback; it shows up on the app's Deploys tab and in the cross-app /approvals page like any other pending approval.
Same guardrails as crashloop auto-rollback: fires at most once per firing episode, and does nothing (leaving the alert to notify alone) when there's no older successful image to fall back to.
Silences, maintenance windows, and noise control
Everything in this section sits between rule evaluation and notification. A rule keeps evaluating, keeps its own firing state, and is always recorded in alert history; these controls only decide whether a notification goes out.
for_duration and consecutive failures. for_duration (threshold, domain health, and other debounced kinds) keeps a rule pending until its condition has held that long. Consecutive failures is a second, tick-based hold: the condition must be true on N evaluation ticks in a row (30 seconds apart) before the rule fires. Set it per rule with consecutive_failures, or for every rule with APP_ALERT_CONSECUTIVE_FAILURES (default 1, meaning off). The streak counter lives in memory, so a control plane restart only delays a rule's first firing.
Silences. A silence has matchers (rule IDs, apps, nodes, rule kinds, severities, labels), a start and end, a creator and a reason. Every matcher you give must match; a list inside one matcher matches any of its entries. A silence needs at least one matcher, so it can never mute everything by accident. Silences last at most 90 days, and one that ends early (End now, alerts silences delete) is kept as history, as are expired ones.
- App matchers apply to app-scoped rules (threshold, crashloop, scheduled task, domain health, backup missing). Platform-wide rules such as certificate expiry have no app and are matched by rule ID, kind, severity or label.
- Node matchers apply to apps placed on that node (by node name or ID).
- Rules carry an optional
severity(info,warning, defaultwarning, orcritical) andlabelsfor silences to match on. - If a rule fires while silenced and is still firing when the silence ends, the held notification is sent then, so you are not left unaware of a live problem. If it resolves while silenced, no "resolved" message is sent for a firing you never heard about. Both cases are recorded in history.
Quick silence. Silence one rule for 1h, 4h or 24h from the rule's row, from the dashboard's Recent alerts card, with levelrail-cli alerts silence <app> <rule-id> --for 4h, or with the silence_alert_rule MCP tool.
Maintenance windows. A recurring silence: a 5-field cron expression for each start, a duration, and an IANA timezone, applied to all alerts, a set of apps, or a set of nodes. The cron is evaluated on the wall clock in that timezone, so "03:00 Europe/Berlin" stays at 03:00 across daylight saving changes and the window keeps its nominal length. Windows are listed with whether they are active now and their next start.
Node-down inhibition. While a node is offline, alerts of apps placed on it are held (recorded as inhibited), so a dead node produces one node-offline alert instead of one per app. Held alerts are released if the app is still firing after the node returns. Platform-wide rules, including the node-offline rule itself, are never inhibited.
Flapping. A rule that fires more than flap_threshold times inside flap_window (defaults APP_ALERT_FLAP_THRESHOLD=5, APP_ALERT_FLAP_WINDOW=30m; 0 disables) is marked flapping. It notifies once with a summary, further fires and resolves are held (recorded as flapping), and one "stable" message is sent when its fires drop to half the threshold. Per-rule overrides: flap_threshold, flap_window.
Grouping and deduplication. With APP_ALERT_GROUP_WINDOW set (for example 2m; default off), firing alerts for the same delivery target and app are buffered and sent as one message listing every alert with a count. A rule that fires twice in the window counts once, and a fire that resolves before the window closes sends nothing. Buffered alerts are flushed on shutdown.
Per-channel rate limit. At most APP_ALERT_CHANNEL_RATE_LIMIT notifications (default 30; 0 disables) per APP_ALERT_CHANNEL_RATE_WINDOW (default 10m) go to one channel. Excess notifications are recorded as ratelimited. Deploy notifications are not counted.
All noise state except silences, windows and history is in memory. A control plane restart forgets pending groups, streaks, flap counters and held notifications; rule firing state itself is persisted, so a still-firing rule does not re-notify after a restart.
Environment variables:
| Variable | Default | Meaning |
|---|---|---|
APP_ALERT_CONSECUTIVE_FAILURES | 1 | ticks a condition must hold before firing (per rule: consecutive_failures) |
APP_ALERT_FLAP_THRESHOLD | 5 | fires within the window that mark a rule flapping (per rule: flap_threshold) |
APP_ALERT_FLAP_WINDOW | 30m | flapping window (per rule: flap_window) |
APP_ALERT_GROUP_WINDOW | off | how long firing alerts are buffered into one message |
APP_ALERT_CHANNEL_RATE_LIMIT | 30 | notifications per channel per rate window |
APP_ALERT_CHANNEL_RATE_WINDOW | 10m | rate limit window |
APP_ALERT_HISTORY_RETENTION | 30d | how long alert history is kept |
Alert history
Every state change and notification decision is recorded: the rule, app, node, severity, the event (fired, resolved, flapping, flap_ended, plus slo_burn_would_rollback for auto-rollback on SLO burn's dry_run mode, see above), and the outcome.
| Outcome | Meaning |
|---|---|
sent | the notification was delivered |
failed | delivery failed; the error is stored |
silenced | a silence or maintenance window matched (silence_id names it) |
inhibited | the app's node was offline |
grouped | folded into one grouped message, or resolved before its group was sent |
ratelimited | the channel rate limit was reached |
flapping | held because the rule is flapping |
skipped | the rule or its channel is disabled, or (on a slo_burn_would_rollback event) dry_run mode logging what it would have done |
Shown per app on /apps/{name}/alerts, globally on /alerts (with outcome and event filters), and available from the CLI (levelrail-cli alerts history), the API (GET /api/v1/alert-history) and the list_alert_history MCP tool. Entries are pruned after APP_ALERT_HISTORY_RETENTION. Who created, changed or removed a silence, window or rule is recorded separately in the generic audit log (GET /api/v1/audit-log), which covers every write request.
What changed before an alert
Every firing app alert, its history entry, diagnose and the dashboard alert rows answer "what changed on this app in the last 30 minutes": deploys (with digest and who), rollbacks, config, domain, scaling and load balancer changes, freeze and maintenance events, and env or secret key names (never values). They come from the app event log, deploy attempts and the audit log, merged newest first by internal/changes.
The most recent change that took effect before the alert (a deploy, rollback, config, env, secret, domain, scaling or load balancer change) is tagged Likely cause. It is a heuristic: it picks the nearest preceding change, not a proven culprit, and never picks a restart, freeze or maintenance event or a failed deploy.
- Notifications (webhook, Slack, Discord, Telegram, email and the rest) carry the top 5 changes, an "and N more" line and a link to the app's alerts page. The generic webhook adds
recent_changesandchanges_linkfields. Discord and PagerDuty messages are capped to their receiver limits with the changes section kept. GET /api/v1/apps/{name}/changes?until=&window=returns the list;GET /api/v1/alert-history?include=changesattaches it to fired entries;GET /api/v1/apps/{name}/diagnoseincludes it asrecent_changes.- CLI:
levelrail-cli apps diagnose <app>andlevelrail-cli alerts history --changes. MCP:diagnose_app_failure, andlist_alert_historywithinclude_changes. - Dashboard: expand a fired row in the alert history or the dashboard's recent alerts card; firing rules on an app's alerts page show the block inline.
| Env var | Default | Meaning |
|---|---|---|
APP_ALERT_CHANGE_WINDOW | 30m | how far back changes are collected |
APP_ALERT_CHANGE_MAX | 20 | cap on entries kept per alert |
APP_DASHBOARD_URL | unset | base URL for notification links; falls back to http://<APP_PUBLIC_HOST>:<port>, and with neither set no link is added |
SLO burn-rate alerts
A slo_burn rule watches a request-based SLO over the app's ingress request metrics: availability (requests without a 5xx) or latency (requests under a limit, rounded down to the nearest histogram bound), with a target such as 99.9 over a 30 day budget window. It uses the multiwindow multi-burn-rate method: a tier fires only when the burn rate is at or above its factor over both its long and its short window, so a brief blip does not page and a real outage does. Burn rate is how many times faster than sustainable the error budget is being spent.
| Tier | Factor | Windows | Severity |
|---|---|---|---|
page_fast | 14.4x | 1h and 5m | critical (page) |
page_slow | 6x | 6h and 30m | critical (page) |
ticket_fast | 3x | 1d and 2h | warning (ticket) |
ticket_slow | 1x | 3d and 6h | warning (ticket) |
An app with no traffic, or fewer than APP_SLO_MIN_REQUESTS requests in a tier's long window, never fires. An app newer than a window is judged on the traffic it has. Hysteresis, silences and maintenance windows apply as for any rule.
Override on the control plane with APP_SLO_<TIER>_FACTOR, APP_SLO_<TIER>_LONG and APP_SLO_<TIER>_SHORT (tier is PAGE_FAST, PAGE_SLOW, TICKET_FAST or TICKET_SLOW), APP_SLO_WINDOW (default 720h), APP_SLO_MIN_REQUESTS (default 10) and APP_SLO_EVAL_INTERVAL (default 1m).
- Create with
levelrail-cli apps alerts create <app> --name NAME --kind slo_burn --slo 99.9(add--slo-latency-ms 300for a latency SLO), or from the Create rule dialog, which previews the error budget remaining and the current burn rates live. levelrail-cli apps alerts slo <app> [--slo 99.9],GET /api/v1/apps/{name}/slo-previewand theget_slo_statusMCP tool show the budget and burn rates without creating a rule.- An app with traffic and no SLO rule gets a suggestion on its alerts page to create a default 99.9% availability SLO.
Notification channels
Channels are global, connect-once destinations (Settings -> Notification channels). Attach them to alert rules by channel_id instead of retyping webhook URLs per rule.
Supported kinds (17 total, map to payload builders in internal/alerting/notify.go): generic, slack, discord, telegram, email, pushover, pagerduty, teams, resend, ntfy, gotify, mattermost, lark, rocketchat, opsgenie, webex, googlechat
For most kinds, notify_url is a webhook URL. A few pack multiple credentials into that one field (e.g., Pushover's user key and app token; PagerDuty's routing key). email is the exception: it sends through the control plane's SMTP sender (Settings -> Email, or env vars APP_SMTP_HOST/APP_SMTP_PORT/APP_SMTP_USERNAME/APP_SMTP_PASSWORD/APP_SMTP_FROM). Returns "email is not configured" if neither path is set up.
Retries
Every kind, including email, retries up to 3 times on a transient failure with the same 500ms-then-1s backoff, just via two different send paths since the transports aren't alike:
HTTP-based kinds (everything except email) use postJSONWithAuth, retrying on:
- Transport errors (DNS, TLS, connection refused, timeout)
- 5xx or 429 responses
Any other status (malformed payload, bad credential, 404'd URL) fails on the first attempt. Retrying inherently-wrong requests only delays surfacing the real problem.
email uses sendEmailWithRetry, retrying on:
- Transport-level failures (DNS, dial, TLS handshake, a client-side timeout, or any non-SMTP backend such as SES)
- SMTP 4xx replies (a transient negative completion per RFC 5321 S4.2.1, e.g. "mailbox busy" or "service not available")
An SMTP 5xx reply (bad recipient, bad auth, policy rejection) fails on the first attempt, the SMTP analogue of an HTTP 4xx: permanent, so retrying it only delays surfacing the real problem.
Test-send
Send a test message through the channel's kind and URL:
POST /api/v1/notification-channels/test- before saving (kind and notify_url in body)POST /api/v1/notification-channels/{id}/test- against an existing channel
Both run synchronously with a 10-second timeout to prevent unresponsive targets from hanging the request. Only the existing-channel variant records a delivery-history row.
Delivery history
GET /api/v1/notification-channels/{id}/deliveries lists every recorded send for a channel, newest first. It captures test sends plus real deploy-outcome and alert-rule dispatches (recorded directly from internal/alerting).
Cursor pagination via ?before (RFC3339 timestamp). Default 50 rows, capped at 200.
Deleting a channel still attached to a rule or deploy-notify target succeeds. The foreign key's ON DELETE SET NULL clears the reference instead of failing (unlike deleting a backup target).
Integration walkthrough
Query one app's CPU over the last hour, bucketed into 5-minute averages:
bashcurl -s -H "Authorization: Bearer $TOKEN" \ "https://your-control-plane/api/v1/apps/my-app/metrics?metric=cpu_percent&step=5m"bashlevelrail-cli apps metrics my-app --metric cpu_percent --since 1h --step 5mResponse:
json{ "metric": "cpu_percent", "points": [ { "timestamp": "2026-09-12T09:00:00Z", "value": 4.2, "count": 20 }, { "timestamp": "2026-09-12T09:05:00Z", "value": 5.1, "count": 20 } ] }Search that app's logs for an error in the last day:
bashcurl -s -H "Authorization: Bearer $TOKEN" \ "https://your-control-plane/api/v1/apps/my-app/logs?from=2026-09-11T00:00:00Z&q=panic"bashlevelrail-cli apps logs my-app --since 24h --q panicConnect a Slack channel and test it:
bashlevelrail-cli channels create --name "on-call" --kind slack --notify-url https://hooks.slack.com/services/... levelrail-cli channels test <id>Create a threshold alert on that app, notifying through the new channel:
bashlevelrail-cli apps alerts create my-app --name "high CPU" --kind threshold \ --metric cpu_percent --comparator ">" --threshold 90 --for-duration 5m \ --channel-id <channel-id>Watch it fire:
levelrail-cli apps alerts list my-appshowsFIRING=trueonce the condition holds for 5 minutes; a delivery row shows up underlevelrail-cli channels deliveries <channel-id>.
API reference
| Method | Path | Ability |
|---|---|---|
GET | /api/v1/apps/{name}/metrics?metric=...&from=...&to=...&step=... | read |
GET | /api/v1/databases/{name}/metrics | read |
GET | /api/v1/nodes/{id}/metrics | root |
GET | /api/v1/apps/resource-usage | read |
GET | /api/v1/apps-metrics | read |
GET | /api/v1/nodes/resource-usage | root |
GET | /api/v1/apps/{name}/logs?from=...&to=...&q=... | read |
GET | /api/v1/apps/{name}/logs/stream (SSE) | read |
GET | /api/v1/apps/{name}/logs/download | read |
GET | /api/v1/databases/{name}/logs | read |
GET | /api/v1/databases/{name}/logs/stream (SSE) | read |
GET | /api/v1/databases/{name}/slow-queries?from=...&to=...&limit=...&offset=... | read |
GET | /api/v1/apps/{name}/log-drain | read |
PUT | /api/v1/apps/{name}/log-drain | write (sensitive) |
DELETE | /api/v1/apps/{name}/log-drain | write (sensitive) |
POST | /api/v1/apps/{name}/alerts | write |
GET | /api/v1/apps/{name}/alerts | read |
PUT | /api/v1/apps/{name}/alerts/{id} | write |
DELETE | /api/v1/apps/{name}/alerts/{id} | write |
POST | /api/v1/apps/{name}/alerts/{id}/silence | write |
GET | /api/v1/apps/{name}/alert-history | read |
GET | /api/v1/alert-silences?include_expired=true | read |
POST | /api/v1/alert-silences | write |
DELETE | /api/v1/alert-silences/{id} (ends the silence now, keeps it in history) | write |
GET | /api/v1/alert-maintenance-windows | read |
POST | /api/v1/alert-maintenance-windows | write |
PUT | /api/v1/alert-maintenance-windows/{id} | write |
DELETE | /api/v1/alert-maintenance-windows/{id} | write |
GET | /api/v1/alert-history?app=...&rule_id=...&outcome=...&event=...&since=...&limit=... | read |
GET | /api/v1/notification-channels | read |
POST | /api/v1/notification-channels | write |
PUT | /api/v1/notification-channels/{id} | write |
DELETE | /api/v1/notification-channels/{id} | write |
POST | /api/v1/notification-channels/test | write |
POST | /api/v1/notification-channels/{id}/test | write |
GET | /api/v1/notification-channels/{id}/deliveries?limit=...&before=... | read |
CLI
levelrail-cli apps metrics <name> --metric NAME [--since 1h | --from RFC3339 --to RFC3339] [--step 60s]
levelrail-cli databases metrics <name> --metric NAME [--since 1h | --from ... --to ...] [--step 60s]
levelrail-cli nodes metrics <id> --metric NAME [--since 1h | --from ... --to ...] [--step 60s]
levelrail-cli apps resource-usage
levelrail-cli nodes resource-usage
levelrail-cli apps logs <name> [--since 1h | --from ... --to ...] [--q PHRASE] [--tail N]
levelrail-cli apps logs <name> --follow
levelrail-cli databases slow-queries <name> [--since 1h | --from ... --to ...] [--limit N] [--offset N]
levelrail-cli apps log-drain get <name>
levelrail-cli apps log-drain set <name> --type http|syslog --target TARGET [--disabled]
levelrail-cli apps log-drain clear <name>
levelrail-cli apps alerts list <app>
levelrail-cli apps alerts create <app> --name NAME --kind threshold --metric METRIC --comparator OP --threshold N [--for-duration 2m] [--channel-id ID]
levelrail-cli apps alerts create <app> --name NAME --kind crashloop --restart-count-threshold N --restart-window DURATION
levelrail-cli apps alerts create <app> --name NAME --kind cert_expiry
levelrail-cli apps alerts create <app> --name NAME --kind patch_status
levelrail-cli apps alerts create <app> --name NAME --kind node_disk_space
levelrail-cli apps alerts create <app> --name NAME --kind node_resource_usage
levelrail-cli apps alerts create <app> --name NAME --kind node_offline
levelrail-cli apps alerts create <app> --name NAME --kind scheduled_task_failure --scheduled-task-id ID --restart-count-threshold N
levelrail-cli apps alerts create <app> --name NAME --kind domain_health [--for-duration 2m]
levelrail-cli apps alerts create <app> --name NAME --kind control_plane_backup_stale [--for-duration 72h]
levelrail-cli apps alerts update <app> <id> --name NAME --kind KIND [flags]
levelrail-cli apps alerts delete <app> <id>
# any create or update also takes: --severity info|warning|critical --consecutive-failures N
# --flap-threshold N --flap-window 30m --label key=value
# a PUT replaces the whole rule, so pass these again on every update
levelrail-cli alerts silences list [--all]
levelrail-cli alerts silences create --for 4h [--rule ID] [--app NAME] [--node NAME] [--kind KIND] [--severity S] [--label k=v] [--reason TEXT]
levelrail-cli alerts silences delete <id>
levelrail-cli alerts silence <app> <rule-id> [--for 1h] [--reason TEXT]
levelrail-cli alerts maintenance list
levelrail-cli alerts maintenance create --name NAME --cron "0 3 * * 0" --duration 2h [--tz Europe/Berlin] [--scope all|app|node] [--target NAME]...
levelrail-cli alerts maintenance update <id> [same flags]
levelrail-cli alerts maintenance delete <id>
levelrail-cli alerts history [--app NAME] [--rule ID] [--outcome OUTCOME] [--event EVENT] [--since TIME] [--limit N]
levelrail-cli channels list
levelrail-cli channels create --name NAME --kind KIND --notify-url URL
levelrail-cli channels update <id> --name NAME --kind KIND [flags]
levelrail-cli channels delete <id>
levelrail-cli channels test <id>
levelrail-cli channels deliveries <id> [--limit N]--kind for channels accepts: generic, slack, discord, telegram, email, pushover, pagerduty, teams, resend, ntfy, gotify, mattermost, lark, rocketchat, opsgenie, webex, googlechat.
Not built yet (deliberate gaps)
Missing metrics collectors:
- Request metrics only cover traffic that passes through this control plane's embedded Caddy (proxy and static routes). WebSocket and long-lived streams are recorded when they end, so their latency is the connection lifetime. Multi-node request federation reuses the
MetricsSourceinterface but has only been exercised in process, since remote nodes do not run their own ingress yet.
Missing node metrics:
- True host-level readings (real free/total CPU or memory) for any node other than the one running the control plane.
GET /api/v1/nodes/{id}/metricsandGET /api/v1/nodes/resource-usageboth sum already-collected per-container samples for CPU/memory usage;internal/agenthas no/procreads today, so a remote node'smemory_total_bytes/disk_used_bytes/disk_total_bytesstay absent rather than wrong. A future per-node agent writing host samples under the samenode:<id>resource-ID format would show up in both endpoints automatically, no handler change needed. - No node reports its own CPU core count anywhere in this codebase, so
nodes/resource-usage's fleet-widetotal_cpu_percentis a raw sum across nodes, not a percentage of fleet CPU capacity the waymemory_used_percent/disk_used_percentare. - Database placement contribution to node-level sums. Databases placed on a node don't appear in summed CPU/memory metrics.
Missing CLI and API features:
- No
levelrail-cli apps logs downloadwrapper (endpoint exists; works viacurl). - No download endpoint for database logs (only apps).
- Slow query log threshold is control-plane-wide (
APP_DATABASE_SLOW_QUERY_THRESHOLD_MS), not configurable per database. - Postgres slow-query parsing only captures the single log line the
duration:LOG entry itself occupies; a query long enough to wrap, or followed by aDETAIL/parameters line, isn't stitched back together. - No API endpoint reconstructs exact log lines a firing crashloop alert attached to its notification (payload only; dashboard links to log view instead).
Fixed configurations:
- Alert evaluation interval (30s) is fixed, not env-configurable (unlike per-kind thresholds).
- No alert-rule-specific change history (visible only in generic
GET /api/v1/audit-log). - SLO burn-rate rules are not built; the noise pipeline has no special handling for them yet.
- Grouping, streak, flap and held-notification state is in memory and does not survive a control plane restart.
- Silences and history are global, not scoped by IAM policy on individual apps: any principal with the base
readorwriteability can list or create them for any app.
See also
- Public status page - opt-in read-only page with component status, uptime bars and incidents
- API Reference - full telemetry endpoint documentation
- Feature Catalog - metrics and logs in the platform overview
- Architecture - telemetry design decisions and phase 2 rationale