Server Beat limits¶
Owner decisions of 2026-09-27: self-healing re-arms itself after sustained health; no quiet hours for server alerts; the external ceiling is at most 1 reboot/hour and 2/day. Values are configuration, measured in production and re-tuned from data (see "Measured data" in the runbook).
| Boundary | Limit | Failure / exhaustion | Measured in production |
|---|---|---|---|
| Beat tick | Every 5 s; a tick is cut off after 4 s; database lock wait 2 s, statement 3 s; no network checks in the tick | The Beat keeps pulsing; db_ok=false in its pulse file tells the Watchdog |
Watchdog max_pulse_gap_seconds |
| Checks | Health, PostgreSQL, containers, OpenBao, Watchdog, Bridge: 30 s; disk and backup: 5 min; TLS: daily; one pending check per scope. A probe that times out or crashes is recorded as FAIL | Incident, bounded action | — |
| Incident lifecycle | Opens after 2 consecutive bad observations; closes after 120 s of continuous health; unsettled checks are re-probed every 30 s, and while unsettled a gap over 3 × 30 s restarts the healthy period (also for the daily TLS check) | One "down" (or "warning") and one "recovered" per incident, also for a flapping check | — |
| Notifications | iPhone push first; e-mail when the push is not confirmed within 60 s, failed, or has no route; e-mail withdraws pushes not yet started; one SMTP deadline of 8 s including name resolution; notify job timeout 20 s; retries every 15 s while a push is pending, otherwise 30 s doubling to 5 min, until delivered; one incident's notifications are delivered in order | Never dropped: a notification that ends as failed is queued again by the Beat; never counted as a repair escalation | — |
| External pulse | Every 30 s; never forwarded for a Beat pulse older than 15 s | External probe decides independently | — |
| On-host Watchdog | Beat dead after 90 s without a pulse file; runner dead after 90 s; database alert when the Beat reports db_ok=false for 90 s; 60 s pause after each restart; max 3 restarts per 15 min per container; at most 2 e-mails per tick, retried every 60 s |
Stop + e-mail; re-arms after 1 h of continuous health (a gap over 60 s between healthy ticks restarts the hour); at most 20 undelivered e-mails, older ones counted and reported | max_pulse_gap_seconds per target |
| Runner: container restart | Skip if already healthy or if its incident is closed; wait at most 60 s for health (poll 3 s; a probe starts only if its full timeout still fits); max 3 per 15 min per target, across incidents; a relapse inside an open incident is repaired again within the same budget | Escalation with the reason (budget spent, or restarts did not heal); re-arms after 1 h of continuous health (a gap over 3 × check cadence restarts the hour) | seconds_to_healthy per restart (last production deploys: healthy within ~6 s) |
| Repair broker | Serves one request at a time; each request carries its caller's deadline (max 60 s); no Docker or systemd operation starts after that deadline or after the caller hung up; a container restart starts only with at least 15 s left | The caller sees a failure; nothing happens late | — |
| Ended repairs | Every Beat tick escalates a repair that ended as failed (lease expired or the runner stopped) while its incident is still open and was not escalated | One escalation per incident | — |
| Migrations | Applied at start on their own connection with a 600 s deadline, before the short observer deadlines apply | Start fails visibly | — |
| Runner: Watchdog restore | Skip if the Watchdog pulse file is fresh; broker call 15 s; verify a fresh pulse within 15 s | Escalation, shared budget rules | — |
| External Watchdog | Probe every minute; 3 failed probes open an incident and send "down"; reboot only if the Beat pulse is also older than 5 min; max 1 reboot/hour, 2/day | Stop + independent SES; re-arms after 1 h of continuous health (a gap over 180 s between probes restarts the hour) | /status last_outage_seconds |
| External alert e-mail | One evaluation: probe 5 s, then reboot 10 s, then at most one digest e-mail 5 s, only if it still fits a 25 s budget; e-mail never runs before the health decision and never blocks pulses | Undelivered alerts stay in a per-incident outbox (max 20 incidents) and go out as one digest on the next evaluation; beyond 20 the oldest are counted and the count and time range are reported in the digest | /status alerts_pending, alerts_dropped |
| Backup | Uncovered application marker and last success older than 30 min; worker pulse older than 5 min or a newer failure also fails | Notification, not a periodic dump | Confirm against the real backup cadence |
| Manifest polling | Cache 300 s, including a missing manifest | No listing on each tick | — |
| Disk | Warning below 15 %; failure below 5 %; every 5 min | Warning incident, then "down" | — |
| TLS | Warning under 14 days; failure under 3 days; daily; connect timeout 4 s | Warning incident, then "down" | — |
| Action timeouts | Container restart 100 s (pre-check 5 s + broker 30 s + grace 60 s); Watchdog restore 35 s; notify 20 s; pulse 10 s; runner lease 3 min | Retry within the shared budget, then escalate | — |
| Retention | Finished Server Beat work 14 days; recovered incidents and sent e-mail records 180 days; hourly, in batches of 5000 | Tick queries stay index lookups (180 days of history without retention: tick 0.07 s) | — |
| Resources | Beat 256 MiB / 0.25 CPU; runner 512 MiB / 0.5 CPU; host units 128 MiB / 15 % CPU each | Runtime limits | — |
Server Beat incident notifications (down, warning, escalated, recovered)
bypass the iPhone severity filter and quiet hours; "recovered" is sent as
attention. bridge.online is informational: a sleeping Mac is normal and
never pages the owner.
Only observed time counts as health: when no probe ran (runner stopped, Watchdog restarted, cron skipped), the next healthy probe starts a new hour. The on-host Watchdog never trusts a healthy period from before its own restart. The external Watchdog closes an incident on confirmed recovery even if the recovery e-mail fails, so the next outage always opens a new incident with its own "down".
Every value is validated when the process starts; an invalid value stops the process with a message instead of failing at incident time.
Configuration reference¶
Runner and Beat (Compose server-beat / server-beat-runner):
| Variable | Default | Meaning |
|---|---|---|
HEYAIRA_SERVER_BEAT_INTERVAL_SECONDS |
5 | Beat tick interval |
HEYAIRA_SERVER_BEAT_RUNNER_POLL_SECONDS |
1 | Runner idle poll |
HEYAIRA_SERVER_BEAT_NOTIFICATION_PROJECT_ID |
— | Project whose iPhone receives pushes (required unless only e-mail is used) |
HEYAIRA_SERVER_BEAT_SES_CONFIG_FILE |
— | Private SES SMTP JSON (host, port, sender, recipient, username, password) |
HEYAIRA_SERVER_BEAT_PUSH_CONFIRM_SECONDS |
60 | Wait for push confirmation before e-mail |
HEYAIRA_SERVER_BEAT_INCIDENT_CONFIRM_OBSERVATIONS |
2 | Consecutive bad observations that open an incident (1-10) |
HEYAIRA_SERVER_BEAT_RECOVERY_STABLE_SECONDS |
120 | Continuous health before "recovered" |
HEYAIRA_SERVER_BEAT_RECHECK_SECONDS |
30 | Re-probe interval while a check is unsettled |
HEYAIRA_SERVER_BEAT_RESTART_ENABLED |
false | Activates repair actions (set by the activation overlay) |
HEYAIRA_SERVER_BEAT_RESTART_CONTAINER |
— | Exact app container restarted by app.healthz |
HEYAIRA_SERVER_BEAT_ALLOWED_CONTAINERS |
— | Comma-separated exact container allowlist (runner and broker) |
HEYAIRA_SERVER_BEAT_RESTART_HEALTH_GRACE_SECONDS |
60 | Health wait after a restart |
HEYAIRA_SERVER_BEAT_RESTART_HEALTH_POLL_SECONDS |
3 | Health poll during the grace |
HEYAIRA_SERVER_BEAT_REPAIR_WINDOW_SECONDS |
900 | Repair budget window |
HEYAIRA_SERVER_BEAT_MAX_REPAIRS |
3 | Repairs per window per target (1-3) |
HEYAIRA_SERVER_BEAT_REPAIR_RESET_SECONDS |
3600 | Continuous health that re-arms a spent budget |
HEYAIRA_SERVER_BEAT_PULSE_DIR |
— | Shared pulse directory (set by the activation overlay) |
HEYAIRA_SERVER_BEAT_WATCHDOG_THRESHOLD_SECONDS |
90 | Watchdog pulse file age that fails watchdog.alive |
HEYAIRA_SERVER_BEAT_BRIDGE_NODE_ID |
— | Bridge node for bridge.online |
HEYAIRA_SERVER_BEAT_BRIDGE_THRESHOLD_SECONDS |
120 | Bridge pulse age that warns |
HEYAIRA_SERVER_HEALTH_URL |
http://server:8000/healthz |
Application health endpoint |
HEYAIRA_SERVER_BEAT_DISK_PATH |
/ |
Filesystem checked by disk.free |
HEYAIRA_SERVER_BEAT_DISK_WARN_PERCENT |
15 | Free space warning |
HEYAIRA_SERVER_BEAT_DISK_FAIL_PERCENT |
5 | Free space failure |
HEYAIRA_SERVER_BEAT_TLS_HOST |
mcp.heyaira.eu |
TLS host |
HEYAIRA_SERVER_BEAT_TLS_PORT |
443 | TLS port |
HEYAIRA_SERVER_BEAT_BACKUP_SCOPE |
prod |
Backup scope checked |
HEYAIRA_SERVER_BEAT_BACKUP_MAX_AGE_SECONDS |
1800 | Uncovered backup age that fails |
HEYAIRA_SERVER_BEAT_EXTERNAL_WATCHDOG_URL |
— | External Worker /pulse URL |
HEYAIRA_SERVER_BEAT_PULSE_TOKEN_FILE |
— | Private pulse token file |
HEYAIRA_SERVER_BEAT_MAINTENANCE_SECONDS |
3600 | Retention run interval |
HEYAIRA_SERVER_BEAT_WORK_RETENTION_DAYS |
14 | Finished work retention |
HEYAIRA_SERVER_BEAT_INCIDENT_RETENTION_DAYS |
180 | Recovered incident retention |
HEYAIRA_REPAIR_BROKER_SOCKET |
— | Broker Unix socket (runner, Watchdog, broker) |
On-host Watchdog (~/.config/heyaira/watchdog.env):
| Variable | Default | Meaning |
|---|---|---|
HEYAIRA_WATCHDOG_STATE_FILE |
— | Persistent state JSON (user state directory, not /run) |
HEYAIRA_SERVER_BEAT_PULSE_DIR |
— | The same host pulse directory the containers mount |
HEYAIRA_WATCHDOG_BEAT_CONTAINER |
— | Exact Beat container |
HEYAIRA_WATCHDOG_RUNNER_CONTAINER |
— | Exact runner container |
HEYAIRA_WATCHDOG_SERVER_BEAT_THRESHOLD_SECONDS |
90 | Beat pulse age that counts as dead |
HEYAIRA_WATCHDOG_RUNNER_THRESHOLD_SECONDS |
90 | Runner pulse age that counts as dead |
HEYAIRA_WATCHDOG_DATABASE_THRESHOLD_SECONDS |
90 | How long db_ok=false lasts before e-mail |
HEYAIRA_WATCHDOG_RESTART_COOLDOWN_SECONDS |
60 | Pause after a restart |
HEYAIRA_WATCHDOG_MAX_RESTARTS |
3 | Restarts per window per container (1-3) |
HEYAIRA_WATCHDOG_RESTART_WINDOW_SECONDS |
900 | Restart budget window |
HEYAIRA_WATCHDOG_REARM_SECONDS |
3600 | Continuous health that re-arms a stop |
HEYAIRA_WATCHDOG_HEALTH_MAX_GAP_SECONDS |
60 | Gap that restarts the healthy period |
HEYAIRA_WATCHDOG_INTERVAL_SECONDS |
5 | Tick interval |
HEYAIRA_SERVER_BEAT_SES_CONFIG_FILE |
— | Private SES SMTP JSON (required: the Watchdog's only channel) |
HEYAIRA_REPAIR_BROKER_SOCKET |
— | Broker socket |
Repair broker (~/.config/heyaira/repair-broker.env): HEYAIRA_REPAIR_BROKER_SOCKET,
HEYAIRA_SERVER_BEAT_ALLOWED_CONTAINERS (app, Beat and runner containers),
HEYAIRA_REPAIR_COMPOSE_PROJECT (production: app, required) and
HEYAIRA_REPAIR_DOCKER_SOCKET (default /var/run/docker.sock; rootless Docker:
/run/user/<uid>/docker.sock).