Skip to content

Server Beat limits

Owner decisions of 2026-09-27: self-healing re-arms itself after sustained health; no quiet hours for server alerts; the external ceiling is at most 1 reboot/hour and 2/day. Values are configuration, measured in production and re-tuned from data (see "Measured data" in the runbook).

Boundary Limit Failure / exhaustion Measured in production
Beat tick Every 5 s; a tick is cut off after 4 s; database lock wait 2 s, statement 3 s; no network checks in the tick The Beat keeps pulsing; db_ok=false in its pulse file tells the Watchdog Watchdog max_pulse_gap_seconds
Checks Health, PostgreSQL, containers, OpenBao, Watchdog, Bridge: 30 s; disk and backup: 5 min; TLS: daily; one pending check per scope. A probe that times out or crashes is recorded as FAIL Incident, bounded action —
Incident lifecycle Opens after 2 consecutive bad observations; closes after 120 s of continuous health; unsettled checks are re-probed every 30 s, and while unsettled a gap over 3 × 30 s restarts the healthy period (also for the daily TLS check) One "down" (or "warning") and one "recovered" per incident, also for a flapping check —
Notifications iPhone push first; e-mail when the push is not confirmed within 60 s, failed, or has no route; e-mail withdraws pushes not yet started; one SMTP deadline of 8 s including name resolution; notify job timeout 20 s; retries every 15 s while a push is pending, otherwise 30 s doubling to 5 min, until delivered; one incident's notifications are delivered in order Never dropped: a notification that ends as failed is queued again by the Beat; never counted as a repair escalation —
External pulse Every 30 s; never forwarded for a Beat pulse older than 15 s External probe decides independently —
On-host Watchdog Beat dead after 90 s without a pulse file; runner dead after 90 s; database alert when the Beat reports db_ok=false for 90 s; 60 s pause after each restart; max 3 restarts per 15 min per container; at most 2 e-mails per tick, retried every 60 s Stop + e-mail; re-arms after 1 h of continuous health (a gap over 60 s between healthy ticks restarts the hour); at most 20 undelivered e-mails, older ones counted and reported max_pulse_gap_seconds per target
Runner: container restart Skip if already healthy or if its incident is closed; wait at most 60 s for health (poll 3 s; a probe starts only if its full timeout still fits); max 3 per 15 min per target, across incidents; a relapse inside an open incident is repaired again within the same budget Escalation with the reason (budget spent, or restarts did not heal); re-arms after 1 h of continuous health (a gap over 3 × check cadence restarts the hour) seconds_to_healthy per restart (last production deploys: healthy within ~6 s)
Repair broker Serves one request at a time; each request carries its caller's deadline (max 60 s); no Docker or systemd operation starts after that deadline or after the caller hung up; a container restart starts only with at least 15 s left The caller sees a failure; nothing happens late —
Ended repairs Every Beat tick escalates a repair that ended as failed (lease expired or the runner stopped) while its incident is still open and was not escalated One escalation per incident —
Migrations Applied at start on their own connection with a 600 s deadline, before the short observer deadlines apply Start fails visibly —
Runner: Watchdog restore Skip if the Watchdog pulse file is fresh; broker call 15 s; verify a fresh pulse within 15 s Escalation, shared budget rules —
External Watchdog Probe every minute; 3 failed probes open an incident and send "down"; reboot only if the Beat pulse is also older than 5 min; max 1 reboot/hour, 2/day Stop + independent SES; re-arms after 1 h of continuous health (a gap over 180 s between probes restarts the hour) /status last_outage_seconds
External alert e-mail One evaluation: probe 5 s, then reboot 10 s, then at most one digest e-mail 5 s, only if it still fits a 25 s budget; e-mail never runs before the health decision and never blocks pulses Undelivered alerts stay in a per-incident outbox (max 20 incidents) and go out as one digest on the next evaluation; beyond 20 the oldest are counted and the count and time range are reported in the digest /status alerts_pending, alerts_dropped
Backup Uncovered application marker and last success older than 30 min; worker pulse older than 5 min or a newer failure also fails Notification, not a periodic dump Confirm against the real backup cadence
Manifest polling Cache 300 s, including a missing manifest No listing on each tick —
Disk Warning below 15 %; failure below 5 %; every 5 min Warning incident, then "down" —
TLS Warning under 14 days; failure under 3 days; daily; connect timeout 4 s Warning incident, then "down" —
Action timeouts Container restart 100 s (pre-check 5 s + broker 30 s + grace 60 s); Watchdog restore 35 s; notify 20 s; pulse 10 s; runner lease 3 min Retry within the shared budget, then escalate —
Retention Finished Server Beat work 14 days; recovered incidents and sent e-mail records 180 days; hourly, in batches of 5000 Tick queries stay index lookups (180 days of history without retention: tick 0.07 s) —
Resources Beat 256 MiB / 0.25 CPU; runner 512 MiB / 0.5 CPU; host units 128 MiB / 15 % CPU each Runtime limits —

Server Beat incident notifications (down, warning, escalated, recovered) bypass the iPhone severity filter and quiet hours; "recovered" is sent as attention. bridge.online is informational: a sleeping Mac is normal and never pages the owner.

Only observed time counts as health: when no probe ran (runner stopped, Watchdog restarted, cron skipped), the next healthy probe starts a new hour. The on-host Watchdog never trusts a healthy period from before its own restart. The external Watchdog closes an incident on confirmed recovery even if the recovery e-mail fails, so the next outage always opens a new incident with its own "down".

Every value is validated when the process starts; an invalid value stops the process with a message instead of failing at incident time.

Configuration reference

Runner and Beat (Compose server-beat / server-beat-runner):

Variable Default Meaning
HEYAIRA_SERVER_BEAT_INTERVAL_SECONDS 5 Beat tick interval
HEYAIRA_SERVER_BEAT_RUNNER_POLL_SECONDS 1 Runner idle poll
HEYAIRA_SERVER_BEAT_NOTIFICATION_PROJECT_ID — Project whose iPhone receives pushes (required unless only e-mail is used)
HEYAIRA_SERVER_BEAT_SES_CONFIG_FILE — Private SES SMTP JSON (host, port, sender, recipient, username, password)
HEYAIRA_SERVER_BEAT_PUSH_CONFIRM_SECONDS 60 Wait for push confirmation before e-mail
HEYAIRA_SERVER_BEAT_INCIDENT_CONFIRM_OBSERVATIONS 2 Consecutive bad observations that open an incident (1-10)
HEYAIRA_SERVER_BEAT_RECOVERY_STABLE_SECONDS 120 Continuous health before "recovered"
HEYAIRA_SERVER_BEAT_RECHECK_SECONDS 30 Re-probe interval while a check is unsettled
HEYAIRA_SERVER_BEAT_RESTART_ENABLED false Activates repair actions (set by the activation overlay)
HEYAIRA_SERVER_BEAT_RESTART_CONTAINER — Exact app container restarted by app.healthz
HEYAIRA_SERVER_BEAT_ALLOWED_CONTAINERS — Comma-separated exact container allowlist (runner and broker)
HEYAIRA_SERVER_BEAT_RESTART_HEALTH_GRACE_SECONDS 60 Health wait after a restart
HEYAIRA_SERVER_BEAT_RESTART_HEALTH_POLL_SECONDS 3 Health poll during the grace
HEYAIRA_SERVER_BEAT_REPAIR_WINDOW_SECONDS 900 Repair budget window
HEYAIRA_SERVER_BEAT_MAX_REPAIRS 3 Repairs per window per target (1-3)
HEYAIRA_SERVER_BEAT_REPAIR_RESET_SECONDS 3600 Continuous health that re-arms a spent budget
HEYAIRA_SERVER_BEAT_PULSE_DIR — Shared pulse directory (set by the activation overlay)
HEYAIRA_SERVER_BEAT_WATCHDOG_THRESHOLD_SECONDS 90 Watchdog pulse file age that fails watchdog.alive
HEYAIRA_SERVER_BEAT_BRIDGE_NODE_ID — Bridge node for bridge.online
HEYAIRA_SERVER_BEAT_BRIDGE_THRESHOLD_SECONDS 120 Bridge pulse age that warns
HEYAIRA_SERVER_HEALTH_URL http://server:8000/healthz Application health endpoint
HEYAIRA_SERVER_BEAT_DISK_PATH / Filesystem checked by disk.free
HEYAIRA_SERVER_BEAT_DISK_WARN_PERCENT 15 Free space warning
HEYAIRA_SERVER_BEAT_DISK_FAIL_PERCENT 5 Free space failure
HEYAIRA_SERVER_BEAT_TLS_HOST mcp.heyaira.eu TLS host
HEYAIRA_SERVER_BEAT_TLS_PORT 443 TLS port
HEYAIRA_SERVER_BEAT_BACKUP_SCOPE prod Backup scope checked
HEYAIRA_SERVER_BEAT_BACKUP_MAX_AGE_SECONDS 1800 Uncovered backup age that fails
HEYAIRA_SERVER_BEAT_EXTERNAL_WATCHDOG_URL — External Worker /pulse URL
HEYAIRA_SERVER_BEAT_PULSE_TOKEN_FILE — Private pulse token file
HEYAIRA_SERVER_BEAT_MAINTENANCE_SECONDS 3600 Retention run interval
HEYAIRA_SERVER_BEAT_WORK_RETENTION_DAYS 14 Finished work retention
HEYAIRA_SERVER_BEAT_INCIDENT_RETENTION_DAYS 180 Recovered incident retention
HEYAIRA_REPAIR_BROKER_SOCKET — Broker Unix socket (runner, Watchdog, broker)

On-host Watchdog (~/.config/heyaira/watchdog.env):

Variable Default Meaning
HEYAIRA_WATCHDOG_STATE_FILE — Persistent state JSON (user state directory, not /run)
HEYAIRA_SERVER_BEAT_PULSE_DIR — The same host pulse directory the containers mount
HEYAIRA_WATCHDOG_BEAT_CONTAINER — Exact Beat container
HEYAIRA_WATCHDOG_RUNNER_CONTAINER — Exact runner container
HEYAIRA_WATCHDOG_SERVER_BEAT_THRESHOLD_SECONDS 90 Beat pulse age that counts as dead
HEYAIRA_WATCHDOG_RUNNER_THRESHOLD_SECONDS 90 Runner pulse age that counts as dead
HEYAIRA_WATCHDOG_DATABASE_THRESHOLD_SECONDS 90 How long db_ok=false lasts before e-mail
HEYAIRA_WATCHDOG_RESTART_COOLDOWN_SECONDS 60 Pause after a restart
HEYAIRA_WATCHDOG_MAX_RESTARTS 3 Restarts per window per container (1-3)
HEYAIRA_WATCHDOG_RESTART_WINDOW_SECONDS 900 Restart budget window
HEYAIRA_WATCHDOG_REARM_SECONDS 3600 Continuous health that re-arms a stop
HEYAIRA_WATCHDOG_HEALTH_MAX_GAP_SECONDS 60 Gap that restarts the healthy period
HEYAIRA_WATCHDOG_INTERVAL_SECONDS 5 Tick interval
HEYAIRA_SERVER_BEAT_SES_CONFIG_FILE — Private SES SMTP JSON (required: the Watchdog's only channel)
HEYAIRA_REPAIR_BROKER_SOCKET — Broker socket

Repair broker (~/.config/heyaira/repair-broker.env): HEYAIRA_REPAIR_BROKER_SOCKET, HEYAIRA_SERVER_BEAT_ALLOWED_CONTAINERS (app, Beat and runner containers), HEYAIRA_REPAIR_COMPOSE_PROJECT (production: app, required) and HEYAIRA_REPAIR_DOCKER_SOCKET (default /var/run/docker.sock; rootless Docker: /run/user/<uid>/docker.sock).