Server Beat runbook¶
Status: implemented and under independent review; not deployed. Limits in SERVER_BEAT_LIMITS.md were approved by the owner on 2026-09-27 and are tuned from measured data. No automatic OpenBao unseal and no OpenBao/PostgreSQL restart, ever.
What runs where¶
server-beat(container): schedules due checks every 5 s and writes its pulse (database and pulse file). It never checks, repairs or notifies.server-beat-runner(container): runs checks and actions from the durable queue, writes its pulse file every 5 s, deletes old history hourly.heyaira-repair-broker(host user unit): the only holder of the Docker socket. Restarts exactly allow-listed containers of the configured Compose project whose service isserver,server-beatorserver-beat-runner, and restarts the Watchdog user unit. Nothing else.heyaira-server-beat-watchdog(host user unit): watches the Beat and runner pulse files, restarts a stale one through the broker, e-mails the owner when that fails or when the Beat reports that PostgreSQL is unreachable. It has no database access and no database credential.- External Worker (Cloudflare): probes the public health URL every minute and, only when the host is gone too, reboots the VPS within its budget.
The Beat and runner sit behind the Compose profile server-beat: a plain
docker compose up or an all-services deploy never starts them. They are
started only by naming them, with the activation overlay.
Activation (separate owner-approved rollout)¶
- Host directories, owned by the rootless-Docker user, mode 0700:
~/.local/state/heyaira/pulses(pulse files),~/.local/state/heyaira/repair(broker socket),~/.local/state/heyaira/watchdog.json(Watchdog state, not under /run). Containers run as root inside rootless Docker, which is the host user, so the files they write are readable by the host units. - Secrets outside Git:
secrets/pulse.tokenandsecrets/ses.json(mode 0600; SES SMTP host, port, sender, recipient, username, password). ~/.config/heyaira/repair-broker.env:HEYAIRA_REPAIR_BROKER_SOCKET=<home>/.local/state/heyaira/repair/broker.sock,HEYAIRA_REPAIR_COMPOSE_PROJECT=app,HEYAIRA_REPAIR_DOCKER_SOCKET=/run/user/<uid>/docker.sock,HEYAIRA_SERVER_BEAT_ALLOWED_CONTAINERS=app-server-1,app-server-beat-1,app-server-beat-runner-1.~/.config/heyaira/watchdog.env:HEYAIRA_WATCHDOG_STATE_FILE,HEYAIRA_SERVER_BEAT_PULSE_DIR(the host pulse directory),HEYAIRA_WATCHDOG_BEAT_CONTAINER=app-server-beat-1,HEYAIRA_WATCHDOG_RUNNER_CONTAINER=app-server-beat-runner-1,HEYAIRA_REPAIR_BROKER_SOCKET,HEYAIRA_SERVER_BEAT_SES_CONFIG_FILE. The Watchdog refuses to start without SES: e-mail is its only channel.- Production
.envfor the containers:HEYAIRA_SERVER_BEAT_NOTIFICATION_PROJECT_ID(explicit; there is no first-project fallback),HEYAIRA_SERVER_BEAT_RESTART_CONTAINER=app-server-1,HEYAIRA_SERVER_BEAT_ALLOWED_CONTAINERS=app-server-1,HEYAIRA_SERVER_BEAT_PULSE_DIRECTORYandHEYAIRA_REPAIR_BROKER_DIRECTORY(host paths),HEYAIRA_SERVER_BEAT_EXTERNAL_WATCHDOG_URL. - Install and start the host units. They run the current release's code
(
/srv/heyaira/current/src) on the host's plainpython3; the Watchdog, broker, pulse and SES modules use the standard library only, so nothing is installed on the host. Copyops/systemd/*.serviceto~/.config/systemd/user/, then:
systemctl --user daemon-reload
systemctl --user enable --now heyaira-repair-broker.service heyaira-server-beat-watchdog.service
systemctl --user status heyaira-repair-broker.service heyaira-server-beat-watchdog.service
- Start the containers with the production overrides plus the activation overlay (absolute path), naming only these services; never recreate OpenBao:
./ops/deploy-heyaira.sh ... --compose-project app \
--compose-override ops/openbao-production-compose.yaml \
--compose-extra /srv/heyaira/config/compose.restart.yaml \
--compose-extra /srv/heyaira/app/ops/server-beat-repair-compose.yaml \
--services "server-beat server-beat-runner"
Before the activation overlay (notify-only operation) there is no broker:
containers.running and watchdog.alive then report "not configured"
(informational, no page) and no repair is attempted.
Every process validates its configuration at start and exits with a message
if something is missing or invalid (for example, a restart container that is
not in the allowlist, repair enabled without a broker or pulse directory, no
notification channel). A crash-looping Beat or runner is seen by the Watchdog;
a crash-looping Watchdog is seen by the runner's watchdog.alive.
Notifications¶
A notification is delivered when the phone confirms the push (delivery
status sent). While the push is only queued the notify job checks again
every 15 s; after 60 s without confirmation, after a failed push, or when
there is no push route (no project, no node, no active phone) the same event
is e-mailed. The push sender lives in the application, so an application
outage is reported by e-mail. Notify jobs retry until delivered (15 s while a
push is pending, otherwise 30 s doubling to 5 min) and never produce a
repair escalation.
The e-mail row is claimed before SMTP and marked sent after it; SMTP has one
8 s deadline for the whole session, name resolution included (at most one
pending lookup per mail host, two in total), and no database lock is held
meanwhile. A crash between SMTP acceptance and the mark can
repeat that one e-mail after the claim expires (at-least-once, stable
Message-ID). Before e-mailing, pushes of the same event that have not started
are withdrawn (status filtered, superseded_by_email) and restored if the
e-mail fails; a push already handed to the push worker may still arrive.
Notifications of one incident are delivered in order: "recovered" waits for a
"down" that is still being retried. A notification that ends as failed for
any reason is queued again by the Beat, however long ago it failed; retention
never deletes a failed notification or an unescalated failed repair of an open
incident.
The on-host Watchdog e-mails directly (same SES file), at most two e-mails per tick; an e-mail that fails is retried every minute; at most 20 wait, older ones are counted and reported in the next e-mail.
Server Beat pushes bypass the iPhone severity filter and quiet hours (owner decision: no quiet hours for server alerts).
Incidents and repair¶
An incident opens after two consecutive bad observations and closes after two minutes of continuous health; unsettled checks are re-probed every 30 s. The owner gets "down" (or "warning" for low disk / expiring TLS) and "recovered" once per incident, also when a check flaps. A failure inside a still-open incident after a healthy observation is repaired again within the same budget, without another page.
Repair actions are enqueued only when repair is activated. The broker never starts a restart after its caller's deadline or after the caller hung up. A repair that ends without the runner (its last lease expired) is escalated by the next Beat tick while its incident is open. A repair skips when the target is already healthy or its incident has closed. When a repair gives up the owner gets one "escalated" that says why: the budget (3 per 15 minutes) is spent, the restarts did not restore health, or the repair could not run (with the error type).
Self-healing re-arms itself (owner decision 2026-09-27). A stop caused by an exhausted budget clears after the target has been continuously healthy for one hour; an unobserved gap (runner 3 × cadence, on-host Watchdog HEYAIRA_WATCHDOG_HEALTH_MAX_GAP_SECONDS, external HEALTH_MAX_GAP_SECONDS) restarts that hour. The external 1/hour and 2/day reboot ceilings still apply after re-arming. Every new outage is a new incident with its own down and recovered notifications. A manual reset (clear the control_repair_budgets row, archive the Watchdog state file) remains possible under change control; never delete open incidents.
External Watchdog¶
Authenticated external pulse is periodic (30 seconds). The runner refuses to forward it when the Beat database pulse is stale. The Worker probes HEALTH_URL every minute; three failed probes open an incident and queue one "down" (also in dry run). A reboot needs the HOST to be gone as well: the Beat pulse must be older than PULSE_REBOOT_AGE_SECONDS (300); an app-only outage belongs to the on-server layers and a needless reboot would seal OpenBao. It reserves budget before OVH I/O: max 1/hour and 2/day. The OVH request is signed first; the last pulse check happens with no pause before the request is sent, and a pulse that arrives during the budget reservation releases it. Exhaustion or an uncertain response stops reboots and queues an alert. OVH_REBOOT_ENABLED=false remains the default.
One evaluation does the probe (5 s), the reboot decision (OVH 10 s) and then
at most one digest e-mail (5 s) only if it fits a 25 s budget; e-mail never
runs before the health decision and never blocks pulses or /status. Undelivered
alerts wait in a per-incident outbox (20 incidents); beyond that the oldest
are counted and the next digest states how many and from when. If /status
shows alerts_pending above zero for more than a few minutes, the SES route
is broken (check the Worker secrets and SES sending status).
Measured data for tuning¶
Every container restart records seconds_to_healthy in its job result and a
server_beat_restart_measured log line; the on-host Watchdog records the
longest observed pulse gap per target (max_pulse_gap_seconds in its state
file); the external Watchdog reports last_outage_seconds on /status.
Independent acceptance¶
Fresh checkout of the exact SHA; scripts/test_postgres.py (zero skipped),
node --test ops/cloudflare/server-beat-watchdog/test/*.mjs, generated
contracts and strict MkDocs. Local tests use real PostgreSQL, HTTP and Unix
transports with synthetic Docker, systemd, SMTP, APNs and OVH boundaries. They
do not prove physical iPhone delivery, real host repair, live e-mail or reboot;
those are separate owner-approved drills after an independent PASS.