HydraIssues

Fleet-level synthetic checks for silently-failing load-bearing paths (release downloads, hydrabackup, mesh handshakes)
open improvement Project: hydrascalerouter Reporter: 11 Aug 2026 18:11

Description

Why

Three load-bearing paths failed silently this arc and went undetected for hours-to-days because nothing actively watched them:

  1. Mesh outage after reboot (#435) — wg-quick@wg0 came up disabled on the district hub after a reboot; the 24-peer mesh was down ~11h. hydraguard logged wg show: exit status 1 every 10s into a journal nobody watched.
  2. Fleet binary downloads broke for hours — the release mirror redirect pointed at a downed Pi; installers/runbooks pulling through releases.experiencenet.com got a black hole. No signal until a human hit it.
  3. hydrabackup failed on 4 hosts for a full day — a hostname rename broke the per-host backup path; nobody noticed backups had stopped.

Common pattern: detection either doesn't exist or exists but has no alert sink.

What #453 / #454 already cover (and why they do NOT catch these)

  • #454 (hydrascalerouter — alert on sustained edge 5xx for registered domains). This is the right mechanism (probe → threshold → auto-file hydraissue) but it only watches router-registered vhosts. It does NOT cover failure 2: (a) releases.experiencenet.com is not a router-registered domain — confirmed against /routes, only hydramirror.experiencenet.com backends appear — so the router never probes it; (b) even if it did, hydrarelease returns a 302 to a mirror (internal/api/server.go:217), not a 5xx, so a redirect to a dead target reads as healthy to an edge-5xx probe.
  • #453 (hydracluster — gate rollouts on node health, alert on stuck jobs / sustained heartbeat loss). Its heartbeat-loss item is adjacent to failure 1, but it is scoped to the rollout dispatch path (pre-flight gate + stuck-job timeout during a rollout). It does not watch standing WireGuard peer-handshake staleness outside a rollout, which is the #435 class. Neither #453 nor #454 covers backups at all.

So all three failures are genuine GAPS relative to #453/#454. This issue is complementary, not a duplicate — it adds the missing synthetic/liveness checks and reuses #454's alert+auto-file plumbing.

The gaps, with concrete evidence

(a) Release-download health. There is no standing external check that a known file actually downloads through releases.experiencenet.com. App-level failover exists (commit 0afaad5, pickHealthyMirror/defaultMirrorProbe in hydrarelease/internal/api/mirror.go) and will 503 if all mirrors are dead, but: the probe is a HEAD to the mirror base URL (mirror.go:107-118), blind to per-file availability / partial sync; it runs lazily per-download with a 10s cache, not as a monitored heartbeat; and because releases.* isn't router-registered, a sustained 503 is invisible to #454. A mirror that is "up" but not serving the requested release path still passes the HEAD and black-holes the client.

(b) hydrabackup success/failure per host. hydrabackup is real (syncs mesh.yaml + snapshots to hydramirror per hydraguard/docs/runbooks/runbook.md:275,433) and is only known to hydracluster as a job category (hydracluster/pkg/api/handlers_api.go:373 → {Category: "infra"}). Nothing tracks per-host last-success time or alerts when a host stops backing up. Failure 3 ran for a full day undetected.

(c) Mesh peer handshake staleness (#435 class). hydraguard already detects the failure (logs wg show errors every 10s) but has nowhere to send it (#435 note: "it needs somewhere to send it"). Nothing alerts when a peer's last-handshake age exceeds a threshold or when the hub peer count drops to zero.

Smallest set of checks that would have caught all three

  1. Synthetic release GET — on an interval, GET https://releases.experiencenet.com/.../<known-small-file> following the 302, assert 200 + expected content-length (or checksum). Catches dead-mirror black holes, sustained 503, and partial-sync mirrors that pass the HEAD probe. (a)
  2. Backup freshness check — per host, assert last hydrabackup success within N hours (read the sync marker/timestamp hydrabackup writes to hydramirror, or a per-host last-success record). Alert on any host past threshold. (b)
  3. Mesh handshake staleness check — on the hub, alert when any configured peer's last-handshake age exceeds a threshold, or peer count hits zero / wg show fails. Give hydraguard's existing detection an alert sink. (c)

Where this should live

The router's ProbeBackends (hydrascalerouter/internal/cli/backendhealth.go, driven from serve.go:199) is today the only fleet-level active prober, and #454 already extends it with threshold + auto-file-hydraissue. Cheapest path: add check (1) as another synthetic probe in that same loop and reuse the #454 alert sink. Checks (2) and (3) read fleet state that lives elsewhere — (2) is most natural against hydracluster's per-host job state, (3) against hydraguard's existing wg show polling — so they can either feed the same router-side alert sink or emit through their own home repo. One tracking issue, three small checks, all routing to the same debounced auto-file path #454 establishes.

Refs: #435 (mesh outage), #453 (rollout gating / heartbeat), #454 (edge-5xx alerting mechanism to reuse).