Three load-bearing paths failed silently this arc and went undetected for hours-to-days because nothing actively watched them:
wg-quick@wg0 came up disabled on the district hub after a reboot; the 24-peer mesh was down ~11h. hydraguard logged wg show: exit status 1 every 10s into a journal nobody watched.releases.experiencenet.com got a black hole. No signal until a human hit it.Common pattern: detection either doesn't exist or exists but has no alert sink.
releases.experiencenet.com is not a router-registered domain — confirmed against /routes, only hydramirror.experiencenet.com backends appear — so the router never probes it; (b) even if it did, hydrarelease returns a 302 to a mirror (internal/api/server.go:217), not a 5xx, so a redirect to a dead target reads as healthy to an edge-5xx probe.So all three failures are genuine GAPS relative to #453/#454. This issue is complementary, not a duplicate — it adds the missing synthetic/liveness checks and reuses #454's alert+auto-file plumbing.
(a) Release-download health. There is no standing external check that a known file actually downloads through releases.experiencenet.com. App-level failover exists (commit 0afaad5, pickHealthyMirror/defaultMirrorProbe in hydrarelease/internal/api/mirror.go) and will 503 if all mirrors are dead, but: the probe is a HEAD to the mirror base URL (mirror.go:107-118), blind to per-file availability / partial sync; it runs lazily per-download with a 10s cache, not as a monitored heartbeat; and because releases.* isn't router-registered, a sustained 503 is invisible to #454. A mirror that is "up" but not serving the requested release path still passes the HEAD and black-holes the client.
(b) hydrabackup success/failure per host. hydrabackup is real (syncs mesh.yaml + snapshots to hydramirror per hydraguard/docs/runbooks/runbook.md:275,433) and is only known to hydracluster as a job category (hydracluster/pkg/api/handlers_api.go:373 → {Category: "infra"}). Nothing tracks per-host last-success time or alerts when a host stops backing up. Failure 3 ran for a full day undetected.
(c) Mesh peer handshake staleness (#435 class). hydraguard already detects the failure (logs wg show errors every 10s) but has nowhere to send it (#435 note: "it needs somewhere to send it"). Nothing alerts when a peer's last-handshake age exceeds a threshold or when the hub peer count drops to zero.
GET https://releases.experiencenet.com/.../<known-small-file> following the 302, assert 200 + expected content-length (or checksum). Catches dead-mirror black holes, sustained 503, and partial-sync mirrors that pass the HEAD probe. (a)wg show fails. Give hydraguard's existing detection an alert sink. (c)The router's ProbeBackends (hydrascalerouter/internal/cli/backendhealth.go, driven from serve.go:199) is today the only fleet-level active prober, and #454 already extends it with threshold + auto-file-hydraissue. Cheapest path: add check (1) as another synthetic probe in that same loop and reuse the #454 alert sink. Checks (2) and (3) read fleet state that lives elsewhere — (2) is most natural against hydracluster's per-host job state, (3) against hydraguard's existing wg show polling — so they can either feed the same router-side alert sink or emit through their own home repo. One tracking issue, three small checks, all routing to the same debounced auto-file path #454 establishes.
Refs: #435 (mesh outage), #453 (rollout gating / heartbeat), #454 (edge-5xx alerting mechanism to reuse).