Split from #447 (item 2 of 4). The renumbering rollout was pushed to pi-node-003 while the node was already unhealthy: systemd-resolved was broken, so hydranode could not fetch execs from hydracluster. The cluster then polled 'grep -q DONE /root/rebind.log' for ~18 hours while the node's heartbeats were ALSO failing. Nobody was notified; #445 was filed by a human noticing the outage.
Required: (a) pre-flight health gate before dispatching a rollout step to a node - DNS resolution works, hydranode exec channel round-trips, incus API responsive; (b) alert (and stop the rollout) when a dispatched job has not completed within N minutes; (c) surface sustained heartbeat loss as an alert instead of a silent status flip.
Real-world evidence for the sustained-heartbeat-loss item (c): the #435 district-hub mesh outage — wg-quick@wg0 came up disabled after a reboot and the 24-peer mesh was down ~11h, with hydraguard logging wg show failures every 10s into a journal with no alert sink. Note the scope boundary: this issue gates the rollout dispatch path (pre-flight gate + stuck-job timeout + heartbeat loss during a rollout). It does NOT cover standing WireGuard peer-handshake staleness outside a rollout, which is the #435 class. That gap (alert when a peer last-handshake age exceeds a threshold / peer count hits zero), plus release-download and hydrabackup monitoring, is tracked separately in #482 so we do not overload the rollout-gating scope here.