Structural follow-up to issue #445 (pi-node-003: incusd deadlocks on startup, 8 production domains down ~18h).
WHAT HAPPENED
The hydraguard mesh renumbering (10.0.0.0/24 -> 10.10.0.0/16) was rolled out ~2026-08-05 21:00. On pi-node-003 the rollout landed HALF-WAY: the wireguard tunnel was renumbered (10.0.0.4 -> 10.10.100.19), but the node-side rebind job (which updates the 8 hydraskin-expose proxy devices in incus) never ran - the node's DNS (systemd-resolved) was broken, so hydranode could not fetch execs from hydracluster, and the cluster just kept polling "grep -q DONE /root/rebind.log" for ~18 hours.
Result: every incus container start succeeded, then its proxy device failed to bind the now-nonexistent 10.0.0.4, incus reverted the start, and the container stop-hook (incusd callhook ... stop) deadlocked against the still-initializing daemon (incus 7.2 bug, all threads in futex_wait). systemd then SIGKILLed incus after waitready(600s)+stop-sigterm(330s) and restarted it, forever. Each ~15.5min cycle leaked a veth pair and a hung "forknet dhcp" process. All 8 production domains were down from Aug 5 21:10 until manual recovery Aug 6 ~15:00. pi-node-004 was unaffected because its rebind completed (proxies listen on 10.10.100.18).
WHY THIS NEEDS A STRUCTURAL FIX (not just the recovery that was done)
PROPOSED STRUCTURAL FIXES
RECOVERY ALREADY APPLIED ON PI-NODE-003 (2026-08-06 ~15:00)
Wedged daemon killed, 16 leaked veth pairs + hung forknet processes cleaned, systemd-resolved restarted (DNS fixed), all 8 hydraskin-expose devices rebound to tcp:10.10.100.19:, verified locally + across the mesh + publicly. /root/rebind.log DONE marker written so the cluster can finish its side. 7/8 domains verified healthy; bxl1.hydramirror.experiencenet.com still returns 503 at the edge (backend registration for that vhost appears stale cluster-side - the app itself is healthy, and bxl1-b on pi-node-004 passes through fine), expected to clear once the cluster processes the rebind completion; if not, the edge backend registration for bxl1.hydramirror needs manual refresh.
UPDATE 2026-08-06 16:00: the bxl1.hydramirror 503 tail is RESOLVED (#448: stale static route on the edge, pinned pre-renumbering; #450: health_path now flows dynamically through hydraskin v0.10.0 -> hydracluster v2.0.102 -> hydrascalerouter v0.10.0, static pins removed). All 8 pi-node-003 domains verified healthy at the edge. Note for the structural fixes above: static edge routes were part of the failure surface - they always win conflicts against node reports and go stale silently; scales should publish via labels only (see hydrascalerouter runbook). The four structural items (atomic rebind, health-gated rollout + alerting, edge 5xx auto-filing, incus >7.2 upgrade) remain open and are the scope of this issue.
SPLIT 2026-08-06: the four structural items are now separate groomed issues: