HydraIssues

[SUPERSEDED by #452-#455] Mesh renumbering rebind not crash-safe: half-applied rebind left pi-node-003 hard-down 18h (needs atomicity, health gating, alerting)
closed bug Project: hydracluster Reporter: anonymous 6 Aug 2026 13:15

Description

Structural follow-up to issue #445 (pi-node-003: incusd deadlocks on startup, 8 production domains down ~18h).

WHAT HAPPENED
The hydraguard mesh renumbering (10.0.0.0/24 -> 10.10.0.0/16) was rolled out ~2026-08-05 21:00. On pi-node-003 the rollout landed HALF-WAY: the wireguard tunnel was renumbered (10.0.0.4 -> 10.10.100.19), but the node-side rebind job (which updates the 8 hydraskin-expose proxy devices in incus) never ran - the node's DNS (systemd-resolved) was broken, so hydranode could not fetch execs from hydracluster, and the cluster just kept polling "grep -q DONE /root/rebind.log" for ~18 hours.

Result: every incus container start succeeded, then its proxy device failed to bind the now-nonexistent 10.0.0.4, incus reverted the start, and the container stop-hook (incusd callhook ... stop) deadlocked against the still-initializing daemon (incus 7.2 bug, all threads in futex_wait). systemd then SIGKILLed incus after waitready(600s)+stop-sigterm(330s) and restarted it, forever. Each ~15.5min cycle leaked a veth pair and a hung "forknet dhcp" process. All 8 production domains were down from Aug 5 21:10 until manual recovery Aug 6 ~15:00. pi-node-004 was unaffected because its rebind completed (proxies listen on 10.10.100.18).

WHY THIS NEEDS A STRUCTURAL FIX (not just the recovery that was done)

  1. Rebind is not transactional or health-gated. The tunnel IP changed while dependent consumers (incus proxy devices) still referenced the old IP. If any step fails or the node is unhealthy, the node is left half-renumbered and hard-down, with no rollback and no alert.
  2. The rebind rollout had no failure alerting. hydracluster polled for DONE for 18h while the node's heartbeats were ALSO failing; nobody was notified. Issue #445 had to be filed by a human noticing the outage.
  3. No edge/domain monitoring: 8 public domains returned 5xx for 18h without an automated issue being raised.
  4. Failure amplification by incus 7.2: a single bad proxy device turns into a daemon-wide deadlock because a container stop during startup-autostart deadlocks incusd (upstream bug; incus 7.3 contains a related fix, commit 629a88537 "forknet: Wait up to 5s for initial DHCP configuration"). Consider upgrading the fleet and/or reporting the stop-hook deadlock upstream.

PROPOSED STRUCTURAL FIXES

  • Make node rebind atomic and ordered: update proxy devices + verify binds BEFORE/atomically-with switching the tunnel address; verify domain health after; roll back to the previous address on failure.
  • Gate rollouts on node health (DNS resolution, hydranode exec channel, incus API responsive) and alert when a rebind job has not completed within N minutes.
  • Alert on sustained heartbeat loss + sustained edge 5xx for registered domains (auto-file a hydraissue).
  • Upgrade incus beyond 7.2 on pi nodes / report the startup-autostart stop-hook deadlock upstream.

RECOVERY ALREADY APPLIED ON PI-NODE-003 (2026-08-06 ~15:00)
Wedged daemon killed, 16 leaked veth pairs + hung forknet processes cleaned, systemd-resolved restarted (DNS fixed), all 8 hydraskin-expose devices rebound to tcp:10.10.100.19:, verified locally + across the mesh + publicly. /root/rebind.log DONE marker written so the cluster can finish its side. 7/8 domains verified healthy; bxl1.hydramirror.experiencenet.com still returns 503 at the edge (backend registration for that vhost appears stale cluster-side - the app itself is healthy, and bxl1-b on pi-node-004 passes through fine), expected to clear once the cluster processes the rebind completion; if not, the edge backend registration for bxl1.hydramirror needs manual refresh.

UPDATE 2026-08-06 16:00: the bxl1.hydramirror 503 tail is RESOLVED (#448: stale static route on the edge, pinned pre-renumbering; #450: health_path now flows dynamically through hydraskin v0.10.0 -> hydracluster v2.0.102 -> hydrascalerouter v0.10.0, static pins removed). All 8 pi-node-003 domains verified healthy at the edge. Note for the structural fixes above: static edge routes were part of the failure surface - they always win conflicts against node reports and go stale silently; scales should publish via labels only (see hydrascalerouter runbook). The four structural items (atomic rebind, health-gated rollout + alerting, edge 5xx auto-filing, incus >7.2 upgrade) remain open and are the scope of this issue.

SPLIT 2026-08-06: the four structural items are now separate groomed issues:

  • #452 (hydracluster): atomic, ordered rebind with rollback
  • #453 (hydracluster): health-gated rollouts + stuck-job and heartbeat-loss alerting
  • #454 (hydrascalerouter): sustained edge 5xx alerting with hydraissue auto-filing
  • #455 (hydraskin): incus upgrade beyond 7.2 / upstream deadlock report
    The recovery and the 503 tail were already resolved (#448, #450). Closing this as the umbrella.

Session Context

Venue
ad6