HydraIssues

Make mesh renumbering rebind atomic and ordered, with rollback
open improvement Project: hydracluster Reporter: anonymous 6 Aug 2026 16:44

Description

Split from #447 (item 1 of 4). During the 2026-08-05 mesh renumbering (10.0.0.0/24 -> 10.10.0.0/16), the hydraguard tunnel address switched before the node-side rebind of the 8 hydraskin-expose proxy devices ran. On pi-node-003 the rebind exec never arrived (broken DNS), leaving the node half-renumbered: every container start bound the now-nonexistent 10.0.0.4 and failed, and the node was hard-down 18h (#445).

Required ordering: update proxy devices and verify they bind BEFORE or atomically with switching the tunnel address; verify domain health at the edge after; roll back to the previous address if any step fails. A rebind must be a transaction with a rollback path, not a fire-and-forget exec plus a DONE-marker poll.