HydraIssues

hydraskin Incus bridge subnet collision between pi-node-001 and pi-node-004-nvme (10.200.86.0/24)
open bug Project: hydraskin Reporter: cederik 4 Aug 2026 11:14

Description

Two live nodes have collided on the same derived Incus bridge subnet:

  • pi-node-001 (node-2b224f6a): hydrabr0 is 10.200.86.1/24, read live from the host. hydracluster's inventory does not even have a skin_bridge_subnet value recorded for this node (the field is simply absent from its API record), so this collision was invisible from hydracluster alone — it only surfaced by reading the live interface on the node.
  • pi-node-004-nvme (node-50ab5309): hydracluster records skin_bridge_subnet: 10.200.86.0/24.

For contrast, pi-node-003-nvme (node-6bf57aed) is on 10.200.95.0/24 — no collision there. This is investigation/filing only per instructions; the derivation itself has NOT been changed and no bridge has been renumbered.

Root cause (hydraskin/internal/install/bridge.go, DeriveBridgeSubnet): the node's /etc/machine-id is hashed with 32-bit FNV-1a, and only hash % 254 of that hash is used, picking the third octet of 10.200.X.0/24 (X in 1..254, wrapping/skipping to avoid the node's own already-configured local networks). That is the entire coordination mechanism — there is only ONE global namespace of 254 possible subnets shared by the whole fleet, and each node picks independently with no lookup of what other nodes already have. hydracluster's own API code (pkg/api/handlers_api.go, SkinBridgeSubnet field comment) documents this explicitly: 'The node derives it and reports it; the cluster only observes. Surfaced so two nodes that hashed to the same subnet can be spotted.' I.e. detection today is 100% manual — someone has to notice two nodes report the same value by reading the API/host state, exactly how this collision was found. There is no automated alert, no uniqueness check at derivation time, and no uniqueness check at enrollment/report time.

I verified this is a genuine hash collision between two distinct machine-ids, not a cloned-image duplicate machine-id (a common Pi provisioning trap): sha256sum of /etc/machine-id differs between pi-node-001, pi-node-003-nvme, and pi-node-004-nvme — three different values. So the two nodes both legitimately hashed into the same 1-in-254 bucket.

Collision probability across the fleet: today only 3 nodes actually carry the hydraskin role (node-2b224f6a/pi-node-001, node-6bf57aed/pi-node-003-nvme, node-50ab5309/pi-node-004-nvme; a 4th Pi, pi-node-002-c/node-5dcde910, exists but currently has no roles assigned). With n=254 equally likely slots, the expected probability of any collision among just k=3 nodes is only 1-(253/254)(252/254) ≈ 1.2% — so hitting a collision this early is a real if unlucky outcome, not something to shrug off as 'basically never happens.' The general birthday-paradox curve for this namespace (P ≈ 1 - e^{-k(k-1)/(2*254)}): 50% probability of at least one collision is reached at roughly k≈19 nodes, and by k≈30 nodes it is roughly 80%. Any fleet growth toward that range will make collisions the norm, not the exception, under the current 254-slot scheme.

Detectability today: none, beyond manually diffing hydracluster's reported skin_bridge_subnet per node against each other (and even that misses cases like pi-node-001 where the field isn't populated at all — its collision was only found by reading the live host). Harmless today because these bridges are node-local and NAT'd (per bridge.go's own doc comment and hydracluster's field comment), but it defeats the purpose of deriving a unique per-node subnet and would become an active routing conflict if anything ever bridges/routes between nodes.

Not fixed here per instructions — renumbering a bridge on a live node moves running containers' addressing and is not a background-agent change. Flagging as a design gap: the 254-slot keyspace is too small for the fleet's expected growth, and there is no server-side (hydracluster) uniqueness check despite hydracluster already being the natural place to arbitrate it (it already receives skin_bridge_subnet in node reports).