HydraIssues

Network recovery escalation reboot-loops venue bodies (GRM outage 2026-09-21, #757)
closed unclassified Project: hydranode Reporter: 21 Sep 2026 11:06

Description

The recovery escalation in pkg/body/recovery.go took the Gallo-Romeins venue body (fluffy, node-11da9ea3) down hard on 2026-09-21: mid-VR-session heartbeat failures escalated to adapter restart (threshold 12) and then reboot (threshold 20, 60s warning, 5-min cooldown). The box has fragile autologin, the network stayed broken, so it reboot-looped from 08:23 UTC on - staff photographed the shutdown dialog appearing seconds after each login (#757).

Design problems, all confirmed in source:

  1. LocalNetworkBroken() (pkg/body/netdiag.go:25) is !GatewayReachable || !UDPOutbound || !DNSWorks. A venue WAN/DNS outage with a perfectly healthy LAN therefore authorizes adapter restarts and reboots. DNS-only failure must never count as local-broken.
  2. rebootMachine() on a body that needs interactive autologin turns one outage into an unbounded reboot loop. Reboot escalation must be opt-in (per role or config), and default OFF for hydrabody/venue nodes. The fleet rule is already "never reboot bodies remotely" - hydranode itself violates it.
  3. restartNetworkAdapter() (recovery_windows.go) does netsh disable, sleep 3s, enable - the disable is PERSISTENT across reboots, so a failed re-enable permanently kills the NIC. Must verify the enable succeeded and retry/roll back.
  4. detectActiveAdapter() returns the first Connected interface - on a body running a Mobile Hotspot plus WireGuard there are several candidates, and flapping the Ethernet that backs the ICS hotspot kills a live headset session (likely what ended the GRM demo).
  5. The escalation keeps firing forever on the 5-min cooldown. It needs a ceiling (e.g. one reboot max, then back off and wait).

Venue impact and on-site mitigation steps are on #757 (shutdown /a, re-enable adapter or disable the HydraNode task until the network is fixed).

Comments (3)

api 21 Sep 2026 11:35

Implemented in hydranode 87153fc, released as v1.10.42 (CI green: https://github.com/cederikdotcom/hydranode/actions/runs/35594603819). Fleet picks it up via auto-update from releases.experiencenet.com.

Changes:

  1. netdiag: DNS removed from the LocalNetworkBroken gate. It now checks gateway reachability + UDP outbound only, so a DNS-only failure (venue WAN outage, healthy LAN) can never authorize destructive recovery. ServerUnreachableOnly now classifies a DNS failure as upstream. DNS is still probed and logged in the diagnostics line.

  2. recovery: the reboot escalation is now OPT-IN via recovery_reboot: true in body.yaml, DEFAULT FALSE. With the flag off (the entire fleet, until someone opts a node in) the escalation ceiling is the adapter restart. Even when enabled there is a hard ceiling of ONE reboot per process lifetime; after that the node logs and stays at diagnostics level. The gate lives in recoveryState.nextAction/evaluate, which is the only rebootMachine caller, so Windows, Linux and macOS are all covered.

  3. recovery_windows: netsh 'set interface X disable' persists across reboots, so a failed re-enable used to permanently kill the NIC. The enable is now verified against the actual admin state (netsh interface show interface) and retried: 3 quick attempts with backoff, then background retries every 30s until the adapter is back, with loud CRITICAL log lines. It never returns leaving the adapter disabled.

  4. detectActiveAdapter no longer grabs the first Connected interface. Virtual/overlay adapters (Wi-Fi Direct, WireGuard, vEthernet, Bluetooth, Loopback, Local Area Connection*) are filtered out, type Dedicated is preferred, and if more than one candidate remains the adapter restart is SKIPPED with a log line instead of guessing.

Unit tests added for the DNS-only classification, the reboot gating and one-reboot ceiling, and the adapter candidate filtering (parsing extracted to recovery_netsh.go, tested on all platforms). go build / vet / test pass; runbook and testbook updated.

Default behavior change: reboots are now opt-in. No node reboots as part of network recovery unless its body.yaml sets recovery_reboot: true. Leaving this issue open pending verification on the live fleet.

api 22 Sep 2026 07:43

VERIFIED LIVE on the fleet. fluffy (node-11da9ea3) auto-updated to v1.10.42 at 07:38 UTC 2026-09-22 and logged at startup: "network recovery: reboot escalation disabled, ceiling is adapter restart". Before the update it reboot-looped every ~11 min on repeated DNS-lookup failures for hydracluster.experiencenet.com (classic DNS-only-failure -> false local-broken -> reboot). Post-update the loop is gone: recovery counter resets on the first successful heartbeat, reboots are opt-in/off, DNS-only failures take the no-action early return. Leaving open only to confirm the same on the other venue bodies as they cycle.

api 22 Sep 2026 14:49

Closing. Fix released as hydranode v1.10.42 and verified end to end: fluffy auto-updated, reboot escalation disabled (startup log confirmed), and the Gallo-Romeins experience ran a live PICO 4 Ultra Enterprise streaming session afterwards (both eyes H.265, 90 Hz). All THREE live venue bodies (gallo-romeins, rupelmonde, cloud-seven) are on v1.10.42. Remaining non-venue nodes (chunky v1.10.39 test/build box, plus offline wobbly-llama and MSI1060) will pick up the fixed agent on their next update cycle - no venue impact. Design changes shipped: reboots opt-in and OFF by default, DNS-only failure triggers no recovery, adapter re-enable verified/retried, safe adapter selection on hotspot/WireGuard boxes, one-reboot-per-lifetime ceiling.