HydraIssues

hydranode body.yaml has a single point of failure; loss bricks the body
closed bug Project: hydranode Reporter: Cederik 11 May 2026 15:15

Description

Problem

hydranode auth state (server_url + token) lives in exactly one place: ~/.hydranode/body.yaml. When running as SYSTEM on Windows that is C:\Windows\System32\config\systemprofile\.hydranode\body.yaml — a path that is routinely touched by Windows updates, AV cleanups, profile resets, and sfc /scannow. If the file is gone, hydranode falls back to enroll-polling mode forever, the body never heartbeats, and from the cluster's perspective the body is offline with no remote-recovery path.

Evidence

Today (2026-05-11) on cosmic-pretzel-98: .hydranode\ directory under systemprofile was simply absent. Only C:\hydranode\\enroll.yaml and the binaries remained. Required physical access to recover (operator at the venue had to write a body.yaml by hand from a token I pulled out of the cluster's nodes.yaml via SSH). Manual recovery worked but is the kind of operation that should never need an operator on site.

Constraints

  • The cluster cannot reach into a body without the body already having a valid token, so the recovery path has to be initiated body-side.
  • WDAC blocks new binary hashes on bodies, so the fix has to be deliverable through the existing hydranode auto-update mechanism.
  • Has to work on Windows + Linux + macOS; the SYSTEM-profile problem is Windows-specific but the durability principle applies everywhere.
  • Must not weaken the security model — the token is still a sensitive credential.

Proposed direction (to be refined in implementation plan)

Defense in depth, not a single new location:

  1. Multiple-location persistence: when hydranode obtains a working token (install, or successful first heartbeat), write it to BOTH ~/.hydranode/body.yaml and a fallback (next to the binary at C:\hydranode\body.yaml on Windows / /etc/hydranode/body.yaml on Linux). Optionally also the Windows registry as a tertiary backup.
  2. Read-with-fallback on startup: env var → CLI flag → primary file → fallback file → registry → enroll mode. First location that succeeds wins; missing primaries get rewritten from a surviving fallback (self-heal).
  3. Cluster-side token reissue for the worst case where every body-side copy is gone: an admin endpoint POST /api/v1/nodes/{id}/reissue-enroll-token that mints a one-shot enrollment token tied to the existing node id, so hydranode reinstall --node-id <id> --enroll-token <one-shot> rebuilds locally without creating a duplicate node record.

Steps 1+2 close the operational gap (single-point-of-failure becomes triple-point-of-failure with auto-repair). Step 3 is for the catastrophic case and replaces the SSH-into-cluster + hand-write-body.yaml dance.

Related

  • Issue #138 (closed) — same root cause class on boom-pickle (different file, hydrabody apps state)
  • Issue #134 (closed) — auto-restart for hydrabody. This issue is the analogous fix for hydranode's auth state.

Fix shipped in hydranode v1.10.25 + v1.10.26 (2026-05-11)

body.yaml is now written to two locations and read with fallback:

  • Primary: ~/.hydranode/body.yaml (existing, SYSTEM profile on Windows)
  • Backup: C:\hydranode\body.yaml (Windows), /etc/hydranode/body.yaml (Linux/macOS) — outside any profile dir

On every hydranode startup: read primary; if missing, read backup, restore the primary from the backup, log [config] primary ... missing, restoring from backup .... If primary is read successfully, mirror it to the backup so they never drift. New helpers in pkg/provider/... — GetBackupConfigPath, loadBodyConfigWithFallback, mirrorIfDifferent. Install path writes both locations up front; uninstall paths already clean both.

v1.10.26 follow-up: opened the log file before loadBodyConfig so the recovery diagnostic actually lands in hydranode.log instead of stderr (which the Windows SYSTEM scheduled task discards).

Verified on cosmic-pretzel-98 (the body that hit the original failure):

  1. Confirmed both files exist with identical hash after auto-update to v1.10.26.
  2. Remove-Item C:\Windows\System32\config\systemprofile\.hydranode\body.yaml (Test-Path = False after).
  3. schtasks /run /tn HydraNode to restart.
  4. New process started at 20:02:49, log shows [config] primary ... missing, restoring from backup C:\hydranode\body.yaml, primary recreated with correct content, heartbeat succeeded at 20:02:50. Total recovery time: 2 seconds.
  5. cosmic stayed online throughout — no operator file-writing needed.

Before: SYSTEM-profile wipe required SSH-into-cluster + hand-writing body.yaml on the body. After: hydranode self-heals within 2 seconds of the next startup. The catastrophic both-gone case is still manual SSH recovery (acceptable, much rarer now).