HydraIssues

hydrarelease updater: backup filename grows unbounded (.backup.backup.backup...) until ENAMETOOLONG
done bug Project: Reporter: 11 May 2026 16:47

Description

Symptom

On the hydraguard hub (141.227.136.199), the auto-updater has been silently failing every 6h for over a month, leaving the host stuck on v1.10.2 while v1.10.7 was available. The hub was only updated today (2026-05-11) via a manual hydraguard update invocation.

Journal evidence (excerpt)

May 11 12:21:11 hydraguard[103029]: [updater] updating 1.9.10 -> 1.10.7
May 11 12:21:11 hydraguard[103029]: Downloading hydraguard v1.10.7 for linux/amd64...
May 11 12:21:11 hydraguard[103029]: [updater] auto-update failed: downloading binary:
  open /usr/local/bin/hydraguard.backup.backup.backup.backup.backup.backup.backup.backup
  .backup.backup.backup.backup.backup.backup.backup.backup.backup.backup.backup.backup
  .backup.backup.backup.backup.backup.backup.backup.backup.backup.backup.backup.backup
  .backup.backup.backup.update: file name too long

The filename has accumulated 35+ chained .backup suffixes — each auto-update cycle appears to append another .backup to whatever the previous attempt left behind, rather than overwriting a fixed-name backup. After roughly 35 cycles the filename hits the filesystem's 255-byte path-component limit and every subsequent update fails the same way.

Likely root cause

In hydrarelease/pkg/updater, the backup-during-download logic is probably something like:

backupPath := binaryPath + ".backup"
os.Rename(binaryPath, backupPath)
// ... download to binaryPath + ".update" ...

But on retry (or repeated download-failure cycles), if binaryPath was renamed to backupPath and then later restored, the in-progress logic seems to use the previous .backup as the basis rather than the canonical install path. Each iteration's 'previous file' is the already-backed-up file, so .backup accrues. Need to trace exactly which path produces this — could be a .update partial that gets renamed-with-suffix on next cycle, could be a backup-of-backup loop on rollback.

Impact

  • Auto-update on every host running the hydrarelease updater is one prolonged failure cycle away from being broken indefinitely.
  • Hub stuck on stale version for ~6 weeks unnoticed.
  • Bug is silent: the updater logs the failure but no alerting wired to it; only discovered while manually updating for a different reason.

Affected hosts (potentially)

Every host using github.com/cederikdotcom/hydrarelease/pkg/updater:

  • hydraguard hub (confirmed broken)
  • hydraheadflatscreen (Windows/macOS/Linux kiosks)
  • hydraneck (manually updated today; unclear if auto-update had been working)
  • Any other service with the updater wired in

Suggested fix

Use a fixed canonical backup path (/usr/local/bin/<name>.backup — always, never with suffix accumulation). On each update cycle: rm -f <name>.backup; mv <name> <name>.backup; download to <name>.update; rename to <name>. Keep last-good backup at exactly <name>.backup, never beyond.

Alternative: timestamped backups in a sibling backups/ directory with a max-N retention policy (same pattern just landed in hydraguard pkg/mesh for mesh.yaml, see commit dad1268).

Verification

  1. Reproduce locally: run the updater in a loop, force the download step to fail, observe filename growth.
  2. After fix: same loop, observe the backup filename stays at <name>.backup regardless of how many cycles run.
  3. Deploy fix and watch one auto-update cycle on the hub succeed end-to-end via journal logs.

Custom Fields

affected_repo
hydrarelease (pkg/updater)
blast_radius
every service using the shared updater library
discovered_via
manual hydraguard update on hub 2026-05-11
related_issue
143-ish (updater wrong service name)

Comments (1)

api 11 May 2026 18:02

Fixed in hydrarelease v1.17.3 (commit d5aab4e).

Added canonicalInstallPath() in pkg/updater/updater.go that strips trailing .backup suffixes from os.Executable() before computing the install + backup targets. The new binary always lands at the canonical name; the .backup file rotates in place rather than chaining. Missed-restart cycles now become benign — the next update overwrites the canonical and the single backup regardless of how many cycles the old process survived.

Unit-tested with 7 cases including the actual hub-observed pathology (35+ chained .backup suffixes).

Propagated to hydraguard via go.mod bump (v1.10.9). Brussels hub now running v1.10.9; future auto-updates rotate cleanly. ~124MB of historical chained binaries from April are still on /usr/local/bin/ (cosmetic only, not growing) and will be removed during the next operator-driven cleanup.