HydraIssues

pkg/updater on Windows: update --force from a separate process does not restart the service and wedges the in-process auto-updater
closed bug Project: hydrarelease Reporter: 20 Aug 2026 12:33

Description

Context

Found 2026-08-20 during the hydrabody v2.0.67 rollout (#537). Applies to every Windows scheduled-task service that uses pkg/updater (hydrabody, hydranode, and friends). Full write-up: hydrarelease docs/runbooks/runbook-updater-library.md, section "Service didn't restart (Windows scheduled task)".

Defect 1: CLI update path never restarts the service

restart_windows.go assumes the updater runs inside the service process: it calls os.Exit(0) and relies on the task's 1-minute recurring TimeTrigger to revive the task with the new binary. The runbook's documented remote deploy path runs " update --force" as a SEPARATE process via hydracluster exec or SSH. There, os.Exit(0) ends the CLI, not the service. The preceding "schtasks /Run" is a no-op because MultipleInstancesPolicy=IgnoreNew ignores it while the task instance runs. Result: binary replaced on disk, service keeps running the old image, and " version" via exec lies about what is running (it reports the on-disk binary).

Defect 2: the forced update wedges the service's own auto-updater

The CLI's update renames the running service's image to .backup. Windows permits renaming a running image but not deleting it. From then on the running old service's PerformUpdate always fails: os.Remove(backupPath) fails silently (error ignored, updater.go ~line 199), and os.Rename(installPath, backupPath) returns "Access is denied". With the cluster signaling the new version each status tick, the log loops once per minute:

body update available: <version>
[updater] triggered check: backing up current version: rename ... Access is denied.

The service never restarts on its own. Observed simultaneously on cosmic-pretzel-98, boom-pickle-38, and chunky-turnip-23; recovered by bouncing the HydraBody task (schtasks /End then /Run) on each.

Proposed fix in pkg/updater

  1. restartService (Windows): detect whether the current process is the service itself. If not (CLI path), run "schtasks /End /TN " followed by "schtasks /Run /TN " instead of os.Exit(0). If in-process, keep the exit-and-let-the-trigger-revive behavior. A simple detection: compare the current PID against the task's running instance, or pass an explicit flag from the serve command.
  2. PerformUpdate: when os.Rename(installPath, backupPath) fails because backupPath is a locked running image, fall back to a fresh name (for example .backup-) instead of failing, and clean stale backups when they become deletable.
  3. Stop ignoring the os.Remove(backupPath) error silently; log it.

Workaround until fixed

After any "update --force" via exec on a Windows node, bounce the task:

schtasks /End /TN <TaskName>; Start-Sleep -Seconds 3; schtasks /Run /TN <TaskName>

Resolution (2026-08-20)

Fixed in hydrarelease pkg/updater v1.20.0 (commit 89126d2):

  1. restartService on Windows now runs schtasks /End then /Run. An in-process caller dies at /End and the task's TimeTrigger revives it; a separate-process caller (update --force via exec or SSH) restarts the task within seconds. Failures return errors instead of hiding behind os.Exit(0).
  2. PerformUpdate falls back to a timestamped .backup- name when .backup is a locked running image, sweeps stale fallbacks after a successful install, logs the previously ignored os.Remove error, and canonicalInstallPath strips the new suffix.

Shipped to the fleet in hydrabody v2.0.68 (dependency bump) and validated end to end with hydrabody v2.0.69, canary on chunky-turnip first:

  • CLI path: update --force via exec on chunky-turnip-23 restarted the task within seconds (fresh process 15:05:52, v2.0.69).
  • In-process path: cosmic-pretzel-98 and boom-pickle-38 self-updated via server trigger; /End + TimeTrigger revival brought both up at 15:08:01 on v2.0.69.
  • No Access-is-denied wedge loops since the fixed updater took over. Single Sunshine instance on every body.

Other services (hydranode v1.17.5, hydraguard, etc.) still ship older updaters and keep the old behavior until they bump hydrarelease; the runbook documents the manual task bounce for them.