HydraIssues

hydrabody Windows scheduled task does not auto-restart on crash; Sunshine prep-cmd then fails with Error 0
closed bug Project: hydrabody Reporter: cederik 30 Apr 2026 14:00

Description

Symptom

Today (2026-04-30) on cosmic-pretzel-98, Sunshine returned Failed to start the specified application (Error 0) to Moonlight for every experience launch, regardless of which experience or which head was streaming. Mercator and Rupelmonde both failed identically with that error popup on the kiosk display.

Root cause

hydrabody had crashed (or otherwise stopped) at some point earlier and the HydraBody scheduled task was reported Status: Ready / Next Run Time: N/A — i.e. not running, not scheduled to retry. The hydrabody process simply was absent.

Sunshine's prep-cmd for every experience in apps.json is:

curl.exe -s -X POST http://localhost:47991/api/v1/stream/started -H "Content-Type: application/json" -d "{\"app\":\"<name>\"}"

That endpoint is served by hydrabody on :47991. With hydrabody down, the curl returns no-listener (TCP RST). Sunshine treats prep-cmd failure as launch failure and reports the generic Error 0 to the kiosk.

Recovery: schtasks /end /tn HydraBody && schtasks /run /tn HydraBody brought hydrabody back, after which streams worked immediately.

Two-part fix

  1. Auto-restart on crash: the HydraBody scheduled task needs If the task fails, restart every X minutes (Settings → restart on failure) configured. hydrabody's installer (internal/cli/install_windows.go or wherever the schtasks /create invocation lives) should add /RI 1 /Z /F or set RestartCount and RestartInterval programmatically via the Task Scheduler API. Also worth setting MultipleInstancesPolicy=IgnoreNew so duplicate launches don't pile up.

  2. Sunshine prep-cmd resilience: even with auto-restart, there's a window during boot/restart where hydrabody is unavailable. Two options:

    • Make the prep-cmd curl non-blocking (treat HTTP error as success) so hydrabody's notification is best-effort, not a launch gate.
    • Or have hydrabody be the one that runs Sunshine entirely (process supervisor relationship), so Sunshine simply isn't running when hydrabody isn't — visitors get a clean "no body" error instead of a confusing per-app launch failure.

The auto-restart is the immediate fix; the prep-cmd resilience is the structural fix.

Verification

Repro: Stop-Process -Name hydrabody -Force on a body, then trigger a stream from any kiosk. Expect: Error 0 popup. After fix: scheduled task should restart hydrabody within ≤1 minute, Sunshine prep-cmd succeeds, stream launches cleanly.

Related

  • Today's session at cosmic surfaced this as a side-effect of investigating venue planting; the reboot we did yesterday for #126 plus the hydrabody operations for that fix may have left the scheduled task in a stuck state.

Recurrence on boom-pickle-38 (rupelmonde) 2026-05-06 → 2026-05-07. Same failure shape: hydrabody crashed (exit code 2), scheduled task moved to Ready / Next Run Time: N/A, body silently offline ~34h until user-visible failure surfaced as #138 (Castleviewer: Failed to find application rupelmonde-castle-viewer). Recovered with schtasks /run /tn HydraBody. Second occurrence on a different body in ~7 days — confirms this is fleet-wide, not cosmic-specific. Bumping priority signal.


Third occurrence on boom-pickle-38 (rupelmonde) 2026-05-11. Body found with hydrabody.exe not running (tasklist empty). Manually restarted via schtasks /run /tn HydraBody. 4 days after the #138 recurrence (which itself was 7 days after the original cosmic-pretzel incident). Reporter filed it as #143; closed as duplicate of this one. Now actively picking up the fix.


Fix shipped in hydrabody v1.11.31 (2026-05-11)

Root cause was the scheduled-task XML: <BootTrigger> had a <Repetition> with <Interval>PT1M</Interval> but no <Duration>. Windows Task Scheduler silently collapses that into a one-shot, so after the boot fire the task moved to Ready / Next Run Time=N/A. <RestartOnFailure> did not save us — it only attaches to an actively-running trigger window, which we never had.

Fix: replace with a recurring <TimeTrigger> (StartBoundary 2020-01-01, Repetition Interval PT1M, Duration P9999D, StopAtDurationEnd false). Combined with the unchanged MultipleInstancesPolicy=IgnoreNew, redundant fires are silently dropped while hydrabody is alive; a crashed hydrabody comes back within ≤60s with no operator intervention.

XML consolidated into pkg/provider/task_xml.go so both ensureInstall (re-applied on every startup, self-heals the live fleet on next auto-update tick) and hydrabody install use the same source. The prep-cmd resilience part of the original two-part fix is deferred — not needed once auto-restart works, since the Sunshine prep-cmd hook only fails during the ≤60s recovery window.

Verified on boom-pickle-38: killed hydrabody (PID 9748) at 14:32:22, new process (PID 12128) up at 14:33:01 — 39 seconds, well within the 60s target. Kiosk mode re-enabled immediately on restart.