HydraIssues

HydraNode scheduled task has the same broken Repetition that #134 fixed for HydraBody
closed bug Project: hydranode Reporter: Cederik 11 May 2026 18:21

Description

Problem

internal/cli/body/install_windows.go:38-45 defines the HydraNode scheduled task with:

<BootTrigger>
  <Repetition>
    <Interval>PT1M</Interval>
    <StopAtDurationEnd>false</StopAtDurationEnd>
  </Repetition>
</BootTrigger>

The <Repetition> block is missing <Duration>. Windows Task Scheduler silently collapses that to a one-shot, so after the single boot fire there's no recurring trigger. <RestartOnFailure> only attaches to a live trigger window so it can't save us. If hydranode crashes, the task moves to Ready / Next Run Time = N/A and stays dead until the machine reboots.

This is structurally identical to issue #134 (closed) which fixed the exact same XML shape on the HydraBody task. The two install paths grew from the same template.

Why it's WORSE than #134

For HydraBody (#134): when hydrabody crashes, hydranode is still alive → the cluster's exec channel still works → an operator can remotely run schtasks /run /tn HydraBody to recover. Survivable.

For HydraNode: when hydranode crashes, the exec channel dies with it. No remote recovery path — the operator MUST physically reach the machine. This is also reflected in our feedback_never_stop_hydranode operational rule.

Combined with #146 (body.yaml could be wiped by Windows updates), this is the single remaining mechanism by which a body can go permanently dark without remote recovery.

Verified in the live fleet

cosmic-pretzel-98 right now: schtasks /query /tn HydraNode /xml shows the broken shape — <BootTrigger> is the only trigger, <Repetition> has no <Duration>. Not yet observed crashing in the wild, but the failure mode is one process exit away.

Expected fix

Apply the same fix as hydrabody v1.11.31:

  1. Replace the <BootTrigger> Repetition with a recurring <TimeTrigger> (<StartBoundary>2020-01-01T00:00:00</StartBoundary>, <Repetition><Interval>PT1M</Interval><Duration>P9999D</Duration><StopAtDurationEnd>false</StopAtDurationEnd></Repetition>). Keep the BootTrigger for first-start-at-boot.
  2. Keep MultipleInstancesPolicy=StopExisting — different from hydrabody's IgnoreNew. With StopExisting, the recurring TimeTrigger would kill and restart hydranode every minute, which is bad. So we MUST switch to IgnoreNew simultaneously, mirroring the hydrabody fix exactly.
  3. ensureInstall-style re-application on every startup is also worth considering for hydranode — hydrabody has it; hydranode currently re-registers only via hydranode install. Without ensureInstall, the existing fleet's broken XML stays broken until the next manual install. Hard requirement for a self-healing rollout.
  4. Tag and ship; auto-update propagates; the new XML self-applies if ensureInstall is added.

Related

  • #134 (closed) — same root cause class on HydraBody task
  • #146 (closed) — the other half of "a body can go dark without remote recovery". Together with this issue, those were the two failure modes that required operator on site.
  • feedback_never_stop_hydranode (memory) — operational rule that exists because of this fragility

Fix shipped in hydranode v1.10.27 (2026-05-11)

Applied the same fix pattern as hydrabody v1.11.31:

  • New <TimeTrigger> with <StartBoundary>2020-01-01T00:00:00</StartBoundary> + Repetition Interval PT1M / Duration P9999D / StopAtDurationEnd false — recurring trigger that fires every minute essentially forever.
  • Kept <BootTrigger> for first-start-at-boot.
  • Changed MultipleInstancesPolicy from StopExisting to IgnoreNew — with a recurring 1-minute trigger, StopExisting would kill and restart hydranode every minute. With IgnoreNew, redundant fires are silently dropped while hydranode is running; a crashed hydranode comes back within ≤60s.
  • XML consolidated into pkg/body/task_xml.go (HydraNodeTaskXML) shared by ensureInstall (re-applied on every startup so existing fleet self-heals on next auto-update tick) and the CLI hydranode install path.

Trade-off (matches hydrabody v1.11.31): the shared updater's post-update schtasks /Run becomes a no-op while hydranode is running, so new binaries land on disk but activate on next restart (boot, crash, or manual cycle). For stability fixes that's fine; for feature updates it's a minor lag.

Verified on cosmic-pretzel-98:

  1. Auto-update to v1.10.27 landed; schtasks /Query /TN HydraNode /XML now shows <TimeTrigger> with Duration P9999D and IgnoreNew policy.
  2. Stop-Process -Name hydranode -Force to kill the running process.
  3. New hydranode process (PID 6264) up within seconds; cosmic stayed online throughout — cluster never saw a missed heartbeat.
  4. No operator intervention required.

This closes the last remaining failure mode where a body could go permanently dark without remote recovery. Combined with hydrabody v1.11.31 (#134) and hydranode v1.10.25-26 (#146), the only operator-on-site case left is the catastrophic both-body.yaml-locations-wiped scenario, which is acceptably rare.