HydraIssues

hydrabody: stale-session watchdog 2-minute timeout kills stream during network-recovery window
closed unclassified Project: hydrabody Reporter: Claude (agent) 21 May 2026 21:29

Description

All 3 iPads disconnected at Nerdland 2026-05-21 due to the stale-session watchdog in pkg/provider/stale_session_watchdog_windows.go:tickStaleSession().


What happened

Network packet loss hit all 3 streams simultaneously. The GameStream TCP control connection on port 48010 (RTSP — the persistent session channel) dropped on all 3 bodies. hydrabodys noClientSince timer reached 2 minutes, called OnStreamEnded for all active sessions, and sent termination reason 0x80030023 to each iPad client.

The v2.0.49 fix (71cf364, closes #300) added port 48010 to the port check and is deployed on all bodies (v2.0.51). It is working correctly in the sense that it now detects the right signal. The problem is not which port is checked — it is what the watchdog does when that signal disappears.

Timeline from iPad logs #336, #337, #338 and the hydrabody log on chunky-turnip-23 (v2.0.51):

  • ~23:10 CEST: all 3 streams started normally, GPU at 44-61 C
  • ~14:12 local: unrecoverable video frames and audio block drops hit all 3 iPads simultaneously (network event)
  • Port 48010 ESTABLISHED connection dropped on all 3 bodies
  • noClientSince timer started on each body
  • ~23:13-23:15 CEST: timer expired, watchdog killed each session
  • iPad Moonlight client received 0x80030023 and attempted immediate re-launch — session already dead
  • MicRelay engine restart failed with AVAudio error 2003329396 (known #320)

Network context: 3 concurrent streams at 25 Mbps each = 75 Mbps through the WireGuard/hydraneck mesh. Issue #328 proposed reducing to 20 Mbps per stream for 3-stream headroom and may not have been applied before this event.


Why the watchdog should NOT have killed these sessions

The watchdog was designed for a specific failure mode: Moonlight crashes on the client side, the Sunshine undo prep-cmd hook never fires, and the experience process keeps rendering indefinitely with no one watching. That is a legitimate stale session.

What happened here is the opposite: the experience was actively rendering (GPU at 97-98%), the Sunshine session was alive, and all three iPad clients were online and trying to reconnect. The TCP connection dropped due to a network event, not because the client went away. Killing the session prevented the reconnect that was already underway.

The watchdog cannot distinguish these two cases using only TCP port state:

Scenario Port 48010 state Correct action
Moonlight crashed, nobody there not ESTABLISHED Kill after timeout
Network blip, client reconnecting not ESTABLISHED (temporarily) Wait, do not kill

The existing feature that can resolve this

Two recent hydracluster features together provide the server-side ground truth the watchdog needs:

  1. Session registry Phase 1 (hydracluster commit 06c3067, issue #332): Records session open/close events from body status transitions. An open session in the registry means hydracluster believes a stream is legitimately in progress.

  2. Heartbeat-based live body tracking (hydracluster commit 47f68c3): Each iPad head heartbeat carries live_body_id — the body it is currently streaming from. As long as the iPad is online and heartbeating with live_body_id pointing at this body, the head has not left the session.

If the watchdog queries hydracluster before killing — GET /api/v1/heads/{head_id} and checks live_body_id, or GET /api/v1/sessions and checks for an open session on this body — it gets a second opinion that survives a brief TCP drop:

  • live_body_id still set to this body: head is alive, trying to reconnect — do not kill
  • live_body_id cleared or head offline: head has actually left — safe to kill
  • hydracluster unreachable: fall back to existing TCP-only logic with a longer timeout

The hydracluster API token is already available to hydrabody (it phones home every tick). This requires no new credentials.


Current code (v2.0.51)

pkg/provider/stale_session_watchdog_windows.go — tickStaleSession():

if sunshineClientConnected() { reset timer; return }
if noClientSince < 2 minutes { return }
// kills all sessions — no server-side check

The 2-minute constant is the only gate between a TCP blip and a dead session.


Fix direction

Option A — Consult hydracluster session registry before killing: when the TCP check fails and the timer is about to expire, call GET /api/v1/heads/{head_id} on hydracluster. If live_body_id is still this body, reset the timer and wait. Only kill when both TCP is gone AND hydracluster confirms the head has left. This is the cleanest fix and leverages the already-deployed session registry.

Option B — Extend the timeout: increase noClientSince from 2 minutes to 5-10 minutes. Simple, no new API call, but still has the fundamental ambiguity between stale sessions and network-recovery windows.

Option A is preferred: it removes the ambiguity entirely, makes the watchdog correct regardless of timeout value, and the required infrastructure (session registry, live_body_id) is already deployed.

Session Context

Venue
cloud-seven / nerdland