All 3 iPads disconnected at Nerdland 2026-05-21 due to the stale-session watchdog in pkg/provider/stale_session_watchdog_windows.go:tickStaleSession().
Network packet loss hit all 3 streams simultaneously. The GameStream TCP control connection on port 48010 (RTSP — the persistent session channel) dropped on all 3 bodies. hydrabodys noClientSince timer reached 2 minutes, called OnStreamEnded for all active sessions, and sent termination reason 0x80030023 to each iPad client.
The v2.0.49 fix (71cf364, closes #300) added port 48010 to the port check and is deployed on all bodies (v2.0.51). It is working correctly in the sense that it now detects the right signal. The problem is not which port is checked — it is what the watchdog does when that signal disappears.
Timeline from iPad logs #336, #337, #338 and the hydrabody log on chunky-turnip-23 (v2.0.51):
Network context: 3 concurrent streams at 25 Mbps each = 75 Mbps through the WireGuard/hydraneck mesh. Issue #328 proposed reducing to 20 Mbps per stream for 3-stream headroom and may not have been applied before this event.
The watchdog was designed for a specific failure mode: Moonlight crashes on the client side, the Sunshine undo prep-cmd hook never fires, and the experience process keeps rendering indefinitely with no one watching. That is a legitimate stale session.
What happened here is the opposite: the experience was actively rendering (GPU at 97-98%), the Sunshine session was alive, and all three iPad clients were online and trying to reconnect. The TCP connection dropped due to a network event, not because the client went away. Killing the session prevented the reconnect that was already underway.
The watchdog cannot distinguish these two cases using only TCP port state:
| Scenario | Port 48010 state | Correct action |
|---|---|---|
| Moonlight crashed, nobody there | not ESTABLISHED | Kill after timeout |
| Network blip, client reconnecting | not ESTABLISHED (temporarily) | Wait, do not kill |
Two recent hydracluster features together provide the server-side ground truth the watchdog needs:
Session registry Phase 1 (hydracluster commit 06c3067, issue #332): Records session open/close events from body status transitions. An open session in the registry means hydracluster believes a stream is legitimately in progress.
Heartbeat-based live body tracking (hydracluster commit 47f68c3): Each iPad head heartbeat carries live_body_id — the body it is currently streaming from. As long as the iPad is online and heartbeating with live_body_id pointing at this body, the head has not left the session.
If the watchdog queries hydracluster before killing — GET /api/v1/heads/{head_id} and checks live_body_id, or GET /api/v1/sessions and checks for an open session on this body — it gets a second opinion that survives a brief TCP drop:
The hydracluster API token is already available to hydrabody (it phones home every tick). This requires no new credentials.
pkg/provider/stale_session_watchdog_windows.go — tickStaleSession():
if sunshineClientConnected() { reset timer; return }
if noClientSince < 2 minutes { return }
// kills all sessions — no server-side check
The 2-minute constant is the only gate between a TCP blip and a dead session.
Option A — Consult hydracluster session registry before killing: when the TCP check fails and the timer is about to expire, call GET /api/v1/heads/{head_id} on hydracluster. If live_body_id is still this body, reset the timer and wait. Only kill when both TCP is gone AND hydracluster confirms the head has left. This is the cleanest fix and leverages the already-deployed session registry.
Option B — Extend the timeout: increase noClientSince from 2 minutes to 5-10 minutes. Simple, no new API call, but still has the fundamental ambiguity between stale sessions and network-recovery windows.
Option A is preferred: it removes the ambiguity entirely, makes the watchdog correct regardless of timeout value, and the required infrastructure (session registry, live_body_id) is already deployed.