HydraIssues

Three iPads stream from one body at Sint-Niklaas: body selection race confirmed in production (204 UUID-mismatch lines), plus a session with no head heartbeat never times out
open bug Project: hydracluster Parent: #663 Reporter: cederik 5 Sep 2026 09:55

Description

Confirmed in production, 2026-09-04, Sint-Niklaas Tourism Office

The venue is now on ethernet and all three iPads want mercator-talks. Staff report that several iPads show the same picture. Confirmed: they were all streaming from ONE body.

This is the same defect as #530, reproduced at the same venue with THREE heads instead of two, plus hard log evidence and two follow-on findings.

Evidence

Each iPad carries its own Moonlight client ID:

Head Node Moonlight client ID LAN IP
ipad-head-map-35 node-fe6e7808 0fd573ddf5489dce 192.168.8.230
ipad-head-map-37 node-e0528a7e d6ce9e0d0c06e3a3 192.168.8.165
ipad-head-map-39 node-a57ee1a7 acdd2f298f9f011a 192.168.8.180

All three IDs heartbeat into ONE session on ONE body, node-4c2be4b0 (cosmic-pretzel-98, Cloud Seven), overwriting each other second by second:

11:43:18 [session] opened body=node-4c2be4b0 name=cosmic-pretzel-98
11:43:18 [session] UUID set from head heartbeat uuid=acdd2f298f9f011a      <- iPad 39
11:43:19 [session] UUID mismatch: body=acdd2f298f9f011a head=d6ce9e0d0c06e3a3   <- iPad 37
11:43:21 [session] UUID mismatch: body=d6ce9e0d0c06e3a3 head=0fd573ddf5489dce   <- iPad 35
11:43:22 [session] UUID mismatch: body=0fd573ddf5489dce head=d6ce9e0d0c06e3a3
...

204 UUID mismatch lines on 2026-09-04, ALL on cosmic-pretzel-98, split 93 / 56 / 55 across the three iPads. Two bursts: 11:38-11:39 and 11:43-11:47 UTC.

Why the session API hides this

sessionStore.active is keyed by body ID (pkg/api/session_store.go:32) and openSession returns early if that body already has a session. Two heads on one body collapse into a single record, attributed to whichever head heartbeat arrived first. GET /api/v1/sessions/history therefore shows a clean, non-overlapping timeline while three iPads are piled on one machine. The UUID mismatch log line is the only trace. Anyone triaging from the API alone will conclude there is no problem.

Mechanism (confirms #530 and #308)

  1. Exclusivity is client-side only. hydraheadipad HydraClusterClient.swift:82 takes bodies.first { streamCount == 0 }. Nothing server-side reserves or rejects.
  2. stream_count is stale for up to 30s. It derives from the body's self-reported StreamStatus, and an idle body posts status on a 30s ticker (internal/cli/run.go:74; it drops to 5s only once already streaming).
  3. All three heads receive an identical, identically ordered list. Verified by querying /api/v1/bodies/eligible as each real head: same two bodies, same order. No same-venue body exists, and gpu_memory_total_mb parses to 0 for all bodies, so the VRAM tiebreaker never fires and the ordering is arbitrary but identical for every head.
  4. Sunshine cannot separate them. Every client on cosmic-pretzel-98 presents the hardcoded Moonlight placeholder 0123456789ABCDEF (verified: exactly one distinct uniqueid in the whole sunshine.log). Sunshine sees one paired client and attaches newcomers to the running session. This is #531.

Ethernet made it worse, not better. Faster and more reliable reconnects mean the three iPads now retry close together, which is exactly when the 30s window bites.

Finding 2: a session no head ever adopts has no head-side timeout

checkSessionHealth (pkg/api/session_watchdog.go) computes:

staleHead := !rec.HeadLastHeartbeatAt.IsZero() && time.Since(rec.HeadLastHeartbeatAt) > headOnlyStaleThreshold

The head-stale check is skipped entirely when the head never heartbeated. Body-stale does not fire either while hydrabody keeps heartbeating. So a session that no head ever adopted can only ever be closed by the streaming -> not streaming edge in handleBodyStatus:370.

Observed on session node-4c2-10 (cosmic-pretzel-98), which lived its whole 2h46m with head_last_heartbeat_at: 0001-01-01T00:00:00Z:

  • 13:33:16 UTC Sunshine CLIENT CONNECTED, session opens 13:33:33
  • 14:17:41 UTC Sunshine CLIENT DISCONNECTED
  • ~15:26 UTC hydrabody local API on the body reports {"count":0,"status":"idle"}, while /api/v1/bodies/eligible still advertises stream_count: 1
  • 16:20:02 UTC session finally closes, end_reason: body_idle

Roughly two hours during which the cluster advertised a free body as busy, excluding it from selection. It resolved on its own; no manual clear was needed. But the watchdog hole is real and means the only recovery path for such a session is that single edge transition. If the edge is missed (cluster restart repopulating the in-memory s.bodyStatus map, so previousStatus.StreamStatus reads ""), nothing closes the session at all.

Note also that previousStatus := s.bodyStatus[nc.node.ID] at handlers_body.go:351 reads and writes the map without holding s.mu, unlike every other access.

Finding 3: capacity shortfall

Sint-Niklaas has NO body of its own. All three iPads reach across to cosmic-pretzel-98 (Cloud Seven) and boom-pickle-38 (Rupelmonde). Two bodies, three iPads, one stream each. Even with the race fixed, the third iPad gets noBodyAvailable. Three concurrent Mercator streams need a third slot, whether that is a third body or multi-stream capacity per #504.

Direction

The central capacity system in #504 already specifies the fix: POST /api/v1/bodies/{id}/slots/claim returning {slot, ports, claim_token, ttl 60s}, with the explicit requirement that "claims must count toward advertised availability (slots_free and legacy stream_count), or the claim-to-stream window re-opens today's race". That is exactly this race. hydracluster should be the allocation authority instead of trusting each head to self-police off a stale count.

Interim mitigation that needs no protocol change: mark the body busy in cluster memory the moment a head reports it picked one, mirroring what handleHeadStreamStop already does for the idle direction (handlers_head.go:560-566).

Related

  • #530 same bug, same venue, two iPads, 2026-08-19. This issue confirms it in production with three.
  • #308 the underlying race; notes Reserve()/Release() exists at pkg/store/store.go:467 but is wired only for hydratestnode.
  • #531 hardcoded Moonlight uniqueid on every iPad; removes Sunshine's own last-line defence.
  • #504 central slot allocator; the designed fix.
  • #84 multi-stream via VDD; the capacity substrate.
  • #116 body selection Phase 1, cross-venue eligibility and client failover.
  • #195 streaming monitor endpoint to detect body assignment conflicts.
  • #182 eligibility ignores installed experiences.

CORRECTION 2026-09-05. This issue previously stated that previousStatus := s.bodyStatus[nc.node.ID] at handlers_body.go:351 reads the map without holding s.mu. That is WRONG: requireNodeToken takes s.mu.Lock() with defer across the whole handler (handlers_body.go:45-46), so handleBodyStatus is fully locked.

The real unlocked accesses are elsewhere and are worse: handleEligibleBodies releases s.mu at handlers_api.go:830 and then reads s.bodyStatus at :877, and handlers_web.go:473 lets the map reference escape the lock to be read during template render at :530. Concurrent map read and write is an unrecoverable Go runtime fatal, and its frequency scales with heads x bodies. The correct copy-under-lock idiom already exists 170 lines away at handlers_web.go:307-309. Fixing these is the FIRST commit of the design in #663, ahead of any allocation work.

Session Context

Venue
sint-niklaas-tourism-office