The venue is now on ethernet and all three iPads want mercator-talks. Staff report that several iPads show the same picture. Confirmed: they were all streaming from ONE body.
This is the same defect as #530, reproduced at the same venue with THREE heads instead of two, plus hard log evidence and two follow-on findings.
Each iPad carries its own Moonlight client ID:
| Head | Node | Moonlight client ID | LAN IP |
|---|---|---|---|
| ipad-head-map-35 | node-fe6e7808 | 0fd573ddf5489dce | 192.168.8.230 |
| ipad-head-map-37 | node-e0528a7e | d6ce9e0d0c06e3a3 | 192.168.8.165 |
| ipad-head-map-39 | node-a57ee1a7 | acdd2f298f9f011a | 192.168.8.180 |
All three IDs heartbeat into ONE session on ONE body, node-4c2be4b0 (cosmic-pretzel-98, Cloud Seven), overwriting each other second by second:
11:43:18 [session] opened body=node-4c2be4b0 name=cosmic-pretzel-98
11:43:18 [session] UUID set from head heartbeat uuid=acdd2f298f9f011a <- iPad 39
11:43:19 [session] UUID mismatch: body=acdd2f298f9f011a head=d6ce9e0d0c06e3a3 <- iPad 37
11:43:21 [session] UUID mismatch: body=d6ce9e0d0c06e3a3 head=0fd573ddf5489dce <- iPad 35
11:43:22 [session] UUID mismatch: body=0fd573ddf5489dce head=d6ce9e0d0c06e3a3
...
204 UUID mismatch lines on 2026-09-04, ALL on cosmic-pretzel-98, split 93 / 56 / 55 across the three iPads. Two bursts: 11:38-11:39 and 11:43-11:47 UTC.
sessionStore.active is keyed by body ID (pkg/api/session_store.go:32) and openSession returns early if that body already has a session. Two heads on one body collapse into a single record, attributed to whichever head heartbeat arrived first. GET /api/v1/sessions/history therefore shows a clean, non-overlapping timeline while three iPads are piled on one machine. The UUID mismatch log line is the only trace. Anyone triaging from the API alone will conclude there is no problem.
HydraClusterClient.swift:82 takes bodies.first { streamCount == 0 }. Nothing server-side reserves or rejects.stream_count is stale for up to 30s. It derives from the body's self-reported StreamStatus, and an idle body posts status on a 30s ticker (internal/cli/run.go:74; it drops to 5s only once already streaming)./api/v1/bodies/eligible as each real head: same two bodies, same order. No same-venue body exists, and gpu_memory_total_mb parses to 0 for all bodies, so the VRAM tiebreaker never fires and the ordering is arbitrary but identical for every head.0123456789ABCDEF (verified: exactly one distinct uniqueid in the whole sunshine.log). Sunshine sees one paired client and attaches newcomers to the running session. This is #531.Ethernet made it worse, not better. Faster and more reliable reconnects mean the three iPads now retry close together, which is exactly when the 30s window bites.
checkSessionHealth (pkg/api/session_watchdog.go) computes:
staleHead := !rec.HeadLastHeartbeatAt.IsZero() && time.Since(rec.HeadLastHeartbeatAt) > headOnlyStaleThreshold
The head-stale check is skipped entirely when the head never heartbeated. Body-stale does not fire either while hydrabody keeps heartbeating. So a session that no head ever adopted can only ever be closed by the streaming -> not streaming edge in handleBodyStatus:370.
Observed on session node-4c2-10 (cosmic-pretzel-98), which lived its whole 2h46m with head_last_heartbeat_at: 0001-01-01T00:00:00Z:
{"count":0,"status":"idle"}, while /api/v1/bodies/eligible still advertises stream_count: 1end_reason: body_idleRoughly two hours during which the cluster advertised a free body as busy, excluding it from selection. It resolved on its own; no manual clear was needed. But the watchdog hole is real and means the only recovery path for such a session is that single edge transition. If the edge is missed (cluster restart repopulating the in-memory s.bodyStatus map, so previousStatus.StreamStatus reads ""), nothing closes the session at all.
Note also that previousStatus := s.bodyStatus[nc.node.ID] at handlers_body.go:351 reads and writes the map without holding s.mu, unlike every other access.
Sint-Niklaas has NO body of its own. All three iPads reach across to cosmic-pretzel-98 (Cloud Seven) and boom-pickle-38 (Rupelmonde). Two bodies, three iPads, one stream each. Even with the race fixed, the third iPad gets noBodyAvailable. Three concurrent Mercator streams need a third slot, whether that is a third body or multi-stream capacity per #504.
The central capacity system in #504 already specifies the fix: POST /api/v1/bodies/{id}/slots/claim returning {slot, ports, claim_token, ttl 60s}, with the explicit requirement that "claims must count toward advertised availability (slots_free and legacy stream_count), or the claim-to-stream window re-opens today's race". That is exactly this race. hydracluster should be the allocation authority instead of trusting each head to self-police off a stale count.
Interim mitigation that needs no protocol change: mark the body busy in cluster memory the moment a head reports it picked one, mirroring what handleHeadStreamStop already does for the idle direction (handlers_head.go:560-566).
Reserve()/Release() exists at pkg/store/store.go:467 but is wired only for hydratestnode.CORRECTION 2026-09-05. This issue previously stated that previousStatus := s.bodyStatus[nc.node.ID] at handlers_body.go:351 reads the map without holding s.mu. That is WRONG: requireNodeToken takes s.mu.Lock() with defer across the whole handler (handlers_body.go:45-46), so handleBodyStatus is fully locked.
The real unlocked accesses are elsewhere and are worse: handleEligibleBodies releases s.mu at handlers_api.go:830 and then reads s.bodyStatus at :877, and handlers_web.go:473 lets the map reference escape the lock to be read during template render at :530. Concurrent map read and write is an unrecoverable Go runtime fatal, and its frequency scales with heads x bodies. The correct copy-under-lock idiom already exists 170 lines away at handlers_web.go:307-309. Fixing these is the FIRST commit of the design in #663, ahead of any allocation work.