#711 is the browsable design. The full document renders there in any browser, no checkout and no admin view needed. This issue keeps it in the plan field too, but plan only renders on the admin detail page, not the public one.
Pushed to origin: design and reasoning on main (merged 2026-09-14, fast-forward, commit 819ea49).
The 50 child issues in the breakdown are deliberately NOT filed. The breakdown is the implementation plan, not a set of tickets.
plan field of this issue, and is also committed tohydracluster/docs/design/body-allocation-and-pairing.md.hydracluster/docs/design/body-allocation-and-pairing-reasoning.md:hydracluster/docs/design/README.md.The repo copy is the source of truth once committed; this issue holds the snapshot.
Two defects have been chased separately for months. They are one defect. This issue is the
single home for both, and for the architecture that closes them.
A design workflow is running now and its output will be attached here. Child issues will be
cut from the design's issue breakdown.
Heads pick their own body and nothing server-side reserves, rejects, or arbitrates.
/api/v1/bodies/eligible returns a list, every head applies bodies.first { streamCount == 0 }
locally, and stream_count is derived from a body heartbeat that can be up to 30s stale.
Confirmed in production 2026-09-04 at Sint-Niklaas: all THREE iPads streamed from
cosmic-pretzel-98 at once and showed the same Mercator. 204 [session] UUID mismatch lines
in one day. Detail and evidence in #660.
hydraheadipad/Sources/HydraHeadiPad/AppState.swift:326:
Always pair to get a fresh cert - no caching. Sunshine can rotate its cert on restart; a
stale cached cert causes silent stream failures... HydraPairSession handles the
alreadyPaired case via unpair->re-pair.
So every stream launch tears down and rebuilds the pairing.
Vendors/hydra-moonlight-ios/Limelight/HydraPairSession.m:128 implements alreadyPaired
as GET /unpair?uniqueid=... followed by a fresh handshake.
The Go head is closer but not clean either. hydraheadflatscreen/pkg/client/gamestream_pair.go:90
DOES detect serverInfo.PairStatus == "1", then unpairs anyway to force a fresh server cert.
The genuinely correct behaviour, detect a valid stored cert and skip the handshake entirely,
exists only in the fork's headless cert-reuse path described at pkg/client/pairing.go:56-62
(HydraExperienceNet v6.1.22+). So the premise that "omarchy does it properly" is HALF true:
Linux has native in-process pairing and a real pair-status check, but the reuse decision
lives in the subprocess fallback, not the native path.
Every iPad ships the same hardcoded Moonlight uniqueid 0123456789ABCDEF (#531). Verified:
exactly one distinct uniqueid appears in a body's entire sunshine.log.
That single fact causes both halves:
sessionStore.active is keyedFix allocation without fixing identity and the allocator cannot tell which head holds what.
Fix pairing without fixing allocation and correctly-paired heads still collide on one body.
streaming -> idle edge strands astaleHead is skipped when the head neversession_watchdog.go).s.bodyStatus is in-memory and lost on restart.Allocation
Reserve()/Release() at pkg/store/store.go:467Pairing and identity
status:false and races the pair sessionRelated design tracks, not re-parented pending the design outcome
POST /api/v1/bodies/{id}/slots/claim with a claim_token and 60s TTL, and already warns thatSint-Niklaas has no body of its own. Three iPads share two remote bodies at Cloud Seven and
Rupelmonde. Fix the race and the third iPad gets a clean noBodyAvailable instead of a
duplicate picture. That is a better failure, not a success. Three concurrent Mercator streams
need a third slot, whether a third body or multi-stream capacity.
SIXTH ARCHITECTURE PROPOSED 2026-09-18 by the owner: #753, "the stream is the unit, not the body slot". Full text in the issue description (browsable) and in hydracluster/docs/design/design-6-stream-as-unit.md, indexed in that directory's README as a proposal rather than a decision.
The five scored designs all take (body, slot) as the unit and model the head's hold as a lease. #753 takes the STREAM as the unit: a stream needs a head and a body, the head attaches to the stream before any body is chosen, the stream then acquires a body, and the head reconnects to its stream rather than re-acquiring a body.
It targets a root cause none of the five name directly. session_store.go:32 is active map[string]*SessionRecord // keyed by body ID, so a session has no identity of its own; it IS the body's current session. Three heads collapsing into one row, the absence of any queue, a body failure killing the visit, and reconnect being /resume on a shared uniqueid are all downstream of that one key.
Checked against the recorded breaks rather than written fresh:
The operational attraction is that migration is per head and per body, which was the owner's second requirement. A body's answer is authoritative regardless of what a head believes, so bodies can flip independently and an old head simply falls through to the next candidate. Cluster-side leases have the opposite property: they need fleet-wide agreement on the lease concept before they bind, which is exactly the mixed-fleet window the reasoning doc attacks in the other designs.
Step 1 is hydracluster-only with no client release: change the key, keep a body index beside it, and the Sint-Niklaas collision becomes representable and countable BEFORE any behaviour changes.
STATUS: proposal, not decision. It has been scored by nobody and attacked by nobody, while the chosen design survived 46 recorded breaks. It should be judged on the same three scorecards and attacked before adoption.