HydraIssues

MASTER: central body allocation authority + real per-device identity (double-booking and pairing thrash are one defect)
open feature Project: hydracluster Reporter: cederik 5 Sep 2026 10:04

Description

Read the design here

#711 is the browsable design. The full document renders there in any browser, no checkout and no admin view needed. This issue keeps it in the plan field too, but plan only renders on the admin detail page, not the public one.

Pushed to origin: design and reasoning on main (merged 2026-09-14, fast-forward, commit 819ea49).

The 50 child issues in the breakdown are deliberately NOT filed. The breakdown is the implementation plan, not a set of tickets.


Where the artefacts live

  • The design is the plan field of this issue, and is also committed to
    hydracluster/docs/design/body-allocation-and-pairing.md.
  • The reasoning behind it is hydracluster/docs/design/body-allocation-and-pairing-reasoning.md:
    the five competing architectures in full, the three judge scorecards, and all 29
    adversarial breaks with the fix for each. Read it before proposing a change to the
    design, so an already-rejected option is not re-proposed.
  • Index: hydracluster/docs/design/README.md.

The repo copy is the source of truth once committed; this issue holds the snapshot.


MASTER: bodies and heads need a central allocation authority and a real device identity

Two defects have been chased separately for months. They are one defect. This issue is the
single home for both, and for the architecture that closes them.

A design workflow is running now and its output will be attached here. Child issues will be
cut from the design's issue breakdown.

The two halves

Half 1: nothing centrally owns capacity

Heads pick their own body and nothing server-side reserves, rejects, or arbitrates.
/api/v1/bodies/eligible returns a list, every head applies bodies.first { streamCount == 0 }
locally, and stream_count is derived from a body heartbeat that can be up to 30s stale.
Confirmed in production 2026-09-04 at Sint-Niklaas: all THREE iPads streamed from
cosmic-pretzel-98 at once and showed the same Mercator. 204 [session] UUID mismatch lines
in one day. Detail and evidence in #660.

Half 2: we re-pair every launch instead of detecting an existing pairing

hydraheadipad/Sources/HydraHeadiPad/AppState.swift:326:

Always pair to get a fresh cert - no caching. Sunshine can rotate its cert on restart; a
stale cached cert causes silent stream failures... HydraPairSession handles the
alreadyPaired case via unpair->re-pair.

So every stream launch tears down and rebuilds the pairing.
Vendors/hydra-moonlight-ios/Limelight/HydraPairSession.m:128 implements alreadyPaired
as GET /unpair?uniqueid=... followed by a fresh handshake.

The Go head is closer but not clean either. hydraheadflatscreen/pkg/client/gamestream_pair.go:90
DOES detect serverInfo.PairStatus == "1", then unpairs anyway to force a fresh server cert.
The genuinely correct behaviour, detect a valid stored cert and skip the handshake entirely,
exists only in the fork's headless cert-reuse path described at pkg/client/pairing.go:56-62
(HydraExperienceNet v6.1.22+). So the premise that "omarchy does it properly" is HALF true:
Linux has native in-process pairing and a real pair-status check, but the reuse decision
lives in the subprocess fallback, not the native path.

Why they are one problem

Every iPad ships the same hardcoded Moonlight uniqueid 0123456789ABCDEF (#531). Verified:
exactly one distinct uniqueid appears in a body's entire sunshine.log.

That single fact causes both halves:

  1. It makes double-booking invisible and unblockable. Sunshine sees one paired client,
    not three, so it cannot refuse the second and third head. Its own exclusivity is defeated.
  2. It makes re-pairing destructive. When iPad A unpairs before re-pairing, it unpairs the
    identity that iPads B and C are also using on that body. Pairing is not merely wasteful,
    it actively breaks the other heads. #371 is the recorded fallout.
  3. It makes the control plane unable to attribute anything. sessionStore.active is keyed
    by body ID, so two heads on one body collapse into a single record and the sessions API
    structurally cannot represent the bad state.

Fix allocation without fixing identity and the allocator cannot tell which head holds what.
Fix pairing without fixing allocation and correctly-paired heads still collide on one body.

Requirements for the architecture

  • Safety must not depend on well-behaved clients. An old, buggy, or offline head must not be
    able to corrupt state or double-book.
  • Level-triggered, not edge-triggered. Today a missed streaming -> idle edge strands a
    session; the watchdog cannot help because staleHead is skipped when the head never
    heartbeated (session_watchdog.go).
  • Survive a control-plane restart. s.bodyStatus is in-memory and lost on restart.
  • Bad states must be REPRESENTABLE and observable. Today's bug hid for weeks because the data
    model could not express it and the only trace was a log line.
  • Pairing must be idempotent: ask whether this device's identity is already trusted and its
    stored cert still valid, and do nothing when it is.
  • Per-device identity, provisioned at enrolment, migrating the deployed fleet off the shared
    uniqueid without a venue visit.
  • Work at capacity 1 today, extend to capacity N per body (#504) without redesign.
  • Respect the hard constraints: WDAC on bodies (no in-place binary overwrite), never reboot a
    body remotely, ops via the hydracluster API or cluster exec only, iPad ships via TestFlight.

Issue map

Allocation

  • #660 production confirmation, Sint-Niklaas, three iPads on one body, full evidence
  • #530 same bug, same venue, two iPads, 2026-08-19
  • #308 the underlying race; notes the unused Reserve()/Release() at pkg/store/store.go:467
  • #195 endpoint to detect body assignment conflicts
  • #182 eligibility ignores installed experiences

Pairing and identity

  • #531 hardcoded Moonlight uniqueid on every iPad. Root enabler of both halves
  • #549 proactive game pairing: head-driven verify-then-pair reconcile loop, district trust mesh
  • #278 port GameStream pairing to Go, eliminate the subprocess path
  • #495 submitPIN ignores Sunshine status:false and races the pair session
  • #371 iPad pairing fails silently when Sunshine named_devices accumulates, wrong UUID in unpair
  • #321 iPad pairing ATS failure when the body IP is non-RFC1918

Related design tracks, not re-parented pending the design outcome

  • #504 central slot allocator and multi-stream bodies. Already specifies
    POST /api/v1/bodies/{id}/slots/claim with a claim_token and 60s TTL, and already warns that
    claims must count toward advertised availability or the race re-opens. The design must either
    absorb this or explicitly supersede it
  • #84 multi-stream via VDD, the capacity substrate
  • #116 body selection Phase 1, cross-venue eligibility and client failover

Capacity, stated plainly

Sint-Niklaas has no body of its own. Three iPads share two remote bodies at Cloud Seven and
Rupelmonde. Fix the race and the third iPad gets a clean noBodyAvailable instead of a
duplicate picture. That is a better failure, not a success. Three concurrent Mercator streams
need a third slot, whether a third body or multi-stream capacity.


SIXTH ARCHITECTURE PROPOSED 2026-09-18 by the owner: #753, "the stream is the unit, not the body slot". Full text in the issue description (browsable) and in hydracluster/docs/design/design-6-stream-as-unit.md, indexed in that directory's README as a proposal rather than a decision.

The five scored designs all take (body, slot) as the unit and model the head's hold as a lease. #753 takes the STREAM as the unit: a stream needs a head and a body, the head attaches to the stream before any body is chosen, the stream then acquires a body, and the head reconnects to its stream rather than re-acquiring a body.

It targets a root cause none of the five name directly. session_store.go:32 is active map[string]*SessionRecord // keyed by body ID, so a session has no identity of its own; it IS the body's current session. Three heads collapsing into one row, the absence of any queue, a body failure killing the visit, and reconnect being /resume on a shared uniqueid are all downstream of that one key.

Checked against the recorded breaks rather than written fresh:

  • DISSOLVES break 2757 (FATAL, reclaim latency and zero fairness). That break found that no head in the fleet retries a 409, because the iPad sets .error on both discovery and pairing failure. A denial is terminal because there is nothing to hold. A stream in seeking_body is something to hold, and waiting streams are an orderable queue.
  • INHERITS break 2749 (FATAL, the unattended stream) unchanged. It makes the leak representable, not bounded. Design 2's 30 minute maxUnattendedStream is still required.
  • CORRECTED BY break 2839 (FATAL, identity). My first shape had the head mint its own stream id, which is exactly the head-asserted pattern that break condemns. Now the authority mints the stream and binds it to a registered device identity, so #753 DEPENDS ON the device-identity registry rather than replacing it.
  • MAKES WORSE break 2741 (FATAL, locality and fairness) if the grant also moves to the body: home_venue_reserve and per-venue limits are policy a local grantor cannot invent. That argues for the split of central discovery and policy with local admission.

The operational attraction is that migration is per head and per body, which was the owner's second requirement. A body's answer is authoritative regardless of what a head believes, so bodies can flip independently and an old head simply falls through to the next candidate. Cluster-side leases have the opposite property: they need fleet-wide agreement on the lease concept before they bind, which is exactly the mixed-fleet window the reasoning doc attacks in the other designs.

Step 1 is hydracluster-only with no client release: change the key, keep a body index beside it, and the Sint-Niklaas collision becomes representable and countable BEFORE any behaviour changes.

STATUS: proposal, not decision. It has been scored by nobody and attacked by nobody, while the chosen design survived 46 recorded breaks. It should be judged on the same three scorecards and attacked before adoption.

Sub-issues (13)

open #753 DESIGN 6 (proposal): the stream is the unit, not the body slot
open #711 DESIGN: body allocation and device pairing (slot leases with fencing epochs, per-device identity)
open #660 Three iPads stream from one body at Sint-Niklaas: body selection race confirmed in production (204 UUID-mismatch lines), plus a session with no head heartbeat never times out
open #549 Proactive game pairing: head-driven verify-then-pair reconcile loop (district trust mesh)
open #531 HydraHeadiPad ships a hardcoded Moonlight uniqueid (0123456789ABCDEF) on every device
open #530 Body selection double-booking: two heads land on one body when launches fall within the heartbeat lag window
open #495 subprocessPairWithSunshine: PIN submit ignores Sunshine status:false and races the pair session
open #371 iPad pairing fails silently when Sunshine named_devices accumulates (wrong UUID in unpair)
open #321 hydraheadipad: pairing fails with App Transport Security error when body IP is non-RFC1918 (11.0.11.24)
open #308 Body selection race condition: two heads can claim the same body simultaneously
open #278 hydraheadflatscreen: port GameStream pairing to Go (eliminate HydraExperienceNet pair subprocess)
open #195 hydracluster: add streaming monitor API endpoint to detect body assignment conflicts
open #182 Body eligibility check ignores installed experiences — body accepts stream it cannot serve

Session Context

Venue
sint-niklaas-tourism-office