HydraIssues

DESIGN: body allocation and device pairing (slot leases with fencing epochs, per-device identity)
open feature Project: hydracluster Parent: #663 Reporter: cederik 14 Sep 2026 10:29

Description

This issue is the browsable copy of the design record. The design is committed to
docs/design/body-allocation-and-pairing.md
on main, commit 819ea49 (merged 2026-09-14). The repo copy is the source of truth; this
issue exists so the design can be read in a browser without a checkout.

The reasoning behind it — five competing architectures in full, three judge
scorecards, and all 29 adversarial breaks with the fix for each — is committed
alongside as body-allocation-and-pairing-reasoning.md. Read that before proposing
a change to this design, so an already-rejected option is not re-proposed.

Master issue: #663. The 50 child issues in the breakdown below are not filed;
the breakdown is the implementation plan, not a set of tickets.


Master Design: Body Allocation and Device Pairing

Status: design, for implementation. Attach to master issue #663. Absorbs #504's slot-claim requirements, generalises #556's XR claim, supersedes the standalone flat-claim proposal in #308. Reparents #660, #530, #308, #195, #182, #531, #549, #278, #495, #321, #137, #84, #116, #195.

Basis: Design 1 (Fenced Slot Leases), with grafts from Designs 2, 3, 4 and 5 as marked, and with every fatal and serious break from the four adversarial passes closed or explicitly declared residual.


Problem

Three iPads at Sint-Niklaas streamed from one body at the same time on 2026-09-04 and all three visitors saw the same picture. The control plane could not see it, the body could not see it, and the API could not represent it. Two separate defects produced that single symptom, and neither can be fixed alone.

The allocation half

There is no central capacity authority. handleEligibleBodies (/home/claude-user/hydracluster/pkg/api/handlers_api.go:804) is a pure read: it filters, sorts and returns a plain JSON array, writes nothing, and reserves nothing. Exclusivity is one line of client-side filtering on one client: bodies.first { ($0.streamCount ?? 0) == 0 } (/home/claude-user/hydraheadipad/Sources/HydraHeadiPad/Services/HydraClusterClient.swift:84). The Go head is weaker still: it decodes StreamCount at /home/claude-user/hydraheadflatscreen/pkg/client/discovery.go:23 and reads it nowhere in the repo, taking the first body with a reachable transport (discovery.go:85). Sunshine's web port on 47990 answers identically whether the body is idle or mid-stream, so reachability is being used as a proxy for availability.

The signal the iPad filters on is stale by up to 30 seconds and can be absent entirely. streamCount is synthesised from the body's self-reported StreamStatus (handlers_api.go:911), which arrives on a 30s idle ticker (/home/claude-user/hydrabody/internal/cli/run.go:74) that only drops to 5s after a tick has already observed streaming (provider.go:158-168), and which is suppressed entirely for up to 45s when the payload is byte-identical (status.go:247). Every head receives an identical, identically ordered list: sameMachine > sameVenue > idle-first > VRAM (handlers_api.go:940), and parseVRAM (handlers_api.go:1031) requires a literal " MiB" substring while hydrabody sends only the GPU name (sysinfo.go:74), so gpu_memory_total_mb is 0 fleet-wide and the last tiebreaker never fires. Convergence on the same first entry is the expected outcome, not a narrow interleaving.

The bad state is structurally unrepresentable. sessionStore.active is map[string]*SessionRecord keyed by body ID (session_store.go:32) and openSession returns early when a record exists, so three heads collapse into one row. The one detector that fires, setSessionUUIDFromHead (session_store.go:133-149), logged [session] UUID mismatch 204 times in one day on body node-4c2be4b0, cycling between three iPad client IDs, and then assigned the new value anyway. That log reaches no API, no event and no metric. The defect ran from #308 (2026-05-20) through #530 (2026-08-19) to #660 (2026-09-04) undetected because no query over the data model could return it.

Two more holes compound it. The watchdog exempts a session no head ever adopted: staleHead is guarded by !rec.HeadLastHeartbeatAt.IsZero() (session_watchdog.go:61), and the body heartbeats every 5s while streaming so the body-side timer never fires either. Session node-4c2-10 held a body for 2h46m with head_last_heartbeat_at of 0001-01-01. And the stop paths fabricate allocator input: handleHeadStreamStop writes s.bodyStatus[bodyID].StreamStatus = "idle" without confirming anything (handlers_head.go:561-566, and handleNodeStreamStop at :1087-1090), so the body's next "streaming" post reads as a rising edge against the cluster's own lie and opens a duplicate session.

The pairing half

Every head tears down and rebuilds its Sunshine pairing on every launch. The iPad states it as policy: "Always pair to get a fresh cert, no caching" (AppState.swift:326-329), and HydraPairSession.alreadyPaired issues GET /unpair?uniqueid=... then re-pairs (HydraPairSession.m:128-156). The wider blast radius is PairManager.finishPairing, which unpairs on every failure path (PairManager.m:57-60), so on flaky venue WiFi a head that never streams revokes two heads that are streaming. The premise that omarchy does this correctly is false: gamestream_pair.go:90 detects PairStatus == "1" and unpairs anyway, discarding the error (_, _ =), and the flatscreen tick loop re-pairs every 30 seconds forever whenever a head holds an assignment and is not streaming (client.go:251-261 plus internal/cli/run.go:66). The only cert-reuse code in the fleet lives in subprocessPairWithSunshine (pairing.go:63,138-147), which is dead on Linux and macOS because readClientCertKey is implemented on both.

All of this runs on one identity. 0123456789ABCDEF and devicename roth are hardcoded in the iPad at HttpManager.m:57, HydraPairSession.m:147 and in the /cancel URL at HydraStreamSession.m:219, and as package constants in the Go head's native path at gamestream_pair.go:24, and hydraheadquest re-pairs the same way (#544). #531 is scoped to hydraheadipad only; that scoping is wrong.

Why they are one problem

The shared identity is what turns a race into the observed symptom. StreamManager.m:121-126 reads _SERVER_BUSY from serverinfo and calls /resume with the shared uniqueid; Sunshine sees the owner returning and hands head B head A's live RTSP session. A perfect allocator would leave that join path intact for any overlap that still occurs. Symmetrically, /cancel?uniqueid=0123456789ABCDEF on head A's exit kills head B's session, and DELETE /heads/{id}/stream with a caller-supplied body_id on any enrolled node's token (server.go:409) lets head A force-terminate head B through the control plane. One head leaving hits the others twice.

The shared identity is also what makes the collision invisible. SUNSHINE_CLIENT_UUID, captured by the prep-cmd hook (stream_hook.go:41-46, httpserver.go:64-90) and reported upward, carries one constant for the whole fleet, so neither the body nor the cluster can attribute a session to a head.

And re-pairing has a scheduled amplifier living inside the body agent: TrimPairedClients calls UnpairAll() when named_certs exceeds 200 (sunshine_provider.go:316-329, called from provider.go:196), which is "safe" only under the re-pair-every-launch model being deleted. Re-pairing drives the count to the threshold; the threshold wipes the fleet's trust; the wipe forces re-pairing.

Finally, capacity. Sint-Niklaas owns no body. Three iPads share two bodies owned by another venue. Fixing the race converts a duplicate picture into an honest refusal. That is a better failure, not a fixed venue, and this document says so in the Capacity section rather than hiding it.


Invariants

Eleven properties. For each: the statement, the mechanism that enforces it, and what enforces it when the client is old, buggy or hostile. "Client" means any head; heads ship on four independent release channels and one of them is behind App Store review, so no invariant may depend on head cooperation.

I1. Slot exclusivity

For every body B and slot index k, at every wall-clock instant, at most one lease exists on (B,k) in a non-terminal state, and at most one admitted Sunshine session exists on (B,k).

Mechanism. A lease is created only by a compare-and-set executed inside hydracluster's existing process-wide s.mu, against durable state, committed with an atomic file write before the HTTP response is written. This is the construction handleCreateXRSession already ships (handlers_xr.go:119-161, comment: "Everything below runs under the store lock so two heads cannot arm the same body"). The staleness of stream_count stops being an input: allocation is a write, and the loser loses the write.

Single-writer, enforced not assumed. Grepping /home/claude-user/hydracluster returns no flock, no LOCK_EX, no lockfile, no pidfile and no generation counter, while store.Save() marshals the entire store from a private in-memory copy and atomically renames over the file (pkg/store/store.go:138-186). Two processes for two seconds during a deploy do not merely lose a lease: each computes epoch = persisted + 1 from its own copy and mints the same epoch, so the fencing token collides. The design therefore adds, as a startup precondition and not a follow-up:

  1. flock(LOCK_EX|LOCK_NB) on the allocation file at process start. Failure means refuse to serve, logging the holder's PID.
  2. A top-level generation: N captured at load and verified-then-incremented inside every save. A save whose on-disk generation has advanced aborts, reloads, and retries the whole read-modify-write.
  3. Epoch allocation independent of the loader's copy: epoch = max(persisted_epoch, body_reported_max_epoch_seen, generation) + 1.

Against a misbehaving client. Six layers, listed in the Architecture section. The two that require nothing of the head are the app-entry gate (an unleased migrated head has no launchable Sunshine app name) and the body admission gate at the prep-cmd Do hook, which counts concurrent launches per slot without reference to any identity and, for a client it does not recognise, calls POST /api/v1/bodies/{id}/admit-observed so the cluster runs the same CAS on the body's behalf.

I2. No unattributed stream

Every live Sunshine session on a body is either backed by a lease naming it, or is adopted into one within one body report interval, or is reported as a conflict. It is never silently unaccounted.

Mechanism. Level-triggered reconciliation on every body status post. A reported session with a recognised lease_id renews that lease. A reported session with no recognised lease is adopted (a lease is synthesised in state adopted, body-renewed) rather than terminated, because terminating unrecognised sessions would tear down every live stream in the fleet on the first restart after a deploy. A reported session whose client does not match the lease holding the slot raises slot_conflict and is revoked by an epoch-fenced, session-scoped directive.

Against a misbehaving client. Adoption is unconditional and happens whether or not the head ever asked for anything, so the census is complete for old heads by construction.

I3. Occupancy is monotone in knowledge

Observed state may release capacity. It may never grant capacity. Absence of knowledge is UNKNOWN, never Idle.

This is Design 2's rule, adopted verbatim as testable code, and it is the fix for the single most dangerous bug found in the leading design. Design 1 as written redefined stream_count to mean "slots held", which drops the bs.StreamStatus == "streaming" term that handlers_api.go:911 has today. A live stream with no lease would then read as free, and such streams are common: hydraheadflatscreen in server-assigned mode never calls discoverBody (only two call sites exist, diagnostics.go:68 with experience="" and localapi.go:215), and liveBodyID() returns "" unless cachedBody != nil, which only the self-service path sets (heartbeat.go:157-168).

Mechanism. Occupancy is the union:

func occupied(b Body, k int) bool {
    if leaseOn(b, k) != nil { return true }                     // durable claim
    obs := s.bodyStatus[b.ID]                                    // derived, may be absent
    if obs.ReportedAt.IsZero() { return true }                   // UNKNOWN -> occupied
    if time.Since(obs.ReportedAt) > bodyReportFresh { return true }
    if k < len(obs.Slots) && obs.Slots[k].State != "idle" { return true }
    if k == 0 {                                                  // legacy body, no slots array
        if obs.StreamStatus == "streaming" { return true }
        if xrEngagedState(obs.XRState) { return true }
        if b.Node.XRSession != nil { return true }                // preserved from :911
    }
    if b.Circuit.State == "open" { return true }
    if b.Node.Status != "online" { return true }
    return false
}

grantable(b,k) additionally requires the home-venue reserve and the venue and org concurrency caps to permit (I10).

Against a misbehaving client. No client can make a body look free. Only a lease release, a body report of idle, or an expiry can.

I4. Fencing: no stale directive can act

Every cluster-to-body directive carries (lease_id, slot, epoch). A body never applies a directive whose epoch is below the highest it has applied for that slot, and never applies a directive older than directiveFreshness.

Mechanism. epoch is monotonic per slot, never reused, bumped on every grant and every terminal transition. The body persists slot_epoch.json, a floor that survives a hydrabody restart, a body reboot, and an allocation file restored from backup. This is Design 1's unique contribution and no other design survives a restored backup.

The floor must be raisable, or it bricks the body. If the cluster's epoch ends up below the body's floor, every admission fails, the head retries, and (in the original design) the revocation escalation fired. The closure has three parts: the cluster re-raises slot.epoch = max(slot.epoch, reported_max_epoch_seen) on every body post, not only the first after boot; a slot whose body has not confirmed the current epoch is not grantable at all; and a body that rejects a directive as sub-floor emits epoch_below_floor, which is the one signal that distinguishes this from ordinary pairing trouble.

Against a misbehaving client. The client is not in the directive path.

I5. Ownership is authenticated, never asserted

A head may act only on its own head id and its own lease id. No caller-supplied body_id is ever an authorisation input.

Mechanism. Claim, renew, release, admission-status and device-identity routes sit behind requireAdminOrSelfNodeToken (selfscope.go, server.go:420, today used on exactly one route). DELETE /api/v1/heads/{id}/stream resolves the body from the lease and ignores the caller-supplied body_id for authorisation.

The claiming read is the exception that must be closed in the handler. GET /api/v1/bodies/eligible is requireAdminOrNodeToken (server.go:355) with an unscoped head_id query parameter, and requireAdminOrSelfNodeToken cannot be applied because it refuses outright when there is no {id} path value. So the scoping goes in handleEligibleBodies itself: if the request carries a node context rather than an admin token, head_id must be empty or exactly the authenticated node's ID, else 403. Both shipping clients already send their own id (HydraClusterClient.swift:75, discovery.go:44), so this costs nothing. Additionally, a read never renews an existing lease, so a third-party read cannot extend one, and implicit-grant rate limits count per authenticated node, not per claimed head_id.

Against a misbehaving client. A head can lose a race. It cannot win one it should have lost, cannot free another head's slot, and cannot terminate another head's stream. Note what this replaces: today PUT /api/v1/heads/{id} and DELETE /api/v1/heads/{id}/stream both accept any enrolled node's token.

I6. One durable identity per device, and no head may name another's

Each head presents exactly one GameStream identity to every body: device_uid = hex(sha256(head_id))[0:16], devicename hydra-<head_id>, and a device-held keypair whose fingerprint is registered with the cluster.

Mechanism. The identity is a name, not a secret; the private key is the secret. device_uid is deterministic so the cluster and a human reading a body's sunshine.log can both recompute it. devicename is the load-bearing choice: go-sunshine's ListClients() returns exactly {UUID, Name} (types.go:39-43) and UnpairClient takes only a uuid (client.go:143-147), so a meaningful name is the only way the body can resolve an observed client to a head. Design 1's claim that "the fingerprint is the capability handle... the body pairs it, the body revokes it" is not implementable: the body can never see a certificate.

The lease therefore names two values in two explicitly separate namespaces:

field namespace who supplies it used by
device_uid GameStream ?uniqueid= query param the head, deterministic from head_id head-side self-scoping of /cancel, /unpair, /resume
sunshine_client_uuid Sunshine named_certs[].uuid, surfaced as SUNSHINE_CLIENT_UUID Sunshine, learned by the cluster from body reports the body's admission set membership test
client_cert_fp sha256 of the client cert DER the head, registered once cluster-side registry, revocation, duplicate detection. Never used at the body.

Against a misbehaving client. Identity registration is first-write-wins per head; a differing fingerprint starts a bounded predecessor/successor overlap rather than silently replacing. Registering a fingerprint already bound to a different, non-denied head returns 409 identity_in_use and raises duplicate_head_identity. Identity is removed from the heartbeat body entirely: the heartbeat may reference the registered identity, never define it, because PUT /api/v1/heads/{id} accepts any enrolled node's token and "re-asserts on every heartbeat" means a device you are trying to revoke overwrites the revocation every 30s.

Cloning. The iPad's client key, cert and p12 are plain files in NSDocumentDirectory (CryptoManager.m:183-192), the Moonlight uniqueid is in NSUserDefaults (DataManager.m:52-58), and enrollmentConfig including the head token is in NSUserDefaults too (QREnrollmentView.swift:120). No NSURLIsExcludedFromBackupKey, kSecAttrAccessible or SecItemAdd appears anywhere in the repo. A device replaced by iCloud restore is therefore a perfect clone of the original, head token included, and every detector in the system keys on a value the two devices share. Closure: move the keypair and the enrolment token to the Keychain with kSecAttrAccessibleAfterFirstUnlockThisDeviceOnly and kSecAttrSynchronizable=false; exclude any residual file from backup; add the server-side backstop that two concurrent heartbeats for one head_id from different source addresses raise duplicate_head_identity and refuse the second claim; and require the claim token, not just the head id, for the "return the existing lease" re-adoption path.

I7. Pairing is additive; no head unpairs another

A head pairs only when it cannot prove it is already paired. A head issues /unpair only for its own device_uid, and only after a proven change in the server's certificate. Nothing else in the fleet unpairs anything except the cluster-directed selective trim.

Mechanism. Verify-then-pair with a three-valued result (Design 4's rule, adopted):

ensurePaired(body) -> ok | UNKNOWN | notPaired
  stored := readHostCert(body.serverUUID)
  if stored == nil: return notPaired
  mTLS GET https://<host>:47984/serverinfo?uniqueid=<device_uid>
      present own client cert; pin `stored` as the server cert
    200 && PairStatus == "1"            -> ok    (no handshake, no PIN, no unpair)
    handshake completed, peer cert != stored -> serverRotated
    200 && PairStatus != "1"            -> notPaired
    anything else (transport error, timeout, TLS alert, captive portal) -> UNKNOWN

The verify must be the mTLS one on 47984 and never the plain-HTTP one on 47989: over plain HTTP PairStatus can only key on the uniqueid, which is the fleet-wide value, so PairStatus == "1" means "somebody paired". This is the exact error at gamestream_pair.go:90 and PairManager.m:41-49.

UNKNOWN fails the launch retryably and mutates nothing. This is the #339 and #551 lesson applied to the pairing path. Without it, a single dropped packet on venue WiFi is classified as a certificate mismatch, the head discards its pinned server cert, and then re-pins whatever comes back from GET http://X:47989/pair?phrase=getservercert (PairManager.m:118) over plain HTTP: a pinned relationship downgraded to trust-on-first-use on an unauthenticated channel, on the strength of a lost packet. The stored cert is also replaced atomically only after phase 5 succeeds, never discarded at the start of a handshake.

Against a misbehaving client. The bodies' trust stores are reconciled to a cluster-supplied desired set. UnpairAll() is deleted. A head that pairs excessively shows up as growth in named_certs on GET /api/v1/nodes/{id}/pairing, with first_seen per entry.

I8. Bounded holds: no slot is held indefinitely without pixels

Every lease has a bound that no party can renew away.

Design 1 stated plainly that "a body that keeps Sunshine streaming with nobody watching holds a slot indefinitely. That is deliberate." That is not acceptable at a two-slot venue, and the leak is silent: OnStreamStarted captures the pid from the prep-cmd Do hook, i.e. before Sunshine has launched the app, so no pid lands in session.Metadata and ProcessAliveCheck returns ActionNone on every subsequent tick (sunshine_provider.go:104-114, watchdog_checks.go:43-46). The crash watchdog fails closed and silently and would look healthy in any log review.

Mechanism. Three independent clocks, none of which the other can extend:

clock applies to renewable by bound
leaseTTL all non-terminal states head heartbeat or body report 90s rolling
startDeadline held/granted/pairing/starting nobody 210s absolute from grant (90s when lease.protocol >= 2)
maxUnattendedStream streaming/adopted with no head renewal nobody 30 min absolute
unknownHold unknown (body unreachable) nobody 10 min absolute

Plus: fix the inert pid capture so the crash watchdog has a second, independent detector (capture the pid on the first tick after the started hook, not inside it).

I9. The violation is representable and queryable

Two heads on one slot is a row, an event and a metric. It is never a log line.

Mechanism. sessionStore.active is re-keyed from body ID to lease id, with (body_id, slot) as an index that may legitimately hold more than one entry. GET /api/v1/slots returns, fleet-wide, desired and observed side by side with an enumerated divergence[] array (Design 4's primitive, promoted from per-body to fleet-wide). setSessionUUIDFromHead's [session] UUID mismatch log becomes the slot_conflict event.

I10. Liveness and fairness

A freed slot is granted to the longest-waiting eligible head, and a venue that owns a body cannot be starved of it by a venue that owns none.

Design 1 had no fairness, liveness or locality invariant, and all three failures follow from that silence. Locality: the eligibility rule is sameVenue || sameOwner (handlers_api.go:864-868) with no reserve, so Sint-Niklaas (0 bodies) can hold 100% of Rupelmonde's hardware and deny Rupelmonde's own visitors, converting a one-venue shortfall into a two-venue outage caused by the fix. Liveness: the iPad's tick() is case .error: break, so a denied head never retries; a freed slot then sits idle until a human taps a screen.

Mechanism. home_venue_reserve (default 1 where slots_total >= 1) caps how many slots on a body may be held by heads whose venue differs from the body's; max_concurrent_grants per venue and per org; a cross-venue lease is preemptible while still pre-stream (held/granted/pairing/starting, never streaming) by a home-venue claim. Durable FIFO admission tickets ordered by (created_at, ticket_id) on the head node, so queue order survives a control-plane restart, with the 10s reconciler tick walking each pool's tickets in order.

Against a misbehaving client. The rate limit counts only granted claims, never refused ones. Design 1's "1 per 10s, 5 per 60s" limiter, applied in a fleet whose retry driver is a human finger and whose error recovery is a tap, refuses the sixth tap at the exact moment a slot frees, renders it as the same error string, and hands the slot to a different iPad with a fresh window. The limiter written to prevent starvation was the mechanism producing it. A throttled claim returns 429 with a distinct typed reason so the head can render "one moment" rather than "no body available".

I11. Teardown safety: refuse, do not kill

The body may refuse a launch on local signals. The body may never kill a live stream on local signals. Killing requires a cluster directive naming (lease_id, epoch).

Four recorded failures make this non-negotiable: #339 (one network blip dropped port 48010 on three bodies and the watchdog killed three live sessions mid-reconnect), #551 (a Sunshine 401 read as idle), #429 (a hydrabody restart plus the gpu-mismatch watchdog killing a live experience three minutes later), #300/#361 (the TCP port watchdog, three generations, removed entirely).

The admission fallback must not be ForceStop. Design 1's stated fallback, if a non-zero prep-cmd Do does not abort the launch, was to "close that slot's app immediately, roughly one second instead of zero". At capacity 1 the only implementation is ForceStop (sunshine_provider.go:189-200), which calls CloseRunningApp() and then marks every active session ended. During a partition, when the body's directive still names the previous holder, the layer built to protect the lease holder becomes the mechanism that terminates the lease holder, repeatedly, once per retry. Closure: per-session teardown ships before the gate is armed; if per-session close is unavailable, the fallback is report only; the fallback is rate-limited to one execution per slot per 60s; and the body does not enforce a directive older than directiveFreshness.


Architecture

The two state spaces

Desired state is durable. A new file, allocation.yaml, sits alongside nodes.yaml and holds slots, epochs, leases, tickets and per-body circuit state. It has its own generation counter and its own flock.

Putting leases in nodes.yaml (Design 1's choice, following the NodeXRSession precedent) is fatal on rollback. store.Load uses yaml.Unmarshal, not UnmarshalStrict, so a binary that does not know the slots: key parses the file happily and drops it, and store.Save() then rewrites the whole file without it. Head heartbeats call st.Save() on every PUT (handlers_head.go:175), which for a streaming iPad is every 5 seconds. A rollback of the emergency fix therefore deletes every lease in the fleet within one heartbeat, and roll-forward double-books every slot at once. A sidecar an old binary never opens makes the rollback lossless. This is a deliberate departure from Design 1's "nodes.yaml is the only durable store", chosen because the rollback risk is highest on exactly the phase most likely to be rolled back.

Defence in depth: nodes.yaml also gains schema_version: 2. An old binary drops it; when the new binary loads a file with the key absent it knows a downgrade happened, enters safe mode (adopt every reported session, grant nothing until every body in the district has reported), and emits downgrade_detected.

# allocation.yaml
generation: 4417
schema_version: 2
bodies:
  node-4c2be4b0:
    slots_total: 1
    home_venue: rupelmonde
    home_venue_reserve: 1
    circuit: {state: closed, failures: 0, until: null}
    slots:
      - index: 0
        epoch: 41
        lease:
          id: lse-9f3c1a
          protocol: 1                 # 1 = legacy (implicit), 2 = explicit claim
          source: implicit            # implicit | explicit | adopted
          mode: flat                  # flat | xr | pair
          head_id: head-ipad-3
          head_venue: sint-niklaas
          device_uid: 3f2a91c4b8e07d55
          sunshine_client_uuid: ""    # learned from the body, never asserted
          experience: Rupelmonde
          app_token: 7c1e...          # 128-bit, published as the app name suffix
          state: pairing
          reason: ""
          token_hash: sha256:...      # HMAC of the claim token; the token is never persisted
          attempt_key: head-ipad-3|Rupelmonde|1757068200
          granted_at: 2026-09-05T10:02:11Z
          renewed_at: 2026-09-05T10:02:41Z    # 15s write quantum
          expires_at: 2026-09-05T10:04:11Z
          start_deadline_at: 2026-09-05T10:05:41Z
tickets:
  head-ipad-5:
    id: tkt-71ba
    pool: bxl1/rupelmonde-owned
    experience: Rupelmonde
    created_at: 2026-09-05T10:03:02Z
    renewed_at: 2026-09-05T10:03:32Z

token_hash, not the token: allocation.yaml is rendered by the admin UI and lands in backups, and this codebase already treats HydraBrainToken that way.

Observed state is derived and disposable. s.bodyStatus gains what it structurally lacks:

type BodyNodeStatus struct {
    // existing: GPU, ProviderStatus, ProviderVersion, StreamStatus,
    //           InstalledExperiences, XR fields
    ReportedAt            time.Time        // stamped SERVER-SIDE on receipt
    BootID                string           // random per hydrabody process start
    ReportSeq             uint64           // monotonic within a boot
    ProviderUptimeSeconds int
    SlotsTotal            int
    Slots                 []SlotObservation // {Index, State, LeaseID, AppliedEpoch,
                                            //  ClientUUID, App, LaunchID, Since}
    MaxEpochSeen          []uint64
    StreamCount           int
    GPUMemoryTotalMB      int
    PairedClients         []ClientInfo      // {UUID, Name}
    PairedDigest          string
    SunshineCertFP        string
}

SkinNodeScales in the same file already has a ReportedAt with a comment explaining why the cluster stamps it server-side (server.go:171). The Raspberry Pi scale inventory has better staleness hygiene than the GPU bodies that run the product.

Freshness is a permanent rule, not a startup window. now - ReportedAt > bodyReportFresh (90s) means UNKNOWN, means occupied, means not grantable, at any time. Design 1 added the field and then wrote no rule; without one, a body whose hydrabody is dead while hydranode keeps heartbeating (n.Status comes from hydranode, never from hydrabody) stays eligible forever.

One reconciler, one code path

func (s *Server) reconcile(alloc *alloc.Store, st *store.Store, now time.Time) (changed bool)

It runs on the existing 10s session_watchdog ticker, and inline under the same s.mu at the end of handleBodyStatus, at the end of the claim handler, and at the end of the claiming read. There is no fast path that can drift from the loop, which is Design 2's anti-drift property, taken without Design 2's 1Hz cost. No second mutex is introduced: requireNodeToken deliberately holds s.mu across whole handlers (handlers_body.go:45-46) while requireAdminOrNodeToken deliberately releases it before next(), and a new lock here is a deadlock waiting to happen.

Per pass, per body:

1. If obs.BootID != last[bodyID].BootID:
       if obs.ProviderUptimeSeconds < 60: do NOT drain (registry not yet repopulated)
       else: mark non-terminal leases `unknown`; re-send full lease set at current epoch
2. slot.epoch = max(slot.epoch, obs.MaxEpochSeen[k])       // raise the floor, always
3. Advance each lease's phase (state table below)
4. Adopt: reported session with no recognised lease -> synthesise `adopted`
5. Admit tickets in FIFO order into slots that grantable() now permits
6. Emit this body's COMPLETE desired slot table on its status response

Step 1's uptime guard is Design 3's and it closes a bug all of Designs 1, 2 and 5 share: hydrabody's session registry is in memory, so its first report after a restart carries sessions: [] while the experience is still rendering on the glass. Draining on that report frees a slot with a live orphan process on it, and the next head is admitted onto it.

Lease state machine

state meaning renewed by exits
held soft hold from a claiming read; nothing confirmed any heartbeat from the holding head -> pairing on a starting heartbeat; -> streaming on body evidence; -> abandoned on a repeat read or an error heartbeat; -> expired at 75s of head silence
granted explicit claim (protocol 2) head renew -> pairing; -> expired at TTL; -> ended on release
pairing head is verifying or pairing head heartbeat -> starting; -> failed(start_timeout) at startDeadline
starting launch issued head heartbeat -> streaming on body evidence; -> failed(start_timeout) at the same, non-reset deadline
streaming body reports a session attributable to this lease body report (primary), head heartbeat (secondary) -> draining; -> failed(unattended) at maxUnattendedStream with no head renewal
adopted live session with no recognised lease; head_id may be unknown body report -> ended on two consecutive idle reports (min 60s); -> failed(unattended) at maxUnattendedStream
draining terminate outstanding; slot still occupied n/a -> ended when the body stops reporting it; -> failed(drain_timeout) at 45s, escalate
unknown body unreachable or offline while lease non-terminal n/a -> re-adopted on the body's return; -> expired at unknownHold (10 min) with a loud event
terminal ended, expired, revoked, abandoned, failed(<reason>) n/a epoch++, slot free

Every terminal state carries a persisted reason and is written to the session history JSONL (#595), not GC'd. Design 2's Reason strings were its best forensic artefact and were then discarded after 120 seconds; at 11:00 the post-mortem of a 09:00 incident needs them.

Who renews, and why the asymmetry. Bodies are on fixed networks and post every 5s while streaming; heads are on flaky venue WiFi and drop (#359: head heartbeats dropping 27 to 60 seconds into a stream, 73% of sessions ending watchdog_timeout). So before streaming, the head renews; from streaming onward the body renews, and a head partitioned from the cluster while still streaming over the venue LAN keeps its slot.

Design 1's handoff could not complete, and that is closed here. Its promotion trigger was "the body reported the slot active with this lease_id", which an unmodified hydrabody never sends. Under this design, a body reporting StreamStatus == "streaming" for a slot that holds a non-terminal lease renews that lease, with or without a lease_id in the report. The cluster must be able to enforce with today's hydrabody and tighten as bodies update; that rule is applied to renewal, not only to admission.

The claiming read, and legacy failover with no client release

GET /api/v1/bodies/eligible keeps its URL, its auth and its response schema, and becomes a claim. Two changes carry the whole legacy story:

  1. Omission, not stream_count: 1. A body held by another head is removed from the list. Marking it busy is invisible to hydraheadflatscreen, which decodes StreamCount and reads it nowhere. Removing it is not. Both shipping clients fail closed under omission.
  2. A soft hold on exactly one body, state: held, counted as occupied from the instant it is written.

Guards that make write-on-read safe:

  • Non-empty experience only. The iPad's Diagnostics panel and the Go head's diagnostics.go:68 both call this endpoint with experience="". Without this guard an operator opening Diagnostics takes a durable hold and denies a visitor the last free body.
  • Scoped head_id (I5).
  • Idempotent on attempt_key = head_id|experience|floor(now/10s). The iPad's request timeout is 15 seconds (HydraClusterClient.swift:16); under venue-open contention the claim can succeed server-side and the response can be lost, orphaning a lease while the head shows an error and retries, each retry adding load. The 10s window collapses a retry into the same hold.
  • Repeat-read means the previous attempt failed. A head that succeeded is streaming and does not re-read (discoverBody is called only from startStream). So a second read from the same head for the same experience, where the previous hold has not been confirmed by a body report, releases that hold, marks it abandoned, demotes that body for that head for bodyDemotionWindow (5 min), counts one failure against the body's circuit breaker, and grants the next-ranked candidate. This is the legacy failover mechanism, and it needs no client release: it turns a human re-tap, or a Go head's next 30s tick, into a candidate advance.
  • Status-driven release. A heartbeat from the holding head reporting error releases the hold immediately. This delivers, server-side and with no iPad release, the release that showError() never sends (showError calls micRelay.stop() and sets .error; only stopStream() and showSessionInterrupted() call notifyStreamStopped).
  • Ranking. Sort becomes sameMachine > free-first > sameVenue > gpuMemoryTotalMB desc > fnv(body_id + head_id). Free-first is promoted above same-venue because today a busy same-venue body outranks an idle cross-venue one, which actively steers the second head at the occupied local machine. The fnv salt is Design 2's, and it breaks identical ordering for any head still choosing client-side. gpu_memory_total_mb is fixed at source: hydrabody already computes it (sysinfo.go:85-87) and already sends it, to hydrabodystatus (bodystatus.go:20,90), never to hydracluster.

Enforcement layers

Six, honestly scoped. Design 1's L3 (credential vending) and L5 (unpair the shared cert at lease end) are deleted; see Decisions.

# Layer Where Stops Needs
E1 Lease CAS hydracluster, under s.mu, durable, fenced the race between two well-behaved heads nothing
E2 Claiming read + omission + assigned-config claim hydracluster old, unmodified heads nothing
E3 App-entry gate hydrabody, via Sunshine apps.json an unleased migrated head naming a launchable app hydrabody + head protocol 2
E4 Body admission gate hydrabody, prep-cmd Do hook a second concurrent launch on a slot; an unrecognised client hydrabody
E5 Head self-scoping all four head repos /resume joining a stranger's session; /cancel killing a sibling; /unpair fratricide head release
E6 Sunshine mTLS Sunshine an unpaired client per-device certs

E6 is not a second-client defence, and Design 1's claim that it is was wrong. The iPads already hold distinct client certificates: generateKeyPairUsingSSL runs under dispatch_once and only when keyPairExists is false (CryptoManager.m:286-289), per install, and #371 observed 294 named_devices entries, which is only possible if each pair appends a distinct certificate. Sunshine already sees many paired clients. Changing the uniqueid does not change how Sunshine counts clients. What actually produced the shared picture is /resume with the shared uniqueid: Sunshine treats it as the owner returning. So the /resume prohibition is promoted to its own numbered layer (E5) and E6 is described for what it is.

E3, stated precisely, because the polarity matters. The cluster computes the desired app set per slot and sends it in the directive:

slot state published Sunshine app entries for that slot
free plain <experience> (so a legacy head can launch and be adopted by E4)
leased, lease.protocol == 1 plain <experience> (legacy heads launch by experience name)
leased, lease.protocol >= 2 <experience>#<app_token> only; plain entry removed
maintenance: true (admin) plain <experience> and Desktop, leasing suspended

The maintenance override exists because removing Desktop unconditionally would delete the operator's only Moonlight-side route into a body that must never be rebooted remotely and is reached only over hydranode exec.

E4, stated precisely. On the stream/started hook:

1. launch counting: if this slot already has an un-ended session, this is a
   second concurrent launch. Refuse. Report second_launch_refused.
   IDENTITY-FREE: works before the identity migration.
2. set membership: is the presenting client (SUNSHINE_CLIENT_UUID, or its
   named_certs name -> head_id) in the current directive's allowed set for
   this slot, and is the directive younger than directiveFreshness (90s)?
     yes -> admit, report
3. unknown client, enforce mode, fresh directive:
     POST /api/v1/bodies/{id}/admit-observed  (2s timeout)
       201 -> admit (cluster minted or adopted a lease)
       409 -> refuse, report
       timeout / unreachable / any ambiguity -> ADMIT and report
4. observe mode: always admit, always report

Rule 3's ambiguity branch is I11. A refusal executes as a non-zero exit from the prep-cmd Do and, only if per-session close is available, a close of that session. Never ForceStop. Rate-limited to one per slot per 60s.

admit-observed (Design 3, the single best trustless primitive in the set) is the only enforcement that reaches hydraheadflatscreen, arrives via hydrabody's own 6h self-update rather than an app store, and needs nothing of the head. Its known weakness is that under a shared uniqueid two heads present the same client_id, so the CAS matches an existing lease. That is why rule 1 exists: launch counting is identity-free and covers the shared-identity era.

How pairing reinforces allocation, and the reverse

  1. Identity makes occupancy attributable. SUNSHINE_CLIENT_UUID -> named_certs[].name -> head_id is the only path by which a body can say which head is streaming, and it only resolves because devicename becomes hydra-<head_id>.
  2. Identity makes E4 and E6 mean something. Set membership needs something to be a member of.
  3. Detect-and-reuse collapses the allocation window from the other side. The iPad reports body_id: nil for the whole .discovering plus .pairing interval, up to the 180s getservercert timeout (HttpManager.m:194), which is exactly the window a second head uses. Verify-then-pair reduces it to one mTLS round trip, which is why startDeadline can drop from 210s to 90s for migrated heads.
  4. The allocator makes pairing serialisable. A pairing handshake becomes a lease with mode: pair, slot: -1, allocated by the same reconciler and refused while any flat/xr lease is active on that body. gamestream_pair.go:117's "another pairing attempt is already in progress", today the fleet's only cross-head collision detector and today flattened into an opaque error string, becomes a typed 409 and a pair_contention event.
  5. Allocation removes pairing's amplifier. Deleting the re-pair makes named_certs growth O(devices) rather than O(sessions), so the trim essentially never fires, so stored certs stay valid, so the reconcile keeps returning "no action".
  6. The migration is coupled in one direction only. Every /unpair call site must be deleted before identities diverge, or the first migrated head reads PairStatus 0, pairs as a new client, and can still unpair the shared identity that every unmigrated head depends on.

Protocol

Cluster, head-facing

POST /api/v1/heads/{id}/slot-claims — auth requireAdminOrSelfNodeToken.

{ "attempt_id": "a-8f21", "district": "bxl1", "venue": "sint-niklaas",
  "experience": "Rupelmonde", "stream_mode": "flat",
  "device_uid": "3f2a91c4b8e07d55", "protocol_version": 2,
  "prefer_body_id": null }
201 {
  "lease": { "id": "lse-9f3c1a", "epoch": 42, "slot": 0,
             "claim_token": "…returned once…",
             "ttl_seconds": 90, "renew_after_seconds": 20,
             "start_deadline_seconds": 90 },
  "body": { "id": "node-4c2be4b0", "name": "boom-pickle-11",
            "hosts": {"lan": "192.168.1.40", "wireguard": "10.10.0.7",
                      "same_machine": false},
            "ports": {"gamestream": 47989, "gamestream_tls": 47984, "web": 47990},
            "app_name": "Rupelmonde#7c1e…",
            "sunshine_cert_fp": "9a1b…", "pair_epoch": 3 },
  "pair": { "required": false, "reason": "cert_valid" }
}
status body meaning
200 same shape idempotent replay of attempt_id, or this head's existing lease re-returned (requires claim_token for re-adoption)
202 {"ticket": {...}, "retry_after_seconds": 5} queued, FIFO position and ETA
409 {"error":"head_busy","lease":{...}} this head already holds a lease
409 {"error":"reserved_for_home_venue","pool":{...}} the reserve blocks this cross-venue claim
409 {"error":"identity_unknown"} / {"error":"identity_revoked"} device identity not registered or revoked
422 {"error":"experience_unavailable"} no fresh body in pool lists it
429 {"error":"claim_throttled","retry_after_seconds":6} granted claims only
503 {"error":"reconciling","retry_after_seconds":5} + Retry-After cluster is in boot safe mode for this pool

POST /api/v1/heads/{id}/slot-claims/{lease_id}/renew — {claim_token} -> 200 {ttl_seconds, epoch} | 410 {"error":"lease_lost","reason":"expired|revoked|superseded"}. 410 is the head's instruction to stop streaming and re-claim. Renewal normally rides the existing heartbeat; this route exists so renewal is not solely dependent on a full-record PUT.

DELETE /api/v1/heads/{id}/slot-claims/{lease_id} — {claim_token, reason} where reason is user_exit | error | abandon | unreachable | interrupted. 204. Scoped by lease id, never by body id.

GET /api/v1/heads/{id}/admission — the head's current lease or ticket, joined server-side so one call answers "where is this head". The 5s poll target for a queued protocol-2 head.

GET /api/v1/bodies/eligible?head_id&district&venue&experience&stream_mode&protocol — unchanged shape plus additive per-body fields slot, slots_total, slots_free, reported_at, gpu_memory_total_mb, pair_epoch, sunshine_cert_fp, lease{id,state,expires_in_s,mine}. stream_count is retained and means occupied slots under the I3 union, so an old iPad's == 0 filter stays correct. Legacy-ness is derived from the request, never from stored state: absence of &protocol=2 means legacy. A protocol_version persisted on the head node would be sticky, and a TestFlight rollback (which keeps app data and therefore the head id) would leave a legacy binary flagged as migrated and unprotected. protocol_version is kept on the head record for reporting and for the Phase-retirement gate only, stamped with protocol_version_seen_at.

Admin-only &include_leased=1 returns omitted bodies with their holders and mutates nothing (Design 5). Without it, reproducing a head's view after the fact requires a mutating call.

PUT /api/v1/heads/{id} — moved to requireAdminOrSelfNodeToken. Carries additive lease_id, lease_state. Renews the lease and the hold. Does not carry device_uid, client_cert_fp or protocol_version as authoritative writes. A regression test asserts that this handler mutates named fields and never replaces the node struct, because the flatscreen client's own comment claims the opposite (heartbeat.go:15-18) and a refactor that made it true would eat every ticket.

GET /api/v1/heads/{id} (head config) — for a head with HeadBodyID set, serving the stream block is a claim point. If a lease on the assigned body cannot be taken, the stream block is omitted and the head falls back to its self-service grid, which it already implements. This is the coverage for server-assigned flatscreen heads, which never call discovery at all and report body_id: "" while genuinely streaming. Gains pairing_targets: [{body_id, hosts[], server_cert_fp}] and per_device_identity: bool.

Cluster, body-facing

POST /api/v1/body/status (hydraclusterapi v0.24.0, additive both directions, both sides tolerate the older peer):

Request adds boot_id, report_seq, provider_uptime_seconds, slots_total, slots[] {index,state,lease_id,launch_id,applied_epoch,client_uuid,app,since}, max_epoch_seen[], stream_count, gpu_memory_total_mb, paired_clients[] {uuid,name}, paired_digest, sunshine_cert_fp, hydrabody_version.

Response adds:

"desired": {
  "epoch_floor": 42,
  "enforce": "observe",
  "maintenance": false,
  "slots": [{ "index": 0, "action": "serve", "lease_id": "lse-9f3c1a",
              "epoch": 42, "head_id": "head-ipad-3",
              "device_uid": "3f2a91c4b8e07d55",
              "sunshine_client_uuid": "…",
              "app_entries": ["Rupelmonde#7c1e…"],
              "protocol": 2 }],
  "issued_at_seconds_ago": 0
},
"trust": { "desired_clients": [{"uuid":"…","name":"hydra-head-ipad-3"}],
           "remove": ["…"], "max_clients": 150 }

The complete desired slot table every tick, not an event. Idempotent, implicitly retried, self-correcting: a dropped response is corrected 5s later. terminate_stream is retained and synthesised for bodies reporting a pre-protocol hydrabody_version. This replaces pendingTerminate, which is body-keyed, untimed, unpersisted and drained unconditionally whether or not the body complied (session_store.go:154-178, handlers_body.go:429).

POST /api/v1/bodies/{id}/admit-observed — node-token auth. {client_uuid, named_cert_name, slot, app_name, launch_id} -> 201 {lease} (adopted, attributed to the head whose device_uid/name matches, else head_id: "unknown"), or 409 {"error":"slot_held","holder":{head_id, device_uid}}. 2s client timeout.

Cluster, identity and trust

POST /api/v1/heads/{id}/device-identity — requireAdminOrSelfNodeToken. {device_uid, cert_fingerprint, cert_pem, not_after} -> 201 {registered_at, supersedes_fingerprint, overlap_until}. Idempotent for the same fingerprint. A different fingerprint starts a 7-day predecessor/successor overlap so a head that rotates while some bodies are offline is not locked out. 409 identity_in_use if the fingerprint is bound to a different, non-denied head; the head responds by regenerating and re-registering, which turns an invisible clone into a loud, self-healing event.

DELETE /api/v1/heads/{id}/device-identity — admin only. Marks the fingerprint revoked, releases any lease held by that device, and adds its uuid and name to every body's trust.remove until each body confirms removal from paired_clients.

Cluster, operator

GET /api/v1/slots?district&venue — the fleet truth call. Per (body, slot): epoch, the lease (id, head, head_venue, state, reason, source, protocol, experience, granted_at, expires_in_seconds, lease_age_seconds, seconds_since_head_renewal), the observed side (client_uuids[], stream_status, slots[], reported_at, age_seconds, boot_id, applied_epoch), and an enumerated divergence[] from {unauthorized_stream, lease_without_stream, epoch_regressed, epoch_below_floor, wrong_client, stale_report, second_launch, adopted_unknown_head, downgrade_detected}.

GET /api/v1/allocation — per pool: queue depth, median wait, denials_last_hour broken out by reason and by both axes (by_head_venue: whose visitors were turned away; by_body_venue: whose hardware is oversubscribed). A bodyless venue has slots_total: 0 of its own, so a single venue_capacity number is undefined for exactly the topology that caused the incident.

GET /api/v1/nodes/{id}/pairing — the trust store view: named_certs with first_seen, evictions performed, server_cert_uuid and its last observed rotation, custodian last-success timestamp.

PUT /api/v1/nodes/{id}/slots — admin. {slots_total, enforce: "observe"|"enforce", maintenance: bool, home_venue_reserve, max_concurrent_grants}.

GET /api/v1/conflicts — live rows where an observed session's client does not match the lease holding its slot, or where an observed client maps to no enrolled head. Closes #195's conflict-detection ask.

Body-local, :47991

The four stream routes (/api/v1/stream/{started,ended,stop,sessions}) are currently the only unauthenticated routes on that port (httpserver.go:47-50) and are precisely the ones that start, end and force-stop streams, while launch, stop, kiosk and debug/screenshot in the same file already require isTrustedNetwork(10.10.0.0/16) OR a bearer token (httpserver.go:182-197). All four move behind that guard, and the two prep-cmd hook routes bind to 127.0.0.1 only. stream/started gains {slot, launch_id, client_uuid} and returns 200 (admit) or 403 {"error":"not_admitted","reason":"second_launch|stale_epoch|no_lease|wrong_client"}.

Timers, and why each value

Name Value Rationale
softHoldTTL 25s initial, 75s rolling while the holder heartbeats 25s covers probe plus handshake start; 75s is 2.5x the flatscreen's flat 30s tick (client.go:93) so one lost beat does not free a slot mid-pair
leaseTTL 90s Three missed 30s beats. Must exceed hydrabody's hardcoded 30s GracePeriod (provider.go:429) or a reconnecting client loses its slot while its session is still reactivatable. Must exceed the 45s dedup ceiling.
renewAfter 20s TTL/4.5: four attempts on the iPad's 5s streaming cadence, three on a 30s cadence
startDeadline 210s legacy, 90s protocol 2 Legacy getservercert alone can block 180s (HttpManager.m:194); a deadline below that re-creates the exact production race from the design's own numbers. Verify-then-pair reduces the migrated case to one round trip. Non-renewable.
drainGrace 45s Strictly greater than the body's 30s reconnect grace plus one 5s post
bodyIdleConfirm 2 consecutive idle reports, min 60s go-sunshine's app-name-keyed registry flips to idle the instant the first head's Undo hook fires (sessions.go:12,27-45); one idle report is not evidence
maxUnattendedStream 30 min Longer than any venue experience; bounds a body-renewed leak
unknownHold 10 min A body unreachable to the cluster may still be serving heads on the venue LAN; do not free the slot early
bodyReportFresh 90s 30s idle cadence plus 45s dedup ceiling equals a 45s worst gap; 90s is two gaps. A permanent rule, not a startup window.
bootSafeMode until every district body has reported, or 180s, whichever first Evidence-based. A pure timer restarts from zero on every boot, so a crash loop with a period shorter than the window suppresses grants indefinitely.
directiveFreshness 90s The body enforces only against a directive younger than this; older means admit-and-report
watchdog tick 10s Reuse session_watchdog's existing ticker and checkXRSessions' slot; no new goroutine
renewWriteQuantum 15s Strictly below the existing per-head 30s full-file rewrite; a persisted renewed_at is at most 15s stale, always in the fail-closed direction
implicitIdempotencyWindow 10s Shorter than the iPad's 15s request timeout, so a lost response is collapsed
bodyDemotionWindow 5 min After an abandoned or unconfirmed hold
circuitTrip / circuitOpen 3 in 10 min / 5 min doubling to 30 min Matches the flatscreen head's own consecutivePairingFails = 3 (client.go:261); 5 min exceeds a Windows boot plus Sunshine start, 30 min is below XR's 15-min-class safety valves
ticketTTL 45s renewed, 10 min absolute A visitor who walks away frees the queue in 45s; nobody waits behind a ghost
pairLeaseTTL 60s 30s getservercert timeout plus PIN retries plus slack
deviceCertOverlap 7 days One venue maintenance cycle
claimThrottle 5 granted claims / 60s Refused claims cost nothing and must not consume budget

Clock discipline

Nothing outside hydracluster compares timestamps.

  • Cluster to head: ttl_seconds, renew_after_seconds, start_deadline_seconds, all relative; expires_at is informational. The head counts down on a monotonic clock.
  • Cluster to body: (lease_id, epoch) and issued_at_seconds_ago. The body never decides expiry. Its only question is set membership (Design 4's rule, adopted).
  • Cluster-internal: one process's time.Now() on values it wrote itself, with one correction that matters. A time.Time round-tripped through gopkg.in/yaml.v3 loses its monotonic reading (verified empirically: m=+0.000209347 present before marshal, absent after unmarshal), so every post-restart comparison is pure wall clock. This bug is already live in checkXRSessions (handlers_xr.go:313 does now.Sub(sess.CreatedAt) against a YAML-loaded value). Closure: on load, clamp any persisted instant more than 60s in the future to now, which shortens a lease rather than extending it; recompute deadlines from a persisted remaining-seconds where one is available; and apply the same clamp to checkXRSessions.

Failure handling

Failure Detection Designed response Recovery
Control-plane restart / redeploy process start Leases and tickets reload from allocation.yaml; every body is UNKNOWN and occupied until it reports; boot safe mode issues no terminate directives and no grants for unreported bodies; unrecognised reported sessions are adopted, never terminated Streaming bodies reconcile in <= 5s; idle bodies within the 45s dedup ceiling; safe mode ends on evidence or at 180s
Two hydracluster processes flock fails at start, or generation mismatch on save Second process refuses to serve; a save whose on-disk generation advanced aborts, reloads and retries Immediate; downgrade_detected/generation_conflict alert
hydracluster rolled back, then forward schema_version absent on load Old binary never opens allocation.yaml, so nothing is erased; on roll-forward the new binary enters safe mode and adopts One reconcile cycle; downgrade_detected event
Crash loop body_unconfirmed counter across boots Persisted last_body_confirmation_at survives restarts, so recovery does not begin from zero each boot; a slot unconfirmed across N boots or M minutes becomes UNKNOWN with a loud alert rather than a silent refusal Bounded; the alert names the body and venue
Head crashes / force-quit / iPad reboot mid-stream head renewals stop; body keeps reporting the session Lease stays streaming (the body renews), so the picture is not killed. On session end the body reports idle twice and the lease ends ~35 to 45s after the real drop. Bounded above by maxUnattendedStream 30 min. Compare today's 2h46m orphan.
Head partitioned from cluster, still streaming head renewals stop, body report continues Not killed. The body vouches Zero interruption
Head partitioned from cluster AND body Sunshine session ends Two idle reports close the lease <= 60s
Head dies during pairing (up to 180s of body_id: nil) hold not renewed, or error heartbeat Hold released on the error status; otherwise expires at 75s of head silence and hard-capped by startDeadline <= 75s typical, 210s worst case legacy
Body reboots new boot_id Non-terminal leases move to unknown (slot stays occupied); on the body's return the lease set is re-sent at the current epoch. Never a remote reboot from us <= unknownHold, usually one report cycle
hydrabody restarts, experience still rendering boot_id change with provider_uptime_seconds < 60 Do not drain. Wait one interval for the registry and orphan scanner to repopulate One interval; closes the #429 shape
Body agent dead, hydranode alive ReportedAt older than 90s UNKNOWN, occupied, excluded from eligibility, body_unconfirmed alert Alert-driven
Clock skew / NTP step across a restart load-time clamp Any persisted instant more than 60s in the future is clamped to now; deadlines shorten, never extend Immediate
Epoch regression from a body max_epoch_seen below the cluster's Treated as lost body state: re-send the full lease set at the current epoch. Never a kill One report cycle
Cluster epoch below the body's floor body rejects a directive as sub-floor epoch_below_floor event; cluster raises slot.epoch from max_epoch_seen on the next post; the slot is not grantable until the body confirms the current epoch One report cycle; no bricking
Terminate dropped or ignored lease stays draining; body still reports the session Directive is the complete desired table, re-sent every post; the acknowledgement is the session's absence from the next report 5s per retry; escalates at drainGrace
Terminate would hit a successor impossible Directives name (lease_id, epoch); the body ignores sub-floor epochs n/a
Sunshine cert rotation mTLS verify completes with a peer cert differing from the pinned one, or the body reports a changed sunshine_cert_fp Head unpairs its own device_uid and runs one handshake; stored cert replaced only after phase 5. Cluster bumps pair_epoch so heads re-verify proactively Background, one handshake, no venue impact
Ambiguous TLS failure (packet loss, captive portal) UNKNOWN Launch fails retryably. Nothing is discarded, nothing is unpaired, no handshake starts Next attempt
named_certs growth paired_clients in the body report; first_seen ledger held by the cluster Selective UnpairClient(uuid) against a cluster-supplied desired set, cap 150 (well below the ~250 at which Sunshine's web API returns /welcome HTML instead of JSON, #371). UnpairAll deleted Continuous
Sunshine clients-list JSON drift unmarshal failure paired_digest: "unknown", loud provider_status, and the cluster refuses to issue removals while the digest is unknown Alert-driven; #521 silently did nothing for weeks
Legacy head, cluster only n/a Claiming read grants a hold and omits held bodies; assigned-config fetch is a claim point; repeat read means failover; error heartbeat means release Immediate, no client release
Legacy head reaches a body anyway E4 launch counting Second concurrent launch on a slot is refused regardless of identity; admit-observed adopts a first launch on a free slot Same request
Thundering herd at venue open claim latency histogram Critical section bounded (parse and candidate build outside the lock, CAS and save inside); idempotency collapses retries; 503 reconciling is typed and carries Retry-After instead of an empty list p99 alert at 60% of the client's 15s timeout
Claim leak (hold burned on an unreachable body) hold expires unconfirmed Body demoted for that head 5 min, one circuit-breaker failure counted; the head's next read gets a different candidate 25 to 75s
Body repeatedly fails to serve 3 start_timeout/launch_failed in 10 min. start_timeout needs no client cooperation, so this works with today's fleet Circuit opens 5 min, doubling to 30. Existing leases untouched; half-open probes with exactly one head Automatic
Capacity exhaustion no grantable slot 202 with a durable FIFO ticket for protocol-2 heads; 409 no_slot_available with an honest reason for legacy heads; denials_last_hour per venue on both axes See Capacity
Cross-venue starvation home_venue_denied counter home_venue_reserve blocks the cross-venue claim; a home-venue claim preempts a cross-venue lease that is still pre-stream Immediate
Cluster unreachable from the body admit-observed timeout Degrade to launch counting plus set membership against the last fresh directive; after directiveFreshness degrade to admit-and-report. Never refuse on ambiguity On reconnect
Duplicate head identity (restored backup) two heartbeats for one head_id from different addresses; or 409 identity_in_use on registration Second claim refused, duplicate_head_identity alert; Keychain storage prevents the clone at source Alert-driven

Observability

The data model represents the violation

sessionStore.active is re-keyed from body ID to lease id, with (body_id, slot) as a non-unique index. Two rows sharing an index is the double-booking, expressed as data. Every terminal lease carries a persisted reason and is appended to the session history JSONL.

One call that shows the truth at 09:00

GET /api/v1/slots?district=bxl1

Every (body, slot) in the district, desired beside observed, with an enumerated divergence[]. divergence: [] across the board is a healthy fleet. Anything else names the class of problem rather than leaving an operator to diff two structures by eye.

Supporting calls: GET /api/v1/allocation for capacity and queues, GET /api/v1/conflicts for live mismatches, GET /api/v1/nodes/{id}/pairing for the trust store with first_seen, GET /api/v1/heads/{id}/admission for "where is this head" as a single joined answer.

Cross-layer correlation

Two mechanisms together make the body's own logs answer "which head, which session, on which body" with no Hydra tooling:

  1. The lease id rides in-band in the Sunshine app name (<experience>#<app_token>), so it appears in sunshine.log. Design 2's idea.
  2. devicename = hydra-<head_id>, so named_devices and sunshine.log name heads instead of showing one roth for the whole fleet. Design 4's idea.

Events

slot_conflict, unauthorized_launch, second_launch_refused, epoch_below_floor, epoch_regressed, pair_contention, duplicate_head_identity, cross_venue_grant, home_venue_denied, lease_unattended_expired, body_unconfirmed, downgrade_detected, generation_conflict, custodian_failed, trim_parse_failed, queue_promotion_wasted.

Events raised by a head that has not yet migrated are tagged expected_legacy: true and counted separately. Without that tag the observe window trains operators to ignore the alarm that gates the flip to enforce.

Metrics

slots_total{venue,body}, slots_occupied{venue,body,cause}, claims_granted_total{venue,source}, claims_denied_total{venue,reason,head_venue,body_venue}, queue_depth{venue}, queue_wait_seconds{venue} (p50/p90), lease_age_seconds, seconds_since_head_renewal, claim_latency_seconds (histogram), body_report_age_seconds{body}, named_certs_count{body}, pair_attempts_total{result}, circuit_open{body}.

Alerts, with thresholds and recipients

Alert Condition Recipient Runbook action
SlotConflict any divergence containing unauthorized_stream or wrong_client for 2 consecutive minutes on-call, page GET /api/v1/slots, identify both heads, release the intruder's lease
SlotLeak any lease with seconds_since_head_renewal > 600 on-call, page Check GET /api/v1/slots; release; investigate the body's session registry
BodyUnconfirmed body_report_age_seconds > 300 on an online node on-call, ticket hydrabody health via cluster exec; check ProviderVersion
VenueDenials claims_denied_total{reason="no_slot_available"} > K/hour for 2 intervals venue owner, ticket Capacity conversation; this is the purchase-order signal
HomeVenueDenied any home_venue_denied venue owner, ticket Review the reserve; a home venue should never be denied its own hardware
EpochBelowFloor any occurrence on-call, page Distinguishes a bookkeeping fault from a pairing fault
CustodianStalled no successful pairing reconcile for a body in 30 min on-call, ticket Exec channel and Sunshine web credentials for that body
CertStoreGrowth named_certs_count > 120 on any body on-call, ticket Selective trim is failing; check trim_parse_failed
ClaimLatency p99 claim_latency_seconds > 9 (60% of the client's 15s timeout) on-call, page Lock contention; the venue is silently converting capacity into error screens
GenerationConflict / DowngradeDetected any occurrence on-call, page Two writers, or a rollback; stop the second process

How 2026-09-04 would have been caught in minutes

At 10:02 the first two iPads claim; the third receives 409 no_slot_available and its denial increments claims_denied_total{venue="sint-niklaas"}. It cannot double-book, so there is nothing to catch. If it reached the body anyway (a stale build, a hardcoded host), E4's launch counting refuses the second concurrent launch on slot 0 and emits second_launch_refused, and the cluster raises slot_conflict. SlotConflict pages after two minutes. GET /api/v1/slots?venue=sint-niklaas returns two rows on (node-4c2be4b0, 0) with divergence: ["unauthorized_stream"] and both head ids. Today the same event produced 204 lines in a journal that no API exposes, and the incident was found weeks later by a human reading it by hand.


Migration

Nine steps. Each is individually shippable and each partial state is better than the previous one.

Step 1 — hydracluster only. Fixes Sint-Niklaas. No client release.

Repo: hydracluster. Channel: server-only, normal CI/CD to the release server. No venue visit, no body binary, no WDAC exposure, no app store.

Contents, in commit order so that any part can be reverted independently:

  1. Crash-class fixes first. handleEligibleBodies reads s.bodyStatus at handlers_api.go:877 after releasing s.mu at :830, and handlers_web.go:473 lets the map reference escape the lock to be read during template render at :530. Concurrent map read and write is an unrecoverable Go runtime fatal error whose frequency scales with heads times bodies, and this step adds readers to the hottest route. The correct copy-under-lock idiom is 170 lines away at handlers_web.go:307-309. Note for the implementer: handlers_body.go:351 is not an unlocked access, contrary to several earlier write-ups; requireNodeToken takes s.mu with defer across the whole handler (handlers_body.go:45-46).
  2. BodyNodeStatus.ReportedAt, stamped server-side, plus the permanent 90s freshness rule in eligibility.
  3. allocation.yaml with flock, generation, schema_version; schema_version: 2 on nodes.yaml with downgrade detection.
  4. Slot table, epochs, leases, tickets; the level-triggered reconciler on the 10s tick and inline under s.mu.
  5. The claiming read: soft hold, omission, non-empty-experience guard, scoped head_id, idempotency on attempt_key, repeat-read failover, status-driven release, free-first sort, fnv salt, 503 reconciling.
  6. The assigned-head config fetch as a claim point.
  7. The I3 occupancy union, preserving the StreamStatus, XRState and n.XRSession terms.
  8. Adoption of unrecognised sessions; provider_uptime_seconds guard (tolerated as absent on old bodies by treating a missing value as "unknown, do not drain on a boot_id change alone").
  9. Delete the fabricated StreamStatus = "idle" writes at handlers_head.go:561-566 and :1087-1090; scope release by lease id.
  10. Delete session_watchdog.go:61's IsZero exemption; every lease has a deadline from grant.
  11. home_venue_reserve, venue and org caps, cross-venue preemption pre-stream.
  12. maxUnattendedStream, unknownHold, the load-time clock clamp (including checkXRSessions).
  13. Re-key sessionStore by lease id; GET /api/v1/slots, /allocation, /conflicts; the event set.
  14. #195: PUT /heads/{id} with stream:{} must also clear stream_url and stream_url_lan.
  15. Halve StartOfflineChecker's ticker relative to its timeout (handlers_body.go:1245), so offline detection is <= 2 minutes rather than <= 4.

Behaviour with the old fleet. Legacy iPads see a one-element list and are exclusive by omission. Legacy Go heads see a list from which held bodies are absent, so their first-reachable pick is a body the cluster chose. Server-assigned Go heads are covered by the config claim point. Nothing on any head changes.

Validation. On bxl1-test / chunky-turnip-23, never a production body: three heads claiming the same experience produce one grant and two refusals; a simulated restart adopts rather than terminates; a simulated rollback loses no leases. In production, watch slot_conflict (expect zero) and claims_denied_total at Sint-Niklaas (expect non-zero, and that is the honest shortfall).

What it does not fix. A head that bypasses the cluster entirely (hardcoded host, cached config) is not stopped until E4 ships. /resume joining is not stopped until E5 and identity ship.

Step 2 — hydracluster only. Pairing custodian.

Repo: hydracluster. Channel: server-only.

The custodian drives a body's localhost Sunshine web API through the existing hydranode exec channel that handleBodySunshinePin already proves works (handlers_head.go:459-464, curl.exe -k --http0.9 -u <district creds> https://localhost:47990/... via execPollHead). It lists clients on a 5-minute tick, persists node.pairing.clients[] with first_seen in nodes.yaml, evicts oldest-first by uuid keeping a floor of 32, alarms at 120, and treats a parse failure as a loud health signal that refuses to evict.

This is Design 5's idea and it is the highest-leverage step in the plan, because it disarms hydrabody's autonomous UnpairAll without a body release, dissolving the ordering dependency that otherwise gates the whole identity migration behind the full 6h body self-update wave. Residual, stated honestly: the 200 threshold is a body-local check the cluster cannot veto, so if the custodian is down and the count crosses 200 the hammer still fires. That is why the alarm is at 120 and why CustodianStalled exists. Also in this step: remove sunshine_username/sunshine_password from head_cache.yaml (config.go:141-150).

Step 3 — hydrabody. Telemetry and edge posts.

Repo: hydrabody. Channel: body self-update via the release mirror, new version at a new path. No in-place overwrite, no reboot, no visit (WDAC-safe). Note the rollback hazard: prep-cmd strings embed os.Executable() (sunshine_provider.go:241-248), so the reconcile must keep working across the version boundary.

  • Post status immediately on the Sunshine started and ended hooks. The ALVR path already proves an out-of-band post is safe under statusMu (postXRStatusNow, alvr_windows.go:606-616). This alone collapses the up-to-30s reporting blackout that the 5s fast tick structurally cannot, because that tick only engages after a tick has already observed streaming.
  • Send stream_count (StreamCount() exists at sunshine_provider.go:91-93 and is never called) and gpu_memory_total_mb (sysinfo.go:85-87, today sent only to hydrabodystatus). This fixes the fleet-wide parseVRAM == 0 at source.
  • boot_id, report_seq, provider_uptime_seconds, sunshine_cert_fp, paired_clients, hydrabody_version.
  • Mutex on p.streamingClientUUID, written from two HTTP goroutines and the tick goroutine and read in reportStatus with no synchronisation today.
  • Authenticate the four :47991 stream routes; bind the hook routes to loopback.
  • Fix the inert ProcessAliveCheck pid capture.

Requires hydraclusterapi v0.24.0, additive, both sides tolerating the older peer.

Step 4 — all four head repos. Removals only. Identity unchanged.

Repos: hydraheadipad (TestFlight), hydraheadflatscreen (macOS and Linux, minutes), hydraheadquest, and the HydraExperienceNet fork.

Delete every /unpair call site: HydraPairSession.m:128-156 (including the branch that returns NSData() as success, which AppState.swift:344 maps to nil and HttpManager.m:376-383 then honours by accepting any self-signed cert, logging "Pairing OK, cert 0 bytes"), PairManager.m:57-60, gamestream_pair.go:90-96, and the Quest equivalent. Add readHostCert on linux and darwin and an iOS BodyTrustStore keyed by serverUUID, never hostname. Implement verify-then-pair with UNKNOWN as a first-class result. Stop the every-30s re-pair at client.go:251-261. Fix submitPIN to parse Sunshine's {"status":false} (#495, present in the native path, not only the subprocess one) and drop the fixed 500ms/300ms PIN delay, which is wrong only because the unpair inserted a variable-length 10s-timeout request between /serverinfo and getservercert.

Ship a pairing_protocol: 2 marker in the heartbeat so the cluster can observe adoption. This step is safe in a mixed fleet by construction because the identity is unchanged. Also: correct docs/platform-parity.md:16 and docs/runbooks/runbook.md:179/:318/:742, which state that Linux uses the subprocess fallback (false since v2.2.0) and codify "pairs before every stream start, there is no cache" as intended design. Those docs are where the "omarchy does it properly" premise came from.

Step 5 — hydracluster plus hydrabody. Identity registry and the server-gated flip.

Cluster: POST/DELETE /heads/{id}/device-identity, the 7-day overlap, 409 identity_in_use, duplicate_head_identity, and per_device_identity: bool on the head config. Body: replace TrimPairedClients/UnpairAll with converge-to-desired using UnpairClient(uuid) against the cluster's trust.desired_clients and trust.remove.

The flip is a server-set flag, not a release note. The cluster sets per_device_identity: true per district only when (a) every hydrabody-role node in that district reports a ProviderVersion at or above the Step 5 version and has reported within the freshness window, and (b) every head in that district reports pairing_protocol >= 2. A body that missed its 6h self-update therefore holds its whole district on the shared identity, visibly and boringly, instead of silently arming a district-wide de-pair. CertStoreGrowth and BodyUnconfirmed chase the stuck body.

Step 6 — all four head repos. Identity and claim.

Per-device device_uid = hex(sha256(head_id))[0:16] and devicename hydra-<head_id>, gated behind the server flag, moved in all four uniqueid sites in one release (HttpManager.m:57, HydraPairSession.m:147, HydraStreamSession.m:219's /cancel URL, Utils.m:16's deviceName, and gamestream_pair.go:24). Write the uniqueid into Moonlight.conf on Go heads, which no code does today, so the agent and the streaming binary stop disagreeing.

iPad: move the client keypair and the enrolment token to the Keychain with kSecAttrAccessibleAfterFirstUnlockThisDeviceOnly and kSecAttrSynchronizable = false; exclude residual files from backup. Claim before pairing, not at .streaming. Renew from the existing 5s streaming tick. Release by lease id on every exit path including showError and showSessionInterrupted. Make cancel() actually cancel the in-flight PairManager NSOperation, which today keeps running with a 180s timeout after the user has left. Make pairing failure non-terminal: retry with jitter and advance to the next candidate rather than landing in .error, which tick() never leaves. Add the queue card (position, ETA, Cancel) and auto-start on promotion. Never /resume a session not attributable to this head's lease.

Go heads: claim-then-probe-then-launch; on 409 advance; release at all five invalidateBodyCache sites and on gamestream_pair.go:117's contention error; report body_id in server-assigned mode too.

The iPad ships via TestFlight, so the fleet runs mixed for days. Every server behaviour stays correct with old heads indefinitely; that is what Steps 1 through 3 buy.

Step 7 — hydrabody plus go-sunshine. E3 and E4.

go-sunshine release re-keying SessionRegistry from app name to (slot, launch_id, client) plus a per-session close, then a hydrabody go.mod bump. Per-lease app entries with app_token; plain-entry polarity per the E3 table; the maintenance override. The E4 admission gate with launch counting, set membership, admit-observed, directive freshness, refuse-not-kill, per-slot rate limit. slot_epoch.json. Replace ForceStop with StopSlot(i). Gate restartSunshine on "last session on this instance" (portrait_vdd_windows.go:11-22 today taskkills the whole Sunshine process on one head's disconnect).

enforce is flipped per body, per venue, starting at Sint-Niklaas, with the precondition that unauthorized_launch events not tagged expected_legacy have been zero for a full operating day.

Step 8 — capacity N (#504), chunky-turnip-23 only.

slots_total: N; N Sunshine instances; Monitors.Count=N on the single Root\MttVDD adapter, never a second adapter instance (#428's mouse bug); slot-scoped FindProcessPID dedup and killExperienceProcesses (today handleLaunch returning already_running: true at httpserver.go:222-231 is the shared-picture bug expressed at the process layer). Gated on #504's own WDAC spike: does WDAC permit a second sunshine.exe path. Capacity stays 1 fleet-wide otherwise, and unaudited head types stay fenced off.

Step 9 — fold XR, then retire legacy.

Fold NodeXRSession into leases with mode: xr behind an equivalence test preserving the shipped behaviour exactly (arm 150s, head-stale 180s, failed hold 60s, ending hold 15 min). Dual-write xr_session in nodes.yaml for one full release so a rollback keeps the Quest path intact. Then, once every head has reported protocol_version >= 2 and pairing_protocol >= 2 for 30 consecutive days, retire legacy mode. Keep the claiming read permanently as a defence against a head that regresses.


Capacity

The honest answer to 3 iPads and 2 bodies

Sint-Niklaas owns no body. Three iPads share two bodies owned by Rupelmonde. After Step 1, the third visitor is refused. That is a better failure than showing them the second visitor's picture, and it is still a worse experience than the bug looked. It must ship together with the capacity signal or venue staff will read the fix as a regression.

What each party sees.

Party Before After
Visitor 1 and 2 at Sint-Niklaas works works
Visitor 3 at Sint-Niklaas, protocol 2 head same picture as visitor 1 "You are 2nd in line, about 4 minutes", with Cancel, auto-start on promotion
Visitor 3, legacy head same picture as visitor 1 "No body available". The visitor must tap again. The queue makes the next tap deterministic and fair rather than a race
Rupelmonde visitor works works. The home_venue_reserve guarantees at least one of their own slots
Operator nothing to see claims_denied_total{head_venue="sint-niklaas"} and queue_wait_seconds
Venue owner anecdote a number, per venue, per hour, on both axes

The legacy tap-again limitation is real and cannot be closed server-side. tick() is case .error: break and every discovery, reachability and pairing failure lands there, so a legacy iPad in .error is inert until a human touches it. There is no cluster response shape that avoids .error when nothing is available. Auto-promotion arrives with the Step 6 iPad build. Until then the design reduces the pain by making reclaim fast: stopStream already awaits notifyStreamStopped, so the slot is free within about a second of the previous visitor exiting, and the FIFO ticket means the waiting head's next claim wins deterministically instead of racing a head that just walked up.

Backpressure knobs, in order of preference: queue in place; redirect to a cross-venue body (already supported, and the reserve bounds the harm to the home venue); cap per venue and per org. Stream quality is deliberately not a backpressure lever: Sunshine is single-stream per instance and there is no half a stream to hand a fourth visitor.

Extending to multi-stream bodies (#504) without redesign

slot is in the wire format, the durable model, the directive and the session key from Step 1. Capacity N is slots_total: N on the node plus the body-side work in Step 7. The legacy stream_count stays meaningful as "occupied slots", so an unmigrated head's == 0 filter keeps it off a partly-busy multi-slot body, and per #504 unaudited head types stay fenced off from multi-capacity bodies entirely.

Four body-side properties are wrong-but-survivable at capacity 1 and become mandatory at N, all in Step 7: the session registry must key on (slot, launch_id, client) rather than app name; ForceStop must become StopSlot(i); restartSunshine must be gated on "last session on this instance"; and process dedup and kill must be slot-scoped rather than image-name-scoped.

Sustained queueing is the signal that a venue needs capacity, and capacity means either a body of its own or a multi-slot body. This design is shaped so the second arrives as a configuration change.


Decisions and rejected alternatives

Chosen, against the leading design

Decision Why
Leases in a sidecar allocation.yaml, not nodes.yaml yaml.Unmarshal is non-strict and Save() rewrites the whole file, so a rollback of the emergency fix erases every lease within one head heartbeat. A file an old binary never opens makes rollback lossless. Departs from Design 1's "nodes.yaml is the only durable store".
Occupancy is the union of lease and observed stream Design 1's "stream_count means slots held" drops the StreamStatus term and advertises live, unleased streams as free. Server-assigned Go heads never claim and report body_id: "", so such streams are common.
L3 (credential vending) deleted It is not a fence. The Sunshine web credentials are district-wide (handlers_head.go:84-89, server.go:493), already cached on disk by the Go head (config.go:141-150, internal/cli/run.go:23-24), readable from any head's config with any node token, and the iPad falls back to the literals "sunshine"/"sunshine" (AppState.swift:331-332). Counting it inflated the apparent depth of the defence.
L5 (unpair the shared cert at lease end) deleted There is no shared certificate; each iPad already holds a distinct one. The body could only guess by uniqueid or wipe every roth entry, on the falling edge of every session, racing the next visitor's handshake. It rearms the fratricide being fixed, as a feature, on the normal path. Replaced by the E3 app-entry gate.
The body's fence is sunshine_client_uuid / named-cert name, not client_cert_fp go-sunshine's ClientInfo is {UUID, Name} (types.go:39-43) and UnpairClient takes a uuid (client.go:143-147). The body can never see a certificate. Design 1's stated fence was unimplementable. The fingerprint stays as the cluster-side registry key.
startDeadline 210s for legacy heads Both leading designs set it below the 180s getservercert window, which reproduced the production race from their own numbers.
Omission, not stream_count: 1 The only server-side change that reaches hydraheadflatscreen, which decodes StreamCount and reads it nowhere.
E6 re-described Sunshine mTLS refuses unpaired clients. It has never been a second-client defence: the iPads already present distinct certs. The /resume prohibition is promoted to its own layer.

Rejected, with reasons

  • Making stream_count fresher as the fix. Shrinks the window, never closes it, and every head still receives an identically ordered list. The edge post ships anyway (Step 3) as an observability and latency improvement, never as the safety mechanism.
  • A fourth session model. handleCreateXRSession is already an atomic, persisted, exclusive, level-swept, restart-surviving, typed-409 claim, and handleEligibleBodies already advertises it as occupancy. This design generalises it (Step 9) so a body has one occupancy concept, not two that must be kept consistent.
  • Reusing store.Reserve/Release verbatim. A genuine token-checked persisted lease, but the granularity is minutes, Reserve mutates without saving, and ExpireReservations has no caller anywhere in the repo: the sweep was written and never wired. The shape is copied; the code is not.
  • Bearer-token head identity. Rejected in #113/#116 as an admin nightmare. Head identity stays the explicit {id} path value or scoped ?head_id=.
  • Manual HeadBodyID pinning as the normal path. Admin override only. Selection flows through eligibility discovery; the claim is an additional server-side step in that same flow.
  • Venue isolation or district parking as the capacity mechanism. #308 offered it as zero-code, but #136 records district and venue as already overloaded as a parking hack (bxl1-test holds a production body), and #113 used district parking only as a test fixture after rejecting a drain feature as too risky.
  • Cluster-orchestrated pairing, sunshine_state.json seeding, cert transport, or a shared district identity. All four assessed and explicitly rejected in #549 in favour of a head-driven verify-then-pair reconcile: smallest delta, no cert transport, no Sunshine-internals coupling, already-paired fleets migrate by doing nothing. Rejecting a shared district identity is the opposite of keeping the shared uniqueid.
  • Body-side kill authority from local signals. Four recorded failures (#339, #551, #429, #300/#361). The body refuses launches; it kills only on a fenced cluster directive.
  • Pair on every launch, or on every tick. hydrabody v2.0.63 shipped always-pair and killed the live stream every 30s because pairWithSunshine ends in stopMoonlightStream; reverted in v2.0.67 (#250).
  • Deleting the subprocess pairing path (#278) before porting reuse natively. The reuse decision lives only in the path #278 removes. Do #278 after Step 4.
  • Inferring reuse from a subprocess exit code inside a 2s window (pairing.go:138-147). Right shape, wrong mechanism, and dead code on both platforms where cert storage is implemented.
  • Keying assignment change-detection on a resolved host address. Learned three times: #98 (LAN/WireGuard flip killed the stream every 60s), #137, hydraheadquest v0.4.0 (kill and relaunch every ~23s). Identity is the lease id; hosts are a candidate set.
  • Parking the claim token on the head record. The flatscreen head PUTs its whole view every 30s from a cached config.
  • A second mutex in hydracluster. requireAdminOrNodeToken deliberately releases s.mu before next() to avoid self-deadlock on Go's non-reentrant mutex.
  • Requiring a synchronous body ack before a grant is valid. Bodies are Windows under WDAC and self-update on a 6h cadence.
  • Replacing nodes.yaml with a database. Right eventually, wrong now; it would block the venue fix behind a storage migration.
  • Citing #371 as evidence of a wrong-UUID unpair bug. That diagnosis was retracted inside #371 itself on 2026-05-30. Its real content is cert-store overflow and the two-pass --creds recovery, and its live lesson here is the ~250-entry web-API degradation threshold.
  • Fixing #531 as filed (iPad only). The same constants are in gamestream_pair.go:24 and hydraheadquest re-pairs the same way.

Residual risks that cannot be closed

  1. /resume before the identity migration. If two heads ever reach the same body during Steps 1 to 5, /resume with the shared uniqueid still joins them. The only defence in that window is the cluster never handing two heads the same body. A head that bypasses the cluster is not stopped. Closed by Step 6 (E5) and Step 7 (E4 launch counting), subject to Open Question 2.
  2. A legacy head mid-pair holds a slot for up to 210 seconds. Bounded, visible (lease_age_seconds), and reduced to 90s per head as Step 6 rolls out.
  3. A legacy iPad in .error requires a human tap. No server-side lever exists.
  4. The custodian cannot veto hydrabody's local 200 threshold until Step 5 lands on a given body.
  5. A body whose reporter is stuck becomes UNKNOWN and therefore unusable. Correct for safety, a capacity loss in practice. BodyUnconfirmed covers it.

Open questions

  1. Does a non-zero prep-cmd Do abort a Sunshine launch? Options: (a) verify on chunky-turnip-23 at bxl1-test before arming E4; (b) assume not. Recommendation: (a). The fallback if it does not abort is per-session close, and if per-session close is unavailable the fallback is report-only. Never ForceStop.
  2. Does /resume re-run the prep-cmd Do hook? This determines whether E4's launch counting sees the production path at all. Recommendation: verify on hardware in the same testbook. If it does not, state plainly that /resume is closed head-side (E5) and identity-side only, and that E4 covers fresh launches.
  3. Which namespace is SUNSHINE_CLIENT_UUID? E4's set membership, the trust reconcile and the conflict detector all depend on it. Recommendation: verify by pairing two clients with distinct certs and reading both the hook environment and named_certs. Block Step 7 on the answer.
  4. device_uid derived from head_id, or a random per-device value? Derived is recomputable by an operator from a sunshine.log and needs no new persisted state, but changes on re-enrolment. Random (the iPad's existing IdManager value, already generated, persisted and reported as moonlight_client_id) survives re-enrolment but is opaque. Recommendation: derived. The certificate, not the name, is what must be unclonable.
  5. Sunshine web credentials. Options: (a) leave district-wide, delete the "sunshine"/"sunshine" fallback and the on-disk cache now; (b) mint per-lease credentials over the exec channel and rotate after pairing. Recommendation: (a) in Step 2, (b) as a later child issue. (b) is real security; (a) is cheap and removes the worst of it.
  6. Where does the enforce flip live: per body, per venue, or per district? Recommendation: stored per body, flipped per venue by an operator, gated on measured zero unauthorized_launch events not tagged expected_legacy for a full operating day. Design 4's lever, with the tagging fix that stops the observe window normalising its own alarm.
  7. Highly available hydracluster. Recommendation: no, not now. flock plus generation enforces a single writer and turns a silent lost update into a loud refusal. A real database is a separate project with its own risk (see what the landregistry SQLite migration cost).
  8. Does Sint-Niklaas get a body, or a multi-slot body at Rupelmonde? This design produces the number that answers it (claims_denied_total and queue_wait_seconds per venue). The decision is the owner's.
  9. Auto-start on queue promotion, or tap-to-confirm? Auto wastes a slot for up to 45s when the visitor has wandered off; confirm strands a queue behind an absent person. Recommendation: auto by default, queue_accept_mode: confirm per venue, with a queue_promotion_wasted counter to inform the choice.

Issue breakdown

Child issues under #663, ordered by dependency. Set session.venue at creation on every one; it cannot be PATCHed later (#577).

# Title Scope
1 Serialise the control plane: flock, generation counter, allocation.yaml sidecar Add an exclusive flock and a generation integer with abort-reload-retry to a new allocation.yaml. Add schema_version to nodes.yaml with downgrade detection and safe mode. Prerequisite for every write below.
2 Fix fatal-class concurrent map access on s.bodyStatus and add ReportedAt Copy-under-lock at handlers_api.go:877 and handlers_web.go:473/:530, using the idiom at handlers_web.go:307. Add server-stamped BodyNodeStatus.ReportedAt and the permanent 90s freshness gate in eligibility.
3 Slot lease store, epochs and the level-triggered reconciler NodeSlotLease, the state machine, persisted reason, epoch allocation as max(persisted, reported, generation)+1, one reconciler called on the 10s tick and inline under s.mu from every mutating handler.
4 Occupancy union rule and eligibility freshness gate Implement occupied() preserving the StreamStatus, XRState and n.XRSession terms. Observed state may release capacity, never grant it. UNKNOWN means occupied.
5 Claiming read: grant on GET /bodies/eligible with omission Soft hold, omission of other heads' bodies, non-empty-experience guard, idempotency on attempt_key, repeat-read failover, status-driven release, free-first sort plus fnv salt, 503 reconciling, admin include_leased.
6 Scope head_id to the authenticated node; move PUT /heads/{id} to self-scope Reject a mismatched head_id in handleEligibleBodies for non-admin callers. Move PUT /heads/{id} and DELETE /heads/{id}/stream to requireAdminOrSelfNodeToken; resolve the body from the lease.
7 Claim on the assigned-head config fetch Serving a stream block to a head with HeadBodyID set takes or renews a lease; omit the block when it cannot be taken. Covers server-assigned flatscreen heads, which never call discovery.
8 Stop fabricating StreamStatus idle; scope release by lease id Delete the unverified "idle" writes at handlers_head.go:561-566 and :1087-1090. Releasing a lease is what frees a body.
9 Delete the watchdog IsZero exemption; absolute deadlines on every lease Remove session_watchdog.go:61's guard. Add startDeadline, maxUnattendedStream, unknownHold, drainGrace, and bodyIdleConfirm.
10 Load-time clock clamp, including checkXRSessions Clamp persisted instants more than 60s in the future to now. Fix the same wall-clock bug already live in checkXRSessions.
11 Home-venue reserve, venue and org caps, cross-venue preemption home_venue_reserve, max_concurrent_grants, preemption of pre-stream cross-venue leases by a home-venue claim, and the cross_venue_grant/home_venue_denied counters.
12 Durable FIFO admission tickets and the denial taxonomy Tickets on the head node ordered by (created_at, ticket_id), admitted by the reconciler. Typed denial reasons on every refusal; throttle counts granted claims only.
13 Re-key sessionStore by lease id and ship GET /api/v1/slots Non-unique (body, slot) index; fleet-wide /slots with the enumerated divergence[]; slot_conflict replacing the UUID-mismatch log line.
14 GET /api/v1/allocation and /conflicts with two-axis denial counters Queue depth, median wait, denials by head venue and by body venue, open circuits. The purchase-order signal.
15 Per-body circuit breaker and unconfirmed-hold demotion Trip on 3 start_timeout/launch_failed in 10 min, open 5 min doubling to 30, half-open with one probe. Demote a body for a head for 5 min after an abandoned hold.
16 #195: PUT /heads/{id} with stream:{} must clear stream_url and stream_url_lan Named by #504 as a blocker; the same stale-assignment family recurred in #98, #137 and hydraheadquest v0.4.0.
17 Server-side pairing custodian over hydranode exec, plus GET /nodes/{id}/pairing List clients over the existing exec channel, persist first_seen, evict oldest-first keeping 32, alarm at 120, refuse to evict on a parse failure. Removes creds from head_cache.yaml.
18 hydraclusterapi v0.24.0: additive body protocol boot_id, report_seq, provider_uptime_seconds, slots[], max_epoch_seen, stream_count, gpu_memory_total_mb, paired_clients, sunshine_cert_fp, hydrabody_version; desired{} and trust{} on the response. Both sides tolerate the older peer.
19 hydrabody: edge-triggered status post on the Sunshine started and ended hooks Immediate out-of-band post under statusMu, following the postXRStatusNow precedent. Keeps the 30s tick as a backstop.
20 hydrabody: report the telemetry it already computes, and fix the streamingClientUUID race stream_count from the existing StreamCount(), gpu_memory_total_mb from the existing sysinfo, plus boot/uptime/cert fields. Mutex on p.streamingClientUUID.
21 hydrabody: authenticate the four :47991 stream routes; loopback-bind the hook routes Bring started, ended, stop and sessions under the isTrustedNetwork-or-bearer guard already used by launch, stop, kiosk and screenshot in the same file.
22 hydrabody: slot_epoch.json floor and directive application Persist the highest applied epoch per slot; refuse sub-floor directives; emit epoch_below_floor; report max_epoch_seen. Never decide expiry locally.
23 Fix hydrabody's inert ProcessAliveCheck pid capture Capture the pid on the first tick after the started hook rather than inside it, so the crash watchdog stops silently returning ActionNone.
24 go-sunshine: re-key SessionRegistry by (slot, launch_id, client) and add per-session close Module release plus a hydrabody go.mod bump. Prerequisite for E4's refusal fallback and for capacity N.
25 hydrabody: per-session teardown replacing ForceStop; gate restartSunshine on last-session StopSlot(i) instead of ending every active session. Portrait teardown must not taskkill the whole Sunshine process on one head's disconnect.
26 hydrabody: E3 per-lease Sunshine app entries with the maintenance override Publish <experience>#<app_token> for a protocol-2 lease and remove the plain entry; keep the plain entry when the slot is free or the lease is legacy; restore plain plus Desktop under maintenance: true.
27 hydrabody: E4 admission gate with launch counting and directive freshness Refuse a second concurrent launch on a slot (identity-free); set membership against a directive younger than 90s; refuse only, never kill; rate-limited to one per slot per 60s.
28 hydracluster: POST /bodies/{id}/admit-observed Run the same CAS on behalf of a client that arrived without a lease. Adopt on a free slot, 409 on a held one. 2s timeout; ambiguity means admit-and-report.
29 hydrabody: selective unpair driven by the cluster's desired set; delete UnpairAll Converge paired_clients to trust.desired_clients with UnpairClient(uuid), cap 150, loud on a clients-list parse failure.
30 Delete every /unpair call site across all four head repos HydraPairSession.m:128-156 including the empty-NSData-as-success branch, PairManager.m:57-60, gamestream_pair.go:90-96, the Quest equivalent. Identity unchanged; safe in a mixed fleet.
31 readHostCert on linux and darwin, plus an iOS BodyTrustStore keyed by serverUUID The missing reader beside the existing writeHostCert. Never key on hostname.
32 Verify-then-pair with UNKNOWN as a first-class result mTLS on 47984 only; ok/UNKNOWN/notPaired; replace the stored cert atomically only after phase 5; unpair self only on a proven peer-cert change.
33 #495: parse Sunshine's status:false in submitPIN; remove the fixed PIN delay Present in the native Go path, not only the subprocess path. Removing the unpair removes the variable-length step that made a fixed delay wrong.
34 Stop hydraheadflatscreen re-pairing every 30 seconds under an assignment Replace the unconditional tick-loop re-pair with ensurePaired.
35 Device identity registry: POST/DELETE /heads/{id}/device-identity First-write-wins with a 7-day predecessor overlap, admin-only revoke that releases leases and propagates trust.remove, 409 identity_in_use, duplicate_head_identity. Remove identity from the heartbeat body.
36 Server-set per_device_identity flag gated on observed ProviderVersion and pairing_protocol Divergence is enabled per district only when every body reports the Step 5 version and every head reports pairing_protocol >= 2.
37 Per-device Moonlight identity in all four head repos device_uid = hex(sha256(head_id))[0:16], devicename hydra-<head_id>, all five hardcoded sites moved in one release, written into Moonlight.conf on Go heads. Behind the server flag.
38 iPad: move the client keypair and enrolment token to the Keychain, excluded from backup kSecAttrAccessibleAfterFirstUnlockThisDeviceOnly, kSecAttrSynchronizable = false, plus NSURLIsExcludedFromBackupKey on residual files. Closes the restored-backup clone.
39 Pairing lease (mode=pair) and non-terminal pairing failure with failover Serialise pairing per body through the same reconciler; turn gamestream_pair.go:117 contention intoa typed 409 and a pair_contention event; make iPad pairing failure retry with jitter and advance to the next candidate instead of landing in .error.
# Title Scope
40 Head-side claim: claim before pair, renew on heartbeat, release on every exit Claim at .discovering, carry the lease through pair and launch, renew from the existing 5s streaming tick, release by lease id on stopStream, showError, showSessionInterrupted and cancel(), and make cancel() actually cancel the in-flight PairManager NSOperation.
41 Head-side: never /resume a session not attributable to own lease On _SERVER_BUSY where currentgame is not this head's lease, release and take the next candidate. Applies to iPad, Go and Quest heads. This is E5 and it is the direct fix for the observed symptom.
42 hydraheadflatscreen: report body_id in server-assigned mode; feed failover back to the server liveBodyID() must cover the assigned path, not only self-service. Emit a release with reason: unreachable at all five invalidateBodyCache sites so eligibility becomes self-correcting.
43 iPad queue card: position, ETA, Cancel, auto-start on promotion Poll GET /heads/{id}/admission every 5s in the existing timer budget. Replaces the terminal .error for the capacity case.
44 Hardware verification testbook at bxl1-test for the four Sunshine inferences On chunky-turnip-23 only: does a non-zero prep-cmd Do abort a launch; does /resume re-run the Do hook; which namespace is SUNSHINE_CLIENT_UUID; is /resume refused for a client that does not own the running session. Blocks issues 27 and 28. Delete streaming sessions after testing.
45 Fold NodeXRSession into leases with mode=xr behind an equivalence test Preserve arm 150s, head-stale 180s, failed hold 60s, ending hold 15 min exactly. Dual-write xr_session in nodes.yaml for one full release so a rollback keeps the Quest path intact.
46 Capacity N on chunky-turnip-23 (#504) slots_total: N, N Sunshine instances, Monitors.Count=N on the single Root\MttVDD adapter, slot-scoped process dedup and kill. Gated on #504's WDAC spike for a second sunshine.exe path. Capacity stays 1 fleet-wide.
47 Alerting: thresholds, recipients and runbook actions Ship the ten alerts in the Observability table as real alert rules with owners, and tag expected_legacy events separately so the observe window does not normalise its own alarm.
48 Docs correction: platform-parity.md and runbook.md docs/platform-parity.md:16 and docs/runbooks/runbook.md:742 state Linux uses the subprocess fallback (false since v2.2.0); :179 and :318 codify "pairs before every stream start, there is no cache" as intended design. Correct both in the same PR as issue 32.
49 Regression test: PUT /heads/{id} must not replace the node struct Assert that handleUpdateHead mutates named fields only, so a future refactor cannot silently eat tickets, identities or protocol markers. The flatscreen client's own comment claims the opposite.
50 Retire legacy mode Once every head has reported protocol_version >= 2 and pairing_protocol >= 2 for 30 consecutive days: stop publishing plain app entries for leased slots, drop terminate_stream synthesis, drop the dual XR write. Keep the claiming read permanently as a defence against a head that regresses.

Dependency order

Issues 1 to 16 are the Step 1 server-only release and must land together, in the numbered order, with 1 and 2 first. Issue 17 is Step 2 and depends only on 1. Issues 18 to 23 are Step 3 and depend on 18. Issues 30 to 34 are Step 4 and depend on nothing server-side, which is why they can ship in parallel with Steps 2 and 3. Issues 29, 35 and 36 are Step 5 and require 17 and 18. Issues 37, 38, 40, 41, 42 and 43 are Step 6 and require 35, 36 and 39. Issues 24 to 28 are Step 7 and require 44's answers plus 18 and 22. Issue 45 is Step 9's first half and should not be attempted during the emergency window. Issue 46 is Step 8 and requires 24 to 27. Issue 50 closes the master.

Session Context

Venue
sint-niklaas-tourism-office