HydraIssues

Proactive game pairing: head-driven verify-then-pair reconcile loop (district trust mesh)
open bug Project: hydraheadflatscreen Parent: #663 Reporter: 25 Aug 2026 13:11

Description

Summary

Investigation (2026-08-25) into whether stream-time game pairing overcomplicates the fleet, and whether pairing should be set up proactively when a body or head joins a district.

Finding: every head type pairs on EVERY stream start (hydraheadflatscreen client.go:251 "Pair before every stream start"; iPad AppState.swift:326 "Always pair... no caching"). Nobody verifies stored trust, so trust is treated as disposable, which is also why hydrabody TrimPairedClients can unpair ALL clients. Cluster-orchestrated pairing (option a), sunshine_state.json seeding (option b), and shared district identity (option c) were assessed and rejected as primary mechanisms; the full analysis with file:line citations lives in the investigation report (sections 1-3), summarized below.

DECISION PROPOSED: proactive trust via a head-driven verify-then-pair reconcile loop. Stream start becomes discovery + launch; stream-time pairing remains only as a rarely-hit self-heal fallback.

4. Recommendation

Proactive pairing via a head-driven reconcile loop, using the existing native pair path, with a cheap verify step. Keep stream-time pairing as the self-heal fallback.

Concretely: each head periodically fetches its eligible-bodies list (the endpoint already exists and already encodes district + venue + owner policy, handlers_api.go:785-899) and, for each body, verifies the pairing (mutual-TLS GET https://host:47984/serverinfo with the client cert presented and the stored server cert pinned; PairStatus == 1 means both directions are healthy). Only on verify failure does it run the existing pairWithSunshine, with jitter and backoff to respect Sunshine's single pair session. Streams then start against an already-trusted body: discovery, launch, no pairing.

Why this over the alternatives:

  1. It is the smallest delta that removes pairing from the critical path. No new cluster orchestration, no cert transport, no Sunshine-internals coupling. The pair mechanism that already works (native, ~1 s, in-process) just runs at a different time.
  2. Verify-then-pair fixes the churn, not just the timing. Today the fleet re-pairs every session because nobody verifies stored trust. A verify step makes stored trust first-class, which also lets the unpair-first behavior (gamestream_pair.go:90-96) and the iOS always-pair rule retire.
  3. It matches the fleet's architecture. It is exactly the tickSunshineApps pattern, pointed at trust instead of apps. No event infrastructure needed. Already-paired fleets migrate by doing nothing: verify passes, loop is idle.

The honest overcomplication answer: today's complexity is roughly half timing, half mechanism. Proactive pairing deletes the timing half (visitor-facing latency, pairing state machine branches, fail-count fallback, always-pair-fresh rules, trim-then-repair coupling). It does not delete the mechanism half: gamestream_pair.go stays, the Windows subprocess path stays until #278 lands native Windows cert storage, and drift handling is new code that did not exist before. The stream-time pair call should remain as a fallback branch (verify at stream start; if broken, pair on demand exactly as today), so a body that was wiped seconds ago still streams. What changes is that the fallback almost never runs.

New failure modes and their answers:

  • Body offline at head-join time: no-op; the reconcile loop retries next tick. This is strictly better than today, where the same body fails at visitor time.
  • Sunshine cert rotation / state wipe / reinstall: verify fails, loop re-pairs within one reconcile interval. Stream-time fallback covers the window.
  • Unpaired drift from TrimPairedClients: change it from unpair-all to selective (unpair only certs not in the district mesh, or oldest-first beyond the cap). With a stable mesh the count equals the head count and never approaches 200; today's count explodes only because web/one-off clients accumulate.
  • District moves (e.g. bxl1-test parking): the eligible list changes next tick; the head pairs with the new district. Stale trust on former-district bodies is inert (certs without network reachability or assignments). Optional hygiene: the head unpairs bodies that left its list, or hydrabody prunes certs it no longer expects.
  • Pair-session collisions (N heads pair one body after its trust wipe): jittered start plus retry-with-backoff; collisions return a clean "another pairing attempt is already in progress" error and resolve in seconds. Off the critical path, this is acceptable.

5. Implementation sketch

Phase 1: hydraheadflatscreen (no cluster or body changes required)

  1. New pkg/client/pair_reconcile.go: a loop on its own ticker (suggest 5 min, plus a run on config fetch and on discovery failure). For each body from GET /api/v1/bodies/eligible?head_id=...:
    • verifyPaired(host): mTLS serverinfo on 47984 using readClientCertKey() and the srvcert stored for the body's UUID; healthy = PairStatus 1 and cert match.
    • On failure and not currently streaming to that body: pairWithSunshine(host) with random 0-30 s jitter.
    • Skip the whole loop while a stream is live on this head (same guard as client.go:258).
  2. startStream (localapi.go:226-239) and the tick pair block (client.go:251-284) become verify-first: verify (fast, no side effects); only pair when verify fails. The pairing status stays as the rare fallback branch.
  3. Drop the unpair-first step for healthy pairs (gamestream_pair.go:90-96) once verify exists; keep it inside the fallback re-pair.
  4. Report paired_bodies (body ID, verified_at, ok/fail) in the existing heartbeat for observability.

Phase 2: hydrabody

  1. Replace TrimPairedClients unpair-all with selective trim via UnpairClient(uuid) (go-sunshine already has it, client.go:144). Desired-set input: head names/uuids from the existing fetchHeadsForVenue loop, extended to district scope.
  2. Optionally report GET /api/clients/list contents in body status so the cluster admin UI can render mesh health (desired vs actual per district).

Phase 3: hydracluster (observability only, no orchestration)

  1. Admin page: district trust-mesh view from head heartbeats + body status. A red cell means "this pair will hit the fallback path". No new pairing endpoints. The existing sunshine-pin exec proxy stays for iOS/emergency use.

Phase 4: other heads

  1. hydraheadipad: same verify-first change in startStream (AppState.swift:324-352); keep HydraPairSession as fallback. Optionally add its own background reconcile.
  2. hydraneckwebrtc: verify-first around its PairWithSunshine.
  3. Windows heads: unchanged behavior but now off the critical path; #278 (native cert storage on Windows, registry-backed) deletes the subprocess path later.

Migration: none needed. Already-paired heads verify clean on the first loop. Bodies keep their current trust stores. Rollout is head-agent version driven, per the normal release flow.

Omarchy marketplace plugin: unaffected by construction. Standalone mode drives stock moonlight-qt with a human-entered PIN and never talks to a cluster (/home/claude-user/omarchy-hydrahead/README.md:16-20, 42-47); the reconcile loop lives in the hydraheadflatscreen agent, which standalone mode does not run. ExperienceNet mode inherits the benefit automatically when it adopts the agent.

6. Open questions

  1. Multi-stream bodies (#504): N Sunshine instances per body means N trust stores and N port sets. Does the mesh pair per instance, or does the slot allocator hand out (host, port-offset) and the head pairs the assigned instance on demand? Suggest: mesh covers instance 0 now; decide with #504.
  2. Mesh scope: "district" in practice means the eligible set (district AND (same venue OR same owner), handlers_api.go:849-861). Pairing exactly the eligible set is proposed here; confirm that cross-venue same-owner bodies should be in the proactive mesh.
  3. Does the pinned Sunshine version's HTTP (47989) serverinfo reliably report per-client PairStatus, or must verify always use mTLS 47984? Assume 47984; confirm on chunky before building.
  4. Stale-trust hygiene on district moves: silently accumulate, head-side unpair on list exit, or body-side prune? Cheapest correct answer is body-side selective trim (Phase 2.5) plus doing nothing head-side.
  5. Interval tuning: 5 min reconcile vs Sunshine wipe windows. Is a body-restart-triggered "verify now" hint (body status change observed via head config fetch) worth the coupling?
  6. iPad heads have no WireGuard interface and reach bodies via the mobilekit hydraneck; confirm 47984/47989 reachability for verify on that path matches today's pair reachability.

Prior art shipped in the omarchy plugin (2026-08-26)

The HydraHead omarchy plugin (omarchy-hydrahead, commit dc7ce80) now implements exactly this issue's core principle at the UI level: its Pair action verifies first and pairs only on failure. Clicking Pair runs a cheap trust check (moonlight list succeeds only against a host that already trusts the client); when trust exists the plugin skips pairing entirely, tells the user "Already paired, no PIN needed", and loads the app list; only a genuinely unpaired host enters the PIN flow. This came from live user testing: the raw pair-always flow surfaced moonlight's confusing "already paired" toast whenever trust already existed.

Takeaways for this issue's implementation: (1) verify-then-pair works and the verify signal (a successful authenticated list/serverinfo) is reliable in practice; (2) the same confusion the reconcile loop prevents fleet-side was independently hit by a human user within minutes of manual testing, which strengthens the case; (3) the plugin's standalone implementation stays UI-triggered, while this issue remains the background reconcile for the managed fleet, so the two do not overlap.