HydraIssues

[SUPERSEDED by #116] Cross-venue body eligibility by owner + WG path validation test
closed feature Project: hydracluster Reporter: cederik 23 Apr 2026 13:27

Description

Context

Validate the WireGuard fallback branch of hydraheadflatscreen end-to-end, via the normal body-selection procedure (no manual HeadBodyID pinning). Target: peppy (Mac mini, bxl1/rupelmonde, node-74adf4f8) streaming from cosmic (Windows body, bxl1/cloud-seven, node-4c2be4b0) — different subnets, so inherently WG-only.

Two gaps block this today:

  1. Selection is venue-strict — handleEligibleBodies at hydracluster/pkg/api/handlers_api.go:742 filters body.venue == requested_venue. Peppy in rupelmonde cannot see cosmic in cloud-seven.
  2. No clean way to skip boom-pickle from selection (drain feature deferred — could break things).

Defer: do NOT run before Friday go-live. Park work afterwards.

Feature A — Cross-venue-by-owner eligibility (hydracluster)

Server-side, new optional ?head_id= query param on /api/v1/bodies/eligible. When present, server looks up the head and filters: district==head.district AND role=hydrabody AND status=online AND (n.Venue==head.venue OR n.Owner==head.owner). Same-venue matches sort first. When head_id absent, preserve current strict-venue behavior for backwards compatibility.

  • NO bearer-token-based head-identity resolution server-side — explicit ?head_id= query param only. Tokens are an admin nightmare.
  • Add same_venue bool to JSON response for client-side debug.
  • Extend sort at handlers_api.go:764-771: same-venue first, idle-first, VRAM-desc.
  • Unit tests: same-venue / same-owner-different-venue / different-owner / different-district / missing head_id.
  • Runbook entry in hydracluster/docs/runbooks/ documenting the cross-venue-by-owner rule.

Feature A — hydraheadflatscreen client

One-line change in pkg/client/discovery.go:24-28 — append &head_id= to the URL. head_id already in Config. Testbook entry in hydraheadflatscreen/docs/testbooks/.

Skip boom-pickle for the test — via district move

Do NOT drain. Instead, temporarily move boom-pickle to the existing bxl1-test district (real district in the catalog, id=bxl1-test, name=Brussels Test, on-premise) for the test window, then back. One POST /api/v1/nodes/node-74f9fbf2/district call each way, zero touch on the boom-pickle machine itself.

Risk to validate BEFORE running on boom-pickle: does hydranode react to a district change? Provisioning is per-role so typically no, but unconfirmed. Either grep hydranode for district-reactive paths first, or test the move on a lower-stakes node.

Deploy

Per feedback_no_manual_deploy: tag + push each repo, GitHub Actions → release server → auto-update. Full kiosk reset on peppy after hydraheadflatscreen update per runbook section 'Full kiosk reset after update (macOS)'.

Test procedure

Phase 0 — preflight: confirm hydracluster new build, peppy agent >=v2.0.25 AND on the new build sending head_id (grep log for bodies/eligible?.*head_id=), wg show peer 10.10.100.12 up, baseline eligibility returns both bodies (boom-pickle first), bxl1-test in districts catalog, record boom-pickle's bxl1/rupelmonde for restore.

Phase 1 — move boom-pickle to bxl1-test: POST /district {"district":"bxl1-test","venue":"rupelmonde"}. Eligibility now returns only cosmic. Full kiosk reset on peppy to clear cachedBody/cachedBodyIP/currentHost (v2.0.25 flip bug caches).

Phase 2 — observe WG branch: tail /tmp/hydraheadflatscreen.log, trigger stream with cosmic experience, expect log literals (from pkg/client/config.go:61,64,68):

  • LAN host 11.0.11.24 not reachable, falling back to WireGuard
  • routing via WireGuard: 10.10.100.12
    Screenshots at 0s/30s/2min/5min/6min. 6-min stability — no 'stream target changed', no 'stopMoonlight', no Moonlight subprocess exit. pairing.yaml cert against 10.10.100.12.

Phase 3 — restore (run even on failure): POST /district back to bxl1/rupelmonde. Full kiosk reset. Stream returns to 'routing via LAN: 10.110.0.51'. Eligibility call shows boom-pickle first again.

Targets (owner visit.flanders)

  • peppy-dumpling-32 node-74adf4f8 bxl1/rupelmonde LAN 10.110.0.53
  • boom-pickle-38 node-74f9fbf2 bxl1/rupelmonde LAN 10.110.0.51 (to be parked in bxl1-test for test window)
  • cosmic-pretzel-98 node-4c2be4b0 bxl1/cloud-seven LAN 11.0.11.24 WG 10.10.100.12

Success criteria

  • Feature A active: eligibility with ?head_id=node-74adf4f8 returns both bodies, boom-pickle first.
  • After move: eligibility returns only cosmic.
  • 'routing via WireGuard: 10.10.100.12' in peppy log via normal selection (no manual pinning).
  • Cosmic content on peppy display >=6 min, 5 screenshot checkpoints, no flip-bug markers.
  • pairing.yaml pairs with 10.10.100.12 during the test.
  • After restore: boom-pickle back to bxl1/rupelmonde, peppy back on LAN.

Follow-ups

  • Testbook in hydraheadflatscreen/docs/testbooks/ + runbook in hydracluster/docs/runbooks/.
  • Drain feature as separate scoped work when safer (would make district-parking obsolete).
  • Consider WG probe in discoverBody so client iterates past unreachable bodies (currently trusts WG IP unconditionally when LAN fails).

Plan file: /home/claude-user/.claude/plans/expressive-sniffing-pearl.md