As of 2026-08-04 there is no browser-based path to a Hydra stream. A user with a link and no installed client cannot see an experience.
What was destroyed today:
| Destroyed | Was | |
|---|---|---|
hydraheadwebstream |
Hetzner server 122096145, cluster node node-d76da797 |
Orchestrator + web UI at hydraheadwebstream.experiencenet.com (46.225.210.188, cax11 arm64), owned :80/:443 with its own autocert cache, v1.10.84, serving HTTP 200 |
hydraneckwebrtc |
Hetzner server 122091469, cluster node node-ca7b537d (46.225.220.240) |
WebRTC relay, controller + worker + coturn, v1.10.90. Controller/worker/coturn on the district server were systemctl disable --now'd earlier |
Correction to the framing this was filed under. These were not two generations. They were two layers of one stack that ran at the same time and depended on each other: hydraheadwebstream serve mode called POST /api/v1/sessions on hydraneckwebrtc for every session and redirected the browser to the returned stream_url. Retiring both removed one path, not two.
The real generational split is inside hydraheadwebstream:
moonlight-web-stream locally on the head device, paired it with Sunshine, and health-checked port 8080. Scaling model: one streamer process per head, configured through HydraHead.d2730ac "Delegate streaming sessions to hydraneckwebrtc" (2026-02-25) and 9f490bd "Delegate experience launch to hydraneckwebrtc" (2026-02-26). hydraheadwebstream became a thin front end — pick experience + district, find a body via hydracluster (idle-first, highest-VRAM-first), look up exe_path/orientation, create a session on the relay, redirect. All streaming state moved to the central relay.Both generations were WebRTC. There was never a non-WebRTC browser transport. The video engine in both is the same component and it is not retired.
hydra-moonlight-web-stream — the Rust fork of MrCreativ3001/moonlight-web-stream that does the real work: Moonlight/GameStream in from Sunshine, WebRTC out to the browser (H.264 default, H.265/AV1 optional; WebSocket + WebCodecs fallback). Last commit 2026-05-20, CI builds linux amd64/arm64 (gnu and musl) plus Windows, publishes to releases.experiencenet.com. It is GPL-3.0-or-later and was always run as a separate process behind an HTTP/WS boundary.
So a third attempt inherits a maintained browser player. What was thrown away is orchestration and relay hosting, not the hard part of getting pixels into a browser.
Also worth stating plainly, so nobody rebuilds it: the native head paths are the authoritative, actively maintained ones and they use direct Moonlight/GameStream with no WebRTC relay — hydraheadflatscreen (Windows/macOS/Linux kiosks, last commit 2026-05-30) and hydraheadipad + hydra-moonlight-ios (last commit 2026-05-22). The browser path's only distinguishing value is zero-install access.
Development on both retired repos stopped on 2026-03-30 and 2026-04-01. The native heads kept shipping through May. Today's server deletions ratified a stall that was already four months old — nothing was killed mid-flight.
"TURN and stuff was too fickle" is accurate but under-describes it. The record is a five-week oscillation in ICE/TURN policy, 2026-02-26 to 2026-03-27, in which the configuration was reversed at least nine times, each reversal justified by a real observed failure. The commits are the only post-mortem that exists.
1. Interface over-gathering on a multi-homed host — the dominant failure.
The relay hosts are simultaneously WireGuard peers (needed to reach bodies on 10.10.x.x) and public WebRTC endpoints. Both coturn and the engine's ICE agent auto-discover all interfaces and advertise unreachable candidates — WireGuard, loopback, IPv6 link-local. Browsers then relay to addresses that go nowhere. The signature is a first disconnect at a consistent ~9 seconds (efdb0a3). The coturn testbook makes "listening on internal/loopback/IPv6 link-local" an explicit FAIL and the runbook marks listening-ip/relay-ip as CRITICAL. This was fixed, un-fixed by accident, and re-fixed repeatedly.
2. The two available fixes were mutually exclusive in the library. This is the actual trap.
NAT 1:1 mapping is the only thing that pins candidates to the public IP and cures (1). Relay-only is the only thing that cures direct-path instability. webrtc-rs refuses to combine them — ErrIneffectiveNat1to1IpMappingSrflx with srflx, ErrIneffectiveNat1to1IpMapping with host — and fails to create the WebRTC answer at all, leaving streams stuck on "starting" (3d9e0d4, 0048d48, 8de91fe). Three separate attempts were made to make them coexist. None worked. The final state (e77fbf2, 2026-03-27) abandoned relay-only and used NAT 1:1 + srflx, on the finding that relay-only caused systematic ICE consent failures at ~40s because the server was gathering on all interfaces and STUN consent responses arrived from the wrong address.
3. Loopback TURN. coturn ran on the same box as the worker, so server-side relay produced a server → coturn → server hairpin on one machine — explicitly described as unstable (7af3188) and it doubles the host's own bandwidth.
4. Relay allocation failures and stale allocations. Under concurrent sessions the server "sometimes fails to get a relay allocation and falls back to direct, which drops after ~12s" (def97f5). coturn error 437: Mismatched allocation from stale state required a coturn restart (runbook diagnostics).
5. Static, long-lived, browser-visible TURN credentials — no TTL, no ephemeral scheme.
The provisioning recipe generates one random credential per node, persists it, writes it into /etc/turnserver.conf as a long-term user= entry, and internal/proxy/inject.go (buildICEServersJS) embeds the username and credential verbatim into the JavaScript served to every browser that opens a stream page. No REST-API time-limited credentials, no per-session scoping, no rotation. Anyone who loaded a stream page held a permanent relay credential for that node. No abuse was observed, but this is a concrete reason a third attempt must not simply restore the old config.
6. A large firewall surface, half of it discovered by failure. UDP+TCP 3478, UDP 40000-40300 (engine media), UDP+TCP 49152-65535 (coturn relay). The TCP half of the relay range was added only after audio silently failed following the Brussels migration — "old Hetzner had no firewall" (9edb6c3). Separately, rmem_max below 4 MiB causes packet loss under video load and ICE disconnects at 10-15s.
7. Some of the thrash was compensating for a bug one layer down. The last ICE commits (616948e then 1046f83) inflate the ICE disconnected timeout to 30-60s, then revert it to 15-25s with: "The signaling race was the root cause of disconnects, not ICE timing." That race was ErrSignalingStateProposedTransitionInvalid in the engine, fixed upstream in moonlight-web-stream v2.13+. A meaningful share of the TURN tuning was chasing a symptom. Do not inherit the old ICE constants — re-measure against a current engine.
8. Browser-side, documented and never resolved. Chrome/Edge check ICE consent every ~5s; on carrier networks with >200ms jitter a missed response drops to disconnected and freezes the cursor for 1-3s. Safari is the most tolerant implementation and was the best client on all networks. Firefox was never validated.
9. One measurement contradicts the whole relay premise. 2215939 (2026-03-26): "Testing showed 169ms jitter is identical on both relay and direct — the TURN relay was adding overhead without benefit for desktop users. Direct (srflx) sessions from desktop lasted 100+ seconds." The shipped policy became desktop → direct, mobile → relay, chosen by a navigator.userAgent regex. Relay was never shown to be the limiting factor for desktop.
experiencenet#69 is referenced by two commits as the tracking issue for the disconnect work; it is titled "Stream issue", filed under hydrabody, closed, empty description.relay_only=true vs the alternative. 2a58860 restores it with the justification "The user confirmed relay feels more stable on mobile" — subjective, and reversed the following day.Each worker wrote per-session telemetry to ~/.hydraneckwebrtc/telemetry.jsonl: ice_candidate_type, time_to_first_disconnect_ms, disconnect_count, jitter_ms, packet_loss_percent, applied_quality_tier, probe_round_trip_time_ms. The dedicated relay's copy died with server 122091469. The district server's copy is inside the /root/.hydraneckwebrtc/ that was deliberately kept. It is very likely the only surviving quantitative record of what failed and how often, and it directly answers the client-mix question above. Copy it somewhere durable before that directory is cleaned or rotated.
D1 — Is a browser path wanted, and for whom?
The native paths are authoritative and healthy. Browser streaming buys exactly one thing: no install. If there is no real demand for remote demos / sales / self-service / unmanaged devices, this issue closes and the DNS cleanup below is all that is needed. If there is, name the audience — the audience determines the NAT answer, and getting D1 wrong is what made D2 unanswerable last time.
D2 — Relay hosting. Four genuinely different options:
D3 — Does the media host stay multi-homed?
Every instance of failure 1 came from a host that was a WireGuard peer and a public WebRTC endpoint. Either candidate pinning becomes an invariant enforced and asserted at startup (refuse to serve if coturn is bound to a non-public address, refuse if public_ip is unset) rather than a config field someone forgets — or the roles are separated so the media host is not on the mesh at all.
D4 — Keep hydra-moonlight-web-stream as the engine?
Default should be yes: maintained, arm64+musl, already carries the Hydra-specific fixes, and the v2.13+ signalling fix retires failure 7. GPL-3.0-or-later means it stays a separate process behind HTTP/WS — a constraint the old design already respected. The alternatives (a native WebRTC sender, or Sunshine → WHIP) are an order of magnitude more work and would discard the one asset that survived.
D5 — One service or the controller/worker split again?
The two-tier split existed to spread sessions across relay hosts. With one district, one relay, and a hard one-session-per-body limit, a controller may be unearned complexity — it also introduced its own bugs (session forwarding, queue promotion, ghost sessions). Note the body limit is not a browser-path problem and will not be fixed by this work: Sunshine uses DXGI Desktop Duplication, so two sessions on one body fight over the same desktop (parallel-streams testbook, 2026-03-21). That is bounded by hydrabody/VDD work — #82, #84, #418.
D6 — What is the measurement plan, before any code?
The stall was five weeks of flipping settings with "the session felt stable" as the signal. Define the pass bar first (e.g. P90 session survives N minutes with zero disconnects across browser × network matrix), reinstate the telemetry schema — it was good — and treat the recovered telemetry.jsonl as the baseline. The coturn-health testbook is reusable as-is and should be a deploy gate.
District ingress (new). hydrascalerouter wraps Traefik on the district server (141.227.136.199, OVHcloud b3-16 Brussels), terminating TLS with ACME HTTP-01 and doing Host() routing, with active health checks. This genuinely removes work: the orchestrator/UI half no longer needs to own :80/:443 or its own autocert cache, which is precisely what hydraheadwebstream did. That half is now a natural scale.
It does nothing for media. Traefik is configured with entrypoints :80 and :443 only (internal/traefik/install.go) — there is no UDP entrypoint, and it is not a load balancer either (registry.Route holds a single Backend; Render emits one server per service).
arm64 scale fleet (new, but less new than it looks). hydraskin turns Pi 5 / N150 boxes into Incus hosts and arm64 is first-class. The engine builds aarch64 gnu+musl, hydraneckwebrtc CI already built linux-arm64, and hydraheadwebstream was already running on an arm64 Hetzner cax11. arm64 is not a blocker. The blocker is different: hydraskin expose publishes TCP only — internal/cli/expose.go hardcodes tcp: in the Incus proxy device. A scale cannot publish UDP 3478 / 40000-40300 / 49152-65535 today.
So under current tooling: orchestrator as a scale — yes. Relay as a scale — no, without adding UDP expose. And even with it, a Pi behind a venue's consumer NAT is the worst possible position to terminate public WebRTC from — it maximises exactly the traversal problem that killed the last attempt.
Does the district hub help NAT traversal? No — and there is a caution. The hub solves relay→body reachability, which already worked over WireGuard and was never the failing leg. The failing leg is browser↔relay, which the hub does not touch. Note also that commit d9b0cd6 named "aggressive NAT rebinding on mobile CGNAT or OVHcloud network behavior" as the suspected cause of the ~9s keepalive loss — and the district hub is an OVHcloud Brussels box. If a relay is placed on or beside the hub, test that suspicion first rather than inheriting it.
hydraheadwebstream.experiencenet.com → destroyed host (46.225.210.188, Hetzner 122096145). Stale A record pointing at an IP the provider will reassign. Remove or repoint.hydraneckwebrtc.experiencenet.com → 141.227.136.199, the live district server, where hydraneckwebrtc, hydraneckwebrtc-controller and coturn are all disabled. A dark domain resolving to an important host. Remove or repoint./root/.hydraneckwebrtc/ before any redeployment. It holds the worker admin token, the controller token, Sunshine credentials, and the TURN username/credential; the TURN secret also exists at /root/.hydraneckwebrtc-turn-credential and inside /etc/turnserver.conf. Given failure 5, the TURN credential should be assumed disclosed to every browser that ever opened a stream page. Treat the directory as an artefact to salvage from and then destroy — not to restart from./root/.hydraneckwebrtc/telemetry.jsonl first (see above) — it is the evidence base for D6.hydraneckwebrtc/docs/testbooks/{parallel-streams,session-protection,single-stream-lifecycle}.md and hydraheadwebstream/docs/testbooks/single-stream.md. These are in git history, so deletion is not sufficient. One of them is a hydrabodystatus token, and hydrabodystatus is still live — rotate that one regardless of what happens to browser streaming.hydraheadflatscreen/CLAUDE.md, hydrabody/CLAUDE.md, hydraneck/docs/runbooks/{overview,traffic-flows,mesh-participation}.md, hydracluster/docs/runbooks/body-recovery.md, and six files under hydrastreamingmonitor/docs/ including testbooks/browser-streaming-e2e.md and runbooks/streaming-role.md, which health-check hydraneckwebrtc.experiencenet.com as part of a "fleet healthy" run. That check now fails permanently.proposed).Decision support only. It deliberately contains no implementation plan — D1 is a real question and may be answered "no". The point is that the next attempt starts from the nine numbered failure modes and the recovered telemetry, rather than from a blank editor and a five-week rediscovery of ErrIneffectiveNat1to1IpMapping.