HydraIssues

ALVR client fails to connect to armed server over the mesh (connect_timeout)
done bug Project: hydrabody Parent: #544 Reporter: cederik 31 Aug 2026 17:42

Description

Single open blocker to first immersive success via the pipeline (parent hydraheadquest #544). As of 2026-08-28 the body arm fully succeeds (rc.6: vrserver up via vrstartup.exe, driver registered, client 0529.client trusted, 'armed waiting for client'), but the ALVR client on quest-head-36afcd never completes its connection within the 2-minute window and the session tears down with detail=connect_timeout; the client sits in its night-sky lobby and reports 'streamer disconnected'.

The identical setup streamed successfully on 2026-08-26 (first immersive stream, museum confirmed), so this is environmental/timing, not a code regression. Aggravators observed 2026-08-28: headset on HOME wifi (192.168.68.x) routed to venue chunky only via the WireGuard mesh; SteamVR boundary-setup prompt interrupts on movement; headset sleeps (screen off) mid-attempt.

Next debug (NOT more taps):

  1. Verify ALVR UDP 9943 (control) and 9944 (stream) actually traverse the mesh chunky (10.10.100.16) -> headset (10.10.200.7). ping/ICMP works; UDP is unverified. Server trusts the client with manual_ips=[10.10.200.7] and dials it.
  2. Raise xr_connect_timeout (BodyConfig xr_connect_timeout_seconds, default 120) so boundary/positioning fumbles do not kill the window.
  3. Confirm whether the ALVR client needs the server (10.10.100.16) added on ITS side, vs relying purely on server-dials-client via manual_ips.
  4. Check ALVR client Rust logs on connect attempt (via hydraadb: incus exec hydraadb -- adb -s 10.10.200.7:5555 logcat).

Body left clean flat-safe after the session (Sunshine up, ALVR driver unregistered, no session).

Debug session 2026-08-31

Findings (headset asleep, so live UDP test deferred):

  • ALVR server config on chunky: stream_protocol Udp, stream_port 9944, client_discovery on, packet_size 1400, client 0529.client trusted with manual_ips [10.10.200.7] (connection_state Streaming, but that is stale from the Aug 26 success).
  • The head tun0 mesh IP is confirmed 10.10.200.7 (matches manual_ips). Headset is on HOME wifi 192.168.68.x, so home->venue only via the mesh.
  • Hypothesis: ALVR packet_size 1400 + 28 (UDP/IP) = 1428 exceeds the 1420 mesh tunnel MTU, forcing IP fragmentation which hub-and-spoke WireGuard tends to drop. Fits the symptom (control connects, brief SteamVR void, then video stream fails and the client drops to lobby). CAVEAT: the same 1400/1420 combo streamed on Aug 26, so this may not be the sole cause; treated as a variable to remove, not a confirmed fix.

Change made (manual, to test; NOT yet in hydrabody): lowered ALVR session.json packet_size 1400 -> 1280 while the dashboard was stopped so it persists into the next arm. If this fixes the connect, productize by setting packet_size for mesh clients in the hydrabody arm sequence (#557).

Next step (needs the headset worn + awake): one clean tap of steamvr-perftest with the lower packet_size, while doing the live UDP diagnostics: from chunky send DF pings and confirm UDP 9943/9944 reach 10.10.200.7 during the connect window, and pull ALVR client Rust logs via hydraadb. Prior data point: on 2026-08-25 DF pings from chunky to the headset passed up to 1392 bytes (path MTU >= 1420).

RESOLVED 2026-09-05 (hydrabody v2.0.70-rc.9)

Root cause was NOT network: alvrClientConnected returned false and bailed early whenever the ALVR dashboard API (127.0.0.1:8082) was unreachable, which it is for the whole armed window right after chain start (same connection-refused as the trust call). The UDP 9944 stream-socket fallback was only reached on a narrow read-error path, so a genuinely connected, streaming client was never detected and every session hit connect_timeout and killed the live stream. Earlier detector fixes (rc.7/rc.8) did nothing because the code never reached them; found by reading the supervise path and confirming ZERO 'connect check' log lines fired.

Fix (rc.9): the UDP 9944 stream-socket check (netstat, foreign addr = client ip) is now the PRIMARY signal and runs every armed tick regardless of the dashboard. CONFIRMED LIVE 2026-09-05: 'connect check: socket=true dashboard=false' detected the connection in 5s, session went active. dashboard stayed false throughout, proving the root cause.

Productize: bake rc.9 into a hydrabody release (currently staging rc.9 on chunky only). Note the trust call also races the dashboard startup (attempt 1/3 refused, 2 succeeds) - fine as-is.

Session Context

Body
chunky-turnip-23