Symptom
ipad-head-cederik could not pair with cosmic-pretzel-98 at cloud-seven (2026-08-20). The app hung in the pairing step on venue Wi-Fi and on 5G. The head sent /pair phrase=getservercert to Sunshine again and again. The PIN reached /api/pin. Pairing never completed. Sunshine logged no pairing error.
Root cause
Two sunshine.exe processes ran at the same time on the body. Both processes listened on ports 47984, 47989, and 47990 (Windows permits the double bind). Windows gave each incoming connection to one of the two processes at random. The getservercert request landed in one process. The PIN landed in the other process. The two in-memory pairing session maps never matched, so pairing could not complete on any network.
The duplicate instances come from a race in hydrabody (pkg/provider/vdd_windows.go, restartSunshine):
- restartSunshine runs taskkill /f /im sunshine.exe, then sleeps 3 seconds, then starts its own instance.
- In that 3-second window, the periodic ensureSunshineRunning tick (sunshine_startup_windows.go) sees that sunshine.exe is not running and starts a second instance.
- On cosmic, the Windows service SunshineService (wrapper sunshinesvc.exe) was also running and respawned a third contender in the same window.
restartSunshine fires after every stream end (the VDD reset path), so a body re-enters the broken state often.
Evidence
- cosmic-pretzel-98: PIDs 3236 (parent hydrabody) and 24896 (parent sunshinesvc), both started 04:01:07, both bound to all three ports. sunshine.log showed interleaved writes from two file offsets with the same timestamps.
- boom-pickle-38: PIDs 18916 and 19120, both started 01:24:31, with SunshineService stopped. This proves the race exists inside hydrabody alone.
- chunky-turnip-23: single instance since 8/3, healthy.
Remediation applied (2026-08-20)
- cosmic-pretzel-98: killed the duplicate, stopped and disabled SunshineService. hydrabody restarted a single instance (PID 19064) in interactive session 1.
- boom-pickle-38: killed both instances. hydrabody restarted a single instance (PID 6284).
Proposed fix in hydrabody
- Add a mutex or in-restart flag so ensureSunshineRunning does not start Sunshine while restartSunshine is in its kill window.
- After taskkill, poll until the process count is zero instead of a fixed 3-second sleep.
- At provision time, disable SunshineService if it exists, so the Windows service wrapper never competes with hydrabody for the Sunshine lifecycle.
- Optional guard: before starting Sunshine, check that ports 47989/47990 are unbound; log a warning and skip the start if another instance already listens.
Resolution (2026-08-20)
Fixed in hydrabody v2.0.67 (commit 62cd084): sunshineLifecycleMu serializes restartSunshine and ensureSunshineRunning, the restart polls for process exit instead of a fixed 3-second sleep, and both start paths disable and stop SunshineService. Runbook entry added: docs/runbooks/runbook.md, section 'Duplicate Sunshine instances'.
Deployed and verified on cosmic-pretzel-98, boom-pickle-38, and chunky-turnip-23: each runs hydrabody v2.0.67 with exactly one sunshine.exe. Note for future updates: the release-server update replaces the binary but the running task keeps the old image; bounce the HydraBody scheduled task (schtasks /End then /Run) to load the new version.