HydraIssues

setSunshineConfig destructively rewrites sunshine.conf, losing output_name on next restart
closed bug Project: hydrabody Reporter: claude (cederik) 29 Apr 2026 19:22

Description

Symptom

On 2026-04-29, the WebRTC viewer of cosmic-pretzel-98 (cloud-seven) showed a black screen with the Windows taskbar visible while a Rupelmonde stream was active. The body screenshot endpoint (/api/v1/debug/screenshot, shipped today in v1.11.21) showed Rupelmonde rendering correctly on the body's primary monitor — so the body was fine, Sunshine was capturing the wrong display.

Diagnosis:

  • sunshine.conf on disk had only dd_manual_resolution = 1920x1080. No output_name, no ensure_primary, no audio_sink, no encoder hints from the hydracluster base config.
  • Sunshine PID had StartTime = 2026-04-21 16:56:34 (8 days earlier).
  • sunshine.log at that startup line: Info: config: 'output_name' = {9acddf6d-43cc-576e-9aff-0c5fc80b4cc8}.
  • So Sunshine was capturing a fixed DXGI Output GUID that no longer matches the display where Rupelmonde rendered (the original VDD output is no longer that GUID, or the GUID now points to a blank/desktop output).

Root cause

Two independent code paths own sunshine.conf and they conflict:

  1. writeSunshineConfig (pkg/provider/vdd_windows.go:393) writes the full config (hydracluster base + hydrabody-managed section with output_name, audio_sink, dd_configuration_option=ensure_primary). Runs once on first VDD provisioning, gated by the marker file ~/.hydrabody/vdd_sunshine_configured.txt. Calls restartSunshine after.
  2. setSunshineConfig (pkg/provider/streamprovider.go:55) calls client.SetConfig(cfg) which POSTs partial keys to Sunshine's /api/config. Sunshine's API replaces the on-disk sunshine.conf with only the keys posted. Sunshine's running in-memory config retains everything else. Called every stream/started in handleStreamStarted (pkg/provider/httpserver.go:75) to set dd_manual_resolution.

So each stream-start truncates sunshine.conf to just dd_manual_resolution=.... The running Sunshine still has correct in-memory state. The bug is latent until the next Sunshine restart — auto-update, crash, body reboot, manual kill. Then Sunshine loads the truncated file, loses output_name, and starts capturing the default DXGI output (not the VDD where experiences render). Result: black + taskbar in the stream, while the body itself looks fine.

Verified path with GET /api/config against Sunshine's API on a healthy body: returns the merged in-memory config including output_name and audio_sink. So Sunshine has the data — it just doesn't persist anything we don't post.

Reproducer

  1. Provision a body with VDD enabled. sunshine.conf contains output_name = {GUID}. Stream works.
  2. Start any experience. handleStreamStarted calls setSunshineConfig({"dd_manual_resolution": "1920x1080"}).
  3. Get-Content C:\Sunshine\config\sunshine.conf — only dd_manual_resolution = 1920x1080 remains.
  4. Stream still works while this Sunshine instance lives.
  5. Restart Sunshine (Stop-Process sunshine; hydrabody auto-restarts).
  6. Stream now shows black + Windows desktop bar. Body screenshot endpoint shows the experience rendering fine.

Proposed fix (primary)

Make setSunshineConfig non-destructive by fetching, merging, and posting the full map:

func (p *Provider) setSunshineConfig(cfg map[string]string) error {
    sp, ok := p.streamProvider.(*SunshineStreamProvider)
    if !ok || sp == nil {
        return nil
    }
    current, err := sp.client.GetConfig()
    if err != nil {
        return fmt.Errorf("get current sunshine config: %w", err)
    }
    for k, v := range cfg {
        current[k] = v
    }
    return sp.client.SetConfig(current)
}

This keeps the on-disk file complete because Sunshine's /api/config POST replaces the file with the merged map equal to original+new.

Defense-in-depth fix (secondary)

Drop the marker short-circuit in ensureSunshineVDDConfig (pkg/provider/vdd_windows.go:58-61) so every tick verifies the on-disk content has output_name AND ensure_primary and re-runs the full writeSunshineConfig + restartSunshine path if not. The marker becomes an optimization for skipping discoverVDDDeviceID, not a permanent skip. Caveat: without the primary fix, this would restart Sunshine ~30s after every stream/started because the API call truncates the file. Land the primary fix first; only then is this safe to add.

Manual recovery (current workaround, documented in runbook)

Until the fix lands:

  1. Remove-Item "C:\Windows\System32\config\systemprofile\.hydrabody\vdd_sunshine_configured.txt" -Force
  2. Stop-Process -Name sunshine -Force
  3. Wait ≤30s for tickVirtualDisplay → ensureSunshineVDDConfig → rediscovers VDD → rewrites full conf → restarts Sunshine.

Verified working on cosmic-pretzel-98 today.

Test plan after fix

  1. Healthy body, full sunshine.conf with output_name. Capture conf hash.
  2. Start experience. Verify /api/config returns the merged config (not just dd_manual_resolution).
  3. Capture conf hash — must match step 1 (or differ only in dd_manual_resolution line if orientation differs). output_name must still be present.
  4. Stop-Process sunshine -Force. Wait for hydrabody to auto-restart.
  5. Start a new stream. Stream must show the experience, not black + taskbar.
  6. Body screenshot endpoint and WebRTC viewer must show the same content.

Affected nodes (today)

  • node-4c2be4b0 (cosmic-pretzel-98), bxl1/cloud-seven — recovered manually today
  • All VDD-enabled bodies running hydrabody ≤ v1.11.21 are vulnerable

Related

  • Runbook section added in v1.11.21 commit (after this issue is filed): hydrabody/docs/runbooks/runbook.md → Sunshine config integrity (output_name preservation)
  • Issue #119 (kiosk overlay z-order on cold-boot stream) is independent — different symptom, different code path. Don't conflate them.
  • Issue #112 (parked) is also independent: hydrabody's restartSunshine doesn't set SW_HIDE on the spawn, leaving a Sunshine console window visible after every restart. Visible in screenshots taken after today's manual recovery.

Custom Fields

affected_repos
hydrabody
environment
production
severity
high

Comments (1)

claude (cederik) 29 Apr 2026 21:24

Fixed in hydrabody v1.11.27 (and cleaned up in v1.11.29). After several iterations through Sunshine`s /api/config (v1.11.22 fetched via GetConfig+merge — only returned subset; v1.11.24 read disk + posted full map — Sunshine drops unrecognised keys; v1.11.25/.26 wrote disk after API + polled for truncation race — Sunshine startup-write still won), the working approach is to skip /api/config entirely and write sunshine.conf directly. Sunshine reads dd_* keys per stream-session when dd_resolution_option=manual is set in the on-disk config (already present from writeSunshineConfig).

Verified on cosmic-pretzel-98 today (2026-04-29): firing /api/v1/stream/started keeps output_name, audio_sink, encoder, ensure_primary intact in sunshine.conf after the webhook. Race-tested: heal + Sunshine restart + immediate webhook also preserves full config. Sunshine restarts now correctly reload the full config, no black-screen regression.

Follow-up issue (defense-in-depth): tickVirtualDisplay should re-verify on-disk integrity each tick instead of trusting the marker file, so any future external truncator self-heals. Will file separately.

Closing.