HydraIssues

[MASTER] Cloud GPU streaming capacity in Belgium for mercator-talks flat streams
open unclassified Project: hydracluster Reporter: 11 Sep 2026 09:42

Description

What

Add extra streaming capacity for mercator-talks on flat heads by renting cloud GPU machines in or near Belgium and joining them to the fleet as bodies in a cloud district.

Why

Venue bodies are a fixed pool. #504 multi-stream slots raise capacity per body, but the number of physical GPUs does not change. Cloud GPU bodies give elastic capacity for peaks and for venues that have heads but no local body.

Requirements

  1. Cloud bodies run the standard body stack (Windows, Sunshine, hydrabody, WireGuard mesh to the HydraGuard hub) or a documented variant, in their own cloud district.
  2. Region and provider selection is driven by measured latency and jitter from the actual venues, not by datacenter location claims.
  3. Part of this issue: a repeatable procedure to test actual latency (RTT, jitter p50/p95/p99, loss) from each hydra venue to each candidate cloud region, before any commitment to a provider.
  4. The test procedure must produce the same metrics as #397/#402 (network-quality telemetry: jitter, loss, reordering, RTT, percentiles) so venue qualification and cloud region qualification use one method. Reuse or extend #42 (latency proximity checker, done) where it fits.
  5. mercator-talks concurrency constraints from #504/#511 (per-slot audio and mic device selection) apply to cloud bodies the same as to venue bodies. #513 hydramark gives the stream_capacity benchmark per machine type.

Deliverable

A reviewed plan on this issue (plan field), then implementation.

Related: #42, #397, #402, #504, #511, #513.

Repo

github.com/cederikdotcom/hydrahyperscaler (v0.1.0) is the program repo: PLAN.md (kept in sync with this plan field), runbooks (smoke-test-grid, venue-latency-campaign, cloud-body-enrol-retire), testbooks, and the campaign results under docs/testbooks/region-qualification/results/.

Sub-issues (14)

open #714 [MASTER] Execution: stream mercator-talks from a hyperscaler
closed #710 [SUPERSEDED by hyperscaler ruling on #695] Provider kill-questions: OVH Windows-on-GPU, Hetzner GEX44, LeaderGPU terms
open #709 Production cloud district bring-up: cloud-<region>-1, direct neck peering, per-venue go-live
open #708 Cloud pilot in bxl1-test-cloud: full body image, hydramark slots, #511 audio, interaction gate G3
open #707 Stack smoke test on a hyperscaler Windows GPU VM: Windows Server + MttVDD + Sunshine on GRID drivers (gate G2b, run FIRST)
open #706 Probe campaign: venue x region x path latency matrix and gate G2
open #705 Reflector fleet: one probe reflector VM per candidate cloud region
done #704 hydraprobe v0.1.0: fleet network-quality probe (RFC 3550 jitter, percentiles, loss, reordering, RTT)
open #513 hydramark: self-owned UE diagnostic probe and benchmark app for multi-stream bodies
open #511 Mercator: per-instance audio device selection via launch args (-MicDevice, -AudioOutDevice)
open #504 Multi-stream bodies: N Sunshine instances on N VDD displays, hydracluster as slot allocator (PoC: A5000 in bxl1-test-2, MacBook Air + iPad, Rupelmonde Castle Viewer)
open #402 Feature Request #397 (revised): Network-quality telemetry...
open #397 Feature Request: In-client network-quality measurement (j...
done #42 Add latency-based proximity checker for nearest district/body