Maximize chat capacity from GPUs that sit idle between VR sessions.
Vision: assign the hydrabrain role to every body. When a body is not VR streaming, its GPU serves chat streams; when a VR session starts (or is about to), inference on that body yields so the headset gets the full GPU.
Steps toward it:
- DONE 2026-09-07: chunky runs llama.cpp with -np 2 (two concurrent streams, 32K context each). llama.cpp is the standard runtime; the sm120 lobo appliance was the 5070 Ti special case.
- Per-GPU stream sizing: pick -np and total -c per node from VRAM (model ~12GB + ~2.5GB KV per 32K slot). Fleet has mixed GPUs; each node needs its own desired doc.
- VR-priority yielding: HydraCluster knows session state. When a body gets selected for a session (or hydrabody arms a stream), withdraw that node from hydrabrain discovery and stop or suspend the runtime; restore after the session ends. Measured on fluffy: game+inference contend hard (GPU 98%, tokens -21%, and the headset frame budget loses too).
- hydrabrainserver already fans out to all discovered providers; guests/users pick a node. Optional later: automatic provider selection instead of a manual picker.
- Watch system RAM: llama.cpp with mmap needs the model resident; bodies with 16GB RAM are tight next to a UE experience.
Constraint reminders: never co-serve chat during a live headset session on the same GPU (fluffy measurement); one HydraScale front door is enough, capacity scales on the provider side.