HydraIssues

Move HydraBrain inference off fluffy to chunky-turnip-23 (A5000, stock llama.cpp runtime)
done unclassified Project: Hydra Parent: #659 Reporter: 7 Sep 2026 14:55

Description

Fluffy is the Gallo-Romeins venue body; lobo contends with the museum experience for GPU (46 to 36 tok/s measured under load). Lobo desired state DISABLED on fluffy (VRAM freed, rollback = enabled:true PUT). Chunky has an RTX A5000 (24GB, sm86) which cannot run the pinned sm120 lobo appliance, so chunky gets an official llama.cpp CUDA runtime with the same pinned Qwen3.8-27B IQ3_S GGUF (URL+size+SHA256). Verify managed readiness, discovery, and end-to-end chat on hydrabrain.experiencenet.com.