HydraIssues

Master: make HydraBrain a first-class HydraCluster role and NimsForest inference provider
done feature Priority: high Project: hydrabrain Reporter: Codex 5 Sep 2026 09:45

Description

Outcome

HydraBrain becomes a first-class HydraCluster machine role. A node assigned the role can be instructed to serve a named local inference workload (initially Lobo on Fluffy), advertise a private authenticated OpenAI-compatible endpoint, report capabilities and health, and be consumed by NimsForest directly or through a Hollow harness.

HydraBrain is infrastructure, not a knowledge system: it does not own prompts, transcripts, memory, RAG, skills, tools, sessions, or organizational knowledge.

Current prototype

  • Node: fluffy-dumpling-87 (node-11da9ea3), Windows 11, RTX 5070 Ti 16 GB, HydraGuard 10.10.100.15.
  • Lobo is installed manually at C:\Lobo\app.
  • Qwen3.8-27B GSQ-RCO IQ3_S is serving with 32K context on loopback 127.0.0.1:18080.
  • Health, model discovery, and /v1/chat/completions have been proven.
  • The process currently runs through the temporary Windows task HydraBrain-Lobo-Test.
  • HydraCluster does not yet recognize hydrabrain; Fluffy's persisted roles remain hydrabody, hydravoice, and hydraguard-air.

Existing Hydra behavior to reuse

  • HydraBody: Windows download/install/version/start/stop/health/restart/reprovision patterns. Do not consume or overload HydraBody's single streaming-provider slot.
  • HydraSkin: expose loopback services only on the HydraGuard mesh address, reject port collisions, report the reachable endpoint, and verify health through the exposed path. Its Incus proxy implementation is Linux-specific and cannot directly expose a native Windows workload.
  • HydraReverseProxy/HydraScaleRouter: public-domain routing is not the default path for inference and currently lacks the authentication/privacy boundary required for prompts.

Required work

HydraBrain repository

Create a small Go service/CLI that runs beside HydraNode and owns inference workload lifecycle. Initial commands: serve, install, uninstall, status, version, and check-update. Read enrollment from the shared HydraNode config. Keep desired workload configuration separate from runtime-observed status.

The node-local gateway must bind only the HydraGuard mesh address, authenticate every inference request, proxy to a loopback-only runtime, preserve streaming responses, bound request sizes/timeouts, and never log prompt/completion bodies.

HydraCluster

  • Add hydrabrain to RoleCatalog and add a Windows provisioning recipe.
  • Add desired HydraBrain workload configuration: runtime, model alias/artifact, profile/context, internal port, and enabled state.
  • Add dedicated status reporting and read-only discovery for healthy HydraBrain providers.
  • Return scoped inference credentials without exposing the HydraCluster admin token.
  • Display role, workload, model, endpoint, health, version, GPU/VRAM, and last report.

NimsForest

  • Add a stateless hydrabrain AIService/provider with configurable base URL, model alias, and credential.
  • Do not use HydraCluster admin credentials.
  • Support health/model validation, streaming, cancellation, timeouts, and usage.
  • Handle llama.cpp/Lobo reasoning_content; a small completion budget may contain reasoning but no visible content.
  • Allow Hollow harness implementations to receive the same endpoint/model/credential during spawn or session binding where the harness supports an OpenAI-compatible provider. Hollow retains session and agent lifecycle ownership.

Integration contract (v1)

Data plane: OpenAI-compatible /v1/chat/completions, /v1/models, and /health over the private HydraGuard mesh.

Advertised model IDs are stable aliases such as lobo-qwen3.8-27b, never absolute Windows file paths. Discovery returns node ID/name, endpoint, protocol, runtime, models/context, health, GPU capacity, and last-seen timestamp.

Security constraints

  • No public route by default.
  • No bind to 0.0.0.0; mesh address only.
  • Bearer authentication and scoped credential rotation.
  • No prompt/completion/body logging.
  • Model artifacts require pinned origin, size, and SHA-256 verification.
  • ZIP extraction must prevent path traversal.
  • Runtime executes with a dedicated identity and least privilege.
  • Explicit request/body/concurrency/time limits.

Acceptance criteria

  1. Assigning hydrabrain to a clean compatible Windows HydraCluster node installs and starts the agent idempotently.
  2. Assigning Lobo plus its model alias provisions verified artifacts and reaches healthy state without manual shell commands.
  3. HydraCluster discovery reports only healthy, recently seen providers and never returns an admin credential.
  4. From an authorized NimsForest deployment on the HydraGuard mesh, direct AIService inference returns visible content and usage.
  5. A supported Hollow harness can use the same HydraBrain provider while Hollow retains session affinity.
  6. Removing/disabling the workload stops exposure and the runtime cleanly.
  7. Reboot and upgrade recovery are automatic and do not block HydraCluster exec.
  8. Tests cover auth failure, mesh-only binding, artifact checksum failure, port collision, unhealthy runtime, streaming cancellation, reasoning-only responses, restart, and removal.
  9. Runbooks document architecture, provisioning, NimsForest handoff, security, operations, and the Fluffy acceptance test.

Out of scope

Central compute leasing/scheduling, knowledge storage, conversation persistence, embeddings/RAG, agent tools/skills, and public anonymous inference.

Sub-issues (5)

open #670 HydraBrain fleet-wide: hydrabrain role on every body, chat streams yield to VR
done #669 Move HydraBrain inference off fluffy to chunky-turnip-23 (A5000, stock llama.cpp runtime)
done #668 HydraBrainServer: mobile layout broken, trial counter too prominent
done #667 HydraBrainServer: guest trial chat, 3 free messages before IAMNim signup
done #664 HydraBrainServer: IAMNim-authenticated HydraScale chat

Custom Fields

prototype_node
node-11da9ea3
related_issue
656
scope
master

Comments (8)

Codex 5 Sep 2026 09:49

Initial scaffold created and pushed: https://github.com/cederikdotcom/hydrabrain at commit 8c34b87. It includes a buildable Go CLI, authenticated allowlisted reverse proxy, strict non-wildcard listener and loopback-upstream validation, upstream credential isolation, request size/time limits, health and capability endpoints, stable model alias configuration, CI, tests, architecture/runbook docs, and a NimsForest integration handoff. go test ./..., go vet ./..., and Windows amd64 cross-build pass. Fluffy still uses the temporary HydraBrain-Lobo-Test task; the scaffold has not been deployed there.

codex 5 Sep 2026 10:45

Implementation slice is complete and published.

Commits / branches:

  • HydraBrain main: 723ad13 — managed desired-state agent, verified artifact installer, process-tree supervision, authenticated mesh gateway, status CLI, CI + release workflow, architecture/runbook/NimsForest handoff.
  • HydraCluster feat/hydrabrain-role: 90d9624 — role catalog, desired config, node-scoped config/status, fresh-ready discovery, dedicated gateway/discovery credentials, token rotation, readiness invalidation, tests and testbook.
  • HydraNode feat/hydrabrain-role: 5c37691 — Windows scheduled-task and Linux systemd role-agent provisioning using hydrabrain run.
  • NimsForest feat/hydrabrain-provider: d62a598 — structured AIService provider, HydraCluster discovery, direct mode, capabilities/model verification, visible-content streaming, reasoning-only diagnosis, and fail-closed cloud behavior.

Published branches:

Verification:

  • Full Go tests and vet pass in all four repositories.
  • HydraBrain Windows amd64 cross-build passes.
  • HydraBrain has a local end-to-end controller/runtime/gateway test.
  • HydraCluster has an HTTP lifecycle test through config → node desired state → ready → discovery → token rotation/invalidation → rediscovery.
  • Reachable-code govulncheck is clear after pinning Go 1.25.13, updating HydraCluster x/net/x/text, and NimsForest NATS to 2.12.6.

Rollout remains intentionally pending: merge/release the three integration branches, configure server.brain_discovery_token, publish HydraBrain, then perform the documented Fluffy cutover while retaining HydraBrain-Lobo-Test as rollback.

codex 5 Sep 2026 10:46

Review PRs opened:

codex 5 Sep 2026 11:01

Architecture correction (2026-09-05):\n\n- Closed HydraNode PR #2 as superseded. A new role does not require HydraNode-specific dispatch, templates, or uninstall catalog entries.\n- HydraBrain main 677fafd now owns install/uninstall and its Windows scheduled task / Linux systemd unit.\n- HydraCluster 41d4791 adds checksum-verifying Windows and Linux recipes that invoke the HydraBrain installer through HydraNode's existing generic recipe executor.\n- HydraCluster also clears removed roles' provisioning state generically, so a later re-assignment offers the recipe again.\n- HydraCluster PR #5 was updated: https://github.com/cederikdotcom/hydracluster/pull/5\n- HydraNode PR #2 is closed: https://github.com/cederikdotcom/hydranode/pull/2\n\nValidation: HydraBrain and HydraCluster full tests and vet pass; HydraBrain Windows amd64 and Linux arm64 cross-builds pass. Fluffy remains on the rollback/manual task until the tagged HydraBrain release and HydraCluster PR are deployed.

codex 5 Sep 2026 11:32

Fluffy live cutover completed (2026-09-05).\n\nLive state:\n- HydraCluster v2.0.110 is running in production. Default recipes are now embedded in the binary; external recipe files are optional overrides.\n- HydraBrain v0.1.0 is published on HydraRelease with verified Windows/Linux assets.\n- fluffy-dumpling-87 (node-11da9ea3) now has the hydrabrain role.\n- Managed task HydraBrain is running; rollback task HydraBrain-Lobo-Test is stopped/Ready and was not deleted.\n- Lobo is owned by HydraBrain and listens only on 127.0.0.1:18080.\n- HydraBrain listens only on 10.10.100.15:18081.\n- Discovery reports ready, stable alias lobo-qwen3.8-27b, 32K context, RTX 5070 Ti, and fresh status.\n\nAcceptance exercised live:\n- anonymous mesh request returned 401;\n- authenticated local completion returned "fluffy managed and healthy";\n- authenticated request from a second HydraGuard node returned "mesh healthy";\n- the temporary credential used in the cross-node exec audit was immediately rotated;\n- discovery disappeared during rotation and returned only after fresh ready status;\n- killing the managed llama-server child caused automatic restart from PID 624 to PID 13896;\n- post-restart completion returned "restart healthy";\n- test prompt/response text was absent from the runtime log.\n\nLive fixes discovered and released:\n- d178378 / v2.0.110 decodes Windows PowerShell byte[] checksum responses.\n- f495d96 embeds default recipes so production deployment needs no manual recipe copy.\n- The new private HydraBrain repo lacked its HydraRelease Actions secret. A restricted one-time bootstrap job published v0.1.0 using the existing authorized pipeline and was removed immediately afterward.\n\nMaster issue remains implementing: this cutover adopts the already reviewed C:\Lobo installation with artifacts: []; clean-node pinned Lobo runtime/model provisioning and the NimsForest provider rollout are still pending.

Codex Hydra infrastructure 6 Sep 2026 07:49

Hydra infrastructure acceptance is complete and deployed (no PR cutover).

Deployed releases:

  • HydraCluster v2.0.112 (2052f88; live-acceptance docs 705d0e2): first-class HydraBrain node panel and role-only reprovision API.
  • HydraBrain v0.4.1 (79d0be6, 3a5f312, 4b91c50; acceptance docs 85f8d06): bounded admission, health withdrawal/recovery, least-privileged Windows runtime, resumable pinned artifacts, and deterministic state.
  • HydraNode v1.10.40 (efed7a0; acceptance docs 0107900): first-class managed-role dispatch plus Linux service/removal support. Fluffy was updated through HydraCluster.

Live Fluffy acceptance:

  • Clean managed-root provisioning verified the checksummed runtime and 11.77 GB model; an interrupted transfer resumed, and a full Windows reboot restored mesh, node agent, brain, discovery, and inference automatically.
  • Role removal/reassignment and the new role-only reprovision path withdraw discovery fail-closed, then progress stopped -> starting -> ready without restarting HydraBody/HydraVoice. The HydraBrain/Lobo PIDs also remained unchanged across the HydraNode update.
  • HydraBrain and Lobo run as limited LocalService. Lobo listens only on 127.0.0.1:18080; the bearer-protected gateway listens only on 10.10.100.15:18081. Anonymous access returns 401.
  • Direct non-streaming, SSE [DONE], and Club's running NimsForest-container inference passed over hydraguard-air with visible content and real usage. Acceptance prompt markers were absent from logs.
  • The admin UI shows model/protocol/private endpoint, verified artifact count, health, agent version, PID, GPU/VRAM/capacity, and last report while omitting credentials, command arguments, environment values, and artifact URLs.
  • The discovery credential was rotated during final security handling; the old value returns 401. HydraCluster, the private mode-0600 handoff, and Club's private mode-0600 consumer config were updated together, and rotated-credential Club inference passed.

Consumer contract remains stable: GET /api/v1/hydrabrains, explicit Fluffy alias, OpenAI-compatible /v1/chat/completions, independent discovery/provider credentials, refresh discovery each turn. Capacity is additive. No pending Hydra interface change is required by #264. Private integration details remain in /tmp/tnbc264-hydra-handoff.md; no credentials are included here.

Codex Hydra infrastructure 6 Sep 2026 08:25

Correction to the prior infrastructure completion note: HydraNode v1.10.40 has been withdrawn. It reintroduced legacy role-specific HydraBrain provisioning alongside the already-correct HydraCluster recipe path, including an incompatible Windows SYSTEM scheduled-task path. The GitHub release is explicitly marked withdrawn and retained only for auditability.

Corrective HydraNode v1.10.41 (a647623, baa7fb2) reverts the v1.10.40 feature and follow-up documentation commits. The v1.10.41 and v1.10.39 source trees are identical. Local tests/vet/build and release run 34021476206 passed; the checksummed release is live as the production latest.

Fleet correction and heartbeat acceptance:

  • Pre-correction baseline: 34 records = 22 online, 7 pre-existing offline, 5 pending enrollment. Eighteen online nodes had reported v1.10.40.
  • All 18 moved forward to v1.10.41 through normal auto-update or HydraCluster's targeted node-update operation. Zero nodes remain on v1.10.40.
  • After multiple two-minute offline-check windows, the same 22 baseline nodes remain online and heartbeating; zero baseline-online nodes were lost. The 7 offline and 5 pending records are unchanged from before the correction and are not claimed as live.
  • One active streaming session remained present through the final fleet check.
  • Fluffy is online on HydraNode v1.10.41. Its HydraBrain v0.4.1 agent and Lobo PIDs did not change across correction, provider discovery remains ready, Club-container inference returned visible content with usage, and the acceptance prompt marker was absent from logs.

HydraBrain remains exclusively recipe-managed. No consumer/discovery contract changed, and #659 remains done.

Codex 6 Sep 2026 09:17

Follow-on #664 is complete and linked here: HydraBrainServer v0.1.1 is live at https://hydrabrain.experiencenet.com as an IAMNim-authenticated HydraScale on pi-node-001. It consumes the existing HydraCluster discovery and HydraBrain OpenAI-chat/capabilities contracts without changing them. Live in-scale acceptance finds Fluffy/Lobo and current machine capacity; no credential or mesh endpoint is exposed to the browser. Repo/runbooks: https://github.com/cederikdotcom/hydrabrainserver