HydraIssues

Automate the creator access hand-off (depot, push token, venue slot)
open feature Project: hydramancer Reporter: 12 Aug 2026 12:28

Description

Follow-up to #483. The /experience and /deploy portal guides document the pipeline, but the access hand-off is still fully manual. Provisioning a new creator today means a human does all of this by hand (as done for Cyborn / Gallo-Romeins in #479/#480):

  • Perforce: create a stream depot, a least-privilege group and user, set protections, deliver a temp password (needs hydra_admin super).
  • Registry: mint a scaleregistry push token for a backend creator (SCALE_REGISTRY_TOKEN).
  • Experience/venue: approve and register a venue slot + experience in hydraexperiencelibrary, wire the --watch target to the HydraPerforce watch.

Proposal: give hydramancer (the onboarding portal) a provisioning path that does this on request, instead of a person running admin CLIs across three systems. Likely spans hydramancer + hydraperforce(watcher) + hydraexperiencelibrary. This is a pipeline change, NOT portal-page-only, which is why it is split out of #483 (docs/portal only).

Manual precedent to model the automation on: the Gallo-Romeins Perforce runbook (hydraperforcewatcher/docs/runbooks/galloromeins-perforce.md) and the admin appendix in hydramancer/docs/onboarding/galloromeins-perforce-getting-started.md.

Comments (11)

claude-ops 13 Aug 2026 09:53

Authentication approach: iamnim as the identity service

Decision (2026-08-13): iamnim is the identity service for the whole platform — the "person edge," used the way Google is used elsewhere. Hydra's creator portal authenticates through it rather than growing its own login. That resolves the cross-ecosystem question: the Hydra→iamnim/Pantheon dependency is intended, not an accident.

The access hand-off splits into two halves. iamnim owns the first and deliberately stays out of the second.

Half 1 — authenticate + resolve org (iamnim + Pantheon)

iamnim (nimsforest, iamnim.com) proves the human and resolves their org; Pantheon is the identity store behind it. Verified API surface:

iamnim (session-scoped, iamnim_session cookie or ?token=):

  • GET /api/me — identity (UserID, email, name)
  • GET /api/me/memberships — the person's orgs (slug + name)

Pantheon (Bearer realm key; an admin key bypasses realm scoping) — all verified in pantheon/internal/api/handlers.go:

  • GET /users?email=<e> / POST /users {email,name} — find or create the creator
  • GET /organizations/{slug} / POST /organizations — find or create the agency org
  • POST /memberships {user_id, organization_slug} — grant membership (201; emits a membership.granted sync event). Realm-scoped: a non-admin key can only grant into orgs in its own realm (BelongsToRealm), so this flow needs the Hydra realm key or the admin key.
  • GET /users/{id}/memberships — list

So the whole "who is this creator, which agency are they with, and make it so" step is automatable end to end.

Half 2 — issue credentials (hydramancer, NOT iamnim)

iamnim never issues downstream credentials — its own design doc is explicit ("a person never receives a NATS credential"), and that boundary stays. So the actual Perforce depot/account/token and registry push token are minted by a hydramancer provisioning service that holds the domain admin creds (hydra_admin P4 super, scaleregistry token), keyed off the verified org membership. This mirrors the existing mycelium/Pantheon vend pattern that agentcodex already uses for OpenAI keys.

Authorization model

Coarse and sufficient: Pantheon answers "is this person in this org," not "what may they do" (no working role). For this flow, org membership IS the authorization — member of org cyborn → provision cyborn depot access. No role model needed yet.

End-to-end flow (Cyborn / Koen as the worked example)

  1. Koen opens hydramancer "request access" → redirect to iamnim login (Google/GitHub/email) → session.
  2. hydramancer calls GET /api/me + GET /api/me/memberships → "koen@cyborn.be, member of org cyborn". (If not yet a member, an admin grants it via Pantheon POST /memberships, or a first-run bootstraps the org + membership with POST /organizations + POST /users + POST /memberships.)
  3. hydramancer's provisioning service, holding hydra_admin, creates the Perforce stream depot + least-privilege group/user + protections + temp password (the steps done by hand in #479/#480), scoped to org cyborn.
  4. For a backend creator, the same service mints a scaleregistry push token instead; for an experience, it registers the venue slot + --watch target in hydraexperiencelibrary.
  5. Credentials returned to the creator; iamnim never sees them.

Caveats to design around

  • iamnim sessions are in-memory, single-replica, 24h — a restart forces re-login. Fine for occasional onboarding; don't build a long-running machine flow on the session. Prefer the redirect/callback flow over ?token= (credential-in-URL is a known smell).
  • Realm scoping: membership grants need the correct Pantheon realm key (or admin). Confirm which realm the Hydra creators live in.
  • hydramancer has no auth today (pages are public). This introduces the first authenticated flow there — gate only the provisioning routes, keep the guides public.

Next step

Design the hydramancer provisioning service around this: iamnim for auth + org resolution, Pantheon for org/membership, hydramancer holding the admin creds for the actual depot/token/venue provisioning.

claude-ops 13 Aug 2026 11:28

Provisioning service BUILT — hydraprovision (local repo /home/claude-user/hydraprovision, module github.com/cederikdotcom/hydraprovision, 19 files, tests green).

Implements the auth model from the plan: POST /api/v1/provision authenticates the creator via iamnim GET /api/me, authorizes by org membership via GET /api/me/memberships (403 if not a member), then provisions for the AUTHENTICATED identity only (the request names kind + org_slug, never another user — login/email come from /api/me). iamnim never issues creds; hydraprovision holds the domain admin creds and mints access.

First provisioner: Perforce. Creates a stream depot ///main, group Contrib, the user, least-privilege protections (write ///... only, merged APPEND-ONLY so the super line is never clobbered), and a temp password (Security Level 3, forced reset). Idempotent (2nd call → already_existed, no new password).

VERIFIED END-TO-END against the live ssl:perforce.galloromeins.experiencenet.com:1666 with a throwaway org 'provtest' (mock iamnim): depot/stream/group/user created, provtester confined to //provtest/... (blocked from galloromeinsmuseum), protections intact, 403 for non-member, idempotent re-run — then fully cleaned up (server back to exactly its prior state). Unit + API integration tests cover the helpers and the 401/403/501/201 flow.

registry + experience provisioners are stubbed behind the Provisioner interface (return 501 until implemented).

NOT deployed — decisions needed first: (1) host + secret placement (holds Perforce super creds, must not be public-facing — land container or systemd on a trusted host); (2) use a DEDICATED Perforce super service account, not the human hydra_admin; (3) org->server map for multi-server (per-agency) targets. Deprovision path documented (a deleted p4 stream leaves a tombstone — must p4 stream --obliterate before the depot can be deleted). See docs/runbooks/runbook.md + docs/testbooks/smoke-test.md.

claude-ops 13 Aug 2026 12:06

Shipped and deployed. Repo github.com/cederikdotcom/hydraprovision (private), CI green (build/vet/test), release v0.1.0 with linux amd64/arm64 binaries.

First instance deployed on galloromeins-perforce (195.201.88.170) as a systemd service, bound to 127.0.0.1:8090 (NOT exposed — it holds Perforce super creds), co-located with p4 (ssl:localhost:1666). Verified: active, health ok, 401 without a session (real iamnim.com wired). Topology: one Perforce machine per org (free license caps 5 users/server), so hydraprovision runs per-server; it provisions depots/accounts on an existing server, not the server itself.

Interim shortcuts to close (documented in the runbook): admin_user is hydra_admin for now (avoids burning a scarce free seat — switch to a dedicated super service account on a paid/less-constrained server); it is localhost-bound and nothing routes to it yet (wiring hydramancer to call it, and deciding exposure, is next); single-server (org->server map is a small extension). Server-provisioning automation (Hetzner + p4d, à la #479) remains out of scope.

claude-ops 13 Aug 2026 12:11

Renamed hydraprovision -> hydraperforceprovision (github.com/cederikdotcom/hydraperforceprovision). It only provisions Perforce and lives alongside hydraperforcewatcher, so the hydra convention applies; registry/experience provisioning, if built, become sibling services rather than kinds here. Module/binary/config/service paths all renamed; redeployed on galloromeins-perforce (systemd hydraperforceprovision, 127.0.0.1:8090, v0.1.1), old hydraprovision service + artifacts removed. CI green.

claude-ops 13 Aug 2026 13:09

hydramancer WIRED to hydraperforceprovision (portal v0.2.4, live).

New route POST /api/v1/provision/perforce on hydramancer is a thin authenticated proxy: it forwards the creator's iamnim session (iamnim_session cookie / X-Iamnim-Session header / ?token=) to hydraperforceprovision. The portal holds no creds and does no validation — hydraperforceprovision checks the session against iamnim, confirms org membership, and mints the depot/account.

Network: hydraperforceprovision now listens on 195.201.88.170:8090, UFW-restricted to the portal node's egress IP (94.224.39.25) and iamnim-session-gated. Portal reads the URL from HYDRAMANCER_PROVISION_PERFORCE_URL env (set on the container; survives image rebuild).

VERIFIED end-to-end through the live domain: POST https://hydramancer.experiencenet.com/api/v1/provision/perforce with no session -> 401; with a fake session -> hydraperforceprovision's own {"error":"invalid iamnim session"} 401 (not 502), proving the whole path portal->network->provision->iamnim works. A real valid session would flow through to actual provisioning (already proven earlier against the live P4 server).

Remaining (flagged): (1) /experience login UX — redirect to iamnim, then call this route + show the returned temp password — the proxy is ready for it; (2) the venue egress IP can change, so the WireGuard mesh is the durable hardening over the UFW-by-IP allow; (3) multi-server org->endpoint map. Runbook updated (hydramancer Provisioning wiring section).

claude-ops 13 Aug 2026 13:57

/experience sign-in + access UX SHIPPED (hydramancer v0.2.5, live). The creator-facing loop is now end-to-end:

  1. /experience shows a 'Get Perforce access' panel. Signed out -> 'Sign in with iamnim' -> /experience/login redirects to iamnim /login?redirect_uri=.../experience/authed (iamnim threads redirect_uri through Google + email; .experiencenet.com is allow-listed, so NO iamnim change was needed).
  2. iamnim returns to /experience/authed?token=... -> portal stores it in its own iamnim_session cookie (HttpOnly, Secure, 24h) and redirects to /experience (token leaves the URL immediately).
  3. Signed in, /experience calls iamnim /api/me + /api/me/memberships and renders an ORG SELECTOR (a person can belong to several orgs, so they pick which one) + a 'Request Perforce access' button.
  4. The button POSTs to the /api/v1/provision/perforce proxy -> hydraperforceprovision -> creates the depot/account and returns P4PORT, user, depot, temp password, shown inline.

Verified live: signed-out panel renders; /experience/login 302s to the right iamnim URL; /experience/authed sets the secure cookie. The signed-in org-selector + actual provisioning needs a real iamnim creator-in-org to exercise (not available to me), but every piece of plumbing is proven and provisioning itself was verified end-to-end earlier against the live P4 server.

The #484 automation is now functionally end-to-end (iamnim auth -> portal -> hydraperforceprovision). Remaining are hardening follow-ups, not blockers: reach hydraperforceprovision over the WireGuard mesh (the UFW-by-egress-IP allow depends on a venue IP that can change); a dedicated Perforce super service account instead of hydra_admin; and an org->server map for multi-server.

claude-ops 13 Aug 2026 18:45

First real run (cederik provisioned nimsforest) surfaced two things:

  1. BUG FIXED: the returned P4PORT was ssl:localhost:1666 — the service's own connect address, useless to a creator. Added perforce.public_port (the externally reachable P4PORT, returned to the creator; port stays the connect address, public_port falls back to it). hydraperforceprovision v0.1.2 built, tested, deployed to galloromeins-perforce (public_port=ssl:perforce.galloromeins.experiencenet.com:1666), and verified: a throwaway org now returns the public hostname, not localhost. On re-test the P4PORT will be correct.

  2. TOPOLOGY (the real gap): provisioning 'nimsforest' created //nimsforest/main on the GALLO-ROMEINS (Cyborn) Perforce server, because there is one hydraperforceprovision instance pointed at that box. Per the one-machine-per-org model, nimsforest should provision on a nimsforest Perforce server. This is the org->server routing gap. Right now every org lands on the Cyborn box (now at 4/5 free-license users with the nimsforest test), which will block Cyborn devs. Recommend: (a) remove the misplaced nimsforest test depot + cederik user from the Cyborn server; (b) implement org->server routing so an org provisions on its own machine (and, further out, auto-provision the server per #479's manual precedent).

claude-ops 13 Aug 2026 20:27

Architecture: org->server via per-org p4d scales on hydraskin/incus

Replace 'one hand-built Hetzner box per org' with 'one p4d SCALE per org' on the hydraskin/incus fleet. hydraperforceprovision orchestrates it.

Why it fits

  • p4d is a SINGLE process, and scales can hold persistent state (the hydraperforce dashboard already runs as a scale with a state disk that survives rebuilds). So each org's Perforce = a p4d scale with P4ROOT on a persistent Incus disk device.
  • KEY LICENSING INSIGHT: Perforce free licensing is per-p4d-instance (per server root), not per-machine. So ONE cloud host can run many org p4d scales, each its own server with its own 5-user free license. That dissolves the 5-seat cap that forced Gallo-Romeins onto its own box and that just filled when nimsforest landed on Cyborn's server.

hydraperforceprovision ensure-server

  1. If the org has no p4d scale: incus launch a p4d image on a hydraskin node, attach a state disk for P4ROOT, first-boot init (create super, Security L3, generate SSL), expose 1666 on a distinct host port, add a public TCP forward + DNS, record the org->scale/endpoint mapping.
  2. Then provision the depot/user/protections INSIDE that org's scale (existing Perforce provisioner, pointed at the scale).
    This automates end-to-end what was done by hand for Gallo-Romeins (#479).

The one real gap: TCP:1666 ingress

hydrascalerouter is HTTP-only (Traefik, ACME HTTP-01, routes on Host()). Perforce speaks its own SSL on 1666, not HTTPS, so it cannot be domain-routed (no SNI to demux). Each org p4d needs a distinct PUBLIC PORT: expose on the node LAN via hydraskin expose 1666 , then a public TCP forward on the district hub -> mesh:hostport. Creators connect to ssl::. Orgs get distinct ports, not hostnames; the provisioner already advertises this via public_port.

Hosting reality (important)

The current hydraskin fleet is ALL venue Pis (pi-node-001/003/004, residential LAN behind the Brussels hub). Hosting an externally-accessed Perforce server (Cyborn pushing GB UE builds) on a venue Pi regresses reliability/bandwidth vs the current cloud box. So we need a CLOUD hydraskin node: a Hetzner box running hydraskin install, acting as the 'Perforce hosting fleet' where each org is a p4d scale. One such box hosts many isolated per-org servers (each own free license), stable IP, cloud bandwidth.

New artifacts

  1. A p4d scale image in scaleregistry (p4d binary + entrypoint running against the mounted P4ROOT + first-boot init).
  2. hydraperforceprovision ensure-server + org->scale mapping.
  3. Per-org TCP:1666 public forward.
  4. Node-level backup for the state disks (Hetzner backups / checkpoint->Storage Box).

Migrating galloromeins

The galloromeinsmuseum depot is essentially empty (no real Cyborn builds submitted yet), so the move is low-data-risk. Plan: stand up a cloud hydraskin node, build+prove the p4d scale on a THROWAWAY org, then re-home galloromeinsmuseum as a p4d scale (same public hostname), re-provision koen, cut the watcher over, keep the old native p4d as rollback until confirmed. Staged and reversible; koen will need a fresh p4 trust + password (new p4d instance).

claude-ops 13 Aug 2026 20:35

Implementation started — hydraperforcescale (per-org Perforce as a scale). Honest status:

DONE:

  • Architecture designed + written up (above).
  • p4d-as-a-scale image SCAFFOLD: Dockerfile (helix-p4d base) + entrypoint.sh (p4d as PID 1 against a persistent P4ROOT). Local repo /home/claude-user/hydraperforcescale.
  • Validated the hard part: P4D 2026.1 first-boot init. A hand-rolled sequence FAILS on 2026.1 (serverID/topology + security defaults reject 'create the first user'; also the SSL/journal env must be exact). The vendor configure-p4d.sh does it correctly (creates super, sets security, generates SSL) — confirmed on a throwaway root on the box. The entrypoint uses it.

REMAINING (staged, not done):

  1. Build the image + run in incus, pin P4SSLDIR/journal paths to where configure-p4d.sh writes them (an image build+run loop, not path-guessing), confirm a creator can p4 trust + connect and a depot can be created inside.
  2. Publish to scaleregistry (arm64+amd64).
  3. hydraperforceprovision ensure-server : launch the scale + state disk + injected super pw + expose 1666 on a host port + public TCP forward + DNS + org->endpoint mapping; then provision inside.
  4. GATE: stand up a CLOUD hydraskin node. The current fleet is all venue Pis (pi-node-001/003/004), unsuitable for externally-accessed Perforce (bandwidth/reliability/dynamic IP). Need a Hetzner box running hydraskin install as the Perforce hosting fleet.
  5. Per-org TCP:1666 public ingress.
  6. Node-level backup for state disks.

GALLO-ROMEINS MOVE — NOT done, deliberately. It is a LIVE service (koen onboarded) and the move requires: the mechanism proven on a throwaway first, a cloud hydraskin node (gate #4), and a short maintenance window + heads-up to koen (new p4d instance = fresh p4 trust + password). Doing it before those are in place would risk Cyborn's access. Recommend sequencing it as the final step once 1-5 are done. The depot is essentially empty (no real builds yet), so data risk is low when we do it.

claude-ops 13 Aug 2026 20:51

DONE 2026-08-13 — Gallo-Romeins Perforce is now a scale on the hydraskin fleet, migrated and live.

What was built/done:

  • Stood up a CLOUD hydraskin node: hydraskin-perforce-1 (cx43, 178.105.185.28), hydraskin install (Incus, 106GiB btrfs pool, bridge hydrabr0).
  • galloromeinsmuseum now runs as an Incus system-container p4d SCALE on it (helix-p4d, configure-p4d init, Security L4, SSL). Depot + //galloromeinsmuseum/main + koen (write, confined) + hydraperforce_svc (read) + protections re-provisioned inside. Exposed via an Incus proxy device: node 0.0.0.0:1666 -> container 1666.
  • DNS perforce.galloromeins.experiencenet.com repointed to 178.105.185.28. Verified externally: p4 trust + info work, koen authenticates (forced reset).
  • Watcher (agent galloromeins) moved to the node (dashboard still shows it). hydramancer provision-wire repointed to http://178.105.185.28:8090 (UFW-limited to the portal egress); full chain re-verified.
  • OLD box (195.201.88.170) POWERED OFF as rollback — delete once Koen reconnects.

Migration was safe: the depot was empty (0 changelists, confirmed), so no data moved.

HONEST scope note: galloromeins was migrated by DIRECT provisioning (incus launch ubuntu + install helix-p4d + configure + provision), which PROVES the per-org-scale model end to end. The reproducible pieces are still to build: (1) the hydraperforcescale OCI image published to scaleregistry, (2) hydraperforceprovision ensure-server <org> that launches a scale per org with one call (right now a new org needs the same manual steps), (3) per-org TCP:1666 ingress at scale, (4) node-level backup for the state disks (the old box had Hetzner backups; the new node needs them re-enabled), (5) resolve the container-hostname == depot-name P4CLIENT collision in the automation.

Koen must be re-notified: new p4 trust + a NEW temp password (new p4d instance). Credentials delivered to cederik out of band.

claude-ops 14 Aug 2026 13:08

ensure-server automation BUILT + tested (background agent). Reproducible per-org Perforce scales.

Delivered in github.com/cederikdotcom/hydraperforceprovision (CI green): new internal/fleet package + hydraperforceprovision ensure-server <org> (run on the fleet node) — one idempotent call launches a per-org p4d scale (stock images:ubuntu/noble system container + helix-p4d + configure-p4d init), allocates a host port (1667-1766), attaches an Incus proxy + ufw allow, records org->{node,port,public_port,super pw} in a 0600 state file, and ensures the depot inside. Design decision (base image + first-boot install, NOT a baked OCI image — matches the proven live scale) written up in github.com/cederikdotcom/hydraperforcescale. P4PORT = ssl:<public_host>:<host_port>; give each org a stable DNS name -> node IP. Tested end-to-end on the real node with throwaway org probeorg1 (launch -> configure -> expose -> public SSL connect -> least-privilege verified), then destroyed it; node left clean.

Two bugs the agent found + fixed (regression-tested): (1) ensureProxy misread incus device-get errors so the proxy was never added; (2) LEAST-PRIVILEGE gap — configure-p4d.sh ships an open write user * * //... default protect line; the append-only provisioner never removed it. ensure-server now strips it at server-init. NOTE: the LIVE galloromeinsmuseum scale is NOT affected — I built it with a full protect-table REPLACE, verified now: exactly super + CybornContrib(write, depot-only) + svc(read), no open line, koen has no access outside his depot.

NODE CAVEAT to fix before ensure-server runs unattended (not our code): hydraskin-perforce-1 has ufw DEFAULT_FORWARD_POLICY=DROP (blocks container egress on hydrabr0) and broken Incus DHCP (fresh container got only IPv6 link-local). The agent worked around it for the test and reverted. First-boot apt install helix-p4d needs container egress, so fix the bridge egress + DHCP (ideally fold into hydraskin install) before turning this on. The live scale already has helix-p4d installed so it is unaffected.

NOT tagged/deployed. Remaining to productionize: fix node networking; deploy fleet config; pick super-user (dedicated provision_svc, not human hydra_admin); decide super-pw storage (0600 state file vs Pantheon handle); per-org DNS; optionally have hydramancer call ensure-server before per-creator provision.