HydraIssues

Self-service experience management for external creators (Cyborn / Gallo-Romeins) via hydramancer, leveraging the hydraexperiencelibrary admin interface
closed unclassified Project: hydramancer Reporter: 15 Aug 2026 11:51

Description

Goal

Give external creators (first: Cyborn / Gallo-Romeins Museum) a self-service view and scoped lifecycle controls over their experiences, delivered through the hydramancer creator portal, by leveraging the admin interface and API that already exist in hydraexperiencelibrary. Today only an operator holding the library admin token can see an experience's state or drive its lifecycle (stage / promote / rollback / pause). This closes that last gap in the pipeline.

Assessment + implementation plan only. No code, deploy, or live-experience change is proposed here.


Backend findings — hydraexperiencelibrary (the interface to leverage)

Auth model: single admin bearer token, NO org/tenant scoping (confirmed)

  • internal/api/middleware.go:10-24 — requireAuth compares the Authorization: Bearer <token> header against one shared s.adminToken. All-or-nothing; no per-org, per-developer, or per-owner filtering anywhere.
  • internal/api/handlers_web.go:244-254 (initWeb) + server.go:99-109 — the admin web UI is gated by hwc.RequireWebAuth (cookie session, same shared admin token via hwc.Config.AdminToken). Also global; no tenant scoping.
  • handleListExperiences (handlers.go:17-24) returns every experience via st.ListExperiences() with zero filtering. There is no code path anywhere that scopes reads or writes by org. The org boundary does not exist in the library and must be enforced upstream (or added).

How an experience ties to an org

  • internal/store/store.go:17-43 — Experience carries Owner and Developer string fields (plus Name, Label, Status, WatchName, build pointers, Districts, etc.). For the target experience: Developer="cyborn", Owner="gallo-romeins-museum".
  • iamnim memberships expose OrganizationSlug / OrganizationName (hydramancer/internal/iamnim). Scoping key = match a member's OrganizationSlug against the experience's Developer and/or Owner. These are free-text today, so a slug-to-field mapping must be made authoritative (see Risks).

Build state exposure (dev / staging / prod)

  • Builds live in a separate builds.yaml store (internal/store/build.go), keyed by experience name + monotonic Number. Experience holds only integer pointers: DevelopmentBuild, StagingBuild, ProductionBuild, PreviousBuild (store.go:25-28).
  • GET /api/v1/experiences/{name} (handlers.go:26-39) returns the experience incl. these pointers.
  • GET /api/v1/experiences/{name}/builds (handlers.go:453-474) lists build records (number, scm_type, scm_ref, build_url, created_at) for one experience.
  • GET /api/v1/pipeline/health?experience=<name> (handlers_pipeline.go:69-158, checkLocalExperiences 264-326) returns a rich per-experience view: status, watch_name, all four build pointers, districts, and the last 10 builds — a strong candidate to back the creator's read view with a single call.

Lifecycle routes — JSON API (all requireAuth, server.go:124-155)

Route Handler Request Response / effect
GET /api/v1/experiences handleListExperiences — all experiences (unscoped)
GET /api/v1/experiences/{name} handleGetExperience — one experience
POST /api/v1/experiences handleCreateExperience name,label,owner,developer,watch_name,districts... creates (operator-only)
PUT /api/v1/experiences/{name} handleUpdateExperience partial patch incl. developer,owner,districts,auto_promote mutates ownership/placement (operator-only)
DELETE /api/v1/experiences/{name} handleDeleteExperience — deletes (operator-only)
POST /api/v1/experiences/{name}/promote handlePromote (handlers.go:197-274) {"to":"staging"|"production"} (empty = auto-detect) dev→staging or staging→prod; sets status; triggers reprovision
POST /api/v1/experiences/{name}/rollback handleRollback (288-320) — prod ← previous build; reprovisions production
POST /api/v1/experiences/{name}/pause handlePause — status→paused (from live/staging/development)
POST /api/v1/experiences/{name}/resume handleResume — status→live (from paused)
POST /api/v1/experiences/{name}/retire handleRetire — status→retired (near-terminal)
POST /api/v1/experiences/{name}/stream handleStream — returns webstream URL for preview
GET /api/v1/experiences/{name}/builds handleListBuilds — build history
GET /api/v1/pipeline/health handlePipelineHealth ?experience= per-experience status + recent builds

Lifecycle routes — admin web (all RequireWebAuth, server.go:102-109)

GET /admin, GET /admin/experiences/{name}, and POST /admin/experiences/{name}/{promote,rollback,pause,resume,retire,stream} (handlers_web.go). The web promote (handleWebPromote, 92-142) ignores any body and always uses autoDetectPromoteTier, unlike the JSON API which honors {"to":...}.

Promotion semantics worth noting

  • handlePromote (handlers.go:225-268): staging copies DevelopmentBuild→StagingBuild and calls reprovisionStaging (only nodes whose ReleaseChannel=="staging"). production shifts Production→Previous, Staging→Production, sets status live, and calls reprovisionProduction (all non-staging nodes in the experience's districts) — this pushes to live public venues.
  • Guard rails already present: "no development build to promote", "already on the same build", "no staging build", "no previous build to rollback".
  • auto_promote=true (handleBuildNotify, handlers.go:422-432) sends every new build straight dev→staging→production automatically.

Build-notify path (FindByWatchName)

POST /api/v1/builds/notify (handlers.go:382-449) is keyed on watch_name, not experience name: st.FindByWatchName(req.WatchName) (store.go:221-229) fans one build out to every experience sharing that watch_name, registers it as a new DevelopmentBuild, flips draft→development, and auto-promotes if the flag is set. This is how a Cyborn Perforce submit becomes a development build. Relevant here because a creator's "development build" state originates from this watch_name match, and because watch_name (like developer/owner) is an untrusted free-text field.


Which actions are safe to expose to a creator

Safe (self-service):

  • Read: their experience(s), build pointers, build history, pipeline/health, stream preview URL.
  • promote to staging (dev→staging) — low blast radius, staging nodes only.
  • pause / resume — reversible status toggles.
  • rollback — reverts production to the previous known-good build (a safety action).

Gate behind policy / operator confirmation (recommended: allow but flag clearly):

  • promote to production — pushes to live public venues. Recommend allowing for the owning creator but making it an explicit, clearly-labelled second action (not the same button as staging), so a creator cannot promote to prod by accident.

Operator-only (never expose to creators):

  • POST /experiences (create), DELETE (delete), PUT (edit developer/owner/districts/watch_name/auto_promote). These control ownership, venue placement, the watch_name fan-out, and auto-promote — i.e. the very fields the authorization boundary depends on. retire is near-terminal (no path back except operator) — keep operator-only or behind a strong confirm.

The central gap and two ways to close it

The library authorizes by a single privileged token and is not org-aware. Something must map "this signed-in person, member of org X" → "may act only on experiences whose developer/owner = X". Two options:

Option A (recommended) — Scope in the hydramancer proxy; library unchanged

Mirror the existing Perforce proxy pattern exactly (hydramancer/internal/api/handlers_provision.go). hydramancer:

  1. Resolves the caller via iamnim /api/me + /api/me/memberships (already wired: server.go:82-93, internal/iamnim).
  2. Holds the library admin token as server-side config (never exposed to the browser), just as the Perforce proxy holds its downstream trust.
  3. For reads: calls GET /api/v1/experiences (or pipeline/health) with the admin token, then filters to experiences whose Developer or Owner is in the caller's org slugs before returning.
  4. For actions: re-checks that {name} belongs to one of the caller's orgs, allows only the safe verb set, then proxies to the library with the admin token.
  • Pros: zero change to hydraexperiencelibrary; reuses the proven proxy + session-forwarding pattern; keeps the privileged token server-side; ships fastest.
  • Cons: the scoping/allow-list logic lives in the portal; the library remains fully-trusting if its token ever leaks; a second consumer would have to re-implement scoping.

Option B — Make the library org-aware

Add org context to the library: either accept a forwarded iamnim session (validate against iamnim like hydraperforceprovision does) or a caller-org claim, and filter list/get and authorize lifecycle by Developer/Owner inside the library itself.

  • Pros: the boundary lives with the data; any consumer is safe; token leak is less catastrophic.
  • Cons: larger change to a stable service; introduces an iamnim dependency into the library; slower; duplicates auth plumbing hydramancer already owns.

Recommendation: Option A now (it matches the platform standard already used for Perforce and needs no change to the library), with Option B noted as the eventual hardening step if a second consumer appears or defense-in-depth is wanted.


Implementation plan (Option A)

hydraexperiencelibrary: no code changes required. (Optional, non-blocking: add an org filter query param to GET /api/v1/experiences later to reduce over-fetching.)

hydramancer:

  1. Config + client: add experienceLibraryURL + experienceLibraryToken to Server (internal/api/server.go), plus a small internal/explibrary client (list, get, builds, promote, rollback, pause, resume, stream) — mirrors internal/iamnim.
  2. Scoping helper: orgSlugs(session) → membership slugs; ownsExperience(exp, slugs) → exp.Developer ∈ slugs || exp.Owner ∈ slugs. Central choke point for the boundary.
  3. Read view: extend webExperience (server.go:82-94) / experience.html with an "Your experiences" panel — for each org slug, list scoped experiences with status + dev/staging/prod build numbers and a "preview" link. Reuse the sign-in gating already there.
  4. Scoped action proxies: new routes POST /api/v1/experiences/{name}/{promote,rollback,pause,resume} on the portal that (a) require an iamnim session, (b) resolve memberships, (c) fetch the target experience via the admin token and verify ownership, (d) permit only the safe verb set (staging-promote as primary; prod-promote as a distinct, clearly-labelled action; pause/resume/rollback), (e) proxy to the library with the admin token, (f) pass status/body straight back — structurally identical to handleProvisionPerforce.
  5. Never expose create/delete/PUT/retire to creators; the admin token stays server-side and is never sent to the browser.
  6. Verify against the real gallo-romeins-museum experience (developer cyborn) with a Cyborn iamnim identity, and confirm a member of another org sees none of it.

Risks / decisions to confirm

  • Free-text ownership fields. Developer/Owner/watch_name are unvalidated strings set by operators via create/PUT. The whole boundary rests on Developer/Owner exactly equalling an iamnim organization_slug. Decide and enforce a canonical slug convention (is it cyborn? gallo-romeins-museum? both map to the Cyborn org?) and treat these fields as security-relevant (operator-only, as above). A membership-to-field mapping table may be needed if slug ≠ field value.
  • Owner vs Developer. Cyborn is the developer; the museum is the owner. Both are distinct orgs. Confirm whether lifecycle control belongs to the developer org (Cyborn, who builds), the owner org (museum, who runs the venue), or both. The plan currently grants control to either — tighten if the museum should not, say, promote to production.
  • Production promote blast radius. Prod promote reprovisions live public venues. Confirm creators may do this self-service or whether it stays operator-gated.
  • Token custody. The library admin token becomes a hydramancer server-side secret. Ensure it is injected via env/land config, never templated into HTML.

Comments (1)

claude-ops 15 Aug 2026 12:02

Duplicate — this is an intermediate investigation artifact from the assessment workflow. The consolidated findings + plan live in #490. Closing as duplicate.