Context
Hydra has three build pipelines that are all instances of one shape:
source -> provision access -> trigger/watch -> build on an environment-specific backend -> artifact -> store -> promote/deploy -> track/notify
- Apple (hydraapplepipeline + hydrapipelinerunnerapple): build on Mac hardware -> IPA -> TestFlight.
- Unreal (hydraunrealengine + hydraunrealengine-server + hydraperforcewatcher): package on Windows -> archive -> mirror -> ExperienceLibrary.
- Container/git (hydragitwatcher, the new #492 path): buildkit-as-a-scale -> OCI image -> scaleregistry -> scale.
A three-agent recon (Apple, Unreal, shared/container) found that the shared pipeline layer ALREADY EXISTS, and that the container path reinvented most of it instead of plugging in. This issue captures the reuse map so we do not build a second build-tracker/runner-manager/promotion model.
The shared layer that already exists (reuse, do not rebuild)
- HydraRelease (module + service): the build/release source of truth. POST /api/v1/builds, GET /api/v1/builds, plus releases promote/rollback. Apple and Unreal feed it; it gives promotion and rollback for free.
- HydraExperienceLibrary: the experience build-record + notify + lifecycle. Its build-notify at POST /api/v1/builds/notify is ALREADY SCM-agnostic ({watch_name, scm_type, scm_ref, build_url}) and drives draft -> development -> staging -> live with promote/rollback/AutoPromote and channel-based reprovision deploy. This model is build-type-neutral; reuse as-is.
- Watcher/detect pattern (hydraperforcewatcher/pkg/watcher): per-target poll loop, last-processed-ref state, idempotent per-item processing, fixed poll -> fetch -> upload -> notify -> record -> push -> cleanup. hydragitwatcher already mirrors this.
- Job/status model + dashboard (hydraunrealengine-server): Job{status: preflight|packaging|uploading|completed|failed|cancelled} + POST/PATCH /api/v1/jobs + a report client (hydraunrealengine/pkg/report). A clean generic build-job tracker.
- Runner/backend-manager (hydrapipelinerunnerapple + server store.Runner): agent phones home, server holds the privileged credential and mints short-lived join tokens, heartbeats drive an online/offline state with a timeout sweeper, multi-instance scaling. Backend-agnostic; this is the template for managing a pool of build workers.
- Shared service skeleton: hydraapi (health), hydramonitor (SSE /api/v1/events), hydraauth (auth middleware), hydraserve (server), hydrawebcomponents (UI), hydraclusterapi (scale-op wire contract), hydrarelease/pkg/updater.
- hydrapipeline: pure-reader dashboard. It renders any service that exposes GET /api/v1/builds or GET /api/v1/jobs via Component.BuildsURL / JobsURL. No new dashboard code is needed to surface a pipeline that follows the convention.
Where the container path (hydragitwatcher) is the odd one out
hydragitwatcher imports only hydraclusterapi + hydrarelease/pkg/updater and reinvents the rest: its own state model, reporter, notifier, and health handler; a separate hydragit.experiencenet.com dashboard. Concretely it does NOT:
- register its builds in HydraRelease (POST /api/v1/builds) - so container builds are invisible to the build/release source of truth and get no promote/rollback.
- feed the SCM-agnostic ExperienceLibrary build-notify - it emits its own shape to hydracluster only.
- expose GET /api/v1/builds - so hydrapipeline cannot render it.
- use the standard skeleton (hydraapi/hydramonitor/hydraauth/hydraserve).
There is also a live contract drift to fix ecosystem-wide: hydraperforcewatcher still emits the legacy changelist field while the receiver already expects scm_type/scm_ref/build_url/watch_name. The container path should emit the modern shape from day one (scm_type:git, scm_ref:, build_url:).
Reuse / plug-in / build-fresh / keep-local map
REUSE (plug in, do not rebuild):
- HydraRelease builds + releases (POST /api/v1/builds; promote/rollback).
- ExperienceLibrary build-notify + draft/development/staging/live promote/rollback + channel-based reprovision deploy.
- The shared skeleton (hydraapi/hydramonitor/hydraauth/hydraserve) and the hydrapipeline dashboard convention (expose GET /api/v1/builds).
- The job/status model + report client pattern (hydraunrealengine-server) if a per-job tracker is wanted beyond HydraRelease.
- The runner/backend-manager pattern (Apple) as the template for a managed build-worker pool.
BUILD FRESH (genuinely new; do NOT copy the Apple/Unreal design here):
- The buildkit build-backend registry + capacity/queue/dispatch. Both Apple and Unreal are capacity-naive (single job per box; dispatch delegated to GitHub/Apple). buildkit gives real concurrency, so the backend pool + queue is the one part to design well.
- Artifact transport = registry push (the image ref becomes build_url) instead of a mirror archive PUT / chunked transfer.
KEEP LOCAL (container-specific, already correct):
- pkg/builder (buildkit driver) and pkg/deployer (scale-op via hydraclusterapi, the opaque node-op relay from #497). The deployer is already the model of correct shared reuse.
Action items (hydragitwatcher)
The one-line summary
The generic pipeline is HydraRelease + ExperienceLibrary + the shared skeleton + the hydrapipeline dashboard. Apple and Unreal already plug in. The container-build path should plug in too; the only genuinely new component it needs is the buildkit build-backend pool + queue.
Cross-references
#492 (git-push-to-deploy), #496 (declarative special-scale config), #497 (control-plane / node-agent split). Related repos: hydrarelease, hydraexperiencelibrary, hydrapipeline, hydraunrealengine-server, hydrapipelinerunnerapple, hydraperforcewatcher, hydragitwatcher.
Builder placement guardrails (learned building rogue, 2026-08-18)
Standing up the container build machine as a scale surfaced concrete constraints:
- Builders need capacity and the right OS, NOT bare metal. hydraskin runs fine on VMs or metal; only VM-ISOLATED workloads need bare metal (Incus VMs need /dev/kvm).
- The container builder does not need a VM or privilege: rootless buildkit (moby/buildkit:rootless) runs as an ordinary unprivileged scale. That is the target shape. Privileged container or VM isolation are stronger-isolation fallbacks, not requirements.
- NEVER place a heavy builder on a small or production node. Attempt 1 put a build VM on a venue Pi (node-50ab5309, 3.9 GiB RAM, ~10 production scales) with a 4 GiB limit; it OOM-killed the build VM (3 OOM events). Production scales survived (Incus contained it) and the VM was removed, but it was close.
- hcloud VMs have no nested virt, so they cannot host Incus VMs (no /dev/kvm) - a builder there must be a container (rootless or privileged), not a VM.
- Build backends are OS-specific and live on separate machines by nature: macOS Mac Minis for the Apple builder, Windows boxes for Unreal, Linux hydraskin nodes for containers. A heterogeneous pool, coordinated by one pattern; do not try to co-locate hydraskin with the macOS Apple builder.
Recommended builder: rootless buildkit as an unprivileged scale on a capable Linux node (RAM + cores), placed away from small/production nodes. Treat builder setup as its own tracked task rather than improvised alongside a deploy.
Rootless buildkit: approach VALIDATED, blocked on build-host networking (2026-08-18)
Rootless buildkit as an unprivileged scale WORKS. In an unprivileged Incus OCI container (security.nesting=true) running buildkitd directly with --oci-worker-no-process-sandbox (skip rootlesskit, which double-userns breaks), buildkitd came up with no privilege and reported platforms linux/amd64 AND linux/arm64 (arm64 via qemu-user-static binfmt registered on the host, which unlike Docker touches no iptables). This is the target builder shape: rootless buildkit scale + host binfmt for cross-arch. Build execution is buildctl inside the scale against the local socket, context pushed in, registry auth in /root/.docker/config.json.
Blocker that stopped the rogue build: the candidate node hydraskin-perforce-1 has no general container egress. The minimal buildkit image has no DHCP client so eth0 got no IPv4, and the node's hydrabr0 shows no outbound MASQUERADE (it was set up for inbound-only Perforce). So the build could not resolve or pull the base image from docker.io. The other candidate (venue Pi node-50ab5309) has working networking but only 3.9 GiB RAM shared with ~10 production scales.
Requirement for the build host: capacity (RAM+cores) AND working container egress (DHCP or static IP + outbound NAT + DNS), ideally a full hydraskin install (which sets REGISTRY_AUTH_FILE=/etc/hydraskin/registry-auth.json, the bridge, and NAT). Neither existing node qualifies as-is. Next step: provision a proper capable hydraskin build node (fresh, hydraskin install gives it networking+auth) OR deliberately fix hydraskin-perforce-1 egress; then run the rootless buildkit scale there and push rogue. rogue remains live on Hydra (v1.0.0) throughout; no production impact.
LOOP PROVEN: rootless buildkit scale built + deployed rogue (2026-08-19)
Full container build-to-deploy loop works, on a dedicated hydraskin build node (hydraskin-build-1, cpx42, hydraexperiencenet). Proven end to end: an unprivileged rootless buildkit built rogue from source, pushed scaleregistry.experiencenet.com/rogue:v1.1.0-amd64, the image launched as a scale with a /data disk, and it served HTTP 200 with the game page and /scores.
Key fixes that made it work (all recorded so we do not rediscover):
- Build host must be a full hydraskin install (gives Incus + NAT bridge + DNS + REGISTRY_AUTH_FILE). The perforce node was not, hence no egress.
- The builder container MUST get its IP via DHCP. The minimal moby/buildkit:rootless image has no DHCP client; forcing a manual static IP gave flaky TCP egress (DNS ok, sustained TCP to docker.io stalled). Fix: run buildkit inside a normal debian container that DHCPs (proper IP + fast egress), and install the buildkit release binaries there.
- Rootless buildkit runs unprivileged with: buildkitd --oci-worker-no-process-sandbox (skip rootlesskit, which the double user-namespace breaks) in a security.nesting=true container. buildctl builds against the local socket; source pushed in via incus file; registry auth in /root/.docker/config.json.
- ARM64 via emulation on an amd64 builder is impractical (a rogue build sat 40+ minutes stuck in the arm64 stage under qemu). The arm64 fleet (the Pis, where the public rogue runs) needs a NATIVE arm64 builder (a cax node). Build native per arch; do not emulate arm64 for real builds.
State: hydraskin-build-1 is the builder node (keep it; it makes sense to have a hydraskin for builders). buildkitd runs in its debian builder container (buildkitd is a manual process, not yet a systemd/entrypoint service - persistence is the remaining polish). The public rogue on the arm64 Pi is unchanged and live (v1.0.0); redeploying it from this pipeline needs an arm64-native builder.
On-demand / elastic builder pool (2026-08-19)
Builders should be ephemeral and on demand, orchestrated by a lean always-on manager. Control/compute split:
MANAGER (control) - lean, always-on, runs as an ordinary unprivileged scale on a Pi (a few MB, I/O bound; it builds nothing). Responsibilities: receive a build request; pick capacity (query hydracluster inventory for a node with spare RAM and the right arch, else decide to spin a node); launch a builder via the platform generic levers (incus launch a buildkit scale through hydracluster generic op-relay #497 with a scoped op-token, OR hcloud API to spin a node); dispatch the job; track status; tear the builder down. It is a SEPARATE small service that USES hydracluster generic ops, not logic added to hydracluster (keep the control plane generic per #497). Holds only credentials as env (#496).
BUILDERS (compute) - heavy, ephemeral, on demand, on capable hosts, pool size 0..N:
- Tier 1 default: ephemeral buildkit SCALE on an existing hydraskin node with spare capacity + matching arch. Launch on build, delete after. Seconds, near-zero idle cost.
- Tier 2 burst/arch: on-demand builder NODE via hcloud (hydraskin install, build, destroy). ~5 min ready, hourly cost, zero idle. Native arm64 = spin a cax node.
Make it clean:
- A purpose-built hydra-builder scale image: buildkit + proper networking (must DHCP, not the static-IP hack) + an entrypoint that runs buildkitd and takes a job. One-step launch, self-configuring, no manual buildkitd.
- Registry-backed layer cache (buildctl --export-cache/--import-cache to scaleregistry) so ephemeral builders warm-start instead of pulling the base image cold every build.
This is the runner/backend-manager pattern from the Apple pipeline (#502 above) with the pool allowed to go to zero. Net: the only always-on cost is a few MB on a Pi; compute is elastic.
Decision: hydraskin-build-1 was TORN DOWN (2026-08-19) since on-demand is the direction and an idle build node is unneeded; the setup is fully reproducible from the notes above (provision + hydraskin install + deploy registry-auth + debian builder container + buildkit binaries). Respin on demand.
Naming: the Linux/container pipeline follows the Apple convention (2026-08-19)
Mirror the Apple two-part shape:
- hydralinuxpipeline = pipeline control service, peer of hydraapplepipeline. Owns git-push-to-deploy, build tracking, status/dashboard. The role hydragitwatcher plays today folds in here (rename/expand hydragitwatcher into it, or keep it as a component under this name).
- hydrapipelinerunnerlinux = builder/runner manager, peer of hydrapipelinerunnerapple. This is the on-demand builder manager: reads the hydraskin overview via hydracluster inventory, evaluates arch + spare capacity, and launches/tears down ephemeral builders (a buildkit scale on a node with room per Tier 1, or an on-demand node per Tier 2).
Key difference from Apple, deliberate: hydrapipelinerunnerapple has a persistent AGENT running on each Mac (long-lived hardware, registered + heartbeated). hydrapipelinerunnerlinux is SERVER-SIDE ONLY: no persistent on-node agent, because Linux builders are ephemeral scales the manager spins and kills on demand. Same name/role in the family, minus the always-on agent half. That is what keeps it lean enough to run as a small always-on scale on a Pi.
Family, consistent with Apple/Unreal:
- hydralinuxpipeline (control) ~ hydraapplepipeline ; hydrapipelinerunnerlinux (builder mgr) ~ hydrapipelinerunnerapple.
Correction: watchers are keyed by VCS, not platform (2026-08-19)
Supersedes the note above that folded hydragitwatcher into hydralinuxpipeline. Watchers are named/scoped by the VERSION CONTROL SYSTEM they observe, not the OS/platform they build for. Three orthogonal axes:
- Watchers (by VCS): hydraperforcewatcher (Perforce), hydragitwatcher (Git). Detect a change, emit the SCM-agnostic build-notify {watch_name, scm_type, scm_ref, build_url}. Agnostic to what is built downstream.
- Pipelines (by platform/artifact): hydraapplepipeline, hydraunrealengine(-server), hydralinuxpipeline. Build + deliver for a specific target.
- Builder managers (by platform): hydrapipelinerunnerapple, hydrapipelinerunnerlinux.
Seam: the SCM-agnostic build-notify decouples watcher from pipeline, so a VCS watcher can feed any platform pipeline. hydragitwatcher therefore stays a SEPARATE git watcher (not part of hydralinuxpipeline) and can feed multiple pipelines: a pushed repo with a Dockerfile routes to hydralinuxpipeline; a pushed Unreal project could route to the Unreal pipeline. Same as hydraperforcewatcher feeding the Unreal pipeline today without being Unreal-specific. hydralinuxpipeline is FED BY hydragitwatcher, it does not contain it.