HydraIssues

hydragon Phase 1: Perforce depot hydragon, watcher, first build in the experience library
open feature Project: hydragon Parent: #615 Reporter: cederik 31 Aug 2026 19:15

Description

Phase 1 of the hydragon MVP (#615). Stand up the Perforce lane and prove the pipeline before any game content exists.

Copy the pattern from #479 (galloromeinsmuseum) and #541 (mercatormetahuman). Runbook goes in hydraperforceprovision.

DECIDED 2026-09-01: hydragon gets its OWN Perforce machine

Do NOT provision on hydraskin-perforce-1. That node has a 106 GiB pool already over-committed at 150 GiB by two live customer scales (#633). hydragon gets a separate machine.

MINIMAL COST SPECIFICATION (user asked for cheapest viable, 2026-09-01):

cx33 plus a 100 GB Hetzner Volume. About 13.25 EUR per month. Verified prices in nbg1, hydraexperiencenet context:

Part Spec Price
cx33 4 cores, 8 GB, 80 GB local 8.49 EUR per month
Volume 100 GB about 4.76 EUR per month
Total about 13.25 EUR per month

Put the depot ON THE VOLUME from day one. Volumes resize up in place, so we never migrate. Local disk cannot grow, which is the whole reason to start on the Volume even though 80 GB local looks like enough today.

Why cx33 and not cheaper: cx23 is 5.49 EUR but gives 2 cores and 4 GB. p4d survives on 4 GB for a small team, but it gets tight during a large submit or a p4 verify. Three euros a month is not worth that risk on the machine holding the source of truth. cx33 is also exactly what galloromeins ran on originally, so it is a proven shape.

Why not the earlier proposal: a cx43 plus a 500 GB Volume is about 40 EUR per month. That was sized for a full Paragon library. Realistic day-one need is the Serath pack, the project source and a few retained builds, roughly 60 GB. Start at 100 GB and grow when the depot tells us to, rather than paying for headroom for a year first.

Sizing note: the Volume needs a free-space alarm from day one. #633 exists because nobody was watching a pool.

This retires the pool blocker and the port-collision question. A dedicated machine also means the watcher instance for hydragon lives there, which sidesteps the single global perforce.port limitation (#631) instead of working around it.

DECIDED 2026-09-01: the dedicated box follows the hydraskin/scale model

Same shape as the other Perforce servers. hydraskin install on the box, then
hydraperforceprovision ensure-server hydragon creates the p4scale-hydragon Incus
container. Consistent tooling, one way to operate every Perforce server we run.

Suggested name hydraskin-perforce-2, matching hydraskin-perforce-1.

THE TRAP: hydraskin would put the pool on the wrong disk

Read this before running hydraskin install. It is exactly how #633 happened.

hydraskin install creates a btrfs loopback pool and sizes it from the free space
on /var/lib, at 75 percent (PoolSizeGB(avail, 75) in
internal/install/install.go). The pool file lands at
/var/lib/incus/storage-pools/default.

On a cx33 that means the pool is sized from the 80 GB local disk, not from the
Volume. About 55 GiB of usable pool, on the disk we deliberately chose not to store the
depot on. The 100 GB Volume sits there unused, and we have rebuilt the same
undersized-shared-pool problem that #633 is open for.

Mounting the Volume at /var/lib/incus does not fix it either: df still reports
/var/lib from the root filesystem, so the size heuristic reads the wrong number even
though the file lands in the right place.

The fix

InitIncus is documented to skip initialisation if a pool already exists. So:

  1. Create the server and attach the Volume.
  2. Format and mount the Volume.
  3. Install Incus, then create the btrfs pool on the Volume, by hand, at the size we
    actually want.
  4. Run hydraskin install. It finds a pool and skips its own creation.
  5. Confirm btrfs quota enable ran on the new pool. Without quotas a size: is
    accepted and enforces nothing, which is the note in hydraskin's own CLAUDE.md.
  6. Then ensure-server hydragon.

Also do on this box, because we know they bite

  • Set manage_ufw in the node configuration. It is absent on hydraskin-perforce-1, so
    ensure-server never opens the port, and the scale looks healthy while being
    unreachable.
  • Give the container an honest quota against the real pool size. No thin
    over-commitment. That is the whole point of a dedicated box.
  • Add a free-space alarm on the pool from day one. Nobody was watching on the other
    node, which is why #633 went unnoticed.
  • One watcher for one p4d on this box, so the single global perforce.port limit
    (#631) never applies here.

Tasks

  1. Provision the hydragon scale and depot with hydraperforceprovision ensure-server, on host port 1668 (1666 is gallo, 1667 is visitflanders, 1668 verified free). Stream depot, mainline //hydragon/main. NOTE: manage_ufw is absent from the node configuration, so ensure-server will NOT open 1668. Add the firewall rule by hand or the scale is unreachable while looking healthy.

  2. CORRECTED 2026-08-31: do NOT set 200 GiB. The shared Incus pool on hydraskin-perforce-1 is 106 GiB with 29.4 GiB used, and galloromeinsmuseum (60 GiB) plus p4scale-visitflanders (90 GiB) already over-commit it. A 200 GiB quota would succeed and mean nothing, and a full pool wedges p4d for Cyborn and Soulmade too (#543, #633). Add storage or use a separate node BEFORE this issue starts. Then size the quota against real capacity. Do not inherit the 5 GiB profile default either.

  3. Create the users and the group. Least privilege: write on //hydragon/... only. Verify no access to the other depots.

  4. Set up the workspace on chunky-turnip-23 (node-b961f1c8): p4 client, stream bound workspace, P4CLIENT set. Watch the depot name against host name collision, which is the galloromeins trap.

  5. Stand up a SECOND hydraperforcewatcher instance for hydragon. The watcher has one global perforce.port, so you cannot add a watch to the running one, and editing its configuration repoints Cyborn's watcher (#631). New instance needs its own configuration, state directory, systemd unit and agent identifier. Put .hydrabuild.yaml at STREAM ROOT: the watcher reads <stream>/.hydrabuild.yaml literally.

  6. Verify the watcher's hydracluster admin token is current. A stale token drops build notifications silently (#609).

  7. Write docs/runbooks/ and docs/testbooks/ entries.

  8. Create the hydragon experience in hydraexperiencelibrary, with a watch name matching the watcher, BEFORE the first submit. A missing experience returns 404, and the watcher drops the build silently and marks the changelist processed (#632).

  9. Check hydramirror free space first. It has 2.52 GB of 40 GB (#630). A trivial artifact fits; a real package does not.

Reality check added 2026-08-31

The watcher PUBLISHES, it does not BUILD. You submit an already packaged Unreal build under Builds/. Nothing in any repository cooks or packages Unreal. "No manual copy" means no manual copy of a packaged build.

Done when

A trivial submit produces a build that appears in hydraexperiencelibrary with no manual copy, and the notification reaches hydracluster.