HydraIssues

hydraperforceprovision: ensure-server hardening from the soulmade rollout
open task Project: hydraperforceprovision Reporter: cederik 20 Aug 2026 19:13

Description

Found while provisioning the soulmade scale on hydraskin-perforce-1 (2026-08-20, see #541). ensure-server worked end to end but needed four manual fixes. Make them part of the tool or the runbook.

  1. ufw rule skipped on re-run. The first run failed at first-boot apt (container had no IPv4: the node was missing "ufw allow in on hydrabr0" + "ufw route allow in/out on hydrabr0"). The re-run completed the install but never applied "ufw allow 1667/tcp", so the scale was unreachable from outside. Fix: make the ufw step converge on every run, and add the bridge rules to "hydraskin install" so the node prerequisite cannot be missed.

  2. Container defaults are too small for p4d. The default profile gives 5GiB root quota and 512MiB memory. Submitting the 35GB Mercator migration failed mid-checkin with "Disk quota exceeded" and wedged the SSL listener. Fixed by hand: root size=90GiB + limits.memory=2GiB on p4scale-soulmade, and root size=60GiB on galloromeinsmuseum (it inherited the same 5GiB bomb). Fix: ensure-server sets root size and limits.memory explicitly, from config fields with sane defaults.

  3. New scales get only the hydra_admin super. Per the GRM pattern (#491) each scale that the provision service targets needs a provision_svc super, and a hydraperforcewatcher agent for the dashboard + build mirror uploads. The soulmade scale has neither yet.

  4. Decommission decision needed: old //soulmade on perforce.visitflanders is migrated (all 1645 digests verified on the new scale) and frozen read-only (Soulmade + Create groups downgraded). Obliterating it would free ~35GB on that box (95 percent full). Needs an explicit go.

Minor: "p4dctl status p4scale" inside the container reports "missing required parameter P4PORT" even while p4d runs fine.