HydraIssues

hydraskin-perforce-1 Incus pool is 106 GiB with 150 GiB of quotas committed across live customer scales
open bug Priority: high Project: hydraskin Reporter: cederik 31 Aug 2026 21:19

Description

Found during hydragon reconnaissance (#615) on 2026-08-31.

The measurement

hydraskin-perforce-1 (178.105.185.28) is a cx43 with a single 152.3 GiB disk. Incus runs on a loop file btrfs pool named default, capped at 106 GiB, with 29.4 GiB used. Every Perforce scale on the node shares it.

The quotas already over-commit the pool:

  • galloromeinsmuseum: 60 GiB
  • p4scale-visitflanders: 90 GiB
  • Total committed: 150 GiB, against a 106 GiB pool.

Why it matters

The quotas are thin. They are promises the pool cannot keep. Today that is fine, because 29.4 GiB is in use. It stops being fine as soon as any tenant grows.

Per #543, a full quota wedges p4d's SSL listener. Because the pool is shared, one tenant filling it takes the others down. Those others are Cyborn's live delivery and Soulmade's. This is a customer-facing outage waiting for a large submit.

hydragon made this visible: #623 asked for a 200 GiB quota, which would take the commitment to 350 GiB against 106 GiB of real storage. That number has been corrected on #615 and #623.

What to do

  1. Decide the real capacity plan. Options: a Hetzner Volume with a second storage pool, growing the loop file pool onto the free space on / (105 GiB free), or a larger server type.
  2. Do not put an Unreal game depot on this node until that lands. hydragon needs its own node or real storage.
  3. Add a pool free-space alarm. Nothing watches it today.
  4. Re-check the existing quotas against whatever capacity actually exists, and stop promising storage the node does not have.

Related: #543, #615, #623, #479, #541.