mirror-a is the only copy of every artifact published since 2026-02-23 and has no peer
to fall back on. Found while diagnosing #407; unrelated to that ticket and more dangerous
than it was.
The disk half is fixed. The redundancy half is not.
mirror-a was at 98% disk — 898 MB free of 38 GB — and had already hit
no space left on device on 2026-07-24 while writing an upload.
Root cause of the pressure was unbounded journald (3.6 GB, no SystemMaxUse) plus 297 MB
of btmp brute-force login records. Both reclaimed; SystemMaxUse=400M now set so it
stays bounded. Free space went 898 MB → 4.6 GB (98% → 88%) without touching any mirror
data.
Separately, max_storage was configured as 50 GB on a 38 GB disk, so the guard in
HandlePutFile could never fire — writes were accepted until the filesystem physically
filled, which is exactly what happened in July. Now set to 34 GB, which sits above the
31.2 GB currently tracked (so publishing keeps working) but below the disk, so the guard
actually engages. Verified after the change: PUT returns 201, GET returns 200.
peers: static: [] and no cluster_url, so peers: 0. Pull-through can never satisfy a
miss, which means a single missing file is terminal.
Worse, the origin (/var/www/releases) stopped receiving artifacts at the Feb 23 mirror
cutover. It still holds a stale partial copy frozen at that date. So for everything
published in the last five months, mirror-a is the sole copy — one 38 GB disk, no backup,
no peer, 2,994 files / 30 GB.
Losing that host loses every release artifact since February.
cluster_url so pull-through canhydrarelease's own workflow still rsyncs todeploy@releases.experiencenet.com:/var/www/releases/hydrarelease — files that areHandleGetFile writes a permanentstatus: failed entry to files.yaml on every unauthenticated miss — 158 entries sovendor/laravel-filemanager/js/script.js,error_log.php). Only persist stubs for authenticated or peer-backed paths.btmp reaching 297 MB indicates sustained SSH brute-force traffic against this host. Not
urgent, but a signal worth looking at.
The redundancy half was fixed earlier. This closes the capacity half.
disk free: 3.5 G -> 23 G (91% -> 36% used)
mirror data: 31 G -> 11 G
tracked: 30.13 -> 10.64 GiB (3286 -> 1195 ready entries)
removed: 1933 files, 19.48 GiB, 528 versions across 28 projects
kept: newest 5 per project and channel
Runway goes from roughly two weeks to about two years at the observed ~258 MB/day.
max_storage was corrected a second time, 34GB -> 32GB. At 34 the guard had 3.92 GiB of headroom against only 3.48 GiB of real free space, so the filesystem would still have filled ~450 MB before the guard fired — the July ENOSPC mode was still reachable. The first correction improved it without closing it.
Now: 21.36 GiB guard headroom against 22.91 GiB real free, so the guard fires first. Note max_storage is read once at startup, so it needs a systemctl restart hydramirror.
hydramirror prune (v1.16.0) — dry run by default, --apply required.
The correctness that mattered was version ordering. Lexically v2.0.9 > v2.0.76, so a naive sort keeps the oldest releases and deletes the newest — on hydrahead, which spanned exactly that range, it would have deleted the current release and retained 2017-era ones. Ordering is numeric per segment, tested against real versions from this mirror.
Refusals, each because these files had no second copy on the machine:
--keep 0 is an error, never "delete everything"releases/ is considered — builds/ (2.4 G), transfers/, backups/ untouchedfailed stubs left to CleanupFailedSequence for each run: force an offsite sync, verify by hash that files about to be deleted are in the backup, dry run, apply, verify. hydrahead went first as a scoped trial (3.65 GiB) before the remaining 26 projects.
Newest kept version of six projects fetched through the mirror, not just checked on disk:
hydrahead v2.0.76 200 11656322 hydranode v1.10.33 200 13400560
hydrabody v2.0.65 200 11431744 hydraskin v0.8.0 200 6697144
hydracluster v2.0.98 200 16662483 nimsforest v0.84.3 200 21610788
A pruned version returns a clean 404. Service active throughout.
Version directories with no files and no store entries — debris from aborted uploads, predating this work. Zero bytes, so no urgency, but it means disk and store had already drifted. prune correctly ignores them, since it only acts on store entries.
Worth a small hydramirror gc or similar eventually. I nearly misread these as prune failing to clean up after itself; the directories prune did empty were correctly removed (v2.0.93, v2.0.92, v2.0.91 are gone).
Retention is manual — nothing schedules prune. A weekly timer at --keep 5 would keep this from recurring, but automatic deletion of single-copy artifacts deserves a deliberate decision rather than being switched on quietly.
Redundancy half is now fixed — and verified
mirror-ais no longer the sole copy. There is now an hourly, checksum-verified offsite copy of the whole artifact store on a Hetzner Storage Box.Nothing under
/var/lib/hydramirrorwas deleted, moved, truncated or overwritten. The sync is read-only at the source and append-only at the destination.hydramirrorwas never stopped or restarted; free space onmirror-awas 3.5 GB before and 3.5 GB after.Why a Storage Box and not a peer
I read the peering code before assuming it would help. It would not have.
peer.Client.Pullonly fires fromHandleGetFileon a local miss — pull-through is on demand only. There is no background replication anywhere in the codebase. A new peer would have started empty and stayed empty except for files somebody happened to request. It would have satisfiedpeers: 0in the health output without ever being a second copy — arguably worse than the current state, because it would have looked fixed.Restoring origin sync from
hydrareleasewas also not it: that workflow only rsyncs hydrarelease's own artifacts to/var/www/releases/hydrarelease, so reinstating it recovers one project out of ~45.I measured the origin rather than trusting either description of it.
/var/www/releaseson46.225.120.7holds 19 GB across 29 projects — so it is not completely frozen at Feb 23 as this ticket states;land v0.87.0landed there today, andhydrareleaseas recently as 2026-07-20. But it has diverged so far that it was never a usable fallback:So the ticket's conclusion was right even though the mechanism was slightly off: a handful of projects still write to the origin, but 98.8% of the store existed in exactly one place. That is the number this work removes.
What was actually done
storage-box-1/ subaccountu520179-sub2,/home/mirror-a/usr/local/bin/hydramirror-offsite-backup.shviahydramirror-offsite.timer(hourly)/usr/local/bin/hydramirror-offsite-verify.sh/var/log/hydramirror-offsite.log/root/.ssh/hydramirror_backup, port 23Deliberate choices:
--delete. The destination is append-only, so data loss onmirror-acan never propagate to the backup.-Hpreserves the hardlinks made byPOST /api/v1/link.config.yamlis not copied verbatim — it goes asconfig.yaml.redactedwithadmin_tokenstripped. The artifacts are irreplaceable; the token is not, and an unencrypted offsite copy of it only widens the blast radius.flockso an hourly run can never overlap a long one.Also hardened, both reversible: daily Storage Box snapshots (04:30 UTC, 10 retained), delete-protection on the Storage Box, and delete + rebuild protection on
mirror-a. Rebuild protection deliberately makes a host-snapshot restore a two-step act, so nobody accidentally rolls the host back and reverts every artifact published since the snapshot.Verification — not "rsync exited 0"
This is the part the ticket rightly cares about.
files.yamlalready stores a SHA256 per file from ingest, so I re-hashed every file on the Storage Box and compared against the digest hydramirror itself computed:Because the comparison is against ingest-time hashes, this also independently proves there is no bit-rot on
mirror-a— all 3122 files still hash to what they hashed to when they were published.Two incidental findings while checking consistency:
guardcheck/probe.txt, areadyentry infiles.yamlwhose file is gone. It is a health probe, not an artifact, so nothing irreplaceable is lost, but the store is carrying one stale row. The verify script reports these separately rather than failing on them.Correction to this ticket's premise
Hetzner automated daily host backups were already enabled on
mirror-a(window 10:00–14:00 UTC, 7 retained, newest 2026-08-03 10:21). So "no backup" was not strictly accurate. But it was never verified, and it is the wrong instrument for this risk: a whole-host image can only be restored by rolling the server back, which reverts every artifact published since the snapshot. The tool for recovering a lost file was also a tool for losing newer ones. It is now a genuine second layer behind a file-level copy, rather than the only layer.The runbook previously asserted these snapshots as "what is backed up" with an unexercised restore procedure. Corrected in cederikdotcom/hydramirror#1.
Capacity — analysis and proposal (nothing executed)
The guard still does not protect what it was meant to protect
From the live health endpoint:
Guard headroom is 3.92 GiB but real free space is 3.48 GiB. The filesystem still fills ~450 MB before the guard can fire, so the July failure mode —
ENOSPCmid-upload instead of a clean507— is still reachable. Loweringmax_storagefrom34GBto32GBcloses that (still comfortably above the 30.08 GiB tracked, so publishing is unaffected today). I did not change it, because it also shortens the runway to a hard publish stop and that is an availability call to make alongside the disk decision below.Runway
August is running ~258 MB/day (0.72 GiB in 3 days), consistent with the ~8.1 GiB/month average over the last three full months. At that rate
mirror-aruns out of physical disk in roughly two weeks.hydrascaleroutershipping ~108 MB per release across two architectures makes this worse, not better.What is actually consuming 30 GiB
27.31 GiB of the 30.08 GiB is
releases/, and 27.23 GiB of that is theproductionchannel. It is almost entirely superseded versions:Worst offenders under keep-newest-5:
hydrahead(75 versions, 3.65 GiB reclaimable),hydrabody(53, 2.93),hydraexperiencenet(47, 2.56),hydraapplepipeline(40, 1.97),hydracluster(76 versions for 1.08).Proposal — keep the newest 5 versions per project+channel. That takes the store from 30.08 GiB to ~10.6 GiB and turns a two-week runway into roughly two years at current growth. I have not deleted anything and am not proposing that anyone do so by hand:
hydramirrorhas no retention code today (store.CleanupFailedonly touchesfailed/pullingstubs), so this needs a realprunesubcommand with--dry-run, tests, and a tag — and it should only run now that a verified second copy exists and can be re-verified afterwards.Alternatives if deleting released artifacts is unacceptable:
cax11→cax21— 40 GB to 80 GB disk, +4.50 EUR/mo. Needs a power-off, and Hetzner disk growth is irreversible (no downgrade afterwards).data_dirmoves onto it, which means relocating the only copy. Safer to attempt now that the offsite copy is verified, but still the riskiest of the three.My recommendation: retention first (free, largest effect, and the Storage Box already holds everything being pruned), disk growth only if the policy is rejected.
Also proposed, not done
HandleGetFilestill persists astatus: failedrow on every unauthenticated miss — now 163 entries, including scanner noise. Unauthenticated misses should not be able to growfiles.yamlwithout limit./var/www/releasesrsync fromhydrarelease'srelease.yml(lines 77–94). It writes files nobody serves and creates a stale partial copy that reads like a backup — exactly the confusion that made this ticket necessary.Cost
Zero incremental. The Storage Box (
storage-box-1, bx11, 1 TiB, 3.20 EUR/mo) already existed and was completely empty — 0 bytes used since it was created in December 2025. The offsite copy uses 23 GiB of 1 TiB, so there is room for retention offload later. No server was created. The proposals above are the only things that would add cost.