HydraIssues

District server (hydraguard-brussels, OVH 141.227.136.199) is a total SPOF — no off-box backup of hub key/mesh/certs
open improvement Project: hydraguard Reporter: 11 Aug 2026 18:11

Description

District server (hydraguard-brussels, OVH 141.227.136.199) is a total single point of failure

Read-only robustness assessment, 2026-08-11. Nothing was changed.

This one OVH box is simultaneously (a) the ONLY WireGuard hub (wg0, 26 peers), (b) the ONLY Traefik/ingress + ACME client for all public domains, and (c) the path public traffic takes to the LAN-private Pi scales. All three die together because they are the same box. There is no second hub, no off-box copy of any load-bearing state, and no DNS failover.

Verified topology

  • hydrascalerouter v0.10.0: 19 routes, all backends on the mesh (http://10.10.100.18:* / 10.10.100.19:* = pi-node-004 / pi-node-003). Backends are only reachable over wg0.
  • wg show wg0: 26 peers = 4 venue guards (AD6/Overijse, cloud-seven, rupelmonde, mobile-kit — site-to-site LANs), ~16 air/neck devices, 3 ipad-heads, and the 3 Pis (10.10.100.17/18/19).
  • DNS: all 19 experience domains and hydraguard.experiencenet.com resolve to the single IP 141.227.136.199 (single A record, no failover). By contrast hydracluster.experiencenet.com→46.224.29.125 and the mirror→46.225.165.81 live elsewhere.

1. Blast radius if it dies

When What breaks Recoverable without the box?
T+0 wg0 hub, Traefik, router API (:8099) all down at once (same host). —
T+0 All 19 public domains + hydraguard.experiencenet.com go dark. books/district/pipeline/perforce/venues/issues/streaming-monitor/nps/unrealengine/mancer/northstar/organization/transfer/mirror/… all offline. DNS still points at the dead IP. No — both ingress and the mesh path to backends are gone.
T+0–90s All 26 WG peers lose handshake: the 4 venue site-to-site tunnels drop; every air/neck/ipad device loses the mesh; the 3 Pis lose their 10.10.100.x address. No — peers just retry a dead endpoint.
survives hydracluster exec / web-shell to the Pis survives. Verified: each Pi runs hydranode/hydraskin holding an outbound TLS session from its own LAN to 46.224.29.125:443 (hydracluster) and 159.69.93.219:443 (scaleregistry) — these do NOT traverse the district hub. YES — this is the escape hatch that makes a rebuild feasible.

Answer to "can the Pis be reached at all if the hub is gone?": Yes, for control/exec — via the Pi→hydracluster outbound channel, which is independent of the district WireGuard hub. Public/experience traffic to the scales, however, is 100% dependent on the hub and is fully down until it is rebuilt.

2. State that lives ONLY on the box — and its backup status (the headline)

State Path Off-box backup?
WG hub private identity /etc/wireguard/hub.key (+ wg0.conf) NONE.
Mesh topology / IPAM /root/.hydraguard/mesh.yaml NONE off-box. Only .bak files and /root/.hydraguard/backups/* — all on the same disk.
20 Let's Encrypt certs /var/lib/hydrascalerouter/acme.json (255 KB) NONE.
Route table /var/lib/hydrascalerouter/routes.json (19 routes) NONE (but derivable from the scale registry).
Traefik static cfg /etc/traefik/traefik.yml NONE.

hydrabackup coverage of this box = zero. The district server runs no hydrabackup, no rsync/restic/borg/s3 — its only cron jobs are stock e2scrub/sysstat. The hydrabackup sidecar does exist and runs on the nbg1 mirror (mirror-a, 46.225.165.81), but its config.yaml has server.domain: '' — it is not pointed at the district server, and no district mesh.yaml/acme.json/hub.key exists anywhere on the mirror. The one load-bearing box is the one box the backup tooling doesn't cover.

Truly unrecoverable-without-it: hub.key. If the box and its disk are lost, a rebuilt hub gets a new public key, forcing manual re-provisioning of all 26 peers — every venue Omada/citymesh guard, every air/neck/ipad device, and all 3 Pis. mesh.yaml (IPAM/topology) is the same single-copy problem. acme.json is re-issuable but LE validation/duplicate rate limits will bite during a scramble.

3. Redundancy options, ranked by effort/risk

(a) Warm-standby second hub + second ingress + DNS failover — HIGH effort/risk. Honest constraints: WireGuard hub-and-spoke does not fail over cleanly — a standby hub must carry the same hub.key (else every peer needs re-keying or a second [Peer]); a standby Traefik needs replicated acme.json or its own certs (LE rate limits); today there is a single A record with no health-checked failover, so DNS work (low TTL + failover provider) is required too. Weeks of work plus ongoing config-drift risk. Defer as a later epic.

(b) Off-box backup of the load-bearing state + a rebuild runbook — LOW effort/risk. ← RECOMMENDED FIRST MOVE. Push hub.key + mesh.yaml + acme.json + routes.json + traefik.yml, encrypted (age/gpg), to the nbg1 mirror on a 6h timer. The cleanest path is to deploy the existing hydrabackup sidecar onto the district server (point server.domain at it) — it is literally built as a "YAML store backup sidecar" and the district box is the single place it is missing. Because hub.key is preserved, this turns an identity-loss + 26-peer re-provisioning event into a ~1-hour restore onto a fresh box with zero peer changes.

(c) Accept the SPOF, write the rebuild runbook only — LOWEST effort. Documents the manual rebuild, but if the disk is truly lost you still eat the full 26-peer re-provisioning because no key copy exists. Strictly worse than (b) for a small extra cost.

Recommendation

Do (b) now: get hub.key + mesh.yaml + acme.json + routes.json + traefik.yml off-box on an automated encrypted timer to the nbg1 mirror (reuse the hydrabackup sidecar already deployed on mirror-a), and write the rebuild runbook (restore state → bring up wg0 → start hydrascalerouter → repoint DNS). This is the cheap, high-value first move that removes the unrecoverable-identity-loss risk this week. Treat the warm-standby hub (a) as a separate, later effort. The surviving hydracluster control plane to the Pis means a rebuild is achievable — but only if the state exists off-box, which today it does not.

Evidence gathered read-only via ssh ubuntu@141.227.136.199 (sudo), the router API on :8099, one hydracluster exec call to pi-node-003, and inspection of scaleregistry (159.69.93.219) + mirror-a (46.225.165.81). No secret values are reproduced here.

Comments (1)

api 11 Aug 2026 18:35

Backup half is DONE and verified. Daily off-box encrypted backup now runs on the district (district-backup.timer, 04:30 UTC) pushing hub key/mesh.yaml/acme.json/service configs to nbg1 (backups/district-brussels/). Hybrid RSA+AES; decryption key on hydracluster (/root/dr/district-backup.key), independent of both the district and nbg1. Verified a full DR round-trip: fetched, decrypted, and confirmed the tarball contains wg0.conf + acme.json + mesh.yaml. Rebuild runbook committed: hydraguard/docs/runbooks/district-server-recovery.md. Remaining (deferred): the warm-standby second hub.