Applied the three low-risk mitigations to all three live Pi nodes, one at a time, strictly non-disruptively: no incusd restart, no reboot, no container stop/start. Order followed: pi-node-001 → pi-node-004 → pi-node-003. incusd was active on every node before and after.
incus config set images.auto_update_interval 0 at server level (space-syntax printed the expected deprecation warning but the value took). Verified with incus config get images.auto_update_interval → 0 on each node. Prior value on all three: empty/unset (Incus default ≈6h). Stops the periodic refresh loop without restarting the daemon or touching containers./etc/systemd/system/incus.service.d/10-start-timeout.conf = [Service] / TimeoutStartSec=1800 (30 min), then systemctl daemon-reload (rc=0). daemon-reload only — incus was NOT restarted; the new timeout applies to the next start.159.69.93.219 scaleregistry.experiencenet.com (releasesandscaleregistry) where not already present. NOTE: this pins the registry IP — if releasesandscaleregistry ever changes address, this line must be updated/removed on all three nodes.| Node | A interval→0 | B drop-in | C hosts | incusd before→after | domain spot-check after |
|---|---|---|---|---|---|
| pi-node-001 (node-2b224f6a) | applied, get=0 | applied, reload rc=0 | applied | active → active | test scales only, none district-routed |
| pi-node-004 (node-50ab5309) | applied, get=0 | applied, reload rc=0 | applied | active → active | issues 200, books 200, hydramancer 200 |
| pi-node-003 (node-6bf57aed) | applied, get=0 | applied, reload rc=0 | applied | active → active | hydrapipeline 200, hydravenues 200, hydradistrict 200, hydraperforce 302 (baseline redirect, unchanged) |
status: ok, routes: 19, traefik: running, backends_reachable: 19 — checked before and after every node. 19/19 backends reachable throughout; no route dropped, no degradation, no revert needed on any node.
On all three nodes prior to change: images.auto_update_interval was empty/unset (Incus default); no pre-existing 10-start-timeout drop-in; no pre-existing /etc/hosts entry for scaleregistry. (Existing 10-registry-auth.conf drop-in left untouched.)
incus-upgrade #455 and the deeper fix (hard registry dependency + short readiness contract) remain outstanding. Leaving #445 open.
— applied via hydracluster exec / web-shell (root), operator cederik
Root-cause post-mortem — pi-node-003 incusd startup deadlock (RESOLVED)
Non-disruptive investigation. Daemon is healthy now (
incusactive, 8/8 containers RUNNING, uptime since 2026-08-06 15:25 CEST). No restart/reload/config-write/container action was taken. Read-only: journal,ls/staton the raft dir, daemon-mediated reads (incus list/image show/warning list), and comparison vs healthy pi-node-004. Livedb.binwas never opened/locked.Root cause (best-supported, with the one honest caveat below)
Trigger = network isolation during the LAN move. scaleregistry.experiencenet.com became DNS-unresolvable — journal shows
dial tcp: lookup scaleregistry.experiencenet.com on 127.0.0.53:53: server misbehaving(systemd-resolved stub, upstream unreachable). 8 of 9 images areauto_update: trueandimages.auto_update_intervalis unset (default 6h), so image auto-update runs at/after startup and periodically. With the registry unreachable, every refresh failed: 248 "Failed to update the image" errors, 00:16–09:14Z on 08-06 — matching the deadlock window.Incus persists refresh failures as DB-backed warnings; those writes go through the single-node cowsql/dqlite raft log. That write activity, plus 52 daemon restarts each SIGKILLed after the 600s start-post readiness probe timed out (
Error: Daemon still not running after 600s timeout,Main process ... code=killed status=9/KILL,unclean termination of a previous run×32), inflated the raft store. Each unclean SIGKILL finalizes the current open segment prematurely, so the crash loop minted many short segments.The reported symptom (reads work, writes hang,
wchan=futex_wait_queue,activatingforever) is the raft commit path stalling: reads are served from local sqlite state and keep working; writes wait on a raft apply that never completes → the daemon never signals ready → systemd start-post kills it at 600s → repeat. In early cycles the daemon reached the image task (logged failures); in later cycles (09:14–15:24, 21 more restart attempts) it logged no image failures and simply hung before readiness — consistent with the hang sitting in dqlite/raft startup itself.Recovery at 15:25Z was clean and fast (55s) once the network/registry came back — replaying the same on-disk log without trouble.
Honest caveat: because that identical 75-segment / ~2718-entry log booted cleanly in 55s the moment the network returned, raft-log size is not independently sufficient to cause the deadlock and is not a standing boot-time hazard — the deadlock required the isolation condition. I could not observe the futex holder directly (incident is resolved; enabling raft-debug/attaching a debugger to the live daemon would risk it), so the exact internal lock is inferred from the read-works/write-hangs signature, which is characteristic of the dqlite raft commit path.
Controlled comparison — pi-node-004 (healthy)
Same Incus 7.2 (aarch64), same
auto_updateconfig, same raft trailing (segments retained back to index 1), and it also restarted on 08-06 (09:35) — but it had network, saw 0 image failures, restarted once, and sits at 9 closed segments (index 1→2049, 2 snapshots). The only variable that differs is network reachability during the restart. This is the strongest evidence the trigger was isolation, not anything intrinsic to 003's store.Answers to the 5 leads
.metafiles). 004 is structurally identical. The ~332-entries-past-snapshot gap is normal (dqlite snapshots ~every 1024 entries). The many segments on 003 come from 52 unclean restarts finalizing short segments (avg ~36 entries/seg vs 004's ~228/seg), not from snapshotting. A large un-snapshotted log does not explain the stall — the log isn't large and replays in 55s.futex_wait_queueis a generic Go/cgo lock wait. I am flagging this as not the mechanism rather than confirming a bug. Deeper confirmation would need raft-debug logging in a maintenance window (unsafe now).auto_update: true, interval default 6h.auto_updateimages + default 6h interval + hard dependency on scaleregistry reachability. If the node reboots/restarts while network-isolated (another LAN move, DNS or registry outage), the same failure-loop + start-post-timeout dynamic can recur. Risk is conditional on (restart) AND (isolation) coinciding.Safe, non-disruptive remediation
Not warranted / do NOT do now: forced raft compaction (log is stable, boots in 55s, old segments purge naturally once the log passes snapshot_index + trailing); any restart/reload. I deliberately did not even run
incus config set/incus image edit, because those are DB writes through the implicated raft path.Recommended, in a short planned maintenance window (priority order):
images.auto_update_intervalto0, or flip the 8 images toauto_update: false. Removes the DB-write flood and the failing startup task during any future isolation. These are DB writes — do them while someone watches the daemon (writes work fine now; low risk, but honor "watch it").skopeo inspectworks before any planned restart.waitreadymore patient) so a slow-but-progressing startup isn't SIGKILLed at 600s and turned into a destructive crash loop. Unit edit + daemon-reload = maintenance window.systemctl stop incus(lets dqlite checkpoint) + db backup will let the store consolidate over subsequent snapshots; explicit compaction is unnecessary.Not checked, to avoid risking the live daemon
db.binwith sqlite (no lock); used only daemon-mediated reads + filestat/ls.incus admin sql.Investigated read-only via the hydracluster exec/web-shell channels. pi-node-003 = node-6bf57aed, pi-node-004 = node-50ab5309.