HydraIssues

pi-node-003: incusd deadlocks on startup (futex_wait_queue), 8 production domains down
open bug Project: hydraskin Reporter: 6 Aug 2026 11:54

Comments (2)

api 7 Aug 2026 21:53

Root-cause post-mortem — pi-node-003 incusd startup deadlock (RESOLVED)

Non-disruptive investigation. Daemon is healthy now (incus active, 8/8 containers RUNNING, uptime since 2026-08-06 15:25 CEST). No restart/reload/config-write/container action was taken. Read-only: journal, ls/stat on the raft dir, daemon-mediated reads (incus list/image show/warning list), and comparison vs healthy pi-node-004. Live db.bin was never opened/locked.

Root cause (best-supported, with the one honest caveat below)

Trigger = network isolation during the LAN move. scaleregistry.experiencenet.com became DNS-unresolvable — journal shows dial tcp: lookup scaleregistry.experiencenet.com on 127.0.0.53:53: server misbehaving (systemd-resolved stub, upstream unreachable). 8 of 9 images are auto_update: true and images.auto_update_interval is unset (default 6h), so image auto-update runs at/after startup and periodically. With the registry unreachable, every refresh failed: 248 "Failed to update the image" errors, 00:16–09:14Z on 08-06 — matching the deadlock window.

Incus persists refresh failures as DB-backed warnings; those writes go through the single-node cowsql/dqlite raft log. That write activity, plus 52 daemon restarts each SIGKILLed after the 600s start-post readiness probe timed out (Error: Daemon still not running after 600s timeout, Main process ... code=killed status=9/KILL, unclean termination of a previous run ×32), inflated the raft store. Each unclean SIGKILL finalizes the current open segment prematurely, so the crash loop minted many short segments.

The reported symptom (reads work, writes hang, wchan=futex_wait_queue, activating forever) is the raft commit path stalling: reads are served from local sqlite state and keep working; writes wait on a raft apply that never completes → the daemon never signals ready → systemd start-post kills it at 600s → repeat. In early cycles the daemon reached the image task (logged failures); in later cycles (09:14–15:24, 21 more restart attempts) it logged no image failures and simply hung before readiness — consistent with the hang sitting in dqlite/raft startup itself.

Recovery at 15:25Z was clean and fast (55s) once the network/registry came back — replaying the same on-disk log without trouble.

Honest caveat: because that identical 75-segment / ~2718-entry log booted cleanly in 55s the moment the network returned, raft-log size is not independently sufficient to cause the deadlock and is not a standing boot-time hazard — the deadlock required the isolation condition. I could not observe the futex holder directly (incident is resolved; enabling raft-debug/attaching a debugger to the live daemon would risk it), so the exact internal lock is inferred from the read-works/write-hangs signature, which is characteristic of the dqlite raft commit path.

Controlled comparison — pi-node-004 (healthy)

Same Incus 7.2 (aarch64), same auto_update config, same raft trailing (segments retained back to index 1), and it also restarted on 08-06 (09:35) — but it had network, saw 0 image failures, restarted once, and sits at 9 closed segments (index 1→2049, 2 snapshots). The only variable that differs is network reachability during the restart. This is the strongest evidence the trigger was isolation, not anything intrinsic to 003's store.

Answers to the 5 leads

  1. Segment count now / growth: 75 closed + 3 open (3 open is normal — 004 also has 3). Growth was entirely incident-era: 55 of 75 closed segments are dated 08-06, 15 on 08-05, ~5 before. Zero new closed segments since recovery (~41h) — the log is stable at rest, not growing. The 58→75 delta is all 08-06, not ongoing.
  2. Snapshots / gap: 003 has 2 snapshots (idx 1024 @08-04, idx 2048 @08-06 05:02Z) — the standard dqlite trailing-of-2 (the earlier "4" counted the .meta files). 004 is structurally identical. The ~332-entries-past-snapshot gap is normal (dqlite snapshots ~every 1024 entries). The many segments on 003 come from 52 unclean restarts finalizing short segments (avg ~36 entries/seg vs 004's ~228/seg), not from snapshotting. A large un-snapshotted log does not explain the stall — the log isn't large and replays in 55s.
  3. Known cowsql futex/replay bug: not supported by this evidence. 75 segments is modest and the same log replays fine; futex_wait_queue is a generic Go/cgo lock wait. I am flagging this as not the mechanism rather than confirming a bug. Deeper confirmation would need raft-debug logging in a maintenance window (unsafe now).
  4. 248 image failures ↔ deadlock: strong correlation. Failures 00:16–09:14Z sit inside the window; the mid-incident snapshot (05:02Z) and 55 same-day segments line up with a write flood. Image auto-refresh does write to the DB on failure (warnings are raft-persisted) and it runs at startup — so during isolation it both bloats the log and adds a failing startup task. Confirmed 8/9 images auto_update: true, interval default 6h.
  5. Recurrence — CONDITIONAL YES. At rest the store is stable (no growth in 41h) and is not a latent time bomb by itself. But the ingredients remain: 8 auto_update images + default 6h interval + hard dependency on scaleregistry reachability. If the node reboots/restarts while network-isolated (another LAN move, DNS or registry outage), the same failure-loop + start-post-timeout dynamic can recur. Risk is conditional on (restart) AND (isolation) coinciding.

Safe, non-disruptive remediation

Not warranted / do NOT do now: forced raft compaction (log is stable, boots in 55s, old segments purge naturally once the log passes snapshot_index + trailing); any restart/reload. I deliberately did not even run incus config set/incus image edit, because those are DB writes through the implicated raft path.

Recommended, in a short planned maintenance window (priority order):

  1. Break the trigger dependency (highest value): set images.auto_update_interval to 0, or flip the 8 images to auto_update: false. Removes the DB-write flood and the failing startup task during any future isolation. These are DB writes — do them while someone watches the daemon (writes work fine now; low risk, but honor "watch it").
  2. Remove the single point of failure for registry name resolution: static hosts entry / local DNS fallback / pull-through cache for scaleregistry.experiencenet.com so a LAN move can't make it "server misbehaving." Operational rule: never reboot pi-node-003 while it can't reach scaleregistry — confirm skopeo inspect works before any planned restart.
  3. Raise the systemd start-post readiness timeout (or make waitready more patient) so a slow-but-progressing startup isn't SIGKILLed at 600s and turned into a destructive crash loop. Unit edit + daemon-reload = maintenance window.
  4. Optional cosmetic: a clean systemctl stop incus (lets dqlite checkpoint) + db backup will let the store consolidate over subsequent snapshots; explicit compaction is unnecessary.

Not checked, to avoid risking the live daemon

  • Did not enable raft/dqlite debug logging or attach a debugger to observe the futex holder.
  • Did not open/copy live db.bin with sqlite (no lock); used only daemon-mediated reads + file stat/ls.
  • Did not run incus admin sql.
  • Did not test-restart to measure current boot time (the whole point is not to restart).

Investigated read-only via the hydracluster exec/web-shell channels. pi-node-003 = node-6bf57aed, pi-node-004 = node-50ab5309.

cederik 9 Aug 2026 14:39

#445 short-term mitigations applied — 2026-08-09

Applied the three low-risk mitigations to all three live Pi nodes, one at a time, strictly non-disruptively: no incusd restart, no reboot, no container stop/start. Order followed: pi-node-001 → pi-node-004 → pi-node-003. incusd was active on every node before and after.

Mitigations

  • A. Disable image auto-update (core fix). incus config set images.auto_update_interval 0 at server level (space-syntax printed the expected deprecation warning but the value took). Verified with incus config get images.auto_update_interval → 0 on each node. Prior value on all three: empty/unset (Incus default ≈6h). Stops the periodic refresh loop without restarting the daemon or touching containers.
  • B. Raise systemd start readiness timeout. Created drop-in /etc/systemd/system/incus.service.d/10-start-timeout.conf = [Service] / TimeoutStartSec=1800 (30 min), then systemctl daemon-reload (rc=0). daemon-reload only — incus was NOT restarted; the new timeout applies to the next start.
  • C. /etc/hosts registry fallback. Appended 159.69.93.219 scaleregistry.experiencenet.com (releasesandscaleregistry) where not already present. NOTE: this pins the registry IP — if releasesandscaleregistry ever changes address, this line must be updated/removed on all three nodes.

Per-node results (A+B+C applied to each; healthy before & after)

Node A interval→0 B drop-in C hosts incusd before→after domain spot-check after
pi-node-001 (node-2b224f6a) applied, get=0 applied, reload rc=0 applied active → active test scales only, none district-routed
pi-node-004 (node-50ab5309) applied, get=0 applied, reload rc=0 applied active → active issues 200, books 200, hydramancer 200
pi-node-003 (node-6bf57aed) applied, get=0 applied, reload rc=0 applied active → active hydrapipeline 200, hydravenues 200, hydradistrict 200, hydraperforce 302 (baseline redirect, unchanged)

District router health (http://127.0.0.1:8099/api/v1/health)

status: ok, routes: 19, traefik: running, backends_reachable: 19 — checked before and after every node. 19/19 backends reachable throughout; no route dropped, no degradation, no revert needed on any node.

Before-state recorded (for A revert if ever required)

On all three nodes prior to change: images.auto_update_interval was empty/unset (Incus default); no pre-existing 10-start-timeout drop-in; no pre-existing /etc/hosts entry for scaleregistry. (Existing 10-registry-auth.conf drop-in left untouched.)

Still open (do not close #445)

incus-upgrade #455 and the deeper fix (hard registry dependency + short readiness contract) remain outstanding. Leaving #445 open.

— applied via hydracluster exec / web-shell (root), operator cederik