HydraIssues

hydracluster will not start when hydradistrict is down, though it tolerates the same failure at runtime
open bug Project: hydracluster Reporter: cederik 6 Aug 2026 15:05

Description

hydracluster refuses to start if hydradistrict is unreachable, even though the same failure is tolerated at runtime. On 2026-08-06 this turned a dependency outage nobody had noticed into a full hydracluster outage, and the recovery required a hand-rolled stub because there is no way to start the service without the dependency.

The inconsistency

The initial fetch is fatal (internal/cli/cluster/serve.go:146-149):

// Initial district fetch
if err := srv.RefreshDistricts(); err != nil {
    return fmt.Errorf("initial district fetch: %w", err)
}
srv.StartDistrictRefresh(60 * time.Second)

The periodic refresh started on the very next line is not (pkg/api/server.go:216-226): it logs "district refresh error" and keeps the previous cache. Within RefreshDistricts the venue fetch is already non-fatal too — it logs a warning and continues (pkg/api/server.go:200-205); only the district List error returns.

So the same failure is fatal at second 0 and survivable at second 60. A fleet manager that already runs fine for hours with a dead hydradistrict should not refuse to boot against one.

What happened

hydracluster ran v2.0.98 from Aug 3 on a district cache fetched at startup. hydradistrict and hydravenues had since stopped (all scales on pi-node-003-nvme were down), but nothing surfaced it, because the cache only refreshes in place and a failed refresh keeps the old data.

A routine update to v2.0.99 restarted the service at 12:33 and the latent outage became a hard failure:

Error: initial district fetch: fetching districts: status 404
hydracluster.service: Main process exited, code=exited, status=1/FAILURE
hydracluster.service: Scheduled restart job, restart counter is at 27.

hydracluster was down until the dependency was worked around. Rolling back would not have helped: v2.0.98 has the identical fatal fetch, so any restart on any version would have crash-looped.

Why it is worse than one service being down

hydrascalerouter derives its routes from hydracluster's /api/v1/nodes, reading each node's reported scales. hydradistrict's own public route is published that way. So:

hydrascalerouter -> hydracluster -> hydradistrict -> (route published by) hydrascalerouter

With hydracluster down, the router cannot refresh and serves a cached route set. It survived only because that cache held. If the router had restarted, or evicted, the two services needed to bring hydracluster back would themselves have been unroutable, with no way in.

Recovery required pointing district_service_url and venue_service_url at a local stub returning [] so the process could boot, then reverting once pi-003 was fixed. That works, but it is a bare process a reboot would have killed, leaving the same crash loop with the workaround gone.

Suggested direction

Make the initial fetch non-fatal: log a warning, start with an empty cache, let the 60s refresh fill it in. That alone restores boot independence and matches the runtime behaviour.

Worth considering alongside:

  • Persist the last-known districts/venues to disk beside nodes.yaml, so a restart comes up with the previous catalogue instead of an empty one, and a dependency outage degrades to stale rather than empty.
  • Surface the state. Today a stale or empty district cache is invisible: the outage went unnoticed for days, and the empty cache is silent in the admin UI. A health field or admin banner showing districts last fetched successfully at T would have caught it.
  • Note the empty-cache side effect: with no districts, POST /api/v1/nodes/{id}/district skips its "unknown district" validation entirely (len(s.districtCache) > 0, pkg/api/handlers_api.go:603), and heads enrolled during the outage got no district/venue at all, because the iPad had nothing to choose from. Degraded mode should be explicit about what it stops validating.

Related

  • Startup also hard-requires district_service_url and venue_service_url to be set at all (serve.go:81-87). Reasonable, but it means there is no supported way to run without them, which is what forced the stub.
  • The pi-003 scales being down was the trigger, not the cause; that has since been fixed.