hydracluster refuses to start if hydradistrict is unreachable, even though the same failure is tolerated at runtime. On 2026-08-06 this turned a dependency outage nobody had noticed into a full hydracluster outage, and the recovery required a hand-rolled stub because there is no way to start the service without the dependency.
The initial fetch is fatal (internal/cli/cluster/serve.go:146-149):
// Initial district fetch
if err := srv.RefreshDistricts(); err != nil {
return fmt.Errorf("initial district fetch: %w", err)
}
srv.StartDistrictRefresh(60 * time.Second)
The periodic refresh started on the very next line is not (pkg/api/server.go:216-226): it logs "district refresh error" and keeps the previous cache. Within RefreshDistricts the venue fetch is already non-fatal too — it logs a warning and continues (pkg/api/server.go:200-205); only the district List error returns.
So the same failure is fatal at second 0 and survivable at second 60. A fleet manager that already runs fine for hours with a dead hydradistrict should not refuse to boot against one.
hydracluster ran v2.0.98 from Aug 3 on a district cache fetched at startup. hydradistrict and hydravenues had since stopped (all scales on pi-node-003-nvme were down), but nothing surfaced it, because the cache only refreshes in place and a failed refresh keeps the old data.
A routine update to v2.0.99 restarted the service at 12:33 and the latent outage became a hard failure:
Error: initial district fetch: fetching districts: status 404
hydracluster.service: Main process exited, code=exited, status=1/FAILURE
hydracluster.service: Scheduled restart job, restart counter is at 27.
hydracluster was down until the dependency was worked around. Rolling back would not have helped: v2.0.98 has the identical fatal fetch, so any restart on any version would have crash-looped.
hydrascalerouter derives its routes from hydracluster's /api/v1/nodes, reading each node's reported scales. hydradistrict's own public route is published that way. So:
hydrascalerouter -> hydracluster -> hydradistrict -> (route published by) hydrascalerouter
With hydracluster down, the router cannot refresh and serves a cached route set. It survived only because that cache held. If the router had restarted, or evicted, the two services needed to bring hydracluster back would themselves have been unroutable, with no way in.
Recovery required pointing district_service_url and venue_service_url at a local stub returning [] so the process could boot, then reverting once pi-003 was fixed. That works, but it is a bare process a reboot would have killed, leaving the same crash loop with the workaround gone.
Make the initial fetch non-fatal: log a warning, start with an empty cache, let the 60s refresh fill it in. That alone restores boot independence and matches the runtime behaviour.
Worth considering alongside:
len(s.districtCache) > 0, pkg/api/handlers_api.go:603), and heads enrolled during the outage got no district/venue at all, because the iPad had nothing to choose from. Degraded mode should be explicit about what it stops validating.