HydraIssues

hydracluster: bundle/embed role recipes with the release so they cannot drift from the binary
open bug Priority: high Project: hydracluster Reporter: 19 Aug 2026 20:06

Description

Role recipes load from a directory on the hydracluster server (/recipes = /root/.hydracluster/recipes) at startup via recipe.Load, but the deploy/update flow only replaces the /usr/local/bin/hydracluster binary. The recipes/ directory is never synced, so it silently drifts from the code the binary expects.

Impact (observed 2026-08-19): the on-disk hydraskin-linux.yaml was a stale version referencing the template var .SkinBridgeSubnet, removed from recipe.Vars in v2.0.105. recipe.Render then errored on every hydraskin node, the provision handler skipped the role, and GET /api/v1/body/provision returned {"roles":null}. Result: NO hydraskin node could auto-provision, fleet-wide and silently. A second drift (verify checked hydraskin-report.timer, but the reporter is now a .service) compounded it. Both fixed by hand-copying the v2.0.105 recipe to the server + restart, plus fixes committed to main - but the drift recurs on the next release.

Proposed fix: ship recipes WITH the release so they cannot drift from the binary. Preferred approach (matches the fleet 'on-disk content COPY'd into the image, not go:embed' convention): bundle recipes/ into the release artifact and have the deploy/update sync it into /recipes on every update, so the recipe set is versioned and shipped alongside the binary rather than hand-managed on the server. Also add a startup check that fails loudly (or flags health-degraded) if a known role's recipe fails to render, so a future drift is visible instead of silent.