HydraIssues

hydracluster reprovision returns ok without re-running the recipe on a repeat call
open bug Project: hydracluster Reporter: 7 Aug 2026 21:23

Description

POST /api/v1/nodes/{id}/reprovision answers {"status":"ok"} with HTTP 200 and then does nothing, if the node was already provisioned recently. The caller has no way to tell the difference between "provisioned" and "ignored".

Found on 2026-08-06 while rolling hydraskin onto pi-node-004-nvme (node-50ab5309).

What happened

  1. First call, to move the node from hydraskin v0.8.0 to v0.10.0. Returned {"status":"ok"}, and it worked: /usr/local/bin/hydraskin was replaced, mtime Aug 6 17:10 CEST, hydraskin version reported v0.10.0.
  2. hydraskin v0.10.1 was released ~10 minutes later.
  3. Second call, same endpoint, same node. Returned {"status":"ok"}, HTTP 200.
  4. Nothing happened. Polled hydraskin version over the node for more than 20 minutes: still v0.10.0. The binary's mtime never moved off Aug 6 17:10, so the recipe's download-and-replace step never ran at all.
  5. Third call, later, same result.

The node was online throughout and hydracluster exec against it worked the whole time, so this is not the node being unreachable.

Why it matters

An endpoint that reports success while silently doing nothing is the worst shape for this. Someone rolling out a fix during an incident gets a green response, believes the node has the new binary, and moves on. The only way to catch it is to independently check the version afterwards, which defeats the point of having the endpoint.

If this is a deliberate debounce or cooldown, it needs to say so: 202 Accepted with a reason, or 429, or a body like {"status":"skipped","reason":"provisioned 8m ago"}. Silently swallowing the request is the part to fix, whether or not the throttle itself is correct.

Workaround used

Ran the recipe's own steps over hydracluster exec instead: resolve the version from releases.experiencenet.com/hydraskin/production/latest.json, download hydraskin-linux-arm64, move it into /usr/local/bin/. Same released artifact, just triggered directly.

Second, related problem

hydracluster nodes provision <id> from the CLI fails with:

Error: node "node-50ab5309" not found

The node plainly exists. nodes provision has no --server / --admin-token flags, unlike hydracluster exec, so it reads the local config and hits an empty local store (hydracluster nodes list prints "No nodes enrolled"). The documented CLI path for provisioning is unusable against the real cluster, which is why the API was called directly.

Either give nodes provision the same --server / --admin-token flags exec has, or make the config carry the admin token so both subcommands resolve the same cluster.

Suggested acceptance

  • A reprovision request that is not acted on must not return a bare success.
  • hydracluster nodes provision <id> works against the live cluster from an operator's workstation.
  • Reprovisioning twice in quick succession either both apply, or the second is refused in a way the caller can detect.

Context: this came out of the hydramancer image redeploy, issue #441.