HydraIssues

HydraGuard air provision should not regenerate keypair for existing peers
closed bug Project: hydraguard Reporter: 23 Mar 2026 14:15

Description

Problem

POST /api/v1/air/provision regenerates the WireGuard keypair every time it is called, even for peers that already exist in mesh.yaml. Combined with auto_apply: true, this silently breaks WireGuard connectivity for existing peers.

This caused repeated key mismatches during the Brussels hub migration — every hub apply or hydracluster provision call triggered a new keypair.

Fix

Make the endpoint truly idempotent: if the peer already exists in mesh.yaml, return the existing config (with the existing private key) instead of generating a new keypair. Only generate a new keypair when the peer is genuinely new.

Files

  • Hydraguard repo: air provision handler

Comments (1)

api 3 Aug 2026 21:03

Fixed in hydraguard v2.3.0, deployed and verified on the hub

/api/v1/air/provision and /api/v1/headipad/provision generated a keypair before checking whether the peer existed, then overwrote its stored public key. With auto_apply: true the new hub config applied immediately, so a call that read like a lookup silently disconnected a peer still holding the old private key.

The fix proposed here was not implementable

The issue asks to "return the existing config (with the existing private key)". That cannot be done: mesh.Air stores only PublicKey. The private key is shown once at creation and never persisted — by design — so there is nothing to hand back.

The closest correct behaviour is to refuse, and say how to proceed deliberately:

  • an existing peer now returns 409 and is left untouched
  • re-keying requires an explicit "force": true, logged as a warning since the device is disconnected until updated
  • key generation moved after the existence check, so nothing is generated unless needed

Same change applied to the head-iPad endpoint, which had the identical bug.

Verified on the live hub

POST /api/v1/air/provision {"name":"chunky-turnip-23"}
HTTP 409
{"error":"air peer \"chunky-turnip-23\" already exists on 10.10.100.16/32; its private
 key was shown only at creation and is not stored. Re-send with \"force\": true to issue
 a new keypair, which will disconnect the peer until the device is updated."}

  key UNCHANGED — peer not disconnected
  wg peers: 24

The API service no longer needs to be kept stopped for this reason

The runbook's standing instruction to keep hydraguard stopped existed because of this bug. That warning has been replaced.

Two things found while fixing it

keygen.Generate was untestable. It shells out to wg genkey, so provisioning could not be tested anywhere without wireguard-tools installed. Now overridable via SetGenerator/ResetGenerator, matching the existing wireguard.SetApplier convention. The regression tests fail against the old behaviour with "the peer's key changed on a call that should have been a no-op".

The runbook's mesh table was wrong and caused a real mistake. It listed peers as air-001, air-tvl-one and so on; the actual IDs are bare (001, tvl-one). Trusting it, I provisioned air-001 against the live hub while testing this fix — which correctly found no such peer and created a new one at 10.10.100.17. Removed and reapplied; the mesh is back to 24 peers. The table has been regenerated from the live mesh, with the command to regenerate it so the next update is a paste rather than a transcription.