HydraIssues

Pulls through incusd fail 'unauthorized' while the same skopeo command succeeds by hand — blocks launching new scales
closed bug Priority: high Project: hydraskin Reporter: 2 Aug 2026 04:27

Description

Symptom

Launching a new scale from scaleregistry.experiencenet.com fails:

Error: Failed to run: skopeo --insecure-policy inspect \
  docker://scaleregistry.experiencenet.com/hydranps:v0.2.3 --no-tags: exit status 1
  (level=fatal msg="... reading manifest v0.2.3 in scaleregistry.experiencenet.com/hydranps: unauthorized")

The identical command run by hand on the same box succeeds:

REGISTRY_AUTH_FILE=/etc/hydraskin/registry-auth.json skopeo --insecure-policy \
  inspect docker://scaleregistry.experiencenet.com/hydranps:v0.2.3 --no-tags
# -> {"Name": "scaleregistry.experiencenet.com/hydranps", "Digest": "sha256:723257a4..."}

Impact: new scales cannot be launched. Already-pulled images and running scales are unaffected — hydranps on pi-node-003 keeps serving normally throughout.

What has been ruled out

Verified on pi-node-003 and pi-node-004 (2026-08-02). The two nodes are byte-identical in every dimension checked, and both fail:

pi-node-003 pi-node-004
incusd HOME /var/lib/incus/ /var/lib/incus/
incusd REGISTRY_AUTH_FILE (live, from /proc/<pid>/environ) set correctly set correctly
skopeo version 1.13.3 1.13.3
credential sha256 424eb4859168 424eb4859168
other auth files on disk none none
manual skopeo works works
pull via incusd fails fails

Also tried, no effect:

  • Placing the credential at incusd's $HOME/.config/containers/auth.json
  • Removing and re-adding the OCI remote after the credential was in place
  • Restarting incus (the drop-in is applied — the variable is confirmed in the live process environment)
  • Different tags (v0.2.3, latest) — both fail via incusd, both succeed manually

So the credential is valid, the file is readable, and the environment variable is present in incusd. Something between incusd and the skopeo child is losing the authentication.

The unexplained part

One pull succeeded. On 2026-08-01, incus launch scaleregistry:hydranps:v0.2.3 on pi-node-003 pulled the image successfully — that is how the running scale exists. The same command on the same node now fails. Nothing about the node's auth configuration changed in between; the credential file is unchanged and still valid.

That inconsistency is the most useful clue and should not be dismissed. Anything that explains only the failures but not the one success is probably the wrong explanation.

First action: the registry has no request logging

hydrascaleregistry logs only startup lines and TLS handshake errors. There is no per-request logging at all, so during these failures the registry logs were silent — which proves nothing, and I initially over-read it as "the request never arrived".

Adding request logging (method, path, auth outcome, source IP) is cheap and is the difference between guessing on the client side and seeing what the server actually decided. Do this first — it likely resolves the diagnosis in one attempt.

Worth logging specifically: whether a request arrived with no Authorization header at all (incusd stripping it) versus arriving with credentials that were rejected. Those point at completely different fixes.

Notes on the documentation

The hydraskin runbook presented the REGISTRY_AUTH_FILE systemd drop-in as the fix for a bare unauthorized. That has been corrected — the drop-in is still required but is demonstrably not sufficient. See docs/runbooks/hydraskin.md.

Done when

  • hydrascaleregistry logs requests with auth outcome
  • The failure is explained, including why the 2026-08-01 pull succeeded
  • A new scale can be launched on pi-node-004 from the registry
  • The fix is in the hydraskin recipe, not applied by hand, so new nodes get it
  • Runbook updated with the actual cause

Comments (1)

api 3 Aug 2026 18:10

Resolved (2026-08-03): the credential was not where skopeo looks

Cause: incusd sanitises the environment of the skopeo processes it spawns, so REGISTRY_AUTH_FILE set on the incus service never reached skopeo. skopeo then fell back to containers/image's default location for root — /run/containers/0/auth.json — found nothing there, and returned a bare unauthorized.

That is why this survived several rounds of diagnosis. Every check pointed at the credential being fine, and it was:

  • the variable was confirmed present in incusd's live environment via /proc/<pid>/environ
  • the credential file was valid and byte-identical across nodes
  • the identical skopeo inspect succeeded when run by hand — because a shell has the environment incusd strips

So the symptom reads as an authentication problem and is actually a path problem. Placing the credential at the runtime path made the pull work immediately, on an image that had failed minutes earlier.

Why the 2026-08-01 pull succeeded

Still not proven, but now explicable: anything that had written /run/containers/0/auth.json that day — a skopeo login, or another tool using containers/image — would have made pulls work until the next reboot cleared /run. That fits the one-off success followed by consistent failure, where nothing about the configuration had changed in between.

Fix

In hydraskin install as of v0.6.0. /run is tmpfs, so the credential is re-placed on every incus start:

# /etc/systemd/system/incus.service.d/10-registry-auth.conf
[Service]
Environment=REGISTRY_AUTH_FILE=/etc/hydraskin/registry-auth.json
ExecStartPre=/bin/sh -c 'mkdir -p /run/containers/0 && if [ -f /etc/hydraskin/registry-auth.json ]; then cp /etc/hydraskin/registry-auth.json /run/containers/0/auth.json && chmod 600 /run/containers/0/auth.json; fi'

The environment variable is kept: it is correct, costs nothing, and would take effect if Incus stops filtering. The if [ -f ] guard means a node without credentials still starts incus.

Verified

On pi-node-004, which had failed every pull all day:

hydraskin install  ->  registry credential placed at /run/containers/0/auth.json
incus image copy scaleregistry:hydravenues:v0.5.6      ->  Image copied successfully
systemctl restart incus                                ->  /run credential re-placed (0600)
incus image copy scaleregistry:hydraorganization:v0.2.6 -> Image copied successfully

Both before and after a restart, so the tmpfs path is genuinely covered rather than working by accident.

Follow-up still worth doing

The request logging from #427 remains worth adding. It would have identified this in one attempt by showing the request arriving with no Authorization header — distinguishing "credential not sent" from "credential rejected", which is exactly the distinction that was missing.

Runbook and hydraskin/CLAUDE.md updated with the mechanism rather than the symptom.