Launching a new scale from scaleregistry.experiencenet.com fails:
Error: Failed to run: skopeo --insecure-policy inspect \
docker://scaleregistry.experiencenet.com/hydranps:v0.2.3 --no-tags: exit status 1
(level=fatal msg="... reading manifest v0.2.3 in scaleregistry.experiencenet.com/hydranps: unauthorized")
The identical command run by hand on the same box succeeds:
REGISTRY_AUTH_FILE=/etc/hydraskin/registry-auth.json skopeo --insecure-policy \
inspect docker://scaleregistry.experiencenet.com/hydranps:v0.2.3 --no-tags
# -> {"Name": "scaleregistry.experiencenet.com/hydranps", "Digest": "sha256:723257a4..."}
Impact: new scales cannot be launched. Already-pulled images and running scales are unaffected — hydranps on pi-node-003 keeps serving normally throughout.
Verified on pi-node-003 and pi-node-004 (2026-08-02). The two nodes are byte-identical in every dimension checked, and both fail:
| pi-node-003 | pi-node-004 | |
|---|---|---|
incusd HOME |
/var/lib/incus/ |
/var/lib/incus/ |
incusd REGISTRY_AUTH_FILE (live, from /proc/<pid>/environ) |
set correctly | set correctly |
| skopeo version | 1.13.3 | 1.13.3 |
| credential sha256 | 424eb4859168 |
424eb4859168 |
| other auth files on disk | none | none |
| manual skopeo | works | works |
| pull via incusd | fails | fails |
Also tried, no effect:
$HOME/.config/containers/auth.jsonv0.2.3, latest) — both fail via incusd, both succeed manuallySo the credential is valid, the file is readable, and the environment variable is present in incusd. Something between incusd and the skopeo child is losing the authentication.
One pull succeeded. On 2026-08-01, incus launch scaleregistry:hydranps:v0.2.3 on pi-node-003 pulled the image successfully — that is how the running scale exists. The same command on the same node now fails. Nothing about the node's auth configuration changed in between; the credential file is unchanged and still valid.
That inconsistency is the most useful clue and should not be dismissed. Anything that explains only the failures but not the one success is probably the wrong explanation.
hydrascaleregistry logs only startup lines and TLS handshake errors. There is no per-request logging at all, so during these failures the registry logs were silent — which proves nothing, and I initially over-read it as "the request never arrived".
Adding request logging (method, path, auth outcome, source IP) is cheap and is the difference between guessing on the client side and seeing what the server actually decided. Do this first — it likely resolves the diagnosis in one attempt.
Worth logging specifically: whether a request arrived with no Authorization header at all (incusd stripping it) versus arriving with credentials that were rejected. Those point at completely different fixes.
The hydraskin runbook presented the REGISTRY_AUTH_FILE systemd drop-in as the fix for a bare unauthorized. That has been corrected — the drop-in is still required but is demonstrably not sufficient. See docs/runbooks/hydraskin.md.
hydrascaleregistry logs requests with auth outcome
Resolved (2026-08-03): the credential was not where skopeo looks
Cause: incusd sanitises the environment of the skopeo processes it spawns, so
REGISTRY_AUTH_FILEset on the incus service never reached skopeo. skopeo then fell back to containers/image's default location for root —/run/containers/0/auth.json— found nothing there, and returned a bareunauthorized.That is why this survived several rounds of diagnosis. Every check pointed at the credential being fine, and it was:
/proc/<pid>/environskopeo inspectsucceeded when run by hand — because a shell has the environment incusd stripsSo the symptom reads as an authentication problem and is actually a path problem. Placing the credential at the runtime path made the pull work immediately, on an image that had failed minutes earlier.
Why the 2026-08-01 pull succeeded
Still not proven, but now explicable: anything that had written
/run/containers/0/auth.jsonthat day — askopeo login, or another tool using containers/image — would have made pulls work until the next reboot cleared/run. That fits the one-off success followed by consistent failure, where nothing about the configuration had changed in between.Fix
In
hydraskin installas of v0.6.0./runis tmpfs, so the credential is re-placed on every incus start:The environment variable is kept: it is correct, costs nothing, and would take effect if Incus stops filtering. The
if [ -f ]guard means a node without credentials still starts incus.Verified
On pi-node-004, which had failed every pull all day:
Both before and after a restart, so the tmpfs path is genuinely covered rather than working by accident.
Follow-up still worth doing
The request logging from #427 remains worth adding. It would have identified this in one attempt by showing the request arriving with no
Authorizationheader — distinguishing "credential not sent" from "credential rejected", which is exactly the distinction that was missing.Runbook and
hydraskin/CLAUDE.mdupdated with the mechanism rather than the symptom.