HydraIssues

hydranode updateService uses role name as systemd unit name — hydraskin updates fail silently, old process keeps running
open bug Project: hydranode Reporter: cederik 4 Aug 2026 11:14

Description

Observed 2026-08-04 on pi-node-004 (node-50ab5309). update-services requested by server triggers hydranode's Body.updateService(role) in pkg/body/service.go. It runs systemctl stop <name> / systemctl start <name> using the role name directly. For role "hydraskin" that unit does not exist: hydraskin's own installer (hydraskin/internal/install/install.go) names its systemd unit hydraskin-report.service, not hydraskin.service.

Journal from pi-node-004:
12:51:52 update-services requested by server
12:51:52 updating service hydraskin from https://releases.experiencenet.com/hydraskin/production/v0.8.0/hydraskin-linux-arm64
12:51:54 error restarting hydraskin: exit status 5 (exit 5 = unit not found)

The systemctl stop call used .Run() with the error discarded entirely, so the stop failure was completely silent; only the start failure logged, and only as a bare 'error restarting' with no indication of the cause.

Result: the v0.8.0 binary downloaded fine to /usr/local/bin/hydraskin, but the running hydraskin-report process (PID 939677) was never touched. readlink /proc/939677/exe showed /usr/local/bin/hydraskin (deleted) — running from an unlinked inode. hydraskin version reported v0.8.0 on disk while the live reporter kept executing v0.7.0 code. The node looked updated and was not. I hand-fixed this one node with systemctl restart hydraskin-report; the underlying defect was untouched and would recur on every hydraskin update on every node.

Fix implemented on branch fix/hydraskin-unit-name-mismatch in hydranode (not merged to default branch, not deployed): added serviceUnitName(role) next to the existing serviceBinaryName(role) in pkg/body/service.go, mapping hydraskin -> hydraskin-report (pass-through for every other role), and routed every systemctl call site — provisionService, deprovisionService, updateService, and writeServiceUnit's unit-file path — through it consistently. Also made the previously-silent stop failure in updateService loud: it now logs the error and combined output instead of discarding it.

Audit notes:

  • service_windows.go has its own updateService/provisionService/deprovisionService but does NOT share this defect: it already resolves role -> scheduled-task name via serviceTaskName()+taskNameOverrides consistently across all three lifecycle functions. No change needed there.
  • provisionService/deprovisionService never actually hit this bug in production because hydraskin is not in the role-provisioning switch in body.go (it has its own installer, not managed by hydranode's provision path) — but they carried the identical role-name-as-unit-name assumption, so the fix was applied there too for consistency and to guard against a future role with the same shape.
  • Added table-driven tests (service_test.go) covering serviceUnitName, including the pass-through case where role == unit, and a test asserting serviceUnitName and serviceBinaryName diverge independently (hydraskin diverges only in unit name; hydraguard-air/hydraneckwebrtc-controller diverge only in binary name).
  • go build ./..., go vet ./..., go test ./... all pass.

Follow-up not implemented: hydraskin already ships a BinaryStale check (hydraskin/internal/install/staleproc.go) that detects exactly this deleted-inode condition via /proc//exe, and hydraskin's own installer already wires it into ConfigureReporter's restart logic. hydranode's updateService has no equivalent post-restart verification for any service it manages — it trusts that a zero exit from systemctl start means the new binary is actually running. Worth considering a generic post-update BinaryStale-style check in hydranode's updateService for all managed services, not just hydraskin.