HydraIssues

[CORRECTED: the Pis were fine] ad6 carries about 10 ms of jitter, on its own LAN as well as to the cloud
open unclassified Project: hydraprobe Parent: #714 Reporter: 14 Sep 2026 21:41

Description

Correction 2026-09-14

The original conclusion on this issue was wrong and is withdrawn. It claimed Raspberry Pi probe clients had an 8 to 14 ms instrument floor. They do not. The evidence was confounded: every Pi sat at ad6 and every Mac sat at cloud-seven or rupelmonde, so platform and venue could not be told apart. Two controls settle it.

Control 1: a Mac at ad6 measures what the Pi measures

cederikmini (macOS, 192.168.68.59, same ad6 LAN as the Pis, pacing error 0.062 ms, kernel-so-timestamp):

Client at ad6 Platform Pacing err p99 RTT p50 IPDV p99
pi-node-001 Pi, kernel-so-timestamping 0.767 ms 15.521 ms 10.386 ms
cederikmini Mac, kernel-so-timestamp 0.062 ms 19.336 ms 10.007 ms

Same venue, same reflector, minutes apart, two different platforms, same answer. The Mac is even slightly SLOWER on RTT. There is no Pi penalty.

Control 2: kernel timestamping changed nothing here

hydraprobe v0.2.0 moved arrival timestamps into the kernel (SO_TIMESTAMPING on Linux, SO_TIMESTAMP on Darwin). On a clean amd64 host that improved IPDV p99 by roughly 2x to 5x, so the change is real and is kept. At ad6 it moved the number from 10.5 to 10.4 ms. Userspace wake-up was never the cause.

What should have caught this sooner

rupelmonde's Mac measured 9.815 ms, the same magnitude as ad6's Pi. That number was already in hand and was scored as a real network failure while the identical ad6 number was scored as an instrument failure. That was inconsistent, and it is what made the wrong conclusion look supported.

What is actually true

ad6 carries about 10 ms of jitter and it is present inside its own LAN.

Path at ad6 Client RTT p50 IPDV p99
to GCP europe-west1 pi-node-001 15.521 ms 10.386 ms
to GCP europe-west1 cederikmini (Mac) 19.336 ms 10.007 ms
to pi-node-004, SAME LAN pi-node-001 0.870 ms 9.445 ms
to pi-node-003, SAME LAN pi-node-001 2.627 ms 10.350 ms
to pi-node-004, SAME LAN cederikmini (Mac) 4.598 ms 9.818 ms
to pi-node-004, reflector held busy pi-node-001 0.900 ms 9.964 ms

Sub-millisecond RTT with 9 to 10 ms of jitter cannot be the WAN. Every client and every reflector combination lands on the same figure, so it is not one machine either. The remaining suspect is the ad6 LAN infrastructure, the Omada gear most likely, and because all venue traffic crosses it the same jitter then rides out to the cloud.

Not yet isolated: which device. Worth checking switch port settings, any energy-efficient-ethernet or power-save feature, and whether the Mac path is wired or wireless (its 4.598 ms LAN RTT suggests wireless, which does not affect the GCP conclusion since both platforms agreed there).

Consequences

  • The Pis are usable probes. No probe replacement is needed. The PLAN 4.3 placement table stands.
  • ad6 is measurable and reaches the FLAT tier (RTT p50 19.336, p99 31.789, IPDV p99 10.007, zero loss, against flat limits of 30, 50 and 15 ms). It does not reach VR.
  • Calibration is still worth keeping, but as knowing your instrument rather than as a Pi ban: record rx_timestamp and pacing error on every run, and confirm a probe against a clean reference before trusting a surprising number. That practice is exactly what produced this correction.
  • ad6's 10 ms is now a venue engineering item, not a measurement artifact. Fixing the LAN could lift ad6 and is worth doing regardless of the cloud programme.

Related: #704, #706, #725, #402.