HydraIssues

Dashboard sign-in fails on tablet and cannot be reproduced server side
done bug Project: hydraorganization Parent: #589 Reporter: cederik 1 Sep 2026 22:43

Description

SYMPTOM

The org dashboard at https://hydraorganization.experiencenet.com/dashboard does not complete sign-in on the owner's tablet. Reported as "not loading" and, after two fixes, still failing. It has not been reproduced from a workstation or from headless Chrome.

This issue exists so the next person does not repeat the elimination below.

VERIFIED WORKING (checked live against production, hydraorganization v0.7.3)

  • Service health 200, no errors in the container log beyond startup.
  • GET /dashboard with no session: 303 to /dashboard/login.
  • GET /dashboard/login: 303 to https://iamnim.com/login with redirect_uri=https://hydraorganization.experiencenet.com/dashboard/authed.
  • The live iamnim login page renders and carries redirect_uri on BOTH the Google button and the magic-link form (checked in the served HTML, not just the repo).
  • iamnim's redirect allowlist accepts .experiencenet.com, so the return leg is permitted.
  • GET /dashboard/authed?token=X: sets iamnim_session HttpOnly Secure SameSite=Lax, sets Referrer-Policy no-referrer, redirects with no token in the Location.
  • The scale can reach iamnim: a junk session cookie is answered 401 by iamnim (a network failure would surface as 502 instead).
  • Cold dashboard render with all four upstreams: 0.5s. Not a timeout.
  • Membership exists: user u_717adf7feb08 (cederik@cederik.com) is a member of visit-flanders and experiencenet in Pantheon realm r_349832f91720.

TWO REAL BUGS FOUND AND FIXED WHILE HUNTING THIS (both presented identically as "page never loads" with NOTHING in the log)

  • v0.7.1, loop 1: a session iamnim REJECTS caused dashboard -> login -> iamnim -> dashboard forever. Reachable in normal operation because iamnim keeps sessions in memory on a single replica, so a restart invalidates every session it issued.
  • v0.7.2, loop 2: the cookie never arriving at all. With no cookie there was no way to distinguish a first visit from a failed sign-in, so it bounced to login forever. /dashboard/authed now redirects to /dashboard?signedin=1; that marker carries no secret and exists only to separate the two cases.

Neither fix resolved the tablet.

INSTRUMENTATION ADDED (v0.7.3)

The auth path now logs, with no secrets and no query strings:

  • auth: return from iamnim, token_present= ua=
  • auth: dashboard hit, cookie_present= just_signed_in= ua=<...>
  • auth: iamnim REJECTED the session cookie (just_signed_in=)
  • auth: signed in as with membership(s)

Read it with: incus console hydraorganization --show-log, on pi-node-004-nvme (node-50ab5309) over the hydracluster exec transport.

NEXT STEP: have the tablet attempt one sign-in, note the time, then read the log. The four lines above identify the failing leg exactly:

  • no "return from iamnim" line at all: the tablet never got back from iamnim, so the problem is upstream (Google sign-in, or the redirect being dropped).
  • return with token_present=false: iamnim redirected back without a session.
  • return with token_present=true, then dashboard hit with cookie_present=false: the tablet refused the cookie. This is the Safari and iPadOS case; the page now says so instead of looping.
  • REJECTED line: iamnim issued a session it then would not accept, most likely an iamnim restart between the two calls.

STILL UNKNOWN

Which tablet and browser. iPad Safari is the leading suspect because it restricts cookies written during a cross-site navigation, but this has not been confirmed. If it is Safari, evaluate whether the session can be established without a cross-site cookie write, for example by having iamnim POST back rather than redirect, or by completing the exchange server side (#591's one-time code makes that possible; the code exchange alone does NOT fix a browser that refuses cookies).

RELATED: #596 (auth), #591 (one-time code handoff), #589 (master).

Comments (3)

claude 2 Sep 2026 08:59

MAJOR NARROWING 2026-09-02. The tablet request NEVER REACHES THE SERVICE.

Evidence: hydraorganization has been up since 2026-09-01 22:42 UTC with no restart (uptime 36931s at the time of checking). The persisted console log at /var/log/incus/hydraorganization/console.log on pi-node-004-nvme contains exactly two entries after startup, both from my own synthetic tests at 22:42:24 and 22:42:25 with user agent TestAgent/1.0. The owner reported trying on the tablet around 08:55 UTC on 2026-09-02. There is NO log line for it.

The log is verified live, not stale: a probe request at 08:58:49 with a unique user agent appeared in the file immediately. So the absence is real, not a logging gap. Every /dashboard hit writes a line before any redirect decision, so a request that arrived would be recorded even if it then bounced.

RULED OUT by this: all the server-side auth logic, both redirect loops, iamnim reachability, membership, and the aggregation path. None of it runs if the request does not arrive.

ALSO CHECKED: TLS is not the difference. hydraorganization.experiencenet.com serves CN=hydraorganization.experiencenet.com -> Lets Encrypt YR1 -> ISRG Root YR, verify return code 0, the same chain shape and the same root as issues.experiencenet.com. DNS resolves to the edge 141.227.136.199.

REMAINING CANDIDATES, in order of likelihood:

  1. The tablet is rendering a CACHED page and issuing no request at all. The dashboard only started sending no-store at v0.6.0; anything the tablet stored before that can still be served from its cache. This exactly matches a phone-side problem seen earlier in this work, where a stale cached build made fixes look like they had not shipped.
  2. The tablet has no route to 141.227.136.199, or its DNS does not resolve the name. Relevant because the iPads normally live on venue LANs and reach the mesh through a hydraneck rather than the open internet.
  3. The tablet never navigated: a kiosk lock, or a tab that failed earlier and was never reloaded.

TWO TESTS TO SEPARATE THESE, both on the tablet:
A. Open https://hydraorganization.experiencenet.com/api/v1/health . Plain JSON, no auth, no redirect. If JSON appears the tablet can reach the service and candidate 2 is out. If it errors, the error text names the cause.
B. Open https://hydraorganization.experiencenet.com/dashboard?cb=1 . A different URL is a different cache entry, so it cannot be served from a stale cache. If this works while /dashboard does not, it is candidate 1 and the fix is clearing that site data.

Either result will show in the log for A and B, so the answer is checkable from both ends.

claude 5 Sep 2026 10:00

CORRECTION AND LIKELY CAUSE, 2026-09-05. The device is a NEW ANDROID TABLET, not an iPad. That kills both earlier theories: a new device has no cache to be stale, and Android is not subject to the Safari cross-site cookie restrictions.

With those gone, the JSON root is the strongest remaining explanation and it fits every observation:

GET / returned {"docs":"/api/v1/health","service":"hydraorganization","version":"..."} to everyone, including browsers. On a new device a person types the bare domain, not /dashboard. They would land on a wall of JSON, which reads exactly as "not working". Critically, NOTHING on the root path was logged, only /dashboard was instrumented, so such a visit produces the silent log I reported earlier. The absence of a log line was evidence of the wrong path being hit, not of the request never arriving.

FIXED in v0.7.4: the root now redirects a browser to /dashboard and keeps the JSON service description for callers asking for it, decided on the Accept header. The redirect is logged, so any future root visit is visible.

Verified live end to end as a browser would do it: http://hydraorganization.experiencenet.com -> 301 https -> 303 /dashboard -> 303 /dashboard/login -> the iamnim sign-in page, HTTP 200. An API client with no Accept header still receives the service JSON.

STILL UNCONFIRMED: whether the owner actually typed the bare domain. This is a hypothesis that fits the evidence, not a reproduction. If the tablet still fails after v0.7.4, the remaining candidates are a mistyped host, no route from that network to the edge 141.227.136.199, or a wrong device clock breaking TLS on a new tablet. All three now produce a visible difference: a root hit logs a line, so if the next attempt still logs nothing the request genuinely is not arriving and the cause is network or DNS rather than anything in this service.

NOTE: traefik at the edge has no access logging (13 lifecycle entries in 3 days), so there is no per-request record one layer up. Enabling it would have answered this in one step and is worth considering.

claude 14 Sep 2026 10:29

CONFIRMED FIXED and closing. The owner reports it working, and the log carries the proof from both ends.

The Android tablet, 2026-09-09 09:50:00, user agent Mozilla/5.0 (Linux; Android 15; Pixel 9) Chrome/135: "auth: root hit by a browser, sending to the dashboard", followed immediately by a dashboard hit. That is the v0.7.4 redirect doing its job on the exact device that was failing. Before v0.7.4 that same visit hit the JSON root, rendered a blob, and logged nothing, which is why the earlier investigation saw silence and wrongly concluded the request was never arriving.

A full successful sign-in is also recorded, 2026-09-14 10:25:52: root hit -> dashboard -> return from iamnim with token_present=true -> dashboard hit with cookie_present=true just_signed_in=true -> "signed in as cederik@cederik.com with 4 membership(s)". Every leg of the chain visible in one trace.

ROOT CAUSE: GET / served a JSON service description to browsers, and that path was not instrumented. Someone typing the bare domain, which is what you do on a new device, got a wall of JSON and left no trace.

THE MISTAKE WORTH KEEPING: I instrumented the path I designed (/dashboard) rather than the path a person actually types (/). The resulting silence in the log was then read as evidence the request never arrived, which sent the investigation after network, DNS and TLS for two rounds. Instrument the entry point a human uses.

Fixes that came out of this hunt and are worth keeping regardless: two genuine sign-in redirect loops (v0.7.1 a session iamnim rejects, v0.7.2 a browser that refuses the cookie), the auth trace itself (v0.7.3), and the root redirect (v0.7.4).