From c5f379b3bb0f96fd4f4c2ed20662277cf489e03d Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Sun, 4 Oct 2026 20:31:44 +0000 Subject: [PATCH] bugs.md: restart proof for the observer diagnostic; Sam's Mac second silence is not a quit Co-Authored-By: Claude Fable 5.1 --- docs/bugs.md | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/docs/bugs.md b/docs/bugs.md index b33b5528a..782c0ed22 100644 --- a/docs/bugs.md +++ b/docs/bugs.md @@ -15,7 +15,7 @@ shard run reported as exit 0, 7a7e873). | 4 Oct 2026 | `vercel env add` from `site/` fails with "Could not retrieve Project Settings" | `site/.vercel/project.json` links the [other-business] team's `igneum` project; the live site (igneum.network, igneum.com, the GitHub integration) is the `igneum` team's project of the same name, which the igneum login reads and the [other-business] link does not | documented in `packaging/README-ship.md` (link, env ls, env add, deploy); no env set | `env ls` from a scratch link to the igneum-team project lists the two names; a branch push produced `igneum-git-` | | 4 Oct 2026 | `tools/jobs.mjs` and `publish-jobs.sh` print the downloads-folder token inside URLs on every run (X24 pattern) | the tokened base URL is echoed as is | the token masked as `` in every printed URL | by eye, this log's own transcript | | 4 Oct 2026 | round 4 X28 and X24, the parts under an hour: `===` on secrets, no HSTS on the relay, the relay token printed by `tools/relay.mjs list` and `watch` | as the review said | c1f59fb: `sameSecret` (timingSafeEqual, `relay/lib/auth.mjs`, test in CI), `Strict-Transport-Security` in `relay/vercel.json`, `/r/` printed (only `url` prints the real one) | relay deployed: key auth 200, wrong key and token 401, token path 200, HSTS header present | -| 4 Oct 2026 | the live feed showed two "checkpoint N locked" events 30 ms apart (1122, 1172, 1230, 1258, 1259 in 400 observer lines), and 708 of 764 locked checkpoints in `live_checkpoints` had `votes_seen` 0 | `finalityTick` read the state, awaited two SQL writes, then set the map; the `finalityLockNotification` handler checked the same map synchronously in between and recorded the lock too; a lock claimed by the notification was never upserted again, so the poll's `votes_seen` never landed | 7de1bdb: the poll claims the state before its first await; a `checkpointDetailed` set makes the poll fill `votes_seen` once | before: 5 duplicates in 400 lines; after the 19:50:33 UTC restart: 21 locks (1297 to 1317), 0 duplicates, 0 write failures; 1300 was notification-first and the poll filled it to 17 votes. The zeros that remain are the node's own count (`finality.rs:770`, its vote map for that hash, empty when the lock came by certificate), not the observer's. Index 1296 appears twice on the feed: it locked inside the restart window, one write per process, a restart-boundary one-off | +| 4 Oct 2026 | the live feed showed two "checkpoint N locked" events 30 ms apart (1122, 1172, 1230, 1258, 1259 in 400 observer lines), and 708 of 764 locked checkpoints in `live_checkpoints` had `votes_seen` 0 | `finalityTick` read the state, awaited two SQL writes, then set the map; the `finalityLockNotification` handler checked the same map synchronously in between and recorded the lock too; a lock claimed by the notification was never upserted again, so the poll's `votes_seen` never landed | 7de1bdb: the poll claims the state before its first await; a `checkpointDetailed` set makes the poll fill `votes_seen` once | before: 5 duplicates in 400 lines; after the 19:50:33 UTC restart: 21 locks (1297 to 1317), 0 duplicates, 0 write failures; 1300 was notification-first and the poll filled it to 17 votes. The zeros that remain are the node's own count (`finality.rs:770`, its vote map for that hash, empty when the lock came by certificate), not the observer's. Index 1296 appears twice on the feed: it locked inside the restart window, one write per process, a restart-boundary one-off At the 20:15:58 restart index 1319 (locked 14 min before) was recorded again: open, the seed should have held it locked; f22870a logs the seeded states and the earlier state on such a record. The 20:24:46 restart seeded 'proposed 500, locked 864' and re-recorded nothing (11 locks, 0 duplicates) | | 4 Oct 2026 | the observer kept running old code after its fix was on master and HEAD had moved past it | `autosync.sh` restarted the observer only when its own fast-forward moved HEAD; a pull by hand (19:35 UTC, HEAD to 8a77b85) bypassed it | autosync compares the checked-out `tools/observer` tree id with a marker written at each restart and restarts on any difference; `autosync.sh check` says what it would do | `check` with no marker: "restart due", exit 3; with the marker equal to the tree: "not due", exit 0; the shared checkout at that moment: due | | 4 Oct 2026 | the Mac card read 0.0 MH/s for a minute while its blocks were still accepted (after the stale mark shipped) | STALE_S 120 s left no margin: one missed 60 s upload (a 120 s gap at 19:45:59 UTC) plus a 30 s STATUS age plus the 10 s cache | STALE_S 180 s (one missed upload is not stale; PC 2's real case was a 1,860 s gap) | test: 150 s is fresh, 240 s stale | | 4 Oct 2026 | a machine stopped on purpose (Sam's Mac, 14:47 UTC) read "silent 4h" on the console, the same as a crash or a lost network | the app logs `quit: stopping the miners, then the node` and `stopped` and uploads once more before it exits (so it does report), but the console never read those lines | 4d208c9: `parseAppTail` (now in `relay/lib/parse.mjs`) sets `stopped {at, reason quit\|update}` when a quit or OTA-install line has no status line after it; the card says "stopped (quit) N ago" with a grey stripe, `tools/console.mjs` says STOPPED | unit test: quit, update, running, and a quit followed by a later status line; live: five running machines show no false stop; a live stopped case is pending the next real quit | @@ -24,4 +24,5 @@ shard run reported as exit 0, 7a7e873). ## Open - Sam's Mac (3a9bf309): CLOSED 19:15 UTC; the console side (stopped versus silent) shipped in 4d208c9. Earlier note: CLOSED 19:15 UTC, the app came back on 0.3.3 and mines (11.4 MH/s): it was stopped by its user for 4.5 h. What stays: the last node upload (14:47:02 UTC) ends with `SIGTERM - shutting down` and `igneumd has stopped`, the miner upload stops at 14:46:56, nothing after: a clean app-driven stop (quit or Stop), not a crash and not a network loss (the stop lines reached the intake). Its app predates the app-log upload (`app ?` on the console), so there is no app stream to say which. The app posts nothing on a clean quit, so the console shows "silent 4h" for a machine that was stopped on purpose. Proposed: one `[ok] app quit by the user` line uploaded before the engine stops, and the console card saying "stopped (quit) at 14:47" instead of "silent". +- Sam's Mac went silent again at 20:20:12 UTC: the app stream ends on a normal status line (8 MH/s, mining), no `quit:` line, the node log ends mid-stream; so a sleep or a lost network, not a quit, and the console's SILENT (not "stopped") is the right reading. The live "stopped" case for the detector is still pending a real quit. - `collect` jobs: the Format-List / Out-String loss did not reproduce (board2 printed `Sum`, board4 printed `Capacity` and `Speed`); only the `$_` case failed. Closed unless it shows again.