igneum/docs/bugs.md
igneum-labs c207c674b9 bugs.md: the observer double-lock fix proven on the live stream (21 locks, 0 duplicates)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 20:02:48 +00:00

7.7 KiB

Bugs found and fixed (bug hunt log)

One line per bug: date, symptom, cause, fix commit, how it was proven. Open items at the end. Started 4 October 2026, 20:20 BST, after two were found by hand half an hour late (three job watchers that never fired, 096b99e; a failed shard run reported as exit 0, 7a7e873).

Date Symptom Cause Fix Proven by
4 Oct 2026 Every ci run on master red since 67bf226 (eleven pushes), unnoticed sim/difficulty/records/testnet-v2-2026-10-04.schedule.log carried a home path; .log was outside the identity scrub's extension list in tools/ci/identity-check.sh (and in the mirror's tools/sync.sh) 2996cca: .log scrubbed like the other text files; the record rewritten with ~; the same list in igneum-public tools/sync.sh (local commit e18256d, not pushed) bash tools/ci/identity-check.sh 0 hits locally; run 37226816xxx on master green
4 Oct 2026 collect-pc1-board3 printed PowerShell parse errors (.Name, .AdapterRAM) the publishing shell expanded $_ inside double quotes to nothing before the command reached the jobs file; nothing to do with Format-List or Out-String (board2 and board4 printed their values) publish-jobs.sh refuses a collect command that pipes into a script block without $_ or $PSItem the eaten form refused with the reason, the single-quoted form published to a test folder
4 Oct 2026 the same job reported done (exit 0) over command exit Some(1) run_collect in app/igneum-app/src/jobrun.rs builds Done from the upload count only; the command's exit code is logged and dropped branch bugfix-collect-exit, 3c36ee4 (app engine; merge by the main session) cargo test --bin igneum-app collect_outcome: the board3 shape (Some(1)) is failed exit 1, Some(0) done, cap timeout, failed uploads still fail
4 Oct 2026 publish-jobs.sh --deploy said "not reachable, differs from the local one, or does not verify yet" after a deploy that had succeeded one check the instant the CLI returned, while the edge still served the previous file; the deploy's own exit status was hidden by || true verify_live: up to --tries (12) checks 5 s apart, each failure names its condition; publish-jobs.sh verify re-checks on its own; a failed deploy stops before the check finished: verify --tries 2 against the live file (try 1 of 2); failed: a local server with an older file ("differs", both publish stamps named) and a closed port ("is not reachable")
4 Oct 2026 console Machines: PC 37ba0461 showed 0.0 MH/s and 0 accepted while its log held an accepted block at 1 MH/s parseLabel in the console API knew nvidia, amd, mac, metal and opencl; the OpenCL fallback on an iGPU is labelled other-<id8>-n and the card was dropped parsers moved to relay/lib/parse.mjs, vendors other and intel added, node --test relay/test/parse.test.mjs in CI the test; the live console after the deploy shows the card
4 Oct 2026 console Machines: a card said "117.2 MH/s now" while its "status" column said 4 m ago (PC 2 during shard run 3: the prover held the GPU and the worker's STATUS line stopped) the card's hash came from the last STATUS line in the tail with no age check; the machine total summed it markStale: a card whose STATUS line is older than 120 s is stale, shown as "last N MH/s" with a red "stale" mark, and left out of the machine total (API, page and tools/console.mjs) the test (58 s fresh, 240 s stale, none stale); the live console after the deploy
4 Oct 2026 vercel env add from site/ fails with "Could not retrieve Project Settings" site/.vercel/project.json links the [other-business] team's igneum project; the live site (igneum.network, igneum.com, the GitHub integration) is the igneum team's project of the same name, which the igneum login reads and the [other-business] link does not documented in packaging/README-ship.md (link, env ls, env add, deploy); no env set env ls from a scratch link to the igneum-team project lists the two names; a branch push produced igneum-git-<branch>
4 Oct 2026 tools/jobs.mjs and publish-jobs.sh print the downloads-folder token inside URLs on every run (X24 pattern) the tokened base URL is echoed as is the token masked as <token> in every printed URL by eye, this log's own transcript
4 Oct 2026 round 4 X28 and X24, the parts under an hour: === on secrets, no HSTS on the relay, the relay token printed by tools/relay.mjs list and watch as the review said c1f59fb: sameSecret (timingSafeEqual, relay/lib/auth.mjs, test in CI), Strict-Transport-Security in relay/vercel.json, /r/<token> printed (only url prints the real one) relay deployed: key auth 200, wrong key and token 401, token path 200, HSTS header present
4 Oct 2026 the live feed showed two "checkpoint N locked" events 30 ms apart (1122, 1172, 1230, 1258, 1259 in 400 observer lines), and 708 of 764 locked checkpoints in live_checkpoints had votes_seen 0 finalityTick read the state, awaited two SQL writes, then set the map; the finalityLockNotification handler checked the same map synchronously in between and recorded the lock too; a lock claimed by the notification was never upserted again, so the poll's votes_seen never landed 7de1bdb: the poll claims the state before its first await; a checkpointDetailed set makes the poll fill votes_seen once before: 5 duplicates in 400 lines; after the 19:50:33 UTC restart: 21 locks (1297 to 1317), 0 duplicates, 0 write failures; 1300 was notification-first and the poll filled it to 17 votes. The zeros that remain are the node's own count (finality.rs:770, its vote map for that hash, empty when the lock came by certificate), not the observer's. Index 1296 appears twice on the feed: it locked inside the restart window, one write per process, a restart-boundary one-off
4 Oct 2026 the observer kept running old code after its fix was on master and HEAD had moved past it autosync.sh restarted the observer only when its own fast-forward moved HEAD; a pull by hand (19:35 UTC, HEAD to 8a77b85) bypassed it autosync compares the checked-out tools/observer tree id with a marker written at each restart and restarts on any difference; autosync.sh check says what it would do check with no marker: "restart due", exit 3; with the marker equal to the tree: "not due", exit 0; the shared checkout at that moment: due
4 Oct 2026 the Mac card read 0.0 MH/s for a minute while its blocks were still accepted (after the stale mark shipped) STALE_S 120 s left no margin: one missed 60 s upload (a 120 s gap at 19:45:59 UTC) plus a 30 s STATUS age plus the 10 s cache STALE_S 180 s (one missed upload is not stale; PC 2's real case was a 1,860 s gap) test: 150 s is fresh, 240 s stale

Open

  • Sam's Mac (3a9bf309): CLOSED 19:15 UTC, the app came back on 0.3.3 and mines (11.4 MH/s): it was stopped by its user for 4.5 h. What stays: the last node upload (14:47:02 UTC) ends with SIGTERM - shutting down and igneumd has stopped, the miner upload stops at 14:46:56, nothing after: a clean app-driven stop (quit or Stop), not a crash and not a network loss (the stop lines reached the intake). Its app predates the app-log upload (app ? on the console), so there is no app stream to say which. The app posts nothing on a clean quit, so the console shows "silent 4h" for a machine that was stopped on purpose. Proposed: one [ok] app quit by the user line uploaded before the engine stops, and the console card saying "stopped (quit) at 14:47" instead of "silent".
  • collect jobs: the Format-List / Out-String loss did not reproduce (board2 printed Sum, board4 printed Capacity and Speed); only the $_ case failed. Closed unless it shows again.