igneum/docs/bugs.md
igneum-labs 79f06b900e bugs.md: the OTA memo row
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 21:29:19 +00:00

15 KiB

Bugs found and fixed (bug hunt log)

One line per bug: date, symptom, cause, fix commit, how it was proven. Open items at the end. Started 4 October 2026, 20:20 BST, after two were found by hand half an hour late (three job watchers that never fired, 096b99e; a failed shard run reported as exit 0, 7a7e873).

Date Symptom Cause Fix Proven by
4 Oct 2026 Every ci run on master red since 67bf226 (eleven pushes), unnoticed sim/difficulty/records/testnet-v2-2026-10-04.schedule.log carried a home path; .log was outside the identity scrub's extension list in tools/ci/identity-check.sh (and in the mirror's tools/sync.sh) 2996cca: .log scrubbed like the other text files; the record rewritten with ~; the same list in igneum-public tools/sync.sh (local commit e18256d, not pushed) bash tools/ci/identity-check.sh 0 hits locally; run 37226816xxx on master green
4 Oct 2026 collect-pc1-board3 printed PowerShell parse errors (.Name, .AdapterRAM) the publishing shell expanded $_ inside double quotes to nothing before the command reached the jobs file; nothing to do with Format-List or Out-String (board2 and board4 printed their values) publish-jobs.sh refuses a collect command that pipes into a script block without $_ or $PSItem the eaten form refused with the reason, the single-quoted form published to a test folder
4 Oct 2026 the same job reported done (exit 0) over command exit Some(1) run_collect in app/igneum-app/src/jobrun.rs builds Done from the upload count only; the command's exit code is logged and dropped branch bugfix-collect-exit, 3c36ee4 (app engine; merge by the main session) cargo test --bin igneum-app collect_outcome: the board3 shape (Some(1)) is failed exit 1, Some(0) done, cap timeout, failed uploads still fail
4 Oct 2026 publish-jobs.sh --deploy said "not reachable, differs from the local one, or does not verify yet" after a deploy that had succeeded one check the instant the CLI returned, while the edge still served the previous file; the deploy's own exit status was hidden by || true verify_live: up to --tries (12) checks 5 s apart, each failure names its condition; publish-jobs.sh verify re-checks on its own; a failed deploy stops before the check finished: verify --tries 2 against the live file (try 1 of 2); failed: a local server with an older file ("differs", both publish stamps named) and a closed port ("is not reachable")
4 Oct 2026 console Machines: PC 37ba0461 showed 0.0 MH/s and 0 accepted while its log held an accepted block at 1 MH/s parseLabel in the console API knew nvidia, amd, mac, metal and opencl; the OpenCL fallback on an iGPU is labelled other-<id8>-n and the card was dropped parsers moved to relay/lib/parse.mjs, vendors other and intel added, node --test relay/test/parse.test.mjs in CI the test; the live console after the deploy shows the card
4 Oct 2026 console Machines: a card said "117.2 MH/s now" while its "status" column said 4 m ago (PC 2 during shard run 3: the prover held the GPU and the worker's STATUS line stopped) the card's hash came from the last STATUS line in the tail with no age check; the machine total summed it markStale: a card whose STATUS line is older than 120 s is stale, shown as "last N MH/s" with a red "stale" mark, and left out of the machine total (API, page and tools/console.mjs) the test (58 s fresh, 240 s stale, none stale); the live console after the deploy
4 Oct 2026 vercel env add from site/ fails with "Could not retrieve Project Settings" site/.vercel/project.json links the [other-business] team's igneum project; the live site (igneum.network, igneum.com, the GitHub integration) is the igneum team's project of the same name, which the igneum login reads and the [other-business] link does not documented in packaging/README-ship.md (link, env ls, env add, deploy); no env set env ls from a scratch link to the igneum-team project lists the two names; a branch push produced igneum-git-<branch>
4 Oct 2026 tools/jobs.mjs and publish-jobs.sh print the downloads-folder token inside URLs on every run (X24 pattern) the tokened base URL is echoed as is the token masked as <token> in every printed URL by eye, this log's own transcript
4 Oct 2026 round 4 X28 and X24, the parts under an hour: === on secrets, no HSTS on the relay, the relay token printed by tools/relay.mjs list and watch as the review said c1f59fb: sameSecret (timingSafeEqual, relay/lib/auth.mjs, test in CI), Strict-Transport-Security in relay/vercel.json, /r/<token> printed (only url prints the real one) relay deployed: key auth 200, wrong key and token 401, token path 200, HSTS header present
4 Oct 2026 the live feed showed two "checkpoint N locked" events 30 ms apart (1122, 1172, 1230, 1258, 1259 in 400 observer lines), and 708 of 764 locked checkpoints in live_checkpoints had votes_seen 0 finalityTick read the state, awaited two SQL writes, then set the map; the finalityLockNotification handler checked the same map synchronously in between and recorded the lock too; a lock claimed by the notification was never upserted again, so the poll's votes_seen never landed 7de1bdb: the poll claims the state before its first await; a checkpointDetailed set makes the poll fill votes_seen once before: 5 duplicates in 400 lines; after the 19:50:33 UTC restart: 21 locks (1297 to 1317), 0 duplicates, 0 write failures; 1300 was notification-first and the poll filled it to 17 votes. The zeros that remain are the node's own count (finality.rs:770, its vote map for that hash, empty when the lock came by certificate), not the observer's. Index 1296 appears twice on the feed: it locked inside the restart window, one write per process, a restart-boundary one-off At the 20:15:58 restart index 1319 (locked 14 min before) was recorded again: open, the seed should have held it locked; f22870a logs the seeded states and the earlier state on such a record. The 20:24:46 restart seeded 'proposed 500, locked 864' and re-recorded nothing (11 locks, 0 duplicates)
4 Oct 2026 the observer kept running old code after its fix was on master and HEAD had moved past it autosync.sh restarted the observer only when its own fast-forward moved HEAD; a pull by hand (19:35 UTC, HEAD to 8a77b85) bypassed it autosync compares the checked-out tools/observer tree id with a marker written at each restart and restarts on any difference; autosync.sh check says what it would do check with no marker: "restart due", exit 3; with the marker equal to the tree: "not due", exit 0; the shared checkout at that moment: due
4 Oct 2026 the Mac card read 0.0 MH/s for a minute while its blocks were still accepted (after the stale mark shipped) STALE_S 120 s left no margin: one missed 60 s upload (a 120 s gap at 19:45:59 UTC) plus a 30 s STATUS age plus the 10 s cache STALE_S 180 s (one missed upload is not stale; PC 2's real case was a 1,860 s gap) test: 150 s is fresh, 240 s stale
4 Oct 2026 a machine stopped on purpose (Sam's Mac, 14:47 UTC) read "silent 4h" on the console, the same as a crash or a lost network the app logs quit: stopping the miners, then the node and stopped and uploads once more before it exits (so it does report), but the console never read those lines 4d208c9: parseAppTail (now in relay/lib/parse.mjs) sets stopped {at, reason quit|update} when a quit or OTA-install line has no status line after it; the card says "stopped (quit) N ago" with a grey stripe, tools/console.mjs says STOPPED unit test: quit, update, running, and a quit followed by a later status line; live: five running machines show no false stop; 20:45:01 UTC Sam's Mac quit for the 0.3.4 install and the card read "STOPPED (update) 3m ago" (the quit line and the [info] installing line in its last upload), while PC 2's earlier canary quit had not been caught because the detector shipped after it
4 Oct 2026 publish-manifest.sh --deploy called the live manifest "verified" after one signature check: any validly signed manifest, including the previous version still at the edge, passed; a failed deploy was hidden by || true one-shot check, signature only, no byte compare verify_live_manifest: up to --tries checks 5 s apart, byte-identical to the folder's file, then the signature, each failure named; --verify-only runs just the check; a failed deploy stops first (the ship tool's own verify step already compared bytes and version, so a cut through tools/ship-app.mjs was covered; a hand publish was not) finished: --verify-only --tries 2 against the live 0.3.3 (try 1); failed: a local server with an altered copy ("differs", both stamps named) and a closed port ("is not reachable")
4 Oct 2026 PC 2 came back on 0.3.4 (the canary) and its card's OTA state read "none" the OTA state parsed only OTA: … lines, and the install lines sit in the previous run's log file; the new run only says update check: 0.3.4 is current (manifest 0.3.4) 797a844 then 2nd commit: every update line the app logs (is available: downloading, downloaded and verified, is ready; it installs at the next safe moment, installing, update to X complete, is marked failed, has no build yet, is current) is parsed and the newest decides the state; the Mac's card had read "0.3.3 is current" from a 70-minute-old check while 0.3.4 was downloaded, verified and staged test: staged beats the older check, a finished run reads current, a newer OTA line wins, failed and updated; live cards after the deploy
4 Oct 2026 the live 0.3.3 Mac bundle's Info.plist said 0.3.2 while the engine reported 0.3.3 (Finder and Get Info showed the wrong version) the 0.3.3 cut was by hand and the plist was not bumped nothing to do: 0.3.4 was cut by tools/ship-app.mjs, which bumps the six version files, and its plist says 0.3.4 PlistBuddy on the staged 0.3.4 bundle
4 Oct 2026 0.3.4 reports igneumd/2.1.0-73fc2fd on Windows and -1b65132 on the Macs; neither is a fork commit (both are igneum-repo commits), so the suffix says nothing about the node source the fork's build-info/build.rs looks for a .git directory; a worktree's .git is a file, so the walk climbs into the vendoring igneum repository and embeds its HEAD; then it reads .git/HEAD as a path, which a worktree lacks fork branch bugfix-build-info-worktree (vendor/igneum-node-bughunt, commits 92e1b030 and its follow-up): the root is any .git, the HEAD file comes from git rev-parse --absolute-git-dir, the ref from --git-common-dir; for the main session to merge into devnet-v4 and finality-fixes (the fork has no remote) cargo build -p kaspa-build-info in the fixed worktree embeds the fork's 92e1b030 and watches the worktree's HEAD; the unfixed devnet-v4 worktree embeds the outer repo's ebce665 and watches the outer repo's files
4 Oct 2026 PC 1's card went from "installing 0.3.4" to "none" while the install was still stuck (the Windows helper bug, 0d123b3) the app uploads at most the last 256 KiB of its log (262,144 chars per upload on PC 1) and a PC logs every block, so an update line leaves the parsed tail in about 30 min 699d0d8: console_ota_memo keeps the last OTA state per machine and app run, written when a parse finds one and read back for the same run when the tail has none; a relaunch starts clean (0855c44 widened the SQL window first, which cannot help against the upload cap) after the deploy all five machines carry memo rows and the cards read installing (PC 1), staged (PC 37ba0461), updated (the rest)

Open

  • INCIDENT 20:45 UTC, CLOSED 21:00: Sam's Mac (3a9bf309) quit for the 0.3.4 install at 20:45:03 and its run began at 21:00:05 (15 min; this Mac took the same bundle in 2 s). Its ota-apply.log (collect job collect-sam-ota-034) shows the helper verified, swapped and saw "0.3.4 is running" within seconds, so the gap is between a process the helper matched and the engine's first log line: a first-launch prompt on that Mac or a first launch that died before logging; the app logs nothing until the engine is up. Open for the main session: the host should log the moment it starts and the helper the pid it matched (app code). Console #309 and #312.

  • Nothing ties a shipped node binary to a fork commit: payload-inputs.json's node_source_commit is the fork HEAD at push time, igneumd --version prints no hash. With the build-info fix the log header carries the fork commit; the ship tool could then compare both platforms' headers with --node-commit (main session).

  • The observer's proving endpoint is the Mac app's node (IGNEUM_EVM_RPC default 127.0.0.1:26800, by design): every app restart or quit blips the proving feed (11 fetch failed lines at 20:57:04 during this Mac's OTA); the observer's own node has no --evm-rpclisten (devnet node owners).

  • ota-apply.sh (embedded in app/igneum-app/src/ota.rs) logs chdir: error retrieving current directory: getcwd because it runs with the old bundle's directory as cwd and moves it aside; harmless; a cd / at the top removes it (app code, for the main session).

  • Sam's Mac (3a9bf309): CLOSED 19:15 UTC; the console side (stopped versus silent) shipped in 4d208c9. Earlier note: CLOSED 19:15 UTC, the app came back on 0.3.3 and mines (11.4 MH/s): it was stopped by its user for 4.5 h. What stays: the last node upload (14:47:02 UTC) ends with SIGTERM - shutting down and igneumd has stopped, the miner upload stops at 14:46:56, nothing after: a clean app-driven stop (quit or Stop), not a crash and not a network loss (the stop lines reached the intake). Its app predates the app-log upload (app ? on the console), so there is no app stream to say which. The app posts nothing on a clean quit, so the console shows "silent 4h" for a machine that was stopped on purpose. Proposed: one [ok] app quit by the user line uploaded before the engine stops, and the console card saying "stopped (quit) at 14:47" instead of "silent".

  • Sam's Mac went silent again at 20:20:12 UTC: the app stream ends on a normal status line (8 MH/s, mining), no quit: line, the node log ends mid-stream; so a sleep or a lost network, not a quit, and the console's SILENT (not "stopped") is the right reading. The live "stopped" case came at 20:45:01 with its 0.3.4 install quit: labelled stopped (update).

  • collect jobs: the Format-List / Out-String loss did not reproduce (board2 printed Sum, board4 printed Capacity and Speed); only the $_ case failed. Closed unless it shows again. | 4 Oct 2026 22:25 BST | Windows update over the air sat on "the installer is starting" (PC 2, 0.3.3 to 0.3.4); no ota-apply.log, no result | the engine spawned powershell.exe with CREATE_NO_WINDOW and DETACHED_PROCESS; with no console PowerShell exits before the script's first line; the same helper launched by a remote job logged, checked the hash and refused as designed | ota.rs spawn_detached: CREATE_NO_WINDOW only (0.3.5) | dry run run-ota-helper-dryrun-pc2 (helper exit 1 with a wrong hash, log and result written); the real proof is the 0.3.5 to 0.3.6 update on a PC |