The first Windows update over the air (0.3.3 to 0.3.4) sat on 'the installer is starting' with no helper log on PC 2. The helper launched by a remote job with the same arguments ran at once. Fix in 0.3.5; the Windows machines take 0.3.5 by hand once more. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
14 KiB
Bugs found and fixed (bug hunt log)
One line per bug: date, symptom, cause, fix commit, how it was proven. Open items at the end. Started 4 October 2026,
20:20 BST, after two were found by hand half an hour late (three job watchers that never fired, 096b99e; a failed
shard run reported as exit 0, 7a7e873).
| Date | Symptom | Cause | Fix | Proven by |
|---|---|---|---|---|
| 4 Oct 2026 | Every ci run on master red since 67bf226 (eleven pushes), unnoticed |
sim/difficulty/records/testnet-v2-2026-10-04.schedule.log carried a home path; .log was outside the identity scrub's extension list in tools/ci/identity-check.sh (and in the mirror's tools/sync.sh) |
2996cca: .log scrubbed like the other text files; the record rewritten with ~; the same list in igneum-public tools/sync.sh (local commit e18256d, not pushed) |
bash tools/ci/identity-check.sh 0 hits locally; run 37226816xxx on master green |
| 4 Oct 2026 | collect-pc1-board3 printed PowerShell parse errors (.Name, .AdapterRAM) |
the publishing shell expanded $_ inside double quotes to nothing before the command reached the jobs file; nothing to do with Format-List or Out-String (board2 and board4 printed their values) |
publish-jobs.sh refuses a collect command that pipes into a script block without $_ or $PSItem |
the eaten form refused with the reason, the single-quoted form published to a test folder |
| 4 Oct 2026 | the same job reported done (exit 0) over command exit Some(1) |
run_collect in app/igneum-app/src/jobrun.rs builds Done from the upload count only; the command's exit code is logged and dropped |
branch bugfix-collect-exit, 3c36ee4 (app engine; merge by the main session) |
cargo test --bin igneum-app collect_outcome: the board3 shape (Some(1)) is failed exit 1, Some(0) done, cap timeout, failed uploads still fail |
| 4 Oct 2026 | publish-jobs.sh --deploy said "not reachable, differs from the local one, or does not verify yet" after a deploy that had succeeded |
one check the instant the CLI returned, while the edge still served the previous file; the deploy's own exit status was hidden by || true |
verify_live: up to --tries (12) checks 5 s apart, each failure names its condition; publish-jobs.sh verify re-checks on its own; a failed deploy stops before the check |
finished: verify --tries 2 against the live file (try 1 of 2); failed: a local server with an older file ("differs", both publish stamps named) and a closed port ("is not reachable") |
| 4 Oct 2026 | console Machines: PC 37ba0461 showed 0.0 MH/s and 0 accepted while its log held an accepted block at 1 MH/s | parseLabel in the console API knew nvidia, amd, mac, metal and opencl; the OpenCL fallback on an iGPU is labelled other-<id8>-n and the card was dropped |
parsers moved to relay/lib/parse.mjs, vendors other and intel added, node --test relay/test/parse.test.mjs in CI |
the test; the live console after the deploy shows the card |
| 4 Oct 2026 | console Machines: a card said "117.2 MH/s now" while its "status" column said 4 m ago (PC 2 during shard run 3: the prover held the GPU and the worker's STATUS line stopped) | the card's hash came from the last STATUS line in the tail with no age check; the machine total summed it | markStale: a card whose STATUS line is older than 120 s is stale, shown as "last N MH/s" with a red "stale" mark, and left out of the machine total (API, page and tools/console.mjs) |
the test (58 s fresh, 240 s stale, none stale); the live console after the deploy |
| 4 Oct 2026 | vercel env add from site/ fails with "Could not retrieve Project Settings" |
site/.vercel/project.json links the [other-business] team's igneum project; the live site (igneum.network, igneum.com, the GitHub integration) is the igneum team's project of the same name, which the igneum login reads and the [other-business] link does not |
documented in packaging/README-ship.md (link, env ls, env add, deploy); no env set |
env ls from a scratch link to the igneum-team project lists the two names; a branch push produced igneum-git-<branch> |
| 4 Oct 2026 | tools/jobs.mjs and publish-jobs.sh print the downloads-folder token inside URLs on every run (X24 pattern) |
the tokened base URL is echoed as is | the token masked as <token> in every printed URL |
by eye, this log's own transcript |
| 4 Oct 2026 | round 4 X28 and X24, the parts under an hour: === on secrets, no HSTS on the relay, the relay token printed by tools/relay.mjs list and watch |
as the review said | c1f59fb: sameSecret (timingSafeEqual, relay/lib/auth.mjs, test in CI), Strict-Transport-Security in relay/vercel.json, /r/<token> printed (only url prints the real one) |
relay deployed: key auth 200, wrong key and token 401, token path 200, HSTS header present |
| 4 Oct 2026 | the live feed showed two "checkpoint N locked" events 30 ms apart (1122, 1172, 1230, 1258, 1259 in 400 observer lines), and 708 of 764 locked checkpoints in live_checkpoints had votes_seen 0 |
finalityTick read the state, awaited two SQL writes, then set the map; the finalityLockNotification handler checked the same map synchronously in between and recorded the lock too; a lock claimed by the notification was never upserted again, so the poll's votes_seen never landed |
7de1bdb: the poll claims the state before its first await; a checkpointDetailed set makes the poll fill votes_seen once |
before: 5 duplicates in 400 lines; after the 19:50:33 UTC restart: 21 locks (1297 to 1317), 0 duplicates, 0 write failures; 1300 was notification-first and the poll filled it to 17 votes. The zeros that remain are the node's own count (finality.rs:770, its vote map for that hash, empty when the lock came by certificate), not the observer's. Index 1296 appears twice on the feed: it locked inside the restart window, one write per process, a restart-boundary one-off At the 20:15:58 restart index 1319 (locked 14 min before) was recorded again: open, the seed should have held it locked; f22870a logs the seeded states and the earlier state on such a record. The 20:24:46 restart seeded 'proposed 500, locked 864' and re-recorded nothing (11 locks, 0 duplicates) |
| 4 Oct 2026 | the observer kept running old code after its fix was on master and HEAD had moved past it | autosync.sh restarted the observer only when its own fast-forward moved HEAD; a pull by hand (19:35 UTC, HEAD to 8a77b85) bypassed it |
autosync compares the checked-out tools/observer tree id with a marker written at each restart and restarts on any difference; autosync.sh check says what it would do |
check with no marker: "restart due", exit 3; with the marker equal to the tree: "not due", exit 0; the shared checkout at that moment: due |
| 4 Oct 2026 | the Mac card read 0.0 MH/s for a minute while its blocks were still accepted (after the stale mark shipped) | STALE_S 120 s left no margin: one missed 60 s upload (a 120 s gap at 19:45:59 UTC) plus a 30 s STATUS age plus the 10 s cache | STALE_S 180 s (one missed upload is not stale; PC 2's real case was a 1,860 s gap) | test: 150 s is fresh, 240 s stale |
| 4 Oct 2026 | a machine stopped on purpose (Sam's Mac, 14:47 UTC) read "silent 4h" on the console, the same as a crash or a lost network | the app logs quit: stopping the miners, then the node and stopped and uploads once more before it exits (so it does report), but the console never read those lines |
4d208c9: parseAppTail (now in relay/lib/parse.mjs) sets stopped {at, reason quit|update} when a quit or OTA-install line has no status line after it; the card says "stopped (quit) N ago" with a grey stripe, tools/console.mjs says STOPPED |
unit test: quit, update, running, and a quit followed by a later status line; live: five running machines show no false stop; 20:45:01 UTC Sam's Mac quit for the 0.3.4 install and the card read "STOPPED (update) 3m ago" (the quit line and the [info] installing line in its last upload), while PC 2's earlier canary quit had not been caught because the detector shipped after it |
| 4 Oct 2026 | publish-manifest.sh --deploy called the live manifest "verified" after one signature check: any validly signed manifest, including the previous version still at the edge, passed; a failed deploy was hidden by || true |
one-shot check, signature only, no byte compare | verify_live_manifest: up to --tries checks 5 s apart, byte-identical to the folder's file, then the signature, each failure named; --verify-only runs just the check; a failed deploy stops first (the ship tool's own verify step already compared bytes and version, so a cut through tools/ship-app.mjs was covered; a hand publish was not) |
finished: --verify-only --tries 2 against the live 0.3.3 (try 1); failed: a local server with an altered copy ("differs", both stamps named) and a closed port ("is not reachable") |
| 4 Oct 2026 | PC 2 came back on 0.3.4 (the canary) and its card's OTA state read "none" | the OTA state parsed only OTA: … lines, and the install lines sit in the previous run's log file; the new run only says update check: 0.3.4 is current (manifest 0.3.4) |
797a844 then 2nd commit: every update line the app logs (is available: downloading, downloaded and verified, is ready; it installs at the next safe moment, installing, update to X complete, is marked failed, has no build yet, is current) is parsed and the newest decides the state; the Mac's card had read "0.3.3 is current" from a 70-minute-old check while 0.3.4 was downloaded, verified and staged |
test: staged beats the older check, a finished run reads current, a newer OTA line wins, failed and updated; live cards after the deploy |
| 4 Oct 2026 | the live 0.3.3 Mac bundle's Info.plist said 0.3.2 while the engine reported 0.3.3 (Finder and Get Info showed the wrong version) |
the 0.3.3 cut was by hand and the plist was not bumped | nothing to do: 0.3.4 was cut by tools/ship-app.mjs, which bumps the six version files, and its plist says 0.3.4 |
PlistBuddy on the staged 0.3.4 bundle |
| 4 Oct 2026 | 0.3.4 reports igneumd/2.1.0-73fc2fd on Windows and -1b65132 on the Macs; neither is a fork commit (both are igneum-repo commits), so the suffix says nothing about the node source |
the fork's build-info/build.rs looks for a .git directory; a worktree's .git is a file, so the walk climbs into the vendoring igneum repository and embeds its HEAD; then it reads .git/HEAD as a path, which a worktree lacks |
fork branch bugfix-build-info-worktree (vendor/igneum-node-bughunt, commits 92e1b030 and its follow-up): the root is any .git, the HEAD file comes from git rev-parse --absolute-git-dir, the ref from --git-common-dir; for the main session to merge into devnet-v4 and finality-fixes (the fork has no remote) |
cargo build -p kaspa-build-info in the fixed worktree embeds the fork's 92e1b030 and watches the worktree's HEAD; the unfixed devnet-v4 worktree embeds the outer repo's ebce665 and watches the outer repo's files |
Open
-
INCIDENT 20:45 UTC, CLOSED 21:00: Sam's Mac (3a9bf309) quit for the 0.3.4 install at 20:45:03 and its run began at 21:00:05 (15 min; this Mac took the same bundle in 2 s). Its
ota-apply.log(collect jobcollect-sam-ota-034) shows the helper verified, swapped and saw "0.3.4 is running" within seconds, so the gap is between a process the helper matched and the engine's first log line: a first-launch prompt on that Mac or a first launch that died before logging; the app logs nothing until the engine is up. Open for the main session: the host should log the moment it starts and the helper the pid it matched (app code). Console #309 and #312. -
Nothing ties a shipped node binary to a fork commit:
payload-inputs.json'snode_source_commitis the fork HEAD at push time,igneumd --versionprints no hash. With the build-info fix the log header carries the fork commit; the ship tool could then compare both platforms' headers with--node-commit(main session). -
The observer's proving endpoint is the Mac app's node (
IGNEUM_EVM_RPCdefault 127.0.0.1:26800, by design): every app restart or quit blips the proving feed (11fetch failedlines at 20:57:04 during this Mac's OTA); the observer's own node has no--evm-rpclisten(devnet node owners). -
ota-apply.sh(embedded inapp/igneum-app/src/ota.rs) logschdir: error retrieving current directory: getcwdbecause it runs with the old bundle's directory as cwd and moves it aside; harmless; acd /at the top removes it (app code, for the main session). -
Sam's Mac (3a9bf309): CLOSED 19:15 UTC; the console side (stopped versus silent) shipped in
4d208c9. Earlier note: CLOSED 19:15 UTC, the app came back on 0.3.3 and mines (11.4 MH/s): it was stopped by its user for 4.5 h. What stays: the last node upload (14:47:02 UTC) ends withSIGTERM - shutting downandigneumd has stopped, the miner upload stops at 14:46:56, nothing after: a clean app-driven stop (quit or Stop), not a crash and not a network loss (the stop lines reached the intake). Its app predates the app-log upload (app ?on the console), so there is no app stream to say which. The app posts nothing on a clean quit, so the console shows "silent 4h" for a machine that was stopped on purpose. Proposed: one[ok] app quit by the userline uploaded before the engine stops, and the console card saying "stopped (quit) at 14:47" instead of "silent". -
Sam's Mac went silent again at 20:20:12 UTC: the app stream ends on a normal status line (8 MH/s, mining), no
quit:line, the node log ends mid-stream; so a sleep or a lost network, not a quit, and the console's SILENT (not "stopped") is the right reading. The live "stopped" case came at 20:45:01 with its 0.3.4 install quit: labelled stopped (update). -
collectjobs: the Format-List / Out-String loss did not reproduce (board2 printedSum, board4 printedCapacityandSpeed); only the$_case failed. Closed unless it shows again. | 4 Oct 2026 22:25 BST | Windows update over the air sat on "the installer is starting" (PC 2, 0.3.3 to 0.3.4); no ota-apply.log, no result | the engine spawned powershell.exe with CREATE_NO_WINDOW and DETACHED_PROCESS; with no console PowerShell exits before the script's first line; the same helper launched by a remote job logged, checked the hash and refused as designed | ota.rs spawn_detached: CREATE_NO_WINDOW only (0.3.5) | dry run run-ota-helper-dryrun-pc2 (helper exit 1 with a wrong hash, log and result written); the real proof is the 0.3.5 to 0.3.6 update on a PC |