release-0.3.6 plan: PC 1 take 2 binaries and hashes, the PC 2 pair read precisely, the finality test-isolation finding

This commit is contained in:
igneum-labs 2026-10-05 09:27:28 +00:00
parent 980b7b7f62
commit bbe3d9cecf

View file

@ -332,7 +332,20 @@ Not changed: the fork's own `Command::new` for the verifier (no flags; the wrapp
|---|---|---|
| 1 | 08:47:41Z | `push-build-inputs.sh` packed the fork (2b6d23ef) and the app (0.3.6), deployed the downloads folder (build-inputs.zip 8,206,587 bytes, e46ae8d74b892097296f3fc3b4a44ce2b0960f6d3578cf00eccb1fed0c82fcbc) and then read the live `build-inputs.sha256` once, straight after "Aliased": the edge still served the previous file, the check failed and nothing was published. A curl a minute later served e46ae8d7. Fix a98be35: six tries, 10 s apart (publish-jobs.sh already retried) |
| 2 | 08:51:35Z | `build-20261005-085135` published and the apps woken (stamp 2026-10-05T08:51:35Z.852c054f). PC 1 runs 0.3.5, which has no waker (job-wake landed after the 0.3.5 cut): the 10-minute poll applies unless Check now is pressed. Started 08:54:43Z. FAILED: linux node exit 101 after 112 s, windows node exit 101 after 129 s; the app built on both (5 s and 6 s), the test stage passed (igneum-miner, igneum-app), nothing uploaded; 4 min of 45. The uploaded report holds only STAGE and RESULT lines, so the full `app/jobs/build-20261005-085135/job.log` (98,325 bytes) was fetched with a collect job (`collect-20261005-090521`, published 09:05:21Z, run 09:08:57Z, woken +193 s). The errors: `kaspad/src/daemon.rs` `no field fees_v1_activation_daa on type OverrideParams`, `no field fees on type Params`, `no method install_fee_params` (8 errors); `kaspa-rpc-service` `no method prewarm_block_template on MiningManagerProxy`. The zip's sources are right (`params.rs` carries the field 26 times, `mining/src/manager.rs` `prewarm_block_template` 5 times). Root cause: PC 1 keeps `/root/igneum-build/target` between jobs, `unzip` restores the Mac's mtimes (params.rs 07:38Z, executor.rs 07:42Z), and take3 had built 0.3.5 at 08:25 to 08:30Z, so cargo judged `kaspa-consensus-core` and `igneum-exec` fresh (neither printed "Compiling") and linked the new `kaspad` and `kaspa-rpc-service` against the cached 0.3.5 crates. Fix: master ea9794d (`push-build-inputs.sh` stamps every staged file before zipping), cherry-picked as c26d6be; and the PC side, 021a715 (`jobbuild.rs` extract stage: `find "$B/src" -type f -exec touch {} +` after the unzip, the zip root and the sibling igneum-pow, with a unit test on the order unzip, stamp, manifest check), which the 0.3.6 app carries so a packer regression cannot repeat this |
| 3 | 09:12:54Z | `build-20261005-091254` (zip 7c63df808331de5893ef93e5d3f2be65d59998668d1f99325751c5fbed033066, 8,206,958 bytes, the stamped tree; the live sha256 check passed on try 1), budget 60 min for the full rebuild, apps woken (stamp 2026-10-05T09:12:54Z.4fa446f5). Result: see below |
| 3 | 09:12:54Z | `build-20261005-091254` (zip 7c63df808331de5893ef93e5d3f2be65d59998668d1f99325751c5fbed033066, 8,206,958 bytes, the stamped tree; the live sha256 check passed on try 1), budget 60 min for the full rebuild, apps woken (stamp 2026-10-05T09:12:54Z.4fa446f5). Started 09:14:00Z (the 10-minute poll), done 09:20:33Z, 393 s, every stage ok: linux 154 s, windows 187 s, test 20 s (igneum-miner and igneum-app, the app's 74 + 25 + 8), pack 2 s, 6 files uploaded (38 MB). `node tools/build-job.mjs run` kept printing "still running" after the SUMMARY said done (its watcher missed the final state; CLAUDE.md's rule on watchers), so the outputs were fetched with `node tools/build-job.mjs fetch build-20261005-091254`: 6 of 6 verified against the PC's sha256 lines, the three Windows PE headers checked, placed under `vendor/igneum-node-036/target-integration/x86_64-pc-windows-gnu/release/`, `app/igneum-app/target/x86_64-pc-windows-gnu/release/` and `infra/cross/out/` |
| Binary (fork 2b6d23ef, PC 1 job build-20261005-091254) | sha256 | Size |
|---|---|---|
| Windows x86-64 igneumd.exe | d08404c20397fefcc02cd56cad0dd4d29b427a468fd89e3189435f7c3431cd7a | 49,971,712 |
| Windows x86-64 igneum-miner.exe | 8bdb4c6e64190b73ae88cd893c3c2efd8fb0e226b43c0c240c1d545206ed7147 | 10,725,888 |
| Windows x86-64 igneum-app.exe (the PC's build; the installer's engine is the GitHub runner's) | 880e7c81acc1c1a532b8c1b244c70c40406f2b8c4e1420392a2dd8edd7f236ed | 2,834,944 |
| Linux x86-64 igneumd (`infra/cross/out/`, for the seed and HiveOS, not in the app) | c24fd2c5e8c4f976e2b3abc55873bca29ae6b97b47046c946f779d3a04947025 | 48,733,480 |
| Linux x86-64 igneum-miner | 038fcf6b133d0af36d17e28fe822c5bcc7055429958dc280e75459c4347f9a53 | 9,607,440 |
| Linux x86-64 igneum-app | db73814f87c07ee3cada1fec5d728062ae8a8fa3a318fdd94b99f1e6113134fb | 2,297,160 |
PC 2's cold build of the same source gave the same igneum-miner (038fcf6b...) and igneum-app (db73814f...) but another
igneumd (d9d227a9... against c24fd2c5...): the daemon's build is not reproducible across the two machines (unverified
why; build paths or timestamps are the usual causes). Not acted on.
Standing rule from the project lead during the cut (CLAUDE.md f378aa0): builds and test suites run on the PCs, the Mac builds only
the macOS binaries. So the fork's suites and the app's tests went to PC 2 as a second build job:
@ -362,7 +375,20 @@ Whether it is pre-existing is being settled the way the coordinator asked: two m
`--node-tests kaspa-consensus` and nothing else, `build-20261005-091739` on release-0.3.5 20139145
(zip `build-inputs-t035.zip` d1dd841d...) and `build-20261005-091739-036` on 2b6d23ef (`build-inputs-t036.zip`
c01a7525...), published 09:17:39Z and 09:18:06Z (the second `add` within the same second had collided with the
first's generated id; `--id` fixed it). Results: see 8i.
first's generated id; `--id` fixed it).
| Job | Tree | Result |
|---|---|---|
| `build-20261005-091739` | release-0.3.5 20139145 (the SUMMARY names it) | started 09:19:09Z, 203 s: Linux node built in 78 s, `cargo test --release -p kaspa-consensus`: 94 passed, 0 failed, 3 ignored, 110 s. The whole suite alone passes on 0.3.5 |
| `build-20261005-091739-036` | 2b6d23ef | NOT a valid run: both zips were packed at 09:17Z, before the first job built 0.3.5 into PC 2's target dir at 09:20Z, so this job's stamps were older than that build and cargo kept the 0.3.5 crates (the same class as 8f attempt 2, now from the two zips of one pair): Linux node exit 101 after 16 s, and its "94 passed" in 2 s came from the cached 0.3.5 test binaries |
So what is known: the two finality tests pass on 0.3.5 with the suite alone, pass on 2b6d23ef when the finality
module runs alone (3b, the Mac), and fail on 2b6d23ef only in the five-package parallel run on a cold cache
(`build-20261005-090600`), where the error is M30's cache queue (`PowCacheQueueFull`, 4 waiting, 2 building), a
per-process limit that a parallel suite exceeds. The owner's call (coordinator, 09:2x BST): a test-isolation
problem, not a consensus regression; recorded here and in the ledger (tests get a per-process PoW cache
directory, 0.3.7); the ship is not blocked. A clean `kaspa-consensus`-alone run on 2b6d23ef is still owed and goes
after the cut (pack the zip after the previous PC 2 job has built, or wait for the 0.3.6 app's extract stamp).
### 8h2. The DMG