diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index b08e1e62e..5a0a69d1d 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -69,7 +69,7 @@ AMD RDNA 4 (the RX 9070 XT) sits at about a seventh of an RTX 5090 on this hash | G3 | the generator soundness suite green, the new scratch tests included | `cargo test` in igneum-pow, `tests/packs.rs`, the Metal fuzz, edge, stats, determinism runs on the v3 class | open | | G4 | the fast-time 3-node network mining across a v3 activation | 0 rejected blocks, 0 forks, every node's first lines show the switch, blocks on both sides of the boundary. Run 1 PASS (21:33 to 21:38 UTC, fork 79bd8e10 + igneum-pow 66eeba3, the mixer-x4 class without era): 3 of 3 nodes print the switch line (active from epoch 3, DAA 150 rounded up to 180 at 60-DAA epochs); templates class 2 for epochs 0 to 2 and class 3 for 3 to 5; 181 blocks before and 124 after DAA 180 (305 total, 3 CPU miners); program ids agree on all 3 miners (e3 v3 5d0dedd9fd9e29a1, e4 e81808dcdb02ce05, e5 06aff9c1d33e7a13); rejected 0/0/0 on miners and nodes; one sink 082fd39ba65df2ff on all three at 304/304/304 blocks; a new (day, class) cache 177 to 235 ms on one core, in-day swap 2 ms. `docs/plans/counter-asic-2-node.md` section 5; summary `docs/plans/counter-asic-2-gate/class-v3-20261005-2133Z-mx4.json`. Run 2 PASS (21:46:36 to 21:51:31 UTC, igneumd and igneum-miner rebuilt on b105a55 = the era and mixer composed class, hot None): every check true; 181 blocks before and 124 after DAA 180; v3 program ids e3 2d278041ba482dba, e4 2ae786d294a8a59d, e5 bc36813df2f41b5f on all three miners (the v2 id for epoch 0 8f8806638d59850f unchanged from run 1: v2 byte-identical on the chain too); rejected 0/0/0; one sink 712c1b212091dcdc at 303/303/303; 3 of 3 switch lines; cache ready v2 178 ms, first v3 181 ms, in-day swap 2 ms | GREEN (runs 1 and 2) | | G5 | the PC-built Windows workers and the Mac workers from the same commit | sha256 of each worker and the commit in the bench log | open | -| G6 | the node change on a fork branch from the 0.3.10 tip 21d4c73c with suites green on PC 2 | the build job id and its SUMMARY line. State 21:27 UTC: fork ca2-v3-node 79bd8e10 (2e464e81 the class switch + ba43cf0f the proving-v1 merge + 79bd8e10 the digest re-pin); Mac: cargo check of the seven crates clean, kaspa-consensus-core 108 + 7, kaspa-pow with igneum-pow 14 (the v3 engine test included); the PC 2 job publishes at 21:45 from the ca2-v3 worktree (the PC's test stage runs kaspa-pow and kaspa-consensus without the igneum-pow feature, so the v3 engine test's evidence is the Mac run). Expected consensus digest for a scratch devnet node with no override file after the flip: c562d70e1428c9789823cc40067623b4767f7c555ce7ff4ea11c1498f013ef6c (0.3.11; 0.3.10's is 9409deda...) | Mac green; PC 2 job build-20261005-215219 published 21:52:19Z (75-minute budget), result pending | +| G6 | the node change on a fork branch from the 0.3.10 tip 21d4c73c with suites green on PC 2 | the build job id and its SUMMARY line. State 21:27 UTC: fork ca2-v3-node 79bd8e10 (2e464e81 the class switch + ba43cf0f the proving-v1 merge + 79bd8e10 the digest re-pin); Mac: cargo check of the seven crates clean, kaspa-consensus-core 108 + 7, kaspa-pow with igneum-pow 14 (the v3 engine test included); the PC 2 job publishes at 21:45 from the ca2-v3 worktree (the PC's test stage runs kaspa-pow and kaspa-consensus without the igneum-pow feature, so the v3 engine test's evidence is the Mac run). Expected consensus digest for a scratch devnet node with no override file after the flip: c562d70e1428c9789823cc40067623b4767f7c555ce7ff4ea11c1498f013ef6c (0.3.11; 0.3.10's is 9409deda...) | Mac green. PC 2 job build-20261005-215219 (21:53:01 to 21:55:48Z, 167 s): the Linux build ok, igneum-app tests 78 + 26 + 8 passed, but kaspa-consensus 96 passed and 1 FAILED: processes::finality::tests::ban_is_decided_by_the_carrying_block_so_nodes_agree_on_every_voter_list, UnexpectedDifficulty(487112384, 487129578) in mine_on_all: the SAME flake the 0.3.10 cut hit on 21d4c73c under the six-package parallel run (release-0.3.10.md section 11: it passes alone, twice). Treated as that cut did: job 2 of 3 runs kaspa-consensus alone, job 3 of 3 the other five crates; G6 is green only when both pass | If any gate fails: stop at that gate, write why in `docs/plans/counter-asic-2-status.md`, do not publish. @@ -101,7 +101,7 @@ The fee switch H = 210,000 arrives about 19:50Z on 6 October. The 0.3.10 node's ## 8a. Proving v1 rides with it -the project lead delegated the proving v1 decisions to its agent (acd4f36bc2c07a4e2). Handoff received 21:05 UTC: fork proving-v1 = ece42979 on 21d4c73c (eb32c645 the feature on protocol 15 / message 75; 2dfad910 the rebased pool test; 3203c8d0 N = 8; ece42979 the digest test edit); Mac unit tests on it: consensus-core 26, exec 8, flows, 0 failed; app proving-v1 FINAL 6dc686a on 5b0d54f (the tier numbers from the miner-on curve, the AMD and Apple "mines and does not prove" line, the CPU path refused under 32 GB of RAM, the 24 GB tier marked "measured on the 32 GB card", the fast-time file at unproven_daa 10; `cargo test -p igneum-app provedefault` 6 of 6 on the Mac). Harness on the final fork tree: `IGNEUM_PV1_BIN=vendor/igneum-node/target-pv1/release tools/lock/with-lock.sh run node tools/proving-v1/net.mjs --secs 1500 --segment 8` gave "RESULT proving v1 harness: PASSED (21 checks) in 244.4 s" at 20:56:45Z. Override fields at publish: proving_v1_activation_daa = tip + 14,400, proving_v1_segment_blocks 8, proving_v1_unproven_daa 600, proving_v1_aggregator_share_bps 1000; they enter the digest only once the activation is set. Pinned guests unchanged. Mixed fleet: a 0.3.10 node peers with a 0.3.11 node at protocol 14 and never receives message 75; it carries segment records as miner bytes and pays nothing for them; before H the fleet is unchanged, after H only 0.3.11 producers carry and pay segment records and the shard split moves to 90/10, so every node must be on 0.3.11 before H (the fee-switch rule, section 7a). The PC 2 suite job on the merged tree (ca2-v3-node + ece42979) is the suite evidence for both halves. The app half also carries the resume fix (assigned 21:47 UTC to the proving agent on its app branch on top of 6dc686a): engine.rs's resume path restarts every enabled card's worker and the pack export if the pack is stale, re-checks within one tick that every enabled card is mining and logs a failure naming the card if not, with a unit test on the state machine (paused -> resumed -> every enabled card mining within one tick) and the known-failed case of PC 2's 21:25:11Z log; the defect left PC 2's miners off for 20 minutes tonight and the Mac's miner off after pause+resume this afternoon. If its commit is not in hand when the app tree is cut, it is first on 0.3.12's list and the status file says so. 0.3.11 carries program_class_v3 AND proving v1 together: one override object, one digest, one publish, the same gates for each half (its fast-time harness green on the final tree, suites on PC 2); both activation heights set at publish by the same rule (tip + 14,400, checked >= 10,800). If one half is not ready when the other is, the ready half ships as 0.3.11 and the other as 0.3.12; the status file says which. +the project lead delegated the proving v1 decisions to its agent (acd4f36bc2c07a4e2). Handoff received 21:05 UTC: fork proving-v1 = ece42979 on 21d4c73c (eb32c645 the feature on protocol 15 / message 75; 2dfad910 the rebased pool test; 3203c8d0 N = 8; ece42979 the digest test edit); Mac unit tests on it: consensus-core 26, exec 8, flows, 0 failed; app proving-v1 FINAL a223ca9 on 5b0d54f (6dc686a plus the resume fix: the resume path re-arms every slot without a live worker, re-exports the pack, and logs " is not mining 90 s after resume"; the known-failed case of PC 2's 21:25:11Z resume is the unit test the_pc2_resume_of_21_25_11z_restarts_under_the_new_rule_and_not_the_old, `cargo test -p igneum-app resume` 3 passed; cause: Cmd::Resume re-armed only faulted slots after stop_miners had cleared every restart_at; the tier numbers from the miner-on curve, the AMD and Apple "mines and does not prove" line, the CPU path refused under 32 GB of RAM, the 24 GB tier marked "measured on the 32 GB card", the fast-time file at unproven_daa 10; `cargo test -p igneum-app provedefault` 6 of 6 on the Mac). Harness on the final fork tree: `IGNEUM_PV1_BIN=vendor/igneum-node/target-pv1/release tools/lock/with-lock.sh run node tools/proving-v1/net.mjs --secs 1500 --segment 8` gave "RESULT proving v1 harness: PASSED (21 checks) in 244.4 s" at 20:56:45Z. Override fields at publish: proving_v1_activation_daa = tip + 14,400, proving_v1_segment_blocks 8, proving_v1_unproven_daa 600, proving_v1_aggregator_share_bps 1000; they enter the digest only once the activation is set. Pinned guests unchanged. Mixed fleet: a 0.3.10 node peers with a 0.3.11 node at protocol 14 and never receives message 75; it carries segment records as miner bytes and pays nothing for them; before H the fleet is unchanged, after H only 0.3.11 producers carry and pay segment records and the shard split moves to 90/10, so every node must be on 0.3.11 before H (the fee-switch rule, section 7a). The PC 2 suite job on the merged tree (ca2-v3-node + ece42979) is the suite evidence for both halves. The app half also carries the resume fix (assigned 21:47 UTC to the proving agent on its app branch on top of 6dc686a): engine.rs's resume path restarts every enabled card's worker and the pack export if the pack is stale, re-checks within one tick that every enabled card is mining and logs a failure naming the card if not, with a unit test on the state machine (paused -> resumed -> every enabled card mining within one tick) and the known-failed case of PC 2's 21:25:11Z log; the defect left PC 2's miners off for 20 minutes tonight and the Mac's miner off after pause+resume this afternoon. If its commit is not in hand when the app tree is cut, it is first on 0.3.12's list and the status file says so. 0.3.11 carries program_class_v3 AND proving v1 together: one override object, one digest, one publish, the same gates for each half (its fast-time harness green on the final tree, suites on PC 2); both activation heights set at publish by the same rule (tip + 14,400, checked >= 10,800). If one half is not ready when the other is, the ready half ships as 0.3.11 and the other as 0.3.12; the status file says which. Next-cut list (not consensus, not in 0.3.11 unless a one-file app change with tests): ota-k2 (branch ota-k2, commit c722579e, "OTA: the second signing key (K2) with revocation": manifest.rs, ota.rs, jobrun.rs, jobs.rs, inputs.rs, ota-sign.rs, engine.rs, the publish scripts, tools/keys, docs/security/keys.md section 4; ships signed with K1; K2's public half is empty until the project lead runs tools/keys/keygen-k2.sh), rig-install (branch rig-install dd632c1, done) with two follow-ups that belong with 0.3.11 if the proving half ships then (a rig that cannot prove defeats the point), else 0.3.12: (i) the signed public manifest names no Linux package, so the installer verifies the HiveOS tarball through the unsigned downloads sidecar behind a flag; fix = `publish-public.sh --hive` adds a `platforms.linux` entry and re-signs (the apps ignore the extra key); (ii) the published Linux package carries no prover binaries and no key-hash / sign-record miner, so the rig's prover unit idles in "setup"; fix = a Linux prover build (sp1 host, pinned guests) in the cross-build set and the package; pool-v0 (its own service, no app change), repro-bench; per-day dataset reuse in the CUDA and OpenCL workers (a Day object split out of Pair: cache freed after the build, dataset kept across prepares of the same day; the Metal worker already does this; first item after the publish, for 0.3.12; until then the integrated tier mines v3 with a restart per epoch); the rig miners' `--exit-on-seed-change` path (exit 42, re-export on restart) replaced by prepare-ahead before any epoch shorter than an hour can be drawn (layer 9 precondition, consequences C20); fork-side pack-loop 05ef0fa3 is merged into the v3 node branch because v3 touches the same miner paths. diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index a30df065b..73351e2f6 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -428,3 +428,9 @@ G4 run 2 (21:46:36 to 21:51:31 UTC, b105a55 = era + mixer, hot None): every chec 21:53. The 0.3.11 suite job is published: build-20261005-215219 to PC 2 (fork 79bd8e10, main b105a55; linux build 30 min, tests 35 min: kaspa-consensus-core, igneum-exec, kaspa-pow, kaspa-consensus, igneum-miner, kaspa-p2p-flows, igneum-app); PC 2 is on 0.3.10 with its 5090 back. The worktree freeze is lifted for the node agent (the run-2 summary commit, then the cache merge). The proving agent's resume fix waits on a test run behind the Mac's held lock slots. 21:55. The 0.3.10 rollout: both PCs mine on 0.3.10 with the rebuilt workers (PC 2 120.6 MH/s at 21:54:33Z, PC 1 141.3); the hand nodes (21:49:38Z, 21:49:50Z) and the seed (21:50:15Z) on 21d4c73c, digest 1f4b4425 everywhere; not yet on 0.3.10: the US laptop 37ba0461 (installer downloaded 21:40:52Z, app not back after 13 min; nothing to drive from here) and Sam's Mac (quit since 20:47Z). Open on PC 2: the prover fails with "CudaClientError: Connect(PermissionDenied)" since the restart (three shards 21:49:56 to 21:50:32Z); the likely cause is the sp1-gpu-server socket handling of the aggregation-cost jobs; the proving agent owns it and publishes a fix after the suite job (PC 2 is the suite job's until it closes). The ca2-v3 tip is 63dabb2 (run-2 summary and doc, the G6 job id recorded). + +## 21:56 gate G6: the first PC 2 job failed on the known kaspa-consensus flake; split re-run + +build-20261005-215219 (21:53:01 to 21:55:48Z): Linux build ok (igneumd 49,164,264 bytes sha256 11979b49..., igneum-miner d25a8270..., igneum-app 68007173...), igneum-app tests 78 + 26 + 8 passed; the node stage exit 101: kaspa-consensus 96 passed, 1 failed, processes::finality::tests::ban_is_decided_by_the_carrying_block_so_nodes_agree_on_every_voter_list, UnexpectedDifficulty(..., 487112384, 487129578) in mine_on_all. This is the flake the 0.3.10 cut met on the same base 21d4c73c under the six-package parallel run (release-0.3.10.md: it passed alone twice, job build-20261005-182804, 97 passed), a timing-dependent difficulty in the test helper, not a v3 change (v3 touches no finality code). Per the gate rule the publish stops here until the suites are green: job 2 of 3 (kaspa-consensus alone) is published now; job 3 of 3 (the other five crates) follows; the flake itself goes on the next-cut list (make mine_on_all deterministic under parallel load). + +The proving v1 app branch is final at a223ca9 (6dc686a plus the resume fix with the PC 2 case as a unit test); rollout plan 8a updated. PC 2's prover fault is the root-socket class (agg-cost-pc2-1 ran the host as root in WSL2 and left /tmp/sp1-cuda-0.sock owned by root; the app's prover has failed every shard since 21:25:24Z); the proving agent's pc2-socket-fix.ps1 (60 s, miners untouched) runs between my two suite jobs.