diff --git a/docs/plans/counter-asic-2-rollout.md b/docs/plans/counter-asic-2-rollout.md index 1e94bc1cf..995cd5957 100644 --- a/docs/plans/counter-asic-2-rollout.md +++ b/docs/plans/counter-asic-2-rollout.md @@ -74,7 +74,7 @@ AMD RDNA 4 (the RX 9070 XT) sits at about a seventh of an RTX 5090 on this hash | G3 | the generator soundness suite green, the new scratch tests included | `cargo test` in igneum-pow, `tests/packs.rs`, the Metal fuzz, edge, stats, determinism runs on the v3 class | open | | G4 | the fast-time 3-node network mining across a v3 activation | 0 rejected blocks, 0 forks, every node's first lines show the switch, blocks on both sides of the boundary. Run 1 PASS (21:33 to 21:38 UTC, fork 79bd8e10 + igneum-pow 66eeba3, the mixer-x4 class without era): 3 of 3 nodes print the switch line (active from epoch 3, DAA 150 rounded up to 180 at 60-DAA epochs); templates class 2 for epochs 0 to 2 and class 3 for 3 to 5; 181 blocks before and 124 after DAA 180 (305 total, 3 CPU miners); program ids agree on all 3 miners (e3 v3 5d0dedd9fd9e29a1, e4 e81808dcdb02ce05, e5 06aff9c1d33e7a13); rejected 0/0/0 on miners and nodes; one sink 082fd39ba65df2ff on all three at 304/304/304 blocks; a new (day, class) cache 177 to 235 ms on one core, in-day swap 2 ms. `docs/plans/counter-asic-2-node.md` section 5; summary `docs/plans/counter-asic-2-gate/class-v3-20261005-2133Z-mx4.json`. Run 2 PASS (21:46:36 to 21:51:31 UTC, igneumd and igneum-miner rebuilt on b105a55 = the era and mixer composed class, hot None): every check true; 181 blocks before and 124 after DAA 180; v3 program ids e3 2d278041ba482dba, e4 2ae786d294a8a59d, e5 bc36813df2f41b5f on all three miners (the v2 id for epoch 0 8f8806638d59850f unchanged from run 1: v2 byte-identical on the chain too); rejected 0/0/0; one sink 712c1b212091dcdc at 303/303/303; 3 of 3 switch lines; cache ready v2 178 ms, first v3 181 ms, in-day swap 2 ms Run 3 PASS on the FINAL class (22:15:04 to 22:19:44Z, binaries from ca2-v3 d233fa1 = x8 + era + the verifier fix, fork 89dfcb95): 182 / 122 blocks around DAA 180, 304 in all; v3 ids e3 a6523b90cff501e3, e4 bc811b3c4b8b1ced, e5 e784541f19cdebe5 on all three miners; epoch 0's v2 id 8f8806638d59850f the same in all three runs; rejected 0/0/0; one sink a9ce45df8beeaf13 at 303/303/303; 3 of 3 switch lines; cache ready v2 179 ms, first v3 191 ms. Summary `docs/plans/counter-asic-2-gate/class-v3-20261005-2215Z-era-mx8.json` | GREEN (runs 1, 2 and 3; run 3 on the final class) | | G5 | the PC-built Windows workers and the Mac workers from the same commit | sha256 of each worker and the commit in the bench log | open | -| G4b | the Mac mines v3: a real Metal miner across a v3 boundary through the miner's `--prepare-packs` flow, and the app passes that flag to the Metal worker | Found 22:21 UTC: `app/igneum-app/src/engine.rs` `miner_args` pushes `--prepare-packs` only when `card.worker != "Metal"` (line 1468 on a223ca9), so the Mac app never hands its Metal worker a prepare pack, and under class v3 the Metal worker compiles v3 only from a prepared pack (servePackProgram): at the first v3 epoch every Mac would answer `need` lines and stop, the 18:23Z outage class. Two closes before the ship, both in hand: (a) the app fix on 0.3.11's app branch (push the flag for every worker with the platform's path separator, a unit test on `miner_args` for a Metal card; assigned to the proving agent on top of a223ca9); (b) gate run 4: the fast-time network with one real Metal miner on the Mac across the activation (the prepare lines on both sides, the worker's `prepared` line for the v3 pack, found or accepted blocks on v3, no `need` or mismatch line; assigned to the node agent). If (a) is not in the app tree at the cut, the ship does not go: a Mac that cannot mine v3 at activation is a fleet outage, and the activation height (tip + 14,400) is not far enough to carry the fix in 0.3.12 safely | (a) DONE: app proving-v1 e0de2ab (prepare_packs_arg() for every worker, `packs/prepare` on macOS; start_miner gives the Metal worker the app data folder as cwd, which was None; unit test every_worker_gets_the_prepare_directory_in_the_platform_form; 113 + 27 + 8 passed). (b) gate run 4 pending | +| G4b | the Mac mines v3: a real Metal miner across a v3 boundary through the miner's `--prepare-packs` flow, and the app passes that flag to the Metal worker | Found 22:21 UTC: `app/igneum-app/src/engine.rs` `miner_args` pushes `--prepare-packs` only when `card.worker != "Metal"` (line 1468 on a223ca9), so the Mac app never hands its Metal worker a prepare pack, and under class v3 the Metal worker compiles v3 only from a prepared pack (servePackProgram): at the first v3 epoch every Mac would answer `need` lines and stop, the 18:23Z outage class. Two closes before the ship, both in hand: (a) the app fix on 0.3.11's app branch (push the flag for every worker with the platform's path separator, a unit test on `miner_args` for a Metal card; assigned to the proving agent on top of a223ca9); (b) gate run 4: the fast-time network with one real Metal miner on the Mac across the activation (the prepare lines on both sides, the worker's `prepared` line for the v3 pack, found or accepted blocks on v3, no `need` or mismatch line; assigned to the node agent). If (a) is not in the app tree at the cut, the ship does not go: a Mac that cannot mine v3 at activation is a fleet outage, and the activation height (tip + 14,400) is not far enough to carry the fix in 0.3.12 safely | GREEN. (b) gate 4 (22:25 to 22:31Z, a real Metal miner on node 0, igneum-bench from ca2-v3 00c55aa): three v3 PREPARE lines with the pack dir and class=v3 era=; the worker's v3 `prepared` lines (252.5 ms the first: program 55.5, dataset 196.9, cache fill 0.9, build 41.1; then 51 and 46 ms with the day resident); every swap "with no pause"; 124 blocks accepted on v3 (301 in the run), cpu re-check mismatched 0, need 0, no mismatch or refusal, no exit 42 or 44; chain 182 / 123 across DAA 180, 0 rejected, one sink. A second outage found and fixed before the run: the Metal worker's serveDataset built every day with the Swift version 2 construction keyed by day only, so a v3 program would have hashed over an x1 dataset; now a v3 prepare builds the day from the pack's memhard.metal and the store keys datasets by (day, class, era), commit 00c55aa | | G6 | the node change on a fork branch from the 0.3.10 tip 21d4c73c with suites green on PC 2 | the build job id and its SUMMARY line. State 21:27 UTC: fork ca2-v3-node 79bd8e10 (2e464e81 the class switch + ba43cf0f the proving-v1 merge + 79bd8e10 the digest re-pin); Mac: cargo check of the seven crates clean, kaspa-consensus-core 108 + 7, kaspa-pow with igneum-pow 14 (the v3 engine test included); the PC 2 job publishes at 21:45 from the ca2-v3 worktree (the PC's test stage runs kaspa-pow and kaspa-consensus without the igneum-pow feature, so the v3 engine test's evidence is the Mac run). Expected consensus digest for a scratch devnet node with no override file after the flip: c562d70e1428c9789823cc40067623b4767f7c555ce7ff4ea11c1498f013ef6c (0.3.11; 0.3.10's is 9409deda...) | Mac green. PC 2 job build-20261005-215219 (21:53:01 to 21:55:48Z, 167 s): the Linux build ok, igneum-app tests 78 + 26 + 8 passed, but kaspa-consensus 96 passed and 1 FAILED: processes::finality::tests::ban_is_decided_by_the_carrying_block_so_nodes_agree_on_every_voter_list, UnexpectedDifficulty(487112384, 487129578) in mine_on_all: the SAME flake the 0.3.10 cut hit on 21d4c73c under the six-package parallel run (release-0.3.10.md section 11: it passes alone, twice). Treated as that cut did: job 2 of 3 build-20261005-215712 (21:57:12 to 21:59:45Z, 122 s): kaspa-consensus alone 97 passed, 0 failed, 3 ignored in 2.10 s, ban_is_decided ... ok; job 3 of 3 build-20261005-220351 (22:04 to 22:06:55Z, 159 s): igneum-exec 17, igneum-miner 18, consensus-core and p2p-flows green, the app 78 + 26 + 8, but kaspa-pow 13 passed and 1 FAILED: igneum::tests::program_class_v3_seeds_hash_their_own_program_over_their_own_cache (consensus/pow/src/igneum.rs:910) compares a v3 program's class to V3_CLASS with era: None, while the merged crate puts the drawn era inside the class (a stale fork test, not a behaviour fault; the PC's stage does run the v3 engine test, so the PC job is the evidence). The fork test is being fixed; job 4 build-20261005-221237 (published 22:12:37Z: main d233fa1 = era + cache + the mixer fix with V3_CLASS = MX8, fork 89dfcb95) ran the five crates and the app on the final tree: every stage ok (22:12:37 to 22:15:44Z): kaspa-consensus-core 108 (2 ignored), igneum-exec 17, kaspa-pow 33 + 14 (the v3 engine test with the igneum-pow feature), igneum-miner 18, kaspa-p2p-flows 7, igneum-app 78 + 26 + 8, 0 failed. With job 2 (kaspa-consensus alone 97) G6 is GREEN on the final tree | GREEN | If any gate fails: stop at that gate, write why in `docs/plans/counter-asic-2-status.md`, do not publish. @@ -104,7 +104,7 @@ Public lines to fix on the integration branch (card-lifetime table 3): `site/ind ## 7a. A dated constraint from the consequences review (C1) -The fee switch H = 210,000 arrives about 19:00Z on 6 October (18:50 to 19:35Z: the devnet read DAA 136,578 at 22:24:31Z on 5 October, 0.909 blocks/s over the stats window and 1.005 DAA/s since 15:40Z; the 19:50Z in fee-switch-devnet.md is up to an hour late). The 0.3.10 node's export RPC carries no daaScore and no feesV1ActivationDaa (the proving v1 fork does), so from H every app prover on 0.3.10 has its shards refused and the devnet's proving goes dark. 0.3.11 must be on every prover before 16:00Z on 6 October; if it is not, the fee switch is republished at H = tip + 86,400 by the fee-switch plan's rule (a digest flip, every node in one sweep). The status file carries the timing against H. +The fee switch H = 210,000 arrives about 18:45Z on 6 October (DAA 137,041 at 22:32:54Z on 5 October; 1.002 DAA/s averaged since 15:40Z; the 19:50Z in fee-switch-devnet.md is an hour late). The 0.3.10 node's export RPC carries no daaScore and no feesV1ActivationDaa (the proving v1 fork does), so from H every app prover on 0.3.10 has its shards refused and the devnet's proving goes dark. 0.3.11 must be on every prover before 16:00Z on 6 October; if it is not, the fee switch is republished at H = tip + 86,400 by the fee-switch plan's rule (a digest flip, every node in one sweep). The status file carries the timing against H. ## 8a. Proving v1 rides with it diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index d7b34fd47..5bb41817b 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -144,7 +144,7 @@ Consequences review (a20f8c09b90016cc7) C1, C10, C11: C1 recorded in the rollout ## 20:29 C1 decided for the morning; one packfile.h fix for 0.3.11 -C1 (the fee switch H = 210,000 at about 19:00Z on 6 October, 18:50 to 19:35Z by the 22:24Z DAA read; a 0.3.10 prover's shards are vetoed from H because its export RPC carries no daaScore and no fee schedule). Decision: (a) 0.3.11 carries the proving v1 fork eb32c645 (on 21d4c73c; protocol 15, message 75) and the app branch b399708 (on 5b0d54f) with H unchanged at 210,000, published as the fleet sweep. THE 16:00Z CHECK on 6 October, for whoever holds the morning: open the console's machine cards; if every prover (PC 1 ae432dc7, PC 2 1ccfe586, the Mac) shows 0.3.11, H stands and nothing is done; if any prover is not on 0.3.11, the publisher republishes H = tip + 86,400 by the fee-switch plan's rule (one digest flip, every node in one sweep: manifest, update-now, hand nodes, seed) before about 19:00Z (18:50 to 19:35Z). The proving agent recommends (b) unless (a) is certain; the check decides it. +C1 (the fee switch H = 210,000 at about 18:45Z on 6 October by the 22:32:54Z DAA read, 1.002 DAA/s since 15:40Z; a 0.3.10 prover's shards are vetoed from H because its export RPC carries no daaScore and no fee schedule). Decision: (a) 0.3.11 carries the proving v1 fork eb32c645 (on 21d4c73c; protocol 15, message 75) and the app branch b399708 (on 5b0d54f) with H unchanged at 210,000, published as the fleet sweep. THE 16:00Z CHECK on 6 October, for whoever holds the morning: open the console's machine cards; if every prover (PC 1 ae432dc7, PC 2 1ccfe586, the Mac) shows 0.3.11, H stands and nothing is done; if any prover is not on 0.3.11, the publisher republishes H = tip + 86,400 by the fee-switch plan's rule (one digest flip, every node in one sweep: manifest, update-now, hand nodes, seed) before about 18:45Z. The proving agent recommends (b) unless (a) is certain; the check decides it. The attempt-0 packfile.h bug (the era agent's find) is the one that took the fleet down at 18:23Z (epoch 34, attempt 1); it is already fixed attempt-aware on pack-loop af983a7 and merged into the 0.3.10 tree. Rule for 0.3.11: one derivation, one test set: the era and cache agents build their workers on that packfile.h and keep only tests that add a case; the host.cu / host.c mh_word derivation fix (new, the era agent's) stays with its own test. @@ -561,3 +561,5 @@ A dry run of ca2-v3 into master in a scratch worktree conflicts in docs/bench-lo Gate 4 (22:25 to 22:31Z, a real Metal miner through the miner's --prepare-packs flow on the fast-time network, igneum-bench from ca2-v3 00c55aa): three v3 prepares with the pack dir and the class and era tokens; the worker's v3 prepared lines (252.5 ms the first: program 55.5, dataset 196.9, cache fill 0.9, build 41.1; then 51 and 46 ms with the day resident); every swap "with no pause"; 124 blocks accepted on v3, cpu re-check mismatched 0, need 0, no mismatch or refusal; chain 182 / 123 across the boundary, 0 rejected, one sink. Found and fixed before the run: the Metal worker's serveDataset built every day with the Swift version 2 construction keyed by day only, so a v3 program would have hashed over an x1 dataset and every found would have been refused by the CPU re-check, the same fleet-outage class as the app's missing flag; now a v3 prepare builds the day from the pack's memhard.metal and the store keys datasets by (day, class, era) (00c55aa). Two Mac outages caught by one gate run that the CPU-miner gates could not see; the rule for the next cut: every worker path (Metal, CUDA, OpenCL) mines across a boundary in the gate network before a class change ships. The integration merges (readwidth 30ff674, then origin/master) are in progress on ca2-v3 with the conflicts resolved by hand keeping both sides. Gates: G1, G2, G3, G4, G4b, G6 GREEN; G5 at the ship's build step. The ship waits only on the merged tip and its checks. + +22:33. H = 210,000 lands about 18:45Z on 6 October (DAA 137,041 at 22:32:54Z, 1.002 DAA/s since 15:40Z); the 16:00Z check keeps 2.75 hours. PC 2: floor-build-3 green at 22:32:29Z (240 s on the warm target; the patched sp1-gpu-server 166,768,224 bytes sha256 5568108b..., v6.8.1 c84ada1e with patch 700173fe, sm_86 sm_89 sm_120, the Go tarball's sha256 matched; the live server untouched, miners never stopped); "go PC 2 sweep" given for floor-sweep-1 (9 points, 8 to 10 min) ahead of the aggregation-cost re-run.