diff --git a/docs/plans/counter-asic-2-node.md b/docs/plans/counter-asic-2-node.md index 717c1af46..24287a9ae 100644 --- a/docs/plans/counter-asic-2-node.md +++ b/docs/plans/counter-asic-2-node.md @@ -108,7 +108,15 @@ Binaries rebuilt on ca2-v3 b105a55 (era layout merged on the mixer: `V3_CLASS` = | Forks | sinks `712c1b212091dcdc` on all three nodes, block counts 303 / 303 / 303, one tip each | | Program and cache ready, one CPU core | v2 epoch 0 178 ms, the first v3 epoch 181 ms, an in-day swap 2 ms | -CPU hash rate across the switch, run 2, miner cpu0 (one thread, the Mac shared with other agents' builds, so approximate): the cumulative rate read 0.024 MH/s through the v2 epochs (30 to 151 s), then fell to 0.021 MH/s cumulative by 271 s (100 s under v3), which puts the v3 interval rate near 0.017 MH/s, about 30 percent under v2 on the CPU interpreter (the era's strided windowed loads and the mixer path). The node has no per-block verify timing line; the era agent's measurement of the CPU verifier (`igneum-pow`, 0.604 ms per warp on readwidth against 1.332 ms on 88dafbc, the mixer's `derive_items` path at m = 1) is the number to fix before the publish, and the before/after goes here when the mixer agent's commit lands. +CPU hash rate across the switch, run 2, miner cpu0 (one thread, the Mac shared with other agents' builds, so approximate): the cumulative rate read 0.024 MH/s through the v2 epochs (30 to 151 s), then fell to 0.021 MH/s cumulative by 271 s (100 s under v3), which puts the v3 interval rate near 0.017 MH/s, about 30 percent under v2 on the CPU interpreter (the era's strided windowed loads and the mixer path). The node has no per-block verify timing line; the CPU verifier is measured in `igneum-pow` (the mixer agent, one M5 Max core, ms per 32-lane unit, same minute, cited from ca2-mixer 1ab8b21's message of 5 October 2026 22:10Z): + +| Path | readwidth | ca2-v3 88dafbc (before the fix) | ca2-v3 d233fa1 (after) | +|---|---|---|---| +| v2 (the live devnet) | 0.607 / 0.610 | 1.332 | 0.609 / 0.611 | +| v3 at x4 | | | 1.238 | +| v3 at x8 (the class) | | | 2.077 (worst cold 2.15) | + +The regression was `memhard::derive_items` at m = 1 (2.2x); the fix dispatches to an out-of-line `derive_items_mask`, one instance per cache size with the line mask a constant. A v3 block costs the node about 3.4x a v2 block to verify (2.08 against 0.61 ms per unit). Epoch 0's v2 id is the same in both runs (`8f8806638d59850f`: a v2 program is untouched by the era code, on the chain as in the packs); the v3 ids differ from run 1 because the era draw is now inside the class. @@ -140,6 +148,6 @@ The PC's test stage builds the fork's test binaries with the plan's feature set, - The GPU workers' class refusal was checked by the C test of `pf_pack_class_ok` and the syntax of both hosts, not by a live worker on a v3 pack: the integration's bit-exact gate (G1) is where a real worker first builds a v3 pack. - `next_pair` keeps the current era seed for the next epoch; an epoch boundary that is also an era boundary (once per 180 days) would prepare the wrong era, and the job line then names the right one, so the worker refuses the prepared pair and the miner prepares again (one wasted compile, no wrong block). The VDF era seed replaces the stand-in before this matters. - Done 22:05Z: ca2-cache rebased as 1950661 fast-forwarded (the first attempt, 2de19e5 on 464d6e1, conflicted in 8 files and was aborted); igneum-pow 53 + 19 tests and the packfile test pass; the fork's kaspa-pow, miner and kaspad check clean against the merged crate (seam unchanged). `V3_CLASS = { era: None, hot: None, ..LoadClass::MX4 }`. -- Owed: the mixer agent's verifier fix (the 2.2x on the v2 path at m = 1). ca2-mixer's tip 54bbfcc (22:10Z) carries the MX8 candidate, tests/mixer.rs, tests/scratch.rs and the measure docs, not the fix; a merge of it into ca2-v3 conflicts in docs/bench-log.md and proto-metal/packbench.swift only (generator.rs and verify.rs auto-merge). Merged when the fix commit arrives, then a rebuild and gate run 3 if asked. +- Done 22:12Z: ca2-mixer 16dfd1e (MX8 candidate, tests) merged at 4e733bb with two one-line fixes the merge needed (3a7fba7 `MX8` gets `era: None, hot: None`; 795472e tests/scratch.rs's `Instr` literals get `win: 0, off: 0`); then ca2-mixer 1ab8b21 (the verifier fix, `V3_CLASS = { era: None, hot: None, ..LoadClass::MX8 }`: class v3 is x8 by the coordinator's decision of 22:05Z) merged at d233fa1. Checks on d233fa1: igneum-pow 53 + 4 + 19 + 7, packfile 0 failures, the fork's `cargo test -p kaspa-pow --features igneum-pow` 14 pass, the release rebuild 1 min 57 s. - Owed (0.3.12, coordinator's ask of 5 October 2026 22:20Z): per-day dataset reuse in the CUDA and OpenCL workers (a `Day` object shared by consecutive pairs, the cache freed after the build); the Metal worker already keys datasets by day. Until then the iGPU tier mines v3 with a dataset rebuild per epoch on those two workers. - Wire compatibility: `RpcPowEpochInfo` gained five fields in its Borsh form (wRPC) and five proto fields (gRPC); the gRPC side reads an old node's zeros as v2 / never / none; the Borsh form is versioned by `GetBlockTemplateResponse` (version 2 carries the whole struct), so a 0.3.11 wRPC client against a 0.3.10 node reads short: the miner uses gRPC, the console reads JSON (serde defaults).