class v5 design page: the M5 Max rows (class v5 +0.23 ms per warp over class v4 on one core, interleaved three times under the measure lock, both a quarter of the 10 ms gate); the time line

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-07 18:33:17 +00:00
parent db056154de
commit 310e61b151

View file

@ -28,6 +28,7 @@
| 12:00 to 12:11 | PASS on every check (section 10) |
| 12:12 to 12:23 | The no-flip case PASS (section 10); both branches pushed; master merged for the build-2 test route (d1285cef) |
| 18:5x to 19:07 | Resumed on the project lead's first-priority order (build class v5 now, on the 0.3.23 line, byte 6): the frozen sub-version 3 (ca3-v4-amend 017e7037) merged into class-v5 and release-0.3.23-node into class-v5-node; the draw's and the acceptance rules' class keys set the state flag aside (the acceptance key had judged a v5 candidate under the pre-amendment rule and the draw diverged at instruction 0; fixed, 70 unit tests green on build-2 at 18:04Z); class v5 pinned to object byte 6 counted exactly (byte 7 is v4 sub-version 3); the gap list as section 13; the FIRST CLASS V5 PACK `proto-cuda/packs-ca3-v5/v5-dn3-epoch0` (Devnet 3's genesis 4020cb43... as epoch and era seed, day 20,733, the genesis state's 11 leaves under root 7e37a9fb...), exported on build-1 at 18:06:39 to 18:06:41Z, program id e5a4ac5978462156, its Metal fingerprint on the M5 Max under the measure lock at 18:07:04 to 18:07:09Z: 82b19cbde8557ea5, the three vector warps bit-exact, the build 29.8 ms GPU; the Metal pack bench takes `leaves.bin` |
| 19:2x to 19:32 | The AP-F4-1 mixer-draw rule and the AP-F1-1 shadow rule written behind the v5 class with their known-failed cases (2f9bf135; the suite running on build-2); the M5 Max rows measured (section 7: +0.23 ms per warp); the node side handed to the node lane, the kits to the v5-kits lane, the families to the attack-pass lane, a 5090 or 4090 fingerprint run of the first pack asked of the fleet lane |
| 13:2x | The coordinator's correction: the hot-set bound is NOT satisfied by sub-version 1 (F8's re-gate: nine of the first 30 seeds over 1.2x, p31 at 29.3x, 0.19 percent hot-set programs; saturated values pass saturation-preserving writers); v5 inherits the gap by merge and takes sub-version 2 by merge when it lands; the AP-F4-1 and AP-F1-1 rules come in the next round; the lane stops here |
| 12:5x to 13:21 | The amended class v4 taken in: ca3-v4-amend 8c728ca3 merged into class-v5 (4e737543) and release-0.3.20-node into class-v5-node (699db5a2); the source rule's class key sets the state flag aside so v5 and its rungs draw under it (a unit test pins the v5 chain draw equal to the amended v4's with no lossy-sourced load, the unamended v3 stream the known-failed case); generator 5 keeps `program_id(5, seed, attempt)` with no sub-version (the hash lane's reading: the suffix separates two streams inside generator 4); class v5 is object byte 6 above the amended v4's 5 (the fork, the harness, the default byte); igneum-pow 67 unit and 39 integration tests on igneum-build-2, the packs test with the re-exported control pack, the flip case PASS again with bytes 6,6,6 (the flip at epoch 8 on 3 of 3, 183 v5 blocks, the stale miner off on 59 of 59, the stateless node stopped at 480, equal roots) |
| 10:4x | The Counter ASIC coordinator relays the project lead's ruling of option A on AP-F8-1: the injecting-source load draw lands in class v4 itself on the hash lane's branch ca3-v4-amend (0.3.19); class v5 takes it by merging that branch once its commit is up, not by a second implementation; the page's section 11 row stands as the record of the bound and its gate |
@ -191,6 +192,7 @@ The same machinery as class v4 and the ladder (counter-asic-3-node.md section 6;
|---|---|---|
| The devnet's state (node 1's exec snapshot, tip 159,357) | 93 records, 8,619 bytes; stream built in 0.1 ms; rebuilt and root-checked in 0.1 ms | `igneum-day-stream`, 08:33 UK, core 40 at nice 19 |
| Verifier per unit, class v4 against class v5 (one 32-lane warp, core 40 at 3.72 GHz, nice 19, the ladder's method; the box at load 70 from other lanes' builds, so every figure is a loaded-box figure: the ladder's rung-0 row read 5.14 alone and 8.77 loaded at load 25) | v4: cold 8.11 ms alone, 8.74 ms with the sibling loaded; average of 50: 7.83 and 8.63. v5 over the devnet's 93 leaves: cold 8.71 alone, 9.03 loaded; average of 50: 8.54 and 8.83. So v5 adds 0.28 ms cold and 0.20 ms on average with the sibling loaded (+3.2 and +2.3 percent), inside the prototype's +0.11 to 0.21 ms row, and both classes stay under the 10 ms gate on the loaded core at load 70 | `tools/class-v5/verify-bench-remote.sh`, 09:37 UK, raw lines under `docs/design/class-v5-bench/20261007T083716Z/` |
| One M5 Max core (the Mac under the measure lock, nice 19, the box rule's own order: class v5 against class v4 interleaved three times, 50 warps each, 19:31 UK, Mac load 8 to 9 from other agents) | class v5 2.535, 2.649, 2.398 ms per warp (average of 50; cold 2.50, 2.99, 2.26); class v4 2.263, 2.224, 2.393 (cold 2.29, 2.25, 2.31): v5 adds 0.23 ms per warp on the average (+10 percent), the order's +0.2 ms; both classes at a quarter of the 10 ms gate on this core, and the 2019-class core by the 2.5x rule reads about 6.3 ms for v5 (O-1.14 still owed). Per tier: a node on any core from a 2019 laptop up keeps its margin; a miner's CPU engine (the harness's) pays the same 10 percent per verified warp and nothing per hash | `igneum-pow bench --program-class v5 --state <stream> --warps 50` on the Mac's own build (rustup 1.99.0; the sub-version 3 igneum-pow at 2f9bf135), raw lines kept in the lane's scratch |
| RTX 4090 (the fleet's pod cv5-2, driver 595.91.07, nvcc 12.8, 450 W limit, the card alone; `proto-newpow/class-v5/bench.cu` on the two pinned packs, 10:23 to 10:30 UK): hash rate, v4 against v5 | v4 63.067, 63.066, 63.064, 63.065 MH/s; v5 63.058, 63.058, 63.052, 63.059 MH/s over 10 batches of 2^24 each (GPU event time): v5 is 0.01 percent under v4, equal within the batch noise; the hash kernel is the same text | `docs/design/class-v5-bench/pod-4090-20261007/run.log` and `run60.log` |
| 4090: watts | the four 20-s windows read 269.4 (v4), 275.8 (v5), 275.9 (v4), 284.2 (v5) W and the four 60-s windows in the order v5, v4, v5, v4 read 284.5, 294.3, 295.8, 297.1 W: a monotone warm-up drift over the eight minutes of the run with the classes interleaved inside it (v4's last window is the highest), so the class moves nothing a window can read; the prototype's control-against-sd1 rows on 6 October read 205 to 208 W on another 4090 at the same 63.08 MH/s (this pod's card draws 4.3 to 4.7 microjoules per hash against that box's 3.27: another host, another power state, the same rate) | the same logs |
| 4090: dataset build with the devnet's leaves (93 x 64 B) | v4 30.56 to 30.58 ms, v5 30.79 to 30.81 ms: +0.24 ms (+0.8 percent) for the cache-resident leaf read; the prototype's +1.43 ms was with a 1 GiB leaf array | the same logs |