From 598fd20f0211aabaa903bcc4594071263525ad89 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Wed, 7 Oct 2026 22:08:58 +0000 Subject: [PATCH] Counter ASIC 4.0 research: 20.2 the verifier rows (AVX2 tile path 0.047 us per tile per unit, mm1430 10.14 ms loaded; the per-load placement derives 11 percent fewer distinct items per warp: a uniformity fault as drawn, with its fix) Co-Authored-By: Claude Fable 5.1 --- docs/analysis/counter-asic-4-research.md | 17 ++++++++++++++--- 1 file changed, 14 insertions(+), 3 deletions(-) diff --git a/docs/analysis/counter-asic-4-research.md b/docs/analysis/counter-asic-4-research.md index 4454bd651..2ad32d23c 100644 --- a/docs/analysis/counter-asic-4-research.md +++ b/docs/analysis/counter-asic-4-research.md @@ -334,9 +334,20 @@ The coordinator's order of 22:5x UK: both new ranks as experimental program clas | The packs | `proto-cuda/packs-ca4/` (also build-1 `/srv/builds/ca4-research/packs/`, the hash lane's kit `igneum-ca4-packs-20261007.zip`) | `mx8_sh256x27` (the control, class v4's shape), `mx8_shl256x27`, `mx8_mm128`, `mx8_mm512`, `mx8_mm1430` (11,440 tiles per hash: the ALU shadow's premium in tiles at the 4090's 0.056 pJ per MAC), all over seed `igneum-genesis`, day 2026-10-03, 1 GiB, generator 2 | every export `OVERALL: PASS` (vectors through the interpreter); the CUDA-against-CPU bit-exactness is the worker's self-test on the card (96 vector lanes through the bound kernel), unrun | | What is NOT built | | the NEON path of the tile verifier (the M5 Max runs scalar); the AVX-VNNI path (the AVX2 path stands in; VNNI would halve its instruction count, approximate); the acceptance rule's view of the per-load placement (the rule interprets the base program only, as for class v4; the sub-version 3 freshness fixpoint runs over base then shadow, which under per-load placement is not the execution order: a class, not a prototype, would re-derive it) | | -### 20.2 The verifier rows (igneum-build-2, EPYC 9454P, core 40 at nice 19, the ladder's method: 50 warps, alone and with the SMT sibling loaded by the same bench; the box under other lanes' load) +### 20.2 The verifier rows (igneum-build-2, EPYC 9454P, core 40 at nice 19 under `lease cores 40,88`, 21:5x to 22:07 UTC; the ladder's method: one cold warp per vector base, alone and with the SMT sibling core 88 running the same bench; the box at load 84 to 95 from other lanes, core 40 at 3.27 GHz) -PENDING: the bench ran under `lease cores` at 21:5x UTC and its rows land here when the lease returns (the box's pool was fully leased and the lease waited). The reading they give: the per-load class should cost the verifier what class v4 costs (the same instructions); the tile classes add 1,024 byte products per tile per warp, about 1.07 microseconds per tile per unit scalar (the 6 October figure), so `mm1430` is about 12 ms scalar on an M5 Max core (FAILS) and the AVX2 row is the one that decides whether a tile class can sit under the 10 ms gate on a 2019-class core. +| Class | Cold ms alone (warp base 0 / 4096 / 1000000) | Cold ms, sibling loaded | Items derived per warp (of 4,096 reads) | Reading | +|---|---|---|---|---| +| `mx8+sh256x27` (class v4's shape, the control) | 5.48 / 5.06 / 5.20 | 9.52 / 9.16 / 9.12 | 4,096 / 4,096 / 4,096 | the ladder's rung 0 read 5.14 and 8.77 on 6 October at load 25; tonight's box is heavier | +| `mx8+shl256x27` (per load) | 5.09 / 4.84 / 4.86 | 8.66 / 8.34 / 8.38 | **3,627 / 3,595 / 3,633** | the same instructions, 7 to 9 percent FASTER: because 11 to 12 percent of the warp's reads hit an item another lane already derived (the finding below) | +| `mx8+mm128` (128 tiles per iteration, no ALU shadow) | 5.33 / 5.13 / 5.08 | 8.59 / 8.37 / 8.25 | 4,096 | AVX2 tile path | +| `mx8+mm512` | 5.31 / 5.12 / 5.12 | 9.30 / 9.15 / 9.03 | 4,096 | passes the gate loaded | +| `mx8+mm1430` (11,440 tiles per hash, the ALU shadow's premium in tiles) | 5.82 / 5.64 / 5.66 | **10.14** / 9.58 / 9.55 | 4,096 | alone 5.8 ms; loaded the base-0 warp is over the gate by 0.14 ms (rung 3's class of miss, at a heavier load than the ladder's run) | +| `mx8+mm1430`, the scalar tile path forced (`IGNEUM_MM8_SCALAR=1`) | 7.60 / 7.39 / 7.35 | not run | 4,096 | the AVX2 path is 4.6x the scalar path on the tile work | + +The tile cost law from these rows: AVX2 (5.82 - 5.33) ms over (1,430 - 128) x 8 = 10,416 extra tiles per unit = 0.047 microseconds per tile per unit; scalar (7.60 - 5.33) / 10,416 = 0.22 microseconds per tile per unit (the 6 October figure of 1.07 was a naive C loop on a loaded core). So a tile class carrying the ALU shadow's premium costs the verifier about 0.5 ms per unit on an AVX2 core and 2.3 ms scalar; against class v4's own 55,296 shadow instructions at 0.1 ns per lane-instruction (about 0.18 ms per unit) the tile block is 3x the verifier cost per unit at the same premium with AVX2, 13x scalar. On the measured box core `mm1430` passes alone (5.8 ms) and misses the loaded gate by 0.14 ms; `mm512` passes loaded. On a 2019-class laptop core by the 2.5x rule: `mm512` about 13 ms loaded (FAILS), `mm128` about 12.8 (the ALU-free base is already 12.6 on that rule's loaded column; the rule itself is the open O-1.14). A NEON path (M-series: scalar tonight) and a VNNI path (halves the AVX2 count, approximate) are the two unbuilt verifier items. + +**A finding on the per-load placement (not a tuning).** With the same base program and the same loads, the per-load class derives 3,595 to 3,633 distinct items per warp where the class v4 shape derives all 4,096: 11 to 12 percent of the warp's 4,096 reads land on an item another lane of the same warp has already read. The cause (read from the construction, not yet from a trace): a sub-block of 16 drawn ALU instructions ends right before the next load, and when its last writer of the load's source register is a `shfl` (29 of 256 shadow instructions are shuffles), two lanes carry a neighbour's value into the same address; in the class v4 shape the base program's own instructions re-randomise every register between a shuffle and a load, and the sub-version 3 source rule (a load's source last written by an injecting op or a rotate, judged over base then shadow) does not see the per-load order. Consequences: a within-warp duplicate read is served from L1 or L2 on the GPU and from a lane buffer on a chip, so it costs neither side DRAM energy but it is 11 percent fewer dependent reads per hash, which is exactly the uniformity bound the attack pass gates (F8's 1.2x on the hot set); the per-load class as drawn FAILS that spirit and is not a candidate as exported. The fix is one rule: a sub-block's instructions that would be the last writer of the next load's source are drawn from the injecting families only (or the sub-block ends one instruction before the load with a rotate), the sub-version 3 freshness fixpoint run over the real execution order; 2 to 3 agent hours plus a re-export and the F8 census at 2^24 on 64 seeds. The card rows of the exported pack still answer the energy question (the same instruction count per hash, 11 percent fewer DRAM reads, so the pack reads slightly FASTER than the control and that part of the delta is the duplicate reads, not the placement). ### 20.3 The card rows (PC 1, the 5090; the hash lane's job) @@ -350,4 +361,4 @@ PENDING. Per pack, unlocked and at the 1,300 knee when the helper answers: MH/s, | per-load shadow (`shl256x27`) | the same by construction, pending the row | the same | the same per joule | USD 200 M (one N5 die or an interposer forced) | pending | | tile block (`mm1430`) | pending | pending (the design target: the ALU shadow's) | at the ALU shadow's premium 2.1x / 1.6x / 1.3x, the k 0.3 column removed | USD 100 M (the tile array is USD 4 of N5) to 200 M with the per-load placement | pending | -The first sentence the coordinator asked for ("if either prototype beats class v4's premium or chip edge on measured rows") cannot be written until the PC 1 rows land; by construction neither lowers the premium (the per-load placement moves capex, the tile block moves the chip's k floor), so the honest expectation is: neither beats class v4's premium; the tile block beats its chip edge only against a pessimistic core (the k 0.3 column), and the per-load placement beats it only on capex. +The first sentence the coordinator asked for ("if either prototype beats class v4's premium or chip edge on measured rows") cannot be written until the PC 1 rows land; the verifier rows are in (20.2: the tile class at the ALU shadow's premium costs the verifier 3x class v4's shadow with AVX2 and misses the loaded gate by 0.14 ms on the box; the per-load class has a uniformity fault as drawn); by construction neither lowers the premium (the per-load placement moves capex, the tile block moves the chip's k floor), so the honest expectation is: neither beats class v4's premium; the tile block beats its chip edge only against a pessimistic core (the k 0.3 column), and the per-load placement beats it only on capex.