Counter ASIC 4.0 research: section 20, the two prototypes built (what, where, gates; the verifier and card rows pending)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-07 21:55:14 +00:00
parent 83ee0dd807
commit 1a7689e1b2

View file

@ -319,3 +319,35 @@ So the inverse lever's best design is the shadow per load: it costs the honest c
- The 5090's tensor free band (R about 2,000) is a scaling from the 4090's measured 40 percent at R = 512 and NVIDIA's claimed peaks.
- The shadow-per-load design's block-size effect at 16 instructions and its compile-ahead are unmeasured; the single-die project cost is the mission lane's N3 row, and a GDDR7 PHY on an N5 die is unpriced.
- The break-even model is the mission lane's (s = 0.30, a two-year life, the rental-equilibrium hash); its emission figures are that model's.
## 20. Third pass: the two prototypes, built (7 October 2026, 21:3x to 22:xx UTC)
The coordinator's order of 22:5x UK: both new ranks as experimental program classes in the kit, behind the pack, no consensus change. Built and gated tonight; the card rows are the hash lane's PC 1 job `run-ca4-pc1-packs-5090-20261007` (queued after the SM-sparse job and the microbench; the knee-lock states wait on PC 1's Power Helper, dead since 21:08 UTC, so the unlocked states land first).
### 20.1 What was built
| Item | Where | What it does | Gate |
|---|---|---|---|
| The per-load shadow, `+shl<S>x<R>` | `igneum-pow` (`ShadowClass.per_load`; `generator.rs`, `verify.rs`, `emit.rs`) | the class v4 shadow's 256 instructions and 27 passes, split into 16 sub-blocks of 16 consecutive instructions, sub-block `j` run 27 times right after the `j`-th load of the base program (the same work per hash, placed inside every read's dependency); id bytes `perload`; emitted in Metal, CUDA and OpenCL; the verifier runs it in the base loop | the base program and the shadow instructions are the class v4 program's draw for draw (unit test); the id differs; the suite green |
| The int8 tile block, `+mm<R>` | `igneum-pow` (`ShadowClass.tiles`, module `mm8`) | `R` `mma.m8n8k16 u8` tiles per iteration after the shadow block, each over two registers with both outputs consumed (the `proto-newpow/mma-shadow` form), descriptors drawn from the stream after the shadow draws; CUDA runs the PTX tile natively, Metal and OpenCL the shuffle-and-byte-product reference; the Rust verifier runs the scalar reference or an AVX2 path (`maddubs` on a 4-way split of B so no 16-bit lane saturates; `IGNEUM_MM8_SCALAR=1` forces scalar); id bytes `mm8/` plus the count | SIMD equal to scalar on 64 seeds and the all-ones edge (unit test, run on the EPYC); the suite green: 64 + 7 + 4 + 19 + 2 + 7 tests, 0 failed, igneum-build-2 21:42 UTC |
| Nothing moves for the chain | the packs test and a byte diff | the pinned v2, v3 and v4 packs are byte-identical; the re-exported `mx8+sh256x27` equals `packs-ca3-shadow/sh256x27` on all seven files | passed |
| The packs | `proto-cuda/packs-ca4/` (also build-1 `/srv/builds/ca4-research/packs/`, the hash lane's kit `igneum-ca4-packs-20261007.zip`) | `mx8_sh256x27` (the control, class v4's shape), `mx8_shl256x27`, `mx8_mm128`, `mx8_mm512`, `mx8_mm1430` (11,440 tiles per hash: the ALU shadow's premium in tiles at the 4090's 0.056 pJ per MAC), all over seed `igneum-genesis`, day 2026-10-03, 1 GiB, generator 2 | every export `OVERALL: PASS` (vectors through the interpreter); the CUDA-against-CPU bit-exactness is the worker's self-test on the card (96 vector lanes through the bound kernel), unrun |
| What is NOT built | | the NEON path of the tile verifier (the M5 Max runs scalar); the AVX-VNNI path (the AVX2 path stands in; VNNI would halve its instruction count, approximate); the acceptance rule's view of the per-load placement (the rule interprets the base program only, as for class v4; the sub-version 3 freshness fixpoint runs over base then shadow, which under per-load placement is not the execution order: a class, not a prototype, would re-derive it) | |
### 20.2 The verifier rows (igneum-build-2, EPYC 9454P, core 40 at nice 19, the ladder's method: 50 warps, alone and with the SMT sibling loaded by the same bench; the box under other lanes' load)
PENDING: the bench ran under `lease cores` at 21:5x UTC and its rows land here when the lease returns (the box's pool was fully leased and the lease waited). The reading they give: the per-load class should cost the verifier what class v4 costs (the same instructions); the tile classes add 1,024 byte products per tile per warp, about 1.07 microseconds per tile per unit scalar (the 6 October figure), so `mm1430` is about 12 ms scalar on an M5 Max core (FAILS) and the AVX2 row is the one that decides whether a tile class can sit under the 10 ms gate on a 2019-class core.
### 20.3 The card rows (PC 1, the 5090; the hash lane's job)
PENDING. Per pack, unlocked and at the 1,300 knee when the helper answers: MH/s, watts, SM MHz, the fingerprint and the self-test verdict. The questions each row answers: `shl256x27` against `sh256x27`: equal watts and a rate inside the block-size effect (16-instruction blocks ran 2.5 to 3.5 percent faster than 256 on 6 October) means the per-load placement is free for the GPU and the capex row of 16.2 stands; `mm128`, `mm512`, `mm1430` against `mx8-genesis`: the premium per tile on the 5090 (the 4090 read 0.056 pJ per MAC at 4,096 tiles per hash), where the rate falls, and whether 11,440 tiles per hash carries the ALU shadow's premium inside the free band; a self-test FAIL on a tile pack is a layout finding, not a tuning.
### 20.4 Into the chip model (rows filled when 20.2 and 20.3 land)
| Row | Honest 5090 energy per hash | Premium over class v3 | Chip at k = 1 / 1.5 / 2 (GDDR7) | Break-even cap, years 1 to 2 | Status |
|---|---|---|---|---|---|
| class v4 shape (`sh256x27`) | 2.34 (knee) / 3.48 (unlocked) microjoules measured | 81.8 W / 143.8 W | 2.1x / 1.6x (k 0.5: 2.9x; k 0.3: 4.1x) | USD 100 M (N5 shadow core) | measured |
| per-load shadow (`shl256x27`) | the same by construction, pending the row | the same | the same per joule | USD 200 M (one N5 die or an interposer forced) | pending |
| tile block (`mm1430`) | pending | pending (the design target: the ALU shadow's) | at the ALU shadow's premium 2.1x / 1.6x / 1.3x, the k 0.3 column removed | USD 100 M (the tile array is USD 4 of N5) to 200 M with the per-load placement | pending |
The first sentence the coordinator asked for ("if either prototype beats class v4's premium or chip edge on measured rows") cannot be written until the PC 1 rows land; by construction neither lowers the premium (the per-load placement moves capex, the tile block moves the chip's k floor), so the honest expectation is: neither beats class v4's premium; the tile block beats its chip edge only against a pessimistic core (the k 0.3 column), and the per-load placement beats it only on capex.