Counter ASIC 4.0 research: 20.3 the card rows of the seven packs (the tile block carries the ALU shadow's premium at the same rate, 14 percent cheaper at the knee; 0.048 to 0.091 pJ per MAC measured on the 5090), 20.4 the chip-model rows, the first sentence on measured rows
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
ac06ff9006
commit
8522caca0e
1 changed files with 36 additions and 9 deletions
|
|
@ -373,9 +373,28 @@ The fix of 20.2a held for distinctness and then met the value-level requirement
|
|||
|
||||
The crypto lane's finding: a product's low bits are biased (P(bit 0) = 1/4, measured exactly), the bias survives the odd stride multiplier, and the stride rotation places the biased bits at address bits R and up, inside the 28-bit item index unless R is 28 or more. The devnet era draws R = 29, which cuts them off, so 31 of 32 devnet-era programs read clean while 6 of 16 drawn-era programs (R from 3 to 22) show a site over 1.04x (2 over 1.2x, the worst 1.51x); under the 2 GiB genesis dataset (D = 29) R = 29 would show it too. The devnet's cleanliness is an accident of its era draw; the chain prevalence is the drawn-era figure; the price to a partial-store chip stays under 0.1 percent of a hash's reads per site, so no chip number moves. The requirement for any class this file proposes (the per-load shadow, the tile block, a re-weighted shadow): (1) a load whose source register's last writer is a product (`mul`, `mulhi`, `mad`) carries biased low bits into the address, and the acceptance must judge it at the VALUE level (the bit bias of the index at the site over the units), not by the index-distinctness ratio alone, which the duplicate-lane test above is; (2) the census of any candidate reads across the drawn eras split by R (3 to 22 against 28 to 31), as 20.2a now does for the per-load class, never the devnet era alone. The per-load class's dynamic rule covers distinctness, not bias; the value-level test is owed and is the same item for class v5's acceptance. Nothing in class v4 or v5 moves on this without main's word.
|
||||
|
||||
### 20.3 The card rows (PC 1, the 5090; the hash lane's job)
|
||||
### 20.3 The card rows (PC 1, the RTX 5090 alone, the installed worker, 04:04 to 04:22 UTC, 8 October 2026; the hash lane's job `run-ca4-pc1-packs-5090-20261008-b`, 250 batches of 2^24 at one warp per block, nvidia-smi at 1 Hz with the three power fields agreeing, the 1,300 MHz lock through the helper)
|
||||
|
||||
PENDING. Per pack, unlocked and at the 1,300 knee when the helper answers: MH/s, watts, SM MHz, the fingerprint and the self-test verdict. The questions each row answers: `shl256x27` against `sh256x27`: equal watts and a rate inside the block-size effect (16-instruction blocks ran 2.5 to 3.5 percent faster than 256 on 6 October) means the per-load placement is free for the GPU and the capex row of 16.2 stands; `mm128`, `mm512`, `mm1430` against `mx8-genesis`: the premium per tile on the 5090 (the 4090 read 0.056 pJ per MAC at 4,096 tiles per hash), where the rate falls, and whether 11,440 tiles per hash carries the ALU shadow's premium inside the free band; a self-test FAIL on a tile pack is a layout finding, not a tuning.
|
||||
Every pack's self-test PASS at both states (the cache, the dataset, the 96 vector lanes through the bound kernel), so the int8 tile's inline PTX compiles under NVRTC 12.8 on sm_120 and is bit-exact against the Rust verifier on the card, as the Metal reference was on the M5 Max; the 2^24 fingerprints are the same at both states and equal the Mac's where the Mac ran them (sh256x27 3d2e8245cc084d07, shl256x27_v2 ee5d7c71180e5ea7, mm128 270e4ae36b37e9a1, mm512 a1c1ff3148d775d1; mm1430 8e9b7066239d35d1 on CUDA, the Mac row not run).
|
||||
|
||||
| Pack | Unlocked (SM 2,842 to 2,865 MHz): MH/s / W / MH/W / microjoules per hash | Premium over mx8 unlocked | At the 1,300 lock (SM 1,290): MH/s / W / MH/W / microjoules | Premium at the lock | Label |
|
||||
|---|---|---|---|---|---|
|
||||
| mx8-genesis (class v3, the control) | 137.54 / 311.0 / 0.442 / 2.26 | | 127.32 / 213.0 / 0.598 / 1.67 | | measured |
|
||||
| mx8_sh256x27 (class v4's shape) | 137.51 / 462.2 / 0.298 / 3.36 | 151.2 W, 1.10 microjoules, 10.8 pJ per counted op | 126.93 / 295.8 / 0.429 / 2.33 | 82.8 W, 0.652 microjoules, 6.4 pJ per op | measured |
|
||||
| mx8_mm128 (1,024 tiles per hash) | 137.45 / 332.9 / 0.413 / 2.42 | 21.9 W: 0.152 pJ per MAC | 127.01 / 217.8 / 0.583 / 1.71 | 4.8 W: 0.036 pJ per MAC | measured |
|
||||
| mx8_mm512 (4,096 tiles) | 137.50 / 369.4 / 0.372 / 2.69 | 58.4 W: 0.101 pJ per MAC | 127.07 / 235.4 / 0.540 / 1.85 | 22.4 W: 0.042 pJ per MAC | measured |
|
||||
| **mx8_mm1430 (11,440 tiles, the ALU shadow's premium in tiles)** | 137.45 / 457.7 / 0.300 / 3.33 | **146.7 W, 1.067 microjoules: 0.091 pJ per MAC**, 0.103 W per tile per iteration, linear within 5 percent | 126.87 / 284.5 / 0.446 / 2.24 | **71.5 W, 0.564 microjoules: 0.048 pJ per MAC** (0.050 W per tile) | measured |
|
||||
| mx8_shl256x27 (the first per-load export, 854050a4293f0615: UNSOUND, energy reading only) | 158.62 / 472.8 / 0.336 | +161.8 W at a rate 15 percent OVER the control (the 11 percent duplicate reads land in L2) | 145.65 / 299.4 / 0.487 | +86.4 W | measured; the construction is dead (20.2a-close) |
|
||||
| mx8_shl256x27_v2 (the fixed export, bd64b207a30413fb: UNSOUND, energy reading only) | 135.90 / 448.3 / 0.303 | +137.3 W, 14 W under the whole-block shape at the same instruction count (the 16-instruction block effect) | 126.04 / 282.9 / 0.446 | +69.9 W, 13 W under the whole block | measured; dead as a class |
|
||||
|
||||
NVRTC compile per pack: mx8 190 ms, sh256x27 269, mm128 331, mm512 869, mm1430 1,239 unlocked and 2,107 at the lock (inside the epoch's compile-ahead budget, the 600-s floor's 38 s). The lock rows read 127 MH/s against the efficiency pass's 134 at the same clock: the installed worker at one warp per block against the app's tuned variant; every ratio here is within one run.
|
||||
|
||||
What the rows say:
|
||||
|
||||
1. **The hash rate is memory-bound on every sound pack at both states** (137.4 to 137.5 unlocked, 126.9 to 127.3 at the lock, within 0.5 percent of the control): 11,440 int8 tiles per hash are free in rate on the 5090, so the 4090's free band (R = 512) extends to at least R = 1,430 on the 5090, as section 15.2 scaled.
|
||||
2. **The tile block carries the ALU shadow's premium at the same hash rate and at a lower cost at the knee**: 146.7 W against 151.2 W unlocked (3 percent under), 71.5 W against 82.8 W at the 1,300 lock (14 percent under). The 5090's energy per MAC at full tile load is 0.091 pJ unlocked and 0.048 pJ at the lock (the 4090's marginal 0.056 pJ at R = 512 sits between), with the fixed cost of the tensor path visible at R = 128 (0.152 pJ unlocked).
|
||||
3. **The GPU's measured cost per MAC is inside the band of what a 5 nm MAC array costs anyone** (NVIDIA's own test chip at 0.46 V: 0.021 pJ per INT4 MAC, about 0.04 to 0.08 per INT8-class MAC; at nominal voltage about 5x that; JSSC 2023 via Dally's slides, claimed), so the chip's `k` on this work reads about 0.9 to 1.7 at the chip's lowest voltage and 4 to 8 at nominal (approximate); the honest centre is at or above 1, where the ALU shadow's honest band is 0.3 to 0.8.
|
||||
4. The per-load placement is cheaper per instruction than the whole block (the 16-instruction block effect: 13 to 14 W under the whole block for the same 55,296 instructions per hash) and the construction is dead (20.2a-close); the first export's 15 percent rate gain is the duplicate reads served from L2, the fault made visible on the card.
|
||||
|
||||
### 20.3a The SM-sparse job's first run (PC 1, 00:22 to 00:41 UTC, 8 October 2026): no variant ran; the knee rows it did read
|
||||
|
||||
|
|
@ -383,12 +402,20 @@ The hash lane's run `run-ca4-pc1-ca4sparse-5090-20261007` (the ca4sparse exe, th
|
|||
|
||||
What the run did read, and keeps (the 5090 at the 1,300 MHz lock, base kernel, measured): class v4 134.26 MH/s at 309.9 W (0.433 MH/W, SM 1,290 MHz), class v3 134.03 at 219.4 W (0.611 MH/W); the class v4 premium 90.5 W at the lock against 133.9 W unlocked on this run (145.3 on the efficiency pass; the unlocked v4 rows ran hot, 465 to 485 W). The premium per counted op at the lock: 90.5 W over 134.26 M x 102,100 = 6.6 pJ (6.5 on the efficiency pass).
|
||||
|
||||
### 20.4 Into the chip model (rows filled when 20.2 and 20.3 land)
|
||||
### 20.4 Into the chip model (the rows measured; the chip side modelled: f = 1 GDDR7 0.466 microjoules per hash, one HBM3 stack 0.321)
|
||||
|
||||
| Row | Honest 5090 energy per hash | Premium over class v3 | Chip at k = 1 / 1.5 / 2 (GDDR7) | Break-even cap, years 1 to 2 | Status |
|
||||
|---|---|---|---|---|---|
|
||||
| class v4 shape (`sh256x27`) | 2.34 (knee) / 3.48 (unlocked) microjoules measured | 81.8 W / 143.8 W | 2.1x / 1.6x (k 0.5: 2.9x; k 0.3: 4.1x) | USD 100 M (N5 shadow core) | measured |
|
||||
| per-load shadow (`shl256x27`) | the same by construction, pending the row | the same | the same per joule | USD 200 M (one N5 die or an interposer forced) | pending |
|
||||
| tile block (`mm1430`) | pending | pending (the design target: the ALU shadow's) | at the ALU shadow's premium 2.1x / 1.6x / 1.3x, the k 0.3 column removed | USD 100 M (the tile array is USD 4 of N5) to 200 M with the per-load placement | pending |
|
||||
Energy per hash = the card's measured; the chip's = `E_mem + k x F`, `F` the measured premium per hash, `k` the chip's cost per forced operation over the 5090's at the same state.
|
||||
|
||||
The first sentence the coordinator asked for ("if either prototype beats class v4's premium or chip edge on measured rows") cannot be written until the PC 1 rows land; the verifier rows are in (20.2: the tile class at the ALU shadow's premium costs the verifier 3x class v4's shadow with AVX2 and misses the loaded gate by 0.14 ms on the box; the per-load class has a uniformity fault as drawn); by construction neither lowers the premium (the per-load placement moves capex, the tile block moves the chip's k floor), so the honest expectation is: neither beats class v4's premium; the tile block beats its chip edge only against a pessimistic core (the k 0.3 column), and the per-load placement beats it only on capex.
|
||||
| Row (the 5090) | Card microjoules | Premium F | Chip edge, GDDR7, at k = 0.3 / 0.5 / 1 / 1.5 / 2 | HBM3 one stack at k = 1 | The honest k band for the work | Break-even cap, years 1 to 2 (the mission lane's model) | Status |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| class v3, unlocked | 2.26 | 0 | 4.85x at every k | 7.0x | | USD 17 M (28 nm controller) | measured card |
|
||||
| class v3, 1,300 lock | 1.67 | 0 | 3.59x | 5.2x | | USD 17 M | measured |
|
||||
| class v4 shape, unlocked | 3.36 | 1.10 | 4.21x / 3.31x / 2.15x / 1.59x / 1.26x | 2.37x | ALU: 0.3 to 0.8 | USD 100 M (an N5 shadow core) | measured; the served figures 2.1x at k = 1, 3.4x pessimistic |
|
||||
| class v4 shape, 1,300 lock | 2.33 | 0.652 | 3.52x / 2.94x / 2.08x / 1.62x / 1.32x | 2.39x | ALU: 0.3 to 0.8 | USD 100 M | measured |
|
||||
| **tile block mm1430, unlocked** | 3.33 | 1.067 | 4.23x / 3.33x / **2.17x** / 1.61x / 1.28x | 2.40x | **int8 MAC: 0.9 to 1.7 at the chip's lowest voltage, 4 to 8 at nominal (approximate)**: the k 0.3 and 0.5 columns are not reachable on this work | USD 100 M (a licensable MMA block is USD 4 of N5; the N5 die is forced as for class v4) | measured card |
|
||||
| **tile block mm1430, 1,300 lock** | 2.24 | 0.564 | 3.46x / 2.98x / **2.17x** / 1.73x / 1.42x | 2.53x | the same | USD 100 M | measured |
|
||||
| the per-load shape (the fixed export) | 3.30 / 2.24 | 0.998 / 0.555 | as the class v4 shape within 5 percent | | ALU | USD 200 M if the sound form exists (16.2) | dead as drawn (20.2a-close) |
|
||||
|
||||
Reading: at `k = 1` the tile block and the ALU shadow give the same 2.1x to 2.2x, and the tile block's premium at the knee is 14 percent lower for it; the difference is the chip's reachable `k`. On the ALU shadow a fixed SIMD array can plausibly reach `k` 0.3 to 0.5 (the served pessimistic 3.4x to 3.5x); on the int8 tile the GPU's own tensor core is the dense array, measured at 0.048 to 0.091 pJ per MAC, and a chip at the same node cannot be 2x to 3x cheaper per MAC at the same voltage, so the chip's downside column collapses: the tile block's edge at the measured premium reads 2.2x at `k = 1` and under 1.8x at any `k` over 1.3, against the ALU shadow's 3.5x at `k` 0.3. That is the finding of the k hunt made measured on the card: **the tile block does not beat class v4's premium (it matches it at the same rate, 14 percent cheaper at the knee) and it beats class v4's chip edge at the pessimistic end (about 2.2x against 3.5x) and not at `k = 1`.** What stands against it as a class: the verifier (AVX2 0.047 microseconds per tile per unit, `mm1430` 5.8 ms alone and 10.14 ms with the sibling loaded on the box's core, a 0.14 ms miss at load 84 to 95; scalar 13x worse; NEON unwritten), the Apple tier (the Metal emulation at 1,024 tiles costs 35 percent of the M5 Max's rate and 4,096 costs 78 percent, so 11,440 is out of reach without an integer matrix path in Metal), the AMD WMMA layout (unverified; the OpenCL reference path emitted but unmeasured on the 9070 XT), and the compile-ahead (1.2 to 2.1 s per pack on the 5090, inside the budget).
|
||||
|
||||
**The first sentence, on measured rows (04:22 UTC, 8 October 2026): neither prototype beats class v4's premium; the tile block matches it at the same hash rate (146.7 W against 151.2 W unlocked, 71.5 against 82.8 W at the 1,300 lock) and beats class v4's chip edge only at the pessimistic end (about 2.2x against 3.5x, because a chip's int8 MAC cannot undercut the 5090's measured 0.048 to 0.091 pJ per MAC the way a fixed datapath undercuts its 6.4 to 10.8 pJ per ALU op), not at k = 1 (2.2x either way); the per-load placement is dead as a construction, and its energy rows read 13 to 14 W under the whole block for the same instructions.** The SM-sparse reading is still owed (three failed runs: the race switch, the race list, the rewrite's capture; the fourth exe checked through NVRTC card-free, queued).
|
||||
|
|
|
|||
Loading…
Reference in a new issue