Counter ASIC 4.0 research: 15.1a the microbench rows (20 probes on the 5090 at both states) and the correction they force: a tile is 32 MACs per lane, not 1,024, so the int8 MAC costs 1.5 to 4 pJ measured (the 6 October 4090 figure 1.8, not 0.056); the tensor shadow is the worse lever (k 0.03 to 0.3); the shuffle at 55.8 pJ per op is the card's dearest instruction (the re-weight's GPU side against it); an L2 hit at 2.4 nJ kills the hot-table lever; the first sentence and the ranking corrected

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-josh 2026-10-08 07:02:04 +01:00
parent 954c4053fc
commit 71fd465b78

View file

@ -249,6 +249,49 @@ The GPU column is the public figure or this project's measurement until the micr
The reading: a GPU lane's cost per operation is overhead (fetch, decode, operand delivery, the banked register file) around a datapath that is a few percent of it, and a chip at the same node builds the same datapath and may drop part of the overhead, which is why every ALU-shaped row reads `k` 0.3 to 0.8. The two rows that read `k` near or above 1 are the ones where the GPU's block is itself a dense array with no per-op overhead to drop (the int8 tensor tile) or a large SRAM whose cost is wires that a chip's SRAM has too (the L2). The tensor tile is the only one of the two with joules to force at a verifiable cost, and only with a SIMD verifier.
### 15.1a The microbench rows (PC 1, the RTX 5090 alone, 05:14 to 05:56 UTC, 8 October 2026 (06:14 to 06:56 UK); the hash lane's job `run-ca4-pc1-microbench-5090-20261008-c`, the ca4mb exe, 20 probes at 60 s each at full residency, "20 probes ran, 0 skipped or failed" at both states; the sampler's mean from 8 s in, the three power fields within 0.2 W; idle 74.7 W unlocked and 60.2 W at the lock; every checksum equal at both states)
The reading per row is (watts minus the `sleep` row) over counted operations per second: picojoules per counted op on this card. The `sleep` row is NOT idle: 120.2 W unlocked and 66.2 at the lock against 74.7 and 60.2 idle, so full residency with every warp spinning on `__nanosleep` costs 45 W at the stock clock before any instruction issues; that residency cost is subtracted from every row below, which makes each figure the op's own marginal.
| Probe (what one counted op is) | Unlocked: W / SM MHz / G ops per s / pJ per op | At the 1,300 lock: W / G ops per s / pJ per op | Label |
|---|---|---|---|
| sleep (full residency, no issue) | 120.2 / 2,872 / 0 / the floor | 66.2 / 0 / the floor | measured |
| int_arx (int32 add, xor, rotate; 4 chains) | 533.9 / 2,816 / 36,600 / **11.3** | 170.9 / 16,819 / **6.2** | measured; the class v4 shadow read 10.8 and 6.4 pJ per counted op on the packs job: the two instruments agree |
| int_mul (int32 multiply-add; 4 chains) | 558.2 / 2,712 / 31,419 / 13.9 | 194.5 / 15,412 / 8.3 | measured |
| int_mulhi (4 chains) | 538.5 / 2,805 / 10,573 / 39.6 | 168.6 / 4,871 / 21.0 | measured (a multi-instruction sequence per op on this card) |
| prmt (byte permute; 4 chains) | 498.4 / 2,812 / 16,933 / 22.3 | 156.6 / 7,882 / 11.5 | measured |
| lop3 (three-input logic; 4 chains) | 454.6 / 2,820 / 13,896 / 24.1 | 149.4 / 6,391 / 13.0 | measured (counted as one op per chain-step; the compiler's fusion is not verified) |
| **shfl (warp shuffle, one xor beside it; 4 chains)** | 512.6 / 2,811 / 7,036 / **55.8** | 161.3 / 3,240 / **29.4** | measured: the most expensive operation on the card per op, 5x the add |
| fp32_fma (4 chains) | 522.7 / 2,714 / 43,671 / 9.2 | 180.7 / 21,863 / 5.2 | measured |
| fp16x2_fma (two half FMAs per instruction) | 365.1 / 2,829 / 48,067 / 5.1 per half FMA | 125.6 / 22,696 / 2.6 | measured |
| **mma_u8_m8n8k16 (int8 MAC inside the tile; 32 per lane per tile, two tiles per step)** | 537.4 / 2,776 / 101,732 / **4.1 per MAC** | 169.7 / 47,555 / **2.2** | measured |
| **mma_s8_m16n8k32 (int8 MAC; 128 per lane per tile)** | 575.0 / 2,629 / 334,227 (80 percent of the dense INT8 peak, approximate) / **1.36 per MAC** | 205.9 / 168,252 / **0.83** | measured: the cheapest counted op on the card, 8x under the ARX path |
| mma_f16_m16n8k16 (FP16 MAC, FP32 accumulate) | 454.5 / 2,819 / 104,798 / 3.2 | 149.7 / 48,264 / 1.7 | measured (not bit-exact across vendors: energy only) |
| mma_bf16_m16n8k16 | 425.3 / 2,820 / 104,953 / 2.9 | 137.8 / 48,320 / 1.5 | measured (energy only) |
| mma_e4m3_m16n8k32 (FP8) | 439.3 / 2,820 / 211,028 / 1.5 | 143.9 / 97,101 / 0.8 | measured (energy only) |
| l2_chase_32m (dependent 4-byte read, L2-resident; per read) | 378.7 / 2,842 / 107.55 G reads per s / **2.4 nJ per read** | 138.3 / 52.13 / 1.4 nJ | measured |
| l2_chase_64m | 381.3 / 2,842 / 107.64 / 2.4 nJ | 140.5 / 52.14 / 1.4 nJ | measured (64 MiB still inside the 5090's L2) |
| l2_indep4_32m (4 independent chains) | 377.0 / 2,842 / 109.42 / 2.3 nJ | 138.6 / 52.96 / 1.4 nJ | measured |
| dram_chase_1g (the hash's own pattern; per read) | 318.4 / 2,850 / 18.17 G reads per s / **10.9 nJ per read** | 200.0 / 15.38 / 8.7 nJ | measured: the whole card's marginal per dependent DRAM read, against the memory system's modelled 2.0 nJ |
| tex_point_u32_32m (per fetch) | 373.0 / 2,842 / 107.65 / 2.3 nJ | 138.6 / 52.20 / 1.4 nJ | measured (the same checksum as the L2 chase: the sampler's address path is the L2 chase) |
| tex_linear_f32_256k (per filtered fetch, 4 chains) | 298.6 / 2,850 / 923.8 / 0.19 nJ | 101.0 / 413.3 / 0.10 nJ | measured (energy only; not bit-exact across vendors) |
**A correction these rows force on sections 15.2, 20.3 and 20.4, and on the 6 October tensor figure.** A `mma.m8n8k16` tile is one warp instruction over 32 lanes: 1,024 multiply-adds per WARP, 32 per lane, so a hash (one lane) does 32 MACs per tile, not 1,024. Section 15.2 and 20.3 multiplied the per-hash tile count by 1,024, and the 6 October figure for the 4090 (new-pow 5.1: "0.056 pJ per multiply-add at 4,096 tiles per hash") carries the same factor of 32: at 63.08 MH/s and 4,096 tiles per hash the 4090 did 8.3 x 10^12 MACs per second, not 2.6 x 10^14, and its 14.7 W is 1.8 pJ per MAC, not 0.056. Re-read with the right count, the packs job's tile rows (20.3) are: `mm1430` 11,440 tiles x 32 = 366,080 MACs per hash, 5.03 x 10^13 MACs per second at 137.45 MH/s, 146.7 W = **2.9 pJ per MAC unlocked** and 71.5 W over 4.64 x 10^13 = **1.5 pJ per MAC at the lock**; the microbench's dependent u8 tile reads 4.1 and 2.2, its s8 m16n8k32 at 80 percent of peak 1.36 and 0.83. The int8 MAC on the 5090 therefore costs 1 to 4 pJ, between a tenth and a third of a 32-bit ALU op (6 to 11 pJ), not a hundredth; and NVIDIA's own 5 nm MAC array (0.04 to 0.1 pJ per INT8-class MAC at 0.46 V, about 0.2 to 0.4 at nominal, claimed) is then 4x to 30x cheaper than the 5090's measured tile, not equal to it. **The chip's `k` on tile work reads about 0.03 to 0.3, below the ALU shadow's 0.3 to 0.8, so the tensor shadow is the WORSE forcing lever, and the k-above-1 candidate of section 15 does not exist on measured rows.** The corrected 20.4 is below; the 6 October 4090 figure is corrected in this file and owed to `docs/analysis/horizon/new-pow.md` 5.1 and `chip-model-v3.md` 5.11 (the Counter lane).
What the other rows say, in the identity's terms (the chip's cost per op from the public figures of 15.1, the GPU's now measured):
| Block | GPU, measured (unlocked / lock) | Chip at 4 to 5 nm, public | k band, corrected | Reading |
|---|---|---|---|---|
| int32 ALU (the class v4 shadow's mix) | 11.3 / 6.2 pJ | 2 to 5 pJ for a SIMD array with its register file (approximate) | 0.3 to 0.8 | the best forcing lever the card has; unchanged |
| warp shuffle | 55.8 / 29.4 pJ | a 32-lane crossbar of a 32-bit word across about 1 mm: about 20 pJ (approximate, the wire figure of 15.1) | 0.4 to 0.7, with the GPU paying 5x the add per op | a shuffle-heavy re-weight raises the GPU's premium per instruction 5x for a k no better than the add's: the re-weight's case is now AGAINST it on the measured GPU side (section 17 rank 5 and the held item) |
| int32 multiply | 13.9 / 8.3 pJ | the same multiplier, 0.5 pJ datapath (claimed) | 0.3 to 0.8 | as the ALU row |
| int8 tile | 2.9 to 4.1 / 1.5 to 2.2 pJ per MAC (u8); 1.36 / 0.83 (s8 m16n8k32) | 0.04 to 0.4 pJ per MAC (claimed) | **0.03 to 0.3** | the worse lever: the chip's array undercuts the GPU's tensor core by 4x to 30x per MAC |
| L2 hit | 2.4 / 1.4 nJ per read | an on-die SRAM read 0.2 to 0.5 nJ (approximate) | **0.1 to 0.3** | the hot-table lever (design 4, the ldcs measurement) is dead on the GPU side: an L2 hit costs the card 5x to 10x what a chip's SRAM costs |
| texture interpolation | 0.19 / 0.10 nJ per fetch | a 9-bit interpolator: picojoules | far under 1 | excluded anyway (not bit-exact across vendors) |
| DRAM dependent read | 10.9 / 8.7 nJ per read (the whole card's marginal) | 2.0 (GDDR7) to 1.2 (HBM3) nJ per read (modelled) | 0.1 to 0.2 on the marginal; this is the `E_card / E_mem` of section 2 seen per read | the premium-free floor, measured per read: at the lock the card pays 8.7 nJ for a read the chip's memory pays 2.0 for |
**The one-sentence answer of section 15, corrected on measured rows: nothing on the 5090 reads k above 1; the int8 tile, the one block the public figures put near parity, is 1 to 4 pJ per MAC measured against a 5 nm array's claimed 0.04 to 0.4, so it is the worst lever of all, and the ALU shadow (k 0.3 to 0.8 on 6 to 11 pJ per op) stays the best forcing work the card has.**
### 15.2 The tensor shadow, re-read on the identity
The 6 October verdict on scheme B ("never as class v5 content") was right for the question it answered: at the 4090's free band (R = 512, 0.23 microjoules) the block forced too few joules and the verifier paid 26x the ALU shadow's cost per joule. The founder's question is a different one: not "is there a cheaper lever" but "is there a lever whose joules the chip cannot undercut". On that question the tensor tile is the best block on the card, because the GPU's marginal 0.056 pJ per MAC is within a factor of about 2 of what a 5 nm MAC array costs anyone (the test chip's 0.04 to 0.1 pJ, claimed), while the ALU shadow's 6.5 to 10.4 pJ per op is 10x to 50x what a fixed SIMD datapath costs. What it would take to make the tensor block carry the SAME premium as the ALU shadow (0.654 microjoules at the lock), on the 5090:
@ -302,10 +345,10 @@ So the inverse lever's best design is the shadow per load: it costs the honest c
|---|---|---|---|---|---|---|---|
| 1 | The operating point as the shipped default | class v3 3.6x; class v4 2.1x at k = 1 | lowers the base (330 to 223 W); the v4 premium 145 to 82 W | none | 2 to 4 | AMD and Apple have no lock | unchanged |
| 2 | The SM-sparse miner kernel | **DEAD on measured rows (20.3b, 05:12 UTC): a quarter of the SMs holds 98 percent of the rate at the same draw; the draw follows the work, not the SM count** | none | none | 1 + 1 + 4 | measured: the energy per hash never falls below base | closed |
| 3 | **The tensor shadow with a SIMD byte-dot verifier** (int8 tiles carrying the premium; a class change) | at the ALU shadow's premium: 2.1x at k = 1, 1.6x at k = 1.5, 1.3x at k = 2; the chip's k 0.3 downside removed | the same as class v4 by construction / the 4070 tensor-bound above about R 600 (approximate) | scalar FAILS; AVX2 and VNNI 2.9 to 5.1 ms on the M5 Max core, 7 to 13 ms on a 2019 core (approximate, unwritten) | 8 to 12 (the SIMD verifier, the AMD layout gate, Apple's emulation row) plus the six gates | the 2019-class core; the AMD layout; Apple loses 5 to 10 percent of rate on emulation; k near 1 is claimed, not measured | NEW this pass: the k-above-1 candidate |
| 3 | **The tensor shadow with a SIMD byte-dot verifier** (int8 tiles carrying the premium; a class change) | **DEAD on measured rows (15.1a, 20.4): the 5090's int8 MAC is 1.5 to 4 pJ measured, a 5 nm array 0.04 to 0.4 claimed, k 0.03 to 0.3, under the ALU shadow's band; at the measured premium the chip keeps 3.5x to 6.7x** | the same as class v4 by construction / the 4070 tensor-bound above about R 600 (approximate) | scalar FAILS; AVX2 and VNNI 2.9 to 5.1 ms on the M5 Max core, 7 to 13 ms on a 2019 core (approximate, unwritten) | 8 to 12 (the SIMD verifier, the AMD layout gate, Apple's emulation row) plus the six gates | the 2019-class core; the AMD layout; Apple loses 5 to 10 percent of rate on emulation; k near 1 is claimed, not measured | NEW this pass: the k-above-1 candidate |
| 4 | The shadow per load (capex) | unchanged in energy | none by construction (block-size effect unmeasured at 16) | unchanged | 4 to 6 plus the gates | compile-ahead at 16 sites; a class change; **the 16 x 27 construction is DEAD (20.2a-close, 22:53 UTC): the acceptance rule in execution order accepts 1.4 percent of its candidates; the sound form (one pass of 432 per load) is undrawn** | NEW: doubles the break-even cap to about USD 200 M, on a construction not yet shown to exist |
| 5 | The ALU shadow re-weighted toward shuffles and multiplies | 2.1x at k = 1; 2.9x at the pessimistic k 0.46 | unchanged | unchanged | 4 to 6 plus the gates | Apple pays shfl 1.91x | down from 3: the tensor tile bounds k better |
| 6 | The L2-resident hot table with cache hints | energy 0.1 microjoules at k 0.7 to 3; capex +USD 15 | 13 W if the hint holds the rate | under 0.5 ms (unmeasured) | 4 for the ldcs measurement (queued) | Metal has no hint; the rate without it | unchanged; a measurement |
| 5 | The ALU shadow re-weighted toward shuffles and multiplies | 2.1x at k = 1; 2.9x at the pessimistic k 0.46 (modelled) | **raised: the 5090 pays 55.8 pJ per shuffle against 11.3 per add (15.1a), so a shuffle-heavy mix at the same instruction count costs the card up to 5x the premium per instruction** | unchanged | 4 to 6 plus the gates | Apple pays shfl 1.91x; the GPU side measured against it | HELD; the measured GPU side is against it |
| 6 | The L2-resident hot table with cache hints | **DEAD on the GPU side (15.1a): an L2 hit costs the 5090 2.4 nJ unlocked and 1.4 at the lock against a chip's SRAM read at 0.2 to 0.5 nJ, k 0.1 to 0.3**; capex +USD 15 | 13 W if the hint holds the rate | under 0.5 ms (unmeasured) | 4 for the ldcs measurement (queued) | Metal has no hint; the rate without it | unchanged; a measurement |
| 7 to 11 | class v5 without the shadow; the refresh as a cost; per-card classes; proof of useful work; memory shaping | as section 9 | | | 0 | dead by arithmetic | unchanged |
## 18. What was built in the second pass: the microbench
@ -381,9 +424,9 @@ Every pack's self-test PASS at both states (the cache, the dataset, the 96 vecto
|---|---|---|---|---|---|
| mx8-genesis (class v3, the control) | 137.54 / 311.0 / 0.442 / 2.26 | | 127.32 / 213.0 / 0.598 / 1.67 | | measured |
| mx8_sh256x27 (class v4's shape) | 137.51 / 462.2 / 0.298 / 3.36 | 151.2 W, 1.10 microjoules, 10.8 pJ per counted op | 126.93 / 295.8 / 0.429 / 2.33 | 82.8 W, 0.652 microjoules, 6.4 pJ per op | measured |
| mx8_mm128 (1,024 tiles per hash) | 137.45 / 332.9 / 0.413 / 2.42 | 21.9 W: 0.152 pJ per MAC | 127.01 / 217.8 / 0.583 / 1.71 | 4.8 W: 0.036 pJ per MAC | measured |
| mx8_mm512 (4,096 tiles) | 137.50 / 369.4 / 0.372 / 2.69 | 58.4 W: 0.101 pJ per MAC | 127.07 / 235.4 / 0.540 / 1.85 | 22.4 W: 0.042 pJ per MAC | measured |
| **mx8_mm1430 (11,440 tiles, the ALU shadow's premium in tiles)** | 137.45 / 457.7 / 0.300 / 3.33 | **146.7 W, 1.067 microjoules: 0.091 pJ per MAC**, 0.103 W per tile per iteration, linear within 5 percent | 126.87 / 284.5 / 0.446 / 2.24 | **71.5 W, 0.564 microjoules: 0.048 pJ per MAC** (0.050 W per tile) | measured |
| mx8_mm128 (1,024 tiles per hash, 32,768 MACs) | 137.45 / 332.9 / 0.413 / 2.42 | 21.9 W: 4.9 pJ per MAC | 127.01 / 217.8 / 0.583 / 1.71 | 4.8 W: 1.2 pJ per MAC | measured (the per-MAC figures corrected 8 October 2026, 06:xx UTC: 32 MACs per lane per tile) |
| mx8_mm512 (4,096 tiles, 131,072 MACs) | 137.50 / 369.4 / 0.372 / 2.69 | 58.4 W: 3.2 pJ per MAC | 127.07 / 235.4 / 0.540 / 1.85 | 22.4 W: 1.3 pJ per MAC | measured (corrected) |
| **mx8_mm1430 (11,440 tiles, 366,080 MACs per hash, the ALU shadow's premium in tiles)** | 137.45 / 457.7 / 0.300 / 3.33 | **146.7 W, 1.067 microjoules: 2.9 pJ per MAC**, 0.103 W per tile per iteration, linear within 5 percent | 126.87 / 284.5 / 0.446 / 2.24 | **71.5 W, 0.564 microjoules: 1.5 pJ per MAC** (0.050 W per tile) | measured (corrected: the first reading divided by 1,024 MACs per lane per tile where the tile gives 32) |
| mx8_shl256x27 (the first per-load export, 854050a4293f0615: UNSOUND, energy reading only) | 158.62 / 472.8 / 0.336 | +161.8 W at a rate 15 percent OVER the control (the 11 percent duplicate reads land in L2) | 145.65 / 299.4 / 0.487 | +86.4 W | measured; the construction is dead (20.2a-close) |
| mx8_shl256x27_v2 (the fixed export, bd64b207a30413fb: UNSOUND, energy reading only) | 135.90 / 448.3 / 0.303 | +137.3 W, 14 W under the whole-block shape at the same instruction count (the 16-instruction block effect) | 126.04 / 282.9 / 0.446 | +69.9 W, 13 W under the whole block | measured; dead as a class |
@ -392,8 +435,8 @@ NVRTC compile per pack: mx8 190 ms, sh256x27 269, mm128 331, mm512 869, mm1430 1
What the rows say:
1. **The hash rate is memory-bound on every sound pack at both states** (137.4 to 137.5 unlocked, 126.9 to 127.3 at the lock, within 0.5 percent of the control): 11,440 int8 tiles per hash are free in rate on the 5090, so the 4090's free band (R = 512) extends to at least R = 1,430 on the 5090, as section 15.2 scaled.
2. **The tile block carries the ALU shadow's premium at the same hash rate and at a lower cost at the knee**: 146.7 W against 151.2 W unlocked (3 percent under), 71.5 W against 82.8 W at the 1,300 lock (14 percent under). The 5090's energy per MAC at full tile load is 0.091 pJ unlocked and 0.048 pJ at the lock (the 4090's marginal 0.056 pJ at R = 512 sits between), with the fixed cost of the tensor path visible at R = 128 (0.152 pJ unlocked).
3. **The GPU's measured cost per MAC is inside the band of what a 5 nm MAC array costs anyone** (NVIDIA's own test chip at 0.46 V: 0.021 pJ per INT4 MAC, about 0.04 to 0.08 per INT8-class MAC; at nominal voltage about 5x that; JSSC 2023 via Dally's slides, claimed), so the chip's `k` on this work reads about 0.9 to 1.7 at the chip's lowest voltage and 4 to 8 at nominal (approximate); the honest centre is at or above 1, where the ALU shadow's honest band is 0.3 to 0.8.
2. **The tile block carries the ALU shadow's premium at the same hash rate and at a lower cost at the knee**: 146.7 W against 151.2 W unlocked (3 percent under), 71.5 W against 82.8 W at the 1,300 lock (14 percent under). The 5090's energy per MAC at full tile load is 2.9 pJ unlocked and 1.5 pJ at the lock (the 4090's 6 October figure, re-read with 32 MACs per lane per tile, is 1.8 pJ), with the fixed cost of the tensor path visible at R = 128 (4.9 pJ unlocked). The microbench (15.1a) reads the same path at 4.1 and 2.2 pJ per MAC on a dependent u8 chain and 1.36 and 0.83 on the wide s8 tile at 80 percent of peak.
3. **The GPU's measured cost per MAC is NOT inside the band of what a 5 nm MAC array costs** (NVIDIA's own test chip at 0.46 V: 0.021 pJ per INT4 MAC, about 0.04 to 0.08 per INT8-class MAC; at nominal voltage about 5x that; JSSC 2023 via Dally's slides, claimed): against the 5090's 1.5 to 2.9 pJ per MAC the chip's `k` on tile work reads about 0.03 to 0.3 (approximate), BELOW the ALU shadow's 0.3 to 0.8. The earlier reading of this row (k 0.9 to 1.7) rested on the factor-of-32 error corrected in 15.1a.
4. The per-load placement is cheaper per instruction than the whole block (the 16-instruction block effect: 13 to 14 W under the whole block for the same 55,296 instructions per hash) and the construction is dead (20.2a-close); the first export's 15 percent rate gain is the duplicate reads served from L2, the fault made visible on the card.
### 20.3a The SM-sparse job's first run (PC 1, 00:22 to 00:41 UTC, 8 October 2026): no variant ran; the knee rows it did read
@ -429,10 +472,10 @@ Energy per hash = the card's measured; the chip's = `E_mem + k x F`, `F` the mea
| class v3, 1,300 lock | 1.67 | 0 | 3.59x | 5.2x | | USD 17 M | measured |
| class v4 shape, unlocked | 3.36 | 1.10 | 4.21x / 3.31x / 2.15x / 1.59x / 1.26x | 2.37x | ALU: 0.3 to 0.8 | USD 100 M (an N5 shadow core) | measured; the served figures 2.1x at k = 1, 3.4x pessimistic |
| class v4 shape, 1,300 lock | 2.33 | 0.652 | 3.52x / 2.94x / 2.08x / 1.62x / 1.32x | 2.39x | ALU: 0.3 to 0.8 | USD 100 M | measured |
| **tile block mm1430, unlocked** | 3.33 | 1.067 | 4.23x / 3.33x / **2.17x** / 1.61x / 1.28x | 2.40x | **int8 MAC: 0.9 to 1.7 at the chip's lowest voltage, 4 to 8 at nominal (approximate)**: the k 0.3 and 0.5 columns are not reachable on this work | USD 100 M (a licensable MMA block is USD 4 of N5; the N5 die is forced as for class v4) | measured card |
| **tile block mm1430, 1,300 lock** | 2.24 | 0.564 | 3.46x / 2.98x / **2.17x** / 1.73x / 1.42x | 2.53x | the same | USD 100 M | measured |
| **tile block mm1430, unlocked** | 3.33 | 1.067 | 4.23x / 3.33x / 2.17x / 1.61x / 1.28x; **at the measured k 0.03 to 0.3: 4.2x to 6.7x** | 2.40x | **int8 MAC: 0.03 to 0.3 (corrected 15.1a)**: the chip's array undercuts the tensor core 4x to 30x per MAC | USD 100 M (a licensable MMA block is USD 4 of N5; the N5 die is forced as for class v4) | measured card; the lever is WORSE than the ALU shadow |
| **tile block mm1430, 1,300 lock** | 2.24 | 0.564 | 3.46x / 2.98x / 2.17x / 1.73x / 1.42x; **at k 0.03 to 0.3: 3.5x to 4.6x** | 2.53x | the same | USD 100 M | measured; worse than the ALU shadow's 3.5x at k 0.3 only at the very bottom of its band, and the band's centre is under 0.1 |
| the per-load shape (the fixed export) | 3.30 / 2.24 | 0.998 / 0.555 | as the class v4 shape within 5 percent | | ALU | USD 200 M if the sound form exists (16.2) | dead as drawn (20.2a-close) |
The k column was priced at this point: section 15.2 set the tile count from the 4090's marginal 0.056 pJ per MAC so that the block would carry the ALU shadow's 0.654 microjoules (11,400 tiles per hash, R about 1,430), and the exported pack is 1,430 tiles per iteration, 11,440 per hash; the 5090 reads 0.091 pJ per MAC unlocked and 0.048 at the lock at that point, so the GPU-side cost of the chip-model-v3 5.11 column (2.1x at k = 1, 1.6x at k = 1.5) is now measured at the premium it was priced for (1.067 microjoules unlocked, 0.564 at the lock, against the ALU shadow's 1.10 and 0.652). The open side is Apple: the Metal emulation costs 35 percent of the M5 Max's rate at 1,024 tiles per hash and 78 percent at 4,096, so the 11,440-tile point is out of the Apple tier's reach without an integer matrix path in Metal; nothing served moves (the chip texts rest on the ALU shadow). Reading: at `k = 1` the tile block and the ALU shadow give the same 2.1x to 2.2x, and the tile block's premium at the knee is 14 percent lower for it; the difference is the chip's reachable `k`. On the ALU shadow a fixed SIMD array can plausibly reach `k` 0.3 to 0.5 (the served pessimistic 3.4x to 3.5x); on the int8 tile the GPU's own tensor core is the dense array, measured at 0.048 to 0.091 pJ per MAC, and a chip at the same node cannot be 2x to 3x cheaper per MAC at the same voltage, so the chip's downside column collapses: the tile block's edge at the measured premium reads 2.2x at `k = 1` and under 1.8x at any `k` over 1.3, against the ALU shadow's 3.5x at `k` 0.3. That is the finding of the k hunt made measured on the card: **the tile block does not beat class v4's premium (it matches it at the same rate, 14 percent cheaper at the knee) and it beats class v4's chip edge at the pessimistic end (about 2.2x against 3.5x) and not at `k = 1`.** What stands against it as a class: the verifier (AVX2 0.047 microseconds per tile per unit, `mm1430` 5.8 ms alone and 10.14 ms with the sibling loaded on the box's core, a 0.14 ms miss at load 84 to 95; scalar 13x worse; NEON unwritten), the Apple tier (the Metal emulation at 1,024 tiles costs 35 percent of the M5 Max's rate and 4,096 costs 78 percent, so 11,440 is out of reach without an integer matrix path in Metal), the AMD WMMA layout (unverified; the OpenCL reference path emitted but unmeasured on the 9070 XT), and the compile-ahead (1.2 to 2.1 s per pack on the 5090, inside the budget).
The k column was priced at this point: section 15.2 set the tile count from the 4090's 6 October figure so that the block would carry the ALU shadow's 0.654 microjoules (R about 1,430), and the exported pack is 1,430 tiles per iteration, 11,440 per hash; the 5090 reads 2.9 pJ per MAC unlocked and 1.5 at the lock at that point (15.1a's correction: 32 MACs per lane per tile), so the GPU-side cost of the chip-model-v3 5.11 column is measured at the premium it was priced for (1.067 microjoules unlocked, 0.564 at the lock, against the ALU shadow's 1.10 and 0.652), and the column's premise (a chip's MAC no cheaper than the GPU's) is false by 4x to 30x on the public chip figures. Reading, corrected: at `k = 1` the tile block and the ALU shadow give the same 2.1x to 2.2x and the tile block's premium at the knee is 14 percent lower for it; but `k = 1` is not where a chip sits on tile work. The GPU's tensor core costs 1.5 to 4 pJ per int8 MAC measured, a 5 nm MAC array 0.04 to 0.4 claimed, so the chip's `k` on tiles is 0.03 to 0.3 against the ALU shadow's 0.3 to 0.8: at the same premium the tile block leaves the chip 3.5x to 6.7x where the ALU shadow leaves it 2.1x to 3.5x. **The tile block does not beat class v4's premium (it matches it, 14 percent cheaper at the knee) and does not beat class v4's chip edge at any k a chip can reach; it is the worse lever, and the 6 October verdict on scheme B stands for the right reason now (the honest card's tensor path is not near the floor; the chip's is).** The Apple side (35 to 78 percent of the M5 Max's rate at 1,024 and 4,096 tiles), the verifier (AVX2 0.047 microseconds per tile per unit, `mm1430` 10.14 ms loaded) and the AMD layout no longer need deciding; nothing served moves (the chip texts rest on the ALU shadow).
**The first sentence, on measured rows (04:22 UTC, 8 October 2026): neither prototype beats class v4's premium; the tile block matches it at the same hash rate (146.7 W against 151.2 W unlocked, 71.5 against 82.8 W at the 1,300 lock) and beats class v4's chip edge only at the pessimistic end (about 2.2x against 3.5x, because a chip's int8 MAC cannot undercut the 5090's measured 0.048 to 0.091 pJ per MAC the way a fixed datapath undercuts its 6.4 to 10.8 pJ per ALU op), not at k = 1 (2.2x either way); the per-load placement is dead as a construction, and its energy rows read 13 to 14 W under the whole block for the same instructions.** The SM-sparse reading landed at 05:12 UTC after three failed runs (the race switch, the race list, the rewrite's capture) and the candidate is dead: a quarter of the SMs holds 98 percent of the rate at the same draw (20.3b).
**The first sentence, on measured rows (06:xx UTC, 8 October 2026, after the microbench's correction): neither prototype beats class v4's premium or its chip edge; the tile block matches the premium at the same hash rate (146.7 W against 151.2 W unlocked, 71.5 against 82.8 W at the 1,300 lock) and is the worse lever against a chip, because the 5090's int8 MAC costs 1.5 to 4 pJ measured where a 5 nm array costs 0.04 to 0.4 claimed (k 0.03 to 0.3, under the ALU shadow's 0.3 to 0.8); the per-load placement is dead as a construction; the SM-sparse kernel is dead on measured rows (the draw follows the work, not the SM count); the L2 hot table is dead on the GPU side (an L2 hit costs 1.4 to 2.4 nJ against a chip's 0.2 to 0.5); and a shuffle-heavy re-weight costs the GPU 5x the add per op (55.8 pJ), so it stays held. The ALU shadow at the operating point's knee is the floor the night leaves: 2.1x at k = 1 for 82 to 90 W on a 5090, measured four times.**