Counter ASIC 4.0 research, second pass: the k-above-1 hunt, the capex lever and the re-ranked list

Sections 15 to 19: every GPU block gaming and AI paid for, priced against the best public 4 to 5 nm chip figure with its k band, verifier cost and the edge at the measured premium (the one-sentence answer: nothing reads k above 1 with certainty; the int8 tensor tile is the only block near or above 1, so the tensor shadow with a SIMD byte-dot verifier is the new rank 3); the capex column added to the chip model (both sides capex-dominated at 7x to 10x their electricity; the per-unit capex wall unreachable by 3x to 7x; the project wall doubled from USD 100 M to about USD 200 M of break-even cap by the shadow-per-load design, the mission lane's method); the microbench of 20 probes documented as built (section 18); the owed list.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-07 21:26:19 +00:00
parent 2b9cef47dc
commit e767701097

View file

@ -219,3 +219,103 @@ The premium cannot reach zero against a chip that is a GPU's memory system witho
Internal: `docs/analysis/chip-model-v3.md` (sections 5 and 6; section 5.10 on ca3-coord), `docs/analysis/latency-shadow-2026-10-06.md`, `docs/plans/counter-asic-3-status.md` (ca3-coord: the efficiency pass, the 4070 and 9070 XT rows), `docs/plans/counter-asic-3.md`, `docs/plans/counter-asic-3-derivation.md`, `docs/plans/counter-asic-2.md`, `docs/design/class-v5-stored-state.md` (class-v5), `docs/design/latency-ladder.md` (ladder and attack-ladder-5a), `docs/analysis/attack-pass/f1-shadow.md` and `f10-ladder.md` (attack-pass), `docs/plans/cryptanalysis/in-house-pass.md` (crypto-engage), `docs/analysis/horizon/new-pow.md` and `algorithm.md`, `docs/analysis/asic-resistance-history.md`, `docs/analysis/scratch-soundness.md`, `docs/bench-log.md`, `docs/evidence.md` row 15.
External (all read 7 October 2026): O'Connor et al., Fine-Grained DRAM, MICRO 2017, https://www.cs.utexas.edu/users/skeckler/pubs/MICRO_2017_Fine_Grained_DRAM.pdf ; Chatterjee et al., subchannels, HPCA 2017, https://www.cs.utexas.edu/users/skeckler/pubs/HPCA_2017_Subchannels.pdf (GUPS about 9 pJ per bit against STREAM about 3, claimed); Horowitz, ISSCC 2014, https://pages.cs.wisc.edu/~markhill/restricted/isscc2014_horowitz_power_scaling.pdf ; Dally, Hot Chips 2023 keynote, https://www.hc2023.hotchips.org/assets/program/conference/day2/Keynote%202/Keynote-NVIDIA_Hardware-for-Deep-Learning.pdf ; Micron GDDR7 4.5 pJ per bit via https://www.club386.com/micron-gddr7-improves-nvidia-gpus-by-up-to-30-in-gaming/ ; Samsung HBM3E 3.9 pJ per bit via https://www.igorslab.de/en/samsung-ends-afterburner-hbm4e-with-325-tb-s-bandwidth-to-exceed-nvidias-requirements/ ; AccelWattch, MICRO 2021, https://paragon.cs.northwestern.edu/papers/2021-MICRO-AccelWattch-Kandiah.pdf ; tensor cores on memory-bound kernels, https://arxiv.org/html/2502.16851v2 ; Shuhai, FCCM 2020, https://arxiv.org/pdf/2005.04324 ; Folded Banks, ISCA 2025, https://dl.acm.org/doi/10.1145/3695053.3731111 ; Li, Reddy, Jacob, MEMSYS 2018, https://terpconnect.umd.edu/~blj/papers/memsys2018-dramsim.pdf ; the Ethash, RandomX, ProgPoW, Equihash, Cuckatoo, Kaspa, Alephium, FishHash, Autolykos, NexaPoW, Xelis and PoSME URLs in sections 5 and 8; the proving URLs in section 7.
## 15. The k below 1 hunt (second pass, the founder's question of 22:4x UK: "a better way to shadow, or a new way, that is easier on GPUs and hard as hell on chips")
Written 21:3x to 21:5x UTC, 7 October 2026. A note on the letter: this file's `k` is the chip core's energy per forced operation over the GPU's, so an operation that is cheaper on the GPU than a chip can match is `k` ABOVE 1 here (the coordinator's brief wrote the same thing as "k below 1" with the ratio the other way up). The identity of section 2 does not change: zero premium is zero forcing at any `k`; what a high `k` buys is the right-hand asymptote `1 / k` and a steep slope at small premium, so a block with `k` near or above 1 makes a GIVEN premium count for more and removes the chip's downside bet. "Hard as hell on chips" is `k` well above 1; "easier on GPUs" is a block whose joules per unit of forcing are low, which is the same thing said twice.
**The one-sentence answer first: on the public figures nothing reads k above 1 with certainty; every block a GPU has, a chip at the same node builds for the same or less energy per operation, and the only block where k reads near or above 1 rather than 0.3 to 0.5 is the int8 tensor tile, because the GPU's tensor core is already a dense MAC array at the node's density, so the best design is the tensor-shaped shadow with a SIMD verifier (section 17, rank 3), and the microbenchmarks built tonight (section 18) replace the public-figure column with measured picojoules per operation on the 5090 when the PC 1 queue reaches them.**
### 15.1 Every GPU block gaming and AI already paid for, priced
The GPU column is the public figure or this project's measurement until the microbench rows land (each probe is named; the row then reads watts minus the sleep floor over counted operations per second, at the unlocked clock and the 1,300 MHz knee). The chip column is the best public figure at a 4 or 5 nm node for the same operation, labelled. `k` is the chip's cost over the GPU's at the same operating point. The verifier column is the CPU's cost to reproduce the operation bit-exactly inside the hash (the x8 verifier does 18 G simple integer ops per second per core and the shadow's law is 0.1 ns per lane-instruction, `latency-shadow-2026-10-06.md` section 4). The last column is the edge if the shadow were made of that operation at the class v4 premium's joules at the 1,400 lock (0.654 microjoules): at ZERO premium every row reads 3.6x (GDDR7) and 5.3x (HBM3), the identity; the row is what the same joules buy against a chip of that `k`.
| GPU block | The 5090's energy per op: public or measured today | Microbench probe | A chip's cost at 4 to 5 nm, public | k band | Bit-exact on CPU, three vendors; verifier cost | Edge at the measured premium (0.654 microjoules at the lock) | Labels |
|---|---|---|---|---|---|---|---|
| Int32 ALU (add, xor, rotate, LOP3, byte permute) | 6.5 pJ per counted op at the 1,400 lock, 10.4 unlocked (the class v4 shadow) | int_arx, lop3, prmt | datapath add 0.06 pJ at 5 nm (mlsysbook citing Horowitz 2014 and Dally 2021, https://mlsysbook.ai/vol1/backmatter/appendix_assumptions.html); the lane's register file and operand delivery dominate on both sides: Dally's 45 nm figures (Hot Chips 2023) are 30 pJ of fetch, decode and operand delivery around a 1.5 pJ HFMA; a SIMD array with a local register file at N5 about 2 to 5 pJ per op | 0.3 to 0.8 | yes; 0.1 ns per lane-op, the shadow's measured law | k 0.5: 2.9x; k 0.8: 2.3x; k 1: 2.1x | GPU measured; chip claimed and approximate |
| Int32 multiply, mulhi | unmeasured marginal (the family step costs of item 6 give the multiply at about the add's issue cost on the 5090) | int_mul, int_mulhi | 32-bit multiply 0.52 pJ at 5 nm datapath (claimed, the same source); the multiplier is the same circuit on both sides, so the overhead ratio is the ALU row's | 0.4 to 0.8 | yes; one op | as the ALU row | claimed |
| Warp shuffle crossbar | the shfl step 0.86x of the add-xor-rotate step on the M5 Max, measured (item 6); the 5090's step cost in the item 6 record; marginal pJ unmeasured | shfl | a 32-lane crossbar is wires: 0.6 pJ per bit-mm at N5 (approximate, chip-model-v3 5.1), a 32-bit word across a 1 mm array about 20 pJ on either die; the algorithm lane's k floor for a shuffle-heavy mix 0.46 (approximate) | 0.5 to 1 | yes; one op | k 0.5: 2.9x; k 1: 2.1x | measured steps; chip approximate |
| Register file bandwidth | not a separable op: the 3 operand reads per lane instruction are most of the 6.5 to 10.4 pJ (Dally: 6 pJ of register file per in-order instruction at 45 nm, claimed) | prmt and lop3 read it (cheap datapath, RF-dominated) | the same banked SRAM at the same node; a chip's lever is operand reuse through a fixed datapath (the 3x factor), which a chained random program denies | about 1 for a random program | the CPU keeps the 32 lanes' registers in L1 (register-major loops) | as the ALU row at k 1: 2.1x | claimed |
| FP32 FMA at the GPU's density | 104.8 TFLOPS FP32 on the 5090 (NVIDIA, claimed: 52 T FMA per second); marginal pJ unmeasured | fp32_fma | FP32 FMA at 5 nm about 0.4 to 0.6 pJ datapath (Horowitz 45 nm: 0.9 add, 3.7 multiply, scaled by the 2.5x to 3x process factor Dally gives per node step, approximate); the same overhead ratio as int | 0.3 to 0.6 | fmaf is IEEE-exact on every vendor with fast-math off and denormals NOT flushed (Apple and AMD flush by default in some modes: a vendor-mode risk the integer hash does not have); one op | k 0.5: 2.9x | claimed, approximate |
| FP16x2 FMA | 2 per instruction; unmeasured | fp16x2_fma | half the FP32 datapath | as FP32 | fma.rn.f16x2 is exact per IEEE half with round-to-nearest, but Apple's and AMD's half paths differ in denormal handling: excluded from the hash content | | approximate |
| **Tensor tile, int8 (u8 m8n8k16, s8 m16n8k32)** | 0.056 pJ per multiply-add marginal at 4,096 tiles per hash on the RTX 4090 (measured 6 October, new-pow 5.1; 0.70 pJ at 64 tiles, the fixed cost amortising); Dally: IMMA 160 pJ per 1,024-MAC instruction at 45 nm (claimed) = 0.16 pJ per MAC, about 0.02 at 5 nm by his process factor | mma_u8_m8n8k16, mma_s8_m16n8k32 | NVIDIA's 5 nm INT4 test chip 95.6 TOPS/W at 0.46 V (JSSC 2023, measured, via Dally's NASEM slides https://www.nationalacademies.org/cdn/materials/9fba0a50-8a84-4dda-8844-e5461844fb28): 0.021 pJ per INT4 MAC at 0.46 V, about 0.1 pJ at 1.05 V (the plot's 5x); an INT8 MAC 2x to 4x an INT4 MAC: 0.04 to 0.08 pJ at the low voltage, 0.2 to 0.4 pJ at nominal; a licensable int8 MMA block is the same circuit as the GPU's | **0.7 to 3 at the same voltage; near 1 is the honest centre** | integer, exact (bit-exact on 2^24 lanes and 1,024 CPU lanes at five rungs, measured); scalar 1.07 microseconds per tile per unit: at the premium's tile count (11,600 tiles per hash, R about 1,450) 12.4 ms, FAILS the gate; a byte-dot verifier (AVX-VNNI vpdpbusd 64 MACs per instruction, NEON udot 16, AVX2 pmaddubsw) 4x to 16x cheaper, 0.8 to 3 ms, approximate and unmeasured | at 0.654 microjoules of tiles: k 0.7: 2.5x; **k 1: 2.1x; k 1.5: 1.6x; k 2: 1.3x**; the chip's downside (k 0.3) is gone because its MAC cannot be 3x cheaper than the GPU's at the same node | GPU measured on a 4090; chip claimed (a test chip at a different voltage); k approximate |
| Tensor tile, FP16 and BF16 (m16n8k16, FP32 accumulate) | the same hardware path; unmeasured | mma_f16_m16n8k16, mma_bf16_m16n8k16 | the same array with floating-point accumulate | about 1 | NOT bit-exact: the order of the 16 products' accumulation inside the tile is unspecified by PTX, so two vendors (or two NVIDIA generations) may differ in the last bit; excluded from the hash content; measured here only for the energy picture | | claimed |
| Tensor tile, FP8 e4m3 (m16n8k32) | sm_89 and later; unmeasured | mma_e4m3_m16n8k32 | the same | about 1 | the same exclusion (floating accumulate), and no AMD RDNA or Apple path | | claimed |
| L2 cache (96 MB on the 5090, 64 MB class on a 4090, 36 on a 4070) | an L2 hit on a GPU is about 0.1 to 0.3 nJ per 32-byte sector (approximate, Horowitz's 1 MB cache at 100 pJ per 64 bits at 45 nm scaled to N5 and a 96 MB array with its crossbar); measured tonight: the l2_chase rows (dependent, the latency) and l2_indep4 (the throughput) | l2_chase_32m, l2_chase_64m, l2_indep4_32m | a 64 MiB SRAM on a 128 mm^2-class die: 0.2 to 0.5 nJ per 64-byte read (chip-model-v3 5.1, approximate); a 3 nm 563 kbit macro reads at 3.89 pJ per access (JSSC 2026, measured, https://sscs.ieee.org/tag/one-cycle-latency-low-leakage-access-mode-1-clm/), so the array's wires, not the cell, set the cost on both sides | **0.7 to 3; the one block where the chip may be WORSE** | the hot items derived lazily on the CPU (unmeasured; the 5 October hot-table packs); the measured rate cost without cache hints 13 to 16 percent on the 5090 (the ldcs job in the hash lane's queue reads it with the hint) | 512 hot reads per hash at 0.2 nJ = 0.1 microjoules: too few joules to force by themselves; at k 1.5 that 0.1 buys 0.1x; the lever is capex-free (section 16) and energy-small | approximate throughout |
| Texture sampler, point sampled | a fetch through the texture cache's address unit; unmeasured | tex_point_u32_32m | an address unit is a few adders; a chip reads SRAM or DRAM directly | about 1 or under (the chip skips the unit) | a point fetch is a load; exact | as the L2 row | approximate |
| Texture sampler, linear interpolation | the interpolation weights are 9-bit fixed point with 8 fractional bits on NVIDIA (CUDA Programming Guide, "Linear Filtering", claimed); unmeasured | tex_linear_f32_256k | a chip's interpolator is one 9-bit multiply-add per fetch, cheaper than the GPU's full sampler | under 1 | NVIDIA's rule is reproducible on a CPU; AMD's and Apple's samplers use their own fraction widths and rounding (vendor-specific, approximate): NOT bit-exact across the three vendors, excluded from the hash content; measured for the energy picture only | | claimed, approximate |
| The rasteriser | not reachable from CUDA, OpenCL or Metal compute; only through a graphics pipeline, whose rasterisation rules (fill convention, sample positions, depth precision) differ by vendor | none | | | NOT bit-exact across vendors by construction; excluded | | |
| DRAM random read (the control) | the hash's own pattern: 17.5 G dependent 4-byte reads per second at about 3.1 nJ incremental per read (modelled, 55 W of memory system) | dram_chase_1g | 2.0 nJ per random 32-byte read on GDDR7, 1.2 on HBM3 (modelled from O'Connor et al. and Samsung's pJ per bit) | 0.4 to 0.65 (the chip's whole case) | exact; the verifier's item derivation | this IS E_mem: 3.6x at zero premium | modelled |
The reading: a GPU lane's cost per operation is overhead (fetch, decode, operand delivery, the banked register file) around a datapath that is a few percent of it, and a chip at the same node builds the same datapath and may drop part of the overhead, which is why every ALU-shaped row reads `k` 0.3 to 0.8. The two rows that read `k` near or above 1 are the ones where the GPU's block is itself a dense array with no per-op overhead to drop (the int8 tensor tile) or a large SRAM whose cost is wires that a chip's SRAM has too (the L2). The tensor tile is the only one of the two with joules to force at a verifiable cost, and only with a SIMD verifier.
### 15.2 The tensor shadow, re-read on the identity
The 6 October verdict on scheme B ("never as class v5 content") was right for the question it answered: at the 4090's free band (R = 512, 0.23 microjoules) the block forced too few joules and the verifier paid 26x the ALU shadow's cost per joule. The founder's question is a different one: not "is there a cheaper lever" but "is there a lever whose joules the chip cannot undercut". On that question the tensor tile is the best block on the card, because the GPU's marginal 0.056 pJ per MAC is within a factor of about 2 of what a 5 nm MAC array costs anyone (the test chip's 0.04 to 0.1 pJ, claimed), while the ALU shadow's 6.5 to 10.4 pJ per op is 10x to 50x what a fixed SIMD datapath costs. What it would take to make the tensor block carry the SAME premium as the ALU shadow (0.654 microjoules at the lock), on the 5090:
| Quantity | Value | Label |
|---|---|---|
| Tiles per hash for 0.654 microjoules at 0.056 pJ per MAC | 11.7 M MACs per hash, 11,400 u8 m8n8k16 tiles (R about 1,430 per iteration) | arithmetic on the 4090's measured marginal; the 5090's marginal is tonight's mma_u8 row |
| Where the 5090's free band ends | the 4090 held its rate to R = 512 at 40 percent of its dense int8 peak (approximate); the 5090's tensor peak is about 1.7x the 4090's (NVIDIA's dense INT8 figures, claimed), so the free band on a 5090 reaches about R 2,000 and 11,400 tiles per hash sits inside it; on a 4070 (about 0.45x the 4090's tensor peak) the same R is tensor-bound and costs rate | approximate |
| The verifier at R 1,430 | scalar 12.4 ms per unit on the M5 Max core (FAILS the 10 ms gate); with AVX-VNNI (64 byte-MACs per instruction) 0.8 ms, NEON udot (16 per instruction) 3 ms, AVX2 pmaddubsw (16 per instruction, every x86 since 2013) 3 ms, plus the x8 base 2.06: 2.9 to 5.1 ms on the M5 Max core, 7 to 13 ms on a 2019-class core by the 2.5x rule; the 2019 core passes only with the AVX2 path and a small R | approximate, unmeasured; the item that decides it |
| The chip's edge at that premium (1,400 lock, GDDR7) | k 0.7: 2.5x; k 1: 2.1x; k 1.5: 1.6x; k 2: 1.3x; the same numbers as the ALU shadow at the same k, with the chip's k 0.3 column (4.1x) removed from the table because an int8 MAC array at 5 nm cannot be 3x cheaper than the GPU's | modelled |
| What it costs the honest cards | the same watts as the ALU shadow by construction (the premium is the design variable); Turing and later NVIDIA and RDNA 3 and later AMD native; Pascal, RDNA 2 and Apple emulate at 10 to 16 ALU steps per tile, which at 11,400 tiles per hash is 110,000 to 180,000 counted ops, over the M5 Max's 130,000 bind point: the Apple tier would lose 5 to 10 percent of rate (approximate) | measured bind point; emulation cost approximate |
| What breaks it | the AMD WMMA fragment layout is unverified (status item 6); Apple has no integer matrix path reachable from the toolchain; the SIMD verifier is unwritten; the 2019-class core may not fit; the chip's k near 1 is the centre of a claimed band, not a measurement | |
So the tensor shadow does not lower the GPU's premium (the founder's "easier on GPUs" is not available: a premium is a premium), it removes the chip's downside bet at the same premium, which is worth about 1x of edge at the pessimistic end (4.1x to about 2.1x at the ALU shadow's k 0.3 against the tile's k near 1). It is ranked below the two miner-only levers because it is a class change with an unwritten verifier and a vendor split, and above the ALU op-mix re-weight because its k floor is better bounded.
## 16. The inverse lever: capex as the wall
The brief: make the chip's capital cost, not its energy, the wall (the state-derived dataset plus an L2-resident hot table plus the era changes), with the break-even row recomputed. The model is the mission lane's (`docs/analysis/mission/future.md` section 2.2, 7 October 2026, model): the maker takes share `s` of the hash; two-year revenue `s x E2 x p` (E2 the two-year emission, 4.18 B IGN in years 1 to 2; p the price); the project pays when `p` is over `p* = C_proj / (s x E2)`, and the market cap at which it pays is `p* x supply` (4.18 B at the end of year 2), so in years 1 to 2 the break-even cap is `C_proj / s`, 3.3x the project cost at s = 0.30. The fleet's own silicon was 2 percent of revenue there and dropped out; the capex column below says whether anything in the hash can make it not drop out.
### 16.1 The capex column
| Chip or card | Silicon and memory per unit | MH/s | USD per MH/s | Capex per MH/s-hour over two years | Electricity per MH/s-hour at USD 0.05 per kWh | Capex over electricity | Label |
|---|---|---|---|---|---|---|---|
| f = 1 GDDR7 chip (chip-model-v3 5.4) | USD 470 (16 devices USD 320, controller USD 50, board USD 100) | 166 | 2.8 | 0.00016 | 0.000023 (0.466 microjoules) | 7x | modelled |
| the same plus a 64 MiB SRAM hot table | +32 mm^2 of N5 at 0.49 mm^2 and USD 0.23 per MB (chip-model-v3 section 2): +USD 15 of die; if it forces the controller onto an N5 die or a second die, +USD 50 to 100 of die and package (approximate) | 166 | 2.9 to 3.5 | 0.00017 to 0.00020 | 0.000028 (the table's reads) | 6x to 7x | modelled, approximate |
| the same plus the class v4 shadow core (30 mm^2 of N5 at N = 100,000) | +USD 25 to 40 | 166 | 3.0 to 3.1 | 0.00017 | 0.000080 (1.59 microjoules at k = 1) | 2x | modelled |
| the same with the shadow per load (section 16.2: the ALU core must sit on the controller's die or across an interposer) | +USD 200 of interposer and package (a one-stack CoWoS class package, chip-model-v3 5.1, approximate) | 166 | 4.3 | 0.00024 | 0.000080 | 3x | approximate |
| RTX 5090 at MSRP | USD 1,999 | 136 | 14.7 | 0.00084 | 0.000084 (1.69 microjoules at the lock); 0.00012 unlocked | 10x (7x) | measured card price; the street price in 2026 is about 2x MSRP, which doubles the capex row |
| RTX 4070 at its tune point | about USD 550 (approximate street) | 31 | 17.7 | 0.00101 | 0.000128 | 8x | approximate |
What the column says: BOTH sides are capex-dominated (7x to 10x their electricity per MH/s-hour), so the chip's economic edge is its capex per MH/s (5x over a 5090 at MSRP), not its energy, and a design that moved the chip's capex would move the number a miner actually computes. But nothing in the hash can move it far: the work per hash is small (512 instructions plus the shadow), so the silicon a chip needs beside its memory is 30 to 100 mm^2 of N5 (USD 25 to 60), and the memory is the same 16 devices the card carries (USD 320). A hot table forces USD 15 to 100 more; the shadow core USD 25 to 40; an interposer USD 200. The chip's capex per MH/s moves from 2.8 to at most about 4.3, against the card's 14.7 to 29. The per-unit capex wall is unreachable by a factor of 3 to 7, for the same reason the energy wall is: the hash's dataset is a card's worth of DRAM and its work is a fraction of a card's logic. The capex wall that exists is the PROJECT cost and the node it forces, and that is already in the break-even model.
### 16.2 The break-even row, recomputed with the capex column
The project cost is what moves the cap; the per-unit capex enters as the fleet term, which was 2 percent of revenue and grows to about 3 percent with the interposer: it still drops out. The rows (s = 0.30, the mission lane's method; the cap is `C_proj / 0.30` in years 1 to 2 and 2.3x more every two years with the emission glide):
| Chip project | What the hash forces | C_proj | Break-even cap, years 1 to 2 | Years 3 to 4 | Years 5 to 6 | Label |
|---|---|---|---|---|---|---|
| 28 nm controller, no shadow (class v3, or class v5 with the shadow dropped) | a GDDR7 controller and PHY; the state-derived dataset adds a leaf buffer (6 KB today) and a link to a node: firmware; the era draws: firmware | USD 5 M | USD 17 M | 48 M | 114 M | model (future.md) |
| 28 nm controller plus an N5 shadow core (class v4 as shipped) | an N5 die for the 14,000-lane ALU array; the two dies exchange 64 B of lane state per iteration: 17.5 G reads/s over 16 reads per iteration x 128 B = 140 GB/s, a PCB or organic-substrate link | USD 30 M | **USD 100 M** | 287 M | 682 M | model |
| the same plus a 64 MiB L2-class hot table | the table joins the N5 die (32 mm^2 more): no node change, USD 15 of die; the project unchanged | USD 30 M to 33 M | USD 100 M to 110 M | 290 M | 690 M | approximate |
| **the shadow per load** (the 256-instruction block split into 16 blocks of 16, one after every load, the same N): the chip's ALU core must sit inside every read's dependency, so the lane state crosses between controller and core twice per read: 17.5 G x 128 B = 2.2 TB/s, an interposer-class link, or one die carrying controller, lanes and PHY at N5 | a single N5 die or a 2.5D package: the mission lane's "N3 single die" row | USD 60 M | **USD 200 M** | 574 M | 1,363 M | approximate (the single-die cost is the mission lane's N3 row; a GDDR7 PHY on N5 is a real but unpriced item) |
| the same plus the tensor shadow (section 15.2) | a licensable int8 MMA block beside the lanes: USD 4 of N5 silicon (algorithm.md 5.2); the project unchanged | USD 60 M | USD 200 M | 574 M | 1,363 M | approximate |
| What the per-unit capex adds in every row | the fleet term, 2 to 3 percent of revenue | | nothing visible | | | model |
So the inverse lever's best design is the shadow per load: it costs the honest cards nothing by construction (the same N, the same instruction mix, a smaller block at 16 sites; the measured block-size effect on the 5090 and the M5 Max was that a 64-instruction block ran 2.5 to 3.5 percent FASTER than 256, latency-shadow sections 3 and 5, so 16 may be faster still or may cost compile-ahead, unmeasured), and it doubles the chip's project cost by forcing the controller and the core onto one advanced die or an interposer, which doubles the break-even cap from about USD 100 M to about USD 200 M in the first two years. It does not reach a capex wall per unit, and it does not change the energy identity (`k` is `k` whichever die the core sits on). It is a class change (the block positions are consensus) and a generator change; 4 to 6 agent hours plus the six gates; its measurement is one pack per card (rate, watts, compile-ahead at 16 sites).
## 17. The ranking, second pass
| Rank | Design | Chip edge per joule (GDDR7; 1,400 lock) | Premium: 5090 / 4070 | Verifier | Agent hours | Breaks on | Change |
|---|---|---|---|---|---|---|---|
| 1 | The operating point as the shipped default | class v3 3.6x; class v4 2.1x at k = 1 | lowers the base (330 to 223 W); the v4 premium 145 to 82 W | none | 2 to 4 | AMD and Apple have no lock | unchanged |
| 2 | The SM-sparse miner kernel | unmeasured; toward 2.8x (v3) and 1.7x (v4) if half the SM-side 99 W is reachable | none | none | 1 + 1 + 4 | the 99 W is clock tree and leakage | unchanged; the PC 1 job is queued |
| 3 | **The tensor shadow with a SIMD byte-dot verifier** (int8 tiles carrying the premium; a class change) | at the ALU shadow's premium: 2.1x at k = 1, 1.6x at k = 1.5, 1.3x at k = 2; the chip's k 0.3 downside removed | the same as class v4 by construction / the 4070 tensor-bound above about R 600 (approximate) | scalar FAILS; AVX2 and VNNI 2.9 to 5.1 ms on the M5 Max core, 7 to 13 ms on a 2019 core (approximate, unwritten) | 8 to 12 (the SIMD verifier, the AMD layout gate, Apple's emulation row) plus the six gates | the 2019-class core; the AMD layout; Apple loses 5 to 10 percent of rate on emulation; k near 1 is claimed, not measured | NEW this pass: the k-above-1 candidate |
| 4 | The shadow per load (capex) | unchanged in energy | none by construction (block-size effect unmeasured at 16) | unchanged | 4 to 6 plus the gates | compile-ahead at 16 sites; a class change | NEW: doubles the break-even cap to about USD 200 M |
| 5 | The ALU shadow re-weighted toward shuffles and multiplies | 2.1x at k = 1; 2.9x at the pessimistic k 0.46 | unchanged | unchanged | 4 to 6 plus the gates | Apple pays shfl 1.91x | down from 3: the tensor tile bounds k better |
| 6 | The L2-resident hot table with cache hints | energy 0.1 microjoules at k 0.7 to 3; capex +USD 15 | 13 W if the hint holds the rate | under 0.5 ms (unmeasured) | 4 for the ldcs measurement (queued) | Metal has no hint; the rate without it | unchanged; a measurement |
| 7 to 11 | class v5 without the shadow; the refresh as a cost; per-card classes; proof of useful work; memory shaping | as section 9 | | | 0 | dead by arithmetic | unchanged |
## 18. What was built in the second pass: the microbench
`proto-cuda/nvrtc/worker.cpp` gains `--microbench [--mb-seconds 60] [--mb-only a,b]`: no pack; 20 probes, one kernel each, compiled by NVRTC one at a time (a form the card or the compiler refuses drops that probe alone, its error on its RESULT row), each run at full residency (the occupancy query's blocks per SM x 256 threads x SMs) for the window, with a UTC start and end stamp per probe for a 1 Hz power sampler and a checksum of the lanes' outputs. The probes: `sleep` (the SM-resident floor: full occupancy, `__nanosleep`, no issue; every other row is read against it), `int_arx`, `int_mul`, `int_mulhi`, `prmt`, `lop3`, `shfl`, `fp32_fma`, `fp16x2_fma`, `mma_u8_m8n8k16`, `mma_s8_m16n8k32`, `mma_f16_m16n8k16`, `mma_bf16_m16n8k16`, `mma_e4m3_m16n8k32`, `l2_chase_32m`, `l2_chase_64m`, `l2_indep4_32m`, `dram_chase_1g` (the hash's own pattern, the control), `tex_point_u32_32m` and `tex_linear_f32_256k` (the sampler's address path and its interpolator, through `cuTexObjectCreate`, loaded optionally). The reading per row: (watts minus the sleep row's watts) over counted operations per second = picojoules per operation on this card at this clock. Gate: mingw cross-compile on igneum-build-1 under `lease pool 4 --class measure --label "ca4 research: worker cross-compile (microbench probes)"`, exit 0 with `-Wall -Wextra` clean, 21:20 UTC; exe `/srv/builds/ca4-research/nvrtc/igneum-worker-cuda-ca4mb.exe`, sha256 49aa60c4b28ed77a54655476d0f07c06a54aa22e4f6f6ccadd4019b84cbf0268; unrun on a card. The job is with the hash lane for PC 1 after the hot-table ldcs job: `--microbench --mb-seconds 60`, the 5090 alone, nvidia-smi at 1 Hz, at the unlocked clock and the 1,300 MHz knee. When its rows land, the GPU column of 15.1 becomes measured and every `k` band narrows from the GPU side; the chip side stays claimed.
## 19. Unverified and owed, second pass
- Every GPU figure in 15.1 except the ALU shadow's, the tensor tile's (a 4090) and the DRAM read's is unmeasured until the microbench rows land; the chip side is public claims at other voltages and nodes; `k` is a band, not a number, in every row.
- The SIMD byte-dot verifier is unwritten; its 4x to 16x is a scaling from instruction widths.
- The 5090's tensor free band (R about 2,000) is a scaling from the 4090's measured 40 percent at R = 512 and NVIDIA's claimed peaks.
- The shadow-per-load design's block-size effect at 16 instructions and its compile-ahead are unmeasured; the single-die project cost is the mission lane's N3 row, and a GDDR7 PHY on an N5 die is unpriced.
- The break-even model is the mission lane's (s = 0.30, a two-year life, the rental-equilibrium hash); its emission figures are that model's.