Class v6: 10.1 SM-sparse measured dead on the 4090, H100 and 5090 (the residual is the clock domain); lane 3's close rows (the capex wall on the corrected project floor: USD 23 M and 340 M a year; W = 16 per chip; the SRAM rows on the synthesised core); 10.0 carries them with the 4090 and H100 rows

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-08 12:20:55 +00:00
parent f0146c2c22
commit 00890f9ccd

View file

@ -294,18 +294,32 @@ The frame now, filled with the defaults, each cell replaced as its lane's row la
| Honest tier (measured joules per hash, class v4) | GDDR7 board (0.466) | HBM3 one stack (0.321) | SRAM die at W = 4 (0.051) | With SM-sparse | W = 16 variant |
|---|---|---|---|---|---|
| RTX 5090 stock (3.36; class v3 2.26) | 3.3x default / 5.8x floor | 3.9x / 7.8x | 5.6x / 21x | no change (default: dead on PC 1, 20.3b: sp43-w32 holds 98.2 percent of rate at the same draw; the rented-pod rows owed by 15:00) | not applied (default: outcome A, the card -47 percent, dead; the rented-5090 row owed by 14:45) |
| RTX 5090 at the 1,300 MHz lock, the record's operating point (2.33 = 1.67 + 0.65) | 2.9x / 4.0x | 3.6x / 5.4x | 3.9x absolute (6.1x in the record's convention, marked) / 14x | no change (same default) | not applied (same default) |
| RTX 5090 stock (3.36; class v3 2.26) | 3.3x default / 5.8x floor | 3.9x / 7.8x | 5.6x / 21x | no change, MEASURED on three more cards (10.1, 13:25 UK): the honest floor at stock is the base row within 2 percent on the 4090 (best shape saves 4 W), the H100 (1.4 percent) and the 5090; the residual is the clock domain, which only the lock takes off | not applied (default: outcome A, the card -47 percent, dead; the rented-5090 row owed by 14:45) |
| RTX 5090 at the 1,300 MHz lock, the record's operating point (2.33 = 1.67 + 0.65) | 2.9x / 4.0x | 3.6x / 5.4x | 3.9x absolute (6.1x in the record's convention, marked) / 14x | no change (the lock pass on PC 1 running; the knee row by 15:00) | not applied (same default) |
| Apple M5 Max (1.40; class v3 0.78; the GPU and DRAM channels) | 1.8x / 2.4x | 2.2x / 3.2x | 3.9x / 8.7x | not applicable (no SM lever on Apple) | not applicable (the M5 Max pays 0 at 64 bytes, measured; the chip +33 percent) |
| RTX 4090 stock (5.005; class v3 3.588; rented pod, 10.1) | 4.3x / 8.7x | 4.9x / 11.6x | 6.6x / 31x | no change (measured: 16 of 128 SMs saves 4 W) | not applied |
| H100 HBM3 at its 700 W cap (2.776, throttled to 1,750 MHz; class v3 1.769; rented pod, 10.1) | 2.9x / 4.8x | 3.4x / 6.4x | 5.0x / 17x | no change (measured: per-SM throughput-bound, 1.4 percent) | not applied |
The honest floor in one line, on the defaults: against the chip anyone can build (the GDDR7 board) the strongest honest tier by joule, the M5 Max, holds the chip to under 2x and a 5090 at its knee to about 3x; against the chip that needs an N2 project (the SRAM die) the holds are about 4x and 4x; every cell's floor-k figure is the worst case if a chip's core costs no more than its bare units.
The capex wall (lane 3, 10.3): below about USD 100 M a year of miner revenue (IGN 0.13 at launch emission) no rational SRAM project starts at any edge; above about USD 2.3 B a year (IGN 2.9) every one does; a USD 100 M project at 30 percent of the hash needs IGN 0.43 to 0.59. The project cost moves the threshold 5x, the chip's edge 1.4x.
The capex wall's two thresholds (lane 3, 10.3, re-folded 13:18 UK on lane 5's corrected project floor of USD 20 to 75 M for the cheapest DRAM-board chip): no rational chip project of any kind below about USD 23 M a year of miner revenue (IGN 0.03 at launch emission, USD 62 K a day, a cap of about USD 60 M); every DRAM-board project at a third of the network above about USD 340 M a year (IGN 0.44, USD 0.93 M a day, a cap of about USD 0.9 B), where the SRAM project also starts. The project cost moves the threshold 5x, the chip's edge 1.4x.
The served line in one sentence, on the defaults: a chip that stores the dataset reaches about 3x per joule against an RTX 5090 at its knee and under 2x against an M5 Max, 4x against both if it is built on an N2 SRAM store for USD 100 M or more, and no rotating layer moves those figures; the layers decide which chip can be built and how long its tape-out lives.
The served line in one sentence, on the defaults: a chip that stores the dataset reaches about 3x per joule against an RTX 5090 at its knee and under 2x against an M5 Max, 4x against both if it is built on an N2 SRAM store for USD 100 M or more, and no rotating layer moves those figures; the layers decide which chip can be built and how long its tape-out lives. The worst case beside it, if a chip's core costs no more than its bare units (the floor k of 10.2): 4x and 2.4x against the DRAM board, 14x and 9x against the SRAM die (lane 3's line: 6x to 10x at the full shadow, never under 2x).
The three class v6 changes it implies: (1) the op mix stays class v4's with the lossy families capped at their base (lane D: 0 of 3,000 eras exhausted under the band; the k lane: mulhi is the card's worst lever by 6x, so no re-weight helps the card); (2) no SM-sparse default (dead on the measured 5090 rows; the per-watt gift of the hot table is the chip's, 20.3c); (3) the dataset schedule 5.5 / 8.5 / 11.5 GiB with the read width pinned at 4 words (the chip's ticket USD 1,500 / 2,500 / 3,000; a tuned 5090 pays 1 to 5 percent per step; about a quarter of today's cards by count per step).
### 10.1 SM-sparse on the honest card (floor lane 1, `docs/analysis/class-v6/floor/sm-sparse.md` on class-v6-floor-sm at 9e478daf, first rows 13:25 UK; the knee row by 15:00, the file by 19:30)
**The 20.3b reading holds on three more cards, and the residual over the DRAM's own energy is the clock domain, not the SMs.** Rented pods at stock (the worker's `--bench`, 250 x 2^24, nvidia-smi at 1 Hz, bit-exact on every row); microjoules per hash, the edge card over chip (GDDR7 0.466, HBM3 0.321, SRAM at lane B's 0.14):
| Card (measured) | Class v3 base | SM-sparse best | What the SM count does | Class v4 |
|---|---|---|---|---|
| RTX 4090 (128 SMs, idle 34.7 W) | 70.63 MH/s at 253.4 W = 3.588 microjoules (7.7x / 11.2x / 25.6x) | 16 SMs 69.95 at 249.6 W = 3.568 (saves 4 W); 32 SMs 70.42 at 251.4; 8 SMs 65.07 at 239.1 = 3.675; 128 SMs x 1 warp 66.28 at 242.7 | the draw over idle is 205 to 219 W at every shape; the knee 16 of 128 SMs | 70.63 at 353.5 W = 5.005 (premium 100 W, 11.1 pJ per counted op) |
| H100 HBM3 (132 SMs, idle 71.3 W) | 254.63 at 450.4 W = 1.769 (3.8x / 5.5x / 12.6x) | the self-tune's sp127-w32 253.2 at 441.6 W = 1.744 (1.4 percent under base); 132 SMs x 8 warps 99.2 percent at the same draw | the rate falls one for one with the SM count (99 SMs 212.7 at 402.8 W, 1.894; 66 SMs 189.2 at 372.9; 33 SMs 123.4 at 291.0, 2.358): per-SM throughput-bound at 1.93 MH/s per SM, not activate-bound; -pl and -lgc refused, -lmc accepted but the memory clock stayed at 2,619 MHz | 251.9 at the 700 W cap = 2.776 (clock throttled to 1,750 MHz); class v5 250.2 at 698.8 W = 2.793 |
| RTX 5090 (PC 1, unlocked) | 20.3b's points: full 2.30 at 313.9 W; a quarter 2.27 at 309.8; an eighth 2.28 at 303.4; a sixteenth 2.74 at 274.5 | the finer ladder (170 to 11 SMs at 32 warps; 170 SMs at 16 to 1 warps; the matched-warp pairs) runs on four rented 5090s and on PC 1 (`run-ca4-pc1-floorsm-5090-20261008`, the unlocked pass done, the 1,300 lock pass running) | 170 x 8 warps 137.0 at 482.9 W against base 137.16 at 452.8 on class v4; 170 x 4 134.9 at 483.9: fewer warps per SM at full SM count does not lower the draw either | base 137.16 at 452.8 W |
The reading for section 2's term: what an idle SM does not save is the die's clock-domain cost (the clock tree, L2 and the crossbar, the memory controllers, leakage at 2,700 to 2,850 MHz); the 4090 at 8 SMs and 65 MH/s still draws 204 W over idle at 2,715 MHz, while the 5090's lock takes 100 W off at the same read rate. **The honest floor at stock per card is the base row within 2 percent; the lever stays rank 1, the operating point.** The patch that ships the auto-tune anyway (`--sm-sparse auto|off|N` and `--sm-hold` on the worker, the tuning-file keys `sm_sparse` and `sm_hold` per card, the pure pieces tested in `emu/variant-test.cpp` on the H100 pod) is on class-v6-floor-sm at 9e478daf; the design's default is off.
### 10.2 The shadow's k from RTL (floor lane 2, `docs/analysis/class-v6/floor/shadow-k.md` on class-v6-floor-k, build-4 and build-3, first synthesised rows 13:3x UK; RTL and flow under `tools/chip-model/rtl`; full by 19:30 UK)
**The first synthesised k is 0.18 at the 5090's lock in the absolute convention at N3, and nothing reads inside the claimed 0.3 to 0.8 band except the unscaled ASAP7 figure at the lock (0.35). These are PER-UNIT FLOORS: no fetch, no decode, no register file beyond an 8-entry window, which is the chip rotation kills (section 1). The coordinator's order (13:5x UK): the headline k for the close is the programmable sequencer-core row (fetch, decode, a 32-register file, the per-era registers, the drawn program), asked of the lane with the M5 Max column; until it lands the default is the record's shadowed rows at k 0.5 with these unit floors beside them as the lower bound, and the GDDR7 board at 4.0x at the lock and 5.8x at stock on the floor k is the WORST CASE the served line must survive, not the reading.** Method: a minimal lane (instruction register, an 8 x 32-bit register window of flops, read muxes, the unit, one write port; no fetch or decode) in Yosys 0.68 plus OpenROAD on ASAP7, routed, SPEF, power from a random-input gate-level VCD with every pin annotated, the TC corner at 0.70 V; so every chip figure is a FLOOR and every k a floor. The node scaling is claimed from TSMC's headline per-node power reductions (N5 x0.70, N3 x0.50, N2 x0.36 of ASAP7; approximate). The GPU side is the research file's 15.1a, measured.
@ -376,6 +390,32 @@ The capex wall on the spec's emission constant (3,168,808,781 base units per DAA
Below about USD 100 M a year of miner revenue (IGN 0.13) no rational SRAM project starts at any edge; above about USD 2.3 B a year (IGN 2.9) every one does; at IGN 0.01 / 0.10 / 1.00 (assumed) a USD 150 M project at 5x and 30 percent reads NPV -148 / -130 / +54 M. **The edge moves p* 1.4x across 3x to 13x; the project cost moves it 5x: the hash does not set the wall, the project cost and the detector do.** Killed with the number: per-era re-fill (0.96 J per era), per-block re-fill (0.96 W on the die against 1.3 to 3.2 percent of GPU rate), straddling atoms (+0.02 nJ on the die, +19 percent of items on the verifier at the gate), a second hot table (the die's own kind of memory, k 0.1 to 0.3). Kept: the per-lane scratch pins the lane (a dependency of the width lever); 3D-stacked SRAM would remove the floor's last effect on joules. The lane's line for the served sentence: about 4x at k 0.5 and 2x at k 1 against a 5090 paying its whole latency shadow, 20x to 60x without it, with the price threshold beside it (about USD 0.4 per IGN at a third of the network, 0.13 for a maker who takes the chain). Owed: the W = 8 rows, the Apple curve past 8 GiB.
The lane's close rows (class-v6-floor-sram 7bc9de4b, 13:18 UK), three of them.
(1) The capex wall re-folded on floor lane 5's corrected project floor (the cheapest DRAM-board chip project USD 20 to 75 M, not the record's 5 M; the same model: ships at 24 months, mines years 3 to 6, a 10 percent discount; an 18-month ship lowers p* by 15 percent):
| Project USD M | Share | p* at 5x / 3x (USD per IGN) | Launch-year miner revenue | Per day | Cap at the end of year 2 | Label |
|---|---|---|---|---|---|---|
| 20 | takes the chain | 0.029 / 0.035 | USD 23 / 27 M a year | 62 / 74 K | 0.06 / 0.07 B | modelled |
| 20 | 0.30 | 0.098 / 0.118 | 75 / 91 M | 207 / 248 K | 0.19 / 0.23 B | modelled |
| 75 | takes the chain | 0.110 / 0.132 | 85 / 102 M | 233 / 279 K | 0.22 / 0.26 B | modelled |
| 75 | 0.30 | 0.368 / 0.441 | 283 / 340 M | 776 / 931 K | 0.72 / 0.86 B | modelled |
The two numbers for the close: no rational chip project of any kind below about USD 23 M a year of miner revenue (IGN 0.03, USD 62 K a day, a cap of about USD 60 M); every DRAM-board project at a third of the network above about USD 340 M a year (IGN 0.44, USD 0.93 M a day, a cap of about USD 0.9 B), where the SRAM project also starts (USD 330 M at 13x, or a whole-chain maker). The correction moves the lower number 4x (from the 100 M of the SRAM-only reading) and the upper not at all.
(2) W = 16 at each chip's own cost (GDDR7 +33 percent on the memory term, HBM3 +24, the die at 560 bits, 0.88 nJ), zero shadow, two card readings (the 5090 passes within 5 percent at its best lane count, or the measured w64 row stands: 71.9 MH/s, 1.89x the energy if the watts hold):
| Chip | W = 1 | W = 16 if the 5090 passes | At the measured w64 row | Class v4 shadow (the synthesised core) | The full shadow | Label |
|---|---|---|---|---|---|---|
| GDDR7 board | 5.1x | 4.4x | 8.3x | 5.1x | 4.7x | modelled chip, measured card |
| HBM3 one stack | 7.5x | 6.7x | 12.7x | 7.2x | 5.9x | modelled chip, measured card |
| SRAM die | 66x | 19x | 36x | 14x | 8.8x | modelled chip, measured card |
W = 16 is the one width that moves the die more than 2x at zero shadow and it costs the DRAM chips 11 to 15 percent; worth it only if the PC 1 row passes, since at the measured w64 row every chip's edge nearly doubles. If it passes, W = 16 replaces W = 8 as the lane's width recommendation (and the design's `never` is lifted by that one measurement, section 10.5).
(3) The shadowed SRAM rows on the k lane's synthesised core (0.11 microjoules of shadow at N3, absolute k about 0.1, the shuffle open; the lane's 2.2b): under class v4 21x at W = 4 and 18x at W = 8 at stock, 14x and 12x at the knee; at the honest cards' whole latency shadow 10x and 9.7x at stock, 6.5x and 6.1x at the knee; an N2 core about 1.2x more. On the synthesised core the shadow is not the whole hold (the record's claimed band read 2x to 6x, marked beside): the memory is a third to a half of the die's energy, the width is worth 1.3x with the shadow on, and the lane's line for the served sentence is 6x to 10x at the full shadow (2x to 4x on the claimed band, marked), never under 2x. These are the per-unit-floor figures of 10.2; the k lane's sequencer-core row re-folds them.
### 10.5 Invention beyond the four layers (floor lane 5, `docs/analysis/class-v6/floor/invention.md` on class-v6-floor-invention, first KEEP/KILL table 13:4x UK; full by 19:30 UK)
**One lever found, conditional on one measurement: the second DRAM sector.** The identity's only `k = 1` work is the memory's own, and the record forced one 32-byte sector per read while the dataset item is 64 bytes. Reading the whole item (W = 16 words; the read-width branch's class w64, fingerprint 836e56e7d496e980, the verifier 0.630 against 0.604 ms) forces a second sector in the same open row: on the chip model's own inputs (909 pJ activate plus 1,150 pJ of movement per sector) the GDDR7 chip's read goes from 2.0 to 3.2 nJ and E_mem from 0.466 to 0.62 microjoules (+33 percent at an unchanged 166 MH/s), one HBM3 stack 0.321 to 0.40 (+24 percent), and on floor lane 3's wire figure the N2 SRAM die 0.036 to 0.126 microjoules (66x to 19x). So the band {1, 4} words is inside one sector and "the chip pays nothing" is right to 32 bytes and wrong at 64; this is the measured reason W = 8 is free for the chip too (section 10.3) and W = 16 is the only width that costs it. The card is the whole question: the 5090 ran w64 at 71.9 MH/s (-47 percent; the exported kernel at full occupancy, four uint4 loads per item, no request hint) while its own 64-byte probe reached 15.7 G reads a second at its best lane count (-13 percent); 589 GB/s is a third of the stream, so the bind is the request path, not the pins. The 9070 XT (-3 percent) and the M5 Max (0) are free at 64 bytes (measured).