Merge counter-asic-4-docs 5f45de99 into master (gate: green on 5f45de99, recorded by tools/ci/pre-push.sh; landed on the box mirror under the exception declared by main: main's ruling, 7 Oct 2026 19:5x UK: the GitHub account is suspended, lanes land on the box mirror's master, the box gate stamp is the verdict; GitHub gets the fast-forward when it answers)
This commit is contained in:
commit
6af12f0871
4 changed files with 995 additions and 1 deletions
|
|
@ -365,7 +365,18 @@ The k column. k is the chip core's energy per op over the GPU's at the same oper
|
|||
|
||||
Correction, 8 October 2026 (the research lane's microbench on the 5090, counter-asic-4-research.md 15.1a and the corrected 20.3 and 20.4 at 71fd465b): the per-MAC figures above are wrong by a factor of 32. A `mma.m8n8k16` tile is 1,024 multiply-adds per warp, 32 per lane, so a hash does 32 MACs per tile, not 1,024; the 4090's "0.056 pJ per MAC" is 1.8 pJ, and the 5090 at the ALU shadow's premium reads 2.9 pJ per MAC unlocked and 1.5 pJ at the 1,300 MHz lock (the packs job, 366,080 MACs per hash; the microbench's dependent u8 tile 4.1 and 2.2, the wide s8 m16n8k32 tile 1.36 and 0.83). Against the same 5 nm MAC array figures (0.04 to 0.4 pJ per INT8-class MAC, claimed) a chip's k on tile work is therefore 0.03 to 0.3, below the ALU shadow's 0.3 to 0.8, not near 1: at the same premium a tensor-shaped shadow leaves the chip 3.5x to 6.7x where the ALU shadow leaves it 2.1x to 3.5x. The tensor-tile column (2.1x at k = 1, 1.6x at k = 1.5) is withdrawn as a candidate; its premise, that a chip's MAC is no cheaper than the GPU's, is false by 4x to 30x on the public figures. The shipped row is unchanged: 2.1x at k = 1 and 3.4x at k about 0.33, the ALU shadow at the operating point's knee, measured four times at 82 to 90 W.
|
||||
|
||||
The capex column. The `f = 1` GDDR7 chip of 5.5 is USD 2.8 per MH/s of silicon and memory, which is USD 0.00016 per MH/s-hour of capex over two years against USD 0.000023 of electricity: capex-dominated 7x, as the 5090 is (USD 14.7 per MH/s at MSRP, 10x). A 64 MiB hot table adds about USD 15 of N5 die, the shadow core USD 25 to 40, an interposer USD 200, so the chip's capex reaches at most about USD 4.3 per MH/s: the per-unit capex wall is unreachable by 3x to 7x, and the break-even market cap moves only through the project cost (the mission lane's model: about USD 100 M with the N5 shadow core, about 200 M if the shadow runs per load and forces one die or an interposer; the per-load form behind that figure, the 16 x 27 placement, was closed on 7 October 2026 at night when it failed the value-level acceptance test across drawn eras, so the 200 M row rests on no construction shown to exist until a sound per-load class, one pass of a 432-instruction sub-block per load, is drawn, accepted and measured). Every figure here is modelled on cited or claimed parts; the research lane's microbench (20 probes, the mma_u8 and l2 rows the ones this model would take) is on PC 1's queue after the hot-table job.
|
||||
The capex column. The `f = 1` GDDR7 chip of 5.5 is USD 2.8 per MH/s of silicon and memory, which is USD 0.00016 per MH/s-hour of capex over two years against USD 0.000023 of electricity: capex-dominated 7x, as the 5090 is (USD 14.7 per MH/s at MSRP, 10x). A 64 MiB hot table adds about USD 15 of N5 die, the shadow core USD 25 to 40, an interposer USD 200, so the chip's capex reaches at most about USD 4.3 per MH/s: the per-unit capex wall is unreachable by 3x to 7x, and the break-even market cap moves only through the project cost (the mission lane's model: about USD 100 M with the N5 shadow core, about 200 M if the shadow runs per load and forces one die or an interposer; the per-load form behind that figure is REOPENED (8 October 2026, 11:0x UTC): the 16 x 27 iterated placement is dead on the no-era census (0.986 to 0.990 rejection per candidate on both instruments), the 7 October night's drawn-era death was an instrument artefact (the per-load acceptance test counted the era window's fixed top index bits as biased, fixed on counter-asic-4 at 10:5x UTC), and the sound form, one pass of a 256-instruction sub-block per load (`mx8+shl4096x1`), accepts 234 of 256 seeds on the no-era census with its drawn-era verdict the invention lane's 15:00 BST read; the 200 M row rests on that form, class v6's layer 5). Every figure here is modelled on cited or claimed parts; the research lane's microbench (20 probes, the mma_u8 and l2 rows the ones this model would take) is on PC 1's queue after the hot-table job.
|
||||
|
||||
### 5.12 Two chips from the hardware-future lane, and the honest denominator (8 October 2026, 11:3x UK, main's order; the figures lane B's, `docs/analysis/class-v6/hardware-future.md` at master 34f63b3c, modelled on this file's method unless marked claimed; the research lane carries them here as the model's own rows)
|
||||
|
||||
| Chip | When | Energy per random read | Rate per die | Energy per hash at zero shadow | Edge per joule against the 5090 at 2.40 microjoules | Against the M5 Max at 0.78 | With the class v4 shadow (F = 1.10 microjoules) at k = 0.5 / k = 1 | Silicon and project | What the four layers of class v6 do to it | Labels |
|
||||
|---|---|---|---|---|---|---|---|---|---|---|
|
||||
| **A 2 GiB SRAM full store on one N2 reticle** (TSMC N2 38 Mb/mm^2 HD SRAM, +11 percent over N3E, claimed, IEEE Spectrum, 12 December 2024; 2 GiB is about 452 mm^2 of macro under the 858 mm^2 reticle) | 2027 to 2028 (an N2 project; wafers about USD 30,000 and booked to 2028, claimed) | 1.0 nJ for a 64-byte read across the array (0.5 to 2.0: the wire term dominates, this file's 5.1 convention) | power-bound: about 2,100 MH/s at 300 W (0.14 microjoules per hash) | 0.14 | **17x (8x to 30x)** | 5.6x | **4.8x / 2.7x**; at the 5090's whole latency shadow (F about 2.0) 3.7x / 2.0x | about USD 400 to 600 of silicon per die, USD 0.25 to 0.4 per MH/s; an N2 project USD 100 M to 500 M and 18 to 24 months (claimed, the 2.5 economics; the mission lane's break-even model then reads a cap of about USD 330 M to 1.7 B in years 1 to 2 at s = 0.30) | layers 1, 3 and 4: nothing (the dataset is the SRAM's content, every draw firmware; the core an N2 core, k 0.3 to 0.5); layer 2 moves its CAPEX, not its joules: two dies at 4 GiB 15x and about USD 1,000, four at 8 GiB 13x and USD 2,000 to 2,500 | modelled; the N2 density and price claimed |
|
||||
| **A custom HBM4E base die** (the controller and PHY in the stack on N3P at 0.75 V, "2x the power efficiency", claimed, TrendForce, 1 December 2025; Micron 2027, SK hynix 2026; a non-hyperscaler customer from about 2028) | 2028 or later | 0.9 to 1.0 nJ | 166 MH/s per stack at the model's activate ceiling (36 at the JEDEC tFAW: the ceiling is unmeasured, 5.3) | 0.18 (0.37) | **14x (6.5x)** | 4.4x (2.1x) | 4.4x / 2.6x | about USD 5 per MH/s (approximate); the stack at a premium: Samsung asking about USD 4 to 5 per Gbit for HBM4 against 1.5 for HBM3E (claimed, 2 October 2026) | untouched by all four layers; the brake until about 2028 is HBM allocation and price, not the hash | modelled; the die and price claimed |
|
||||
| Per-bank processing in memory (Samsung HBM-PIM, SK hynix GDDR6-AiM and AiMX, Samsung LPDDR5X-PIM at Hot Chips 25 August 2026, LPDDR6-PIM in JEDEC work) | now to 2027 | the host memory's own | | | **under 1x: structurally blind to the hash's dependent random reads** | | | | a per-bank unit sees its own bank only: 1.6 percent of a hash's reads land in-bank at a 2 GiB dataset on a 32 MB bank, 0.4 percent at 8 GiB; every other read crosses the channel as it does for a controller chip; layer 2's growth lowers the share further | modelled (arithmetic on the bank size) |
|
||||
| The honest denominator | | | | the M5 Max at 0.78 microjoules per hash (the GPU and DRAM channels, measured 6 October), 3.1x the 5090 per joule; the 5090 at its knee 1.67 (measured 7 October) | | | | | against the best honest joule every chip edge is 2x to 3x smaller than against the 5090's stock point: the SRAM die 5.6x and the base die 3.6x to 4.4x at zero shadow, about 3x with the shadow at k = 0.5 | measured cards |
|
||||
|
||||
Reading, as the research lane reads it for the class v6 synthesis: the number this file's 5.6 verdict carried ("over 2x", the GDDR7 board at 5.1x) is no longer the strongest chip in the five-year window. The SRAM full store on merchant N2 reaches 17x at zero shadow and 2.7x to 4.8x with the class v4 shadow, and none of the rotating layers reaches it; what holds it is money and time (an N2 project at USD 100 M to 500 M, 18 to 24 months, the break-even cap in the hundreds of millions to about USD 1.7 B) and the one lever the chain has on it is capex through the dataset's size (layer 2's floor and its schedule, priced in the class v6 synthesis). The served chip line is reviewed against these rows at 20:00 UK today in the synthesis's last section; nothing served moves on this section.
|
||||
|
||||
## 6. The per-day derivation (item 2)
|
||||
|
||||
|
|
|
|||
495
docs/analysis/counter-asic-4-research.md
Normal file
495
docs/analysis/counter-asic-4-research.md
Normal file
|
|
@ -0,0 +1,495 @@
|
|||
# Counter ASIC 4.0 research: can the chip's per-joule edge be held near 2x without the GPU paying an energy premium?
|
||||
|
||||
7 October 2026, 20:27 to 21:0x UTC (21:27 to 22:0x UK), branch `counter-asic-4`, the Counter ASIC 4.0 research lane. Written for the founder's question of this evening: class v4 (the latency shadow, about 100,000 counted integer operations per hash hidden in the memory wait) brings the strongest chip's per-joule edge over an RTX 5090 from 5x to 9x down to 2.1x to 3.9x, and costs the 5090 145 W at its stock clock and 88 W at a 1,400 MHz core lock for the same hash rate. The brief: find the design that keeps the chip's edge at or under about 2x per joule WITHOUT the GPU paying an energy premium, or prove that no such design exists and say what the floor is.
|
||||
|
||||
Every number carries a label: **measured** (a card on a named night, the file named), **modelled** (arithmetic on the chip model's cited memory figures, `docs/analysis/chip-model-v3.md` section 5), **claimed** (a vendor's or an author's figure, URL given), **approximate** (from memory or a scaling estimate). Nothing in this file changes a consensus object, a pack or a parameter; the one code change (section 12) is a miner-side kernel variant that is opt-in by name and never raced by default.
|
||||
|
||||
## 0. One page
|
||||
|
||||
**The answer: no hash-side design reaches a chip edge of about 2x at zero premium, and the reason is an identity, not a missing idea.** Against the chip that matters (the stored-dataset chip, "a GPU's memory system without the GPU", chip-model-v3 section 5.5) the per-joule edge is
|
||||
|
||||
edge = (E_card + F) / (E_mem + k F)
|
||||
|
||||
where `E_card` is the honest card's whole-card energy per hash on class v3, `F` the premium per hash the card pays for forced work, `E_mem` the chip's memory-system energy per hash (0.466 microjoules on GDDR7, 0.321 on one HBM3 stack, modelled), and `k` the chip core's energy per forced operation over the card's. A lever with `F = 0` is a lever with `k F = 0`: every joule the chip must spend on forced work is a joule the card spends first. So at zero premium the edge is `E_card / E_mem`, which no hash change touches, and the ONLY things that move it are the card's own watts and the chip's memory choice.
|
||||
|
||||
The floor, on tonight's measured rows (the hash lane's efficiency pass, PC 2, 7 October 2026, `docs/plans/counter-asic-3-status.md` on branch ca3-coord, section "the efficiency pass"):
|
||||
|
||||
| Honest card, class v3 (no shadow) | Energy per hash | Chip edge at zero premium, GDDR7 / one HBM3 stack / eight stacks | Label |
|
||||
|---|---|---|---|
|
||||
| RTX 5090, core unlocked (136.59 MH/s at 330.2 W) | 2.42 microjoules | 5.2x / 7.5x / 9.2x | measured card, modelled chip |
|
||||
| RTX 5090, app stock point, 5 October (124 MH/s at 290 W) | 2.34 | 5.0x / 7.3x / 8.9x | measured, modelled |
|
||||
| **RTX 5090, 1,400 MHz core lock (134.68 MH/s at 228.0 W)** | **1.69** | **3.6x / 5.3x / 6.5x** | measured, modelled |
|
||||
| RTX 5090, the knee (the hash lane's second pass, 20:5x UTC): class v3 best at 1,300 MHz, 134.62 MH/s at 223.3 W; class v4 best at 1,200 MHz, 133.80 MH/s at 305.1 W; the premium at the best points 81.8 W; the floor below 1,100 MHz unmeasured (a third pass runs) | 1.66 (v3); 2.28 (v4) | 3.6x / 5.2x / 6.3x on v3; class v4 at k = 1 (6.1 pJ per counted op) 2.1x | measured, modelled |
|
||||
| RTX 5090, the bound any operating point can reach: idle 74 W (PC 2) plus the memory system's 55 W at 17.5 G reads/s, 135 MH/s | 0.96 | 2.1x / 3.0x / 3.7x | approximate (idle measured; the memory share modelled) |
|
||||
| RTX 4070 at its Ember tune point (30.95 MH/s at 79.5 W) | 2.57 | 5.5x / 8.0x / 9.8x | measured, modelled |
|
||||
| RX 9070 XT (about 18.6 MH/s at about 304 W) | 10.6 | 23x / 33x / 40x | approximate card figure |
|
||||
| Apple M5 Max, GPU plus DRAM channels (27.08 MH/s at 21.0 W) | 0.78 | 1.7x / 2.4x / 3.0x | measured, modelled |
|
||||
|
||||
So the premium-free floor against the GDDR7 chip is 3.6x on a 5090 locked at 1,400 MHz, about 2x at the bound set by that card's own idle and memory power, and 1.7x on an Apple GPU today with no shadow at all. The 2x line at zero premium is a property of the honest card's idle, not of the hash. The shadow buys the rest only with a premium, and only if the chip's core is no better than about half a GPU lane per op: at the 1,400 lock the measured premium (88.3 W, 0.654 microjoules, 6.5 pJ per counted op) gives 2.1x at `k = 1`, 2.9x at a core of 3.3 pJ per op, 4.1x at 1 pJ, and below about 1.8 pJ per op (`k` under 0.28) the shadow turns against us, because it then dilutes the card more than the chip.
|
||||
|
||||
**The ranked designs (section 9 has every column):**
|
||||
|
||||
1. **The operating point as the shipped default** (miner-only; core lock at the knee plus undervolt where the driver allows; AMD through ADLX or nothing; Apple has no lever). Premium: none by construction, it lowers `E_card`. Measured: v3 2.42 to 1.69 microjoules, the v4 premium 145 to 88 W, the class v4 edge at `k = 1` 2.1x. 2 to 4 agent hours (already ordered into Ember Tune for 0.3.24). Breaks: a locked card's free shadow shrinks (about 150,000 ops at 1,400 MHz), so ladder rung 2 is not free on it.
|
||||
2. **The SM-sparse miner kernel** (miner-only; the hash on a fraction of the SMs with 32 warps per block and the rest clock-gated). Reads the unmeasured 100 W between the 5090's idle-plus-memory (about 130 W) and its 228 W at the lock; if half is SM activity the premium-free edge falls toward 2.5x. Built tonight as worker variants `sp<N>-w<W>` (section 12), one PC 1 job owed. Breaks if the 100 W is clock tree and leakage.
|
||||
3. **The shadow kept at rung 0, its op mix re-weighted toward the families where a chip gains least over a GPU lane** (shuffle and multiply heavy). Same premium, the pessimistic-core edge 3.4x to about 2.9x, `k = 1` unchanged at 2.1x. 4 to 6 agent hours plus the six gates (a class change under the 95 percent rule). Held until the knee and the SM-sparse rows are read (the coordinator, 21:5x UK).
|
||||
|
||||
Dead by arithmetic (sections 3 to 7): dropping the shadow after class v5 (v5 leaves the chip at 5.1x to 9.1x), the per-window refresh as a cost (a 1 GiB rebuild is about 0.1 J on a chip core), extra reads or longer chains (both sides scale together), memory-system shaping (the same DRAM on both sides), per-card adaptive classes (a chip picks the class that maximises its own reward per joule, so a menu never beats the card's best single class), proof of useful work (the useful fraction is bounded at 8 percent of the hash budget at 1 GH/s and falls with scale; no ZK ASIC has beaten a GPU per joule on real hardware in the public record; every deployed useful-work scheme was gamed), and the tensor block (`k` near 1 but 0.056 pJ per multiply-add, so there are no joules in it to force without 26x the verifier cost).
|
||||
|
||||
## 1. The measured facts this rests on
|
||||
|
||||
| Fact | Value | Label | Source |
|
||||
|---|---|---|---|
|
||||
| RTX 5090 class v3 and v4 across the core-clock grid, the same hash rate | unlocked: v3 136.59 MH/s at 330.2 W, v4 136.84 at 475.5 (premium 145.3 W); 1,400 MHz lock: v3 134.68 at 228.0 W, v4 134.98 at 316.3 (premium 88.3 W); the rate memory-bound and flat from 2,850 to 1,400 MHz (136.8 to 135.0); the knee below 1,400 being read now | measured, PC 2, 7 October 2026 | `docs/plans/counter-asic-3-status.md` (ca3-coord), the efficiency pass table |
|
||||
| The shadow's marginal energy per counted op on the 5090 | 10.4 pJ unlocked (145.3 W over 136.84 M hashes x 102,100 ops); 6.5 pJ at the 1,400 lock (88.3 W over 134.98 M x 102,100); 10.2 to 13.2 pJ on the 6 October ladder under the 431 W cap | measured (arithmetic on measured rows) | the same; `docs/analysis/latency-shadow-2026-10-06.md` section 5 |
|
||||
| The 5090's ALU budget and where the shadow binds | 45.2 T op/s at 3,050 MHz; about 332,000 ops per hash before compute binds at 136 MH/s; at a 1,400 MHz lock the budget scales to about 20.7 T op/s and the bind point to about 150,000 ops | measured budget; the scaling approximate | `docs/benchmarks/repro.md` 2.2; latency-shadow section 5 item 1 |
|
||||
| RTX 5090 idle | 73.8 W at 862 MHz (PC 2, 6 October); 90.6 W (PC 1, 5 October) | measured | latency-shadow section 5; the Ember readbacks in the same file |
|
||||
| RTX 4070 at its tune point (1,860 MHz lock, 160 W cap) | class v3 30.95 MH/s at 79.5 W (2.57 microjoules); class v4 3.51 microjoules, 109 W: the owner pays 30 W more; the rate holds to about 200,000 ops | measured, PC 1, 6 October | status file item 8, the 4070 rows; `docs/analysis/horizon/algorithm.md` 5.1 |
|
||||
| RX 9070 XT | about 18.6 MH/s at about 304 W (10.6 microjoules); the shadow costs it no rate at any rung (+3.6 percent at rung 3); its watts under class v4 OWED | approximate (watts from memory); rate measured | algorithm.md 5.1; `docs/design/latency-ladder.md` section 8 |
|
||||
| Apple M5 Max, GPU plus DRAM channels | class v3 27.08 MH/s at 21.0 W (0.78 microjoules); class v4 at 100,000 ops 1.40 microjoules (+16 W, -1.5 percent of rate); marginal 6.9 pJ per counted op | measured, 6 October | latency-shadow section 3 |
|
||||
| The stored-dataset chip (f = 1) | GDDR7, the 5090's own 16 devices without the GPU: 166.4 MH/s at 77.6 W, 0.466 microjoules; one HBM3 stack 83.6 MH/s at 26.8 W, 0.321; eight stacks 666 MH/s at 174 W, 0.262; a node and executor beside it about 85 W (approximate) | modelled (activate-bound ceilings, 2.0 and 1.2 nJ per random 32-byte read, static and controller allowances) | chip-model-v3 sections 5.3 to 5.5 and 5.10 (the Counter lane's class v5 rows, relayed 20:5x UTC) |
|
||||
| Class v5 at zero shadow | GDDR7 5.1x per joule against the 5090 at 2.40 microjoules (2.5x with a node per chip); one HBM3 stack 7.2x (1.8x with a node per chip); eight stacks 9.1x (6.2x); the leaves ship at 16.5 KB/s to 10,000 members today, so a farm shares one node | modelled | chip-model-v3 section 5.10; `docs/design/class-v5-stored-state.md` 2a.2 |
|
||||
| The tensor-shaped block (mm8) on an RTX 4090 | free in rate to 4,096 tiles per hash; 2.9 to 14.7 W; 0.70 falling to 0.056 pJ per multiply-add; verifier +1.1 microseconds per step per unit scalar | measured, 6 October | `docs/analysis/horizon/new-pow.md` 5.1 |
|
||||
| Proving on the 5090 | a shard at the adopted v1 budget (4.7 M SP1 cycles) 4.3 s alone, 13.2 s beside the miner; the prover alone peaks at 190.6 W on empty shards; a GPU-proved block's useful fraction bound 8 percent of network hashes at 1 GH/s | measured; the bound is arithmetic | `docs/bench-log.md` "proving v1" and "the S_p curve"; new-pow section 6 |
|
||||
| The verifier gate | class v4 rung 0 8.77 ms cold with the sibling loaded on the reference core (5.14 alone); rungs 0 to 2 admissible, rung 3 out by 0.08 ms; class v5 adds 0.20 to 0.28 ms | measured, igneum-build-1, 6 and 7 October | latency-ladder section 5; class-v5 section 7 |
|
||||
|
||||
## 2. The identity, and what it forbids
|
||||
|
||||
Write the card's energy per hash on class v3 as `E_card = E_idle + E_memsys + E_wait`: the idle floor spread over the hashes (74 W over 135 MH/s is 0.55 microjoules on the 5090), the memory system's own reads and standby (about 55 W, 0.41 microjoules, modelled), and the SMs, fabric and clock tree while they wait on loads (the rest: 228 - 74 - 55 = 99 W, 0.73 microjoules at the lock, unmeasured as a breakdown). The chip pays `E_mem` (its memory, controller and static: the whole 0.466) and nothing else. Add forced work `W` ops per hash at `e_gpu` per op on the card and `e_chip` on the chip, `k = e_chip / e_gpu`, `F = W e_gpu`:
|
||||
|
||||
| Quantity | Formula | 5090 at the 1,400 lock, GDDR7 chip | 5090 unlocked, GDDR7 | 5090 lock, one HBM3 stack |
|
||||
|---|---|---|---|---|
|
||||
| Edge at zero premium | `E_card / E_mem` | 3.63x | 5.19x | 5.27x |
|
||||
| Edge with the premium `F` | `(E_card + F) / (E_mem + k F)` | at F = 0.654 (88 W): 2.07x at k = 1 (e_chip 6.5 pJ); 2.92x at 3.3 pJ; 3.50x at 2.0 pJ; 4.12x at 1.0 pJ | at F = 1.062 (145 W): 2.27x at k = 1 (10.4 pJ); 4.33x at 3.3 pJ; 6.12x at 1.0 pJ | 2.38x at k = 1; 3.56x at 3.3 pJ; 5.54x at 1.0 pJ |
|
||||
| The asymptote at infinite premium | `1 / k` | 1x at k = 1, 2x at k = 0.5 | the same | the same |
|
||||
| The premium that reaches 2x at k = 1 | `E_card - 2 E_mem` | 0.761 microjoules, 103 W | 1.486, 203 W | 1.051, 142 W |
|
||||
| The premium that reaches 2x at k at or under 0.5 | none | | | |
|
||||
| The slope at F = 0 | `(E_mem - k E_card) / E_mem^2` per microjoule | k = 1: -5.65x per microjoule; k = 0.5: -1.75; k = 0.3: -0.19; k = 0.275: 0 | k = 1: -9.0; k = 0.19: 0 | k = 1: -13.3; k = 0.19: 0 |
|
||||
|
||||
Three consequences, each labelled arithmetic on the rows above:
|
||||
|
||||
1. **Zero premium is zero forcing.** `F = 0` gives `k F = 0` for every `k`. A design that costs the GPU nothing extra costs the chip nothing extra. The only exceptions would be levers with an asymmetry the identity does not see: work the card already pays for in `E_idle + E_wait` that the chip does not pay for. Sections 4 and 5 look for one and find none that a consensus rule can check (the GPU's L2 is the nearest, section 4 row 5, and it costs rate).
|
||||
2. **The shadow's worth is a bet on `k`, and the bet gets narrower at the knee.** At the 1,400 lock the shadow buys 1.5x of edge at `k = 1`, 0.7x at a 3.3 pJ core, nothing at 1.8 pJ, and below that it costs us. The reference the project has used for a pessimistic core (Bitmain's withdrawn X9 at about a third of a desktop CPU's energy per RandomX hash, claimed; latency-ladder section 5a) is a ratio against a CPU core, and a CPU core spends about 100 pJ or more per instruction against a GPU lane's 6.5 to 10.4 pJ measured here (Horowitz 2014 gives 70 pJ per in-order instruction at 45 nm, Dally 2023 gives 250 pJ per out-of-order ARM instruction; both claimed). A chip three times better than a CPU core on a random program is a chip WORSE than a GPU lane per op. So the X9 implies nothing about `k` against a GPU; the honest band for a chip maker who builds a SIMD array like a GPU's is `k` 0.5 to 1 (approximate), with the N5 fixed-datapath floor (0.19 pJ of datapath before register file, operand wires and instruction fetch; chip-model-v3 5.1) as the theoretical bound that nobody has shown on a random 32-lane program with shuffles. The honest headline stays 2.1x at `k = 1`, and the project should stop quoting the X9 as a `k` of 0.33 against a GPU.
|
||||
3. **What moves the premium-free floor is `E_card`, and the hash cannot touch it.** The 5090's own idle (74 to 91 W measured) plus its memory system (55 W modelled) bound the floor near 2.0x on GDDR7 at any operating point; the Apple GPU at 21 W shows a GPU designed for low idle sitting at 1.7x with no shadow. The measurements that read how much of the 99 W of `E_wait` is reachable are the knee sweep (running) and the SM-sparse kernel (section 12).
|
||||
|
||||
## 3. Thread 1: class v5 with the shadow at zero
|
||||
|
||||
The question: once class v5 is live (the dataset built from the chain's execution state, refreshed per window), does the shadow still buy anything? Re-run of the chip model with v5 on and the shadow at zero (ladder rung below 0, which is class v3's energy, since no rung below 0 exists and dropping the shadow is the class object moving back to v3 under the 95 percent rule):
|
||||
|
||||
| Chip under class v5, shadow at zero | Rate, watts | Energy per hash | Edge over the 5090 at 2.40 microjoules (the model's denominator) | Edge at the 1,400 lock (1.69) | With a node and executor per chip (85 W, approximate) | Label |
|
||||
|---|---|---|---|---|---|---|
|
||||
| GDDR7, 16 devices | 166 MH/s at 78 W | 0.466 | 5.1x | 3.6x | 0.977 microjoules: 2.5x (1.7x at the lock) | modelled (the Counter lane, chip-model-v3 5.10) |
|
||||
| HBM3, one stack | 84 MH/s at 28 W | 0.321 | 7.2x | 5.3x | 1.34: 1.8x (1.3x) | modelled |
|
||||
| HBM3, eight stacks | 666 MH/s at 175 W | 0.262 | 9.1x | 6.5x | 0.389: 6.2x (4.3x) | modelled |
|
||||
|
||||
What v5 adds to the chip's bill is the window's leaves, 64 bytes per state record, 5,952 bytes at today's devnet state and at most the dataset's size at the sample cap (class-v5 2a.2), delivered over a link: 16.5 KB/s to 10,000 members today; the WAN line is about 700,000 state records (45 MB per member per hour at 1 Gbit/s), a LAN farm is never bounded. So the "node per chip" column is not reachable by construction: one node serves a farm, as the v5 design states itself. The arithmetic the founder asked for: under v5 alone the strongest chip keeps 5.1x (GDDR7) to 9.1x (eight HBM3 stacks) at the stock point and 3.6x to 6.5x at the lock, all over 2x. **The answer is not "drop the shadow after v5."** What v5 does remove is the f = 0 recompute chip as a category and the pool miner who holds nothing but the day key (class-v5 section 1); its merit is there, not against the memory-system chip.
|
||||
|
||||
A construction that WOULD force a node per card does not exist below the state size: everything the lottery derives is derived from a seed and the state, so the only bytes a central node cannot compress away are the state's, and today's state is 6 KB. The refresh cadence does not help either (section 4 row 7).
|
||||
|
||||
## 4. Thread 2: energy-shaped work instead of op-shaped work
|
||||
|
||||
The brief's hypothesis: some operations cost a GPU almost nothing extra in energy because the hardware already pays for them during the memory wait. The identity says what to look for: an operation class with `k` as high as possible (a chip can improve on the GPU least) at a cost `e_gpu` that is REAL joules (so the forcing exists). A class with tiny `e_gpu` forces nothing; a class with tiny `k` forces the card and spares the chip. Per class, the figures:
|
||||
|
||||
| Operation class | Card's marginal energy per op | A chip's cost to match it | k band | Premium per 0.1x of edge bought at the 1,400 lock, GDDR7 (from the slope at F = 0) | Verifier cost | Reading | Labels |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| Random 32-bit integer ALU program (the class v4 shadow: add, xor, mul, mad, rotates, sub, mulhi, or) | 5090: 10.4 pJ unlocked, 6.5 pJ at the 1,400 lock; M5 Max 6.9 pJ | a SIMD array at N5: 0.19 pJ of datapath plus register file, operand network and instruction fetch; the model's columns 3.3 pJ (k 0.3 of 11) to 11 pJ | 0.5 to 1 honest band; 0.15 to 0.3 the attacker's claim (unshown) | k = 1: 0.018 microjoules (2.4 W) per 0.1x; k = 0.5: 0.057 (7.7 W); k = 0.3: 0.53 (71 W) | 3.2 microseconds per 1,000 shadow instructions per warp on the M5 Max core; rung 0 fits with 1.2 ms to spare on the loaded reference core | the right lever IF k is near 1; its value collapses below k = 0.5 and reverses at 0.28 | card measured; chip approximate; slope arithmetic |
|
||||
| Shuffle-heavy mix (shfl, shfla through the warp crossbar) | the shfl step 0.86x of the add-xor-rotate chain on the M5 Max, 1.13x for rotr; the 5090's per-family step costs in the item 6 record; AMD shfla 0.75 to 0.84 | a crossbar is the same wire physics on any die; the algorithm lane's k floor rises from 0.32 to 0.46 with a shuffle-heavy weight table (approximate) | 0.46 to 1 | at k = 0.46: 0.075 microjoules (10 W) per 0.1x | the same law; shfl is one op in the interpreter | the content to prefer inside the shadow: the same premium, less of it reaches a chip | measured steps; k floor approximate (algorithm.md 5.3) |
|
||||
| 32-bit multiply heavy | a 32 x 32 multiplier is 0.52 pJ at N5 datapath on both sides (claimed, mlsysbook citing Horowitz and Dally); the GPU's overhead around it is the operand delivery | the same multiplier | about 0.5 to 1 | as the shuffle row | one op | second choice behind shuffles; the saturation and lossy-source rules of sub-version 3 already bound what a multiply can do to the read map | claimed figures |
|
||||
| Tensor-shaped int8 tiles (mm8) | 0.056 pJ per multiply-add at 4,096 tiles per hash on the 4090 (57 pJ per 1,024-MAC tile); 14.7 W in all | a licensable int8 MMA block: the same circuit | 0.5 to 1 | to reach the shadow's 0.654 microjoules the block needs about 11,500 tiles per hash (R about 1,450); the 4090 binds near R 3,000 | scalar 1.1 microseconds per step per unit: about 10.6 ms at R = 1,450 on the M5 Max core, FAILS the gate alone; a byte-dot verifier (AVX-VNNI, NEON udot, AVX2 pmaddubsw) 4x to 16x cheaper, approximate, unmeasured | k is good and the joules are too cheap: no forcing without the verifier paying 26x per joule forced (new-pow 5.1 item 3). Kept as reserve R8 for datapath diversity | measured block; verifier scaling approximate |
|
||||
| L2-resident hot reads (a second table sized to GPU L2, read beside the dataset) | an L2 hit on a 96 MB L2: about 0.1 to 0.3 nJ per 32-byte sector (approximate: Horowitz's 1 MB cache at 100 pJ per 64 bits at 45 nm, scaled to N5 and a 96 MB array with its crossbar); measured in the added form the table cost the 5090 13 to 16 percent of RATE because the dataset stream evicts it (5 October, no cache-policy hints) | a 64 MiB SRAM read on a 128 mm^2-class array: 0.2 to 0.5 nJ (chip-model-v3 5.1, approximate) | 0.7 to 3: the one class where a chip may be WORSE than the GPU per op, because a big SRAM array is not cheaper than a GPU's L2 slice | at k = 1.5: 0.008 microjoules (1 W) per 0.1x, if the hot reads were free in rate | the hot table's items are derived lazily on the CPU (unmeasured per unit) | worth ONE measurement, not a claim: the 5 October rows used no cache hints; CUDA `createpolicy.fractional.L2::evict_last` for the hot table and `ld.global.L2::evict_first` for the dataset stream, AMD's slc bits; nothing on Metal, so Apple pays rate. If the rate holds with hints, this is the only lever with k possibly over 1 | approximate throughout |
|
||||
| Dependent pointer chasing: more reads per hash, a longer chain | the card's incremental energy per read about 3.1 nJ (55 W over 17.5 G reads/s, modelled); its overhead power is per second, so a longer chain lowers both rates together | 2.0 nJ per read | the edge is unchanged: both sides' energy per hash scale by the same factor (E_card and E_mem are both per-second powers over a rate that moves together) | none bought at any premium | +0.13 ms per 128 reads per unit, approximate | not a lever (chip-model-v3 5.7 row "latency") | modelled |
|
||||
| Extra random reads at the same bandwidth (wider loads, w16, w32) | the 5090 fetches a 32-byte sector per 4-byte read already; w16 moved the 5090 2.7 percent and the chip not at all; w64 made the 5090 bandwidth-bound (71.9 MH/s) | the same 32-byte atom per channel access | 1 | nothing to force: the chip already pays the sector | none | not a lever; the 4-byte decision stands | measured (read-width) |
|
||||
| Per-lane live state (a scratch, register-resident permutations across the hash) | a chip keeps 1,172 lanes at 55 ns of controller latency against the 5090's 7,262 at 415 ns; lane state is 64 B to 320 B | SRAM at picojoules per access | not applicable: lane state is not an energy term on either side | none | | `docs/analysis/scratch-soundness.md`: the live state is bounded by the read-modify-write count and sits in a chip's SRAM at under 5 percent of its mirror | measured bound |
|
||||
| The dataset refresh as the cost (class v5's rebuild per window, per epoch or per block) | a 1 GiB rebuild is 2^24 items at about 9,360 ops each, 157 G ops: 32 ms on a 4090, 13.4 ms on the 5090 | 157 G ops at 1 to 3 pJ per op is 0.16 to 0.5 J per rebuild; a rebuild per block is 0.16 to 0.5 W against 78 W of hashing; the write is 1 GiB per second against 1.8 TB/s | the refresh is 0.1 percent of the chip's item traffic | nothing | a rebuild per second costs the honest 4090 3 percent of its hash time (rate, not joules) | dead by arithmetic; the refresh cadence is a liveness tool (proof of following), not an energy lever | arithmetic on measured build times and approximate chip energies |
|
||||
|
||||
What the table says: the founder's instinct is right that the shadow spends energy because it uses ALUs, and the public energy figures agree on why a GPU lane is expensive per op (Dally, Hot Chips 2023, claimed: fetch, decode and operand delivery about 30 pJ per lane instruction at 45 nm against 0.1 pJ for the add itself; a tensor instruction amortises that overhead over 1,024 multiply-adds, which is exactly why the mm8 block is cheap for the card and useless as a forcing lever). The classes where the card's energy is "already paid" during the wait (lane state, in-flight reads, the memory burst) are classes where the chip's cost is zero or equal, so they force nothing. The one class where a chip might pay more than the GPU per op is the GPU's own L2, and the one measurement of it lost rate; a cache-hinted re-run is the open item. Everything else reduces to choosing the shadow's content for the highest `k` at a fixed premium, which is design 3.
|
||||
|
||||
## 5. Thread 3: memory-system binding beyond random reads
|
||||
|
||||
Row-buffer-miss patterns, bank-conflict shaping, refresh-aware access and burst-sized items were each checked against the identity. The hash already opens a row per read (random 4-byte reads over 1 GiB: a row hit rate near zero on both sides); a 32-byte item is already the GDDR7 channel burst (the 5090 measures a 32-byte sector per read; the 9070 XT a 64-byte line, which is AMD's cost, not the chip's); bank parallelism is the same 1,024 banks on the 5090's board and on a chip built from the same 16 devices; refresh is the DRAM's. **There is no asymmetry to shape: the chip and the card have the same DRAM, the same row cycle (tRC about 45 ns across DDR4, GDDR and HBM; Li, Reddy and Jacob, MEMSYS 2018, claimed) and the same activate window.** The two levers a memory system offers are both the chip's: a lower energy per read (HBM's 1.2 nJ against GDDR7's 2.0 through 0.8 pJ per bit of interposer I/O against a PCB, modelled; fine-grained DRAM would cut the 909 pJ activation further, O'Connor et al. MICRO 2017, https://www.cs.utexas.edu/~skeckler/pubs/MICRO_2017_Fine_Grained_DRAM.pdf, claimed) and more activates per second per watt (AMD's "Folded Banks", ISCA 2025, 6.7x the irregular bandwidth from 8x the activate parallelism, https://dl.acm.org/doi/10.1145/3695053.3731111, claimed). Both move `E_mem` down, so the premium-free edge can only rise with memory generations, which is the HBM4 point the Horizon lane made (algorithm.md proposal 5).
|
||||
|
||||
What the Ethash and RandomX record measured (the literature sub-agent's pass tonight, 7 October 2026; every figure claimed by its vendor unless marked):
|
||||
|
||||
| Algorithm, what it relied on | The chip, its date and figures | The GPU at the time | Chip edge per joule | Where the resistance failed | Sources |
|
||||
|---|---|---|---|---|---|
|
||||
| Ethash: 4 GB+ DAG, 64 random 128-byte reads per hash, bandwidth-bound | Antminer E3 (Jul 2018) 190 MH/s at 760 W; Innosilicon A10 Pro (2020) 500 at 950; Linzhi Phoenix (Dec 2020) 2,733 MH/s at about 3,000 W (F2Pool test, measured); Jasminer X4 (Nov 2021) 2,500 at 1,200; Antminer E9 Pro (2023) 3,680 at 2,200; Jasminer X16-P (2024) 5,800 at 1,900 | RTX 3080 about 100 MH/s at 220 W (0.45 MH/W, measured) | E3 0.55x; A10 Pro 1.2x; Phoenix 2.0x; X4 4.6x; E9 Pro 3.7x; X16-P 6.8x | the chips won by moving the same bytes at lower energy per bit (custom controllers, on-package and on-die memory: Linzhi's 72 mixers and 2.8 Tb/s of internal memory); nothing in the hash's logic | https://www.theblock.co/post/88622/questions-new-ethash-asic-ethereum ; https://f2pool.io/mining/pow-round-up/20201228-pow-round-up/ ; https://www.asicminervalue.com/miners/jasminer/x4 ; https://pool.kryptex.com/ms/device/asic/jasminer/x16-p ; https://ethereum-magicians.org/t/progpow-audit-delay-issue/3309?page=4 |
|
||||
| RandomX: 2 GiB dataset, one 64-byte read per iteration, 8 chained random programs, CPU-shaped | Antminer X5 (Sep 2023) 212 kH/s at 1,350 W; Antminer X9 announced 26 Dec 2025 at 1,000 kH/s and 2,472 W, withdrawn May 2026 before any unit shipped; Pinecone R1X 1,200 kH/s at 2,055 W (unverified) | Ryzen 9 3950X 12.9 kH/s at 105 W (measured); GPUs 18x worse per joule than the CPU | X5 1.3x; X9 3.3x (claimed, never shipped); R1X 4.7x (unverified) | against its honest device the latency-bound random program held a chip to 1.3x shipped and 3x to 5x claimed after seven years; RandomX v2 testnet forked 5 October 2026 | https://github.com/tevador/RandomX/blob/master/doc/design.md ; https://pool.kryptex.com/articles/antminer-x5-en ; https://github.com/monero-project/monero/issues/10270 ; https://oneminers.com/blogs/news/whatever-happened-to-the-antminer-x9-bitmain-monero-miner |
|
||||
| ProgPoW and KawPoW: random math per period on the GPU's datapath, 256-byte DAG loads | none shipped as of 2026 | RTX 4090 65 MH/s at 330 W | the authors' claim 1.1x to 1.2x; Rao's audit (Sep 2019): conventional chips gain little, memory-intensive algorithms face an "imminent threat" in general | the resistance held; the energy per hash on GPUs is the price (KawPoW 0.2 MH/W against Ethash 0.45 on the same cards) | https://github.com/ifdefelse/ProgPOW/blob/master/README.md ; https://leastauthority.com/?p=759 |
|
||||
| Equihash 200,9: a 144 MB sort | Antminer Z9 mini (May 2018) 10 kSol/s at 247 to 300 W; Z15 (Jun 2020) 420 kSol/s at 1,510 W; Z15 Pro (2023) 840 kSol/s at 2,780 W | Vega 64 506 Sol/s at about 250 W | 17x to 151x | the working set fitted on-die and the sort pipelines; Zcash voted 45 to 19 against resistance (Jun 2018) | https://overclock3d.net/news/misc/bitmain-launches-their-antminer-z9-mini-zcash-mining-asic/ ; https://support.bitmain.com/hc/en-us/articles/19259431024281 |
|
||||
| Cuckatoo32: a 512 MB SRAM random walk | iPollo G1 (Dec 2020) 36 graphs/s at 2,800 W | RTX 3090 1.0 to 1.43 graphs/s at 285 W | 2.6x to 3.7x | the random walk stayed latency-bound even in SRAM: the smallest edge of any compute-shaped chip | https://pool.kryptex.com/device/asic/ipollo/g1 ; https://forum.grin.mw/t/mining-hardware-comparision/9011 |
|
||||
| kHeavyHash, Blake3, SHA (compute only) | IceRiver KS0 to Antminer KS5 Pro (2023 to 2024); Antminer AL1 (2024) | RTX 4090 | 160x to 700x | nothing memory-bound to hold them | https://asicminervalue.com/miners/bitmain/antminer-ks5-pro-21th ; https://pool.kryptex.com/device/asic/bitmain/antminer-al1 |
|
||||
|
||||
The reading for Igneum: Igneum's hash is in the Ethash and RandomX class (DRAM-bound), and the Ethash chips' 2x to 7x is exactly `E_card / E_mem` with a better memory system; the Cuckatoo and RandomX rows show what latency-bound work does to a chip's logic edge (3x against a CPU, under 1.3x shipped). The identity predicts both.
|
||||
|
||||
## 6. Thread 4: per-card adaptive work
|
||||
|
||||
The question: could the protocol admit per-class programs where the shadow is scaled to the card's own memory-to-compute ratio, without letting a chip pick the cheapest? **No, and the proof is one line.** Let the menu be a set of classes `c`, each with a difficulty weight `w(c)` (reward per hash under that class). An honest card's best reward per joule is `max_c w(c) / E_card(c)`, a chip's is `max_c w(c) / E_chip(c)`. The chip's edge under free choice is the ratio of the two maxima. Since the chip may pick the card's own best class `c*`, `max_c w(c) / E_chip(c) >= w(c*) / E_chip(c*)`, so
|
||||
|
||||
edge(menu) >= w(c*) / E_chip(c*) over w(c*) / E_card(c*) = edge(c*),
|
||||
|
||||
the edge at the card's best single class. A menu can never beat the best single class for the honest card, and whenever any class on the menu is cheaper for the chip than `c*` is, the menu is worse. The chain cannot verify hardware, only values, so no rule can bind a class to a card type (new-pow 3.2 makes the same point about timings and capacities). The ladder (`docs/design/latency-ladder.md`) is the construction that survives this: one network-wide `N`, stepped by miner signal within a verifier-bounded genesis list, which the attack-pass lane's F10 row gated (monotone, 89 percent does not move it, no two-step jump). What a per-card ratio CAN do is sit in the miner, not the protocol: a card picks its own clock, grid and kernel shape (designs 1 and 2) for the one network class.
|
||||
|
||||
## 7. Thread 5: proof of useful work
|
||||
|
||||
Igneum's miners are its provers. Could the proving replace or fund the shadow?
|
||||
|
||||
| Quantity | Value | Label | Source |
|
||||
|---|---|---|---|
|
||||
| Energy per proven shard on the 5090 tonight (the adopted v1 budget, 4.7 M SP1 cycles) | 4.3 s alone at up to 190.6 W: at most about 820 J per shard, at least about 5,700 cycles per joule (174 microjoules per cycle); beside the miner 13.2 s | measured time and the prover's peak watts (the watts for the v1 shard itself unmeasured, the empty-shard peak is the bound) | bench-log "the S_p curve" and "proving v1 step 1" |
|
||||
| The hash for comparison | 590,000 hashes per joule at the 1,400 lock | measured | section 1 |
|
||||
| What the chain needs proven | 1 block per second, 2 shards per block at the v1 budget: about 1.6 kW of 5090 proving network-wide, independent of the hash rate | arithmetic | the same |
|
||||
| The useful fraction of a proving-as-lottery scheme | proving work per segment over network hashes per segment: 8 percent at 1 GH/s, 0.08 percent at 100 GH/s; gas sets the numerator, the security budget the denominator; a verifier that recomputes the piece or checks a 32 to 40 ms proof cannot sit under the 10 ms gate | arithmetic on measured figures | new-pow sections 3.1 and 6 (scheme A, NEVER) |
|
||||
| Public GPU zkVM proving rates | Airbender 21.8 MHz of RISC-V cycles on an H100 and 9.7 MHz on a 4090 (about 16 to 46 J per million cycles at TDP); SP1 Turbo 3.45 MHz on an H100; RISC Zero 1.1 MHz; SP1 Hypercube 16 RTX 5090s for 99.7 percent of Ethereum blocks under 12 s (Nov 2025), "12.5x GPU efficiency in 6 months" | claimed | https://zksync.io/airbender ; https://blog.succinct.xyz/real-time-proving-16-gpus/ |
|
||||
| Proving chips | Cysic ZK Air "comparable to 10 RTX 4090", ZK Pro "50", no watts published, shipping slipped to 2026; Ingonyama's ZPU paper model "about 13x the efficiency of an A40" with no silicon; Fabric Cryptography's VPU "orders of magnitude", no number, no watts, still pre-order; academic ASICs (SZKP, NoCap, UniZK, zkSpeed) 46x to 800x against CPUs or a 2017 V100, all simulation, bandwidth-hungry (2 to 3 TB/s of HBM) | claimed | https://docs.cysic.xyz/hardware-products/zk-asic-products/ ; https://www.ingonyama.com/oldblogs/zpu-the-zero-knowledge-processing-unit ; https://www.fabriccryptography.com/ ; https://arxiv.org/html/2408.05890v1 ; https://arxiv.org/html/2504.06211v1 |
|
||||
| The one measured non-GPU ZK data point | Jane Street's Hardcaml MSM, ZPrize FPGA winner: 2^26 MSM in 5.08 s at 52 W, about 264 J; the GPU winner on an A40 about 0.56 s per MSM, at most about 170 J at TDP: the FPGA lost on energy | measured times; energy approximate | https://blog.janestreet.com/zero-knowledge-fpgas-hardcaml/ ; https://github.com/matter-labs/z-prize-msm-gpu-combined |
|
||||
| What happened to useful-work schemes | Aleo dropped PoSW consensus after "a small number of provers developed specialized hardware" and rewrote the puzzle when miners optimised MSM/NTT (2022 to 2024); Qubic's rented CPUs took 27 to 51 percent of Monero's hash and reorged it (Aug 2025); Boundless PoVW's token fell 96 percent from its high and RISC Zero closed its hosted prover (Dec 2025); Pearl cuPOW's 112 MW did "zero useful AI computation" | measured events | https://equilibrium.co/writing/aleo-mainnet-launch-reflecting-on-the-journey ; https://aleo.org/post/prover-incentives-2-retrospective/ ; https://www.theblock.co/post/364496/qubics-monero-hashrate-controversy ; https://blockeden.xyz/blog/2026/01/14/boundless-risc-zero-decentralized-proof-market-zk/ ; https://arxiv.org/abs/2606.04819 |
|
||||
|
||||
Does it change the question? **No, in two ways.** (1) The useful fraction is bounded by gas, not by the hash budget: at any scale past 1 GH/s the chain cannot spend its security budget on proving because there is not enough proving to spend it on, and the per-hash verifier cannot check a proof's piece under 10 ms. (2) Proving is compute-shaped (NTT, hashing over 31-bit fields, MSM), the shape in which chips reach 46x to 800x in simulation and the shape that Aleo's provers specialised against; making the lottery out of it hands the chip the lottery. Proving chips do not exist as measured products today (no vendor publishes watts; the one measured FPGA lost to the GPU), and the public GPU proving software moved 12.5x in six months, which a fixed pipeline cannot follow; so for 2026 to 2028 a proving chip's credible edge over a 5090 is 5x to 30x per joule (approximate, from the simulation claims against obsolete GPUs discounted by the software trend), which is worse than the hash's 2.1x to 3.6x. The 80/20 split stands; the energy spent on proving is defensible only because paying users demand the proofs.
|
||||
|
||||
## 8. Thread 6: the literature since 2023
|
||||
|
||||
| Item | What it did | Cost to GPUs in energy | What chips did | Date, source |
|
||||
|---|---|---|---|---|
|
||||
| Kaspa's chip turn | kept kHeavyHash; "ASIC resistance only delays the unavoidable" | none (the GPUs left) | IceRiver KS0 162x per joule over a 4090, Antminer KS5 Pro 700x; GPU share negligible within months | 2023 to 2024; https://kaspa.org/?p=45054 ; https://www.asicminervalue.com/miners/iceriver/ks0 |
|
||||
| Karlsen, Iron Fish to FishHash (Ethash-derived, 4.6 to 5 GB DAG) | the Kaspa forks that wanted GPUs back took an Ethash-class DAG | 4090 82.5 MH/s at 260 W (0.32 MH/W, measured) | none has followed; the Ethash chips' 2x to 7x is the precedent waiting | Apr and Sep 2024; https://pool.kryptex.com/articles/ironfish-hardfork-en |
|
||||
| Alephium Blake3 | compute | n/a | Antminer AL1 about 200x per joule | 2024; https://pool.kryptex.com/device/asic/bitmain/antminer-al1 |
|
||||
| Autolykos v2 (Ergo) | a 2 GB table growing 5 percent per 51,200 blocks to about 66 GB | 4090 265 MH/s at 240 W (1.1 MH/W) | none through 2026 | https://docs.ergoplatform.com/mining/asic/ |
|
||||
| NexaPoW (Schnorr per attempt) | "useful ASICs" | 4090 0.84 MH/W | Dragonball A21 1.89 MH/W: 2.2x | Dec 2024; https://pool.kryptex.com/articles/dragonball-a21-en |
|
||||
| XelisHash v2 and v3 | ChaCha8 plus Blake3 scratchpad, sequential | GPU-mined | none; forked v2 for FPGA resistance (Jul 2024), v3 Dec 2025 | https://github.com/xelis-project/xelis-hash |
|
||||
| RandomX v2.0 | released 25 to 30 Mar 2026, testnet fork 5 Oct 2026 | CPU-mined | the response to the X5 and the announced X9 | https://www.coinpro.ch/en/?p=42081 |
|
||||
| PoSME (Condrey, arXiv 2604.15751, 17 Apr 2026; IETF draft May 2026) | latency-bound pointer chasing, 32 MiB to 2 GiB arena; claims an HBM3 bound of 1.6x to 2.0x, wafer-scale SRAM 5x to 28x, verifier about 6 ms | GPUs 14x to 19x slower than a CPU (claimed) | no deployment | https://arxiv.org/abs/2604.15751 |
|
||||
| Bandwidth-hard functions from random permutations | formal bandwidth hardness | theory | | https://arxiv.org/pdf/2207.11519 (2022) |
|
||||
| Qubic uPoW on Monero | rented CPU fleets at 27 to 51 percent of hash | economic, not silicon | a six-block reorg, 12 Aug 2025 | https://coinmetrics.substack.com/p/state-of-the-network-issue-326 |
|
||||
|
||||
Nothing since 2023 found a lever outside the identity. The latency-bound family (RandomX, PoSME, Cuckatoo, Igneum) holds a chip's LOGIC edge to about 1.3x to 3x against the honest device; the memory-side edge is `E_card / E_mem` everywhere, and the one project that has lived with it longest (Monero) answers with re-tunes, not with a trick.
|
||||
|
||||
## 9. The ranked list of designs
|
||||
|
||||
Each row: (a) the chip's per-joule edge over the 5090 with its labels, (b) the 5090's and the 4070's energy premium over class v3, (c) the verifier cost against the 10 ms gate, (d) build cost in agent hours, (e) what breaks it. The chip is the f = 1 GDDR7 chip unless named; HBM3 one stack in brackets.
|
||||
|
||||
| Rank | Design | (a) Chip edge per joule over the 5090 | (b) Premium over class v3: 5090 / 4070 | (c) Verifier | (d) Agent hours | (e) What breaks it |
|
||||
|---|---|---|---|---|---|---|
|
||||
| 1 | **The operating point as the shipped default**: core lock at the knee plus undervolt where the driver allows, per card model, in Ember Tune (miner-only, no consensus) | class v3 at the 1,400 lock 3.6x (5.3x), measured card, modelled chip; at the knee below 1,400 lower (being read); class v4 at the lock 2.1x at k = 1, 2.9x at a 3.3 pJ core (measured premium, modelled chip) | lowers the base: 330 to 228 W on v3; the v4 premium 145 to 88 W (both measured); the 4070 already sits at its tune point (79.5 W; its premium 30 W measured) | none | 2 to 4 (the Ember core-clock knob is ordered for 0.3.24; add the undervolt line, the AMD ADLX line, the Apple "no lever" line) | AMD has no `-lgc` (ADLX or nothing); Apple has no lever; a locked 5090's free shadow is about 150,000 ops, so ladder rung 2 costs it rate; the knee may sit where the driver refuses the lock |
|
||||
| 2 | **The SM-sparse miner kernel**: the hash on about N of the SMs at 32 warps per block, the rest clock-gated (miner-only; worker variants `sp<N>-w<W>`, section 12) | unmeasured: reads the 99 W of `E_wait` at the lock; if half is SM activity, class v3 at the lock about 2.8x (4.0x) and class v4 at k = 1 about 1.7x; if none, unchanged | lowers the base by what it finds; no premium | none | 1 (built tonight) + 1 (one PC 1 job) + 3 to 4 to ship as the default kernel through the tuning file and the Devnet 2 gate | the overhead is clock tree and leakage, not SMs (the knee sweep shows this first); fewer SMs cannot hold 17.5 G reads in flight (MSHR limits), so the rate falls with the watts; AMD and Apple need their own variants |
|
||||
| 3 | **The shadow at rung 0 with a high-k op mix** (shuffle and multiply heavy weight table; a class change under the 95 percent rule) | at the 1,400 lock and the same premium: 2.1x at k = 1 (unchanged), 2.9x at the pessimistic core instead of 3.4x (k floor 0.32 to 0.46, approximate) | unchanged: 88 W locked, 145 W unlocked / 30 W | unchanged (shfl is one op; about 8.8 ms loaded at rung 0) | 4 to 6 (the weight-table knob, packs, three-vendor fingerprints) plus the six gates across lanes | Apple pays shfl at 1.91x per op (under 1 percent of its ALU time at rung 0, argued); held by the coordinator until designs 1 and 2 are read |
|
||||
| 4 | **The L2-resident hot table with cache-policy hints** (a 32 to 64 MiB table beside the dataset; dataset loads evict-first, hot loads evict-last) | if the rate holds: the one lever with k possibly over 1 (a 64 MiB SRAM read on a chip 0.2 to 0.5 nJ against an L2 hit 0.1 to 0.3 nJ, both approximate): 512 hot reads per hash is about 0.1 microjoules on both sides, class v3 at the lock about 3.2x (4.4x) | about 0.1 microjoules (13 W) if the hints hold the table; the 5 October rows without hints cost 13 to 16 percent of RATE on the 5090 and 16 to 20 on the 9070 XT | the hot items derived lazily per unit: unmeasured, likely under 0.5 ms | 4 for the measurement (a hinted variant of the 5 October hot-table packs on PC 2); 10 or more if adopted (a class change) | Metal has no cache hints (Apple pays rate); AMD's slc bits unverified; the chip's SRAM is $12 of die at 64 MiB, so the lever is energy only; it is a measurement, not a result |
|
||||
| 5 | **The tensor block with a byte-dot verifier** (mm8 at R about 1,450, the shadow's premium in tiles) | k 0.5 to 1 (the same circuit): at the shadow's premium 2.1x at k = 1, 2.9x at k = 0.5 (modelled on the 4090's measured block energy) | the same 0.65 microjoules by construction / the 4070 untested | scalar FAILS (about 10.6 ms at R = 1,450 on the M5 Max core); VNNI, udot or pmaddubsw 4x to 16x cheaper (approximate): 1 to 3 ms, inside the gate on a 2019 core only with AVX2 | 8 to 12 (the two-output block exists in `proto-newpow/mma-shadow`; the SIMD verifier, the AMD WMMA layout gate, Apple's emulation measurement) | Pascal, RDNA 2 and every Apple card emulate at 10 to 16 steps per tile; the AMD fragment layout is unverified; buys nothing over design 3 at equal premium except a better-bounded k |
|
||||
| 6 | **Class v5 with the shadow dropped** | 5.1x (7.2x), modelled; 3.6x (5.3x) at the lock | none / none | -0.65 ms (the shadow gone), +0.2 ms (v5) | 0 | fails the brief: the memory-system chip is untouched; v5 is shipped for its own reasons (the recompute chip and the stateless pool miner gone) |
|
||||
| 7 | **The refresh as the cost** (per epoch or per block rebuild) | unchanged (the rebuild is 0.1 to 0.5 J on a chip core) | the 4090 loses 3 percent of hash time at one rebuild per second / the same | none | 0 | dead by arithmetic |
|
||||
| 8 | **Per-card adaptive classes** | at least the edge of the card's best single class (the proof in section 6) | | | 0 | dead by the maximum argument; the ladder is what survives |
|
||||
| 9 | **Proof of useful work** (proving in the lottery) | a proving chip's credible 5x to 30x (approximate) | | a proof piece cannot be checked under 10 ms | 0 | dead by the gas bound and the Aleo precedent |
|
||||
| 10 | **Memory-system shaping** (row, bank, refresh, burst; more reads; wider loads) | unchanged | none | none | 0 | no asymmetry: the same DRAM both sides |
|
||||
|
||||
The three strategic answers that are not designs, for the record: (i) accept chips (the Kaspa turn: the chip edge is then the point, and the honest number to publish is the chip's USD 0.00021 per MH/s-hour against an owned 5090's 0.00113, algorithm.md 5.1); (ii) hold the line with the shadow, the ladder and the share-pattern detector, paying the premium at the knee and saying so; (iii) retire the premium by card-side efficiency (designs 1 and 2) and keep the shadow as the floor-setter. This file recommends (iii) with (ii)'s honesty.
|
||||
|
||||
## 10. The recommendation, in five sentences
|
||||
|
||||
The premium cannot reach zero against a chip that is a GPU's memory system without the GPU, because every joule we force the chip to spend is a joule the card spends first, at a ratio set by silicon; what zero premium buys is the card's own whole-card energy over the chip's memory energy, which is 3.6x on a 5090 locked at 1,400 MHz, about 2x at that card's idle-plus-memory bound, and 1.7x on an Apple GPU today. The shadow is therefore kept at rung 0 as the floor-setter, with its measured price stated at the knee (88 W at 1,400 MHz, lower if the knee is lower) and its measured worth (2.1x at a chip core equal to a GPU lane, 2.9x at a core three times better, nothing below 1.8 pJ per op), and its content re-weighted toward shuffles and multiplies once designs 1 and 2 are read, because that is the only way to raise the chip's `k` at the same premium. The lever that moves every row with no consensus change is the honest card's operating point, so the knee lock and the undervolt ship as the default in Ember Tune for every NVIDIA card model, with the AMD and Apple rows stated honestly as "ADLX or nothing" and "no lever". The SM-sparse kernel built tonight is the one measurement that can lower the premium-free floor further on the same card, and it decides itself in one PC 1 job: if a quarter of the SMs holds the rate at much lower watts it ships as the miner's default kernel, if the watts track the rate it is dead. Class v5 ships for its own reasons and does not retire the shadow; the public text should say the floor and the premium as numbers, and should stop quoting the withdrawn X9 as a chip core a third as costly as a GPU lane, since that ratio was against a CPU.
|
||||
|
||||
## 11. Consequences per tier
|
||||
|
||||
| Tier | What this file means tonight | What is being done |
|
||||
|---|---|---|
|
||||
| Home miner, one 8 GB card (NVIDIA; 4060 Ti class, 3.81 microjoules rented and untuned) | No chip exists; against the modelled GDDR7 chip the untuned card sits at 8.2x per joule on class v3 and about 3.1x on class v4 at k = 1 (algorithm.md 5.1). The two levers that move its row are in the miner, not the hash: the Ember tune (10 to 30 percent per joule measured on the 4070) and, if it reads well, the SM-sparse kernel. Class v4's premium on a small card is about 30 W at the tune point (the 4070's row) | design 1 in 0.3.24; design 2's PC 1 job |
|
||||
| One 12 GB card (4070, 5070) | the 4070 at its tune point is 5.5x (v3) and about 2.2x (v4, k = 1) against the GDDR7 chip, measured card, modelled chip; its v4 premium is 30 W (79.5 to 109 W) for no rate | the same |
|
||||
| One 16 GB card (RX 9070 XT) | 23x per joule behind the GDDR7 chip on class v3 (approximate watts); the shadow costs it no rate (+3.6 percent at rung 3) and its watts under v4 are OWED; AMD has no clock-lock path in the app (ADLX line or nothing) | the AMD watts row; the ADLX line in Ember Tune's text |
|
||||
| One 24 or 32 GB card (5090 class) | the measured rows of this file: 5.2x unlocked, 3.6x at the 1,400 lock on class v3; class v4 2.1x at k = 1 for 88 W at the lock; a 5090 owner who locks the core pays 316 W instead of 476 W for 1.4 percent of rate | the knee sweep (running), the Ember knob, the SM-sparse job |
|
||||
| Apple (M-series) | the honest best per joule the project owns: 1.7x from the GDDR7 chip with no shadow, 0.9x at class v4 and k = 1; no clock lever; a shuffle-heavy mix costs it 1.91x per shfl op (under 1 percent of its ALU time at rung 0, argued) | nothing to do for Apple in designs 1 and 2; design 3 is checked against its 5 percent rule |
|
||||
| A rig | a rig's bill is watts: the lock takes a 5090 rig from 476 to 316 W per card on class v4 (-34 percent of electricity) for 1.4 percent of rate; the shadow's premium at the knee is the rig's price for the 2.1x floor | design 1 as the default |
|
||||
| A pool user | nothing changes in shares or payout from any design here; designs 3 to 5 would be class changes announced by the 95 percent signal | |
|
||||
| A node operator (the verifier) | designs 1 and 2 cost nothing; design 3 nothing; design 4 under 0.5 ms (unmeasured); design 5 needs a SIMD byte-dot path to fit the gate | |
|
||||
| A chip | its edge is `E_card / E_mem` plus whatever the shadow forces at its own `k`: 3.6x to 5.2x on GDDR7 with no shadow depending on the card's operating point, 2.1x with the shadow at k = 1; a core cheaper than 1.8 pJ per counted op would turn the shadow in its favour at the knee, which no public figure shows for a random 32-lane program | the `k` question stays with the external brief (algorithm.md proposal 8) |
|
||||
| The public claim | the honest sentence is the floor and the premium as numbers: "a chip that stores the dataset keeps the card's whole-card energy over its memory energy, 3.6x on a locked 5090 and 1.7x on an Apple GPU with no extra work; the shadow holds it to 2.1x against a core as good as a GPU lane for 88 W on a 5090 at the knee" | to the Counter lane for the chip texts once the knee is a number (the coordinator, 21:5x UK) |
|
||||
|
||||
## 12. What was built tonight: the SM-sparse worker variant
|
||||
|
||||
`proto-cuda/nvrtc/worker.cpp` (this branch) gains opt-in variants `sp<N>` and `sp<N>-w<W>` (N persistent blocks of W warps, W default 32): the pack's bound kernel text is rewritten at compile time with exact anchors (the race's existing mechanism, `variantSource`): `__global__ void igneum_hash_bound(...)` becomes `__device__ __forceinline__ void igneum_hash_bound_unit(..., uint32_t gid)` with its `gid` line removed, and a persistent wrapper of the kernel's own name loops the unit over the dispatch's nonces with stride `gridDim.x * blockDim.x`, taking a trailing `uint32_t nonces` argument. `launchHash` launches it as a grid of N blocks; `raceTime` carries the entry's shape while timing it; the winner installs it. The variants are made on demand by name (`parseSparseVariant`) and are not in the catalogue, so a default race never runs them and no shipped worker changes behaviour. Refused on a persistent (variant-5) pack and when combined with a `__launch_bounds__` minBlocks variant. The rewrite was mirrored in Python on the class v4 devnet pack and reads as intended; the exe was cross-compiled on igneum-build-1 under `lease pool 4 --class measure --label "ca4 research: worker cross-compile (SM-sparse variant)"` (20:48 UTC; `/srv/builds/ca4-research/nvrtc/igneum-worker-cuda-ca4sparse.exe`, sha256 ee8d0e70dd101f125f42c0c7cf07481a794ee18a1317acf68561b37ea18d72be, no icon or version block: a scratch exe for one job, not a shipping one). Not yet run on a card: the first run's self-test (the vector warps through the bound kernel and the 2^24 fingerprint equal to the base kernel's) is its gate, and an NVRTC compile error shows as the variant dropped with "compile:" in the race line. The job is with the hash lane for PC 1 after the knee sweep (rungs sp170, sp85, sp43, sp21, sp11 at the 1,400 lock and unlocked, class v3 and v4 packs; MH/s, watts, SM MHz, fingerprint per row).
|
||||
|
||||
## 13. Unverified and owed
|
||||
|
||||
- The knee: read at 20:5x UTC as 1,300 MHz on class v3 (223.3 W, 1.66 microjoules) and 1,200 MHz on class v4 (305.1 W, 2.28), the premium 81.8 W at the best points; every "at the lock" row in this file is the 1,400 row and moves by under 3 percent to the knee (3.63x to 3.56x on GDDR7; the class v4 edge at k = 1 stays 2.1x). The floor below 1,100 MHz is unmeasured (a third pass runs to it).
|
||||
- The breakdown of the 5090's 228 W at the lock into idle, memory and SM activity is unmeasured; the SM-sparse job reads it.
|
||||
- The RX 9070 XT's watts under class v4; the 2019-class verifier core (O-1.14); the M5 Max package watts.
|
||||
- The chip side is the model (chip-model-v3 section 5): activate-bound ceilings, 2.0 and 1.2 nJ per random read, the static and controller allowances, the 85 W node; no chip has been measured. The `k` band 0.5 to 1 for a SIMD-array core is approximate; the N5 datapath floor is a bound, not a design.
|
||||
- The L2-hit and 64 MiB SRAM energies of design 4 are scalings from Horowitz 2014 and chip-model-v3 5.1 (approximate); the cache-hinted hot table is unmeasured.
|
||||
- The tensor block's SIMD verifier cost is a 4x to 16x scaling, unmeasured.
|
||||
- Every external figure is the vendor's or the author's claim at its URL, read 7 October 2026 by the literature sub-agents of this lane; where a page refused the fetch the figure came through a search summary and is marked claimed.
|
||||
|
||||
## 14. Sources
|
||||
|
||||
Internal: `docs/analysis/chip-model-v3.md` (sections 5 and 6; section 5.10 on ca3-coord), `docs/analysis/latency-shadow-2026-10-06.md`, `docs/plans/counter-asic-3-status.md` (ca3-coord: the efficiency pass, the 4070 and 9070 XT rows), `docs/plans/counter-asic-3.md`, `docs/plans/counter-asic-3-derivation.md`, `docs/plans/counter-asic-2.md`, `docs/design/class-v5-stored-state.md` (class-v5), `docs/design/latency-ladder.md` (ladder and attack-ladder-5a), `docs/analysis/attack-pass/f1-shadow.md` and `f10-ladder.md` (attack-pass), `docs/plans/cryptanalysis/in-house-pass.md` (crypto-engage), `docs/analysis/horizon/new-pow.md` and `algorithm.md`, `docs/analysis/asic-resistance-history.md`, `docs/analysis/scratch-soundness.md`, `docs/bench-log.md`, `docs/evidence.md` row 15.
|
||||
|
||||
External (all read 7 October 2026): O'Connor et al., Fine-Grained DRAM, MICRO 2017, https://www.cs.utexas.edu/users/skeckler/pubs/MICRO_2017_Fine_Grained_DRAM.pdf ; Chatterjee et al., subchannels, HPCA 2017, https://www.cs.utexas.edu/users/skeckler/pubs/HPCA_2017_Subchannels.pdf (GUPS about 9 pJ per bit against STREAM about 3, claimed); Horowitz, ISSCC 2014, https://pages.cs.wisc.edu/~markhill/restricted/isscc2014_horowitz_power_scaling.pdf ; Dally, Hot Chips 2023 keynote, https://www.hc2023.hotchips.org/assets/program/conference/day2/Keynote%202/Keynote-NVIDIA_Hardware-for-Deep-Learning.pdf ; Micron GDDR7 4.5 pJ per bit via https://www.club386.com/micron-gddr7-improves-nvidia-gpus-by-up-to-30-in-gaming/ ; Samsung HBM3E 3.9 pJ per bit via https://www.igorslab.de/en/samsung-ends-afterburner-hbm4e-with-325-tb-s-bandwidth-to-exceed-nvidias-requirements/ ; AccelWattch, MICRO 2021, https://paragon.cs.northwestern.edu/papers/2021-MICRO-AccelWattch-Kandiah.pdf ; tensor cores on memory-bound kernels, https://arxiv.org/html/2502.16851v2 ; Shuhai, FCCM 2020, https://arxiv.org/pdf/2005.04324 ; Folded Banks, ISCA 2025, https://dl.acm.org/doi/10.1145/3695053.3731111 ; Li, Reddy, Jacob, MEMSYS 2018, https://terpconnect.umd.edu/~blj/papers/memsys2018-dramsim.pdf ; the Ethash, RandomX, ProgPoW, Equihash, Cuckatoo, Kaspa, Alephium, FishHash, Autolykos, NexaPoW, Xelis and PoSME URLs in sections 5 and 8; the proving URLs in section 7.
|
||||
|
||||
## 15. The k below 1 hunt (second pass, the founder's question of 22:4x UK: "a better way to shadow, or a new way, that is easier on GPUs and hard as hell on chips")
|
||||
|
||||
Written 21:3x to 21:5x UTC, 7 October 2026. A note on the letter: this file's `k` is the chip core's energy per forced operation over the GPU's, so an operation that is cheaper on the GPU than a chip can match is `k` ABOVE 1 here (the coordinator's brief wrote the same thing as "k below 1" with the ratio the other way up). The identity of section 2 does not change: zero premium is zero forcing at any `k`; what a high `k` buys is the right-hand asymptote `1 / k` and a steep slope at small premium, so a block with `k` near or above 1 makes a GIVEN premium count for more and removes the chip's downside bet. "Hard as hell on chips" is `k` well above 1; "easier on GPUs" is a block whose joules per unit of forcing are low, which is the same thing said twice.
|
||||
|
||||
**The one-sentence answer first: on the public figures nothing reads k above 1 with certainty; every block a GPU has, a chip at the same node builds for the same or less energy per operation, and the only block where k reads near or above 1 rather than 0.3 to 0.5 is the int8 tensor tile, because the GPU's tensor core is already a dense MAC array at the node's density, so the best design is the tensor-shaped shadow with a SIMD verifier (section 17, rank 3), and the microbenchmarks built tonight (section 18) replace the public-figure column with measured picojoules per operation on the 5090 when the PC 1 queue reaches them.**
|
||||
|
||||
### 15.1 Every GPU block gaming and AI already paid for, priced
|
||||
|
||||
The GPU column is the public figure or this project's measurement until the microbench rows land (each probe is named; the row then reads watts minus the sleep floor over counted operations per second, at the unlocked clock and the 1,300 MHz knee). The chip column is the best public figure at a 4 or 5 nm node for the same operation, labelled. `k` is the chip's cost over the GPU's at the same operating point. The verifier column is the CPU's cost to reproduce the operation bit-exactly inside the hash (the x8 verifier does 18 G simple integer ops per second per core and the shadow's law is 0.1 ns per lane-instruction, `latency-shadow-2026-10-06.md` section 4). The last column is the edge if the shadow were made of that operation at the class v4 premium's joules at the 1,400 lock (0.654 microjoules): at ZERO premium every row reads 3.6x (GDDR7) and 5.3x (HBM3), the identity; the row is what the same joules buy against a chip of that `k`.
|
||||
|
||||
| GPU block | The 5090's energy per op: public or measured today | Microbench probe | A chip's cost at 4 to 5 nm, public | k band | Bit-exact on CPU, three vendors; verifier cost | Edge at the measured premium (0.654 microjoules at the lock) | Labels |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| Int32 ALU (add, xor, rotate, LOP3, byte permute) | 6.5 pJ per counted op at the 1,400 lock, 10.4 unlocked (the class v4 shadow) | int_arx, lop3, prmt | datapath add 0.06 pJ at 5 nm (mlsysbook citing Horowitz 2014 and Dally 2021, https://mlsysbook.ai/vol1/backmatter/appendix_assumptions.html); the lane's register file and operand delivery dominate on both sides: Dally's 45 nm figures (Hot Chips 2023) are 30 pJ of fetch, decode and operand delivery around a 1.5 pJ HFMA; a SIMD array with a local register file at N5 about 2 to 5 pJ per op | 0.3 to 0.8 | yes; 0.1 ns per lane-op, the shadow's measured law | k 0.5: 2.9x; k 0.8: 2.3x; k 1: 2.1x | GPU measured; chip claimed and approximate |
|
||||
| Int32 multiply, mulhi | unmeasured marginal (the family step costs of item 6 give the multiply at about the add's issue cost on the 5090) | int_mul, int_mulhi | 32-bit multiply 0.52 pJ at 5 nm datapath (claimed, the same source); the multiplier is the same circuit on both sides, so the overhead ratio is the ALU row's | 0.4 to 0.8 | yes; one op | as the ALU row | claimed |
|
||||
| Warp shuffle crossbar | the shfl step 0.86x of the add-xor-rotate step on the M5 Max, measured (item 6); the 5090's step cost in the item 6 record; marginal pJ unmeasured | shfl | a 32-lane crossbar is wires: 0.6 pJ per bit-mm at N5 (approximate, chip-model-v3 5.1), a 32-bit word across a 1 mm array about 20 pJ on either die; the algorithm lane's k floor for a shuffle-heavy mix 0.46 (approximate) | 0.5 to 1 | yes; one op | k 0.5: 2.9x; k 1: 2.1x | measured steps; chip approximate |
|
||||
| Register file bandwidth | not a separable op: the 3 operand reads per lane instruction are most of the 6.5 to 10.4 pJ (Dally: 6 pJ of register file per in-order instruction at 45 nm, claimed) | prmt and lop3 read it (cheap datapath, RF-dominated) | the same banked SRAM at the same node; a chip's lever is operand reuse through a fixed datapath (the 3x factor), which a chained random program denies | about 1 for a random program | the CPU keeps the 32 lanes' registers in L1 (register-major loops) | as the ALU row at k 1: 2.1x | claimed |
|
||||
| FP32 FMA at the GPU's density | 104.8 TFLOPS FP32 on the 5090 (NVIDIA, claimed: 52 T FMA per second); marginal pJ unmeasured | fp32_fma | FP32 FMA at 5 nm about 0.4 to 0.6 pJ datapath (Horowitz 45 nm: 0.9 add, 3.7 multiply, scaled by the 2.5x to 3x process factor Dally gives per node step, approximate); the same overhead ratio as int | 0.3 to 0.6 | fmaf is IEEE-exact on every vendor with fast-math off and denormals NOT flushed (Apple and AMD flush by default in some modes: a vendor-mode risk the integer hash does not have); one op | k 0.5: 2.9x | claimed, approximate |
|
||||
| FP16x2 FMA | 2 per instruction; unmeasured | fp16x2_fma | half the FP32 datapath | as FP32 | fma.rn.f16x2 is exact per IEEE half with round-to-nearest, but Apple's and AMD's half paths differ in denormal handling: excluded from the hash content | | approximate |
|
||||
| **Tensor tile, int8 (u8 m8n8k16, s8 m16n8k32)** | 0.056 pJ per multiply-add marginal at 4,096 tiles per hash on the RTX 4090 (measured 6 October, new-pow 5.1; 0.70 pJ at 64 tiles, the fixed cost amortising); Dally: IMMA 160 pJ per 1,024-MAC instruction at 45 nm (claimed) = 0.16 pJ per MAC, about 0.02 at 5 nm by his process factor | mma_u8_m8n8k16, mma_s8_m16n8k32 | NVIDIA's 5 nm INT4 test chip 95.6 TOPS/W at 0.46 V (JSSC 2023, measured, via Dally's NASEM slides https://www.nationalacademies.org/cdn/materials/9fba0a50-8a84-4dda-8844-e5461844fb28): 0.021 pJ per INT4 MAC at 0.46 V, about 0.1 pJ at 1.05 V (the plot's 5x); an INT8 MAC 2x to 4x an INT4 MAC: 0.04 to 0.08 pJ at the low voltage, 0.2 to 0.4 pJ at nominal; a licensable int8 MMA block is the same circuit as the GPU's | **0.7 to 3 at the same voltage; near 1 is the honest centre** | integer, exact (bit-exact on 2^24 lanes and 1,024 CPU lanes at five rungs, measured); scalar 1.07 microseconds per tile per unit: at the premium's tile count (11,600 tiles per hash, R about 1,450) 12.4 ms, FAILS the gate; a byte-dot verifier (AVX-VNNI vpdpbusd 64 MACs per instruction, NEON udot 16, AVX2 pmaddubsw) 4x to 16x cheaper, 0.8 to 3 ms, approximate and unmeasured | at 0.654 microjoules of tiles: k 0.7: 2.5x; **k 1: 2.1x; k 1.5: 1.6x; k 2: 1.3x**; the chip's downside (k 0.3) is gone because its MAC cannot be 3x cheaper than the GPU's at the same node | GPU measured on a 4090; chip claimed (a test chip at a different voltage); k approximate |
|
||||
| Tensor tile, FP16 and BF16 (m16n8k16, FP32 accumulate) | the same hardware path; unmeasured | mma_f16_m16n8k16, mma_bf16_m16n8k16 | the same array with floating-point accumulate | about 1 | NOT bit-exact: the order of the 16 products' accumulation inside the tile is unspecified by PTX, so two vendors (or two NVIDIA generations) may differ in the last bit; excluded from the hash content; measured here only for the energy picture | | claimed |
|
||||
| Tensor tile, FP8 e4m3 (m16n8k32) | sm_89 and later; unmeasured | mma_e4m3_m16n8k32 | the same | about 1 | the same exclusion (floating accumulate), and no AMD RDNA or Apple path | | claimed |
|
||||
| L2 cache (96 MB on the 5090, 64 MB class on a 4090, 36 on a 4070) | an L2 hit on a GPU is about 0.1 to 0.3 nJ per 32-byte sector (approximate, Horowitz's 1 MB cache at 100 pJ per 64 bits at 45 nm scaled to N5 and a 96 MB array with its crossbar); measured tonight: the l2_chase rows (dependent, the latency) and l2_indep4 (the throughput) | l2_chase_32m, l2_chase_64m, l2_indep4_32m | a 64 MiB SRAM on a 128 mm^2-class die: 0.2 to 0.5 nJ per 64-byte read (chip-model-v3 5.1, approximate); a 3 nm 563 kbit macro reads at 3.89 pJ per access (JSSC 2026, measured, https://sscs.ieee.org/tag/one-cycle-latency-low-leakage-access-mode-1-clm/), so the array's wires, not the cell, set the cost on both sides | **0.7 to 3; the one block where the chip may be WORSE** | the hot items derived lazily on the CPU (unmeasured; the 5 October hot-table packs); the measured rate cost without cache hints 13 to 16 percent on the 5090 (the ldcs job in the hash lane's queue reads it with the hint) | 512 hot reads per hash at 0.2 nJ = 0.1 microjoules: too few joules to force by themselves; at k 1.5 that 0.1 buys 0.1x; the lever is capex-free (section 16) and energy-small | approximate throughout |
|
||||
| Texture sampler, point sampled | a fetch through the texture cache's address unit; unmeasured | tex_point_u32_32m | an address unit is a few adders; a chip reads SRAM or DRAM directly | about 1 or under (the chip skips the unit) | a point fetch is a load; exact | as the L2 row | approximate |
|
||||
| Texture sampler, linear interpolation | the interpolation weights are 9-bit fixed point with 8 fractional bits on NVIDIA (CUDA Programming Guide, "Linear Filtering", claimed); unmeasured | tex_linear_f32_256k | a chip's interpolator is one 9-bit multiply-add per fetch, cheaper than the GPU's full sampler | under 1 | NVIDIA's rule is reproducible on a CPU; AMD's and Apple's samplers use their own fraction widths and rounding (vendor-specific, approximate): NOT bit-exact across the three vendors, excluded from the hash content; measured for the energy picture only | | claimed, approximate |
|
||||
| The rasteriser | not reachable from CUDA, OpenCL or Metal compute; only through a graphics pipeline, whose rasterisation rules (fill convention, sample positions, depth precision) differ by vendor | none | | | NOT bit-exact across vendors by construction; excluded | | |
|
||||
| DRAM random read (the control) | the hash's own pattern: 17.5 G dependent 4-byte reads per second at about 3.1 nJ incremental per read (modelled, 55 W of memory system) | dram_chase_1g | 2.0 nJ per random 32-byte read on GDDR7, 1.2 on HBM3 (modelled from O'Connor et al. and Samsung's pJ per bit) | 0.4 to 0.65 (the chip's whole case) | exact; the verifier's item derivation | this IS E_mem: 3.6x at zero premium | modelled |
|
||||
|
||||
The reading: a GPU lane's cost per operation is overhead (fetch, decode, operand delivery, the banked register file) around a datapath that is a few percent of it, and a chip at the same node builds the same datapath and may drop part of the overhead, which is why every ALU-shaped row reads `k` 0.3 to 0.8. The two rows that read `k` near or above 1 are the ones where the GPU's block is itself a dense array with no per-op overhead to drop (the int8 tensor tile) or a large SRAM whose cost is wires that a chip's SRAM has too (the L2). The tensor tile is the only one of the two with joules to force at a verifiable cost, and only with a SIMD verifier.
|
||||
|
||||
### 15.1a The microbench rows (PC 1, the RTX 5090 alone, 05:14 to 05:56 UTC, 8 October 2026 (06:14 to 06:56 UK); the hash lane's job `run-ca4-pc1-microbench-5090-20261008-c`, the ca4mb exe, 20 probes at 60 s each at full residency, "20 probes ran, 0 skipped or failed" at both states; the sampler's mean from 8 s in, the three power fields within 0.2 W; idle 74.7 W unlocked and 60.2 W at the lock; every checksum equal at both states)
|
||||
|
||||
The reading per row is (watts minus the `sleep` row) over counted operations per second: picojoules per counted op on this card. The `sleep` row is NOT idle: 120.2 W unlocked and 66.2 at the lock against 74.7 and 60.2 idle, so full residency with every warp spinning on `__nanosleep` costs 45 W at the stock clock before any instruction issues; that residency cost is subtracted from every row below, which makes each figure the op's own marginal.
|
||||
|
||||
| Probe (what one counted op is) | Unlocked: W / SM MHz / G ops per s / pJ per op | At the 1,300 lock: W / G ops per s / pJ per op | Label |
|
||||
|---|---|---|---|
|
||||
| sleep (full residency, no issue) | 120.2 / 2,872 / 0 / the floor | 66.2 / 0 / the floor | measured |
|
||||
| int_arx (int32 add, xor, rotate; 4 chains) | 533.9 / 2,816 / 36,600 / **11.3** | 170.9 / 16,819 / **6.2** | measured; the class v4 shadow read 10.8 and 6.4 pJ per counted op on the packs job: the two instruments agree |
|
||||
| int_mul (int32 multiply-add; 4 chains) | 558.2 / 2,712 / 31,419 / 13.9 | 194.5 / 15,412 / 8.3 | measured |
|
||||
| int_mulhi (4 chains) | 538.5 / 2,805 / 10,573 / 39.6 | 168.6 / 4,871 / 21.0 | measured (a multi-instruction sequence per op on this card) |
|
||||
| prmt (byte permute; 4 chains) | 498.4 / 2,812 / 16,933 / 22.3 | 156.6 / 7,882 / 11.5 | measured |
|
||||
| lop3 (three-input logic; 4 chains) | 454.6 / 2,820 / 13,896 / 24.1 | 149.4 / 6,391 / 13.0 | measured (counted as one op per chain-step; the compiler's fusion is not verified) |
|
||||
| **shfl (warp shuffle, one xor beside it; 4 chains)** | 512.6 / 2,811 / 7,036 / **55.8** | 161.3 / 3,240 / **29.4** | measured: the most expensive operation on the card per op, 5x the add |
|
||||
| fp32_fma (4 chains) | 522.7 / 2,714 / 43,671 / 9.2 | 180.7 / 21,863 / 5.2 | measured |
|
||||
| fp16x2_fma (two half FMAs per instruction) | 365.1 / 2,829 / 48,067 / 5.1 per half FMA | 125.6 / 22,696 / 2.6 | measured |
|
||||
| **mma_u8_m8n8k16 (int8 MAC inside the tile; 32 per lane per tile, two tiles per step)** | 537.4 / 2,776 / 101,732 / **4.1 per MAC** | 169.7 / 47,555 / **2.2** | measured |
|
||||
| **mma_s8_m16n8k32 (int8 MAC; 128 per lane per tile)** | 575.0 / 2,629 / 334,227 (80 percent of the dense INT8 peak, approximate) / **1.36 per MAC** | 205.9 / 168,252 / **0.83** | measured: the cheapest counted op on the card, 8x under the ARX path |
|
||||
| mma_f16_m16n8k16 (FP16 MAC, FP32 accumulate) | 454.5 / 2,819 / 104,798 / 3.2 | 149.7 / 48,264 / 1.7 | measured (not bit-exact across vendors: energy only) |
|
||||
| mma_bf16_m16n8k16 | 425.3 / 2,820 / 104,953 / 2.9 | 137.8 / 48,320 / 1.5 | measured (energy only) |
|
||||
| mma_e4m3_m16n8k32 (FP8) | 439.3 / 2,820 / 211,028 / 1.5 | 143.9 / 97,101 / 0.8 | measured (energy only) |
|
||||
| l2_chase_32m (dependent 4-byte read, L2-resident; per read) | 378.7 / 2,842 / 107.55 G reads per s / **2.4 nJ per read** | 138.3 / 52.13 / 1.4 nJ | measured |
|
||||
| l2_chase_64m | 381.3 / 2,842 / 107.64 / 2.4 nJ | 140.5 / 52.14 / 1.4 nJ | measured (64 MiB still inside the 5090's L2) |
|
||||
| l2_indep4_32m (4 independent chains) | 377.0 / 2,842 / 109.42 / 2.3 nJ | 138.6 / 52.96 / 1.4 nJ | measured |
|
||||
| dram_chase_1g (the hash's own pattern; per read) | 318.4 / 2,850 / 18.17 G reads per s / **10.9 nJ per read** | 200.0 / 15.38 / 8.7 nJ | measured: the whole card's marginal per dependent DRAM read, against the memory system's modelled 2.0 nJ |
|
||||
| tex_point_u32_32m (per fetch) | 373.0 / 2,842 / 107.65 / 2.3 nJ | 138.6 / 52.20 / 1.4 nJ | measured (the same checksum as the L2 chase: the sampler's address path is the L2 chase) |
|
||||
| tex_linear_f32_256k (per filtered fetch, 4 chains) | 298.6 / 2,850 / 923.8 / 0.19 nJ | 101.0 / 413.3 / 0.10 nJ | measured (energy only; not bit-exact across vendors) |
|
||||
|
||||
**A correction these rows force on sections 15.2, 20.3 and 20.4, and on the 6 October tensor figure.** A `mma.m8n8k16` tile is one warp instruction over 32 lanes: 1,024 multiply-adds per WARP, 32 per lane, so a hash (one lane) does 32 MACs per tile, not 1,024. Section 15.2 and 20.3 multiplied the per-hash tile count by 1,024, and the 6 October figure for the 4090 (new-pow 5.1: "0.056 pJ per multiply-add at 4,096 tiles per hash") carries the same factor of 32: at 63.08 MH/s and 4,096 tiles per hash the 4090 did 8.3 x 10^12 MACs per second, not 2.6 x 10^14, and its 14.7 W is 1.8 pJ per MAC, not 0.056. Re-read with the right count, the packs job's tile rows (20.3) are: `mm1430` 11,440 tiles x 32 = 366,080 MACs per hash, 5.03 x 10^13 MACs per second at 137.45 MH/s, 146.7 W = **2.9 pJ per MAC unlocked** and 71.5 W over 4.64 x 10^13 = **1.5 pJ per MAC at the lock**; the microbench's dependent u8 tile reads 4.1 and 2.2, its s8 m16n8k32 at 80 percent of peak 1.36 and 0.83. The int8 MAC on the 5090 therefore costs 1 to 4 pJ, between a tenth and a third of a 32-bit ALU op (6 to 11 pJ), not a hundredth; and NVIDIA's own 5 nm MAC array (0.04 to 0.1 pJ per INT8-class MAC at 0.46 V, about 0.2 to 0.4 at nominal, claimed) is then 4x to 30x cheaper than the 5090's measured tile, not equal to it. **The chip's `k` on tile work reads about 0.03 to 0.3, below the ALU shadow's 0.3 to 0.8, so the tensor shadow is the WORSE forcing lever, and the k-above-1 candidate of section 15 does not exist on measured rows.** The corrected 20.4 is below; the 6 October 4090 figure is corrected in this file and owed to `docs/analysis/horizon/new-pow.md` 5.1 and `chip-model-v3.md` 5.11 (the Counter lane).
|
||||
|
||||
What the other rows say, in the identity's terms (the chip's cost per op from the public figures of 15.1, the GPU's now measured):
|
||||
|
||||
| Block | GPU, measured (unlocked / lock) | Chip at 4 to 5 nm, public | k band, corrected | Reading |
|
||||
|---|---|---|---|---|
|
||||
| int32 ALU (the class v4 shadow's mix) | 11.3 / 6.2 pJ | 2 to 5 pJ for a SIMD array with its register file (approximate) | 0.3 to 0.8 | the best forcing lever the card has; unchanged |
|
||||
| warp shuffle | 55.8 / 29.4 pJ | a 32-lane crossbar of a 32-bit word across about 1 mm: about 20 pJ (approximate, the wire figure of 15.1) | 0.4 to 0.7, with the GPU paying 5x the add per op | a shuffle-heavy re-weight raises the GPU's premium per instruction 5x for a k no better than the add's: the re-weight's case is now AGAINST it on the measured GPU side (section 17 rank 5 and the held item) |
|
||||
| int32 multiply | 13.9 / 8.3 pJ | the same multiplier, 0.5 pJ datapath (claimed) | 0.3 to 0.8 | as the ALU row |
|
||||
| int8 tile | 2.9 to 4.1 / 1.5 to 2.2 pJ per MAC (u8); 1.36 / 0.83 (s8 m16n8k32) | 0.04 to 0.4 pJ per MAC (claimed) | **0.03 to 0.3** | the worse lever: the chip's array undercuts the GPU's tensor core by 4x to 30x per MAC |
|
||||
| L2 hit | 2.4 / 1.4 nJ per read | an on-die SRAM read 0.2 to 0.5 nJ (approximate) | **0.1 to 0.3** | the hot-table lever (design 4, the ldcs measurement) is dead on the GPU side: an L2 hit costs the card 5x to 10x what a chip's SRAM costs |
|
||||
| texture interpolation | 0.19 / 0.10 nJ per fetch | a 9-bit interpolator: picojoules | far under 1 | excluded anyway (not bit-exact across vendors) |
|
||||
| DRAM dependent read | 10.9 / 8.7 nJ per read (the whole card's marginal) | 2.0 (GDDR7) to 1.2 (HBM3) nJ per read (modelled) | 0.1 to 0.2 on the marginal; this is the `E_card / E_mem` of section 2 seen per read | the premium-free floor, measured per read: at the lock the card pays 8.7 nJ for a read the chip's memory pays 2.0 for |
|
||||
|
||||
**The one-sentence answer of section 15, corrected on measured rows: nothing on the 5090 reads k above 1; the int8 tile, the one block the public figures put near parity, is 1 to 4 pJ per MAC measured against a 5 nm array's claimed 0.04 to 0.4, so it is the worst lever of all, and the ALU shadow (k 0.3 to 0.8 on 6 to 11 pJ per op) stays the best forcing work the card has.**
|
||||
|
||||
### 15.1b The op-mix re-weight's hold, restated on measured rows (06:0x UTC, 8 October 2026)
|
||||
|
||||
The re-opening condition set at 05:2x UTC (the 5090's shuffle and multiply at or under the add's picojoules per op) is not met and is now a measurement: shfl 55.8 pJ against the ARX op's 11.3 (4.9x), mul 13.9 (1.2x), mulhi 39.6 (3.5x). A shuffle-heavy weight table (the algorithm lane's shfl 14 and shfla 8 of 75 against class v4's shfl 8) therefore costs the GPU MORE per instruction at the same count, so the served 3.4x stands on a measured basis and the re-weight stays held. What the chip's `k` on shuffle-heavy work would have to be for 2.9x to be the honest pessimistic column, on the identity at the 1,300 knee (card 2.33 microjoules, `E_mem` 0.466): if the premium stayed at the ALU shadow's 0.652 microjoules (the same watts, fewer instructions), 2.9x needs `E_mem + k F = 0.80`, `k = 0.52`; if the instruction count stayed and the premium rose with the mix (about 29 percent of the counted ops at 4.9x the cost: the premium about 2.1x, 1.4 microjoules, the card 3.1), 2.9x needs `k = 0.42`. So the re-weight's pessimistic column is honest only if a chip pays 0.4 to 0.5 of the GPU's 55.8 pJ per shuffle, 22 to 28 pJ for a 32-lane crossbar move of a 32-bit word, which is above the wire figure of 15.1 (about 20 pJ across 1 mm, approximate) and unmeasured; and the GPU side of that bargain is a premium up to 2x higher per instruction. On measured rows the re-weight is not a candidate; it would re-price only against a measured chip crossbar.
|
||||
|
||||
### 15.2 The tensor shadow, re-read on the identity
|
||||
|
||||
The 6 October verdict on scheme B ("never as class v5 content") was right for the question it answered: at the 4090's free band (R = 512, 0.23 microjoules) the block forced too few joules and the verifier paid 26x the ALU shadow's cost per joule. The founder's question is a different one: not "is there a cheaper lever" but "is there a lever whose joules the chip cannot undercut". On that question the tensor tile is the best block on the card, because the GPU's marginal 0.056 pJ per MAC is within a factor of about 2 of what a 5 nm MAC array costs anyone (the test chip's 0.04 to 0.1 pJ, claimed), while the ALU shadow's 6.5 to 10.4 pJ per op is 10x to 50x what a fixed SIMD datapath costs. What it would take to make the tensor block carry the SAME premium as the ALU shadow (0.654 microjoules at the lock), on the 5090:
|
||||
|
||||
| Quantity | Value | Label |
|
||||
|---|---|---|
|
||||
| Tiles per hash for 0.654 microjoules at 0.056 pJ per MAC | 11.7 M MACs per hash, 11,400 u8 m8n8k16 tiles (R about 1,430 per iteration) | arithmetic on the 4090's measured marginal; the 5090's marginal is tonight's mma_u8 row |
|
||||
| Where the 5090's free band ends | the 4090 held its rate to R = 512 at 40 percent of its dense int8 peak (approximate); the 5090's tensor peak is about 1.7x the 4090's (NVIDIA's dense INT8 figures, claimed), so the free band on a 5090 reaches about R 2,000 and 11,400 tiles per hash sits inside it; on a 4070 (about 0.45x the 4090's tensor peak) the same R is tensor-bound and costs rate | approximate |
|
||||
| The verifier at R 1,430 | scalar 12.4 ms per unit on the M5 Max core (FAILS the 10 ms gate); with AVX-VNNI (64 byte-MACs per instruction) 0.8 ms, NEON udot (16 per instruction) 3 ms, AVX2 pmaddubsw (16 per instruction, every x86 since 2013) 3 ms, plus the x8 base 2.06: 2.9 to 5.1 ms on the M5 Max core, 7 to 13 ms on a 2019-class core by the 2.5x rule; the 2019 core passes only with the AVX2 path and a small R | approximate, unmeasured; the item that decides it |
|
||||
| The chip's edge at that premium (1,400 lock, GDDR7) | k 0.7: 2.5x; k 1: 2.1x; k 1.5: 1.6x; k 2: 1.3x; the same numbers as the ALU shadow at the same k, with the chip's k 0.3 column (4.1x) removed from the table because an int8 MAC array at 5 nm cannot be 3x cheaper than the GPU's | modelled |
|
||||
| What it costs the honest cards | the same watts as the ALU shadow by construction (the premium is the design variable); Turing and later NVIDIA and RDNA 3 and later AMD native; Pascal, RDNA 2 and Apple emulate at 10 to 16 ALU steps per tile, which at 11,400 tiles per hash is 110,000 to 180,000 counted ops, over the M5 Max's 130,000 bind point: the Apple tier would lose 5 to 10 percent of rate (approximate) | measured bind point; emulation cost approximate |
|
||||
| What breaks it | the AMD WMMA fragment layout is unverified (status item 6); Apple has no integer matrix path reachable from the toolchain; the SIMD verifier is unwritten; the 2019-class core may not fit; the chip's k near 1 is the centre of a claimed band, not a measurement | |
|
||||
|
||||
So the tensor shadow does not lower the GPU's premium (the founder's "easier on GPUs" is not available: a premium is a premium), it removes the chip's downside bet at the same premium, which is worth about 1x of edge at the pessimistic end (4.1x to about 2.1x at the ALU shadow's k 0.3 against the tile's k near 1). It is ranked below the two miner-only levers because it is a class change with an unwritten verifier and a vendor split, and above the ALU op-mix re-weight because its k floor is better bounded.
|
||||
|
||||
## 16. The inverse lever: capex as the wall
|
||||
|
||||
The brief: make the chip's capital cost, not its energy, the wall (the state-derived dataset plus an L2-resident hot table plus the era changes), with the break-even row recomputed. The model is the mission lane's (`docs/analysis/mission/future.md` section 2.2, 7 October 2026, model): the maker takes share `s` of the hash; two-year revenue `s x E2 x p` (E2 the two-year emission, 4.18 B IGN in years 1 to 2; p the price); the project pays when `p` is over `p* = C_proj / (s x E2)`, and the market cap at which it pays is `p* x supply` (4.18 B at the end of year 2), so in years 1 to 2 the break-even cap is `C_proj / s`, 3.3x the project cost at s = 0.30. The fleet's own silicon was 2 percent of revenue there and dropped out; the capex column below says whether anything in the hash can make it not drop out.
|
||||
|
||||
### 16.1 The capex column
|
||||
|
||||
| Chip or card | Silicon and memory per unit | MH/s | USD per MH/s | Capex per MH/s-hour over two years | Electricity per MH/s-hour at USD 0.05 per kWh | Capex over electricity | Label |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| f = 1 GDDR7 chip (chip-model-v3 5.4) | USD 470 (16 devices USD 320, controller USD 50, board USD 100) | 166 | 2.8 | 0.00016 | 0.000023 (0.466 microjoules) | 7x | modelled |
|
||||
| the same plus a 64 MiB SRAM hot table | +32 mm^2 of N5 at 0.49 mm^2 and USD 0.23 per MB (chip-model-v3 section 2): +USD 15 of die; if it forces the controller onto an N5 die or a second die, +USD 50 to 100 of die and package (approximate) | 166 | 2.9 to 3.5 | 0.00017 to 0.00020 | 0.000028 (the table's reads) | 6x to 7x | modelled, approximate |
|
||||
| the same plus the class v4 shadow core (30 mm^2 of N5 at N = 100,000) | +USD 25 to 40 | 166 | 3.0 to 3.1 | 0.00017 | 0.000080 (1.59 microjoules at k = 1) | 2x | modelled |
|
||||
| the same with the shadow per load (section 16.2: the ALU core must sit on the controller's die or across an interposer) | +USD 200 of interposer and package (a one-stack CoWoS class package, chip-model-v3 5.1, approximate) | 166 | 4.3 | 0.00024 | 0.000080 | 3x | approximate |
|
||||
| RTX 5090 at MSRP | USD 1,999 | 136 | 14.7 | 0.00084 | 0.000084 (1.69 microjoules at the lock); 0.00012 unlocked | 10x (7x) | measured card price; the street price in 2026 is about 2x MSRP, which doubles the capex row |
|
||||
| RTX 4070 at its tune point | about USD 550 (approximate street) | 31 | 17.7 | 0.00101 | 0.000128 | 8x | approximate |
|
||||
|
||||
What the column says: BOTH sides are capex-dominated (7x to 10x their electricity per MH/s-hour), so the chip's economic edge is its capex per MH/s (5x over a 5090 at MSRP), not its energy, and a design that moved the chip's capex would move the number a miner actually computes. But nothing in the hash can move it far: the work per hash is small (512 instructions plus the shadow), so the silicon a chip needs beside its memory is 30 to 100 mm^2 of N5 (USD 25 to 60), and the memory is the same 16 devices the card carries (USD 320). A hot table forces USD 15 to 100 more; the shadow core USD 25 to 40; an interposer USD 200. The chip's capex per MH/s moves from 2.8 to at most about 4.3, against the card's 14.7 to 29. The per-unit capex wall is unreachable by a factor of 3 to 7, for the same reason the energy wall is: the hash's dataset is a card's worth of DRAM and its work is a fraction of a card's logic. The capex wall that exists is the PROJECT cost and the node it forces, and that is already in the break-even model.
|
||||
|
||||
### 16.2 The break-even row, recomputed with the capex column
|
||||
|
||||
The project cost is what moves the cap; the per-unit capex enters as the fleet term, which was 2 percent of revenue and grows to about 3 percent with the interposer: it still drops out. The rows (s = 0.30, the mission lane's method; the cap is `C_proj / 0.30` in years 1 to 2 and 2.3x more every two years with the emission glide):
|
||||
|
||||
| Chip project | What the hash forces | C_proj | Break-even cap, years 1 to 2 | Years 3 to 4 | Years 5 to 6 | Label |
|
||||
|---|---|---|---|---|---|---|
|
||||
| 28 nm controller, no shadow (class v3, or class v5 with the shadow dropped) | a GDDR7 controller and PHY; the state-derived dataset adds a leaf buffer (6 KB today) and a link to a node: firmware; the era draws: firmware | USD 5 M | USD 17 M | 48 M | 114 M | model (future.md) |
|
||||
| 28 nm controller plus an N5 shadow core (class v4 as shipped) | an N5 die for the 14,000-lane ALU array; the two dies exchange 64 B of lane state per iteration: 17.5 G reads/s over 16 reads per iteration x 128 B = 140 GB/s, a PCB or organic-substrate link | USD 30 M | **USD 100 M** | 287 M | 682 M | model |
|
||||
| the same plus a 64 MiB L2-class hot table | the table joins the N5 die (32 mm^2 more): no node change, USD 15 of die; the project unchanged | USD 30 M to 33 M | USD 100 M to 110 M | 290 M | 690 M | approximate |
|
||||
| **the shadow per load** (the 256-instruction block split into 16 blocks of 16, one after every load, the same N): the chip's ALU core must sit inside every read's dependency, so the lane state crosses between controller and core twice per read: 17.5 G x 128 B = 2.2 TB/s, an interposer-class link, or one die carrying controller, lanes and PHY at N5 | a single N5 die or a 2.5D package: the mission lane's "N3 single die" row | USD 60 M | **USD 200 M** | 574 M | 1,363 M | approximate (the single-die cost is the mission lane's N3 row; a GDDR7 PHY on N5 is a real but unpriced item) |
|
||||
| the same plus the tensor shadow (section 15.2) | a licensable int8 MMA block beside the lanes: USD 4 of N5 silicon (algorithm.md 5.2); the project unchanged | USD 60 M | USD 200 M | 574 M | 1,363 M | approximate |
|
||||
| What the per-unit capex adds in every row | the fleet term, 2 to 3 percent of revenue | | nothing visible | | | model |
|
||||
|
||||
So the inverse lever's best design is the shadow per load: it costs the honest cards nothing by construction (the same N, the same instruction mix, a smaller block at 16 sites; the measured block-size effect on the 5090 and the M5 Max was that a 64-instruction block ran 2.5 to 3.5 percent FASTER than 256, latency-shadow sections 3 and 5, so 16 may be faster still or may cost compile-ahead, unmeasured), and it doubles the chip's project cost by forcing the controller and the core onto one advanced die or an interposer, which doubles the break-even cap from about USD 100 M to about USD 200 M in the first two years. It does not reach a capex wall per unit, and it does not change the energy identity (`k` is `k` whichever die the core sits on). It is a class change (the block positions are consensus) and a generator change; 4 to 6 agent hours plus the six gates; its measurement is one pack per card (rate, watts, compile-ahead at 16 sites).
|
||||
|
||||
## 17. The ranking, second pass
|
||||
|
||||
| Rank | Design | Chip edge per joule (GDDR7; 1,400 lock) | Premium: 5090 / 4070 | Verifier | Agent hours | Breaks on | Change |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| 1 | The operating point as the shipped default | class v3 3.6x; class v4 2.1x at k = 1 | lowers the base (330 to 223 W); the v4 premium 145 to 82 W | none | 2 to 4 | AMD and Apple have no lock | unchanged |
|
||||
| 2 | The SM-sparse miner kernel | **DEAD on measured rows (20.3b, 05:12 UTC): a quarter of the SMs holds 98 percent of the rate at the same draw; the draw follows the work, not the SM count** | none | none | 1 + 1 + 4 | measured: the energy per hash never falls below base | closed |
|
||||
| 3 | **The tensor shadow with a SIMD byte-dot verifier** (int8 tiles carrying the premium; a class change) | **DEAD on measured rows (15.1a, 20.4): the 5090's int8 MAC is 1.5 to 4 pJ measured, a 5 nm array 0.04 to 0.4 claimed, k 0.03 to 0.3, under the ALU shadow's band; at the measured premium the chip keeps 3.5x to 6.7x** | the same as class v4 by construction / the 4070 tensor-bound above about R 600 (approximate) | scalar FAILS; AVX2 and VNNI 2.9 to 5.1 ms on the M5 Max core, 7 to 13 ms on a 2019 core (approximate, unwritten) | 8 to 12 (the SIMD verifier, the AMD layout gate, Apple's emulation row) plus the six gates | the 2019-class core; the AMD layout; Apple loses 5 to 10 percent of rate on emulation; k near 1 is claimed, not measured | NEW this pass: the k-above-1 candidate |
|
||||
| 4 | The shadow per load (capex) | unchanged in energy | none by construction (the sound form's 5090 rows in the hash lane's v6 job) | the sound form 8.28 to 8.82 ms loaded on the reference core (measured, the invention lane) | 4 to 6 plus the gates | **REOPENED (8 October, 11:0x UTC): the 16 x 27 iterated form is dead on the no-era census (0.989 rejection); a drawn-era death before the 10:5x fix was an instrument artefact (the window's fixed bits); the sound form `mx8+shl4096x1` accepts 234 of 256 seeds no-era and its drawn-era verdict is the invention lane's 15:00 BST read** | doubles the break-even cap to about USD 200 M on the sound form (class v6 layer 5) |
|
||||
| 5 | The ALU shadow re-weighted toward shuffles and multiplies | 2.1x at k = 1; 2.9x at the pessimistic k 0.46 (modelled) | **raised: the 5090 pays 55.8 pJ per shuffle against 11.3 per add (15.1a), so a shuffle-heavy mix at the same instruction count costs the card up to 5x the premium per instruction** | unchanged | 4 to 6 plus the gates | Apple pays shfl 1.91x; the GPU side measured against it | HELD; the measured GPU side is against it |
|
||||
| 6 | The L2-resident hot table with cache hints | **DEAD on the GPU side (15.1a): an L2 hit costs the 5090 2.4 nJ unlocked and 1.4 at the lock against a chip's SRAM read at 0.2 to 0.5 nJ, k 0.1 to 0.3**; capex +USD 15 | 13 W if the hint holds the rate | under 0.5 ms (unmeasured) | 4 for the ldcs measurement (queued) | Metal has no hint; the rate without it | unchanged; a measurement |
|
||||
| 7 to 11 | class v5 without the shadow; the refresh as a cost; per-card classes; proof of useful work; memory shaping | as section 9 | | | 0 | dead by arithmetic | unchanged |
|
||||
|
||||
## 18. What was built in the second pass: the microbench
|
||||
|
||||
`proto-cuda/nvrtc/worker.cpp` gains `--microbench [--mb-seconds 60] [--mb-only a,b]`: no pack; 20 probes, one kernel each, compiled by NVRTC one at a time (a form the card or the compiler refuses drops that probe alone, its error on its RESULT row), each run at full residency (the occupancy query's blocks per SM x 256 threads x SMs) for the window, with a UTC start and end stamp per probe for a 1 Hz power sampler and a checksum of the lanes' outputs. The probes: `sleep` (the SM-resident floor: full occupancy, `__nanosleep`, no issue; every other row is read against it), `int_arx`, `int_mul`, `int_mulhi`, `prmt`, `lop3`, `shfl`, `fp32_fma`, `fp16x2_fma`, `mma_u8_m8n8k16`, `mma_s8_m16n8k32`, `mma_f16_m16n8k16`, `mma_bf16_m16n8k16`, `mma_e4m3_m16n8k32`, `l2_chase_32m`, `l2_chase_64m`, `l2_indep4_32m`, `dram_chase_1g` (the hash's own pattern, the control), `tex_point_u32_32m` and `tex_linear_f32_256k` (the sampler's address path and its interpolator, through `cuTexObjectCreate`, loaded optionally). The reading per row: (watts minus the sleep row's watts) over counted operations per second = picojoules per operation on this card at this clock. Gate: mingw cross-compile on igneum-build-1 under `lease pool 4 --class measure --label "ca4 research: worker cross-compile (microbench probes)"`, exit 0 with `-Wall -Wextra` clean, 21:20 UTC; exe `/srv/builds/ca4-research/nvrtc/igneum-worker-cuda-ca4mb.exe`, sha256 49aa60c4b28ed77a54655476d0f07c06a54aa22e4f6f6ccadd4019b84cbf0268; unrun on a card. The job is with the hash lane for PC 1 after the hot-table ldcs job: `--microbench --mb-seconds 60`, the 5090 alone, nvidia-smi at 1 Hz, at the unlocked clock and the 1,300 MHz knee. When its rows land, the GPU column of 15.1 becomes measured and every `k` band narrows from the GPU side; the chip side stays claimed.
|
||||
|
||||
## 19. Unverified and owed, second pass
|
||||
|
||||
- Every GPU figure in 15.1 except the ALU shadow's, the tensor tile's (a 4090) and the DRAM read's is unmeasured until the microbench rows land; the chip side is public claims at other voltages and nodes; `k` is a band, not a number, in every row.
|
||||
- The SIMD byte-dot verifier is unwritten; its 4x to 16x is a scaling from instruction widths.
|
||||
- The 5090's tensor free band (R about 2,000) is a scaling from the 4090's measured 40 percent at R = 512 and NVIDIA's claimed peaks.
|
||||
- The shadow-per-load design's block-size effect at 16 instructions and its compile-ahead are unmeasured; the single-die project cost is the mission lane's N3 row, and a GDDR7 PHY on an N5 die is unpriced.
|
||||
- The break-even model is the mission lane's (s = 0.30, a two-year life, the rental-equilibrium hash); its emission figures are that model's.
|
||||
|
||||
## 20. Third pass: the two prototypes, built (7 October 2026, 21:3x to 22:xx UTC)
|
||||
|
||||
The coordinator's order of 22:5x UK: both new ranks as experimental program classes in the kit, behind the pack, no consensus change. Built and gated tonight; the card rows are the hash lane's PC 1 job `run-ca4-pc1-packs-5090-20261007` (queued after the SM-sparse job and the microbench; the knee-lock states wait on PC 1's Power Helper, dead since 21:08 UTC, so the unlocked states land first).
|
||||
|
||||
### 20.1 What was built
|
||||
|
||||
| Item | Where | What it does | Gate |
|
||||
|---|---|---|---|
|
||||
| The per-load shadow, `+shl<S>x<R>` | `igneum-pow` (`ShadowClass.per_load`; `generator.rs`, `verify.rs`, `emit.rs`) | the class v4 shadow's 256 instructions and 27 passes, split into 16 sub-blocks of 16 consecutive instructions, sub-block `j` run 27 times right after the `j`-th load of the base program (the same work per hash, placed inside every read's dependency); id bytes `perload`; emitted in Metal, CUDA and OpenCL; the verifier runs it in the base loop | the base program and the shadow instructions are the class v4 program's draw for draw (unit test); the id differs; the suite green |
|
||||
| The int8 tile block, `+mm<R>` | `igneum-pow` (`ShadowClass.tiles`, module `mm8`) | `R` `mma.m8n8k16 u8` tiles per iteration after the shadow block, each over two registers with both outputs consumed (the `proto-newpow/mma-shadow` form), descriptors drawn from the stream after the shadow draws; CUDA runs the PTX tile natively, Metal and OpenCL the shuffle-and-byte-product reference; the Rust verifier runs the scalar reference or an AVX2 path (`maddubs` on a 4-way split of B so no 16-bit lane saturates; `IGNEUM_MM8_SCALAR=1` forces scalar); id bytes `mm8/` plus the count | SIMD equal to scalar on 64 seeds and the all-ones edge (unit test, run on the EPYC); the suite green: 64 + 7 + 4 + 19 + 2 + 7 tests, 0 failed, igneum-build-2 21:42 UTC |
|
||||
| Nothing moves for the chain | the packs test and a byte diff | the pinned v2, v3 and v4 packs are byte-identical; the re-exported `mx8+sh256x27` equals `packs-ca3-shadow/sh256x27` on all seven files | passed |
|
||||
| The packs | `proto-cuda/packs-ca4/` (also build-1 `/srv/builds/ca4-research/packs/`, the hash lane's kit `igneum-ca4-packs-20261007.zip`) | `mx8_sh256x27` (the control, class v4's shape), `mx8_shl256x27`, `mx8_mm128`, `mx8_mm512`, `mx8_mm1430` (11,440 tiles per hash: the ALU shadow's premium in tiles at the 4090's 0.056 pJ per MAC), all over seed `igneum-genesis`, day 2026-10-03, 1 GiB, generator 2 | every export `OVERALL: PASS` (vectors through the interpreter); the CUDA-against-CPU bit-exactness is the worker's self-test on the card (96 vector lanes through the bound kernel), unrun |
|
||||
| What is NOT built | | the NEON path of the tile verifier (the M5 Max runs scalar); the AVX-VNNI path (the AVX2 path stands in; VNNI would halve its instruction count, approximate); the acceptance rule's view of the per-load placement (the rule interprets the base program only, as for class v4; the sub-version 3 freshness fixpoint runs over base then shadow, which under per-load placement is not the execution order: a class, not a prototype, would re-derive it) | |
|
||||
|
||||
### 20.2 The verifier rows (igneum-build-2, EPYC 9454P, core 40 at nice 19 under `lease cores 40,88`, 21:5x to 22:07 UTC; the ladder's method: one cold warp per vector base, alone and with the SMT sibling core 88 running the same bench; the box at load 84 to 95 from other lanes, core 40 at 3.27 GHz)
|
||||
|
||||
| Class | Cold ms alone (warp base 0 / 4096 / 1000000) | Cold ms, sibling loaded | Items derived per warp (of 4,096 reads) | Reading |
|
||||
|---|---|---|---|---|
|
||||
| `mx8+sh256x27` (class v4's shape, the control) | 5.48 / 5.06 / 5.20 | 9.52 / 9.16 / 9.12 | 4,096 / 4,096 / 4,096 | the ladder's rung 0 read 5.14 and 8.77 on 6 October at load 25; tonight's box is heavier |
|
||||
| `mx8+shl256x27` (per load) | 5.09 / 4.84 / 4.86 | 8.66 / 8.34 / 8.38 | **3,627 / 3,595 / 3,633** | the same instructions, 7 to 9 percent FASTER: because 11 to 12 percent of the warp's reads hit an item another lane already derived (the finding below) |
|
||||
| `mx8+mm128` (128 tiles per iteration, no ALU shadow) | 5.33 / 5.13 / 5.08 | 8.59 / 8.37 / 8.25 | 4,096 | AVX2 tile path |
|
||||
| `mx8+mm512` | 5.31 / 5.12 / 5.12 | 9.30 / 9.15 / 9.03 | 4,096 | passes the gate loaded |
|
||||
| `mx8+mm1430` (11,440 tiles per hash, the ALU shadow's premium in tiles) | 5.82 / 5.64 / 5.66 | **10.14** / 9.58 / 9.55 | 4,096 | alone 5.8 ms; loaded the base-0 warp is over the gate by 0.14 ms (rung 3's class of miss, at a heavier load than the ladder's run) |
|
||||
| `mx8+mm1430`, the scalar tile path forced (`IGNEUM_MM8_SCALAR=1`) | 7.60 / 7.39 / 7.35 | not run | 4,096 | the AVX2 path is 4.6x the scalar path on the tile work |
|
||||
|
||||
The tile cost law from these rows: AVX2 (5.82 - 5.33) ms over (1,430 - 128) x 8 = 10,416 extra tiles per unit = 0.047 microseconds per tile per unit; scalar (7.60 - 5.33) / 10,416 = 0.22 microseconds per tile per unit (the 6 October figure of 1.07 was a naive C loop on a loaded core). So a tile class carrying the ALU shadow's premium costs the verifier about 0.5 ms per unit on an AVX2 core and 2.3 ms scalar; against class v4's own 55,296 shadow instructions at 0.1 ns per lane-instruction (about 0.18 ms per unit) the tile block is 3x the verifier cost per unit at the same premium with AVX2, 13x scalar. On the measured box core `mm1430` passes alone (5.8 ms) and misses the loaded gate by 0.14 ms; `mm512` passes loaded. On a 2019-class laptop core by the 2.5x rule: `mm512` about 13 ms loaded (FAILS), `mm128` about 12.8 (the ALU-free base is already 12.6 on that rule's loaded column; the rule itself is the open O-1.14). A NEON path (M-series: scalar tonight) and a VNNI path (halves the AVX2 count, approximate) are the two unbuilt verifier items.
|
||||
|
||||
**A finding on the per-load placement (not a tuning).** With the same base program and the same loads, the per-load class derives 3,595 to 3,633 distinct items per warp where the class v4 shape derives all 4,096: 11 to 12 percent of the warp's 4,096 reads land on an item another lane of the same warp has already read. The cause (read from the construction, not yet from a trace): a sub-block of 16 drawn ALU instructions ends right before the next load, and when its last writer of the load's source register is a `shfl` (29 of 256 shadow instructions are shuffles), two lanes carry a neighbour's value into the same address; in the class v4 shape the base program's own instructions re-randomise every register between a shuffle and a load, and the sub-version 3 source rule (a load's source last written by an injecting op or a rotate, judged over base then shadow) does not see the per-load order. Consequences: a within-warp duplicate read is served from L1 or L2 on the GPU and from a lane buffer on a chip, so it costs neither side DRAM energy but it is 11 percent fewer dependent reads per hash, which is exactly the uniformity bound the attack pass gates (F8's 1.2x on the hot set); the per-load class as drawn FAILS that spirit and is not a candidate as exported. The fix is one rule: a sub-block's instructions that would be the last writer of the next load's source are drawn from the injecting families only (or the sub-block ends one instruction before the load with a rotate), the sub-version 3 freshness fixpoint run over the real execution order; 2 to 3 agent hours plus a re-export and the F8 census at 2^24 on 64 seeds. The card rows of the exported pack still answer the energy question (the same instruction count per hash, 11 percent fewer DRAM reads, so the pack reads slightly FASTER than the control and that part of the delta is the duplicate reads, not the placement).
|
||||
|
||||
### 20.2a The per-load fault, fixed (22:16 to 22:33 UTC)
|
||||
|
||||
| Reading | Clock (UTC) | Value |
|
||||
|---|---|---|
|
||||
| The known-failed record (the first export, id 854050a4293f0615), traced | 22:16 | 10,728 distinct items of 12,288 over three units (the class v4 shape 12,286); 1,482 same-iteration duplicate lanes at sites 8, 10 and 15 only; the sub-block last writers of those loads' sources `mul`, `mul`, `mul`; shuffles and rotates clean |
|
||||
| The mechanism, from the 64-seed census of the first rule | 22:16 | 29 of 64 seeds failing, up to 620 duplicate lanes a seed, load sources collapsed to 1 to 17 distinct values in 32 lanes; the last base writers named `mulhi`, `mul`, `or`, `rotl`, `load`: a lossy BASE writer followed by 27 passes of a 16-instruction map collapses the register before the next load, so a last-writer rule alone does not cover it |
|
||||
| The fix, committed (2f718001) | 22:24 | two layers: (1) generator: a per-load sub-block instruction writing the NEXT load's source is redrawn from the injecting families when its op is `mul`, `mulhi` or `or`; (2) `accept.rs`: the dynamic test steps the per-load sub-blocks inside `run_unit` in the order the class executes, and a new rejection `DuplicateLanes` (this class only) refuses a candidate whose load reads one address in two lanes of a unit; a rejected candidate redraws the attempt |
|
||||
| After, the trace | 22:23 | 12,287 distinct of 12,288, 0 duplicate lanes on the genesis seed (accepted at attempt 3) |
|
||||
| After, the 64-seed census (2 units each on a second dataset, 16,384 load rows) | 22:23 | 1 duplicate pair in all (seed ca4-census/49, site 3): the chance floor of a 2^24 index space (32 x 31 / 2 / 2^24 per row, about 0.5 pairs expected; the class v4 shape's own trace carries 2 of 12,288 from the same floor) |
|
||||
| After, the drawn-era split (16 eras `igneum-era-test/0..15` over the per-load class, 2 units each) | 22:32 | R under 28: 12 eras, 3,072 rows, 0 duplicate pairs; R 28 and up: 4 eras, 1,024 rows, 0 pairs; every era accepted at attempt 3 |
|
||||
| The suite | 22:23 | 64 + 2 + 7 + 4 + 19 + 2 + 7 passed, 0 failed (igneum-build-2) |
|
||||
| The fixed pack `mx8_shl256x27_v2` | 22:29 | attempt 3, id bd64b207a30413fb, OVERALL PASS; on build-1 and in the hash lane's kit beside the first export (both run on PC 1, the row names its directory and id) |
|
||||
| Metal fingerprints (M5 Max, `packbench`, 3 batches of 2^24, under the measure lock) | 22:30 | `mx8_sh256x27` 3d2e8245cc084d07 (the 6 October fingerprint, 27.01 MH/s); `mx8_shl256x27_v2` ee5d7c71180e5ea7, vectors 3 of 3 standalone and in batch, 26.88 MH/s (-0.5 percent against the control: the per-load placement costs the Apple GPU nothing); `mx8_mm128` 270e4ae36b37e9a1, 3 of 3, **17.67 MH/s (-35 percent)**; `mx8_mm512` a1c1ff3148d775d1, 3 of 3, **5.98 MH/s (-78 percent)**, compile 15.0 s |
|
||||
|
||||
Two readings from the Metal row that move section 15.2's tile verdict. First, the reference tile (`mm8_ref`: 12 index shuffles and 32 byte products per lane) is bit-exact against the Rust verifier on all three vector warps of both tile packs, so the tile semantics are pinned on two implementations before any card runs the PTX form. Second, the Apple cost is far above the "10 ALU steps per tile" estimate: 1,024 tiles per hash cost the M5 Max 35 percent of its rate and 4,096 tiles 78 percent, so a tile shadow at the ALU shadow's premium (11,440 tiles per hash) would take the Apple tier out entirely unless Metal gains an integer matrix path reachable from the toolchain (`mpp::tensor_ops::matmul2d` with `uchar` operands, unverified). The tile shadow therefore stands as a class only with an Apple exemption nobody has designed, which moves it from rank 3 to beside rank 5 until that path is measured; the k question it answers is unchanged.
|
||||
|
||||
### 20.2a-close The per-load construction: REOPENED (8 October 2026, 11:0x UTC, the coordinator's wording): the 22:5x UTC death of the 16 x 27 form stands on the no-era census (both instruments: 0.986 to 0.990 rejection per candidate); a drawn-era read of any per-load form before the 10:5x UTC fix measured the instrument (the window's fixed top bits counted as biased); the sound form (one pass of a 256-instruction sub-block per load, `mx8+shl4096x1`) accepts 234 of 256 seeds on the no-era census and its verdict on drawn eras is the invention lane's 15:00 BST read
|
||||
|
||||
The fix of 20.2a held for distinctness and then met the value-level requirement of 20.2b, and the construction did not survive it. With the acceptance rule stepping the per-load sub-blocks in the order the class executes and judging both the duplicate lanes at a load row and the one-count of every index bit per site over the 64 units (6-sigma band), the 16 x 27 per-load class accepts 22 of 1,621 candidates over 64 seeds (1.4 percent); 42 of 64 seeds exhaust the chain's 32 attempts, which on the chain is an epoch without a program. The first failing test per candidate: a biased index bit 775, duplicate lanes 643, the base rule's lane-constant site 110, (b) 43, (a) 28. The genesis seed accepts none of 32; candidate 0 of the class carries index bit 0 set in 40 of 1,024 addresses at site 0 (z 29.5; the sub-block writer a `rotl` of a `mad` result). The structural reason, read from the record: 27 passes of a 16-instruction map immediately before a load is a tight iteration of a small function, and whatever lossy or product arithmetic it carries (or inherits from the base writer before it) collapses or biases the load's address register before any base instruction can re-randomise it; the class v4 shape places the same 6,912 instructions after instruction 63, where the next iteration's 64 base instructions and 16 loads stand between the block and every load. Both exports (854050a4293f0615, bd64b207a30413fb) were accepted only because the rule did not model the placement; their PC 1 rows stay as the energy reading of the placement, labelled "unsound construction, energy reading only". Design 4 (the shadow per load) and the USD 200 M capex row of 16.2 therefore rest on a construction that does not exist yet; the sound form to try is one pass of a 432-instruction sub-block per load (or 64 x 7), the same N, where the sub-block is a program segment rather than an iterated map; it is a new class to draw, accept and measure (kernel text 6,912 lines per iteration against the 1,024-line block's measured 17 percent on the M5 Max, so the footprint is its own gate), not tonight's. The class v4 shape across the same 16 drawn eras reads 0 duplicate pairs and carries the adv-cache-2 product bit at address bit R exactly in 14 of 17 eras on this pre-amendment generator (one-count 250 or 780 of 1,024, z 15 to 19; sites with `mul`, `rotl` and `or` writers), which is that lane's finding replicated by a second instrument.
|
||||
|
||||
### 20.2a-correction (8 October 2026, 10:5x UTC, the invention lane's reading of the per-load tests)
|
||||
|
||||
The 20.2a-close figures (22 of 1,621 candidates accepted, 42 of 64 seeds exhausting 32 attempts) are the NO-ERA figures: the census ran `mx8+shl256x27` with no era draw, so the index was the plain `x AND mask` and every bit was fair to judge. The instrument as first written was wrong under an era: `BiasedIndexBit` judged every index bit, and under an era a narrow-window site's top bits are fixed by design (`verify::window`), so under any drawn era every per-load form read 0 of 256 for the instrument's reason, not the construction's; the invention lane's census (build-1, this crate at 5984ffab) found it. Fixed in this branch at 10:5x UTC: the test judges only the bits inside each site's window mask. The invention lane's own no-era census confirms the construction's figure for the iterated forms (0.989 to 0.990 rejection per candidate for the 16-instruction iterated forms, the 16 x 27 form dead) and finds the sound form: `mx8+shl4096x1` (one pass of a 256-instruction sub-block after every load) accepts 234 of 256 seeds within 32 attempts at 0.927 rejection per candidate (P(exhaust at 256) about 4e-9), the constant-work ladder flat at 0.927 to 0.949 from a 36-instruction sub-block up; the remaining 0.93 is the product's low-bit law at the load's source (bit 0 in 1,563 of 2,329 bias rejections), which class v4's (a') dataflow rule removes at the draw and the per-load class never applied: the named fix. Its verifier on core 40 with core 88 loaded 8.28 to 8.82 ms against 8.33 to 8.63 for class v4's shape in the same minutes. The class v6 document carries the sound form as layer 5 (`docs/design/class-v6-rotating-family.md` section 7c) with its 5090 rows owed from the hash lane's v6 job.
|
||||
|
||||
### 20.2b A named requirement for every CA4 prototype: no biased product bits in an address (the crypto lane's adv-cache-2 reading, 7 October 2026, 22:1x UTC)
|
||||
|
||||
The crypto lane's finding: a product's low bits are biased (P(bit 0) = 1/4, measured exactly), the bias survives the odd stride multiplier, and the stride rotation places the biased bits at address bits R and up, inside the 28-bit item index unless R is 28 or more. The devnet era draws R = 29, which cuts them off, so 31 of 32 devnet-era programs read clean while 6 of 16 drawn-era programs (R from 3 to 22) show a site over 1.04x (2 over 1.2x, the worst 1.51x); under the 2 GiB genesis dataset (D = 29) R = 29 would show it too. The devnet's cleanliness is an accident of its era draw; the chain prevalence is the drawn-era figure; the price to a partial-store chip stays under 0.1 percent of a hash's reads per site, so no chip number moves. The requirement for any class this file proposes (the per-load shadow, the tile block, a re-weighted shadow): (1) a load whose source register's last writer is a product (`mul`, `mulhi`, `mad`) carries biased low bits into the address, and the acceptance must judge it at the VALUE level (the bit bias of the index at the site over the units), not by the index-distinctness ratio alone, which the duplicate-lane test above is; (2) the census of any candidate reads across the drawn eras split by R (3 to 22 against 28 to 31), as 20.2a now does for the per-load class, never the devnet era alone. The per-load class's dynamic rule covers distinctness, not bias; the value-level test is owed and is the same item for class v5's acceptance. Nothing in class v4 or v5 moves on this without main's word.
|
||||
|
||||
### 20.3 The card rows (PC 1, the RTX 5090 alone, the installed worker, 04:04:33 to 04:22:54 UTC, 8 October 2026 (05:04 to 05:22 UK); the hash lane's job `run-ca4-pc1-packs-5090-20261008-b`, 250 batches of 2^24 at one warp per block, nvidia-smi at 1 Hz with the three power fields agreeing, the 1,300 MHz lock through the helper)
|
||||
|
||||
Every pack's self-test PASS at both states (the cache, the dataset, the 96 vector lanes through the bound kernel), so the int8 tile's inline PTX compiles under NVRTC 12.8 on sm_120 and is bit-exact against the Rust verifier on the card, as the Metal reference was on the M5 Max; the 2^24 fingerprints are the same at both states and equal the Mac's where the Mac ran them (sh256x27 3d2e8245cc084d07, shl256x27_v2 ee5d7c71180e5ea7, mm128 270e4ae36b37e9a1, mm512 a1c1ff3148d775d1; mm1430 8e9b7066239d35d1 on CUDA, the Mac row not run).
|
||||
|
||||
| Pack | Unlocked (SM 2,842 to 2,865 MHz): MH/s / W / MH/W / microjoules per hash | Premium over mx8 unlocked | At the 1,300 lock (SM 1,290): MH/s / W / MH/W / microjoules | Premium at the lock | Label |
|
||||
|---|---|---|---|---|---|
|
||||
| mx8-genesis (class v3, the control) | 137.54 / 311.0 / 0.442 / 2.26 | | 127.32 / 213.0 / 0.598 / 1.67 | | measured |
|
||||
| mx8_sh256x27 (class v4's shape) | 137.51 / 462.2 / 0.298 / 3.36 | 151.2 W, 1.10 microjoules, 10.8 pJ per counted op | 126.93 / 295.8 / 0.429 / 2.33 | 82.8 W, 0.652 microjoules, 6.4 pJ per op | measured |
|
||||
| mx8_mm128 (1,024 tiles per hash, 32,768 MACs) | 137.45 / 332.9 / 0.413 / 2.42 | 21.9 W: 4.9 pJ per MAC | 127.01 / 217.8 / 0.583 / 1.71 | 4.8 W: 1.2 pJ per MAC | measured (the per-MAC figures corrected 8 October 2026, 06:xx UTC: 32 MACs per lane per tile) |
|
||||
| mx8_mm512 (4,096 tiles, 131,072 MACs) | 137.50 / 369.4 / 0.372 / 2.69 | 58.4 W: 3.2 pJ per MAC | 127.07 / 235.4 / 0.540 / 1.85 | 22.4 W: 1.3 pJ per MAC | measured (corrected) |
|
||||
| **mx8_mm1430 (11,440 tiles, 366,080 MACs per hash, the ALU shadow's premium in tiles)** | 137.45 / 457.7 / 0.300 / 3.33 | **146.7 W, 1.067 microjoules: 2.9 pJ per MAC**, 0.103 W per tile per iteration, linear within 5 percent | 126.87 / 284.5 / 0.446 / 2.24 | **71.5 W, 0.564 microjoules: 1.5 pJ per MAC** (0.050 W per tile) | measured (corrected: the first reading divided by 1,024 MACs per lane per tile where the tile gives 32) |
|
||||
| mx8_shl256x27 (the first per-load export, 854050a4293f0615: UNSOUND, energy reading only) | 158.62 / 472.8 / 0.336 | +161.8 W at a rate 15 percent OVER the control (the 11 percent duplicate reads land in L2) | 145.65 / 299.4 / 0.487 | +86.4 W | measured; the construction is dead (20.2a-close) |
|
||||
| mx8_shl256x27_v2 (the fixed export, bd64b207a30413fb: UNSOUND, energy reading only) | 135.90 / 448.3 / 0.303 | +137.3 W, 14 W under the whole-block shape at the same instruction count (the 16-instruction block effect) | 126.04 / 282.9 / 0.446 | +69.9 W, 13 W under the whole block | measured; dead as a class |
|
||||
|
||||
NVRTC compile per pack: mx8 190 ms, sh256x27 269, mm128 331, mm512 869, mm1430 1,239 unlocked and 2,107 at the lock (inside the epoch's compile-ahead budget, the 600-s floor's 38 s). The lock rows read 127 MH/s against the efficiency pass's 134 at the same clock: the installed worker at one warp per block against the app's tuned variant; every ratio here is within one run.
|
||||
|
||||
What the rows say:
|
||||
|
||||
1. **The hash rate is memory-bound on every sound pack at both states** (137.4 to 137.5 unlocked, 126.9 to 127.3 at the lock, within 0.5 percent of the control): 11,440 int8 tiles per hash are free in rate on the 5090, so the 4090's free band (R = 512) extends to at least R = 1,430 on the 5090, as section 15.2 scaled.
|
||||
2. **The tile block carries the ALU shadow's premium at the same hash rate and at a lower cost at the knee**: 146.7 W against 151.2 W unlocked (3 percent under), 71.5 W against 82.8 W at the 1,300 lock (14 percent under). The 5090's energy per MAC at full tile load is 2.9 pJ unlocked and 1.5 pJ at the lock (the 4090's 6 October figure, re-read with 32 MACs per lane per tile, is 1.8 pJ), with the fixed cost of the tensor path visible at R = 128 (4.9 pJ unlocked). The microbench (15.1a) reads the same path at 4.1 and 2.2 pJ per MAC on a dependent u8 chain and 1.36 and 0.83 on the wide s8 tile at 80 percent of peak.
|
||||
3. **The GPU's measured cost per MAC is NOT inside the band of what a 5 nm MAC array costs** (NVIDIA's own test chip at 0.46 V: 0.021 pJ per INT4 MAC, about 0.04 to 0.08 per INT8-class MAC; at nominal voltage about 5x that; JSSC 2023 via Dally's slides, claimed): against the 5090's 1.5 to 2.9 pJ per MAC the chip's `k` on tile work reads about 0.03 to 0.3 (approximate), BELOW the ALU shadow's 0.3 to 0.8. The earlier reading of this row (k 0.9 to 1.7) rested on the factor-of-32 error corrected in 15.1a.
|
||||
4. The per-load placement is cheaper per instruction than the whole block (the 16-instruction block effect: 13 to 14 W under the whole block for the same 55,296 instructions per hash) and the construction is dead (20.2a-close); the first export's 15 percent rate gain is the duplicate reads served from L2, the fault made visible on the card.
|
||||
|
||||
### 20.3a The SM-sparse job's first run (PC 1, 00:22 to 00:41 UTC, 8 October 2026): no variant ran; the knee rows it did read
|
||||
|
||||
The hash lane's run `run-ca4-pc1-ca4sparse-5090-20261007` (the ca4sparse exe, the 5090 alone, class v3 and v4 packs, unlocked and at the 1,300 MHz lock, 48 rows, exit 0, every fingerprint equal to the Mac's, self-test PASS) served the BASE kernel on every row: the worker's `--bench` path turned the race off and built the pair without one, so `--variant sp43-w32` was parsed and never applied ("race 0 ms variant base" on all 48 rows, no NVRTC error because NVRTC never saw the variant). The numbers confirm it: every sparse row's rate equals base (v4 137.0 to 137.1 MH/s unlocked, 133.8 to 134.3 at the lock; v3 136.7 to 136.8 and 133.8 to 134.0) and the draw drifts with heat, not with the SM count (v4 unlocked 465.5 W at 69 C on the first row rising to 485.2 W at 76 C on the sixth; v3 331.6 falling to 319.2 W as the card cooled from 71 to 62 C). The SM-sparse question is unanswered by this run. The fix (worker.cpp, this branch, 8 October 2026, 00:xx UTC): a `--bench` with `--variant <name>` runs the pinned race (base and the named variant only, no timing, the variant installed whatever its speed) and the RESULT line carries `variant=`, `sparse_blocks=` and `block_warps=`; a run whose served variant is not the requested one prints a `variant_not_installed` line, which is the known-failed case of this fix. The rerun is the hash lane's queue slot.
|
||||
|
||||
What the run did read, and keeps (the 5090 at the 1,300 MHz lock, base kernel, measured): class v4 134.26 MH/s at 309.9 W (0.433 MH/W, SM 1,290 MHz), class v3 134.03 at 219.4 W (0.611 MH/W); the class v4 premium 90.5 W at the lock against 133.9 W unlocked on this run (145.3 on the efficiency pass; the unlocked v4 rows ran hot, 465 to 485 W). The premium per counted op at the lock: 90.5 W over 134.26 M x 102,100 = 6.6 pJ (6.5 on the efficiency pass).
|
||||
|
||||
### 20.3b The SM-sparse reading (PC 1, the 5090 alone, 04:35:49 to 05:12:48 UTC, 8 October 2026 (05:35 to 06:12 UK); the hash lane's job `run-ca4-pc1-ca4sparse-5090-20261008-c` on the fourth exe, sha256 a4550202...): the candidate is dead
|
||||
|
||||
The run as a gate: the card-free check on the card's own exe read `variants=2 names=base,sp43-w32`, `rewrite=applied`, `call="igneum_hash_bound_unit(ds, out, baseNonce, mask, iw, gid)" names_match=1`, `nvrtc64_120_0.dll sm_120 compiled=1 image_bytes=31384`; every sparse row served its variant (`served=sp<N>-w32 sparse_blocks=N`, no "compile:" text) and every fingerprint on every row equals the Mac's (class v4 e370fb2080b7dbb1, class v3 90f794dd556f7a3b), so the rewritten persistent kernel is bit-exact. 32-second rows, the three power fields agreeing, idle 74 W; the drift check at the end repeats the unlocked rows within 1 W and 0.2 MH/s.
|
||||
|
||||
| Variant (blocks of 32 warps, about one per SM) | Class v4 unlocked: MH/s / W / MH/W | Class v3 unlocked | Class v4 at the 1,300 lock | Class v3 at the 1,300 lock | Label |
|
||||
|---|---|---|---|---|---|
|
||||
| base (the plain grid, one warp per block) | 137.07 / 450.8 / 0.304 | 136.77 / 313.9 / 0.436 | 134.31 / 301.0 / 0.446 | 134.04 / 211.4 / 0.634 | measured |
|
||||
| sp170-w32 (the persistent shape on every SM) | 136.09 / 461.2 / 0.295 | 136.02 / 313.3 / 0.434 | 128.95 / 297.6 / 0.433 | 129.20 / 208.2 / 0.621 | measured |
|
||||
| sp85-w32 (half the SMs) | 136.29 / 461.7 / 0.295 | 136.87 / 310.7 / 0.441 | 125.37 / 290.1 / 0.432 | 129.80 / 207.8 / 0.625 | measured |
|
||||
| **sp43-w32 (a quarter)** | **134.58 / 460.1 / 0.293** | **136.55 / 309.8 / 0.441** | 64.33 / 208.2 / 0.309 | 126.25 / 205.4 / 0.615 | measured |
|
||||
| sp21-w32 (an eighth) | 70.75 / 327.4 / 0.216 | 132.82 / 303.4 / 0.438 | 31.48 / 155.2 / 0.203 | 84.78 / 174.9 / 0.485 | measured |
|
||||
| sp11-w32 (a sixteenth) | 37.34 / 250.9 / 0.149 | 100.14 / 274.5 / 0.365 | 16.55 / 116.7 / 0.142 | 44.57 / 147.5 / 0.302 | measured |
|
||||
|
||||
The two readings the design asked for. (1) The shape itself: sp170-w32 against base costs 0.7 percent of rate on class v4 and 0.5 on v3, +10 W on v4 and 0 W on v3: within noise. (2) Watts minus idle per MH/s against base: class v4 unlocked base 2.75 W per MH/s, sp43 2.87, sp21 3.58, sp11 4.74; class v3 unlocked base 1.75, sp43 1.73, sp21 1.73, sp11 2.00. A quarter of the SMs holds 98.2 percent of the class v4 rate at the SAME draw (460 W against 451) and 99.8 percent of the class v3 rate at 4 W less; the draw falls only when the rate falls, and the energy per hash never falls below base on either class. **So the SM-side 99 W of section 2 is not activity that idle SMs would save: the card's draw follows the work, not the SM count; the ALU work per hash costs the same on 43 SMs as on 170, and the 43 SMs run it at the same energy.** The candidate is dead by its own rule (watts track the work one for one), and the lock rows say the same from the other side (class v4 sp43 at 1,300 MHz is compute-bound at 64 MH/s: the shadow's ops are real throughput work). The one number kept for the chip model: the class v4 premium at sp43 unlocked, 150.3 W over class v3 at a held rate, equal to the full-card premium (136.9 W base, 151.2 on the packs job), so the shadow's energy does not depend on how many SMs carry it; it is the energy of the ops.
|
||||
|
||||
What this closes in the ranking: rank 2 (the SM-sparse miner kernel) is dead; the premium-free floor of section 0 stays the card's idle plus its memory system plus whatever `E_wait` is at the knee, and the only miner-side lever on `E_card` is the operating point (rank 1). The 5090's four runs of this night (three failed, one read) also gave four repeats of the knee pass: the class v4 premium 133.9 to 151.2 W unlocked and 82.8 to 90.5 W at 1,300 MHz, with every one of the three power fields agreeing within 0.2 W.
|
||||
|
||||
### 20.3c The hot-table rows (PC 1, the 5090, the hash lane's job `run-ca4-pc1-hot-ldcs-5090-20261008b`, 36 rows all PASS, bench-log "8 October 2026, the hot-table packs on the RTX 5090", master c09dfee4 at 07:12 UTC): a record, and the call
|
||||
|
||||
(1) `--variant ldcs` (streaming dataset loads) equals base on every pack in both states within 0.1 MH/s and 1 W: the cache-policy hint is a dead lever, as 15.1a's L2 row already said. (2) At the 1,300 lock the replaced-form hot packs lose 2 percent of rate where mx8 loses 7.5 (hot32k4 147.4 to 144.4 MH/s, hot64k8 164.2 to 160.9, mx8 137.7 to 127.3) while every pack drops a third of its watts (319 to 215 W): the hot packs are latency-bound on the table. (3) Per watt at the lock: hot64k8 0.734 MH/W, hot32k4 0.671, hot64k4 0.645, hot96k4 0.635, hot64k2 0.630, the mx8 control 0.602 (equal to the efficiency pass's 0.60, so the rows are comparable), the added-form packs 0.52 to 0.54.
|
||||
|
||||
The call, on the identity: a hash 22 percent cheaper on the GPU (hot64k8, 1.36 against 1.66 microjoules at the lock) is a LOSS for resistance in the replaced form, not a gain, because the saving sits in the one path where the chip is cheaper still. The GPU saves 0.30 microjoules by moving 64 of its 128 reads from DRAM (8.7 nJ per read measured, 15.1a) to L2 (1.4 nJ): 4.7 nJ per replaced read; the f = 1 chip moves the same 64 reads from GDDR7 (2.0 nJ, modelled) to on-die SRAM (0.2 to 0.5 nJ, approximate), saving about 1.6 nJ per read, 0.10 microjoules of its 0.466, and, because its rate is bound by the memory's activate ceiling (166 MH/s on the 5090's 16 devices), halving its DRAM reads per hash doubles its rate on the same memory and halves its capex per MH/s. The edge per joule at the lock: mx8 1.66 / 0.466 = 3.6x; hot64k8 about 1.36 / 0.36 = 3.8x (modelled chip, approximate); per dollar the chip's 2.8 USD per MH/s becomes about 1.5. This is the 5 October finding (the replaced form "helps the on-die-cache chip, x1.33 at k = 4", counter-asic-2.md decision 5) re-read against the stored-dataset chip with measured GPU energies: the replaced form is a per-watt gift to every miner and a larger one to the chip. The added form ("a" packs, 0.52 to 0.54 MH/W) costs the GPU rate for no chip cost, as chip-model-v3 section 3 said. Nothing moves.
|
||||
|
||||
### 20.4 Into the chip model (the rows measured; the chip side modelled: f = 1 GDDR7 0.466 microjoules per hash, one HBM3 stack 0.321)
|
||||
|
||||
Energy per hash = the card's measured; the chip's = `E_mem + k x F`, `F` the measured premium per hash, `k` the chip's cost per forced operation over the 5090's at the same state.
|
||||
|
||||
| Row (the 5090) | Card microjoules | Premium F | Chip edge, GDDR7, at k = 0.3 / 0.5 / 1 / 1.5 / 2 | HBM3 one stack at k = 1 | The honest k band for the work | Break-even cap, years 1 to 2 (the mission lane's model) | Status |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| class v3, unlocked | 2.26 | 0 | 4.85x at every k | 7.0x | | USD 17 M (28 nm controller) | measured card |
|
||||
| class v3, 1,300 lock | 1.67 | 0 | 3.59x | 5.2x | | USD 17 M | measured |
|
||||
| class v4 shape, unlocked | 3.36 | 1.10 | 4.21x / 3.31x / 2.15x / 1.59x / 1.26x | 2.37x | ALU: 0.3 to 0.8 | USD 100 M (an N5 shadow core) | measured; the served figures 2.1x at k = 1, 3.4x pessimistic |
|
||||
| class v4 shape, 1,300 lock | 2.33 | 0.652 | 3.52x / 2.94x / 2.08x / 1.62x / 1.32x | 2.39x | ALU: 0.3 to 0.8 | USD 100 M | measured |
|
||||
| **tile block mm1430, unlocked** | 3.33 | 1.067 | 4.23x / 3.33x / 2.17x / 1.61x / 1.28x; **at the measured k 0.03 to 0.3: 4.2x to 6.7x** | 2.40x | **int8 MAC: 0.03 to 0.3 (corrected 15.1a)**: the chip's array undercuts the tensor core 4x to 30x per MAC | USD 100 M (a licensable MMA block is USD 4 of N5; the N5 die is forced as for class v4) | measured card; the lever is WORSE than the ALU shadow |
|
||||
| **tile block mm1430, 1,300 lock** | 2.24 | 0.564 | 3.46x / 2.98x / 2.17x / 1.73x / 1.42x; **at k 0.03 to 0.3: 3.5x to 4.6x** | 2.53x | the same | USD 100 M | measured; worse than the ALU shadow's 3.5x at k 0.3 only at the very bottom of its band, and the band's centre is under 0.1 |
|
||||
| the per-load shape (the fixed export) | 3.30 / 2.24 | 0.998 / 0.555 | as the class v4 shape within 5 percent | | ALU | USD 200 M if the sound form exists (16.2) | dead as drawn (20.2a-close) |
|
||||
|
||||
The k column was priced at this point: section 15.2 set the tile count from the 4090's 6 October figure so that the block would carry the ALU shadow's 0.654 microjoules (R about 1,430), and the exported pack is 1,430 tiles per iteration, 11,440 per hash; the 5090 reads 2.9 pJ per MAC unlocked and 1.5 at the lock at that point (15.1a's correction: 32 MACs per lane per tile), so the GPU-side cost of the chip-model-v3 5.11 column is measured at the premium it was priced for (1.067 microjoules unlocked, 0.564 at the lock, against the ALU shadow's 1.10 and 0.652), and the column's premise (a chip's MAC no cheaper than the GPU's) is false by 4x to 30x on the public chip figures. Reading, corrected: at `k = 1` the tile block and the ALU shadow give the same 2.1x to 2.2x and the tile block's premium at the knee is 14 percent lower for it; but `k = 1` is not where a chip sits on tile work. The GPU's tensor core costs 1.5 to 4 pJ per int8 MAC measured, a 5 nm MAC array 0.04 to 0.4 claimed, so the chip's `k` on tiles is 0.03 to 0.3 against the ALU shadow's 0.3 to 0.8: at the same premium the tile block leaves the chip 3.5x to 6.7x where the ALU shadow leaves it 2.1x to 3.5x. **The tile block does not beat class v4's premium (it matches it, 14 percent cheaper at the knee) and does not beat class v4's chip edge at any k a chip can reach; it is the worse lever, and the 6 October verdict on scheme B stands for the right reason now (the honest card's tensor path is not near the floor; the chip's is).** The Apple side (35 to 78 percent of the M5 Max's rate at 1,024 and 4,096 tiles), the verifier (AVX2 0.047 microseconds per tile per unit, `mm1430` 10.14 ms loaded) and the AMD layout no longer need deciding; nothing served moves (the chip texts rest on the ALU shadow).
|
||||
|
||||
**The first sentence, on measured rows (06:xx UTC, 8 October 2026, after the microbench's correction): neither prototype beats class v4's premium or its chip edge; the tile block matches the premium at the same hash rate (146.7 W against 151.2 W unlocked, 71.5 against 82.8 W at the 1,300 lock) and is the worse lever against a chip, because the 5090's int8 MAC costs 1.5 to 4 pJ measured where a 5 nm array costs 0.04 to 0.4 claimed (k 0.03 to 0.3, under the ALU shadow's 0.3 to 0.8); the per-load placement is dead as a construction; the SM-sparse kernel is dead on measured rows (the draw follows the work, not the SM count); the L2 hot table is dead on the GPU side (an L2 hit costs 1.4 to 2.4 nJ against a chip's 0.2 to 0.5); and a shuffle-heavy re-weight costs the GPU 5x the add per op (55.8 pJ), so it stays held. The ALU shadow at the operating point's knee is the floor the night leaves: 2.1x at k = 1 for 82 to 90 W on a 5090, measured four times.**
|
||||
484
docs/design/class-v6-rotating-family.md
Normal file
484
docs/design/class-v6-rotating-family.md
Normal file
|
|
@ -0,0 +1,484 @@
|
|||
# Class v6, the rotating family: four layers that change the hash on a schedule no release carries
|
||||
|
||||
8 October 2026, 10:0x to 17:xx UTC (11:0x to 18:xx UK), branch `counter-asic-4`, the Counter ASIC 4.0 research lane under the Counter ASIC coordinator, the hash lane for the measured per-tier rows. Opened on the founder's question of 11:0x UK: "can we add in any more layers? class rotating? things that would render an ASIC useless as soon as it dropped." Status: a design document and a gate plan; no consensus code this week; nothing here touches the devnet, the testnet object or any served number. Every figure carries a label: **measured** (a card on a named job), **modelled** (arithmetic on the chip model's cited figures, `docs/analysis/chip-model-v3.md`), **claimed** (a vendor's or an author's figure), **approximate** (from memory or a scaling). Outline and layer table at 10:5x UTC; the hash lane's rows due 16:00 UK; the document complete by 18:00 UK; at each clock the default lands with the gaps named.
|
||||
|
||||
## 0. For the founder tonight: what each layer does to a chip on its release day, and what it costs a 5090, a 5070 Ti and an M5 Max
|
||||
|
||||
**The lead, as main ordered it (11:2x UK): the strongest chip in the five-year window is not a DRAM chip and no rotating layer reaches it.** Lane B's reading (`docs/analysis/class-v6/hardware-future.md`, master 34f63b3c; carried into the chip model as section 5.12): a 2 GiB SRAM full store on one reticle of merchant N2 (452 mm^2 of macro at 38 Mb/mm^2, claimed) reads about 2,100 MH/s per die at 300 W, 0.14 microjoules per hash, 17x the 5090 at its stock point per joule at zero shadow (8x to 30x on the read-energy band), 5.6x the M5 Max; with the class v4 shadow 4.8x at a core half as costly as a GPU lane and 2.7x at `k = 1`, and at the honest card's whole latency shadow 3.7x and 2.0x; USD 400 to 600 of silicon per die (USD 0.25 to 0.4 per MH/s), an N2 project of USD 100 M to 500 M and 18 to 24 months (claimed), which on the mission lane's model is a break-even cap of about USD 330 M to 1.7 B in the chain's first two years. Beside it a custom HBM4E base die (2028 or later) at 6.5x to 14x, untouched by all four layers; per-bank processing in memory structurally blind to the hash (1.6 percent of reads in-bank at 2 GiB). All modelled on the chip model's method; the GPU side measured. Against the SRAM store at the hash's own width floor lane 3 reads 66x at zero shadow (lane B's 17x was the 64-byte row) and 2.7x to 3.0x at k = 1 and 5.0x to 5.7x at k = 0.5 with the shadow (section 10.3), so the honest range with the shadow is 3x to 6x and the shadow is the whole hold; and the one layer that answers it is layer 2 as a FLOOR that grows faster than SRAM cost falls: each doubling of the floor adds a die (4 GiB two dies at about USD 1,000 and 15x, 8 GiB four dies at USD 2,000 to 2,500 and 13x) while every card tier pays in device memory and, on Apple, in rate. The schedule's cost is priced in section 3 (lane A's per-tier table and main's candidate schedule, 6 GiB at the v6 epoch, 10 two years on, 14 at four, land there by 17:00 UK) and the served chip line is reviewed against this reading in section 9 at 20:00.
|
||||
|
||||
The honest frame first, from last night's measured close (`docs/analysis/counter-asic-4-research.md` sections 2 and 15.1a). Two chips exist in the model. A **fixed-function chip** wires one hash: the mixer's rounds, the op mix's lane ratios, the read width, the program length and the block shape are silicon. A **GPU-like chip** stores the dataset in commodity DRAM and runs the hour's program on a programmable integer core (the `f = 1` chip of chip-model-v3 section 5, the chip anyone builds); for it every drawn parameter is firmware. The four layers below render the FIRST kind useless on the day a parameter leaves its wired value, and move nothing for the second kind except the size of the core it must carry and the capex that forces. What the second kind keeps is the identity of section 2: at zero shadow premium the card's whole-card energy over the chip's memory energy (3.6x on a 5090 at its knee, measured card, modelled chip), and with the class v4 shadow 2.1x at a core as good as a GPU lane, which the k lane's first RTL rows (10.2, 13:3x UK) say no chip's UNITS are: at N3 a bare unit pays about 0.18 of the locked 5090's pJ per op (a per-unit floor, synthesised on ASAP7 and scaled; no fetch, decode or register file, which is the part rotation forces a chip to carry), so the DRAM chip with the shadow reads up to 4.0x at the knee and 5.8x at stock as the worst case, with the sequencer-core row owed as the headline and k 0.5 (2.9x at the knee) the default until it lands. No rotation changes that; the layers change which chip can be built and how long its tape-out lives.
|
||||
|
||||
| Layer | A fixed-function chip on its release day | A GPU-like chip on its release day | RTX 5090 (measured where stated) | RTX 5070 Ti (MEASURED stock, a rented Vast pod, 10:08 to 10:11 UTC: class v3 78.69 MH/s at 140.8 W, class v4 78.78 at 224.0 W, the premium 83.2 W = 10.3 pJ per counted op, fingerprints equal to the Mac's; no core-lock grid, the host refused -lgc) | Apple M5 Max (measured where stated) |
|
||||
|---|---|---|---|---|---|
|
||||
| 1. Per-era draws of the parameters now fixed by release (mixer round count, op-mix weights, read width, program length, shadow block shape), from chain state | dead at the first era whose draw leaves its wired value: a wired x8 mixer at an x4 era (x16 is out of the band: the verifier measured 11.4 ms with the sibling loaded for the whole row) recomputes nothing right; a 4-byte-granularity controller at a 16-byte era moves 4x the bytes; a core sized for 100,000 ops at a 200,000 era runs at half rate; the tape-out lives one era (180 days, layer 3) | firmware; the core sized for the band's top (200,000 ops: about 60 mm^2 of N5 instead of 30, USD 25 to 40 more per chip, modelled); `k` unchanged; capex per MH/s +10 to 20 percent (modelled) | rate: 0 across the band while latency-bound (the ladder's rungs 0 to 2: 0 and -2.7 percent at the 431 W cap, measured); watts: the premium follows N and the mix (6.2 to 11.3 pJ per counted op measured; a shuffle-heavy draw up to 5x per op, so the band excludes it); the verifier +0.2 to +1.5 ms per era draw (measured ladder, estimated mixer) | rate 0 at class v4 (78.69 to 78.78 MH/s, memory-bound, measured); the class v4 premium 83.2 W at stock (10.3 pJ per counted op, the 5090's 10.8), 224 W under its 300 W limit with no throttle, so the N band's top (200,000 ops) costs about 170 W of premium at stock by the per-op figure and the card's limit binds first (approximate); the knee rows need a host that allows -lgc | rate -3.3 to -10 points across the N band (measured ladder rungs 0 to 2), 0 for the mixer and the width; the Apple tier sets the band's top (130,000 ops at the 5 percent rule) |
|
||||
| 2. The state-derived dataset's size tracks chain-state growth with a floor (class v5's leaves scaled by the state, never below the 1.13.3 schedule) | a chip with fixed memory ages out when the dataset passes it: one HBM3 stack 24 GB, the 5090's board 32 GB; the time-memory curve (chip-model 5.4) says the excess must be recomputed at 6.3 nJ per item against 2.0 per read, so its energy per hash rises with the overflow | the same memory limit; a chip buys DRAM a card cannot (24 GB stacks at USD 200, modelled), so it ages out LAST: the 8, 12 and 16 GB card tiers go first | fine to 32 GB: the 2 GiB genesis dataset plus the state's leaves; at a used chain's 5 GB state (class-v5 section 3) the dataset is capped at the sample size, 2 GiB | 16 GB: fine to the 8 GiB step (year 12 on the schedule); the state floor does not move it | 36 to 128 GB unified: fine to the 16 GiB step; the daily build grows with the size (13 to 30 ms measured at 1 GiB) |
|
||||
| 3. Scheduled family epochs by height, every 180 days by default, no release (the reserve R0 to R8 of spec 1.13.2 unlocking by height, then rotating) | a chip without the family's datapath loses its weight of the mix at the unlock (4 points of 79) or emulates it at the vendor penalty (1.5x to 2.4x per op, measured on the cards); a chip taped out against one family set is a GPU-like chip or dead | pre-wires every family for about USD 4 of N5 silicon (algorithm.md 5.2, modelled); moves the per-joule edge under 10 percent per family | measured family step costs: shfla 1.53x, perm 1.30, mm8 2.43 the add-xor-rotate step; under 1 percent of rate at 4 points | pending; the same families on the same silicon generation | measured: shfla 1.91x, perm 1.13 emulated, mm8 emulated at 1.6x per dot4; under 1 percent of rate at 4 points |
|
||||
| 4. The acceptance floor (c''') and the F8 uniformity test generalised to every era's draw, with a redraw on failure (plus the per-site largest-bucket bound and the value-level bias test as the next class's two tests) | nothing on the chip; it is what makes layers 1 and 3 safe without per-era cryptanalysis | nothing | the generator's attempts per seed (today about 30 at 2.4 percent rejection under (c'''); a redraw costs nothing on a card) | the same | the same |
|
||||
|
||||
Per tier against the stored-dataset chip (GDDR7, 0.466 microjoules per hash, modelled), at stock: the 5090 at 2.26 microjoules (class v3) and 3.36 (class v4) reads 4.9x and 2.2x at `k = 1`; the 5070 Ti at 1.79 and 2.84 microjoules reads 3.8x and 1.9x at `k = 1` (a card whose stock point is nearer its knee); the M5 Max at 0.78 and 1.40 (GPU and DRAM channels) reads 1.7x and 0.9x. The layers do not move these; the operating point and the shadow do.
|
||||
|
||||
What a fully general chip still gets, in one line: the stored-dataset chip with a programmable core at the top of the band keeps 3.6x at zero premium and 2.1x at `k = 1` on a 5090 at its knee (section 2 of the research file), and the four layers cost it about USD 30 to 60 per chip of extra silicon and a project that must be N5-class from the first tape-out (the USD 100 M break-even cap of the mission lane's model); the chips of 2027 on (lane B, section 7a: a custom HBM4E base die, DRAM on logic, a 2 GiB SRAM store on one N2 reticle at 17x per joule at zero shadow and USD 0.25 to 0.4 per MH/s) are touched by none of the four layers except layer 2's capex, and against the SRAM chip the public 2x needs the honest card's whole latency shadow at a core as good as a GPU lane.
|
||||
|
||||
### 0.1 The per-tier table for the schedule decision (main's order, 11:4x UK: the year each tier falls off under the candidate schedule, and the share of today's measured cards)
|
||||
|
||||
The candidate schedule main named: a floor of 6 GiB at the class v6 epoch, 10 GiB two years on, 14 GiB at four years, each step by height like a class epoch, the schedule a consensus field (lane A prices it in section 3 by 17:00 UK). The device memory a tier needs is the dataset plus the cache (512 MiB from 4 GiB, 1 GiB from 8) plus the measured 0.4 GiB working set and about 0.5 GiB of driver and app (the hash lane's rows); a Mac holds about half its unified memory for the GPU under macOS, the display and the node (the card-lifetime table's share rule).
|
||||
|
||||
| Tier | Device memory | Needs at 6 / 10 / 14 GiB | Falls off the candidate schedule at | Falls off the 1.13.3 schedule at (year) | The Apple rate cost, measured today | Share of today's measured cards | Label |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| RTX 5090, 32 GB | 32 | 7.4 / 11.9 / 15.9 GiB | never inside the schedule (about 29 GiB) | 54 | | the bench table's row; the fleet census unread | modelled from the measured working set |
|
||||
| RTX 5070 Ti, 16 GB | 16 | 7.4 / 11.9 / 15.9 | at the 14 GiB step (four years on): 15.9 of 16 with no headroom | 23 (about 13.5 GiB) | | unread | modelled |
|
||||
| A 16 GB GPU (the 9070 XT class) | 16 | the same | the 14 GiB step | 23 | | unread | modelled |
|
||||
| A 12 GB card (the 4070 measured) | 12 | 7.4 / 11.9 / 15.9 | at the 10 GiB step (two years on): 11.9 of 12 with no headroom | 15 (about 9.5 GiB) | | unread | modelled |
|
||||
| An 8 GB card | 8 | 7.4 fits with 0.6 of headroom / out / out | at the 10 GiB step; at 6 GiB it keeps mining with 0.6 GiB spare (the 6 GB 2060 tier drops at the v6 epoch) | 12 | | unread | modelled |
|
||||
| Apple, 16 GB unified (about 8 GiB for the GPU) | 16 | 7.4 fits at the v6 epoch with no headroom / out / out | at the 10 GiB step (two years on) | 12 (about 8 GiB) | the rate -12 percent at 2 GiB, -20 at 4, -22 at 8 against 1 GiB on the M5 Max (measured 10:40 UTC); a 16 GB Mac at 6 GiB pays the 8 GiB row's cost or more | unread | modelled memory, measured rate |
|
||||
| Apple, 32 GB unified (about 16 GiB) | 32 | 7.4 / 11.9 / 15.9 | holds every step of the candidate schedule, 15.9 of about 16 at the 14 GiB step with no headroom | 28 | the same measured curve | unread | modelled memory, measured rate |
|
||||
| Apple, 64 GB and up (about 32 GiB) | 64 | fits | holds every step | 60 | the same measured curve (the M5 Max itself) | unread | modelled memory, measured rate |
|
||||
|
||||
Lane A's check of these steps under the standing 75 percent rule (section 3.4): 6 GiB does not fit the 8 GB tier (78 to 82 percent of the card), 10 GiB retires the 12 GB tier and the 16 GB Mac at year 2, 14 GiB retires the 16 GB tier at year 4; the schedule that drops the tiers in the order the note described is 5.5 / 8 / 11 GiB, each step retiring about a quarter of today's measured consumer cards by count. The Apple column is the one the founder decides on: the M5 Max is the honest best per joule (0.78 microjoules on its GPU and DRAM channels, 3.1x the 5090) and the Mac tier is a large audience; a 16 GB Mac is out at the 10 GiB step and every Mac pays the measured rate cost of a larger working set (-12 to -22 percent) before any memory limit, which the NVIDIA cards pay too at their knee by the 12:5x UK measurement (the 5090 at the 1,300 lock: -4.9 / -11.2 / -14.1 percent of rate at 2 / 4 / 8 GiB, 4 / 8 / 10 percent more energy per hash; unlocked -2.8 / -3.8 / -4.3; section 3.3). The share of today's measured cards per tier is unread until the fleet's census is sent; the bench table's rows name the cards, not their count.
|
||||
|
||||
## 1. Where the four layers sit in the design as it stands
|
||||
|
||||
| Item | Today (spec 01, the class system) | What class v6 adds |
|
||||
|---|---|---|
|
||||
| The era draw (1.13.1) | one SplitMix64 stream from the 1-hour VDF's `E_n`; draws the op-weight perturbation (B = 2 points, proposed), the fold rotations, the table layout and the working-set window (layers 4 and 8 of Counter ASIC 2.0, IN); `epoch_len` and the ladder's draws consumed and not used (set by signal) | layer 1: five more draws of the same stream, in a fixed order, each within a genesis-fixed band, each consumed whether used or not |
|
||||
| The mixer (1.8.5, `mixer_mult` 8 under class v3) | fixed at genesis "so the verify budget holds" | layer 1 draws it in {4, 8, 16} (the x4 and x8 rows measured, x16 the hash lane's row or the estimate), the verifier's budget the bound |
|
||||
| The read width (1.13.1 draw 1, `allowed = {1}`) | pinned to 4 bytes by the 5 October decision (w64 made the 5090 bandwidth-bound) | layer 1 draws from {1, 4} words (w16 within 2.7 percent on the 5090 and the 9070 XT, measured); never 16 words |
|
||||
| The shadow (class v4: 256 x 27, rung 0 of the ladder) | N moves by miner signal within the ladder | layer 1 draws the block shape (64 to 256 instructions, measured; never 1,024) and the op-mix weights within the band; N stays the ladder's (signal, not draw) |
|
||||
| The dataset (1.13.3) | 2 GiB plus 0.5 GiB a year; class v5's leaves (64 B per state record) | layer 2: the leaf count tracks the state with the schedule as the floor; the sample cap stays at the dataset size |
|
||||
| The reserve (1.13.2) | R1..R8 unlock one per era, by height | layer 3: the unlock cadence a genesis constant (180 days) and the set rotating after the reserve is exhausted (the retired family returns in a fixed cycle) |
|
||||
| The acceptance rule (1.4.6; sub-version 3's (a') (c') (c'''); F8's census as a gate) | judged on the genesis parameters; the census run per class by hand | layer 4: the rule and the census parameterised by the era draw; the generator redraws a candidate that fails any of the four tests |
|
||||
|
||||
## 2. Layer 1: per-era draws of the released parameters
|
||||
|
||||
The band a card sees across the draw's range, from the hash lane's cost rows (12:0x UK; every row labelled): a 5090 136 plus or minus 4 MH/s and 312 to about 460 W stock (223 to 300 W at the knee), a 5070 Ti 79 plus or minus 2 MH/s and 141 to 224 W, an M5 Max 27.7 to 28.3 MH/s and about 60 to 96 W (GPU power approximate); the band's width is the program-length and op-mix draws (watts), not the mixer or the read width (rate). The two re-weighted shadow packs run on the 5090 this afternoon and replace the op-mix line's modelled figures when they land. The table:
|
||||
|
||||
| Parameter | Band (genesis) | Why that band (the measured rows that set it) | Chip rows: fixed-function / GPU-like | Per tier: 5090 / 5070 Ti / M5 Max | Verifier | Open number |
|
||||
|---|---|---|---|---|---|---|
|
||||
| Mixer applications per round `m` | **{4, 8} at genesis; 16 in the list as `admissible: false` (the ladder's rule: a flag in the genesis list, flipped only by the 90 percent upgrade path once the 2019-class core measures it)** | **The card pays nothing for the draw, MEASURED (the hash lane's kit b on PC 1, 12:49 to 12:58 UK, generator 2 with the era, the multiplier the only difference, 250 batches per row, fingerprints PASS): x8 137.73 MH/s at 320.2 W unlocked and 127.44 at 212.0 W at the 1,300 lock; x16 137.72 at 320.1 and 127.46 at 211.7; within 0.1 MH/s and 0.5 W at either state, item derivation hidden under the read chain, so the verifier's 1.86x per doubling is the draw's whole cost.** x4 and x8 measured (mixer-x4.md 6.4: 1.92 and 2.79 ms per unit on a loaded M5 Max core; the build unmoved on every discrete card, latency-bound); **x16 MEASURED (today's class v3 stream, the same id f5e904bc5d148926 for x8 and x16). Cold alone: build-1 core 4 x8 4.67 ms per warp, x16 8.69 (the hash lane, 10:13 UTC); build-3 core 4 (AX102, a faster core class) x8 3.48, x16 6.53 (the build-server lane, 11:01 UTC). With the SMT sibling loaded for the row's whole length (the method that counts: a sibling bench looped for the row, not run once, which ends in about a second and leaves the verify phase sibling-idle; the hash lane's earlier 5.41 / 10.11 / 9.04 to 9.10 ms rows were of that kind and are withdrawn): build-3 x8 6.24, x16 11.44 ms, about 1.8x on both classes. The multiplier is 1.86x on the verifier (the mixer IS the verifier's cost); chip-model-v3's estimate (9 ms on a 2019-class core) stands within the cold rows** | the `f = 0` recompute chip's rate halves per doubling (0.31x bare at x8, 0.16x at x16, modelled); the `f = 1` chip unmoved (it recomputes nothing) | rate 0 / 0 / 0 (latency-bound, measured at x4 and x8); the daily build 42 / pending / 29 ms at x8 (measured); the x16 build on the 5090 rides the 12:45 UK job | +1.9x per doubling measured (x8 to x16); **x16 at 11.4 ms loaded is over the 10 ms gate on the measured row, so `admissible: false` stands on the measurement, not only on the chip model's margin; the O-1.14 laptop run can only tighten it** | adv-mixer-3's margin (every statistic clean from k = 1, the SAT ladder k = 1 solved, k = 2..4 timeout): x4 keeps the margin by that report's reading; the band {4, 8} is the measured one |
|
||||
| Op-mix weights (the ten non-load families) | **B = 4 points on the injecting families only (add, sub, xor, mad, shfl, rotl, rotr); the lossy families (or, mul, mulhi) capped at their base, the ring-A rule `or + mul + mulhi` at most the table's 18 plus B, which keeps the per-candidate rejection r under 0.85 (r^256 under 1e-18)**; the shuffle weight capped at its class v4 value on the energy side | lane D's coverage run (4,900 drawn eras on three boxes, 11:3x UK, measured): at the lossy corner of B = 4 (or, mul and mulhi all at +4, 30 of 75 lossy against 18) r is 0.956 (the shipped 0.681) and 8 of 663 eras exhaust the 256-attempt cap (mean attempt 18.6, max 252), so 1.2 percent of that corner's epochs would take the last-resort program, which fails rule (a) in 9 percent of seeds (adv-accept-3); the random stratum with every weight drawn reads r 0.718, max attempt 107, 0 exhaustions in 1,444. The energy side (15.1a): shfl 55.8 pJ per op, mulhi 39.6, prmt 22.3, lop3 24.1, mul 13.9, arx 11.3; a shuffle-heavy table raises the premium per instruction up to 2x (modelled), a multiply-heavy one 1.1x | a chip that specialised its lane ratio loses the ratio; a general core nothing | the premium per instruction moves with the mix within the capped band (the cost rows: at N = 100,000 the 5090's +146 W stock becomes up to +160 W multiply-heavy; shuffle-heavy excluded by the cap); the 5070 Ti scaled by 83/146, the M5 Max by 16/146; the two packs' measured rows replace this when they land. **First row MEASURED (kit b, PC 1, 12:5x UK): the shuffle-heavy table (shfl 13 of 48 against the stock 5) on class mx8 without a shadow block reads 137.65 MH/s at 312.9 W unlocked (7 W under the stock table's 320.2) and 127.33 at 212.4 at the lock (level): the multiply-heavy table (mulhi 11, mad 8, mul 5 of 48 against the stock 9, 4, 1; shfl 1 against 5; kit c, 13:03 to 13:06 UK) reads 137.77 at 313.5 W unlocked and 128.42 at 211.8 W at the lock, so the three tables sit within 7 W unlocked and 0.6 W at the lock: on the 48-op base program the weight move is a 2 percent term whichever way it goes; the row layer 1 needs is the same two tables inside the shadow block (mx8+sh256x27, kit d, about 13:45 UK), and the microbench arithmetic stays the default until it lands** | +0 ms | the known-failed test: the 8 of 663 exhaustions at the uncapped lossy corner (61 of 5,000 on the 17:00 cut), which the capped band must read as 0: **MEASURED 0 of 3,000 under the band (lane D's 17:00 cut, 5.1b; r 0.595, mean attempt 1.47, max 59)** |
|
||||
| Read width `W` | **pinned at 4 words (16 bytes) at genesis, not drawn** (floor lane 3, 10.3: the width is the only wire lever on the SRAM die, 66x at 4 bytes against 44x at 16 at zero shadow; 8 words after the owed PC 1 and Mac rows; 16 never) | w4 and w16 measured 5 October: the 5090 139.8 against 136.1 MH/s, the 9070 XT 17.90 against 18.15, the M5 Max within 1 percent; w64 bandwidth-bound (71.9 MH/s on the 5090) | the SRAM die's energy per read rises with the bits moved (0.25 to 0.38 nJ); the DRAM chip's toward the card's | 0 / 0 / 0 (measured) | 0 | the W = 8 rows (owed) |
|
||||
| Shadow block shape | 64 to 256 instructions per block, the pass count the ladder's | 64-instruction blocks ran 2.5 to 3.5 percent FASTER than 256 on the 5090 and the M5 Max (measured 6 October); 1,024 cost the M5 Max 17 percent | nothing for any chip (the work is the same) | +2.5 to 0 percent / pending / +2.5 to 0 | 0 | none |
|
||||
| **The index fold (a ring-A design rule of layer 1, every era's draw passes through it): `load_index` folds a product's low bits before the stride rotation, so no era's R lands a biased product bit on an address bit** | every era; lane D's coverage (11:3x UK): the class is at HALF the family's epochs, not a corner: 52 to 58 percent of accepted programs in every stratum carry one site whose address bit R (or R+1, R+2) is biased over 6 sigma at 2^20, 33 to 40 percent over 100 sigma, the worst z 1,024, on programs (c''') passes; the F8 tail's bucket excess is the same mechanism at scale (+94 to +128 sigma at the biased bit) | lane D's family harness (build-1, 10:21 UTC, 16 drawn eras, every layer-1 parameter from the era's stream, the chain draw through the real rule, reads at the rule's own 2^20 sample): the index-bit bias fires HARD on 7 of 16 eras, abs z 130 to 511 at one site, every one at address bit R or R+1 (a product's bit 0 at P = 0.25 on era 15, R = 25, z -511; a product's bit 1 at 3/8 on era 13; an or-shaped source at 5/8 on era 6); the other 9 clean under abs z 3.8; the (c''') ratio sees none of it (0.9954 to 1.0000): adv-cache-2's era-stride class measured on class v5 accepted programs at the acceptance's own sample | a chip holding the favoured half of that site's window serves 75 percent of its reads instead of 50: about 1.6 percent of a hash's reads at f = 1/2 for one site, zero at f = 1 (the partial store already costs 1.26x the ops, chip-model 5.4): the f = 1 verdict does not move; the row is an auditor's flag on "uniform random reads", not a chip lever | nothing: the fold is one xor-rotate on the address path, measured as 0 on every card by the era-layout rows (the index form is the era draw's own) | 0 | the fold's form in `load_index` (fold the product's low bits before the rotation) and its vectors; a 6-sigma REFUSAL in layer 4 is not the lever: it would redraw about 40 percent of epochs (7 of 16 eras); the known-failed test is lane D's 7 of 16 eras at the rule's sample, which must read 0 of 16 with the fold |
|
||||
| Program length N | not drawn: the ladder's signal (latency-ladder.md) | an unconditional draw retires the Apple tier at 200,000 (measured -10 percent) | the core sized for the ladder's admissible top (rung 2, 199,600 ops) | the ladder's rows | the ladder's rows | none |
|
||||
|
||||
## 3. Layer 2: the dataset's size tracks the chain state with a floor, so fixed-memory silicon ages out
|
||||
|
||||
### 3.1 The rule
|
||||
|
||||
Class v5 (`docs/design/class-v5-stored-state.md`) already keys every item to a leaf of the chain's execution state and caps the leaf count at the dataset size (2^24 items at 1 GiB, 2^25 at the designed 2 GiB), sampling the state when it is larger. Layer 2 turns the cap into the size: the day's item count is the larger of the 1.13.3 schedule (2 GiB plus 0.5 GiB a year, the floor) and the state's record count times 64 bytes, rounded up to the power of two the index mapping needs (or the multiply-shift mapping of 1.13.3 option (a), which takes any size), with a genesis-fixed ceiling at the schedule's power-of-two step for the year (2 GiB to year 4, 4 to year 12, 8 to year 28, the cache doubling with it) so the state can only bring a step forward, never add one, the verifier's lazy derivation stays bounded, and every honest tier's lifetime is the card-lifetime table's at worst (lane A's finding: the ceiling per tier per year is the number to fix before the rate). Everything below the floor is today's design; everything above it is the chain's own growth deciding the memory a miner must hold.
|
||||
|
||||
| State size (records) | Leaves | Dataset under layer 2 | Who holds it | Label |
|
||||
|---|---|---|---|---|
|
||||
| 93 (the devnet today) | 6 KB | the floor: 2 GiB at genesis | every card from 8 GB up | measured state, designed floor |
|
||||
| 10 M (a used chain's accounts alone, class-v5 section 3) | 640 MB | the floor still: 2 GiB (the leaves are a fold into the items, not the items) | the same | arithmetic |
|
||||
| 61 M (10 M accounts, 50 M slots, 1 M code chunks) | 3.9 GB | 4 GiB (the next power of two above the leaves), the year-4 step reached by state instead of by the calendar | 8 GB cards at 75 percent of memory hold 6 GB: the dataset plus the cache (512 MiB) and the scratch fits; 4 GB cards are out | arithmetic on the card-lifetime table |
|
||||
| 250 M (an Ethereum-class state, approximate) | 16 GB | 16 GiB | 24 and 32 GB cards and Apple 64 GB; the 8, 12 and 16 GB tiers out | approximate |
|
||||
|
||||
### 3.2 The chip rows
|
||||
|
||||
The time-memory curve of chip-model-v3 section 5.4 is the whole argument: a stored item costs the chip 1.2 to 2.0 nJ (one random read) and a recomputed item 6.3 nJ and 9,360 ops, so every memory system's cheapest point is `f = 1` and a chip whose DRAM is smaller than the dataset falls back along the curve for the overflow. With the microbench's measured card side beside it:
|
||||
|
||||
| Chip memory | Holds up to | At a 16 GiB dataset | Energy per hash against the 5090's 8.7 nJ per dependent read at the lock (measured, whole card) | Label |
|
||||
|---|---|---|---|---|
|
||||
| GDDR7, 16 devices of 2 GB (the 5090's board without the GPU) | 32 GB | fits | 128 reads x 2.0 nJ = 0.26 microjoules plus static: unchanged | modelled |
|
||||
| One HBM3 stack, 24 GB | 24 GB | fits | 0.32 microjoules: unchanged | modelled |
|
||||
| A 16 GB chip (a cheaper board: 8 devices) | 16 GB | the dataset equals its memory; at 32 GiB it recomputes half: 64 x 2.0 + 64 x 6.3 nJ = 0.53 microjoules against 0.26, the rate halved by the recompute, the chip's capex per MH/s doubled | modelled |
|
||||
| A fixed SRAM mirror (the `f = 0` chip's 256 MiB cache) | the cache only | the cache doubles with the dataset (option C): the 128 mm^2 mirror becomes 255 at year 4, 510 at year 12 (chip-model section 2), at a used chain's state the same by layer 2 in the first year | modelled |
|
||||
|
||||
Reading: layer 2 does not move the chip anyone builds, because a chip buys DRAM a card cannot (a 24 GB HBM3 stack at about USD 200, modelled; a 32 GB GDDR7 board at USD 320) and the card tiers are smaller than the chip's memory at every step. What it does is honest and worth saying plainly: **it retires home cards before chips.** On the card-lifetime table (option (b) steps) the 4 GB tier ends at the 4 GiB step, the 8 GB tier at the 8 GiB step, the 12 and 16 GB tiers at 16 GiB; under layer 2 those steps arrive when the state does, not when the calendar does, so a chain that takes off is a chain whose 8 GB miners leave in its first years. The `f = 0` recompute chip, which nobody builds, is the one chip layer 2 kills outright (its mirror doubles with the state). The floor keeps the schedule's lower bound; a ceiling (the next cache doubling) keeps the verifier's bound.
|
||||
|
||||
### 3.3 Per tier
|
||||
|
||||
The hash lane's VRAM rows (12:0x UK, modelled from the measured 0.4 GiB working set plus about 0.5 GiB of driver and app): the dataset needs 3.2, 5.4 and 9.9 GiB of device memory at the floor, 2x and 4x; a 12 GB card falls off at about 9.5 GiB (year 15 on the 1.13.3 schedule), a 16 GB GPU at about 13.5 GiB (year 23), a 16 GB unified Mac at about 8 GiB (year 12), the 5090 at about 29 GiB (year 54). **The DRAM-read cost per hash on the NVIDIA cards is NOT size-independent at the knee, MEASURED (the hash lane's kit b, PC 1, 12:49 to 12:58 UK, the pinned class v3 program 73bcbfe8 at 2, 4 and 8 GiB against the 1 GiB control 137.65 MH/s at 312.2 W unlocked and 127.39 at 212.6 W at the 1,300 MHz lock; 250 batches per row, fingerprints PASS): unlocked 133.86 at 315.7 W (-2.8 percent), 132.42 at 317.2 (-3.8), 131.75 at 319.2 (-4.3); at the lock 121.14 at 210.3 W (-4.9 percent), 113.06 at 204.5 (-11.2), 109.39 at 201.6 (-14.1); MH/W at the lock 0.599, 0.576, 0.553, 0.543, which is 4 / 8 / 10 percent more energy per hash at 2 / 4 / 8 GiB.** The lane's earlier reading (2 MiB pages keep the TLB's reach past 8 GiB) holds unlocked, where the card hides most of the page-walk term in its slack; the latency-bound regime at the lock exposes it. So each step of the schedule costs a tuned 5090 about 4 to 5 percent per hash while the chip's joules do not move (floor lane 3: its ticket goes USD 1,500 / 2,500 / 3,000 at 5.5 / 8.5 / 11.5 GiB), and every chip edge against a card at its knee rises by 4 to 11 percent across the schedule; the honest sentence for the schedule decision is USD 1,000 of chip ticket per step for about 1 to 5 percent of the tuned 5090's energy and about a quarter of today's measured cards by count. **On the M5 Max it is not size-independent, measured by this lane at 10:40 UTC under the Mac's measure lock (Metal packbench, the hash lane's class v3 packs at 2^28 to 2^31 words, the same seed and era, 3 batches of 2^24, vectors 3 of 3 and fingerprints per pack): 26.48 MH/s at 1 GiB (footprint 1,664 MiB, the build 32 ms), 23.26 at 2 GiB (-12.2 percent; 2,688 MiB; 54 ms), 21.31 at 4 GiB (-19.5 percent; 4,736 MiB; 94 ms), 20.61 at 8 GiB (-22.2 percent; 8,832 MiB; 193 ms).** The Apple GPU's dependent random read costs more time as the working set grows past its page reach (approximate reading: a TLB-reach effect on unified LPDDR5X; the power channels were not sampled this run, so the joules per hash move by at least the rate's share), which is a real per-tier cost of layer 2 that the NVIDIA model does not show: at an 8 GiB floor the Apple tier mines 22 percent slower per card than at 1 GiB, before any memory limit. The table below carries it.
|
||||
|
||||
| Tier | At the floor (today to year 4) | At a 4 GiB state-driven step | At 16 GiB | Label |
|
||||
|---|---|---|---|---|
|
||||
| RTX 5090, 32 GB | the daily build 13.4 ms at 1 GiB (measured), about 27 ms at 2 GiB; nothing else moves | about 54 ms; nothing else | about 220 ms a day; fits | measured at 1 GiB, scaled |
|
||||
| RTX 5070 Ti, 16 GB | fits | fits (the card-lifetime 16 GB row: room to the 16 GiB step) | OUT (the dataset equals the card) | the table, approximate |
|
||||
| Apple M5 Max, 36 to 128 GB unified | the build 32 ms at 1 GiB, 54 at 2 (measured); the rate -12 percent at 2 GiB (measured) | fits; the rate -19.5 percent at 4 GiB (measured) | fits on 64 GB and up at the 50 percent share; the rate -22 percent at 8 GiB (measured); a 36 GB machine is out at 16 GiB | measured 10:40 UTC |
|
||||
| 8 and 12 GB cards | fit | 8 GB fits at 75 percent (6 GB usable against 4.5 GB of working set), 12 GB fits | OUT | the table |
|
||||
| A pool user | the pool ships the leaves, 64 B per record (16.5 KB/s to 10,000 members today; the WAN line at about 700,000 records, class-v5 2a.2) | the same | a 16 GB leaf array per member per window: above the WAN line; the pool serves the built dataset or the member builds from the state it holds | class-v5's arithmetic |
|
||||
| A node | one leaf pass per window (100 ns per record: 6 s on one core at 61 M records, 0.2 s on 32 threads) | the same | the same | class-v5's measured rate |
|
||||
|
||||
### 3.4 The dataset-floor schedule, priced (lane A, `docs/analysis/class-v6/history.md` section 4.2a, master 59963461 (branch commit 84b30841), 10:45 UTC; the method card-lifetime-2026-10-05.md's room per tier, 75 percent of card memory usable and 50 percent of Apple unified, the non-dataset working set 254 to 479 MiB best and 600 to 1,500 worst; the population the bench table's 32 measured consumer cards by count, the fleet's hashrate-weighted census owed; the arithmetic script-checked)
|
||||
|
||||
Main's candidate steps checked first under the standing budget rule: **6 GiB does NOT fit the 8 GB tier** (6,398 to 6,744 MiB, 78 to 82 percent of the card; it fits only at a headless-rig reading of about 85 percent); 10 GiB retires the 10, 11 and 12 GB tiers AND the Apple 16 GB laptop at year 2 (the 12 GB card at 86 to 89 percent); 14 GiB retires the 16 GB tier at year 4 (89 to 92 percent), leaving 24 GB and above. That is not the tier order the note described, so the table carries the schedule that drops the tiers in that order:
|
||||
|
||||
| Step | Dataset floor | What falls (share of today's measured consumer cards, by count) | What holds | Label |
|
||||
|---|---|---|---|---|
|
||||
| the class v6 epoch | 5.5 GiB | the 6 GB RTX 2060 (3 percent) | the 8 GB tier at 72 / 76 percent of the card (the worst case one point over the rule) | modelled on the measured working sets |
|
||||
| two years on | 8 GiB | the 8 GB tier (22 percent: RTX 3070, 3070 Ti, 3060 Ti, 4060, 4060 Ti 8 GB, 5060, 2070 Super), with it the 10 GB RTX 3080, the Apple 16 GB laptop, and in practice the 11 GB 1080 Ti and 2080 Ti at 100 percent of usable | the 12 GB tier at 69 / 72 percent | modelled |
|
||||
| four years on | 11 GiB | the 12 GB tier (22 percent: 4070, 4070 Ti, 4070 Super, 5070, 3080 Ti, 3060, Arc B580) | the 16 GB tier at 70 / 73 percent; Apple 32 GB at 71 percent of its half | modelled |
|
||||
| the fixed schedule's 16 GiB step, whenever the state brings it forward | 16 GiB | the 16 GB tier (28 percent) | 24 GB and above (16 percent of the measured consumer cards plus every datacentre part) through every step | modelled |
|
||||
|
||||
A step every two years retires about a quarter of today's measured consumer cards each time; on the Apple side every Mac pays the measured rate cost of the larger working set on top (section 3.3: -12 percent at 2 GiB, -20 at 4, -22 at 8 on the M5 Max).
|
||||
|
||||
Against the chips, with the number: (1) a hybrid-bonded or soldered chip sized at launch (the Jasminer X4 shape, the E3 shape) dies at the first step it cannot carry, as the E3 did 20 to 27 months after shipping; a chip sized to the floor with 8 GB dies at the year-2 step on either schedule, one with 16 GB holds through year 4; but the schedule is a consensus field read at genesis, so a maker sizes to the step it wants to survive (USD 160 more of GDDR7 on a USD 470 part; about 3x the 2021 memory die area for the Jasminer shape, approximate), and the E3 died because it was sized to the card fleet's limit, not to a schedule; (2) the f = 1 GDDR7 chip's 32 GB board pays USD 0 through 16 GiB and keeps 5.1x per joule and USD 2.8 per MH/s at every step; (3) the SRAM store pays capex only, at the cache doubling (USD 46 per die at 256 MiB to 111 at 512 MiB, 306 at 1 GiB for the recompute chip's mirror; lane B's full store a die per 2 GiB: USD 400 to 600 each). **The sentence for the close: the schedule is a fleet-retirement rule with a chip tax of USD 0 to 160 per unit on the chips that matter, and it kills only a chip whose maker ignores a public consensus field; if adopted, the 75 percent rule and a 6 GiB floor cannot both stand for the 8 GB tier (5.5 GiB keeps it inside the rule in the best case; about 85 percent is what makes 6 GiB fit).**
|
||||
|
||||
## 4. Layer 3: scheduled family epochs by height, every 180 days by default, no release
|
||||
|
||||
### 4.1 The rule
|
||||
|
||||
Spec 1.13.2 already unlocks reserve family `n` at the start of era `n` with weight `W_new` taken proportionally from the live families; the eras are the six-month `E_n` of the 1-hour VDF. Layer 3 fixes two things the reserve leaves open: the cadence is a genesis constant (`family_epoch_daa`, 180 days of DAA seconds by default, independent of the era draw so a family can rotate without an era), and the set rotates after the reserve is exhausted (at family epoch `n` past the reserve's end, the live set is the genesis eleven plus the reserve entries whose index is in a fixed cycle over the reserve, so a retired family returns on a fixed schedule and a chip can never wait one out). The weights move by the 1.13.2 rule (proportional). No release carries any of it; the emitter and the verifier hold every family from genesis, with the per-vendor conformance vectors of 1.15 for each.
|
||||
|
||||
### 4.2 The chip rows (the reserve's measured and modelled figures, algorithm.md 5.2 and counter-asic-3-reserve.md)
|
||||
|
||||
| Family | What a chip must add (N5, approximate) | Energy per op at the N5 floor | The honest cards' step cost (measured, ratio to the add-xor-rotate step): Apple / NVIDIA / AMD | What it does to a chip without it on the unlock day |
|
||||
|---|---|---|---|---|
|
||||
| R1 shfla (lane + delta) | a 32-lane crossbar per warp, 3 to 6 adders per lane | 1.00 pJ | 1.91 / 1.53 / 0.75 to 0.84 | at 4 points of 79 the chip without a crossbar emulates through its existing xor-shuffle path or loses 5 percent of the mix's work per op |
|
||||
| R2 perm (byte permute) | a 4x4 byte crossbar, 1 to 2 adders | 0.10 | 1.13 emulated / 1.30 / 1.73 to 1.93 emulated | under 1 percent |
|
||||
| R3 popc and clz | a popcount tree and a priority encoder, 1.5 to 3 adders | 0.10 | 0.87 and 1.01 / 1.50 and 1.63 / 0.91 to 1.30 | under 1 percent |
|
||||
| R4 bfe, R5 shl and shr, R6 sel, R7 andn | 0.1 to 0.3 adders each | 0.02 to 0.06 | 0.75 to 1.54 | 0 |
|
||||
| R8 mm8 (the int8 tile) | a u8 MAC tile per warp, about 100 adders per lane, licensable | 1.60 (the N5 floor); the GPU's own tile measured at 1.5 to 4 pJ per MAC (the research file's 15.1a) | Apple emulated at 1.6x per dot4 (10 steps per tile), NVIDIA 2.43, AMD 1.68 to 1.83 native (layout unverified) | nothing: a chip's MAC array is cheaper than the GPU's (k 0.03 to 0.3, measured GPU side); R8 is kept for datapath diversity, never for joules |
|
||||
| All eight pre-wired | about 8 to 14 adders per lane, about USD 4 of N5 on a 14,000-lane array | | | the GPU-like chip pays USD 4 once and is never surprised |
|
||||
|
||||
Reading: against the `f = 1` chip every family moves the per-joule edge under 10 percent (a family changes 4 points of 79 in the shadow mix at `N x k x 6.4 to 11.3 pJ`), and a chip that pre-wires the reserve pays USD 4; the layer's whole value is against a chip taped out without a family (a fixed-function datapath), which loses the family's share of the work on the unlock day or emulates it at the measured vendor penalties. The rotation after exhaustion closes the one gap the reserve has: a chip that waits for a family to retire.
|
||||
|
||||
### 4.3 Per tier and the known-failed case
|
||||
|
||||
The re-tune per family epoch (the hash lane's rows): a card with a lever (the 5090, the 5070 Ti, the 4070) re-runs Ember's search, 11 to 12 minutes at stock watts for about 15 s of hashing lost per 180 days (0.0001 percent; the 5090's 12 minutes at +235 W is 0.05 kWh, about 1 p); a card without one (the 9070 XT in 0.3.20, every Apple machine) runs the 3-minute baseline and loses nothing.
|
||||
|
||||
The cost of a live family is its step cost at its weight, checked per vendor at the unlock rehearsal (the 5 percent rule of 1.13.2): at 4 points the worst measured vendor (Apple on shfla, 1.91x per op) pays under 1 percent of its ALU time, which on a latency-bound card is 0 rate; the daily build and the verifier are unmoved (one op per instruction, under 0.01 ms per warp). The 5090 and the 5070 Ti pay the NVIDIA column (1.26 to 1.63x per op on a family's 4 points: 0 rate); the M5 Max pays the Apple column (0 rate at 4 points; the mm8 emulation at 1.6x per dot4 stays under the 8x bound). The known-failed case: a kernel built without the live family refuses at packcheck (the program id carries the family set through the class's allowed list, as `program_id_class` carries the era's `allowed[3]`), and the fast-time harness crosses one family epoch with three nodes and one stale miner, the stale miner's blocks rejected from the first block of the new epoch (the class signal harness's shape, `infra/fast-time/class-v5-signal.mjs`).
|
||||
|
||||
## 5. Layer 4: the acceptance rule and the census as a function of the era draw
|
||||
|
||||
### 5.1 The rule
|
||||
|
||||
Every test the generator applies to a candidate program is today a function of the program and the genesis parameters: (a) the stale-source rule, (b) the injecting-write rule, (c) the dynamic test over 64 units on the closed-form stand-in (constant bits, saturation, lane-constant sites, output bias, the distinct-address sum), sub-version 3's (a') freshness fixpoint, (c') saturation per site and (c''') the distinct ratio floor (0.995). Layer 4 makes the rule take the era draw as an input (the mixer multiplier, the weights, the width and the block shape of layer 1; the live family set of layer 3; the dataset size of layer 2) and adds the four tests the night's work named, so that layers 1 and 3 need no per-era cryptanalysis: every era's programs are drawn against the same tests, and a candidate that fails any is redrawn from the next stream values, exactly as today's attempts are.
|
||||
|
||||
| Test | Today | Under layer 4 | The number it rests on | The known-failed case |
|
||||
|---|---|---|---|---|
|
||||
| (c''') the distinct-item ratio floor | 0.995 over the 64 units at the genesis width and mixer | the same floor evaluated with the era's width (a 16-byte load touches one item too) and dataset size; 2.435 percent of candidates under it today, attempts +3.4 percent | class-v5 section 14 (measured census of 4,600 candidates) | a candidate below 0.995 on the exemplar seed 100767 is refused |
|
||||
| F8's uniformity (the largest 64-line bucket within 6 sigma over 2^28 derivations; the top 0.1 percent of items within 1.2x of the window model over 2^24 nonces on 64 seeds) | a gate run by hand per class on the attack board | run by the census tool per era draw at genesis (the band's corners plus 64 random eras) and by the node's acceptance as the per-site version below; a draw whose corner fails is excluded from the band | f8-uniform.md section 7: +4.84 sigma against the control's +4.18, PASS; the tail p4, p8, p10, p34 attributed this morning as per-site bucket concentration at a narrow-window site | the `quarter-lines` and `const-item` plants fire at +75.97 and +92,682 sigma (f8-uniform.md 2.1) |
|
||||
| The per-site largest-bucket bound (this morning's attribution) | named for a next class | per load site, the largest 256-item bucket over the units' addresses, stated in SIGMA against its own window's Poisson expectation, never as a ratio: ratios of 2.2 to 2.4 at full-window sites are the CLEAN maximum (65,536 Poisson(16) buckets read +4.4 sigma), so a ratio bound would refuse clean programs; lane D's rebuild carries the sigma column | AP-F8-1's tail: p10 1.50x, p8 1.38x, p34 1.25x, p4 1.22x, each a narrow-window site's bucket; the sigma figures from lane D's family-gate.md (17:00 UK) | the four tail seeds must be refused on sigma; the 60 passing seeds accepted |
|
||||
| The value-level bias test (adv-cache-2; the research file's 20.2b) | named for a next class | per load site, the one-count of every index bit over the 64 units within 6 sigma of n / 2 (the per-load prototype's `BiasedIndexBit`, built last night, measured as the record); AS A TEST ONLY AFTER the index fold of layer 1 is in, because on today's `load_index` it would redraw about 40 percent of epochs (lane D: 7 of 16 drawn eras at abs z 130 to 511); with the fold in, the test is the guard that the fold holds | a product's low bits at P(bit 0) = 1/4 placed at address bit R by the stride rotation; 6 of 16 drawn eras over 1.04x, 14 of 17 eras flagged by the instrument on the pre-amendment generator | a program whose site is sourced by a product under an era with R under 28 is refused; the devnet era's R = 29 is not relied on |
|
||||
| The duplicate-lane test (the research file's 20.2a) | built for the per-load class only | kept as a per-load-only test unless a drawn block shape ever places shadow work between loads (layer 1 does not: the block shape is the size, the placement stays after instruction 63) | 1,482 duplicate lanes on the per-load record, 0 to 2 on every sound class | the per-load candidate 0 is refused |
|
||||
|
||||
### 5.1a The three rings (lane D, `docs/analysis/class-v6/family-gate.md` on master at fbe62a8b, 10:40 UTC; the design as the harness measured it)
|
||||
|
||||
| Ring | Who runs it, when | What it checks | Cost |
|
||||
|---|---|---|---|
|
||||
| A | the chain, per era, at the cut | band membership of every layer-1 draw; the day-key cost rule (AP-F4-1); the rotation class recorded (whether R leaves a product's low bits inside the index at the era's D); the lossy-share bound on the weights (`or + mul + mulhi` at most the table's 18 plus B, so r stays under 0.85 and r^256 under 1e-18); a failure consumes the era stream's next block (a redraw) up to a cap, then the base table as the last resort, so the draw is total | microseconds |
|
||||
| B | the chain, per epoch | the acceptance rule keyed on the family's shape (`is_family_shape`: the shadow block in {64, 128, 256} at 6,912 per iteration, m in {4, 8, 16}, the mix one-hot on the drawn width) in the draw's source rule and the acceptance alike; (c'') and (c''') with the expectation divided by the width (at W = 4 words the index space is a quarter, else every width-4 program is refused); the per-site largest-256-item-bucket excess in sigma and the per-site index-bit one-count in sigma on the same 2^20 pass | about 2.2 s per chosen candidate |
|
||||
| C | offline, per drawn era | the attempts census (r, parts, exhaustion), the F8-form census at 2^24 on 64 seeds against the window model at the era's D, the stand-in gap at the era's D, the exhaustion count at 10^4 to 10^5 seeds per lossy corner | box-hours |
|
||||
|
||||
The sampling bound: n drawn eras all passing bound the failing fraction at 3/n at 95 percent (190 eras for 1 in 64, 3,067 for 2^-10); the union bound over the tests' miss rates is bounded by their known-failed CASE counts (9 hot sets give the (c''') floor a miss rate under 0.33), so the proof of testing carries per test the count of fired cases, and tightening is by cases from the adversarial tails, not by eras. The record: one JSON per drawn era (the family id, the draw, the pinned binary sha, the box, the seeds, every verdict with its statistic, threshold, log sha256 and core-seconds) plus a family summary (eras per stratum, the worst era per test, the refuse rate per band, the bound's arithmetic, the corner cells not reached). The mixer m is a per-FAMILY axis (no per-era statistic sees it; the index tests run on the closed form), closed by adv-mixer-3's ladder at m = 4 (2 of 4 applications of margin against 6 of 8) and the x16 verifier row; the per-load placement has one value in the band; W = 16 is excluded. Main's word (11:4x UK): lane D's band is layer 1's op-mix band and the index fold with the bias test as its guard is layer 1's rule, not a next-class note.
|
||||
|
||||
### 5.1b The 17:00 cut, measured (lane D, `docs/analysis/class-v6/family-gate.md` section 6 on the mirror's master at 238100b0, 13:13 UK; 24,000 drawn eras of the family on build-1 and build-3, 54 core-hours, one era per seed with every layer-1 parameter from the era's own stream, the acceptance keyed on the family's shapes; census rows, logs, scripts and the harness diff under `docs/analysis/class-v6/logs/`)
|
||||
|
||||
What the cut settles, each row measured unless marked:
|
||||
|
||||
| Finding | The numbers | What it moves in this document |
|
||||
|---|---|---|
|
||||
| The op-mix band is settled by measurement | B = 4 with or, mul and mulhi free to rise exhausts the 256-attempt cap in 1.2 percent of eras (61 of 5,000 at the lossy corner); with or, mul and mulhi never raised above their base (section 2's band) r = 0.595, mean attempt 1.47, max 59, 0 exhausted in 3,000. The lossy-share curve, complete at 3,000 eras per point (13:5x UK): r = 0.80, 0.88, 0.92, 0.96 at +1 to +4 points on or, mul and mulhi; exhaustion 0, 0.07, 0.20, 1.10 percent of eras, against independent-attempt estimates of 4e-25, 3e-15, 9e-10, 1e-5, so the per-era correlation is 10^5 to 10^10 above the geometric figure and the band's edge is the measured +2, not the arithmetic's | the layer 1 op-mix row's known-failed test now reads 0 of 3,000 under the band (it owed a 0); the cap on the lossy families stays at the base, with +2 points the most the band could ever open to |
|
||||
| What breaks at the lossy corner | class v5's last-resort scan passes at its first or second candidate on every exhausted era seen, so the corner costs liveness time, not an unchecked program; the per-era exhaustion is about 1,000x the independent-attempt estimate because one era's attempts share its weight table, which is the independence the scan's 1e-300 assumes | section 5.2's bound: the attempts within an era are not independent draws; the bound is per era from the census, not r^256 |
|
||||
| The (c''') floor needs a per-width calibration | at width 4 (expectation divided by the width) it refuses 4.5 to 6.5 percent of candidates against 0.6 to 1.0 percent at width 1 | ring B's row: with W pinned at 4 at genesis (section 10.3) it is one measurement, being taken now (the harness's next build records the pre-floor spread of every candidate at width 4 under the band over 3,000 eras plus the width-1 control); the 09:00 report states either a single width-4 floor at the shipped clean-refusal rate (about 2.4 percent of candidates) or the sigma-over-expectation form, whichever keeps the known-failed hot sets refused at the lower clean cost; the default here is the sigma form |
|
||||
| The shape axis moves the draw's cost, the mixer axis is invisible | 64 x 108: 1.8 attempts; 256 x 27: 3.8, through (a')'s fixpoint over the block; r 0.714 to 0.731 across m | m is closed per family by adv-mixer-3's ladder at m = 4 and the x16 verifier row, not by drawing eras (section 5.1a's reading stands, now measured); the block-shape row of layer 1 carries the attempt cost |
|
||||
| The era-stride bias at the bit level | 48 to 58 percent of accepted programs in every stratum carry one site biased at over 6 sigma at 2^20; 73 percent of those at address bit R exactly (12 percent at R+1, 5 at R+2: the product law's bits 0, 1, 2 through `rotl(x*M, R)`); a third over 100 sigma, worst z 1,024; under R in 28..31 the over-100-sigma share falls from 37.7 to 7.5 percent; the bucket statistic sees the same mechanism (p99 +6.8 sigma on bit-clean eras, +44 on biased ones) | one value-level test covers both; the remedy is structural, the index fold of the product's low bits in `load_index` before the rotation (layer 1's rule, section 2), not a per-epoch refusal that would redraw half the epochs; the live-dataset price per site (ring C) is being read now (the F8 census at 2^24 on 64 seeds at two band points; 31 of 64 seeds PASS so far at the shape-256 point, no test fired) |
|
||||
| The bound arithmetic on the counts | 10,000 random eras and 3,000 band eras passing bound the failing fraction on the ring-B tests at 3.0e-4 and 1.0e-3 at 95 percent; the floors' miss rates on the live-dataset classes rest on 9 and 5 known-failed cases (under 0.33 and 0.60), tightened by cases from the adversarial tails, not by eras | section 5.1a's sampling bound now has its measured n; the gate-record JSON per era lands with the full report |
|
||||
|
||||
Owed from lane D by 09:00 UK tomorrow: the per-width floor, the bucket bound in sigma, the finished lossy curve, the two ring-C live rows, the gate-record JSON per era.
|
||||
|
||||
### 5.2 The cost of the redraw, and its bound
|
||||
|
||||
Lane D's measured coverage (11:3x UK; the full table in family-gate.md at 17:00): the (c''') refuse rate per stratum 2.56 percent of candidates in the random stratum, 3.41 at shape 64, 0.42 at the lossy corner; the rule-of-three coverage line a failing fraction under 3e-4 at 95 percent at 10,000 eras. The hash lane's draw costs: the base rules reject about two thirds of raw candidates per attempt (milliseconds), (c'') about 4 percent at 2.2 s per pass on a box core (2.3 s per epoch draw per node, 4 to 5 s on slower cores), (c''') 2.435 percent in the same pass, the bucket bound 1 to 3 percent (estimate) and the value-level test about the same cost again inside the same histogram; expected attempts per accepted program stay about 2.1, and a parameter set under which 32 attempts fail is redrawn.
|
||||
|
||||
Attempts per seed are the price. Today's rule under sub-version 3 accepts a candidate at about 30 attempts per seed on the devnet stream (the hash lane's census: 0 of 24,631 seeds exhausted, max attempt 29 against the 256 cap). Each added test raises the rejection rate by its own fraction; the two new tests' fractions on the pre-amendment generator were 4.9 percent (the per-site bound, the tail's four seeds of 64) and up to 14 of 17 eras on the value-level test, which is the one that needs the generator's own fix (the sub-version 3 source rule already refuses a load sourced by a product on the base program; the value-level test catches what the static rule misses). The bound the chain keeps: 256 attempts, the last-resort draw, and the census's exhaustion count per era at genesis (0 of 10^6 is the standing requirement); a band corner whose exhaustion count is not 0 of 10^6 is excluded from the band at genesis, which is the mechanism that lets layer 1 draw without per-era cryptanalysis.
|
||||
|
||||
## 6. The gate plan: the family analysed as a family
|
||||
|
||||
The attack board (F1 to F10: the shadow's compressibility, the mixer's structure, the cache's recompute, the weak day, the X9 anchor, the verifier, the era draw, uniformity, grinding, the ladder) and the in-house pass's nine lanes (adv-mixer 1 to 3, adv-cache 1 to 3, adv-accept 1 to 3) were each run against one class with its genesis parameters. Under class v6 each harness takes the era draw as an input and the census walks the band:
|
||||
|
||||
| Gate | What runs | The window | The pass line | The known-failed case |
|
||||
|---|---|---|---|---|
|
||||
| G-band (layer 1) | every F row and every adv lane at the band's corners (m in {4, 8, 16} x W in {1, 4} x the weight extremes x the block shapes 64 and 256) plus 64 random eras of the stream | the testnet period, on the pool with --class measure | every row's own line at every corner (F8 1.2x, the 6-sigma buckets, adv-mixer-3's "clean from k = 1", adv-cache-2's hot-set share under 0.1 percent) | the quarter-lines and const-item plants at every corner; a corner that fails any row is out of the band |
|
||||
| G-size (layer 2) | the hot-set and recompute lanes at 2, 4, 8 and 16 GiB with the sample rule; the verifier's lazy bound at each cache doubling | the same | the verifier under 10 ms with the sibling loaded at every size (the ladder's method); the chip curve monotone | a stale-state hasher at a larger state reads 0 of 32 lanes |
|
||||
| G-family (layer 3) | the family-live 5 percent run per vendor at each reserve entry's weight (Metal, CUDA, OpenCL on NVIDIA and AMD), the conformance vectors of 1.15 per family, one family epoch crossed on the fast-time harness with a stale miner | the same | every vendor within 5 percent with the family live; the stale miner's blocks rejected from the first block | the stale kernel refused at packcheck |
|
||||
| G-accept (layer 4) | the generalised rule against the per-load record (candidate 0 refused), the tail's four seeds (refused) and the 60 passing seeds (accepted); the exhaustion census per corner (0 of 10^6) | the same | the hand census equals the rule's verdicts seed for seed | the two plants above |
|
||||
| G-vectors | every vendor's fingerprint equal on one pack per corner (the sub-version 3 pairing method: Metal, CUDA, Apple OpenCL, AMD OpenCL) | before any object | 96 of 96 lanes and the 2^24 fingerprint per corner | a corner's pack refused by a worker on the wrong class |
|
||||
|
||||
Hours (agent, never weeks): the generator's five draws and the band constants 6 to 8; the acceptance rule parameterised plus the two new tests 6 to 8 (the per-load prototype carries the two tests' code); the census tool over the band 4; the family cadence and rotation in the fork 6 to 8 (the class signal's shape); the fast-time harnesses 4 per layer; the attack board re-run over the band is pool time, about 20 corners x the board's 3 to 6 hours each, pipelined on both boxes over the testnet period.
|
||||
|
||||
## 7. What a fully general chip still gets
|
||||
|
||||
The identity of the research file's section 2, with the night's measured rows: `edge = (E_card + F) / (E_mem + k F)`. None of the four layers enters it except through `F` (the premium the card pays) and `k` (the chip core's cost per op over the card's), and the layers move neither for a GPU-like chip:
|
||||
|
||||
| Layer | Its effect on a GPU-like chip's edge | Its effect on that chip's capex | The honest line |
|
||||
|---|---|---|---|
|
||||
| 1 | none on `k` (firmware); `F` moves with the draw and the card pays it first; the core sized for the band's top | +USD 25 to 40 of N5 per chip; the project N5-class from the first tape-out: USD 30 M, a break-even cap of about USD 100 M (the mission lane's model) | a chip that is a GPU's memory system plus a programmable core is drawn against by nothing here |
|
||||
| 2 | none while the dataset fits its DRAM (24 to 32 GB against the card tiers' 8 to 32) | none until 16 GiB; then +USD 200 per HBM3 stack | retires cards before chips |
|
||||
| 3 | under 10 percent per family; USD 4 pre-wired | +USD 4 | a datapath taped out against one family set dies; a general one does not |
|
||||
| 4 | none | none | the safety of 1 and 3, not a lever |
|
||||
| All four, against a 5090 at its knee | **3.6x at zero premium, 2.1x at `k = 1` with the class v4 shadow (measured card, modelled chip), unchanged** | about +USD 30 to 60 per chip; the N5 project forced | "useless as soon as it dropped" is true of a fixed-function ASIC and false of the chip anyone builds; what holds the general chip is the price per joule of the honest card's own operating point and the shadow's premium, as last night's close said |
|
||||
|
||||
## 7a. The hardware future beside the four layers (lane B, `docs/analysis/class-v6/hardware-future.md` on master at 34f63b3c, 10:25 UTC; every chip figure modelled on chip-model-v3's method; the first cut, the full report by 09:00 UK tomorrow)
|
||||
|
||||
The chip rows of sections 2 to 7 price the GDDR7 board and one HBM3 stack. Lane B's table carries the memory systems a chip could buy from 2027 on, per joule at zero shadow against the 5090's 2.40 microjoules (the M5 Max's 0.78 in brackets) and with the class v4 shadow at the premium F = 1.10 at `k = 0.5 / k = 1`:
|
||||
|
||||
| Memory system (when) | Energy per random read, modelled | Edge at zero shadow | With the shadow, k 0.5 / 1 | What the four layers do to it |
|
||||
|---|---|---|---|---|
|
||||
| GDDR7 board, 28 nm controller (the record) | 2.0 nJ | 5.1x (1.7x) | 3.3x / 2.1x | layer 1 forces the N5 core (the project to USD 30 M); nothing else |
|
||||
| HBM3E, one stack | 1.2 nJ | 7.5x (2.4x) | 3.8x / 2.4x | the same |
|
||||
| HBM4, one stack, 2027 to 2028 | 1.0 to 1.1 nJ | 4.5x to 11x (the activate ceiling doubles if tFAW is per channel, unmeasured) | 4.4x / 2.6x | the same |
|
||||
| A custom HBM4E base die (controller and PHY in the stack, N3P; a non-hyperscaler from about 2028) | 0.9 to 1.0 nJ | 6.5x to 14x (2.1x to 4.4x) | 4.4x / 2.6x | beats all four layers: untouched by any of them |
|
||||
| DRAM on logic (FGDRAM-class 256-byte rows, hybrid bonded), 2029 to 2031 | 0.5 to 0.7 nJ | 12x to 20x (5.2x) | 5.1x / 2.8x | beats all four layers |
|
||||
| **An SRAM full store on one N2 reticle, 2 GiB (452 mm^2 of macro, USD 400 to 600 of silicon)** | 1.0 nJ (0.5 to 2.0) | power-bound at about 2,100 MH/s per die at 300 W: 17x (8x to 30x; 5.6x against the M5 Max), USD 0.25 to 0.4 per MH/s | 4.8x / 2.7x; at the 5090's whole latency shadow (F about 2.0) 3.7x / 2.0x | beats layers 1, 3 and 4; **layer 2 moves its capex, not its joules** (two dies at 4 GiB 15x and USD 1,000; four at 8 GiB 13x and USD 2,000 to 2,500) |
|
||||
| LPDDR6 controller chip | 1.5 to 2.0 nJ | 4x to 5x (1.3x to 1.6x) | | the honest SoC tier's own memory |
|
||||
| Per-bank PIM, UPMEM, an FPGA with HBM2e, wafer-scale, CXL, optical | | under 1x or no path (PIM is blind: 1.6 percent of reads in-bank at 2 GiB on a 32 MB bank, 0.4 at 8 GiB) | | |
|
||||
|
||||
What this changes in the design, taken into the layers:
|
||||
|
||||
1. **Layer 1's program-length band is sized against the N2 SRAM chip, not the GDDR7 board.** The lower bound is the length that holds the record's 2.1x today (rung 0, 102,100 ops); the upper bound is the honest cards' full latency shadow (the 5090 about 330,000 ops unlocked and 150,000 at the lock, the M5 Max 290,000 by its budget and 130,000 by the 5 percent rule, the 9070 XT 650,000, all measured or budgeted in the ladder's rows), re-based at each family epoch; N still moves by the ladder's signal inside that band, never by an unconditional draw. The public "2x" is not reachable against the SRAM chip at any `k` under 1 (2.0x needs the 5090's whole shadow at `k = 1`), which is the honest line section 0 now carries.
|
||||
2. **Layer 2's floor is a card-lifetime decision, not a chip lever**: the SRAM chip pays capex, not joules, for a larger dataset (USD 400 to 600 per 2 GiB die), and the real processing-near-memory threat is the custom base die, which no layer touches; the brake until about 2028 is HBM allocation and price (claimed: Samsung asking USD 4 to 5 per Gbit for HBM4 against 1.5 for HBM3E, 2 October 2026). The spec's schedule stands; the dataset stays inside 16 GB unified memory (8 GiB at the top of the schedule does).
|
||||
3. **The denominator: resistance is stated against the best honest joule** (the M5 Max at 0.78 microjoules, the 5090 at its knee at 1.67), where every chip edge is 2x to 3x smaller than against the 5090's stock point; section 0's per-tier line already reads so.
|
||||
|
||||
## 7b. The history beside the four layers (lane A, `docs/analysis/class-v6/history.md` on master at 4c58ad65, 10:26 UTC; the first cut, the full report by 09:00 UK tomorrow)
|
||||
|
||||
Lane A's verdicts, taken into the layers where they bind:
|
||||
|
||||
| Finding | What it does to this document |
|
||||
|---|---|
|
||||
| The chip that stores the dataset (every Ethash chip: E3 1.0x, Linzhi 2.1x, Jasminer X4 5.1x by DRAM hybrid-bonded onto a 40 nm logic die, E9 Pro 4.1x) is closed by none of the four layers; every per-era draw and every family epoch is firmware to it; the number v6 inherits is 5.1x on GDDR7 at zero premium, 3.6x at the 5090's knee, 2.1x with the class v4 shadow at `k = 1` | section 0's frame and section 7's table say the same; the Jasminer X4 is the measured precedent for the hybrid-bonded row of 7a |
|
||||
| Layer 2 as declared is not an anti-chip rate: at 2 GiB plus 0.5 GiB a year a 32 GB chip board lasts 60 years and the 8 GB card is out at year 12; the Ethash E3 is the only chip a growth rule ever killed (DDR3 at the fleet's own 4 GB limit, 20 to 27 months after shipping); the value of layer 2 is its floor (above every SRAM die) and a CEILING per honest tier; if the state grows as Ethereum's did the dataset outgrows every card in a decade (approximate), so the ceiling in GB per tier per year is the number to fix before the rate | layer 2's rule gains the ceiling as its first constant: the dataset never exceeds the schedule's power-of-two step for the year (2 GiB to year 4, 4 to year 12, 8 to year 28), whatever the state does, so the state can only bring a step forward, never add one; section 3.2's reading ("retires cards before chips") stands and is now the reason the ceiling exists |
|
||||
| Layer 3 closes the fork (Grin's tweaks worked only as hard forks on a lane built to die; Monero's chips were back inside 4 months of v8; X16R's drawn order cost the FPGA nothing) and taxes a sequencer chip only die area (Rao: ProgPoW's whole inner loop about 1 M gates, 0.025 mm^2 at 10 nm), because a reserve readable at genesis is built in from day one | section 4.2's "USD 4 pre-wired" row is the same fact with the history's gate count beside it; the random item-derivation program (R0) is the one reserve item unknowable at genesis and keeps its place at the head of the order |
|
||||
| Layer 4 is the one layer acting on a mechanism that beat hashes after launch (Kik's 64-bit seed, Dinur-Nadler on MTP, AP-F8-1); the window-model null changes with any draw of read width, program length or windows, so the census re-derives per era (2.2 s per candidate), and the shadow-written residue (about 1.0004x) needs its own ceiling per shadow placement | section 5.1's F8 row gains the per-era null: the census tool derives the window model from the era's own draws before judging; the residue ceiling is a constant of the block shape (64 and 256 are the two shapes, each with its own) |
|
||||
| Corrections to the 5 October history: the Antminer X9 withdrawn in May 2026 with zero units shipped; RandomX v2 shipped 25 March 2026 (program 256 to 384, CFROUND 16x rarer, AES in the loop, prefetch 2 ahead, +52.9 percent work; activation pending in Monero PR 10038); Vorick's "survives forks at under 5x" chip unverified | this document quotes none of the three; the research file's section 2 already retired the X9 as a `k` figure |
|
||||
| The ranked list: layer 2's floor and ceiling per tier first; the clock and the detector (two instruments: nonce pattern for a chip, share-by-key-and-template for a concentrated fleet, which is what found Qubic until it randomised); keep the read width out of the draw unless a width other than 4 B is measured latency-bound on every vendor; bound the op-mix draw by the per-vendor energy table; order the reserve by what a sequencer cannot fold into firmware, mm8 last, R0 the one unknowable item; a null per drawn parameter for layer 4; the per-load placement as research, not a draw value | taken as the order of section 6's gates; the read width: w16 is measured within 2.7 percent on the 5090 and the 9070 XT and within 1 percent on the M5 Max (latency-bound on all three), so {1, 4} words stays in the band with that measurement as its warrant; the per-load placement is already out of the draw (section 2) and dead as a construction (the research file 20.2a-close) |
|
||||
|
||||
## 9. The served chip line reviewed against the measured and modelled rows (for main's word at the 15:45 UK close; nothing served moves on this document)
|
||||
|
||||
The served sentence (`docs/plans/counter-asic-3-public-text-2026-10-07.md`, the homepage paragraph and row 17): "At launch the strongest chip in our public model reaches 2.1x per joule against an RTX 5090 with a core as good as a GPU lane, 3.4x with one three times better, under class v4 from the first block; a 5090 locked at its knee pays 82 W for that shadow work. Class v5 then makes the dataset the chain's own state, so a chip that stores it or recomputes it is wrong on every item. Without class v4 the same chip would reach 5x to 9x."
|
||||
|
||||
| Clause | Its basis today | Measured, modelled or claimed | Stands, or what it needs |
|
||||
|---|---|---|---|
|
||||
| "the strongest chip in our public model" | the GDDR7 stored-dataset chip of chip-model-v3 5.5, the 2027-on chips of 5.12 now in the model | the SRAM store (17x at zero shadow, 2.7x to 4.8x with it) and the custom base die (6.5x to 14x) are modelled in the same file since today | the clause is no longer true of the public model as written: the strongest chip in it is the SRAM store, five years out, with the project cost and the clock beside it |
|
||||
| "2.1x per joule against an RTX 5090 with a core as good as a GPU lane" | the class v4 shadow at the 5090's knee: 82.8 to 90.5 W premium measured four times, 6.4 to 6.6 pJ per counted op; the chip's memory 0.466 microjoules modelled | measured card, modelled chip | stands for the GDDR7 chip; against the SRAM store the same clause reads 2.7x at `k = 1` and 2.0x only at the card's whole latency shadow |
|
||||
| "3.4x with one three times better" | `k` 0.33 for an ALU-shaped core: an estimate from datapath and wire figures (approximate); the microbench bounds the GPU side (11.3 pJ per ARX op stock, 6.2 at the knee) and the re-weight's hold stands. **The k lane's first RTL rows (10.2, 13:3x UK): k 0.18 at the lock at N3 (0.35 on unscaled ASAP7), every chip figure a floor; the GDDR7 chip with the shadow 4.0x at the knee, 5.8x at stock, so "three times better" is the unscaled-ASAP7 unit case and a bare unit at N3 is five to six times better; the sequencer core a rotating family forces sits between, its row owed as the headline; the served figure must survive 4.0x at the knee as the worst case** | modelled; the chip side never measured | stands as the pessimistic column for the GDDR7 chip; the SRAM store's pessimistic column is 4.8x (`k` 0.5 on an N2 core, lane B) |
|
||||
| "a 5090 locked at its knee pays 82 W for that shadow work" | the efficiency pass: 81.8 W at the best points, 82.8 on the packs job, 90.5 on the sparse job's base rows | measured | stands |
|
||||
| "Class v5 then makes the dataset the chain's own state, so a chip that stores it or recomputes it is wrong on every item" | class-v5-stored-state.md: a stateless or stale chip is wrong on every item; a chip that holds the state (one node per farm, the leaves at 16.5 KB/s) is not | designed, measured on the harness | the clause overstates: "a chip that does not follow the chain is wrong on every item" is the true form; a chip that stores the dataset and follows the chain is unmoved (chip-model 5.10) |
|
||||
| "Without class v4 the same chip would reach 5x to 9x" | chip-model 5.4: 5.1x GDDR7 to 9.2x eight HBM3 stacks at zero shadow | modelled | stands for the DRAM chips; the SRAM store reads 17x |
|
||||
| The served "random reads" sentences (uniform random reads over the dataset) | lane D and adv-cache-2: the reads spread over the whole dataset; a bit-level bias at one site in about half of epochs; 1.6 percent of a hash's reads to a half store, nothing to a full store; the next class folds it out | measured | CORRECTED today by the audit lane on those figures (main's word, 11:4x UK); not pending |
|
||||
| The convention behind every "x per joule with the shadow" figure | the record's convention (chip-model-v3 and the served text): `k` is the chip core's op cost as a fraction of the GPU's at the same operating point, so when the 5090 locks to 6.2 pJ the chip's core is priced at 6.2 k pJ too; the absolute convention: the chip's core costs its own pJ per op (the k lane's RTL figure), the GPU's its measured pJ at its point | measured GPU side; the chip side claimed until the k lane's RTL rows | the record's convention flatters every chip row at the lock by up to 1.6x (floor lane 3, 13:06 UK: the SRAM die 6.1x against 3.8x at k 0.5, 3.3x against 2.0x at k 1); the served line's figures are at the knee, so they are in the record's convention and need re-reading in the absolute one before main's word; the close carries both, absolute first |
|
||||
| The schedule's cost beside the SRAM store's edge (main's order) | section 3: the floor and ceiling per tier; main's candidate schedule (6, 10, 14 GiB) priced by lane A by 17:00; the Apple rate cost measured today (-12 / -20 / -22 percent at 2 / 4 / 8 GiB) | measured and modelled | lands at 17:00 |
|
||||
|
||||
The wording this lane proposes for main's word, if the SRAM-store reading stands at the 15:45 UK close (the synthesis on the mirror's master by 16:30; the decision itself stays the founder's): "At launch the strongest chip we can price today, a memory-controller chip that stores the dataset, reaches 2.1x per joule against an RTX 5090 at its knee with a core as good as a GPU lane and 3.4x with one three times better, under class v4 from the first block; a 5090 locked at its knee pays 82 W for that shadow work. The strongest chip we can model for 2027 to 2028, a 2 GiB SRAM store on one 2 nm die, would reach 3x to 5x with that same shadow and 17x without it, at an N2 project of USD 100 M to 500 M and 18 months or more; the dataset's size is the one lever on it, and the schedule says how it grows. Class v5 makes the dataset the chain's own state, so a chip that does not follow the chain is wrong on every item. The model and every measurement are public." Every number in it is labelled in this document and the chip model; nothing served moves on this lane's word.
|
||||
|
||||
## 7c. Beyond the four: the invention lane's layers 5 to 7 and its rejected list (lane C, `docs/analysis/class-v6/invention.md` on master at a9f03598, 10:48 UTC; the census script `tools/attack/v6-invention/v6inv-census.sh`)
|
||||
|
||||
| Layer | What it is | The chip rows | The card cost | Status |
|
||||
|---|---|---|---|---|
|
||||
| 5. The shadow placed per load, in its sound form (`mx8+shl4096x1`: one pass of a 256-instruction sub-block after every load, the same 4,096 shadow instructions per iteration as class v4) | the capex lever of the research file's section 16.2: the chip's core sits inside every read's dependency, so controller and core share one N5-class die or an interposer | the project about USD 60 M against 30 M, the break-even cap about USD 200 M against 100 M (modelled); `k` unchanged | measured on build-1 (10:4x UTC, this crate): 234 of 256 seeds accept within 32 attempts at 0.927 rejection per candidate (P(exhaust at 256) about 4e-9); the verifier 8.28 to 8.82 ms on core 40 with core 88 loaded against 8.33 to 8.63 for class v4's shape; **the 5090 rows MEASURED (the hash lane's v6 job, 11:14 to 11:20 UTC, every pack PASS at both states): mx8-genesis 137.65 MH/s at 312.2 W unlocked and 127.39 at 212.6 W at 1,300 (0.599 MH/W); class v4's shape mx8_sh256x27 137.62 at 464.6 and 126.99 at 296.8 (0.428); the sound per-load form mx8_shl4096x1 135.99 at 428.6 and 126.08 at 271.8 (0.464); mx8_shl2304x3 135.85 at 483.6 and 125.88 at 303.4 (0.415). The rate within 1.3 percent of the control on both forms; the premium over class v3: shl4096x1 116.4 W unlocked and 59.2 W at the lock against the whole block's 152.4 and 84.2, so at class v4's instruction count the one-pass per-load form costs the 5090 24 percent LESS unlocked and 30 percent less at the knee (the 16-instruction block effect of 6 October, now with a sound construction); shl2304x3 (three passes of 2,304) costs 19 W MORE unlocked and 6.6 W more at the lock than the whole block, so the saving is the one-pass shape, not the placement**; the Apple footprint of a 4,096-line block owed (the 1,024-line block cost the M5 Max 17 percent on 6 October) | a candidate class; the 16 x 27 iterated form stays dead (20.2a-close of the research file, corrected to the no-era figure) |
|
||||
| 6. The register-file width drawn per era in {8, 16, 32} | the link tax on layer 5: 4.5 to 9 TB/s of die-to-die traffic closes the interposer branch (modelled, the link figures approximate) | forces the single die | about 0 rate on every card by the occupancy arithmetic (unmeasured) | research |
|
||||
| 7. Warp-uniform data-dependent block selection (B drawn sub-blocks, one selected per iteration by a warp-folded register, no divergence) | moves the FPGA lane only (a per-program bitstream must carry every block) | nothing against the `f = 1` chip | B capped by the Apple compile footprint | research |
|
||||
| The reserve ordered by hardware orthogonality inside layer 3 | shfla first, the int8 tile last | section 4.2's order | | taken |
|
||||
|
||||
Rejected by the invention lane with the number (its section 2): per-lane data-dependent branches, reads tied to the shard proof per block, randomised memory topology, the VRAM ratchet as a lever (layer 2's floor stands as written), pool-sampled witnesses, prover-gated eligibility (1.6 kW of proving network-wide at any hash rate: 0.7 percent of the hash at 100 GH/s), time-locked commitments beyond the era VDF, derived reads in the hash, row-straddling reads, scratch draws, the refresh per block. No candidate moves the per-joule identity.
|
||||
|
||||
## 10. The floor: the stored-dataset chip's per-joule floor and everything behind it (the five floor lanes of 12:55 UK; their rows by 15:30 UK (the coordinator's clock, 13:30 UK; the earlier 19:30 defaults became 15:30), their documents under `docs/analysis/class-v6/floor/` by 19:30; this section closes at 15:45 UK on every row in by 15:30, each missing row's default stated and the row owed; later rows are amendments with their own time)
|
||||
|
||||
The question the founder set: the stored-dataset chip pays the DRAM's own energy per random read and nothing else; the honest card pays that plus everything its silicon spends around the read; the floor is the ratio, and every term in it is a lever. The identity (the research file's section 2): at zero shadow `edge = E_card / E_mem`; with the shadow `(E_card + F) / (E_mem + k F)`. The lanes each take one term.
|
||||
|
||||
| Lane (branch) | The term | What this document already holds (measured) | The default if the lane's rows are not in by 15:30 UK | The row owed |
|
||||
|---|---|---|---|---|
|
||||
| SM-sparse (`class-v6-floor-sm`) | `E_card` above `E_mem`: the honest card's watts toward the DRAM's own at the activate ceiling, on rented 5090, 4090 and H100 and PC 1 through the hash lane's jobs, with a worker patch | the research file's 20.3b (PC 1, 7 October night, the fourth exe): a quarter of the SMs holds 98.2 percent of the class v4 rate at the SAME draw (460 against 451 W) and 99.8 percent of class v3 at 4 W less; watts minus idle per MH/s never falls below base; the draw follows the work, not the SM count; at the 1,300 lock the sparse shapes collapse. The candidate was closed on that row | the 20.3b reading stands: the honest card's premium-free floor is its idle plus its memory system plus whatever the SMs spend waiting (99 W of 228 at the lock, measured as the residual, not reached by idling SMs); the 5090 at its knee 1.66 microjoules against the GDDR7 chip's 0.466: 3.6x | a measured breakdown of the 99 W that an idle SM does not save (clock tree, L2, fabric), and whether a different occupancy shape (fewer warps per SM at full SM count) moves it |
|
||||
| The shadow's k (`class-v6-floor-k`) | `k` from RTL synthesis (Yosys plus OpenROAD on build-4) replacing every claimed chip-side k; the op mix that maximises k | the GPU side measured (the research file's 15.1a): ARX 11.3 / 6.2 pJ per counted op, mul 13.9 / 8.3, mulhi 39.6 / 21.0, prmt 22.3 / 11.5, lop3 24.1 / 13.0, shfl 55.8 / 29.4, fp32 FMA 9.2 / 5.2, the int8 tile 1.4 to 4.1 per MAC; the chip side was claimed only (a 5 nm SIMD array 2 to 5 pJ per op, approximate) until the lane's first rows (10.2, 13:3x UK): synthesised on ASAP7, k 0.18 at the lock in the absolute convention at N3 (0.35 unscaled), the ARX families 0.17 to 0.18, mad 0.20, mulhi 0.03 | the default stays the record's rows at k 0.5 (the DRAM chip 2.9x at the knee) with the unit floors as the lower bound: the GDDR7 chip with the class v4 shadow at 4.0x at the lock and 5.8x at stock (absolute, floor k) is the worst case the served line must survive, not the reading; the sequencer-core row with the M5 Max column is the headline when it lands; the op mix held at class v4's (15.1b: the shuffle at 4.9x the add on the GPU side; mulhi the worst lever on the card by 6x) | the synthesised pJ per op per family for a 32-lane SIMD core at the chosen node, the resulting k per family, and the mix that maximises k at a fixed GPU premium |
|
||||
| The SRAM full-store die (`class-v6-floor-sram`) | `E_mem` for the chip lane B named strongest: the wire-energy lever, the dataset schedule that keeps the store above USD 5,000 of silicon through 2031, the capex wall recomputed with rotation | lane B's rows (7a, chip-model 5.12): 2 GiB on one N2 reticle, 1.0 nJ per read (0.5 to 2.0), 17x at zero shadow, 2.7x to 4.8x with it, USD 400 to 600 per die; lane A's schedule (3.4): 5.5 / 8 / 11 GiB retiring a quarter of today's consumer cards per step; the M5 Max's measured rate cost of the larger working set (3.3) | lane B's and lane A's rows stand; the schedule decision is the founder's with the per-tier table (0.1) | the schedule in GiB per year that keeps the store at USD 5,000 or more of silicon through 2031 on the SRAM cost curve, and what it costs each tier by year |
|
||||
| The honest denominator (`class-v6-floor-denominator`) | `E_card` per tier with every software knob (the core lock, the undervolt, the memory clock, the occupancy, the block shape); the Ember tier table; FIRST TABLE IN (10.4, 13:4x UK): the honest NVIDIA floor is the 16 GB Blackwell card at its knee (the 5080 2.06 measured), the only lever the operating point | the knee rows (the research file's 20.3a and 20.3b; the cost rows): the 5090 at 1,300 MHz 1.66 to 1.69 microjoules, the 4070 at its tune 2.57, the 5070 Ti stock 1.79, the M5 Max 0.78 (GPU and DRAM channels); the operating point is rank 1 of the research file | those rows stand as the per-tier floor; the Ember knob ships as ordered for 0.3.24 | the per-tier best point with every knob, the Ember table, and the AMD and Apple lines (ADLX or nothing; no lever) |
|
||||
| Invention beyond the four (`class-v6-floor-invention`) | the tensor core's k re-read, the RT core, controller plus PHY cost, proof of latency, era-driven address mapping, the literature since 2023 | the research file: the tensor tile's measured 1.4 to 4.1 pJ per MAC against a 5 nm array's claimed 0.04 to 0.4 (k 0.03 to 0.3, the worse lever); the L2 hit at 1.4 to 2.4 nJ against a chip's SRAM (k 0.1 to 0.3); the texture interpolator excluded (not bit-exact across vendors); the literature table (sections 5 and 8) | those readings stand; no new lever | anything that reads k above 1 on measured rows, which nothing has |
|
||||
|
||||
The close (15:45 UK; the synthesis on the mirror's master by 16:30): the honest floor per tier against each chip row, the recommended class v6 changes (the op mix, the SM-sparse default, the dataset schedule), the served chip line as a measured range, and what stays for v7, written here from the lanes' rows as they land.
|
||||
|
||||
### 10.0 The close at 15:45 UK (the founder's standing rule, 13:3x UK: anything doable in 30 minutes to 2 hours gets a clock inside that window; so 15:45 is THE close, not a provisional: the floor lands on every row in by 15:30, and the lock rows and the k lane's remaining unit rows go in below as amendments, each with its own time; there is no 20:00 event. Lane 3's 7bc9de4b capex numbers are final for the close; the SRAM die headlines the k lane's synthesised sequencer-core row the moment it lands, by 15:00, with its 6x to 10x at the full shadow kept beside it marked as the floor-k worst case)
|
||||
|
||||
The table, filled with the defaults, each cell replaced as its lane's row lands before 15:30 (the k lane's sequencer-core row by 15:00; the SM-sparse 5090 knee row by 15:00; lane 5's W = 16 row on a rented 5090 by 14:45; the denominator lane's table); after 15:45 each later row is an amendment under this section with its own time. The edge is whole-card joules per hash over whole-chip joules per hash; the card's energy is measured (class v4, the shadow on), the chip's memory energy modelled (chip-model-v3 5.5, lane B's 5.12 and lane 3's wire figure for the SRAM die at the genesis width of 4 words, 0.051 microjoules per hash), and the chip's shadow is priced two ways in every cell: first at the DEFAULT k of 0.5 in the record's convention (half the card's own shadow premium: 0.55 microjoules at stock, 0.33 at the lock, 0.31 on the M5 Max; marked `default`), then at the k lane's per-unit floor (0.11 microjoules per hash at N3, absolute, every unit a floor; marked `floor`), the worst case the served line must survive; the sequencer-core row replaces the default when it lands.
|
||||
|
||||
| Honest tier (measured joules per hash, class v4) | GDDR7 board (0.466) | HBM3 one stack (0.321) | SRAM die at W = 4 (0.051) | With SM-sparse | W = 16 variant |
|
||||
|---|---|---|---|---|---|
|
||||
| RTX 5090 stock (3.36; class v3 2.26) | 3.3x default / 5.8x floor | 3.9x / 7.8x | 5.6x / 21x | 1.3 to 3.9 percent at best, MEASURED (10.1, 14:10 UK: four rented 5090s, the ceiling held to 43 of 170 SMs; the 4090 saves 4 W at 16 SMs, the H100 1.4 percent); the residual is the clock domain, which only the lock takes off; the fraction a chip cannot strip is 16 to 19 percent of the 5090's stock joules, 25 percent at the lock | not applied (default: outcome A, the card -47 percent, dead; the rented-5090 row owed by 14:45) |
|
||||
| RTX 5090 at the 1,300 MHz lock, the record's operating point (2.33 = 1.67 + 0.65) | 2.9x / 4.0x | 3.6x / 5.4x | 3.9x absolute (6.1x in the record's convention, marked) / 14x | no change (the lock pass on PC 1 running; the knee row by 15:00) | not applied (same default) |
|
||||
| Apple M5 Max (1.40; class v3 0.78; the GPU and DRAM channels) | 1.8x / 2.4x | 2.2x / 3.2x | 3.9x / 8.7x | not applicable (no SM lever on Apple) | not applicable (the M5 Max pays 0 at 64 bytes, measured; the chip +33 percent) |
|
||||
| RTX 5080 at its 1,100 MHz lock, the honest NVIDIA floor (2.06 class v4 measured: 71.20 MH/s at 146.6 W; class v5 2.10; floor lane 4, 10.4) | 2.6x / 3.6x | 3.0x / 4.8x | 3.6x / 13x | no change | not applied |
|
||||
| RTX 4090 stock (5.005; class v3 3.588; rented pod, 10.1) | 4.3x / 8.7x | 4.9x / 11.6x | 6.6x / 31x | no change (measured: 16 of 128 SMs saves 4 W) | not applied |
|
||||
| H100 HBM3 at its 700 W cap (2.776, throttled to 1,750 MHz; class v3 1.769; rented pod, 10.1) | 2.9x / 4.8x | 3.4x / 6.4x | 5.0x / 17x | no change (measured: per-SM throughput-bound, 1.4 percent) | not applied |
|
||||
|
||||
The honest floor in one line, on the defaults: against the chip anyone can build (the GDDR7 board) the strongest honest tier by joule, the M5 Max, holds the chip to under 2x, the 16 GB Blackwell card at its knee (the 5080 measured, the 5070 Ti and 5070 modelled; floor lane 4's finding that the honest NVIDIA floor is not the 5090) to about 2.6x and a 5090 at its knee to about 3x; against the chip that needs an N2 project (the SRAM die) the holds are about 4x and 4x; every cell's floor-k figure is the worst case if a chip's core costs no more than its bare units.
|
||||
|
||||
The capex wall's two thresholds (lane 3, 10.3, re-folded 13:18 UK on lane 5's corrected project floor of USD 20 to 75 M for the cheapest DRAM-board chip): no rational chip project of any kind below about USD 23 M a year of miner revenue (IGN 0.03 at launch emission, USD 62 K a day, a cap of about USD 60 M); every DRAM-board project at a third of the network above about USD 340 M a year (IGN 0.44, USD 0.93 M a day, a cap of about USD 0.9 B), where the SRAM project also starts. The project cost moves the threshold 5x, the chip's edge 1.4x.
|
||||
|
||||
The served line in one sentence, on the defaults: a chip that stores the dataset reaches about 3x per joule against an RTX 5090 at its knee and under 2x against an M5 Max, 4x against both if it is built on an N2 SRAM store for USD 100 M or more, and no rotating layer moves those figures; the layers decide which chip can be built and how long its tape-out lives. The worst case beside it, if a chip's core costs no more than its bare units (the floor k of 10.2): 4x and 2.4x against the DRAM board, 14x and 9x against the SRAM die (lane 3's line: 6x to 10x at the full shadow, never under 2x).
|
||||
|
||||
The three class v6 changes it implies: (1) the op mix stays class v4's with the lossy families capped at their base (lane D: 0 of 3,000 eras exhausted under the band; the k lane: mulhi is the card's worst lever by 6x, so no re-weight helps the card); (2) no SM-sparse default (dead on the measured 5090 rows; the per-watt gift of the hot table is the chip's, 20.3c); (3) the dataset schedule 5.5 / 8.5 / 11.5 GiB with the read width pinned at 4 words (the chip's ticket USD 1,500 / 2,500 / 3,000; a tuned 5090 pays 1 to 5 percent per step; about a quarter of today's cards by count per step).
|
||||
|
||||
### 10.1 SM-sparse on the honest card (floor lane 1, `docs/analysis/class-v6/floor/sm-sparse.md` on class-v6-floor-sm at 9e478daf, first rows 13:25 UK; the knee row by 15:00, the file by 15:30)
|
||||
|
||||
**The 20.3b reading holds on three more cards, and the residual over the DRAM's own energy is the clock domain, not the SMs.** Rented pods at stock (the worker's `--bench`, 250 x 2^24, nvidia-smi at 1 Hz, bit-exact on every row); microjoules per hash, the edge card over chip (GDDR7 0.466, HBM3 0.321, SRAM at lane B's 0.14):
|
||||
|
||||
| Card (measured) | Class v3 base | SM-sparse best | What the SM count does | Class v4 |
|
||||
|---|---|---|---|---|
|
||||
| RTX 4090 (128 SMs, idle 34.7 W) | 70.63 MH/s at 253.4 W = 3.588 microjoules (7.7x / 11.2x / 25.6x) | 16 SMs 69.95 at 249.6 W = 3.568 (saves 4 W); 32 SMs 70.42 at 251.4; 8 SMs 65.07 at 239.1 = 3.675; 128 SMs x 1 warp 66.28 at 242.7 | the draw over idle is 205 to 219 W at every shape; the knee 16 of 128 SMs | 70.63 at 353.5 W = 5.005 (premium 100 W, 11.1 pJ per counted op) |
|
||||
| H100 HBM3 (132 SMs, idle 71.3 W) | 254.63 at 450.4 W = 1.769 (3.8x / 5.5x / 12.6x) | the self-tune's sp127-w32 253.2 at 441.6 W = 1.744 (1.4 percent under base); 132 SMs x 8 warps 99.2 percent at the same draw | the rate falls one for one with the SM count (99 SMs 212.7 at 402.8 W, 1.894; 66 SMs 189.2 at 372.9; 33 SMs 123.4 at 291.0, 2.358): per-SM throughput-bound at 1.93 MH/s per SM, not activate-bound; -pl and -lgc refused, -lmc accepted but the memory clock stayed at 2,619 MHz | 251.9 at the 700 W cap = 2.776 (clock throttled to 1,750 MHz); class v5 250.2 at 698.8 W = 2.793 |
|
||||
| RTX 5090 (PC 1, unlocked) | 20.3b's points: full 2.30 at 313.9 W; a quarter 2.27 at 309.8; an eighth 2.28 at 303.4; a sixteenth 2.74 at 274.5 | the finer ladder (170 to 11 SMs at 32 warps; 170 SMs at 16 to 1 warps; the matched-warp pairs) runs on four rented 5090s and on PC 1 (`run-ca4-pc1-floorsm-5090-20261008`, the unlocked pass done, the 1,300 lock pass running) | 170 x 8 warps 137.0 at 482.9 W against base 137.16 at 452.8 on class v4; 170 x 4 134.9 at 483.9: fewer warps per SM at full SM count does not lower the draw either | base 137.16 at 452.8 W |
|
||||
|
||||
The reading for section 2's term: what an idle SM does not save is the die's clock-domain cost (the clock tree, L2 and the crossbar, the memory controllers, leakage at 2,700 to 2,850 MHz); the 4090 at 8 SMs and 65 MH/s still draws 204 W over idle at 2,715 MHz, while the 5090's lock takes 100 W off at the same read rate. **The honest floor at stock per card is the base row within 2 percent; the lever stays rank 1, the operating point.** The patch that ships the auto-tune anyway (`--sm-sparse auto|off|N` and `--sm-hold` on the worker, the tuning-file keys `sm_sparse` and `sm_hold` per card, the pure pieces tested in `emu/variant-test.cpp` on the H100 pod) is on class-v6-floor-sm at 9e478daf; the design's default is off.
|
||||
|
||||
The 5090 ladder (14:10 UK; four rented 5090s on four hosts, base rows 2.08 to 2.49 microjoules at 142 MH/s by board; the ladder as a percent of each host's own base; bit-exact; placement checked by `%smid`). **At stock the 5090 holds class v3's ceiling down to 43 of 170 SMs (99.5 percent) and 28 (98.3), falls at 21 (96.6) and 16 (91); the occupancy saves 1.3 to 3.9 percent of energy per hash by host and never more; class v4's knee is 43 SMs by rate (compute-bound below it) with no saving.** Host c (idle 7 W), class v3, 32 warps per block, one block per SM:
|
||||
|
||||
| SMs | MH/s (percent of base) | W | Microjoules per hash | GDDR7 / HBM3 / SRAM (lane B's 0.14) |
|
||||
|---|---|---|---|---|
|
||||
| 170 | 142.6 (100) | 296.4 | 2.079 | 4.46x / 6.48x / 14.9x |
|
||||
| 85 | 142.4 (99.9) | 296.0 | 2.078 | |
|
||||
| 43 | 141.8 (99.4) | 291.1 | 2.053 | 4.41x / 6.40x / 14.7x |
|
||||
| 28 | 140.3 (98.3) | 288.6 | 2.058 | |
|
||||
| 21 | 137.8 (96.6) | 288.1 | 2.091 | |
|
||||
| 16 | 130.0 (91.1) | 280.4 | 2.157 | |
|
||||
| 11 | 100.9 (70.7) | 255.5 | 2.533 | 5.4x |
|
||||
|
||||
Host e (idle 12 W): base 142.0 at 352.7 W = 2.485; 43 SMs 141.2 at 337.4 = 2.389 (the 3.9 percent case); 21 SMs 137.2 at 334.9 = 2.440. Full SM count with fewer warps (host c): 8 warps per SM 142.4 at 293.8 W = 2.063; 2 warps 139.3 at 288.7; 1 warp 119.6 at 270.3 = 2.260. Matched warps (43 x 32 against 170 x 8): 2.053 against 2.063, equal: the warps in flight set the rate and nothing in the occupancy sets the watts. Class v4 at stock: host c base 142.6 at 450.0 W (the cap) = 3.157; PC 1 unlocked today: base 137.2 at 452.8 W, 43 SMs 134.5 at 470 W, 36 SMs 120.3, 28 SMs 94.4, 21 SMs 70.7, 16 SMs 54.0 (the shadow compute-bound under 43 SMs; the pass drifted 62 to 75 C). The 1,300 lock ladder on PC 1 (`run-ca4-pc1-floorsm-5090-20261008`) is the amendment if not in by 15:30; the default is 20.3b's lock points.
|
||||
|
||||
The decomposition for section 2's term (the H100 microbench, full residency at the boost clock, measured): idle 80 W; resident SMs issuing nothing 136.6 W (+57 W, the SM clock domain); the ALU probe 437.9; the dependent DRAM chase 450.3 W at 32.2 G reads a second (11.5 nJ per read whole-card); the hash 450.4 W at 32.6 G reads a second. The memory path on the GPU side is 313 of the 370 W over idle (9.6 nJ per read above the HBM's modelled 1.2). The 4090 and 5090 probes land by 15:00; PC 1's memory-clock ladder at the 1,300 lock is queued (`run-ca4-pc1-memclk-5090-20261008`). **The fraction a chip cannot strip** (the memory side's own on chip-model 5.3's rows: DRAM plus static plus a controller, about 0.40 microjoules per hash on GDDR7 at 128 reads, 0.47 on the 4090's GDDR6X at its rate, 0.23 on HBM3): the 5090 at stock 16 to 19 percent, at the 1,300 lock 25 percent, the 4090 13 percent, the H100 13 percent. Everything else is the card's silicon around the read and a chip strips it; the honest floor's real number per tier is that fraction, and the only honest-side lever on the rest is the clock (rank 1).
|
||||
|
||||
|
||||
### 10.2 The shadow's k from RTL (floor lane 2, `docs/analysis/class-v6/floor/shadow-k.md` on class-v6-floor-k, build-4 and build-3, first synthesised rows 13:3x UK; RTL and flow under `tools/chip-model/rtl`; full by 15:30 UK (the earlier 19:30 clock pulled by the coordinator at 13:30))
|
||||
|
||||
**The first synthesised k is 0.18 at the 5090's lock in the absolute convention at N3, and nothing reads inside the claimed 0.3 to 0.8 band except the unscaled ASAP7 figure at the lock (0.35). These are PER-UNIT FLOORS: no fetch, no decode, no register file beyond an 8-entry window, which is the chip rotation kills (section 1). The coordinator's order (13:5x UK): the headline k for the close is the programmable sequencer-core row (fetch, decode, a 32-register file, the per-era registers, the drawn program), asked of the lane with the M5 Max column; until it lands the default is the record's shadowed rows at k 0.5 with these unit floors beside them as the lower bound, and the GDDR7 board at 4.0x at the lock and 5.8x at stock on the floor k is the WORST CASE the served line must survive, not the reading.** Method: a minimal lane (instruction register, an 8 x 32-bit register window of flops, read muxes, the unit, one write port; no fetch or decode) in Yosys 0.68 plus OpenROAD on ASAP7, routed, SPEF, power from a random-input gate-level VCD with every pin annotated, the TC corner at 0.70 V; so every chip figure is a FLOOR and every k a floor. The node scaling is claimed from TSMC's headline per-node power reductions (N5 x0.70, N3 x0.50, N2 x0.36 of ASAP7; approximate). The GPU side is the research file's 15.1a, measured.
|
||||
|
||||
| Family (chip RTL) | pJ per op ASAP7 (synthesised) | pJ per op N3 (scaled, claimed) | 5090 pJ per op stock / lock (measured) | k absolute, N3 chip against stock / lock | k at ASAP7 unscaled against the lock |
|
||||
|---|---|---|---|---|---|
|
||||
| ARX add | 2.16 | 1.09 | 11.3 / 6.2 | 0.096 / 0.18 | 0.35 |
|
||||
| ARX xor | 2.09 | 1.05 | 11.3 / 6.2 | 0.093 / 0.17 | 0.34 |
|
||||
| ARX rotl (imm) / rotr (reg) | 2.15 / 2.13 | 1.08 / 1.07 | 11.3 / 6.2 | 0.095 / 0.17 | 0.35 |
|
||||
| ARX sub | 2.14 | 1.08 | 11.3 / 6.2 | 0.095 / 0.17 | 0.35 |
|
||||
| or (lossy; values saturate, activity falls) | 1.45 | 0.73 | 11.3 / 6.2 | 0.065 / 0.12 | 0.23 |
|
||||
| ARX random mix | 2.24 | 1.13 | 11.3 / 6.2 | 0.10 / 0.18 | 0.36 |
|
||||
| mul (32 x 32 low) | 1.35 | 0.68 | 13.9 / 8.3 | 0.049 / 0.082 | 0.16 |
|
||||
| mulhi (high word, the same multiplier) | 1.35 | 0.68 | 39.6 / 21.0 | 0.017 / 0.032 | 0.064 |
|
||||
| mad (3 reads, mul plus add) | 3.29 | 1.66 | 13.9 / 8.3 | 0.12 / 0.20 | 0.40 |
|
||||
| index fold (mul by M, rotl R, masks) | 2.37 | 1.19 | 13.9 / 8.3 | 0.086 / 0.14 | 0.29 |
|
||||
| prmt (byte permute) | 1.27 | 0.64 | 22.3 / 11.5 | 0.029 / 0.056 | 0.11 |
|
||||
| lop3 (8-bit LUT) | 1.42 | 0.72 | 24.1 / 13.0 | 0.030 / 0.055 | 0.11 |
|
||||
|
||||
The ranking by k, highest first (hardest for a chip; absolute, N3 against the lock): mad 0.20, the ARX families 0.17 to 0.18, the fold 0.14, or 0.12, mul 0.08, prmt and lop3 0.06, mulhi 0.03. mulhi is the worst lever on the card by 6x: the 5090 pays 21 pJ for a high word that costs the chip the same 0.68 pJ as a low word, which is the measured reason the op-mix band keeps mulhi capped at its base (section 2) and the reason the research file's 15.1b hold on a re-weight stands from the chip side as well. The class v4 mix costs the chip about 1.0 to 1.2 pJ per op at N3 (2.0 to 2.4 at ASAP7), the shuffle row pending.
|
||||
|
||||
What it does to the GDDR7 stored-dataset board (E_mem 0.466 microjoules, modelled) with the class v4 shadow (102,100 ops per hash): absolute, the chip's shadow is 102,100 x about 1.1 pJ = 0.11 microjoules at N3, E_chip about 0.58, so the edge reads **4.0x at the lock (the card 2.33) and 5.8x at stock (3.36)**, against the served 2.1x at k = 1 and 3.4x at k 0.33. In the record's convention it is k 0.18 at the lock, 2.33 / (0.466 + 0.18 x 0.652) = 4.0x, the same number there, and at stock k 0.10, 3.36 / (0.466 + 0.10 x 1.10) = 5.8x. **The shadow buys the card about 0.5x to 1x of edge out of the 5.2x and 3.6x zero-shadow figures, not the 1.5x the served k = 1 row implies.** This is the row the served line must survive (section 9): the served "2.1x" is the k = 1 reading; the unit floor says a core's units are five to six times cheaper than that at N3, and the sequencer core that a rotating family forces (fetch, decode, register file, the drawn program's control) sits between the two, where the lane's next row puts it; until then the default reading is k 0.5 (the DRAM chip 2.9x at the knee in the record's convention) with 4.0x as the floor-k worst case. Pending from the lane by 15:30 UK (later rows as amendments): the 32-lane shuffle butterfly and the general crossbar (in place now), the 8 KB scratch as a flop array, the int8 8x8x8 tile, the mix optimiser over the layer-1 band, and the re-fold of lane 3's SRAM rows at W = 4 and 8 in both conventions.
|
||||
|
||||
### 10.3 The SRAM die and the dataset floor (floor lane 3, `docs/analysis/class-v6/floor/sram-and-floor.md` on class-v6-floor-sram at 708d01b4, first reading 12:0x UTC; full by 15:30 UK (the earlier 19:30 clock pulled by the coordinator at 13:30); everything chip-side modelled on lane B's wire figure, 1.3 pJ per bit plus 0.1 nJ per macro access and 0.05 nJ of controller; the GPU side measured)
|
||||
|
||||
**The read width is the only wire lever on the SRAM die, and the floor is a ticket, not a joule.** Lane B's 1.0 nJ was the 64-byte row; the die reads what the hash asks for, and the wire scales with the bits moved while the macro access does not:
|
||||
|
||||
**The convention rule for every shadowed row (the coordinator's order, 13:1x UK): the headline is the ABSOLUTE convention, pJ per op on each side at the operating point (the die's core costs what it costs; the GPU's op at its own clock, 11.3 pJ stock and 6.2 at the 1,300 lock on the 5090); the record's convention (k against the GPU's op cost at the same operating point, so the die's core gets cheaper when the card locks) is carried beside it, marked, and it flatters the die at the lock by 1.6x (6.1x against 3.8x at k 0.5; 3.3x against 2.0x at k 1).** The shadowed columns in the table below are the lane's first reading at the card's stock point, where the two conventions coincide. The lock rows, both conventions, absolute first (floor lane 3's 2.2a at d550dadd, 13:09 UK; the card at the lock 1.67 plus 0.65 microjoules measured, the die's E_hash0 from its 2.2; the absolute die pays k x 1.10 microjoules whatever the card does, the record's k x 0.65):
|
||||
|
||||
| W words | Absolute, k 0.5 / 1 (the headline) | The record's convention, k 0.5 / 1 (marked: flatters the die by 1.6x) | Label |
|
||||
|---|---|---|---|
|
||||
| 1 | 4.0x / 2.0x | 6.4x / 3.4x | modelled chip, measured card |
|
||||
| 4 (genesis) | 3.8x / 2.0x | 6.1x / 3.3x | modelled chip, measured card |
|
||||
| 8 | 3.7x / 2.0x | 5.8x / 3.2x | modelled chip, measured card |
|
||||
| 16 | 3.4x / 1.9x | 5.2x / 3.0x | modelled chip, measured card |
|
||||
|
||||
The k lane's RTL rows (floor lane 2, in absolute) replace the k axis of this table when they land; until then the honest shadowed reading of the SRAM die against a 5090 at its knee is 2.0x at k 1 and 3.7x to 4.0x at k 0.5, whatever the read width.
|
||||
|
||||
| W words | Bits per read | Chip E_read | GH/s per die (300 W) | Zero shadow against the 5090 | With the class v4 shadow, k 0.5 / 1 (stock point; the record's convention, to be re-folded by the k lane in absolute) | At the card's whole shadow (F 2.0), k 0.5 / 1 (same convention note) | The honest card's cost | Label |
|
||||
|---|---|---|---|---|---|---|---|---|
|
||||
| 1 (4 bytes, today) | 80 | 0.25 nJ | 8.3 | **66x** | 5.7x / 3.0x | 4.1x / 2.1x | 0 | modelled chip; measured card |
|
||||
| 4 (16 bytes) | 176 | 0.38 | 5.6 | 44x | 5.6x / 2.9x | 4.0x / 2.1x | the 5090 +2.7 percent, the 9070 XT -1.4, the M5 Max within 1 (measured 5 October) | measured card |
|
||||
| 8 (32 bytes, the GDDR7 sector the 5090 fetches anyway) | 304 | 0.55 | 3.9 | 31x | 5.4x / 2.9x | 4.0x / 2.0x | free by the measured rows (13:06 UK: the 5090's ceiling about 18 G sectors a second, w16 moved 573 GB/s of sectors at 139.8 MH/s and w64 589 at 71.9; an aligned 32-byte read is one sector per load as w16; AMD moves its 64 B line either way; the M5 Max 1.03 at both w16 and w64); one PC 1 row to confirm, and the crate's width set is {1, 4, 16} words, so a pack at 8 needs a generator and emitter line first | measured card (inferred at 8) |
|
||||
| 16 (64 bytes, lane B's record) | 560 | 0.88 | 2.4 | 19x (the record's 17x) | 5.0x / 2.7x | 3.8x / 2.0x | the 5090 -47 percent: dead | measured card |
|
||||
|
||||
Meaning: at the hash's own width the die is 3.5x stronger than the record said at zero shadow; with the shadow on, every row sits at 5.0x to 5.7x (k 0.5) and 2.7x to 3.0x (k 1): **the shadow is the whole hold, the memory moves it 0.2x to 0.7x.** The fold is dst-keyed (`verify::fold_words`), so a wide read cannot be pre-folded; dependent chains per step leave the ratio unchanged (per-read on both sides); banking plus a sequencer cannot localise (uniform on a window of at least 256 MiB, the 16 sites alternating; moving the lane costs about 290 bits, which caps the wire lever at about W = 8); the per-site window shrink buys the die and the DRAM chip nothing; a hop at the honest widths is 0.04 to 0.15 nJ, so a ten-die store still reads 25x to 58x at zero shadow. This moves layer 1's width row: **pin W = 4 (16 bytes) at genesis (measured free on all three vendors; the die from 66x to 44x), W = 8 (31x; 5.4x at k 0.5) once the generator and emitter carry it and one PC 1 row confirms, never 16.**
|
||||
|
||||
The floor as a ticket: SRAM is flat at about USD 250 per GiB to 2031 (density +6 to 11 percent per node against dearer wafers; claimed, approximate), so USD 5,000 of silicon per store is 20 GiB at N2 and about 21 GiB in 2031, which retires every card under 32 GB and every Mac under 64 GB; the constraint and "fewest cards" cannot both hold. The per-MH/s does not rise with the floor (every die powered: USD 0.25 to 0.5 per MH/s at any size); the floor raises the minimum ticket only. The honest card's side of the floor is measured (the hash lane's kit b, section 3.3): the 5090 at its knee pays 4 / 8 / 10 percent more energy per hash at 2 / 4 / 8 GiB, so every chip row's edge against a card at its knee rises by 4 to 11 percent across the schedule while the die's joules do not move; the schedule's sentence is USD 1,000 of chip ticket per step for 1 to 5 percent of the tuned 5090's energy and about a quarter of today's cards by count. The k convention matters at the lock: the record's (k against the GPU's op cost at the same operating point) reads the SRAM die at 6.1x (k 0.5) and 3.3x (k 1); the absolute convention (the die's core costs what it costs, k against the stock 11.3 pJ per op) reads 3.8x and 2.0x; the record's flatters the die by 1.6x at the lock, so the served line in section 9 and the close carry the absolute rows as the headline and the record's beside them marked.
|
||||
|
||||
| Floor | Reticles 2026 / 2031 | USD of silicon | Tiers out (worst / best working set) | Label |
|
||||
|---|---|---|---|---|
|
||||
| 5.5 GiB (the v6 epoch) | 3 / 3 | 1,500 | the 8 GB tier in the worst case only / none | modelled |
|
||||
| 8.5 GiB (year 2) | 5 / 4 | 2,500 | 8 GB, 12 GB cache-resident, Apple 16 GB / 8 GB, Apple 16 (the 12 GB tier holds at 73 to 75 percent) | modelled |
|
||||
| 11.5 GiB (year 4) | 6 / 5 | 3,000 | 8, 12, 16 GB cache-resident, Apple 16 / 8, 12, Apple 16 (the 16 GB tier holds at 73 to 75 percent) | modelled |
|
||||
| 16 GiB | 8 / 7 | 4,000 | adds 16 GB, 24 GB cache-resident, Apple 32 | modelled |
|
||||
| 20 GiB (USD 5,000) | 10 / 9 | 5,000 | everything under 32 GB; Macs under 64 GB | modelled |
|
||||
|
||||
The lane's recommended schedule: 5.5 / 8.5 / 11.5 GiB (each just under a tier's room and just over a reticle multiple; 8.5 forces a fifth die where 8.0 does not, at no tier cost with the cache freed; the non-power-of-two floors need the multiply-shift mapping of 1.13.3 option (a)); the Apple rate cost on top, -21 to -25 percent (measured to 8 GiB, section 3.3); the state sizes assumed 92 M / 143 M / 193 M records, the 16 GiB step by state above 268 M.
|
||||
|
||||
The capex wall on the spec's emission constant (3,168,808,781 base units per DAA second, 10^8 per IGN, 80 percent to miners: 0.77 / 0.80 / 0.40 / 0.40 / 0.20 / 0.20 B IGN to miners in years 1 to 6; the mission lane's E2 4.18 B is not this constant's figure and is carried beside it); a project starting at launch, shipping at 24 months, mining years 3 to 6, a 10 percent discount; rotation costs it nothing (firmware); the floor's fleet term drops out (under two dies for the whole network at every price):
|
||||
|
||||
| Project USD M | Share of the hash | p* at an edge of 3x / 5x / 13x (USD per IGN) | Launch-year miner revenue at p* | Label |
|
||||
|---|---|---|---|---|
|
||||
| 100 | 0.30 | 0.59 / 0.49 / 0.43 | USD 330 to 450 M a year (0.9 to 1.2 M a day) | modelled |
|
||||
| 100 | 1.00 (takes the chain) | 0.18 / 0.15 / 0.13 | 98 to 136 M a year | modelled |
|
||||
| 250 | 0.30 | 1.47 / 1.23 / 1.06 | 0.8 to 1.1 B a year | modelled |
|
||||
| 500 | 0.30 | 2.94 / 2.45 / 2.12 | 1.6 to 2.3 B a year | modelled |
|
||||
|
||||
Below about USD 100 M a year of miner revenue (IGN 0.13) no rational SRAM project starts at any edge; above about USD 2.3 B a year (IGN 2.9) every one does; at IGN 0.01 / 0.10 / 1.00 (assumed) a USD 150 M project at 5x and 30 percent reads NPV -148 / -130 / +54 M. **The edge moves p* 1.4x across 3x to 13x; the project cost moves it 5x: the hash does not set the wall, the project cost and the detector do.** Killed with the number: per-era re-fill (0.96 J per era), per-block re-fill (0.96 W on the die against 1.3 to 3.2 percent of GPU rate), straddling atoms (+0.02 nJ on the die, +19 percent of items on the verifier at the gate), a second hot table (the die's own kind of memory, k 0.1 to 0.3). Kept: the per-lane scratch pins the lane (a dependency of the width lever); 3D-stacked SRAM would remove the floor's last effect on joules. The lane's line for the served sentence: about 4x at k 0.5 and 2x at k 1 against a 5090 paying its whole latency shadow, 20x to 60x without it, with the price threshold beside it (about USD 0.4 per IGN at a third of the network, 0.13 for a maker who takes the chain). Owed: the W = 8 rows, the Apple curve past 8 GiB.
|
||||
|
||||
The lane's close rows (class-v6-floor-sram 7bc9de4b, 13:18 UK), three of them.
|
||||
|
||||
(1) The capex wall re-folded on floor lane 5's corrected project floor (the cheapest DRAM-board chip project USD 20 to 75 M, not the record's 5 M; the same model: ships at 24 months, mines years 3 to 6, a 10 percent discount; an 18-month ship lowers p* by 15 percent):
|
||||
|
||||
| Project USD M | Share | p* at 5x / 3x (USD per IGN) | Launch-year miner revenue | Per day | Cap at the end of year 2 | Label |
|
||||
|---|---|---|---|---|---|---|
|
||||
| 20 | takes the chain | 0.029 / 0.035 | USD 23 / 27 M a year | 62 / 74 K | 0.06 / 0.07 B | modelled |
|
||||
| 20 | 0.30 | 0.098 / 0.118 | 75 / 91 M | 207 / 248 K | 0.19 / 0.23 B | modelled |
|
||||
| 75 | takes the chain | 0.110 / 0.132 | 85 / 102 M | 233 / 279 K | 0.22 / 0.26 B | modelled |
|
||||
| 75 | 0.30 | 0.368 / 0.441 | 283 / 340 M | 776 / 931 K | 0.72 / 0.86 B | modelled |
|
||||
|
||||
The two numbers for the close: no rational chip project of any kind below about USD 23 M a year of miner revenue (IGN 0.03, USD 62 K a day, a cap of about USD 60 M); every DRAM-board project at a third of the network above about USD 340 M a year (IGN 0.44, USD 0.93 M a day, a cap of about USD 0.9 B), where the SRAM project also starts (USD 330 M at 13x, or a whole-chain maker). The correction moves the lower number 4x (from the 100 M of the SRAM-only reading) and the upper not at all.
|
||||
|
||||
(2) W = 16 at each chip's own cost (GDDR7 +33 percent on the memory term, HBM3 +24, the die at 560 bits, 0.88 nJ), zero shadow, two card readings (the 5090 passes within 5 percent at its best lane count, or the measured w64 row stands: 71.9 MH/s, 1.89x the energy if the watts hold):
|
||||
|
||||
| Chip | W = 1 | W = 16 if the 5090 passes | At the measured w64 row | Class v4 shadow (the synthesised core) | The full shadow | Label |
|
||||
|---|---|---|---|---|---|---|
|
||||
| GDDR7 board | 5.1x | 4.4x | 8.3x | 5.1x | 4.7x | modelled chip, measured card |
|
||||
| HBM3 one stack | 7.5x | 6.7x | 12.7x | 7.2x | 5.9x | modelled chip, measured card |
|
||||
| SRAM die | 66x | 19x | 36x | 14x | 8.8x | modelled chip, measured card |
|
||||
|
||||
W = 16 is the one width that moves the die more than 2x at zero shadow and it costs the DRAM chips 11 to 15 percent; worth it only if the PC 1 row passes, since at the measured w64 row every chip's edge nearly doubles. If it passes, W = 16 replaces W = 8 as the lane's width recommendation (and the design's `never` is lifted by that one measurement, section 10.5).
|
||||
|
||||
(3) The shadowed SRAM rows on the k lane's synthesised core (0.11 microjoules of shadow at N3, absolute k about 0.1, the shuffle open; the lane's 2.2b): under class v4 21x at W = 4 and 18x at W = 8 at stock, 14x and 12x at the knee; at the honest cards' whole latency shadow 10x and 9.7x at stock, 6.5x and 6.1x at the knee; an N2 core about 1.2x more. On the synthesised core the shadow is not the whole hold (the record's claimed band read 2x to 6x, marked beside): the memory is a third to a half of the die's energy, the width is worth 1.3x with the shadow on, and the lane's line for the served sentence is 6x to 10x at the full shadow (2x to 4x on the claimed band, marked), never under 2x. These are the per-unit-floor figures of 10.2; the k lane's sequencer-core row re-folds them.
|
||||
|
||||
|
||||
### 10.4 The honest denominator per tier (floor lane 4, `docs/analysis/class-v6/floor/denominator.md` on class-v6-floor-denominator, first table 13:4x UK; the rented sweep by 15:00, the final table by 15:15; the Ember tier table `app/igneum-app/tiers/class-v5-tiers.json`, 30 card classes, 5 measured, test 10 of 10 on build-3)
|
||||
|
||||
**The honest NVIDIA floor is the 16 GB Blackwell card at its knee, not the 5090.** Class v5 is class v4's shape plus the state leaves (0.0 percent of rate, +2.0 percent of watts at the 5090's knee, measured, the v5lock job). The chip columns are the chip's class v5 energy at its shadow priced in absolute picojoules per forced op (3.2 pJ, the record's k 0.5 at the 5090's knee; 6.4 pJ, k 1): GDDR7 0.79 / 1.11, HBM3 one stack 0.65 / 0.97, N2 SRAM 0.46 / 0.79 microjoules per hash (the SRAM column at lane B's 64-byte row; at the genesis width, 10.3, the die's memory term is 0.051 and the column reads higher).
|
||||
|
||||
| Tier | Card, the point | Class v5 microjoules per hash | Label | GDDR7 k 0.5 / k 1 | HBM3 | SRAM |
|
||||
|---|---|---|---|---|---|---|
|
||||
| 32 GB | 5090, 1,300 MHz lock (Balanced) | 2.37 (v4 2.32 measured at 134.76 MH/s, 312.5 W) | measured | 3.0x / 2.1x | 3.6x / 2.4x | 5.0x / 3.0x |
|
||||
| 32 GB | 5090, 1,200 MHz (Efficiency, rate -2.2 percent) | 2.33 | measured | 3.0x / 2.1x | 3.6x / 2.4x | 5.0x / 3.0x |
|
||||
| 16 GB | 5080, 1,100 MHz lock | 2.10 (v4 2.06 measured: 71.20 MH/s at 146.6 W) | measured | 2.7x / 1.9x | 3.3x / 2.2x | 4.5x / 2.7x |
|
||||
| 16 GB | 5070 Ti at its knee | 1.70 (stock v4 2.84 measured; the Blackwell shape) | modelled | 2.2x / 1.5x | 2.6x / 1.8x | 3.7x / 2.2x |
|
||||
| 12 GB | 5070 at its knee | 1.75 (stock v4 2.99 measured) | modelled | 2.2x / 1.6x | 2.7x / 1.8x | 3.8x / 2.2x |
|
||||
| 12 GB | 4070 at its tune (1,860 MHz, 50 percent cap) | 3.58 (v4 3.51 measured: 31.08 at 109.0 W) | measured | 4.5x / 3.2x | 5.6x / 3.7x | 7.7x / 4.5x |
|
||||
| 24 GB | 4090 at an Ada knee | 3.64 (stock v4 5.0 measured by lane 1; the 4070's shape) | modelled, stock measured | 4.6x / 3.3x | 5.6x / 3.8x | 7.8x / 4.6x |
|
||||
| 24 GB | 3090 at the Ampere cap | 6.4 (the cap is the only lever, 10 to 15 percent) | modelled | 8.1x / 5.7x | 9.9x / 6.6x | 13.8x / 8.1x |
|
||||
| 16 GB AMD | 9070 XT, ADLX -500 MHz / -30 percent | 8.1 (v4 7.9 measured: 18.9 MH/s at 149.3 W) | measured | 10.3x / 7.3x | 12.6x / 8.4x | 17.5x / 10.3x |
|
||||
| DC | H100 at a lock (an owned host) | 2.0 (stock v4 2.77 measured) | modelled | 2.5x / 1.8x | 3.1x / 2.1x | 4.3x / 2.5x |
|
||||
| Apple | M5 Max, no lever (the GPU and DRAM meter) | 1.43 (v4 1.40 measured) | measured | 1.8x / 1.3x | 2.2x / 1.5x | 3.1x / 1.8x |
|
||||
| Apple LPDDR6, five years out | a Max-class SoC on JESD209-6 | 1.18 at the meter (0.27 DRAM, 0.35 GPU, 0.56 shadow) | modelled | 1.5x / 1.06x | 1.8x / 1.2x | 2.5x / 1.5x |
|
||||
|
||||
What it means: against the 16 GB Blackwell card at its knee (2.1 measured on the 5080; 1.7 to 2.1 modelled on the 5070 Ti and 5070) the GDDR7 chip is under 2x at k 1 and the SRAM die 2.7x; against the M5 Max 1.3x and 1.8x; against the LPDDR6 Apple part five years out the GDDR7 chip is at parity at k 1 and the SRAM die 1.5x. The only software lever is the operating point (the lock on Blackwell and Ada, the cap on Ampere, the ADLX offsets on AMD, nothing on Apple or Intel); the occupancy is worth 1 to 2 percent (10.1), the memory clock untouched, the block shape and the cache hint closed. The lane's per-tier scoring rule: size the shadow by the Apple tier's 5 percent point as the 2.0 rule already does and keep N at 100,000 (sizing to the Apple ceiling of 130,000 costs every honest miner 8 to 15 percent of electricity for about 0.15x of edge: the 5090's GDDR7 edge at k 1 goes from 2.13x to 1.95x); do not score the acceptance floor per tier (one network-wide number). The re-measure rule after a class flip: a stored set is stale under a new class, the stored point is the provisional start of the re-measure, max is stock and never stale; the measured v4 to v5 flip moved the 5090's knee by nothing. Owed and labelled: every Ada and Ampere knee is modelled because no rented host allows -lgc (14 pods today, every one refused); the rented stock rows are measured; the sweep adds class v5 watts measured on 16 card classes and a memory-clock try on each.
|
||||
|
||||
### 10.5 Invention beyond the four layers (floor lane 5, `docs/analysis/class-v6/floor/invention.md` on class-v6-floor-invention, first KEEP/KILL table 13:4x UK; full by 15:30 UK (the earlier 19:30 clock pulled by the coordinator at 13:30))
|
||||
|
||||
**One lever found, conditional on one measurement: the second DRAM sector.** The identity's only `k = 1` work is the memory's own, and the record forced one 32-byte sector per read while the dataset item is 64 bytes. Reading the whole item (W = 16 words; the read-width branch's class w64, fingerprint 836e56e7d496e980, the verifier 0.630 against 0.604 ms) forces a second sector in the same open row: on the chip model's own inputs (909 pJ activate plus 1,150 pJ of movement per sector) the GDDR7 chip's read goes from 2.0 to 3.2 nJ and E_mem from 0.466 to 0.62 microjoules (+33 percent at an unchanged 166 MH/s), one HBM3 stack 0.321 to 0.40 (+24 percent), and on floor lane 3's wire figure the N2 SRAM die 0.036 to 0.126 microjoules (66x to 19x). So the band {1, 4} words is inside one sector and "the chip pays nothing" is right to 32 bytes and wrong at 64; this is the measured reason W = 8 is free for the chip too (section 10.3) and W = 16 is the only width that costs it. The card is the whole question: the 5090 ran w64 at 71.9 MH/s (-47 percent; the exported kernel at full occupancy, four uint4 loads per item, no request hint) while its own 64-byte probe reached 15.7 G reads a second at its best lane count (-13 percent); 589 GB/s is a third of the stream, so the bind is the request path, not the pins. The 9070 XT (-3 percent) and the M5 Max (0) are free at 64 bytes (measured).
|
||||
|
||||
| Candidate | Verdict | Number (5090 knee, GDDR7 chip) | Label |
|
||||
|---|---|---|---|
|
||||
| The second sector, W = 16 words (w64) | KEEP rank 1, conditional | outcome A (the exported form, -47 percent of rate): 5.3x, dead; B (the probe's -13 percent): 3.4x; C (a one-request form within 5 percent of the 4-byte rate): 3.0x to 3.2x at zero shadow, 2.7x at k 0.5, 2.0x at k 1; the chip +33 percent whatever the card does; the M5 Max and the 9070 XT pay nothing | card measured at two occupancies (unlocked, no watts); chip modelled |
|
||||
| The job that decides it | one PC 1 run, under an hour; asked of the hash lane at 13:5x UK, after kit d | the w64 pack plus the w4 control, the race's occupancy variants, a variant with `ld.global.L2::64B` on the first load (a variantSource anchor rewrite), both clocks, nvidia-smi at 1 Hz; the pass line 95 percent of the 4-byte rate, the kill line under it | owed; default if not run by 15:30 UK: outcome A stands, W = 16 stays `never` |
|
||||
| Tensor tiles | KILL as content; the k band corrected | merchant silicon at a nominal 0.30 to 0.56 pJ per INT8 MAC (MTIA v2 0.51, AI 100 Ultra 0.34, B200 0.44, MI355X 0.56, M4 ANE 0.30 measured; the H800 loop 0.34 measured); against the class's tile (1.5 pJ at the knee) k 0.2 to 0.4, against the dense tile (0.83) 0.4 to 0.7: never above the ALU band; "3x to 30x" needs an INT4 layer at 0.46 V (the VSQ chip reads 0.052 at its nominal 0.67 V); Apple -35 percent at 1,024 tiles kills it | claimed, measured |
|
||||
| The RT core | KILL | traversal implementation-defined (Vulkan, DXR, the NVIDIA forum 14 Oct 2024, Blender's Cycles); the BVH opaque; dedicated units 4 to 33 nJ per ray against 290 to 750 measured board-level on an RTX 2080: k 0.01 to 0.2 | claimed, measured |
|
||||
| Channels per dollar, controller plus PHY | KILL as a lever; KEEP as a model correction (rank 2) | no 28 nm GDDR7 PHY exists (Cadence N3 2024, Innosilicon FinFET; the oldest GDDR6 PHY 12 nm): the record's USD 5 M project and 17 M cap row is unbuildable; the floor is 12 nm GDDR6 at USD 20 to 30 M (cap 70 to 100 M) or N5 GDDR7 at USD 50 to 75 M (cap 170 to 250 M); SK hynix ISSCC 2024 PAM3 I/O 1.10 TX plus 0.76 RX pJ per bit measured, E_mem moves under 0.06 | claimed, modelled |
|
||||
| Proof of latency | KILL | the chain checks values, never time: a 1 s block against a 0.4 microsecond round trip (2.5 M per block); the farm's node is on its LAN; a sequential prefix favours the chip about 4x; no sub-RTT PoW in the literature | measured latency, modelled |
|
||||
| A drawn address map | KILL | firmware, a few hundred gates; no prefetch on a dependent chain; 0 to the chip, 0.8 to 3.2 percent of spread to the cards; literature null | measured spread |
|
||||
| The literature since 2023 | nothing outside the identity | new rows: iPollo V2H 0.14 J per MH, about 13x a 5090 on Ethash (vendor; the precedent's top, not 6.8x); Jasminer X16-Q 5.7x; Pinecone R1X unverified; Qubic pivoted to Doge in April 2026; no random-read joule measured on any current DRAM | claimed |
|
||||
| Own: video decode, row straddle, scratch in DRAM, independent chains per hash | KILL | a decoder-bound card and a ms verifier; 3.1x at -39 percent of rate (dominated); 465 KB sits in L2; the 5090's read rate plateaus over lanes (expected 0 to 5 percent) | measured, modelled |
|
||||
| Own: W = 32 (two items) | hold for v7 behind rank 1 | the chip bandwidth-bound at 110 MH/s and 1.02 microjoules; 2.4x only if the card held its rate, which at 128 B the pins forbid (-20 percent at best) | modelled |
|
||||
|
||||
Nothing reads k above 1 on a measured GPU figure; the second sector reads k about 0.75 to 0.8 on GDDR7 (1.15 nJ of DRAM movement at k = 1 plus the card's fabric share), the highest on the table. Section 10's default (no new lever) is right unless the w64 job passes.
|
||||
|
||||
## 8. Unverified and owed
|
||||
|
||||
- Lane D's family harness (the family gate's measured coverage: the 10,000-era random stratum and the three corner strata, the lossy cap, width 4, shape 64) lands in `family-gate.md` by 17:00 UK and is layer 4's coverage table; this document's layer 4 cites it where it is named and does not restate it.
|
||||
|
||||
- The 5070 Ti tier: no card owned; the hash lane's rented row or the 5070's row scaled (approximate).
|
||||
- The x16 mixer: the verifier and build rows, or the estimate.
|
||||
- The op-mix band: the two re-weighted packs, or the microbench arithmetic.
|
||||
- Every chip figure is the model's; no chip has been measured.
|
||||
|
|
@ -15,6 +15,10 @@ docs/analysis/ci-failures-2026-10-06.md
|
|||
# 7 October 2026: the last research round (the mission lanes and the closed list): internal research written for the
|
||||
# owner, quoting his words and the operations record; the public spec mirror carries none of it
|
||||
docs/analysis/mission
|
||||
# 7 October 2026 (night): the Counter ASIC 4.0 research record (an internal research document: the founder's words, box names, lane records)
|
||||
docs/analysis/counter-asic-4-research.md
|
||||
# 8 October 2026: the class v6 design record (an internal research document: the founder's words, lane records)
|
||||
docs/design/class-v6-rotating-family.md
|
||||
# 7 October 2026: the in-house adversarial cryptanalysis pass (crypto-engage and the adv-* lanes): rule set, plans, reports and
|
||||
# copied logs are research and operations documents, not public export; the public text is the served sentence main landed.
|
||||
docs/analysis/cryptanalysis
|
||||
|
|
|
|||
Loading…
Reference in a new issue