Class v6 10.0f: the external review's five corrections (the programmable chip's 3-year life with the per-life table, the lifetime headline withdrawn; the profitability surface in place of the single threshold; the dataset claim narrowed to a stale machine; proving an opportunity not an exclusion; the served sentence as the review words it; Monero an observed comparison; the window k not a lower bound); the 32-lane window core as the chip assessed (about 2.6x a node ahead, 2.2x node-for-node)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-08 14:30:45 +00:00
parent 44eacd733d
commit ad00108853

View file

@ -261,7 +261,7 @@ The served sentence (`docs/plans/counter-asic-3-public-text-2026-10-07.md`, the
| The convention behind every "x per joule with the shadow" figure | the record's convention (chip-model-v3 and the served text): `k` is the chip core's op cost as a fraction of the GPU's at the same operating point, so when the 5090 locks to 6.2 pJ the chip's core is priced at 6.2 k pJ too; the absolute convention: the chip's core costs its own pJ per op (the k lane's RTL figure), the GPU's its measured pJ at its point | measured GPU side; the chip side claimed until the k lane's RTL rows | the record's convention flatters every chip row at the lock by up to 1.6x (floor lane 3, 13:06 UK: the SRAM die 6.1x against 3.8x at k 0.5, 3.3x against 2.0x at k 1); the served line's figures are at the knee, so they are in the record's convention and need re-reading in the absolute one before main's word; the close carries both, absolute first |
| The schedule's cost beside the SRAM store's edge (main's order) | section 3: the floor and ceiling per tier; main's candidate schedule (6, 10, 14 GiB) priced by lane A by 17:00; the Apple rate cost measured today (-12 / -20 / -22 percent at 2 / 4 / 8 GiB) | measured and modelled | lands at 17:00 |
The wording this lane proposes for main's word, if the SRAM-store reading stands at the 15:45 UK close (the synthesis on the mirror's master by 16:30; the decision itself stays the founder's): "At launch the strongest chip we can price today, a memory-controller chip that stores the dataset, reaches 2.1x per joule against an RTX 5090 at its knee with a core as good as a GPU lane and 3.4x with one three times better, under class v4 from the first block; a 5090 locked at its knee pays 82 W for that shadow work. The strongest chip we can model for 2027 to 2028, a 2 GiB SRAM store on one 2 nm die, would reach 3x to 5x with that same shadow and 17x without it, at an N2 project of USD 100 M to 500 M and 18 months or more; the dataset's size is the one lever on it, and the schedule says how it grows. Class v5 makes the dataset the chain's own state, so a chip that does not follow the chain is wrong on every item. The model and every measurement are public." Every number in it is labelled in this document and the chip model; nothing served moves on this lane's word.
The wording the external review gives (10.0f item 5) is the one proposed for main's word; this lane's earlier proposal, kept for the record, if the SRAM-store reading stands at the 15:45 UK close (the synthesis on the mirror's master by 16:30; the decision itself stays the founder's): "At launch the strongest chip we can price today, a memory-controller chip that stores the dataset, reaches 2.1x per joule against an RTX 5090 at its knee with a core as good as a GPU lane and 3.4x with one three times better, under class v4 from the first block; a 5090 locked at its knee pays 82 W for that shadow work. The strongest chip we can model for 2027 to 2028, a 2 GiB SRAM store on one 2 nm die, would reach 3x to 5x with that same shadow and 17x without it, at an N2 project of USD 100 M to 500 M and 18 months or more; the dataset's size is the one lever on it, and the schedule says how it grows. Class v5 makes the dataset the chain's own state, so a chip that does not follow the chain is wrong on every item. The model and every measurement are public." Every number in it is labelled in this document and the chip model; nothing served moves on this lane's word.
## 7c. Beyond the four: the invention lane's layers 5 to 7 and its rejected list (lane C, `docs/analysis/class-v6/invention.md` on master at a9f03598, 10:48 UTC; the census script `tools/attack/v6-invention/v6inv-census.sh`)
@ -305,11 +305,11 @@ The honest floor in one line, on the synthesised core: against the chip anyone c
The capex wall's two thresholds (lane 3, 10.3, re-folded 13:18 UK on lane 5's corrected project floor of USD 20 to 75 M for the cheapest DRAM-board chip): no rational chip project of any kind below about USD 23 M a year of miner revenue (IGN 0.03 at launch emission, USD 62 K a day, a cap of about USD 60 M); every DRAM-board project at a third of the network above about USD 340 M a year (IGN 0.44, USD 0.93 M a day, a cap of about USD 0.9 B), where the SRAM project also starts. The project cost moves the threshold 5x, the chip's edge 1.4x.
The served line in one sentence, on the synthesised core with the 64-register window (10.0e), in the two units: a chip that stores the dataset reaches about 2.2x to 2.4x per joule against the honest NVIDIA tiers at their knee (the 5080 2.2x, the 5090 2.4x with a chip a node ahead; 2.0x against a chip on the GPU's own node, 2.8x two nodes ahead; 2.8x on the 5090 without the window) and under 2x only against the Apple tier (the M5 Max 1.5x), about 4x and 2.6x if it is built on an N2 SRAM store for USD 100 M or more, and under 1x per hash over its 180-day class life only above about USD 300 M a year of miner revenue (10.0a), with Monero's measured 1.0x to 1.5x beside it; no rotating layer moves the per-joule figures, and layer 3's class life is what sets the per-hash one; the layers decide which chip can be built and how long its tape-out lives. The worst case beside it, if a chip's core costs no more than its bare units (the floor k of 10.2): 4x and 2.4x against the DRAM board, 14x and 9x against the SRAM die (lane 3's line: 6x to 10x at the full shadow, never under 2x). In the other convention, the card's whole 330,000-op latency shadow priced on the chip at the synthesised core (lane 3's 10.3 table): the SRAM die 3.5x against a 5090 at stock and 2.2x at its knee (4.8x and 3.0x with a 2 nm core), the DRAM board 2.6x at stock. The record's claimed band (k 0.3 to 0.8) is confirmed from the chip side by the core row (0.56 at the lock), so the served 2.1x at k = 1 is the conservative end and 2.8x the measured-core reading.
The served line as the external review words it is in 10.0f and overrides this paragraph's wording. This lane's earlier sentence, kept for the record: a chip that stores the dataset reaches about 2.2x to 2.4x per joule against the honest NVIDIA tiers at their knee (the 5080 2.2x, the 5090 2.4x with a chip a node ahead; 2.0x against a chip on the GPU's own node, 2.8x two nodes ahead; 2.8x on the 5090 without the window) and under 2x only against the Apple tier (the M5 Max 1.5x), about 4x and 2.6x if it is built on an N2 SRAM store for USD 100 M or more, and under 1x per hash over its 180-day class life only above about USD 300 M a year of miner revenue (10.0a), with Monero's measured 1.0x to 1.5x beside it; no rotating layer moves the per-joule figures, and layer 3's class life is what sets the per-hash one; the layers decide which chip can be built and how long its tape-out lives. The worst case beside it, if a chip's core costs no more than its bare units (the floor k of 10.2): 4x and 2.4x against the DRAM board, 14x and 9x against the SRAM die (lane 3's line: 6x to 10x at the full shadow, never under 2x). In the other convention, the card's whole 330,000-op latency shadow priced on the chip at the synthesised core (lane 3's 10.3 table): the SRAM die 3.5x against a 5090 at stock and 2.2x at its knee (4.8x and 3.0x with a 2 nm core), the DRAM board 2.6x at stock. The record's claimed band (k 0.3 to 0.8) is confirmed from the chip side by the core row (0.56 at the lock), so the served 2.1x at k = 1 is the conservative end and 2.8x the measured-core reading.
The three class v6 changes it implies: (1) the op mix moves to the band's best mix as the genesis table (add 16, xor 14, mad 12, rotl 11, sub 10, rotr 10, shfl 4, mul 4, mulhi 2, or 0; the k lane's optimiser, 15:2x UK: +42 percent of k_eff on the unit floors, +15 percent on the core; the shuffle at its floor because the chip's butterfly is k 0.011 to 0.021, the lowest of every family; lane D: 0 of 3,000 eras exhausted under the band), subject to one acceptance pass of that table through lane D's harness, the census lane's passed table the default; (2) no SM-sparse default (measured dead on the 5090, 4090 and H100 at 1.3 to 3.9 percent at best; `--sm-sparse auto` ships off by default, on in the Efficiency and Balanced tiers at its measured 1 to 2.5 percent, 10.1 item 6; the per-watt gift of the hot table is the chip's, 20.3c); (3) the dataset schedule 5.5 / 8.5 / 11.5 GiB with the read width pinned at 4 words (8 measured not free, 16 never); and a fourth the sweep adds, (4) the 64-register window per lane as the core shape (k 0.78 at the lock against 0.56; 10.0c, the edge table 10.0e), its GPU side modelled until the generator carries a 64-entry init and fold rule (a half-day line) and a pack is measured (the chip's ticket USD 1,500 / 2,500 / 3,000; a tuned 5090 pays 1 to 5 percent per step; about a quarter of today's cards by count per step).
#### 10.0a The second headline unit: cost per hash over the chip's life (the founder's question, 14:1x UK; modelled on measured card rows)
#### 10.0a The second unit: cost per hash over the chip's life (the founder's question, 14:1x UK; modelled on measured card rows; SUPERSEDED as a headline by 10.0f item 1: the 180-day life below is the fixed-lane chip's row, and the programmable chip's 3-year life is the default)
Per joule is the chip's edge while it runs; per hash over its life adds the project it must pay back inside one class life, which layer 3 fixes at 180 days (section 4). The model, every input labelled: the GPU is a 5090 at its 1,300 MHz knee on class v4 (2.33 microjoules per hash and 127 MH/s, measured), bought at USD 2,000 (claimed, approximate) and written off over three years with no residual (conservative against the card: it has a resale value and a chip has none when its class dies), electricity at USD 0.10 per kWh (assumed): capex 1.7e-13 USD per hash plus 0.65e-13 of electricity = **2.3e-13 USD per hash**. The chip is the cheapest DRAM-board project (floor lane 5's corrected floor, USD 20 M at 12 nm GDDR6 to 75 M at N5 GDDR7, claimed), written off over one 180-day class life, at E_chip 0.82 microjoules (the synthesised core, 10.2; 0.23e-13 USD per hash of electricity). The chip's lifetime cost per hash is below the GPU's only when its fleet's hashes over the 180 days exceed C / (2.3e-13 - 0.23e-13): **6.2 TH/s for 180 days at USD 20 M, 23 TH/s at USD 75 M** (49,000 and 180,000 knee-locked 5090s' worth). A fleet that size at a third of the network means a network of 19 to 70 TH/s, and a network holds that hash only if its miner revenue covers the GPU's own cost per hash on the whole of it (else the honest miners leave): **USD 140 M a year at the USD 20 M project, USD 500 M at the USD 75 M project; about USD 300 M at the midpoint, which is lane 3's USD 340 M per-joule threshold read the other way.** Below that revenue the chip's cost per hash over its life is above the GPU's at every per-joule edge in the close table.
@ -321,7 +321,7 @@ At lane 3's three prices (assumed; year-1 emission 0.77 B IGN to miners; the chi
| 0.10 | 77 M | 10.6 TH/s | 3.2 TH/s | 1.9x / 6.7x | modelled |
| 1.00 | 770 M | 106 TH/s | 32 TH/s | 0.28x / 0.77x | modelled |
Read with 10.0's per-joule table: at IGN 0.10 the chip is 2.8x per joule at the knee and 1.9x to 6.7x per hash over its life; at IGN 1.00 both units favour the chip (0.28x to 0.77x per hash). **The served sentence in the two units: under 3x per joule at the knee, under 1x per hash over its life only above about USD 300 M a year of miner revenue (IGN 0.4).** What breaks it: a class life longer than 180 days (the chip's capex spreads; layer 3's epoch is the lever), a chip that survives a class flip (the rotating layers' whole purpose is that it does not), a GPU price above USD 2,000 or a GPU life under three years (both move the GPU's figure up, in the chip's favour, by at most 1.5x on the capex term).
Read with 10.0's per-joule table: at IGN 0.10 the chip is 2.8x per joule at the knee and 1.9x to 6.7x per hash over its life; at IGN 1.00 both units favour the chip (0.28x to 0.77x per hash). The sentence this gave (under 1x per hash over its life only above about USD 300 M a year) holds for a fixed-lane chip only; see 10.0f. What breaks it: a class life longer than 180 days (the chip's capex spreads; layer 3's epoch is the lever), a chip that survives a class flip (the rotating layers' whole purpose is that it does not), a GPU price above USD 2,000 or a GPU life under three years (both move the GPU's figure up, in the chip's favour, by at most 1.5x on the capex term).
#### 10.0b The precedents, sourced (the founder's question; floor lane 5's literature pass, 14:1x UK, every figure with its URL, read 8 October 2026; labels as marked; this replaces this lane's first row of 14:0x, which had RandomX chip-free, a premise the pass corrected)
@ -361,7 +361,7 @@ The chip's class v4 shadow on the window core is 102,100 x 4.8 pJ = 0.49 microjo
| Apple M5 Max (1.40) | **1.5x** | 1.7x | 2.6x | the same |
| RTX 5090 stock (3.36), for the record | 3.5x | 4.1x | 6.2x | the same |
**Stated plainly: with the window the GPU-tier floor is about 2.2x to 2.4x per joule against the chip anyone can build, and under 2x only against the Apple tier (1.5x); Monero's RandomX measured 1.0x to 1.5x beside it (10.0b, the Antminer X5 after 46 months). The lifetime cost-per-hash sentence stays the headline: under 1x per hash over a 180-day class life only above about USD 300 M a year of miner revenue (10.0a).** Against the SRAM die the window reads 3.8x to 4.3x on the NVIDIA floors and 2.6x on the M5 Max, with the N2 project's ticket (10.3) and its capex wall beside it.
**Stated plainly: with the window the GPU-tier floor is about 2.2x to 2.4x per joule against the chip anyone can build (about 2.6x on the 5090 and 2.4x on the 5080 against the 32-lane window core the adversary would build, 10.0f), and under 2x only against the Apple tier (1.5x); Monero's RandomX measured 1.0x to 1.5x beside it (10.0b, the Antminer X5 after 46 months; an observed comparison carrying its devices and operating points, not a ceiling). The lifetime cost-per-hash sentence is withdrawn as a headline by the external review (10.0f item 1): it held for a fixed-lane chip only; for the programmable chip at a 3-year life the per-hash unit reads under 1x above about USD 23 M to 83 M a year.** Against the SRAM die the window reads 3.8x to 4.3x on the NVIDIA floors and 2.6x on the M5 Max, with the N2 project's ticket (10.3) and its capex wall beside it.
The node column (the k lane, 15:0x UK; the factors claimed from TSMC's per-node headlines: N7 to N5 x0.70, N5 to N3E x0.72, N3E to N2 x0.72; ASAP7 is a predictive 7 nm-class PDK so "N5" is ASAP7 x0.70; the 5090 and 4090 are TSMC 4N, N5 class; the M5 Max N3). Absolute; the GDDR7 board at the lock = 2.33 / (0.466 + 102,100 x e_chip):
@ -381,6 +381,29 @@ The 32-lane shuffle (the butterfly over the 1 KB window, routed with SPEF): 1.24
The mix optimiser (`tools/chip-model/rtl/flow/mix.py`, the layer-1 band: B = 4 on the injecting families, the lossy families at or under base, or + mul + mulhi at most 22, the shuffle capped at 8; exhaustive over the corners): on the unit floors at N3 and the lock the class v4 mix reads k_eff 0.097 and the best mix in the band 0.137 (+42 percent): **add 16, xor 14, mad 12, rotl 11, sub 10, rotr 10, shfl 4, mul 4, mulhi 2, or 0 (sum 83)**; the worst in the band 0.074 (shuffle and mulhi heavy). The census lane's re-weight (add 13, xor 11, mul 6, mad 10, shfl 8, rotl 8, sub 7, mulhi 2, rotr 6, or 4; PASS on 256 seeds) sits at about 0.115, two thirds of the way. On the core the per-op overhead compresses the spread: the best mix lifts the core's k about 15 percent (0.56 to about 0.64 at the lock on the 8-lane core; the window core 0.78 to about 0.9). The card's side of that mix (kit d, section 2): fewer shuffles and less mulhi lower the card's block premium too (the multiply-heavy table 23 percent under the shuffle-heavy one at the knee), so the edge moves about 0.1x in the card's favour at the knee (2.8x to about 2.7x on the base core, 2.4x to about 2.3x with the window; arithmetic on the measured and synthesised rows). **The close's change (1) is amended: the op mix moves from "class v4's, held" to the band's best mix as the genesis table, subject to one acceptance pass of that exact table through lane D's family harness (the census lane's neighbouring re-weight passed 256 of 256 both ways; the default if the pass is not run is the census lane's table at k_eff 0.115).**
#### 10.0f The external review's five corrections, taken (the coordinator, 15:4x UK; the founder accepts the review; this section overrides the close's wording above where they differ)
**(1) Lifetime.** A programmable chip survives family epochs on firmware (the litepaper and ledger M32 already concede it), so the close's headline "under 1x per hash over its 180-day life" is true only for a fixed-lane chip and comes out. The chip to assess is programmable, its productive life assumed 3 years (the default), shown at 0.5, 1, 2 and 3 years; a family transition earns an obsolescence benefit ONLY where a demonstrated loss of competitiveness exists, which means naming the physical resource the next family needs that the chip cannot supply economically (constants, rotations, op order and frequencies do not count; a register window the chip did not build, a dataset larger than its board, a read width wider than its sector would). The 10.0a model re-run per life (the same inputs: the knee-locked 5090 at 2.3e-13 USD per hash all-in; the chip at E_chip 0.82 microjoules, 0.23e-13 of electricity; the project USD 20 M to 75 M; the fleet at a third of a network sized to the GPU's own cost per hash):
| Productive life | Break-even chip fleet (USD 20 M / 75 M project) | Miner revenue a year above which the chip is under 1x per hash (20 M / 75 M) | Label |
|---|---|---|---|
| 0.5 year (a fixed-lane chip that dies at the family epoch) | 6.2 / 23 TH/s | USD 140 M / 500 M | modelled |
| 1 year | 3.1 / 11.6 TH/s | 70 M / 250 M | modelled |
| 2 years | 1.55 / 5.8 TH/s | 35 M / 125 M | modelled |
| 3 years (the programmable chip, the default) | 1.03 / 3.9 TH/s | 23 M / 83 M | modelled |
So for the programmable chip the lifetime unit says: above about USD 23 M to 83 M a year of miner revenue the chip's cost per hash is under the GPU's, which is lane 3's lower capex wall read the other way, and the per-hash unit adds nothing to the per-joule one beyond that; the headline is the per-joule figure and the economics below.
**(2) Economics.** The single USD 340 M threshold comes out. In its place a profitability surface (lane 3's model is the base, its 180-day life one row of the surface; asked of lane 3, the row owed): NPV with the development cost, the initial fleet capital, the captured share q, the operating margin, electricity, the pre-production period with zero revenue, discounting and residual value; the operator and the manufacturer modelled separately (self-mining, hardware sales, hybrid; the development cost shared across entrants; a first design against a revision); GPU miners allowed to enter and exit. The result is stated as the set of conditions under which development is attractive, never as "the chain stays below X". Until the surface lands, the honest statement from lane 3's rows is conditional: a USD 20 M DRAM-board project taking the whole chain is attractive above about USD 23 M a year of miner revenue at a 10 percent discount; the same project at a third of the hash above about USD 75 M to 91 M; a USD 75 M project at a third above about USD 280 M to 340 M; a USD 100 M SRAM project at a third above about USD 330 M to 450 M; every figure moves 5x with the project cost and 1.4x with the chip's edge, and a longer productive life than the 180 days assumed there lowers each by the ratio of the lives.
**(3) Dataset.** The class v5 claim is narrowed: a chain-state dataset makes a STALE machine wrong on every item; it does not exclude a specialised machine with a host and external memory that keeps the dataset current. What layer 2 prices instead is the update bandwidth, the sync and the storage that keeping it current costs: the state delta per block (the touched leaves times 64 bytes; at the devnet's 150 tx/s and a few leaves per transaction, of the order of 100 KB a second, approximate), the lazy derivation of the touched items per epoch on the host (the verifier's own cost, under 10 ms per warp), and a board that holds the schedule's size (5.5 / 8.5 / 11.5 GiB) beside the chip; against that the layer's measured honest-side cost (section 3.3). The chip rows of section 3.2 stand for the stateless chip only; the hosted chip's extra cost is the host, the link and the board, priced in 10.3's ticket terms (owed as a row, approximate).
**(4) Proving.** Useful proving does not bind mining to a GPU: an ASIC-plus-GPU operator is inside the adversary model, and proving is an opportunity for GPU owners, not an exclusion. Thread 5 of the research file (proof of useful work) and 10.5's "proof of latency" kill stand with that reading.
**(5) The served sentence, as the review words it:** "Class v6 retains the 64-register window. Current modelling estimates a 2.2x to 2.4x energy-efficiency advantage for the strongest specialised designs assessed against the GPU tier (2.0x on the GPU's own node). The long-program and select-tree proposals were rejected. Economic resistance depends on development cost, deployment economics and productive hardware lifetime; family transitions receive an obsolescence benefit only where a loss of competitiveness is demonstrated; programmable multi-epoch designs are included in the assessment." Two marks the review adds: the Monero figure (1.0x to 1.5x) carries its devices, software, operating points and power boundaries and is an observed comparison, not a ceiling; and the register-window k (0.78 on the 8-lane core) is synthesis-derived and NOT a lower bound (the flop-array lesson of 10.0c: a chip maker's own implementation can come in under it).
**The chip to assess (the coordinator, 15:4x UK, on the k lane's 32-lane rows):** the adversary's design is free (lane count, clock, pipeline, banking), never the GPU-shaped one. The imem and sequencer amortise across lanes, so a chip maker picks the wider core: the 32-lane 32-register core reads k 0.45 at N3 (0.63 node-for-node) against the 8-lane 0.56, and the register file is the one per-lane term that does not amortise (+2.8 pJ from 32 to 64 entries). So the chip assessed is the 32-lane core with the 64-register window, about k 0.63 at N3 and 0.88 node-for-node (pending the adversarial re-optimised row by 18:00 UK), and the GDDR7 board's edge at the lock moves to **about 2.6x on the 5090 a node ahead and about 2.2x node-for-node** (the 5080 about 2.4x and 2.0x; the M5 Max about 1.6x and 1.4x; arithmetic on the measured card rows and the synthesised k, scaling claimed). The shuffle is the lowest-k family (0.021 routed); the optimiser's weights (add 16, xor 14, mad 12, rotl 11, sub 10, rotr 10, shfl 4, mul 4, mulhi 2, or 0) replace the census lane's re-weight as the class v6 mix target once censused (lane D's pass asked).
#### 10.0d The rotation schedule the close adopts (the rotation lane, `docs/design/class-rotation-four-layers.md` on class-v6-rotation at bd43f808, build-3, gate green, 14:4x UK; one line per layer; both of this document's constraints held: the 180-day family epoch not shorter, W = 4 not drawn)
| Layer | Boundaries a year | What it draws, from where | Exposure per boundary (this document's units) | Chip | Label |