Class v6: 10.3 the SRAM die and the dataset floor (floor lane 3: the read width the only wire lever, 66x at 4 bytes at zero shadow, the shadow the whole hold at 2.7x to 3.0x at k = 1; W pinned at 4 words at genesis; the floor as a ticket, 5.5 / 8.5 / 11.5 GiB; the capex wall's p* table)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

Documents-only replay of 7aa9f1898 (counter-asic-4) for the box mirror master
This commit is contained in:
igneum-labs 2026-10-08 12:03:48 +00:00
parent b3774f4de9
commit 761f2cfc7b

View file

@ -4,7 +4,7 @@
## 0. For the founder tonight: what each layer does to a chip on its release day, and what it costs a 5090, a 5070 Ti and an M5 Max
**The lead, as main ordered it (11:2x UK): the strongest chip in the five-year window is not a DRAM chip and no rotating layer reaches it.** Lane B's reading (`docs/analysis/class-v6/hardware-future.md`, master 34f63b3c; carried into the chip model as section 5.12): a 2 GiB SRAM full store on one reticle of merchant N2 (452 mm^2 of macro at 38 Mb/mm^2, claimed) reads about 2,100 MH/s per die at 300 W, 0.14 microjoules per hash, 17x the 5090 at its stock point per joule at zero shadow (8x to 30x on the read-energy band), 5.6x the M5 Max; with the class v4 shadow 4.8x at a core half as costly as a GPU lane and 2.7x at `k = 1`, and at the honest card's whole latency shadow 3.7x and 2.0x; USD 400 to 600 of silicon per die (USD 0.25 to 0.4 per MH/s), an N2 project of USD 100 M to 500 M and 18 to 24 months (claimed), which on the mission lane's model is a break-even cap of about USD 330 M to 1.7 B in the chain's first two years. Beside it a custom HBM4E base die (2028 or later) at 6.5x to 14x, untouched by all four layers; per-bank processing in memory structurally blind to the hash (1.6 percent of reads in-bank at 2 GiB). All modelled on the chip model's method; the GPU side measured. Against the SRAM store the honest range with the shadow is 3x to 5x, and the one layer that answers it is layer 2 as a FLOOR that grows faster than SRAM cost falls: each doubling of the floor adds a die (4 GiB two dies at about USD 1,000 and 15x, 8 GiB four dies at USD 2,000 to 2,500 and 13x) while every card tier pays in device memory and, on Apple, in rate. The schedule's cost is priced in section 3 (lane A's per-tier table and main's candidate schedule, 6 GiB at the v6 epoch, 10 two years on, 14 at four, land there by 17:00 UK) and the served chip line is reviewed against this reading in section 9 at 20:00.
**The lead, as main ordered it (11:2x UK): the strongest chip in the five-year window is not a DRAM chip and no rotating layer reaches it.** Lane B's reading (`docs/analysis/class-v6/hardware-future.md`, master 34f63b3c; carried into the chip model as section 5.12): a 2 GiB SRAM full store on one reticle of merchant N2 (452 mm^2 of macro at 38 Mb/mm^2, claimed) reads about 2,100 MH/s per die at 300 W, 0.14 microjoules per hash, 17x the 5090 at its stock point per joule at zero shadow (8x to 30x on the read-energy band), 5.6x the M5 Max; with the class v4 shadow 4.8x at a core half as costly as a GPU lane and 2.7x at `k = 1`, and at the honest card's whole latency shadow 3.7x and 2.0x; USD 400 to 600 of silicon per die (USD 0.25 to 0.4 per MH/s), an N2 project of USD 100 M to 500 M and 18 to 24 months (claimed), which on the mission lane's model is a break-even cap of about USD 330 M to 1.7 B in the chain's first two years. Beside it a custom HBM4E base die (2028 or later) at 6.5x to 14x, untouched by all four layers; per-bank processing in memory structurally blind to the hash (1.6 percent of reads in-bank at 2 GiB). All modelled on the chip model's method; the GPU side measured. Against the SRAM store at the hash's own width floor lane 3 reads 66x at zero shadow (lane B's 17x was the 64-byte row) and 2.7x to 3.0x at k = 1 and 5.0x to 5.7x at k = 0.5 with the shadow (section 10.3), so the honest range with the shadow is 3x to 6x and the shadow is the whole hold; and the one layer that answers it is layer 2 as a FLOOR that grows faster than SRAM cost falls: each doubling of the floor adds a die (4 GiB two dies at about USD 1,000 and 15x, 8 GiB four dies at USD 2,000 to 2,500 and 13x) while every card tier pays in device memory and, on Apple, in rate. The schedule's cost is priced in section 3 (lane A's per-tier table and main's candidate schedule, 6 GiB at the v6 epoch, 10 two years on, 14 at four, land there by 17:00 UK) and the served chip line is reviewed against this reading in section 9 at 20:00.
The honest frame first, from last night's measured close (`docs/analysis/counter-asic-4-research.md` sections 2 and 15.1a). Two chips exist in the model. A **fixed-function chip** wires one hash: the mixer's rounds, the op mix's lane ratios, the read width, the program length and the block shape are silicon. A **GPU-like chip** stores the dataset in commodity DRAM and runs the hour's program on a programmable integer core (the `f = 1` chip of chip-model-v3 section 5, the chip anyone builds); for it every drawn parameter is firmware. The four layers below render the FIRST kind useless on the day a parameter leaves its wired value, and move nothing for the second kind except the size of the core it must carry and the capex that forces. What the second kind keeps is the identity of section 2: at zero shadow premium the card's whole-card energy over the chip's memory energy (3.6x on a 5090 at its knee, measured card, modelled chip), and with the class v4 shadow 2.1x at a core as good as a GPU lane. No rotation changes that; the layers change which chip can be built and how long its tape-out lives.
@ -56,7 +56,7 @@ The band a card sees across the draw's range, from the hash lane's cost rows (12
|---|---|---|---|---|---|---|
| Mixer applications per round `m` | **{4, 8} at genesis; 16 in the list as `admissible: false` (the ladder's rule: a flag in the genesis list, flipped only by the 90 percent upgrade path once the 2019-class core measures it)** | x4 and x8 measured (mixer-x4.md 6.4: 1.92 and 2.79 ms per unit on a loaded M5 Max core; the build unmoved on every discrete card, latency-bound); **x16 MEASURED (today's class v3 stream, the same id f5e904bc5d148926 for x8 and x16). Cold alone: build-1 core 4 x8 4.67 ms per warp, x16 8.69 (the hash lane, 10:13 UTC); build-3 core 4 (AX102, a faster core class) x8 3.48, x16 6.53 (the build-server lane, 11:01 UTC). With the SMT sibling loaded for the row's whole length (the method that counts: a sibling bench looped for the row, not run once, which ends in about a second and leaves the verify phase sibling-idle; the hash lane's earlier 5.41 / 10.11 / 9.04 to 9.10 ms rows were of that kind and are withdrawn): build-3 x8 6.24, x16 11.44 ms, about 1.8x on both classes. The multiplier is 1.86x on the verifier (the mixer IS the verifier's cost); chip-model-v3's estimate (9 ms on a 2019-class core) stands within the cold rows** | the `f = 0` recompute chip's rate halves per doubling (0.31x bare at x8, 0.16x at x16, modelled); the `f = 1` chip unmoved (it recomputes nothing) | rate 0 / 0 / 0 (latency-bound, measured at x4 and x8); the daily build 42 / pending / 29 ms at x8 (measured); the x16 build on the 5090 rides the 12:45 UK job | +1.9x per doubling measured (x8 to x16); **x16 at 11.4 ms loaded is over the 10 ms gate on the measured row, so `admissible: false` stands on the measurement, not only on the chip model's margin; the O-1.14 laptop run can only tighten it** | adv-mixer-3's margin (every statistic clean from k = 1, the SAT ladder k = 1 solved, k = 2..4 timeout): x4 keeps the margin by that report's reading; the band {4, 8} is the measured one |
| Op-mix weights (the ten non-load families) | **B = 4 points on the injecting families only (add, sub, xor, mad, shfl, rotl, rotr); the lossy families (or, mul, mulhi) capped at their base, the ring-A rule `or + mul + mulhi` at most the table's 18 plus B, which keeps the per-candidate rejection r under 0.85 (r^256 under 1e-18)**; the shuffle weight capped at its class v4 value on the energy side | lane D's coverage run (4,900 drawn eras on three boxes, 11:3x UK, measured): at the lossy corner of B = 4 (or, mul and mulhi all at +4, 30 of 75 lossy against 18) r is 0.956 (the shipped 0.681) and 8 of 663 eras exhaust the 256-attempt cap (mean attempt 18.6, max 252), so 1.2 percent of that corner's epochs would take the last-resort program, which fails rule (a) in 9 percent of seeds (adv-accept-3); the random stratum with every weight drawn reads r 0.718, max attempt 107, 0 exhaustions in 1,444. The energy side (15.1a): shfl 55.8 pJ per op, mulhi 39.6, prmt 22.3, lop3 24.1, mul 13.9, arx 11.3; a shuffle-heavy table raises the premium per instruction up to 2x (modelled), a multiply-heavy one 1.1x | a chip that specialised its lane ratio loses the ratio; a general core nothing | the premium per instruction moves with the mix within the capped band (the cost rows: at N = 100,000 the 5090's +146 W stock becomes up to +160 W multiply-heavy; shuffle-heavy excluded by the cap); the 5070 Ti scaled by 83/146, the M5 Max by 16/146; the two packs' measured rows replace this when they land | +0 ms | the known-failed test: the 8 of 663 exhaustions at the uncapped lossy corner, which the capped band must read as 0 |
| Read width `W` | {1, 4} words | w4 and w16 measured 5 October: the 5090 139.8 against 136.1 MH/s, the 9070 XT 17.90 against 18.15; w64 bandwidth-bound (71.9 MH/s on the 5090) | a 4-byte-granularity controller moves 4x the bytes at a w16 era (the Ren-Devadas lever reversed); the `f = 1` chip's energy per read rises toward the card's | 0 / pending / 0 (measured w16 on the M5 Max: within 1 percent) | 0 | none |
| Read width `W` | **pinned at 4 words (16 bytes) at genesis, not drawn** (floor lane 3, 10.3: the width is the only wire lever on the SRAM die, 66x at 4 bytes against 44x at 16 at zero shadow; 8 words after the owed PC 1 and Mac rows; 16 never) | w4 and w16 measured 5 October: the 5090 139.8 against 136.1 MH/s, the 9070 XT 17.90 against 18.15, the M5 Max within 1 percent; w64 bandwidth-bound (71.9 MH/s on the 5090) | the SRAM die's energy per read rises with the bits moved (0.25 to 0.38 nJ); the DRAM chip's toward the card's | 0 / 0 / 0 (measured) | 0 | the W = 8 rows (owed) |
| Shadow block shape | 64 to 256 instructions per block, the pass count the ladder's | 64-instruction blocks ran 2.5 to 3.5 percent FASTER than 256 on the 5090 and the M5 Max (measured 6 October); 1,024 cost the M5 Max 17 percent | nothing for any chip (the work is the same) | +2.5 to 0 percent / pending / +2.5 to 0 | 0 | none |
| **The index fold (a ring-A design rule of layer 1, every era's draw passes through it): `load_index` folds a product's low bits before the stride rotation, so no era's R lands a biased product bit on an address bit** | every era; lane D's coverage (11:3x UK): the class is at HALF the family's epochs, not a corner: 52 to 58 percent of accepted programs in every stratum carry one site whose address bit R (or R+1, R+2) is biased over 6 sigma at 2^20, 33 to 40 percent over 100 sigma, the worst z 1,024, on programs (c''') passes; the F8 tail's bucket excess is the same mechanism at scale (+94 to +128 sigma at the biased bit) | lane D's family harness (build-1, 10:21 UTC, 16 drawn eras, every layer-1 parameter from the era's stream, the chain draw through the real rule, reads at the rule's own 2^20 sample): the index-bit bias fires HARD on 7 of 16 eras, abs z 130 to 511 at one site, every one at address bit R or R+1 (a product's bit 0 at P = 0.25 on era 15, R = 25, z -511; a product's bit 1 at 3/8 on era 13; an or-shaped source at 5/8 on era 6); the other 9 clean under abs z 3.8; the (c''') ratio sees none of it (0.9954 to 1.0000): adv-cache-2's era-stride class measured on class v5 accepted programs at the acceptance's own sample | a chip holding the favoured half of that site's window serves 75 percent of its reads instead of 50: about 1.6 percent of a hash's reads at f = 1/2 for one site, zero at f = 1 (the partial store already costs 1.26x the ops, chip-model 5.4): the f = 1 verdict does not move; the row is an auditor's flag on "uniform random reads", not a chip lever | nothing: the fold is one xor-rotate on the address path, measured as 0 on every card by the era-layout rows (the index form is the era draw's own) | 0 | the fold's form in `load_index` (fold the product's low bits before the rotation) and its vectors; a 6-sigma REFUSAL in layer 4 is not the lever: it would redraw about 40 percent of epochs (7 of 16 eras); the known-failed test is lane D's 7 of 16 eras at the rule's sample, which must read 0 of 16 with the fold |
| Program length N | not drawn: the ladder's signal (latency-ladder.md) | an unconditional draw retires the Apple tier at 200,000 (measured -10 percent) | the core sized for the ladder's admissible top (rung 2, 199,600 ops) | the ladder's rows | the ladder's rows | none |
@ -272,6 +272,42 @@ The question the founder set: the stored-dataset chip pays the DRAM's own energy
The close (20:00): the honest floor per tier against each chip row, the recommended class v6 changes (the op mix, the SM-sparse default, the dataset schedule), the served chip line as a measured range, and what stays for v7, written here from the lanes' rows as they land.
### 10.3 The SRAM die and the dataset floor (floor lane 3, `docs/analysis/class-v6/floor/sram-and-floor.md` on class-v6-floor-sram at 708d01b4, first reading 12:0x UTC; full by 19:30 UK; everything chip-side modelled on lane B's wire figure, 1.3 pJ per bit plus 0.1 nJ per macro access and 0.05 nJ of controller; the GPU side measured)
**The read width is the only wire lever on the SRAM die, and the floor is a ticket, not a joule.** Lane B's 1.0 nJ was the 64-byte row; the die reads what the hash asks for, and the wire scales with the bits moved while the macro access does not:
| W words | Bits per read | Chip E_read | GH/s per die (300 W) | Zero shadow against the 5090 | With the class v4 shadow, k 0.5 / 1 | At the card's whole shadow (F 2.0), k 0.5 / 1 | The honest card's cost | Label |
|---|---|---|---|---|---|---|---|---|
| 1 (4 bytes, today) | 80 | 0.25 nJ | 8.3 | **66x** | 5.7x / 3.0x | 4.1x / 2.1x | 0 | modelled chip; measured card |
| 4 (16 bytes) | 176 | 0.38 | 5.6 | 44x | 5.6x / 2.9x | 4.0x / 2.1x | the 5090 +2.7 percent, the 9070 XT -1.4, the M5 Max within 1 (measured 5 October) | measured card |
| 8 (32 bytes, the GDDR7 sector the 5090 fetches anyway) | 304 | 0.55 | 3.9 | 31x | 5.4x / 2.9x | 4.0x / 2.0x | unmeasured (one PC 1 and one Mac job owed) | modelled |
| 16 (64 bytes, lane B's record) | 560 | 0.88 | 2.4 | 19x (the record's 17x) | 5.0x / 2.7x | 3.8x / 2.0x | the 5090 -47 percent: dead | measured card |
Meaning: at the hash's own width the die is 3.5x stronger than the record said at zero shadow; with the shadow on, every row sits at 5.0x to 5.7x (k 0.5) and 2.7x to 3.0x (k 1): **the shadow is the whole hold, the memory moves it 0.2x to 0.7x.** The fold is dst-keyed (`verify::fold_words`), so a wide read cannot be pre-folded; dependent chains per step leave the ratio unchanged (per-read on both sides); banking plus a sequencer cannot localise (uniform on a window of at least 256 MiB, the 16 sites alternating; moving the lane costs about 290 bits, which caps the wire lever at about W = 8); the per-site window shrink buys the die and the DRAM chip nothing; a hop at the honest widths is 0.04 to 0.15 nJ, so a ten-die store still reads 25x to 58x at zero shadow. This moves layer 1's width row: **pin W = 4 (16 bytes) at genesis (measured free on all three vendors; the die from 66x to 44x), W = 8 after the two owed jobs, never 16.**
The floor as a ticket: SRAM is flat at about USD 250 per GiB to 2031 (density +6 to 11 percent per node against dearer wafers; claimed, approximate), so USD 5,000 of silicon per store is 20 GiB at N2 and about 21 GiB in 2031, which retires every card under 32 GB and every Mac under 64 GB; the constraint and "fewest cards" cannot both hold. The per-MH/s does not rise with the floor (every die powered: USD 0.25 to 0.5 per MH/s at any size); the floor raises the minimum ticket only.
| Floor | Reticles 2026 / 2031 | USD of silicon | Tiers out (worst / best working set) | Label |
|---|---|---|---|---|
| 5.5 GiB (the v6 epoch) | 3 / 3 | 1,500 | the 8 GB tier in the worst case only / none | modelled |
| 8.5 GiB (year 2) | 5 / 4 | 2,500 | 8 GB, 12 GB cache-resident, Apple 16 GB / 8 GB, Apple 16 (the 12 GB tier holds at 73 to 75 percent) | modelled |
| 11.5 GiB (year 4) | 6 / 5 | 3,000 | 8, 12, 16 GB cache-resident, Apple 16 / 8, 12, Apple 16 (the 16 GB tier holds at 73 to 75 percent) | modelled |
| 16 GiB | 8 / 7 | 4,000 | adds 16 GB, 24 GB cache-resident, Apple 32 | modelled |
| 20 GiB (USD 5,000) | 10 / 9 | 5,000 | everything under 32 GB; Macs under 64 GB | modelled |
The lane's recommended schedule: 5.5 / 8.5 / 11.5 GiB (each just under a tier's room and just over a reticle multiple; 8.5 forces a fifth die where 8.0 does not, at no tier cost with the cache freed; the non-power-of-two floors need the multiply-shift mapping of 1.13.3 option (a)); the Apple rate cost on top, -21 to -25 percent (measured to 8 GiB, section 3.3); the state sizes assumed 92 M / 143 M / 193 M records, the 16 GiB step by state above 268 M.
The capex wall on the spec's emission constant (3,168,808,781 base units per DAA second, 10^8 per IGN, 80 percent to miners: 0.77 / 0.80 / 0.40 / 0.40 / 0.20 / 0.20 B IGN to miners in years 1 to 6; the mission lane's E2 4.18 B is not this constant's figure and is carried beside it); a project starting at launch, shipping at 24 months, mining years 3 to 6, a 10 percent discount; rotation costs it nothing (firmware); the floor's fleet term drops out (under two dies for the whole network at every price):
| Project USD M | Share of the hash | p* at an edge of 3x / 5x / 13x (USD per IGN) | Launch-year miner revenue at p* | Label |
|---|---|---|---|---|
| 100 | 0.30 | 0.59 / 0.49 / 0.43 | USD 330 to 450 M a year (0.9 to 1.2 M a day) | modelled |
| 100 | 1.00 (takes the chain) | 0.18 / 0.15 / 0.13 | 98 to 136 M a year | modelled |
| 250 | 0.30 | 1.47 / 1.23 / 1.06 | 0.8 to 1.1 B a year | modelled |
| 500 | 0.30 | 2.94 / 2.45 / 2.12 | 1.6 to 2.3 B a year | modelled |
Below about USD 100 M a year of miner revenue (IGN 0.13) no rational SRAM project starts at any edge; above about USD 2.3 B a year (IGN 2.9) every one does; at IGN 0.01 / 0.10 / 1.00 (assumed) a USD 150 M project at 5x and 30 percent reads NPV -148 / -130 / +54 M. **The edge moves p* 1.4x across 3x to 13x; the project cost moves it 5x: the hash does not set the wall, the project cost and the detector do.** Killed with the number: per-era re-fill (0.96 J per era), per-block re-fill (0.96 W on the die against 1.3 to 3.2 percent of GPU rate), straddling atoms (+0.02 nJ on the die, +19 percent of items on the verifier at the gate), a second hot table (the die's own kind of memory, k 0.1 to 0.3). Kept: the per-lane scratch pins the lane (a dependency of the width lever); 3D-stacked SRAM would remove the floor's last effect on joules. The lane's line for the served sentence: about 4x at k 0.5 and 2x at k 1 against a 5090 paying its whole latency shadow, 20x to 60x without it, with the price threshold beside it (about USD 0.4 per IGN at a third of the network, 0.13 for a maker who takes the chain). Owed: the W = 8 rows, the Apple curve past 8 GiB.
## 8. Unverified and owed
- Lane D's family harness (the family gate's measured coverage: the 10,000-era random stratum and the three corner strata, the lossy cap, width 4, shape 64) lands in `family-gate.md` by 17:00 UK and is layer 4's coverage table; this document's layer 4 cites it where it is named and does not restate it.