From 65ef3ebcf39587964ed91f7baa768fc463cec771 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Thu, 8 Oct 2026 12:08:18 +0000 Subject: [PATCH] Class v6: the k convention rule for the floor and the served-line review (absolute rows the headline, the record's beside them marked; the k lane's rows re-fold lane 3's shadowed rows before the close) Co-Authored-By: Claude Fable 5.1 Documents-only replay of ac0fe332f (counter-asic-4) for the box mirror master --- docs/design/class-v6-rotating-family.md | 7 +++++-- 1 file changed, 5 insertions(+), 2 deletions(-) diff --git a/docs/design/class-v6-rotating-family.md b/docs/design/class-v6-rotating-family.md index 40a11a332..0d0ba6ad4 100644 --- a/docs/design/class-v6-rotating-family.md +++ b/docs/design/class-v6-rotating-family.md @@ -243,6 +243,7 @@ The served sentence (`docs/plans/counter-asic-3-public-text-2026-10-07.md`, the | "Class v5 then makes the dataset the chain's own state, so a chip that stores it or recomputes it is wrong on every item" | class-v5-stored-state.md: a stateless or stale chip is wrong on every item; a chip that holds the state (one node per farm, the leaves at 16.5 KB/s) is not | designed, measured on the harness | the clause overstates: "a chip that does not follow the chain is wrong on every item" is the true form; a chip that stores the dataset and follows the chain is unmoved (chip-model 5.10) | | "Without class v4 the same chip would reach 5x to 9x" | chip-model 5.4: 5.1x GDDR7 to 9.2x eight HBM3 stacks at zero shadow | modelled | stands for the DRAM chips; the SRAM store reads 17x | | The served "random reads" sentences (uniform random reads over the dataset) | lane D and adv-cache-2: the reads spread over the whole dataset; a bit-level bias at one site in about half of epochs; 1.6 percent of a hash's reads to a half store, nothing to a full store; the next class folds it out | measured | CORRECTED today by the audit lane on those figures (main's word, 11:4x UK); not pending | +| The convention behind every "x per joule with the shadow" figure | the record's convention (chip-model-v3 and the served text): `k` is the chip core's op cost as a fraction of the GPU's at the same operating point, so when the 5090 locks to 6.2 pJ the chip's core is priced at 6.2 k pJ too; the absolute convention: the chip's core costs its own pJ per op (the k lane's RTL figure), the GPU's its measured pJ at its point | measured GPU side; the chip side claimed until the k lane's RTL rows | the record's convention flatters every chip row at the lock by up to 1.6x (floor lane 3, 13:06 UK: the SRAM die 6.1x against 3.8x at k 0.5, 3.3x against 2.0x at k 1); the served line's figures are at the knee, so they are in the record's convention and need re-reading in the absolute one before main's word; the close carries both, absolute first | | The schedule's cost beside the SRAM store's edge (main's order) | section 3: the floor and ceiling per tier; main's candidate schedule (6, 10, 14 GiB) priced by lane A by 17:00; the Apple rate cost measured today (-12 / -20 / -22 percent at 2 / 4 / 8 GiB) | measured and modelled | lands at 17:00 | The wording this lane proposes for main's word, if the SRAM-store reading stands at 20:00: "At launch the strongest chip we can price today, a memory-controller chip that stores the dataset, reaches 2.1x per joule against an RTX 5090 at its knee with a core as good as a GPU lane and 3.4x with one three times better, under class v4 from the first block; a 5090 locked at its knee pays 82 W for that shadow work. The strongest chip we can model for 2027 to 2028, a 2 GiB SRAM store on one 2 nm die, would reach 3x to 5x with that same shadow and 17x without it, at an N2 project of USD 100 M to 500 M and 18 months or more; the dataset's size is the one lever on it, and the schedule says how it grows. Class v5 makes the dataset the chain's own state, so a chip that does not follow the chain is wrong on every item. The model and every measurement are public." Every number in it is labelled in this document and the chip model; nothing served moves on this lane's word. @@ -276,7 +277,9 @@ The close (20:00): the honest floor per tier against each chip row, the recommen **The read width is the only wire lever on the SRAM die, and the floor is a ticket, not a joule.** Lane B's 1.0 nJ was the 64-byte row; the die reads what the hash asks for, and the wire scales with the bits moved while the macro access does not: -| W words | Bits per read | Chip E_read | GH/s per die (300 W) | Zero shadow against the 5090 | With the class v4 shadow, k 0.5 / 1 | At the card's whole shadow (F 2.0), k 0.5 / 1 | The honest card's cost | Label | +**The convention rule for every shadowed row (the coordinator's order, 13:1x UK): the headline is the ABSOLUTE convention, pJ per op on each side at the operating point (the die's core costs what it costs; the GPU's op at its own clock, 11.3 pJ stock and 6.2 at the 1,300 lock on the 5090); the record's convention (k against the GPU's op cost at the same operating point, so the die's core gets cheaper when the card locks) is carried beside it, marked, and it flatters the die at the lock by 1.6x (6.1x against 3.8x at k 0.5; 3.3x against 2.0x at k 1).** The shadowed columns in the table below are the lane's first reading at the card's stock point, where the two conventions coincide; the k lane's synthesised rows (floor lane 2, coming in absolute) re-fold them before the close, and the lock rows then appear in both conventions with the absolute one first. + +| W words | Bits per read | Chip E_read | GH/s per die (300 W) | Zero shadow against the 5090 | With the class v4 shadow, k 0.5 / 1 (stock point; the record's convention, to be re-folded by the k lane in absolute) | At the card's whole shadow (F 2.0), k 0.5 / 1 (same convention note) | The honest card's cost | Label | |---|---|---|---|---|---|---|---|---| | 1 (4 bytes, today) | 80 | 0.25 nJ | 8.3 | **66x** | 5.7x / 3.0x | 4.1x / 2.1x | 0 | modelled chip; measured card | | 4 (16 bytes) | 176 | 0.38 | 5.6 | 44x | 5.6x / 2.9x | 4.0x / 2.1x | the 5090 +2.7 percent, the 9070 XT -1.4, the M5 Max within 1 (measured 5 October) | measured card | @@ -285,7 +288,7 @@ The close (20:00): the honest floor per tier against each chip row, the recommen Meaning: at the hash's own width the die is 3.5x stronger than the record said at zero shadow; with the shadow on, every row sits at 5.0x to 5.7x (k 0.5) and 2.7x to 3.0x (k 1): **the shadow is the whole hold, the memory moves it 0.2x to 0.7x.** The fold is dst-keyed (`verify::fold_words`), so a wide read cannot be pre-folded; dependent chains per step leave the ratio unchanged (per-read on both sides); banking plus a sequencer cannot localise (uniform on a window of at least 256 MiB, the 16 sites alternating; moving the lane costs about 290 bits, which caps the wire lever at about W = 8); the per-site window shrink buys the die and the DRAM chip nothing; a hop at the honest widths is 0.04 to 0.15 nJ, so a ten-die store still reads 25x to 58x at zero shadow. This moves layer 1's width row: **pin W = 4 (16 bytes) at genesis (measured free on all three vendors; the die from 66x to 44x), W = 8 (31x; 5.4x at k 0.5) once the generator and emitter carry it and one PC 1 row confirms, never 16.** -The floor as a ticket: SRAM is flat at about USD 250 per GiB to 2031 (density +6 to 11 percent per node against dearer wafers; claimed, approximate), so USD 5,000 of silicon per store is 20 GiB at N2 and about 21 GiB in 2031, which retires every card under 32 GB and every Mac under 64 GB; the constraint and "fewest cards" cannot both hold. The per-MH/s does not rise with the floor (every die powered: USD 0.25 to 0.5 per MH/s at any size); the floor raises the minimum ticket only. The honest card's side of the floor is measured (the hash lane's kit b, section 3.3): the 5090 at its knee pays 4 / 8 / 10 percent more energy per hash at 2 / 4 / 8 GiB, so every chip row's edge against a card at its knee rises by 4 to 11 percent across the schedule while the die's joules do not move; the schedule's sentence is USD 1,000 of chip ticket per step for 1 to 5 percent of the tuned 5090's energy and about a quarter of today's cards by count. The k convention matters at the lock: the record's (k against the GPU's op cost at the same operating point) reads the SRAM die at 6.1x (k 0.5) and 3.3x (k 1); the absolute convention (the die's core costs what it costs, k against the stock 11.3 pJ per op) reads 3.8x and 2.0x; the record's flatters the die by 1.6x at the lock, so the served line in section 9 uses the absolute one. +The floor as a ticket: SRAM is flat at about USD 250 per GiB to 2031 (density +6 to 11 percent per node against dearer wafers; claimed, approximate), so USD 5,000 of silicon per store is 20 GiB at N2 and about 21 GiB in 2031, which retires every card under 32 GB and every Mac under 64 GB; the constraint and "fewest cards" cannot both hold. The per-MH/s does not rise with the floor (every die powered: USD 0.25 to 0.5 per MH/s at any size); the floor raises the minimum ticket only. The honest card's side of the floor is measured (the hash lane's kit b, section 3.3): the 5090 at its knee pays 4 / 8 / 10 percent more energy per hash at 2 / 4 / 8 GiB, so every chip row's edge against a card at its knee rises by 4 to 11 percent across the schedule while the die's joules do not move; the schedule's sentence is USD 1,000 of chip ticket per step for 1 to 5 percent of the tuned 5090's energy and about a quarter of today's cards by count. The k convention matters at the lock: the record's (k against the GPU's op cost at the same operating point) reads the SRAM die at 6.1x (k 0.5) and 3.3x (k 1); the absolute convention (the die's core costs what it costs, k against the stock 11.3 pJ per op) reads 3.8x and 2.0x; the record's flatters the die by 1.6x at the lock, so the served line in section 9 and the close carry the absolute rows as the headline and the record's beside them marked. | Floor | Reticles 2026 / 2031 | USD of silicon | Tiers out (worst / best working set) | Label | |---|---|---|---|---|