diff --git a/docs/design/class-v6-rotating-family.md b/docs/design/class-v6-rotating-family.md index a45e63b66..f1c4ce477 100644 --- a/docs/design/class-v6-rotating-family.md +++ b/docs/design/class-v6-rotating-family.md @@ -307,7 +307,7 @@ The capex wall's two thresholds (lane 3, 10.3, re-folded 13:18 UK on lane 5's co The served line as the external review words it is in 10.0f and overrides this paragraph's wording. This lane's earlier sentence, kept for the record: a chip that stores the dataset reaches about 2.2x to 2.4x per joule against the honest NVIDIA tiers at their knee (the 5080 2.2x, the 5090 2.4x with a chip a node ahead; 2.0x against a chip on the GPU's own node, 2.8x two nodes ahead; 2.8x on the 5090 without the window) and under 2x only against the Apple tier (the M5 Max 1.5x), about 4x and 2.6x if it is built on an N2 SRAM store for USD 100 M or more, and under 1x per hash over its 180-day class life only above about USD 300 M a year of miner revenue (10.0a), with Monero's measured 1.0x to 1.5x beside it; no rotating layer moves the per-joule figures, and layer 3's class life is what sets the per-hash one; the layers decide which chip can be built and how long its tape-out lives. The worst case beside it, if a chip's core costs no more than its bare units (the floor k of 10.2): 4x and 2.4x against the DRAM board, 14x and 9x against the SRAM die (lane 3's line: 6x to 10x at the full shadow, never under 2x). In the other convention, the card's whole 330,000-op latency shadow priced on the chip at the synthesised core (lane 3's 10.3 table): the SRAM die 3.5x against a 5090 at stock and 2.2x at its knee (4.8x and 3.0x with a 2 nm core), the DRAM board 2.6x at stock. The record's claimed band (k 0.3 to 0.8) is confirmed from the chip side by the core row (0.56 at the lock), so the served 2.1x at k = 1 is the conservative end and 2.8x the measured-core reading. -The three class v6 changes it implies: (1) the op mix moves to the band's best mix as the genesis table (add 16, xor 14, mad 12, rotl 11, sub 10, rotr 10, shfl 4, mul 4, mulhi 2, or 0; the k lane's optimiser, 15:2x UK: +42 percent of k_eff on the unit floors, +15 percent on the core; the shuffle at its floor because the chip's butterfly is k 0.011 to 0.021, the lowest of every family; lane D: 0 of 3,000 eras exhausted under the band), subject to one acceptance pass of that table through lane D's harness, the census lane's passed table the default; (2) no SM-sparse default (measured dead on the 5090, 4090 and H100 at 1.3 to 3.9 percent at best; `--sm-sparse auto` ships off by default, on in the Efficiency and Balanced tiers at its measured 1 to 2.5 percent, 10.1 item 6; the per-watt gift of the hot table is the chip's, 20.3c); (3) the dataset schedule 5.5 / 8.5 / 11.5 GiB with the read width pinned at 4 words (8 measured not free, 16 never); and a fourth the sweep adds, (4) the 64-register window per lane as the core shape (k 0.78 at the lock against 0.56; 10.0c, the edge table 10.0e), its GPU side modelled until the generator carries a 64-entry init and fold rule (a half-day line) and a pack is measured (the chip's ticket USD 1,500 / 2,500 / 3,000; a tuned 5090 pays 1 to 5 percent per step; about a quarter of today's cards by count per step). +The three class v6 changes it implies: (1) the op mix moves to the band's best mix as the genesis table (add 16, xor 14, mad 12, rotl 11, sub 10, rotr 10, shfl 4, mul 4, mulhi 2, or 0; the k lane's optimiser, 15:2x UK: +42 percent of k_eff on the unit floors, +15 percent on the core; the shuffle at its floor because the chip's butterfly is k 0.011 to 0.021, the lowest of every family; lane D: 0 of 3,000 eras exhausted under the band), subject to one acceptance pass of that table through lane D's harness, the census lane's passed table the default; (2) no SM-sparse default (measured dead on the 5090, 4090 and H100 at 1.3 to 3.9 percent at best; `--sm-sparse auto` ships off by default, on in the Efficiency and Balanced tiers at its measured 1 to 2.5 percent, 10.1 item 6; the per-watt gift of the hot table is the chip's, 20.3c); (3) the dataset schedule 5.5 / 8.5 / 11.5 GiB with the read width pinned at 4 words (8 measured not free, 16 never); and a fourth the sweep adds, (4) the 64-register window per lane as the core shape (k 0.78 at the lock against 0.56; 10.0c, the edge table 10.0e), its GPU side modelled until the generator carries a 64-entry init and fold rule (a half-day line) and a pack is measured, with the LIVENESS RULE main confirms (15:56 UK): the fold forming each load address consumes all 64 registers, so the result depends on the whole window and a chip cannot shrink it; a liveness tool that checks every drawn program for that dependency is an acceptance test beside the census (the chip's ticket USD 1,500 / 2,500 / 3,000; a tuned 5090 pays 1 to 5 percent per step; about a quarter of today's cards by count per step). #### 10.0a The second unit: cost per hash over the chip's life (the founder's question, 14:1x UK; modelled on measured card rows; SUPERSEDED as a headline by 10.0f item 1: the 180-day life below is the fixed-lane chip's row, and the programmable chip's 3-year life is the default) @@ -435,7 +435,7 @@ The conditions, read off the surface, which are the economic-resistance statemen **(6) RandomX** generates several programs per hash; "Monero's hash never changes" is wrong as a statement of dynamic work and reads instead "RandomX's rules have been stable since 2019; its programs vary per hash". The X5 figure (6.37 J per kH at the wall, Bitmain's specification) is an observed comparison against a stated CPU measurement, not a ceiling. 10.0b's row is read with this. -**(7) The next research programme**, named in the close: the connected-state experiment, the mixed integer and FP32 candidate, and the multi-family programmable adversary, three lanes launched at 15:5x UK with their clocks (their rows land as amendments under this section). Rotation is not expected to deliver the missing joule; it is the response capability of item (1). +**(7) The next research programme**, named in the close: the connected-state experiment, the mixed integer and FP32 candidate, and the multi-family programmable adversary, three lanes launched at 15:5x UK (connected state a4d3518190ad011fc, mixed FP32 aa943688eaa06c538, the multi-family adversary a1a9876a88f5a72fc) with rows from 18:30 UK for class v7, not tonight's cut; their rows land as amendments under this section, labelled. Rotation is not expected to deliver the missing joule; it is the response capability of item (1). #### 10.0h The served text (for the site audit lane, served as written; each number's label; what is withdrawn)