diff --git a/docs/design/class-v6-rotating-family.md b/docs/design/class-v6-rotating-family.md index 277c6df41..7d7e79f4a 100644 --- a/docs/design/class-v6-rotating-family.md +++ b/docs/design/class-v6-rotating-family.md @@ -8,7 +8,7 @@ The honest frame first, from last night's measured close (`docs/analysis/counter | Layer | A fixed-function chip on its release day | A GPU-like chip on its release day | RTX 5090 (measured where stated) | RTX 5070 Ti (MEASURED stock, a rented Vast pod, 10:08 to 10:11 UTC: class v3 78.69 MH/s at 140.8 W, class v4 78.78 at 224.0 W, the premium 83.2 W = 10.3 pJ per counted op, fingerprints equal to the Mac's; no core-lock grid, the host refused -lgc) | Apple M5 Max (measured where stated) | |---|---|---|---|---|---| -| 1. Per-era draws of the parameters now fixed by release (mixer round count, op-mix weights, read width, program length, shadow block shape), from chain state | dead at the first era whose draw leaves its wired value: a wired x8 mixer at an x16 era recomputes nothing right; a 4-byte-granularity controller at a 16-byte era moves 4x the bytes; a core sized for 100,000 ops at a 200,000 era runs at half rate; the tape-out lives one era (180 days, layer 3) | firmware; the core sized for the band's top (200,000 ops: about 60 mm^2 of N5 instead of 30, USD 25 to 40 more per chip, modelled); `k` unchanged; capex per MH/s +10 to 20 percent (modelled) | rate: 0 across the band while latency-bound (the ladder's rungs 0 to 2: 0 and -2.7 percent at the 431 W cap, measured); watts: the premium follows N and the mix (6.2 to 11.3 pJ per counted op measured; a shuffle-heavy draw up to 5x per op, so the band excludes it); the verifier +0.2 to +1.5 ms per era draw (measured ladder, estimated mixer) | rate 0 at class v4 (78.69 to 78.78 MH/s, memory-bound, measured); the class v4 premium 83.2 W at stock (10.3 pJ per counted op, the 5090's 10.8), 224 W under its 300 W limit with no throttle, so the N band's top (200,000 ops) costs about 170 W of premium at stock by the per-op figure and the card's limit binds first (approximate); the knee rows need a host that allows -lgc | rate -3.3 to -10 points across the N band (measured ladder rungs 0 to 2), 0 for the mixer and the width; the Apple tier sets the band's top (130,000 ops at the 5 percent rule) | +| 1. Per-era draws of the parameters now fixed by release (mixer round count, op-mix weights, read width, program length, shadow block shape), from chain state | dead at the first era whose draw leaves its wired value: a wired x8 mixer at an x4 era (x16 is out of the band: the verifier measured 10.11 ms loaded) recomputes nothing right; a 4-byte-granularity controller at a 16-byte era moves 4x the bytes; a core sized for 100,000 ops at a 200,000 era runs at half rate; the tape-out lives one era (180 days, layer 3) | firmware; the core sized for the band's top (200,000 ops: about 60 mm^2 of N5 instead of 30, USD 25 to 40 more per chip, modelled); `k` unchanged; capex per MH/s +10 to 20 percent (modelled) | rate: 0 across the band while latency-bound (the ladder's rungs 0 to 2: 0 and -2.7 percent at the 431 W cap, measured); watts: the premium follows N and the mix (6.2 to 11.3 pJ per counted op measured; a shuffle-heavy draw up to 5x per op, so the band excludes it); the verifier +0.2 to +1.5 ms per era draw (measured ladder, estimated mixer) | rate 0 at class v4 (78.69 to 78.78 MH/s, memory-bound, measured); the class v4 premium 83.2 W at stock (10.3 pJ per counted op, the 5090's 10.8), 224 W under its 300 W limit with no throttle, so the N band's top (200,000 ops) costs about 170 W of premium at stock by the per-op figure and the card's limit binds first (approximate); the knee rows need a host that allows -lgc | rate -3.3 to -10 points across the N band (measured ladder rungs 0 to 2), 0 for the mixer and the width; the Apple tier sets the band's top (130,000 ops at the 5 percent rule) | | 2. The state-derived dataset's size tracks chain-state growth with a floor (class v5's leaves scaled by the state, never below the 1.13.3 schedule) | a chip with fixed memory ages out when the dataset passes it: one HBM3 stack 24 GB, the 5090's board 32 GB; the time-memory curve (chip-model 5.4) says the excess must be recomputed at 6.3 nJ per item against 2.0 per read, so its energy per hash rises with the overflow | the same memory limit; a chip buys DRAM a card cannot (24 GB stacks at USD 200, modelled), so it ages out LAST: the 8, 12 and 16 GB card tiers go first | fine to 32 GB: the 2 GiB genesis dataset plus the state's leaves; at a used chain's 5 GB state (class-v5 section 3) the dataset is capped at the sample size, 2 GiB | 16 GB: fine to the 8 GiB step (year 12 on the schedule); the state floor does not move it | 36 to 128 GB unified: fine to the 16 GiB step; the daily build grows with the size (13 to 30 ms measured at 1 GiB) | | 3. Scheduled family epochs by height, every 180 days by default, no release (the reserve R0 to R8 of spec 1.13.2 unlocking by height, then rotating) | a chip without the family's datapath loses its weight of the mix at the unlock (4 points of 79) or emulates it at the vendor penalty (1.5x to 2.4x per op, measured on the cards); a chip taped out against one family set is a GPU-like chip or dead | pre-wires every family for about USD 4 of N5 silicon (algorithm.md 5.2, modelled); moves the per-joule edge under 10 percent per family | measured family step costs: shfla 1.53x, perm 1.30, mm8 2.43 the add-xor-rotate step; under 1 percent of rate at 4 points | pending; the same families on the same silicon generation | measured: shfla 1.91x, perm 1.13 emulated, mm8 emulated at 1.6x per dot4; under 1 percent of rate at 4 points | | 4. The acceptance floor (c''') and the F8 uniformity test generalised to every era's draw, with a redraw on failure (plus the per-site largest-bucket bound and the value-level bias test as the next class's two tests) | nothing on the chip; it is what makes layers 1 and 3 safe without per-era cryptanalysis | nothing | the generator's attempts per seed (today about 30 at 2.4 percent rejection under (c'''); a redraw costs nothing on a card) | the same | the same | @@ -35,7 +35,7 @@ PENDING the hash lane's rows (16:00 UK): the x16 mixer verifier and build; the t | Parameter | Band (genesis) | Why that band (the measured rows that set it) | Chip rows: fixed-function / GPU-like | Per tier: 5090 / 5070 Ti / M5 Max | Verifier | Open number | |---|---|---|---|---|---|---| -| Mixer applications per round `m` | {4, 8, 16} | x4 and x8 measured (mixer-x4.md 6.4: 1.92 and 2.79 ms per unit on a loaded M5 Max core; the build unmoved on every discrete card, latency-bound); x16 about 3.7 ms per unit on the M5 Max core, 9 ms on a 2019-class core (chip-model-v3 section 3 item 1, estimate) | the `f = 0` recompute chip's rate halves per doubling (0.31x bare at x8, 0.16x at x16, modelled); the `f = 1` chip unmoved (it recomputes nothing) | rate 0 / 0 / 0 (latency-bound, measured at x4 and x8); the daily build 42 / pending / 29 ms at x8 (measured) | +0.6 to +1.5 ms per doubling (measured x4, x8; x16 the row) | the x16 row; adv-mixer-3's margin (every statistic clean from k = 1, the SAT ladder k = 1 solved, k = 2..4 timeout): x4 keeps the margin by that report's reading, x16 widens it | +| Mixer applications per round `m` | **{4, 8}; 16 inadmissible on the loaded reference core** | x4 and x8 measured (mixer-x4.md 6.4: 1.92 and 2.79 ms per unit on a loaded M5 Max core; the build unmoved on every discrete card, latency-bound); **x16 MEASURED 11:13 UK on build-1 (core 4, the frozen crate, today's class v3 stream, the same id f5e904bc5d148926 for x8 and x16): x8 4.672 ms per warp cold alone and 5.413 with the sibling loaded; x16 8.688 and 10.112 ms; the multiplier is 1.86x on the verifier (the mixer IS the verifier's cost: 4.0 of the 8.7 ms), the sibling 16 percent on both; chip-model-v3's estimate (9 ms on a 2019-class core) stands within 4 percent on a current server core** | the `f = 0` recompute chip's rate halves per doubling (0.31x bare at x8, 0.16x at x16, modelled); the `f = 1` chip unmoved (it recomputes nothing) | rate 0 / 0 / 0 (latency-bound, measured at x4 and x8); the daily build 42 / pending / 29 ms at x8 (measured); the x16 build on the 5090 rides the 12:45 UK job | +1.9x per doubling measured (x8 to x16); x16 over the 10 ms loaded gate by 0.11 ms on the reference core, the ladder's rung-3 class of miss, so the band's top is x8 and x16 is `admissible: false` in the genesis list unless a quiet-core re-run admits it | adv-mixer-3's margin (every statistic clean from k = 1, the SAT ladder k = 1 solved, k = 2..4 timeout): x4 keeps the margin by that report's reading; the band {4, 8} is the measured one | | Op-mix weights (the ten non-load families) | each weight within B points of table 1.4.2, B = 2 today; the proposal: B = 4 with the shuffle and mulhi weights capped at their class v4 values | the microbench (15.1a): shfl 55.8 pJ per op, mulhi 39.6, prmt 22.3, lop3 24.1, mul 13.9, arx 11.3 at the stock clock; a draw that raises shfl from 8 to 14 of 75 raises the premium per instruction by about 35 percent (modelled on those rows) | a chip that specialised its lane ratio loses the ratio; a general core nothing | the premium per instruction moves with the mix: pending the two packs | +0 ms | the two re-weighted packs' rows | | Read width `W` | {1, 4} words | w4 and w16 measured 5 October: the 5090 139.8 against 136.1 MH/s, the 9070 XT 17.90 against 18.15; w64 bandwidth-bound (71.9 MH/s on the 5090) | a 4-byte-granularity controller moves 4x the bytes at a w16 era (the Ren-Devadas lever reversed); the `f = 1` chip's energy per read rises toward the card's | 0 / pending / 0 (measured w16 on the M5 Max: within 1 percent) | 0 | none | | Shadow block shape | 64 to 256 instructions per block, the pass count the ladder's | 64-instruction blocks ran 2.5 to 3.5 percent FASTER than 256 on the 5090 and the M5 Max (measured 6 October); 1,024 cost the M5 Max 17 percent | nothing for any chip (the work is the same) | +2.5 to 0 percent / pending / +2.5 to 0 | 0 | none |