diff --git a/docs/design/class-v6-rotating-family.md b/docs/design/class-v6-rotating-family.md index 318e18fcf..7dd27f271 100644 --- a/docs/design/class-v6-rotating-family.md +++ b/docs/design/class-v6-rotating-family.md @@ -55,7 +55,7 @@ The band a card sees across the draw's range, from the hash lane's cost rows (12 | Parameter | Band (genesis) | Why that band (the measured rows that set it) | Chip rows: fixed-function / GPU-like | Per tier: 5090 / 5070 Ti / M5 Max | Verifier | Open number | |---|---|---|---|---|---|---| | Mixer applications per round `m` | **{4, 8} at genesis; 16 in the list as `admissible: false` (the ladder's rule: a flag in the genesis list, flipped only by the 90 percent upgrade path once the 2019-class core measures it)** | **The card pays nothing for the draw, MEASURED (the hash lane's kit b on PC 1, 12:49 to 12:58 UK, generator 2 with the era, the multiplier the only difference, 250 batches per row, fingerprints PASS): x8 137.73 MH/s at 320.2 W unlocked and 127.44 at 212.0 W at the 1,300 lock; x16 137.72 at 320.1 and 127.46 at 211.7; within 0.1 MH/s and 0.5 W at either state, item derivation hidden under the read chain, so the verifier's 1.86x per doubling is the draw's whole cost.** x4 and x8 measured (mixer-x4.md 6.4: 1.92 and 2.79 ms per unit on a loaded M5 Max core; the build unmoved on every discrete card, latency-bound); **x16 MEASURED (today's class v3 stream, the same id f5e904bc5d148926 for x8 and x16). Cold alone: build-1 core 4 x8 4.67 ms per warp, x16 8.69 (the hash lane, 10:13 UTC); build-3 core 4 (AX102, a faster core class) x8 3.48, x16 6.53 (the build-server lane, 11:01 UTC). With the SMT sibling loaded for the row's whole length (the method that counts: a sibling bench looped for the row, not run once, which ends in about a second and leaves the verify phase sibling-idle; the hash lane's earlier 5.41 / 10.11 / 9.04 to 9.10 ms rows were of that kind and are withdrawn): build-3 x8 6.24, x16 11.44 ms, about 1.8x on both classes. The multiplier is 1.86x on the verifier (the mixer IS the verifier's cost); chip-model-v3's estimate (9 ms on a 2019-class core) stands within the cold rows** | the `f = 0` recompute chip's rate halves per doubling (0.31x bare at x8, 0.16x at x16, modelled); the `f = 1` chip unmoved (it recomputes nothing) | rate 0 / 0 / 0 (latency-bound, measured at x4 and x8); the daily build 42 / pending / 29 ms at x8 (measured); the x16 build on the 5090 rides the 12:45 UK job | +1.9x per doubling measured (x8 to x16); **x16 at 11.4 ms loaded is over the 10 ms gate on the measured row, so `admissible: false` stands on the measurement, not only on the chip model's margin; the O-1.14 laptop run can only tighten it** | adv-mixer-3's margin (every statistic clean from k = 1, the SAT ladder k = 1 solved, k = 2..4 timeout): x4 keeps the margin by that report's reading; the band {4, 8} is the measured one | -| Op-mix weights (the ten non-load families) | **B = 4 points on the injecting families only (add, sub, xor, mad, shfl, rotl, rotr); the lossy families (or, mul, mulhi) capped at their base, the ring-A rule `or + mul + mulhi` at most the table's 18 plus B, which keeps the per-candidate rejection r under 0.85 (r^256 under 1e-18)**; the shuffle weight capped at its class v4 value on the energy side | lane D's coverage run (4,900 drawn eras on three boxes, 11:3x UK, measured): at the lossy corner of B = 4 (or, mul and mulhi all at +4, 30 of 75 lossy against 18) r is 0.956 (the shipped 0.681) and 8 of 663 eras exhaust the 256-attempt cap (mean attempt 18.6, max 252), so 1.2 percent of that corner's epochs would take the last-resort program, which fails rule (a) in 9 percent of seeds (adv-accept-3); the random stratum with every weight drawn reads r 0.718, max attempt 107, 0 exhaustions in 1,444. The energy side (15.1a): shfl 55.8 pJ per op, mulhi 39.6, prmt 22.3, lop3 24.1, mul 13.9, arx 11.3; a shuffle-heavy table raises the premium per instruction up to 2x (modelled), a multiply-heavy one 1.1x | a chip that specialised its lane ratio loses the ratio; a general core nothing | the premium per instruction moves with the mix within the capped band (the cost rows: at N = 100,000 the 5090's +146 W stock becomes up to +160 W multiply-heavy; shuffle-heavy excluded by the cap); the 5070 Ti scaled by 83/146, the M5 Max by 16/146; the two packs' measured rows replace this when they land. **First row MEASURED (kit b, PC 1, 12:5x UK): the shuffle-heavy table (shfl 13 of 48 against the stock 5) on class mx8 without a shadow block reads 137.65 MH/s at 312.9 W unlocked (7 W under the stock table's 320.2) and 127.33 at 212.4 at the lock (level): the multiply-heavy table (mulhi 11, mad 8, mul 5 of 48 against the stock 9, 4, 1; shfl 1 against 5; kit c, 13:03 to 13:06 UK) reads 137.77 at 313.5 W unlocked and 128.42 at 211.8 W at the lock, so the three tables sit within 7 W unlocked and 0.6 W at the lock: on the 48-op base program the weight move is a 2 percent term whichever way it goes; the row layer 1 needs is the same two tables inside the shadow block (mx8+sh256x27, kit d, about 13:45 UK), and the microbench arithmetic stays the default until it lands. The acceptance side of a re-weight (the census lane, d20eb04bd, 14:1x UK, measured): W = 4 at weights 13, 11, 6, 10, 8, 8, 7, 2, 6, 4 (mulhi and mul down, mad and four ARX families up) 256 of 256 both ways at r 0.407, (c''') 1.2 percent, F8-form 0.999 to 1.008, the product law at bit 2 in 75 of 256 programs; both W = 16 and the re-weight together 256 of 256 at r 0.426, (c''') 3.4 percent. So a re-weight is clean on acceptance; the hold on it is the energy side (15.1b) and now the chip side too (10.2: mulhi is the card's worst lever by 6x, and a mix that lowers mulhi helps the card against the chip only if the k lane's sweep says so)** | +0 ms | the known-failed test: the 8 of 663 exhaustions at the uncapped lossy corner (61 of 5,000 on the 17:00 cut), which the capped band must read as 0: **MEASURED 0 of 3,000 under the band (lane D's 17:00 cut, 5.1b; r 0.595, mean attempt 1.47, max 59)** | +| Op-mix weights (the ten non-load families) | **B = 4 points on the injecting families only (add, sub, xor, mad, shfl, rotl, rotr); the lossy families (or, mul, mulhi) capped at their base, the ring-A rule `or + mul + mulhi` at most the table's 18 plus B, which keeps the per-candidate rejection r under 0.85 (r^256 under 1e-18)**; the shuffle weight capped at its class v4 value on the energy side | lane D's coverage run (4,900 drawn eras on three boxes, 11:3x UK, measured): at the lossy corner of B = 4 (or, mul and mulhi all at +4, 30 of 75 lossy against 18) r is 0.956 (the shipped 0.681) and 8 of 663 eras exhaust the 256-attempt cap (mean attempt 18.6, max 252), so 1.2 percent of that corner's epochs would take the last-resort program, which fails rule (a) in 9 percent of seeds (adv-accept-3); the random stratum with every weight drawn reads r 0.718, max attempt 107, 0 exhaustions in 1,444. The energy side (15.1a): shfl 55.8 pJ per op, mulhi 39.6, prmt 22.3, lop3 24.1, mul 13.9, arx 11.3; a shuffle-heavy table raises the premium per instruction up to 2x (modelled), a multiply-heavy one 1.1x | a chip that specialised its lane ratio loses the ratio; a general core nothing | the premium per instruction moves with the mix within the capped band (the cost rows: at N = 100,000 the 5090's +146 W stock becomes up to +160 W multiply-heavy; shuffle-heavy excluded by the cap); the 5070 Ti scaled by 83/146, the M5 Max by 16/146; the two packs' measured rows replace this when they land. **First row MEASURED (kit b, PC 1, 12:5x UK): the shuffle-heavy table (shfl 13 of 48 against the stock 5) on class mx8 without a shadow block reads 137.65 MH/s at 312.9 W unlocked (7 W under the stock table's 320.2) and 127.33 at 212.4 at the lock (level): the multiply-heavy table (mulhi 11, mad 8, mul 5 of 48 against the stock 9, 4, 1; shfl 1 against 5; kit c, 13:03 to 13:06 UK) reads 137.77 at 313.5 W unlocked and 128.42 at 211.8 W at the lock, so the three tables sit within 7 W unlocked and 0.6 W at the lock: on the 48-op base program the weight move is a 2 percent term whichever way it goes; the row layer 1 needs is the same two tables inside the shadow block, MEASURED (kit d, PC 1's 5090, 14:51 to 15:0x UK, the v5 kit's CUDA worker, 250 x 2^24 per row, all PASS; the caveat: this job's rows run on the kit worker, whose base reads 120.0 MH/s unlocked against the installed worker's 137.6 this morning, so the like-for-like comparison is within the job and the comparison to the morning's stock table is in watts only). Unlocked: the w4 base 119.95 MH/s at 308.2 W; the multiply-heavy table in the block (mulhi 11, mad 8, mul 5, shfl 1 of 48) 118.28 at 412.8 W, a block premium of 104.6 W (0.88 microjoules per hash); the shuffle-heavy table (shfl 13 of 48) 118.38 at 463.7 W, 155.5 W (1.31); at the 1,300 lock: the base 100.55 at 190.5 W, multiply-heavy 98.66 at 252.5 W (62.0 W, 0.63), shuffle-heavy 98.27 at 271.1 W (80.6 W, 0.82); the morning's stock-table premiums on the installed worker 152.4 W unlocked and 84.2 W locked. **Inside the shadow block the weight table is a 30 percent lever on the block's energy: the multiply-heavy table costs the card 33 percent less than the shuffle-heavy one per hash unlocked and 23 percent less at the knee at a flat rate (within 0.4 percent), and the shuffle-heavy table costs the same as the stock table in watts (within 3 W).** The microbench's arithmetic had the sign right at the multiply-heavy end (1.1x modelled against about 0.7x measured: the mulhi and mad units are cheaper per op than the ARX mix inside a dependent chain on Blackwell) and the shuffle-heavy end reads 1.0x against the 2x modelled (the shuffles are already most of the stock table's cost). What it does to the band: a multiply-heavy table lowers the card's premium by a quarter, but the k lane's rows (10.2) price mul and mulhi on the chip at k 0.03 to 0.08 against the ARX families' 0.18, so the chip's premium falls further than the card's and the edge rises; the cap on the lossy families at their base stands on both the acceptance side (lane D) and the energy side (this row). The shuffle cap on the energy side is now measured as costless to the card (shuffle-heavy = stock in watts), so if the k lane's butterfly row reads above the ARX k the shuffle weight can rise inside B = 4 at no card cost; that row is owed. The acceptance side of a re-weight (the census lane, d20eb04bd, 14:1x UK, measured): W = 4 at weights 13, 11, 6, 10, 8, 8, 7, 2, 6, 4 (mulhi and mul down, mad and four ARX families up) 256 of 256 both ways at r 0.407, (c''') 1.2 percent, F8-form 0.999 to 1.008, the product law at bit 2 in 75 of 256 programs; both W = 16 and the re-weight together 256 of 256 at r 0.426, (c''') 3.4 percent. So a re-weight is clean on acceptance; the hold on it is the energy side (15.1b) and now the chip side too (10.2: mulhi is the card's worst lever by 6x, and a mix that lowers mulhi helps the card against the chip only if the k lane's sweep says so)** | +0 ms | the known-failed test: the 8 of 663 exhaustions at the uncapped lossy corner (61 of 5,000 on the 17:00 cut), which the capped band must read as 0: **MEASURED 0 of 3,000 under the band (lane D's 17:00 cut, 5.1b; r 0.595, mean attempt 1.47, max 59)** | | Read width `W` | **pinned at 4 words (16 bytes) at genesis, not drawn** (floor lane 3, 10.3: the width is the only wire lever on the SRAM die, 66x at 4 bytes against 44x at 16 at zero shadow; 8 words after the owed PC 1 and Mac rows; 16 never) | w4 and w16 measured 5 October: the 5090 139.8 against 136.1 MH/s, the 9070 XT 17.90 against 18.15, the M5 Max within 1 percent; w64 bandwidth-bound (71.9 MH/s on the 5090) | the SRAM die's energy per read rises with the bits moved (0.25 to 0.38 nJ); the DRAM chip's toward the card's | 0 / 0 / 0 (measured) | 0 | the W = 8 rows (owed). The acceptance side of W = 16, for the record (the census lane, `docs/analysis/class-v6/census-w16-mix.md` on class-v6-census at d20eb04bd, 14:1x UK, measured): 256 of 256 accepted without an era at r 0.704 (the control 0.668) and 256 of 256 across eras 0 to 7, 0 instrument refusals, (c''') 1.4 percent, F8-form 0.999 to 1.006 of uniform, the verifier +0.6 percent; and the structural note that at W = 16 the product's low bits 0 to 3 are alignment and never enter the address (the control carries bit 0 biased in 116 of 256 programs, W = 4 moves it to bit 2 in 74 to 75) while the era-stride bit R stays on every width until the index fold. So W = 16 is clean on acceptance and dead on energy (10.5): the width is decided by the card's second sector, not by the generator | | Shadow block shape | 64 to 256 instructions per block, the pass count the ladder's | 64-instruction blocks ran 2.5 to 3.5 percent FASTER than 256 on the 5090 and the M5 Max (measured 6 October); 1,024 cost the M5 Max 17 percent | nothing for any chip (the work is the same) | +2.5 to 0 percent / pending / +2.5 to 0 | 0 | none | | **The index fold (a ring-A design rule of layer 1, every era's draw passes through it): `load_index` folds a product's low bits before the stride rotation, so no era's R lands a biased product bit on an address bit** | every era; lane D's coverage (11:3x UK): the class is at HALF the family's epochs, not a corner: 52 to 58 percent of accepted programs in every stratum carry one site whose address bit R (or R+1, R+2) is biased over 6 sigma at 2^20, 33 to 40 percent over 100 sigma, the worst z 1,024, on programs (c''') passes; the F8 tail's bucket excess is the same mechanism at scale (+94 to +128 sigma at the biased bit) | lane D's family harness (build-1, 10:21 UTC, 16 drawn eras, every layer-1 parameter from the era's stream, the chain draw through the real rule, reads at the rule's own 2^20 sample): the index-bit bias fires HARD on 7 of 16 eras, abs z 130 to 511 at one site, every one at address bit R or R+1 (a product's bit 0 at P = 0.25 on era 15, R = 25, z -511; a product's bit 1 at 3/8 on era 13; an or-shaped source at 5/8 on era 6); the other 9 clean under abs z 3.8; the (c''') ratio sees none of it (0.9954 to 1.0000): adv-cache-2's era-stride class measured on class v5 accepted programs at the acceptance's own sample | a chip holding the favoured half of that site's window serves 75 percent of its reads instead of 50: about 1.6 percent of a hash's reads at f = 1/2 for one site, zero at f = 1 (the partial store already costs 1.26x the ops, chip-model 5.4): the f = 1 verdict does not move; the row is an auditor's flag on "uniform random reads", not a chip lever | nothing: the fold is one xor-rotate on the address path, measured as 0 on every card by the era-layout rows (the index form is the era draw's own) | 0 | the fold's form in `load_index` (fold the product's low bits before the rotation) and its vectors; a 6-sigma REFUSAL in layer 4 is not the lever: it would redraw about 40 percent of epochs (7 of 16 eras); the known-failed test is lane D's 7 of 16 eras at the rule's sample, which must read 0 of 16 with the fold | @@ -294,7 +294,7 @@ The table, filled with the defaults, each cell replaced as its lane's row lands | Honest tier (measured joules per hash, class v4) | GDDR7 board (0.466): core / floor | HBM3 one stack (0.321): core / floor | SRAM die at W = 4 (0.051): core / floor | With SM-sparse | W = 16 variant | |---|---|---|---|---|---| -| RTX 5090 stock (3.36; class v3 2.26) | 4.1x core / 5.8x floor | 5.0x / 7.8x | 8.2x / 21x | 1.3 to 3.9 percent at best, MEASURED (10.1, 14:10 UK: four rented 5090s, the ceiling held to 43 of 170 SMs; the 4090 saves 4 W at 16 SMs, the H100 1.4 percent); the residual is the clock domain, which only the lock takes off; the fraction a chip cannot strip is 16 to 19 percent of the 5090's stock joules, 25 percent at the lock | not applied: KILLED on measured rows (10.5, 13:47 UK): the hinted form holds the rate on Ada but the card pays the second sector at 8 to 9 nJ, energy per hash +34 percent (4090), +33 (H100), the chip +33 (GDDR7); the edge does not move; the 5090 measured too (13:40 UK: the hinted sparse form holds the rate at +2.3 percent for +88 W, energy +25 percent, the GDDR7 edge 4.7x to 4.4x at zero shadow and unchanged with the shadow); the PC 1 knee row a formality | +| RTX 5090 stock (3.36; class v3 2.26) | 4.1x core / 5.8x floor | 5.0x / 7.8x | 8.2x / 21x | 1.3 to 3.9 percent at best, MEASURED (10.1, 14:10 UK: four rented 5090s, the ceiling held to 43 of 170 SMs; the 4090 saves 4 W at 16 SMs, the H100 1.4 percent); the residual is the clock domain, which only the lock takes off; the fraction a chip cannot strip is 16 to 19 percent of the 5090's stock joules, 25 percent at the lock | not applied: KILLED on measured rows (10.5, 13:47 UK): the hinted form holds the rate on Ada but the card pays the second sector at 8 to 9 nJ, energy per hash +34 percent (4090), +33 (H100), the chip +33 (GDDR7); the edge does not move; the 5090 measured too (13:40 UK: the hinted sparse form holds the rate at +2.3 percent for +88 W, energy +25 percent, the GDDR7 edge 4.7x to 4.4x at zero shadow and unchanged with the shadow); the PC 1 knee row measured in kit d (15:0x UK): the hinted w64-l2 pack 37.59 MH/s at 164.1 W at the lock, 0.229 MH/W against w4's 0.528; dead at the knee too | | RTX 5090 at the 1,300 MHz lock, the record's operating point (2.33 = 1.67 + 0.65) | 2.8x / 4.0x | 3.4x / 5.4x | 5.7x / 14x | no change, MEASURED (10.1 item 2: class v4 at the lock saves nothing on the full grid, 2.25 microjoules at 302 W on PC 1 today; the class v3 lock row stays 20.3b's, today's pass partial) | not applied (same default) | | Apple M5 Max (1.40; class v3 0.78; the GPU and DRAM channels) | 1.7x / 2.4x | 2.1x / 3.2x | 3.4x / 8.7x | not applicable (no SM lever on Apple) | not applicable (the M5 Max pays 0 at 64 bytes, measured; the chip +33 percent) | | RTX 5080 at its 1,100 MHz lock, the honest NVIDIA floor (2.06 class v4 measured: 71.20 MH/s at 146.6 W; class v5 2.10; floor lane 4, 10.4) | 2.5x / 3.6x | 3.0x / 4.8x | 5.0x / 13x | no change | not applied |