diff --git a/docs/design/class-v6-rotating-family.md b/docs/design/class-v6-rotating-family.md index 2086edb2d..471668615 100644 --- a/docs/design/class-v6-rotating-family.md +++ b/docs/design/class-v6-rotating-family.md @@ -56,7 +56,7 @@ The band a card sees across the draw's range, from the hash lane's cost rows (12 |---|---|---|---|---|---|---| | Mixer applications per round `m` | **{4, 8} at genesis; 16 in the list as `admissible: false` (the ladder's rule: a flag in the genesis list, flipped only by the 90 percent upgrade path once the 2019-class core measures it)** | **The card pays nothing for the draw, MEASURED (the hash lane's kit b on PC 1, 12:49 to 12:58 UK, generator 2 with the era, the multiplier the only difference, 250 batches per row, fingerprints PASS): x8 137.73 MH/s at 320.2 W unlocked and 127.44 at 212.0 W at the 1,300 lock; x16 137.72 at 320.1 and 127.46 at 211.7; within 0.1 MH/s and 0.5 W at either state, item derivation hidden under the read chain, so the verifier's 1.86x per doubling is the draw's whole cost.** x4 and x8 measured (mixer-x4.md 6.4: 1.92 and 2.79 ms per unit on a loaded M5 Max core; the build unmoved on every discrete card, latency-bound); **x16 MEASURED (today's class v3 stream, the same id f5e904bc5d148926 for x8 and x16). Cold alone: build-1 core 4 x8 4.67 ms per warp, x16 8.69 (the hash lane, 10:13 UTC); build-3 core 4 (AX102, a faster core class) x8 3.48, x16 6.53 (the build-server lane, 11:01 UTC). With the SMT sibling loaded for the row's whole length (the method that counts: a sibling bench looped for the row, not run once, which ends in about a second and leaves the verify phase sibling-idle; the hash lane's earlier 5.41 / 10.11 / 9.04 to 9.10 ms rows were of that kind and are withdrawn): build-3 x8 6.24, x16 11.44 ms, about 1.8x on both classes. The multiplier is 1.86x on the verifier (the mixer IS the verifier's cost); chip-model-v3's estimate (9 ms on a 2019-class core) stands within the cold rows** | the `f = 0` recompute chip's rate halves per doubling (0.31x bare at x8, 0.16x at x16, modelled); the `f = 1` chip unmoved (it recomputes nothing) | rate 0 / 0 / 0 (latency-bound, measured at x4 and x8); the daily build 42 / pending / 29 ms at x8 (measured); the x16 build on the 5090 rides the 12:45 UK job | +1.9x per doubling measured (x8 to x16); **x16 at 11.4 ms loaded is over the 10 ms gate on the measured row, so `admissible: false` stands on the measurement, not only on the chip model's margin; the O-1.14 laptop run can only tighten it** | adv-mixer-3's margin (every statistic clean from k = 1, the SAT ladder k = 1 solved, k = 2..4 timeout): x4 keeps the margin by that report's reading; the band {4, 8} is the measured one | | Op-mix weights (the ten non-load families) | **B = 4 points on the injecting families only (add, sub, xor, mad, shfl, rotl, rotr); the lossy families (or, mul, mulhi) capped at their base, the ring-A rule `or + mul + mulhi` at most the table's 18 plus B, which keeps the per-candidate rejection r under 0.85 (r^256 under 1e-18)**; the shuffle weight capped at its class v4 value on the energy side | lane D's coverage run (4,900 drawn eras on three boxes, 11:3x UK, measured): at the lossy corner of B = 4 (or, mul and mulhi all at +4, 30 of 75 lossy against 18) r is 0.956 (the shipped 0.681) and 8 of 663 eras exhaust the 256-attempt cap (mean attempt 18.6, max 252), so 1.2 percent of that corner's epochs would take the last-resort program, which fails rule (a) in 9 percent of seeds (adv-accept-3); the random stratum with every weight drawn reads r 0.718, max attempt 107, 0 exhaustions in 1,444. The energy side (15.1a): shfl 55.8 pJ per op, mulhi 39.6, prmt 22.3, lop3 24.1, mul 13.9, arx 11.3; a shuffle-heavy table raises the premium per instruction up to 2x (modelled), a multiply-heavy one 1.1x | a chip that specialised its lane ratio loses the ratio; a general core nothing | the premium per instruction moves with the mix within the capped band (the cost rows: at N = 100,000 the 5090's +146 W stock becomes up to +160 W multiply-heavy; shuffle-heavy excluded by the cap); the 5070 Ti scaled by 83/146, the M5 Max by 16/146; the two packs' measured rows replace this when they land. **First row MEASURED (kit b, PC 1, 12:5x UK): the shuffle-heavy table (shfl 13 of 48 against the stock 5) on class mx8 without a shadow block reads 137.65 MH/s at 312.9 W unlocked (7 W under the stock table's 320.2) and 127.33 at 212.4 at the lock (level): the multiply-heavy table (mulhi 11, mad 8, mul 5 of 48 against the stock 9, 4, 1; shfl 1 against 5; kit c, 13:03 to 13:06 UK) reads 137.77 at 313.5 W unlocked and 128.42 at 211.8 W at the lock, so the three tables sit within 7 W unlocked and 0.6 W at the lock: on the 48-op base program the weight move is a 2 percent term whichever way it goes; the row layer 1 needs is the same two tables inside the shadow block, MEASURED (kit d, PC 1's 5090, 14:51 to 15:0x UK, the v5 kit's CUDA worker, 250 x 2^24 per row, all PASS; the caveat: this job's rows run on the kit worker, whose base reads 120.0 MH/s unlocked against the installed worker's 137.6 this morning, so the like-for-like comparison is within the job and the comparison to the morning's stock table is in watts only). Unlocked: the w4 base 119.95 MH/s at 308.2 W; the multiply-heavy table in the block (mulhi 11, mad 8, mul 5, shfl 1 of 48) 118.28 at 412.8 W, a block premium of 104.6 W (0.88 microjoules per hash); the shuffle-heavy table (shfl 13 of 48) 118.38 at 463.7 W, 155.5 W (1.31); at the 1,300 lock: the base 100.55 at 190.5 W, multiply-heavy 98.66 at 252.5 W (62.0 W, 0.63), shuffle-heavy 98.27 at 271.1 W (80.6 W, 0.82); the morning's stock-table premiums on the installed worker 152.4 W unlocked and 84.2 W locked. **Inside the shadow block the weight table is a 30 percent lever on the block's energy: the multiply-heavy table costs the card 33 percent less than the shuffle-heavy one per hash unlocked and 23 percent less at the knee at a flat rate (within 0.4 percent), and the shuffle-heavy table costs the same as the stock table in watts (within 3 W).** The microbench's arithmetic had the sign right at the multiply-heavy end (1.1x modelled against about 0.7x measured: the mulhi and mad units are cheaper per op than the ARX mix inside a dependent chain on Blackwell) and the shuffle-heavy end reads 1.0x against the 2x modelled (the shuffles are already most of the stock table's cost). What it does to the band: a multiply-heavy table lowers the card's premium by a quarter, but the k lane's rows (10.2) price mul and mulhi on the chip at k 0.03 to 0.08 against the ARX families' 0.18, so the chip's premium falls further than the card's and the edge rises; the cap on the lossy families at their base stands on both the acceptance side (lane D) and the energy side (this row). The shuffle cap on the energy side is now measured as costless to the card (shuffle-heavy = stock in watts), so if the k lane's butterfly row reads above the ARX k the shuffle weight can rise inside B = 4 at no card cost; that row is owed. The acceptance side of a re-weight (the census lane, d20eb04bd, 14:1x UK, measured): W = 4 at weights 13, 11, 6, 10, 8, 8, 7, 2, 6, 4 (mulhi and mul down, mad and four ARX families up) 256 of 256 both ways at r 0.407, (c''') 1.2 percent, F8-form 0.999 to 1.008, the product law at bit 2 in 75 of 256 programs; both W = 16 and the re-weight together 256 of 256 at r 0.426, (c''') 3.4 percent. So a re-weight is clean on acceptance; the hold on it is the energy side (15.1b) and now the chip side too (10.2: mulhi is the card's worst lever by 6x, and a mix that lowers mulhi helps the card against the chip only if the k lane's sweep says so)** | +0 ms | the known-failed test: the 8 of 663 exhaustions at the uncapped lossy corner (61 of 5,000 on the 17:00 cut), which the capped band must read as 0: **MEASURED 0 of 3,000 under the band (lane D's 17:00 cut, 5.1b; r 0.595, mean attempt 1.47, max 59)** | -| Read width `W` | **pinned at 4 words (16 bytes) at genesis, not drawn** (floor lane 3, 10.3: the width is the only wire lever on the SRAM die, 66x at 4 bytes against 44x at 16 at zero shadow; 8 words after the owed PC 1 and Mac rows; 16 never) | w4 and w16 measured 5 October: the 5090 139.8 against 136.1 MH/s, the 9070 XT 17.90 against 18.15, the M5 Max within 1 percent; w64 bandwidth-bound (71.9 MH/s on the 5090) | the SRAM die's energy per read rises with the bits moved (0.25 to 0.38 nJ); the DRAM chip's toward the card's | 0 / 0 / 0 (measured) | 0 | the W = 8 rows (owed). The acceptance side of W = 16, for the record (the census lane, `docs/analysis/class-v6/census-w16-mix.md` on class-v6-census at d20eb04bd, 14:1x UK, measured): 256 of 256 accepted without an era at r 0.704 (the control 0.668) and 256 of 256 across eras 0 to 7, 0 instrument refusals, (c''') 1.4 percent, F8-form 0.999 to 1.006 of uniform, the verifier +0.6 percent; and the structural note that at W = 16 the product's low bits 0 to 3 are alignment and never enter the address (the control carries bit 0 biased in 116 of 256 programs, W = 4 moves it to bit 2 in 74 to 75) while the era-stride bit R stays on every width until the index fold. So W = 16 is clean on acceptance and dead on energy (10.5): the width is decided by the card's second sector, not by the generator | +| Read width `W` | **pinned at 4 words (16 bytes) at genesis, not drawn** (floor lane 3, 10.3: the width is the only wire lever on the SRAM die, 66x at 4 bytes against 44x at 16 at zero shadow; 8 words measured NOT free on the 5090 at stock, +4.8 percent of energy for 0.2x of the die's shadowed edge, so it does not pin unless its lock row reverses the term; 16 never) | w4 and w16 measured 5 October: the 5090 139.8 against 136.1 MH/s, the 9070 XT 17.90 against 18.15, the M5 Max within 1 percent; w64 bandwidth-bound (71.9 MH/s on the 5090) | the SRAM die's energy per read rises with the bits moved (0.25 to 0.38 nJ); the DRAM chip's toward the card's | 0 / 0 / 0 (measured) | 0 | the W = 8 rows (owed). The acceptance side of W = 16, for the record (the census lane, `docs/analysis/class-v6/census-w16-mix.md` on class-v6-census at d20eb04bd, 14:1x UK, measured): 256 of 256 accepted without an era at r 0.704 (the control 0.668) and 256 of 256 across eras 0 to 7, 0 instrument refusals, (c''') 1.4 percent, F8-form 0.999 to 1.006 of uniform, the verifier +0.6 percent; and the structural note that at W = 16 the product's low bits 0 to 3 are alignment and never enter the address (the control carries bit 0 biased in 116 of 256 programs, W = 4 moves it to bit 2 in 74 to 75) while the era-stride bit R stays on every width until the index fold. So W = 16 is clean on acceptance and dead on energy (10.5): the width is decided by the card's second sector, not by the generator | | Shadow block shape | 64 to 256 instructions per block, the pass count the ladder's | 64-instruction blocks ran 2.5 to 3.5 percent FASTER than 256 on the 5090 and the M5 Max (measured 6 October); 1,024 cost the M5 Max 17 percent | nothing for any chip (the work is the same) | +2.5 to 0 percent / pending / +2.5 to 0 | 0 | none | | **The index fold (a ring-A design rule of layer 1, every era's draw passes through it): `load_index` folds a product's low bits before the stride rotation, so no era's R lands a biased product bit on an address bit** | every era; lane D's coverage (11:3x UK): the class is at HALF the family's epochs, not a corner: 52 to 58 percent of accepted programs in every stratum carry one site whose address bit R (or R+1, R+2) is biased over 6 sigma at 2^20, 33 to 40 percent over 100 sigma, the worst z 1,024, on programs (c''') passes; the F8 tail's bucket excess is the same mechanism at scale (+94 to +128 sigma at the biased bit) | lane D's family harness (build-1, 10:21 UTC, 16 drawn eras, every layer-1 parameter from the era's stream, the chain draw through the real rule, reads at the rule's own 2^20 sample): the index-bit bias fires HARD on 7 of 16 eras, abs z 130 to 511 at one site, every one at address bit R or R+1 (a product's bit 0 at P = 0.25 on era 15, R = 25, z -511; a product's bit 1 at 3/8 on era 13; an or-shaped source at 5/8 on era 6); the other 9 clean under abs z 3.8; the (c''') ratio sees none of it (0.9954 to 1.0000): adv-cache-2's era-stride class measured on class v5 accepted programs at the acceptance's own sample | a chip holding the favoured half of that site's window serves 75 percent of its reads instead of 50: about 1.6 percent of a hash's reads at f = 1/2 for one site, zero at f = 1 (the partial store already costs 1.26x the ops, chip-model 5.4): the f = 1 verdict does not move; the row is an auditor's flag on "uniform random reads", not a chip lever | nothing: the fold is one xor-rotate on the address path, measured as 0 on every card by the era-layout rows (the index form is the era draw's own) | 0 | the fold's form in `load_index` (fold the product's low bits before the rotation) and its vectors; a 6-sigma REFUSAL in layer 4 is not the lever: it would redraw about 40 percent of epochs (7 of 16 eras); the known-failed test is lane D's 7 of 16 eras at the rule's sample, which must read 0 of 16 with the fold | | Program length N | not drawn: the ladder's signal (latency-ladder.md) | an unconditional draw retires the Apple tier at 200,000 (measured -10 percent). **The price of a heavier shadow, MEASURED (the founder's 1.5x test, the 1p5x-knobs lane on a rented 5090 and 4090 at stock, 14:28 to 14:39 UK, 250 x 2^24, nvidia-smi 1 Hz, one fingerprint on both cards): class v5 genesis (55,296 shadow ops per hash) against a 1,024-instruction block at 27 passes on the same state (hl-k3-sh1024, generator 5, 221,184 ops): the 5090 140.83 MH/s at 442.7 W (3.14 microjoules) against 111.07 at 551.1 W (4.96; 110.0 at the 575 W limit over 60 s, 5.23); the 4090 62.41 at 279.3 W (4.48) against 62.64 at 439.2 W (7.01): 4x the shadow instructions cost 1.58x the energy per hash on both architectures (0.40x per instruction against the 256-block, the per-pass overhead amortised), power-bound on the 5090 (21 percent fewer MH/s), rate-neutral on the 4090 (+61 percent watts). This replaces the hash lane's modelled +290 W at N = 200,000 with a measured point** | the core sized for the ladder's admissible top (rung 2, 199,600 ops); the chip pays k of the heavier shadow, the card all of it | the ladder's rows; at 4x the shadow the 5090 +1.58x per hash (measured) | the ladder's rows | none | @@ -307,7 +307,7 @@ The capex wall's two thresholds (lane 3, 10.3, re-folded 13:18 UK on lane 5's co The served line in one sentence, on the synthesised core with the 64-register window (10.0e), in the two units: a chip that stores the dataset reaches about 2.2x to 2.4x per joule against the honest NVIDIA tiers at their knee (the 5080 2.2x, the 5090 2.4x with a chip a node ahead; 2.0x against a chip on the GPU's own node, 2.8x two nodes ahead; 2.8x on the 5090 without the window) and under 2x only against the Apple tier (the M5 Max 1.5x), about 4x and 2.6x if it is built on an N2 SRAM store for USD 100 M or more, and under 1x per hash over its 180-day class life only above about USD 300 M a year of miner revenue (10.0a), with Monero's measured 1.0x to 1.5x beside it; no rotating layer moves the per-joule figures, and layer 3's class life is what sets the per-hash one; the layers decide which chip can be built and how long its tape-out lives. The worst case beside it, if a chip's core costs no more than its bare units (the floor k of 10.2): 4x and 2.4x against the DRAM board, 14x and 9x against the SRAM die (lane 3's line: 6x to 10x at the full shadow, never under 2x). In the other convention, the card's whole 330,000-op latency shadow priced on the chip at the synthesised core (lane 3's 10.3 table): the SRAM die 3.5x against a 5090 at stock and 2.2x at its knee (4.8x and 3.0x with a 2 nm core), the DRAM board 2.6x at stock. The record's claimed band (k 0.3 to 0.8) is confirmed from the chip side by the core row (0.56 at the lock), so the served 2.1x at k = 1 is the conservative end and 2.8x the measured-core reading. -The three class v6 changes it implies: (1) the op mix stays class v4's with the lossy families capped at their base (lane D: 0 of 3,000 eras exhausted under the band; the k lane: mulhi is the card's worst lever by 6x, so no re-weight helps the card); (2) no SM-sparse default (measured dead on the 5090, 4090 and H100 at 1.3 to 3.9 percent at best; `--sm-sparse auto` ships off by default, on in the Efficiency and Balanced tiers at its measured 1 to 2.5 percent, 10.1 item 6; the per-watt gift of the hot table is the chip's, 20.3c); (3) the dataset schedule 5.5 / 8.5 / 11.5 GiB with the read width pinned at 4 words; and a fourth the sweep adds, (4) the 64-register window per lane as the core shape (k 0.78 at the lock against 0.56; 10.0c, the edge table 10.0e), its GPU side modelled until the generator carries a 64-entry init and fold rule (a half-day line) and a pack is measured (the chip's ticket USD 1,500 / 2,500 / 3,000; a tuned 5090 pays 1 to 5 percent per step; about a quarter of today's cards by count per step). +The three class v6 changes it implies: (1) the op mix stays class v4's with the lossy families capped at their base (lane D: 0 of 3,000 eras exhausted under the band; the k lane: mulhi is the card's worst lever by 6x, so no re-weight helps the card); (2) no SM-sparse default (measured dead on the 5090, 4090 and H100 at 1.3 to 3.9 percent at best; `--sm-sparse auto` ships off by default, on in the Efficiency and Balanced tiers at its measured 1 to 2.5 percent, 10.1 item 6; the per-watt gift of the hot table is the chip's, 20.3c); (3) the dataset schedule 5.5 / 8.5 / 11.5 GiB with the read width pinned at 4 words (8 measured not free, 16 never); and a fourth the sweep adds, (4) the 64-register window per lane as the core shape (k 0.78 at the lock against 0.56; 10.0c, the edge table 10.0e), its GPU side modelled until the generator carries a 64-entry init and fold rule (a half-day line) and a pack is measured (the chip's ticket USD 1,500 / 2,500 / 3,000; a tuned 5090 pays 1 to 5 percent per step; about a quarter of today's cards by count per step). #### 10.0a The second headline unit: cost per hash over the chip's life (the founder's question, 14:1x UK; modelled on measured card rows) @@ -477,10 +477,10 @@ The k lane's RTL rows (floor lane 2, in absolute) replace the k axis of this tab |---|---|---|---|---|---|---|---|---| | 1 (4 bytes, today) | 80 | 0.25 nJ | 8.3 | **66x** | 5.7x / 3.0x | 4.1x / 2.1x | 0 | modelled chip; measured card | | 4 (16 bytes) | 176 | 0.38 | 5.6 | 44x | 5.6x / 2.9x | 4.0x / 2.1x | the 5090 +2.7 percent, the 9070 XT -1.4, the M5 Max within 1 (measured 5 October) | measured card | -| 8 (32 bytes, the GDDR7 sector the 5090 fetches anyway) | 304 | 0.55 | 3.9 | 31x | 5.4x / 2.9x | 4.0x / 2.0x | free by the measured rows (13:06 UK: the 5090's ceiling about 18 G sectors a second, w16 moved 573 GB/s of sectors at 139.8 MH/s and w64 589 at 71.9; an aligned 32-byte read is one sector per load as w16; AMD moves its 64 B line either way; the M5 Max 1.03 at both w16 and w64); one PC 1 row to confirm, and the crate's width set is {1, 4, 16} words, so a pack at 8 needs a generator and emitter line first | measured card (inferred at 8) | +| 8 (32 bytes) | 304 | 0.55 | 3.9 | 31x (32.5x net) | 5.4x / 2.9x | 4.0x / 2.0x | NOT free, MEASURED (the hash lane on PC 1's 5090 at stock under the class v5 state term, 15:03 to 15:06 UK, 250 x 2^24, every row PASS): the rate held to 0.01 MH/s (117.54 against 117.54 at W = 4) but the energy rose 4.8 percent (451.5 W against 430.9; 3.84 against 3.67 microjoules), the pair's second sector costing the card about 1.2 nJ a read; the 13:06 interpolation (one sector per load as w16) is withdrawn by lane 3 on this row | measured card | | 16 (64 bytes, lane B's record) | 560 | 0.88 | 2.4 | 19x (the record's 17x) | 5.0x / 2.7x | 3.8x / 2.0x | the 5090 -47 percent: dead | measured card | -Meaning: at the hash's own width the die is 3.5x stronger than the record said at zero shadow; with the shadow on, every row sits at 5.0x to 5.7x (k 0.5) and 2.7x to 3.0x (k 1): **the shadow is the whole hold, the memory moves it 0.2x to 0.7x.** The fold is dst-keyed (`verify::fold_words`), so a wide read cannot be pre-folded; dependent chains per step leave the ratio unchanged (per-read on both sides); banking plus a sequencer cannot localise (uniform on a window of at least 256 MiB, the 16 sites alternating; moving the lane costs about 290 bits, which caps the wire lever at about W = 8); the per-site window shrink buys the die and the DRAM chip nothing; a hop at the honest widths is 0.04 to 0.15 nJ, so a ten-die store still reads 25x to 58x at zero shadow. This moves layer 1's width row: **pin W = 4 (16 bytes) at genesis (measured free on all three vendors; the die from 66x to 44x), W = 8 (31x; 5.4x at k 0.5) once the generator and emitter carry it and one PC 1 row confirms, never 16.** +Meaning: at the hash's own width the die is 3.5x stronger than the record said at zero shadow; with the shadow on, every row sits at 5.0x to 5.7x (k 0.5) and 2.7x to 3.0x (k 1): **the shadow is the whole hold, the memory moves it 0.2x to 0.7x.** The fold is dst-keyed (`verify::fold_words`), so a wide read cannot be pre-folded; dependent chains per step leave the ratio unchanged (per-read on both sides); banking plus a sequencer cannot localise (uniform on a window of at least 256 MiB, the 16 sites alternating; moving the lane costs about 290 bits, which caps the wire lever at about W = 8); the per-site window shrink buys the die and the DRAM chip nothing; a hop at the honest widths is 0.04 to 0.15 nJ, so a ten-die store still reads 25x to 58x at zero shadow. This moves layer 1's width row: **pin W = 4 (16 bytes) at genesis (measured free on all three vendors; the die from 66x to 44x); W = 8 does NOT pin (measured +4.8 percent of the card's energy at stock for 0.2x of the die's shadowed edge, lane 3's 083ea1e4 at 15:15 UK; its 1,300 MHz lock row, lost to PC 1's app restart and republished for about 16:30, is the one amendment that could reverse it); W = 16 never.** The floor as a ticket: SRAM is flat at about USD 250 per GiB to 2031 (density +6 to 11 percent per node against dearer wafers; claimed, approximate), so USD 5,000 of silicon per store is 20 GiB at N2 and about 21 GiB in 2031, which retires every card under 32 GB and every Mac under 64 GB; the constraint and "fewest cards" cannot both hold. The per-MH/s does not rise with the floor (every die powered: USD 0.25 to 0.5 per MH/s at any size); the floor raises the minimum ticket only. The honest card's side of the floor is measured (the hash lane's kit b, section 3.3): the 5090 at its knee pays 4 / 8 / 10 percent more energy per hash at 2 / 4 / 8 GiB, so every chip row's edge against a card at its knee rises by 4 to 11 percent across the schedule while the die's joules do not move; the schedule's sentence is USD 1,000 of chip ticket per step for 1 to 5 percent of the tuned 5090's energy and about a quarter of today's cards by count. The k convention matters at the lock: the record's (k against the GPU's op cost at the same operating point) reads the SRAM die at 6.1x (k 0.5) and 3.3x (k 1); the absolute convention (the die's core costs what it costs, k against the stock 11.3 pJ per op) reads 3.8x and 2.0x; the record's flatters the die by 1.6x at the lock, so the served line in section 9 and the close carry the absolute rows as the headline and the record's beside them marked. @@ -534,7 +534,7 @@ That table was the rate question; the energy question answered it (the lane's bb | HBM3 | 8.4x (9.0x), up from 7.5x | 8.4x | 6.7x | 7.5x / 7.2x / 5.9x | modelled chip, measured card | | SRAM die | 24x (26x), from 66x | 16.7x | 9.9x | 31x / 17.9x / 9.7x | modelled chip, measured card | -W = 16 costs the honest card 25 to 34 percent of its energy so that the DRAM chips' edges rise and the SRAM die lands exactly where W = 8 puts it for free with the shadow on; the only thing it buys is the die's zero-shadow number, which no served line carries. **W = 8 is the width; "never 16" stands on measured rows at stock on the 5090, 4090, H100 and 3090**; the 5090's knee row landed (kit d, 15:0x UK: 2.3x the energy per hash at the lock) and closes it; nothing on W = 16 is owed. +W = 16 costs the honest card 25 to 34 percent of its energy so that the DRAM chips' edges rise and the SRAM die lands exactly where W = 8 puts it for free with the shadow on; the only thing it buys is the die's zero-shadow number, which no served line carries. **W = 4 is the width (W = 8 withdrawn on the 15:0x measurement: +4.8 percent of the card's energy for 0.2x of the die's shadowed edge); "never 16" stands on measured rows at stock on the 5090, 4090, H100 and 3090**; the 5090's knee row landed (kit d, 15:0x UK: 2.3x the energy per hash at the lock) and closes it; nothing on W = 16 is owed. (3) The shadowed SRAM rows on the k lane's synthesised core (0.11 microjoules of shadow at N3, absolute k about 0.1, the shuffle open; the lane's 2.2b): under class v4 21x at W = 4 and 18x at W = 8 at stock, 14x and 12x at the knee; at the honest cards' whole latency shadow 10x and 9.7x at stock, 6.5x and 6.1x at the knee; an N2 core about 1.2x more. On the synthesised core the shadow is not the whole hold (the record's claimed band read 2x to 6x, marked beside): the memory is a third to a half of the die's energy, the width is worth 1.3x with the shadow on, and the lane's line for the served sentence is 6x to 10x at the full shadow (2x to 4x on the claimed band, marked), never under 2x. These are the per-unit-floor figures of 10.2, now the marked worst case; the core-row fold follows.