diff --git a/docs/analysis/class-v6/floor/shadow-k.md b/docs/analysis/class-v6/floor/shadow-k.md index fb9544ffd..bf294bc95 100644 --- a/docs/analysis/class-v6/floor/shadow-k.md +++ b/docs/analysis/class-v6/floor/shadow-k.md @@ -205,6 +205,54 @@ table compresses to 1.7x to 2.8x: the honest SoC's own operating point is most o 1's honest 5090 at 1.0 microjoules (assumed; its rows were not in by 18:00 UK) would put the GDDR7 board at 1.7x to 2.3x and the strongest chips at 2.6x to 4.2x. +## 6. The ranking and the mixed draw + +### 6.1 Families ranked by k, highest first (the hardest for a chip), absolute at N3 against the 5090's lock + +| Rank | Family | Chip pJ/op N3 (unit floor) | 5090 pJ/op at the lock | k (floor) | In the class draw? | +|---|---|---|---|---|---| +| 1 | mad | 1.66 | 8.3 | 0.20 | yes (8 of 75) | +| 2 | add, sub, rotl, rotr, xor | 1.05 to 1.09 | 6.2 | 0.17 to 0.18 | yes (12, 6, 7, 6, 10) | +| 3 | the index fold | 1.19 | 8.3 | 0.14 | every load | +| 4 | or | 0.73 | 6.2 | 0.12 | yes (4) | +| 5 | mul | 0.68 | 8.3 | 0.082 | yes (8) | +| 6 | prmt, lop3 | 0.64, 0.72 | 11.5, 13.0 | 0.056, 0.055 | not drawn (RTL rows only) | +| 7 | mulhi | 0.68 | 21.0 | 0.032 | yes (6) | +| 8 | 32-lane shuffle (butterfly) | ROW_SHFL_K | 29.4 | pending | yes (8) | +| 9 | L1 scratch read (8 KB flop array) against the card's L2 hit | ROW_SCRATCH_K | 1,400 | pending | not drawn | +| 10 | int8 8x8x8 tile, per MAC | ROW_TILE_K | 2.2 | pending | not drawn (the tensor lever is dead on other grounds) | + +The order is set by the card's price, not the chip's: the chip pays 0.6 to 1.7 pJ for everything, and the card +pays 6.2 for an add, 8.3 for a multiply, 21 for a high word and 29 for a shuffle. So the families the card pays +MOST for (mulhi, shuffle) are the ones a chip undercuts most, and the forcing work is the plain ARX and mad the +card does cheapest. This is the measured form of 15.1b's hold on the shuffle-heavy re-weight. + +### 6.2 The mixed draw that maximises the expected k + +The objective at a FIXED GPU premium: `k_eff(w) = sum w_i e_chip_i / sum w_i e_gpu_i` (the chip's energy for the +drawn program over the card's for the same program; at a fixed `F` the chip pays `k_eff x F`). The band (layer +1, the research lane's 17:00 reading of lane D's cut): B = 4 points on the injecting families only (add, sub, +xor, mad, shfl, rotl, rotr), the lossy families (or, mul, mulhi) at or under their base, `or + mul + mulhi` at most +18 + B, the shuffle capped at its class v4 weight; the index fold on every address and the F8 uniformity floor are +not functions of the weights and do not move. `tools/chip-model/rtl/flow/mix.py` searches the band exhaustively +over the corners (the objective is a ratio of linear forms, so the optimum is at a corner of each family's range). + +On the unit floors (N3, the lock; the shuffle row pending, so computed over the other 67 points of 75): + +| Mix | add, xor, mul, mad, shfl, rotl, sub, mulhi, rotr, or | k_eff (floors) | The card's pJ per op at the lock | k_eff on the core (floor + the core's 2.4 pJ per op overhead at N3) | +|---|---|---|---|---| +| class v4 (the base) | 12, 10, 8, 8, 8, 7, 6, 6, 6, 4 | 0.130 | 8.0 | 0.43 | +| the census lane's re-weight (d20eb04bd, PASS on 256 seeds, no era and eras 0 to 7, W = 4 and W = 16) | 13, 11, 6, 10, 8, 8, 7, 2, 6, 4 | 0.153 (+18 percent) | 7.1 | 0.49 (+14 percent) | +| the band's corner (every injecting family at +4 except the shuffle at its floor; lossy at base minus 4) | 16, 14, 4, 12, 4, 11, 10, 2, 10, 0 | MIX_BEST | | | + +Reading: the re-weight moves six points off the chip's two easiest families (mulhi, mul) onto the hardest (mad, +the ARX families) and lifts the expected k by about a seventh on the core; its acceptance is the census lane's +(0 exhausted, F8 0.999 to 1.008 of uniform, the verifier within 4 percent); its energy cost on the card is a +premium per instruction 11 percent LOWER (fewer mulhi), so at a fixed premium the program is 12 percent longer. +The shuffle's own row decides whether it goes to its floor: on the card it is the dearest op (29.4 pJ at the lock) +and on the chip a butterfly over the window, so its k is the lowest of the drawn families unless the synthesis says +otherwise. The recommendation stands as the census lane's draw until the shuffle row lands (an amendment). + ## 7. Sources - The GPU side: `docs/analysis/counter-asic-4-research.md` 15.1a (the 5090 microbench, 8 October 2026, measured)