shadow-k: the ranking and the mixed draw (the census lane's re-weight on the unit floors and the core)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-josh 2026-10-08 14:10:39 +01:00
parent ed7d2626d4
commit e921670571

View file

@ -205,6 +205,54 @@ table compresses to 1.7x to 2.8x: the honest SoC's own operating point is most o
1's honest 5090 at 1.0 microjoules (assumed; its rows were not in by 18:00 UK) would put the GDDR7 board at 1.7x
to 2.3x and the strongest chips at 2.6x to 4.2x.
## 6. The ranking and the mixed draw
### 6.1 Families ranked by k, highest first (the hardest for a chip), absolute at N3 against the 5090's lock
| Rank | Family | Chip pJ/op N3 (unit floor) | 5090 pJ/op at the lock | k (floor) | In the class draw? |
|---|---|---|---|---|---|
| 1 | mad | 1.66 | 8.3 | 0.20 | yes (8 of 75) |
| 2 | add, sub, rotl, rotr, xor | 1.05 to 1.09 | 6.2 | 0.17 to 0.18 | yes (12, 6, 7, 6, 10) |
| 3 | the index fold | 1.19 | 8.3 | 0.14 | every load |
| 4 | or | 0.73 | 6.2 | 0.12 | yes (4) |
| 5 | mul | 0.68 | 8.3 | 0.082 | yes (8) |
| 6 | prmt, lop3 | 0.64, 0.72 | 11.5, 13.0 | 0.056, 0.055 | not drawn (RTL rows only) |
| 7 | mulhi | 0.68 | 21.0 | 0.032 | yes (6) |
| 8 | 32-lane shuffle (butterfly) | ROW_SHFL_K | 29.4 | pending | yes (8) |
| 9 | L1 scratch read (8 KB flop array) against the card's L2 hit | ROW_SCRATCH_K | 1,400 | pending | not drawn |
| 10 | int8 8x8x8 tile, per MAC | ROW_TILE_K | 2.2 | pending | not drawn (the tensor lever is dead on other grounds) |
The order is set by the card's price, not the chip's: the chip pays 0.6 to 1.7 pJ for everything, and the card
pays 6.2 for an add, 8.3 for a multiply, 21 for a high word and 29 for a shuffle. So the families the card pays
MOST for (mulhi, shuffle) are the ones a chip undercuts most, and the forcing work is the plain ARX and mad the
card does cheapest. This is the measured form of 15.1b's hold on the shuffle-heavy re-weight.
### 6.2 The mixed draw that maximises the expected k
The objective at a FIXED GPU premium: `k_eff(w) = sum w_i e_chip_i / sum w_i e_gpu_i` (the chip's energy for the
drawn program over the card's for the same program; at a fixed `F` the chip pays `k_eff x F`). The band (layer
1, the research lane's 17:00 reading of lane D's cut): B = 4 points on the injecting families only (add, sub,
xor, mad, shfl, rotl, rotr), the lossy families (or, mul, mulhi) at or under their base, `or + mul + mulhi` at most
18 + B, the shuffle capped at its class v4 weight; the index fold on every address and the F8 uniformity floor are
not functions of the weights and do not move. `tools/chip-model/rtl/flow/mix.py` searches the band exhaustively
over the corners (the objective is a ratio of linear forms, so the optimum is at a corner of each family's range).
On the unit floors (N3, the lock; the shuffle row pending, so computed over the other 67 points of 75):
| Mix | add, xor, mul, mad, shfl, rotl, sub, mulhi, rotr, or | k_eff (floors) | The card's pJ per op at the lock | k_eff on the core (floor + the core's 2.4 pJ per op overhead at N3) |
|---|---|---|---|---|
| class v4 (the base) | 12, 10, 8, 8, 8, 7, 6, 6, 6, 4 | 0.130 | 8.0 | 0.43 |
| the census lane's re-weight (d20eb04bd, PASS on 256 seeds, no era and eras 0 to 7, W = 4 and W = 16) | 13, 11, 6, 10, 8, 8, 7, 2, 6, 4 | 0.153 (+18 percent) | 7.1 | 0.49 (+14 percent) |
| the band's corner (every injecting family at +4 except the shuffle at its floor; lossy at base minus 4) | 16, 14, 4, 12, 4, 11, 10, 2, 10, 0 | MIX_BEST | | |
Reading: the re-weight moves six points off the chip's two easiest families (mulhi, mul) onto the hardest (mad,
the ARX families) and lifts the expected k by about a seventh on the core; its acceptance is the census lane's
(0 exhausted, F8 0.999 to 1.008 of uniform, the verifier within 4 percent); its energy cost on the card is a
premium per instruction 11 percent LOWER (fewer mulhi), so at a fixed premium the program is 12 percent longer.
The shuffle's own row decides whether it goes to its floor: on the card it is the dearest op (29.4 pJ at the lock)
and on the chip a butterfly over the window, so its k is the lowest of the drawn families unless the synthesis says
otherwise. The recommendation stands as the census lane's draw until the shuffle row lands (an amendment).
## 7. Sources
- The GPU side: `docs/analysis/counter-asic-4-research.md` 15.1a (the 5090 microbench, 8 October 2026, measured)