shadow-k: the ranking and the mixed draw (the census lane's re-weight on the unit floors and the core)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
ed7d2626d4
commit
e921670571
1 changed files with 48 additions and 0 deletions
|
|
@ -205,6 +205,54 @@ table compresses to 1.7x to 2.8x: the honest SoC's own operating point is most o
|
|||
1's honest 5090 at 1.0 microjoules (assumed; its rows were not in by 18:00 UK) would put the GDDR7 board at 1.7x
|
||||
to 2.3x and the strongest chips at 2.6x to 4.2x.
|
||||
|
||||
## 6. The ranking and the mixed draw
|
||||
|
||||
### 6.1 Families ranked by k, highest first (the hardest for a chip), absolute at N3 against the 5090's lock
|
||||
|
||||
| Rank | Family | Chip pJ/op N3 (unit floor) | 5090 pJ/op at the lock | k (floor) | In the class draw? |
|
||||
|---|---|---|---|---|---|
|
||||
| 1 | mad | 1.66 | 8.3 | 0.20 | yes (8 of 75) |
|
||||
| 2 | add, sub, rotl, rotr, xor | 1.05 to 1.09 | 6.2 | 0.17 to 0.18 | yes (12, 6, 7, 6, 10) |
|
||||
| 3 | the index fold | 1.19 | 8.3 | 0.14 | every load |
|
||||
| 4 | or | 0.73 | 6.2 | 0.12 | yes (4) |
|
||||
| 5 | mul | 0.68 | 8.3 | 0.082 | yes (8) |
|
||||
| 6 | prmt, lop3 | 0.64, 0.72 | 11.5, 13.0 | 0.056, 0.055 | not drawn (RTL rows only) |
|
||||
| 7 | mulhi | 0.68 | 21.0 | 0.032 | yes (6) |
|
||||
| 8 | 32-lane shuffle (butterfly) | ROW_SHFL_K | 29.4 | pending | yes (8) |
|
||||
| 9 | L1 scratch read (8 KB flop array) against the card's L2 hit | ROW_SCRATCH_K | 1,400 | pending | not drawn |
|
||||
| 10 | int8 8x8x8 tile, per MAC | ROW_TILE_K | 2.2 | pending | not drawn (the tensor lever is dead on other grounds) |
|
||||
|
||||
The order is set by the card's price, not the chip's: the chip pays 0.6 to 1.7 pJ for everything, and the card
|
||||
pays 6.2 for an add, 8.3 for a multiply, 21 for a high word and 29 for a shuffle. So the families the card pays
|
||||
MOST for (mulhi, shuffle) are the ones a chip undercuts most, and the forcing work is the plain ARX and mad the
|
||||
card does cheapest. This is the measured form of 15.1b's hold on the shuffle-heavy re-weight.
|
||||
|
||||
### 6.2 The mixed draw that maximises the expected k
|
||||
|
||||
The objective at a FIXED GPU premium: `k_eff(w) = sum w_i e_chip_i / sum w_i e_gpu_i` (the chip's energy for the
|
||||
drawn program over the card's for the same program; at a fixed `F` the chip pays `k_eff x F`). The band (layer
|
||||
1, the research lane's 17:00 reading of lane D's cut): B = 4 points on the injecting families only (add, sub,
|
||||
xor, mad, shfl, rotl, rotr), the lossy families (or, mul, mulhi) at or under their base, `or + mul + mulhi` at most
|
||||
18 + B, the shuffle capped at its class v4 weight; the index fold on every address and the F8 uniformity floor are
|
||||
not functions of the weights and do not move. `tools/chip-model/rtl/flow/mix.py` searches the band exhaustively
|
||||
over the corners (the objective is a ratio of linear forms, so the optimum is at a corner of each family's range).
|
||||
|
||||
On the unit floors (N3, the lock; the shuffle row pending, so computed over the other 67 points of 75):
|
||||
|
||||
| Mix | add, xor, mul, mad, shfl, rotl, sub, mulhi, rotr, or | k_eff (floors) | The card's pJ per op at the lock | k_eff on the core (floor + the core's 2.4 pJ per op overhead at N3) |
|
||||
|---|---|---|---|---|
|
||||
| class v4 (the base) | 12, 10, 8, 8, 8, 7, 6, 6, 6, 4 | 0.130 | 8.0 | 0.43 |
|
||||
| the census lane's re-weight (d20eb04bd, PASS on 256 seeds, no era and eras 0 to 7, W = 4 and W = 16) | 13, 11, 6, 10, 8, 8, 7, 2, 6, 4 | 0.153 (+18 percent) | 7.1 | 0.49 (+14 percent) |
|
||||
| the band's corner (every injecting family at +4 except the shuffle at its floor; lossy at base minus 4) | 16, 14, 4, 12, 4, 11, 10, 2, 10, 0 | MIX_BEST | | |
|
||||
|
||||
Reading: the re-weight moves six points off the chip's two easiest families (mulhi, mul) onto the hardest (mad,
|
||||
the ARX families) and lifts the expected k by about a seventh on the core; its acceptance is the census lane's
|
||||
(0 exhausted, F8 0.999 to 1.008 of uniform, the verifier within 4 percent); its energy cost on the card is a
|
||||
premium per instruction 11 percent LOWER (fewer mulhi), so at a fixed premium the program is 12 percent longer.
|
||||
The shuffle's own row decides whether it goes to its floor: on the card it is the dearest op (29.4 pJ at the lock)
|
||||
and on the chip a butterfly over the window, so its k is the lowest of the drawn families unless the synthesis says
|
||||
otherwise. The recommendation stands as the census lane's draw until the shuffle row lands (an amendment).
|
||||
|
||||
## 7. Sources
|
||||
|
||||
- The GPU side: `docs/analysis/counter-asic-4-research.md` 15.1a (the 5090 microbench, 8 October 2026, measured)
|
||||
|
|
|
|||
Loading…
Reference in a new issue