shadow-k: the 32-lane core (5.55 pJ per lane-op, k 0.45 at the lock), the shuffle row (1.24 pJ, k 0.021), the mix optimiser's best corner

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

Documents-only replay of bfccc26ed (bfccc26ed4) for the box mirror master
This commit is contained in:
igneum-labs 2026-10-08 14:28:57 +00:00
parent 7e4ba45831
commit ba6763e74f

View file

@ -102,7 +102,7 @@ stock and at the 1,300 MHz lock, and 6.9 pJ per counted op on the Apple M5 Max (
| index fold (x M, rotl R, masks; era constants in registers) | 14,965 | 2.37 | 1.66 | 1.19 | 0.86 | 13.9 / 8.3 (read against int_mul) | 0.086 / 0.14 | | 0.29 | |
| prmt (byte permute) | 6,744 | 1.27 | 0.89 | 0.64 | 0.46 | 22.3 / 11.5 | 0.029 / 0.056 | | 0.11 | |
| lop3 (8-bit truth table) | 7,274 | 1.42 | 0.99 | 0.72 | 0.52 | 24.1 / 13.0 | 0.030 / 0.055 | | 0.11 | |
| 32-lane xor-mask shuffle (butterfly over the 1 KB window), per lane-op | ROW_SHFL |
| 32-lane xor-mask shuffle (butterfly over the 1 KB window), per lane-op | 181,580 | 1.24 | 0.87 | 0.63 | 0.45 | 55.8 / 29.4 | 0.011 / 0.021 | | 0.042 | routed with SPEF; 15:1x UK |
| 32-lane general crossbar, per lane-op | ROW_XBAR |
| 8 KB scratch, one random read (flop array: the pessimistic form) | ROW_SCRATCH |
| int8 8x8x8 tile, per MAC | ROW_TILE |
@ -139,7 +139,7 @@ draws 13.9 mW, the run phase 36.9 mW for 8 lanes at 1.5 ns).
| of which the sequential term (register file, imem and IR clock pins, no clock gating) | | | 2.4 | 1.7 | 1.2 | 0.9 | | | | | |
| of which the units, the read muxes and the butterfly | | | 4.5 | 3.1 | 2.3 | 1.6 | | | | | |
| core, 8 lanes, 32 registers | placed and routed, SPEF | ROW_CORE8_PLACED |
| core, 32 lanes, 32 registers | ROW_CORE32 |
| core, 32 lanes, 32 registers | synthesis only (steady state from 150 and 400 run cycles) | 600,381 | 5.55 | 3.9 | 2.8 | 2.0 | 11.3 / 6.2 / 6.9 | 0.25 / 0.45 / 0.41 | 0.18 / 0.32 / 0.29 | 0.90 | synthesised; 15:2x UK |
| core, 32 lanes, 16 registers | synthesis only | 443,258 | 4.2 | 2.9 | 2.1 | 1.5 | 11.3 / 6.2 / 6.9 | 0.18 / 0.34 / 0.30 | 0.13 / 0.24 / 0.22 | 0.68 | synthesised; one run length, about plus or minus 10 percent; 15:0x UK |
| the bare ARX lane (section 3, the floor) | routed | 11,631 | 2.2 | 1.5 | 1.1 | 0.8 | 11.3 / 6.2 / 6.9 | 0.10 / 0.18 / 0.16 | 0.07 / 0.13 / 0.11 | 0.35 | the lower bound |
@ -168,7 +168,7 @@ half a day; the select tree is a new instruction kind), knob 2 has no GPU side (
| (3) 1,024-entry imem, the program drawn at 1,024, as built (a flop array) | 324,543 | 12.0 | 6.0 | 8.4 | 6.1 | 4.4 | 0.53 / 0.97 / 0.87 | measured: the 1,024 block at 27 passes costs the 5090 1.58x the energy per hash (4.96 against 3.14 microjoules at stock, 111 against 141 MH/s at the 575 W cap) and the 4090 1.57x (7.01 against 4.48, the rate held at 62.5 MH/s), for 4x the shadow instructions: 0.40x per instruction |
| (3) the same with the imem as a 4 KB SRAM macro shared by the lanes (2 to 4 pJ per 32-bit read, approximate) | | about 7.2 | about 1.9 | 5.0 | 3.6 | 2.6 | about 0.32 / 0.58 / 0.52 | the same |
| (4) the drawn select tree (the era's 16-entry op permutation ahead of decode; every unit evaluated every cycle, as in the base) | 186,870 | 6.85 | 2.4 | 4.8 | 3.4 | 2.5 | 0.30 / 0.55 / 0.50 | the units' microbench sum (approximate) |
| (2) 32 lanes, 32 registers (the butterfly across 32; the imem amortised over 32) | ROW_CORE32 |
| (2) 32 lanes, 32 registers (the butterfly across 32; the imem amortised over 32) | 600,381 | 5.55 | 1.5 | 3.9 | 2.8 | 2.0 | 0.25 / 0.45 / 0.41 | no GPU knob |
| (2') 32 lanes, 16 registers (the register-file sensitivity the other way; one run length of 150 cycles, the load phase subtracted at the 8-lane ratio, about plus or minus 10 percent) | 443,258 | 4.2 | 2.3 | 2.9 | 2.1 | 1.5 | 0.18 / 0.34 / 0.30 | no GPU knob |
| (5) all four together (32 lanes, 64 registers, 1,024 imem, the select tree) | ROW_CORE32ALL |
@ -269,7 +269,7 @@ to 2.3x and the strongest chips at 2.6x to 4.2x.
| 5 | mul | 0.68 | 8.3 | 0.082 | yes (8) |
| 6 | prmt, lop3 | 0.64, 0.72 | 11.5, 13.0 | 0.056, 0.055 | not drawn (RTL rows only) |
| 7 | mulhi | 0.68 | 21.0 | 0.032 | yes (6) |
| 8 | 32-lane shuffle (butterfly) | ROW_SHFL_K | 29.4 | pending | yes (8) |
| 8 | 32-lane shuffle (butterfly) | 0.63 | 29.4 | 0.021 | yes (8) |
| 9 | L1 scratch read (8 KB flop array) against the card's L2 hit | ROW_SCRATCH_K | 1,400 | pending | not drawn |
| 10 | int8 8x8x8 tile, per MAC | ROW_TILE_K | 2.2 | pending | not drawn (the tensor lever is dead on other grounds) |
@ -288,13 +288,14 @@ xor, mad, shfl, rotl, rotr), the lossy families (or, mul, mulhi) at or under the
not functions of the weights and do not move. `tools/chip-model/rtl/flow/mix.py` searches the band exhaustively
over the corners (the objective is a ratio of linear forms, so the optimum is at a corner of each family's range).
On the unit floors (N3, the lock; the shuffle row pending, so computed over the other 67 points of 75):
On the unit floors (N3, the lock; the shuffle row in at 0.63 pJ, its k 0.021 the lowest of the drawn families):
| Mix | add, xor, mul, mad, shfl, rotl, sub, mulhi, rotr, or | k_eff (floors) | The card's pJ per op at the lock | k_eff on the core (floor + the core's 2.4 pJ per op overhead at N3) |
|---|---|---|---|---|
| class v4 (the base) | 12, 10, 8, 8, 8, 7, 6, 6, 6, 4 | 0.130 | 8.0 | 0.43 |
| the census lane's re-weight (d20eb04bd, PASS on 256 seeds, no era and eras 0 to 7, W = 4 and W = 16) | 13, 11, 6, 10, 8, 8, 7, 2, 6, 4 | 0.153 (+18 percent) | 7.1 | 0.49 (+14 percent) |
| the band's corner (every injecting family at +4 except the shuffle at its floor; lossy at base minus 4) | 16, 14, 4, 12, 4, 11, 10, 2, 10, 0 | MIX_BEST | | |
| class v4 (the base) | 12, 10, 8, 8, 8, 7, 6, 6, 6, 4 | 0.097 (with the shuffle row in: 0.130 over the other 67 points) | 10.3 | 0.43 |
| the census lane's re-weight (d20eb04bd, PASS on 256 seeds, no era and eras 0 to 7, W = 4 and W = 16) | 13, 11, 6, 10, 8, 8, 7, 2, 6, 4 | 0.115 (+19 percent; 0.153 over the 67 points) | 9.0 | 0.49 (+14 percent) |
| the band's best corner (mix.py, exhaustive: every injecting family at +4, the shuffle at its floor, the lossy families at base minus 4) | 16, 14, 4, 12, 4, 11, 10, 2, 10, 0 | 0.137 (+42 percent) | 8.2 | about 0.51 (+19 percent) |
| the band's worst corner (shuffle and mulhi heavy) | 8, 6, 4, 4, 8, 3, 2, 6, 2, 0 | 0.074 | 12.3 | about 0.40 |
Reading: the re-weight moves six points off the chip's two easiest families (mulhi, mul) onto the hardest (mad,
the ARX families) and lifts the expected k by about a seventh on the core; its acceptance is the census lane's
@ -302,7 +303,7 @@ the ARX families) and lifts the expected k by about a seventh on the core; its a
premium per instruction 11 percent LOWER (fewer mulhi), so at a fixed premium the program is 12 percent longer.
The shuffle's own row decides whether it goes to its floor: on the card it is the dearest op (29.4 pJ at the lock)
and on the chip a butterfly over the window, so its k is the lowest of the drawn families unless the synthesis says
otherwise. The recommendation stands as the census lane's draw until the shuffle row lands (an amendment).
otherwise. The recommendation: the band's best corner (add 16, xor 14, mad 12, rotl 11, sub 10, rotr 10, shfl 4, mul 4, mulhi 2, or 0) if its census passes (the census lane's harness, about five box-minutes per candidate); the census lane's draw as the passed fallback. Either way the shuffle goes to its floor: the card pays 29.4 pJ for a move the chip does for 0.6.
## 7. Sources