shadow-k: the placed gated 64-register window core (9.45 pJ per lane-op ASAP7, k 0.77 at the lock at N3; the board 2.0x node-for-node, 2.45x a node ahead; the window's placed residual +0.23 of k)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-josh 2026-10-08 21:33:11 +01:00
parent 14d4f3a32e
commit 86e5b0fb24

View file

@ -237,13 +237,16 @@ lane-op at ASAP7. The forms a chip maker would use:
| flops with the register-file clock gated (one of 64 registers written per cycle; the ICG cells allowed back in and inferred by Yosys, 320 gates) | 6.2 (225,441 cells; sequential 0.2) | 3.1 | 0.50 (0.70 node-for-node) | synthesised 16:0x UK; a model |
| the same gating on the 32-register base, for the penalty | 4.5 (156,833 cells; sequential 0.15) | 2.3 | 0.37 (0.51 node-for-node) | synthesised 16:0x UK; a model |
| the gated 32-register base PLACED AND ROUTED (SPEF, clock tree, 379,633 cells; steady state solved from 150 and 600 run cycles) | 6.7 (+49 percent over synthesis) | 3.4 | 0.54 (0.76 node-for-node) | placed 19:5x UK; the GDDR7 board at the lock 2.4x node-for-node, 2.8x a node ahead: the morning's headline figures to the digit |
| the gated 64-register window core PLACED AND ROUTED (563,339 cells; parasitics estimated from global routing, the SPEF lost to a full disk; steady state solved from 150 and 600 run cycles; plus or minus 15 percent) | 9.45 (+52 percent over synthesis; N5 6.6, N2 3.4) | 4.8 | 0.77 (1.07 node-for-node, 0.55 at N2) | placed 21:3x UK; the GDDR7 board at the lock 2.0x node-for-node, 2.45x a node ahead, 2.9x two ahead; the window's residual against the placed base +2.75 pJ per lane-op, +0.23 of k at the lock |
| latch-based register file (the clocked element halved; about 30 percent under the gated flop file, approximate) | about 0.7 x the gated row | | | modelled |
| SRAM-banked state shared across time-multiplexed lanes (one execution port serving many lanes' streams in turn, each lane's 256 bytes in a bank): modelled 8.5 to 10.5 from the access energies; BUILT as `core_tm` (8 lanes x 64 registers in banks, one port, round-robin, ungated): 8.8 pJ per lane-op synthesised (173,426 cells), against the SIMD ungated 9.7 | 8.8 (built) | 4.4 | 0.71 | synthesised 16:2x UK: port sharing saves 0.9 pJ of units and the bank-select muxes take most of it back; NOT the lever; the multi-family adversary lane's macro-window core reads 5.9 (its FakeRAM term modelled 2.0 to 7.0 pJ per access), within 5 percent of the gated flop row, so the file's form is not the lever either |
| values recomputed instead of kept | not available: 95 percent of the window is necessary (above) | | | measured on drawn programs |
The defence, then, is the penalty that remains after the adversary's best form: the gated 64-register file
against the gated 32-register file (the two synthesised rows above when they land; measured: 6.2 against 4.5 pJ per lane-op at ASAP7, a penalty of 1.7 pJ, 0.84 at N3, +0.13 of k at the lock,
0.37 to 0.50 at N3 and 0.51 to 0.70 node-for-node; the analytic estimate had been 1.5 pJ). Two corrections this
against the gated 32-register file (the two synthesised rows above when they land; measured: 6.2 against 4.5 pJ per lane-op at ASAP7 synthesised, a penalty of 1.7 pJ; PLACED 9.45 against 6.7, a
penalty of 2.75 pJ, 1.4 at N3, +0.23 of k at the lock, 0.54 to 0.77 at N3 and 0.76 to 1.07 node-for-node; the
analytic estimate had been 1.5 pJ). So on placed rows the window takes the GDDR7 board at the lock from 2.4x to
2.0x node-for-node and from 2.8x to 2.45x a node ahead, for at most 5 percent per load on the card. Two corrections this
forces: the honest adversary's BASE core is the gated one (k 0.37 at N3, 0.51 node-for-node), under the ungated
0.56 and 0.78 of section 4, which are the GPU-shaped core a maker would not build; and placement costs more than
the +20 to +40 percent estimated (the ungated placed base reads 11.3 against 6.9 pJ: wires plus a 2.5 pJ clock tree