shadow-k: the time-multiplexed core built (8.8 pJ, not the lever), the connected-state variant on the gated window core (6.3 pJ per lane-op, k 0.51 at the lock at N3), the lanes' agreement on the window's cost

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

Documents-only replay of 08447ca97 (86e5b0fb24) for the box mirror master
This commit is contained in:
igneum-labs 2026-10-08 15:23:25 +00:00
parent 13617b891c
commit 3aa976cd1d

View file

@ -237,7 +237,7 @@ lane-op at ASAP7. The forms a chip maker would use:
| flops with the register-file clock gated (one of 64 registers written per cycle; the ICG cells allowed back in and inferred by Yosys, 320 gates) | 6.2 (225,441 cells; sequential 0.2) | 3.1 | 0.50 (0.70 node-for-node) | synthesised 16:0x UK; a model |
| the same gating on the 32-register base, for the penalty | 4.5 (156,833 cells; sequential 0.15) | 2.3 | 0.37 (0.51 node-for-node) | synthesised 16:0x UK; a model |
| latch-based register file (the clocked element halved; about 30 percent under the gated flop file, approximate) | about 0.7 x the gated row | | | modelled |
| SRAM-banked state shared across time-multiplexed lanes (one execution port serving many lanes' streams in turn, each lane's 256 bytes in a bank of a few KB): three 32-bit operand reads and one write per lane-op from a small macro at about 1.0 to 1.5 pJ per 32-bit access at 7 nm (approximate: the JSSC 2026 3 nm macro reads at 3.9 pJ per access at 563 kbit; a few-KB bank is a third of that; claimed), plus the units and the read network (the base core's 4.5 pJ combinational term) | about 8.5 to 10.5 | 4.3 to 5.3 | 0.69 to 0.85 | modelled: NOT cheaper than the gated flop file; the GPU's own register file is SRAM-banked because it is 256 KB per SM, and at 256 bytes per lane flops with gating win |
| SRAM-banked state shared across time-multiplexed lanes (one execution port serving many lanes' streams in turn, each lane's 256 bytes in a bank): modelled 8.5 to 10.5 from the access energies; BUILT as `core_tm` (8 lanes x 64 registers in banks, one port, round-robin, ungated): 8.8 pJ per lane-op synthesised (173,426 cells), against the SIMD ungated 9.7 | 8.8 (built) | 4.4 | 0.71 | synthesised 16:2x UK: port sharing saves 0.9 pJ of units and the bank-select muxes take most of it back; NOT the lever; the multi-family adversary lane's macro-window core reads 5.9 (its FakeRAM term modelled 2.0 to 7.0 pJ per access), within 5 percent of the gated flop row, so the file's form is not the lever either |
| values recomputed instead of kept | not available: 95 percent of the window is necessary (above) | | | measured on drawn programs |
The defence, then, is the penalty that remains after the adversary's best form: the gated 64-register file
@ -274,6 +274,27 @@ the latency-bound chain. The sound class form is the full chain (the arithmetic-
class string `+reg64c`, pack hl-v6-win with `check_window_liveness` in its suite). The chip's side (this section's
gated rows) therefore carries the whole defence.
### 4d. The connected-state variant (cs64s27x16, the connected-state lane's structure) on the adversary's core
The connected-state lane's program (a 64-register window; per step a load whose address register is the previous
block's last dst, the word landing in m_j, then a 27-instruction block whose first instruction reads m_j and every
later one draws its src from the block's last four dsts, the last instruction injecting; 16 steps of text, 448
instructions, 16 passes per block; its liveness tool: 63 of 64 live at every address, about 11 registers in the
per-step dependent chain, about 20 touched per block) priced on the gated 64-register core with a 512-entry imem,
the program drawn by those rules in the testbench (`CS` mode of `tb_core_common.vh`), synthesis-only:
| Row | Cells | pJ per lane-op ASAP7 | N5 | N3 | N2 | k at the lock N5 / N3 / N2 | k at stock N5 / N3 |
|---|---|---|---|---|---|---|---|
| cs64s27x16 on the gated 64-register core, 512 imem | 253,059 | 6.3 | 4.4 | 3.2 | 2.3 | 0.71 / 0.51 / 0.37 | 0.39 / 0.28 |
| the class v4 draw on the gated 64-register core, 256 imem (4c) | 225,441 | 6.2 | 4.3 | 3.1 | 2.2 | 0.70 / 0.50 / 0.36 | 0.38 / 0.27 |
The chip's shadow for the program is 55,296 x 3.2 pJ = 0.18 microjoules per hash at N3 (0.24 node-for-node)
against 0.13 for the genesis window on the same core; the card pays +0.6 percent for the window on the 5090 (the
connected-state lane's measurement). The structure's other knobs do not reach the chip: the 16-pass loop and the
448 text cost the shared imem about 0.1 pJ per lane-op, hot-set banking is not needed (the gated file charges only
the written register), and the chain's width sets lane count, which is free. On the GDDR7 board at the lock the
window moves the chip's edge by about 1.1x (3.6x to 3.3x node-for-node), under the 1.25x gate that lane set.
## 5. The chip edge at the measured k
`E_chip = E_mem + N_ops x e_chip` (absolute: the chip's shadow cost is 102,100 x 3.5 pJ = 0.36 microjoules per hash