From 08447ca97e4b76809b3ec87ef1d750a46b18a191 Mon Sep 17 00:00:00 2001 From: igneum-josh Date: Thu, 8 Oct 2026 16:23:25 +0100 Subject: [PATCH] shadow-k: the time-multiplexed core built (8.8 pJ, not the lever), the connected-state variant on the gated window core (6.3 pJ per lane-op, k 0.51 at the lock at N3), the lanes' agreement on the window's cost Co-Authored-By: Claude Fable 5.1 --- docs/analysis/class-v6/floor/shadow-k.md | 23 ++++++++++++++++++++++- 1 file changed, 22 insertions(+), 1 deletion(-) diff --git a/docs/analysis/class-v6/floor/shadow-k.md b/docs/analysis/class-v6/floor/shadow-k.md index 11465d38d..34f822a4d 100644 --- a/docs/analysis/class-v6/floor/shadow-k.md +++ b/docs/analysis/class-v6/floor/shadow-k.md @@ -237,7 +237,7 @@ lane-op at ASAP7. The forms a chip maker would use: | flops with the register-file clock gated (one of 64 registers written per cycle; the ICG cells allowed back in and inferred by Yosys, 320 gates) | 6.2 (225,441 cells; sequential 0.2) | 3.1 | 0.50 (0.70 node-for-node) | synthesised 16:0x UK; a model | | the same gating on the 32-register base, for the penalty | 4.5 (156,833 cells; sequential 0.15) | 2.3 | 0.37 (0.51 node-for-node) | synthesised 16:0x UK; a model | | latch-based register file (the clocked element halved; about 30 percent under the gated flop file, approximate) | about 0.7 x the gated row | | | modelled | -| SRAM-banked state shared across time-multiplexed lanes (one execution port serving many lanes' streams in turn, each lane's 256 bytes in a bank of a few KB): three 32-bit operand reads and one write per lane-op from a small macro at about 1.0 to 1.5 pJ per 32-bit access at 7 nm (approximate: the JSSC 2026 3 nm macro reads at 3.9 pJ per access at 563 kbit; a few-KB bank is a third of that; claimed), plus the units and the read network (the base core's 4.5 pJ combinational term) | about 8.5 to 10.5 | 4.3 to 5.3 | 0.69 to 0.85 | modelled: NOT cheaper than the gated flop file; the GPU's own register file is SRAM-banked because it is 256 KB per SM, and at 256 bytes per lane flops with gating win | +| SRAM-banked state shared across time-multiplexed lanes (one execution port serving many lanes' streams in turn, each lane's 256 bytes in a bank): modelled 8.5 to 10.5 from the access energies; BUILT as `core_tm` (8 lanes x 64 registers in banks, one port, round-robin, ungated): 8.8 pJ per lane-op synthesised (173,426 cells), against the SIMD ungated 9.7 | 8.8 (built) | 4.4 | 0.71 | synthesised 16:2x UK: port sharing saves 0.9 pJ of units and the bank-select muxes take most of it back; NOT the lever; the multi-family adversary lane's macro-window core reads 5.9 (its FakeRAM term modelled 2.0 to 7.0 pJ per access), within 5 percent of the gated flop row, so the file's form is not the lever either | | values recomputed instead of kept | not available: 95 percent of the window is necessary (above) | | | measured on drawn programs | The defence, then, is the penalty that remains after the adversary's best form: the gated 64-register file @@ -274,6 +274,27 @@ the latency-bound chain. The sound class form is the full chain (the arithmetic- class string `+reg64c`, pack hl-v6-win with `check_window_liveness` in its suite). The chip's side (this section's gated rows) therefore carries the whole defence. +### 4d. The connected-state variant (cs64s27x16, the connected-state lane's structure) on the adversary's core + +The connected-state lane's program (a 64-register window; per step a load whose address register is the previous +block's last dst, the word landing in m_j, then a 27-instruction block whose first instruction reads m_j and every +later one draws its src from the block's last four dsts, the last instruction injecting; 16 steps of text, 448 +instructions, 16 passes per block; its liveness tool: 63 of 64 live at every address, about 11 registers in the +per-step dependent chain, about 20 touched per block) priced on the gated 64-register core with a 512-entry imem, +the program drawn by those rules in the testbench (`CS` mode of `tb_core_common.vh`), synthesis-only: + +| Row | Cells | pJ per lane-op ASAP7 | N5 | N3 | N2 | k at the lock N5 / N3 / N2 | k at stock N5 / N3 | +|---|---|---|---|---|---|---|---| +| cs64s27x16 on the gated 64-register core, 512 imem | 253,059 | 6.3 | 4.4 | 3.2 | 2.3 | 0.71 / 0.51 / 0.37 | 0.39 / 0.28 | +| the class v4 draw on the gated 64-register core, 256 imem (4c) | 225,441 | 6.2 | 4.3 | 3.1 | 2.2 | 0.70 / 0.50 / 0.36 | 0.38 / 0.27 | + +The chip's shadow for the program is 55,296 x 3.2 pJ = 0.18 microjoules per hash at N3 (0.24 node-for-node) +against 0.13 for the genesis window on the same core; the card pays +0.6 percent for the window on the 5090 (the +connected-state lane's measurement). The structure's other knobs do not reach the chip: the 16-pass loop and the +448 text cost the shared imem about 0.1 pJ per lane-op, hot-set banking is not needed (the gated file charges only +the written register), and the chain's width sets lane count, which is free. On the GDDR7 board at the lock the +window moves the chip's edge by about 1.1x (3.6x to 3.3x node-for-node), under the 1.25x gate that lane set. + ## 5. The chip edge at the measured k `E_chip = E_mem + N_ops x e_chip` (absolute: the chip's shadow cost is 102,100 x 3.5 pJ = 0.36 microjoules per hash