shadow-k: the time-multiplexed core built (8.8 pJ, not the lever), the connected-state variant on the gated window core (6.3 pJ per lane-op, k 0.51 at the lock at N3), the lanes' agreement on the window's cost
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Documents-only replay of08447ca97(86e5b0fb24) for the box mirror master
This commit is contained in:
parent
13617b891c
commit
3aa976cd1d
1 changed files with 22 additions and 1 deletions
|
|
@ -237,7 +237,7 @@ lane-op at ASAP7. The forms a chip maker would use:
|
|||
| flops with the register-file clock gated (one of 64 registers written per cycle; the ICG cells allowed back in and inferred by Yosys, 320 gates) | 6.2 (225,441 cells; sequential 0.2) | 3.1 | 0.50 (0.70 node-for-node) | synthesised 16:0x UK; a model |
|
||||
| the same gating on the 32-register base, for the penalty | 4.5 (156,833 cells; sequential 0.15) | 2.3 | 0.37 (0.51 node-for-node) | synthesised 16:0x UK; a model |
|
||||
| latch-based register file (the clocked element halved; about 30 percent under the gated flop file, approximate) | about 0.7 x the gated row | | | modelled |
|
||||
| SRAM-banked state shared across time-multiplexed lanes (one execution port serving many lanes' streams in turn, each lane's 256 bytes in a bank of a few KB): three 32-bit operand reads and one write per lane-op from a small macro at about 1.0 to 1.5 pJ per 32-bit access at 7 nm (approximate: the JSSC 2026 3 nm macro reads at 3.9 pJ per access at 563 kbit; a few-KB bank is a third of that; claimed), plus the units and the read network (the base core's 4.5 pJ combinational term) | about 8.5 to 10.5 | 4.3 to 5.3 | 0.69 to 0.85 | modelled: NOT cheaper than the gated flop file; the GPU's own register file is SRAM-banked because it is 256 KB per SM, and at 256 bytes per lane flops with gating win |
|
||||
| SRAM-banked state shared across time-multiplexed lanes (one execution port serving many lanes' streams in turn, each lane's 256 bytes in a bank): modelled 8.5 to 10.5 from the access energies; BUILT as `core_tm` (8 lanes x 64 registers in banks, one port, round-robin, ungated): 8.8 pJ per lane-op synthesised (173,426 cells), against the SIMD ungated 9.7 | 8.8 (built) | 4.4 | 0.71 | synthesised 16:2x UK: port sharing saves 0.9 pJ of units and the bank-select muxes take most of it back; NOT the lever; the multi-family adversary lane's macro-window core reads 5.9 (its FakeRAM term modelled 2.0 to 7.0 pJ per access), within 5 percent of the gated flop row, so the file's form is not the lever either |
|
||||
| values recomputed instead of kept | not available: 95 percent of the window is necessary (above) | | | measured on drawn programs |
|
||||
|
||||
The defence, then, is the penalty that remains after the adversary's best form: the gated 64-register file
|
||||
|
|
@ -274,6 +274,27 @@ the latency-bound chain. The sound class form is the full chain (the arithmetic-
|
|||
class string `+reg64c`, pack hl-v6-win with `check_window_liveness` in its suite). The chip's side (this section's
|
||||
gated rows) therefore carries the whole defence.
|
||||
|
||||
### 4d. The connected-state variant (cs64s27x16, the connected-state lane's structure) on the adversary's core
|
||||
|
||||
The connected-state lane's program (a 64-register window; per step a load whose address register is the previous
|
||||
block's last dst, the word landing in m_j, then a 27-instruction block whose first instruction reads m_j and every
|
||||
later one draws its src from the block's last four dsts, the last instruction injecting; 16 steps of text, 448
|
||||
instructions, 16 passes per block; its liveness tool: 63 of 64 live at every address, about 11 registers in the
|
||||
per-step dependent chain, about 20 touched per block) priced on the gated 64-register core with a 512-entry imem,
|
||||
the program drawn by those rules in the testbench (`CS` mode of `tb_core_common.vh`), synthesis-only:
|
||||
|
||||
| Row | Cells | pJ per lane-op ASAP7 | N5 | N3 | N2 | k at the lock N5 / N3 / N2 | k at stock N5 / N3 |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| cs64s27x16 on the gated 64-register core, 512 imem | 253,059 | 6.3 | 4.4 | 3.2 | 2.3 | 0.71 / 0.51 / 0.37 | 0.39 / 0.28 |
|
||||
| the class v4 draw on the gated 64-register core, 256 imem (4c) | 225,441 | 6.2 | 4.3 | 3.1 | 2.2 | 0.70 / 0.50 / 0.36 | 0.38 / 0.27 |
|
||||
|
||||
The chip's shadow for the program is 55,296 x 3.2 pJ = 0.18 microjoules per hash at N3 (0.24 node-for-node)
|
||||
against 0.13 for the genesis window on the same core; the card pays +0.6 percent for the window on the 5090 (the
|
||||
connected-state lane's measurement). The structure's other knobs do not reach the chip: the 16-pass loop and the
|
||||
448 text cost the shared imem about 0.1 pJ per lane-op, hot-set banking is not needed (the gated file charges only
|
||||
the written register), and the chain's width sets lane count, which is free. On the GDDR7 board at the lock the
|
||||
window moves the chip's edge by about 1.1x (3.6x to 3.3x node-for-node), under the 1.25x gate that lane set.
|
||||
|
||||
## 5. The chip edge at the measured k
|
||||
|
||||
`E_chip = E_mem + N_ops x e_chip` (absolute: the chip's shadow cost is 102,100 x 3.5 pJ = 0.36 microjoules per hash
|
||||
|
|
|
|||
Loading…
Reference in a new issue