shadow-k: the adversary's gated rows (4.5 and 6.2 pJ per lane-op ASAP7; the window's residual +0.13 of k), the placed ungated base (11.3 pJ), the ICG simulation model

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-josh 2026-10-08 16:04:44 +01:00
parent 26844fde7f
commit 9737ffb260
4 changed files with 26 additions and 7 deletions

View file

@ -138,7 +138,7 @@ draws 13.9 mW, the run phase 36.9 mW for 8 lanes at 1.5 ns).
| core, 8 lanes, 32 registers | synthesis only (no wires, no clock tree) | 186,443 | 6.9 | 4.8 | 3.5 | 2.5 | 11.3 / 6.2 / 6.9 | 0.31 / 0.56 / 0.50 | 0.22 / 0.40 / 0.36 | 1.1 | synthesised; 14:0x UK |
| of which the sequential term (register file, imem and IR clock pins, no clock gating) | | | 2.4 | 1.7 | 1.2 | 0.9 | | | | | |
| of which the units, the read muxes and the butterfly | | | 4.5 | 3.1 | 2.3 | 1.6 | | | | | |
| core, 8 lanes, 32 registers | placed and routed, SPEF | ROW_CORE8_PLACED |
| core, 8 lanes, 32 registers, ungated | placed and routed, SPEF, clock tree (one run length of 300 cycles, the load phase subtracted, about plus or minus 10 percent) | 444,478 | 11.3 | 7.9 | 5.6 | 4.1 | 11.3 / 6.2 / 6.9 | 0.50 / 0.91 / 0.82 | 0.36 / 0.66 / 0.59 | 1.8 | placed 16:0x UK on a rented pod; +64 percent over synthesis (wires, and a clock tree of 2.5 pJ per lane-op that gating removes) |
| core, 32 lanes, 32 registers | synthesis only (steady state from 150 and 400 run cycles) | 600,381 | 5.55 | 3.9 | 2.8 | 2.0 | 11.3 / 6.2 / 6.9 | 0.25 / 0.45 / 0.41 | 0.18 / 0.32 / 0.29 | 0.90 | synthesised; 15:2x UK |
| core, 32 lanes, 16 registers | synthesis only | 443,258 | 4.2 | 2.9 | 2.1 | 1.5 | 11.3 / 6.2 / 6.9 | 0.18 / 0.34 / 0.30 | 0.13 / 0.24 / 0.22 | 0.68 | synthesised; one run length, about plus or minus 10 percent; 15:0x UK |
| the bare ARX lane (section 3, the floor) | routed | 11,631 | 2.2 | 1.5 | 1.1 | 0.8 | 11.3 / 6.2 / 6.9 | 0.10 / 0.18 / 0.16 | 0.07 / 0.13 / 0.11 | 0.35 | the lower bound |
@ -234,16 +234,19 @@ lane-op at ASAP7. The forms a chip maker would use:
| Form of the 64-register state (per lane, 256 bytes) | pJ per lane-op ASAP7 | N3 | k at the lock (N3) | Label |
|---|---|---|---|---|
| flops, no clock gating, 64:1 read muxes (the 4a row) | 9.7 | 4.9 | 0.78 | synthesised; a model |
| flops with the register-file clock gated (one of 64 registers written per cycle; the ICG cells allowed back in and inferred by Yosys) | ROW_CORE8R64G | | | synthesised; a model |
| the same gating on the 32-register base, for the penalty | ROW_CORE8G | | | synthesised; a model |
| flops with the register-file clock gated (one of 64 registers written per cycle; the ICG cells allowed back in and inferred by Yosys, 320 gates) | 6.2 (225,441 cells; sequential 0.2) | 3.1 | 0.50 (0.70 node-for-node) | synthesised 16:0x UK; a model |
| the same gating on the 32-register base, for the penalty | 4.5 (156,833 cells; sequential 0.15) | 2.3 | 0.37 (0.51 node-for-node) | synthesised 16:0x UK; a model |
| latch-based register file (the clocked element halved; about 30 percent under the gated flop file, approximate) | about 0.7 x the gated row | | | modelled |
| SRAM-banked state shared across time-multiplexed lanes (one execution port serving many lanes' streams in turn, each lane's 256 bytes in a bank of a few KB): three 32-bit operand reads and one write per lane-op from a small macro at about 1.0 to 1.5 pJ per 32-bit access at 7 nm (approximate: the JSSC 2026 3 nm macro reads at 3.9 pJ per access at 563 kbit; a few-KB bank is a third of that; claimed), plus the units and the read network (the base core's 4.5 pJ combinational term) | about 8.5 to 10.5 | 4.3 to 5.3 | 0.69 to 0.85 | modelled: NOT cheaper than the gated flop file; the GPU's own register file is SRAM-banked because it is 256 KB per SM, and at 256 bytes per lane flops with gating win |
| values recomputed instead of kept | not available: 95 percent of the window is necessary (above) | | | measured on drawn programs |
The defence, then, is the penalty that remains after the adversary's best form: the gated 64-register file
against the gated 32-register file (the two synthesised rows above when they land; the analytic estimate from
the ungated rows is 9.7 - 2.5 = 7.2 against 6.9 - 1.2 = 5.7 pJ per lane-op at ASAP7, a penalty of about 1.5 pJ,
0.75 at N3, +0.12 of k at the lock, approximate until the gated rows replace it). On the 32-lane core the same
against the gated 32-register file (the two synthesised rows above when they land; measured: 6.2 against 4.5 pJ per lane-op at ASAP7, a penalty of 1.7 pJ, 0.84 at N3, +0.13 of k at the lock,
0.37 to 0.50 at N3 and 0.51 to 0.70 node-for-node; the analytic estimate had been 1.5 pJ). Two corrections this
forces: the honest adversary's BASE core is the gated one (k 0.37 at N3, 0.51 node-for-node), under the ungated
0.56 and 0.78 of section 4, which are the GPU-shaped core a maker would not build; and placement costs more than
the +20 to +40 percent estimated (the ungated placed base reads 11.3 against 6.9 pJ: wires plus a 2.5 pJ clock tree
that gating removes), so the placed gated rows (on a rented pod, 17:30 UK) are the figures to serve. On the 32-lane core the same
penalty applies per lane (the register file does not amortise), so the window moves the 32-lane core from k 0.45
to about 0.57 at the lock at N3 (0.63 to about 0.80 node-for-node).

View file

@ -0,0 +1,12 @@
// Behavioural model of the ASAP7 integrated clock gate (latch_posedge_precontrol: the enable is latched while the
// clock is low, the gated clock is CLK AND the latched enable OR test). The liberty carries no function for it.
module ICGx1_ASAP7_75t_R(input CLK, input ENA, input SE, output GCLK);
reg en = 0;
always @(CLK or ENA or SE) if (!CLK) en = ENA | SE;
assign GCLK = CLK & en;
endmodule
module ICGx2_ASAP7_75t_R(input CLK, input ENA, input SE, output GCLK);
reg en = 0;
always @(CLK or ENA or SE) if (!CLK) en = ENA | SE;
assign GCLK = CLK & en;
endmodule

View file

@ -4,6 +4,7 @@ set -euo pipefail
name=$1; tag=$2; shift 2
out=/work/sim/$name
simcells=$(yosys-config --datdir)/simcells.v
iverilog -g2005 -I /work/tb -o $out/sim_$tag $out/sim_net.v /work/tb/tb_$(sed -n "s/^$name \([^ ]*\) .*/\1/p" /work/flow/designs.txt).v $simcells
tbfile=${TB:-/work/tb/tb_$(sed -n "s/^$name \([^ ]*\) .*/\1/p" /work/flow/designs.txt).v}
iverilog -g2005 -I /work/tb -o $out/sim_$tag $out/sim_net.v /work/flow/asap7_icg_model.v $tbfile $simcells
( cd $out && vvp -n sim_$tag "$@" | tee sim_$tag.log && mv dump.vcd $tag.vcd )
ls -la $out/$tag.vcd

View file

@ -0,0 +1,3 @@
`define TOP core_v6_8
`define HALF 750
`include "tb_core_legacy.vh"