shadow-k: row (2'), the 32-lane 16-register core (4.2 pJ per lane-op ASAP7, k 0.34 at the lock at N3)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-josh 2026-10-08 14:55:35 +01:00
parent 05275ca6a9
commit 3e2a202feb

View file

@ -140,7 +140,7 @@ draws 13.9 mW, the run phase 36.9 mW for 8 lanes at 1.5 ns).
| of which the units, the read muxes and the butterfly | | | 4.5 | 3.1 | 2.3 | 1.6 | | | | | | | of which the units, the read muxes and the butterfly | | | 4.5 | 3.1 | 2.3 | 1.6 | | | | | |
| core, 8 lanes, 32 registers | placed and routed, SPEF | ROW_CORE8_PLACED | | core, 8 lanes, 32 registers | placed and routed, SPEF | ROW_CORE8_PLACED |
| core, 32 lanes, 32 registers | ROW_CORE32 | | core, 32 lanes, 32 registers | ROW_CORE32 |
| core, 32 lanes, 16 registers | ROW_CORE32R16 | | core, 32 lanes, 16 registers | synthesis only | 443,258 | 4.2 | 2.9 | 2.1 | 1.5 | 11.3 / 6.2 / 6.9 | 0.18 / 0.34 / 0.30 | 0.13 / 0.24 / 0.22 | 0.68 | synthesised; one run length, about plus or minus 10 percent; 15:0x UK |
| the bare ARX lane (section 3, the floor) | routed | 11,631 | 2.2 | 1.5 | 1.1 | 0.8 | 11.3 / 6.2 / 6.9 | 0.10 / 0.18 / 0.16 | 0.07 / 0.13 / 0.11 | 0.35 | the lower bound | | the bare ARX lane (section 3, the floor) | routed | 11,631 | 2.2 | 1.5 | 1.1 | 0.8 | 11.3 / 6.2 / 6.9 | 0.10 / 0.18 / 0.16 | 0.07 / 0.13 / 0.11 | 0.35 | the lower bound |
Reading: fetch, decode, a 32-register file and the full unit set cost a chip 3.1x the bare lane, and the core's Reading: fetch, decode, a 32-register file and the full unit set cost a chip 3.1x the bare lane, and the core's
@ -169,7 +169,7 @@ half a day; the select tree is a new instruction kind), knob 2 has no GPU side (
| (3) the same with the imem as a 4 KB SRAM macro shared by the lanes (2 to 4 pJ per 32-bit read, approximate) | | about 7.2 | about 1.9 | 5.0 | 3.6 | 2.6 | about 0.32 / 0.58 / 0.52 | the same | | (3) the same with the imem as a 4 KB SRAM macro shared by the lanes (2 to 4 pJ per 32-bit read, approximate) | | about 7.2 | about 1.9 | 5.0 | 3.6 | 2.6 | about 0.32 / 0.58 / 0.52 | the same |
| (4) the drawn select tree (the era's 16-entry op permutation ahead of decode; every unit evaluated every cycle, as in the base) | 186,870 | 6.85 | 2.4 | 4.8 | 3.4 | 2.5 | 0.30 / 0.55 / 0.50 | the units' microbench sum (approximate) | | (4) the drawn select tree (the era's 16-entry op permutation ahead of decode; every unit evaluated every cycle, as in the base) | 186,870 | 6.85 | 2.4 | 4.8 | 3.4 | 2.5 | 0.30 / 0.55 / 0.50 | the units' microbench sum (approximate) |
| (2) 32 lanes, 32 registers (the butterfly across 32; the imem amortised over 32) | ROW_CORE32 | | (2) 32 lanes, 32 registers (the butterfly across 32; the imem amortised over 32) | ROW_CORE32 |
| (2') 32 lanes, 16 registers (the register-file sensitivity the other way) | ROW_CORE32R16 | | (2') 32 lanes, 16 registers (the register-file sensitivity the other way; one run length of 150 cycles, the load phase subtracted at the 8-lane ratio, about plus or minus 10 percent) | 443,258 | 4.2 | 2.3 | 2.9 | 2.1 | 1.5 | 0.18 / 0.34 / 0.30 | no GPU knob |
| (5) all four together (32 lanes, 64 registers, 1,024 imem, the select tree) | ROW_CORE32ALL | | (5) all four together (32 lanes, 64 registers, 1,024 imem, the select tree) | ROW_CORE32ALL |
Reading, for the founder's "under 2x at the lock" (which needs k near 0.9 on the GDDR7 board): the 64-register Reading, for the founder's "under 2x at the lock" (which needs k near 0.9 on the GDDR7 board): the 64-register