diff --git a/docs/analysis/class-v6/floor/shadow-k.md b/docs/analysis/class-v6/floor/shadow-k.md index 6ea36c5ba..e324d7436 100644 --- a/docs/analysis/class-v6/floor/shadow-k.md +++ b/docs/analysis/class-v6/floor/shadow-k.md @@ -104,7 +104,7 @@ stock and at the 1,300 MHz lock, and 6.9 pJ per counted op on the Apple M5 Max ( | lop3 (8-bit truth table) | 7,274 | 1.42 | 0.99 | 0.72 | 0.52 | 24.1 / 13.0 | 0.030 / 0.055 | | 0.11 | | | 32-lane xor-mask shuffle (butterfly over the 1 KB window), per lane-op | 181,580 | 1.24 | 0.87 | 0.63 | 0.45 | 55.8 / 29.4 | 0.011 / 0.021 | | 0.042 | routed with SPEF; 15:1x UK | | 32-lane general crossbar, per lane-op | ROW_XBAR | -| 8 KB scratch, one random read (flop array: the pessimistic form) | ROW_SCRATCH | +| 8 KB scratch, one random 32-bit read of a 2,048 x 32 flop array (the pessimistic form of a chip's L1; an SRAM macro reads lower) | 868,159 | 207 | 145 | 104 | 75 | 2,400 / 1,400 per L2 hit | 0.043 / 0.074 | | 0.15 | routed with SPEF, 19:4x UK; the card's shared-memory read is unmeasured (owed) | | int8 8x8x8 tile, per MAC | ROW_TILE | Reading the floors: a lane's add costs the chip about 2.2 pJ at ASAP7 and 1.1 at N3, against the 5090's 6.2 @@ -236,6 +236,7 @@ lane-op at ASAP7. The forms a chip maker would use: | flops, no clock gating, 64:1 read muxes (the 4a row) | 9.7 | 4.9 | 0.78 | synthesised; a model | | flops with the register-file clock gated (one of 64 registers written per cycle; the ICG cells allowed back in and inferred by Yosys, 320 gates) | 6.2 (225,441 cells; sequential 0.2) | 3.1 | 0.50 (0.70 node-for-node) | synthesised 16:0x UK; a model | | the same gating on the 32-register base, for the penalty | 4.5 (156,833 cells; sequential 0.15) | 2.3 | 0.37 (0.51 node-for-node) | synthesised 16:0x UK; a model | +| the gated 32-register base PLACED AND ROUTED (SPEF, clock tree, 379,633 cells; steady state solved from 150 and 600 run cycles) | 6.7 (+49 percent over synthesis) | 3.4 | 0.54 (0.76 node-for-node) | placed 19:5x UK; the GDDR7 board at the lock 2.4x node-for-node, 2.8x a node ahead: the morning's headline figures to the digit | | latch-based register file (the clocked element halved; about 30 percent under the gated flop file, approximate) | about 0.7 x the gated row | | | modelled | | SRAM-banked state shared across time-multiplexed lanes (one execution port serving many lanes' streams in turn, each lane's 256 bytes in a bank): modelled 8.5 to 10.5 from the access energies; BUILT as `core_tm` (8 lanes x 64 registers in banks, one port, round-robin, ungated): 8.8 pJ per lane-op synthesised (173,426 cells), against the SIMD ungated 9.7 | 8.8 (built) | 4.4 | 0.71 | synthesised 16:2x UK: port sharing saves 0.9 pJ of units and the bank-select muxes take most of it back; NOT the lever; the multi-family adversary lane's macro-window core reads 5.9 (its FakeRAM term modelled 2.0 to 7.0 pJ per access), within 5 percent of the gated flop row, so the file's form is not the lever either | | values recomputed instead of kept | not available: 95 percent of the window is necessary (above) | | | measured on drawn programs | @@ -388,7 +389,7 @@ to 2.3x and the strongest chips at 2.6x to 4.2x. | 6 | prmt, lop3 | 0.64, 0.72 | 11.5, 13.0 | 0.056, 0.055 | not drawn (RTL rows only) | | 7 | mulhi | 0.68 | 21.0 | 0.032 | yes (6) | | 8 | 32-lane shuffle (butterfly) | 0.63 | 29.4 | 0.021 | yes (8) | -| 9 | L1 scratch read (8 KB flop array) against the card's L2 hit | ROW_SCRATCH_K | 1,400 | pending | not drawn | +| 9 | L1 scratch read (8 KB flop array) against the card's L2 hit | 104 (pJ per read) | 1,400 | 0.074 | not drawn | | 10 | int8 8x8x8 tile, per MAC | ROW_TILE_K | 2.2 | pending | not drawn (the tensor lever is dead on other grounds) | The order is set by the card's price, not the chip's: the chip pays 0.6 to 1.7 pJ for everything, and the card diff --git a/tools/chip-model/rtl/Makefile b/tools/chip-model/rtl/Makefile index bd774375a..f960ebb52 100644 --- a/tools/chip-model/rtl/Makefile +++ b/tools/chip-model/rtl/Makefile @@ -43,17 +43,17 @@ SIMS_shfl := mix SIMS_xbar := mix SIMS_scratch := mix SIMS_tile := mix -SIMS_core8 := mix mixld:+loads=1 -SIMS_core32 := mix mixld:+loads=1 -SIMS_core32r16 := mix mixld:+loads=1 -SIMS_core8r64 := mix +SIMS_core8 := s150:+cycles=150 s600:+cycles=600 +SIMS_core32 := s150:+cycles=150 s600:+cycles=600 +SIMS_core32r16 := s150:+cycles=150 s600:+cycles=600 +SIMS_core8r64 := s150:+cycles=150 s600:+cycles=600 SIMS_core8i1k := mix SIMS_core8sel := mix SIMS_core32all := mix -SIMS_core8g := mix -SIMS_core8r64g := mix +SIMS_core8g := s150:+cycles=150 s600:+cycles=600 +SIMS_core8r64g := s150:+cycles=150 s600:+cycles=600 SIMS_coretm := mix -SIMS_cs64 := mix +SIMS_cs64 := s150:+cycles=150 s600:+cycles=600 SIMS_fp32 := mix fadd:+op=0 fmul:+op=1 ffma:+op=2 fcvt:+op=3 .PHONY: rows table clean