From ed7d2626d49daed4ca2bc38a46286211a8e691ca Mon Sep 17 00:00:00 2001 From: igneum-josh Date: Thu, 8 Oct 2026 14:05:13 +0100 Subject: [PATCH] shadow-k: the programmable core's synthesis-only row (6.9 pJ per lane-op ASAP7, k 0.56 at the lock at N3), the edge table in the absolute convention, the lease-routed Makefile Co-Authored-By: Claude Fable 5.1 --- docs/analysis/class-v6/floor/shadow-k.md | 147 ++++++++++++++++++++++ tools/chip-model/rtl/Makefile | 9 +- tools/chip-model/rtl/flow/collect.py | 12 +- tools/chip-model/rtl/flow/edge.py | 37 +++--- tools/chip-model/rtl/flow/power.tcl | 1 + tools/chip-model/rtl/tb/tb_core_common.vh | 2 +- 6 files changed, 186 insertions(+), 22 deletions(-) diff --git a/docs/analysis/class-v6/floor/shadow-k.md b/docs/analysis/class-v6/floor/shadow-k.md index c722a26d3..fb9544ffd 100644 --- a/docs/analysis/class-v6/floor/shadow-k.md +++ b/docs/analysis/class-v6/floor/shadow-k.md @@ -75,3 +75,150 @@ lanes, approximate), the clock tree beyond the block's own, and the result's mov the card side the 15.1a figure is the whole card's marginal per counted op (the sleep floor subtracted), which includes the card's own fetch, decode, operand collection and register file. So the `k` here is the chip's datapath-and-window cost over the card's whole-lane cost: a FLOOR on the chip's cost and so a floor on `k`. + +## 3. The rows: per-unit floors (the minimal lane per family) + +Every row: ASAP7 routed, SPEF, OpenSTA `report_power` under the random-input gate-level VCD (every pin annotated, +0 unannotated), the TC corner (0.70 V). "pJ/op" is total power (internal + switching + leakage) times the clock +period over the ops per cycle. The propagated-0.5 cross-check is in `table.csv`; it agrees within 2x on the +logic lanes and overestimates the multiplier lanes 50x (OpenSTA's statistical propagation through a multiplier +is not a measurement), so the VCD row is the row. The GPU side is 15.1a: pJ per counted op on the RTX 5090 at +stock and at the 1,300 MHz lock, and 6.9 pJ per counted op on the Apple M5 Max (the class v4 shadow, measured). +`k` is absolute: the chip's own pJ per op at the node over the card's at its point. + +| Family (chip RTL) | Cells (with fill) | pJ/op ASAP7 | pJ/op N5 | pJ/op N3 | pJ/op N2 | 5090 pJ/op stock / lock | k, N3 chip vs 5090 stock / lock | k, N3 vs M5 Max 6.9 | k, ASAP7 unscaled vs lock | Label | +|---|---|---|---|---|---|---|---|---|---|---| +| ARX add | 11,631 | 2.16 | 1.51 | 1.09 | 0.78 | 11.3 / 6.2 | 0.096 / 0.18 | 0.16 | 0.35 | synthesised; N3 and N2 scaled (claimed) | +| ARX sub | | 2.14 | 1.50 | 1.08 | 0.78 | 11.3 / 6.2 | 0.095 / 0.17 | 0.16 | 0.35 | | +| ARX xor | | 2.09 | 1.46 | 1.05 | 0.76 | 11.3 / 6.2 | 0.093 / 0.17 | 0.15 | 0.34 | | +| ARX rotl (immediate) | | 2.15 | 1.50 | 1.08 | 0.78 | 11.3 / 6.2 | 0.095 / 0.17 | 0.16 | 0.35 | | +| ARX rotr (by register) | | 2.13 | 1.49 | 1.07 | 0.77 | 11.3 / 6.2 | 0.095 / 0.17 | 0.16 | 0.34 | | +| or (lossy; the values saturate to ones and the activity falls) | | 1.45 | 1.01 | 0.73 | 0.53 | 11.3 / 6.2 | 0.065 / 0.12 | 0.11 | 0.23 | | +| ARX random mix | | 2.24 | 1.57 | 1.13 | 0.81 | 11.3 / 6.2 | 0.10 / 0.18 | 0.16 | 0.36 | | +| mul (32 x 32, low word) | 19,603 | 1.35 | 0.94 | 0.68 | 0.49 | 13.9 / 8.3 | 0.049 / 0.082 | | 0.16 | | +| mulhi (high word; the same multiplier) | | 1.35 | 0.94 | 0.68 | 0.49 | 39.6 / 21.0 | 0.017 / 0.032 | | 0.064 | | +| mad (src x src2 + dst; three reads) | | 3.29 | 2.30 | 1.66 | 1.19 | 13.9 / 8.3 | 0.12 / 0.20 | | 0.40 | | +| mul random mix | | 1.53 | 1.07 | 0.77 | 0.56 | 13.9 / 8.3 | 0.055 / 0.093 | | 0.18 | | +| index fold (x M, rotl R, masks; era constants in registers) | 14,965 | 2.37 | 1.66 | 1.19 | 0.86 | 13.9 / 8.3 (read against int_mul) | 0.086 / 0.14 | | 0.29 | | +| prmt (byte permute) | 6,744 | 1.27 | 0.89 | 0.64 | 0.46 | 22.3 / 11.5 | 0.029 / 0.056 | | 0.11 | | +| lop3 (8-bit truth table) | 7,274 | 1.42 | 0.99 | 0.72 | 0.52 | 24.1 / 13.0 | 0.030 / 0.055 | | 0.11 | | +| 32-lane xor-mask shuffle (butterfly over the 1 KB window), per lane-op | ROW_SHFL | +| 32-lane general crossbar, per lane-op | ROW_XBAR | +| 8 KB scratch, one random read (flop array: the pessimistic form) | ROW_SCRATCH | +| int8 8x8x8 tile, per MAC | ROW_TILE | + +Reading the floors: a lane's add costs the chip about 2.2 pJ at ASAP7 and 1.1 at N3, against the 5090's 6.2 +at the lock and 11.3 at stock, so even the floor is not a tenth of the card's cost at the knee, and the claimed +"2 to 5 pJ for a SIMD array at N5" (15.1) was the right order for the unscaled lane with nothing around it. The +multiplier is the cheapest unit per op relative to the card (the 5090 pays 8.3 pJ for mul and 21 for mulhi, the +chip 0.68 for either, because the high word falls out of the same array), and mad is the dearest for the chip +(three register reads and two units). The fold costs a chip one multiply and a rotate: 1.2 pJ per address at N3, +or 0.15 nJ per hash over 128 loads, a third of a percent of the GDDR7 board's 0.466 microjoules. + +## 4. The headline row: the programmable sequencer core + +The per-unit lanes of section 3 are floors for a chip that cannot exist under class v6: layers 1 and 3 (per-era +op-mix, program-length and read-width draws, family epochs) kill any fixed lane, so the chip that competes is a +programmable core. `core_v6` (`tools/chip-model/rtl/rtl/core_v6.v`) is the minimal in-order SIMD sequencer that +executes the drawn class v6 shadow program: a 256 x 32-bit instruction memory (a flop array) sized to the drawn +program, a program counter wrapping at the era's drawn length, fetch into an instruction register, decode, the +era's parameter registers (M, R, WM, OFF, MASK, N), and per lane a 32 x 32-bit register file (flops) with three +read ports and one write port and every class unit: add, sub, xor, or, rotl, rotr, mul, mulhi, mad, prmt, lop3, +the xor-mask shuffle across the lanes (a butterfly), and the load (the index fold on the address path, the returned +word written next cycle). One instruction per cycle for every lane. The program is 256 instructions drawn with the +class v4 weights (`NONLOAD_WEIGHTS`), with a variant at one load per 16 instructions. Two builds: 8 lanes (the fast +row) and 32 lanes (the imem amortised over 32), and a 32-lane build with a 16-register file (the sensitivity). + +The activity: the gate-level random-input VCD of the synthesised netlist; two run lengths (150 and 800 cycles) +bracket the 260-cycle program-load phase, and solving the pair gives the steady-state run power (the load phase +draws 13.9 mW, the run phase 36.9 mW for 8 lanes at 1.5 ns). + +| Row | Stage | Cells | pJ per lane-op ASAP7 | N5 | N3 | N2 | 5090 stock / lock / M5 Max pJ per op | k absolute at N3 vs stock / lock / M5 Max | k at N2 | k unscaled ASAP7 vs lock | Label | +|---|---|---|---|---|---|---|---|---|---|---|---| +| core, 8 lanes, 32 registers | synthesis only (no wires, no clock tree) | 186,443 | 6.9 | 4.8 | 3.5 | 2.5 | 11.3 / 6.2 / 6.9 | 0.31 / 0.56 / 0.50 | 0.22 / 0.40 / 0.36 | 1.1 | synthesised; 14:0x UK | +| of which the sequential term (register file, imem and IR clock pins, no clock gating) | | | 2.4 | 1.7 | 1.2 | 0.9 | | | | | | +| of which the units, the read muxes and the butterfly | | | 4.5 | 3.1 | 2.3 | 1.6 | | | | | | +| core, 8 lanes, 32 registers | placed and routed, SPEF | ROW_CORE8_PLACED | +| core, 32 lanes, 32 registers | ROW_CORE32 | +| core, 32 lanes, 16 registers | ROW_CORE32R16 | +| the bare ARX lane (section 3, the floor) | routed | 11,631 | 2.2 | 1.5 | 1.1 | 0.8 | 11.3 / 6.2 / 6.9 | 0.10 / 0.18 / 0.16 | 0.07 / 0.13 / 0.11 | 0.35 | the lower bound | + +Reading: fetch, decode, a 32-register file and the full unit set cost a chip 3.1x the bare lane, and the core's +k at the 1,300 lock is 0.56 at N3 (0.40 at N2), inside the record's claimed 0.3 to 0.8 band and at its centre, with +the bare lane's 0.18 as the lower bound. Two corrections pull opposite ways: placement adds wires and a clock tree +(+20 to +40 percent on a design like this, approximate; the placed rows read it) and a chip maker gates the +register-file clock (one of 32 registers is written per cycle; gating removes about 2.0 of the 2.4 pJ sequential +term, approximate), so the net figure for the re-fold is 7.0 / 4.9 / 3.5 / 2.5 pJ per lane-op at ASAP7 / N5 / N3 / +N2 (the synthesis-only figure within the rounding). The imem is amortised over 8 lanes in this row and over 32 in +the core32 row. + +## 5. The chip edge at the measured k + +`E_chip = E_mem + N_ops x e_chip` (absolute: the chip's shadow cost is 102,100 x 3.5 pJ = 0.36 microjoules per hash +at N3, 0.26 at N2, whatever the card does); the record's convention `E_mem + k x F` beside it with `k` read at the +card's point (it coincides at the lock by construction and is the same at stock here because both are the same +arithmetic on the same ops; it differs when the chip's cost is held fixed while the card's point moves, which is +what the SRAM lane found flattered the die 1.6x). Cards: the 5090 at stock (3.36 microjoules, F 1.10) and at the +lock (2.33, F 0.652), the M5 Max (1.40 on class v4, 0.78 on class v3, 6.9 pJ per op). Chips: the record's GDDR7 +board and the hardware-future lane's rows (their `E_mem` at zero shadow). The record's rows that this recomputes +are hardware-future.md section 5: "2.6x to 2.8x at k = 1 and 4.4x to 5.1x at k = 0.5" for the strongest DRAM chips +against the stock 5090 (the GDDR7 board's own row there is 2.1x and 3.3x). + +Chip shadow per hash, absolute: N3 0.357 microjoules, N2 0.255. +| Chip (E_mem, microjoules) | Card row | Zero shadow | Absolute, N3 core | Absolute, N2 core | Record convention at the measured k (N3) | at k = 1 | at k = 0.5 | +|---|---|---|---|---|---|---|---| +| GDDR7 board, 28 nm controller (the record) (0.466) | 5090 stock (class v4) | 4.8x | 4.1x | 4.7x | 4.1x (k 0.32) | 2.1x | 3.3x | +| GDDR7 board, 28 nm controller (the record) (0.466) | 5090 at the 1,300 MHz lock | 3.6x | 2.8x | 3.2x | 2.8x (k 0.55) | 2.1x | 2.9x | +| GDDR7 board, 28 nm controller (the record) (0.466) | M5 Max (class v4 1.40 microjoules, class v3 0.78; 6.9 pJ per op; F 0.62) | 1.7x | 1.7x | 1.9x | 1.8x (k 0.51) | 1.3x | 1.8x | +| HBM3E one stack (0.321) | 5090 stock (class v4) | 7.0x | 5.0x | 5.8x | 5.0x (k 0.32) | 2.4x | 3.9x | +| HBM3E one stack (0.321) | 5090 at the 1,300 MHz lock | 5.2x | 3.4x | 4.0x | 3.4x (k 0.55) | 2.4x | 3.6x | +| HBM3E one stack (0.321) | M5 Max (class v4 1.40 microjoules, class v3 0.78; 6.9 pJ per op; F 0.62) | 2.4x | 2.1x | 2.4x | 2.2x (k 0.51) | 1.5x | 2.2x | +| HBM4 one stack, N12 base die (0.22) | 5090 stock (class v4) | 10.3x | 5.8x | 7.1x | 5.8x (k 0.32) | 2.5x | 4.4x | +| HBM4 one stack, N12 base die (0.22) | 5090 at the 1,300 MHz lock | 7.6x | 4.0x | 4.9x | 4.0x (k 0.55) | 2.7x | 4.3x | +| HBM4 one stack, N12 base die (0.22) | M5 Max (class v4 1.40 microjoules, class v3 0.78; 6.9 pJ per op; F 0.62) | 3.5x | 2.4x | 2.9x | 2.6x (k 0.51) | 1.7x | 2.6x | +| custom HBM4E base die, N3P (0.18) | 5090 stock (class v4) | 12.6x | 6.3x | 7.7x | 6.3x (k 0.32) | 2.6x | 4.6x | +| custom HBM4E base die, N3P (0.18) | 5090 at the 1,300 MHz lock | 9.3x | 4.3x | 5.4x | 4.3x (k 0.55) | 2.8x | 4.6x | +| custom HBM4E base die, N3P (0.18) | M5 Max (class v4 1.40 microjoules, class v3 0.78; 6.9 pJ per op; F 0.62) | 4.3x | 2.6x | 3.2x | 2.8x (k 0.51) | 1.7x | 2.9x | +| DRAM on logic, hybrid bonded (2029 to 2031) (0.15) | 5090 stock (class v4) | 15.1x | 6.6x | 8.3x | 6.6x (k 0.32) | 2.7x | 4.8x | +| DRAM on logic, hybrid bonded (2029 to 2031) (0.15) | 5090 at the 1,300 MHz lock | 11.2x | 4.6x | 5.7x | 4.6x (k 0.55) | 2.9x | 4.9x | +| DRAM on logic, hybrid bonded (2029 to 2031) (0.15) | M5 Max (class v4 1.40 microjoules, class v3 0.78; 6.9 pJ per op; F 0.62) | 5.2x | 2.8x | 3.5x | 3.0x (k 0.51) | 1.8x | 3.0x | +| SRAM full store, one N2 reticle (0.14) | 5090 stock (class v4) | 16.1x | 6.8x | 8.5x | 6.8x (k 0.32) | 2.7x | 4.9x | +| SRAM full store, one N2 reticle (0.14) | 5090 at the 1,300 MHz lock | 12.0x | 4.7x | 5.9x | 4.7x (k 0.55) | 2.9x | 5.0x | +| SRAM full store, one N2 reticle (0.14) | M5 Max (class v4 1.40 microjoules, class v3 0.78; 6.9 pJ per op; F 0.62) | 5.6x | 2.8x | 3.5x | 3.1x (k 0.51) | 1.8x | 3.1x | + +Floor lane 1 assumption (not in by 18:00 UK): the honest 5090 at 1.0 microjoules per hash at zero shadow. +| Chip | Premium kept at 0.652 (card 1.652): absolute N3 / N2 | Premium scaled with the card, 0.390 (card 1.390): absolute N3 / N2 | zero shadow | +|---|---|---|---| +| GDDR7 board, 28 nm controller (the record) | 2.0x / 2.3x | 1.7x / 1.9x | 2.1x | +| HBM3E one stack | 2.4x / 2.9x | 2.0x / 2.4x | 3.1x | +| HBM4 one stack, N12 base die | 2.9x / 3.5x | 2.4x / 2.9x | 4.5x | +| custom HBM4E base die, N3P | 3.1x / 3.8x | 2.6x / 3.2x | 5.6x | +| DRAM on logic, hybrid bonded (2029 to 2031) | 3.3x / 4.1x | 2.7x / 3.4x | 6.7x | +| SRAM full store, one N2 reticle | 3.3x / 4.2x | 2.8x / 3.5x | 7.1x | + +What the rows say. (1) At the knee the GDDR7 board keeps 2.8x with the shadow on (3.2x at an N2 core) against +3.6x at zero shadow: the class v4 shadow buys the honest card 0.8x of edge, not the 1.5x the k = 1 row served and +not the 0.1x the bare-lane floor would give. (2) The strongest DRAM chips read 4.3x to 4.7x at the knee and 6.3x to +6.8x at stock against the 5090 (the record's 2.6x to 2.8x at k = 1 and 4.4x to 5.1x at k = 0.5 bracket the stock +figure and understate the knee one, because the record's k was read at stock). (3) Against the M5 Max the whole +table compresses to 1.7x to 2.8x: the honest SoC's own operating point is most of the resistance. (4) Floor lane +1's honest 5090 at 1.0 microjoules (assumed; its rows were not in by 18:00 UK) would put the GDDR7 board at 1.7x +to 2.3x and the strongest chips at 2.6x to 4.2x. + +## 7. Sources + +- The GPU side: `docs/analysis/counter-asic-4-research.md` 15.1a (the 5090 microbench, 8 October 2026, measured) + and 20.3 (the packs job: 10.8 / 6.4 pJ per counted op on the class v4 shadow); the M5 Max 6.9 pJ per counted + op from the coordinator's order of 8 October 2026 (the latency-shadow record). +- ASAP7: L. T. Clark et al., "ASAP7: A 7-nm finFET predictive process design kit", Microelectronics Journal 53 + (2016); the ORFS platform files (`flow/platforms/asap7`, the 7.5-track RVT library, TC corner 0.70 V / 0 C). +- The flow: OpenROAD-flow-scripts (docker image `openroad/orfs:latest`, Yosys 0.68, OpenROAD and OpenSTA with + `read_vcd`); iverilog 12 for the gate-level simulation. +- Node scaling (claimed): TSMC N5 "30 percent lower power at the same speed" against N7 (TSMC technology page, + https://www.tsmc.com/english/dedicatedFoundry/technology/logic/l_5nm); N3E "25 to 30 percent lower power" + against N5 (https://www.tsmc.com/english/dedicatedFoundry/technology/logic/l_3nm); N2 "25 to 30 percent lower + power" against N3E (https://www.tsmc.com/english/dedicatedFoundry/technology/logic/l_2nm). Read 8 October 2026. +- The chip model's memory rows: `docs/analysis/chip-model-v3.md` 5.5 (the GDDR7 board, 0.466 microjoules) and + `docs/analysis/class-v6/hardware-future.md` section 5 (the strongest DRAM and SRAM rows). +- The band: `docs/design/class-v6-rotating-family.md` section 2 (layer 1) and `igneum-pow/src/generator.rs` + `NONLOAD_WEIGHTS`. diff --git a/tools/chip-model/rtl/Makefile b/tools/chip-model/rtl/Makefile index e3bc96c38..974a433bd 100644 --- a/tools/chip-model/rtl/Makefile +++ b/tools/chip-model/rtl/Makefile @@ -13,7 +13,14 @@ WORK ?= $(abspath .) ORFS_IMG ?= openroad/orfs:latest SIM_IMG ?= orfs-sim:latest UID_GID := $(shell id -u):$(shell id -g) -DOCKER := docker run --rm -u $(UID_GID) -e HOME=/tmp -v $(WORK):/work +# Every run on a box goes through the lease tool (the coordinator's rule, 8 October 2026): THREADS cores from the +# bounded pool at nice 19; the docker container is pinned to the leased set and ORFS gets NUM_CORES=THREADS. +THREADS ?= 24 +LEASE ?= /srv/builds/_bin/lease +LEASE_ON ?= $(shell test -x $(LEASE) && echo 1) +LEASEPFX = $(if $(LEASE_ON),$(LEASE) pool $(THREADS) --label "floor-k shadow-k" --owner floor-k --nice 19 --,) +CPUSET = $(if $(LEASE_ON),--cpuset-cpus {cpuset},) +DOCKER := $(LEASEPFX) docker run --rm -u $(UID_GID) -e HOME=/tmp -e NUM_CORES=$(THREADS) $(CPUSET) -v $(WORK):/work ORFS := $(DOCKER) -w /OpenROAD-flow-scripts/flow $(ORFS_IMG) SIM := $(DOCKER) -w /work $(SIM_IMG) diff --git a/tools/chip-model/rtl/flow/collect.py b/tools/chip-model/rtl/flow/collect.py index 4c9bbb4e6..01d9fa21b 100644 --- a/tools/chip-model/rtl/flow/collect.py +++ b/tools/chip-model/rtl/flow/collect.py @@ -62,13 +62,19 @@ for d in OPS: row[f'k_{node}_lock'] = (pj * s / gpu[1]) if gpu[1] else None row[f'k_{node}_unlocked'] = (pj * s / gpu[0]) if gpu[0] else None row['gpu_pJ_unlocked'] = gpu[0]; row['gpu_pJ_lock'] = gpu[1] + # the Apple M5 Max: 6.9 pJ per counted op measured on the class v4 shadow (the honest tier's top); only the + # ARX-class and the core rows have a measured M5 Max figure + m5 = 6.9 if (d in ('arx', 'core8', 'core32', 'core32r16')) else None + row['m5_pJ'] = m5 + for node, sc in SCALE.items(): + row[f'k_{node}_m5'] = (pj * sc / m5) if m5 else None out.append(row) if out: with open(os.path.join(work, 'table.csv'), 'w', newline='') as f: w = csv.DictWriter(f, fieldnames=list(out[0].keys())); w.writeheader(); w.writerows(out) -print('| Family | Sim | Period ps | Cells | Ops/cycle | Total W | Leak W | VCD annotated | pJ/op ASAP7 | pJ/op N5 | pJ/op N3 | pJ/op N2 | 5090 pJ/op unlocked / lock | k at N3 (lock) | k at N2 (lock) | k at N3 (unlocked) |') -print('|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|') +print('| Family | Sim | Period ps | Cells | Ops/cycle | Total W | Leak W | VCD annotated | pJ/op ASAP7 | pJ/op N5 | pJ/op N3 | pJ/op N2 | 5090 pJ/op unlocked / lock | k at N3 (lock) | k at N2 (lock) | k at N3 (unlocked) | k at N3 vs M5 Max (6.9) |') +print('|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|') for r in out: f = lambda v, n=3: ('' if v is None else (f'{v:.{n}g}' if isinstance(v, float) else str(v))) - print(f"| {r['family']} | {r['sim']} | {f(r['period_ps'])} | {r['cells']} | {r['ops_per_cycle']} | {f(r['total_W'])} | {f(r['leak_W'])} | {r['annotated_pct']} | {f(r['pJ_asap7'])} | {f(r['pJ_N5'])} | {f(r['pJ_N3'])} | {f(r['pJ_N2'])} | {f(r['gpu_pJ_unlocked'])} / {f(r['gpu_pJ_lock'])} | {f(r['k_N3_lock'])} | {f(r['k_N2_lock'])} | {f(r['k_N3_unlocked'])} |") + print(f"| {r['family']} | {r['sim']} | {f(r['period_ps'])} | {r['cells']} | {r['ops_per_cycle']} | {f(r['total_W'])} | {f(r['leak_W'])} | {r['annotated_pct']} | {f(r['pJ_asap7'])} | {f(r['pJ_N5'])} | {f(r['pJ_N3'])} | {f(r['pJ_N2'])} | {f(r['gpu_pJ_unlocked'])} / {f(r['gpu_pJ_lock'])} | {f(r['k_N3_lock'])} | {f(r['k_N2_lock'])} | {f(r['k_N3_unlocked'])} | {f(r['k_N3_m5'])} |") diff --git a/tools/chip-model/rtl/flow/edge.py b/tools/chip-model/rtl/flow/edge.py index cf4542d07..1363f339d 100644 --- a/tools/chip-model/rtl/flow/edge.py +++ b/tools/chip-model/rtl/flow/edge.py @@ -1,23 +1,26 @@ #!/usr/bin/env python3 -"""The chip edge rows recomputed at the measured k. Usage: edge.py """ +"""The chip edge rows at the synthesised chip cost. Usage: edge.py [] +Absolute convention: E_chip = E_mem + N_ops x e_chip (the chip's shadow cost does not depend on the card's point). +The record's convention beside it: E_chip = E_mem + k x F with k = e_chip / e_gpu at the card's point.""" import sys -k_v4 = float(sys.argv[1]); k_best = float(sys.argv[2]) -# the record's rows: card microjoules per hash and the premium F (counter-asic-4-research.md 20.4; the GDDR7 -# board E_mem 0.466; hardware-future.md section 5 for the strongest DRAM chips) -cards = [('5090 unlocked (class v4 shape)', 3.36, 1.10), ('5090 at the 1,300 MHz lock', 2.33, 0.652)] +e3 = float(sys.argv[1]); e2 = float(sys.argv[2]) if len(sys.argv) > 2 else e3 * 0.72 +NOPS = 102100 # class v4 counted ops per hash (the packs job) +# the 5090 rows: (name, card microjoules per hash, premium F, pJ per counted op at that point) +cards = [('5090 stock (class v4)', 3.36, 1.10, 10.8), ('5090 at the 1,300 MHz lock', 2.33, 0.652, 6.4)] +honest = [('M5 Max (class v4 1.40 microjoules, class v3 0.78; 6.9 pJ per op; F 0.62)', 1.40, 0.62, 6.9)] chips = [('GDDR7 board, 28 nm controller (the record)', 0.466), ('HBM3E one stack', 0.321), ('HBM4 one stack, N12 base die', 0.22), ('custom HBM4E base die, N3P', 0.18), ('DRAM on logic, hybrid bonded (2029 to 2031)', 0.15), ('SRAM full store, one N2 reticle', 0.14)] -def edge(card, F, mem, k): return (card) / (mem + k * F) -print('| Chip | 5090 row | Card uJ | F uJ | Edge at k = 1 | at k = 0.5 | at k = 0.3 | at the measured k (class v4 mix) | at the measured k (best mix) | at zero shadow |') -print('|---|---|---|---|---|---|---|---|---|---|') +def edge_abs(card, mem, e): return card / (mem + NOPS * e * 1e-6) +def edge_rec(card, F, mem, k): return card / (mem + k * F) +print(f'Chip shadow per hash, absolute: N3 {NOPS*e3*1e-6:.3f} microjoules, N2 {NOPS*e2*1e-6:.3f}.') +print('| Chip (E_mem, microjoules) | Card row | Zero shadow | Absolute, N3 core | Absolute, N2 core | Record convention at the measured k (N3) | at k = 1 | at k = 0.5 |') +print('|---|---|---|---|---|---|---|---|') for cn, mem in chips: - for rn, card, F in cards: - print(f'| {cn} | {rn} | {card:.2f} | {F:.3f} | {edge(card,F,mem,1):.2f}x | {edge(card,F,mem,0.5):.2f}x | {edge(card,F,mem,0.3):.2f}x | {edge(card,F,mem,k_v4):.2f}x (k {k_v4:.3f}) | {edge(card,F,mem,k_best):.2f}x (k {k_best:.3f}) | {(card-F)/mem:.2f}x |') -# the honest-card variants: floor lane 1's assumed 1.0 microjoules per hash at zero shadow on the 5090 -print('\nFloor lane 1 assumption: the honest 5090 at 1.0 microjoules per hash at zero shadow (not in by the default time).') -print('| Chip | Premium kept at the lock\'s 0.652 | Premium scaled with the card (0.652 x 1.0 / 1.67 = 0.390) |') -print('|---|---|---|') + for rn, card, F, egpu in cards + honest: + k3 = e3 / egpu + print(f'| {cn} ({mem}) | {rn} | {(card-F)/mem:.1f}x | {edge_abs(card,mem,e3):.1f}x | {edge_abs(card,mem,e2):.1f}x | {edge_rec(card,F,mem,k3):.1f}x (k {k3:.2f}) | {edge_rec(card,F,mem,1):.1f}x | {edge_rec(card,F,mem,0.5):.1f}x |') +print('\nFloor lane 1 assumption (not in by 18:00 UK): the honest 5090 at 1.0 microjoules per hash at zero shadow.') +print('| Chip | Premium kept at 0.652 (card 1.652): absolute N3 / N2 | Premium scaled with the card, 0.390 (card 1.390): absolute N3 / N2 | zero shadow |') +print('|---|---|---|---|') for cn, mem in chips: - a = edge(1.0 + 0.652, 0.652, mem, k_v4); b = edge(1.0 + 0.390, 0.390, mem, k_v4) - a1 = edge(1.0 + 0.652, 0.652, mem, 1.0); b1 = edge(1.0 + 0.390, 0.390, mem, 1.0) - print(f'| {cn} | {a:.2f}x at the measured k ({a1:.2f}x at k = 1; {1.0/mem:.2f}x at zero shadow) | {b:.2f}x at the measured k ({b1:.2f}x at k = 1) |') + print(f'| {cn} | {edge_abs(1.652,mem,e3):.1f}x / {edge_abs(1.652,mem,e2):.1f}x | {edge_abs(1.390,mem,e3):.1f}x / {edge_abs(1.390,mem,e2):.1f}x | {1.0/mem:.1f}x |') diff --git a/tools/chip-model/rtl/flow/power.tcl b/tools/chip-model/rtl/flow/power.tcl index 3e6e7f807..44c19072c 100644 --- a/tools/chip-model/rtl/flow/power.tcl +++ b/tools/chip-model/rtl/flow/power.tcl @@ -18,6 +18,7 @@ set_power_activity -input_port rst -activity 0 -duty 0 report_power foreach vcd [glob -nocomplain /work/sim/$::env(DESIGN_NICKNAME)/*.vcd] { set tag [file rootname [file tail $vcd]] + if { $tag == "dump" } { continue } puts "FLOORK === POWER_VCD $tag ===" read_vcd -scope tb/dut $vcd if { [info commands report_activity_annotation] != "" } { report_activity_annotation } diff --git a/tools/chip-model/rtl/tb/tb_core_common.vh b/tools/chip-model/rtl/tb/tb_core_common.vh index 551ba6b4d..c1dac7f2c 100644 --- a/tools/chip-model/rtl/tb/tb_core_common.vh +++ b/tools/chip-model/rtl/tb/tb_core_common.vh @@ -13,7 +13,7 @@ module tb; endfunction always #`HALF clk = ~clk; initial begin - if (!$value$plusargs("cycles=%d", cycles)) cycles = 4000; + if (!$value$plusargs("cycles=%d", cycles)) cycles = 300; // the core VCDs run to gigabytes per thousand cycles if (!$value$plusargs("loads=%d", loads)) loads = 0; // 1: one load in 16 instructions $dumpfile("dump.vcd"); $dumpvars(0, tb.dut); repeat (4) @(negedge clk); rst = 0;