shadow-k: the programmable core's synthesis-only row (6.9 pJ per lane-op ASAP7, k 0.56 at the lock at N3), the edge table in the absolute convention, the lease-routed Makefile
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
49c49c340d
commit
ed7d2626d4
6 changed files with 186 additions and 22 deletions
|
|
@ -75,3 +75,150 @@ lanes, approximate), the clock tree beyond the block's own, and the result's mov
|
|||
the card side the 15.1a figure is the whole card's marginal per counted op (the sleep floor subtracted), which
|
||||
includes the card's own fetch, decode, operand collection and register file. So the `k` here is the chip's
|
||||
datapath-and-window cost over the card's whole-lane cost: a FLOOR on the chip's cost and so a floor on `k`.
|
||||
|
||||
## 3. The rows: per-unit floors (the minimal lane per family)
|
||||
|
||||
Every row: ASAP7 routed, SPEF, OpenSTA `report_power` under the random-input gate-level VCD (every pin annotated,
|
||||
0 unannotated), the TC corner (0.70 V). "pJ/op" is total power (internal + switching + leakage) times the clock
|
||||
period over the ops per cycle. The propagated-0.5 cross-check is in `table.csv`; it agrees within 2x on the
|
||||
logic lanes and overestimates the multiplier lanes 50x (OpenSTA's statistical propagation through a multiplier
|
||||
is not a measurement), so the VCD row is the row. The GPU side is 15.1a: pJ per counted op on the RTX 5090 at
|
||||
stock and at the 1,300 MHz lock, and 6.9 pJ per counted op on the Apple M5 Max (the class v4 shadow, measured).
|
||||
`k` is absolute: the chip's own pJ per op at the node over the card's at its point.
|
||||
|
||||
| Family (chip RTL) | Cells (with fill) | pJ/op ASAP7 | pJ/op N5 | pJ/op N3 | pJ/op N2 | 5090 pJ/op stock / lock | k, N3 chip vs 5090 stock / lock | k, N3 vs M5 Max 6.9 | k, ASAP7 unscaled vs lock | Label |
|
||||
|---|---|---|---|---|---|---|---|---|---|---|
|
||||
| ARX add | 11,631 | 2.16 | 1.51 | 1.09 | 0.78 | 11.3 / 6.2 | 0.096 / 0.18 | 0.16 | 0.35 | synthesised; N3 and N2 scaled (claimed) |
|
||||
| ARX sub | | 2.14 | 1.50 | 1.08 | 0.78 | 11.3 / 6.2 | 0.095 / 0.17 | 0.16 | 0.35 | |
|
||||
| ARX xor | | 2.09 | 1.46 | 1.05 | 0.76 | 11.3 / 6.2 | 0.093 / 0.17 | 0.15 | 0.34 | |
|
||||
| ARX rotl (immediate) | | 2.15 | 1.50 | 1.08 | 0.78 | 11.3 / 6.2 | 0.095 / 0.17 | 0.16 | 0.35 | |
|
||||
| ARX rotr (by register) | | 2.13 | 1.49 | 1.07 | 0.77 | 11.3 / 6.2 | 0.095 / 0.17 | 0.16 | 0.34 | |
|
||||
| or (lossy; the values saturate to ones and the activity falls) | | 1.45 | 1.01 | 0.73 | 0.53 | 11.3 / 6.2 | 0.065 / 0.12 | 0.11 | 0.23 | |
|
||||
| ARX random mix | | 2.24 | 1.57 | 1.13 | 0.81 | 11.3 / 6.2 | 0.10 / 0.18 | 0.16 | 0.36 | |
|
||||
| mul (32 x 32, low word) | 19,603 | 1.35 | 0.94 | 0.68 | 0.49 | 13.9 / 8.3 | 0.049 / 0.082 | | 0.16 | |
|
||||
| mulhi (high word; the same multiplier) | | 1.35 | 0.94 | 0.68 | 0.49 | 39.6 / 21.0 | 0.017 / 0.032 | | 0.064 | |
|
||||
| mad (src x src2 + dst; three reads) | | 3.29 | 2.30 | 1.66 | 1.19 | 13.9 / 8.3 | 0.12 / 0.20 | | 0.40 | |
|
||||
| mul random mix | | 1.53 | 1.07 | 0.77 | 0.56 | 13.9 / 8.3 | 0.055 / 0.093 | | 0.18 | |
|
||||
| index fold (x M, rotl R, masks; era constants in registers) | 14,965 | 2.37 | 1.66 | 1.19 | 0.86 | 13.9 / 8.3 (read against int_mul) | 0.086 / 0.14 | | 0.29 | |
|
||||
| prmt (byte permute) | 6,744 | 1.27 | 0.89 | 0.64 | 0.46 | 22.3 / 11.5 | 0.029 / 0.056 | | 0.11 | |
|
||||
| lop3 (8-bit truth table) | 7,274 | 1.42 | 0.99 | 0.72 | 0.52 | 24.1 / 13.0 | 0.030 / 0.055 | | 0.11 | |
|
||||
| 32-lane xor-mask shuffle (butterfly over the 1 KB window), per lane-op | ROW_SHFL |
|
||||
| 32-lane general crossbar, per lane-op | ROW_XBAR |
|
||||
| 8 KB scratch, one random read (flop array: the pessimistic form) | ROW_SCRATCH |
|
||||
| int8 8x8x8 tile, per MAC | ROW_TILE |
|
||||
|
||||
Reading the floors: a lane's add costs the chip about 2.2 pJ at ASAP7 and 1.1 at N3, against the 5090's 6.2
|
||||
at the lock and 11.3 at stock, so even the floor is not a tenth of the card's cost at the knee, and the claimed
|
||||
"2 to 5 pJ for a SIMD array at N5" (15.1) was the right order for the unscaled lane with nothing around it. The
|
||||
multiplier is the cheapest unit per op relative to the card (the 5090 pays 8.3 pJ for mul and 21 for mulhi, the
|
||||
chip 0.68 for either, because the high word falls out of the same array), and mad is the dearest for the chip
|
||||
(three register reads and two units). The fold costs a chip one multiply and a rotate: 1.2 pJ per address at N3,
|
||||
or 0.15 nJ per hash over 128 loads, a third of a percent of the GDDR7 board's 0.466 microjoules.
|
||||
|
||||
## 4. The headline row: the programmable sequencer core
|
||||
|
||||
The per-unit lanes of section 3 are floors for a chip that cannot exist under class v6: layers 1 and 3 (per-era
|
||||
op-mix, program-length and read-width draws, family epochs) kill any fixed lane, so the chip that competes is a
|
||||
programmable core. `core_v6` (`tools/chip-model/rtl/rtl/core_v6.v`) is the minimal in-order SIMD sequencer that
|
||||
executes the drawn class v6 shadow program: a 256 x 32-bit instruction memory (a flop array) sized to the drawn
|
||||
program, a program counter wrapping at the era's drawn length, fetch into an instruction register, decode, the
|
||||
era's parameter registers (M, R, WM, OFF, MASK, N), and per lane a 32 x 32-bit register file (flops) with three
|
||||
read ports and one write port and every class unit: add, sub, xor, or, rotl, rotr, mul, mulhi, mad, prmt, lop3,
|
||||
the xor-mask shuffle across the lanes (a butterfly), and the load (the index fold on the address path, the returned
|
||||
word written next cycle). One instruction per cycle for every lane. The program is 256 instructions drawn with the
|
||||
class v4 weights (`NONLOAD_WEIGHTS`), with a variant at one load per 16 instructions. Two builds: 8 lanes (the fast
|
||||
row) and 32 lanes (the imem amortised over 32), and a 32-lane build with a 16-register file (the sensitivity).
|
||||
|
||||
The activity: the gate-level random-input VCD of the synthesised netlist; two run lengths (150 and 800 cycles)
|
||||
bracket the 260-cycle program-load phase, and solving the pair gives the steady-state run power (the load phase
|
||||
draws 13.9 mW, the run phase 36.9 mW for 8 lanes at 1.5 ns).
|
||||
|
||||
| Row | Stage | Cells | pJ per lane-op ASAP7 | N5 | N3 | N2 | 5090 stock / lock / M5 Max pJ per op | k absolute at N3 vs stock / lock / M5 Max | k at N2 | k unscaled ASAP7 vs lock | Label |
|
||||
|---|---|---|---|---|---|---|---|---|---|---|---|
|
||||
| core, 8 lanes, 32 registers | synthesis only (no wires, no clock tree) | 186,443 | 6.9 | 4.8 | 3.5 | 2.5 | 11.3 / 6.2 / 6.9 | 0.31 / 0.56 / 0.50 | 0.22 / 0.40 / 0.36 | 1.1 | synthesised; 14:0x UK |
|
||||
| of which the sequential term (register file, imem and IR clock pins, no clock gating) | | | 2.4 | 1.7 | 1.2 | 0.9 | | | | | |
|
||||
| of which the units, the read muxes and the butterfly | | | 4.5 | 3.1 | 2.3 | 1.6 | | | | | |
|
||||
| core, 8 lanes, 32 registers | placed and routed, SPEF | ROW_CORE8_PLACED |
|
||||
| core, 32 lanes, 32 registers | ROW_CORE32 |
|
||||
| core, 32 lanes, 16 registers | ROW_CORE32R16 |
|
||||
| the bare ARX lane (section 3, the floor) | routed | 11,631 | 2.2 | 1.5 | 1.1 | 0.8 | 11.3 / 6.2 / 6.9 | 0.10 / 0.18 / 0.16 | 0.07 / 0.13 / 0.11 | 0.35 | the lower bound |
|
||||
|
||||
Reading: fetch, decode, a 32-register file and the full unit set cost a chip 3.1x the bare lane, and the core's
|
||||
k at the 1,300 lock is 0.56 at N3 (0.40 at N2), inside the record's claimed 0.3 to 0.8 band and at its centre, with
|
||||
the bare lane's 0.18 as the lower bound. Two corrections pull opposite ways: placement adds wires and a clock tree
|
||||
(+20 to +40 percent on a design like this, approximate; the placed rows read it) and a chip maker gates the
|
||||
register-file clock (one of 32 registers is written per cycle; gating removes about 2.0 of the 2.4 pJ sequential
|
||||
term, approximate), so the net figure for the re-fold is 7.0 / 4.9 / 3.5 / 2.5 pJ per lane-op at ASAP7 / N5 / N3 /
|
||||
N2 (the synthesis-only figure within the rounding). The imem is amortised over 8 lanes in this row and over 32 in
|
||||
the core32 row.
|
||||
|
||||
## 5. The chip edge at the measured k
|
||||
|
||||
`E_chip = E_mem + N_ops x e_chip` (absolute: the chip's shadow cost is 102,100 x 3.5 pJ = 0.36 microjoules per hash
|
||||
at N3, 0.26 at N2, whatever the card does); the record's convention `E_mem + k x F` beside it with `k` read at the
|
||||
card's point (it coincides at the lock by construction and is the same at stock here because both are the same
|
||||
arithmetic on the same ops; it differs when the chip's cost is held fixed while the card's point moves, which is
|
||||
what the SRAM lane found flattered the die 1.6x). Cards: the 5090 at stock (3.36 microjoules, F 1.10) and at the
|
||||
lock (2.33, F 0.652), the M5 Max (1.40 on class v4, 0.78 on class v3, 6.9 pJ per op). Chips: the record's GDDR7
|
||||
board and the hardware-future lane's rows (their `E_mem` at zero shadow). The record's rows that this recomputes
|
||||
are hardware-future.md section 5: "2.6x to 2.8x at k = 1 and 4.4x to 5.1x at k = 0.5" for the strongest DRAM chips
|
||||
against the stock 5090 (the GDDR7 board's own row there is 2.1x and 3.3x).
|
||||
|
||||
Chip shadow per hash, absolute: N3 0.357 microjoules, N2 0.255.
|
||||
| Chip (E_mem, microjoules) | Card row | Zero shadow | Absolute, N3 core | Absolute, N2 core | Record convention at the measured k (N3) | at k = 1 | at k = 0.5 |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| GDDR7 board, 28 nm controller (the record) (0.466) | 5090 stock (class v4) | 4.8x | 4.1x | 4.7x | 4.1x (k 0.32) | 2.1x | 3.3x |
|
||||
| GDDR7 board, 28 nm controller (the record) (0.466) | 5090 at the 1,300 MHz lock | 3.6x | 2.8x | 3.2x | 2.8x (k 0.55) | 2.1x | 2.9x |
|
||||
| GDDR7 board, 28 nm controller (the record) (0.466) | M5 Max (class v4 1.40 microjoules, class v3 0.78; 6.9 pJ per op; F 0.62) | 1.7x | 1.7x | 1.9x | 1.8x (k 0.51) | 1.3x | 1.8x |
|
||||
| HBM3E one stack (0.321) | 5090 stock (class v4) | 7.0x | 5.0x | 5.8x | 5.0x (k 0.32) | 2.4x | 3.9x |
|
||||
| HBM3E one stack (0.321) | 5090 at the 1,300 MHz lock | 5.2x | 3.4x | 4.0x | 3.4x (k 0.55) | 2.4x | 3.6x |
|
||||
| HBM3E one stack (0.321) | M5 Max (class v4 1.40 microjoules, class v3 0.78; 6.9 pJ per op; F 0.62) | 2.4x | 2.1x | 2.4x | 2.2x (k 0.51) | 1.5x | 2.2x |
|
||||
| HBM4 one stack, N12 base die (0.22) | 5090 stock (class v4) | 10.3x | 5.8x | 7.1x | 5.8x (k 0.32) | 2.5x | 4.4x |
|
||||
| HBM4 one stack, N12 base die (0.22) | 5090 at the 1,300 MHz lock | 7.6x | 4.0x | 4.9x | 4.0x (k 0.55) | 2.7x | 4.3x |
|
||||
| HBM4 one stack, N12 base die (0.22) | M5 Max (class v4 1.40 microjoules, class v3 0.78; 6.9 pJ per op; F 0.62) | 3.5x | 2.4x | 2.9x | 2.6x (k 0.51) | 1.7x | 2.6x |
|
||||
| custom HBM4E base die, N3P (0.18) | 5090 stock (class v4) | 12.6x | 6.3x | 7.7x | 6.3x (k 0.32) | 2.6x | 4.6x |
|
||||
| custom HBM4E base die, N3P (0.18) | 5090 at the 1,300 MHz lock | 9.3x | 4.3x | 5.4x | 4.3x (k 0.55) | 2.8x | 4.6x |
|
||||
| custom HBM4E base die, N3P (0.18) | M5 Max (class v4 1.40 microjoules, class v3 0.78; 6.9 pJ per op; F 0.62) | 4.3x | 2.6x | 3.2x | 2.8x (k 0.51) | 1.7x | 2.9x |
|
||||
| DRAM on logic, hybrid bonded (2029 to 2031) (0.15) | 5090 stock (class v4) | 15.1x | 6.6x | 8.3x | 6.6x (k 0.32) | 2.7x | 4.8x |
|
||||
| DRAM on logic, hybrid bonded (2029 to 2031) (0.15) | 5090 at the 1,300 MHz lock | 11.2x | 4.6x | 5.7x | 4.6x (k 0.55) | 2.9x | 4.9x |
|
||||
| DRAM on logic, hybrid bonded (2029 to 2031) (0.15) | M5 Max (class v4 1.40 microjoules, class v3 0.78; 6.9 pJ per op; F 0.62) | 5.2x | 2.8x | 3.5x | 3.0x (k 0.51) | 1.8x | 3.0x |
|
||||
| SRAM full store, one N2 reticle (0.14) | 5090 stock (class v4) | 16.1x | 6.8x | 8.5x | 6.8x (k 0.32) | 2.7x | 4.9x |
|
||||
| SRAM full store, one N2 reticle (0.14) | 5090 at the 1,300 MHz lock | 12.0x | 4.7x | 5.9x | 4.7x (k 0.55) | 2.9x | 5.0x |
|
||||
| SRAM full store, one N2 reticle (0.14) | M5 Max (class v4 1.40 microjoules, class v3 0.78; 6.9 pJ per op; F 0.62) | 5.6x | 2.8x | 3.5x | 3.1x (k 0.51) | 1.8x | 3.1x |
|
||||
|
||||
Floor lane 1 assumption (not in by 18:00 UK): the honest 5090 at 1.0 microjoules per hash at zero shadow.
|
||||
| Chip | Premium kept at 0.652 (card 1.652): absolute N3 / N2 | Premium scaled with the card, 0.390 (card 1.390): absolute N3 / N2 | zero shadow |
|
||||
|---|---|---|---|
|
||||
| GDDR7 board, 28 nm controller (the record) | 2.0x / 2.3x | 1.7x / 1.9x | 2.1x |
|
||||
| HBM3E one stack | 2.4x / 2.9x | 2.0x / 2.4x | 3.1x |
|
||||
| HBM4 one stack, N12 base die | 2.9x / 3.5x | 2.4x / 2.9x | 4.5x |
|
||||
| custom HBM4E base die, N3P | 3.1x / 3.8x | 2.6x / 3.2x | 5.6x |
|
||||
| DRAM on logic, hybrid bonded (2029 to 2031) | 3.3x / 4.1x | 2.7x / 3.4x | 6.7x |
|
||||
| SRAM full store, one N2 reticle | 3.3x / 4.2x | 2.8x / 3.5x | 7.1x |
|
||||
|
||||
What the rows say. (1) At the knee the GDDR7 board keeps 2.8x with the shadow on (3.2x at an N2 core) against
|
||||
3.6x at zero shadow: the class v4 shadow buys the honest card 0.8x of edge, not the 1.5x the k = 1 row served and
|
||||
not the 0.1x the bare-lane floor would give. (2) The strongest DRAM chips read 4.3x to 4.7x at the knee and 6.3x to
|
||||
6.8x at stock against the 5090 (the record's 2.6x to 2.8x at k = 1 and 4.4x to 5.1x at k = 0.5 bracket the stock
|
||||
figure and understate the knee one, because the record's k was read at stock). (3) Against the M5 Max the whole
|
||||
table compresses to 1.7x to 2.8x: the honest SoC's own operating point is most of the resistance. (4) Floor lane
|
||||
1's honest 5090 at 1.0 microjoules (assumed; its rows were not in by 18:00 UK) would put the GDDR7 board at 1.7x
|
||||
to 2.3x and the strongest chips at 2.6x to 4.2x.
|
||||
|
||||
## 7. Sources
|
||||
|
||||
- The GPU side: `docs/analysis/counter-asic-4-research.md` 15.1a (the 5090 microbench, 8 October 2026, measured)
|
||||
and 20.3 (the packs job: 10.8 / 6.4 pJ per counted op on the class v4 shadow); the M5 Max 6.9 pJ per counted
|
||||
op from the coordinator's order of 8 October 2026 (the latency-shadow record).
|
||||
- ASAP7: L. T. Clark et al., "ASAP7: A 7-nm finFET predictive process design kit", Microelectronics Journal 53
|
||||
(2016); the ORFS platform files (`flow/platforms/asap7`, the 7.5-track RVT library, TC corner 0.70 V / 0 C).
|
||||
- The flow: OpenROAD-flow-scripts (docker image `openroad/orfs:latest`, Yosys 0.68, OpenROAD and OpenSTA with
|
||||
`read_vcd`); iverilog 12 for the gate-level simulation.
|
||||
- Node scaling (claimed): TSMC N5 "30 percent lower power at the same speed" against N7 (TSMC technology page,
|
||||
https://www.tsmc.com/english/dedicatedFoundry/technology/logic/l_5nm); N3E "25 to 30 percent lower power"
|
||||
against N5 (https://www.tsmc.com/english/dedicatedFoundry/technology/logic/l_3nm); N2 "25 to 30 percent lower
|
||||
power" against N3E (https://www.tsmc.com/english/dedicatedFoundry/technology/logic/l_2nm). Read 8 October 2026.
|
||||
- The chip model's memory rows: `docs/analysis/chip-model-v3.md` 5.5 (the GDDR7 board, 0.466 microjoules) and
|
||||
`docs/analysis/class-v6/hardware-future.md` section 5 (the strongest DRAM and SRAM rows).
|
||||
- The band: `docs/design/class-v6-rotating-family.md` section 2 (layer 1) and `igneum-pow/src/generator.rs`
|
||||
`NONLOAD_WEIGHTS`.
|
||||
|
|
|
|||
|
|
@ -13,7 +13,14 @@ WORK ?= $(abspath .)
|
|||
ORFS_IMG ?= openroad/orfs:latest
|
||||
SIM_IMG ?= orfs-sim:latest
|
||||
UID_GID := $(shell id -u):$(shell id -g)
|
||||
DOCKER := docker run --rm -u $(UID_GID) -e HOME=/tmp -v $(WORK):/work
|
||||
# Every run on a box goes through the lease tool (the coordinator's rule, 8 October 2026): THREADS cores from the
|
||||
# bounded pool at nice 19; the docker container is pinned to the leased set and ORFS gets NUM_CORES=THREADS.
|
||||
THREADS ?= 24
|
||||
LEASE ?= /srv/builds/_bin/lease
|
||||
LEASE_ON ?= $(shell test -x $(LEASE) && echo 1)
|
||||
LEASEPFX = $(if $(LEASE_ON),$(LEASE) pool $(THREADS) --label "floor-k shadow-k" --owner floor-k --nice 19 --,)
|
||||
CPUSET = $(if $(LEASE_ON),--cpuset-cpus {cpuset},)
|
||||
DOCKER := $(LEASEPFX) docker run --rm -u $(UID_GID) -e HOME=/tmp -e NUM_CORES=$(THREADS) $(CPUSET) -v $(WORK):/work
|
||||
ORFS := $(DOCKER) -w /OpenROAD-flow-scripts/flow $(ORFS_IMG)
|
||||
SIM := $(DOCKER) -w /work $(SIM_IMG)
|
||||
|
||||
|
|
|
|||
|
|
@ -62,13 +62,19 @@ for d in OPS:
|
|||
row[f'k_{node}_lock'] = (pj * s / gpu[1]) if gpu[1] else None
|
||||
row[f'k_{node}_unlocked'] = (pj * s / gpu[0]) if gpu[0] else None
|
||||
row['gpu_pJ_unlocked'] = gpu[0]; row['gpu_pJ_lock'] = gpu[1]
|
||||
# the Apple M5 Max: 6.9 pJ per counted op measured on the class v4 shadow (the honest tier's top); only the
|
||||
# ARX-class and the core rows have a measured M5 Max figure
|
||||
m5 = 6.9 if (d in ('arx', 'core8', 'core32', 'core32r16')) else None
|
||||
row['m5_pJ'] = m5
|
||||
for node, sc in SCALE.items():
|
||||
row[f'k_{node}_m5'] = (pj * sc / m5) if m5 else None
|
||||
out.append(row)
|
||||
|
||||
if out:
|
||||
with open(os.path.join(work, 'table.csv'), 'w', newline='') as f:
|
||||
w = csv.DictWriter(f, fieldnames=list(out[0].keys())); w.writeheader(); w.writerows(out)
|
||||
print('| Family | Sim | Period ps | Cells | Ops/cycle | Total W | Leak W | VCD annotated | pJ/op ASAP7 | pJ/op N5 | pJ/op N3 | pJ/op N2 | 5090 pJ/op unlocked / lock | k at N3 (lock) | k at N2 (lock) | k at N3 (unlocked) |')
|
||||
print('|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|')
|
||||
print('| Family | Sim | Period ps | Cells | Ops/cycle | Total W | Leak W | VCD annotated | pJ/op ASAP7 | pJ/op N5 | pJ/op N3 | pJ/op N2 | 5090 pJ/op unlocked / lock | k at N3 (lock) | k at N2 (lock) | k at N3 (unlocked) | k at N3 vs M5 Max (6.9) |')
|
||||
print('|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|')
|
||||
for r in out:
|
||||
f = lambda v, n=3: ('' if v is None else (f'{v:.{n}g}' if isinstance(v, float) else str(v)))
|
||||
print(f"| {r['family']} | {r['sim']} | {f(r['period_ps'])} | {r['cells']} | {r['ops_per_cycle']} | {f(r['total_W'])} | {f(r['leak_W'])} | {r['annotated_pct']} | {f(r['pJ_asap7'])} | {f(r['pJ_N5'])} | {f(r['pJ_N3'])} | {f(r['pJ_N2'])} | {f(r['gpu_pJ_unlocked'])} / {f(r['gpu_pJ_lock'])} | {f(r['k_N3_lock'])} | {f(r['k_N2_lock'])} | {f(r['k_N3_unlocked'])} |")
|
||||
print(f"| {r['family']} | {r['sim']} | {f(r['period_ps'])} | {r['cells']} | {r['ops_per_cycle']} | {f(r['total_W'])} | {f(r['leak_W'])} | {r['annotated_pct']} | {f(r['pJ_asap7'])} | {f(r['pJ_N5'])} | {f(r['pJ_N3'])} | {f(r['pJ_N2'])} | {f(r['gpu_pJ_unlocked'])} / {f(r['gpu_pJ_lock'])} | {f(r['k_N3_lock'])} | {f(r['k_N2_lock'])} | {f(r['k_N3_unlocked'])} | {f(r['k_N3_m5'])} |")
|
||||
|
|
|
|||
|
|
@ -1,23 +1,26 @@
|
|||
#!/usr/bin/env python3
|
||||
"""The chip edge rows recomputed at the measured k. Usage: edge.py <k_eff_v4> <k_eff_best>"""
|
||||
"""The chip edge rows at the synthesised chip cost. Usage: edge.py <chip pJ per lane-op at N3> [<at N2>]
|
||||
Absolute convention: E_chip = E_mem + N_ops x e_chip (the chip's shadow cost does not depend on the card's point).
|
||||
The record's convention beside it: E_chip = E_mem + k x F with k = e_chip / e_gpu at the card's point."""
|
||||
import sys
|
||||
k_v4 = float(sys.argv[1]); k_best = float(sys.argv[2])
|
||||
# the record's rows: card microjoules per hash and the premium F (counter-asic-4-research.md 20.4; the GDDR7
|
||||
# board E_mem 0.466; hardware-future.md section 5 for the strongest DRAM chips)
|
||||
cards = [('5090 unlocked (class v4 shape)', 3.36, 1.10), ('5090 at the 1,300 MHz lock', 2.33, 0.652)]
|
||||
e3 = float(sys.argv[1]); e2 = float(sys.argv[2]) if len(sys.argv) > 2 else e3 * 0.72
|
||||
NOPS = 102100 # class v4 counted ops per hash (the packs job)
|
||||
# the 5090 rows: (name, card microjoules per hash, premium F, pJ per counted op at that point)
|
||||
cards = [('5090 stock (class v4)', 3.36, 1.10, 10.8), ('5090 at the 1,300 MHz lock', 2.33, 0.652, 6.4)]
|
||||
honest = [('M5 Max (class v4 1.40 microjoules, class v3 0.78; 6.9 pJ per op; F 0.62)', 1.40, 0.62, 6.9)]
|
||||
chips = [('GDDR7 board, 28 nm controller (the record)', 0.466), ('HBM3E one stack', 0.321), ('HBM4 one stack, N12 base die', 0.22),
|
||||
('custom HBM4E base die, N3P', 0.18), ('DRAM on logic, hybrid bonded (2029 to 2031)', 0.15), ('SRAM full store, one N2 reticle', 0.14)]
|
||||
def edge(card, F, mem, k): return (card) / (mem + k * F)
|
||||
print('| Chip | 5090 row | Card uJ | F uJ | Edge at k = 1 | at k = 0.5 | at k = 0.3 | at the measured k (class v4 mix) | at the measured k (best mix) | at zero shadow |')
|
||||
print('|---|---|---|---|---|---|---|---|---|---|')
|
||||
def edge_abs(card, mem, e): return card / (mem + NOPS * e * 1e-6)
|
||||
def edge_rec(card, F, mem, k): return card / (mem + k * F)
|
||||
print(f'Chip shadow per hash, absolute: N3 {NOPS*e3*1e-6:.3f} microjoules, N2 {NOPS*e2*1e-6:.3f}.')
|
||||
print('| Chip (E_mem, microjoules) | Card row | Zero shadow | Absolute, N3 core | Absolute, N2 core | Record convention at the measured k (N3) | at k = 1 | at k = 0.5 |')
|
||||
print('|---|---|---|---|---|---|---|---|')
|
||||
for cn, mem in chips:
|
||||
for rn, card, F in cards:
|
||||
print(f'| {cn} | {rn} | {card:.2f} | {F:.3f} | {edge(card,F,mem,1):.2f}x | {edge(card,F,mem,0.5):.2f}x | {edge(card,F,mem,0.3):.2f}x | {edge(card,F,mem,k_v4):.2f}x (k {k_v4:.3f}) | {edge(card,F,mem,k_best):.2f}x (k {k_best:.3f}) | {(card-F)/mem:.2f}x |')
|
||||
# the honest-card variants: floor lane 1's assumed 1.0 microjoules per hash at zero shadow on the 5090
|
||||
print('\nFloor lane 1 assumption: the honest 5090 at 1.0 microjoules per hash at zero shadow (not in by the default time).')
|
||||
print('| Chip | Premium kept at the lock\'s 0.652 | Premium scaled with the card (0.652 x 1.0 / 1.67 = 0.390) |')
|
||||
print('|---|---|---|')
|
||||
for rn, card, F, egpu in cards + honest:
|
||||
k3 = e3 / egpu
|
||||
print(f'| {cn} ({mem}) | {rn} | {(card-F)/mem:.1f}x | {edge_abs(card,mem,e3):.1f}x | {edge_abs(card,mem,e2):.1f}x | {edge_rec(card,F,mem,k3):.1f}x (k {k3:.2f}) | {edge_rec(card,F,mem,1):.1f}x | {edge_rec(card,F,mem,0.5):.1f}x |')
|
||||
print('\nFloor lane 1 assumption (not in by 18:00 UK): the honest 5090 at 1.0 microjoules per hash at zero shadow.')
|
||||
print('| Chip | Premium kept at 0.652 (card 1.652): absolute N3 / N2 | Premium scaled with the card, 0.390 (card 1.390): absolute N3 / N2 | zero shadow |')
|
||||
print('|---|---|---|---|')
|
||||
for cn, mem in chips:
|
||||
a = edge(1.0 + 0.652, 0.652, mem, k_v4); b = edge(1.0 + 0.390, 0.390, mem, k_v4)
|
||||
a1 = edge(1.0 + 0.652, 0.652, mem, 1.0); b1 = edge(1.0 + 0.390, 0.390, mem, 1.0)
|
||||
print(f'| {cn} | {a:.2f}x at the measured k ({a1:.2f}x at k = 1; {1.0/mem:.2f}x at zero shadow) | {b:.2f}x at the measured k ({b1:.2f}x at k = 1) |')
|
||||
print(f'| {cn} | {edge_abs(1.652,mem,e3):.1f}x / {edge_abs(1.652,mem,e2):.1f}x | {edge_abs(1.390,mem,e3):.1f}x / {edge_abs(1.390,mem,e2):.1f}x | {1.0/mem:.1f}x |')
|
||||
|
|
|
|||
|
|
@ -18,6 +18,7 @@ set_power_activity -input_port rst -activity 0 -duty 0
|
|||
report_power
|
||||
foreach vcd [glob -nocomplain /work/sim/$::env(DESIGN_NICKNAME)/*.vcd] {
|
||||
set tag [file rootname [file tail $vcd]]
|
||||
if { $tag == "dump" } { continue }
|
||||
puts "FLOORK === POWER_VCD $tag ==="
|
||||
read_vcd -scope tb/dut $vcd
|
||||
if { [info commands report_activity_annotation] != "" } { report_activity_annotation }
|
||||
|
|
|
|||
|
|
@ -13,7 +13,7 @@ module tb;
|
|||
endfunction
|
||||
always #`HALF clk = ~clk;
|
||||
initial begin
|
||||
if (!$value$plusargs("cycles=%d", cycles)) cycles = 4000;
|
||||
if (!$value$plusargs("cycles=%d", cycles)) cycles = 300; // the core VCDs run to gigabytes per thousand cycles
|
||||
if (!$value$plusargs("loads=%d", loads)) loads = 0; // 1: one load in 16 instructions
|
||||
$dumpfile("dump.vcd"); $dumpvars(0, tb.dut);
|
||||
repeat (4) @(negedge clk); rst = 0;
|
||||
|
|
|
|||
Loading…
Reference in a new issue