shadow-k: the programmable core's synthesis-only row (6.9 pJ per lane-op ASAP7, k 0.56 at the lock at N3), the edge table in the absolute convention, the lease-routed Makefile

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-josh 2026-10-08 14:05:13 +01:00
parent 49c49c340d
commit ed7d2626d4
6 changed files with 186 additions and 22 deletions

View file

@ -75,3 +75,150 @@ lanes, approximate), the clock tree beyond the block's own, and the result's mov
the card side the 15.1a figure is the whole card's marginal per counted op (the sleep floor subtracted), which
includes the card's own fetch, decode, operand collection and register file. So the `k` here is the chip's
datapath-and-window cost over the card's whole-lane cost: a FLOOR on the chip's cost and so a floor on `k`.
## 3. The rows: per-unit floors (the minimal lane per family)
Every row: ASAP7 routed, SPEF, OpenSTA `report_power` under the random-input gate-level VCD (every pin annotated,
0 unannotated), the TC corner (0.70 V). "pJ/op" is total power (internal + switching + leakage) times the clock
period over the ops per cycle. The propagated-0.5 cross-check is in `table.csv`; it agrees within 2x on the
logic lanes and overestimates the multiplier lanes 50x (OpenSTA's statistical propagation through a multiplier
is not a measurement), so the VCD row is the row. The GPU side is 15.1a: pJ per counted op on the RTX 5090 at
stock and at the 1,300 MHz lock, and 6.9 pJ per counted op on the Apple M5 Max (the class v4 shadow, measured).
`k` is absolute: the chip's own pJ per op at the node over the card's at its point.
| Family (chip RTL) | Cells (with fill) | pJ/op ASAP7 | pJ/op N5 | pJ/op N3 | pJ/op N2 | 5090 pJ/op stock / lock | k, N3 chip vs 5090 stock / lock | k, N3 vs M5 Max 6.9 | k, ASAP7 unscaled vs lock | Label |
|---|---|---|---|---|---|---|---|---|---|---|
| ARX add | 11,631 | 2.16 | 1.51 | 1.09 | 0.78 | 11.3 / 6.2 | 0.096 / 0.18 | 0.16 | 0.35 | synthesised; N3 and N2 scaled (claimed) |
| ARX sub | | 2.14 | 1.50 | 1.08 | 0.78 | 11.3 / 6.2 | 0.095 / 0.17 | 0.16 | 0.35 | |
| ARX xor | | 2.09 | 1.46 | 1.05 | 0.76 | 11.3 / 6.2 | 0.093 / 0.17 | 0.15 | 0.34 | |
| ARX rotl (immediate) | | 2.15 | 1.50 | 1.08 | 0.78 | 11.3 / 6.2 | 0.095 / 0.17 | 0.16 | 0.35 | |
| ARX rotr (by register) | | 2.13 | 1.49 | 1.07 | 0.77 | 11.3 / 6.2 | 0.095 / 0.17 | 0.16 | 0.34 | |
| or (lossy; the values saturate to ones and the activity falls) | | 1.45 | 1.01 | 0.73 | 0.53 | 11.3 / 6.2 | 0.065 / 0.12 | 0.11 | 0.23 | |
| ARX random mix | | 2.24 | 1.57 | 1.13 | 0.81 | 11.3 / 6.2 | 0.10 / 0.18 | 0.16 | 0.36 | |
| mul (32 x 32, low word) | 19,603 | 1.35 | 0.94 | 0.68 | 0.49 | 13.9 / 8.3 | 0.049 / 0.082 | | 0.16 | |
| mulhi (high word; the same multiplier) | | 1.35 | 0.94 | 0.68 | 0.49 | 39.6 / 21.0 | 0.017 / 0.032 | | 0.064 | |
| mad (src x src2 + dst; three reads) | | 3.29 | 2.30 | 1.66 | 1.19 | 13.9 / 8.3 | 0.12 / 0.20 | | 0.40 | |
| mul random mix | | 1.53 | 1.07 | 0.77 | 0.56 | 13.9 / 8.3 | 0.055 / 0.093 | | 0.18 | |
| index fold (x M, rotl R, masks; era constants in registers) | 14,965 | 2.37 | 1.66 | 1.19 | 0.86 | 13.9 / 8.3 (read against int_mul) | 0.086 / 0.14 | | 0.29 | |
| prmt (byte permute) | 6,744 | 1.27 | 0.89 | 0.64 | 0.46 | 22.3 / 11.5 | 0.029 / 0.056 | | 0.11 | |
| lop3 (8-bit truth table) | 7,274 | 1.42 | 0.99 | 0.72 | 0.52 | 24.1 / 13.0 | 0.030 / 0.055 | | 0.11 | |
| 32-lane xor-mask shuffle (butterfly over the 1 KB window), per lane-op | ROW_SHFL |
| 32-lane general crossbar, per lane-op | ROW_XBAR |
| 8 KB scratch, one random read (flop array: the pessimistic form) | ROW_SCRATCH |
| int8 8x8x8 tile, per MAC | ROW_TILE |
Reading the floors: a lane's add costs the chip about 2.2 pJ at ASAP7 and 1.1 at N3, against the 5090's 6.2
at the lock and 11.3 at stock, so even the floor is not a tenth of the card's cost at the knee, and the claimed
"2 to 5 pJ for a SIMD array at N5" (15.1) was the right order for the unscaled lane with nothing around it. The
multiplier is the cheapest unit per op relative to the card (the 5090 pays 8.3 pJ for mul and 21 for mulhi, the
chip 0.68 for either, because the high word falls out of the same array), and mad is the dearest for the chip
(three register reads and two units). The fold costs a chip one multiply and a rotate: 1.2 pJ per address at N3,
or 0.15 nJ per hash over 128 loads, a third of a percent of the GDDR7 board's 0.466 microjoules.
## 4. The headline row: the programmable sequencer core
The per-unit lanes of section 3 are floors for a chip that cannot exist under class v6: layers 1 and 3 (per-era
op-mix, program-length and read-width draws, family epochs) kill any fixed lane, so the chip that competes is a
programmable core. `core_v6` (`tools/chip-model/rtl/rtl/core_v6.v`) is the minimal in-order SIMD sequencer that
executes the drawn class v6 shadow program: a 256 x 32-bit instruction memory (a flop array) sized to the drawn
program, a program counter wrapping at the era's drawn length, fetch into an instruction register, decode, the
era's parameter registers (M, R, WM, OFF, MASK, N), and per lane a 32 x 32-bit register file (flops) with three
read ports and one write port and every class unit: add, sub, xor, or, rotl, rotr, mul, mulhi, mad, prmt, lop3,
the xor-mask shuffle across the lanes (a butterfly), and the load (the index fold on the address path, the returned
word written next cycle). One instruction per cycle for every lane. The program is 256 instructions drawn with the
class v4 weights (`NONLOAD_WEIGHTS`), with a variant at one load per 16 instructions. Two builds: 8 lanes (the fast
row) and 32 lanes (the imem amortised over 32), and a 32-lane build with a 16-register file (the sensitivity).
The activity: the gate-level random-input VCD of the synthesised netlist; two run lengths (150 and 800 cycles)
bracket the 260-cycle program-load phase, and solving the pair gives the steady-state run power (the load phase
draws 13.9 mW, the run phase 36.9 mW for 8 lanes at 1.5 ns).
| Row | Stage | Cells | pJ per lane-op ASAP7 | N5 | N3 | N2 | 5090 stock / lock / M5 Max pJ per op | k absolute at N3 vs stock / lock / M5 Max | k at N2 | k unscaled ASAP7 vs lock | Label |
|---|---|---|---|---|---|---|---|---|---|---|---|
| core, 8 lanes, 32 registers | synthesis only (no wires, no clock tree) | 186,443 | 6.9 | 4.8 | 3.5 | 2.5 | 11.3 / 6.2 / 6.9 | 0.31 / 0.56 / 0.50 | 0.22 / 0.40 / 0.36 | 1.1 | synthesised; 14:0x UK |
| of which the sequential term (register file, imem and IR clock pins, no clock gating) | | | 2.4 | 1.7 | 1.2 | 0.9 | | | | | |
| of which the units, the read muxes and the butterfly | | | 4.5 | 3.1 | 2.3 | 1.6 | | | | | |
| core, 8 lanes, 32 registers | placed and routed, SPEF | ROW_CORE8_PLACED |
| core, 32 lanes, 32 registers | ROW_CORE32 |
| core, 32 lanes, 16 registers | ROW_CORE32R16 |
| the bare ARX lane (section 3, the floor) | routed | 11,631 | 2.2 | 1.5 | 1.1 | 0.8 | 11.3 / 6.2 / 6.9 | 0.10 / 0.18 / 0.16 | 0.07 / 0.13 / 0.11 | 0.35 | the lower bound |
Reading: fetch, decode, a 32-register file and the full unit set cost a chip 3.1x the bare lane, and the core's
k at the 1,300 lock is 0.56 at N3 (0.40 at N2), inside the record's claimed 0.3 to 0.8 band and at its centre, with
the bare lane's 0.18 as the lower bound. Two corrections pull opposite ways: placement adds wires and a clock tree
(+20 to +40 percent on a design like this, approximate; the placed rows read it) and a chip maker gates the
register-file clock (one of 32 registers is written per cycle; gating removes about 2.0 of the 2.4 pJ sequential
term, approximate), so the net figure for the re-fold is 7.0 / 4.9 / 3.5 / 2.5 pJ per lane-op at ASAP7 / N5 / N3 /
N2 (the synthesis-only figure within the rounding). The imem is amortised over 8 lanes in this row and over 32 in
the core32 row.
## 5. The chip edge at the measured k
`E_chip = E_mem + N_ops x e_chip` (absolute: the chip's shadow cost is 102,100 x 3.5 pJ = 0.36 microjoules per hash
at N3, 0.26 at N2, whatever the card does); the record's convention `E_mem + k x F` beside it with `k` read at the
card's point (it coincides at the lock by construction and is the same at stock here because both are the same
arithmetic on the same ops; it differs when the chip's cost is held fixed while the card's point moves, which is
what the SRAM lane found flattered the die 1.6x). Cards: the 5090 at stock (3.36 microjoules, F 1.10) and at the
lock (2.33, F 0.652), the M5 Max (1.40 on class v4, 0.78 on class v3, 6.9 pJ per op). Chips: the record's GDDR7
board and the hardware-future lane's rows (their `E_mem` at zero shadow). The record's rows that this recomputes
are hardware-future.md section 5: "2.6x to 2.8x at k = 1 and 4.4x to 5.1x at k = 0.5" for the strongest DRAM chips
against the stock 5090 (the GDDR7 board's own row there is 2.1x and 3.3x).
Chip shadow per hash, absolute: N3 0.357 microjoules, N2 0.255.
| Chip (E_mem, microjoules) | Card row | Zero shadow | Absolute, N3 core | Absolute, N2 core | Record convention at the measured k (N3) | at k = 1 | at k = 0.5 |
|---|---|---|---|---|---|---|---|
| GDDR7 board, 28 nm controller (the record) (0.466) | 5090 stock (class v4) | 4.8x | 4.1x | 4.7x | 4.1x (k 0.32) | 2.1x | 3.3x |
| GDDR7 board, 28 nm controller (the record) (0.466) | 5090 at the 1,300 MHz lock | 3.6x | 2.8x | 3.2x | 2.8x (k 0.55) | 2.1x | 2.9x |
| GDDR7 board, 28 nm controller (the record) (0.466) | M5 Max (class v4 1.40 microjoules, class v3 0.78; 6.9 pJ per op; F 0.62) | 1.7x | 1.7x | 1.9x | 1.8x (k 0.51) | 1.3x | 1.8x |
| HBM3E one stack (0.321) | 5090 stock (class v4) | 7.0x | 5.0x | 5.8x | 5.0x (k 0.32) | 2.4x | 3.9x |
| HBM3E one stack (0.321) | 5090 at the 1,300 MHz lock | 5.2x | 3.4x | 4.0x | 3.4x (k 0.55) | 2.4x | 3.6x |
| HBM3E one stack (0.321) | M5 Max (class v4 1.40 microjoules, class v3 0.78; 6.9 pJ per op; F 0.62) | 2.4x | 2.1x | 2.4x | 2.2x (k 0.51) | 1.5x | 2.2x |
| HBM4 one stack, N12 base die (0.22) | 5090 stock (class v4) | 10.3x | 5.8x | 7.1x | 5.8x (k 0.32) | 2.5x | 4.4x |
| HBM4 one stack, N12 base die (0.22) | 5090 at the 1,300 MHz lock | 7.6x | 4.0x | 4.9x | 4.0x (k 0.55) | 2.7x | 4.3x |
| HBM4 one stack, N12 base die (0.22) | M5 Max (class v4 1.40 microjoules, class v3 0.78; 6.9 pJ per op; F 0.62) | 3.5x | 2.4x | 2.9x | 2.6x (k 0.51) | 1.7x | 2.6x |
| custom HBM4E base die, N3P (0.18) | 5090 stock (class v4) | 12.6x | 6.3x | 7.7x | 6.3x (k 0.32) | 2.6x | 4.6x |
| custom HBM4E base die, N3P (0.18) | 5090 at the 1,300 MHz lock | 9.3x | 4.3x | 5.4x | 4.3x (k 0.55) | 2.8x | 4.6x |
| custom HBM4E base die, N3P (0.18) | M5 Max (class v4 1.40 microjoules, class v3 0.78; 6.9 pJ per op; F 0.62) | 4.3x | 2.6x | 3.2x | 2.8x (k 0.51) | 1.7x | 2.9x |
| DRAM on logic, hybrid bonded (2029 to 2031) (0.15) | 5090 stock (class v4) | 15.1x | 6.6x | 8.3x | 6.6x (k 0.32) | 2.7x | 4.8x |
| DRAM on logic, hybrid bonded (2029 to 2031) (0.15) | 5090 at the 1,300 MHz lock | 11.2x | 4.6x | 5.7x | 4.6x (k 0.55) | 2.9x | 4.9x |
| DRAM on logic, hybrid bonded (2029 to 2031) (0.15) | M5 Max (class v4 1.40 microjoules, class v3 0.78; 6.9 pJ per op; F 0.62) | 5.2x | 2.8x | 3.5x | 3.0x (k 0.51) | 1.8x | 3.0x |
| SRAM full store, one N2 reticle (0.14) | 5090 stock (class v4) | 16.1x | 6.8x | 8.5x | 6.8x (k 0.32) | 2.7x | 4.9x |
| SRAM full store, one N2 reticle (0.14) | 5090 at the 1,300 MHz lock | 12.0x | 4.7x | 5.9x | 4.7x (k 0.55) | 2.9x | 5.0x |
| SRAM full store, one N2 reticle (0.14) | M5 Max (class v4 1.40 microjoules, class v3 0.78; 6.9 pJ per op; F 0.62) | 5.6x | 2.8x | 3.5x | 3.1x (k 0.51) | 1.8x | 3.1x |
Floor lane 1 assumption (not in by 18:00 UK): the honest 5090 at 1.0 microjoules per hash at zero shadow.
| Chip | Premium kept at 0.652 (card 1.652): absolute N3 / N2 | Premium scaled with the card, 0.390 (card 1.390): absolute N3 / N2 | zero shadow |
|---|---|---|---|
| GDDR7 board, 28 nm controller (the record) | 2.0x / 2.3x | 1.7x / 1.9x | 2.1x |
| HBM3E one stack | 2.4x / 2.9x | 2.0x / 2.4x | 3.1x |
| HBM4 one stack, N12 base die | 2.9x / 3.5x | 2.4x / 2.9x | 4.5x |
| custom HBM4E base die, N3P | 3.1x / 3.8x | 2.6x / 3.2x | 5.6x |
| DRAM on logic, hybrid bonded (2029 to 2031) | 3.3x / 4.1x | 2.7x / 3.4x | 6.7x |
| SRAM full store, one N2 reticle | 3.3x / 4.2x | 2.8x / 3.5x | 7.1x |
What the rows say. (1) At the knee the GDDR7 board keeps 2.8x with the shadow on (3.2x at an N2 core) against
3.6x at zero shadow: the class v4 shadow buys the honest card 0.8x of edge, not the 1.5x the k = 1 row served and
not the 0.1x the bare-lane floor would give. (2) The strongest DRAM chips read 4.3x to 4.7x at the knee and 6.3x to
6.8x at stock against the 5090 (the record's 2.6x to 2.8x at k = 1 and 4.4x to 5.1x at k = 0.5 bracket the stock
figure and understate the knee one, because the record's k was read at stock). (3) Against the M5 Max the whole
table compresses to 1.7x to 2.8x: the honest SoC's own operating point is most of the resistance. (4) Floor lane
1's honest 5090 at 1.0 microjoules (assumed; its rows were not in by 18:00 UK) would put the GDDR7 board at 1.7x
to 2.3x and the strongest chips at 2.6x to 4.2x.
## 7. Sources
- The GPU side: `docs/analysis/counter-asic-4-research.md` 15.1a (the 5090 microbench, 8 October 2026, measured)
and 20.3 (the packs job: 10.8 / 6.4 pJ per counted op on the class v4 shadow); the M5 Max 6.9 pJ per counted
op from the coordinator's order of 8 October 2026 (the latency-shadow record).
- ASAP7: L. T. Clark et al., "ASAP7: A 7-nm finFET predictive process design kit", Microelectronics Journal 53
(2016); the ORFS platform files (`flow/platforms/asap7`, the 7.5-track RVT library, TC corner 0.70 V / 0 C).
- The flow: OpenROAD-flow-scripts (docker image `openroad/orfs:latest`, Yosys 0.68, OpenROAD and OpenSTA with
`read_vcd`); iverilog 12 for the gate-level simulation.
- Node scaling (claimed): TSMC N5 "30 percent lower power at the same speed" against N7 (TSMC technology page,
https://www.tsmc.com/english/dedicatedFoundry/technology/logic/l_5nm); N3E "25 to 30 percent lower power"
against N5 (https://www.tsmc.com/english/dedicatedFoundry/technology/logic/l_3nm); N2 "25 to 30 percent lower
power" against N3E (https://www.tsmc.com/english/dedicatedFoundry/technology/logic/l_2nm). Read 8 October 2026.
- The chip model's memory rows: `docs/analysis/chip-model-v3.md` 5.5 (the GDDR7 board, 0.466 microjoules) and
`docs/analysis/class-v6/hardware-future.md` section 5 (the strongest DRAM and SRAM rows).
- The band: `docs/design/class-v6-rotating-family.md` section 2 (layer 1) and `igneum-pow/src/generator.rs`
`NONLOAD_WEIGHTS`.

View file

@ -13,7 +13,14 @@ WORK ?= $(abspath .)
ORFS_IMG ?= openroad/orfs:latest
SIM_IMG ?= orfs-sim:latest
UID_GID := $(shell id -u):$(shell id -g)
DOCKER := docker run --rm -u $(UID_GID) -e HOME=/tmp -v $(WORK):/work
# Every run on a box goes through the lease tool (the coordinator's rule, 8 October 2026): THREADS cores from the
# bounded pool at nice 19; the docker container is pinned to the leased set and ORFS gets NUM_CORES=THREADS.
THREADS ?= 24
LEASE ?= /srv/builds/_bin/lease
LEASE_ON ?= $(shell test -x $(LEASE) && echo 1)
LEASEPFX = $(if $(LEASE_ON),$(LEASE) pool $(THREADS) --label "floor-k shadow-k" --owner floor-k --nice 19 --,)
CPUSET = $(if $(LEASE_ON),--cpuset-cpus {cpuset},)
DOCKER := $(LEASEPFX) docker run --rm -u $(UID_GID) -e HOME=/tmp -e NUM_CORES=$(THREADS) $(CPUSET) -v $(WORK):/work
ORFS := $(DOCKER) -w /OpenROAD-flow-scripts/flow $(ORFS_IMG)
SIM := $(DOCKER) -w /work $(SIM_IMG)

View file

@ -62,13 +62,19 @@ for d in OPS:
row[f'k_{node}_lock'] = (pj * s / gpu[1]) if gpu[1] else None
row[f'k_{node}_unlocked'] = (pj * s / gpu[0]) if gpu[0] else None
row['gpu_pJ_unlocked'] = gpu[0]; row['gpu_pJ_lock'] = gpu[1]
# the Apple M5 Max: 6.9 pJ per counted op measured on the class v4 shadow (the honest tier's top); only the
# ARX-class and the core rows have a measured M5 Max figure
m5 = 6.9 if (d in ('arx', 'core8', 'core32', 'core32r16')) else None
row['m5_pJ'] = m5
for node, sc in SCALE.items():
row[f'k_{node}_m5'] = (pj * sc / m5) if m5 else None
out.append(row)
if out:
with open(os.path.join(work, 'table.csv'), 'w', newline='') as f:
w = csv.DictWriter(f, fieldnames=list(out[0].keys())); w.writeheader(); w.writerows(out)
print('| Family | Sim | Period ps | Cells | Ops/cycle | Total W | Leak W | VCD annotated | pJ/op ASAP7 | pJ/op N5 | pJ/op N3 | pJ/op N2 | 5090 pJ/op unlocked / lock | k at N3 (lock) | k at N2 (lock) | k at N3 (unlocked) |')
print('|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|')
print('| Family | Sim | Period ps | Cells | Ops/cycle | Total W | Leak W | VCD annotated | pJ/op ASAP7 | pJ/op N5 | pJ/op N3 | pJ/op N2 | 5090 pJ/op unlocked / lock | k at N3 (lock) | k at N2 (lock) | k at N3 (unlocked) | k at N3 vs M5 Max (6.9) |')
print('|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|')
for r in out:
f = lambda v, n=3: ('' if v is None else (f'{v:.{n}g}' if isinstance(v, float) else str(v)))
print(f"| {r['family']} | {r['sim']} | {f(r['period_ps'])} | {r['cells']} | {r['ops_per_cycle']} | {f(r['total_W'])} | {f(r['leak_W'])} | {r['annotated_pct']} | {f(r['pJ_asap7'])} | {f(r['pJ_N5'])} | {f(r['pJ_N3'])} | {f(r['pJ_N2'])} | {f(r['gpu_pJ_unlocked'])} / {f(r['gpu_pJ_lock'])} | {f(r['k_N3_lock'])} | {f(r['k_N2_lock'])} | {f(r['k_N3_unlocked'])} |")
print(f"| {r['family']} | {r['sim']} | {f(r['period_ps'])} | {r['cells']} | {r['ops_per_cycle']} | {f(r['total_W'])} | {f(r['leak_W'])} | {r['annotated_pct']} | {f(r['pJ_asap7'])} | {f(r['pJ_N5'])} | {f(r['pJ_N3'])} | {f(r['pJ_N2'])} | {f(r['gpu_pJ_unlocked'])} / {f(r['gpu_pJ_lock'])} | {f(r['k_N3_lock'])} | {f(r['k_N2_lock'])} | {f(r['k_N3_unlocked'])} | {f(r['k_N3_m5'])} |")

View file

@ -1,23 +1,26 @@
#!/usr/bin/env python3
"""The chip edge rows recomputed at the measured k. Usage: edge.py <k_eff_v4> <k_eff_best>"""
"""The chip edge rows at the synthesised chip cost. Usage: edge.py <chip pJ per lane-op at N3> [<at N2>]
Absolute convention: E_chip = E_mem + N_ops x e_chip (the chip's shadow cost does not depend on the card's point).
The record's convention beside it: E_chip = E_mem + k x F with k = e_chip / e_gpu at the card's point."""
import sys
k_v4 = float(sys.argv[1]); k_best = float(sys.argv[2])
# the record's rows: card microjoules per hash and the premium F (counter-asic-4-research.md 20.4; the GDDR7
# board E_mem 0.466; hardware-future.md section 5 for the strongest DRAM chips)
cards = [('5090 unlocked (class v4 shape)', 3.36, 1.10), ('5090 at the 1,300 MHz lock', 2.33, 0.652)]
e3 = float(sys.argv[1]); e2 = float(sys.argv[2]) if len(sys.argv) > 2 else e3 * 0.72
NOPS = 102100 # class v4 counted ops per hash (the packs job)
# the 5090 rows: (name, card microjoules per hash, premium F, pJ per counted op at that point)
cards = [('5090 stock (class v4)', 3.36, 1.10, 10.8), ('5090 at the 1,300 MHz lock', 2.33, 0.652, 6.4)]
honest = [('M5 Max (class v4 1.40 microjoules, class v3 0.78; 6.9 pJ per op; F 0.62)', 1.40, 0.62, 6.9)]
chips = [('GDDR7 board, 28 nm controller (the record)', 0.466), ('HBM3E one stack', 0.321), ('HBM4 one stack, N12 base die', 0.22),
('custom HBM4E base die, N3P', 0.18), ('DRAM on logic, hybrid bonded (2029 to 2031)', 0.15), ('SRAM full store, one N2 reticle', 0.14)]
def edge(card, F, mem, k): return (card) / (mem + k * F)
print('| Chip | 5090 row | Card uJ | F uJ | Edge at k = 1 | at k = 0.5 | at k = 0.3 | at the measured k (class v4 mix) | at the measured k (best mix) | at zero shadow |')
print('|---|---|---|---|---|---|---|---|---|---|')
def edge_abs(card, mem, e): return card / (mem + NOPS * e * 1e-6)
def edge_rec(card, F, mem, k): return card / (mem + k * F)
print(f'Chip shadow per hash, absolute: N3 {NOPS*e3*1e-6:.3f} microjoules, N2 {NOPS*e2*1e-6:.3f}.')
print('| Chip (E_mem, microjoules) | Card row | Zero shadow | Absolute, N3 core | Absolute, N2 core | Record convention at the measured k (N3) | at k = 1 | at k = 0.5 |')
print('|---|---|---|---|---|---|---|---|')
for cn, mem in chips:
for rn, card, F in cards:
print(f'| {cn} | {rn} | {card:.2f} | {F:.3f} | {edge(card,F,mem,1):.2f}x | {edge(card,F,mem,0.5):.2f}x | {edge(card,F,mem,0.3):.2f}x | {edge(card,F,mem,k_v4):.2f}x (k {k_v4:.3f}) | {edge(card,F,mem,k_best):.2f}x (k {k_best:.3f}) | {(card-F)/mem:.2f}x |')
# the honest-card variants: floor lane 1's assumed 1.0 microjoules per hash at zero shadow on the 5090
print('\nFloor lane 1 assumption: the honest 5090 at 1.0 microjoules per hash at zero shadow (not in by the default time).')
print('| Chip | Premium kept at the lock\'s 0.652 | Premium scaled with the card (0.652 x 1.0 / 1.67 = 0.390) |')
print('|---|---|---|')
for rn, card, F, egpu in cards + honest:
k3 = e3 / egpu
print(f'| {cn} ({mem}) | {rn} | {(card-F)/mem:.1f}x | {edge_abs(card,mem,e3):.1f}x | {edge_abs(card,mem,e2):.1f}x | {edge_rec(card,F,mem,k3):.1f}x (k {k3:.2f}) | {edge_rec(card,F,mem,1):.1f}x | {edge_rec(card,F,mem,0.5):.1f}x |')
print('\nFloor lane 1 assumption (not in by 18:00 UK): the honest 5090 at 1.0 microjoules per hash at zero shadow.')
print('| Chip | Premium kept at 0.652 (card 1.652): absolute N3 / N2 | Premium scaled with the card, 0.390 (card 1.390): absolute N3 / N2 | zero shadow |')
print('|---|---|---|---|')
for cn, mem in chips:
a = edge(1.0 + 0.652, 0.652, mem, k_v4); b = edge(1.0 + 0.390, 0.390, mem, k_v4)
a1 = edge(1.0 + 0.652, 0.652, mem, 1.0); b1 = edge(1.0 + 0.390, 0.390, mem, 1.0)
print(f'| {cn} | {a:.2f}x at the measured k ({a1:.2f}x at k = 1; {1.0/mem:.2f}x at zero shadow) | {b:.2f}x at the measured k ({b1:.2f}x at k = 1) |')
print(f'| {cn} | {edge_abs(1.652,mem,e3):.1f}x / {edge_abs(1.652,mem,e2):.1f}x | {edge_abs(1.390,mem,e3):.1f}x / {edge_abs(1.390,mem,e2):.1f}x | {1.0/mem:.1f}x |')

View file

@ -18,6 +18,7 @@ set_power_activity -input_port rst -activity 0 -duty 0
report_power
foreach vcd [glob -nocomplain /work/sim/$::env(DESIGN_NICKNAME)/*.vcd] {
set tag [file rootname [file tail $vcd]]
if { $tag == "dump" } { continue }
puts "FLOORK === POWER_VCD $tag ==="
read_vcd -scope tb/dut $vcd
if { [info commands report_activity_annotation] != "" } { report_activity_annotation }

View file

@ -13,7 +13,7 @@ module tb;
endfunction
always #`HALF clk = ~clk;
initial begin
if (!$value$plusargs("cycles=%d", cycles)) cycles = 4000;
if (!$value$plusargs("cycles=%d", cycles)) cycles = 300; // the core VCDs run to gigabytes per thousand cycles
if (!$value$plusargs("loads=%d", loads)) loads = 0; // 1: one load in 16 instructions
$dumpfile("dump.vcd"); $dumpvars(0, tb.dut);
repeat (4) @(negedge clk); rst = 0;