multi-family adversary: the mf core RTL (SRAM window and imem, every bank entry, operand isolation, 5-phase slot), the flow, the collector, the board and hash models, the first rows (synthesis only, 8 lanes, full and base)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

Documents-only replay of a807edd84 (2db208912e5fc0656930aaabd7db74ec2a68ca8e) for the box mirror master; left on the branch: tools/chip-model/mf/Makefile tools/chip-model/mf/flow/board.py tools/chip-model/mf/flow/collect.py tools/chip-model/mf/flow/gl2sim.sh tools/chip-model/mf/flow/hash.py tools/chip-model/mf/flow/mf32base.mk tools/chip-model/mf/flow/mf32base.sdc tools/chip-model/mf/flow/mf32full.mk tools/chip-model/mf/flow/mf32full.sdc tools/chip-model/mf/flow/mf8base.mk tools/chip-model/mf/flow/mf8base.sdc tools/chip-model/mf/flow/mf8full.mk tools/chip-model/mf/flow/mf8full.sdc tools/chip-model/mf/flow/power.tcl tools/chip-model/mf/flow/sim.sh tools/chip-model/mf/results/board-mf8full-n3.md tools/chip-model/mf/results/board-mf8full-n5.md tools/chip-model/mf/results/hash-mf8base.md tools/chip-model/mf/results/hash-mf8full.md tools/chip-model/mf/results/out/mf8base/logs/asap7/mf8base/base/1_synth.json tools/chip-model/mf/results/out/mf8base/logs/asap7/mf8base/base/1_synth.log tools/chip-model/mf/results/out/mf8base/logs/asap7/mf8base/base/power_synth.log tools/chip-model/mf/results/out/mf8base/objects/asap7/mf8base/base/copyright.txt tools/chip-model/mf/results/out/mf8base/reports/asap7/mf8base/base/synth_check.txt tools/chip-model/mf/results/out/mf8base/reports/asap7/mf8base/base/synth_stat.txt tools/chip-model/mf/results/out/mf8base/results/asap7/mf8base/base/mem.json tools/chip-model/mf/results/out/mf8full/logs/asap7/mf8full/base/1_synth.json tools/chip-model/mf/results/out/mf8full/logs/asap7/mf8full/base/1_synth.log tools/chip-model/mf/results/out/mf8full/logs/asap7/mf8full/base/power_synth.log tools/chip-model/mf/results/out/mf8full/objects/asap7/mf8full/base/copyright.txt tools/chip-model/mf/results/out/mf8full/reports/asap7/mf8full/base/synth_check.txt tools/chip-model/mf/results/out/mf8full/reports/asap7/mf8full/base/synth_stat.txt tools/chip-model/mf/results/out/mf8full/results/asap7/mf8full/base/mem.json tools/chip-model/mf/results/sim/mf8base/gl2sim.log tools/chip-model/mf/results/sim/mf8base/sim_add_s2000.log tools/chip-model/mf/results/sim/mf8base/sim_add_s500.log tools/chip-model/mf/results/sim/mf8base/sim_load_s2000.log tools/chip-model/mf/results/sim/mf8base/sim_load_s500.log tools/chip-model/mf/results/sim/mf8base/sim_mad_s2000.log tools/chip-model/mf/results/sim/mf8base/sim_mad_s500.log tools/chip-model/mf/results/sim/mf8base/sim_mix_s2000.log tools/chip-model/mf/results/sim/mf8base/sim_mix_s500.log tools/chip-model/mf/results/sim/mf8base/sim_mixld_s2000.log tools/chip-model/mf/results/sim/mf8base/sim_mixld_s500.log tools/chip-model/mf/results/sim/mf8base/sim_mul_s2000.log tools/chip-model/mf/results/sim/mf8base/sim_mul_s500.log tools/chip-model/mf/results/sim/mf8base/sim_mulhi_s2000.log tools/chip-model/mf/results/sim/mf8base/sim_mulhi_s500.log tools/chip-model/mf/results/sim/mf8base/sim_or_s2000.log tools/chip-model/mf/results/sim/mf8base/sim_or_s500.log tools/chip-model/mf/results/sim/mf8base/sim_rotl_s2000.log tools/chip-model/mf/results/sim/mf8base/sim_rotl_s500.log tools/chip-model/mf/results/sim/mf8base/sim_rotr_s2000.log tools/chip-model/mf/results/sim/mf8base/sim_rotr_s500.log tools/chip-model/mf/results/sim/mf8base/sim_shfl_s2000.log tools/chip-model/mf/results/sim/mf8base/sim_shfl_s500.log tools/chip-model/mf/results/sim/mf8base/sim_sub_s2000.log tools/chip-model/mf/results/sim/mf8base/sim_sub_s500.log tools/chip-model/mf/results/sim/mf8base/sim_xor_s2000.log tools/chip-model/mf/results/sim/mf8base/sim_xor_s500.log tools/chip-model/mf/results/sim/mf8full/gl2sim.log tools/chip-model/mf/results/sim/mf8full/sim_add_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_add_s500.log tools/chip-model/mf/results/sim/mf8full/sim_andn_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_andn_s500.log tools/chip-model/mf/results/sim/mf8full/sim_bfe_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_bfe_s500.log tools/chip-model/mf/results/sim/mf8full/sim_clz_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_clz_s500.log tools/chip-model/mf/results/sim/mf8full/sim_fwd_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_fwd_s500.log tools/chip-model/mf/results/sim/mf8full/sim_load_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_load_s500.log tools/chip-model/mf/results/sim/mf8full/sim_lop3_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_lop3_s500.log tools/chip-model/mf/results/sim/mf8full/sim_mad_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_mad_s500.log tools/chip-model/mf/results/sim/mf8full/sim_mix1_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_mix1_s500.log tools/chip-model/mf/results/sim/mf8full/sim_mix2_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_mix2_s500.log tools/chip-model/mf/results/sim/mf8full/sim_mix64_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_mix64_s500.log tools/chip-model/mf/results/sim/mf8full/sim_mix_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_mix_s500.log tools/chip-model/mf/results/sim/mf8full/sim_mixld_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_mixld_s500.log tools/chip-model/mf/results/sim/mf8full/sim_mixw4_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_mixw4_s500.log tools/chip-model/mf/results/sim/mf8full/sim_mm8_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_mm8_s500.log tools/chip-model/mf/results/sim/mf8full/sim_mul_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_mul_s500.log tools/chip-model/mf/results/sim/mf8full/sim_mulhi_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_mulhi_s500.log tools/chip-model/mf/results/sim/mf8full/sim_or_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_or_s500.log tools/chip-model/mf/results/sim/mf8full/sim_popc_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_popc_s500.log tools/chip-model/mf/results/sim/mf8full/sim_prmt_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_prmt_s500.log tools/chip-model/mf/results/sim/mf8full/sim_rotl_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_rotl_s500.log tools/chip-model/mf/results/sim/mf8full/sim_rotr_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_rotr_s500.log tools/chip-model/mf/results/sim/mf8full/sim_sel_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_sel_s500.log tools/chip-model/mf/results/sim/mf8full/sim_shfl_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_shfl_s500.log tools/chip-model/mf/results/sim/mf8full/sim_shfla_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_shfla_s500.log tools/chip-model/mf/results/sim/mf8full/sim_shl_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_shl_s500.log tools/chip-model/mf/results/sim/mf8full/sim_shr_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_shr_s500.log tools/chip-model/mf/results/sim/mf8full/sim_sub_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_sub_s500.log tools/chip-model/mf/results/sim/mf8full/sim_xor_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_xor_s500.log tools/chip-model/mf/results/table.csv tools/chip-model/mf/results/table.md tools/chip-model/mf/rtl/core_mf.v tools/chip-model/mf/rtl/mf32base.v tools/chip-model/mf/rtl/mf32full.v tools/chip-model/mf/rtl/mf8base.v tools/chip-model/mf/rtl/mf8full.v tools/chip-model/mf/tb/sram_models.v tools/chip-model/mf/tb/tb_mf32base.v tools/chip-model/mf/tb/tb_mf32full.v tools/chip-model/mf/tb/tb_mf8base.v tools/chip-model/mf/tb/tb_mf8full.v tools/chip-model/mf/tb/tb_mf_common.vh
This commit is contained in:
igneum-labs 2026-10-08 15:20:03 +00:00
parent 59915e86df
commit 01452ff1ea

View file

@ -0,0 +1,407 @@
# The multi-family adversary: one programmable chip against the 180-day family bank (8 October 2026)
Branch `class-v6-adversary` from the box mirror's master (c86a7e23), the adversary lane of the second external review
(the founder's 15:5x BST acceptance). The question the review put: does the 180-day family calendar add cost to a chip,
or only change its firmware? The answer is built as the adversary would build it: ONE programmable design that executes
every entry of the bank (`docs/analysis/rotation/layer3-family-bank.md`: 18 op families, 3 dataset atoms, the fold form
with drawn constants, 3 block shapes, the 64-register window, the per-era parameter draws), with the designer free to
choose lanes, clock, pipeline, banking, time-multiplexing, memory organisation and the memory arrangement itself, and
priced on a complete machine. Every chip figure is synthesis and placement of RTL written for this lane (Yosys 0.68 and
OpenROAD, the ORFS image, ASAP7, run on rented CPU hosts, never on the Mac, never on a Hetzner box); every GPU figure is
the record's measurement (`docs/analysis/counter-asic-4-research.md` 15.1a, the floor programme's tier table in the
design document 10.4). Labels as the record uses them: measured, modelled, claimed, approximate. The RTL, testbenches,
flow and collector are under `tools/chip-model/mf/`; the k lane's rows (`floor/shadow-k.md`) are cited, never re-run.
The standing rules this file is written under (the review, 15:4x BST): the adversary's design is free; a synthesised
k is a model, never a lower bound; a family transition is credited with an obsolescence benefit only where a loss of
competitiveness is demonstrated; the stress life is 3 years, shown at 0.5, 1, 2 and 3; the cohort is the discrete-GPU
population, Apple reported and not headlined; every defence is scored after the adversary re-optimises.
## 0. One page
(filled from the rows below at each cut; section 8 carries the three-part statement)
## 1. The design: what the adversary builds
### 1.1 What it must execute
| Bank entry | What the core does with it | Silicon or firmware |
|---|---|---|
| G1 to G7: add, sub, xor, or, rotl, rotr, and the lossy or | the ALU group, one unit per lane, operands isolated | silicon once; the mix weights are the program |
| G8, G9, G6: mul, mulhi, mad | one 32 x 32 multiplier per lane; mulhi is the high word of the same product; mad adds the third read | silicon once |
| G10 shfl (the 32-lane xor-mask shuffle) | a log2(LANES)-stage butterfly across the core's lanes, the mask from the instruction | silicon once per core |
| R1 shfla (lane + delta) | a general LANES:1 crossbar per lane, the delta from a register | silicon once per core (the one network a butterfly cannot emulate in one op) |
| R2 perm (prmt) | the 4-of-8 byte selector | silicon once |
| R3 popc and clz | a popcount tree and a priority encoder | silicon once |
| R4 bfe, R5 shl and shr, R6 sel, R7 andn | a shifter, a mask, a select, an and-not | silicon once (fractions of an adder each) |
| R8 mm8 (the int8 tile) | a u8 dot4 accumulate per lane (4 MACs per op; the card's m8n8k16 tile is 32 MACs per lane, so 8 chip ops per card tile) | silicon once |
| lop3 | the 8-bit truth table | silicon once |
| the load and the fold form | the lane's own multiplier computes `x * M`, then the rotate and the three masks with the era's constants in registers; the returned word writes the destination | silicon once; M, R, WM, OFF, MASK are registers written at the era |
| W = 4 wide reads | the fold step `x = rotl(x, r) * M ^ w` as an instruction (fwd), three per load | firmware |
| the 64-register window | 64 x 32 bits per lane in an SRAM macro | silicon once (the macro) |
| the 3 block shapes (64, 128, 256) | the program-length register; the imem holds 256 | firmware |
| the drawn select tree | a 32-entry op permutation ahead of decode, written at the era | firmware (a 160-bit register) |
| the op-mix band, the fold constants | the program and five registers | firmware |
| the 3 dataset atoms (mixer x4, x8, dr368) | the per-window dataset build runs on the same core as a program (the atoms are straight-line ARX and multiply code over 16 registers); the per-hash path never executes an atom | firmware; the build's cost is section 7 |
### 1.2 The microarchitecture (the designer's choices)
- **The register state is an SRAM macro, not flops.** One FakeRAM2.0 `fakeram7_64x256` (64 words x 256 bits, single
port) per 8 lanes: the lanes are SIMD, every lane reads the same register index, so one 256-bit access serves eight
lanes' 32-bit reads. The macro's LEF and Liberty come with the ORFS ASAP7 platform (ABKGroup FakeRAM2.0, 7 nm,
0.70 V; area 1,517 um^2 for the 64 x 256, 365 um^2 for the 256 x 34); the placement, the wiring and the clock tree
see the macro as a real block. Its dynamic energy is NOT taken from the FakeRAM Liberty (a placeholder, 1.345 per
clock edge identical for every size, "VALUES NOT REALISTIC" in the generator's own config); section 2.3 replaces it
with a published macro energy band and the gate-level simulation supplies the exact access counts.
- **The instruction memory is two `fakeram7_256x34` macros** shared by every lane of the core (40 bits used of 68).
- **A single-port macro is time-multiplexed over a 5-phase slot:** read dst, read src, read src2 only when the op
needs it (mad, lop3, sel, mm8, shfla), execute, write back. A core retires LANES lane-ops per 5 cycles; throughput is
bought with lane count (die area), the cheapest resource the chip has, not with ports.
- **Every unit's operands are isolated** (AND-gated by its own select), so an unused unit does not toggle: adding a
family's unit costs leakage and a wider result mux, not switching on every op. The k lane's core evaluates every unit
every cycle; this is one of the reasons its figure is not a lower bound.
- **The era's draws are registers** (fold constants, program length, the select tree); a family epoch writes them.
- Clock 1,500 ps (667 MHz) at the TC corner, as the k lane; nothing pipelined beyond the slot; two builds: 8 lanes
(placed and routed) and 32 lanes (synthesised), each in two variants: `full` (every bank entry) and `base` (the 10
genesis families with the load and the fold, the same microarchitecture): the difference between the two is what the
bank adds to the chip.
### 1.3 What the card pays for the same instruction (the GPU side, measured)
The 5090's pJ per counted op at stock and at the 1,300 MHz lock (15.1a): add-class 11.3 / 6.2, mul and mad 13.9 / 8.3,
mulhi 39.6 / 21.0, prmt 22.3 / 11.5, lop3 24.1 / 13.0, shfl 55.8 / 29.4, the u8 tile 4.1 / 2.2 per MAC. The reserve
families by the measured NVIDIA step-cost ratio to the add step (design 4.2): shfla 1.53, popc 1.50, clz 1.63, bfe
1.54, shl and shr 0.75, sel and andn about 1.0 (approximate), mm8 by the tile row (32 MACs per lane per instruction).
## 2. Method
### 2.1 The flow
ORFS on ASAP7 (7.5-track RVT, TC corner 0.70 V, NLDM), the default flow: Yosys with ABC, floorplan at 40 percent
utilisation (30 for the 32-lane core) with the macros placed by the flow's macro placer under the BLOCKS power grid,
global and detailed placement, CTS, global and detailed routing, OpenRCX parasitics. Power is OpenSTA `report_power`
under the VCD of a random-input gate-level simulation of the netlist (iverilog; every instruction field drawn by
`$random`, the window initialised with random words, a random returned word on every load), with a propagated 0.5
activity as the cross-check. Synthesis-only rows (no wires, no clock tree) are marked; placed rows carry the SPEF.
### 2.2 The activity and the steady state
Each row is a tag: the op field fixed per family (`+fam=K`), or a drawn program: the class v4 draw over the 10 genesis
families (`mix`), the same with one load in 16 (`mixld`), the draw with two reserve families live at 4 points each
(`mix1`: shfla and mm8, the two dearest), every reserve family live at 4 points (`mix2`, the bank's bound, not a legal
draw), the W = 4 form (`mixw4`), the 64-instruction shape (`mix64`). Two run lengths per tag (500 and 2,000 clocks, 100
and 400 slots) bracket the 326-clock load phase (reset, the era's registers, the 256-word program, the 64-word window
init) and the run-phase power is solved from the pair, as the k lane does.
### 2.3 The SRAM macro energy (the one modelled term on the chip side)
The FakeRAM Liberty's internal power is a placeholder, so the collector removes the macros' Liberty-attributed power
(reported separately per VCD with `report_power -instances`) and adds a modelled access energy times the exact
access count from the simulation (3 reads and 1 write per slot for a three-operand op, 2 reads and 1 write otherwise,
per 8 lanes; 2 imem reads per slot per core). The band: a 64 x 256 single-port macro at a 7 nm class node 3.5 pJ per
256-bit access (2.0 to 7.0); a 256 x 34 macro 1.5 pJ per access (0.8 to 3.0). Sources: Horowitz, ISSCC 2014 (45 nm:
an 8 KB SRAM read of 64 bits 10 pJ, 32 KB 20 pJ), scaled by the bits moved and by the energy-per-bit reduction from
45 nm to a 7 nm class node (about 0.1x to 0.2x, approximate, the same generation scaling the record applies to logic);
the k lane's "2 to 4 pJ per 32-bit read, approximate" for a 4 KB imem; CACTI-class estimates for a 16 Kbit macro at
7 nm (0.01 to 0.03 pJ per bit read, approximate). Every row carries the band; the low end is near a flop array with
perfect clock gating, the high end a conservative compiler macro. The macros' switching on their output nets (256
bits into the lanes' latches) stays in the logic figure, measured.
### 2.4 Node scaling (claimed) and what the method leaves out
ASAP7 is a predictive 7 nm-class PDK; the row is stated at ASAP7 and scaled by the foundry's headline per-node
power reductions at the same speed (the k lane's factors, every one claimed): N5 = 0.70, N3 = 0.50, N2 = 0.36 of
ASAP7. Node-for-node against the 5090 (TSMC 4N, N5 class) is the N5 column; a node ahead is N3. Left out on the chip
side: the memory controller's queueing logic and the lane's address output (priced in the board model as the
controller die), test and clock distribution beyond the block; on the card side the 15.1a figure is the whole card's
marginal per counted op, which includes fetch, decode, operand collection and the register file, so the comparison
is the chip's whole lane (fetch, decode, window, units, network) against the card's whole lane.
## 3. The rows: pJ per lane-op per family on the base core (synthesis only, 8 lanes)
Synthesis only (no wires, no clock tree), 8 lanes, 1,500 ps, ASAP7 TC; "pJ logic" is OpenSTA's figure for the
standard cells under the VCD with the FakeRAM placeholder removed (its sequential and combinational parts beside it);
"pJ SRAM" the modelled macro term at the simulated access count (low / nominal / high, section 2.3); the per-op
figure is logic plus the nominal SRAM term, the band in brackets; the 5090 column is 15.1a (the reserve families
by the measured step ratio; the mix rows against the card's 10.3 pJ per op on the class v4 draw at the lock, 18.8
unlocked by the ARX ratio, approximate). The full core: 95,678 cells and 3 macros (the base core 66,973 and 3
macros: the bank adds 43 percent of the standard cells, 0.13 mW of leakage per 8 lanes, and 11 percent to the
energy of the class v4 draw on the same microarchitecture, the wider result mux and the leakage of the idle units).
The full core (every bank entry):
| Design | Stage | Family | Cells | Logic W | Leak W | pJ logic (seq / comb) | pJ SRAM (low / nom / high) | pJ/lane-op ASAP7 | N5 | N3 | N2 | 5090 pJ/op unlocked / lock | k N5 lock | k N3 lock | k N3 unlocked |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| mf8full | power_synth | mix | 95678 | 0.00514 | 0.000395 | 4.82 (0.88 / 3.6) | 0.99 / 1.76 / 3.52 | 6.58 (5.81 to 8.34) | 4.61 | 3.32 | 2.39 | 18.8 / 10.3 | 0.447 | 0.322 | 0.176 |
| mf8full | power_synth | mixld | 95678 | 0.00513 | 0.000395 | 4.81 (0.88 / 3.6) | 0.986 / 1.75 / 3.5 | 6.56 (5.79 to 8.31) | 4.59 | 3.31 | 2.38 | 18.8 / 10.3 | 0.446 | 0.321 | 0.176 |
| mf8full | power_synth | mix1 | 95678 | 0.00495 | 0.000395 | 4.64 (0.89 / 3.4) | 1.01 / 1.8 / 3.59 | 6.43 (5.65 to 8.23) | 4.5 | 3.24 | 2.33 | 18.8 / 10.3 | 0.437 | 0.315 | 0.172 |
| mf8full | power_synth | mix2 | 95678 | 0.00436 | 0.000396 | 4.09 (0.88 / 2.8) | 1 / 1.78 / 3.56 | 5.87 (5.09 to 7.65) | 4.11 | 2.96 | 2.13 | 18.8 / 10.3 | 0.399 | 0.287 | 0.157 |
| mf8full | power_synth | mixw4 | 95678 | 0.00508 | 0.000395 | 4.76 (0.88 / 3.5) | 0.977 / 1.73 / 3.47 | 6.49 (5.74 to 8.23) | 4.55 | 3.27 | 2.36 | 18.8 / 10.3 | 0.441 | 0.318 | 0.174 |
| mf8full | power_synth | mix64 | 95678 | 0.00488 | 0.000395 | 4.58 (0.88 / 3.3) | 0.993 / 1.76 / 3.53 | 6.34 (5.57 to 8.1) | 4.44 | 3.19 | 2.3 | 18.8 / 10.3 | 0.431 | 0.31 | 0.17 |
| mf8full | power_synth | add | 95678 | 0.003 | 0.000396 | 2.81 (0.86 / 1.6) | 0.95 / 1.69 / 3.38 | 4.5 (3.76 to 6.18) | 3.15 | 2.27 | 1.63 | 11.3 / 6.2 | 0.508 | 0.366 | 0.201 |
| mf8full | power_synth | sub | 95678 | 0.00299 | 0.000396 | 2.8 (0.86 / 1.6) | 0.95 / 1.69 / 3.38 | 4.49 (3.75 to 6.18) | 3.14 | 2.26 | 1.63 | 11.3 / 6.2 | 0.507 | 0.365 | 0.2 |
| mf8full | power_synth | xor | 95678 | 0.00308 | 0.000396 | 2.89 (0.86 / 1.7) | 0.95 / 1.69 / 3.38 | 4.57 (3.84 to 6.26) | 3.2 | 2.3 | 1.66 | 11.3 / 6.2 | 0.516 | 0.372 | 0.204 |
| mf8full | power_synth | or | 95678 | 0.00183 | 0.000397 | 1.72 (0.84 / 0.5) | 0.95 / 1.69 / 3.38 | 3.41 (2.67 to 5.09) | 2.38 | 1.72 | 1.24 | 11.3 / 6.2 | 0.385 | 0.277 | 0.152 |
| mf8full | power_synth | rotl | 95678 | 0.00318 | 0.000396 | 2.99 (0.86 / 1.8) | 0.95 / 1.69 / 3.38 | 4.67 (3.94 to 6.36) | 3.27 | 2.36 | 1.7 | 11.3 / 6.2 | 0.528 | 0.38 | 0.208 |
| mf8full | power_synth | rotr | 95678 | 0.003 | 0.000396 | 2.81 (0.86 / 1.6) | 0.95 / 1.69 / 3.38 | 4.5 (3.76 to 6.18) | 3.15 | 2.27 | 1.63 | 11.3 / 6.2 | 0.508 | 0.366 | 0.201 |
| mf8full | power_synth | mul | 95678 | 0.00265 | 0.000396 | 2.48 (0.85 / 1.3) | 0.95 / 1.69 / 3.38 | 4.17 (3.43 to 5.86) | 2.92 | 2.1 | 1.51 | 13.9 / 8.3 | 0.352 | 0.253 | 0.151 |
| mf8full | power_synth | mulhi | 95678 | 0.00249 | 0.000396 | 2.33 (0.85 / 1.1) | 0.95 / 1.69 / 3.38 | 4.02 (3.28 to 5.71) | 2.81 | 2.03 | 1.46 | 39.6 / 21 | 0.134 | 0.0965 | 0.0512 |
| mf8full | power_synth | mad | 95678 | 0.00562 | 0.000395 | 5.26 (0.91 / 4) | 1.2 / 2.12 / 4.25 | 7.39 (6.46 to 9.51) | 5.17 | 3.72 | 2.68 | 13.9 / 8.3 | 0.623 | 0.449 | 0.268 |
| mf8full | power_synth | shfl | 95678 | 0.00266 | 0.000397 | 2.49 (0.86 / 1.3) | 0.95 / 1.69 / 3.38 | 4.18 (3.44 to 5.87) | 2.93 | 2.11 | 1.52 | 55.8 / 29.4 | 0.0995 | 0.0716 | 0.0377 |
| mf8full | power_synth | load | 95678 | 0.00433 | 0.000395 | 4.06 (0.86 / 2.8) | 0.95 / 1.69 / 3.38 | 5.75 (5.01 to 7.44) | 4.03 | 2.9 | 2.09 | 13.9 / 8.3 | 0.485 | 0.349 | 0.209 |
| mf8full | power_synth | fwd | 95678 | 0.004 | 0.000395 | 3.75 (0.86 / 2.5) | 0.95 / 1.69 / 3.38 | 5.44 (4.7 to 7.13) | 3.81 | 2.74 | 1.97 | 13.9 / 8.3 | 0.459 | 0.33 | 0.197 |
| mf8full | power_synth | prmt | 95678 | 0.00255 | 0.000397 | 2.39 (0.86 / 1.2) | 0.95 / 1.69 / 3.38 | 4.07 (3.34 to 5.76) | 2.85 | 2.05 | 1.48 | 22.3 / 11.5 | 0.248 | 0.179 | 0.0921 |
| mf8full | power_synth | lop3 | 95678 | 0.003 | 0.000397 | 2.81 (0.91 / 1.5) | 1.2 / 2.12 / 4.25 | 4.94 (4.01 to 7.06) | 3.46 | 2.49 | 1.79 | 24.1 / 13 | 0.266 | 0.191 | 0.103 |
| mf8full | power_synth | shfla | 95678 | 0.00302 | 0.000397 | 2.83 (0.9 / 1.6) | 1.2 / 2.12 / 4.25 | 4.95 (4.03 to 7.08) | 3.47 | 2.5 | 1.8 | 17.3 / 9.49 | 0.365 | 0.263 | 0.144 |
| mf8full | power_synth | popc | 95678 | 0.0017 | 0.000397 | 1.59 (0.85 / 0.38) | 0.95 / 1.69 / 3.38 | 3.28 (2.54 to 4.97) | 2.3 | 1.65 | 1.19 | 17 / 9.3 | 0.247 | 0.178 | 0.0976 |
| mf8full | power_synth | clz | 95678 | 0.00173 | 0.000397 | 1.62 (0.85 / 0.41) | 0.95 / 1.69 / 3.38 | 3.31 (2.57 to 5) | 2.32 | 1.67 | 1.2 | 18.4 / 10.1 | 0.229 | 0.165 | 0.0906 |
| mf8full | power_synth | bfe | 95678 | 0.00182 | 0.000397 | 1.71 (0.85 / 0.49) | 0.95 / 1.69 / 3.38 | 3.39 (2.66 to 5.08) | 2.38 | 1.71 | 1.23 | 17.4 / 9.55 | 0.249 | 0.179 | 0.0983 |
| mf8full | power_synth | shl | 95678 | 0.00234 | 0.000397 | 2.19 (0.86 / 0.98) | 0.95 / 1.69 / 3.38 | 3.88 (3.14 to 5.57) | 2.71 | 1.95 | 1.41 | 8.48 / 4.65 | 0.584 | 0.42 | 0.231 |
| mf8full | power_synth | shr | 95678 | 0.00221 | 0.000397 | 2.07 (0.85 / 0.84) | 0.95 / 1.69 / 3.38 | 3.76 (3.02 to 5.45) | 2.63 | 1.89 | 1.36 | 8.48 / 4.65 | 0.566 | 0.407 | 0.224 |
| mf8full | power_synth | sel | 95678 | 0.0028 | 0.000397 | 2.63 (0.91 / 1.4) | 1.2 / 2.12 / 4.25 | 4.75 (3.83 to 6.88) | 3.33 | 2.4 | 1.72 | 11.3 / 6.2 | 0.537 | 0.386 | 0.212 |
| mf8full | power_synth | andn | 95678 | 0.00234 | 0.000397 | 2.19 (0.86 / 0.98) | 0.95 / 1.69 / 3.38 | 3.88 (3.14 to 5.57) | 2.72 | 1.96 | 1.41 | 11.3 / 6.2 | 0.438 | 0.315 | 0.173 |
| mf8full | power_synth | mm8 | 95678 | 0.0036 | 0.000396 | 3.37 (0.91 / 2.1) | 1.2 / 2.12 / 4.25 | 5.5 (4.57 to 7.62) | 3.85 | 2.77 | 2 | 16.4 / 8.8 | 0.437 | 0.315 | 0.169 |
The base core (the 10 genesis families on the same microarchitecture; the comparator):
| Design | Stage | Family | Cells | Logic W | Leak W | pJ logic (seq / comb) | pJ SRAM (low / nom / high) | pJ/lane-op ASAP7 | N5 | N3 | N2 | 5090 pJ/op unlocked / lock | k N5 lock | k N3 lock | k N3 unlocked |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| mf8base | power_synth | mix | 66973 | 0.00443 | 0.000263 | 4.15 (0.88 / 3) | 0.99 / 1.76 / 3.52 | 5.91 (5.14 to 7.66) | 4.13 | 2.98 | 2.14 | 18.8 / 10.3 | 0.401 | 0.289 | 0.158 |
| mf8base | power_synth | mixld | 66973 | 0.00439 | 0.000263 | 4.12 (0.88 / 3) | 0.986 / 1.75 / 3.5 | 5.87 (5.1 to 7.62) | 4.11 | 2.96 | 2.13 | 18.8 / 10.3 | 0.399 | 0.287 | 0.157 |
| mf8base | power_synth | add | 66973 | 0.00246 | 0.000264 | 2.31 (0.86 / 1.2) | 0.95 / 1.69 / 3.38 | 4 (3.26 to 5.68) | 2.8 | 2.01 | 1.45 | 11.3 / 6.2 | 0.451 | 0.325 | 0.178 |
| mf8base | power_synth | sub | 66973 | 0.00245 | 0.000264 | 2.29 (0.86 / 1.2) | 0.95 / 1.69 / 3.38 | 3.98 (3.24 to 5.67) | 2.79 | 2.01 | 1.44 | 11.3 / 6.2 | 0.449 | 0.324 | 0.178 |
| mf8base | power_synth | xor | 66973 | 0.00247 | 0.000264 | 2.32 (0.86 / 1.2) | 0.95 / 1.69 / 3.38 | 4 (3.27 to 5.69) | 2.8 | 2.02 | 1.45 | 11.3 / 6.2 | 0.452 | 0.326 | 0.179 |
| mf8base | power_synth | or | 66973 | 0.0016 | 0.000265 | 1.5 (0.84 / 0.42) | 0.95 / 1.69 / 3.38 | 3.19 (2.45 to 4.87) | 2.23 | 1.61 | 1.16 | 11.3 / 6.2 | 0.36 | 0.259 | 0.142 |
| mf8base | power_synth | rotl | 66973 | 0.00259 | 0.000264 | 2.43 (0.86 / 1.3) | 0.95 / 1.69 / 3.38 | 4.11 (3.38 to 5.8) | 2.88 | 2.07 | 1.49 | 11.3 / 6.2 | 0.464 | 0.334 | 0.183 |
| mf8base | power_synth | rotr | 66973 | 0.00251 | 0.000264 | 2.35 (0.86 / 1.3) | 0.95 / 1.69 / 3.38 | 4.04 (3.3 to 5.73) | 2.83 | 2.04 | 1.47 | 11.3 / 6.2 | 0.456 | 0.328 | 0.18 |
| mf8base | power_synth | mul | 66973 | 0.00229 | 0.000264 | 2.14 (0.85 / 1.1) | 0.95 / 1.69 / 3.38 | 3.83 (3.09 to 5.52) | 2.68 | 1.93 | 1.39 | 13.9 / 8.3 | 0.323 | 0.233 | 0.139 |
| mf8base | power_synth | mulhi | 66973 | 0.00217 | 0.000264 | 2.03 (0.84 / 0.96) | 0.95 / 1.69 / 3.38 | 3.72 (2.98 to 5.41) | 2.61 | 1.88 | 1.35 | 39.6 / 21 | 0.124 | 0.0893 | 0.0474 |
| mf8base | power_synth | mad | 66973 | 0.00472 | 0.000263 | 4.43 (0.9 / 3.3) | 1.2 / 2.12 / 4.25 | 6.55 (5.62 to 8.67) | 4.58 | 3.3 | 2.38 | 13.9 / 8.3 | 0.552 | 0.398 | 0.237 |
| mf8base | power_synth | shfl | 66973 | 0.00192 | 0.000265 | 1.8 (0.86 / 0.7) | 0.95 / 1.69 / 3.38 | 3.49 (2.75 to 5.17) | 2.44 | 1.76 | 1.27 | 55.8 / 29.4 | 0.083 | 0.0598 | 0.0315 |
| mf8base | power_synth | load | 66973 | 0.00359 | 0.000263 | 3.36 (0.86 / 2.3) | 0.95 / 1.69 / 3.38 | 5.05 (4.31 to 6.74) | 3.53 | 2.54 | 1.83 | 13.9 / 8.3 | 0.426 | 0.307 | 0.183 |
Reading the rows. (1) The ALU-group families cost the chip 3.1 to 3.3 pJ per lane-op at N5 against the card's 6.2
at the lock (k 0.51 to 0.53); or, popc, clz, bfe, prmt 2.3 to 2.9 (k 0.23 to 0.38); the multiply 2.9 (k 0.35);
mulhi 2.8 against the card's 21 (k 0.13); the shuffle 2.9 against 29.4 (k 0.10); mad is the dearest op for the
chip at 5.2 (three reads and two units, k 0.62) and the shifters the highest k (0.57 to 0.58) because the card
does them cheapest. (2) The SRAM term is 1.7 pJ nominal per lane-op (0.95 to 3.4): a third of the row; the
sequential term (the phase latches, the instruction register, the counters at five edges per op) 0.86 pJ; the rest
is the units and the macro output nets. (3) The reserve families are cheaper for the chip than the genesis
families: every live-family mix sits under the class v4 draw, and the bound with all eight live reads 11 percent
under it, because the card's dearest instructions (mulhi, shfl, mad) are the genesis ones.
## 4. The whole-hash energy and k per family, node-for-node and a node ahead
The whole-hash shadow is 102,612 chip ops (the 102,100 counted shadow ops of the record plus the 512-instruction
base program; W = 4 adds 384 fold steps). The chip's cost is absolute (its own pJ at its node); the card's premium
on the same draw is the measured 0.652 microjoules at the lock, moved by the live family's measured step ratio at 4
points of 79 (modelled). Node-for-node is N5 (the 5090's own class), a node ahead N3; every factor claimed.
design mf8full stage power_synth; the class v4 draw on the core: 6.58 pJ per lane-op ASAP7, 4.61 N5, 3.32 N3 (band 4.07 to 5.84 at N5)
| Family live (4 points of 79) or draw | Chip pJ per op N5 / N3 (band) | 5090 pJ per op at the lock | k N5 / N3 | Chip shadow per hash, microjoules N5 / N3 (the mix with the family live) | Change vs the class v4 draw | The card's premium per hash at the lock (modelled from the step ratio) | Change |
|---|---|---|---|---|---|---|---|
| the class v4 draw (10 genesis families) | 4.61 / 3.32 (4.07 to 5.84) | 10.3 | 0.45 / 0.32 | 0.473 / 0.340 | +0.0 percent | 0.652 | +0.0 percent |
| add (unit row, every instruction) | 3.15 / 2.27 (2.63 to 4.33) | 6.2 | 0.51 / 0.37 | 0.323 / 0.233 | -31.7 percent | 0.392 | -39.8 percent |
| sub (unit row, every instruction) | 3.14 / 2.26 (2.63 to 4.32) | 6.2 | 0.51 / 0.36 | 0.322 / 0.232 | -31.8 percent | 0.392 | -39.8 percent |
| xor (unit row, every instruction) | 3.20 / 2.30 (2.68 to 4.38) | 6.2 | 0.52 / 0.37 | 0.328 / 0.236 | -30.5 percent | 0.392 | -39.8 percent |
| or (unit row, every instruction) | 2.38 / 1.72 (1.87 to 3.57) | 6.2 | 0.38 / 0.28 | 0.245 / 0.176 | -48.2 percent | 0.392 | -39.8 percent |
| rotl (unit row, every instruction) | 3.27 / 2.36 (2.75 to 4.45) | 6.2 | 0.53 / 0.38 | 0.336 / 0.242 | -29.0 percent | 0.392 | -39.8 percent |
| rotr (unit row, every instruction) | 3.15 / 2.27 (2.63 to 4.33) | 6.2 | 0.51 / 0.37 | 0.323 / 0.233 | -31.7 percent | 0.392 | -39.8 percent |
| mul (unit row, every instruction) | 2.92 / 2.10 (2.40 to 4.10) | 8.3 | 0.35 / 0.25 | 0.299 / 0.216 | -36.7 percent | 0.525 | -19.4 percent |
| mulhi (unit row, every instruction) | 2.81 / 2.03 (2.30 to 3.99) | 21.0 | 0.13 / 0.10 | 0.289 / 0.208 | -38.9 percent | 1.329 | +103.9 percent |
| mad (unit row, every instruction) | 5.17 / 3.72 (4.52 to 6.66) | 8.3 | 0.62 / 0.45 | 0.531 / 0.382 | +12.3 percent | 0.525 | -19.4 percent |
| shfl (unit row, every instruction) | 2.93 / 2.11 (2.41 to 4.11) | 29.4 | 0.10 / 0.07 | 0.300 / 0.216 | -36.5 percent | 1.861 | +185.4 percent |
| load (unit row, every instruction) | 4.03 / 2.90 (3.51 to 5.21) | 8.3 | 0.48 / 0.35 | 0.413 / 0.297 | -12.6 percent | 0.525 | -19.4 percent |
| prmt live | 2.85 / 2.05 (2.34 to 4.03) | 11.5 | 0.25 / 0.18 | 0.464 / 0.334 | -1.9 percent | 0.662 | +1.5 percent |
| lop3 live | 3.46 / 2.49 (2.81 to 4.94) | 13.0 | 0.27 / 0.19 | 0.467 / 0.336 | -1.3 percent | 0.688 | +5.6 percent |
| shfla live | 3.47 / 2.50 (2.82 to 4.95) | 9.5 | 0.37 / 0.26 | 0.467 / 0.336 | -1.3 percent | 0.669 | +2.7 percent |
| popc live | 2.30 / 1.65 (1.78 to 3.48) | 9.3 | 0.25 / 0.18 | 0.461 / 0.332 | -2.5 percent | 0.669 | +2.5 percent |
| clz live | 2.32 / 1.67 (1.80 to 3.50) | 10.1 | 0.23 / 0.17 | 0.461 / 0.332 | -2.5 percent | 0.673 | +3.2 percent |
| bfe live | 2.38 / 1.71 (1.86 to 3.56) | 9.5 | 0.25 / 0.18 | 0.461 / 0.332 | -2.5 percent | 0.670 | +2.7 percent |
| shl live | 2.71 / 1.95 (2.20 to 3.90) | 4.7 | 0.58 / 0.42 | 0.463 / 0.333 | -2.1 percent | 0.644 | -1.3 percent |
| shr live | 2.63 / 1.89 (2.12 to 3.81) | 4.7 | 0.57 / 0.41 | 0.462 / 0.333 | -2.2 percent | 0.644 | -1.3 percent |
| sel live | 3.33 / 2.40 (2.68 to 4.81) | 6.2 | 0.54 / 0.39 | 0.466 / 0.336 | -1.4 percent | 0.652 | +0.0 percent |
| andn live | 2.72 / 1.96 (2.20 to 3.90) | 6.2 | 0.44 / 0.32 | 0.463 / 0.333 | -2.1 percent | 0.652 | +0.0 percent |
| mm8 live | 3.85 / 2.77 (3.20 to 5.34) | 8.8 | 0.44 / 0.31 | 0.469 / 0.337 | -0.8 percent | 0.699 | +7.2 percent |
| fwd (unit row, every instruction) | 3.81 / 2.74 (3.29 to 4.99) | 8.3 | 0.46 / 0.33 | 0.391 / 0.281 | -17.3 percent | 0.525 | -19.4 percent |
| the draw with shfla and mm8 live (measured mix) | 4.50 / 3.24 (3.96 to 5.76) | 10.3 | 0.44 / 0.31 | 0.462 / 0.333 | -2.2 percent | 0.652 | +0.0 percent |
| every reserve family live at 4 points (the bound, measured mix) | 4.11 / 2.96 (3.56 to 5.35) | 10.3 | 0.40 / 0.29 | 0.422 / 0.304 | -10.8 percent | 0.652 | +0.0 percent |
| the draw with one load in 16 | 4.59 / 3.31 (4.06 to 5.82) | 10.3 | 0.45 / 0.32 | 0.471 / 0.339 | -0.3 percent | 0.652 | +0.0 percent |
| W = 4: three fold steps per load | 4.55 / 3.27 (4.02 to 5.76) | 10.3 | 0.44 / 0.32 | 0.466 / 0.336 | -1.3 percent | 0.652 | +0.0 percent |
| the 64-instruction shape | 4.44 / 3.19 (3.90 to 5.67) | 10.3 | 0.43 / 0.31 | 0.455 / 0.328 | -3.7 percent | 0.652 | +0.0 percent |
The whole-hash edge per joule with the class v4 shadow on the core (absolute: the chip pays its own pJ whatever the card does):
| Memory (E_mem, microjoules) | Chip E_hash N5 / N3 | vs 5090 lock 2.33 | vs 5090 stock 3.36 | vs 5080 lock 2.06 | vs M5 Max 1.40 |
|---|---|---|---|---|---|
| GDDR7 board (0.466) | 0.939 / 0.806 | 2.5x / 2.9x | 3.6x / 4.2x | 2.2x / 2.6x | 1.5x / 1.7x |
| HBM3 one stack (0.321) | 0.794 / 0.661 | 2.9x / 3.5x | 4.2x / 5.1x | 2.6x / 3.1x | 1.8x / 2.1x |
| SRAM N2 die at W = 1 (0.036) | 0.509 / 0.376 | 4.6x / 6.2x | 6.6x / 8.9x | 4.1x / 5.5x | 2.8x / 3.7x |
The same edge on the base core (the genesis-only comparator): 2.6x / 3.0x microjoules per hash
on the GDDR7 board at N5 / N3, 3.8x / 4.4x against the 5090 at its lock; the bank costs the
adversary 5 percent of its edge (0.939 against 0.890 microjoules on the GDDR7 board, 2.5x against 2.6x), which is
the whole answer to the review's question in one number: the 180-day calendar costs a chip that carries the bank
about 11 percent of its shadow energy and 43 percent of its core cells, and no epoch costs it a part.
## 5. The placed rows (8 lanes, routed, SPEF) and the 32-lane core
ROWS_PLACED
## 6. The board: joules per valid hash and USD per sustained MH/s for the complete machine
The complete machine: the memory devices at their modelled random-read energy and activate ceiling (chip-model-v3
5.3, unmeasured; the 5090 reaches 82 percent of the GDDR7 figure, the sustained fraction here), the controller and
PHY die, the core die sized to retire the hash's ops at the memory's sustained rate (lanes at 667 MHz over 5 phases;
0.002 mm^2 per lane at N5 from the 8-lane core's floorplan at 40 percent utilisation scaled x0.55, approximate; USD
0.36 per mm^2 of N5 from sram-mirror's yield model; 50 uW of leakage per lane, the synthesis figure), power delivery
(PSU 92 percent, VRM 90 percent), cooling (3 percent), a board and assembly (USD 200), and one full node per 100
machines (85 W and USD 1,500 shared: the state-derived dataset's host, priced as the review's rule 3 asks). The
"high" case takes the memory's lower ceiling (the 5090's measured 17.5 G on GDDR7; the JEDEC tFAW floor on HBM3),
the high read energy and the high SRAM term together. GPU rows: the 5090 at its lock 2.33 microjoules at 134.76
MH/s, USD 1,999 plus USD 150 of rig share (USD 15.9 per MH/s; at the USD 3,000 street price 23.4); the 5080 at its
lock 2.06 at 71.20 MH/s, USD 999 plus 150 (USD 16.1 per MH/s); the discrete-GPU cohort by count (design 3.4 and
10.4: 8 GB 22 percent, 12 GB 22, 16 GB 28, 24 GB and up 16, the rest 10 and 11 GB) at its tuned points about 3.6
microjoules (2.5 to 4.5, approximate: the 5070 and 5070 Ti 1.7 to 1.75 modelled, the 4070 3.58 measured, the 4090
3.64, the 3090 6.4, the 9070 XT 8.1 measured) and about USD 18 per MH/s (14 to 28); the M5 Max 1.40 at 27.9 MH/s is
reported in section 4, not headlined.
Node-for-node (N5), the full core at 4.61 pJ per lane-op:
core 4.61 pJ per lane-op (band 4.07 to 5.84), 102612 ops per hash, leakage 50 uW per lane, 0.002 mm^2 per lane at N5
| Memory arrangement | Case | Sustained MH/s | Machine W | microjoules per hash (whole machine) | Core lanes | Core mm^2 (N5) | Capex USD | USD per MH/s | vs 5090 lock 2.33 (J / USD) | vs 5080 lock 2.06 | vs cohort 3.6 / USD 18 |
|---|---|---|---|---|---|---|---|---|---|---|---|
| GDDR7 board, 16 devices, 64 channels (the 5090 memory without the GPU) | nominal | 136 | 175 | 1.280 | 104,960 | 210 | 661 | 4.84 | 1.8x / 3.3x | 1.6x / 3.3x | 2.8x / 3.7x |
| GDDR7 board, 16 devices, 64 channels (the 5090 memory without the GPU) | low | 136 | 154 | 1.132 | 104,960 | 210 | 661 | 4.84 | 2.1x / 3.3x | 1.8x / 3.3x | 3.2x / 3.7x |
| GDDR7 board, 16 devices, 64 channels (the 5090 memory without the GPU) | high | 112 | 180 | 1.603 | 86,235 | 172 | 647 | 5.77 | 1.5x / 2.8x | 1.3x / 2.8x | 2.2x / 3.1x |
| HBM3 one stack (activate ceiling unmeasured; JEDEC tFAW floor 2.3 G) | nominal | 71 | 77 | 1.084 | 54,656 | 109 | 704 | 9.91 | 2.1x / 1.6x | 1.9x / 1.6x | 3.3x / 1.8x |
| HBM3 one stack (activate ceiling unmeasured; JEDEC tFAW floor 2.3 G) | low | 71 | 70 | 0.984 | 54,656 | 109 | 704 | 9.91 | 2.4x / 1.6x | 2.1x / 1.6x | 3.7x / 1.8x |
| HBM3 one stack (activate ceiling unmeasured; JEDEC tFAW floor 2.3 G) | high | 15 | 34 | 2.228 | 11,748 | 23 | 673 | 44.09 | 1.0x / 0.4x | 0.9x / 0.4x | 1.6x / 0.4x |
| HBM3 eight stacks (an H100-class package) | nominal | 566 | 559 | 0.987 | 435,713 | 871 | 3,079 | 5.44 | 2.4x / 2.9x | 2.1x / 3.0x | 3.6x / 3.3x |
| HBM3 eight stacks (an H100-class package) | low | 566 | 502 | 0.886 | 435,713 | 871 | 3,079 | 5.44 | 2.6x / 2.9x | 2.3x / 3.0x | 4.1x / 3.3x |
| HBM3 eight stacks (an H100-class package) | high | 122 | 217 | 1.772 | 93,987 | 188 | 2,833 | 23.18 | 1.3x / 0.7x | 1.2x / 0.7x | 2.0x / 0.8x |
| SRAM full store, one N2 reticle, 2 GiB at W = 1 (the strongest five-year chip) | nominal | 535 | 400 | 0.747 | 411,225 | 822 | 1,211 | 2.27 | 3.1x / 7.0x | 2.8x / 7.1x | 4.8x / 7.9x |
| SRAM full store, one N2 reticle, 2 GiB at W = 1 (the strongest five-year chip) | low | 609 | 403 | 0.662 | 468,572 | 937 | 1,252 | 2.06 | 3.5x / 7.8x | 3.1x / 7.8x | 5.4x / 8.8x |
| SRAM full store, one N2 reticle, 2 GiB at W = 1 (the strongest five-year chip) | high | 419 | 394 | 0.940 | 322,466 | 645 | 1,147 | 2.74 | 2.5x / 5.8x | 2.2x / 5.9x | 3.8x / 6.6x |
Lifetime cost in USD per TH (10^12 hashes), capex spread over the life plus electricity at USD 0.08 per kWh:
| Machine | 0.5 y | 1 y | 2 y | 3 y | of which electricity |
|---|---|---|---|---|---|
| 5090 at the lock | 1.063 | 0.557 | 0.305 | 0.220 | 0.0518 |
| 5080 at the lock | 1.069 | 0.557 | 0.302 | 0.216 | 0.0458 |
| cohort card | 1.222 | 0.651 | 0.365 | 0.270 | 0.0800 |
| GDDR7 board, 16 devices, 64 channels chip | 0.335 | 0.182 | 0.105 | 0.080 | 0.0284 |
| HBM3 one stack chip | 0.653 | 0.338 | 0.181 | 0.129 | 0.0241 |
| HBM3 eight stacks chip | 0.367 | 0.194 | 0.108 | 0.079 | 0.0219 |
| SRAM full store, one N2 reticle, 2 GiB at W = 1 chip | 0.160 | 0.088 | 0.053 | 0.041 | 0.0166 |
A node ahead (N3), the same machine:
|---|---|---|---|---|---|---|---|---|---|---|---|
| GDDR7 board, 16 devices, 64 channels (the 5090 memory without the GPU) | nominal | 136 | 150 | 1.101 | 104,960 | 147 | 638 | 4.67 | 2.1x / 3.4x | 1.9x / 3.5x | 3.3x / 3.9x |
| GDDR7 board, 16 devices, 64 channels (the 5090 memory without the GPU) | low | 136 | 133 | 0.972 | 104,960 | 147 | 638 | 4.67 | 2.4x / 3.4x | 2.1x / 3.5x | 3.7x / 3.9x |
| GDDR7 board, 16 devices, 64 channels (the 5090 memory without the GPU) | high | 112 | 155 | 1.380 | 86,235 | 121 | 628 | 5.61 | 1.7x / 2.8x | 1.5x / 2.9x | 2.6x / 3.2x |
| HBM3 one stack (activate ceiling unmeasured; JEDEC tFAW floor 2.3 G) | nominal | 71 | 64 | 0.905 | 54,656 | 77 | 693 | 9.75 | 2.6x / 1.6x | 2.3x / 1.7x | 4.0x / 1.8x |
| HBM3 one stack (activate ceiling unmeasured; JEDEC tFAW floor 2.3 G) | low | 71 | 59 | 0.824 | 54,656 | 77 | 693 | 9.75 | 2.8x / 1.6x | 2.5x / 1.7x | 4.4x / 1.8x |
| HBM3 one stack (activate ceiling unmeasured; JEDEC tFAW floor 2.3 G) | high | 15 | 31 | 2.004 | 11,748 | 16 | 671 | 43.93 | 1.2x / 0.4x | 1.0x / 0.4x | 1.8x / 0.4x |
| HBM3 eight stacks (an H100-class package) | nominal | 566 | 458 | 0.808 | 435,713 | 610 | 2,985 | 5.27 | 2.9x / 3.0x | 2.5x / 3.1x | 4.5x / 3.4x |
| HBM3 eight stacks (an H100-class package) | low | 566 | 411 | 0.726 | 435,713 | 610 | 2,985 | 5.27 | 3.2x / 3.0x | 2.8x / 3.1x | 5.0x / 3.4x |
| HBM3 eight stacks (an H100-class package) | high | 122 | 189 | 1.548 | 93,987 | 132 | 2,812 | 23.02 | 1.5x / 0.7x | 1.3x / 0.7x | 2.3x / 0.8x |
| SRAM full store, one N2 reticle, 2 GiB at W = 1 (the strongest five-year chip) | nominal | 724 | 398 | 0.550 | 557,288 | 780 | 1,196 | 1.65 | 4.2x / 9.7x | 3.7x / 9.8x | 6.5x / 10.9x |
| SRAM full store, one N2 reticle, 2 GiB at W = 1 (the strongest five-year chip) | low | 828 | 402 | 0.485 | 636,578 | 891 | 1,236 | 1.49 | 4.8x / 10.7x | 4.2x / 10.8x | 7.4x / 12.1x |
Reading. (1) The complete GDDR7 machine reads 1.8x the 5090 at its lock per joule node-for-node (1.5x to 2.1x) and
2.1x a node ahead (1.7x to 2.4x); against the 5080 at its lock 1.6x (1.3x to 1.8x) and 1.9x; against the cohort
card 2.8x and 3.3x. Per dollar of capex it is 3.3x the 5090 at MSRP (4.8x at the street price) and 3.7x the cohort,
because the board carries the same USD 320 of memory and USD 76 of core silicon where the card carries a 750 mm^2
GPU. (2) The HBM3 one-stack machine is no better per joule than GDDR7 once the controller, the host and the power
train are in, and its dollars per MH/s are twice the card's at the modelled ceiling and 2.7x worse than the card at
the JEDEC tFAW floor: the "same DRAM" assumption is tested and the adversary's cheapest memory IS the card's memory,
so the GDDR7 board is the machine the economics must answer. (3) The N2 SRAM die at the hash's own width is the
strongest machine at 3.1x per joule node-for-node (2.5x to 3.5x) and 4.2x a node ahead, and 7x per dollar of
silicon, with the project cost (USD 100 M to 500 M, claimed) as its only hold. (4) Over a 3-year life the GDDR7
machine's cost per TH is 0.080 USD against the 5090's 0.220 and the cohort's 0.270 (2.7x to 3.4x); at a 1-year life
3.1x; electricity is a quarter of the machine's lifetime cost and a fifth of the card's, so the per-joule edge is
the smaller half of the economic edge and the capex per MH/s the larger. The profitability surface the review asks
for (development, fleet capital, share, margin, electricity, pre-production, discounting, residual, operator and
manufacturer apart) takes these per-TH rows as its inputs; it is the economics lane's, not this file's.
## 7. The lifetime answer: what each transition needs, and what is credited
The rule: a transition is credited only where the next entry needs a physical resource this design cannot supply
economically; firmware changes (a program, a register, a weight) are credited zero. "Performance loss" is the change in
the chip's energy per hash when the entry is live at its draw weight (4 points of 79 for a reserve family; the whole
program for an atom or a shape), from the rows of sections 3 and 4; the GPU's own change on the same draw is beside it.
| Transition (the next epoch draws it) | Physical resource it needs | Does this design hold it? | Chip loss at the draw weight | The card's change on the same draw (measured ratio) | Credited obsolescence benefit |
|---|---|---|---|---|---|
| G1 to G7 (the ARX group) at any band weight | the ALU group | yes | 0 to +3 percent of the shadow energy across the band (the unit rows 3.1 to 3.3 pJ at N5 against the draw's 4.6) | 1.00 | 0 |
| G8, G9, G6 (mul, mulhi, mad) at any band weight | the multiplier and the third read | yes | -1 to +2 percent at B = 4 (mul 2.9, mulhi 2.8, mad 5.2 pJ at N5) | mul 1.1, mulhi about 3.4 (the card's dearest ALU op) | 0 |
| G10 shfl at its cap (8 points) | the lane butterfly | yes (per core) | 0 (2.9 pJ at N5, under the draw's mean) | 4.9x the add per op | 0 (the card pays 29.4 pJ for the move the chip pays about 1) |
| R1 shfla live (4 points) | a general lane crossbar | yes (per core; the one network a butterfly cannot emulate in one op) | 0 (2.9 pJ at N5, under the draw's mean)A | 1.53 (NVIDIA), 1.91 (Apple) | 0 |
| R2 perm live | a byte selector | yes | -1.9 percent | 1.30 | 0 |
| R3 popc and clz live | a popcount tree, a priority encoder | yes | -2.5 percent | 1.50 and 1.63 | 0 |
| R4 to R7 (bfe, shl and shr, sel, andn) live | a shifter, a mask, a select, an and-not | yes | -2.1 to -1.4 percent | 0.75 to 1.54 | 0 |
| R8 mm8 live | a u8 dot4 per lane (8 chip ops per card tile) | yes | -0.8 percent (3.9 pJ per dp4a at N5, 1.0 pJ per MAC, against the card's 2.2 pJ per MAC at the lock) | the tile: 2.43 the add step per card instruction (32 MACs) | 0 (the chip's MAC is cheaper than the card's by 4x to 30x on the public figures; this is the card's loss, not the chip's) |
| the op-mix band draw (B = 4 on injecting families) | nothing: the program | yes | within the family rows above | within 11 percent per instruction (shadow-k 6.2) | 0 |
| the fold constants draw | five registers | yes | 0 | 0 | 0 |
| the block shape draw (64, 128, 256) | the program-length register; 256 words of imem | yes | -3.7 percent at 64 (the imem term; 0 at 128 and 256) | 64 ran 2.5 to 3.5 percent faster than 256 on the 5090 and the M5 Max (measured) | 0 |
| the read-width draw W = 4 (16 bytes) | three fwd instructions per load (firmware); the same DRAM sector | yes | +0.3 percent per hash (384 fold steps at 3.8 pJ) | within 2.7 percent on the 5090 and the 9070 XT, within 1 on the M5 Max (measured) | 0 |
| the 64-register window | the 64 x 256 macro per 8 lanes | yes (built in) | 0 (it is the base) | modelled: occupancy to about half on a 5090 or 4090, the rate expected to hold (the hash lane) | 0 |
| the dataset atom draw A1 / A2 / A3 (mixer x4, x8, dr368) | nothing per hash; the per-window build is a program on the same core (section 7.1) | yes | 0 per hash; the build 0.72 J per window per machine at 1 GiB (0.2 mW averaged), 1.45 J at 2 GiB | the card's rate unmoved by the atom (within 0.1 MH/s on the 5090, measured) | 0 |
| a new op family outside the 18 (a bank refresh, a release) | a unit the die lacks | no: the chip emulates it from the 18 at the vendor penalty, as the cards do (1.5x to 2.4x per op, measured on the cards) or loses its 4 points | at 4 points of 79: at most 4 / 79 x (penalty - 1) of the shadow energy, about 2 to 6 percent (modelled) | the same emulation on every card that predates the release, 0 on a card with the native op | 0 unless the family is one the 18 cannot emulate; none proposed is |
| a new read atom W = 8 (32 bytes, `admissible: false` today) | seven fwd instructions per load; the same GDDR7 sector; on an SRAM die +0.3 nJ of wire per read (floor lane 3) | yes | +0.7 percent per hash (896 fold steps) | free by the measured rows (one sector per load on NVIDIA) | 0 |
| the dataset floor step (layer 2: 5.5 / 8.5 / 11.5 GiB) | device memory: 3 to 6 GDDR7 devices more, or 3 to 6 N2 reticles on the SRAM die | yes on DRAM (USD 60 to 120 more); the SRAM die's ticket rises USD 1,000 per step | 0 per joule on DRAM; the SRAM die's USD per MH/s unmoved (every die powered) | the tuned 5090 pays 4, 8 and 10 percent more energy per hash at 2, 4 and 8 GiB (measured) | 0 per joule; a capex ticket on the SRAM die only (floor lane 3) |
Reading, before the numbers: nothing in the bank asks for a resource the design lacks, because the bank is public at
genesis and its whole op-family set is a few adders per lane and two networks per core. What the bank does to this
chip is the per-op cost of carrying the unit set (sections 3 and 5: the full core against the base core on the same
microarchitecture) and the leakage of the units that are not live, and that is the number the credit must come from.
### 7.1 The per-window dataset build on the chip (the atoms' only cost)
Under class v5 the dataset is rebuilt from the chain's state every window (3,600 s). A 1 GiB build is 157 G ops
(measured as 13.4 ms on a 5090; the record's row 7). On this core at 4.61 pJ per lane-op (N5) that is 0.72 J per
window per machine, 0.0002 W averaged over the window, against a machine of hundreds of watts: 0.0001 percent of
the machine's energy, the same for every atom within the atom's op count (x4 half of x8, dr368 about x4's). The
chip's node is the host the board model carries (one full node per 100 machines: 85 W and USD 1,500 shared); a
specialised machine with a host keeping the dataset current is inside the adversary model by the review's rule (3),
and this is its price: under 1 W and USD 15 per machine.
## 8. Energy resistance, economic resistance and response capability, stated separately
STATEMENT
## 9. Sources and what is owed
- The GPU side: `docs/analysis/counter-asic-4-research.md` 15.1a (the 5090 microbench, 8 October 2026, measured);
the floor programme's tier table (`docs/design/class-v6-rotating-family.md` 10.4: the 5090 at the 1,300 lock 2.33
microjoules at 134.76 MH/s and 312.5 W, the 5080 at the 1,100 lock 2.06 at 71.20 MH/s and 146.6 W, stock 3.36 and
3.48, the 4070, 4090, 3090 and 9070 XT rows, the M5 Max 1.40); the reserve families' step costs, design 4.2
(measured on the cards); the card population by count, design 3.4 (approximate).
- The chip side: this lane's RTL and flow (`tools/chip-model/mf/`), ORFS on ASAP7 (Clark et al., Microelectronics
Journal 2016; the 7.5-track RVT library, TC corner), FakeRAM2.0 macros (ABKGroup, the ORFS platform's
`fakeram7_64x256` and `fakeram7_256x34`, LEF and Liberty; the generator's own config marks its values "not
realistic", so only area, pins and placement are taken from it); the k lane's rows (`floor/shadow-k.md`, the
shuffle row 1.24 pJ per lane-op routed, relayed 16:0x BST) cited as published.
- The SRAM access energy band: Horowitz, "Computing's energy problem (and what we can do about it)", ISSCC 2014
(45 nm: 8 KB SRAM 10 pJ per 64-bit read, 32 KB 20 pJ); the scaling to a 7 nm class node approximate; the k lane's
imem figure (2 to 4 pJ per 32-bit read, approximate); CACTI-class estimates (approximate).
- Node scaling (claimed): TSMC's technology pages for N5, N3E and N2, read 8 October 2026 (shadow-k section 7).
- The board: `docs/analysis/chip-model-v3.md` 5.3 and 5.5 (the GDDR7 and HBM3 random-read engines, the activate
ceilings, the static and controller allowances, the prices, all modelled or claimed; the 5090 at 82 percent of the
GDDR7 ceiling, measured); `floor/sram-and-floor.md` 2.1 (the SRAM die at W = 1, modelled); the power delivery,
cooling and host allowances are this file's (approximate; the sensitivity is stated beside each).
- The dataset build: the record's row 7 (a 1 GiB rebuild 157 G ops, 13.4 ms on a 5090, measured).
Owed: the k lane's crossbar, scratch and tile rows (in place and route at 16:0x BST); the placed 32-lane core (the
8-lane core is placed; the 32-lane row is synthesis only); a real PDK memory compiler's figure for the two macros
(FakeRAM gives area and pins only); the 64-register window's GPU cost measured (the hash lane's generator line);
the HBM3 activate ceiling (unmeasured, the AWS F2 hour); the profitability surface of the review's rule (2) over the
lifetime rows of section 6, which this file states as cost per TH and leaves the NPV to the economics lane.