multi-family adversary: the mf core RTL (SRAM window and imem, every bank entry, operand isolation, 5-phase slot), the flow, the collector, the board and hash models, the first rows (synthesis only, 8 lanes, full and base)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Documents-only replay of a807edd84 (2db208912e5fc0656930aaabd7db74ec2a68ca8e) for the box mirror master; left on the branch: tools/chip-model/mf/Makefile tools/chip-model/mf/flow/board.py tools/chip-model/mf/flow/collect.py tools/chip-model/mf/flow/gl2sim.sh tools/chip-model/mf/flow/hash.py tools/chip-model/mf/flow/mf32base.mk tools/chip-model/mf/flow/mf32base.sdc tools/chip-model/mf/flow/mf32full.mk tools/chip-model/mf/flow/mf32full.sdc tools/chip-model/mf/flow/mf8base.mk tools/chip-model/mf/flow/mf8base.sdc tools/chip-model/mf/flow/mf8full.mk tools/chip-model/mf/flow/mf8full.sdc tools/chip-model/mf/flow/power.tcl tools/chip-model/mf/flow/sim.sh tools/chip-model/mf/results/board-mf8full-n3.md tools/chip-model/mf/results/board-mf8full-n5.md tools/chip-model/mf/results/hash-mf8base.md tools/chip-model/mf/results/hash-mf8full.md tools/chip-model/mf/results/out/mf8base/logs/asap7/mf8base/base/1_synth.json tools/chip-model/mf/results/out/mf8base/logs/asap7/mf8base/base/1_synth.log tools/chip-model/mf/results/out/mf8base/logs/asap7/mf8base/base/power_synth.log tools/chip-model/mf/results/out/mf8base/objects/asap7/mf8base/base/copyright.txt tools/chip-model/mf/results/out/mf8base/reports/asap7/mf8base/base/synth_check.txt tools/chip-model/mf/results/out/mf8base/reports/asap7/mf8base/base/synth_stat.txt tools/chip-model/mf/results/out/mf8base/results/asap7/mf8base/base/mem.json tools/chip-model/mf/results/out/mf8full/logs/asap7/mf8full/base/1_synth.json tools/chip-model/mf/results/out/mf8full/logs/asap7/mf8full/base/1_synth.log tools/chip-model/mf/results/out/mf8full/logs/asap7/mf8full/base/power_synth.log tools/chip-model/mf/results/out/mf8full/objects/asap7/mf8full/base/copyright.txt tools/chip-model/mf/results/out/mf8full/reports/asap7/mf8full/base/synth_check.txt tools/chip-model/mf/results/out/mf8full/reports/asap7/mf8full/base/synth_stat.txt tools/chip-model/mf/results/out/mf8full/results/asap7/mf8full/base/mem.json tools/chip-model/mf/results/sim/mf8base/gl2sim.log tools/chip-model/mf/results/sim/mf8base/sim_add_s2000.log tools/chip-model/mf/results/sim/mf8base/sim_add_s500.log tools/chip-model/mf/results/sim/mf8base/sim_load_s2000.log tools/chip-model/mf/results/sim/mf8base/sim_load_s500.log tools/chip-model/mf/results/sim/mf8base/sim_mad_s2000.log tools/chip-model/mf/results/sim/mf8base/sim_mad_s500.log tools/chip-model/mf/results/sim/mf8base/sim_mix_s2000.log tools/chip-model/mf/results/sim/mf8base/sim_mix_s500.log tools/chip-model/mf/results/sim/mf8base/sim_mixld_s2000.log tools/chip-model/mf/results/sim/mf8base/sim_mixld_s500.log tools/chip-model/mf/results/sim/mf8base/sim_mul_s2000.log tools/chip-model/mf/results/sim/mf8base/sim_mul_s500.log tools/chip-model/mf/results/sim/mf8base/sim_mulhi_s2000.log tools/chip-model/mf/results/sim/mf8base/sim_mulhi_s500.log tools/chip-model/mf/results/sim/mf8base/sim_or_s2000.log tools/chip-model/mf/results/sim/mf8base/sim_or_s500.log tools/chip-model/mf/results/sim/mf8base/sim_rotl_s2000.log tools/chip-model/mf/results/sim/mf8base/sim_rotl_s500.log tools/chip-model/mf/results/sim/mf8base/sim_rotr_s2000.log tools/chip-model/mf/results/sim/mf8base/sim_rotr_s500.log tools/chip-model/mf/results/sim/mf8base/sim_shfl_s2000.log tools/chip-model/mf/results/sim/mf8base/sim_shfl_s500.log tools/chip-model/mf/results/sim/mf8base/sim_sub_s2000.log tools/chip-model/mf/results/sim/mf8base/sim_sub_s500.log tools/chip-model/mf/results/sim/mf8base/sim_xor_s2000.log tools/chip-model/mf/results/sim/mf8base/sim_xor_s500.log tools/chip-model/mf/results/sim/mf8full/gl2sim.log tools/chip-model/mf/results/sim/mf8full/sim_add_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_add_s500.log tools/chip-model/mf/results/sim/mf8full/sim_andn_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_andn_s500.log tools/chip-model/mf/results/sim/mf8full/sim_bfe_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_bfe_s500.log tools/chip-model/mf/results/sim/mf8full/sim_clz_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_clz_s500.log tools/chip-model/mf/results/sim/mf8full/sim_fwd_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_fwd_s500.log tools/chip-model/mf/results/sim/mf8full/sim_load_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_load_s500.log tools/chip-model/mf/results/sim/mf8full/sim_lop3_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_lop3_s500.log tools/chip-model/mf/results/sim/mf8full/sim_mad_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_mad_s500.log tools/chip-model/mf/results/sim/mf8full/sim_mix1_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_mix1_s500.log tools/chip-model/mf/results/sim/mf8full/sim_mix2_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_mix2_s500.log tools/chip-model/mf/results/sim/mf8full/sim_mix64_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_mix64_s500.log tools/chip-model/mf/results/sim/mf8full/sim_mix_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_mix_s500.log tools/chip-model/mf/results/sim/mf8full/sim_mixld_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_mixld_s500.log tools/chip-model/mf/results/sim/mf8full/sim_mixw4_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_mixw4_s500.log tools/chip-model/mf/results/sim/mf8full/sim_mm8_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_mm8_s500.log tools/chip-model/mf/results/sim/mf8full/sim_mul_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_mul_s500.log tools/chip-model/mf/results/sim/mf8full/sim_mulhi_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_mulhi_s500.log tools/chip-model/mf/results/sim/mf8full/sim_or_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_or_s500.log tools/chip-model/mf/results/sim/mf8full/sim_popc_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_popc_s500.log tools/chip-model/mf/results/sim/mf8full/sim_prmt_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_prmt_s500.log tools/chip-model/mf/results/sim/mf8full/sim_rotl_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_rotl_s500.log tools/chip-model/mf/results/sim/mf8full/sim_rotr_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_rotr_s500.log tools/chip-model/mf/results/sim/mf8full/sim_sel_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_sel_s500.log tools/chip-model/mf/results/sim/mf8full/sim_shfl_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_shfl_s500.log tools/chip-model/mf/results/sim/mf8full/sim_shfla_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_shfla_s500.log tools/chip-model/mf/results/sim/mf8full/sim_shl_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_shl_s500.log tools/chip-model/mf/results/sim/mf8full/sim_shr_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_shr_s500.log tools/chip-model/mf/results/sim/mf8full/sim_sub_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_sub_s500.log tools/chip-model/mf/results/sim/mf8full/sim_xor_s2000.log tools/chip-model/mf/results/sim/mf8full/sim_xor_s500.log tools/chip-model/mf/results/table.csv tools/chip-model/mf/results/table.md tools/chip-model/mf/rtl/core_mf.v tools/chip-model/mf/rtl/mf32base.v tools/chip-model/mf/rtl/mf32full.v tools/chip-model/mf/rtl/mf8base.v tools/chip-model/mf/rtl/mf8full.v tools/chip-model/mf/tb/sram_models.v tools/chip-model/mf/tb/tb_mf32base.v tools/chip-model/mf/tb/tb_mf32full.v tools/chip-model/mf/tb/tb_mf8base.v tools/chip-model/mf/tb/tb_mf8full.v tools/chip-model/mf/tb/tb_mf_common.vh
This commit is contained in:
parent
59915e86df
commit
01452ff1ea
1 changed files with 407 additions and 0 deletions
407
docs/analysis/class-v6/multi-family-adversary.md
Normal file
407
docs/analysis/class-v6/multi-family-adversary.md
Normal file
|
|
@ -0,0 +1,407 @@
|
|||
# The multi-family adversary: one programmable chip against the 180-day family bank (8 October 2026)
|
||||
|
||||
Branch `class-v6-adversary` from the box mirror's master (c86a7e23), the adversary lane of the second external review
|
||||
(the founder's 15:5x BST acceptance). The question the review put: does the 180-day family calendar add cost to a chip,
|
||||
or only change its firmware? The answer is built as the adversary would build it: ONE programmable design that executes
|
||||
every entry of the bank (`docs/analysis/rotation/layer3-family-bank.md`: 18 op families, 3 dataset atoms, the fold form
|
||||
with drawn constants, 3 block shapes, the 64-register window, the per-era parameter draws), with the designer free to
|
||||
choose lanes, clock, pipeline, banking, time-multiplexing, memory organisation and the memory arrangement itself, and
|
||||
priced on a complete machine. Every chip figure is synthesis and placement of RTL written for this lane (Yosys 0.68 and
|
||||
OpenROAD, the ORFS image, ASAP7, run on rented CPU hosts, never on the Mac, never on a Hetzner box); every GPU figure is
|
||||
the record's measurement (`docs/analysis/counter-asic-4-research.md` 15.1a, the floor programme's tier table in the
|
||||
design document 10.4). Labels as the record uses them: measured, modelled, claimed, approximate. The RTL, testbenches,
|
||||
flow and collector are under `tools/chip-model/mf/`; the k lane's rows (`floor/shadow-k.md`) are cited, never re-run.
|
||||
|
||||
The standing rules this file is written under (the review, 15:4x BST): the adversary's design is free; a synthesised
|
||||
k is a model, never a lower bound; a family transition is credited with an obsolescence benefit only where a loss of
|
||||
competitiveness is demonstrated; the stress life is 3 years, shown at 0.5, 1, 2 and 3; the cohort is the discrete-GPU
|
||||
population, Apple reported and not headlined; every defence is scored after the adversary re-optimises.
|
||||
|
||||
## 0. One page
|
||||
|
||||
(filled from the rows below at each cut; section 8 carries the three-part statement)
|
||||
|
||||
## 1. The design: what the adversary builds
|
||||
|
||||
### 1.1 What it must execute
|
||||
|
||||
| Bank entry | What the core does with it | Silicon or firmware |
|
||||
|---|---|---|
|
||||
| G1 to G7: add, sub, xor, or, rotl, rotr, and the lossy or | the ALU group, one unit per lane, operands isolated | silicon once; the mix weights are the program |
|
||||
| G8, G9, G6: mul, mulhi, mad | one 32 x 32 multiplier per lane; mulhi is the high word of the same product; mad adds the third read | silicon once |
|
||||
| G10 shfl (the 32-lane xor-mask shuffle) | a log2(LANES)-stage butterfly across the core's lanes, the mask from the instruction | silicon once per core |
|
||||
| R1 shfla (lane + delta) | a general LANES:1 crossbar per lane, the delta from a register | silicon once per core (the one network a butterfly cannot emulate in one op) |
|
||||
| R2 perm (prmt) | the 4-of-8 byte selector | silicon once |
|
||||
| R3 popc and clz | a popcount tree and a priority encoder | silicon once |
|
||||
| R4 bfe, R5 shl and shr, R6 sel, R7 andn | a shifter, a mask, a select, an and-not | silicon once (fractions of an adder each) |
|
||||
| R8 mm8 (the int8 tile) | a u8 dot4 accumulate per lane (4 MACs per op; the card's m8n8k16 tile is 32 MACs per lane, so 8 chip ops per card tile) | silicon once |
|
||||
| lop3 | the 8-bit truth table | silicon once |
|
||||
| the load and the fold form | the lane's own multiplier computes `x * M`, then the rotate and the three masks with the era's constants in registers; the returned word writes the destination | silicon once; M, R, WM, OFF, MASK are registers written at the era |
|
||||
| W = 4 wide reads | the fold step `x = rotl(x, r) * M ^ w` as an instruction (fwd), three per load | firmware |
|
||||
| the 64-register window | 64 x 32 bits per lane in an SRAM macro | silicon once (the macro) |
|
||||
| the 3 block shapes (64, 128, 256) | the program-length register; the imem holds 256 | firmware |
|
||||
| the drawn select tree | a 32-entry op permutation ahead of decode, written at the era | firmware (a 160-bit register) |
|
||||
| the op-mix band, the fold constants | the program and five registers | firmware |
|
||||
| the 3 dataset atoms (mixer x4, x8, dr368) | the per-window dataset build runs on the same core as a program (the atoms are straight-line ARX and multiply code over 16 registers); the per-hash path never executes an atom | firmware; the build's cost is section 7 |
|
||||
|
||||
### 1.2 The microarchitecture (the designer's choices)
|
||||
|
||||
- **The register state is an SRAM macro, not flops.** One FakeRAM2.0 `fakeram7_64x256` (64 words x 256 bits, single
|
||||
port) per 8 lanes: the lanes are SIMD, every lane reads the same register index, so one 256-bit access serves eight
|
||||
lanes' 32-bit reads. The macro's LEF and Liberty come with the ORFS ASAP7 platform (ABKGroup FakeRAM2.0, 7 nm,
|
||||
0.70 V; area 1,517 um^2 for the 64 x 256, 365 um^2 for the 256 x 34); the placement, the wiring and the clock tree
|
||||
see the macro as a real block. Its dynamic energy is NOT taken from the FakeRAM Liberty (a placeholder, 1.345 per
|
||||
clock edge identical for every size, "VALUES NOT REALISTIC" in the generator's own config); section 2.3 replaces it
|
||||
with a published macro energy band and the gate-level simulation supplies the exact access counts.
|
||||
- **The instruction memory is two `fakeram7_256x34` macros** shared by every lane of the core (40 bits used of 68).
|
||||
- **A single-port macro is time-multiplexed over a 5-phase slot:** read dst, read src, read src2 only when the op
|
||||
needs it (mad, lop3, sel, mm8, shfla), execute, write back. A core retires LANES lane-ops per 5 cycles; throughput is
|
||||
bought with lane count (die area), the cheapest resource the chip has, not with ports.
|
||||
- **Every unit's operands are isolated** (AND-gated by its own select), so an unused unit does not toggle: adding a
|
||||
family's unit costs leakage and a wider result mux, not switching on every op. The k lane's core evaluates every unit
|
||||
every cycle; this is one of the reasons its figure is not a lower bound.
|
||||
- **The era's draws are registers** (fold constants, program length, the select tree); a family epoch writes them.
|
||||
- Clock 1,500 ps (667 MHz) at the TC corner, as the k lane; nothing pipelined beyond the slot; two builds: 8 lanes
|
||||
(placed and routed) and 32 lanes (synthesised), each in two variants: `full` (every bank entry) and `base` (the 10
|
||||
genesis families with the load and the fold, the same microarchitecture): the difference between the two is what the
|
||||
bank adds to the chip.
|
||||
|
||||
### 1.3 What the card pays for the same instruction (the GPU side, measured)
|
||||
|
||||
The 5090's pJ per counted op at stock and at the 1,300 MHz lock (15.1a): add-class 11.3 / 6.2, mul and mad 13.9 / 8.3,
|
||||
mulhi 39.6 / 21.0, prmt 22.3 / 11.5, lop3 24.1 / 13.0, shfl 55.8 / 29.4, the u8 tile 4.1 / 2.2 per MAC. The reserve
|
||||
families by the measured NVIDIA step-cost ratio to the add step (design 4.2): shfla 1.53, popc 1.50, clz 1.63, bfe
|
||||
1.54, shl and shr 0.75, sel and andn about 1.0 (approximate), mm8 by the tile row (32 MACs per lane per instruction).
|
||||
|
||||
## 2. Method
|
||||
|
||||
### 2.1 The flow
|
||||
|
||||
ORFS on ASAP7 (7.5-track RVT, TC corner 0.70 V, NLDM), the default flow: Yosys with ABC, floorplan at 40 percent
|
||||
utilisation (30 for the 32-lane core) with the macros placed by the flow's macro placer under the BLOCKS power grid,
|
||||
global and detailed placement, CTS, global and detailed routing, OpenRCX parasitics. Power is OpenSTA `report_power`
|
||||
under the VCD of a random-input gate-level simulation of the netlist (iverilog; every instruction field drawn by
|
||||
`$random`, the window initialised with random words, a random returned word on every load), with a propagated 0.5
|
||||
activity as the cross-check. Synthesis-only rows (no wires, no clock tree) are marked; placed rows carry the SPEF.
|
||||
|
||||
### 2.2 The activity and the steady state
|
||||
|
||||
Each row is a tag: the op field fixed per family (`+fam=K`), or a drawn program: the class v4 draw over the 10 genesis
|
||||
families (`mix`), the same with one load in 16 (`mixld`), the draw with two reserve families live at 4 points each
|
||||
(`mix1`: shfla and mm8, the two dearest), every reserve family live at 4 points (`mix2`, the bank's bound, not a legal
|
||||
draw), the W = 4 form (`mixw4`), the 64-instruction shape (`mix64`). Two run lengths per tag (500 and 2,000 clocks, 100
|
||||
and 400 slots) bracket the 326-clock load phase (reset, the era's registers, the 256-word program, the 64-word window
|
||||
init) and the run-phase power is solved from the pair, as the k lane does.
|
||||
|
||||
### 2.3 The SRAM macro energy (the one modelled term on the chip side)
|
||||
|
||||
The FakeRAM Liberty's internal power is a placeholder, so the collector removes the macros' Liberty-attributed power
|
||||
(reported separately per VCD with `report_power -instances`) and adds a modelled access energy times the exact
|
||||
access count from the simulation (3 reads and 1 write per slot for a three-operand op, 2 reads and 1 write otherwise,
|
||||
per 8 lanes; 2 imem reads per slot per core). The band: a 64 x 256 single-port macro at a 7 nm class node 3.5 pJ per
|
||||
256-bit access (2.0 to 7.0); a 256 x 34 macro 1.5 pJ per access (0.8 to 3.0). Sources: Horowitz, ISSCC 2014 (45 nm:
|
||||
an 8 KB SRAM read of 64 bits 10 pJ, 32 KB 20 pJ), scaled by the bits moved and by the energy-per-bit reduction from
|
||||
45 nm to a 7 nm class node (about 0.1x to 0.2x, approximate, the same generation scaling the record applies to logic);
|
||||
the k lane's "2 to 4 pJ per 32-bit read, approximate" for a 4 KB imem; CACTI-class estimates for a 16 Kbit macro at
|
||||
7 nm (0.01 to 0.03 pJ per bit read, approximate). Every row carries the band; the low end is near a flop array with
|
||||
perfect clock gating, the high end a conservative compiler macro. The macros' switching on their output nets (256
|
||||
bits into the lanes' latches) stays in the logic figure, measured.
|
||||
|
||||
### 2.4 Node scaling (claimed) and what the method leaves out
|
||||
|
||||
ASAP7 is a predictive 7 nm-class PDK; the row is stated at ASAP7 and scaled by the foundry's headline per-node
|
||||
power reductions at the same speed (the k lane's factors, every one claimed): N5 = 0.70, N3 = 0.50, N2 = 0.36 of
|
||||
ASAP7. Node-for-node against the 5090 (TSMC 4N, N5 class) is the N5 column; a node ahead is N3. Left out on the chip
|
||||
side: the memory controller's queueing logic and the lane's address output (priced in the board model as the
|
||||
controller die), test and clock distribution beyond the block; on the card side the 15.1a figure is the whole card's
|
||||
marginal per counted op, which includes fetch, decode, operand collection and the register file, so the comparison
|
||||
is the chip's whole lane (fetch, decode, window, units, network) against the card's whole lane.
|
||||
|
||||
## 3. The rows: pJ per lane-op per family on the base core (synthesis only, 8 lanes)
|
||||
|
||||
Synthesis only (no wires, no clock tree), 8 lanes, 1,500 ps, ASAP7 TC; "pJ logic" is OpenSTA's figure for the
|
||||
standard cells under the VCD with the FakeRAM placeholder removed (its sequential and combinational parts beside it);
|
||||
"pJ SRAM" the modelled macro term at the simulated access count (low / nominal / high, section 2.3); the per-op
|
||||
figure is logic plus the nominal SRAM term, the band in brackets; the 5090 column is 15.1a (the reserve families
|
||||
by the measured step ratio; the mix rows against the card's 10.3 pJ per op on the class v4 draw at the lock, 18.8
|
||||
unlocked by the ARX ratio, approximate). The full core: 95,678 cells and 3 macros (the base core 66,973 and 3
|
||||
macros: the bank adds 43 percent of the standard cells, 0.13 mW of leakage per 8 lanes, and 11 percent to the
|
||||
energy of the class v4 draw on the same microarchitecture, the wider result mux and the leakage of the idle units).
|
||||
|
||||
The full core (every bank entry):
|
||||
|
||||
| Design | Stage | Family | Cells | Logic W | Leak W | pJ logic (seq / comb) | pJ SRAM (low / nom / high) | pJ/lane-op ASAP7 | N5 | N3 | N2 | 5090 pJ/op unlocked / lock | k N5 lock | k N3 lock | k N3 unlocked |
|
||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
|
||||
| mf8full | power_synth | mix | 95678 | 0.00514 | 0.000395 | 4.82 (0.88 / 3.6) | 0.99 / 1.76 / 3.52 | 6.58 (5.81 to 8.34) | 4.61 | 3.32 | 2.39 | 18.8 / 10.3 | 0.447 | 0.322 | 0.176 |
|
||||
| mf8full | power_synth | mixld | 95678 | 0.00513 | 0.000395 | 4.81 (0.88 / 3.6) | 0.986 / 1.75 / 3.5 | 6.56 (5.79 to 8.31) | 4.59 | 3.31 | 2.38 | 18.8 / 10.3 | 0.446 | 0.321 | 0.176 |
|
||||
| mf8full | power_synth | mix1 | 95678 | 0.00495 | 0.000395 | 4.64 (0.89 / 3.4) | 1.01 / 1.8 / 3.59 | 6.43 (5.65 to 8.23) | 4.5 | 3.24 | 2.33 | 18.8 / 10.3 | 0.437 | 0.315 | 0.172 |
|
||||
| mf8full | power_synth | mix2 | 95678 | 0.00436 | 0.000396 | 4.09 (0.88 / 2.8) | 1 / 1.78 / 3.56 | 5.87 (5.09 to 7.65) | 4.11 | 2.96 | 2.13 | 18.8 / 10.3 | 0.399 | 0.287 | 0.157 |
|
||||
| mf8full | power_synth | mixw4 | 95678 | 0.00508 | 0.000395 | 4.76 (0.88 / 3.5) | 0.977 / 1.73 / 3.47 | 6.49 (5.74 to 8.23) | 4.55 | 3.27 | 2.36 | 18.8 / 10.3 | 0.441 | 0.318 | 0.174 |
|
||||
| mf8full | power_synth | mix64 | 95678 | 0.00488 | 0.000395 | 4.58 (0.88 / 3.3) | 0.993 / 1.76 / 3.53 | 6.34 (5.57 to 8.1) | 4.44 | 3.19 | 2.3 | 18.8 / 10.3 | 0.431 | 0.31 | 0.17 |
|
||||
| mf8full | power_synth | add | 95678 | 0.003 | 0.000396 | 2.81 (0.86 / 1.6) | 0.95 / 1.69 / 3.38 | 4.5 (3.76 to 6.18) | 3.15 | 2.27 | 1.63 | 11.3 / 6.2 | 0.508 | 0.366 | 0.201 |
|
||||
| mf8full | power_synth | sub | 95678 | 0.00299 | 0.000396 | 2.8 (0.86 / 1.6) | 0.95 / 1.69 / 3.38 | 4.49 (3.75 to 6.18) | 3.14 | 2.26 | 1.63 | 11.3 / 6.2 | 0.507 | 0.365 | 0.2 |
|
||||
| mf8full | power_synth | xor | 95678 | 0.00308 | 0.000396 | 2.89 (0.86 / 1.7) | 0.95 / 1.69 / 3.38 | 4.57 (3.84 to 6.26) | 3.2 | 2.3 | 1.66 | 11.3 / 6.2 | 0.516 | 0.372 | 0.204 |
|
||||
| mf8full | power_synth | or | 95678 | 0.00183 | 0.000397 | 1.72 (0.84 / 0.5) | 0.95 / 1.69 / 3.38 | 3.41 (2.67 to 5.09) | 2.38 | 1.72 | 1.24 | 11.3 / 6.2 | 0.385 | 0.277 | 0.152 |
|
||||
| mf8full | power_synth | rotl | 95678 | 0.00318 | 0.000396 | 2.99 (0.86 / 1.8) | 0.95 / 1.69 / 3.38 | 4.67 (3.94 to 6.36) | 3.27 | 2.36 | 1.7 | 11.3 / 6.2 | 0.528 | 0.38 | 0.208 |
|
||||
| mf8full | power_synth | rotr | 95678 | 0.003 | 0.000396 | 2.81 (0.86 / 1.6) | 0.95 / 1.69 / 3.38 | 4.5 (3.76 to 6.18) | 3.15 | 2.27 | 1.63 | 11.3 / 6.2 | 0.508 | 0.366 | 0.201 |
|
||||
| mf8full | power_synth | mul | 95678 | 0.00265 | 0.000396 | 2.48 (0.85 / 1.3) | 0.95 / 1.69 / 3.38 | 4.17 (3.43 to 5.86) | 2.92 | 2.1 | 1.51 | 13.9 / 8.3 | 0.352 | 0.253 | 0.151 |
|
||||
| mf8full | power_synth | mulhi | 95678 | 0.00249 | 0.000396 | 2.33 (0.85 / 1.1) | 0.95 / 1.69 / 3.38 | 4.02 (3.28 to 5.71) | 2.81 | 2.03 | 1.46 | 39.6 / 21 | 0.134 | 0.0965 | 0.0512 |
|
||||
| mf8full | power_synth | mad | 95678 | 0.00562 | 0.000395 | 5.26 (0.91 / 4) | 1.2 / 2.12 / 4.25 | 7.39 (6.46 to 9.51) | 5.17 | 3.72 | 2.68 | 13.9 / 8.3 | 0.623 | 0.449 | 0.268 |
|
||||
| mf8full | power_synth | shfl | 95678 | 0.00266 | 0.000397 | 2.49 (0.86 / 1.3) | 0.95 / 1.69 / 3.38 | 4.18 (3.44 to 5.87) | 2.93 | 2.11 | 1.52 | 55.8 / 29.4 | 0.0995 | 0.0716 | 0.0377 |
|
||||
| mf8full | power_synth | load | 95678 | 0.00433 | 0.000395 | 4.06 (0.86 / 2.8) | 0.95 / 1.69 / 3.38 | 5.75 (5.01 to 7.44) | 4.03 | 2.9 | 2.09 | 13.9 / 8.3 | 0.485 | 0.349 | 0.209 |
|
||||
| mf8full | power_synth | fwd | 95678 | 0.004 | 0.000395 | 3.75 (0.86 / 2.5) | 0.95 / 1.69 / 3.38 | 5.44 (4.7 to 7.13) | 3.81 | 2.74 | 1.97 | 13.9 / 8.3 | 0.459 | 0.33 | 0.197 |
|
||||
| mf8full | power_synth | prmt | 95678 | 0.00255 | 0.000397 | 2.39 (0.86 / 1.2) | 0.95 / 1.69 / 3.38 | 4.07 (3.34 to 5.76) | 2.85 | 2.05 | 1.48 | 22.3 / 11.5 | 0.248 | 0.179 | 0.0921 |
|
||||
| mf8full | power_synth | lop3 | 95678 | 0.003 | 0.000397 | 2.81 (0.91 / 1.5) | 1.2 / 2.12 / 4.25 | 4.94 (4.01 to 7.06) | 3.46 | 2.49 | 1.79 | 24.1 / 13 | 0.266 | 0.191 | 0.103 |
|
||||
| mf8full | power_synth | shfla | 95678 | 0.00302 | 0.000397 | 2.83 (0.9 / 1.6) | 1.2 / 2.12 / 4.25 | 4.95 (4.03 to 7.08) | 3.47 | 2.5 | 1.8 | 17.3 / 9.49 | 0.365 | 0.263 | 0.144 |
|
||||
| mf8full | power_synth | popc | 95678 | 0.0017 | 0.000397 | 1.59 (0.85 / 0.38) | 0.95 / 1.69 / 3.38 | 3.28 (2.54 to 4.97) | 2.3 | 1.65 | 1.19 | 17 / 9.3 | 0.247 | 0.178 | 0.0976 |
|
||||
| mf8full | power_synth | clz | 95678 | 0.00173 | 0.000397 | 1.62 (0.85 / 0.41) | 0.95 / 1.69 / 3.38 | 3.31 (2.57 to 5) | 2.32 | 1.67 | 1.2 | 18.4 / 10.1 | 0.229 | 0.165 | 0.0906 |
|
||||
| mf8full | power_synth | bfe | 95678 | 0.00182 | 0.000397 | 1.71 (0.85 / 0.49) | 0.95 / 1.69 / 3.38 | 3.39 (2.66 to 5.08) | 2.38 | 1.71 | 1.23 | 17.4 / 9.55 | 0.249 | 0.179 | 0.0983 |
|
||||
| mf8full | power_synth | shl | 95678 | 0.00234 | 0.000397 | 2.19 (0.86 / 0.98) | 0.95 / 1.69 / 3.38 | 3.88 (3.14 to 5.57) | 2.71 | 1.95 | 1.41 | 8.48 / 4.65 | 0.584 | 0.42 | 0.231 |
|
||||
| mf8full | power_synth | shr | 95678 | 0.00221 | 0.000397 | 2.07 (0.85 / 0.84) | 0.95 / 1.69 / 3.38 | 3.76 (3.02 to 5.45) | 2.63 | 1.89 | 1.36 | 8.48 / 4.65 | 0.566 | 0.407 | 0.224 |
|
||||
| mf8full | power_synth | sel | 95678 | 0.0028 | 0.000397 | 2.63 (0.91 / 1.4) | 1.2 / 2.12 / 4.25 | 4.75 (3.83 to 6.88) | 3.33 | 2.4 | 1.72 | 11.3 / 6.2 | 0.537 | 0.386 | 0.212 |
|
||||
| mf8full | power_synth | andn | 95678 | 0.00234 | 0.000397 | 2.19 (0.86 / 0.98) | 0.95 / 1.69 / 3.38 | 3.88 (3.14 to 5.57) | 2.72 | 1.96 | 1.41 | 11.3 / 6.2 | 0.438 | 0.315 | 0.173 |
|
||||
| mf8full | power_synth | mm8 | 95678 | 0.0036 | 0.000396 | 3.37 (0.91 / 2.1) | 1.2 / 2.12 / 4.25 | 5.5 (4.57 to 7.62) | 3.85 | 2.77 | 2 | 16.4 / 8.8 | 0.437 | 0.315 | 0.169 |
|
||||
|
||||
The base core (the 10 genesis families on the same microarchitecture; the comparator):
|
||||
|
||||
| Design | Stage | Family | Cells | Logic W | Leak W | pJ logic (seq / comb) | pJ SRAM (low / nom / high) | pJ/lane-op ASAP7 | N5 | N3 | N2 | 5090 pJ/op unlocked / lock | k N5 lock | k N3 lock | k N3 unlocked |
|
||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
|
||||
| mf8base | power_synth | mix | 66973 | 0.00443 | 0.000263 | 4.15 (0.88 / 3) | 0.99 / 1.76 / 3.52 | 5.91 (5.14 to 7.66) | 4.13 | 2.98 | 2.14 | 18.8 / 10.3 | 0.401 | 0.289 | 0.158 |
|
||||
| mf8base | power_synth | mixld | 66973 | 0.00439 | 0.000263 | 4.12 (0.88 / 3) | 0.986 / 1.75 / 3.5 | 5.87 (5.1 to 7.62) | 4.11 | 2.96 | 2.13 | 18.8 / 10.3 | 0.399 | 0.287 | 0.157 |
|
||||
| mf8base | power_synth | add | 66973 | 0.00246 | 0.000264 | 2.31 (0.86 / 1.2) | 0.95 / 1.69 / 3.38 | 4 (3.26 to 5.68) | 2.8 | 2.01 | 1.45 | 11.3 / 6.2 | 0.451 | 0.325 | 0.178 |
|
||||
| mf8base | power_synth | sub | 66973 | 0.00245 | 0.000264 | 2.29 (0.86 / 1.2) | 0.95 / 1.69 / 3.38 | 3.98 (3.24 to 5.67) | 2.79 | 2.01 | 1.44 | 11.3 / 6.2 | 0.449 | 0.324 | 0.178 |
|
||||
| mf8base | power_synth | xor | 66973 | 0.00247 | 0.000264 | 2.32 (0.86 / 1.2) | 0.95 / 1.69 / 3.38 | 4 (3.27 to 5.69) | 2.8 | 2.02 | 1.45 | 11.3 / 6.2 | 0.452 | 0.326 | 0.179 |
|
||||
| mf8base | power_synth | or | 66973 | 0.0016 | 0.000265 | 1.5 (0.84 / 0.42) | 0.95 / 1.69 / 3.38 | 3.19 (2.45 to 4.87) | 2.23 | 1.61 | 1.16 | 11.3 / 6.2 | 0.36 | 0.259 | 0.142 |
|
||||
| mf8base | power_synth | rotl | 66973 | 0.00259 | 0.000264 | 2.43 (0.86 / 1.3) | 0.95 / 1.69 / 3.38 | 4.11 (3.38 to 5.8) | 2.88 | 2.07 | 1.49 | 11.3 / 6.2 | 0.464 | 0.334 | 0.183 |
|
||||
| mf8base | power_synth | rotr | 66973 | 0.00251 | 0.000264 | 2.35 (0.86 / 1.3) | 0.95 / 1.69 / 3.38 | 4.04 (3.3 to 5.73) | 2.83 | 2.04 | 1.47 | 11.3 / 6.2 | 0.456 | 0.328 | 0.18 |
|
||||
| mf8base | power_synth | mul | 66973 | 0.00229 | 0.000264 | 2.14 (0.85 / 1.1) | 0.95 / 1.69 / 3.38 | 3.83 (3.09 to 5.52) | 2.68 | 1.93 | 1.39 | 13.9 / 8.3 | 0.323 | 0.233 | 0.139 |
|
||||
| mf8base | power_synth | mulhi | 66973 | 0.00217 | 0.000264 | 2.03 (0.84 / 0.96) | 0.95 / 1.69 / 3.38 | 3.72 (2.98 to 5.41) | 2.61 | 1.88 | 1.35 | 39.6 / 21 | 0.124 | 0.0893 | 0.0474 |
|
||||
| mf8base | power_synth | mad | 66973 | 0.00472 | 0.000263 | 4.43 (0.9 / 3.3) | 1.2 / 2.12 / 4.25 | 6.55 (5.62 to 8.67) | 4.58 | 3.3 | 2.38 | 13.9 / 8.3 | 0.552 | 0.398 | 0.237 |
|
||||
| mf8base | power_synth | shfl | 66973 | 0.00192 | 0.000265 | 1.8 (0.86 / 0.7) | 0.95 / 1.69 / 3.38 | 3.49 (2.75 to 5.17) | 2.44 | 1.76 | 1.27 | 55.8 / 29.4 | 0.083 | 0.0598 | 0.0315 |
|
||||
| mf8base | power_synth | load | 66973 | 0.00359 | 0.000263 | 3.36 (0.86 / 2.3) | 0.95 / 1.69 / 3.38 | 5.05 (4.31 to 6.74) | 3.53 | 2.54 | 1.83 | 13.9 / 8.3 | 0.426 | 0.307 | 0.183 |
|
||||
|
||||
Reading the rows. (1) The ALU-group families cost the chip 3.1 to 3.3 pJ per lane-op at N5 against the card's 6.2
|
||||
at the lock (k 0.51 to 0.53); or, popc, clz, bfe, prmt 2.3 to 2.9 (k 0.23 to 0.38); the multiply 2.9 (k 0.35);
|
||||
mulhi 2.8 against the card's 21 (k 0.13); the shuffle 2.9 against 29.4 (k 0.10); mad is the dearest op for the
|
||||
chip at 5.2 (three reads and two units, k 0.62) and the shifters the highest k (0.57 to 0.58) because the card
|
||||
does them cheapest. (2) The SRAM term is 1.7 pJ nominal per lane-op (0.95 to 3.4): a third of the row; the
|
||||
sequential term (the phase latches, the instruction register, the counters at five edges per op) 0.86 pJ; the rest
|
||||
is the units and the macro output nets. (3) The reserve families are cheaper for the chip than the genesis
|
||||
families: every live-family mix sits under the class v4 draw, and the bound with all eight live reads 11 percent
|
||||
under it, because the card's dearest instructions (mulhi, shfl, mad) are the genesis ones.
|
||||
|
||||
|
||||
## 4. The whole-hash energy and k per family, node-for-node and a node ahead
|
||||
|
||||
The whole-hash shadow is 102,612 chip ops (the 102,100 counted shadow ops of the record plus the 512-instruction
|
||||
base program; W = 4 adds 384 fold steps). The chip's cost is absolute (its own pJ at its node); the card's premium
|
||||
on the same draw is the measured 0.652 microjoules at the lock, moved by the live family's measured step ratio at 4
|
||||
points of 79 (modelled). Node-for-node is N5 (the 5090's own class), a node ahead N3; every factor claimed.
|
||||
|
||||
design mf8full stage power_synth; the class v4 draw on the core: 6.58 pJ per lane-op ASAP7, 4.61 N5, 3.32 N3 (band 4.07 to 5.84 at N5)
|
||||
|
||||
| Family live (4 points of 79) or draw | Chip pJ per op N5 / N3 (band) | 5090 pJ per op at the lock | k N5 / N3 | Chip shadow per hash, microjoules N5 / N3 (the mix with the family live) | Change vs the class v4 draw | The card's premium per hash at the lock (modelled from the step ratio) | Change |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| the class v4 draw (10 genesis families) | 4.61 / 3.32 (4.07 to 5.84) | 10.3 | 0.45 / 0.32 | 0.473 / 0.340 | +0.0 percent | 0.652 | +0.0 percent |
|
||||
| add (unit row, every instruction) | 3.15 / 2.27 (2.63 to 4.33) | 6.2 | 0.51 / 0.37 | 0.323 / 0.233 | -31.7 percent | 0.392 | -39.8 percent |
|
||||
| sub (unit row, every instruction) | 3.14 / 2.26 (2.63 to 4.32) | 6.2 | 0.51 / 0.36 | 0.322 / 0.232 | -31.8 percent | 0.392 | -39.8 percent |
|
||||
| xor (unit row, every instruction) | 3.20 / 2.30 (2.68 to 4.38) | 6.2 | 0.52 / 0.37 | 0.328 / 0.236 | -30.5 percent | 0.392 | -39.8 percent |
|
||||
| or (unit row, every instruction) | 2.38 / 1.72 (1.87 to 3.57) | 6.2 | 0.38 / 0.28 | 0.245 / 0.176 | -48.2 percent | 0.392 | -39.8 percent |
|
||||
| rotl (unit row, every instruction) | 3.27 / 2.36 (2.75 to 4.45) | 6.2 | 0.53 / 0.38 | 0.336 / 0.242 | -29.0 percent | 0.392 | -39.8 percent |
|
||||
| rotr (unit row, every instruction) | 3.15 / 2.27 (2.63 to 4.33) | 6.2 | 0.51 / 0.37 | 0.323 / 0.233 | -31.7 percent | 0.392 | -39.8 percent |
|
||||
| mul (unit row, every instruction) | 2.92 / 2.10 (2.40 to 4.10) | 8.3 | 0.35 / 0.25 | 0.299 / 0.216 | -36.7 percent | 0.525 | -19.4 percent |
|
||||
| mulhi (unit row, every instruction) | 2.81 / 2.03 (2.30 to 3.99) | 21.0 | 0.13 / 0.10 | 0.289 / 0.208 | -38.9 percent | 1.329 | +103.9 percent |
|
||||
| mad (unit row, every instruction) | 5.17 / 3.72 (4.52 to 6.66) | 8.3 | 0.62 / 0.45 | 0.531 / 0.382 | +12.3 percent | 0.525 | -19.4 percent |
|
||||
| shfl (unit row, every instruction) | 2.93 / 2.11 (2.41 to 4.11) | 29.4 | 0.10 / 0.07 | 0.300 / 0.216 | -36.5 percent | 1.861 | +185.4 percent |
|
||||
| load (unit row, every instruction) | 4.03 / 2.90 (3.51 to 5.21) | 8.3 | 0.48 / 0.35 | 0.413 / 0.297 | -12.6 percent | 0.525 | -19.4 percent |
|
||||
| prmt live | 2.85 / 2.05 (2.34 to 4.03) | 11.5 | 0.25 / 0.18 | 0.464 / 0.334 | -1.9 percent | 0.662 | +1.5 percent |
|
||||
| lop3 live | 3.46 / 2.49 (2.81 to 4.94) | 13.0 | 0.27 / 0.19 | 0.467 / 0.336 | -1.3 percent | 0.688 | +5.6 percent |
|
||||
| shfla live | 3.47 / 2.50 (2.82 to 4.95) | 9.5 | 0.37 / 0.26 | 0.467 / 0.336 | -1.3 percent | 0.669 | +2.7 percent |
|
||||
| popc live | 2.30 / 1.65 (1.78 to 3.48) | 9.3 | 0.25 / 0.18 | 0.461 / 0.332 | -2.5 percent | 0.669 | +2.5 percent |
|
||||
| clz live | 2.32 / 1.67 (1.80 to 3.50) | 10.1 | 0.23 / 0.17 | 0.461 / 0.332 | -2.5 percent | 0.673 | +3.2 percent |
|
||||
| bfe live | 2.38 / 1.71 (1.86 to 3.56) | 9.5 | 0.25 / 0.18 | 0.461 / 0.332 | -2.5 percent | 0.670 | +2.7 percent |
|
||||
| shl live | 2.71 / 1.95 (2.20 to 3.90) | 4.7 | 0.58 / 0.42 | 0.463 / 0.333 | -2.1 percent | 0.644 | -1.3 percent |
|
||||
| shr live | 2.63 / 1.89 (2.12 to 3.81) | 4.7 | 0.57 / 0.41 | 0.462 / 0.333 | -2.2 percent | 0.644 | -1.3 percent |
|
||||
| sel live | 3.33 / 2.40 (2.68 to 4.81) | 6.2 | 0.54 / 0.39 | 0.466 / 0.336 | -1.4 percent | 0.652 | +0.0 percent |
|
||||
| andn live | 2.72 / 1.96 (2.20 to 3.90) | 6.2 | 0.44 / 0.32 | 0.463 / 0.333 | -2.1 percent | 0.652 | +0.0 percent |
|
||||
| mm8 live | 3.85 / 2.77 (3.20 to 5.34) | 8.8 | 0.44 / 0.31 | 0.469 / 0.337 | -0.8 percent | 0.699 | +7.2 percent |
|
||||
| fwd (unit row, every instruction) | 3.81 / 2.74 (3.29 to 4.99) | 8.3 | 0.46 / 0.33 | 0.391 / 0.281 | -17.3 percent | 0.525 | -19.4 percent |
|
||||
| the draw with shfla and mm8 live (measured mix) | 4.50 / 3.24 (3.96 to 5.76) | 10.3 | 0.44 / 0.31 | 0.462 / 0.333 | -2.2 percent | 0.652 | +0.0 percent |
|
||||
| every reserve family live at 4 points (the bound, measured mix) | 4.11 / 2.96 (3.56 to 5.35) | 10.3 | 0.40 / 0.29 | 0.422 / 0.304 | -10.8 percent | 0.652 | +0.0 percent |
|
||||
| the draw with one load in 16 | 4.59 / 3.31 (4.06 to 5.82) | 10.3 | 0.45 / 0.32 | 0.471 / 0.339 | -0.3 percent | 0.652 | +0.0 percent |
|
||||
| W = 4: three fold steps per load | 4.55 / 3.27 (4.02 to 5.76) | 10.3 | 0.44 / 0.32 | 0.466 / 0.336 | -1.3 percent | 0.652 | +0.0 percent |
|
||||
| the 64-instruction shape | 4.44 / 3.19 (3.90 to 5.67) | 10.3 | 0.43 / 0.31 | 0.455 / 0.328 | -3.7 percent | 0.652 | +0.0 percent |
|
||||
|
||||
The whole-hash edge per joule with the class v4 shadow on the core (absolute: the chip pays its own pJ whatever the card does):
|
||||
| Memory (E_mem, microjoules) | Chip E_hash N5 / N3 | vs 5090 lock 2.33 | vs 5090 stock 3.36 | vs 5080 lock 2.06 | vs M5 Max 1.40 |
|
||||
|---|---|---|---|---|---|
|
||||
| GDDR7 board (0.466) | 0.939 / 0.806 | 2.5x / 2.9x | 3.6x / 4.2x | 2.2x / 2.6x | 1.5x / 1.7x |
|
||||
| HBM3 one stack (0.321) | 0.794 / 0.661 | 2.9x / 3.5x | 4.2x / 5.1x | 2.6x / 3.1x | 1.8x / 2.1x |
|
||||
| SRAM N2 die at W = 1 (0.036) | 0.509 / 0.376 | 4.6x / 6.2x | 6.6x / 8.9x | 4.1x / 5.5x | 2.8x / 3.7x |
|
||||
|
||||
The same edge on the base core (the genesis-only comparator): 2.6x / 3.0x microjoules per hash
|
||||
on the GDDR7 board at N5 / N3, 3.8x / 4.4x against the 5090 at its lock; the bank costs the
|
||||
adversary 5 percent of its edge (0.939 against 0.890 microjoules on the GDDR7 board, 2.5x against 2.6x), which is
|
||||
the whole answer to the review's question in one number: the 180-day calendar costs a chip that carries the bank
|
||||
about 11 percent of its shadow energy and 43 percent of its core cells, and no epoch costs it a part.
|
||||
|
||||
|
||||
## 5. The placed rows (8 lanes, routed, SPEF) and the 32-lane core
|
||||
|
||||
ROWS_PLACED
|
||||
|
||||
## 6. The board: joules per valid hash and USD per sustained MH/s for the complete machine
|
||||
|
||||
The complete machine: the memory devices at their modelled random-read energy and activate ceiling (chip-model-v3
|
||||
5.3, unmeasured; the 5090 reaches 82 percent of the GDDR7 figure, the sustained fraction here), the controller and
|
||||
PHY die, the core die sized to retire the hash's ops at the memory's sustained rate (lanes at 667 MHz over 5 phases;
|
||||
0.002 mm^2 per lane at N5 from the 8-lane core's floorplan at 40 percent utilisation scaled x0.55, approximate; USD
|
||||
0.36 per mm^2 of N5 from sram-mirror's yield model; 50 uW of leakage per lane, the synthesis figure), power delivery
|
||||
(PSU 92 percent, VRM 90 percent), cooling (3 percent), a board and assembly (USD 200), and one full node per 100
|
||||
machines (85 W and USD 1,500 shared: the state-derived dataset's host, priced as the review's rule 3 asks). The
|
||||
"high" case takes the memory's lower ceiling (the 5090's measured 17.5 G on GDDR7; the JEDEC tFAW floor on HBM3),
|
||||
the high read energy and the high SRAM term together. GPU rows: the 5090 at its lock 2.33 microjoules at 134.76
|
||||
MH/s, USD 1,999 plus USD 150 of rig share (USD 15.9 per MH/s; at the USD 3,000 street price 23.4); the 5080 at its
|
||||
lock 2.06 at 71.20 MH/s, USD 999 plus 150 (USD 16.1 per MH/s); the discrete-GPU cohort by count (design 3.4 and
|
||||
10.4: 8 GB 22 percent, 12 GB 22, 16 GB 28, 24 GB and up 16, the rest 10 and 11 GB) at its tuned points about 3.6
|
||||
microjoules (2.5 to 4.5, approximate: the 5070 and 5070 Ti 1.7 to 1.75 modelled, the 4070 3.58 measured, the 4090
|
||||
3.64, the 3090 6.4, the 9070 XT 8.1 measured) and about USD 18 per MH/s (14 to 28); the M5 Max 1.40 at 27.9 MH/s is
|
||||
reported in section 4, not headlined.
|
||||
|
||||
Node-for-node (N5), the full core at 4.61 pJ per lane-op:
|
||||
|
||||
core 4.61 pJ per lane-op (band 4.07 to 5.84), 102612 ops per hash, leakage 50 uW per lane, 0.002 mm^2 per lane at N5
|
||||
|
||||
| Memory arrangement | Case | Sustained MH/s | Machine W | microjoules per hash (whole machine) | Core lanes | Core mm^2 (N5) | Capex USD | USD per MH/s | vs 5090 lock 2.33 (J / USD) | vs 5080 lock 2.06 | vs cohort 3.6 / USD 18 |
|
||||
|---|---|---|---|---|---|---|---|---|---|---|---|
|
||||
| GDDR7 board, 16 devices, 64 channels (the 5090 memory without the GPU) | nominal | 136 | 175 | 1.280 | 104,960 | 210 | 661 | 4.84 | 1.8x / 3.3x | 1.6x / 3.3x | 2.8x / 3.7x |
|
||||
| GDDR7 board, 16 devices, 64 channels (the 5090 memory without the GPU) | low | 136 | 154 | 1.132 | 104,960 | 210 | 661 | 4.84 | 2.1x / 3.3x | 1.8x / 3.3x | 3.2x / 3.7x |
|
||||
| GDDR7 board, 16 devices, 64 channels (the 5090 memory without the GPU) | high | 112 | 180 | 1.603 | 86,235 | 172 | 647 | 5.77 | 1.5x / 2.8x | 1.3x / 2.8x | 2.2x / 3.1x |
|
||||
| HBM3 one stack (activate ceiling unmeasured; JEDEC tFAW floor 2.3 G) | nominal | 71 | 77 | 1.084 | 54,656 | 109 | 704 | 9.91 | 2.1x / 1.6x | 1.9x / 1.6x | 3.3x / 1.8x |
|
||||
| HBM3 one stack (activate ceiling unmeasured; JEDEC tFAW floor 2.3 G) | low | 71 | 70 | 0.984 | 54,656 | 109 | 704 | 9.91 | 2.4x / 1.6x | 2.1x / 1.6x | 3.7x / 1.8x |
|
||||
| HBM3 one stack (activate ceiling unmeasured; JEDEC tFAW floor 2.3 G) | high | 15 | 34 | 2.228 | 11,748 | 23 | 673 | 44.09 | 1.0x / 0.4x | 0.9x / 0.4x | 1.6x / 0.4x |
|
||||
| HBM3 eight stacks (an H100-class package) | nominal | 566 | 559 | 0.987 | 435,713 | 871 | 3,079 | 5.44 | 2.4x / 2.9x | 2.1x / 3.0x | 3.6x / 3.3x |
|
||||
| HBM3 eight stacks (an H100-class package) | low | 566 | 502 | 0.886 | 435,713 | 871 | 3,079 | 5.44 | 2.6x / 2.9x | 2.3x / 3.0x | 4.1x / 3.3x |
|
||||
| HBM3 eight stacks (an H100-class package) | high | 122 | 217 | 1.772 | 93,987 | 188 | 2,833 | 23.18 | 1.3x / 0.7x | 1.2x / 0.7x | 2.0x / 0.8x |
|
||||
| SRAM full store, one N2 reticle, 2 GiB at W = 1 (the strongest five-year chip) | nominal | 535 | 400 | 0.747 | 411,225 | 822 | 1,211 | 2.27 | 3.1x / 7.0x | 2.8x / 7.1x | 4.8x / 7.9x |
|
||||
| SRAM full store, one N2 reticle, 2 GiB at W = 1 (the strongest five-year chip) | low | 609 | 403 | 0.662 | 468,572 | 937 | 1,252 | 2.06 | 3.5x / 7.8x | 3.1x / 7.8x | 5.4x / 8.8x |
|
||||
| SRAM full store, one N2 reticle, 2 GiB at W = 1 (the strongest five-year chip) | high | 419 | 394 | 0.940 | 322,466 | 645 | 1,147 | 2.74 | 2.5x / 5.8x | 2.2x / 5.9x | 3.8x / 6.6x |
|
||||
|
||||
Lifetime cost in USD per TH (10^12 hashes), capex spread over the life plus electricity at USD 0.08 per kWh:
|
||||
| Machine | 0.5 y | 1 y | 2 y | 3 y | of which electricity |
|
||||
|---|---|---|---|---|---|
|
||||
| 5090 at the lock | 1.063 | 0.557 | 0.305 | 0.220 | 0.0518 |
|
||||
| 5080 at the lock | 1.069 | 0.557 | 0.302 | 0.216 | 0.0458 |
|
||||
| cohort card | 1.222 | 0.651 | 0.365 | 0.270 | 0.0800 |
|
||||
| GDDR7 board, 16 devices, 64 channels chip | 0.335 | 0.182 | 0.105 | 0.080 | 0.0284 |
|
||||
| HBM3 one stack chip | 0.653 | 0.338 | 0.181 | 0.129 | 0.0241 |
|
||||
| HBM3 eight stacks chip | 0.367 | 0.194 | 0.108 | 0.079 | 0.0219 |
|
||||
| SRAM full store, one N2 reticle, 2 GiB at W = 1 chip | 0.160 | 0.088 | 0.053 | 0.041 | 0.0166 |
|
||||
|
||||
A node ahead (N3), the same machine:
|
||||
|
||||
|---|---|---|---|---|---|---|---|---|---|---|---|
|
||||
| GDDR7 board, 16 devices, 64 channels (the 5090 memory without the GPU) | nominal | 136 | 150 | 1.101 | 104,960 | 147 | 638 | 4.67 | 2.1x / 3.4x | 1.9x / 3.5x | 3.3x / 3.9x |
|
||||
| GDDR7 board, 16 devices, 64 channels (the 5090 memory without the GPU) | low | 136 | 133 | 0.972 | 104,960 | 147 | 638 | 4.67 | 2.4x / 3.4x | 2.1x / 3.5x | 3.7x / 3.9x |
|
||||
| GDDR7 board, 16 devices, 64 channels (the 5090 memory without the GPU) | high | 112 | 155 | 1.380 | 86,235 | 121 | 628 | 5.61 | 1.7x / 2.8x | 1.5x / 2.9x | 2.6x / 3.2x |
|
||||
| HBM3 one stack (activate ceiling unmeasured; JEDEC tFAW floor 2.3 G) | nominal | 71 | 64 | 0.905 | 54,656 | 77 | 693 | 9.75 | 2.6x / 1.6x | 2.3x / 1.7x | 4.0x / 1.8x |
|
||||
| HBM3 one stack (activate ceiling unmeasured; JEDEC tFAW floor 2.3 G) | low | 71 | 59 | 0.824 | 54,656 | 77 | 693 | 9.75 | 2.8x / 1.6x | 2.5x / 1.7x | 4.4x / 1.8x |
|
||||
| HBM3 one stack (activate ceiling unmeasured; JEDEC tFAW floor 2.3 G) | high | 15 | 31 | 2.004 | 11,748 | 16 | 671 | 43.93 | 1.2x / 0.4x | 1.0x / 0.4x | 1.8x / 0.4x |
|
||||
| HBM3 eight stacks (an H100-class package) | nominal | 566 | 458 | 0.808 | 435,713 | 610 | 2,985 | 5.27 | 2.9x / 3.0x | 2.5x / 3.1x | 4.5x / 3.4x |
|
||||
| HBM3 eight stacks (an H100-class package) | low | 566 | 411 | 0.726 | 435,713 | 610 | 2,985 | 5.27 | 3.2x / 3.0x | 2.8x / 3.1x | 5.0x / 3.4x |
|
||||
| HBM3 eight stacks (an H100-class package) | high | 122 | 189 | 1.548 | 93,987 | 132 | 2,812 | 23.02 | 1.5x / 0.7x | 1.3x / 0.7x | 2.3x / 0.8x |
|
||||
| SRAM full store, one N2 reticle, 2 GiB at W = 1 (the strongest five-year chip) | nominal | 724 | 398 | 0.550 | 557,288 | 780 | 1,196 | 1.65 | 4.2x / 9.7x | 3.7x / 9.8x | 6.5x / 10.9x |
|
||||
| SRAM full store, one N2 reticle, 2 GiB at W = 1 (the strongest five-year chip) | low | 828 | 402 | 0.485 | 636,578 | 891 | 1,236 | 1.49 | 4.8x / 10.7x | 4.2x / 10.8x | 7.4x / 12.1x |
|
||||
|
||||
Reading. (1) The complete GDDR7 machine reads 1.8x the 5090 at its lock per joule node-for-node (1.5x to 2.1x) and
|
||||
2.1x a node ahead (1.7x to 2.4x); against the 5080 at its lock 1.6x (1.3x to 1.8x) and 1.9x; against the cohort
|
||||
card 2.8x and 3.3x. Per dollar of capex it is 3.3x the 5090 at MSRP (4.8x at the street price) and 3.7x the cohort,
|
||||
because the board carries the same USD 320 of memory and USD 76 of core silicon where the card carries a 750 mm^2
|
||||
GPU. (2) The HBM3 one-stack machine is no better per joule than GDDR7 once the controller, the host and the power
|
||||
train are in, and its dollars per MH/s are twice the card's at the modelled ceiling and 2.7x worse than the card at
|
||||
the JEDEC tFAW floor: the "same DRAM" assumption is tested and the adversary's cheapest memory IS the card's memory,
|
||||
so the GDDR7 board is the machine the economics must answer. (3) The N2 SRAM die at the hash's own width is the
|
||||
strongest machine at 3.1x per joule node-for-node (2.5x to 3.5x) and 4.2x a node ahead, and 7x per dollar of
|
||||
silicon, with the project cost (USD 100 M to 500 M, claimed) as its only hold. (4) Over a 3-year life the GDDR7
|
||||
machine's cost per TH is 0.080 USD against the 5090's 0.220 and the cohort's 0.270 (2.7x to 3.4x); at a 1-year life
|
||||
3.1x; electricity is a quarter of the machine's lifetime cost and a fifth of the card's, so the per-joule edge is
|
||||
the smaller half of the economic edge and the capex per MH/s the larger. The profitability surface the review asks
|
||||
for (development, fleet capital, share, margin, electricity, pre-production, discounting, residual, operator and
|
||||
manufacturer apart) takes these per-TH rows as its inputs; it is the economics lane's, not this file's.
|
||||
|
||||
|
||||
## 7. The lifetime answer: what each transition needs, and what is credited
|
||||
|
||||
The rule: a transition is credited only where the next entry needs a physical resource this design cannot supply
|
||||
economically; firmware changes (a program, a register, a weight) are credited zero. "Performance loss" is the change in
|
||||
the chip's energy per hash when the entry is live at its draw weight (4 points of 79 for a reserve family; the whole
|
||||
program for an atom or a shape), from the rows of sections 3 and 4; the GPU's own change on the same draw is beside it.
|
||||
|
||||
| Transition (the next epoch draws it) | Physical resource it needs | Does this design hold it? | Chip loss at the draw weight | The card's change on the same draw (measured ratio) | Credited obsolescence benefit |
|
||||
|---|---|---|---|---|---|
|
||||
| G1 to G7 (the ARX group) at any band weight | the ALU group | yes | 0 to +3 percent of the shadow energy across the band (the unit rows 3.1 to 3.3 pJ at N5 against the draw's 4.6) | 1.00 | 0 |
|
||||
| G8, G9, G6 (mul, mulhi, mad) at any band weight | the multiplier and the third read | yes | -1 to +2 percent at B = 4 (mul 2.9, mulhi 2.8, mad 5.2 pJ at N5) | mul 1.1, mulhi about 3.4 (the card's dearest ALU op) | 0 |
|
||||
| G10 shfl at its cap (8 points) | the lane butterfly | yes (per core) | 0 (2.9 pJ at N5, under the draw's mean) | 4.9x the add per op | 0 (the card pays 29.4 pJ for the move the chip pays about 1) |
|
||||
| R1 shfla live (4 points) | a general lane crossbar | yes (per core; the one network a butterfly cannot emulate in one op) | 0 (2.9 pJ at N5, under the draw's mean)A | 1.53 (NVIDIA), 1.91 (Apple) | 0 |
|
||||
| R2 perm live | a byte selector | yes | -1.9 percent | 1.30 | 0 |
|
||||
| R3 popc and clz live | a popcount tree, a priority encoder | yes | -2.5 percent | 1.50 and 1.63 | 0 |
|
||||
| R4 to R7 (bfe, shl and shr, sel, andn) live | a shifter, a mask, a select, an and-not | yes | -2.1 to -1.4 percent | 0.75 to 1.54 | 0 |
|
||||
| R8 mm8 live | a u8 dot4 per lane (8 chip ops per card tile) | yes | -0.8 percent (3.9 pJ per dp4a at N5, 1.0 pJ per MAC, against the card's 2.2 pJ per MAC at the lock) | the tile: 2.43 the add step per card instruction (32 MACs) | 0 (the chip's MAC is cheaper than the card's by 4x to 30x on the public figures; this is the card's loss, not the chip's) |
|
||||
| the op-mix band draw (B = 4 on injecting families) | nothing: the program | yes | within the family rows above | within 11 percent per instruction (shadow-k 6.2) | 0 |
|
||||
| the fold constants draw | five registers | yes | 0 | 0 | 0 |
|
||||
| the block shape draw (64, 128, 256) | the program-length register; 256 words of imem | yes | -3.7 percent at 64 (the imem term; 0 at 128 and 256) | 64 ran 2.5 to 3.5 percent faster than 256 on the 5090 and the M5 Max (measured) | 0 |
|
||||
| the read-width draw W = 4 (16 bytes) | three fwd instructions per load (firmware); the same DRAM sector | yes | +0.3 percent per hash (384 fold steps at 3.8 pJ) | within 2.7 percent on the 5090 and the 9070 XT, within 1 on the M5 Max (measured) | 0 |
|
||||
| the 64-register window | the 64 x 256 macro per 8 lanes | yes (built in) | 0 (it is the base) | modelled: occupancy to about half on a 5090 or 4090, the rate expected to hold (the hash lane) | 0 |
|
||||
| the dataset atom draw A1 / A2 / A3 (mixer x4, x8, dr368) | nothing per hash; the per-window build is a program on the same core (section 7.1) | yes | 0 per hash; the build 0.72 J per window per machine at 1 GiB (0.2 mW averaged), 1.45 J at 2 GiB | the card's rate unmoved by the atom (within 0.1 MH/s on the 5090, measured) | 0 |
|
||||
| a new op family outside the 18 (a bank refresh, a release) | a unit the die lacks | no: the chip emulates it from the 18 at the vendor penalty, as the cards do (1.5x to 2.4x per op, measured on the cards) or loses its 4 points | at 4 points of 79: at most 4 / 79 x (penalty - 1) of the shadow energy, about 2 to 6 percent (modelled) | the same emulation on every card that predates the release, 0 on a card with the native op | 0 unless the family is one the 18 cannot emulate; none proposed is |
|
||||
| a new read atom W = 8 (32 bytes, `admissible: false` today) | seven fwd instructions per load; the same GDDR7 sector; on an SRAM die +0.3 nJ of wire per read (floor lane 3) | yes | +0.7 percent per hash (896 fold steps) | free by the measured rows (one sector per load on NVIDIA) | 0 |
|
||||
| the dataset floor step (layer 2: 5.5 / 8.5 / 11.5 GiB) | device memory: 3 to 6 GDDR7 devices more, or 3 to 6 N2 reticles on the SRAM die | yes on DRAM (USD 60 to 120 more); the SRAM die's ticket rises USD 1,000 per step | 0 per joule on DRAM; the SRAM die's USD per MH/s unmoved (every die powered) | the tuned 5090 pays 4, 8 and 10 percent more energy per hash at 2, 4 and 8 GiB (measured) | 0 per joule; a capex ticket on the SRAM die only (floor lane 3) |
|
||||
|
||||
Reading, before the numbers: nothing in the bank asks for a resource the design lacks, because the bank is public at
|
||||
genesis and its whole op-family set is a few adders per lane and two networks per core. What the bank does to this
|
||||
chip is the per-op cost of carrying the unit set (sections 3 and 5: the full core against the base core on the same
|
||||
microarchitecture) and the leakage of the units that are not live, and that is the number the credit must come from.
|
||||
|
||||
### 7.1 The per-window dataset build on the chip (the atoms' only cost)
|
||||
|
||||
Under class v5 the dataset is rebuilt from the chain's state every window (3,600 s). A 1 GiB build is 157 G ops
|
||||
(measured as 13.4 ms on a 5090; the record's row 7). On this core at 4.61 pJ per lane-op (N5) that is 0.72 J per
|
||||
window per machine, 0.0002 W averaged over the window, against a machine of hundreds of watts: 0.0001 percent of
|
||||
the machine's energy, the same for every atom within the atom's op count (x4 half of x8, dr368 about x4's). The
|
||||
chip's node is the host the board model carries (one full node per 100 machines: 85 W and USD 1,500 shared); a
|
||||
specialised machine with a host keeping the dataset current is inside the adversary model by the review's rule (3),
|
||||
and this is its price: under 1 W and USD 15 per machine.
|
||||
|
||||
|
||||
## 8. Energy resistance, economic resistance and response capability, stated separately
|
||||
|
||||
STATEMENT
|
||||
|
||||
## 9. Sources and what is owed
|
||||
|
||||
- The GPU side: `docs/analysis/counter-asic-4-research.md` 15.1a (the 5090 microbench, 8 October 2026, measured);
|
||||
the floor programme's tier table (`docs/design/class-v6-rotating-family.md` 10.4: the 5090 at the 1,300 lock 2.33
|
||||
microjoules at 134.76 MH/s and 312.5 W, the 5080 at the 1,100 lock 2.06 at 71.20 MH/s and 146.6 W, stock 3.36 and
|
||||
3.48, the 4070, 4090, 3090 and 9070 XT rows, the M5 Max 1.40); the reserve families' step costs, design 4.2
|
||||
(measured on the cards); the card population by count, design 3.4 (approximate).
|
||||
- The chip side: this lane's RTL and flow (`tools/chip-model/mf/`), ORFS on ASAP7 (Clark et al., Microelectronics
|
||||
Journal 2016; the 7.5-track RVT library, TC corner), FakeRAM2.0 macros (ABKGroup, the ORFS platform's
|
||||
`fakeram7_64x256` and `fakeram7_256x34`, LEF and Liberty; the generator's own config marks its values "not
|
||||
realistic", so only area, pins and placement are taken from it); the k lane's rows (`floor/shadow-k.md`, the
|
||||
shuffle row 1.24 pJ per lane-op routed, relayed 16:0x BST) cited as published.
|
||||
- The SRAM access energy band: Horowitz, "Computing's energy problem (and what we can do about it)", ISSCC 2014
|
||||
(45 nm: 8 KB SRAM 10 pJ per 64-bit read, 32 KB 20 pJ); the scaling to a 7 nm class node approximate; the k lane's
|
||||
imem figure (2 to 4 pJ per 32-bit read, approximate); CACTI-class estimates (approximate).
|
||||
- Node scaling (claimed): TSMC's technology pages for N5, N3E and N2, read 8 October 2026 (shadow-k section 7).
|
||||
- The board: `docs/analysis/chip-model-v3.md` 5.3 and 5.5 (the GDDR7 and HBM3 random-read engines, the activate
|
||||
ceilings, the static and controller allowances, the prices, all modelled or claimed; the 5090 at 82 percent of the
|
||||
GDDR7 ceiling, measured); `floor/sram-and-floor.md` 2.1 (the SRAM die at W = 1, modelled); the power delivery,
|
||||
cooling and host allowances are this file's (approximate; the sensitivity is stated beside each).
|
||||
- The dataset build: the record's row 7 (a 1 GiB rebuild 157 G ops, 13.4 ms on a 5090, measured).
|
||||
|
||||
Owed: the k lane's crossbar, scratch and tile rows (in place and route at 16:0x BST); the placed 32-lane core (the
|
||||
8-lane core is placed; the 32-lane row is synthesis only); a real PDK memory compiler's figure for the two macros
|
||||
(FakeRAM gives area and pins only); the 64-register window's GPU cost measured (the hash lane's generator line);
|
||||
the HBM3 activate ceiling (unmeasured, the AWS F2 hour); the profitability surface of the review's rule (2) over the
|
||||
lifetime rows of section 6, which this file states as cost per TH and leaves the NPV to the economics lane.
|
||||
|
||||
Loading…
Reference in a new issue