diff --git a/docs/analysis/class-v6/multi-family-adversary.md b/docs/analysis/class-v6/multi-family-adversary.md new file mode 100644 index 000000000..ba51d13e9 --- /dev/null +++ b/docs/analysis/class-v6/multi-family-adversary.md @@ -0,0 +1,407 @@ +# The multi-family adversary: one programmable chip against the 180-day family bank (8 October 2026) + +Branch `class-v6-adversary` from the box mirror's master (c86a7e23), the adversary lane of the second external review +(the founder's 15:5x BST acceptance). The question the review put: does the 180-day family calendar add cost to a chip, +or only change its firmware? The answer is built as the adversary would build it: ONE programmable design that executes +every entry of the bank (`docs/analysis/rotation/layer3-family-bank.md`: 18 op families, 3 dataset atoms, the fold form +with drawn constants, 3 block shapes, the 64-register window, the per-era parameter draws), with the designer free to +choose lanes, clock, pipeline, banking, time-multiplexing, memory organisation and the memory arrangement itself, and +priced on a complete machine. Every chip figure is synthesis and placement of RTL written for this lane (Yosys 0.68 and +OpenROAD, the ORFS image, ASAP7, run on rented CPU hosts, never on the Mac, never on a Hetzner box); every GPU figure is +the record's measurement (`docs/analysis/counter-asic-4-research.md` 15.1a, the floor programme's tier table in the +design document 10.4). Labels as the record uses them: measured, modelled, claimed, approximate. The RTL, testbenches, +flow and collector are under `tools/chip-model/mf/`; the k lane's rows (`floor/shadow-k.md`) are cited, never re-run. + +The standing rules this file is written under (the review, 15:4x BST): the adversary's design is free; a synthesised +k is a model, never a lower bound; a family transition is credited with an obsolescence benefit only where a loss of +competitiveness is demonstrated; the stress life is 3 years, shown at 0.5, 1, 2 and 3; the cohort is the discrete-GPU +population, Apple reported and not headlined; every defence is scored after the adversary re-optimises. + +## 0. One page + +(filled from the rows below at each cut; section 8 carries the three-part statement) + +## 1. The design: what the adversary builds + +### 1.1 What it must execute + +| Bank entry | What the core does with it | Silicon or firmware | +|---|---|---| +| G1 to G7: add, sub, xor, or, rotl, rotr, and the lossy or | the ALU group, one unit per lane, operands isolated | silicon once; the mix weights are the program | +| G8, G9, G6: mul, mulhi, mad | one 32 x 32 multiplier per lane; mulhi is the high word of the same product; mad adds the third read | silicon once | +| G10 shfl (the 32-lane xor-mask shuffle) | a log2(LANES)-stage butterfly across the core's lanes, the mask from the instruction | silicon once per core | +| R1 shfla (lane + delta) | a general LANES:1 crossbar per lane, the delta from a register | silicon once per core (the one network a butterfly cannot emulate in one op) | +| R2 perm (prmt) | the 4-of-8 byte selector | silicon once | +| R3 popc and clz | a popcount tree and a priority encoder | silicon once | +| R4 bfe, R5 shl and shr, R6 sel, R7 andn | a shifter, a mask, a select, an and-not | silicon once (fractions of an adder each) | +| R8 mm8 (the int8 tile) | a u8 dot4 accumulate per lane (4 MACs per op; the card's m8n8k16 tile is 32 MACs per lane, so 8 chip ops per card tile) | silicon once | +| lop3 | the 8-bit truth table | silicon once | +| the load and the fold form | the lane's own multiplier computes `x * M`, then the rotate and the three masks with the era's constants in registers; the returned word writes the destination | silicon once; M, R, WM, OFF, MASK are registers written at the era | +| W = 4 wide reads | the fold step `x = rotl(x, r) * M ^ w` as an instruction (fwd), three per load | firmware | +| the 64-register window | 64 x 32 bits per lane in an SRAM macro | silicon once (the macro) | +| the 3 block shapes (64, 128, 256) | the program-length register; the imem holds 256 | firmware | +| the drawn select tree | a 32-entry op permutation ahead of decode, written at the era | firmware (a 160-bit register) | +| the op-mix band, the fold constants | the program and five registers | firmware | +| the 3 dataset atoms (mixer x4, x8, dr368) | the per-window dataset build runs on the same core as a program (the atoms are straight-line ARX and multiply code over 16 registers); the per-hash path never executes an atom | firmware; the build's cost is section 7 | + +### 1.2 The microarchitecture (the designer's choices) + +- **The register state is an SRAM macro, not flops.** One FakeRAM2.0 `fakeram7_64x256` (64 words x 256 bits, single + port) per 8 lanes: the lanes are SIMD, every lane reads the same register index, so one 256-bit access serves eight + lanes' 32-bit reads. The macro's LEF and Liberty come with the ORFS ASAP7 platform (ABKGroup FakeRAM2.0, 7 nm, + 0.70 V; area 1,517 um^2 for the 64 x 256, 365 um^2 for the 256 x 34); the placement, the wiring and the clock tree + see the macro as a real block. Its dynamic energy is NOT taken from the FakeRAM Liberty (a placeholder, 1.345 per + clock edge identical for every size, "VALUES NOT REALISTIC" in the generator's own config); section 2.3 replaces it + with a published macro energy band and the gate-level simulation supplies the exact access counts. +- **The instruction memory is two `fakeram7_256x34` macros** shared by every lane of the core (40 bits used of 68). +- **A single-port macro is time-multiplexed over a 5-phase slot:** read dst, read src, read src2 only when the op + needs it (mad, lop3, sel, mm8, shfla), execute, write back. A core retires LANES lane-ops per 5 cycles; throughput is + bought with lane count (die area), the cheapest resource the chip has, not with ports. +- **Every unit's operands are isolated** (AND-gated by its own select), so an unused unit does not toggle: adding a + family's unit costs leakage and a wider result mux, not switching on every op. The k lane's core evaluates every unit + every cycle; this is one of the reasons its figure is not a lower bound. +- **The era's draws are registers** (fold constants, program length, the select tree); a family epoch writes them. +- Clock 1,500 ps (667 MHz) at the TC corner, as the k lane; nothing pipelined beyond the slot; two builds: 8 lanes + (placed and routed) and 32 lanes (synthesised), each in two variants: `full` (every bank entry) and `base` (the 10 + genesis families with the load and the fold, the same microarchitecture): the difference between the two is what the + bank adds to the chip. + +### 1.3 What the card pays for the same instruction (the GPU side, measured) + +The 5090's pJ per counted op at stock and at the 1,300 MHz lock (15.1a): add-class 11.3 / 6.2, mul and mad 13.9 / 8.3, +mulhi 39.6 / 21.0, prmt 22.3 / 11.5, lop3 24.1 / 13.0, shfl 55.8 / 29.4, the u8 tile 4.1 / 2.2 per MAC. The reserve +families by the measured NVIDIA step-cost ratio to the add step (design 4.2): shfla 1.53, popc 1.50, clz 1.63, bfe +1.54, shl and shr 0.75, sel and andn about 1.0 (approximate), mm8 by the tile row (32 MACs per lane per instruction). + +## 2. Method + +### 2.1 The flow + +ORFS on ASAP7 (7.5-track RVT, TC corner 0.70 V, NLDM), the default flow: Yosys with ABC, floorplan at 40 percent +utilisation (30 for the 32-lane core) with the macros placed by the flow's macro placer under the BLOCKS power grid, +global and detailed placement, CTS, global and detailed routing, OpenRCX parasitics. Power is OpenSTA `report_power` +under the VCD of a random-input gate-level simulation of the netlist (iverilog; every instruction field drawn by +`$random`, the window initialised with random words, a random returned word on every load), with a propagated 0.5 +activity as the cross-check. Synthesis-only rows (no wires, no clock tree) are marked; placed rows carry the SPEF. + +### 2.2 The activity and the steady state + +Each row is a tag: the op field fixed per family (`+fam=K`), or a drawn program: the class v4 draw over the 10 genesis +families (`mix`), the same with one load in 16 (`mixld`), the draw with two reserve families live at 4 points each +(`mix1`: shfla and mm8, the two dearest), every reserve family live at 4 points (`mix2`, the bank's bound, not a legal +draw), the W = 4 form (`mixw4`), the 64-instruction shape (`mix64`). Two run lengths per tag (500 and 2,000 clocks, 100 +and 400 slots) bracket the 326-clock load phase (reset, the era's registers, the 256-word program, the 64-word window +init) and the run-phase power is solved from the pair, as the k lane does. + +### 2.3 The SRAM macro energy (the one modelled term on the chip side) + +The FakeRAM Liberty's internal power is a placeholder, so the collector removes the macros' Liberty-attributed power +(reported separately per VCD with `report_power -instances`) and adds a modelled access energy times the exact +access count from the simulation (3 reads and 1 write per slot for a three-operand op, 2 reads and 1 write otherwise, +per 8 lanes; 2 imem reads per slot per core). The band: a 64 x 256 single-port macro at a 7 nm class node 3.5 pJ per +256-bit access (2.0 to 7.0); a 256 x 34 macro 1.5 pJ per access (0.8 to 3.0). Sources: Horowitz, ISSCC 2014 (45 nm: +an 8 KB SRAM read of 64 bits 10 pJ, 32 KB 20 pJ), scaled by the bits moved and by the energy-per-bit reduction from +45 nm to a 7 nm class node (about 0.1x to 0.2x, approximate, the same generation scaling the record applies to logic); +the k lane's "2 to 4 pJ per 32-bit read, approximate" for a 4 KB imem; CACTI-class estimates for a 16 Kbit macro at +7 nm (0.01 to 0.03 pJ per bit read, approximate). Every row carries the band; the low end is near a flop array with +perfect clock gating, the high end a conservative compiler macro. The macros' switching on their output nets (256 +bits into the lanes' latches) stays in the logic figure, measured. + +### 2.4 Node scaling (claimed) and what the method leaves out + +ASAP7 is a predictive 7 nm-class PDK; the row is stated at ASAP7 and scaled by the foundry's headline per-node +power reductions at the same speed (the k lane's factors, every one claimed): N5 = 0.70, N3 = 0.50, N2 = 0.36 of +ASAP7. Node-for-node against the 5090 (TSMC 4N, N5 class) is the N5 column; a node ahead is N3. Left out on the chip +side: the memory controller's queueing logic and the lane's address output (priced in the board model as the +controller die), test and clock distribution beyond the block; on the card side the 15.1a figure is the whole card's +marginal per counted op, which includes fetch, decode, operand collection and the register file, so the comparison +is the chip's whole lane (fetch, decode, window, units, network) against the card's whole lane. + +## 3. The rows: pJ per lane-op per family on the base core (synthesis only, 8 lanes) + +Synthesis only (no wires, no clock tree), 8 lanes, 1,500 ps, ASAP7 TC; "pJ logic" is OpenSTA's figure for the +standard cells under the VCD with the FakeRAM placeholder removed (its sequential and combinational parts beside it); +"pJ SRAM" the modelled macro term at the simulated access count (low / nominal / high, section 2.3); the per-op +figure is logic plus the nominal SRAM term, the band in brackets; the 5090 column is 15.1a (the reserve families +by the measured step ratio; the mix rows against the card's 10.3 pJ per op on the class v4 draw at the lock, 18.8 +unlocked by the ARX ratio, approximate). The full core: 95,678 cells and 3 macros (the base core 66,973 and 3 +macros: the bank adds 43 percent of the standard cells, 0.13 mW of leakage per 8 lanes, and 11 percent to the +energy of the class v4 draw on the same microarchitecture, the wider result mux and the leakage of the idle units). + +The full core (every bank entry): + +| Design | Stage | Family | Cells | Logic W | Leak W | pJ logic (seq / comb) | pJ SRAM (low / nom / high) | pJ/lane-op ASAP7 | N5 | N3 | N2 | 5090 pJ/op unlocked / lock | k N5 lock | k N3 lock | k N3 unlocked | +|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---| +| mf8full | power_synth | mix | 95678 | 0.00514 | 0.000395 | 4.82 (0.88 / 3.6) | 0.99 / 1.76 / 3.52 | 6.58 (5.81 to 8.34) | 4.61 | 3.32 | 2.39 | 18.8 / 10.3 | 0.447 | 0.322 | 0.176 | +| mf8full | power_synth | mixld | 95678 | 0.00513 | 0.000395 | 4.81 (0.88 / 3.6) | 0.986 / 1.75 / 3.5 | 6.56 (5.79 to 8.31) | 4.59 | 3.31 | 2.38 | 18.8 / 10.3 | 0.446 | 0.321 | 0.176 | +| mf8full | power_synth | mix1 | 95678 | 0.00495 | 0.000395 | 4.64 (0.89 / 3.4) | 1.01 / 1.8 / 3.59 | 6.43 (5.65 to 8.23) | 4.5 | 3.24 | 2.33 | 18.8 / 10.3 | 0.437 | 0.315 | 0.172 | +| mf8full | power_synth | mix2 | 95678 | 0.00436 | 0.000396 | 4.09 (0.88 / 2.8) | 1 / 1.78 / 3.56 | 5.87 (5.09 to 7.65) | 4.11 | 2.96 | 2.13 | 18.8 / 10.3 | 0.399 | 0.287 | 0.157 | +| mf8full | power_synth | mixw4 | 95678 | 0.00508 | 0.000395 | 4.76 (0.88 / 3.5) | 0.977 / 1.73 / 3.47 | 6.49 (5.74 to 8.23) | 4.55 | 3.27 | 2.36 | 18.8 / 10.3 | 0.441 | 0.318 | 0.174 | +| mf8full | power_synth | mix64 | 95678 | 0.00488 | 0.000395 | 4.58 (0.88 / 3.3) | 0.993 / 1.76 / 3.53 | 6.34 (5.57 to 8.1) | 4.44 | 3.19 | 2.3 | 18.8 / 10.3 | 0.431 | 0.31 | 0.17 | +| mf8full | power_synth | add | 95678 | 0.003 | 0.000396 | 2.81 (0.86 / 1.6) | 0.95 / 1.69 / 3.38 | 4.5 (3.76 to 6.18) | 3.15 | 2.27 | 1.63 | 11.3 / 6.2 | 0.508 | 0.366 | 0.201 | +| mf8full | power_synth | sub | 95678 | 0.00299 | 0.000396 | 2.8 (0.86 / 1.6) | 0.95 / 1.69 / 3.38 | 4.49 (3.75 to 6.18) | 3.14 | 2.26 | 1.63 | 11.3 / 6.2 | 0.507 | 0.365 | 0.2 | +| mf8full | power_synth | xor | 95678 | 0.00308 | 0.000396 | 2.89 (0.86 / 1.7) | 0.95 / 1.69 / 3.38 | 4.57 (3.84 to 6.26) | 3.2 | 2.3 | 1.66 | 11.3 / 6.2 | 0.516 | 0.372 | 0.204 | +| mf8full | power_synth | or | 95678 | 0.00183 | 0.000397 | 1.72 (0.84 / 0.5) | 0.95 / 1.69 / 3.38 | 3.41 (2.67 to 5.09) | 2.38 | 1.72 | 1.24 | 11.3 / 6.2 | 0.385 | 0.277 | 0.152 | +| mf8full | power_synth | rotl | 95678 | 0.00318 | 0.000396 | 2.99 (0.86 / 1.8) | 0.95 / 1.69 / 3.38 | 4.67 (3.94 to 6.36) | 3.27 | 2.36 | 1.7 | 11.3 / 6.2 | 0.528 | 0.38 | 0.208 | +| mf8full | power_synth | rotr | 95678 | 0.003 | 0.000396 | 2.81 (0.86 / 1.6) | 0.95 / 1.69 / 3.38 | 4.5 (3.76 to 6.18) | 3.15 | 2.27 | 1.63 | 11.3 / 6.2 | 0.508 | 0.366 | 0.201 | +| mf8full | power_synth | mul | 95678 | 0.00265 | 0.000396 | 2.48 (0.85 / 1.3) | 0.95 / 1.69 / 3.38 | 4.17 (3.43 to 5.86) | 2.92 | 2.1 | 1.51 | 13.9 / 8.3 | 0.352 | 0.253 | 0.151 | +| mf8full | power_synth | mulhi | 95678 | 0.00249 | 0.000396 | 2.33 (0.85 / 1.1) | 0.95 / 1.69 / 3.38 | 4.02 (3.28 to 5.71) | 2.81 | 2.03 | 1.46 | 39.6 / 21 | 0.134 | 0.0965 | 0.0512 | +| mf8full | power_synth | mad | 95678 | 0.00562 | 0.000395 | 5.26 (0.91 / 4) | 1.2 / 2.12 / 4.25 | 7.39 (6.46 to 9.51) | 5.17 | 3.72 | 2.68 | 13.9 / 8.3 | 0.623 | 0.449 | 0.268 | +| mf8full | power_synth | shfl | 95678 | 0.00266 | 0.000397 | 2.49 (0.86 / 1.3) | 0.95 / 1.69 / 3.38 | 4.18 (3.44 to 5.87) | 2.93 | 2.11 | 1.52 | 55.8 / 29.4 | 0.0995 | 0.0716 | 0.0377 | +| mf8full | power_synth | load | 95678 | 0.00433 | 0.000395 | 4.06 (0.86 / 2.8) | 0.95 / 1.69 / 3.38 | 5.75 (5.01 to 7.44) | 4.03 | 2.9 | 2.09 | 13.9 / 8.3 | 0.485 | 0.349 | 0.209 | +| mf8full | power_synth | fwd | 95678 | 0.004 | 0.000395 | 3.75 (0.86 / 2.5) | 0.95 / 1.69 / 3.38 | 5.44 (4.7 to 7.13) | 3.81 | 2.74 | 1.97 | 13.9 / 8.3 | 0.459 | 0.33 | 0.197 | +| mf8full | power_synth | prmt | 95678 | 0.00255 | 0.000397 | 2.39 (0.86 / 1.2) | 0.95 / 1.69 / 3.38 | 4.07 (3.34 to 5.76) | 2.85 | 2.05 | 1.48 | 22.3 / 11.5 | 0.248 | 0.179 | 0.0921 | +| mf8full | power_synth | lop3 | 95678 | 0.003 | 0.000397 | 2.81 (0.91 / 1.5) | 1.2 / 2.12 / 4.25 | 4.94 (4.01 to 7.06) | 3.46 | 2.49 | 1.79 | 24.1 / 13 | 0.266 | 0.191 | 0.103 | +| mf8full | power_synth | shfla | 95678 | 0.00302 | 0.000397 | 2.83 (0.9 / 1.6) | 1.2 / 2.12 / 4.25 | 4.95 (4.03 to 7.08) | 3.47 | 2.5 | 1.8 | 17.3 / 9.49 | 0.365 | 0.263 | 0.144 | +| mf8full | power_synth | popc | 95678 | 0.0017 | 0.000397 | 1.59 (0.85 / 0.38) | 0.95 / 1.69 / 3.38 | 3.28 (2.54 to 4.97) | 2.3 | 1.65 | 1.19 | 17 / 9.3 | 0.247 | 0.178 | 0.0976 | +| mf8full | power_synth | clz | 95678 | 0.00173 | 0.000397 | 1.62 (0.85 / 0.41) | 0.95 / 1.69 / 3.38 | 3.31 (2.57 to 5) | 2.32 | 1.67 | 1.2 | 18.4 / 10.1 | 0.229 | 0.165 | 0.0906 | +| mf8full | power_synth | bfe | 95678 | 0.00182 | 0.000397 | 1.71 (0.85 / 0.49) | 0.95 / 1.69 / 3.38 | 3.39 (2.66 to 5.08) | 2.38 | 1.71 | 1.23 | 17.4 / 9.55 | 0.249 | 0.179 | 0.0983 | +| mf8full | power_synth | shl | 95678 | 0.00234 | 0.000397 | 2.19 (0.86 / 0.98) | 0.95 / 1.69 / 3.38 | 3.88 (3.14 to 5.57) | 2.71 | 1.95 | 1.41 | 8.48 / 4.65 | 0.584 | 0.42 | 0.231 | +| mf8full | power_synth | shr | 95678 | 0.00221 | 0.000397 | 2.07 (0.85 / 0.84) | 0.95 / 1.69 / 3.38 | 3.76 (3.02 to 5.45) | 2.63 | 1.89 | 1.36 | 8.48 / 4.65 | 0.566 | 0.407 | 0.224 | +| mf8full | power_synth | sel | 95678 | 0.0028 | 0.000397 | 2.63 (0.91 / 1.4) | 1.2 / 2.12 / 4.25 | 4.75 (3.83 to 6.88) | 3.33 | 2.4 | 1.72 | 11.3 / 6.2 | 0.537 | 0.386 | 0.212 | +| mf8full | power_synth | andn | 95678 | 0.00234 | 0.000397 | 2.19 (0.86 / 0.98) | 0.95 / 1.69 / 3.38 | 3.88 (3.14 to 5.57) | 2.72 | 1.96 | 1.41 | 11.3 / 6.2 | 0.438 | 0.315 | 0.173 | +| mf8full | power_synth | mm8 | 95678 | 0.0036 | 0.000396 | 3.37 (0.91 / 2.1) | 1.2 / 2.12 / 4.25 | 5.5 (4.57 to 7.62) | 3.85 | 2.77 | 2 | 16.4 / 8.8 | 0.437 | 0.315 | 0.169 | + +The base core (the 10 genesis families on the same microarchitecture; the comparator): + +| Design | Stage | Family | Cells | Logic W | Leak W | pJ logic (seq / comb) | pJ SRAM (low / nom / high) | pJ/lane-op ASAP7 | N5 | N3 | N2 | 5090 pJ/op unlocked / lock | k N5 lock | k N3 lock | k N3 unlocked | +|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---| +| mf8base | power_synth | mix | 66973 | 0.00443 | 0.000263 | 4.15 (0.88 / 3) | 0.99 / 1.76 / 3.52 | 5.91 (5.14 to 7.66) | 4.13 | 2.98 | 2.14 | 18.8 / 10.3 | 0.401 | 0.289 | 0.158 | +| mf8base | power_synth | mixld | 66973 | 0.00439 | 0.000263 | 4.12 (0.88 / 3) | 0.986 / 1.75 / 3.5 | 5.87 (5.1 to 7.62) | 4.11 | 2.96 | 2.13 | 18.8 / 10.3 | 0.399 | 0.287 | 0.157 | +| mf8base | power_synth | add | 66973 | 0.00246 | 0.000264 | 2.31 (0.86 / 1.2) | 0.95 / 1.69 / 3.38 | 4 (3.26 to 5.68) | 2.8 | 2.01 | 1.45 | 11.3 / 6.2 | 0.451 | 0.325 | 0.178 | +| mf8base | power_synth | sub | 66973 | 0.00245 | 0.000264 | 2.29 (0.86 / 1.2) | 0.95 / 1.69 / 3.38 | 3.98 (3.24 to 5.67) | 2.79 | 2.01 | 1.44 | 11.3 / 6.2 | 0.449 | 0.324 | 0.178 | +| mf8base | power_synth | xor | 66973 | 0.00247 | 0.000264 | 2.32 (0.86 / 1.2) | 0.95 / 1.69 / 3.38 | 4 (3.27 to 5.69) | 2.8 | 2.02 | 1.45 | 11.3 / 6.2 | 0.452 | 0.326 | 0.179 | +| mf8base | power_synth | or | 66973 | 0.0016 | 0.000265 | 1.5 (0.84 / 0.42) | 0.95 / 1.69 / 3.38 | 3.19 (2.45 to 4.87) | 2.23 | 1.61 | 1.16 | 11.3 / 6.2 | 0.36 | 0.259 | 0.142 | +| mf8base | power_synth | rotl | 66973 | 0.00259 | 0.000264 | 2.43 (0.86 / 1.3) | 0.95 / 1.69 / 3.38 | 4.11 (3.38 to 5.8) | 2.88 | 2.07 | 1.49 | 11.3 / 6.2 | 0.464 | 0.334 | 0.183 | +| mf8base | power_synth | rotr | 66973 | 0.00251 | 0.000264 | 2.35 (0.86 / 1.3) | 0.95 / 1.69 / 3.38 | 4.04 (3.3 to 5.73) | 2.83 | 2.04 | 1.47 | 11.3 / 6.2 | 0.456 | 0.328 | 0.18 | +| mf8base | power_synth | mul | 66973 | 0.00229 | 0.000264 | 2.14 (0.85 / 1.1) | 0.95 / 1.69 / 3.38 | 3.83 (3.09 to 5.52) | 2.68 | 1.93 | 1.39 | 13.9 / 8.3 | 0.323 | 0.233 | 0.139 | +| mf8base | power_synth | mulhi | 66973 | 0.00217 | 0.000264 | 2.03 (0.84 / 0.96) | 0.95 / 1.69 / 3.38 | 3.72 (2.98 to 5.41) | 2.61 | 1.88 | 1.35 | 39.6 / 21 | 0.124 | 0.0893 | 0.0474 | +| mf8base | power_synth | mad | 66973 | 0.00472 | 0.000263 | 4.43 (0.9 / 3.3) | 1.2 / 2.12 / 4.25 | 6.55 (5.62 to 8.67) | 4.58 | 3.3 | 2.38 | 13.9 / 8.3 | 0.552 | 0.398 | 0.237 | +| mf8base | power_synth | shfl | 66973 | 0.00192 | 0.000265 | 1.8 (0.86 / 0.7) | 0.95 / 1.69 / 3.38 | 3.49 (2.75 to 5.17) | 2.44 | 1.76 | 1.27 | 55.8 / 29.4 | 0.083 | 0.0598 | 0.0315 | +| mf8base | power_synth | load | 66973 | 0.00359 | 0.000263 | 3.36 (0.86 / 2.3) | 0.95 / 1.69 / 3.38 | 5.05 (4.31 to 6.74) | 3.53 | 2.54 | 1.83 | 13.9 / 8.3 | 0.426 | 0.307 | 0.183 | + +Reading the rows. (1) The ALU-group families cost the chip 3.1 to 3.3 pJ per lane-op at N5 against the card's 6.2 +at the lock (k 0.51 to 0.53); or, popc, clz, bfe, prmt 2.3 to 2.9 (k 0.23 to 0.38); the multiply 2.9 (k 0.35); +mulhi 2.8 against the card's 21 (k 0.13); the shuffle 2.9 against 29.4 (k 0.10); mad is the dearest op for the +chip at 5.2 (three reads and two units, k 0.62) and the shifters the highest k (0.57 to 0.58) because the card +does them cheapest. (2) The SRAM term is 1.7 pJ nominal per lane-op (0.95 to 3.4): a third of the row; the +sequential term (the phase latches, the instruction register, the counters at five edges per op) 0.86 pJ; the rest +is the units and the macro output nets. (3) The reserve families are cheaper for the chip than the genesis +families: every live-family mix sits under the class v4 draw, and the bound with all eight live reads 11 percent +under it, because the card's dearest instructions (mulhi, shfl, mad) are the genesis ones. + + +## 4. The whole-hash energy and k per family, node-for-node and a node ahead + +The whole-hash shadow is 102,612 chip ops (the 102,100 counted shadow ops of the record plus the 512-instruction +base program; W = 4 adds 384 fold steps). The chip's cost is absolute (its own pJ at its node); the card's premium +on the same draw is the measured 0.652 microjoules at the lock, moved by the live family's measured step ratio at 4 +points of 79 (modelled). Node-for-node is N5 (the 5090's own class), a node ahead N3; every factor claimed. + +design mf8full stage power_synth; the class v4 draw on the core: 6.58 pJ per lane-op ASAP7, 4.61 N5, 3.32 N3 (band 4.07 to 5.84 at N5) + +| Family live (4 points of 79) or draw | Chip pJ per op N5 / N3 (band) | 5090 pJ per op at the lock | k N5 / N3 | Chip shadow per hash, microjoules N5 / N3 (the mix with the family live) | Change vs the class v4 draw | The card's premium per hash at the lock (modelled from the step ratio) | Change | +|---|---|---|---|---|---|---|---| +| the class v4 draw (10 genesis families) | 4.61 / 3.32 (4.07 to 5.84) | 10.3 | 0.45 / 0.32 | 0.473 / 0.340 | +0.0 percent | 0.652 | +0.0 percent | +| add (unit row, every instruction) | 3.15 / 2.27 (2.63 to 4.33) | 6.2 | 0.51 / 0.37 | 0.323 / 0.233 | -31.7 percent | 0.392 | -39.8 percent | +| sub (unit row, every instruction) | 3.14 / 2.26 (2.63 to 4.32) | 6.2 | 0.51 / 0.36 | 0.322 / 0.232 | -31.8 percent | 0.392 | -39.8 percent | +| xor (unit row, every instruction) | 3.20 / 2.30 (2.68 to 4.38) | 6.2 | 0.52 / 0.37 | 0.328 / 0.236 | -30.5 percent | 0.392 | -39.8 percent | +| or (unit row, every instruction) | 2.38 / 1.72 (1.87 to 3.57) | 6.2 | 0.38 / 0.28 | 0.245 / 0.176 | -48.2 percent | 0.392 | -39.8 percent | +| rotl (unit row, every instruction) | 3.27 / 2.36 (2.75 to 4.45) | 6.2 | 0.53 / 0.38 | 0.336 / 0.242 | -29.0 percent | 0.392 | -39.8 percent | +| rotr (unit row, every instruction) | 3.15 / 2.27 (2.63 to 4.33) | 6.2 | 0.51 / 0.37 | 0.323 / 0.233 | -31.7 percent | 0.392 | -39.8 percent | +| mul (unit row, every instruction) | 2.92 / 2.10 (2.40 to 4.10) | 8.3 | 0.35 / 0.25 | 0.299 / 0.216 | -36.7 percent | 0.525 | -19.4 percent | +| mulhi (unit row, every instruction) | 2.81 / 2.03 (2.30 to 3.99) | 21.0 | 0.13 / 0.10 | 0.289 / 0.208 | -38.9 percent | 1.329 | +103.9 percent | +| mad (unit row, every instruction) | 5.17 / 3.72 (4.52 to 6.66) | 8.3 | 0.62 / 0.45 | 0.531 / 0.382 | +12.3 percent | 0.525 | -19.4 percent | +| shfl (unit row, every instruction) | 2.93 / 2.11 (2.41 to 4.11) | 29.4 | 0.10 / 0.07 | 0.300 / 0.216 | -36.5 percent | 1.861 | +185.4 percent | +| load (unit row, every instruction) | 4.03 / 2.90 (3.51 to 5.21) | 8.3 | 0.48 / 0.35 | 0.413 / 0.297 | -12.6 percent | 0.525 | -19.4 percent | +| prmt live | 2.85 / 2.05 (2.34 to 4.03) | 11.5 | 0.25 / 0.18 | 0.464 / 0.334 | -1.9 percent | 0.662 | +1.5 percent | +| lop3 live | 3.46 / 2.49 (2.81 to 4.94) | 13.0 | 0.27 / 0.19 | 0.467 / 0.336 | -1.3 percent | 0.688 | +5.6 percent | +| shfla live | 3.47 / 2.50 (2.82 to 4.95) | 9.5 | 0.37 / 0.26 | 0.467 / 0.336 | -1.3 percent | 0.669 | +2.7 percent | +| popc live | 2.30 / 1.65 (1.78 to 3.48) | 9.3 | 0.25 / 0.18 | 0.461 / 0.332 | -2.5 percent | 0.669 | +2.5 percent | +| clz live | 2.32 / 1.67 (1.80 to 3.50) | 10.1 | 0.23 / 0.17 | 0.461 / 0.332 | -2.5 percent | 0.673 | +3.2 percent | +| bfe live | 2.38 / 1.71 (1.86 to 3.56) | 9.5 | 0.25 / 0.18 | 0.461 / 0.332 | -2.5 percent | 0.670 | +2.7 percent | +| shl live | 2.71 / 1.95 (2.20 to 3.90) | 4.7 | 0.58 / 0.42 | 0.463 / 0.333 | -2.1 percent | 0.644 | -1.3 percent | +| shr live | 2.63 / 1.89 (2.12 to 3.81) | 4.7 | 0.57 / 0.41 | 0.462 / 0.333 | -2.2 percent | 0.644 | -1.3 percent | +| sel live | 3.33 / 2.40 (2.68 to 4.81) | 6.2 | 0.54 / 0.39 | 0.466 / 0.336 | -1.4 percent | 0.652 | +0.0 percent | +| andn live | 2.72 / 1.96 (2.20 to 3.90) | 6.2 | 0.44 / 0.32 | 0.463 / 0.333 | -2.1 percent | 0.652 | +0.0 percent | +| mm8 live | 3.85 / 2.77 (3.20 to 5.34) | 8.8 | 0.44 / 0.31 | 0.469 / 0.337 | -0.8 percent | 0.699 | +7.2 percent | +| fwd (unit row, every instruction) | 3.81 / 2.74 (3.29 to 4.99) | 8.3 | 0.46 / 0.33 | 0.391 / 0.281 | -17.3 percent | 0.525 | -19.4 percent | +| the draw with shfla and mm8 live (measured mix) | 4.50 / 3.24 (3.96 to 5.76) | 10.3 | 0.44 / 0.31 | 0.462 / 0.333 | -2.2 percent | 0.652 | +0.0 percent | +| every reserve family live at 4 points (the bound, measured mix) | 4.11 / 2.96 (3.56 to 5.35) | 10.3 | 0.40 / 0.29 | 0.422 / 0.304 | -10.8 percent | 0.652 | +0.0 percent | +| the draw with one load in 16 | 4.59 / 3.31 (4.06 to 5.82) | 10.3 | 0.45 / 0.32 | 0.471 / 0.339 | -0.3 percent | 0.652 | +0.0 percent | +| W = 4: three fold steps per load | 4.55 / 3.27 (4.02 to 5.76) | 10.3 | 0.44 / 0.32 | 0.466 / 0.336 | -1.3 percent | 0.652 | +0.0 percent | +| the 64-instruction shape | 4.44 / 3.19 (3.90 to 5.67) | 10.3 | 0.43 / 0.31 | 0.455 / 0.328 | -3.7 percent | 0.652 | +0.0 percent | + +The whole-hash edge per joule with the class v4 shadow on the core (absolute: the chip pays its own pJ whatever the card does): +| Memory (E_mem, microjoules) | Chip E_hash N5 / N3 | vs 5090 lock 2.33 | vs 5090 stock 3.36 | vs 5080 lock 2.06 | vs M5 Max 1.40 | +|---|---|---|---|---|---| +| GDDR7 board (0.466) | 0.939 / 0.806 | 2.5x / 2.9x | 3.6x / 4.2x | 2.2x / 2.6x | 1.5x / 1.7x | +| HBM3 one stack (0.321) | 0.794 / 0.661 | 2.9x / 3.5x | 4.2x / 5.1x | 2.6x / 3.1x | 1.8x / 2.1x | +| SRAM N2 die at W = 1 (0.036) | 0.509 / 0.376 | 4.6x / 6.2x | 6.6x / 8.9x | 4.1x / 5.5x | 2.8x / 3.7x | + +The same edge on the base core (the genesis-only comparator): 2.6x / 3.0x microjoules per hash +on the GDDR7 board at N5 / N3, 3.8x / 4.4x against the 5090 at its lock; the bank costs the +adversary 5 percent of its edge (0.939 against 0.890 microjoules on the GDDR7 board, 2.5x against 2.6x), which is +the whole answer to the review's question in one number: the 180-day calendar costs a chip that carries the bank +about 11 percent of its shadow energy and 43 percent of its core cells, and no epoch costs it a part. + + +## 5. The placed rows (8 lanes, routed, SPEF) and the 32-lane core + +ROWS_PLACED + +## 6. The board: joules per valid hash and USD per sustained MH/s for the complete machine + +The complete machine: the memory devices at their modelled random-read energy and activate ceiling (chip-model-v3 +5.3, unmeasured; the 5090 reaches 82 percent of the GDDR7 figure, the sustained fraction here), the controller and +PHY die, the core die sized to retire the hash's ops at the memory's sustained rate (lanes at 667 MHz over 5 phases; +0.002 mm^2 per lane at N5 from the 8-lane core's floorplan at 40 percent utilisation scaled x0.55, approximate; USD +0.36 per mm^2 of N5 from sram-mirror's yield model; 50 uW of leakage per lane, the synthesis figure), power delivery +(PSU 92 percent, VRM 90 percent), cooling (3 percent), a board and assembly (USD 200), and one full node per 100 +machines (85 W and USD 1,500 shared: the state-derived dataset's host, priced as the review's rule 3 asks). The +"high" case takes the memory's lower ceiling (the 5090's measured 17.5 G on GDDR7; the JEDEC tFAW floor on HBM3), +the high read energy and the high SRAM term together. GPU rows: the 5090 at its lock 2.33 microjoules at 134.76 +MH/s, USD 1,999 plus USD 150 of rig share (USD 15.9 per MH/s; at the USD 3,000 street price 23.4); the 5080 at its +lock 2.06 at 71.20 MH/s, USD 999 plus 150 (USD 16.1 per MH/s); the discrete-GPU cohort by count (design 3.4 and +10.4: 8 GB 22 percent, 12 GB 22, 16 GB 28, 24 GB and up 16, the rest 10 and 11 GB) at its tuned points about 3.6 +microjoules (2.5 to 4.5, approximate: the 5070 and 5070 Ti 1.7 to 1.75 modelled, the 4070 3.58 measured, the 4090 +3.64, the 3090 6.4, the 9070 XT 8.1 measured) and about USD 18 per MH/s (14 to 28); the M5 Max 1.40 at 27.9 MH/s is +reported in section 4, not headlined. + +Node-for-node (N5), the full core at 4.61 pJ per lane-op: + +core 4.61 pJ per lane-op (band 4.07 to 5.84), 102612 ops per hash, leakage 50 uW per lane, 0.002 mm^2 per lane at N5 + +| Memory arrangement | Case | Sustained MH/s | Machine W | microjoules per hash (whole machine) | Core lanes | Core mm^2 (N5) | Capex USD | USD per MH/s | vs 5090 lock 2.33 (J / USD) | vs 5080 lock 2.06 | vs cohort 3.6 / USD 18 | +|---|---|---|---|---|---|---|---|---|---|---|---| +| GDDR7 board, 16 devices, 64 channels (the 5090 memory without the GPU) | nominal | 136 | 175 | 1.280 | 104,960 | 210 | 661 | 4.84 | 1.8x / 3.3x | 1.6x / 3.3x | 2.8x / 3.7x | +| GDDR7 board, 16 devices, 64 channels (the 5090 memory without the GPU) | low | 136 | 154 | 1.132 | 104,960 | 210 | 661 | 4.84 | 2.1x / 3.3x | 1.8x / 3.3x | 3.2x / 3.7x | +| GDDR7 board, 16 devices, 64 channels (the 5090 memory without the GPU) | high | 112 | 180 | 1.603 | 86,235 | 172 | 647 | 5.77 | 1.5x / 2.8x | 1.3x / 2.8x | 2.2x / 3.1x | +| HBM3 one stack (activate ceiling unmeasured; JEDEC tFAW floor 2.3 G) | nominal | 71 | 77 | 1.084 | 54,656 | 109 | 704 | 9.91 | 2.1x / 1.6x | 1.9x / 1.6x | 3.3x / 1.8x | +| HBM3 one stack (activate ceiling unmeasured; JEDEC tFAW floor 2.3 G) | low | 71 | 70 | 0.984 | 54,656 | 109 | 704 | 9.91 | 2.4x / 1.6x | 2.1x / 1.6x | 3.7x / 1.8x | +| HBM3 one stack (activate ceiling unmeasured; JEDEC tFAW floor 2.3 G) | high | 15 | 34 | 2.228 | 11,748 | 23 | 673 | 44.09 | 1.0x / 0.4x | 0.9x / 0.4x | 1.6x / 0.4x | +| HBM3 eight stacks (an H100-class package) | nominal | 566 | 559 | 0.987 | 435,713 | 871 | 3,079 | 5.44 | 2.4x / 2.9x | 2.1x / 3.0x | 3.6x / 3.3x | +| HBM3 eight stacks (an H100-class package) | low | 566 | 502 | 0.886 | 435,713 | 871 | 3,079 | 5.44 | 2.6x / 2.9x | 2.3x / 3.0x | 4.1x / 3.3x | +| HBM3 eight stacks (an H100-class package) | high | 122 | 217 | 1.772 | 93,987 | 188 | 2,833 | 23.18 | 1.3x / 0.7x | 1.2x / 0.7x | 2.0x / 0.8x | +| SRAM full store, one N2 reticle, 2 GiB at W = 1 (the strongest five-year chip) | nominal | 535 | 400 | 0.747 | 411,225 | 822 | 1,211 | 2.27 | 3.1x / 7.0x | 2.8x / 7.1x | 4.8x / 7.9x | +| SRAM full store, one N2 reticle, 2 GiB at W = 1 (the strongest five-year chip) | low | 609 | 403 | 0.662 | 468,572 | 937 | 1,252 | 2.06 | 3.5x / 7.8x | 3.1x / 7.8x | 5.4x / 8.8x | +| SRAM full store, one N2 reticle, 2 GiB at W = 1 (the strongest five-year chip) | high | 419 | 394 | 0.940 | 322,466 | 645 | 1,147 | 2.74 | 2.5x / 5.8x | 2.2x / 5.9x | 3.8x / 6.6x | + +Lifetime cost in USD per TH (10^12 hashes), capex spread over the life plus electricity at USD 0.08 per kWh: +| Machine | 0.5 y | 1 y | 2 y | 3 y | of which electricity | +|---|---|---|---|---|---| +| 5090 at the lock | 1.063 | 0.557 | 0.305 | 0.220 | 0.0518 | +| 5080 at the lock | 1.069 | 0.557 | 0.302 | 0.216 | 0.0458 | +| cohort card | 1.222 | 0.651 | 0.365 | 0.270 | 0.0800 | +| GDDR7 board, 16 devices, 64 channels chip | 0.335 | 0.182 | 0.105 | 0.080 | 0.0284 | +| HBM3 one stack chip | 0.653 | 0.338 | 0.181 | 0.129 | 0.0241 | +| HBM3 eight stacks chip | 0.367 | 0.194 | 0.108 | 0.079 | 0.0219 | +| SRAM full store, one N2 reticle, 2 GiB at W = 1 chip | 0.160 | 0.088 | 0.053 | 0.041 | 0.0166 | + +A node ahead (N3), the same machine: + +|---|---|---|---|---|---|---|---|---|---|---|---| +| GDDR7 board, 16 devices, 64 channels (the 5090 memory without the GPU) | nominal | 136 | 150 | 1.101 | 104,960 | 147 | 638 | 4.67 | 2.1x / 3.4x | 1.9x / 3.5x | 3.3x / 3.9x | +| GDDR7 board, 16 devices, 64 channels (the 5090 memory without the GPU) | low | 136 | 133 | 0.972 | 104,960 | 147 | 638 | 4.67 | 2.4x / 3.4x | 2.1x / 3.5x | 3.7x / 3.9x | +| GDDR7 board, 16 devices, 64 channels (the 5090 memory without the GPU) | high | 112 | 155 | 1.380 | 86,235 | 121 | 628 | 5.61 | 1.7x / 2.8x | 1.5x / 2.9x | 2.6x / 3.2x | +| HBM3 one stack (activate ceiling unmeasured; JEDEC tFAW floor 2.3 G) | nominal | 71 | 64 | 0.905 | 54,656 | 77 | 693 | 9.75 | 2.6x / 1.6x | 2.3x / 1.7x | 4.0x / 1.8x | +| HBM3 one stack (activate ceiling unmeasured; JEDEC tFAW floor 2.3 G) | low | 71 | 59 | 0.824 | 54,656 | 77 | 693 | 9.75 | 2.8x / 1.6x | 2.5x / 1.7x | 4.4x / 1.8x | +| HBM3 one stack (activate ceiling unmeasured; JEDEC tFAW floor 2.3 G) | high | 15 | 31 | 2.004 | 11,748 | 16 | 671 | 43.93 | 1.2x / 0.4x | 1.0x / 0.4x | 1.8x / 0.4x | +| HBM3 eight stacks (an H100-class package) | nominal | 566 | 458 | 0.808 | 435,713 | 610 | 2,985 | 5.27 | 2.9x / 3.0x | 2.5x / 3.1x | 4.5x / 3.4x | +| HBM3 eight stacks (an H100-class package) | low | 566 | 411 | 0.726 | 435,713 | 610 | 2,985 | 5.27 | 3.2x / 3.0x | 2.8x / 3.1x | 5.0x / 3.4x | +| HBM3 eight stacks (an H100-class package) | high | 122 | 189 | 1.548 | 93,987 | 132 | 2,812 | 23.02 | 1.5x / 0.7x | 1.3x / 0.7x | 2.3x / 0.8x | +| SRAM full store, one N2 reticle, 2 GiB at W = 1 (the strongest five-year chip) | nominal | 724 | 398 | 0.550 | 557,288 | 780 | 1,196 | 1.65 | 4.2x / 9.7x | 3.7x / 9.8x | 6.5x / 10.9x | +| SRAM full store, one N2 reticle, 2 GiB at W = 1 (the strongest five-year chip) | low | 828 | 402 | 0.485 | 636,578 | 891 | 1,236 | 1.49 | 4.8x / 10.7x | 4.2x / 10.8x | 7.4x / 12.1x | + +Reading. (1) The complete GDDR7 machine reads 1.8x the 5090 at its lock per joule node-for-node (1.5x to 2.1x) and +2.1x a node ahead (1.7x to 2.4x); against the 5080 at its lock 1.6x (1.3x to 1.8x) and 1.9x; against the cohort +card 2.8x and 3.3x. Per dollar of capex it is 3.3x the 5090 at MSRP (4.8x at the street price) and 3.7x the cohort, +because the board carries the same USD 320 of memory and USD 76 of core silicon where the card carries a 750 mm^2 +GPU. (2) The HBM3 one-stack machine is no better per joule than GDDR7 once the controller, the host and the power +train are in, and its dollars per MH/s are twice the card's at the modelled ceiling and 2.7x worse than the card at +the JEDEC tFAW floor: the "same DRAM" assumption is tested and the adversary's cheapest memory IS the card's memory, +so the GDDR7 board is the machine the economics must answer. (3) The N2 SRAM die at the hash's own width is the +strongest machine at 3.1x per joule node-for-node (2.5x to 3.5x) and 4.2x a node ahead, and 7x per dollar of +silicon, with the project cost (USD 100 M to 500 M, claimed) as its only hold. (4) Over a 3-year life the GDDR7 +machine's cost per TH is 0.080 USD against the 5090's 0.220 and the cohort's 0.270 (2.7x to 3.4x); at a 1-year life +3.1x; electricity is a quarter of the machine's lifetime cost and a fifth of the card's, so the per-joule edge is +the smaller half of the economic edge and the capex per MH/s the larger. The profitability surface the review asks +for (development, fleet capital, share, margin, electricity, pre-production, discounting, residual, operator and +manufacturer apart) takes these per-TH rows as its inputs; it is the economics lane's, not this file's. + + +## 7. The lifetime answer: what each transition needs, and what is credited + +The rule: a transition is credited only where the next entry needs a physical resource this design cannot supply +economically; firmware changes (a program, a register, a weight) are credited zero. "Performance loss" is the change in +the chip's energy per hash when the entry is live at its draw weight (4 points of 79 for a reserve family; the whole +program for an atom or a shape), from the rows of sections 3 and 4; the GPU's own change on the same draw is beside it. + +| Transition (the next epoch draws it) | Physical resource it needs | Does this design hold it? | Chip loss at the draw weight | The card's change on the same draw (measured ratio) | Credited obsolescence benefit | +|---|---|---|---|---|---| +| G1 to G7 (the ARX group) at any band weight | the ALU group | yes | 0 to +3 percent of the shadow energy across the band (the unit rows 3.1 to 3.3 pJ at N5 against the draw's 4.6) | 1.00 | 0 | +| G8, G9, G6 (mul, mulhi, mad) at any band weight | the multiplier and the third read | yes | -1 to +2 percent at B = 4 (mul 2.9, mulhi 2.8, mad 5.2 pJ at N5) | mul 1.1, mulhi about 3.4 (the card's dearest ALU op) | 0 | +| G10 shfl at its cap (8 points) | the lane butterfly | yes (per core) | 0 (2.9 pJ at N5, under the draw's mean) | 4.9x the add per op | 0 (the card pays 29.4 pJ for the move the chip pays about 1) | +| R1 shfla live (4 points) | a general lane crossbar | yes (per core; the one network a butterfly cannot emulate in one op) | 0 (2.9 pJ at N5, under the draw's mean)A | 1.53 (NVIDIA), 1.91 (Apple) | 0 | +| R2 perm live | a byte selector | yes | -1.9 percent | 1.30 | 0 | +| R3 popc and clz live | a popcount tree, a priority encoder | yes | -2.5 percent | 1.50 and 1.63 | 0 | +| R4 to R7 (bfe, shl and shr, sel, andn) live | a shifter, a mask, a select, an and-not | yes | -2.1 to -1.4 percent | 0.75 to 1.54 | 0 | +| R8 mm8 live | a u8 dot4 per lane (8 chip ops per card tile) | yes | -0.8 percent (3.9 pJ per dp4a at N5, 1.0 pJ per MAC, against the card's 2.2 pJ per MAC at the lock) | the tile: 2.43 the add step per card instruction (32 MACs) | 0 (the chip's MAC is cheaper than the card's by 4x to 30x on the public figures; this is the card's loss, not the chip's) | +| the op-mix band draw (B = 4 on injecting families) | nothing: the program | yes | within the family rows above | within 11 percent per instruction (shadow-k 6.2) | 0 | +| the fold constants draw | five registers | yes | 0 | 0 | 0 | +| the block shape draw (64, 128, 256) | the program-length register; 256 words of imem | yes | -3.7 percent at 64 (the imem term; 0 at 128 and 256) | 64 ran 2.5 to 3.5 percent faster than 256 on the 5090 and the M5 Max (measured) | 0 | +| the read-width draw W = 4 (16 bytes) | three fwd instructions per load (firmware); the same DRAM sector | yes | +0.3 percent per hash (384 fold steps at 3.8 pJ) | within 2.7 percent on the 5090 and the 9070 XT, within 1 on the M5 Max (measured) | 0 | +| the 64-register window | the 64 x 256 macro per 8 lanes | yes (built in) | 0 (it is the base) | modelled: occupancy to about half on a 5090 or 4090, the rate expected to hold (the hash lane) | 0 | +| the dataset atom draw A1 / A2 / A3 (mixer x4, x8, dr368) | nothing per hash; the per-window build is a program on the same core (section 7.1) | yes | 0 per hash; the build 0.72 J per window per machine at 1 GiB (0.2 mW averaged), 1.45 J at 2 GiB | the card's rate unmoved by the atom (within 0.1 MH/s on the 5090, measured) | 0 | +| a new op family outside the 18 (a bank refresh, a release) | a unit the die lacks | no: the chip emulates it from the 18 at the vendor penalty, as the cards do (1.5x to 2.4x per op, measured on the cards) or loses its 4 points | at 4 points of 79: at most 4 / 79 x (penalty - 1) of the shadow energy, about 2 to 6 percent (modelled) | the same emulation on every card that predates the release, 0 on a card with the native op | 0 unless the family is one the 18 cannot emulate; none proposed is | +| a new read atom W = 8 (32 bytes, `admissible: false` today) | seven fwd instructions per load; the same GDDR7 sector; on an SRAM die +0.3 nJ of wire per read (floor lane 3) | yes | +0.7 percent per hash (896 fold steps) | free by the measured rows (one sector per load on NVIDIA) | 0 | +| the dataset floor step (layer 2: 5.5 / 8.5 / 11.5 GiB) | device memory: 3 to 6 GDDR7 devices more, or 3 to 6 N2 reticles on the SRAM die | yes on DRAM (USD 60 to 120 more); the SRAM die's ticket rises USD 1,000 per step | 0 per joule on DRAM; the SRAM die's USD per MH/s unmoved (every die powered) | the tuned 5090 pays 4, 8 and 10 percent more energy per hash at 2, 4 and 8 GiB (measured) | 0 per joule; a capex ticket on the SRAM die only (floor lane 3) | + +Reading, before the numbers: nothing in the bank asks for a resource the design lacks, because the bank is public at +genesis and its whole op-family set is a few adders per lane and two networks per core. What the bank does to this +chip is the per-op cost of carrying the unit set (sections 3 and 5: the full core against the base core on the same +microarchitecture) and the leakage of the units that are not live, and that is the number the credit must come from. + +### 7.1 The per-window dataset build on the chip (the atoms' only cost) + +Under class v5 the dataset is rebuilt from the chain's state every window (3,600 s). A 1 GiB build is 157 G ops +(measured as 13.4 ms on a 5090; the record's row 7). On this core at 4.61 pJ per lane-op (N5) that is 0.72 J per +window per machine, 0.0002 W averaged over the window, against a machine of hundreds of watts: 0.0001 percent of +the machine's energy, the same for every atom within the atom's op count (x4 half of x8, dr368 about x4's). The +chip's node is the host the board model carries (one full node per 100 machines: 85 W and USD 1,500 shared); a +specialised machine with a host keeping the dataset current is inside the adversary model by the review's rule (3), +and this is its price: under 1 W and USD 15 per machine. + + +## 8. Energy resistance, economic resistance and response capability, stated separately + +STATEMENT + +## 9. Sources and what is owed + +- The GPU side: `docs/analysis/counter-asic-4-research.md` 15.1a (the 5090 microbench, 8 October 2026, measured); + the floor programme's tier table (`docs/design/class-v6-rotating-family.md` 10.4: the 5090 at the 1,300 lock 2.33 + microjoules at 134.76 MH/s and 312.5 W, the 5080 at the 1,100 lock 2.06 at 71.20 MH/s and 146.6 W, stock 3.36 and + 3.48, the 4070, 4090, 3090 and 9070 XT rows, the M5 Max 1.40); the reserve families' step costs, design 4.2 + (measured on the cards); the card population by count, design 3.4 (approximate). +- The chip side: this lane's RTL and flow (`tools/chip-model/mf/`), ORFS on ASAP7 (Clark et al., Microelectronics + Journal 2016; the 7.5-track RVT library, TC corner), FakeRAM2.0 macros (ABKGroup, the ORFS platform's + `fakeram7_64x256` and `fakeram7_256x34`, LEF and Liberty; the generator's own config marks its values "not + realistic", so only area, pins and placement are taken from it); the k lane's rows (`floor/shadow-k.md`, the + shuffle row 1.24 pJ per lane-op routed, relayed 16:0x BST) cited as published. +- The SRAM access energy band: Horowitz, "Computing's energy problem (and what we can do about it)", ISSCC 2014 + (45 nm: 8 KB SRAM 10 pJ per 64-bit read, 32 KB 20 pJ); the scaling to a 7 nm class node approximate; the k lane's + imem figure (2 to 4 pJ per 32-bit read, approximate); CACTI-class estimates (approximate). +- Node scaling (claimed): TSMC's technology pages for N5, N3E and N2, read 8 October 2026 (shadow-k section 7). +- The board: `docs/analysis/chip-model-v3.md` 5.3 and 5.5 (the GDDR7 and HBM3 random-read engines, the activate + ceilings, the static and controller allowances, the prices, all modelled or claimed; the 5090 at 82 percent of the + GDDR7 ceiling, measured); `floor/sram-and-floor.md` 2.1 (the SRAM die at W = 1, modelled); the power delivery, + cooling and host allowances are this file's (approximate; the sensitivity is stated beside each). +- The dataset build: the record's row 7 (a 1 GiB rebuild 157 G ops, 13.4 ms on a 5090, measured). + +Owed: the k lane's crossbar, scratch and tile rows (in place and route at 16:0x BST); the placed 32-lane core (the +8-lane core is placed; the 32-lane row is synthesis only); a real PDK memory compiler's figure for the two macros +(FakeRAM gives area and pins only); the 64-register window's GPU cost measured (the hash lane's generator line); +the HBM3 activate ceiling (unmeasured, the AWS F2 hour); the profitability surface of the review's rule (2) over the +lifetime rows of section 6, which this file states as cost per TH and leaves the NPV to the economics lane. +