igneum/docs/analysis/class-v6/floor/shadow-k.md
2026-10-08 20:42:29 +00:00

42 KiB

The shadow's k from RTL (floor lane 2, 8 October 2026)

Branch class-v6-floor-k from counter-asic-4 at 7618e729. Every chip-side number in this file comes from synthesis and place-and-route of RTL written for this lane (Yosys 0.68 plus OpenROAD, the ORFS docker image openroad/orfs:latest, the ASAP7 predictive PDK, run on igneum-build-4), with switching activity from a random-input gate-level simulation (iverilog on the routed netlist). Every GPU-side number is the record's measurement (docs/analysis/counter-asic-4-research.md 15.1a: PC 1, the RTX 5090, 20 probes at 60 s, 8 October 2026) and is not re-estimated here. The RTL, testbenches, flow configs and a Makefile that reproduces every row are under tools/chip-model/rtl/.

0. One page

(filled from the rows below; see section 2 for the table and section 4 for the edge)

1. The question and the identity

The chip's energy per hash is E_chip = E_mem + k x F (research file section 2): E_mem the memory system's reads, controller and static (0.466 microjoules per hash on the record's GDDR7 board, modelled), F the shadow's premium on the card (measured: 1.10 microjoules unlocked, 0.652 at the 1,300 MHz lock on the 5090 for class v4's 102,100 counted ops per hash), and k = e_chip / e_gpu the chip core's energy per forced op over the card's. The card's side of k is measured; the chip's side has been a claimed band (ALU 0.3 to 0.8, int8 tile 0.03 to 0.3, L2 hit 0.1 to 0.3, shuffle 0.4 to 0.7). This file replaces the claimed side with a synthesised one, family by family, and reads the mix that maximises the expected k under the invention lane's sound-form rules.

2. Method

2.1 What was built (RTL, tools/chip-model/rtl/rtl/)

Each family is a minimal chip-side shadow core: the smallest circuit a chip maker would have to build to run that family's instructions bit-exactly at one op per cycle per lane. The common lane shape is an instruction register, an 8 x 32-bit register window in flops (the card's lane holds its working set in a register file too; 8 registers is the class program's window), two or three read ports through muxes, the functional unit, and one write port. Nothing is shared across lanes and nothing is pipelined beyond one stage, so the figure is the datapath plus the minimum operand delivery: a floor for the chip, which makes the k it gives a floor too.

Family Module What one op is Ops per cycle Clock set (ps)
ARX lane_arx int32 add, sub, xor, or, rotl by immediate, rotr by register (op drawn per cycle; rows per op by fixing the op field) 1 1,000
MUL lane_mul 32 x 32 to 64: mul (low word), mulhi (high word), mad (src x src2 + dst) 1 1,500
PRMT lane_prmt byte permute: 4 output bytes from the 8 bytes of two registers by a 16-bit selector, sign-replicate bit as on the card 1 1,000
LOP3 lane_lop3 three-input logic by an 8-bit truth table, bitwise 1 1,000
FOLD lane_fold the index fold ((rotl(x * M, R) & WM) | OFF) & MASK with the era's constants held in registers 1 1,500
SHFL shfl32 the 32-lane xor-mask shuffle r[dst][lane] ^= r[src][lane ^ m] over a 1 KB window (32 lanes x 8 x 32 bits): a 5-stage butterfly 32 lane-ops 1,000
XBAR xbar32 the general 32-lane crossbar over the same window (any source lane per lane, a 5-bit select each): the upper bound on a shuffle network 32 lane-ops 1,000
SCRATCH scratch8k one random 32-bit read of a 2,048 x 32 (8 KB) flop array, a write on one cycle in eight: the chip's L1 as the pessimistic flop form 1 read 1,500
TILE tile8 the int8 8x8x8 tile C += A x B with int32 accumulators: 512 MACs per cycle 512 MACs 2,000

2.2 The flow

ORFS on ASAP7 (7.5-track RVT cells, the TC corner: 0.70 V, 0 C, NLDM), the default flow end to end: Yosys synthesis with ABC, floorplan at 40 percent utilisation, global and detailed placement, CTS, global and detailed routing, parasitic extraction (OpenRCX, the platform's rules). Power is OpenSTA's report_power on the routed design with its SPEF, under two activities: (a) the VCD of a random-input simulation of the routed netlist (ASAP7 has no cell Verilog models in the image, so Yosys builds the cell bodies from the liberty functions and writes the netlist as primitives; iverilog runs 3,000 to 4,000 cycles with every instruction field drawn by $random each cycle and a random 32-bit load into the window every 16th cycle), and (b) a propagated 0.5 activity on every input as the cross-check. The energy per op is the total power (internal, switching and leakage) times the clock period over the ops per cycle. Leakage is reported beside it; at these clocks it is a few percent.

2.3 Node scaling (approximate; every factor claimed from the foundry's own headline)

ASAP7 is a predictive 7 nm-class FinFET PDK (ASU and ARM, Clark et al., Microelectronics Journal 2016), not a foundry node, so the row is first stated at ASAP7 and then scaled by the foundry's published per-node power reductions at the same speed: N7 to N5 x0.70 (TSMC: "30 percent lower power"), N5 to N3E x0.72 (TSMC: "25 to 30 percent lower power", the midpoint), N3E to N2 x0.72 (TSMC: "25 to 30 percent lower power", the midpoint). So N5 = 0.70, N3 = 0.50 and N2 = 0.36 of the ASAP7 figure. These are the foundry's claims for a whole design at a fixed frequency, and a shadow core at a low clock could run at a lower voltage still; the N2 column is therefore the chip's best case from this method, not its floor. Sources in section 7.

2.4 What the method leaves out, on both sides

On the chip side the figure omits instruction fetch and decode (a chip would run the program from a small SRAM or a decoded instruction cache shared by many lanes, about 1 to 2 pJ per lane-instruction amortised across 32 lanes, approximate), the clock tree beyond the block's own, and the result's move to a memory address unit. On the card side the 15.1a figure is the whole card's marginal per counted op (the sleep floor subtracted), which includes the card's own fetch, decode, operand collection and register file. So the k here is the chip's datapath-and-window cost over the card's whole-lane cost: a FLOOR on the chip's cost and so a floor on k.

3. The rows: per-unit floors (the minimal lane per family)

Every row: ASAP7 routed, SPEF, OpenSTA report_power under the random-input gate-level VCD (every pin annotated, 0 unannotated), the TC corner (0.70 V). "pJ/op" is total power (internal + switching + leakage) times the clock period over the ops per cycle. The propagated-0.5 cross-check is in table.csv; it agrees within 2x on the logic lanes and overestimates the multiplier lanes 50x (OpenSTA's statistical propagation through a multiplier is not a measurement), so the VCD row is the row. The GPU side is 15.1a: pJ per counted op on the RTX 5090 at stock and at the 1,300 MHz lock, and 6.9 pJ per counted op on the Apple M5 Max (the class v4 shadow, measured). k is absolute: the chip's own pJ per op at the node over the card's at its point.

Family (chip RTL) Cells (with fill) pJ/op ASAP7 pJ/op N5 pJ/op N3 pJ/op N2 5090 pJ/op stock / lock k, N3 chip vs 5090 stock / lock k, N3 vs M5 Max 6.9 k, ASAP7 unscaled vs lock Label
ARX add 11,631 2.16 1.51 1.09 0.78 11.3 / 6.2 0.096 / 0.18 0.16 0.35 synthesised; N3 and N2 scaled (claimed)
ARX sub 2.14 1.50 1.08 0.78 11.3 / 6.2 0.095 / 0.17 0.16 0.35
ARX xor 2.09 1.46 1.05 0.76 11.3 / 6.2 0.093 / 0.17 0.15 0.34
ARX rotl (immediate) 2.15 1.50 1.08 0.78 11.3 / 6.2 0.095 / 0.17 0.16 0.35
ARX rotr (by register) 2.13 1.49 1.07 0.77 11.3 / 6.2 0.095 / 0.17 0.16 0.34
or (lossy; the values saturate to ones and the activity falls) 1.45 1.01 0.73 0.53 11.3 / 6.2 0.065 / 0.12 0.11 0.23
ARX random mix 2.24 1.57 1.13 0.81 11.3 / 6.2 0.10 / 0.18 0.16 0.36
mul (32 x 32, low word) 19,603 1.35 0.94 0.68 0.49 13.9 / 8.3 0.049 / 0.082 0.16
mulhi (high word; the same multiplier) 1.35 0.94 0.68 0.49 39.6 / 21.0 0.017 / 0.032 0.064
mad (src x src2 + dst; three reads) 3.29 2.30 1.66 1.19 13.9 / 8.3 0.12 / 0.20 0.40
mul random mix 1.53 1.07 0.77 0.56 13.9 / 8.3 0.055 / 0.093 0.18
index fold (x M, rotl R, masks; era constants in registers) 14,965 2.37 1.66 1.19 0.86 13.9 / 8.3 (read against int_mul) 0.086 / 0.14 0.29
prmt (byte permute) 6,744 1.27 0.89 0.64 0.46 22.3 / 11.5 0.029 / 0.056 0.11
lop3 (8-bit truth table) 7,274 1.42 0.99 0.72 0.52 24.1 / 13.0 0.030 / 0.055 0.11
32-lane xor-mask shuffle (butterfly over the 1 KB window), per lane-op 181,580 1.24 0.87 0.63 0.45 55.8 / 29.4 0.011 / 0.021 0.042 routed with SPEF; 15:1x UK
32-lane general crossbar, per lane-op ROW_XBAR
8 KB scratch, one random 32-bit read of a 2,048 x 32 flop array (the pessimistic form of a chip's L1; an SRAM macro reads lower) 868,159 207 145 104 75 2,400 / 1,400 per L2 hit 0.043 / 0.074 0.15 routed with SPEF, 19:4x UK; the card's shared-memory read is unmeasured (owed)
int8 8x8x8 tile, per MAC ROW_TILE

Reading the floors: a lane's add costs the chip about 2.2 pJ at ASAP7 and 1.1 at N3, against the 5090's 6.2 at the lock and 11.3 at stock, so even the floor is not a tenth of the card's cost at the knee, and the claimed "2 to 5 pJ for a SIMD array at N5" (15.1) was the right order for the unscaled lane with nothing around it. The multiplier is the cheapest unit per op relative to the card (the 5090 pays 8.3 pJ for mul and 21 for mulhi, the chip 0.68 for either, because the high word falls out of the same array), and mad is the dearest for the chip (three register reads and two units). The fold costs a chip one multiply and a rotate: 1.2 pJ per address at N3, or 0.15 nJ per hash over 128 loads, a third of a percent of the GDDR7 board's 0.466 microjoules.

4. The headline row: the programmable sequencer core

The per-unit lanes of section 3 are floors for a chip that cannot exist under class v6: layers 1 and 3 (per-era op-mix, program-length and read-width draws, family epochs) kill any fixed lane, so the chip that competes is a programmable core. core_v6 (tools/chip-model/rtl/rtl/core_v6.v) is the minimal in-order SIMD sequencer that executes the drawn class v6 shadow program: a 256 x 32-bit instruction memory (a flop array) sized to the drawn program, a program counter wrapping at the era's drawn length, fetch into an instruction register, decode, the era's parameter registers (M, R, WM, OFF, MASK, N), and per lane a 32 x 32-bit register file (flops) with three read ports and one write port and every class unit: add, sub, xor, or, rotl, rotr, mul, mulhi, mad, prmt, lop3, the xor-mask shuffle across the lanes (a butterfly), and the load (the index fold on the address path, the returned word written next cycle). One instruction per cycle for every lane. The program is 256 instructions drawn with the class v4 weights (NONLOAD_WEIGHTS), with a variant at one load per 16 instructions. Two builds: 8 lanes (the fast row) and 32 lanes (the imem amortised over 32), and a 32-lane build with a 16-register file (the sensitivity).

The activity: the gate-level random-input VCD of the synthesised netlist; two run lengths (150 and 800 cycles) bracket the 260-cycle program-load phase, and solving the pair gives the steady-state run power (the load phase draws 13.9 mW, the run phase 36.9 mW for 8 lanes at 1.5 ns).

Row Stage Cells pJ per lane-op ASAP7 N5 N3 N2 5090 stock / lock / M5 Max pJ per op k absolute at N3 vs stock / lock / M5 Max k at N2 k unscaled ASAP7 vs lock Label
core, 8 lanes, 32 registers synthesis only (no wires, no clock tree) 186,443 6.9 4.8 3.5 2.5 11.3 / 6.2 / 6.9 0.31 / 0.56 / 0.50 0.22 / 0.40 / 0.36 1.1 synthesised; 14:0x UK
of which the sequential term (register file, imem and IR clock pins, no clock gating) 2.4 1.7 1.2 0.9
of which the units, the read muxes and the butterfly 4.5 3.1 2.3 1.6
core, 8 lanes, 32 registers, ungated placed and routed, SPEF, clock tree (one run length of 300 cycles, the load phase subtracted, about plus or minus 10 percent) 444,478 11.3 7.9 5.6 4.1 11.3 / 6.2 / 6.9 0.50 / 0.91 / 0.82 0.36 / 0.66 / 0.59 1.8 placed 16:0x UK on a rented pod; +64 percent over synthesis (wires, and a clock tree of 2.5 pJ per lane-op that gating removes)
core, 32 lanes, 32 registers synthesis only (steady state from 150 and 400 run cycles) 600,381 5.55 3.9 2.8 2.0 11.3 / 6.2 / 6.9 0.25 / 0.45 / 0.41 0.18 / 0.32 / 0.29 0.90 synthesised; 15:2x UK
core, 32 lanes, 16 registers synthesis only 443,258 4.2 2.9 2.1 1.5 11.3 / 6.2 / 6.9 0.18 / 0.34 / 0.30 0.13 / 0.24 / 0.22 0.68 synthesised; one run length, about plus or minus 10 percent; 15:0x UK
the bare ARX lane (section 3, the floor) routed 11,631 2.2 1.5 1.1 0.8 11.3 / 6.2 / 6.9 0.10 / 0.18 / 0.16 0.07 / 0.13 / 0.11 0.35 the lower bound

Reading: fetch, decode, a 32-register file and the full unit set cost a chip 3.1x the bare lane, and the core's k at the 1,300 lock is 0.56 at N3 (0.40 at N2), inside the record's claimed 0.3 to 0.8 band and at its centre, with the bare lane's 0.18 as the lower bound. Two corrections pull opposite ways: placement adds wires and a clock tree (+20 to +40 percent on a design like this, approximate; the placed rows read it) and a chip maker gates the register-file clock (one of 32 registers is written per cycle; gating removes about 2.0 of the 2.4 pJ sequential term, approximate), so the net figure for the re-fold is 7.0 / 4.9 / 3.5 / 2.5 pJ per lane-op at ASAP7 / N5 / N3 / N2 (the synthesis-only figure within the rounding). The imem is amortised over 8 lanes in this row and over 32 in the core32 row.

4a. The design sweep: what a class could add to the core's cost (the coordinator's order, 14:3x UK; rows 14:5x)

Synthesis-only (no wires, no clock tree), 8 lanes unless stated, the same corner and scaling; the steady-state run power solved from two run lengths (150 and 600 cycles after the program load). The GPU side per knob is the hash lane's: knob 3 measured on a rented 5090 and 4090 (RunPod, 14:28 to 14:39 UK), knobs 1 and 4 modelled until a generator line exists (a 64-entry window is a new ISA: an init rule and a fold rule for the extra registers, about half a day; the select tree is a new instruction kind), knob 2 has no GPU side (the warp's shuffle already spans 32 lanes).

Variant Cells pJ per lane-op ASAP7 of which clocking (RF, imem, IR; no gating) N5 N3 N2 k at N3 vs 5090 stock / lock / M5 Max GPU side
base: 32 registers, 256-entry imem (the headline row) 186,443 6.9 2.4 4.8 3.5 2.5 0.31 / 0.56 / 0.50 measured (class v4)
(1) 64-register file (a 40-bit instruction word) 267,731 9.7 3.75 6.8 4.9 3.5 0.43 / 0.78 / 0.70 new ISA; 64 live registers takes a 5090 or 4090 thread to about 110 of 255, occupancy to about half; under the latency-bound chain the rate is expected to hold and the energy to move little (the hash lane, modelled); the unmeasured term is the per-lane register traffic
(3) 1,024-entry imem, the program drawn at 1,024, as built (a flop array) 324,543 12.0 6.0 8.4 6.1 4.4 0.53 / 0.97 / 0.87 measured: the 1,024 block at 27 passes costs the 5090 1.58x the energy per hash (4.96 against 3.14 microjoules at stock, 111 against 141 MH/s at the 575 W cap) and the 4090 1.57x (7.01 against 4.48, the rate held at 62.5 MH/s), for 4x the shadow instructions: 0.40x per instruction
(3) the same with the imem as a 4 KB SRAM macro shared by the lanes (2 to 4 pJ per 32-bit read, approximate) about 7.2 about 1.9 5.0 3.6 2.6 about 0.32 / 0.58 / 0.52 the same
(4) the drawn select tree (the era's 16-entry op permutation ahead of decode; every unit evaluated every cycle, as in the base) 186,870 6.85 2.4 4.8 3.4 2.5 0.30 / 0.55 / 0.50 the units' microbench sum (approximate)
(2) 32 lanes, 32 registers (the butterfly across 32; the imem amortised over 32) 600,381 5.55 1.5 3.9 2.8 2.0 0.25 / 0.45 / 0.41 no GPU knob
(2') 32 lanes, 16 registers (the register-file sensitivity the other way; one run length of 150 cycles, the load phase subtracted at the 8-lane ratio, about plus or minus 10 percent) 443,258 4.2 2.3 2.9 2.1 1.5 0.18 / 0.34 / 0.30 no GPU knob
(5) all four together (32 lanes, 64 registers, 1,024 imem, the select tree) ROW_CORE32ALL

Reading, for the founder's "under 2x at the lock" (which needs k near 0.9 on the GDDR7 board): the 64-register window is the one robust knob, because its cost is per lane and a chip cannot share it (+2.8 pJ per lane-op at ASAP7, +0.22 of k at the lock). The long block's cost is instruction memory, which a chip shares across its lanes as SRAM, so most of its 0.97 as built is the flop array's clock and the honest figure is about 0.6; the card meanwhile pays 1.58x the energy per hash for it (measured), so on the GDDR7 board at stock the chip reads 4.96 / (0.466 + 408,400 x 3.6 pJ) = 2.6x (1.7x on the flop-array row, which is not a chip anyone builds). The select tree costs the chip nothing because every unit already evaluates every cycle in the base core. (1) + (3) together reach about 10 pJ per lane-op at ASAP7 with the SRAM imem (5.0 at N3, k about 0.81 at the lock, 0.45 at stock), 14.8 as built (k about 1.2); so k 0.85 is reached only on the flop-array reading, and the DRAM board under 2x at the lock needs the window, the long block and the knee together and holds only while the chip's imem stays unamortised, which it does not. The GPU pays nothing for the window until occupancy binds, 1.58x per hash for the long block (0.40x per forcing instruction), and nothing for the select tree.

4b. The node column: which part of the edge is the node (the coordinator's order, 15:0x UK)

The record compares a chip scaled to N3 or N2 against a TSMC 4N card (N5 class; the M5 Max is N3), so part of the edge is the node. Factors as in 2.3, every one claimed. Absolute; the GDDR7 board at the lock = 2.33 / (0.466 + 102,100 x e_chip).

Core pJ per lane-op ASAP7 / N5 / N3 / N2 k at the lock, N5 / N3 / N2 GDDR7 board at the lock, N5 / N3 / N2
base (32 registers, 256 imem) 6.9 / 4.8 / 3.5 / 2.5 0.78 / 0.56 / 0.40 2.4x / 2.8x / 3.2x
64-register file 9.7 / 6.8 / 4.9 / 3.5 1.09 / 0.78 / 0.56 2.0x / 2.4x / 2.8x
all four together ROW_CORE32ALL_NODE

One line: of the 2.8x at k 0.56, the N5-to-N3 node step is worth 0.4x (2.4x node-for-node, a factor of 1.17, claimed); the rest is the memory system (3.6x at zero shadow at the lock, modelled) less what the class v4 shadow takes back on the card's own node (3.6x to 2.4x), which is the design. Node-for-node the base core already sits at k 0.78 and the 64-register core at 1.09, so "near 0.9" is reached node-for-node by the window alone; what it does not survive is the node step a chip project would buy (an N3 core gives back 0.4x, an N2 core 0.8x).

4c. The adversary's 64-register core: is the window a defence? (the coordinator's order, 15:3x UK)

Every row in this section is a MODEL of a chip core, never a lower bound on what a chip maker can build; the synthesis gives the cost of the circuit as drawn, and a better circuit is always possible.

The live-state analysis (tools/chip-model/rtl/flow/livestate.py, run on build-3: programs drawn as the core testbench draws them, the class v4 op weights, a load on one instruction in 16 as the dependent memory wait, dst and src uniform over the window, the result fold reading every register at the end of the block; 64 drawn programs, 1,024 waits per row):

Window R Live values at a wait (mean, min to max) Of which necessary (reach a later address or the result, transitively) Dead writes per block
8 (the class ISA) 7.0 of 8 (6 to 7) 6.9 2.9 percent
32 30.5 of 32 (29 to 31) 30.0 2.9 percent
64 61.7 of 64 (59 to 63) 61.0 2.8 percent
64 at a 1,024-instruction block 61.5 of 64 (58 to 63) 60.6 3.1 percent

So under a fold that reads every register, 95 percent of the window is live AND necessary across every memory wait: the adversary cannot shrink the state it keeps by liveness, and recomputing a value instead of keeping it costs the dependent chain that produced it (every value feeds the result transitively). The window is a defence ONLY because of the fold rule; a fold that read 8 of the 64 registers would let the chip drop the rest (the dead fraction would rise toward the fraction never read before the fold), so the fold-reads-all rule is the design rule that goes with the window.

What the adversary can do with the state it must keep is make it cheaper per access, not smaller. The GPU-shaped row (4a, 64 registers in flops, every flop clocked every cycle, three 64:1 read muxes) is 9.7 pJ per lane-op at ASAP7. The forms a chip maker would use:

Form of the 64-register state (per lane, 256 bytes) pJ per lane-op ASAP7 N3 k at the lock (N3) Label
flops, no clock gating, 64:1 read muxes (the 4a row) 9.7 4.9 0.78 synthesised; a model
flops with the register-file clock gated (one of 64 registers written per cycle; the ICG cells allowed back in and inferred by Yosys, 320 gates) 6.2 (225,441 cells; sequential 0.2) 3.1 0.50 (0.70 node-for-node) synthesised 16:0x UK; a model
the same gating on the 32-register base, for the penalty 4.5 (156,833 cells; sequential 0.15) 2.3 0.37 (0.51 node-for-node) synthesised 16:0x UK; a model
the gated 32-register base PLACED AND ROUTED (SPEF, clock tree, 379,633 cells; steady state solved from 150 and 600 run cycles) 6.7 (+49 percent over synthesis) 3.4 0.54 (0.76 node-for-node) placed 19:5x UK; the GDDR7 board at the lock 2.4x node-for-node, 2.8x a node ahead: the morning's headline figures to the digit
the gated 64-register window core PLACED AND ROUTED (563,339 cells; parasitics estimated from global routing, the SPEF lost to a full disk; steady state solved from 150 and 600 run cycles; plus or minus 15 percent) 9.45 (+52 percent over synthesis; N5 6.6, N2 3.4) 4.8 0.77 (1.07 node-for-node, 0.55 at N2) placed 21:3x UK; the GDDR7 board at the lock 2.0x node-for-node, 2.45x a node ahead, 2.9x two ahead; the window's residual against the placed base +2.75 pJ per lane-op, +0.23 of k at the lock
latch-based register file (the clocked element halved; about 30 percent under the gated flop file, approximate) about 0.7 x the gated row modelled
SRAM-banked state shared across time-multiplexed lanes (one execution port serving many lanes' streams in turn, each lane's 256 bytes in a bank): modelled 8.5 to 10.5 from the access energies; BUILT as core_tm (8 lanes x 64 registers in banks, one port, round-robin, ungated): 8.8 pJ per lane-op synthesised (173,426 cells), against the SIMD ungated 9.7 8.8 (built) 4.4 0.71 synthesised 16:2x UK: port sharing saves 0.9 pJ of units and the bank-select muxes take most of it back; NOT the lever; the multi-family adversary lane's macro-window core reads 5.9 (its FakeRAM term modelled 2.0 to 7.0 pJ per access), within 5 percent of the gated flop row, so the file's form is not the lever either
values recomputed instead of kept not available: 95 percent of the window is necessary (above) measured on drawn programs

The defence, then, is the penalty that remains after the adversary's best form: the gated 64-register file against the gated 32-register file (the two synthesised rows above when they land; measured: 6.2 against 4.5 pJ per lane-op at ASAP7 synthesised, a penalty of 1.7 pJ; PLACED 9.45 against 6.7, a penalty of 2.75 pJ, 1.4 at N3, +0.23 of k at the lock, 0.54 to 0.77 at N3 and 0.76 to 1.07 node-for-node; the analytic estimate had been 1.5 pJ). So on placed rows the window takes the GDDR7 board at the lock from 2.4x to 2.0x node-for-node and from 2.8x to 2.45x a node ahead, for at most 5 percent per load on the card. Two corrections this forces: the honest adversary's BASE core is the gated one (k 0.37 at N3, 0.51 node-for-node), under the ungated 0.56 and 0.78 of section 4, which are the GPU-shaped core a maker would not build; and placement costs more than the +20 to +40 percent estimated (the ungated placed base reads 11.3 against 6.9 pJ: wires plus a 2.5 pJ clock tree that gating removes), so the placed gated rows (on a rented pod, 17:30 UK) are the figures to serve. On the 32-lane core the same penalty applies per lane (the register file does not amortise), so the window moves the 32-lane core from k 0.45 to about 0.57 at the lock at N3 (0.63 to about 0.80 node-for-node).

The GPU side (the hash lane's hand): the compiled allocation of the 64-register measurement pack (ptxas registers per thread, local-memory spill bytes, occupancy) and the rate beside the 8-register base, clock 18:00 UK; until then the modelled reading stands: about 110 of 255 registers per thread, occupancy about half, the rate expected to hold under the latency-bound chain (the 5090 hides about 330,000 ops per hash before compute binds) and the energy to move little, the per-lane register traffic the unmeasured term.

The measured GPU side (the hash lane, 16:1x to 16:4x UK, RunPod secure pods, driver 580, the kit worker, 250 x 2^24, nvidia-smi 1 Hz; ptxas from nvcc 12.8 -Xptxas -v on the pack's kernel; pods destroyed, USD 1.22):

Card, pack MH/s W microjoules per hash registers per thread (ptxas) spill blocks per SM (occupancy) per load Label
5090, the base (mx8-devnet-epoch0) 141.74 303.1 2.139 30 0 B 24 (4,080 warps) 16.7 nJ measured
5090, the window, arithmetic-only (hl-reg64, twice the base's work by construction) 80.38 308.6 3.839 96 0 B 20 (83 percent) 15.0 nJ measured: level per unit of work
5090, the window, full chain (hl-reg64c: every load's address mixes all 64 registers) 70.96 320.3 4.513 88 0 B 20 (83 percent) 17.6 nJ (+5 percent) measured
4090, the base 62.67 208.9 3.333 29 0 B 24 26.0 nJ measured
4090, the window, arithmetic-only 31.57 210.3 6.663 104 0 B 16 (67 percent) 26.0 nJ measured: level
4090, the window, full chain 31.38 216.5 6.898 87 0 B 20 (83 percent) 27.0 nJ (+4 percent) measured

So the card's side of the window defence is at most 5 percent per load: no spill on either card in either form, 88 to 104 registers per thread, occupancy 67 to 83 percent, and the rate per unit of work held within 5 percent under the latency-bound chain. The sound class form is the full chain (the arithmetic-only fold fails the liveness rule; class string +reg64c, pack hl-v6-win with check_window_liveness in its suite). The chip's side (this section's gated rows) therefore carries the whole defence.

4d. The connected-state variant (cs64s27x16, the connected-state lane's structure) on the adversary's core

The connected-state lane's program (a 64-register window; per step a load whose address register is the previous block's last dst, the word landing in m_j, then a 27-instruction block whose first instruction reads m_j and every later one draws its src from the block's last four dsts, the last instruction injecting; 16 steps of text, 448 instructions, 16 passes per block; its liveness tool: 63 of 64 live at every address, about 11 registers in the per-step dependent chain, about 20 touched per block) priced on the gated 64-register core with a 512-entry imem, the program drawn by those rules in the testbench (CS mode of tb_core_common.vh), synthesis-only:

Row Cells pJ per lane-op ASAP7 N5 N3 N2 k at the lock N5 / N3 / N2 k at stock N5 / N3
cs64s27x16 on the gated 64-register core, 512 imem 253,059 6.3 4.4 3.2 2.3 0.71 / 0.51 / 0.37 0.39 / 0.28
the class v4 draw on the gated 64-register core, 256 imem (4c) 225,441 6.2 4.3 3.1 2.2 0.70 / 0.50 / 0.36 0.38 / 0.27

The chip's shadow for the program is 55,296 x 3.2 pJ = 0.18 microjoules per hash at N3 (0.24 node-for-node) against 0.13 for the genesis window on the same core; the card pays +0.6 percent for the window on the 5090 (the connected-state lane's measurement). The structure's other knobs do not reach the chip: the 16-pass loop and the 448 text cost the shared imem about 0.1 pJ per lane-op, hot-set banking is not needed (the gated file charges only the written register), and the chain's width sets lane count, which is free. On the GDDR7 board at the lock the window moves the chip's edge by about 1.1x (3.6x to 3.3x node-for-node), under the 1.25x gate that lane set.

4e. The mixed-resource lane's FP32 units (class-v6-mixedfp) on the adversary's lane

The mixed-resource lane's candidate adds four FP32 families (fadd, fmul, ffma, fcvt) to the shadow's draw, every result injected by xor, with inputs masked to a 7-bit exponent range (never zero, denormal, NaN or Inf). The adversary's simplified units (rtl/fp32_units.v): an FMA with the 24 x 24 mantissa multiplier, a 100-bit alignment window, a full normaliser and RNE; a separate adder and multiplier; the int32 to float converter; the exponent path narrowed to the range; no NaN, Inf, denormal or flag logic. The lane: an 8 x 32-bit window, the four units, d ^= bits(result). Routed with SPEF, random-input VCD, 42,936 cells, a 2 ns clock.

Op (every unit evaluating each cycle: an UPPER bound per op, no operand isolation) pJ per op ASAP7 N5 N3 N2 5090 fp32_fma stock / lock k at N3 vs stock / lock
fadd 6.6 4.6 3.3 2.4 9.2 / 5.2 0.36 / 0.64
fmul 7.0 4.9 3.5 2.5 9.2 / 5.2 0.38 / 0.68
ffma 6.7 4.7 3.4 2.4 9.2 / 5.2 0.37 / 0.65
fcvt 6.9 4.8 3.5 2.5 9.2 / 5.2 (cvt unmeasured on the card) 0.38 / 0.67
random mix 7.1 4.9 3.6 2.6 0.39 / 0.68

Reading: the four read alike because all four units switch every cycle on the same operands, so each row is the upper bound for its op (a chip isolates the idle units; by cell share about ffma 3.5 to 4, fmul 2.5, fadd 2, fcvt 1 pJ at ASAP7, approximate). Even on the upper bound the FP family is the chip's dearest per op relative to the card: k 0.65 at N3 at the lock against 0.18 for the integer ARX lane floor, because the card does an FMA for 5.2 pJ (under its own int add at 6.2) while the chip's multiply, alignment and normaliser cost about three int ops. On the units' floors the shadow's k_eff rises from 0.097 (class v4) to about 0.14 at the fp12 mix and 0.17 at fp24 (0.20 and 0.24 with isolation taken as half), the core's per-op overhead on top. The GPU-cost budget (10 percent of energy per hash) is the binding side, and the vendor-rounding question is the class's, not the chip's.

5. The chip edge at the measured k

E_chip = E_mem + N_ops x e_chip (absolute: the chip's shadow cost is 102,100 x 3.5 pJ = 0.36 microjoules per hash at N3, 0.26 at N2, whatever the card does); the record's convention E_mem + k x F beside it with k read at the card's point (it coincides at the lock by construction and is the same at stock here because both are the same arithmetic on the same ops; it differs when the chip's cost is held fixed while the card's point moves, which is what the SRAM lane found flattered the die 1.6x). Cards: the 5090 at stock (3.36 microjoules, F 1.10) and at the lock (2.33, F 0.652), the M5 Max (1.40 on class v4, 0.78 on class v3, 6.9 pJ per op). Chips: the record's GDDR7 board and the hardware-future lane's rows (their E_mem at zero shadow). The record's rows that this recomputes are hardware-future.md section 5: "2.6x to 2.8x at k = 1 and 4.4x to 5.1x at k = 0.5" for the strongest DRAM chips against the stock 5090 (the GDDR7 board's own row there is 2.1x and 3.3x).

Chip shadow per hash, absolute: N3 0.357 microjoules, N2 0.255.

Chip (E_mem, microjoules) Card row Zero shadow Absolute, N3 core Absolute, N2 core Record convention at the measured k (N3) at k = 1 at k = 0.5
GDDR7 board, 28 nm controller (the record) (0.466) 5090 stock (class v4) 4.8x 4.1x 4.7x 4.1x (k 0.32) 2.1x 3.3x
GDDR7 board, 28 nm controller (the record) (0.466) 5090 at the 1,300 MHz lock 3.6x 2.8x 3.2x 2.8x (k 0.55) 2.1x 2.9x
GDDR7 board, 28 nm controller (the record) (0.466) M5 Max (class v4 1.40 microjoules, class v3 0.78; 6.9 pJ per op; F 0.62) 1.7x 1.7x 1.9x 1.8x (k 0.51) 1.3x 1.8x
HBM3E one stack (0.321) 5090 stock (class v4) 7.0x 5.0x 5.8x 5.0x (k 0.32) 2.4x 3.9x
HBM3E one stack (0.321) 5090 at the 1,300 MHz lock 5.2x 3.4x 4.0x 3.4x (k 0.55) 2.4x 3.6x
HBM3E one stack (0.321) M5 Max (class v4 1.40 microjoules, class v3 0.78; 6.9 pJ per op; F 0.62) 2.4x 2.1x 2.4x 2.2x (k 0.51) 1.5x 2.2x
HBM4 one stack, N12 base die (0.22) 5090 stock (class v4) 10.3x 5.8x 7.1x 5.8x (k 0.32) 2.5x 4.4x
HBM4 one stack, N12 base die (0.22) 5090 at the 1,300 MHz lock 7.6x 4.0x 4.9x 4.0x (k 0.55) 2.7x 4.3x
HBM4 one stack, N12 base die (0.22) M5 Max (class v4 1.40 microjoules, class v3 0.78; 6.9 pJ per op; F 0.62) 3.5x 2.4x 2.9x 2.6x (k 0.51) 1.7x 2.6x
custom HBM4E base die, N3P (0.18) 5090 stock (class v4) 12.6x 6.3x 7.7x 6.3x (k 0.32) 2.6x 4.6x
custom HBM4E base die, N3P (0.18) 5090 at the 1,300 MHz lock 9.3x 4.3x 5.4x 4.3x (k 0.55) 2.8x 4.6x
custom HBM4E base die, N3P (0.18) M5 Max (class v4 1.40 microjoules, class v3 0.78; 6.9 pJ per op; F 0.62) 4.3x 2.6x 3.2x 2.8x (k 0.51) 1.7x 2.9x
DRAM on logic, hybrid bonded (2029 to 2031) (0.15) 5090 stock (class v4) 15.1x 6.6x 8.3x 6.6x (k 0.32) 2.7x 4.8x
DRAM on logic, hybrid bonded (2029 to 2031) (0.15) 5090 at the 1,300 MHz lock 11.2x 4.6x 5.7x 4.6x (k 0.55) 2.9x 4.9x
DRAM on logic, hybrid bonded (2029 to 2031) (0.15) M5 Max (class v4 1.40 microjoules, class v3 0.78; 6.9 pJ per op; F 0.62) 5.2x 2.8x 3.5x 3.0x (k 0.51) 1.8x 3.0x
SRAM full store, one N2 reticle (0.14) 5090 stock (class v4) 16.1x 6.8x 8.5x 6.8x (k 0.32) 2.7x 4.9x
SRAM full store, one N2 reticle (0.14) 5090 at the 1,300 MHz lock 12.0x 4.7x 5.9x 4.7x (k 0.55) 2.9x 5.0x
SRAM full store, one N2 reticle (0.14) M5 Max (class v4 1.40 microjoules, class v3 0.78; 6.9 pJ per op; F 0.62) 5.6x 2.8x 3.5x 3.1x (k 0.51) 1.8x 3.1x

Floor lane 1 assumption (not in by 18:00 UK): the honest 5090 at 1.0 microjoules per hash at zero shadow.

Chip Premium kept at 0.652 (card 1.652): absolute N3 / N2 Premium scaled with the card, 0.390 (card 1.390): absolute N3 / N2 zero shadow
GDDR7 board, 28 nm controller (the record) 2.0x / 2.3x 1.7x / 1.9x 2.1x
HBM3E one stack 2.4x / 2.9x 2.0x / 2.4x 3.1x
HBM4 one stack, N12 base die 2.9x / 3.5x 2.4x / 2.9x 4.5x
custom HBM4E base die, N3P 3.1x / 3.8x 2.6x / 3.2x 5.6x
DRAM on logic, hybrid bonded (2029 to 2031) 3.3x / 4.1x 2.7x / 3.4x 6.7x
SRAM full store, one N2 reticle 3.3x / 4.2x 2.8x / 3.5x 7.1x

What the rows say. (1) At the knee the GDDR7 board keeps 2.8x with the shadow on (3.2x at an N2 core) against 3.6x at zero shadow: the class v4 shadow buys the honest card 0.8x of edge, not the 1.5x the k = 1 row served and not the 0.1x the bare-lane floor would give. (2) The strongest DRAM chips read 4.3x to 4.7x at the knee and 6.3x to 6.8x at stock against the 5090 (the record's 2.6x to 2.8x at k = 1 and 4.4x to 5.1x at k = 0.5 bracket the stock figure and understate the knee one, because the record's k was read at stock). (3) Against the M5 Max the whole table compresses to 1.7x to 2.8x: the honest SoC's own operating point is most of the resistance. (4) Floor lane 1's honest 5090 at 1.0 microjoules (assumed; its rows were not in by 18:00 UK) would put the GDDR7 board at 1.7x to 2.3x and the strongest chips at 2.6x to 4.2x.

6. The ranking and the mixed draw

6.1 Families ranked by k, highest first (the hardest for a chip), absolute at N3 against the 5090's lock

Rank Family Chip pJ/op N3 (unit floor) 5090 pJ/op at the lock k (floor) In the class draw?
1 mad 1.66 8.3 0.20 yes (8 of 75)
2 add, sub, rotl, rotr, xor 1.05 to 1.09 6.2 0.17 to 0.18 yes (12, 6, 7, 6, 10)
3 the index fold 1.19 8.3 0.14 every load
4 or 0.73 6.2 0.12 yes (4)
5 mul 0.68 8.3 0.082 yes (8)
6 prmt, lop3 0.64, 0.72 11.5, 13.0 0.056, 0.055 not drawn (RTL rows only)
7 mulhi 0.68 21.0 0.032 yes (6)
8 32-lane shuffle (butterfly) 0.63 29.4 0.021 yes (8)
9 L1 scratch read (8 KB flop array) against the card's L2 hit 104 (pJ per read) 1,400 0.074 not drawn
10 int8 8x8x8 tile, per MAC ROW_TILE_K 2.2 pending not drawn (the tensor lever is dead on other grounds)

The order is set by the card's price, not the chip's: the chip pays 0.6 to 1.7 pJ for everything, and the card pays 6.2 for an add, 8.3 for a multiply, 21 for a high word and 29 for a shuffle. So the families the card pays MOST for (mulhi, shuffle) are the ones a chip undercuts most, and the forcing work is the plain ARX and mad the card does cheapest. This is the measured form of 15.1b's hold on the shuffle-heavy re-weight.

6.2 The mixed draw that maximises the expected k

The objective at a FIXED GPU premium: k_eff(w) = sum w_i e_chip_i / sum w_i e_gpu_i (the chip's energy for the drawn program over the card's for the same program; at a fixed F the chip pays k_eff x F). The band (layer 1, the research lane's 17:00 reading of lane D's cut): B = 4 points on the injecting families only (add, sub, xor, mad, shfl, rotl, rotr), the lossy families (or, mul, mulhi) at or under their base, or + mul + mulhi at most 18 + B, the shuffle capped at its class v4 weight; the index fold on every address and the F8 uniformity floor are not functions of the weights and do not move. tools/chip-model/rtl/flow/mix.py searches the band exhaustively over the corners (the objective is a ratio of linear forms, so the optimum is at a corner of each family's range).

On the unit floors (N3, the lock; the shuffle row in at 0.63 pJ, its k 0.021 the lowest of the drawn families):

Mix add, xor, mul, mad, shfl, rotl, sub, mulhi, rotr, or k_eff (floors) The card's pJ per op at the lock k_eff on the core (floor + the core's 2.4 pJ per op overhead at N3)
class v4 (the base) 12, 10, 8, 8, 8, 7, 6, 6, 6, 4 0.097 (with the shuffle row in: 0.130 over the other 67 points) 10.3 0.43
the census lane's re-weight (d20eb04bd, PASS on 256 seeds, no era and eras 0 to 7, W = 4 and W = 16) 13, 11, 6, 10, 8, 8, 7, 2, 6, 4 0.115 (+19 percent; 0.153 over the 67 points) 9.0 0.49 (+14 percent)
the band's best corner (mix.py, exhaustive: every injecting family at +4, the shuffle at its floor, the lossy families at base minus 4) 16, 14, 4, 12, 4, 11, 10, 2, 10, 0 0.137 (+42 percent) 8.2 about 0.51 (+19 percent)
the band's worst corner (shuffle and mulhi heavy) 8, 6, 4, 4, 8, 3, 2, 6, 2, 0 0.074 12.3 about 0.40

Reading: the re-weight moves six points off the chip's two easiest families (mulhi, mul) onto the hardest (mad, the ARX families) and lifts the expected k by about a seventh on the core; its acceptance is the census lane's (0 exhausted, F8 0.999 to 1.008 of uniform, the verifier within 4 percent); its energy cost on the card is a premium per instruction 11 percent LOWER (fewer mulhi), so at a fixed premium the program is 12 percent longer. The shuffle's own row decides whether it goes to its floor: on the card it is the dearest op (29.4 pJ at the lock) and on the chip a butterfly over the window, so its k is the lowest of the drawn families unless the synthesis says otherwise. The recommendation: the band's best corner (add 16, xor 14, mad 12, rotl 11, sub 10, rotr 10, shfl 4, mul 4, mulhi 2, or 0) if its census passes (the census lane's harness, about five box-minutes per candidate); the census lane's draw as the passed fallback. Either way the shuffle goes to its floor: the card pays 29.4 pJ for a move the chip does for 0.6.

7. Sources

  • The GPU side: docs/analysis/counter-asic-4-research.md 15.1a (the 5090 microbench, 8 October 2026, measured) and 20.3 (the packs job: 10.8 / 6.4 pJ per counted op on the class v4 shadow); the M5 Max 6.9 pJ per counted op from the coordinator's order of 8 October 2026 (the latency-shadow record).
  • ASAP7: L. T. Clark et al., "ASAP7: A 7-nm finFET predictive process design kit", Microelectronics Journal 53 (2016); the ORFS platform files (flow/platforms/asap7, the 7.5-track RVT library, TC corner 0.70 V / 0 C).
  • The flow: OpenROAD-flow-scripts (docker image openroad/orfs:latest, Yosys 0.68, OpenROAD and OpenSTA with read_vcd); iverilog 12 for the gate-level simulation.
  • Node scaling (claimed): TSMC N5 "30 percent lower power at the same speed" against N7 (TSMC technology page, https://www.tsmc.com/english/dedicatedFoundry/technology/logic/l_5nm); N3E "25 to 30 percent lower power" against N5 (https://www.tsmc.com/english/dedicatedFoundry/technology/logic/l_3nm); N2 "25 to 30 percent lower power" against N3E (https://www.tsmc.com/english/dedicatedFoundry/technology/logic/l_2nm). Read 8 October 2026.
  • The chip model's memory rows: docs/analysis/chip-model-v3.md 5.5 (the GDDR7 board, 0.466 microjoules) and docs/analysis/class-v6/hardware-future.md section 5 (the strongest DRAM and SRAM rows).
  • The band: docs/design/class-v6-rotating-family.md section 2 (layer 1) and igneum-pow/src/generator.rs NONLOAD_WEIGHTS.