chip-model: shadow-k collector, mix optimiser, edge rows; the report's method sections

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-josh 2026-10-08 13:05:02 +01:00
parent 9009ed0f40
commit 6fab785c51
7 changed files with 171 additions and 2 deletions

View file

@ -0,0 +1,77 @@
# The shadow's k from RTL (floor lane 2, 8 October 2026)
Branch `class-v6-floor-k` from `counter-asic-4` at 7618e729. Every chip-side number in this file comes from
synthesis and place-and-route of RTL written for this lane (Yosys 0.68 plus OpenROAD, the ORFS docker image
`openroad/orfs:latest`, the ASAP7 predictive PDK, run on igneum-build-4), with switching activity from a
random-input gate-level simulation (iverilog on the routed netlist). Every GPU-side number is the record's
measurement (`docs/analysis/counter-asic-4-research.md` 15.1a: PC 1, the RTX 5090, 20 probes at 60 s, 8
October 2026) and is not re-estimated here. The RTL, testbenches, flow configs and a Makefile that reproduces
every row are under `tools/chip-model/rtl/`.
## 0. One page
(filled from the rows below; see section 2 for the table and section 4 for the edge)
## 1. The question and the identity
The chip's energy per hash is `E_chip = E_mem + k x F` (research file section 2): `E_mem` the memory system's
reads, controller and static (0.466 microjoules per hash on the record's GDDR7 board, modelled), `F` the shadow's
premium on the card (measured: 1.10 microjoules unlocked, 0.652 at the 1,300 MHz lock on the 5090 for class
v4's 102,100 counted ops per hash), and `k = e_chip / e_gpu` the chip core's energy per forced op over the
card's. The card's side of `k` is measured; the chip's side has been a claimed band (ALU 0.3 to 0.8, int8 tile
0.03 to 0.3, L2 hit 0.1 to 0.3, shuffle 0.4 to 0.7). This file replaces the claimed side with a synthesised one,
family by family, and reads the mix that maximises the expected `k` under the invention lane's sound-form rules.
## 2. Method
### 2.1 What was built (RTL, `tools/chip-model/rtl/rtl/`)
Each family is a minimal chip-side shadow core: the smallest circuit a chip maker would have to build to run that
family's instructions bit-exactly at one op per cycle per lane. The common lane shape is an instruction register,
an 8 x 32-bit register window in flops (the card's lane holds its working set in a register file too; 8 registers
is the class program's window), two or three read ports through muxes, the functional unit, and one write port.
Nothing is shared across lanes and nothing is pipelined beyond one stage, so the figure is the datapath plus the
minimum operand delivery: a floor for the chip, which makes the `k` it gives a floor too.
| Family | Module | What one op is | Ops per cycle | Clock set (ps) |
|---|---|---|---|---|
| ARX | `lane_arx` | int32 add, sub, xor, or, rotl by immediate, rotr by register (op drawn per cycle; rows per op by fixing the op field) | 1 | 1,000 |
| MUL | `lane_mul` | 32 x 32 to 64: mul (low word), mulhi (high word), mad (src x src2 + dst) | 1 | 1,500 |
| PRMT | `lane_prmt` | byte permute: 4 output bytes from the 8 bytes of two registers by a 16-bit selector, sign-replicate bit as on the card | 1 | 1,000 |
| LOP3 | `lane_lop3` | three-input logic by an 8-bit truth table, bitwise | 1 | 1,000 |
| FOLD | `lane_fold` | the index fold `((rotl(x * M, R) & WM) \| OFF) & MASK` with the era's constants held in registers | 1 | 1,500 |
| SHFL | `shfl32` | the 32-lane xor-mask shuffle `r[dst][lane] ^= r[src][lane ^ m]` over a 1 KB window (32 lanes x 8 x 32 bits): a 5-stage butterfly | 32 lane-ops | 1,000 |
| XBAR | `xbar32` | the general 32-lane crossbar over the same window (any source lane per lane, a 5-bit select each): the upper bound on a shuffle network | 32 lane-ops | 1,000 |
| SCRATCH | `scratch8k` | one random 32-bit read of a 2,048 x 32 (8 KB) flop array, a write on one cycle in eight: the chip's L1 as the pessimistic flop form | 1 read | 1,500 |
| TILE | `tile8` | the int8 8x8x8 tile `C += A x B` with int32 accumulators: 512 MACs per cycle | 512 MACs | 2,000 |
### 2.2 The flow
ORFS on ASAP7 (7.5-track RVT cells, the TC corner: 0.70 V, 0 C, NLDM), the default flow end to end: Yosys
synthesis with ABC, floorplan at 40 percent utilisation, global and detailed placement, CTS, global and detailed
routing, parasitic extraction (OpenRCX, the platform's rules). Power is OpenSTA's `report_power` on the routed
design with its SPEF, under two activities: (a) the VCD of a random-input simulation of the routed netlist
(ASAP7 has no cell Verilog models in the image, so Yosys builds the cell bodies from the liberty functions and
writes the netlist as primitives; iverilog runs 3,000 to 4,000 cycles with every instruction field drawn by
`$random` each cycle and a random 32-bit load into the window every 16th cycle), and (b) a propagated 0.5
activity on every input as the cross-check. The energy per op is the total power (internal, switching and leakage)
times the clock period over the ops per cycle. Leakage is reported beside it; at these clocks it is a few percent.
### 2.3 Node scaling (approximate; every factor claimed from the foundry's own headline)
ASAP7 is a predictive 7 nm-class FinFET PDK (ASU and ARM, Clark et al., Microelectronics Journal 2016), not a
foundry node, so the row is first stated at ASAP7 and then scaled by the foundry's published per-node power
reductions at the same speed: N7 to N5 x0.70 (TSMC: "30 percent lower power"), N5 to N3E x0.72 (TSMC: "25 to 30
percent lower power", the midpoint), N3E to N2 x0.72 (TSMC: "25 to 30 percent lower power", the midpoint). So
N5 = 0.70, N3 = 0.50 and N2 = 0.36 of the ASAP7 figure. These are the foundry's claims for a whole design at a
fixed frequency, and a shadow core at a low clock could run at a lower voltage still; the N2 column is therefore
the chip's best case from this method, not its floor. Sources in section 7.
### 2.4 What the method leaves out, on both sides
On the chip side the figure omits instruction fetch and decode (a chip would run the program from a small SRAM
or a decoded instruction cache shared by many lanes, about 1 to 2 pJ per lane-instruction amortised across 32
lanes, approximate), the clock tree beyond the block's own, and the result's move to a memory address unit. On
the card side the 15.1a figure is the whole card's marginal per counted op (the sleep floor subtracted), which
includes the card's own fetch, decode, operand collection and register file. So the `k` here is the chip's
datapath-and-window cost over the card's whole-lane cost: a FLOOR on the chip's cost and so a floor on `k`.

View file

@ -34,10 +34,11 @@ def parse_log(path):
tag = head.strip().replace('POWER_VCD ', 'vcd:').replace('POWER_PROPAGATED_0.5', 'prop')
# OpenSTA report_power: "Total <int> <sw> <leak> <total> 100.0%" in watts
tm = re.search(r'^Total\s+([\d.eE+-]+)\s+([\d.eE+-]+)\s+([\d.eE+-]+)\s+([\d.eE+-]+)', body, re.M)
ann = re.search(r'(\d+)\s*\(\s*([\d.]+)%\)\s*annotated', body)
ann = re.search(r'Annotated (\d+) pin activities', body)
unann = re.search(r'unannotated\s+(\d+)', body)
if tm:
rows[tag] = dict(internal=float(tm.group(1)), switching=float(tm.group(2)), leakage=float(tm.group(3)),
total=float(tm.group(4)), annotated=ann.group(2) if ann else '')
total=float(tm.group(4)), annotated=(f"{ann.group(1)} pins, {unann.group(1) if unann else '?'} unannotated") if ann else '')
return period, cells, rows
out = []

View file

@ -0,0 +1,23 @@
#!/usr/bin/env python3
"""The chip edge rows recomputed at the measured k. Usage: edge.py <k_eff_v4> <k_eff_best>"""
import sys
k_v4 = float(sys.argv[1]); k_best = float(sys.argv[2])
# the record's rows: card microjoules per hash and the premium F (counter-asic-4-research.md 20.4; the GDDR7
# board E_mem 0.466; hardware-future.md section 5 for the strongest DRAM chips)
cards = [('5090 unlocked (class v4 shape)', 3.36, 1.10), ('5090 at the 1,300 MHz lock', 2.33, 0.652)]
chips = [('GDDR7 board, 28 nm controller (the record)', 0.466), ('HBM3E one stack', 0.321), ('HBM4 one stack, N12 base die', 0.22),
('custom HBM4E base die, N3P', 0.18), ('DRAM on logic, hybrid bonded (2029 to 2031)', 0.15), ('SRAM full store, one N2 reticle', 0.14)]
def edge(card, F, mem, k): return (card) / (mem + k * F)
print('| Chip | 5090 row | Card uJ | F uJ | Edge at k = 1 | at k = 0.5 | at k = 0.3 | at the measured k (class v4 mix) | at the measured k (best mix) | at zero shadow |')
print('|---|---|---|---|---|---|---|---|---|---|')
for cn, mem in chips:
for rn, card, F in cards:
print(f'| {cn} | {rn} | {card:.2f} | {F:.3f} | {edge(card,F,mem,1):.2f}x | {edge(card,F,mem,0.5):.2f}x | {edge(card,F,mem,0.3):.2f}x | {edge(card,F,mem,k_v4):.2f}x (k {k_v4:.3f}) | {edge(card,F,mem,k_best):.2f}x (k {k_best:.3f}) | {(card-F)/mem:.2f}x |')
# the honest-card variants: floor lane 1's assumed 1.0 microjoules per hash at zero shadow on the 5090
print('\nFloor lane 1 assumption: the honest 5090 at 1.0 microjoules per hash at zero shadow (not in by the default time).')
print('| Chip | Premium kept at the lock\'s 0.652 | Premium scaled with the card (0.652 x 1.0 / 1.67 = 0.390) |')
print('|---|---|---|')
for cn, mem in chips:
a = edge(1.0 + 0.652, 0.652, mem, k_v4); b = edge(1.0 + 0.390, 0.390, mem, k_v4)
a1 = edge(1.0 + 0.652, 0.652, mem, 1.0); b1 = edge(1.0 + 0.390, 0.390, mem, 1.0)
print(f'| {cn} | {a:.2f}x at the measured k ({a1:.2f}x at k = 1; {1.0/mem:.2f}x at zero shadow) | {b:.2f}x at the measured k ({b1:.2f}x at k = 1) |')

View file

@ -0,0 +1,65 @@
#!/usr/bin/env python3
"""The op mix that maximises the chip's effective k at a FIXED GPU premium, inside the class v6 layer-1 band.
k_eff(mix) = sum_i w_i e_chip_i / sum_i w_i e_gpu_i : the chip's energy for the drawn program over the card's
for the same program. At a fixed premium F on the card the chip pays k_eff x F, so the mix that maximises k_eff
is the one that forces the most. The band (docs/design/class-v6-rotating-family.md section 2, layer 1):
B = 4 points on the injecting families only (add, sub, xor, mad, shfl, rotl, rotr), each within base - B .. base + B;
the lossy families (or, mul, mulhi) capped at their base (base - B .. base); or + mul + mulhi at most 18 + B;
the shuffle weight capped at its class v4 value (8). The index fold and the F8 floor are not functions of the
weights and do not move. Usage: mix.py <table.csv> [node=N3] [state=lock|unlocked]"""
import csv, sys, itertools
path = sys.argv[1]; node = 'N3'; state = 'lock'
for a in sys.argv[2:]:
k, v = a.split('=');
if k == 'node': node = v
if k == 'state': state = v
rows = {(r['family'], r['sim']): r for r in csv.DictReader(open(path))}
def chip(fam, sim='vcd:mix'):
r = rows.get((fam, sim)) or rows.get((fam, 'vcd:mix'))
return float(r[f'pJ_{node}']) if r else None
# the 5090 measured pJ per counted op (15.1a), per family of the generator's draw
gpu = {'add': (11.3, 6.2), 'sub': (11.3, 6.2), 'xor': (11.3, 6.2), 'or': (11.3, 6.2), 'rotl': (11.3, 6.2), 'rotr': (11.3, 6.2),
'mul': (13.9, 8.3), 'mad': (13.9, 8.3), 'mulhi': (39.6, 21.0), 'shfl': (55.8, 29.4)}
si = 0 if state == 'unlocked' else 1
e_gpu = {f: v[si] for f, v in gpu.items()}
e_chip = {'add': chip('arx', 'vcd:add'), 'sub': chip('arx', 'vcd:sub'), 'xor': chip('arx', 'vcd:xor'), 'or': chip('arx', 'vcd:or'),
'rotl': chip('arx', 'vcd:rotl'), 'rotr': chip('arx', 'vcd:rotr'),
'mul': chip('mul', 'vcd:mul'), 'mad': chip('mul', 'vcd:mad'), 'mulhi': chip('mul', 'vcd:mulhi'),
'shfl': chip('shfl')}
base = {'add': 12, 'xor': 10, 'mul': 8, 'mad': 8, 'shfl': 8, 'rotl': 7, 'sub': 6, 'mulhi': 6, 'rotr': 6, 'or': 4}
B = 4
inject = ['add', 'sub', 'xor', 'mad', 'shfl', 'rotl', 'rotr']; lossy = ['or', 'mul', 'mulhi']
missing = [f for f, v in e_chip.items() if v is None]
if missing:
print('missing chip rows:', missing); sys.exit(1)
k_fam = {f: e_chip[f] / e_gpu[f] for f in base}
def keff(w): return sum(w[f] * e_chip[f] for f in w) / sum(w[f] * e_gpu[f] for f in w)
print(f'node {node}, 5090 state {state}')
print('| Family | Base weight | 5090 pJ/op | Chip pJ/op | k |')
print('|---|---|---|---|---|')
for f in sorted(base, key=lambda f: -k_fam[f]):
print(f'| {f} | {base[f]} | {e_gpu[f]} | {e_chip[f]:.3g} | {k_fam[f]:.3f} |')
print(f'\nclass v4 mix: k_eff = {keff(base):.4f}')
# exhaustive over the band: each family at one of the allowed values (steps of 1 point)
ranges = {}
for f in base:
lo = max(0, base[f] - B)
hi = base[f] + B if f in inject else base[f]
if f == 'shfl': hi = min(hi, 8)
ranges[f] = list(range(lo, hi + 1))
best = None; worst = None
fams = list(base)
# reduce the search: injecting families only take their extremes and base (the objective is a ratio of linear forms,
# monotone in each weight), lossy families all values
cand = {f: ([ranges[f][0], base[f], ranges[f][-1]] if f in inject else ranges[f]) for f in fams}
for combo in itertools.product(*[cand[f] for f in fams]):
w = dict(zip(fams, combo))
if w['or'] + w['mul'] + w['mulhi'] > 18 + B: continue
if sum(w.values()) == 0: continue
v = keff(w)
if best is None or v > best[0]: best = (v, dict(w))
if worst is None or v < worst[0]: worst = (v, dict(w))
for name, (v, w) in (('best', best), ('worst', worst)):
print(f'{name} mix in the band: k_eff = {v:.4f}: ' + ', '.join(f'{f} {w[f]}' for f in fams) + f' (sum {sum(w.values())})')

View file

@ -13,3 +13,4 @@ export CORNER = TC
export SKIP_LAST_GASP = 1
export WORK_HOME = /work/out/scratch
export SYNTH_MEMORY_MAX_BITS = 131072

View file

@ -13,3 +13,4 @@ export CORNER = TC
export SKIP_LAST_GASP = 1
export WORK_HOME = /work/out/shfl
export SYNTH_MEMORY_MAX_BITS = 131072

View file

@ -13,3 +13,4 @@ export CORNER = TC
export SKIP_LAST_GASP = 1
export WORK_HOME = /work/out/xbar
export SYNTH_MEMORY_MAX_BITS = 131072