chip-model: shadow-k collector, mix optimiser, edge rows; the report's method sections
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
9009ed0f40
commit
6fab785c51
7 changed files with 171 additions and 2 deletions
77
docs/analysis/class-v6/floor/shadow-k.md
Normal file
77
docs/analysis/class-v6/floor/shadow-k.md
Normal file
|
|
@ -0,0 +1,77 @@
|
|||
# The shadow's k from RTL (floor lane 2, 8 October 2026)
|
||||
|
||||
Branch `class-v6-floor-k` from `counter-asic-4` at 7618e729. Every chip-side number in this file comes from
|
||||
synthesis and place-and-route of RTL written for this lane (Yosys 0.68 plus OpenROAD, the ORFS docker image
|
||||
`openroad/orfs:latest`, the ASAP7 predictive PDK, run on igneum-build-4), with switching activity from a
|
||||
random-input gate-level simulation (iverilog on the routed netlist). Every GPU-side number is the record's
|
||||
measurement (`docs/analysis/counter-asic-4-research.md` 15.1a: PC 1, the RTX 5090, 20 probes at 60 s, 8
|
||||
October 2026) and is not re-estimated here. The RTL, testbenches, flow configs and a Makefile that reproduces
|
||||
every row are under `tools/chip-model/rtl/`.
|
||||
|
||||
## 0. One page
|
||||
|
||||
(filled from the rows below; see section 2 for the table and section 4 for the edge)
|
||||
|
||||
## 1. The question and the identity
|
||||
|
||||
The chip's energy per hash is `E_chip = E_mem + k x F` (research file section 2): `E_mem` the memory system's
|
||||
reads, controller and static (0.466 microjoules per hash on the record's GDDR7 board, modelled), `F` the shadow's
|
||||
premium on the card (measured: 1.10 microjoules unlocked, 0.652 at the 1,300 MHz lock on the 5090 for class
|
||||
v4's 102,100 counted ops per hash), and `k = e_chip / e_gpu` the chip core's energy per forced op over the
|
||||
card's. The card's side of `k` is measured; the chip's side has been a claimed band (ALU 0.3 to 0.8, int8 tile
|
||||
0.03 to 0.3, L2 hit 0.1 to 0.3, shuffle 0.4 to 0.7). This file replaces the claimed side with a synthesised one,
|
||||
family by family, and reads the mix that maximises the expected `k` under the invention lane's sound-form rules.
|
||||
|
||||
## 2. Method
|
||||
|
||||
### 2.1 What was built (RTL, `tools/chip-model/rtl/rtl/`)
|
||||
|
||||
Each family is a minimal chip-side shadow core: the smallest circuit a chip maker would have to build to run that
|
||||
family's instructions bit-exactly at one op per cycle per lane. The common lane shape is an instruction register,
|
||||
an 8 x 32-bit register window in flops (the card's lane holds its working set in a register file too; 8 registers
|
||||
is the class program's window), two or three read ports through muxes, the functional unit, and one write port.
|
||||
Nothing is shared across lanes and nothing is pipelined beyond one stage, so the figure is the datapath plus the
|
||||
minimum operand delivery: a floor for the chip, which makes the `k` it gives a floor too.
|
||||
|
||||
| Family | Module | What one op is | Ops per cycle | Clock set (ps) |
|
||||
|---|---|---|---|---|
|
||||
| ARX | `lane_arx` | int32 add, sub, xor, or, rotl by immediate, rotr by register (op drawn per cycle; rows per op by fixing the op field) | 1 | 1,000 |
|
||||
| MUL | `lane_mul` | 32 x 32 to 64: mul (low word), mulhi (high word), mad (src x src2 + dst) | 1 | 1,500 |
|
||||
| PRMT | `lane_prmt` | byte permute: 4 output bytes from the 8 bytes of two registers by a 16-bit selector, sign-replicate bit as on the card | 1 | 1,000 |
|
||||
| LOP3 | `lane_lop3` | three-input logic by an 8-bit truth table, bitwise | 1 | 1,000 |
|
||||
| FOLD | `lane_fold` | the index fold `((rotl(x * M, R) & WM) \| OFF) & MASK` with the era's constants held in registers | 1 | 1,500 |
|
||||
| SHFL | `shfl32` | the 32-lane xor-mask shuffle `r[dst][lane] ^= r[src][lane ^ m]` over a 1 KB window (32 lanes x 8 x 32 bits): a 5-stage butterfly | 32 lane-ops | 1,000 |
|
||||
| XBAR | `xbar32` | the general 32-lane crossbar over the same window (any source lane per lane, a 5-bit select each): the upper bound on a shuffle network | 32 lane-ops | 1,000 |
|
||||
| SCRATCH | `scratch8k` | one random 32-bit read of a 2,048 x 32 (8 KB) flop array, a write on one cycle in eight: the chip's L1 as the pessimistic flop form | 1 read | 1,500 |
|
||||
| TILE | `tile8` | the int8 8x8x8 tile `C += A x B` with int32 accumulators: 512 MACs per cycle | 512 MACs | 2,000 |
|
||||
|
||||
### 2.2 The flow
|
||||
|
||||
ORFS on ASAP7 (7.5-track RVT cells, the TC corner: 0.70 V, 0 C, NLDM), the default flow end to end: Yosys
|
||||
synthesis with ABC, floorplan at 40 percent utilisation, global and detailed placement, CTS, global and detailed
|
||||
routing, parasitic extraction (OpenRCX, the platform's rules). Power is OpenSTA's `report_power` on the routed
|
||||
design with its SPEF, under two activities: (a) the VCD of a random-input simulation of the routed netlist
|
||||
(ASAP7 has no cell Verilog models in the image, so Yosys builds the cell bodies from the liberty functions and
|
||||
writes the netlist as primitives; iverilog runs 3,000 to 4,000 cycles with every instruction field drawn by
|
||||
`$random` each cycle and a random 32-bit load into the window every 16th cycle), and (b) a propagated 0.5
|
||||
activity on every input as the cross-check. The energy per op is the total power (internal, switching and leakage)
|
||||
times the clock period over the ops per cycle. Leakage is reported beside it; at these clocks it is a few percent.
|
||||
|
||||
### 2.3 Node scaling (approximate; every factor claimed from the foundry's own headline)
|
||||
|
||||
ASAP7 is a predictive 7 nm-class FinFET PDK (ASU and ARM, Clark et al., Microelectronics Journal 2016), not a
|
||||
foundry node, so the row is first stated at ASAP7 and then scaled by the foundry's published per-node power
|
||||
reductions at the same speed: N7 to N5 x0.70 (TSMC: "30 percent lower power"), N5 to N3E x0.72 (TSMC: "25 to 30
|
||||
percent lower power", the midpoint), N3E to N2 x0.72 (TSMC: "25 to 30 percent lower power", the midpoint). So
|
||||
N5 = 0.70, N3 = 0.50 and N2 = 0.36 of the ASAP7 figure. These are the foundry's claims for a whole design at a
|
||||
fixed frequency, and a shadow core at a low clock could run at a lower voltage still; the N2 column is therefore
|
||||
the chip's best case from this method, not its floor. Sources in section 7.
|
||||
|
||||
### 2.4 What the method leaves out, on both sides
|
||||
|
||||
On the chip side the figure omits instruction fetch and decode (a chip would run the program from a small SRAM
|
||||
or a decoded instruction cache shared by many lanes, about 1 to 2 pJ per lane-instruction amortised across 32
|
||||
lanes, approximate), the clock tree beyond the block's own, and the result's move to a memory address unit. On
|
||||
the card side the 15.1a figure is the whole card's marginal per counted op (the sleep floor subtracted), which
|
||||
includes the card's own fetch, decode, operand collection and register file. So the `k` here is the chip's
|
||||
datapath-and-window cost over the card's whole-lane cost: a FLOOR on the chip's cost and so a floor on `k`.
|
||||
|
|
@ -34,10 +34,11 @@ def parse_log(path):
|
|||
tag = head.strip().replace('POWER_VCD ', 'vcd:').replace('POWER_PROPAGATED_0.5', 'prop')
|
||||
# OpenSTA report_power: "Total <int> <sw> <leak> <total> 100.0%" in watts
|
||||
tm = re.search(r'^Total\s+([\d.eE+-]+)\s+([\d.eE+-]+)\s+([\d.eE+-]+)\s+([\d.eE+-]+)', body, re.M)
|
||||
ann = re.search(r'(\d+)\s*\(\s*([\d.]+)%\)\s*annotated', body)
|
||||
ann = re.search(r'Annotated (\d+) pin activities', body)
|
||||
unann = re.search(r'unannotated\s+(\d+)', body)
|
||||
if tm:
|
||||
rows[tag] = dict(internal=float(tm.group(1)), switching=float(tm.group(2)), leakage=float(tm.group(3)),
|
||||
total=float(tm.group(4)), annotated=ann.group(2) if ann else '')
|
||||
total=float(tm.group(4)), annotated=(f"{ann.group(1)} pins, {unann.group(1) if unann else '?'} unannotated") if ann else '')
|
||||
return period, cells, rows
|
||||
|
||||
out = []
|
||||
|
|
|
|||
23
tools/chip-model/rtl/flow/edge.py
Normal file
23
tools/chip-model/rtl/flow/edge.py
Normal file
|
|
@ -0,0 +1,23 @@
|
|||
#!/usr/bin/env python3
|
||||
"""The chip edge rows recomputed at the measured k. Usage: edge.py <k_eff_v4> <k_eff_best>"""
|
||||
import sys
|
||||
k_v4 = float(sys.argv[1]); k_best = float(sys.argv[2])
|
||||
# the record's rows: card microjoules per hash and the premium F (counter-asic-4-research.md 20.4; the GDDR7
|
||||
# board E_mem 0.466; hardware-future.md section 5 for the strongest DRAM chips)
|
||||
cards = [('5090 unlocked (class v4 shape)', 3.36, 1.10), ('5090 at the 1,300 MHz lock', 2.33, 0.652)]
|
||||
chips = [('GDDR7 board, 28 nm controller (the record)', 0.466), ('HBM3E one stack', 0.321), ('HBM4 one stack, N12 base die', 0.22),
|
||||
('custom HBM4E base die, N3P', 0.18), ('DRAM on logic, hybrid bonded (2029 to 2031)', 0.15), ('SRAM full store, one N2 reticle', 0.14)]
|
||||
def edge(card, F, mem, k): return (card) / (mem + k * F)
|
||||
print('| Chip | 5090 row | Card uJ | F uJ | Edge at k = 1 | at k = 0.5 | at k = 0.3 | at the measured k (class v4 mix) | at the measured k (best mix) | at zero shadow |')
|
||||
print('|---|---|---|---|---|---|---|---|---|---|')
|
||||
for cn, mem in chips:
|
||||
for rn, card, F in cards:
|
||||
print(f'| {cn} | {rn} | {card:.2f} | {F:.3f} | {edge(card,F,mem,1):.2f}x | {edge(card,F,mem,0.5):.2f}x | {edge(card,F,mem,0.3):.2f}x | {edge(card,F,mem,k_v4):.2f}x (k {k_v4:.3f}) | {edge(card,F,mem,k_best):.2f}x (k {k_best:.3f}) | {(card-F)/mem:.2f}x |')
|
||||
# the honest-card variants: floor lane 1's assumed 1.0 microjoules per hash at zero shadow on the 5090
|
||||
print('\nFloor lane 1 assumption: the honest 5090 at 1.0 microjoules per hash at zero shadow (not in by the default time).')
|
||||
print('| Chip | Premium kept at the lock\'s 0.652 | Premium scaled with the card (0.652 x 1.0 / 1.67 = 0.390) |')
|
||||
print('|---|---|---|')
|
||||
for cn, mem in chips:
|
||||
a = edge(1.0 + 0.652, 0.652, mem, k_v4); b = edge(1.0 + 0.390, 0.390, mem, k_v4)
|
||||
a1 = edge(1.0 + 0.652, 0.652, mem, 1.0); b1 = edge(1.0 + 0.390, 0.390, mem, 1.0)
|
||||
print(f'| {cn} | {a:.2f}x at the measured k ({a1:.2f}x at k = 1; {1.0/mem:.2f}x at zero shadow) | {b:.2f}x at the measured k ({b1:.2f}x at k = 1) |')
|
||||
65
tools/chip-model/rtl/flow/mix.py
Normal file
65
tools/chip-model/rtl/flow/mix.py
Normal file
|
|
@ -0,0 +1,65 @@
|
|||
#!/usr/bin/env python3
|
||||
"""The op mix that maximises the chip's effective k at a FIXED GPU premium, inside the class v6 layer-1 band.
|
||||
|
||||
k_eff(mix) = sum_i w_i e_chip_i / sum_i w_i e_gpu_i : the chip's energy for the drawn program over the card's
|
||||
for the same program. At a fixed premium F on the card the chip pays k_eff x F, so the mix that maximises k_eff
|
||||
is the one that forces the most. The band (docs/design/class-v6-rotating-family.md section 2, layer 1):
|
||||
B = 4 points on the injecting families only (add, sub, xor, mad, shfl, rotl, rotr), each within base - B .. base + B;
|
||||
the lossy families (or, mul, mulhi) capped at their base (base - B .. base); or + mul + mulhi at most 18 + B;
|
||||
the shuffle weight capped at its class v4 value (8). The index fold and the F8 floor are not functions of the
|
||||
weights and do not move. Usage: mix.py <table.csv> [node=N3] [state=lock|unlocked]"""
|
||||
import csv, sys, itertools
|
||||
|
||||
path = sys.argv[1]; node = 'N3'; state = 'lock'
|
||||
for a in sys.argv[2:]:
|
||||
k, v = a.split('=');
|
||||
if k == 'node': node = v
|
||||
if k == 'state': state = v
|
||||
rows = {(r['family'], r['sim']): r for r in csv.DictReader(open(path))}
|
||||
def chip(fam, sim='vcd:mix'):
|
||||
r = rows.get((fam, sim)) or rows.get((fam, 'vcd:mix'))
|
||||
return float(r[f'pJ_{node}']) if r else None
|
||||
# the 5090 measured pJ per counted op (15.1a), per family of the generator's draw
|
||||
gpu = {'add': (11.3, 6.2), 'sub': (11.3, 6.2), 'xor': (11.3, 6.2), 'or': (11.3, 6.2), 'rotl': (11.3, 6.2), 'rotr': (11.3, 6.2),
|
||||
'mul': (13.9, 8.3), 'mad': (13.9, 8.3), 'mulhi': (39.6, 21.0), 'shfl': (55.8, 29.4)}
|
||||
si = 0 if state == 'unlocked' else 1
|
||||
e_gpu = {f: v[si] for f, v in gpu.items()}
|
||||
e_chip = {'add': chip('arx', 'vcd:add'), 'sub': chip('arx', 'vcd:sub'), 'xor': chip('arx', 'vcd:xor'), 'or': chip('arx', 'vcd:or'),
|
||||
'rotl': chip('arx', 'vcd:rotl'), 'rotr': chip('arx', 'vcd:rotr'),
|
||||
'mul': chip('mul', 'vcd:mul'), 'mad': chip('mul', 'vcd:mad'), 'mulhi': chip('mul', 'vcd:mulhi'),
|
||||
'shfl': chip('shfl')}
|
||||
base = {'add': 12, 'xor': 10, 'mul': 8, 'mad': 8, 'shfl': 8, 'rotl': 7, 'sub': 6, 'mulhi': 6, 'rotr': 6, 'or': 4}
|
||||
B = 4
|
||||
inject = ['add', 'sub', 'xor', 'mad', 'shfl', 'rotl', 'rotr']; lossy = ['or', 'mul', 'mulhi']
|
||||
missing = [f for f, v in e_chip.items() if v is None]
|
||||
if missing:
|
||||
print('missing chip rows:', missing); sys.exit(1)
|
||||
k_fam = {f: e_chip[f] / e_gpu[f] for f in base}
|
||||
def keff(w): return sum(w[f] * e_chip[f] for f in w) / sum(w[f] * e_gpu[f] for f in w)
|
||||
print(f'node {node}, 5090 state {state}')
|
||||
print('| Family | Base weight | 5090 pJ/op | Chip pJ/op | k |')
|
||||
print('|---|---|---|---|---|')
|
||||
for f in sorted(base, key=lambda f: -k_fam[f]):
|
||||
print(f'| {f} | {base[f]} | {e_gpu[f]} | {e_chip[f]:.3g} | {k_fam[f]:.3f} |')
|
||||
print(f'\nclass v4 mix: k_eff = {keff(base):.4f}')
|
||||
# exhaustive over the band: each family at one of the allowed values (steps of 1 point)
|
||||
ranges = {}
|
||||
for f in base:
|
||||
lo = max(0, base[f] - B)
|
||||
hi = base[f] + B if f in inject else base[f]
|
||||
if f == 'shfl': hi = min(hi, 8)
|
||||
ranges[f] = list(range(lo, hi + 1))
|
||||
best = None; worst = None
|
||||
fams = list(base)
|
||||
# reduce the search: injecting families only take their extremes and base (the objective is a ratio of linear forms,
|
||||
# monotone in each weight), lossy families all values
|
||||
cand = {f: ([ranges[f][0], base[f], ranges[f][-1]] if f in inject else ranges[f]) for f in fams}
|
||||
for combo in itertools.product(*[cand[f] for f in fams]):
|
||||
w = dict(zip(fams, combo))
|
||||
if w['or'] + w['mul'] + w['mulhi'] > 18 + B: continue
|
||||
if sum(w.values()) == 0: continue
|
||||
v = keff(w)
|
||||
if best is None or v > best[0]: best = (v, dict(w))
|
||||
if worst is None or v < worst[0]: worst = (v, dict(w))
|
||||
for name, (v, w) in (('best', best), ('worst', worst)):
|
||||
print(f'{name} mix in the band: k_eff = {v:.4f}: ' + ', '.join(f'{f} {w[f]}' for f in fams) + f' (sum {sum(w.values())})')
|
||||
|
|
@ -13,3 +13,4 @@ export CORNER = TC
|
|||
export SKIP_LAST_GASP = 1
|
||||
export WORK_HOME = /work/out/scratch
|
||||
|
||||
export SYNTH_MEMORY_MAX_BITS = 131072
|
||||
|
|
|
|||
|
|
@ -13,3 +13,4 @@ export CORNER = TC
|
|||
export SKIP_LAST_GASP = 1
|
||||
export WORK_HOME = /work/out/shfl
|
||||
|
||||
export SYNTH_MEMORY_MAX_BITS = 131072
|
||||
|
|
|
|||
|
|
@ -13,3 +13,4 @@ export CORNER = TC
|
|||
export SKIP_LAST_GASP = 1
|
||||
export WORK_HOME = /work/out/xbar
|
||||
|
||||
export SYNTH_MEMORY_MAX_BITS = 131072
|
||||
|
|
|
|||
Loading…
Reference in a new issue