From ccd5a366c13d2e69272035bad478c645288bb957 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Thu, 8 Oct 2026 11:59:10 +0000 Subject: [PATCH 1/9] chip-model: the shadow-k floor lane's RTL, testbenches and ORFS flow (ASAP7) for every class v6 mixer family Co-Authored-By: Claude Fable 5.1 Documents-only replay of 9009ed0f4 (bfccc26ed4d299ad35920bda4aa2af0cc5efbe01) for the box mirror master --- tools/chip-model/rtl/Makefile | 60 +++++++++++++++++++++ tools/chip-model/rtl/flow/arx.mk | 15 ++++++ tools/chip-model/rtl/flow/arx.sdc | 10 ++++ tools/chip-model/rtl/flow/collect.py | 72 +++++++++++++++++++++++++ tools/chip-model/rtl/flow/designs.txt | 9 ++++ tools/chip-model/rtl/flow/fold.mk | 15 ++++++ tools/chip-model/rtl/flow/fold.sdc | 10 ++++ tools/chip-model/rtl/flow/gl2sim.sh | 21 ++++++++ tools/chip-model/rtl/flow/lop3.mk | 15 ++++++ tools/chip-model/rtl/flow/lop3.sdc | 10 ++++ tools/chip-model/rtl/flow/mul.mk | 15 ++++++ tools/chip-model/rtl/flow/mul.sdc | 10 ++++ tools/chip-model/rtl/flow/power.tcl | 23 ++++++++ tools/chip-model/rtl/flow/prmt.mk | 15 ++++++ tools/chip-model/rtl/flow/prmt.sdc | 10 ++++ tools/chip-model/rtl/flow/scratch.mk | 15 ++++++ tools/chip-model/rtl/flow/scratch.sdc | 10 ++++ tools/chip-model/rtl/flow/shfl.mk | 15 ++++++ tools/chip-model/rtl/flow/shfl.sdc | 10 ++++ tools/chip-model/rtl/flow/sim.sh | 9 ++++ tools/chip-model/rtl/flow/tile.mk | 15 ++++++ tools/chip-model/rtl/flow/tile.sdc | 10 ++++ tools/chip-model/rtl/flow/xbar.mk | 15 ++++++ tools/chip-model/rtl/flow/xbar.sdc | 10 ++++ tools/chip-model/rtl/rtl/lane_arx.v | 38 +++++++++++++ tools/chip-model/rtl/rtl/lane_common.vh | 5 ++ tools/chip-model/rtl/rtl/lane_fold.v | 32 +++++++++++ tools/chip-model/rtl/rtl/lane_lop3.v | 25 +++++++++ tools/chip-model/rtl/rtl/lane_mul.v | 36 +++++++++++++ tools/chip-model/rtl/rtl/lane_prmt.v | 29 ++++++++++ tools/chip-model/rtl/rtl/scratch8k.v | 19 +++++++ tools/chip-model/rtl/rtl/shfl32.v | 36 +++++++++++++ tools/chip-model/rtl/rtl/tile8.v | 33 ++++++++++++ tools/chip-model/rtl/rtl/xbar32.v | 29 ++++++++++ tools/chip-model/rtl/tb/tb_lane_arx.v | 20 +++++++ tools/chip-model/rtl/tb/tb_lane_fold.v | 21 ++++++++ tools/chip-model/rtl/tb/tb_lane_lop3.v | 18 +++++++ tools/chip-model/rtl/tb/tb_lane_mul.v | 20 +++++++ tools/chip-model/rtl/tb/tb_lane_prmt.v | 18 +++++++ tools/chip-model/rtl/tb/tb_scratch8k.v | 17 ++++++ tools/chip-model/rtl/tb/tb_shfl32.v | 18 +++++++ tools/chip-model/rtl/tb/tb_tile8.v | 18 +++++++ tools/chip-model/rtl/tb/tb_xbar32.v | 18 +++++++ 43 files changed, 869 insertions(+) create mode 100644 tools/chip-model/rtl/Makefile create mode 100644 tools/chip-model/rtl/flow/arx.mk create mode 100644 tools/chip-model/rtl/flow/arx.sdc create mode 100644 tools/chip-model/rtl/flow/collect.py create mode 100644 tools/chip-model/rtl/flow/designs.txt create mode 100644 tools/chip-model/rtl/flow/fold.mk create mode 100644 tools/chip-model/rtl/flow/fold.sdc create mode 100755 tools/chip-model/rtl/flow/gl2sim.sh create mode 100644 tools/chip-model/rtl/flow/lop3.mk create mode 100644 tools/chip-model/rtl/flow/lop3.sdc create mode 100644 tools/chip-model/rtl/flow/mul.mk create mode 100644 tools/chip-model/rtl/flow/mul.sdc create mode 100644 tools/chip-model/rtl/flow/power.tcl create mode 100644 tools/chip-model/rtl/flow/prmt.mk create mode 100644 tools/chip-model/rtl/flow/prmt.sdc create mode 100644 tools/chip-model/rtl/flow/scratch.mk create mode 100644 tools/chip-model/rtl/flow/scratch.sdc create mode 100644 tools/chip-model/rtl/flow/shfl.mk create mode 100644 tools/chip-model/rtl/flow/shfl.sdc create mode 100755 tools/chip-model/rtl/flow/sim.sh create mode 100644 tools/chip-model/rtl/flow/tile.mk create mode 100644 tools/chip-model/rtl/flow/tile.sdc create mode 100644 tools/chip-model/rtl/flow/xbar.mk create mode 100644 tools/chip-model/rtl/flow/xbar.sdc create mode 100644 tools/chip-model/rtl/rtl/lane_arx.v create mode 100644 tools/chip-model/rtl/rtl/lane_common.vh create mode 100644 tools/chip-model/rtl/rtl/lane_fold.v create mode 100644 tools/chip-model/rtl/rtl/lane_lop3.v create mode 100644 tools/chip-model/rtl/rtl/lane_mul.v create mode 100644 tools/chip-model/rtl/rtl/lane_prmt.v create mode 100644 tools/chip-model/rtl/rtl/scratch8k.v create mode 100644 tools/chip-model/rtl/rtl/shfl32.v create mode 100644 tools/chip-model/rtl/rtl/tile8.v create mode 100644 tools/chip-model/rtl/rtl/xbar32.v create mode 100644 tools/chip-model/rtl/tb/tb_lane_arx.v create mode 100644 tools/chip-model/rtl/tb/tb_lane_fold.v create mode 100644 tools/chip-model/rtl/tb/tb_lane_lop3.v create mode 100644 tools/chip-model/rtl/tb/tb_lane_mul.v create mode 100644 tools/chip-model/rtl/tb/tb_lane_prmt.v create mode 100644 tools/chip-model/rtl/tb/tb_scratch8k.v create mode 100644 tools/chip-model/rtl/tb/tb_shfl32.v create mode 100644 tools/chip-model/rtl/tb/tb_tile8.v create mode 100644 tools/chip-model/rtl/tb/tb_xbar32.v diff --git a/tools/chip-model/rtl/Makefile b/tools/chip-model/rtl/Makefile new file mode 100644 index 000000000..05b786486 --- /dev/null +++ b/tools/chip-model/rtl/Makefile @@ -0,0 +1,60 @@ +# Shadow-k floor lane: RTL -> ASAP7 (Yosys + OpenROAD, the ORFS docker image) -> pJ per op from a +# random-input gate-level simulation. Reproduces every row of docs/analysis/class-v6/floor/shadow-k.md. +# +# make row-arx synthesise, place, route, simulate and report one family +# make rows every family (serially; use -j for parallel on the box) +# make table collect every power log into table.md and table.csv +# +# Runs on build-4 or build-3 (docker, the build user in the docker group). Never on the Mac. +# Images: openroad/orfs:latest (yosys 0.68, OpenROAD) and orfs-sim:latest (the same plus iverilog, +# built once with: docker run --name b -u root openroad/orfs:latest bash -c 'apt-get update && apt-get install -y iverilog' && docker commit b orfs-sim:latest). + +WORK ?= $(abspath .) +ORFS_IMG ?= openroad/orfs:latest +SIM_IMG ?= orfs-sim:latest +UID_GID := $(shell id -u):$(shell id -g) +DOCKER := docker run --rm -u $(UID_GID) -e HOME=/tmp -v $(WORK):/work +ORFS := $(DOCKER) -w /OpenROAD-flow-scripts/flow $(ORFS_IMG) +SIM := $(DOCKER) -w /work $(SIM_IMG) + +DESIGNS := arx mul prmt lop3 fold shfl xbar scratch tile +top = $(shell sed -n 's/^$(1) \([^ ]*\) .*/\1/p' flow/designs.txt) + +# per-family simulation tags (the op field fixed per row where the family has several ops) +SIMS_arx := mix add:+op=0 sub:+op=1 xor:+op=2 or:+op=3 rotl:+op=4 rotr:+op=5 +SIMS_mul := mix mul:+op=0 mulhi:+op=1 mad:+op=2 +SIMS_prmt := mix +SIMS_lop3 := mix +SIMS_fold := mix +SIMS_shfl := mix +SIMS_xbar := mix +SIMS_scratch := mix +SIMS_tile := mix + +.PHONY: rows table clean + +rows: $(addprefix row-,$(DESIGNS)) + +row-%: power-% + @echo "row $* done: $(WORK)/out/$*/logs/asap7/$*/base/power.log" + +flow-%: + mkdir -p out/$* logs + $(ORFS) make DESIGN_CONFIG=/work/flow/$*.mk 2>&1 | tee logs/flow-$*.log + test -f out/$*/results/asap7/$*/base/6_final.v + +sim-%: flow-% + mkdir -p sim/$* + $(SIM) bash /work/flow/gl2sim.sh $* $(call top,$*) 2>&1 | tee logs/gl2sim-$*.log + for t in $(SIMS_$*); do tag=$${t%%:*}; args=$${t#*:}; [ "$$args" = "$$t" ] && args=""; \ + $(SIM) bash /work/flow/sim.sh $* $$tag $$args 2>&1 | tee logs/sim-$*-$$tag.log; done + +power-%: sim-% + $(ORFS) make DESIGN_CONFIG=/work/flow/$*.mk RUN_SCRIPT=/work/flow/power.tcl RUN_LOG_NAME_STEM=power run 2>&1 | tee logs/power-$*.log + +table: + python3 flow/collect.py $(WORK) > table.md + @echo wrote table.md and table.csv + +clean: + rm -rf out sim logs table.md table.csv diff --git a/tools/chip-model/rtl/flow/arx.mk b/tools/chip-model/rtl/flow/arx.mk new file mode 100644 index 000000000..ca1b61a86 --- /dev/null +++ b/tools/chip-model/rtl/flow/arx.mk @@ -0,0 +1,15 @@ +# ORFS design config for the arx shadow-core family (top lane_arx), ASAP7. +export PLATFORM = asap7 +export DESIGN_NAME = lane_arx +export DESIGN_NICKNAME = arx +export VERILOG_FILES = /work/rtl/lane_arx.v +export VERILOG_INCLUDE_DIRS = /work/rtl +export SDC_FILE = /work/flow/arx.sdc +export CORE_UTILIZATION = 40 +export CORE_ASPECT_RATIO = 1 +export CORE_MARGIN = 0.5 +export PLACE_DENSITY = 0.55 +export CORNER = TC +export SKIP_LAST_GASP = 1 +export WORK_HOME = /work/out/arx + diff --git a/tools/chip-model/rtl/flow/arx.sdc b/tools/chip-model/rtl/flow/arx.sdc new file mode 100644 index 000000000..81f424f33 --- /dev/null +++ b/tools/chip-model/rtl/flow/arx.sdc @@ -0,0 +1,10 @@ +current_design lane_arx +set clk_name core_clock +set clk_port_name clk +set clk_period 1000 +set clk_io_pct 0.2 +set clk_port [get_ports $clk_port_name] +create_clock -name $clk_name -period $clk_period $clk_port +set non_clock_inputs [all_inputs -no_clocks] +set_input_delay [expr $clk_period * $clk_io_pct] -clock $clk_name $non_clock_inputs +set_output_delay [expr $clk_period * $clk_io_pct] -clock $clk_name [all_outputs] diff --git a/tools/chip-model/rtl/flow/collect.py b/tools/chip-model/rtl/flow/collect.py new file mode 100644 index 000000000..6ed151f21 --- /dev/null +++ b/tools/chip-model/rtl/flow/collect.py @@ -0,0 +1,72 @@ +#!/usr/bin/env python3 +"""Collect the ORFS power logs into the shadow-k table: pJ per op per family at ASAP7 (TC corner), +scaled to N5, N3 and N2, against the 5090's measured pJ per counted op (counter-asic-4-research.md 15.1a). +Usage: collect.py (writes table.csv beside, prints table.md).""" +import re, sys, os, csv + +work = sys.argv[1] if len(sys.argv) > 1 else '.' +# ops per cycle per design (the per-op divisor) and the GPU row each family is read against +OPS = {'arx': 1, 'mul': 1, 'prmt': 1, 'lop3': 1, 'fold': 1, 'shfl': 32, 'xbar': 32, 'scratch': 1, 'tile': 512} +# 5090 measured pJ per counted op: (unlocked, at the 1,300 MHz lock); 15.1a +GPU = { + 'arx:mix': (11.3, 6.2), 'arx:add': (11.3, 6.2), 'arx:sub': (11.3, 6.2), 'arx:xor': (11.3, 6.2), 'arx:or': (11.3, 6.2), + 'arx:rotl': (11.3, 6.2), 'arx:rotr': (11.3, 6.2), + 'mul:mix': (13.9, 8.3), 'mul:mul': (13.9, 8.3), 'mul:mad': (13.9, 8.3), 'mul:mulhi': (39.6, 21.0), + 'prmt:mix': (22.3, 11.5), 'lop3:mix': (24.1, 13.0), + 'fold:mix': (13.9, 8.3), # the fold is a multiply plus a rotate and masks: read against int_mul + 'shfl:mix': (55.8, 29.4), 'xbar:mix': (55.8, 29.4), + 'scratch:mix': (2400.0, 1400.0), # the card's L2 hit (no shared-memory probe measured: owed) + 'tile:mix': (4.1, 2.2), # dependent u8 m8n8k16 per MAC; the wide s8 tile reads 1.36 / 0.83 +} +# per-node energy scaling from ASAP7 (a 7 nm-class predictive PDK at 0.70 V), approximate and claimed: +# N7 -> N5 x0.70 (TSMC: "30 percent lower power at the same speed"), N5 -> N3E x0.72 (TSMC: 25 to 30 percent), +# N3E -> N2 x0.72 (TSMC: 25 to 30 percent). Sources in shadow-k.md section 3. +SCALE = {'ASAP7': 1.0, 'N5': 0.70, 'N3': 0.70 * 0.72, 'N2': 0.70 * 0.72 * 0.72} + +def parse_log(path): + txt = open(path, errors='replace').read() + period = None + m = re.search(r'FLOORK clock_period_ps ([\d.]+)', txt); period = float(m.group(1)) if m else None + cells = re.search(r'FLOORK cells (\d+)', txt); cells = int(cells.group(1)) if cells else None + rows = {} + for sec in re.split(r'FLOORK === ', txt)[1:]: + head, body = sec.split(' ===', 1) + tag = head.strip().replace('POWER_VCD ', 'vcd:').replace('POWER_PROPAGATED_0.5', 'prop') + # OpenSTA report_power: "Total 100.0%" in watts + tm = re.search(r'^Total\s+([\d.eE+-]+)\s+([\d.eE+-]+)\s+([\d.eE+-]+)\s+([\d.eE+-]+)', body, re.M) + ann = re.search(r'(\d+)\s*\(\s*([\d.]+)%\)\s*annotated', body) + if tm: + rows[tag] = dict(internal=float(tm.group(1)), switching=float(tm.group(2)), leakage=float(tm.group(3)), + total=float(tm.group(4)), annotated=ann.group(2) if ann else '') + return period, cells, rows + +out = [] +for d in OPS: + log = os.path.join(work, 'out', d, 'logs', 'asap7', d, 'base', 'power.log') + if not os.path.exists(log): + continue + period, cells, rows = parse_log(log) + for tag, r in rows.items(): + sub = tag.split(':')[1] if ':' in tag else 'prop' + key = f'{d}:{sub}' if sub != 'prop' else f'{d}:mix' + gpu = GPU.get(key, (None, None)) + pj = r['total'] * period * 1e-12 / OPS[d] * 1e12 # W * s / ops -> pJ + pj_dyn = (r['internal'] + r['switching']) * period / OPS[d] + row = dict(family=d, sim=tag, period_ps=period, cells=cells, ops_per_cycle=OPS[d], + total_W=r['total'], leak_W=r['leakage'], annotated_pct=r['annotated'], + pJ_asap7=pj, pJ_dyn_asap7=pj_dyn) + for node, s in SCALE.items(): + row[f'pJ_{node}'] = pj * s + row[f'k_{node}_lock'] = (pj * s / gpu[1]) if gpu[1] else None + row[f'k_{node}_unlocked'] = (pj * s / gpu[0]) if gpu[0] else None + row['gpu_pJ_unlocked'] = gpu[0]; row['gpu_pJ_lock'] = gpu[1] + out.append(row) + +if out: + with open(os.path.join(work, 'table.csv'), 'w', newline='') as f: + w = csv.DictWriter(f, fieldnames=list(out[0].keys())); w.writeheader(); w.writerows(out) +print('| Family | Sim | Period ps | Cells | Ops/cycle | Total W | Leak W | VCD annotated | pJ/op ASAP7 | pJ/op N5 | pJ/op N3 | pJ/op N2 | 5090 pJ/op unlocked / lock | k at N3 (lock) | k at N2 (lock) | k at N3 (unlocked) |') +print('|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|') +for r in out: + f = lambda v, n=3: ('' if v is None else (f'{v:.{n}g}' if isinstance(v, float) else str(v))) + print(f"| {r['family']} | {r['sim']} | {f(r['period_ps'])} | {r['cells']} | {r['ops_per_cycle']} | {f(r['total_W'])} | {f(r['leak_W'])} | {r['annotated_pct']} | {f(r['pJ_asap7'])} | {f(r['pJ_N5'])} | {f(r['pJ_N3'])} | {f(r['pJ_N2'])} | {f(r['gpu_pJ_unlocked'])} / {f(r['gpu_pJ_lock'])} | {f(r['k_N3_lock'])} | {f(r['k_N2_lock'])} | {f(r['k_N3_unlocked'])} |") diff --git a/tools/chip-model/rtl/flow/designs.txt b/tools/chip-model/rtl/flow/designs.txt new file mode 100644 index 000000000..133bafae3 --- /dev/null +++ b/tools/chip-model/rtl/flow/designs.txt @@ -0,0 +1,9 @@ +arx lane_arx 1000 +mul lane_mul 1500 +prmt lane_prmt 1000 +lop3 lane_lop3 1000 +fold lane_fold 1500 +shfl shfl32 1000 +xbar xbar32 1000 +scratch scratch8k 1500 +tile tile8 2000 diff --git a/tools/chip-model/rtl/flow/fold.mk b/tools/chip-model/rtl/flow/fold.mk new file mode 100644 index 000000000..d5fe9d932 --- /dev/null +++ b/tools/chip-model/rtl/flow/fold.mk @@ -0,0 +1,15 @@ +# ORFS design config for the fold shadow-core family (top lane_fold), ASAP7. +export PLATFORM = asap7 +export DESIGN_NAME = lane_fold +export DESIGN_NICKNAME = fold +export VERILOG_FILES = /work/rtl/lane_fold.v +export VERILOG_INCLUDE_DIRS = /work/rtl +export SDC_FILE = /work/flow/fold.sdc +export CORE_UTILIZATION = 40 +export CORE_ASPECT_RATIO = 1 +export CORE_MARGIN = 0.5 +export PLACE_DENSITY = 0.55 +export CORNER = TC +export SKIP_LAST_GASP = 1 +export WORK_HOME = /work/out/fold + diff --git a/tools/chip-model/rtl/flow/fold.sdc b/tools/chip-model/rtl/flow/fold.sdc new file mode 100644 index 000000000..9dad665a3 --- /dev/null +++ b/tools/chip-model/rtl/flow/fold.sdc @@ -0,0 +1,10 @@ +current_design lane_fold +set clk_name core_clock +set clk_port_name clk +set clk_period 1500 +set clk_io_pct 0.2 +set clk_port [get_ports $clk_port_name] +create_clock -name $clk_name -period $clk_period $clk_port +set non_clock_inputs [all_inputs -no_clocks] +set_input_delay [expr $clk_period * $clk_io_pct] -clock $clk_name $non_clock_inputs +set_output_delay [expr $clk_period * $clk_io_pct] -clock $clk_name [all_outputs] diff --git a/tools/chip-model/rtl/flow/gl2sim.sh b/tools/chip-model/rtl/flow/gl2sim.sh new file mode 100755 index 000000000..17c94928f --- /dev/null +++ b/tools/chip-model/rtl/flow/gl2sim.sh @@ -0,0 +1,21 @@ +#!/usr/bin/env bash +# gl2sim.sh : netlist -> sim netlist with cell bodies from the liberty. +set -euo pipefail +name=$1; top=$2 +net=/work/out/$name/results/asap7/$name/base/6_final.v +lib=/OpenROAD-flow-scripts/flow/platforms/asap7/lib/NLDM +out=/work/sim/$name; mkdir -p $out +cat > $out/gl2sim.ys <.mk RUN_SCRIPT=/work/flow/power.tcl run +source $::env(SCRIPTS_DIR)/load.tcl +load_design 6_final.odb 6_final.sdc +set spef $::env(RESULTS_DIR)/6_final.spef +if { [file exists $spef] } { read_spef $spef } else { estimate_parasitics -global_routing } +puts "FLOORK clock_period_ps [expr [get_property [lindex [all_clocks] 0] period]]" +puts "FLOORK cells [llength [get_cells *]]" +report_tns +report_wns +puts "FLOORK === POWER_PROPAGATED_0.5 ===" +set_power_activity -input -activity 0.5 -duty 0.5 +set_power_activity -input_port rst -activity 0 -duty 0 +report_power +foreach vcd [glob -nocomplain /work/sim/$::env(DESIGN_NICKNAME)/*.vcd] { + set tag [file rootname [file tail $vcd]] + puts "FLOORK === POWER_VCD $tag ===" + read_vcd -scope tb/dut $vcd + if { [info commands report_activity_annotation] != "" } { report_activity_annotation } + report_power +} +puts "FLOORK done" diff --git a/tools/chip-model/rtl/flow/prmt.mk b/tools/chip-model/rtl/flow/prmt.mk new file mode 100644 index 000000000..e17ad1bde --- /dev/null +++ b/tools/chip-model/rtl/flow/prmt.mk @@ -0,0 +1,15 @@ +# ORFS design config for the prmt shadow-core family (top lane_prmt), ASAP7. +export PLATFORM = asap7 +export DESIGN_NAME = lane_prmt +export DESIGN_NICKNAME = prmt +export VERILOG_FILES = /work/rtl/lane_prmt.v +export VERILOG_INCLUDE_DIRS = /work/rtl +export SDC_FILE = /work/flow/prmt.sdc +export CORE_UTILIZATION = 40 +export CORE_ASPECT_RATIO = 1 +export CORE_MARGIN = 0.5 +export PLACE_DENSITY = 0.55 +export CORNER = TC +export SKIP_LAST_GASP = 1 +export WORK_HOME = /work/out/prmt + diff --git a/tools/chip-model/rtl/flow/prmt.sdc b/tools/chip-model/rtl/flow/prmt.sdc new file mode 100644 index 000000000..dabb0c22f --- /dev/null +++ b/tools/chip-model/rtl/flow/prmt.sdc @@ -0,0 +1,10 @@ +current_design lane_prmt +set clk_name core_clock +set clk_port_name clk +set clk_period 1000 +set clk_io_pct 0.2 +set clk_port [get_ports $clk_port_name] +create_clock -name $clk_name -period $clk_period $clk_port +set non_clock_inputs [all_inputs -no_clocks] +set_input_delay [expr $clk_period * $clk_io_pct] -clock $clk_name $non_clock_inputs +set_output_delay [expr $clk_period * $clk_io_pct] -clock $clk_name [all_outputs] diff --git a/tools/chip-model/rtl/flow/scratch.mk b/tools/chip-model/rtl/flow/scratch.mk new file mode 100644 index 000000000..a8fdb0197 --- /dev/null +++ b/tools/chip-model/rtl/flow/scratch.mk @@ -0,0 +1,15 @@ +# ORFS design config for the scratch shadow-core family (top scratch8k), ASAP7. +export PLATFORM = asap7 +export DESIGN_NAME = scratch8k +export DESIGN_NICKNAME = scratch +export VERILOG_FILES = /work/rtl/scratch8k.v +export VERILOG_INCLUDE_DIRS = /work/rtl +export SDC_FILE = /work/flow/scratch.sdc +export CORE_UTILIZATION = 40 +export CORE_ASPECT_RATIO = 1 +export CORE_MARGIN = 0.5 +export PLACE_DENSITY = 0.55 +export CORNER = TC +export SKIP_LAST_GASP = 1 +export WORK_HOME = /work/out/scratch + diff --git a/tools/chip-model/rtl/flow/scratch.sdc b/tools/chip-model/rtl/flow/scratch.sdc new file mode 100644 index 000000000..822efb6cb --- /dev/null +++ b/tools/chip-model/rtl/flow/scratch.sdc @@ -0,0 +1,10 @@ +current_design scratch8k +set clk_name core_clock +set clk_port_name clk +set clk_period 1500 +set clk_io_pct 0.2 +set clk_port [get_ports $clk_port_name] +create_clock -name $clk_name -period $clk_period $clk_port +set non_clock_inputs [all_inputs -no_clocks] +set_input_delay [expr $clk_period * $clk_io_pct] -clock $clk_name $non_clock_inputs +set_output_delay [expr $clk_period * $clk_io_pct] -clock $clk_name [all_outputs] diff --git a/tools/chip-model/rtl/flow/shfl.mk b/tools/chip-model/rtl/flow/shfl.mk new file mode 100644 index 000000000..a73d97541 --- /dev/null +++ b/tools/chip-model/rtl/flow/shfl.mk @@ -0,0 +1,15 @@ +# ORFS design config for the shfl shadow-core family (top shfl32), ASAP7. +export PLATFORM = asap7 +export DESIGN_NAME = shfl32 +export DESIGN_NICKNAME = shfl +export VERILOG_FILES = /work/rtl/shfl32.v +export VERILOG_INCLUDE_DIRS = /work/rtl +export SDC_FILE = /work/flow/shfl.sdc +export CORE_UTILIZATION = 40 +export CORE_ASPECT_RATIO = 1 +export CORE_MARGIN = 0.5 +export PLACE_DENSITY = 0.55 +export CORNER = TC +export SKIP_LAST_GASP = 1 +export WORK_HOME = /work/out/shfl + diff --git a/tools/chip-model/rtl/flow/shfl.sdc b/tools/chip-model/rtl/flow/shfl.sdc new file mode 100644 index 000000000..530484342 --- /dev/null +++ b/tools/chip-model/rtl/flow/shfl.sdc @@ -0,0 +1,10 @@ +current_design shfl32 +set clk_name core_clock +set clk_port_name clk +set clk_period 1000 +set clk_io_pct 0.2 +set clk_port [get_ports $clk_port_name] +create_clock -name $clk_name -period $clk_period $clk_port +set non_clock_inputs [all_inputs -no_clocks] +set_input_delay [expr $clk_period * $clk_io_pct] -clock $clk_name $non_clock_inputs +set_output_delay [expr $clk_period * $clk_io_pct] -clock $clk_name [all_outputs] diff --git a/tools/chip-model/rtl/flow/sim.sh b/tools/chip-model/rtl/flow/sim.sh new file mode 100755 index 000000000..f7c7cbdd6 --- /dev/null +++ b/tools/chip-model/rtl/flow/sim.sh @@ -0,0 +1,9 @@ +#!/usr/bin/env bash +# sim.sh [plusargs...] : gate-level random-input simulation -> /work/sim//.vcd +set -euo pipefail +name=$1; tag=$2; shift 2 +out=/work/sim/$name +simcells=$(yosys-config --datdir)/simcells.v +iverilog -g2005 -o $out/sim_$tag $out/sim_net.v /work/tb/tb_$(sed -n "s/^$name \([^ ]*\) .*/\1/p" /work/flow/designs.txt).v $simcells +( cd $out && vvp -n sim_$tag "$@" | tee sim_$tag.log && mv dump.vcd $tag.vcd ) +ls -la $out/$tag.vcd diff --git a/tools/chip-model/rtl/flow/tile.mk b/tools/chip-model/rtl/flow/tile.mk new file mode 100644 index 000000000..69f85f03c --- /dev/null +++ b/tools/chip-model/rtl/flow/tile.mk @@ -0,0 +1,15 @@ +# ORFS design config for the tile shadow-core family (top tile8), ASAP7. +export PLATFORM = asap7 +export DESIGN_NAME = tile8 +export DESIGN_NICKNAME = tile +export VERILOG_FILES = /work/rtl/tile8.v +export VERILOG_INCLUDE_DIRS = /work/rtl +export SDC_FILE = /work/flow/tile.sdc +export CORE_UTILIZATION = 40 +export CORE_ASPECT_RATIO = 1 +export CORE_MARGIN = 0.5 +export PLACE_DENSITY = 0.55 +export CORNER = TC +export SKIP_LAST_GASP = 1 +export WORK_HOME = /work/out/tile + diff --git a/tools/chip-model/rtl/flow/tile.sdc b/tools/chip-model/rtl/flow/tile.sdc new file mode 100644 index 000000000..2bac92473 --- /dev/null +++ b/tools/chip-model/rtl/flow/tile.sdc @@ -0,0 +1,10 @@ +current_design tile8 +set clk_name core_clock +set clk_port_name clk +set clk_period 2000 +set clk_io_pct 0.2 +set clk_port [get_ports $clk_port_name] +create_clock -name $clk_name -period $clk_period $clk_port +set non_clock_inputs [all_inputs -no_clocks] +set_input_delay [expr $clk_period * $clk_io_pct] -clock $clk_name $non_clock_inputs +set_output_delay [expr $clk_period * $clk_io_pct] -clock $clk_name [all_outputs] diff --git a/tools/chip-model/rtl/flow/xbar.mk b/tools/chip-model/rtl/flow/xbar.mk new file mode 100644 index 000000000..01026b194 --- /dev/null +++ b/tools/chip-model/rtl/flow/xbar.mk @@ -0,0 +1,15 @@ +# ORFS design config for the xbar shadow-core family (top xbar32), ASAP7. +export PLATFORM = asap7 +export DESIGN_NAME = xbar32 +export DESIGN_NICKNAME = xbar +export VERILOG_FILES = /work/rtl/xbar32.v +export VERILOG_INCLUDE_DIRS = /work/rtl +export SDC_FILE = /work/flow/xbar.sdc +export CORE_UTILIZATION = 40 +export CORE_ASPECT_RATIO = 1 +export CORE_MARGIN = 0.5 +export PLACE_DENSITY = 0.55 +export CORNER = TC +export SKIP_LAST_GASP = 1 +export WORK_HOME = /work/out/xbar + diff --git a/tools/chip-model/rtl/flow/xbar.sdc b/tools/chip-model/rtl/flow/xbar.sdc new file mode 100644 index 000000000..509dd19ff --- /dev/null +++ b/tools/chip-model/rtl/flow/xbar.sdc @@ -0,0 +1,10 @@ +current_design xbar32 +set clk_name core_clock +set clk_port_name clk +set clk_period 1000 +set clk_io_pct 0.2 +set clk_port [get_ports $clk_port_name] +create_clock -name $clk_name -period $clk_period $clk_port +set non_clock_inputs [all_inputs -no_clocks] +set_input_delay [expr $clk_period * $clk_io_pct] -clock $clk_name $non_clock_inputs +set_output_delay [expr $clk_period * $clk_io_pct] -clock $clk_name [all_outputs] diff --git a/tools/chip-model/rtl/rtl/lane_arx.v b/tools/chip-model/rtl/rtl/lane_arx.v new file mode 100644 index 000000000..be590f9a1 --- /dev/null +++ b/tools/chip-model/rtl/rtl/lane_arx.v @@ -0,0 +1,38 @@ +// 32-bit int ARX lane: add, sub, xor, or, rotl by immediate, rotr by register (the class v4/v6 families +// add, sub, xor, or, rotl, rotr). op: 0 add 1 sub 2 xor 3 or 4 rotl-imm 5 rotr-var 6 add 7 xor. +`include "lane_common.vh" +module lane_arx( + input clk, input rst, + input [2:0] op, input [2:0] dst, input [2:0] src, input [4:0] rot, + input ld_en, input [31:0] ld_val, + output [31:0] out); + reg [31:0] rf [0:7]; + reg [2:0] op_q, dst_q, src_q; reg [4:0] rot_q; reg ld_q; reg [31:0] ld_val_q; + integer i; + always @(posedge clk) begin + if (rst) begin op_q <= 0; dst_q <= 0; src_q <= 0; rot_q <= 0; ld_q <= 0; ld_val_q <= 0; end + else begin op_q <= op; dst_q <= dst; src_q <= src; rot_q <= rot; ld_q <= ld_en; ld_val_q <= ld_val; end + end + wire [31:0] d = rf[dst_q]; + wire [31:0] s = rf[src_q]; + wire [4:0] rn = (rot_q == 5'd0) ? 5'd1 : rot_q; // rotate by 1..31 + wire [4:0] sn = (s[4:0] == 5'd0) ? 5'd1 : s[4:0]; + reg [31:0] res; + always @* begin + case (op_q) + 3'd0: res = d + s; + 3'd1: res = d - s; + 3'd2: res = d ^ s; + 3'd3: res = d | s; + 3'd4: res = `ROTL32(d, rn); + 3'd5: res = `ROTR32(d, sn); + 3'd6: res = d + s; + default: res = d ^ s; + endcase + end + always @(posedge clk) begin + if (rst) begin for (i = 0; i < 8; i = i + 1) rf[i] <= 32'h9e3779b9 * (i + 1); end + else rf[dst_q] <= ld_q ? ld_val_q : res; + end + assign out = res; +endmodule diff --git a/tools/chip-model/rtl/rtl/lane_common.vh b/tools/chip-model/rtl/rtl/lane_common.vh new file mode 100644 index 000000000..c977a0d4c --- /dev/null +++ b/tools/chip-model/rtl/rtl/lane_common.vh @@ -0,0 +1,5 @@ +// Shared helpers for the shadow-core lanes (class v5/v6 mixer draw space). +// A lane = an instruction register, an 8 x 32-bit register window (flops), two or three +// read ports through muxes, one functional unit, one write port. One op per cycle. +`define ROTL32(x, n) (((x) << (n)) | ((x) >> (32 - (n)))) +`define ROTR32(x, n) (((x) >> (n)) | ((x) << (32 - (n)))) diff --git a/tools/chip-model/rtl/rtl/lane_fold.v b/tools/chip-model/rtl/rtl/lane_fold.v new file mode 100644 index 000000000..5b1fbc890 --- /dev/null +++ b/tools/chip-model/rtl/rtl/lane_fold.v @@ -0,0 +1,32 @@ +// The index fold (class v6 layer 1, the ring-A rule): idx = ((rotl(x * M, R) & WM) | OFF) & MASK. +// M, R, WM, OFF and MASK are the era's constants, held in registers (loaded by cfg_en, then static). +// One fold per cycle on a register read; the index is written back to the lane's address register. +`include "lane_common.vh" +module lane_fold( + input clk, input rst, + input [2:0] dst, input [2:0] src, + input cfg_en, input [31:0] cfg_m, input [4:0] cfg_r, input [31:0] cfg_wm, input [31:0] cfg_off, input [31:0] cfg_mask, + input ld_en, input [31:0] ld_val, + output [31:0] out); + reg [31:0] rf [0:7]; + reg [2:0] dst_q, src_q; reg ld_q; reg [31:0] ld_val_q; + reg [31:0] m_q, wm_q, off_q, mask_q; reg [4:0] r_q; + integer i; + always @(posedge clk) begin + if (rst) begin dst_q <= 0; src_q <= 0; ld_q <= 0; ld_val_q <= 0; m_q <= 32'h9e3779b1; r_q <= 5'd13; wm_q <= 32'h0fffffc0; off_q <= 32'h3; mask_q <= 32'h0fffffff; end + else begin + dst_q <= dst; src_q <= src; ld_q <= ld_en; ld_val_q <= ld_val; + if (cfg_en) begin m_q <= cfg_m | 32'h1; r_q <= cfg_r; wm_q <= cfg_wm; off_q <= cfg_off; mask_q <= cfg_mask; end + end + end + wire [31:0] x = rf[src_q]; + wire [31:0] prod = x * m_q; + wire [4:0] rn = (r_q == 5'd0) ? 5'd1 : r_q; + wire [31:0] y = `ROTL32(prod, rn); + wire [31:0] res = ((y & wm_q) | off_q) & mask_q; + always @(posedge clk) begin + if (rst) begin for (i = 0; i < 8; i = i + 1) rf[i] <= 32'h9e3779b9 * (i + 1) + 4; end + else rf[dst_q] <= ld_q ? ld_val_q : (res ^ rf[dst_q]); + end + assign out = res; +endmodule diff --git a/tools/chip-model/rtl/rtl/lane_lop3.v b/tools/chip-model/rtl/rtl/lane_lop3.v new file mode 100644 index 000000000..68c1c8690 --- /dev/null +++ b/tools/chip-model/rtl/rtl/lane_lop3.v @@ -0,0 +1,25 @@ +// Three-input logic lane (LOP3): an 8-bit truth table over (d, s, s2), bitwise. +`include "lane_common.vh" +module lane_lop3( + input clk, input rst, + input [2:0] dst, input [2:0] src, input [2:0] src2, input [7:0] lut, + input ld_en, input [31:0] ld_val, + output [31:0] out); + reg [31:0] rf [0:7]; + reg [2:0] dst_q, src_q, src2_q; reg [7:0] lut_q; reg ld_q; reg [31:0] ld_val_q; + integer i, b; + always @(posedge clk) begin + if (rst) begin dst_q <= 0; src_q <= 0; src2_q <= 0; lut_q <= 0; ld_q <= 0; ld_val_q <= 0; end + else begin dst_q <= dst; src_q <= src; src2_q <= src2; lut_q <= lut; ld_q <= ld_en; ld_val_q <= ld_val; end + end + wire [31:0] d = rf[dst_q]; + wire [31:0] s = rf[src_q]; + wire [31:0] s2 = rf[src2_q]; + reg [31:0] res; + always @* for (b = 0; b < 32; b = b + 1) res[b] = lut_q[{d[b], s[b], s2[b]}]; + always @(posedge clk) begin + if (rst) begin for (i = 0; i < 8; i = i + 1) rf[i] <= 32'h9e3779b9 * (i + 1) + 3; end + else rf[dst_q] <= ld_q ? ld_val_q : res; + end + assign out = res; +endmodule diff --git a/tools/chip-model/rtl/rtl/lane_mul.v b/tools/chip-model/rtl/rtl/lane_mul.v new file mode 100644 index 000000000..0b8b0736c --- /dev/null +++ b/tools/chip-model/rtl/rtl/lane_mul.v @@ -0,0 +1,36 @@ +// 32 x 32 multiplier lane: mul (low 32), mulhi (high 32 of the 64-bit product), mad (src*src2 + dst, low 32). +// op: 0 mul 1 mulhi 2 mad 3 mul. The mad reads three registers. +`include "lane_common.vh" +module lane_mul( + input clk, input rst, + input [1:0] op, input [2:0] dst, input [2:0] src, input [2:0] src2, + input ld_en, input [31:0] ld_val, + output [31:0] out); + reg [31:0] rf [0:7]; + reg [1:0] op_q; reg [2:0] dst_q, src_q, src2_q; reg ld_q; reg [31:0] ld_val_q; + integer i; + always @(posedge clk) begin + if (rst) begin op_q <= 0; dst_q <= 0; src_q <= 0; src2_q <= 0; ld_q <= 0; ld_val_q <= 0; end + else begin op_q <= op; dst_q <= dst; src_q <= src; src2_q <= src2; ld_q <= ld_en; ld_val_q <= ld_val; end + end + wire [31:0] d = rf[dst_q]; + wire [31:0] s = rf[src_q]; + wire [31:0] s2 = rf[src2_q]; + wire mad = (op_q == 2'd2); + wire [31:0] ma = mad ? s : d; + wire [31:0] mb = mad ? s2 : s; + wire [63:0] p = ma * mb; + reg [31:0] res; + always @* begin + case (op_q) + 2'd1: res = p[63:32]; + 2'd2: res = p[31:0] + d; + default: res = p[31:0]; + endcase + end + always @(posedge clk) begin + if (rst) begin for (i = 0; i < 8; i = i + 1) rf[i] <= 32'h9e3779b9 * (i + 1) + 1; end + else rf[dst_q] <= ld_q ? ld_val_q : res; + end + assign out = res; +endmodule diff --git a/tools/chip-model/rtl/rtl/lane_prmt.v b/tools/chip-model/rtl/rtl/lane_prmt.v new file mode 100644 index 000000000..d8d472f28 --- /dev/null +++ b/tools/chip-model/rtl/rtl/lane_prmt.v @@ -0,0 +1,29 @@ +// Byte-permute lane (PRMT): four output bytes, each one of the eight bytes of {s, d}, by a 16-bit selector +// (4 bits per output byte; bit 3 = sign-replicate as on the card). +`include "lane_common.vh" +module lane_prmt( + input clk, input rst, + input [2:0] dst, input [2:0] src, input [15:0] sel, + input ld_en, input [31:0] ld_val, + output [31:0] out); + reg [31:0] rf [0:7]; + reg [2:0] dst_q, src_q; reg [15:0] sel_q; reg ld_q; reg [31:0] ld_val_q; + integer i; + always @(posedge clk) begin + if (rst) begin dst_q <= 0; src_q <= 0; sel_q <= 0; ld_q <= 0; ld_val_q <= 0; end + else begin dst_q <= dst; src_q <= src; sel_q <= sel; ld_q <= ld_en; ld_val_q <= ld_val; end + end + wire [31:0] d = rf[dst_q]; + wire [31:0] s = rf[src_q]; + wire [63:0] bytes = {s, d}; + function [7:0] pick; input [63:0] b; input [3:0] k; + reg [7:0] v; + begin v = b[8*k[2:0] +: 8]; pick = k[3] ? {8{v[7]}} : v; end + endfunction + wire [31:0] res = {pick(bytes, sel_q[15:12]), pick(bytes, sel_q[11:8]), pick(bytes, sel_q[7:4]), pick(bytes, sel_q[3:0])}; + always @(posedge clk) begin + if (rst) begin for (i = 0; i < 8; i = i + 1) rf[i] <= 32'h9e3779b9 * (i + 1) + 2; end + else rf[dst_q] <= ld_q ? ld_val_q : res; + end + assign out = res; +endmodule diff --git a/tools/chip-model/rtl/rtl/scratch8k.v b/tools/chip-model/rtl/rtl/scratch8k.v new file mode 100644 index 000000000..8df1c1346 --- /dev/null +++ b/tools/chip-model/rtl/rtl/scratch8k.v @@ -0,0 +1,19 @@ +// 8 KB random-read scratch (2,048 x 32-bit) as a flop array: one random read per cycle (the op), a write on +// one cycle in eight. A flop array is the pessimistic form of the chip's L1; an SRAM macro reads lower. +module scratch8k( + input clk, input rst, + input [10:0] raddr, input we, input [10:0] waddr, input [31:0] wdata, + output reg [31:0] rdata); + reg [31:0] mem [0:2047]; + reg [10:0] raddr_q, waddr_q; reg we_q; reg [31:0] wdata_q; + integer i; + always @(posedge clk) begin + if (rst) begin raddr_q <= 0; waddr_q <= 0; we_q <= 0; wdata_q <= 0; end + else begin raddr_q <= raddr; waddr_q <= waddr; we_q <= we; wdata_q <= wdata; end + end + always @(posedge clk) begin + if (rst) begin for (i = 0; i < 2048; i = i + 1) mem[i] <= 32'h9e3779b9 * (i + 1); end + else if (we_q) mem[waddr_q] <= wdata_q; + end + always @(posedge clk) rdata <= rst ? 32'd0 : mem[raddr_q]; +endmodule diff --git a/tools/chip-model/rtl/rtl/shfl32.v b/tools/chip-model/rtl/rtl/shfl32.v new file mode 100644 index 000000000..1c88fbf45 --- /dev/null +++ b/tools/chip-model/rtl/rtl/shfl32.v @@ -0,0 +1,36 @@ +// 32-lane xor-mask shuffle over a 1 KB register window (32 lanes x 8 x 32 bits): +// r[dst][lane] ^= r[src][lane ^ m] for every lane, one instruction per cycle (32 lane-ops). +// The network is a 5-stage butterfly (one 2:1 mux per bit per stage). +`include "lane_common.vh" +module shfl32( + input clk, input rst, + input [2:0] dst, input [2:0] src, input [4:0] m, + input ld_en, input [4:0] ld_lane, input [31:0] ld_val, + output [31:0] out); + reg [31:0] rf [0:255]; // rf[lane*8 + reg] + reg [2:0] dst_q, src_q; reg [4:0] m_q, ld_lane_q; reg ld_q; reg [31:0] ld_val_q; + integer i, l; + always @(posedge clk) begin + if (rst) begin dst_q <= 0; src_q <= 0; m_q <= 0; ld_q <= 0; ld_lane_q <= 0; ld_val_q <= 0; end + else begin dst_q <= dst; src_q <= src; m_q <= m; ld_q <= ld_en; ld_lane_q <= ld_lane; ld_val_q <= ld_val; end + end + // explicit butterfly + wire [31:0] b0 [0:31]; wire [31:0] b1 [0:31]; wire [31:0] b2 [0:31]; wire [31:0] b3 [0:31]; wire [31:0] b4 [0:31]; wire [31:0] b5 [0:31]; + genvar g; + generate for (g = 0; g < 32; g = g + 1) begin : bf + assign b0[g] = rf[g*8 + src_q]; + assign b1[g] = m_q[0] ? b0[g ^ 1] : b0[g]; + assign b2[g] = m_q[1] ? b1[g ^ 2] : b1[g]; + assign b3[g] = m_q[2] ? b2[g ^ 4] : b2[g]; + assign b4[g] = m_q[3] ? b3[g ^ 8] : b3[g]; + assign b5[g] = m_q[4] ? b4[g ^ 16] : b4[g]; + end endgenerate + always @(posedge clk) begin + if (rst) begin for (i = 0; i < 256; i = i + 1) rf[i] <= 32'h9e3779b9 * (i + 1) + 5; end + else begin + for (l = 0; l < 32; l = l + 1) + rf[l*8 + dst_q] <= (ld_q && ld_lane_q == l) ? ld_val_q : (rf[l*8 + dst_q] ^ b5[l]); + end + end + assign out = b5[0] ^ b5[17]; +endmodule diff --git a/tools/chip-model/rtl/rtl/tile8.v b/tools/chip-model/rtl/rtl/tile8.v new file mode 100644 index 000000000..5c6b2e94a --- /dev/null +++ b/tools/chip-model/rtl/rtl/tile8.v @@ -0,0 +1,33 @@ +// int8 8x8x8 tile multiply: C[8][8] (int32) += A[8][8] (int8) x B[8][8] (int8): 512 MACs per cycle. +// A and B are loaded from the inputs every cycle; the accumulators clear on clr. +module tile8( + input clk, input rst, input clr, + input [511:0] a_in, input [511:0] b_in, + input [2:0] sel_r, input [2:0] sel_c, + output [31:0] out); + reg [511:0] a_q, b_q; reg clr_q; reg [2:0] sr_q, sc_q; + reg [31:0] c [0:63]; + integer i; + always @(posedge clk) begin + if (rst) begin a_q <= 0; b_q <= 0; clr_q <= 1; sr_q <= 0; sc_q <= 0; end + else begin a_q <= a_in; b_q <= b_in; clr_q <= clr; sr_q <= sel_r; sc_q <= sel_c; end + end + genvar r, cc, k; + generate for (r = 0; r < 8; r = r + 1) begin : row + for (cc = 0; cc < 8; cc = cc + 1) begin : col + wire signed [19:0] dot; + wire signed [15:0] p [0:7]; + for (k = 0; k < 8; k = k + 1) begin : mk + wire signed [7:0] av = a_q[(r*8 + k)*8 +: 8]; + wire signed [7:0] bv = b_q[(k*8 + cc)*8 +: 8]; + assign p[k] = av * bv; + end + assign dot = p[0] + p[1] + p[2] + p[3] + p[4] + p[5] + p[6] + p[7]; + always @(posedge clk) begin + if (rst || clr_q) c[r*8 + cc] <= 32'd0; + else c[r*8 + cc] <= c[r*8 + cc] + {{12{dot[19]}}, dot}; + end + end + end endgenerate + assign out = c[{sr_q, sc_q}]; +endmodule diff --git a/tools/chip-model/rtl/rtl/xbar32.v b/tools/chip-model/rtl/rtl/xbar32.v new file mode 100644 index 000000000..52b7f7398 --- /dev/null +++ b/tools/chip-model/rtl/rtl/xbar32.v @@ -0,0 +1,29 @@ +// 32-lane general crossbar over the same 1 KB window: each lane picks any source lane (a 5-bit select per +// lane), the upper bound on a shuffle's network cost. r[dst][lane] ^= r[src][sel[lane]]. +`include "lane_common.vh" +module xbar32( + input clk, input rst, + input [2:0] dst, input [2:0] src, input [159:0] sel, + input ld_en, input [4:0] ld_lane, input [31:0] ld_val, + output [31:0] out); + reg [31:0] rf [0:255]; + reg [2:0] dst_q, src_q; reg [159:0] sel_q; reg [4:0] ld_lane_q; reg ld_q; reg [31:0] ld_val_q; + integer i, l; + always @(posedge clk) begin + if (rst) begin dst_q <= 0; src_q <= 0; sel_q <= 0; ld_q <= 0; ld_lane_q <= 0; ld_val_q <= 0; end + else begin dst_q <= dst; src_q <= src; sel_q <= sel; ld_q <= ld_en; ld_lane_q <= ld_lane; ld_val_q <= ld_val; end + end + wire [31:0] s [0:31]; + wire [31:0] x [0:31]; + genvar g; + generate for (g = 0; g < 32; g = g + 1) begin : xb + assign s[g] = rf[g*8 + src_q]; + assign x[g] = s[sel_q[5*g +: 5]]; + end endgenerate + always @(posedge clk) begin + if (rst) begin for (i = 0; i < 256; i = i + 1) rf[i] <= 32'h9e3779b9 * (i + 1) + 6; end + else for (l = 0; l < 32; l = l + 1) + rf[l*8 + dst_q] <= (ld_q && ld_lane_q == l) ? ld_val_q : (rf[l*8 + dst_q] ^ x[l]); + end + assign out = x[0] ^ x[17]; +endmodule diff --git a/tools/chip-model/rtl/tb/tb_lane_arx.v b/tools/chip-model/rtl/tb/tb_lane_arx.v new file mode 100644 index 000000000..b4b618ab7 --- /dev/null +++ b/tools/chip-model/rtl/tb/tb_lane_arx.v @@ -0,0 +1,20 @@ +`timescale 1ps/1ps +module tb; + reg clk = 0, rst = 1; reg [2:0] op = 0, dst = 0, src = 0; reg [4:0] rot = 0; reg ld_en = 0; reg [31:0] ld_val = 0; + wire [31:0] out; + lane_arx dut(.clk(clk), .rst(rst), .op(op), .dst(dst), .src(src), .rot(rot), .ld_en(ld_en), .ld_val(ld_val), .out(out)); + integer n, fixed_op, cycles; reg [31:0] acc = 0; + always #500 clk = ~clk; + initial begin + if (!$value$plusargs("op=%d", fixed_op)) fixed_op = -1; + if (!$value$plusargs("cycles=%d", cycles)) cycles = 4000; + $dumpfile("dump.vcd"); $dumpvars(0, tb.dut); + repeat (4) @(negedge clk); rst = 0; + for (n = 0; n < cycles; n = n + 1) begin + @(negedge clk); + op = (fixed_op < 0) ? $random : fixed_op; dst = $random; src = $random; rot = $random; + ld_en = (($random & 15) == 0); ld_val = $random; acc = acc ^ out; + end + $display("CHECKSUM %08x", acc); $finish; + end +endmodule diff --git a/tools/chip-model/rtl/tb/tb_lane_fold.v b/tools/chip-model/rtl/tb/tb_lane_fold.v new file mode 100644 index 000000000..644de5980 --- /dev/null +++ b/tools/chip-model/rtl/tb/tb_lane_fold.v @@ -0,0 +1,21 @@ +`timescale 1ps/1ps +module tb; + reg clk = 0, rst = 1; reg [2:0] dst = 0, src = 0; reg cfg_en = 0; reg [31:0] cfg_m = 0, cfg_wm = 0, cfg_off = 0, cfg_mask = 0; reg [4:0] cfg_r = 0; + reg ld_en = 0; reg [31:0] ld_val = 0; wire [31:0] out; + lane_fold dut(.clk(clk), .rst(rst), .dst(dst), .src(src), .cfg_en(cfg_en), .cfg_m(cfg_m), .cfg_r(cfg_r), .cfg_wm(cfg_wm), .cfg_off(cfg_off), .cfg_mask(cfg_mask), .ld_en(ld_en), .ld_val(ld_val), .out(out)); + integer n, cycles; reg [31:0] acc = 0; + always #750 clk = ~clk; + initial begin + if (!$value$plusargs("cycles=%d", cycles)) cycles = 4000; + $dumpfile("dump.vcd"); $dumpvars(0, tb.dut); + repeat (4) @(negedge clk); rst = 0; + // one era draw: the constants are written once and then static, as on the chip + @(negedge clk); cfg_en = 1; cfg_m = 32'h9e3779b1; cfg_r = 5'd13; cfg_wm = 32'h0fffffc0; cfg_off = 32'h3; cfg_mask = 32'h0fffffff; + @(negedge clk); cfg_en = 0; + for (n = 0; n < cycles; n = n + 1) begin + @(negedge clk); + dst = $random; src = $random; ld_en = (($random & 3) == 0); ld_val = $random; acc = acc ^ out; + end + $display("CHECKSUM %08x", acc); $finish; + end +endmodule diff --git a/tools/chip-model/rtl/tb/tb_lane_lop3.v b/tools/chip-model/rtl/tb/tb_lane_lop3.v new file mode 100644 index 000000000..3bccc1670 --- /dev/null +++ b/tools/chip-model/rtl/tb/tb_lane_lop3.v @@ -0,0 +1,18 @@ +`timescale 1ps/1ps +module tb; + reg clk = 0, rst = 1; reg [2:0] dst = 0, src = 0, src2 = 0; reg [7:0] lut = 0; reg ld_en = 0; reg [31:0] ld_val = 0; + wire [31:0] out; + lane_lop3 dut(.clk(clk), .rst(rst), .dst(dst), .src(src), .src2(src2), .lut(lut), .ld_en(ld_en), .ld_val(ld_val), .out(out)); + integer n, cycles; reg [31:0] acc = 0; + always #500 clk = ~clk; + initial begin + if (!$value$plusargs("cycles=%d", cycles)) cycles = 4000; + $dumpfile("dump.vcd"); $dumpvars(0, tb.dut); + repeat (4) @(negedge clk); rst = 0; + for (n = 0; n < cycles; n = n + 1) begin + @(negedge clk); + dst = $random; src = $random; src2 = $random; lut = $random; ld_en = (($random & 15) == 0); ld_val = $random; acc = acc ^ out; + end + $display("CHECKSUM %08x", acc); $finish; + end +endmodule diff --git a/tools/chip-model/rtl/tb/tb_lane_mul.v b/tools/chip-model/rtl/tb/tb_lane_mul.v new file mode 100644 index 000000000..f5f37f7a0 --- /dev/null +++ b/tools/chip-model/rtl/tb/tb_lane_mul.v @@ -0,0 +1,20 @@ +`timescale 1ps/1ps +module tb; + reg clk = 0, rst = 1; reg [1:0] op = 0; reg [2:0] dst = 0, src = 0, src2 = 0; reg ld_en = 0; reg [31:0] ld_val = 0; + wire [31:0] out; + lane_mul dut(.clk(clk), .rst(rst), .op(op), .dst(dst), .src(src), .src2(src2), .ld_en(ld_en), .ld_val(ld_val), .out(out)); + integer n, fixed_op, cycles; reg [31:0] acc = 0; + always #750 clk = ~clk; + initial begin + if (!$value$plusargs("op=%d", fixed_op)) fixed_op = -1; + if (!$value$plusargs("cycles=%d", cycles)) cycles = 4000; + $dumpfile("dump.vcd"); $dumpvars(0, tb.dut); + repeat (4) @(negedge clk); rst = 0; + for (n = 0; n < cycles; n = n + 1) begin + @(negedge clk); + op = (fixed_op < 0) ? $random : fixed_op; dst = $random; src = $random; src2 = $random; + ld_en = (($random & 15) == 0); ld_val = $random; acc = acc ^ out; + end + $display("CHECKSUM %08x", acc); $finish; + end +endmodule diff --git a/tools/chip-model/rtl/tb/tb_lane_prmt.v b/tools/chip-model/rtl/tb/tb_lane_prmt.v new file mode 100644 index 000000000..30764b26c --- /dev/null +++ b/tools/chip-model/rtl/tb/tb_lane_prmt.v @@ -0,0 +1,18 @@ +`timescale 1ps/1ps +module tb; + reg clk = 0, rst = 1; reg [2:0] dst = 0, src = 0; reg [15:0] sel = 0; reg ld_en = 0; reg [31:0] ld_val = 0; + wire [31:0] out; + lane_prmt dut(.clk(clk), .rst(rst), .dst(dst), .src(src), .sel(sel), .ld_en(ld_en), .ld_val(ld_val), .out(out)); + integer n, cycles; reg [31:0] acc = 0; + always #500 clk = ~clk; + initial begin + if (!$value$plusargs("cycles=%d", cycles)) cycles = 4000; + $dumpfile("dump.vcd"); $dumpvars(0, tb.dut); + repeat (4) @(negedge clk); rst = 0; + for (n = 0; n < cycles; n = n + 1) begin + @(negedge clk); + dst = $random; src = $random; sel = $random & 16'h7777; ld_en = (($random & 15) == 0); ld_val = $random; acc = acc ^ out; + end + $display("CHECKSUM %08x", acc); $finish; + end +endmodule diff --git a/tools/chip-model/rtl/tb/tb_scratch8k.v b/tools/chip-model/rtl/tb/tb_scratch8k.v new file mode 100644 index 000000000..ed118a378 --- /dev/null +++ b/tools/chip-model/rtl/tb/tb_scratch8k.v @@ -0,0 +1,17 @@ +`timescale 1ps/1ps +module tb; + reg clk = 0, rst = 1; reg [10:0] raddr = 0, waddr = 0; reg we = 0; reg [31:0] wdata = 0; wire [31:0] rdata; + scratch8k dut(.clk(clk), .rst(rst), .raddr(raddr), .we(we), .waddr(waddr), .wdata(wdata), .rdata(rdata)); + integer n, cycles; reg [31:0] acc = 0; + always #750 clk = ~clk; + initial begin + if (!$value$plusargs("cycles=%d", cycles)) cycles = 3000; + $dumpfile("dump.vcd"); $dumpvars(0, tb.dut); + repeat (4) @(negedge clk); rst = 0; + for (n = 0; n < cycles; n = n + 1) begin + @(negedge clk); + raddr = $random; waddr = $random; we = (($random & 7) == 0); wdata = $random; acc = acc ^ rdata; + end + $display("CHECKSUM %08x", acc); $finish; + end +endmodule diff --git a/tools/chip-model/rtl/tb/tb_shfl32.v b/tools/chip-model/rtl/tb/tb_shfl32.v new file mode 100644 index 000000000..2cf58a69f --- /dev/null +++ b/tools/chip-model/rtl/tb/tb_shfl32.v @@ -0,0 +1,18 @@ +`timescale 1ps/1ps +module tb; + reg clk = 0, rst = 1; reg [2:0] dst = 0, src = 0; reg [4:0] m = 0; reg ld_en = 0; reg [4:0] ld_lane = 0; reg [31:0] ld_val = 0; + wire [31:0] out; + shfl32 dut(.clk(clk), .rst(rst), .dst(dst), .src(src), .m(m), .ld_en(ld_en), .ld_lane(ld_lane), .ld_val(ld_val), .out(out)); + integer n, cycles; reg [31:0] acc = 0; + always #500 clk = ~clk; + initial begin + if (!$value$plusargs("cycles=%d", cycles)) cycles = 3000; + $dumpfile("dump.vcd"); $dumpvars(0, tb.dut); + repeat (4) @(negedge clk); rst = 0; + for (n = 0; n < cycles; n = n + 1) begin + @(negedge clk); + dst = $random; src = $random; m = $random; ld_en = 1; ld_lane = $random; ld_val = $random; acc = acc ^ out; + end + $display("CHECKSUM %08x", acc); $finish; + end +endmodule diff --git a/tools/chip-model/rtl/tb/tb_tile8.v b/tools/chip-model/rtl/tb/tb_tile8.v new file mode 100644 index 000000000..f9b66a2dc --- /dev/null +++ b/tools/chip-model/rtl/tb/tb_tile8.v @@ -0,0 +1,18 @@ +`timescale 1ps/1ps +module tb; + reg clk = 0, rst = 1, clr = 0; reg [511:0] a_in = 0, b_in = 0; reg [2:0] sel_r = 0, sel_c = 0; wire [31:0] out; + tile8 dut(.clk(clk), .rst(rst), .clr(clr), .a_in(a_in), .b_in(b_in), .sel_r(sel_r), .sel_c(sel_c), .out(out)); + integer n, j, cycles; reg [31:0] acc = 0; + always #1000 clk = ~clk; + initial begin + if (!$value$plusargs("cycles=%d", cycles)) cycles = 3000; + $dumpfile("dump.vcd"); $dumpvars(0, tb.dut); + repeat (4) @(negedge clk); rst = 0; + for (n = 0; n < cycles; n = n + 1) begin + @(negedge clk); + for (j = 0; j < 16; j = j + 1) begin a_in[j*32 +: 32] = $random; b_in[j*32 +: 32] = $random; end + clr = (($random & 63) == 0); sel_r = $random; sel_c = $random; acc = acc ^ out; + end + $display("CHECKSUM %08x", acc); $finish; + end +endmodule diff --git a/tools/chip-model/rtl/tb/tb_xbar32.v b/tools/chip-model/rtl/tb/tb_xbar32.v new file mode 100644 index 000000000..6440b2340 --- /dev/null +++ b/tools/chip-model/rtl/tb/tb_xbar32.v @@ -0,0 +1,18 @@ +`timescale 1ps/1ps +module tb; + reg clk = 0, rst = 1; reg [2:0] dst = 0, src = 0; reg [159:0] sel = 0; reg ld_en = 0; reg [4:0] ld_lane = 0; reg [31:0] ld_val = 0; + wire [31:0] out; + xbar32 dut(.clk(clk), .rst(rst), .dst(dst), .src(src), .sel(sel), .ld_en(ld_en), .ld_lane(ld_lane), .ld_val(ld_val), .out(out)); + integer n, cycles; reg [31:0] acc = 0; + always #500 clk = ~clk; + initial begin + if (!$value$plusargs("cycles=%d", cycles)) cycles = 3000; + $dumpfile("dump.vcd"); $dumpvars(0, tb.dut); + repeat (4) @(negedge clk); rst = 0; + for (n = 0; n < cycles; n = n + 1) begin + @(negedge clk); + dst = $random; src = $random; sel = {$random, $random, $random, $random, $random}; ld_en = 1; ld_lane = $random; ld_val = $random; acc = acc ^ out; + end + $display("CHECKSUM %08x", acc); $finish; + end +endmodule From 7784bdc7d17dc062cf99903c25817936689e0de2 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Thu, 8 Oct 2026 12:05:02 +0000 Subject: [PATCH 2/9] chip-model: shadow-k collector, mix optimiser, edge rows; the report's method sections Co-Authored-By: Claude Fable 5.1 Documents-only replay of 6fab785c5 (bfccc26ed4d299ad35920bda4aa2af0cc5efbe01) for the box mirror master --- docs/analysis/class-v6/floor/shadow-k.md | 77 ++++++++++++++++++++++++ tools/chip-model/rtl/flow/collect.py | 5 +- tools/chip-model/rtl/flow/edge.py | 23 +++++++ tools/chip-model/rtl/flow/mix.py | 65 ++++++++++++++++++++ tools/chip-model/rtl/flow/scratch.mk | 1 + tools/chip-model/rtl/flow/shfl.mk | 1 + tools/chip-model/rtl/flow/xbar.mk | 1 + 7 files changed, 171 insertions(+), 2 deletions(-) create mode 100644 docs/analysis/class-v6/floor/shadow-k.md create mode 100644 tools/chip-model/rtl/flow/edge.py create mode 100644 tools/chip-model/rtl/flow/mix.py diff --git a/docs/analysis/class-v6/floor/shadow-k.md b/docs/analysis/class-v6/floor/shadow-k.md new file mode 100644 index 000000000..c722a26d3 --- /dev/null +++ b/docs/analysis/class-v6/floor/shadow-k.md @@ -0,0 +1,77 @@ +# The shadow's k from RTL (floor lane 2, 8 October 2026) + +Branch `class-v6-floor-k` from `counter-asic-4` at 7618e729. Every chip-side number in this file comes from +synthesis and place-and-route of RTL written for this lane (Yosys 0.68 plus OpenROAD, the ORFS docker image +`openroad/orfs:latest`, the ASAP7 predictive PDK, run on igneum-build-4), with switching activity from a +random-input gate-level simulation (iverilog on the routed netlist). Every GPU-side number is the record's +measurement (`docs/analysis/counter-asic-4-research.md` 15.1a: PC 1, the RTX 5090, 20 probes at 60 s, 8 +October 2026) and is not re-estimated here. The RTL, testbenches, flow configs and a Makefile that reproduces +every row are under `tools/chip-model/rtl/`. + +## 0. One page + +(filled from the rows below; see section 2 for the table and section 4 for the edge) + +## 1. The question and the identity + +The chip's energy per hash is `E_chip = E_mem + k x F` (research file section 2): `E_mem` the memory system's +reads, controller and static (0.466 microjoules per hash on the record's GDDR7 board, modelled), `F` the shadow's +premium on the card (measured: 1.10 microjoules unlocked, 0.652 at the 1,300 MHz lock on the 5090 for class +v4's 102,100 counted ops per hash), and `k = e_chip / e_gpu` the chip core's energy per forced op over the +card's. The card's side of `k` is measured; the chip's side has been a claimed band (ALU 0.3 to 0.8, int8 tile +0.03 to 0.3, L2 hit 0.1 to 0.3, shuffle 0.4 to 0.7). This file replaces the claimed side with a synthesised one, +family by family, and reads the mix that maximises the expected `k` under the invention lane's sound-form rules. + +## 2. Method + +### 2.1 What was built (RTL, `tools/chip-model/rtl/rtl/`) + +Each family is a minimal chip-side shadow core: the smallest circuit a chip maker would have to build to run that +family's instructions bit-exactly at one op per cycle per lane. The common lane shape is an instruction register, +an 8 x 32-bit register window in flops (the card's lane holds its working set in a register file too; 8 registers +is the class program's window), two or three read ports through muxes, the functional unit, and one write port. +Nothing is shared across lanes and nothing is pipelined beyond one stage, so the figure is the datapath plus the +minimum operand delivery: a floor for the chip, which makes the `k` it gives a floor too. + +| Family | Module | What one op is | Ops per cycle | Clock set (ps) | +|---|---|---|---|---| +| ARX | `lane_arx` | int32 add, sub, xor, or, rotl by immediate, rotr by register (op drawn per cycle; rows per op by fixing the op field) | 1 | 1,000 | +| MUL | `lane_mul` | 32 x 32 to 64: mul (low word), mulhi (high word), mad (src x src2 + dst) | 1 | 1,500 | +| PRMT | `lane_prmt` | byte permute: 4 output bytes from the 8 bytes of two registers by a 16-bit selector, sign-replicate bit as on the card | 1 | 1,000 | +| LOP3 | `lane_lop3` | three-input logic by an 8-bit truth table, bitwise | 1 | 1,000 | +| FOLD | `lane_fold` | the index fold `((rotl(x * M, R) & WM) \| OFF) & MASK` with the era's constants held in registers | 1 | 1,500 | +| SHFL | `shfl32` | the 32-lane xor-mask shuffle `r[dst][lane] ^= r[src][lane ^ m]` over a 1 KB window (32 lanes x 8 x 32 bits): a 5-stage butterfly | 32 lane-ops | 1,000 | +| XBAR | `xbar32` | the general 32-lane crossbar over the same window (any source lane per lane, a 5-bit select each): the upper bound on a shuffle network | 32 lane-ops | 1,000 | +| SCRATCH | `scratch8k` | one random 32-bit read of a 2,048 x 32 (8 KB) flop array, a write on one cycle in eight: the chip's L1 as the pessimistic flop form | 1 read | 1,500 | +| TILE | `tile8` | the int8 8x8x8 tile `C += A x B` with int32 accumulators: 512 MACs per cycle | 512 MACs | 2,000 | + +### 2.2 The flow + +ORFS on ASAP7 (7.5-track RVT cells, the TC corner: 0.70 V, 0 C, NLDM), the default flow end to end: Yosys +synthesis with ABC, floorplan at 40 percent utilisation, global and detailed placement, CTS, global and detailed +routing, parasitic extraction (OpenRCX, the platform's rules). Power is OpenSTA's `report_power` on the routed +design with its SPEF, under two activities: (a) the VCD of a random-input simulation of the routed netlist +(ASAP7 has no cell Verilog models in the image, so Yosys builds the cell bodies from the liberty functions and +writes the netlist as primitives; iverilog runs 3,000 to 4,000 cycles with every instruction field drawn by +`$random` each cycle and a random 32-bit load into the window every 16th cycle), and (b) a propagated 0.5 +activity on every input as the cross-check. The energy per op is the total power (internal, switching and leakage) +times the clock period over the ops per cycle. Leakage is reported beside it; at these clocks it is a few percent. + +### 2.3 Node scaling (approximate; every factor claimed from the foundry's own headline) + +ASAP7 is a predictive 7 nm-class FinFET PDK (ASU and ARM, Clark et al., Microelectronics Journal 2016), not a +foundry node, so the row is first stated at ASAP7 and then scaled by the foundry's published per-node power +reductions at the same speed: N7 to N5 x0.70 (TSMC: "30 percent lower power"), N5 to N3E x0.72 (TSMC: "25 to 30 +percent lower power", the midpoint), N3E to N2 x0.72 (TSMC: "25 to 30 percent lower power", the midpoint). So +N5 = 0.70, N3 = 0.50 and N2 = 0.36 of the ASAP7 figure. These are the foundry's claims for a whole design at a +fixed frequency, and a shadow core at a low clock could run at a lower voltage still; the N2 column is therefore +the chip's best case from this method, not its floor. Sources in section 7. + +### 2.4 What the method leaves out, on both sides + +On the chip side the figure omits instruction fetch and decode (a chip would run the program from a small SRAM +or a decoded instruction cache shared by many lanes, about 1 to 2 pJ per lane-instruction amortised across 32 +lanes, approximate), the clock tree beyond the block's own, and the result's move to a memory address unit. On +the card side the 15.1a figure is the whole card's marginal per counted op (the sleep floor subtracted), which +includes the card's own fetch, decode, operand collection and register file. So the `k` here is the chip's +datapath-and-window cost over the card's whole-lane cost: a FLOOR on the chip's cost and so a floor on `k`. diff --git a/tools/chip-model/rtl/flow/collect.py b/tools/chip-model/rtl/flow/collect.py index 6ed151f21..a8232b337 100644 --- a/tools/chip-model/rtl/flow/collect.py +++ b/tools/chip-model/rtl/flow/collect.py @@ -34,10 +34,11 @@ def parse_log(path): tag = head.strip().replace('POWER_VCD ', 'vcd:').replace('POWER_PROPAGATED_0.5', 'prop') # OpenSTA report_power: "Total 100.0%" in watts tm = re.search(r'^Total\s+([\d.eE+-]+)\s+([\d.eE+-]+)\s+([\d.eE+-]+)\s+([\d.eE+-]+)', body, re.M) - ann = re.search(r'(\d+)\s*\(\s*([\d.]+)%\)\s*annotated', body) + ann = re.search(r'Annotated (\d+) pin activities', body) + unann = re.search(r'unannotated\s+(\d+)', body) if tm: rows[tag] = dict(internal=float(tm.group(1)), switching=float(tm.group(2)), leakage=float(tm.group(3)), - total=float(tm.group(4)), annotated=ann.group(2) if ann else '') + total=float(tm.group(4)), annotated=(f"{ann.group(1)} pins, {unann.group(1) if unann else '?'} unannotated") if ann else '') return period, cells, rows out = [] diff --git a/tools/chip-model/rtl/flow/edge.py b/tools/chip-model/rtl/flow/edge.py new file mode 100644 index 000000000..cf4542d07 --- /dev/null +++ b/tools/chip-model/rtl/flow/edge.py @@ -0,0 +1,23 @@ +#!/usr/bin/env python3 +"""The chip edge rows recomputed at the measured k. Usage: edge.py """ +import sys +k_v4 = float(sys.argv[1]); k_best = float(sys.argv[2]) +# the record's rows: card microjoules per hash and the premium F (counter-asic-4-research.md 20.4; the GDDR7 +# board E_mem 0.466; hardware-future.md section 5 for the strongest DRAM chips) +cards = [('5090 unlocked (class v4 shape)', 3.36, 1.10), ('5090 at the 1,300 MHz lock', 2.33, 0.652)] +chips = [('GDDR7 board, 28 nm controller (the record)', 0.466), ('HBM3E one stack', 0.321), ('HBM4 one stack, N12 base die', 0.22), + ('custom HBM4E base die, N3P', 0.18), ('DRAM on logic, hybrid bonded (2029 to 2031)', 0.15), ('SRAM full store, one N2 reticle', 0.14)] +def edge(card, F, mem, k): return (card) / (mem + k * F) +print('| Chip | 5090 row | Card uJ | F uJ | Edge at k = 1 | at k = 0.5 | at k = 0.3 | at the measured k (class v4 mix) | at the measured k (best mix) | at zero shadow |') +print('|---|---|---|---|---|---|---|---|---|---|') +for cn, mem in chips: + for rn, card, F in cards: + print(f'| {cn} | {rn} | {card:.2f} | {F:.3f} | {edge(card,F,mem,1):.2f}x | {edge(card,F,mem,0.5):.2f}x | {edge(card,F,mem,0.3):.2f}x | {edge(card,F,mem,k_v4):.2f}x (k {k_v4:.3f}) | {edge(card,F,mem,k_best):.2f}x (k {k_best:.3f}) | {(card-F)/mem:.2f}x |') +# the honest-card variants: floor lane 1's assumed 1.0 microjoules per hash at zero shadow on the 5090 +print('\nFloor lane 1 assumption: the honest 5090 at 1.0 microjoules per hash at zero shadow (not in by the default time).') +print('| Chip | Premium kept at the lock\'s 0.652 | Premium scaled with the card (0.652 x 1.0 / 1.67 = 0.390) |') +print('|---|---|---|') +for cn, mem in chips: + a = edge(1.0 + 0.652, 0.652, mem, k_v4); b = edge(1.0 + 0.390, 0.390, mem, k_v4) + a1 = edge(1.0 + 0.652, 0.652, mem, 1.0); b1 = edge(1.0 + 0.390, 0.390, mem, 1.0) + print(f'| {cn} | {a:.2f}x at the measured k ({a1:.2f}x at k = 1; {1.0/mem:.2f}x at zero shadow) | {b:.2f}x at the measured k ({b1:.2f}x at k = 1) |') diff --git a/tools/chip-model/rtl/flow/mix.py b/tools/chip-model/rtl/flow/mix.py new file mode 100644 index 000000000..835f332cc --- /dev/null +++ b/tools/chip-model/rtl/flow/mix.py @@ -0,0 +1,65 @@ +#!/usr/bin/env python3 +"""The op mix that maximises the chip's effective k at a FIXED GPU premium, inside the class v6 layer-1 band. + +k_eff(mix) = sum_i w_i e_chip_i / sum_i w_i e_gpu_i : the chip's energy for the drawn program over the card's +for the same program. At a fixed premium F on the card the chip pays k_eff x F, so the mix that maximises k_eff +is the one that forces the most. The band (docs/design/class-v6-rotating-family.md section 2, layer 1): +B = 4 points on the injecting families only (add, sub, xor, mad, shfl, rotl, rotr), each within base - B .. base + B; +the lossy families (or, mul, mulhi) capped at their base (base - B .. base); or + mul + mulhi at most 18 + B; +the shuffle weight capped at its class v4 value (8). The index fold and the F8 floor are not functions of the +weights and do not move. Usage: mix.py [node=N3] [state=lock|unlocked]""" +import csv, sys, itertools + +path = sys.argv[1]; node = 'N3'; state = 'lock' +for a in sys.argv[2:]: + k, v = a.split('='); + if k == 'node': node = v + if k == 'state': state = v +rows = {(r['family'], r['sim']): r for r in csv.DictReader(open(path))} +def chip(fam, sim='vcd:mix'): + r = rows.get((fam, sim)) or rows.get((fam, 'vcd:mix')) + return float(r[f'pJ_{node}']) if r else None +# the 5090 measured pJ per counted op (15.1a), per family of the generator's draw +gpu = {'add': (11.3, 6.2), 'sub': (11.3, 6.2), 'xor': (11.3, 6.2), 'or': (11.3, 6.2), 'rotl': (11.3, 6.2), 'rotr': (11.3, 6.2), + 'mul': (13.9, 8.3), 'mad': (13.9, 8.3), 'mulhi': (39.6, 21.0), 'shfl': (55.8, 29.4)} +si = 0 if state == 'unlocked' else 1 +e_gpu = {f: v[si] for f, v in gpu.items()} +e_chip = {'add': chip('arx', 'vcd:add'), 'sub': chip('arx', 'vcd:sub'), 'xor': chip('arx', 'vcd:xor'), 'or': chip('arx', 'vcd:or'), + 'rotl': chip('arx', 'vcd:rotl'), 'rotr': chip('arx', 'vcd:rotr'), + 'mul': chip('mul', 'vcd:mul'), 'mad': chip('mul', 'vcd:mad'), 'mulhi': chip('mul', 'vcd:mulhi'), + 'shfl': chip('shfl')} +base = {'add': 12, 'xor': 10, 'mul': 8, 'mad': 8, 'shfl': 8, 'rotl': 7, 'sub': 6, 'mulhi': 6, 'rotr': 6, 'or': 4} +B = 4 +inject = ['add', 'sub', 'xor', 'mad', 'shfl', 'rotl', 'rotr']; lossy = ['or', 'mul', 'mulhi'] +missing = [f for f, v in e_chip.items() if v is None] +if missing: + print('missing chip rows:', missing); sys.exit(1) +k_fam = {f: e_chip[f] / e_gpu[f] for f in base} +def keff(w): return sum(w[f] * e_chip[f] for f in w) / sum(w[f] * e_gpu[f] for f in w) +print(f'node {node}, 5090 state {state}') +print('| Family | Base weight | 5090 pJ/op | Chip pJ/op | k |') +print('|---|---|---|---|---|') +for f in sorted(base, key=lambda f: -k_fam[f]): + print(f'| {f} | {base[f]} | {e_gpu[f]} | {e_chip[f]:.3g} | {k_fam[f]:.3f} |') +print(f'\nclass v4 mix: k_eff = {keff(base):.4f}') +# exhaustive over the band: each family at one of the allowed values (steps of 1 point) +ranges = {} +for f in base: + lo = max(0, base[f] - B) + hi = base[f] + B if f in inject else base[f] + if f == 'shfl': hi = min(hi, 8) + ranges[f] = list(range(lo, hi + 1)) +best = None; worst = None +fams = list(base) +# reduce the search: injecting families only take their extremes and base (the objective is a ratio of linear forms, +# monotone in each weight), lossy families all values +cand = {f: ([ranges[f][0], base[f], ranges[f][-1]] if f in inject else ranges[f]) for f in fams} +for combo in itertools.product(*[cand[f] for f in fams]): + w = dict(zip(fams, combo)) + if w['or'] + w['mul'] + w['mulhi'] > 18 + B: continue + if sum(w.values()) == 0: continue + v = keff(w) + if best is None or v > best[0]: best = (v, dict(w)) + if worst is None or v < worst[0]: worst = (v, dict(w)) +for name, (v, w) in (('best', best), ('worst', worst)): + print(f'{name} mix in the band: k_eff = {v:.4f}: ' + ', '.join(f'{f} {w[f]}' for f in fams) + f' (sum {sum(w.values())})') diff --git a/tools/chip-model/rtl/flow/scratch.mk b/tools/chip-model/rtl/flow/scratch.mk index a8fdb0197..06d205dcd 100644 --- a/tools/chip-model/rtl/flow/scratch.mk +++ b/tools/chip-model/rtl/flow/scratch.mk @@ -13,3 +13,4 @@ export CORNER = TC export SKIP_LAST_GASP = 1 export WORK_HOME = /work/out/scratch +export SYNTH_MEMORY_MAX_BITS = 131072 diff --git a/tools/chip-model/rtl/flow/shfl.mk b/tools/chip-model/rtl/flow/shfl.mk index a73d97541..20b18916c 100644 --- a/tools/chip-model/rtl/flow/shfl.mk +++ b/tools/chip-model/rtl/flow/shfl.mk @@ -13,3 +13,4 @@ export CORNER = TC export SKIP_LAST_GASP = 1 export WORK_HOME = /work/out/shfl +export SYNTH_MEMORY_MAX_BITS = 131072 diff --git a/tools/chip-model/rtl/flow/xbar.mk b/tools/chip-model/rtl/flow/xbar.mk index 01026b194..612f4f154 100644 --- a/tools/chip-model/rtl/flow/xbar.mk +++ b/tools/chip-model/rtl/flow/xbar.mk @@ -13,3 +13,4 @@ export CORNER = TC export SKIP_LAST_GASP = 1 export WORK_HOME = /work/out/xbar +export SYNTH_MEMORY_MAX_BITS = 131072 From c31ff8a73e4d3f77c6b3c4ecce521735dcbffe07 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Thu, 8 Oct 2026 12:18:54 +0000 Subject: [PATCH 3/9] chip-model: the programmable shadow core (core_v6: imem, fetch, decode, per-lane register file, the class units, era registers), 8, 32 and 32-lane 16-register builds, synthesis-only power path Co-Authored-By: Claude Fable 5.1 Documents-only replay of 49c49c340 (bfccc26ed4d299ad35920bda4aa2af0cc5efbe01) for the box mirror master --- tools/chip-model/rtl/Makefile | 13 ++- tools/chip-model/rtl/flow/collect.py | 5 +- tools/chip-model/rtl/flow/core32.mk | 15 +++ tools/chip-model/rtl/flow/core32.sdc | 10 ++ tools/chip-model/rtl/flow/core32r16.mk | 15 +++ tools/chip-model/rtl/flow/core32r16.sdc | 10 ++ tools/chip-model/rtl/flow/core8.mk | 15 +++ tools/chip-model/rtl/flow/core8.sdc | 10 ++ tools/chip-model/rtl/flow/designs.txt | 3 + tools/chip-model/rtl/flow/gl2sim.sh | 2 +- tools/chip-model/rtl/flow/power.tcl | 7 +- tools/chip-model/rtl/flow/sim.sh | 2 +- tools/chip-model/rtl/rtl/core_v6.v | 113 +++++++++++++++++++++ tools/chip-model/rtl/rtl/core_v6_32.v | 7 ++ tools/chip-model/rtl/rtl/core_v6_32r16.v | 7 ++ tools/chip-model/rtl/rtl/core_v6_8.v | 7 ++ tools/chip-model/rtl/tb/tb_core_common.vh | 34 +++++++ tools/chip-model/rtl/tb/tb_core_v6_32.v | 3 + tools/chip-model/rtl/tb/tb_core_v6_32r16.v | 3 + tools/chip-model/rtl/tb/tb_core_v6_8.v | 3 + 20 files changed, 277 insertions(+), 7 deletions(-) create mode 100644 tools/chip-model/rtl/flow/core32.mk create mode 100644 tools/chip-model/rtl/flow/core32.sdc create mode 100644 tools/chip-model/rtl/flow/core32r16.mk create mode 100644 tools/chip-model/rtl/flow/core32r16.sdc create mode 100644 tools/chip-model/rtl/flow/core8.mk create mode 100644 tools/chip-model/rtl/flow/core8.sdc create mode 100644 tools/chip-model/rtl/rtl/core_v6.v create mode 100644 tools/chip-model/rtl/rtl/core_v6_32.v create mode 100644 tools/chip-model/rtl/rtl/core_v6_32r16.v create mode 100644 tools/chip-model/rtl/rtl/core_v6_8.v create mode 100644 tools/chip-model/rtl/tb/tb_core_common.vh create mode 100644 tools/chip-model/rtl/tb/tb_core_v6_32.v create mode 100644 tools/chip-model/rtl/tb/tb_core_v6_32r16.v create mode 100644 tools/chip-model/rtl/tb/tb_core_v6_8.v diff --git a/tools/chip-model/rtl/Makefile b/tools/chip-model/rtl/Makefile index 05b786486..e3bc96c38 100644 --- a/tools/chip-model/rtl/Makefile +++ b/tools/chip-model/rtl/Makefile @@ -17,7 +17,7 @@ DOCKER := docker run --rm -u $(UID_GID) -e HOME=/tmp -v $(WORK):/work ORFS := $(DOCKER) -w /OpenROAD-flow-scripts/flow $(ORFS_IMG) SIM := $(DOCKER) -w /work $(SIM_IMG) -DESIGNS := arx mul prmt lop3 fold shfl xbar scratch tile +DESIGNS := arx mul prmt lop3 fold shfl xbar scratch tile core8 core32 core32r16 top = $(shell sed -n 's/^$(1) \([^ ]*\) .*/\1/p' flow/designs.txt) # per-family simulation tags (the op field fixed per row where the family has several ops) @@ -30,6 +30,9 @@ SIMS_shfl := mix SIMS_xbar := mix SIMS_scratch := mix SIMS_tile := mix +SIMS_core8 := mix mixld:+loads=1 +SIMS_core32 := mix mixld:+loads=1 +SIMS_core32r16 := mix mixld:+loads=1 .PHONY: rows table clean @@ -52,6 +55,14 @@ sim-%: flow-% power-%: sim-% $(ORFS) make DESIGN_CONFIG=/work/flow/$*.mk RUN_SCRIPT=/work/flow/power.tcl RUN_LOG_NAME_STEM=power run 2>&1 | tee logs/power-$*.log +# synthesis-only row (no placement, no parasitics): the 14:45 fallback +synth-%: + mkdir -p out/$* logs sim/$*-synth + $(ORFS) make DESIGN_CONFIG=/work/flow/$*.mk synth 2>&1 | tee logs/synth-$*.log + NET=1_2_yosys.v $(SIM) bash -c 'NET=1_2_yosys.v bash /work/flow/gl2sim.sh $* $(call top,$*)' 2>&1 | tee logs/gl2sim-$*-synth.log + $(SIM) bash /work/flow/sim.sh $* synth 2>&1 | tee logs/sim-$*-synth.log + $(ORFS) make DESIGN_CONFIG=/work/flow/$*.mk RUN_SCRIPT=/work/flow/power.tcl RUN_LOG_NAME_STEM=power_synth FLOORK_ODB=1_synth.odb FLOORK_SDC=1_synth.sdc run 2>&1 | tee logs/power-$*-synth.log + table: python3 flow/collect.py $(WORK) > table.md @echo wrote table.md and table.csv diff --git a/tools/chip-model/rtl/flow/collect.py b/tools/chip-model/rtl/flow/collect.py index a8232b337..4c9bbb4e6 100644 --- a/tools/chip-model/rtl/flow/collect.py +++ b/tools/chip-model/rtl/flow/collect.py @@ -6,7 +6,7 @@ import re, sys, os, csv work = sys.argv[1] if len(sys.argv) > 1 else '.' # ops per cycle per design (the per-op divisor) and the GPU row each family is read against -OPS = {'arx': 1, 'mul': 1, 'prmt': 1, 'lop3': 1, 'fold': 1, 'shfl': 32, 'xbar': 32, 'scratch': 1, 'tile': 512} +OPS = {'arx': 1, 'mul': 1, 'prmt': 1, 'lop3': 1, 'fold': 1, 'shfl': 32, 'xbar': 32, 'scratch': 1, 'tile': 512, 'core8': 8, 'core32': 32} # 5090 measured pJ per counted op: (unlocked, at the 1,300 MHz lock); 15.1a GPU = { 'arx:mix': (11.3, 6.2), 'arx:add': (11.3, 6.2), 'arx:sub': (11.3, 6.2), 'arx:xor': (11.3, 6.2), 'arx:or': (11.3, 6.2), @@ -16,7 +16,8 @@ GPU = { 'fold:mix': (13.9, 8.3), # the fold is a multiply plus a rotate and masks: read against int_mul 'shfl:mix': (55.8, 29.4), 'xbar:mix': (55.8, 29.4), 'scratch:mix': (2400.0, 1400.0), # the card's L2 hit (no shared-memory probe measured: owed) - 'tile:mix': (4.1, 2.2), # dependent u8 m8n8k16 per MAC; the wide s8 tile reads 1.36 / 0.83 + 'tile:mix': (4.1, 2.2), + 'core8:mix': (11.3, 6.2), 'core8:mixld': (11.3, 6.2), 'core32:mix': (11.3, 6.2), 'core32:mixld': (11.3, 6.2), # the class v4 draw: read against int_arx (the packs job read the whole mix at 10.8 / 6.4) # dependent u8 m8n8k16 per MAC; the wide s8 tile reads 1.36 / 0.83 } # per-node energy scaling from ASAP7 (a 7 nm-class predictive PDK at 0.70 V), approximate and claimed: # N7 -> N5 x0.70 (TSMC: "30 percent lower power at the same speed"), N5 -> N3E x0.72 (TSMC: 25 to 30 percent), diff --git a/tools/chip-model/rtl/flow/core32.mk b/tools/chip-model/rtl/flow/core32.mk new file mode 100644 index 000000000..9ec9f6dfa --- /dev/null +++ b/tools/chip-model/rtl/flow/core32.mk @@ -0,0 +1,15 @@ +# ORFS design config for the programmable shadow core (core32: core_v6_32), ASAP7. +export PLATFORM = asap7 +export DESIGN_NAME = core_v6_32 +export DESIGN_NICKNAME = core32 +export VERILOG_FILES = /work/rtl/core_v6_32.v +export VERILOG_INCLUDE_DIRS = /work/rtl +export SDC_FILE = /work/flow/core32.sdc +export CORE_UTILIZATION = 40 +export CORE_ASPECT_RATIO = 1 +export CORE_MARGIN = 0.5 +export PLACE_DENSITY = 0.55 +export CORNER = TC +export SKIP_LAST_GASP = 1 +export WORK_HOME = /work/out/core32 +export SYNTH_MEMORY_MAX_BITS = 2000000 diff --git a/tools/chip-model/rtl/flow/core32.sdc b/tools/chip-model/rtl/flow/core32.sdc new file mode 100644 index 000000000..118104793 --- /dev/null +++ b/tools/chip-model/rtl/flow/core32.sdc @@ -0,0 +1,10 @@ +current_design core_v6_32 +set clk_name core_clock +set clk_port_name clk +set clk_period 1500 +set clk_io_pct 0.2 +set clk_port [get_ports $clk_port_name] +create_clock -name $clk_name -period $clk_period $clk_port +set non_clock_inputs [all_inputs -no_clocks] +set_input_delay [expr $clk_period * $clk_io_pct] -clock $clk_name $non_clock_inputs +set_output_delay [expr $clk_period * $clk_io_pct] -clock $clk_name [all_outputs] diff --git a/tools/chip-model/rtl/flow/core32r16.mk b/tools/chip-model/rtl/flow/core32r16.mk new file mode 100644 index 000000000..a1b779ece --- /dev/null +++ b/tools/chip-model/rtl/flow/core32r16.mk @@ -0,0 +1,15 @@ +# ORFS design config for the programmable shadow core (core32r16: core_v6_32r16), ASAP7. +export PLATFORM = asap7 +export DESIGN_NAME = core_v6_32r16 +export DESIGN_NICKNAME = core32r16 +export VERILOG_FILES = /work/rtl/core_v6_32r16.v +export VERILOG_INCLUDE_DIRS = /work/rtl +export SDC_FILE = /work/flow/core32r16.sdc +export CORE_UTILIZATION = 40 +export CORE_ASPECT_RATIO = 1 +export CORE_MARGIN = 0.5 +export PLACE_DENSITY = 0.55 +export CORNER = TC +export SKIP_LAST_GASP = 1 +export WORK_HOME = /work/out/core32r16 +export SYNTH_MEMORY_MAX_BITS = 2000000 diff --git a/tools/chip-model/rtl/flow/core32r16.sdc b/tools/chip-model/rtl/flow/core32r16.sdc new file mode 100644 index 000000000..d4aaf339a --- /dev/null +++ b/tools/chip-model/rtl/flow/core32r16.sdc @@ -0,0 +1,10 @@ +current_design core_v6_32r16 +set clk_name core_clock +set clk_port_name clk +set clk_period 1500 +set clk_io_pct 0.2 +set clk_port [get_ports $clk_port_name] +create_clock -name $clk_name -period $clk_period $clk_port +set non_clock_inputs [all_inputs -no_clocks] +set_input_delay [expr $clk_period * $clk_io_pct] -clock $clk_name $non_clock_inputs +set_output_delay [expr $clk_period * $clk_io_pct] -clock $clk_name [all_outputs] diff --git a/tools/chip-model/rtl/flow/core8.mk b/tools/chip-model/rtl/flow/core8.mk new file mode 100644 index 000000000..306449501 --- /dev/null +++ b/tools/chip-model/rtl/flow/core8.mk @@ -0,0 +1,15 @@ +# ORFS design config for the programmable shadow core (core8: core_v6_8), ASAP7. +export PLATFORM = asap7 +export DESIGN_NAME = core_v6_8 +export DESIGN_NICKNAME = core8 +export VERILOG_FILES = /work/rtl/core_v6_8.v +export VERILOG_INCLUDE_DIRS = /work/rtl +export SDC_FILE = /work/flow/core8.sdc +export CORE_UTILIZATION = 40 +export CORE_ASPECT_RATIO = 1 +export CORE_MARGIN = 0.5 +export PLACE_DENSITY = 0.55 +export CORNER = TC +export SKIP_LAST_GASP = 1 +export WORK_HOME = /work/out/core8 +export SYNTH_MEMORY_MAX_BITS = 2000000 diff --git a/tools/chip-model/rtl/flow/core8.sdc b/tools/chip-model/rtl/flow/core8.sdc new file mode 100644 index 000000000..10f3be2bd --- /dev/null +++ b/tools/chip-model/rtl/flow/core8.sdc @@ -0,0 +1,10 @@ +current_design core_v6_8 +set clk_name core_clock +set clk_port_name clk +set clk_period 1500 +set clk_io_pct 0.2 +set clk_port [get_ports $clk_port_name] +create_clock -name $clk_name -period $clk_period $clk_port +set non_clock_inputs [all_inputs -no_clocks] +set_input_delay [expr $clk_period * $clk_io_pct] -clock $clk_name $non_clock_inputs +set_output_delay [expr $clk_period * $clk_io_pct] -clock $clk_name [all_outputs] diff --git a/tools/chip-model/rtl/flow/designs.txt b/tools/chip-model/rtl/flow/designs.txt index 133bafae3..329ec661e 100644 --- a/tools/chip-model/rtl/flow/designs.txt +++ b/tools/chip-model/rtl/flow/designs.txt @@ -7,3 +7,6 @@ shfl shfl32 1000 xbar xbar32 1000 scratch scratch8k 1500 tile tile8 2000 +core8 core_v6_8 1500 +core32 core_v6_32 1500 +core32r16 core_v6_32r16 1500 diff --git a/tools/chip-model/rtl/flow/gl2sim.sh b/tools/chip-model/rtl/flow/gl2sim.sh index 17c94928f..a80c15115 100755 --- a/tools/chip-model/rtl/flow/gl2sim.sh +++ b/tools/chip-model/rtl/flow/gl2sim.sh @@ -2,7 +2,7 @@ # gl2sim.sh : netlist -> sim netlist with cell bodies from the liberty. set -euo pipefail name=$1; top=$2 -net=/work/out/$name/results/asap7/$name/base/6_final.v +net=/work/out/$name/results/asap7/$name/base/${NET:-6_final.v} lib=/OpenROAD-flow-scripts/flow/platforms/asap7/lib/NLDM out=/work/sim/$name; mkdir -p $out cat > $out/gl2sim.ys <.mk RUN_SCRIPT=/work/flow/power.tcl run source $::env(SCRIPTS_DIR)/load.tcl -load_design 6_final.odb 6_final.sdc +set odb [expr {[info exists ::env(FLOORK_ODB)] ? $::env(FLOORK_ODB) : "6_final.odb"}] +set sdc [expr {[info exists ::env(FLOORK_SDC)] ? $::env(FLOORK_SDC) : "6_final.sdc"}] +load_design $odb $sdc +puts "FLOORK stage $odb" set spef $::env(RESULTS_DIR)/6_final.spef -if { [file exists $spef] } { read_spef $spef } else { estimate_parasitics -global_routing } +if { $odb == "6_final.odb" && [file exists $spef] } { read_spef $spef } elseif { $odb == "6_final.odb" } { estimate_parasitics -global_routing } elseif { [string match "3_*" $odb] || [string match "4_*" $odb] } { estimate_parasitics -placement } else { puts "FLOORK no parasitics (synthesis only)" } puts "FLOORK clock_period_ps [expr [get_property [lindex [all_clocks] 0] period]]" puts "FLOORK cells [llength [get_cells *]]" report_tns diff --git a/tools/chip-model/rtl/flow/sim.sh b/tools/chip-model/rtl/flow/sim.sh index f7c7cbdd6..88b05666d 100755 --- a/tools/chip-model/rtl/flow/sim.sh +++ b/tools/chip-model/rtl/flow/sim.sh @@ -4,6 +4,6 @@ set -euo pipefail name=$1; tag=$2; shift 2 out=/work/sim/$name simcells=$(yosys-config --datdir)/simcells.v -iverilog -g2005 -o $out/sim_$tag $out/sim_net.v /work/tb/tb_$(sed -n "s/^$name \([^ ]*\) .*/\1/p" /work/flow/designs.txt).v $simcells +iverilog -g2005 -I /work/tb -o $out/sim_$tag $out/sim_net.v /work/tb/tb_$(sed -n "s/^$name \([^ ]*\) .*/\1/p" /work/flow/designs.txt).v $simcells ( cd $out && vvp -n sim_$tag "$@" | tee sim_$tag.log && mv dump.vcd $tag.vcd ) ls -la $out/$tag.vcd diff --git a/tools/chip-model/rtl/rtl/core_v6.v b/tools/chip-model/rtl/rtl/core_v6.v new file mode 100644 index 000000000..618df15bf --- /dev/null +++ b/tools/chip-model/rtl/rtl/core_v6.v @@ -0,0 +1,113 @@ +// The programmable shadow core: the minimal in-order SIMD core that executes a class v6 shadow program as +// drawn. Per core: an instruction memory sized to the drawn program length (256 x 32-bit, a flop array), a +// program counter that wraps at the era's drawn length, fetch into an instruction register, decode, the era's +// parameter registers (fold constants M, R, WM, OFF, MASK; program length N). Per lane: a 32 x 32-bit register +// file (flops) with two or three read ports (mad reads three) and one write port, and the class's units: add, sub, +// xor, or, rotl by immediate, rotr by register, mul, mulhi, mad, prmt, lop3, the xor-mask shuffle across the +// lanes (a log2(LANES)-stage butterfly), and the load (the index fold on the address path, the returned word +// written on the next cycle). One instruction per cycle for every lane: LANES lane-ops per cycle. +// +// Instruction word: op[3:0] dst[8:4] src[13:9] src2[18:14] imm[23:19] aux[31:24]. +// 0 add 1 sub 2 xor 3 or 4 rotl(imm) 5 rotr(src) 6 mul 7 mulhi 8 mad 9 shfl(imm mask) 10 prmt(aux,aux) +// 11 lop3(aux lut) 12 load(fold(src) -> addr; dst <= returned word) 13 add 14 xor 15 sub +`include "lane_common.vh" +module core_v6 #(parameter LANES = 32, parameter LOG_LANES = 5, parameter REGS = 32, parameter LOG_REGS = 5) ( + input clk, input rst, input run, + input prog_we, input [7:0] prog_addr, input [31:0] prog_data, + input cfg_en, input [7:0] cfg_n, input [31:0] cfg_m, input [4:0] cfg_r, input [31:0] cfg_wm, input [31:0] cfg_off, input [31:0] cfg_mask, + input [31:0] ld_val, + output [31:0] addr, output [31:0] out); + // instruction memory and the sequencer + reg [31:0] imem [0:255]; + reg [7:0] pc; reg [7:0] n_q; reg [31:0] ir; + reg [31:0] m_q, wm_q, off_q, mask_q; reg [4:0] r_q; + integer i, l; + always @(posedge clk) begin + if (prog_we) imem[prog_addr] <= prog_data; + if (rst) begin pc <= 0; ir <= 0; n_q <= 8'd255; m_q <= 32'h9e3779b1; r_q <= 5'd13; wm_q <= 32'h0fffffc0; off_q <= 3; mask_q <= 32'h0fffffff; end + else begin + if (cfg_en) begin n_q <= cfg_n; m_q <= cfg_m | 1; r_q <= cfg_r; wm_q <= cfg_wm; off_q <= cfg_off; mask_q <= cfg_mask; end + if (run) begin ir <= imem[pc]; pc <= (pc == n_q) ? 8'd0 : pc + 8'd1; end + end + end + // decode + wire [3:0] op = ir[3:0]; wire [LOG_REGS-1:0] dst = ir[4 +: LOG_REGS]; wire [LOG_REGS-1:0] src = ir[9 +: LOG_REGS]; wire [LOG_REGS-1:0] src2 = ir[14 +: LOG_REGS]; + wire [4:0] imm = ir[23:19]; wire [7:0] aux = ir[31:24]; + wire is_load = (op == 4'd12); + wire [4:0] rn = (imm == 0) ? 5'd1 : imm; + wire [LOG_LANES-1:0] smask = imm[LOG_LANES-1:0]; + // lanes + reg [31:0] rf [0:LANES*REGS-1]; + wire [31:0] d_v [0:LANES-1]; wire [31:0] s_v [0:LANES-1]; wire [31:0] s2_v [0:LANES-1]; + wire [31:0] shin [0:LANES-1]; wire [31:0] shout [0:LANES-1]; + wire [31:0] res [0:LANES-1]; + genvar g, st; + generate for (g = 0; g < LANES; g = g + 1) begin : ln + assign d_v[g] = rf[g*REGS + dst]; + assign s_v[g] = rf[g*REGS + src]; + assign s2_v[g] = rf[g*REGS + src2]; + assign shin[g] = s_v[g]; + end endgenerate + // the butterfly shuffle network across the lanes + wire [31:0] bf [0:LOG_LANES][0:LANES-1]; + generate + for (g = 0; g < LANES; g = g + 1) begin : bf0 + assign bf[0][g] = shin[g]; + end + for (st = 0; st < LOG_LANES; st = st + 1) begin : bfs + for (g = 0; g < LANES; g = g + 1) begin : bfl + assign bf[st+1][g] = smask[st] ? bf[st][g ^ (1 << st)] : bf[st][g]; + end + end + for (g = 0; g < LANES; g = g + 1) begin : bfo + assign shout[g] = bf[LOG_LANES][g]; + end + endgenerate + // the units per lane + function [7:0] pick; input [63:0] b; input [3:0] k; reg [7:0] v; + begin v = b[8*k[2:0] +: 8]; pick = k[3] ? {8{v[7]}} : v; end + endfunction + generate for (g = 0; g < LANES; g = g + 1) begin : un + wire [31:0] d = d_v[g]; wire [31:0] s = s_v[g]; wire [31:0] s2 = s2_v[g]; + wire [4:0] sn = (s[4:0] == 0) ? 5'd1 : s[4:0]; + wire mad = (op == 4'd8); + wire [63:0] p = (mad ? s : d) * (mad ? s2 : s); + wire [63:0] bytes = {s, d}; wire [15:0] sel = {aux, aux}; + wire [31:0] prm = {pick(bytes, sel[15:12]), pick(bytes, sel[11:8]), pick(bytes, sel[7:4]), pick(bytes, sel[3:0])}; + reg [31:0] lp; integer b; + always @* for (b = 0; b < 32; b = b + 1) lp[b] = aux[{d[b], s[b], s2[b]}]; + reg [31:0] r; + always @* begin + case (op) + 4'd0, 4'd13: r = d + s; + 4'd1, 4'd15: r = d - s; + 4'd2, 4'd14: r = d ^ s; + 4'd3: r = d | s; + 4'd4: r = `ROTL32(d, rn); + 4'd5: r = `ROTR32(d, sn); + 4'd6: r = p[31:0]; + 4'd7: r = p[63:32]; + 4'd8: r = p[31:0] + d; + 4'd9: r = d ^ shout[g]; + 4'd10: r = prm; + 4'd11: r = lp; + default: r = ld_val ^ (32'h9e3779b9 * (g + 1)); // the returned word (lane-salted by the testbench's bus) + endcase + end + assign res[g] = r; + end endgenerate + // the address path: the fold of lane 0's source (one address per lane on a real part; lane 0 drives the port) + wire [31:0] fx = s_v[0] * m_q; + wire [4:0] frn = (r_q == 0) ? 5'd1 : r_q; + wire [31:0] fy = `ROTL32(fx, frn); + assign addr = is_load ? (((fy & wm_q) | off_q) & mask_q) : 32'd0; + // writeback + always @(posedge clk) begin + if (rst) begin for (i = 0; i < LANES*REGS; i = i + 1) rf[i] <= 32'h9e3779b9 * (i + 1); end + else if (run) for (l = 0; l < LANES; l = l + 1) rf[l*REGS + dst] <= res[l]; + end + // keep every lane alive: the xor over the lanes' results + reg [31:0] red; integer q; + always @* begin red = 0; for (q = 0; q < LANES; q = q + 1) red = red ^ res[q]; end + assign out = red; +endmodule diff --git a/tools/chip-model/rtl/rtl/core_v6_32.v b/tools/chip-model/rtl/rtl/core_v6_32.v new file mode 100644 index 000000000..4b4f0e4a7 --- /dev/null +++ b/tools/chip-model/rtl/rtl/core_v6_32.v @@ -0,0 +1,7 @@ +`include "core_v6.v" +module core_v6_32(input clk, input rst, input run, input prog_we, input [7:0] prog_addr, input [31:0] prog_data, + input cfg_en, input [7:0] cfg_n, input [31:0] cfg_m, input [4:0] cfg_r, input [31:0] cfg_wm, input [31:0] cfg_off, input [31:0] cfg_mask, + input [31:0] ld_val, output [31:0] addr, output [31:0] out); + core_v6 #(.LANES(32), .LOG_LANES(5)) c(.clk(clk), .rst(rst), .run(run), .prog_we(prog_we), .prog_addr(prog_addr), .prog_data(prog_data), + .cfg_en(cfg_en), .cfg_n(cfg_n), .cfg_m(cfg_m), .cfg_r(cfg_r), .cfg_wm(cfg_wm), .cfg_off(cfg_off), .cfg_mask(cfg_mask), .ld_val(ld_val), .addr(addr), .out(out)); +endmodule diff --git a/tools/chip-model/rtl/rtl/core_v6_32r16.v b/tools/chip-model/rtl/rtl/core_v6_32r16.v new file mode 100644 index 000000000..96fe03d67 --- /dev/null +++ b/tools/chip-model/rtl/rtl/core_v6_32r16.v @@ -0,0 +1,7 @@ +`include "core_v6.v" +module core_v6_32r16(input clk, input rst, input run, input prog_we, input [7:0] prog_addr, input [31:0] prog_data, + input cfg_en, input [7:0] cfg_n, input [31:0] cfg_m, input [4:0] cfg_r, input [31:0] cfg_wm, input [31:0] cfg_off, input [31:0] cfg_mask, + input [31:0] ld_val, output [31:0] addr, output [31:0] out); + core_v6 #(.LANES(32), .LOG_LANES(5), .REGS(16), .LOG_REGS(4)) c(.clk(clk), .rst(rst), .run(run), .prog_we(prog_we), .prog_addr(prog_addr), .prog_data(prog_data), + .cfg_en(cfg_en), .cfg_n(cfg_n), .cfg_m(cfg_m), .cfg_r(cfg_r), .cfg_wm(cfg_wm), .cfg_off(cfg_off), .cfg_mask(cfg_mask), .ld_val(ld_val), .addr(addr), .out(out)); +endmodule diff --git a/tools/chip-model/rtl/rtl/core_v6_8.v b/tools/chip-model/rtl/rtl/core_v6_8.v new file mode 100644 index 000000000..c9e21b232 --- /dev/null +++ b/tools/chip-model/rtl/rtl/core_v6_8.v @@ -0,0 +1,7 @@ +`include "core_v6.v" +module core_v6_8(input clk, input rst, input run, input prog_we, input [7:0] prog_addr, input [31:0] prog_data, + input cfg_en, input [7:0] cfg_n, input [31:0] cfg_m, input [4:0] cfg_r, input [31:0] cfg_wm, input [31:0] cfg_off, input [31:0] cfg_mask, + input [31:0] ld_val, output [31:0] addr, output [31:0] out); + core_v6 #(.LANES(8), .LOG_LANES(3)) c(.clk(clk), .rst(rst), .run(run), .prog_we(prog_we), .prog_addr(prog_addr), .prog_data(prog_data), + .cfg_en(cfg_en), .cfg_n(cfg_n), .cfg_m(cfg_m), .cfg_r(cfg_r), .cfg_wm(cfg_wm), .cfg_off(cfg_off), .cfg_mask(cfg_mask), .ld_val(ld_val), .addr(addr), .out(out)); +endmodule diff --git a/tools/chip-model/rtl/tb/tb_core_common.vh b/tools/chip-model/rtl/tb/tb_core_common.vh new file mode 100644 index 000000000..551ba6b4d --- /dev/null +++ b/tools/chip-model/rtl/tb/tb_core_common.vh @@ -0,0 +1,34 @@ +// shared body for the core testbenches: `TOP and `LANES set by the including file +`timescale 1ps/1ps +module tb; + reg clk = 0, rst = 1, run = 0, prog_we = 0, cfg_en = 0; reg [7:0] prog_addr = 0, cfg_n = 0; reg [31:0] prog_data = 0, cfg_m = 0, cfg_wm = 0, cfg_off = 0, cfg_mask = 0, ld_val = 0; reg [4:0] cfg_r = 0; + wire [31:0] addr, out; + `TOP dut(.clk(clk), .rst(rst), .run(run), .prog_we(prog_we), .prog_addr(prog_addr), .prog_data(prog_data), .cfg_en(cfg_en), .cfg_n(cfg_n), .cfg_m(cfg_m), .cfg_r(cfg_r), .cfg_wm(cfg_wm), .cfg_off(cfg_off), .cfg_mask(cfg_mask), .ld_val(ld_val), .addr(addr), .out(out)); + integer n, cycles, loads, k, roll; reg [31:0] acc = 0; reg [3:0] opc; reg [31:0] w; + // the class v4 draw: add 12, xor 10, mul 8, mad 8, shfl 8, rotl 7, sub 6, mulhi 6, rotr 6, or 4 (sum 75) + function [3:0] draw_op; input integer r; integer x; + begin x = r % 75; if (x < 0) x = -x; + draw_op = (x < 12) ? 4'd0 : (x < 22) ? 4'd2 : (x < 30) ? 4'd6 : (x < 38) ? 4'd8 : (x < 46) ? 4'd9 : (x < 53) ? 4'd4 : (x < 59) ? 4'd1 : (x < 65) ? 4'd7 : (x < 71) ? 4'd5 : 4'd3; + end + endfunction + always #`HALF clk = ~clk; + initial begin + if (!$value$plusargs("cycles=%d", cycles)) cycles = 4000; + if (!$value$plusargs("loads=%d", loads)) loads = 0; // 1: one load in 16 instructions + $dumpfile("dump.vcd"); $dumpvars(0, tb.dut); + repeat (4) @(negedge clk); rst = 0; + // the era draw: constants and the program length + @(negedge clk); cfg_en = 1; cfg_n = 8'd255; cfg_m = 32'h9e3779b1; cfg_r = 5'd13; cfg_wm = 32'h0fffffc0; cfg_off = 3; cfg_mask = 32'h0fffffff; + @(negedge clk); cfg_en = 0; + // the program: 256 instructions drawn with the class v4 weights + for (k = 0; k < 256; k = k + 1) begin + @(negedge clk); w = $random; opc = (loads && (k % 16 == 15)) ? 4'd12 : draw_op($random); + prog_we = 1; prog_addr = k; prog_data = {w[31:4], opc}; + end + @(negedge clk); prog_we = 0; run = 1; + for (n = 0; n < cycles; n = n + 1) begin + @(negedge clk); ld_val = $random; acc = acc ^ out ^ addr; + end + $display("CHECKSUM %08x", acc); $finish; + end +endmodule diff --git a/tools/chip-model/rtl/tb/tb_core_v6_32.v b/tools/chip-model/rtl/tb/tb_core_v6_32.v new file mode 100644 index 000000000..032d82423 --- /dev/null +++ b/tools/chip-model/rtl/tb/tb_core_v6_32.v @@ -0,0 +1,3 @@ +`define TOP core_v6_32 +`define HALF 750 +`include "tb_core_common.vh" diff --git a/tools/chip-model/rtl/tb/tb_core_v6_32r16.v b/tools/chip-model/rtl/tb/tb_core_v6_32r16.v new file mode 100644 index 000000000..f2084f017 --- /dev/null +++ b/tools/chip-model/rtl/tb/tb_core_v6_32r16.v @@ -0,0 +1,3 @@ +`define TOP core_v6_32r16 +`define HALF 750 +`include "tb_core_common.vh" diff --git a/tools/chip-model/rtl/tb/tb_core_v6_8.v b/tools/chip-model/rtl/tb/tb_core_v6_8.v new file mode 100644 index 000000000..367b83e59 --- /dev/null +++ b/tools/chip-model/rtl/tb/tb_core_v6_8.v @@ -0,0 +1,3 @@ +`define TOP core_v6_8 +`define HALF 750 +`include "tb_core_common.vh" From a97605d7c0d7694ed5b9841c0eaa613b3c15c8bf Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Thu, 8 Oct 2026 13:05:13 +0000 Subject: [PATCH 4/9] shadow-k: the programmable core's synthesis-only row (6.9 pJ per lane-op ASAP7, k 0.56 at the lock at N3), the edge table in the absolute convention, the lease-routed Makefile Co-Authored-By: Claude Fable 5.1 Documents-only replay of ed7d2626d (bfccc26ed4d299ad35920bda4aa2af0cc5efbe01) for the box mirror master --- docs/analysis/class-v6/floor/shadow-k.md | 147 ++++++++++++++++++++++ tools/chip-model/rtl/Makefile | 9 +- tools/chip-model/rtl/flow/collect.py | 12 +- tools/chip-model/rtl/flow/edge.py | 37 +++--- tools/chip-model/rtl/flow/power.tcl | 1 + tools/chip-model/rtl/tb/tb_core_common.vh | 2 +- 6 files changed, 186 insertions(+), 22 deletions(-) diff --git a/docs/analysis/class-v6/floor/shadow-k.md b/docs/analysis/class-v6/floor/shadow-k.md index c722a26d3..fb9544ffd 100644 --- a/docs/analysis/class-v6/floor/shadow-k.md +++ b/docs/analysis/class-v6/floor/shadow-k.md @@ -75,3 +75,150 @@ lanes, approximate), the clock tree beyond the block's own, and the result's mov the card side the 15.1a figure is the whole card's marginal per counted op (the sleep floor subtracted), which includes the card's own fetch, decode, operand collection and register file. So the `k` here is the chip's datapath-and-window cost over the card's whole-lane cost: a FLOOR on the chip's cost and so a floor on `k`. + +## 3. The rows: per-unit floors (the minimal lane per family) + +Every row: ASAP7 routed, SPEF, OpenSTA `report_power` under the random-input gate-level VCD (every pin annotated, +0 unannotated), the TC corner (0.70 V). "pJ/op" is total power (internal + switching + leakage) times the clock +period over the ops per cycle. The propagated-0.5 cross-check is in `table.csv`; it agrees within 2x on the +logic lanes and overestimates the multiplier lanes 50x (OpenSTA's statistical propagation through a multiplier +is not a measurement), so the VCD row is the row. The GPU side is 15.1a: pJ per counted op on the RTX 5090 at +stock and at the 1,300 MHz lock, and 6.9 pJ per counted op on the Apple M5 Max (the class v4 shadow, measured). +`k` is absolute: the chip's own pJ per op at the node over the card's at its point. + +| Family (chip RTL) | Cells (with fill) | pJ/op ASAP7 | pJ/op N5 | pJ/op N3 | pJ/op N2 | 5090 pJ/op stock / lock | k, N3 chip vs 5090 stock / lock | k, N3 vs M5 Max 6.9 | k, ASAP7 unscaled vs lock | Label | +|---|---|---|---|---|---|---|---|---|---|---| +| ARX add | 11,631 | 2.16 | 1.51 | 1.09 | 0.78 | 11.3 / 6.2 | 0.096 / 0.18 | 0.16 | 0.35 | synthesised; N3 and N2 scaled (claimed) | +| ARX sub | | 2.14 | 1.50 | 1.08 | 0.78 | 11.3 / 6.2 | 0.095 / 0.17 | 0.16 | 0.35 | | +| ARX xor | | 2.09 | 1.46 | 1.05 | 0.76 | 11.3 / 6.2 | 0.093 / 0.17 | 0.15 | 0.34 | | +| ARX rotl (immediate) | | 2.15 | 1.50 | 1.08 | 0.78 | 11.3 / 6.2 | 0.095 / 0.17 | 0.16 | 0.35 | | +| ARX rotr (by register) | | 2.13 | 1.49 | 1.07 | 0.77 | 11.3 / 6.2 | 0.095 / 0.17 | 0.16 | 0.34 | | +| or (lossy; the values saturate to ones and the activity falls) | | 1.45 | 1.01 | 0.73 | 0.53 | 11.3 / 6.2 | 0.065 / 0.12 | 0.11 | 0.23 | | +| ARX random mix | | 2.24 | 1.57 | 1.13 | 0.81 | 11.3 / 6.2 | 0.10 / 0.18 | 0.16 | 0.36 | | +| mul (32 x 32, low word) | 19,603 | 1.35 | 0.94 | 0.68 | 0.49 | 13.9 / 8.3 | 0.049 / 0.082 | | 0.16 | | +| mulhi (high word; the same multiplier) | | 1.35 | 0.94 | 0.68 | 0.49 | 39.6 / 21.0 | 0.017 / 0.032 | | 0.064 | | +| mad (src x src2 + dst; three reads) | | 3.29 | 2.30 | 1.66 | 1.19 | 13.9 / 8.3 | 0.12 / 0.20 | | 0.40 | | +| mul random mix | | 1.53 | 1.07 | 0.77 | 0.56 | 13.9 / 8.3 | 0.055 / 0.093 | | 0.18 | | +| index fold (x M, rotl R, masks; era constants in registers) | 14,965 | 2.37 | 1.66 | 1.19 | 0.86 | 13.9 / 8.3 (read against int_mul) | 0.086 / 0.14 | | 0.29 | | +| prmt (byte permute) | 6,744 | 1.27 | 0.89 | 0.64 | 0.46 | 22.3 / 11.5 | 0.029 / 0.056 | | 0.11 | | +| lop3 (8-bit truth table) | 7,274 | 1.42 | 0.99 | 0.72 | 0.52 | 24.1 / 13.0 | 0.030 / 0.055 | | 0.11 | | +| 32-lane xor-mask shuffle (butterfly over the 1 KB window), per lane-op | ROW_SHFL | +| 32-lane general crossbar, per lane-op | ROW_XBAR | +| 8 KB scratch, one random read (flop array: the pessimistic form) | ROW_SCRATCH | +| int8 8x8x8 tile, per MAC | ROW_TILE | + +Reading the floors: a lane's add costs the chip about 2.2 pJ at ASAP7 and 1.1 at N3, against the 5090's 6.2 +at the lock and 11.3 at stock, so even the floor is not a tenth of the card's cost at the knee, and the claimed +"2 to 5 pJ for a SIMD array at N5" (15.1) was the right order for the unscaled lane with nothing around it. The +multiplier is the cheapest unit per op relative to the card (the 5090 pays 8.3 pJ for mul and 21 for mulhi, the +chip 0.68 for either, because the high word falls out of the same array), and mad is the dearest for the chip +(three register reads and two units). The fold costs a chip one multiply and a rotate: 1.2 pJ per address at N3, +or 0.15 nJ per hash over 128 loads, a third of a percent of the GDDR7 board's 0.466 microjoules. + +## 4. The headline row: the programmable sequencer core + +The per-unit lanes of section 3 are floors for a chip that cannot exist under class v6: layers 1 and 3 (per-era +op-mix, program-length and read-width draws, family epochs) kill any fixed lane, so the chip that competes is a +programmable core. `core_v6` (`tools/chip-model/rtl/rtl/core_v6.v`) is the minimal in-order SIMD sequencer that +executes the drawn class v6 shadow program: a 256 x 32-bit instruction memory (a flop array) sized to the drawn +program, a program counter wrapping at the era's drawn length, fetch into an instruction register, decode, the +era's parameter registers (M, R, WM, OFF, MASK, N), and per lane a 32 x 32-bit register file (flops) with three +read ports and one write port and every class unit: add, sub, xor, or, rotl, rotr, mul, mulhi, mad, prmt, lop3, +the xor-mask shuffle across the lanes (a butterfly), and the load (the index fold on the address path, the returned +word written next cycle). One instruction per cycle for every lane. The program is 256 instructions drawn with the +class v4 weights (`NONLOAD_WEIGHTS`), with a variant at one load per 16 instructions. Two builds: 8 lanes (the fast +row) and 32 lanes (the imem amortised over 32), and a 32-lane build with a 16-register file (the sensitivity). + +The activity: the gate-level random-input VCD of the synthesised netlist; two run lengths (150 and 800 cycles) +bracket the 260-cycle program-load phase, and solving the pair gives the steady-state run power (the load phase +draws 13.9 mW, the run phase 36.9 mW for 8 lanes at 1.5 ns). + +| Row | Stage | Cells | pJ per lane-op ASAP7 | N5 | N3 | N2 | 5090 stock / lock / M5 Max pJ per op | k absolute at N3 vs stock / lock / M5 Max | k at N2 | k unscaled ASAP7 vs lock | Label | +|---|---|---|---|---|---|---|---|---|---|---|---| +| core, 8 lanes, 32 registers | synthesis only (no wires, no clock tree) | 186,443 | 6.9 | 4.8 | 3.5 | 2.5 | 11.3 / 6.2 / 6.9 | 0.31 / 0.56 / 0.50 | 0.22 / 0.40 / 0.36 | 1.1 | synthesised; 14:0x UK | +| of which the sequential term (register file, imem and IR clock pins, no clock gating) | | | 2.4 | 1.7 | 1.2 | 0.9 | | | | | | +| of which the units, the read muxes and the butterfly | | | 4.5 | 3.1 | 2.3 | 1.6 | | | | | | +| core, 8 lanes, 32 registers | placed and routed, SPEF | ROW_CORE8_PLACED | +| core, 32 lanes, 32 registers | ROW_CORE32 | +| core, 32 lanes, 16 registers | ROW_CORE32R16 | +| the bare ARX lane (section 3, the floor) | routed | 11,631 | 2.2 | 1.5 | 1.1 | 0.8 | 11.3 / 6.2 / 6.9 | 0.10 / 0.18 / 0.16 | 0.07 / 0.13 / 0.11 | 0.35 | the lower bound | + +Reading: fetch, decode, a 32-register file and the full unit set cost a chip 3.1x the bare lane, and the core's +k at the 1,300 lock is 0.56 at N3 (0.40 at N2), inside the record's claimed 0.3 to 0.8 band and at its centre, with +the bare lane's 0.18 as the lower bound. Two corrections pull opposite ways: placement adds wires and a clock tree +(+20 to +40 percent on a design like this, approximate; the placed rows read it) and a chip maker gates the +register-file clock (one of 32 registers is written per cycle; gating removes about 2.0 of the 2.4 pJ sequential +term, approximate), so the net figure for the re-fold is 7.0 / 4.9 / 3.5 / 2.5 pJ per lane-op at ASAP7 / N5 / N3 / +N2 (the synthesis-only figure within the rounding). The imem is amortised over 8 lanes in this row and over 32 in +the core32 row. + +## 5. The chip edge at the measured k + +`E_chip = E_mem + N_ops x e_chip` (absolute: the chip's shadow cost is 102,100 x 3.5 pJ = 0.36 microjoules per hash +at N3, 0.26 at N2, whatever the card does); the record's convention `E_mem + k x F` beside it with `k` read at the +card's point (it coincides at the lock by construction and is the same at stock here because both are the same +arithmetic on the same ops; it differs when the chip's cost is held fixed while the card's point moves, which is +what the SRAM lane found flattered the die 1.6x). Cards: the 5090 at stock (3.36 microjoules, F 1.10) and at the +lock (2.33, F 0.652), the M5 Max (1.40 on class v4, 0.78 on class v3, 6.9 pJ per op). Chips: the record's GDDR7 +board and the hardware-future lane's rows (their `E_mem` at zero shadow). The record's rows that this recomputes +are hardware-future.md section 5: "2.6x to 2.8x at k = 1 and 4.4x to 5.1x at k = 0.5" for the strongest DRAM chips +against the stock 5090 (the GDDR7 board's own row there is 2.1x and 3.3x). + +Chip shadow per hash, absolute: N3 0.357 microjoules, N2 0.255. +| Chip (E_mem, microjoules) | Card row | Zero shadow | Absolute, N3 core | Absolute, N2 core | Record convention at the measured k (N3) | at k = 1 | at k = 0.5 | +|---|---|---|---|---|---|---|---| +| GDDR7 board, 28 nm controller (the record) (0.466) | 5090 stock (class v4) | 4.8x | 4.1x | 4.7x | 4.1x (k 0.32) | 2.1x | 3.3x | +| GDDR7 board, 28 nm controller (the record) (0.466) | 5090 at the 1,300 MHz lock | 3.6x | 2.8x | 3.2x | 2.8x (k 0.55) | 2.1x | 2.9x | +| GDDR7 board, 28 nm controller (the record) (0.466) | M5 Max (class v4 1.40 microjoules, class v3 0.78; 6.9 pJ per op; F 0.62) | 1.7x | 1.7x | 1.9x | 1.8x (k 0.51) | 1.3x | 1.8x | +| HBM3E one stack (0.321) | 5090 stock (class v4) | 7.0x | 5.0x | 5.8x | 5.0x (k 0.32) | 2.4x | 3.9x | +| HBM3E one stack (0.321) | 5090 at the 1,300 MHz lock | 5.2x | 3.4x | 4.0x | 3.4x (k 0.55) | 2.4x | 3.6x | +| HBM3E one stack (0.321) | M5 Max (class v4 1.40 microjoules, class v3 0.78; 6.9 pJ per op; F 0.62) | 2.4x | 2.1x | 2.4x | 2.2x (k 0.51) | 1.5x | 2.2x | +| HBM4 one stack, N12 base die (0.22) | 5090 stock (class v4) | 10.3x | 5.8x | 7.1x | 5.8x (k 0.32) | 2.5x | 4.4x | +| HBM4 one stack, N12 base die (0.22) | 5090 at the 1,300 MHz lock | 7.6x | 4.0x | 4.9x | 4.0x (k 0.55) | 2.7x | 4.3x | +| HBM4 one stack, N12 base die (0.22) | M5 Max (class v4 1.40 microjoules, class v3 0.78; 6.9 pJ per op; F 0.62) | 3.5x | 2.4x | 2.9x | 2.6x (k 0.51) | 1.7x | 2.6x | +| custom HBM4E base die, N3P (0.18) | 5090 stock (class v4) | 12.6x | 6.3x | 7.7x | 6.3x (k 0.32) | 2.6x | 4.6x | +| custom HBM4E base die, N3P (0.18) | 5090 at the 1,300 MHz lock | 9.3x | 4.3x | 5.4x | 4.3x (k 0.55) | 2.8x | 4.6x | +| custom HBM4E base die, N3P (0.18) | M5 Max (class v4 1.40 microjoules, class v3 0.78; 6.9 pJ per op; F 0.62) | 4.3x | 2.6x | 3.2x | 2.8x (k 0.51) | 1.7x | 2.9x | +| DRAM on logic, hybrid bonded (2029 to 2031) (0.15) | 5090 stock (class v4) | 15.1x | 6.6x | 8.3x | 6.6x (k 0.32) | 2.7x | 4.8x | +| DRAM on logic, hybrid bonded (2029 to 2031) (0.15) | 5090 at the 1,300 MHz lock | 11.2x | 4.6x | 5.7x | 4.6x (k 0.55) | 2.9x | 4.9x | +| DRAM on logic, hybrid bonded (2029 to 2031) (0.15) | M5 Max (class v4 1.40 microjoules, class v3 0.78; 6.9 pJ per op; F 0.62) | 5.2x | 2.8x | 3.5x | 3.0x (k 0.51) | 1.8x | 3.0x | +| SRAM full store, one N2 reticle (0.14) | 5090 stock (class v4) | 16.1x | 6.8x | 8.5x | 6.8x (k 0.32) | 2.7x | 4.9x | +| SRAM full store, one N2 reticle (0.14) | 5090 at the 1,300 MHz lock | 12.0x | 4.7x | 5.9x | 4.7x (k 0.55) | 2.9x | 5.0x | +| SRAM full store, one N2 reticle (0.14) | M5 Max (class v4 1.40 microjoules, class v3 0.78; 6.9 pJ per op; F 0.62) | 5.6x | 2.8x | 3.5x | 3.1x (k 0.51) | 1.8x | 3.1x | + +Floor lane 1 assumption (not in by 18:00 UK): the honest 5090 at 1.0 microjoules per hash at zero shadow. +| Chip | Premium kept at 0.652 (card 1.652): absolute N3 / N2 | Premium scaled with the card, 0.390 (card 1.390): absolute N3 / N2 | zero shadow | +|---|---|---|---| +| GDDR7 board, 28 nm controller (the record) | 2.0x / 2.3x | 1.7x / 1.9x | 2.1x | +| HBM3E one stack | 2.4x / 2.9x | 2.0x / 2.4x | 3.1x | +| HBM4 one stack, N12 base die | 2.9x / 3.5x | 2.4x / 2.9x | 4.5x | +| custom HBM4E base die, N3P | 3.1x / 3.8x | 2.6x / 3.2x | 5.6x | +| DRAM on logic, hybrid bonded (2029 to 2031) | 3.3x / 4.1x | 2.7x / 3.4x | 6.7x | +| SRAM full store, one N2 reticle | 3.3x / 4.2x | 2.8x / 3.5x | 7.1x | + +What the rows say. (1) At the knee the GDDR7 board keeps 2.8x with the shadow on (3.2x at an N2 core) against +3.6x at zero shadow: the class v4 shadow buys the honest card 0.8x of edge, not the 1.5x the k = 1 row served and +not the 0.1x the bare-lane floor would give. (2) The strongest DRAM chips read 4.3x to 4.7x at the knee and 6.3x to +6.8x at stock against the 5090 (the record's 2.6x to 2.8x at k = 1 and 4.4x to 5.1x at k = 0.5 bracket the stock +figure and understate the knee one, because the record's k was read at stock). (3) Against the M5 Max the whole +table compresses to 1.7x to 2.8x: the honest SoC's own operating point is most of the resistance. (4) Floor lane +1's honest 5090 at 1.0 microjoules (assumed; its rows were not in by 18:00 UK) would put the GDDR7 board at 1.7x +to 2.3x and the strongest chips at 2.6x to 4.2x. + +## 7. Sources + +- The GPU side: `docs/analysis/counter-asic-4-research.md` 15.1a (the 5090 microbench, 8 October 2026, measured) + and 20.3 (the packs job: 10.8 / 6.4 pJ per counted op on the class v4 shadow); the M5 Max 6.9 pJ per counted + op from the coordinator's order of 8 October 2026 (the latency-shadow record). +- ASAP7: L. T. Clark et al., "ASAP7: A 7-nm finFET predictive process design kit", Microelectronics Journal 53 + (2016); the ORFS platform files (`flow/platforms/asap7`, the 7.5-track RVT library, TC corner 0.70 V / 0 C). +- The flow: OpenROAD-flow-scripts (docker image `openroad/orfs:latest`, Yosys 0.68, OpenROAD and OpenSTA with + `read_vcd`); iverilog 12 for the gate-level simulation. +- Node scaling (claimed): TSMC N5 "30 percent lower power at the same speed" against N7 (TSMC technology page, + https://www.tsmc.com/english/dedicatedFoundry/technology/logic/l_5nm); N3E "25 to 30 percent lower power" + against N5 (https://www.tsmc.com/english/dedicatedFoundry/technology/logic/l_3nm); N2 "25 to 30 percent lower + power" against N3E (https://www.tsmc.com/english/dedicatedFoundry/technology/logic/l_2nm). Read 8 October 2026. +- The chip model's memory rows: `docs/analysis/chip-model-v3.md` 5.5 (the GDDR7 board, 0.466 microjoules) and + `docs/analysis/class-v6/hardware-future.md` section 5 (the strongest DRAM and SRAM rows). +- The band: `docs/design/class-v6-rotating-family.md` section 2 (layer 1) and `igneum-pow/src/generator.rs` + `NONLOAD_WEIGHTS`. diff --git a/tools/chip-model/rtl/Makefile b/tools/chip-model/rtl/Makefile index e3bc96c38..974a433bd 100644 --- a/tools/chip-model/rtl/Makefile +++ b/tools/chip-model/rtl/Makefile @@ -13,7 +13,14 @@ WORK ?= $(abspath .) ORFS_IMG ?= openroad/orfs:latest SIM_IMG ?= orfs-sim:latest UID_GID := $(shell id -u):$(shell id -g) -DOCKER := docker run --rm -u $(UID_GID) -e HOME=/tmp -v $(WORK):/work +# Every run on a box goes through the lease tool (the coordinator's rule, 8 October 2026): THREADS cores from the +# bounded pool at nice 19; the docker container is pinned to the leased set and ORFS gets NUM_CORES=THREADS. +THREADS ?= 24 +LEASE ?= /srv/builds/_bin/lease +LEASE_ON ?= $(shell test -x $(LEASE) && echo 1) +LEASEPFX = $(if $(LEASE_ON),$(LEASE) pool $(THREADS) --label "floor-k shadow-k" --owner floor-k --nice 19 --,) +CPUSET = $(if $(LEASE_ON),--cpuset-cpus {cpuset},) +DOCKER := $(LEASEPFX) docker run --rm -u $(UID_GID) -e HOME=/tmp -e NUM_CORES=$(THREADS) $(CPUSET) -v $(WORK):/work ORFS := $(DOCKER) -w /OpenROAD-flow-scripts/flow $(ORFS_IMG) SIM := $(DOCKER) -w /work $(SIM_IMG) diff --git a/tools/chip-model/rtl/flow/collect.py b/tools/chip-model/rtl/flow/collect.py index 4c9bbb4e6..01d9fa21b 100644 --- a/tools/chip-model/rtl/flow/collect.py +++ b/tools/chip-model/rtl/flow/collect.py @@ -62,13 +62,19 @@ for d in OPS: row[f'k_{node}_lock'] = (pj * s / gpu[1]) if gpu[1] else None row[f'k_{node}_unlocked'] = (pj * s / gpu[0]) if gpu[0] else None row['gpu_pJ_unlocked'] = gpu[0]; row['gpu_pJ_lock'] = gpu[1] + # the Apple M5 Max: 6.9 pJ per counted op measured on the class v4 shadow (the honest tier's top); only the + # ARX-class and the core rows have a measured M5 Max figure + m5 = 6.9 if (d in ('arx', 'core8', 'core32', 'core32r16')) else None + row['m5_pJ'] = m5 + for node, sc in SCALE.items(): + row[f'k_{node}_m5'] = (pj * sc / m5) if m5 else None out.append(row) if out: with open(os.path.join(work, 'table.csv'), 'w', newline='') as f: w = csv.DictWriter(f, fieldnames=list(out[0].keys())); w.writeheader(); w.writerows(out) -print('| Family | Sim | Period ps | Cells | Ops/cycle | Total W | Leak W | VCD annotated | pJ/op ASAP7 | pJ/op N5 | pJ/op N3 | pJ/op N2 | 5090 pJ/op unlocked / lock | k at N3 (lock) | k at N2 (lock) | k at N3 (unlocked) |') -print('|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|') +print('| Family | Sim | Period ps | Cells | Ops/cycle | Total W | Leak W | VCD annotated | pJ/op ASAP7 | pJ/op N5 | pJ/op N3 | pJ/op N2 | 5090 pJ/op unlocked / lock | k at N3 (lock) | k at N2 (lock) | k at N3 (unlocked) | k at N3 vs M5 Max (6.9) |') +print('|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|') for r in out: f = lambda v, n=3: ('' if v is None else (f'{v:.{n}g}' if isinstance(v, float) else str(v))) - print(f"| {r['family']} | {r['sim']} | {f(r['period_ps'])} | {r['cells']} | {r['ops_per_cycle']} | {f(r['total_W'])} | {f(r['leak_W'])} | {r['annotated_pct']} | {f(r['pJ_asap7'])} | {f(r['pJ_N5'])} | {f(r['pJ_N3'])} | {f(r['pJ_N2'])} | {f(r['gpu_pJ_unlocked'])} / {f(r['gpu_pJ_lock'])} | {f(r['k_N3_lock'])} | {f(r['k_N2_lock'])} | {f(r['k_N3_unlocked'])} |") + print(f"| {r['family']} | {r['sim']} | {f(r['period_ps'])} | {r['cells']} | {r['ops_per_cycle']} | {f(r['total_W'])} | {f(r['leak_W'])} | {r['annotated_pct']} | {f(r['pJ_asap7'])} | {f(r['pJ_N5'])} | {f(r['pJ_N3'])} | {f(r['pJ_N2'])} | {f(r['gpu_pJ_unlocked'])} / {f(r['gpu_pJ_lock'])} | {f(r['k_N3_lock'])} | {f(r['k_N2_lock'])} | {f(r['k_N3_unlocked'])} | {f(r['k_N3_m5'])} |") diff --git a/tools/chip-model/rtl/flow/edge.py b/tools/chip-model/rtl/flow/edge.py index cf4542d07..1363f339d 100644 --- a/tools/chip-model/rtl/flow/edge.py +++ b/tools/chip-model/rtl/flow/edge.py @@ -1,23 +1,26 @@ #!/usr/bin/env python3 -"""The chip edge rows recomputed at the measured k. Usage: edge.py """ +"""The chip edge rows at the synthesised chip cost. Usage: edge.py [] +Absolute convention: E_chip = E_mem + N_ops x e_chip (the chip's shadow cost does not depend on the card's point). +The record's convention beside it: E_chip = E_mem + k x F with k = e_chip / e_gpu at the card's point.""" import sys -k_v4 = float(sys.argv[1]); k_best = float(sys.argv[2]) -# the record's rows: card microjoules per hash and the premium F (counter-asic-4-research.md 20.4; the GDDR7 -# board E_mem 0.466; hardware-future.md section 5 for the strongest DRAM chips) -cards = [('5090 unlocked (class v4 shape)', 3.36, 1.10), ('5090 at the 1,300 MHz lock', 2.33, 0.652)] +e3 = float(sys.argv[1]); e2 = float(sys.argv[2]) if len(sys.argv) > 2 else e3 * 0.72 +NOPS = 102100 # class v4 counted ops per hash (the packs job) +# the 5090 rows: (name, card microjoules per hash, premium F, pJ per counted op at that point) +cards = [('5090 stock (class v4)', 3.36, 1.10, 10.8), ('5090 at the 1,300 MHz lock', 2.33, 0.652, 6.4)] +honest = [('M5 Max (class v4 1.40 microjoules, class v3 0.78; 6.9 pJ per op; F 0.62)', 1.40, 0.62, 6.9)] chips = [('GDDR7 board, 28 nm controller (the record)', 0.466), ('HBM3E one stack', 0.321), ('HBM4 one stack, N12 base die', 0.22), ('custom HBM4E base die, N3P', 0.18), ('DRAM on logic, hybrid bonded (2029 to 2031)', 0.15), ('SRAM full store, one N2 reticle', 0.14)] -def edge(card, F, mem, k): return (card) / (mem + k * F) -print('| Chip | 5090 row | Card uJ | F uJ | Edge at k = 1 | at k = 0.5 | at k = 0.3 | at the measured k (class v4 mix) | at the measured k (best mix) | at zero shadow |') -print('|---|---|---|---|---|---|---|---|---|---|') +def edge_abs(card, mem, e): return card / (mem + NOPS * e * 1e-6) +def edge_rec(card, F, mem, k): return card / (mem + k * F) +print(f'Chip shadow per hash, absolute: N3 {NOPS*e3*1e-6:.3f} microjoules, N2 {NOPS*e2*1e-6:.3f}.') +print('| Chip (E_mem, microjoules) | Card row | Zero shadow | Absolute, N3 core | Absolute, N2 core | Record convention at the measured k (N3) | at k = 1 | at k = 0.5 |') +print('|---|---|---|---|---|---|---|---|') for cn, mem in chips: - for rn, card, F in cards: - print(f'| {cn} | {rn} | {card:.2f} | {F:.3f} | {edge(card,F,mem,1):.2f}x | {edge(card,F,mem,0.5):.2f}x | {edge(card,F,mem,0.3):.2f}x | {edge(card,F,mem,k_v4):.2f}x (k {k_v4:.3f}) | {edge(card,F,mem,k_best):.2f}x (k {k_best:.3f}) | {(card-F)/mem:.2f}x |') -# the honest-card variants: floor lane 1's assumed 1.0 microjoules per hash at zero shadow on the 5090 -print('\nFloor lane 1 assumption: the honest 5090 at 1.0 microjoules per hash at zero shadow (not in by the default time).') -print('| Chip | Premium kept at the lock\'s 0.652 | Premium scaled with the card (0.652 x 1.0 / 1.67 = 0.390) |') -print('|---|---|---|') + for rn, card, F, egpu in cards + honest: + k3 = e3 / egpu + print(f'| {cn} ({mem}) | {rn} | {(card-F)/mem:.1f}x | {edge_abs(card,mem,e3):.1f}x | {edge_abs(card,mem,e2):.1f}x | {edge_rec(card,F,mem,k3):.1f}x (k {k3:.2f}) | {edge_rec(card,F,mem,1):.1f}x | {edge_rec(card,F,mem,0.5):.1f}x |') +print('\nFloor lane 1 assumption (not in by 18:00 UK): the honest 5090 at 1.0 microjoules per hash at zero shadow.') +print('| Chip | Premium kept at 0.652 (card 1.652): absolute N3 / N2 | Premium scaled with the card, 0.390 (card 1.390): absolute N3 / N2 | zero shadow |') +print('|---|---|---|---|') for cn, mem in chips: - a = edge(1.0 + 0.652, 0.652, mem, k_v4); b = edge(1.0 + 0.390, 0.390, mem, k_v4) - a1 = edge(1.0 + 0.652, 0.652, mem, 1.0); b1 = edge(1.0 + 0.390, 0.390, mem, 1.0) - print(f'| {cn} | {a:.2f}x at the measured k ({a1:.2f}x at k = 1; {1.0/mem:.2f}x at zero shadow) | {b:.2f}x at the measured k ({b1:.2f}x at k = 1) |') + print(f'| {cn} | {edge_abs(1.652,mem,e3):.1f}x / {edge_abs(1.652,mem,e2):.1f}x | {edge_abs(1.390,mem,e3):.1f}x / {edge_abs(1.390,mem,e2):.1f}x | {1.0/mem:.1f}x |') diff --git a/tools/chip-model/rtl/flow/power.tcl b/tools/chip-model/rtl/flow/power.tcl index 3e6e7f807..44c19072c 100644 --- a/tools/chip-model/rtl/flow/power.tcl +++ b/tools/chip-model/rtl/flow/power.tcl @@ -18,6 +18,7 @@ set_power_activity -input_port rst -activity 0 -duty 0 report_power foreach vcd [glob -nocomplain /work/sim/$::env(DESIGN_NICKNAME)/*.vcd] { set tag [file rootname [file tail $vcd]] + if { $tag == "dump" } { continue } puts "FLOORK === POWER_VCD $tag ===" read_vcd -scope tb/dut $vcd if { [info commands report_activity_annotation] != "" } { report_activity_annotation } diff --git a/tools/chip-model/rtl/tb/tb_core_common.vh b/tools/chip-model/rtl/tb/tb_core_common.vh index 551ba6b4d..c1dac7f2c 100644 --- a/tools/chip-model/rtl/tb/tb_core_common.vh +++ b/tools/chip-model/rtl/tb/tb_core_common.vh @@ -13,7 +13,7 @@ module tb; endfunction always #`HALF clk = ~clk; initial begin - if (!$value$plusargs("cycles=%d", cycles)) cycles = 4000; + if (!$value$plusargs("cycles=%d", cycles)) cycles = 300; // the core VCDs run to gigabytes per thousand cycles if (!$value$plusargs("loads=%d", loads)) loads = 0; // 1: one load in 16 instructions $dumpfile("dump.vcd"); $dumpvars(0, tb.dut); repeat (4) @(negedge clk); rst = 0; From 069d4f032f279945d8909a175f3b52f925162794 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Thu, 8 Oct 2026 13:10:39 +0000 Subject: [PATCH 5/9] shadow-k: the ranking and the mixed draw (the census lane's re-weight on the unit floors and the core) Co-Authored-By: Claude Fable 5.1 Documents-only replay of e92167057 (bfccc26ed4d299ad35920bda4aa2af0cc5efbe01) for the box mirror master --- docs/analysis/class-v6/floor/shadow-k.md | 48 ++++++++++++++++++++++++ 1 file changed, 48 insertions(+) diff --git a/docs/analysis/class-v6/floor/shadow-k.md b/docs/analysis/class-v6/floor/shadow-k.md index fb9544ffd..bf294bc95 100644 --- a/docs/analysis/class-v6/floor/shadow-k.md +++ b/docs/analysis/class-v6/floor/shadow-k.md @@ -205,6 +205,54 @@ table compresses to 1.7x to 2.8x: the honest SoC's own operating point is most o 1's honest 5090 at 1.0 microjoules (assumed; its rows were not in by 18:00 UK) would put the GDDR7 board at 1.7x to 2.3x and the strongest chips at 2.6x to 4.2x. +## 6. The ranking and the mixed draw + +### 6.1 Families ranked by k, highest first (the hardest for a chip), absolute at N3 against the 5090's lock + +| Rank | Family | Chip pJ/op N3 (unit floor) | 5090 pJ/op at the lock | k (floor) | In the class draw? | +|---|---|---|---|---|---| +| 1 | mad | 1.66 | 8.3 | 0.20 | yes (8 of 75) | +| 2 | add, sub, rotl, rotr, xor | 1.05 to 1.09 | 6.2 | 0.17 to 0.18 | yes (12, 6, 7, 6, 10) | +| 3 | the index fold | 1.19 | 8.3 | 0.14 | every load | +| 4 | or | 0.73 | 6.2 | 0.12 | yes (4) | +| 5 | mul | 0.68 | 8.3 | 0.082 | yes (8) | +| 6 | prmt, lop3 | 0.64, 0.72 | 11.5, 13.0 | 0.056, 0.055 | not drawn (RTL rows only) | +| 7 | mulhi | 0.68 | 21.0 | 0.032 | yes (6) | +| 8 | 32-lane shuffle (butterfly) | ROW_SHFL_K | 29.4 | pending | yes (8) | +| 9 | L1 scratch read (8 KB flop array) against the card's L2 hit | ROW_SCRATCH_K | 1,400 | pending | not drawn | +| 10 | int8 8x8x8 tile, per MAC | ROW_TILE_K | 2.2 | pending | not drawn (the tensor lever is dead on other grounds) | + +The order is set by the card's price, not the chip's: the chip pays 0.6 to 1.7 pJ for everything, and the card +pays 6.2 for an add, 8.3 for a multiply, 21 for a high word and 29 for a shuffle. So the families the card pays +MOST for (mulhi, shuffle) are the ones a chip undercuts most, and the forcing work is the plain ARX and mad the +card does cheapest. This is the measured form of 15.1b's hold on the shuffle-heavy re-weight. + +### 6.2 The mixed draw that maximises the expected k + +The objective at a FIXED GPU premium: `k_eff(w) = sum w_i e_chip_i / sum w_i e_gpu_i` (the chip's energy for the +drawn program over the card's for the same program; at a fixed `F` the chip pays `k_eff x F`). The band (layer +1, the research lane's 17:00 reading of lane D's cut): B = 4 points on the injecting families only (add, sub, +xor, mad, shfl, rotl, rotr), the lossy families (or, mul, mulhi) at or under their base, `or + mul + mulhi` at most +18 + B, the shuffle capped at its class v4 weight; the index fold on every address and the F8 uniformity floor are +not functions of the weights and do not move. `tools/chip-model/rtl/flow/mix.py` searches the band exhaustively +over the corners (the objective is a ratio of linear forms, so the optimum is at a corner of each family's range). + +On the unit floors (N3, the lock; the shuffle row pending, so computed over the other 67 points of 75): + +| Mix | add, xor, mul, mad, shfl, rotl, sub, mulhi, rotr, or | k_eff (floors) | The card's pJ per op at the lock | k_eff on the core (floor + the core's 2.4 pJ per op overhead at N3) | +|---|---|---|---|---| +| class v4 (the base) | 12, 10, 8, 8, 8, 7, 6, 6, 6, 4 | 0.130 | 8.0 | 0.43 | +| the census lane's re-weight (d20eb04bd, PASS on 256 seeds, no era and eras 0 to 7, W = 4 and W = 16) | 13, 11, 6, 10, 8, 8, 7, 2, 6, 4 | 0.153 (+18 percent) | 7.1 | 0.49 (+14 percent) | +| the band's corner (every injecting family at +4 except the shuffle at its floor; lossy at base minus 4) | 16, 14, 4, 12, 4, 11, 10, 2, 10, 0 | MIX_BEST | | | + +Reading: the re-weight moves six points off the chip's two easiest families (mulhi, mul) onto the hardest (mad, +the ARX families) and lifts the expected k by about a seventh on the core; its acceptance is the census lane's +(0 exhausted, F8 0.999 to 1.008 of uniform, the verifier within 4 percent); its energy cost on the card is a +premium per instruction 11 percent LOWER (fewer mulhi), so at a fixed premium the program is 12 percent longer. +The shuffle's own row decides whether it goes to its floor: on the card it is the dearest op (29.4 pJ at the lock) +and on the chip a butterfly over the window, so its k is the lowest of the drawn families unless the synthesis says +otherwise. The recommendation stands as the census lane's draw until the shuffle row lands (an amendment). + ## 7. Sources - The GPU side: `docs/analysis/counter-asic-4-research.md` 15.1a (the 5090 microbench, 8 October 2026, measured) From 44bb0b406d26cc9b2974f17feee0005bb283c703 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Thu, 8 Oct 2026 13:51:23 +0000 Subject: [PATCH 6/9] shadow-k: the design sweep (64 registers k 0.78, 1,024 imem 0.97 as built and about 0.6 with SRAM, select tree 0.55), the sweep variants' RTL and configs, the synthesis-row driver, the steady-state solve in the collector Co-Authored-By: Claude Fable 5.1 Documents-only replay of 036c1aded (bfccc26ed4d299ad35920bda4aa2af0cc5efbe01) for the box mirror master --- docs/analysis/class-v6/floor/shadow-k.md | 33 +++++++++++++++++++++ tools/chip-model/rtl/Makefile | 9 ++++-- tools/chip-model/rtl/flow/collect.py | 34 +++++++++++++++++----- tools/chip-model/rtl/flow/core32all.mk | 15 ++++++++++ tools/chip-model/rtl/flow/core32all.sdc | 10 +++++++ tools/chip-model/rtl/flow/core8i1k.mk | 15 ++++++++++ tools/chip-model/rtl/flow/core8i1k.sdc | 10 +++++++ tools/chip-model/rtl/flow/core8r64.mk | 15 ++++++++++ tools/chip-model/rtl/flow/core8r64.sdc | 10 +++++++ tools/chip-model/rtl/flow/core8sel.mk | 15 ++++++++++ tools/chip-model/rtl/flow/core8sel.sdc | 10 +++++++ tools/chip-model/rtl/flow/designs.txt | 4 +++ tools/chip-model/rtl/flow/synthrow.sh | 14 +++++++++ tools/chip-model/rtl/rtl/core_v6.v | 29 +++++++++++------- tools/chip-model/rtl/rtl/core_v6_32.v | 6 ++-- tools/chip-model/rtl/rtl/core_v6_32all.v | 7 +++++ tools/chip-model/rtl/rtl/core_v6_32r16.v | 6 ++-- tools/chip-model/rtl/rtl/core_v6_8.v | 6 ++-- tools/chip-model/rtl/rtl/core_v6_8i1k.v | 7 +++++ tools/chip-model/rtl/rtl/core_v6_8r64.v | 7 +++++ tools/chip-model/rtl/rtl/core_v6_8sel.v | 7 +++++ tools/chip-model/rtl/tb/tb_core_common.vh | 16 +++++----- tools/chip-model/rtl/tb/tb_core_legacy.vh | 34 ++++++++++++++++++++++ tools/chip-model/rtl/tb/tb_core_v6_32.v | 2 +- tools/chip-model/rtl/tb/tb_core_v6_32all.v | 6 ++++ tools/chip-model/rtl/tb/tb_core_v6_32r16.v | 3 ++ tools/chip-model/rtl/tb/tb_core_v6_8.v | 3 ++ tools/chip-model/rtl/tb/tb_core_v6_8i1k.v | 6 ++++ tools/chip-model/rtl/tb/tb_core_v6_8r64.v | 6 ++++ tools/chip-model/rtl/tb/tb_core_v6_8sel.v | 6 ++++ 30 files changed, 313 insertions(+), 38 deletions(-) create mode 100644 tools/chip-model/rtl/flow/core32all.mk create mode 100644 tools/chip-model/rtl/flow/core32all.sdc create mode 100644 tools/chip-model/rtl/flow/core8i1k.mk create mode 100644 tools/chip-model/rtl/flow/core8i1k.sdc create mode 100644 tools/chip-model/rtl/flow/core8r64.mk create mode 100644 tools/chip-model/rtl/flow/core8r64.sdc create mode 100644 tools/chip-model/rtl/flow/core8sel.mk create mode 100644 tools/chip-model/rtl/flow/core8sel.sdc create mode 100755 tools/chip-model/rtl/flow/synthrow.sh create mode 100644 tools/chip-model/rtl/rtl/core_v6_32all.v create mode 100644 tools/chip-model/rtl/rtl/core_v6_8i1k.v create mode 100644 tools/chip-model/rtl/rtl/core_v6_8r64.v create mode 100644 tools/chip-model/rtl/rtl/core_v6_8sel.v create mode 100644 tools/chip-model/rtl/tb/tb_core_legacy.vh create mode 100644 tools/chip-model/rtl/tb/tb_core_v6_32all.v create mode 100644 tools/chip-model/rtl/tb/tb_core_v6_8i1k.v create mode 100644 tools/chip-model/rtl/tb/tb_core_v6_8r64.v create mode 100644 tools/chip-model/rtl/tb/tb_core_v6_8sel.v diff --git a/docs/analysis/class-v6/floor/shadow-k.md b/docs/analysis/class-v6/floor/shadow-k.md index bf294bc95..966f3df64 100644 --- a/docs/analysis/class-v6/floor/shadow-k.md +++ b/docs/analysis/class-v6/floor/shadow-k.md @@ -152,6 +152,39 @@ term, approximate), so the net figure for the re-fold is 7.0 / 4.9 / 3.5 / 2.5 p N2 (the synthesis-only figure within the rounding). The imem is amortised over 8 lanes in this row and over 32 in the core32 row. +### 4a. The design sweep: what a class could add to the core's cost (the coordinator's order, 14:3x UK; rows 14:5x) + +Synthesis-only (no wires, no clock tree), 8 lanes unless stated, the same corner and scaling; the steady-state run +power solved from two run lengths (150 and 600 cycles after the program load). The GPU side per knob is the hash +lane's: knob 3 measured on a rented 5090 and 4090 (RunPod, 14:28 to 14:39 UK), knobs 1 and 4 modelled until a +generator line exists (a 64-entry window is a new ISA: an init rule and a fold rule for the extra registers, about +half a day; the select tree is a new instruction kind), knob 2 has no GPU side (the warp's shuffle already spans +32 lanes). + +| Variant | Cells | pJ per lane-op ASAP7 | of which clocking (RF, imem, IR; no gating) | N5 | N3 | N2 | k at N3 vs 5090 stock / lock / M5 Max | GPU side | +|---|---|---|---|---|---|---|---|---| +| base: 32 registers, 256-entry imem (the headline row) | 186,443 | 6.9 | 2.4 | 4.8 | 3.5 | 2.5 | 0.31 / 0.56 / 0.50 | measured (class v4) | +| (1) 64-register file (a 40-bit instruction word) | 267,731 | 9.7 | 3.75 | 6.8 | 4.9 | 3.5 | 0.43 / 0.78 / 0.70 | new ISA; 64 live registers takes a 5090 or 4090 thread to about 110 of 255, occupancy to about half; under the latency-bound chain the rate is expected to hold and the energy to move little (the hash lane, modelled); the unmeasured term is the per-lane register traffic | +| (3) 1,024-entry imem, the program drawn at 1,024, as built (a flop array) | 324,543 | 12.0 | 6.0 | 8.4 | 6.1 | 4.4 | 0.53 / 0.97 / 0.87 | measured: the 1,024 block at 27 passes costs the 5090 1.58x the energy per hash (4.96 against 3.14 microjoules at stock, 111 against 141 MH/s at the 575 W cap) and the 4090 1.57x (7.01 against 4.48, the rate held at 62.5 MH/s), for 4x the shadow instructions: 0.40x per instruction | +| (3) the same with the imem as a 4 KB SRAM macro shared by the lanes (2 to 4 pJ per 32-bit read, approximate) | | about 7.2 | about 1.9 | 5.0 | 3.6 | 2.6 | about 0.32 / 0.58 / 0.52 | the same | +| (4) the drawn select tree (the era's 16-entry op permutation ahead of decode; every unit evaluated every cycle, as in the base) | 186,870 | 6.85 | 2.4 | 4.8 | 3.4 | 2.5 | 0.30 / 0.55 / 0.50 | the units' microbench sum (approximate) | +| (2) 32 lanes, 32 registers (the butterfly across 32; the imem amortised over 32) | ROW_CORE32 | +| (2') 32 lanes, 16 registers (the register-file sensitivity the other way) | ROW_CORE32R16 | +| (5) all four together (32 lanes, 64 registers, 1,024 imem, the select tree) | ROW_CORE32ALL | + +Reading, for the founder's "under 2x at the lock" (which needs k near 0.9 on the GDDR7 board): the 64-register +window is the one robust knob, because its cost is per lane and a chip cannot share it (+2.8 pJ per lane-op at +ASAP7, +0.22 of k at the lock). The long block's cost is instruction memory, which a chip shares across its lanes +as SRAM, so most of its 0.97 as built is the flop array's clock and the honest figure is about 0.6; the card +meanwhile pays 1.58x the energy per hash for it (measured), so on the GDDR7 board at stock the chip reads +4.96 / (0.466 + 408,400 x 3.6 pJ) = 2.6x (1.7x on the flop-array row, which is not a chip anyone builds). The +select tree costs the chip nothing because every unit already evaluates every cycle in the base core. (1) + (3) +together reach about 10 pJ per lane-op at ASAP7 with the SRAM imem (5.0 at N3, k about 0.81 at the lock, 0.45 at +stock), 14.8 as built (k about 1.2); so k 0.85 is reached only on the flop-array reading, and the DRAM board +under 2x at the lock needs the window, the long block and the knee together and holds only while the chip's imem +stays unamortised, which it does not. The GPU pays nothing for the window until occupancy binds, 1.58x per hash for +the long block (0.40x per forcing instruction), and nothing for the select tree. + ## 5. The chip edge at the measured k `E_chip = E_mem + N_ops x e_chip` (absolute: the chip's shadow cost is 102,100 x 3.5 pJ = 0.36 microjoules per hash diff --git a/tools/chip-model/rtl/Makefile b/tools/chip-model/rtl/Makefile index 974a433bd..dea28cc08 100644 --- a/tools/chip-model/rtl/Makefile +++ b/tools/chip-model/rtl/Makefile @@ -24,7 +24,7 @@ DOCKER := $(LEASEPFX) docker run --rm -u $(UID_GID) -e HOME=/tmp -e NUM_CORES= ORFS := $(DOCKER) -w /OpenROAD-flow-scripts/flow $(ORFS_IMG) SIM := $(DOCKER) -w /work $(SIM_IMG) -DESIGNS := arx mul prmt lop3 fold shfl xbar scratch tile core8 core32 core32r16 +DESIGNS := arx mul prmt lop3 fold shfl xbar scratch tile core8 core32 core32r16 core8r64 core8i1k core8sel core32all top = $(shell sed -n 's/^$(1) \([^ ]*\) .*/\1/p' flow/designs.txt) # per-family simulation tags (the op field fixed per row where the family has several ops) @@ -40,6 +40,10 @@ SIMS_tile := mix SIMS_core8 := mix mixld:+loads=1 SIMS_core32 := mix mixld:+loads=1 SIMS_core32r16 := mix mixld:+loads=1 +SIMS_core8r64 := mix +SIMS_core8i1k := mix +SIMS_core8sel := mix +SIMS_core32all := mix .PHONY: rows table clean @@ -67,7 +71,8 @@ synth-%: mkdir -p out/$* logs sim/$*-synth $(ORFS) make DESIGN_CONFIG=/work/flow/$*.mk synth 2>&1 | tee logs/synth-$*.log NET=1_2_yosys.v $(SIM) bash -c 'NET=1_2_yosys.v bash /work/flow/gl2sim.sh $* $(call top,$*)' 2>&1 | tee logs/gl2sim-$*-synth.log - $(SIM) bash /work/flow/sim.sh $* synth 2>&1 | tee logs/sim-$*-synth.log + $(SIM) bash /work/flow/sim.sh $* s150 +cycles=150 2>&1 | tee logs/sim-$*-s150.log + $(SIM) bash /work/flow/sim.sh $* s600 +cycles=600 2>&1 | tee logs/sim-$*-s600.log $(ORFS) make DESIGN_CONFIG=/work/flow/$*.mk RUN_SCRIPT=/work/flow/power.tcl RUN_LOG_NAME_STEM=power_synth FLOORK_ODB=1_synth.odb FLOORK_SDC=1_synth.sdc run 2>&1 | tee logs/power-$*-synth.log table: diff --git a/tools/chip-model/rtl/flow/collect.py b/tools/chip-model/rtl/flow/collect.py index 01d9fa21b..842842019 100644 --- a/tools/chip-model/rtl/flow/collect.py +++ b/tools/chip-model/rtl/flow/collect.py @@ -6,7 +6,7 @@ import re, sys, os, csv work = sys.argv[1] if len(sys.argv) > 1 else '.' # ops per cycle per design (the per-op divisor) and the GPU row each family is read against -OPS = {'arx': 1, 'mul': 1, 'prmt': 1, 'lop3': 1, 'fold': 1, 'shfl': 32, 'xbar': 32, 'scratch': 1, 'tile': 512, 'core8': 8, 'core32': 32} +OPS = {'arx': 1, 'mul': 1, 'prmt': 1, 'lop3': 1, 'fold': 1, 'shfl': 32, 'xbar': 32, 'scratch': 1, 'tile': 512, 'core8': 8, 'core32': 32, 'core32r16': 32, 'core8r64': 8, 'core8i1k': 8, 'core8sel': 8, 'core32all': 32} # 5090 measured pJ per counted op: (unlocked, at the 1,300 MHz lock); 15.1a GPU = { 'arx:mix': (11.3, 6.2), 'arx:add': (11.3, 6.2), 'arx:sub': (11.3, 6.2), 'arx:xor': (11.3, 6.2), 'arx:or': (11.3, 6.2), @@ -17,7 +17,7 @@ GPU = { 'shfl:mix': (55.8, 29.4), 'xbar:mix': (55.8, 29.4), 'scratch:mix': (2400.0, 1400.0), # the card's L2 hit (no shared-memory probe measured: owed) 'tile:mix': (4.1, 2.2), - 'core8:mix': (11.3, 6.2), 'core8:mixld': (11.3, 6.2), 'core32:mix': (11.3, 6.2), 'core32:mixld': (11.3, 6.2), # the class v4 draw: read against int_arx (the packs job read the whole mix at 10.8 / 6.4) # dependent u8 m8n8k16 per MAC; the wide s8 tile reads 1.36 / 0.83 + 'core8:mix': (11.3, 6.2), 'core8:mixld': (11.3, 6.2), 'core32:mix': (11.3, 6.2), 'core32:mixld': (11.3, 6.2), 'core32r16:mix': (11.3, 6.2), 'core8r64:mix': (11.3, 6.2), 'core8i1k:mix': (11.3, 6.2), 'core8sel:mix': (11.3, 6.2), 'core32all:mix': (11.3, 6.2), # the class v4 draw: read against int_arx (the packs job read the whole mix at 10.8 / 6.4) # dependent u8 m8n8k16 per MAC; the wide s8 tile reads 1.36 / 0.83 } # per-node energy scaling from ASAP7 (a 7 nm-class predictive PDK at 0.70 V), approximate and claimed: # N7 -> N5 x0.70 (TSMC: "30 percent lower power at the same speed"), N5 -> N3E x0.72 (TSMC: 25 to 30 percent), @@ -42,15 +42,35 @@ def parse_log(path): total=float(tm.group(4)), annotated=(f"{ann.group(1)} pins, {unann.group(1) if unann else '?'} unannotated") if ann else '') return period, cells, rows +LOAD_CYCLES = 260 # the core testbenches: 4 reset + 2 config + 256 program-load cycles before the run (1,028 at NPROG 1024) +def steady(rows, period, short, long_, cs, cl, loadc): + """Solve the run-phase power from two run lengths: P_i = (loadc x b + c_i x a) / (loadc + c_i).""" + out = {} + for key in ('internal', 'switching', 'leakage', 'total'): + Ps, Pl = rows[short][key], rows[long_][key] + A = Ps * (loadc + cs); B = Pl * (loadc + cl) + a = (B - A) / (cl - cs) + out[key] = a + out['annotated'] = rows[long_]['annotated'] + '; steady state solved from the two run lengths' + return out out = [] for d in OPS: - log = os.path.join(work, 'out', d, 'logs', 'asap7', d, 'base', 'power.log') - if not os.path.exists(log): + found = None + for stem in ('power_synth', 'power_synth_short', 'power'): + log = os.path.join(work, 'out', d, 'logs', 'asap7', d, 'base', stem + '.log') + if os.path.exists(log): + found = log + if stem == 'power': break + if not found: continue - period, cells, rows = parse_log(log) + period, cells, rows = parse_log(found) + pairs = [('vcd:s150', 'vcd:s600', 150, 600), ('vcd:short', 'vcd:synth', 150, 800), ('vcd:s150', 'vcd:s400', 150, 400)] + for sh, lg, cs, cl in pairs: + if sh in rows and lg in rows: + rows['vcd:steady'] = steady(rows, period, sh, lg, cs, cl, 1028 if d in ('core8i1k', 'core32all') else LOAD_CYCLES) for tag, r in rows.items(): sub = tag.split(':')[1] if ':' in tag else 'prop' - key = f'{d}:{sub}' if sub != 'prop' else f'{d}:mix' + key = f'{d}:{sub}' if sub in ('add','sub','xor','or','rotl','rotr','mul','mulhi','mad','mixld') else f'{d}:mix' gpu = GPU.get(key, (None, None)) pj = r['total'] * period * 1e-12 / OPS[d] * 1e12 # W * s / ops -> pJ pj_dyn = (r['internal'] + r['switching']) * period / OPS[d] @@ -64,7 +84,7 @@ for d in OPS: row['gpu_pJ_unlocked'] = gpu[0]; row['gpu_pJ_lock'] = gpu[1] # the Apple M5 Max: 6.9 pJ per counted op measured on the class v4 shadow (the honest tier's top); only the # ARX-class and the core rows have a measured M5 Max figure - m5 = 6.9 if (d in ('arx', 'core8', 'core32', 'core32r16')) else None + m5 = 6.9 if (d == 'arx' or d.startswith('core')) else None row['m5_pJ'] = m5 for node, sc in SCALE.items(): row[f'k_{node}_m5'] = (pj * sc / m5) if m5 else None diff --git a/tools/chip-model/rtl/flow/core32all.mk b/tools/chip-model/rtl/flow/core32all.mk new file mode 100644 index 000000000..c027bf40e --- /dev/null +++ b/tools/chip-model/rtl/flow/core32all.mk @@ -0,0 +1,15 @@ +# ORFS design config for the programmable shadow core (core8: core_v6_32all), ASAP7. +export PLATFORM = asap7 +export DESIGN_NAME = core_v6_32all +export DESIGN_NICKNAME = core32all +export VERILOG_FILES = /work/rtl/core_v6_32all.v +export VERILOG_INCLUDE_DIRS = /work/rtl +export SDC_FILE = /work/flow/core32all.sdc +export CORE_UTILIZATION = 40 +export CORE_ASPECT_RATIO = 1 +export CORE_MARGIN = 0.5 +export PLACE_DENSITY = 0.55 +export CORNER = TC +export SKIP_LAST_GASP = 1 +export WORK_HOME = /work/out/core32all +export SYNTH_MEMORY_MAX_BITS = 2000000 diff --git a/tools/chip-model/rtl/flow/core32all.sdc b/tools/chip-model/rtl/flow/core32all.sdc new file mode 100644 index 000000000..13993d335 --- /dev/null +++ b/tools/chip-model/rtl/flow/core32all.sdc @@ -0,0 +1,10 @@ +current_design core_v6_32all +set clk_name core_clock +set clk_port_name clk +set clk_period 1500 +set clk_io_pct 0.2 +set clk_port [get_ports $clk_port_name] +create_clock -name $clk_name -period $clk_period $clk_port +set non_clock_inputs [all_inputs -no_clocks] +set_input_delay [expr $clk_period * $clk_io_pct] -clock $clk_name $non_clock_inputs +set_output_delay [expr $clk_period * $clk_io_pct] -clock $clk_name [all_outputs] diff --git a/tools/chip-model/rtl/flow/core8i1k.mk b/tools/chip-model/rtl/flow/core8i1k.mk new file mode 100644 index 000000000..bbfe890c8 --- /dev/null +++ b/tools/chip-model/rtl/flow/core8i1k.mk @@ -0,0 +1,15 @@ +# ORFS design config for the programmable shadow core (core8: core_v6_8i1k), ASAP7. +export PLATFORM = asap7 +export DESIGN_NAME = core_v6_8i1k +export DESIGN_NICKNAME = core8i1k +export VERILOG_FILES = /work/rtl/core_v6_8i1k.v +export VERILOG_INCLUDE_DIRS = /work/rtl +export SDC_FILE = /work/flow/core8i1k.sdc +export CORE_UTILIZATION = 40 +export CORE_ASPECT_RATIO = 1 +export CORE_MARGIN = 0.5 +export PLACE_DENSITY = 0.55 +export CORNER = TC +export SKIP_LAST_GASP = 1 +export WORK_HOME = /work/out/core8i1k +export SYNTH_MEMORY_MAX_BITS = 2000000 diff --git a/tools/chip-model/rtl/flow/core8i1k.sdc b/tools/chip-model/rtl/flow/core8i1k.sdc new file mode 100644 index 000000000..fa6e34e57 --- /dev/null +++ b/tools/chip-model/rtl/flow/core8i1k.sdc @@ -0,0 +1,10 @@ +current_design core_v6_8i1k +set clk_name core_clock +set clk_port_name clk +set clk_period 1500 +set clk_io_pct 0.2 +set clk_port [get_ports $clk_port_name] +create_clock -name $clk_name -period $clk_period $clk_port +set non_clock_inputs [all_inputs -no_clocks] +set_input_delay [expr $clk_period * $clk_io_pct] -clock $clk_name $non_clock_inputs +set_output_delay [expr $clk_period * $clk_io_pct] -clock $clk_name [all_outputs] diff --git a/tools/chip-model/rtl/flow/core8r64.mk b/tools/chip-model/rtl/flow/core8r64.mk new file mode 100644 index 000000000..ff5d73224 --- /dev/null +++ b/tools/chip-model/rtl/flow/core8r64.mk @@ -0,0 +1,15 @@ +# ORFS design config for the programmable shadow core (core8: core_v6_8r64), ASAP7. +export PLATFORM = asap7 +export DESIGN_NAME = core_v6_8r64 +export DESIGN_NICKNAME = core8r64 +export VERILOG_FILES = /work/rtl/core_v6_8r64.v +export VERILOG_INCLUDE_DIRS = /work/rtl +export SDC_FILE = /work/flow/core8r64.sdc +export CORE_UTILIZATION = 40 +export CORE_ASPECT_RATIO = 1 +export CORE_MARGIN = 0.5 +export PLACE_DENSITY = 0.55 +export CORNER = TC +export SKIP_LAST_GASP = 1 +export WORK_HOME = /work/out/core8r64 +export SYNTH_MEMORY_MAX_BITS = 2000000 diff --git a/tools/chip-model/rtl/flow/core8r64.sdc b/tools/chip-model/rtl/flow/core8r64.sdc new file mode 100644 index 000000000..ffb92a9cd --- /dev/null +++ b/tools/chip-model/rtl/flow/core8r64.sdc @@ -0,0 +1,10 @@ +current_design core_v6_8r64 +set clk_name core_clock +set clk_port_name clk +set clk_period 1500 +set clk_io_pct 0.2 +set clk_port [get_ports $clk_port_name] +create_clock -name $clk_name -period $clk_period $clk_port +set non_clock_inputs [all_inputs -no_clocks] +set_input_delay [expr $clk_period * $clk_io_pct] -clock $clk_name $non_clock_inputs +set_output_delay [expr $clk_period * $clk_io_pct] -clock $clk_name [all_outputs] diff --git a/tools/chip-model/rtl/flow/core8sel.mk b/tools/chip-model/rtl/flow/core8sel.mk new file mode 100644 index 000000000..07f9f6ebb --- /dev/null +++ b/tools/chip-model/rtl/flow/core8sel.mk @@ -0,0 +1,15 @@ +# ORFS design config for the programmable shadow core (core8: core_v6_8sel), ASAP7. +export PLATFORM = asap7 +export DESIGN_NAME = core_v6_8sel +export DESIGN_NICKNAME = core8sel +export VERILOG_FILES = /work/rtl/core_v6_8sel.v +export VERILOG_INCLUDE_DIRS = /work/rtl +export SDC_FILE = /work/flow/core8sel.sdc +export CORE_UTILIZATION = 40 +export CORE_ASPECT_RATIO = 1 +export CORE_MARGIN = 0.5 +export PLACE_DENSITY = 0.55 +export CORNER = TC +export SKIP_LAST_GASP = 1 +export WORK_HOME = /work/out/core8sel +export SYNTH_MEMORY_MAX_BITS = 2000000 diff --git a/tools/chip-model/rtl/flow/core8sel.sdc b/tools/chip-model/rtl/flow/core8sel.sdc new file mode 100644 index 000000000..36f995570 --- /dev/null +++ b/tools/chip-model/rtl/flow/core8sel.sdc @@ -0,0 +1,10 @@ +current_design core_v6_8sel +set clk_name core_clock +set clk_port_name clk +set clk_period 1500 +set clk_io_pct 0.2 +set clk_port [get_ports $clk_port_name] +create_clock -name $clk_name -period $clk_period $clk_port +set non_clock_inputs [all_inputs -no_clocks] +set_input_delay [expr $clk_period * $clk_io_pct] -clock $clk_name $non_clock_inputs +set_output_delay [expr $clk_period * $clk_io_pct] -clock $clk_name [all_outputs] diff --git a/tools/chip-model/rtl/flow/designs.txt b/tools/chip-model/rtl/flow/designs.txt index 329ec661e..af4f9566e 100644 --- a/tools/chip-model/rtl/flow/designs.txt +++ b/tools/chip-model/rtl/flow/designs.txt @@ -10,3 +10,7 @@ tile tile8 2000 core8 core_v6_8 1500 core32 core_v6_32 1500 core32r16 core_v6_32r16 1500 +core8r64 core_v6_8r64 1500 +core8i1k core_v6_8i1k 1500 +core8sel core_v6_8sel 1500 +core32all core_v6_32all 1500 diff --git a/tools/chip-model/rtl/flow/synthrow.sh b/tools/chip-model/rtl/flow/synthrow.sh new file mode 100755 index 000000000..1b53a543f --- /dev/null +++ b/tools/chip-model/rtl/flow/synthrow.sh @@ -0,0 +1,14 @@ +#!/usr/bin/env bash +# synthrow.sh [cycles-short] [cycles-long] : the synthesis-only power row for a design whose +# ORFS synthesis has already run (1_2_yosys.v and 1_synth.odb present): gate-level sim at two run lengths, then +# OpenSTA power on the synthesised netlist (no wires). Every step under the lease tool when present. +set -euo pipefail +name=$1; cs=${2:-150}; cl=${3:-600} +cd "$(dirname "$0")/.." +WORK=$(pwd); top=$(sed -n "s/^$name \([^ ]*\) .*/\1/p" flow/designs.txt) +UG=$(id -u):$(id -g) +if [ -x /srv/builds/_bin/lease ]; then L="/srv/builds/_bin/lease pool 8 --label floor-k-synthrow-$name --owner floor-k --nice 19 --"; else L=""; fi +$L docker run --rm -u $UG -e HOME=/tmp -e NET=1_2_yosys.v -v $WORK:/work -w /work orfs-sim:latest bash /work/flow/gl2sim.sh $name $top +$L docker run --rm -u $UG -e HOME=/tmp -v $WORK:/work -w /work orfs-sim:latest bash /work/flow/sim.sh $name s$cs +cycles=$cs +$L docker run --rm -u $UG -e HOME=/tmp -v $WORK:/work -w /work orfs-sim:latest bash /work/flow/sim.sh $name s$cl +cycles=$cl +$L docker run --rm -u $UG -e HOME=/tmp -e NUM_CORES=8 -v $WORK:/work -w /OpenROAD-flow-scripts/flow openroad/orfs:latest make DESIGN_CONFIG=/work/flow/$name.mk RUN_SCRIPT=/work/flow/power.tcl RUN_LOG_NAME_STEM=power_synth FLOORK_ODB=1_synth.odb FLOORK_SDC=1_synth.sdc run diff --git a/tools/chip-model/rtl/rtl/core_v6.v b/tools/chip-model/rtl/rtl/core_v6.v index 618df15bf..da57559ca 100644 --- a/tools/chip-model/rtl/rtl/core_v6.v +++ b/tools/chip-model/rtl/rtl/core_v6.v @@ -11,28 +11,35 @@ // 0 add 1 sub 2 xor 3 or 4 rotl(imm) 5 rotr(src) 6 mul 7 mulhi 8 mad 9 shfl(imm mask) 10 prmt(aux,aux) // 11 lop3(aux lut) 12 load(fold(src) -> addr; dst <= returned word) 13 add 14 xor 15 sub `include "lane_common.vh" -module core_v6 #(parameter LANES = 32, parameter LOG_LANES = 5, parameter REGS = 32, parameter LOG_REGS = 5) ( +module core_v6 #(parameter LANES = 32, parameter LOG_LANES = 5, parameter REGS = 32, parameter LOG_REGS = 5, + parameter IW = 32, parameter IMEM_LOG = 8, parameter SELTREE = 0) ( input clk, input rst, input run, - input prog_we, input [7:0] prog_addr, input [31:0] prog_data, - input cfg_en, input [7:0] cfg_n, input [31:0] cfg_m, input [4:0] cfg_r, input [31:0] cfg_wm, input [31:0] cfg_off, input [31:0] cfg_mask, + input prog_we, input [9:0] prog_addr, input [IW-1:0] prog_data, + input cfg_en, input [9:0] cfg_n, input [63:0] cfg_sel, input [31:0] cfg_m, input [4:0] cfg_r, input [31:0] cfg_wm, input [31:0] cfg_off, input [31:0] cfg_mask, input [31:0] ld_val, output [31:0] addr, output [31:0] out); // instruction memory and the sequencer - reg [31:0] imem [0:255]; - reg [7:0] pc; reg [7:0] n_q; reg [31:0] ir; + localparam IMEM = 1 << IMEM_LOG; + reg [IW-1:0] imem [0:IMEM-1]; + reg [IMEM_LOG-1:0] pc; reg [IMEM_LOG-1:0] n_q; reg [IW-1:0] ir; + reg [63:0] sel_q; // the era's drawn op permutation (SELTREE = 1): 16 x 4-bit op codes reg [31:0] m_q, wm_q, off_q, mask_q; reg [4:0] r_q; integer i, l; always @(posedge clk) begin - if (prog_we) imem[prog_addr] <= prog_data; - if (rst) begin pc <= 0; ir <= 0; n_q <= 8'd255; m_q <= 32'h9e3779b1; r_q <= 5'd13; wm_q <= 32'h0fffffc0; off_q <= 3; mask_q <= 32'h0fffffff; end + if (prog_we) imem[prog_addr[IMEM_LOG-1:0]] <= prog_data; + if (rst) begin pc <= 0; ir <= 0; n_q <= {IMEM_LOG{1'b1}}; sel_q <= 64'hfedcba9876543210; m_q <= 32'h9e3779b1; r_q <= 5'd13; wm_q <= 32'h0fffffc0; off_q <= 3; mask_q <= 32'h0fffffff; end else begin - if (cfg_en) begin n_q <= cfg_n; m_q <= cfg_m | 1; r_q <= cfg_r; wm_q <= cfg_wm; off_q <= cfg_off; mask_q <= cfg_mask; end - if (run) begin ir <= imem[pc]; pc <= (pc == n_q) ? 8'd0 : pc + 8'd1; end + if (cfg_en) begin n_q <= cfg_n[IMEM_LOG-1:0]; sel_q <= cfg_sel; m_q <= cfg_m | 1; r_q <= cfg_r; wm_q <= cfg_wm; off_q <= cfg_off; mask_q <= cfg_mask; end + if (run) begin ir <= imem[pc]; pc <= (pc == n_q) ? {IMEM_LOG{1'b0}} : pc + 1'b1; end end end // decode - wire [3:0] op = ir[3:0]; wire [LOG_REGS-1:0] dst = ir[4 +: LOG_REGS]; wire [LOG_REGS-1:0] src = ir[9 +: LOG_REGS]; wire [LOG_REGS-1:0] src2 = ir[14 +: LOG_REGS]; - wire [4:0] imm = ir[23:19]; wire [7:0] aux = ir[31:24]; + // the drawn select tree (SELTREE = 1): the op code is remapped through the era's 16-entry permutation before + // decode, so the datapath's select structure is the era's draw, not a fixed table a chip could hard-wire + wire [3:0] op_raw = ir[3:0]; + wire [3:0] op = SELTREE ? sel_q[op_raw*4 +: 4] : op_raw; + wire [LOG_REGS-1:0] dst = ir[4 +: LOG_REGS]; wire [LOG_REGS-1:0] src = ir[4+LOG_REGS +: LOG_REGS]; wire [LOG_REGS-1:0] src2 = ir[4+2*LOG_REGS +: LOG_REGS]; + wire [4:0] imm = ir[4+3*LOG_REGS +: 5]; wire [7:0] aux = ir[9+3*LOG_REGS +: 8]; wire is_load = (op == 4'd12); wire [4:0] rn = (imm == 0) ? 5'd1 : imm; wire [LOG_LANES-1:0] smask = imm[LOG_LANES-1:0]; diff --git a/tools/chip-model/rtl/rtl/core_v6_32.v b/tools/chip-model/rtl/rtl/core_v6_32.v index 4b4f0e4a7..a12f2b262 100644 --- a/tools/chip-model/rtl/rtl/core_v6_32.v +++ b/tools/chip-model/rtl/rtl/core_v6_32.v @@ -1,7 +1,7 @@ `include "core_v6.v" -module core_v6_32(input clk, input rst, input run, input prog_we, input [7:0] prog_addr, input [31:0] prog_data, - input cfg_en, input [7:0] cfg_n, input [31:0] cfg_m, input [4:0] cfg_r, input [31:0] cfg_wm, input [31:0] cfg_off, input [31:0] cfg_mask, +module core_v6_32(input clk, input rst, input run, input prog_we, input [9:0] prog_addr, input [32-1:0] prog_data, + input cfg_en, input [9:0] cfg_n, input [63:0] cfg_sel, input [31:0] cfg_m, input [4:0] cfg_r, input [31:0] cfg_wm, input [31:0] cfg_off, input [31:0] cfg_mask, input [31:0] ld_val, output [31:0] addr, output [31:0] out); core_v6 #(.LANES(32), .LOG_LANES(5)) c(.clk(clk), .rst(rst), .run(run), .prog_we(prog_we), .prog_addr(prog_addr), .prog_data(prog_data), - .cfg_en(cfg_en), .cfg_n(cfg_n), .cfg_m(cfg_m), .cfg_r(cfg_r), .cfg_wm(cfg_wm), .cfg_off(cfg_off), .cfg_mask(cfg_mask), .ld_val(ld_val), .addr(addr), .out(out)); + .cfg_en(cfg_en), .cfg_n(cfg_n), .cfg_sel(cfg_sel), .cfg_m(cfg_m), .cfg_r(cfg_r), .cfg_wm(cfg_wm), .cfg_off(cfg_off), .cfg_mask(cfg_mask), .ld_val(ld_val), .addr(addr), .out(out)); endmodule diff --git a/tools/chip-model/rtl/rtl/core_v6_32all.v b/tools/chip-model/rtl/rtl/core_v6_32all.v new file mode 100644 index 000000000..918bdb392 --- /dev/null +++ b/tools/chip-model/rtl/rtl/core_v6_32all.v @@ -0,0 +1,7 @@ +`include "core_v6.v" +module core_v6_32all(input clk, input rst, input run, input prog_we, input [9:0] prog_addr, input [40-1:0] prog_data, + input cfg_en, input [9:0] cfg_n, input [63:0] cfg_sel, input [31:0] cfg_m, input [4:0] cfg_r, input [31:0] cfg_wm, input [31:0] cfg_off, input [31:0] cfg_mask, + input [31:0] ld_val, output [31:0] addr, output [31:0] out); + core_v6 #(.LANES(32), .LOG_LANES(5), .REGS(64), .LOG_REGS(6), .IW(40), .IMEM_LOG(10), .SELTREE(1)) c(.clk(clk), .rst(rst), .run(run), .prog_we(prog_we), .prog_addr(prog_addr), .prog_data(prog_data), + .cfg_en(cfg_en), .cfg_n(cfg_n), .cfg_sel(cfg_sel), .cfg_m(cfg_m), .cfg_r(cfg_r), .cfg_wm(cfg_wm), .cfg_off(cfg_off), .cfg_mask(cfg_mask), .ld_val(ld_val), .addr(addr), .out(out)); +endmodule diff --git a/tools/chip-model/rtl/rtl/core_v6_32r16.v b/tools/chip-model/rtl/rtl/core_v6_32r16.v index 96fe03d67..43528b44c 100644 --- a/tools/chip-model/rtl/rtl/core_v6_32r16.v +++ b/tools/chip-model/rtl/rtl/core_v6_32r16.v @@ -1,7 +1,7 @@ `include "core_v6.v" -module core_v6_32r16(input clk, input rst, input run, input prog_we, input [7:0] prog_addr, input [31:0] prog_data, - input cfg_en, input [7:0] cfg_n, input [31:0] cfg_m, input [4:0] cfg_r, input [31:0] cfg_wm, input [31:0] cfg_off, input [31:0] cfg_mask, +module core_v6_32r16(input clk, input rst, input run, input prog_we, input [9:0] prog_addr, input [32-1:0] prog_data, + input cfg_en, input [9:0] cfg_n, input [63:0] cfg_sel, input [31:0] cfg_m, input [4:0] cfg_r, input [31:0] cfg_wm, input [31:0] cfg_off, input [31:0] cfg_mask, input [31:0] ld_val, output [31:0] addr, output [31:0] out); core_v6 #(.LANES(32), .LOG_LANES(5), .REGS(16), .LOG_REGS(4)) c(.clk(clk), .rst(rst), .run(run), .prog_we(prog_we), .prog_addr(prog_addr), .prog_data(prog_data), - .cfg_en(cfg_en), .cfg_n(cfg_n), .cfg_m(cfg_m), .cfg_r(cfg_r), .cfg_wm(cfg_wm), .cfg_off(cfg_off), .cfg_mask(cfg_mask), .ld_val(ld_val), .addr(addr), .out(out)); + .cfg_en(cfg_en), .cfg_n(cfg_n), .cfg_sel(cfg_sel), .cfg_m(cfg_m), .cfg_r(cfg_r), .cfg_wm(cfg_wm), .cfg_off(cfg_off), .cfg_mask(cfg_mask), .ld_val(ld_val), .addr(addr), .out(out)); endmodule diff --git a/tools/chip-model/rtl/rtl/core_v6_8.v b/tools/chip-model/rtl/rtl/core_v6_8.v index c9e21b232..77c025d50 100644 --- a/tools/chip-model/rtl/rtl/core_v6_8.v +++ b/tools/chip-model/rtl/rtl/core_v6_8.v @@ -1,7 +1,7 @@ `include "core_v6.v" -module core_v6_8(input clk, input rst, input run, input prog_we, input [7:0] prog_addr, input [31:0] prog_data, - input cfg_en, input [7:0] cfg_n, input [31:0] cfg_m, input [4:0] cfg_r, input [31:0] cfg_wm, input [31:0] cfg_off, input [31:0] cfg_mask, +module core_v6_8(input clk, input rst, input run, input prog_we, input [9:0] prog_addr, input [32-1:0] prog_data, + input cfg_en, input [9:0] cfg_n, input [63:0] cfg_sel, input [31:0] cfg_m, input [4:0] cfg_r, input [31:0] cfg_wm, input [31:0] cfg_off, input [31:0] cfg_mask, input [31:0] ld_val, output [31:0] addr, output [31:0] out); core_v6 #(.LANES(8), .LOG_LANES(3)) c(.clk(clk), .rst(rst), .run(run), .prog_we(prog_we), .prog_addr(prog_addr), .prog_data(prog_data), - .cfg_en(cfg_en), .cfg_n(cfg_n), .cfg_m(cfg_m), .cfg_r(cfg_r), .cfg_wm(cfg_wm), .cfg_off(cfg_off), .cfg_mask(cfg_mask), .ld_val(ld_val), .addr(addr), .out(out)); + .cfg_en(cfg_en), .cfg_n(cfg_n), .cfg_sel(cfg_sel), .cfg_m(cfg_m), .cfg_r(cfg_r), .cfg_wm(cfg_wm), .cfg_off(cfg_off), .cfg_mask(cfg_mask), .ld_val(ld_val), .addr(addr), .out(out)); endmodule diff --git a/tools/chip-model/rtl/rtl/core_v6_8i1k.v b/tools/chip-model/rtl/rtl/core_v6_8i1k.v new file mode 100644 index 000000000..61758a7b9 --- /dev/null +++ b/tools/chip-model/rtl/rtl/core_v6_8i1k.v @@ -0,0 +1,7 @@ +`include "core_v6.v" +module core_v6_8i1k(input clk, input rst, input run, input prog_we, input [9:0] prog_addr, input [32-1:0] prog_data, + input cfg_en, input [9:0] cfg_n, input [63:0] cfg_sel, input [31:0] cfg_m, input [4:0] cfg_r, input [31:0] cfg_wm, input [31:0] cfg_off, input [31:0] cfg_mask, + input [31:0] ld_val, output [31:0] addr, output [31:0] out); + core_v6 #(.LANES(8), .LOG_LANES(3), .IMEM_LOG(10)) c(.clk(clk), .rst(rst), .run(run), .prog_we(prog_we), .prog_addr(prog_addr), .prog_data(prog_data), + .cfg_en(cfg_en), .cfg_n(cfg_n), .cfg_sel(cfg_sel), .cfg_m(cfg_m), .cfg_r(cfg_r), .cfg_wm(cfg_wm), .cfg_off(cfg_off), .cfg_mask(cfg_mask), .ld_val(ld_val), .addr(addr), .out(out)); +endmodule diff --git a/tools/chip-model/rtl/rtl/core_v6_8r64.v b/tools/chip-model/rtl/rtl/core_v6_8r64.v new file mode 100644 index 000000000..134032fce --- /dev/null +++ b/tools/chip-model/rtl/rtl/core_v6_8r64.v @@ -0,0 +1,7 @@ +`include "core_v6.v" +module core_v6_8r64(input clk, input rst, input run, input prog_we, input [9:0] prog_addr, input [40-1:0] prog_data, + input cfg_en, input [9:0] cfg_n, input [63:0] cfg_sel, input [31:0] cfg_m, input [4:0] cfg_r, input [31:0] cfg_wm, input [31:0] cfg_off, input [31:0] cfg_mask, + input [31:0] ld_val, output [31:0] addr, output [31:0] out); + core_v6 #(.LANES(8), .LOG_LANES(3), .REGS(64), .LOG_REGS(6), .IW(40)) c(.clk(clk), .rst(rst), .run(run), .prog_we(prog_we), .prog_addr(prog_addr), .prog_data(prog_data), + .cfg_en(cfg_en), .cfg_n(cfg_n), .cfg_sel(cfg_sel), .cfg_m(cfg_m), .cfg_r(cfg_r), .cfg_wm(cfg_wm), .cfg_off(cfg_off), .cfg_mask(cfg_mask), .ld_val(ld_val), .addr(addr), .out(out)); +endmodule diff --git a/tools/chip-model/rtl/rtl/core_v6_8sel.v b/tools/chip-model/rtl/rtl/core_v6_8sel.v new file mode 100644 index 000000000..b769156c7 --- /dev/null +++ b/tools/chip-model/rtl/rtl/core_v6_8sel.v @@ -0,0 +1,7 @@ +`include "core_v6.v" +module core_v6_8sel(input clk, input rst, input run, input prog_we, input [9:0] prog_addr, input [32-1:0] prog_data, + input cfg_en, input [9:0] cfg_n, input [63:0] cfg_sel, input [31:0] cfg_m, input [4:0] cfg_r, input [31:0] cfg_wm, input [31:0] cfg_off, input [31:0] cfg_mask, + input [31:0] ld_val, output [31:0] addr, output [31:0] out); + core_v6 #(.LANES(8), .LOG_LANES(3), .SELTREE(1)) c(.clk(clk), .rst(rst), .run(run), .prog_we(prog_we), .prog_addr(prog_addr), .prog_data(prog_data), + .cfg_en(cfg_en), .cfg_n(cfg_n), .cfg_sel(cfg_sel), .cfg_m(cfg_m), .cfg_r(cfg_r), .cfg_wm(cfg_wm), .cfg_off(cfg_off), .cfg_mask(cfg_mask), .ld_val(ld_val), .addr(addr), .out(out)); +endmodule diff --git a/tools/chip-model/rtl/tb/tb_core_common.vh b/tools/chip-model/rtl/tb/tb_core_common.vh index c1dac7f2c..38e969897 100644 --- a/tools/chip-model/rtl/tb/tb_core_common.vh +++ b/tools/chip-model/rtl/tb/tb_core_common.vh @@ -1,10 +1,10 @@ // shared body for the core testbenches: `TOP and `LANES set by the including file `timescale 1ps/1ps module tb; - reg clk = 0, rst = 1, run = 0, prog_we = 0, cfg_en = 0; reg [7:0] prog_addr = 0, cfg_n = 0; reg [31:0] prog_data = 0, cfg_m = 0, cfg_wm = 0, cfg_off = 0, cfg_mask = 0, ld_val = 0; reg [4:0] cfg_r = 0; + reg clk = 0, rst = 1, run = 0, prog_we = 0, cfg_en = 0; reg [9:0] prog_addr = 0, cfg_n = 0; reg [`IW-1:0] prog_data = 0; reg [63:0] cfg_sel = 0; reg [31:0] cfg_m = 0, cfg_wm = 0, cfg_off = 0, cfg_mask = 0, ld_val = 0; reg [4:0] cfg_r = 0; wire [31:0] addr, out; - `TOP dut(.clk(clk), .rst(rst), .run(run), .prog_we(prog_we), .prog_addr(prog_addr), .prog_data(prog_data), .cfg_en(cfg_en), .cfg_n(cfg_n), .cfg_m(cfg_m), .cfg_r(cfg_r), .cfg_wm(cfg_wm), .cfg_off(cfg_off), .cfg_mask(cfg_mask), .ld_val(ld_val), .addr(addr), .out(out)); - integer n, cycles, loads, k, roll; reg [31:0] acc = 0; reg [3:0] opc; reg [31:0] w; + `TOP dut(.clk(clk), .rst(rst), .run(run), .prog_we(prog_we), .prog_addr(prog_addr), .prog_data(prog_data), .cfg_en(cfg_en), .cfg_n(cfg_n), .cfg_sel(cfg_sel), .cfg_m(cfg_m), .cfg_r(cfg_r), .cfg_wm(cfg_wm), .cfg_off(cfg_off), .cfg_mask(cfg_mask), .ld_val(ld_val), .addr(addr), .out(out)); + integer n, cycles, loads, k, roll; reg [31:0] acc = 0; reg [3:0] opc; reg [63:0] w; // the class v4 draw: add 12, xor 10, mul 8, mad 8, shfl 8, rotl 7, sub 6, mulhi 6, rotr 6, or 4 (sum 75) function [3:0] draw_op; input integer r; integer x; begin x = r % 75; if (x < 0) x = -x; @@ -18,12 +18,12 @@ module tb; $dumpfile("dump.vcd"); $dumpvars(0, tb.dut); repeat (4) @(negedge clk); rst = 0; // the era draw: constants and the program length - @(negedge clk); cfg_en = 1; cfg_n = 8'd255; cfg_m = 32'h9e3779b1; cfg_r = 5'd13; cfg_wm = 32'h0fffffc0; cfg_off = 3; cfg_mask = 32'h0fffffff; + @(negedge clk); cfg_en = 1; cfg_n = `NPROG - 1; cfg_sel = `SEL; cfg_m = 32'h9e3779b1; cfg_r = 5'd13; cfg_wm = 32'h0fffffc0; cfg_off = 3; cfg_mask = 32'h0fffffff; @(negedge clk); cfg_en = 0; - // the program: 256 instructions drawn with the class v4 weights - for (k = 0; k < 256; k = k + 1) begin - @(negedge clk); w = $random; opc = (loads && (k % 16 == 15)) ? 4'd12 : draw_op($random); - prog_we = 1; prog_addr = k; prog_data = {w[31:4], opc}; + // the program: `NPROG instructions drawn with the class v4 weights + for (k = 0; k < `NPROG; k = k + 1) begin + @(negedge clk); w = {$random, $random}; opc = (loads && (k % 16 == 15)) ? 4'd12 : draw_op($random); + prog_we = 1; prog_addr = k; prog_data = {w[`IW-1:4], opc}; end @(negedge clk); prog_we = 0; run = 1; for (n = 0; n < cycles; n = n + 1) begin diff --git a/tools/chip-model/rtl/tb/tb_core_legacy.vh b/tools/chip-model/rtl/tb/tb_core_legacy.vh new file mode 100644 index 000000000..4fecabfee --- /dev/null +++ b/tools/chip-model/rtl/tb/tb_core_legacy.vh @@ -0,0 +1,34 @@ +// shared body for the core testbenches: `TOP and `LANES set by the including file +`timescale 1ps/1ps +module tb; + reg clk = 0, rst = 1, run = 0, prog_we = 0, cfg_en = 0; reg [7:0] prog_addr = 0, cfg_n = 0; reg [31:0] prog_data = 0, cfg_m = 0, cfg_wm = 0, cfg_off = 0, cfg_mask = 0, ld_val = 0; reg [4:0] cfg_r = 0; + wire [31:0] addr, out; + `TOP dut(.clk(clk), .rst(rst), .run(run), .prog_we(prog_we), .prog_addr(prog_addr), .prog_data(prog_data), .cfg_en(cfg_en), .cfg_n(cfg_n), .cfg_m(cfg_m), .cfg_r(cfg_r), .cfg_wm(cfg_wm), .cfg_off(cfg_off), .cfg_mask(cfg_mask), .ld_val(ld_val), .addr(addr), .out(out)); + integer n, cycles, loads, k, roll; reg [31:0] acc = 0; reg [3:0] opc; reg [31:0] w; + // the class v4 draw: add 12, xor 10, mul 8, mad 8, shfl 8, rotl 7, sub 6, mulhi 6, rotr 6, or 4 (sum 75) + function [3:0] draw_op; input integer r; integer x; + begin x = r % 75; if (x < 0) x = -x; + draw_op = (x < 12) ? 4'd0 : (x < 22) ? 4'd2 : (x < 30) ? 4'd6 : (x < 38) ? 4'd8 : (x < 46) ? 4'd9 : (x < 53) ? 4'd4 : (x < 59) ? 4'd1 : (x < 65) ? 4'd7 : (x < 71) ? 4'd5 : 4'd3; + end + endfunction + always #`HALF clk = ~clk; + initial begin + if (!$value$plusargs("cycles=%d", cycles)) cycles = 300; + if (!$value$plusargs("loads=%d", loads)) loads = 0; // 1: one load in 16 instructions + $dumpfile("dump.vcd"); $dumpvars(0, tb.dut); + repeat (4) @(negedge clk); rst = 0; + // the era draw: constants and the program length + @(negedge clk); cfg_en = 1; cfg_n = 8'd255; cfg_m = 32'h9e3779b1; cfg_r = 5'd13; cfg_wm = 32'h0fffffc0; cfg_off = 3; cfg_mask = 32'h0fffffff; + @(negedge clk); cfg_en = 0; + // the program: 256 instructions drawn with the class v4 weights + for (k = 0; k < 256; k = k + 1) begin + @(negedge clk); w = $random; opc = (loads && (k % 16 == 15)) ? 4'd12 : draw_op($random); + prog_we = 1; prog_addr = k; prog_data = {w[31:4], opc}; + end + @(negedge clk); prog_we = 0; run = 1; + for (n = 0; n < cycles; n = n + 1) begin + @(negedge clk); ld_val = $random; acc = acc ^ out ^ addr; + end + $display("CHECKSUM %08x", acc); $finish; + end +endmodule diff --git a/tools/chip-model/rtl/tb/tb_core_v6_32.v b/tools/chip-model/rtl/tb/tb_core_v6_32.v index 032d82423..2c63f5096 100644 --- a/tools/chip-model/rtl/tb/tb_core_v6_32.v +++ b/tools/chip-model/rtl/tb/tb_core_v6_32.v @@ -1,3 +1,3 @@ `define TOP core_v6_32 `define HALF 750 -`include "tb_core_common.vh" +`include "tb_core_legacy.vh" diff --git a/tools/chip-model/rtl/tb/tb_core_v6_32all.v b/tools/chip-model/rtl/tb/tb_core_v6_32all.v new file mode 100644 index 000000000..9509ee182 --- /dev/null +++ b/tools/chip-model/rtl/tb/tb_core_v6_32all.v @@ -0,0 +1,6 @@ +`define TOP core_v6_32all +`define HALF 750 +`define IW 40 +`define NPROG 1024 +`define SEL 64'hc5e7092b4d6f81a3 +`include "tb_core_common.vh" diff --git a/tools/chip-model/rtl/tb/tb_core_v6_32r16.v b/tools/chip-model/rtl/tb/tb_core_v6_32r16.v index f2084f017..696d54346 100644 --- a/tools/chip-model/rtl/tb/tb_core_v6_32r16.v +++ b/tools/chip-model/rtl/tb/tb_core_v6_32r16.v @@ -1,3 +1,6 @@ `define TOP core_v6_32r16 `define HALF 750 +`define IW 32 +`define NPROG 256 +`define SEL 64'hfedcba9876543210 `include "tb_core_common.vh" diff --git a/tools/chip-model/rtl/tb/tb_core_v6_8.v b/tools/chip-model/rtl/tb/tb_core_v6_8.v index 367b83e59..b45dfef48 100644 --- a/tools/chip-model/rtl/tb/tb_core_v6_8.v +++ b/tools/chip-model/rtl/tb/tb_core_v6_8.v @@ -1,3 +1,6 @@ `define TOP core_v6_8 `define HALF 750 +`define IW 32 +`define NPROG 256 +`define SEL 64'hfedcba9876543210 `include "tb_core_common.vh" diff --git a/tools/chip-model/rtl/tb/tb_core_v6_8i1k.v b/tools/chip-model/rtl/tb/tb_core_v6_8i1k.v new file mode 100644 index 000000000..738a4815b --- /dev/null +++ b/tools/chip-model/rtl/tb/tb_core_v6_8i1k.v @@ -0,0 +1,6 @@ +`define TOP core_v6_8i1k +`define HALF 750 +`define IW 32 +`define NPROG 1024 +`define SEL 64'hfedcba9876543210 +`include "tb_core_common.vh" diff --git a/tools/chip-model/rtl/tb/tb_core_v6_8r64.v b/tools/chip-model/rtl/tb/tb_core_v6_8r64.v new file mode 100644 index 000000000..5efbee6f6 --- /dev/null +++ b/tools/chip-model/rtl/tb/tb_core_v6_8r64.v @@ -0,0 +1,6 @@ +`define TOP core_v6_8r64 +`define HALF 750 +`define IW 40 +`define NPROG 256 +`define SEL 64'hfedcba9876543210 +`include "tb_core_common.vh" diff --git a/tools/chip-model/rtl/tb/tb_core_v6_8sel.v b/tools/chip-model/rtl/tb/tb_core_v6_8sel.v new file mode 100644 index 000000000..65094edc8 --- /dev/null +++ b/tools/chip-model/rtl/tb/tb_core_v6_8sel.v @@ -0,0 +1,6 @@ +`define TOP core_v6_8sel +`define HALF 750 +`define IW 32 +`define NPROG 256 +`define SEL 64'hc5e7092b4d6f81a3 +`include "tb_core_common.vh" From f5e1045318b9b3892e76ab429e9d85e3f658c5cf Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Thu, 8 Oct 2026 13:53:42 +0000 Subject: [PATCH 7/9] shadow-k: the node column (node-for-node the base core is k 0.78 and the GDDR7 board 2.4x at the lock) Co-Authored-By: Claude Fable 5.1 Documents-only replay of 05275ca6a (bfccc26ed4d299ad35920bda4aa2af0cc5efbe01) for the box mirror master --- docs/analysis/class-v6/floor/shadow-k.md | 18 ++++++++++++++++++ 1 file changed, 18 insertions(+) diff --git a/docs/analysis/class-v6/floor/shadow-k.md b/docs/analysis/class-v6/floor/shadow-k.md index 966f3df64..392090f12 100644 --- a/docs/analysis/class-v6/floor/shadow-k.md +++ b/docs/analysis/class-v6/floor/shadow-k.md @@ -185,6 +185,24 @@ under 2x at the lock needs the window, the long block and the knee together and stays unamortised, which it does not. The GPU pays nothing for the window until occupancy binds, 1.58x per hash for the long block (0.40x per forcing instruction), and nothing for the select tree. +### 4b. The node column: which part of the edge is the node (the coordinator's order, 15:0x UK) + +The record compares a chip scaled to N3 or N2 against a TSMC 4N card (N5 class; the M5 Max is N3), so part of +the edge is the node. Factors as in 2.3, every one claimed. Absolute; the GDDR7 board at the lock = 2.33 / +(0.466 + 102,100 x e_chip). + +| Core | pJ per lane-op ASAP7 / N5 / N3 / N2 | k at the lock, N5 / N3 / N2 | GDDR7 board at the lock, N5 / N3 / N2 | +|---|---|---|---| +| base (32 registers, 256 imem) | 6.9 / 4.8 / 3.5 / 2.5 | 0.78 / 0.56 / 0.40 | 2.4x / 2.8x / 3.2x | +| 64-register file | 9.7 / 6.8 / 4.9 / 3.5 | 1.09 / 0.78 / 0.56 | 2.0x / 2.4x / 2.8x | +| all four together | ROW_CORE32ALL_NODE | + +One line: of the 2.8x at k 0.56, the N5-to-N3 node step is worth 0.4x (2.4x node-for-node, a factor of 1.17, +claimed); the rest is the memory system (3.6x at zero shadow at the lock, modelled) less what the class v4 shadow +takes back on the card's own node (3.6x to 2.4x), which is the design. Node-for-node the base core already sits at +k 0.78 and the 64-register core at 1.09, so "near 0.9" is reached node-for-node by the window alone; what it does +not survive is the node step a chip project would buy (an N3 core gives back 0.4x, an N2 core 0.8x). + ## 5. The chip edge at the measured k `E_chip = E_mem + N_ops x e_chip` (absolute: the chip's shadow cost is 102,100 x 3.5 pJ = 0.36 microjoules per hash From b039ded1948807a548857208b238ffd0c736548a Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Thu, 8 Oct 2026 13:55:35 +0000 Subject: [PATCH 8/9] shadow-k: row (2'), the 32-lane 16-register core (4.2 pJ per lane-op ASAP7, k 0.34 at the lock at N3) Co-Authored-By: Claude Fable 5.1 Documents-only replay of 3e2a202fe (bfccc26ed4d299ad35920bda4aa2af0cc5efbe01) for the box mirror master --- docs/analysis/class-v6/floor/shadow-k.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/docs/analysis/class-v6/floor/shadow-k.md b/docs/analysis/class-v6/floor/shadow-k.md index 392090f12..1bc1591e0 100644 --- a/docs/analysis/class-v6/floor/shadow-k.md +++ b/docs/analysis/class-v6/floor/shadow-k.md @@ -140,7 +140,7 @@ draws 13.9 mW, the run phase 36.9 mW for 8 lanes at 1.5 ns). | of which the units, the read muxes and the butterfly | | | 4.5 | 3.1 | 2.3 | 1.6 | | | | | | | core, 8 lanes, 32 registers | placed and routed, SPEF | ROW_CORE8_PLACED | | core, 32 lanes, 32 registers | ROW_CORE32 | -| core, 32 lanes, 16 registers | ROW_CORE32R16 | +| core, 32 lanes, 16 registers | synthesis only | 443,258 | 4.2 | 2.9 | 2.1 | 1.5 | 11.3 / 6.2 / 6.9 | 0.18 / 0.34 / 0.30 | 0.13 / 0.24 / 0.22 | 0.68 | synthesised; one run length, about plus or minus 10 percent; 15:0x UK | | the bare ARX lane (section 3, the floor) | routed | 11,631 | 2.2 | 1.5 | 1.1 | 0.8 | 11.3 / 6.2 / 6.9 | 0.10 / 0.18 / 0.16 | 0.07 / 0.13 / 0.11 | 0.35 | the lower bound | Reading: fetch, decode, a 32-register file and the full unit set cost a chip 3.1x the bare lane, and the core's @@ -169,7 +169,7 @@ half a day; the select tree is a new instruction kind), knob 2 has no GPU side ( | (3) the same with the imem as a 4 KB SRAM macro shared by the lanes (2 to 4 pJ per 32-bit read, approximate) | | about 7.2 | about 1.9 | 5.0 | 3.6 | 2.6 | about 0.32 / 0.58 / 0.52 | the same | | (4) the drawn select tree (the era's 16-entry op permutation ahead of decode; every unit evaluated every cycle, as in the base) | 186,870 | 6.85 | 2.4 | 4.8 | 3.4 | 2.5 | 0.30 / 0.55 / 0.50 | the units' microbench sum (approximate) | | (2) 32 lanes, 32 registers (the butterfly across 32; the imem amortised over 32) | ROW_CORE32 | -| (2') 32 lanes, 16 registers (the register-file sensitivity the other way) | ROW_CORE32R16 | +| (2') 32 lanes, 16 registers (the register-file sensitivity the other way; one run length of 150 cycles, the load phase subtracted at the 8-lane ratio, about plus or minus 10 percent) | 443,258 | 4.2 | 2.3 | 2.9 | 2.1 | 1.5 | 0.18 / 0.34 / 0.30 | no GPU knob | | (5) all four together (32 lanes, 64 registers, 1,024 imem, the select tree) | ROW_CORE32ALL | Reading, for the founder's "under 2x at the lock" (which needs k near 0.9 on the GDDR7 board): the 64-register From e0a81517b532ca6a5a87b46051ba0775af6cf941 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Thu, 8 Oct 2026 14:28:57 +0000 Subject: [PATCH 9/9] shadow-k: the 32-lane core (5.55 pJ per lane-op, k 0.45 at the lock), the shuffle row (1.24 pJ, k 0.021), the mix optimiser's best corner Co-Authored-By: Claude Fable 5.1 Documents-only replay of bfccc26ed (bfccc26ed4d299ad35920bda4aa2af0cc5efbe01) for the box mirror master --- docs/analysis/class-v6/floor/shadow-k.md | 19 ++++++++++--------- 1 file changed, 10 insertions(+), 9 deletions(-) diff --git a/docs/analysis/class-v6/floor/shadow-k.md b/docs/analysis/class-v6/floor/shadow-k.md index 1bc1591e0..a5c561c52 100644 --- a/docs/analysis/class-v6/floor/shadow-k.md +++ b/docs/analysis/class-v6/floor/shadow-k.md @@ -102,7 +102,7 @@ stock and at the 1,300 MHz lock, and 6.9 pJ per counted op on the Apple M5 Max ( | index fold (x M, rotl R, masks; era constants in registers) | 14,965 | 2.37 | 1.66 | 1.19 | 0.86 | 13.9 / 8.3 (read against int_mul) | 0.086 / 0.14 | | 0.29 | | | prmt (byte permute) | 6,744 | 1.27 | 0.89 | 0.64 | 0.46 | 22.3 / 11.5 | 0.029 / 0.056 | | 0.11 | | | lop3 (8-bit truth table) | 7,274 | 1.42 | 0.99 | 0.72 | 0.52 | 24.1 / 13.0 | 0.030 / 0.055 | | 0.11 | | -| 32-lane xor-mask shuffle (butterfly over the 1 KB window), per lane-op | ROW_SHFL | +| 32-lane xor-mask shuffle (butterfly over the 1 KB window), per lane-op | 181,580 | 1.24 | 0.87 | 0.63 | 0.45 | 55.8 / 29.4 | 0.011 / 0.021 | | 0.042 | routed with SPEF; 15:1x UK | | 32-lane general crossbar, per lane-op | ROW_XBAR | | 8 KB scratch, one random read (flop array: the pessimistic form) | ROW_SCRATCH | | int8 8x8x8 tile, per MAC | ROW_TILE | @@ -139,7 +139,7 @@ draws 13.9 mW, the run phase 36.9 mW for 8 lanes at 1.5 ns). | of which the sequential term (register file, imem and IR clock pins, no clock gating) | | | 2.4 | 1.7 | 1.2 | 0.9 | | | | | | | of which the units, the read muxes and the butterfly | | | 4.5 | 3.1 | 2.3 | 1.6 | | | | | | | core, 8 lanes, 32 registers | placed and routed, SPEF | ROW_CORE8_PLACED | -| core, 32 lanes, 32 registers | ROW_CORE32 | +| core, 32 lanes, 32 registers | synthesis only (steady state from 150 and 400 run cycles) | 600,381 | 5.55 | 3.9 | 2.8 | 2.0 | 11.3 / 6.2 / 6.9 | 0.25 / 0.45 / 0.41 | 0.18 / 0.32 / 0.29 | 0.90 | synthesised; 15:2x UK | | core, 32 lanes, 16 registers | synthesis only | 443,258 | 4.2 | 2.9 | 2.1 | 1.5 | 11.3 / 6.2 / 6.9 | 0.18 / 0.34 / 0.30 | 0.13 / 0.24 / 0.22 | 0.68 | synthesised; one run length, about plus or minus 10 percent; 15:0x UK | | the bare ARX lane (section 3, the floor) | routed | 11,631 | 2.2 | 1.5 | 1.1 | 0.8 | 11.3 / 6.2 / 6.9 | 0.10 / 0.18 / 0.16 | 0.07 / 0.13 / 0.11 | 0.35 | the lower bound | @@ -168,7 +168,7 @@ half a day; the select tree is a new instruction kind), knob 2 has no GPU side ( | (3) 1,024-entry imem, the program drawn at 1,024, as built (a flop array) | 324,543 | 12.0 | 6.0 | 8.4 | 6.1 | 4.4 | 0.53 / 0.97 / 0.87 | measured: the 1,024 block at 27 passes costs the 5090 1.58x the energy per hash (4.96 against 3.14 microjoules at stock, 111 against 141 MH/s at the 575 W cap) and the 4090 1.57x (7.01 against 4.48, the rate held at 62.5 MH/s), for 4x the shadow instructions: 0.40x per instruction | | (3) the same with the imem as a 4 KB SRAM macro shared by the lanes (2 to 4 pJ per 32-bit read, approximate) | | about 7.2 | about 1.9 | 5.0 | 3.6 | 2.6 | about 0.32 / 0.58 / 0.52 | the same | | (4) the drawn select tree (the era's 16-entry op permutation ahead of decode; every unit evaluated every cycle, as in the base) | 186,870 | 6.85 | 2.4 | 4.8 | 3.4 | 2.5 | 0.30 / 0.55 / 0.50 | the units' microbench sum (approximate) | -| (2) 32 lanes, 32 registers (the butterfly across 32; the imem amortised over 32) | ROW_CORE32 | +| (2) 32 lanes, 32 registers (the butterfly across 32; the imem amortised over 32) | 600,381 | 5.55 | 1.5 | 3.9 | 2.8 | 2.0 | 0.25 / 0.45 / 0.41 | no GPU knob | | (2') 32 lanes, 16 registers (the register-file sensitivity the other way; one run length of 150 cycles, the load phase subtracted at the 8-lane ratio, about plus or minus 10 percent) | 443,258 | 4.2 | 2.3 | 2.9 | 2.1 | 1.5 | 0.18 / 0.34 / 0.30 | no GPU knob | | (5) all four together (32 lanes, 64 registers, 1,024 imem, the select tree) | ROW_CORE32ALL | @@ -269,7 +269,7 @@ to 2.3x and the strongest chips at 2.6x to 4.2x. | 5 | mul | 0.68 | 8.3 | 0.082 | yes (8) | | 6 | prmt, lop3 | 0.64, 0.72 | 11.5, 13.0 | 0.056, 0.055 | not drawn (RTL rows only) | | 7 | mulhi | 0.68 | 21.0 | 0.032 | yes (6) | -| 8 | 32-lane shuffle (butterfly) | ROW_SHFL_K | 29.4 | pending | yes (8) | +| 8 | 32-lane shuffle (butterfly) | 0.63 | 29.4 | 0.021 | yes (8) | | 9 | L1 scratch read (8 KB flop array) against the card's L2 hit | ROW_SCRATCH_K | 1,400 | pending | not drawn | | 10 | int8 8x8x8 tile, per MAC | ROW_TILE_K | 2.2 | pending | not drawn (the tensor lever is dead on other grounds) | @@ -288,13 +288,14 @@ xor, mad, shfl, rotl, rotr), the lossy families (or, mul, mulhi) at or under the not functions of the weights and do not move. `tools/chip-model/rtl/flow/mix.py` searches the band exhaustively over the corners (the objective is a ratio of linear forms, so the optimum is at a corner of each family's range). -On the unit floors (N3, the lock; the shuffle row pending, so computed over the other 67 points of 75): +On the unit floors (N3, the lock; the shuffle row in at 0.63 pJ, its k 0.021 the lowest of the drawn families): | Mix | add, xor, mul, mad, shfl, rotl, sub, mulhi, rotr, or | k_eff (floors) | The card's pJ per op at the lock | k_eff on the core (floor + the core's 2.4 pJ per op overhead at N3) | |---|---|---|---|---| -| class v4 (the base) | 12, 10, 8, 8, 8, 7, 6, 6, 6, 4 | 0.130 | 8.0 | 0.43 | -| the census lane's re-weight (d20eb04bd, PASS on 256 seeds, no era and eras 0 to 7, W = 4 and W = 16) | 13, 11, 6, 10, 8, 8, 7, 2, 6, 4 | 0.153 (+18 percent) | 7.1 | 0.49 (+14 percent) | -| the band's corner (every injecting family at +4 except the shuffle at its floor; lossy at base minus 4) | 16, 14, 4, 12, 4, 11, 10, 2, 10, 0 | MIX_BEST | | | +| class v4 (the base) | 12, 10, 8, 8, 8, 7, 6, 6, 6, 4 | 0.097 (with the shuffle row in: 0.130 over the other 67 points) | 10.3 | 0.43 | +| the census lane's re-weight (d20eb04bd, PASS on 256 seeds, no era and eras 0 to 7, W = 4 and W = 16) | 13, 11, 6, 10, 8, 8, 7, 2, 6, 4 | 0.115 (+19 percent; 0.153 over the 67 points) | 9.0 | 0.49 (+14 percent) | +| the band's best corner (mix.py, exhaustive: every injecting family at +4, the shuffle at its floor, the lossy families at base minus 4) | 16, 14, 4, 12, 4, 11, 10, 2, 10, 0 | 0.137 (+42 percent) | 8.2 | about 0.51 (+19 percent) | +| the band's worst corner (shuffle and mulhi heavy) | 8, 6, 4, 4, 8, 3, 2, 6, 2, 0 | 0.074 | 12.3 | about 0.40 | Reading: the re-weight moves six points off the chip's two easiest families (mulhi, mul) onto the hardest (mad, the ARX families) and lifts the expected k by about a seventh on the core; its acceptance is the census lane's @@ -302,7 +303,7 @@ the ARX families) and lifts the expected k by about a seventh on the core; its a premium per instruction 11 percent LOWER (fewer mulhi), so at a fixed premium the program is 12 percent longer. The shuffle's own row decides whether it goes to its floor: on the card it is the dearest op (29.4 pJ at the lock) and on the chip a butterfly over the window, so its k is the lowest of the drawn families unless the synthesis says -otherwise. The recommendation stands as the census lane's draw until the shuffle row lands (an amendment). +otherwise. The recommendation: the band's best corner (add 16, xor 14, mad 12, rotl 11, sub 10, rotr 10, shfl 4, mul 4, mulhi 2, or 0) if its census passes (the census lane's harness, about five box-minutes per candidate); the census lane's draw as the passed fallback. Either way the shuffle goes to its floor: the card pays 29.4 pJ for a move the chip does for 0.6. ## 7. Sources