diff --git a/docs/analysis/class-v6/floor/shadow-k.md b/docs/analysis/class-v6/floor/shadow-k.md new file mode 100644 index 000000000..a5c561c52 --- /dev/null +++ b/docs/analysis/class-v6/floor/shadow-k.md @@ -0,0 +1,324 @@ +# The shadow's k from RTL (floor lane 2, 8 October 2026) + +Branch `class-v6-floor-k` from `counter-asic-4` at 7618e729. Every chip-side number in this file comes from +synthesis and place-and-route of RTL written for this lane (Yosys 0.68 plus OpenROAD, the ORFS docker image +`openroad/orfs:latest`, the ASAP7 predictive PDK, run on igneum-build-4), with switching activity from a +random-input gate-level simulation (iverilog on the routed netlist). Every GPU-side number is the record's +measurement (`docs/analysis/counter-asic-4-research.md` 15.1a: PC 1, the RTX 5090, 20 probes at 60 s, 8 +October 2026) and is not re-estimated here. The RTL, testbenches, flow configs and a Makefile that reproduces +every row are under `tools/chip-model/rtl/`. + +## 0. One page + +(filled from the rows below; see section 2 for the table and section 4 for the edge) + +## 1. The question and the identity + +The chip's energy per hash is `E_chip = E_mem + k x F` (research file section 2): `E_mem` the memory system's +reads, controller and static (0.466 microjoules per hash on the record's GDDR7 board, modelled), `F` the shadow's +premium on the card (measured: 1.10 microjoules unlocked, 0.652 at the 1,300 MHz lock on the 5090 for class +v4's 102,100 counted ops per hash), and `k = e_chip / e_gpu` the chip core's energy per forced op over the +card's. The card's side of `k` is measured; the chip's side has been a claimed band (ALU 0.3 to 0.8, int8 tile +0.03 to 0.3, L2 hit 0.1 to 0.3, shuffle 0.4 to 0.7). This file replaces the claimed side with a synthesised one, +family by family, and reads the mix that maximises the expected `k` under the invention lane's sound-form rules. + +## 2. Method + +### 2.1 What was built (RTL, `tools/chip-model/rtl/rtl/`) + +Each family is a minimal chip-side shadow core: the smallest circuit a chip maker would have to build to run that +family's instructions bit-exactly at one op per cycle per lane. The common lane shape is an instruction register, +an 8 x 32-bit register window in flops (the card's lane holds its working set in a register file too; 8 registers +is the class program's window), two or three read ports through muxes, the functional unit, and one write port. +Nothing is shared across lanes and nothing is pipelined beyond one stage, so the figure is the datapath plus the +minimum operand delivery: a floor for the chip, which makes the `k` it gives a floor too. + +| Family | Module | What one op is | Ops per cycle | Clock set (ps) | +|---|---|---|---|---| +| ARX | `lane_arx` | int32 add, sub, xor, or, rotl by immediate, rotr by register (op drawn per cycle; rows per op by fixing the op field) | 1 | 1,000 | +| MUL | `lane_mul` | 32 x 32 to 64: mul (low word), mulhi (high word), mad (src x src2 + dst) | 1 | 1,500 | +| PRMT | `lane_prmt` | byte permute: 4 output bytes from the 8 bytes of two registers by a 16-bit selector, sign-replicate bit as on the card | 1 | 1,000 | +| LOP3 | `lane_lop3` | three-input logic by an 8-bit truth table, bitwise | 1 | 1,000 | +| FOLD | `lane_fold` | the index fold `((rotl(x * M, R) & WM) \| OFF) & MASK` with the era's constants held in registers | 1 | 1,500 | +| SHFL | `shfl32` | the 32-lane xor-mask shuffle `r[dst][lane] ^= r[src][lane ^ m]` over a 1 KB window (32 lanes x 8 x 32 bits): a 5-stage butterfly | 32 lane-ops | 1,000 | +| XBAR | `xbar32` | the general 32-lane crossbar over the same window (any source lane per lane, a 5-bit select each): the upper bound on a shuffle network | 32 lane-ops | 1,000 | +| SCRATCH | `scratch8k` | one random 32-bit read of a 2,048 x 32 (8 KB) flop array, a write on one cycle in eight: the chip's L1 as the pessimistic flop form | 1 read | 1,500 | +| TILE | `tile8` | the int8 8x8x8 tile `C += A x B` with int32 accumulators: 512 MACs per cycle | 512 MACs | 2,000 | + +### 2.2 The flow + +ORFS on ASAP7 (7.5-track RVT cells, the TC corner: 0.70 V, 0 C, NLDM), the default flow end to end: Yosys +synthesis with ABC, floorplan at 40 percent utilisation, global and detailed placement, CTS, global and detailed +routing, parasitic extraction (OpenRCX, the platform's rules). Power is OpenSTA's `report_power` on the routed +design with its SPEF, under two activities: (a) the VCD of a random-input simulation of the routed netlist +(ASAP7 has no cell Verilog models in the image, so Yosys builds the cell bodies from the liberty functions and +writes the netlist as primitives; iverilog runs 3,000 to 4,000 cycles with every instruction field drawn by +`$random` each cycle and a random 32-bit load into the window every 16th cycle), and (b) a propagated 0.5 +activity on every input as the cross-check. The energy per op is the total power (internal, switching and leakage) +times the clock period over the ops per cycle. Leakage is reported beside it; at these clocks it is a few percent. + +### 2.3 Node scaling (approximate; every factor claimed from the foundry's own headline) + +ASAP7 is a predictive 7 nm-class FinFET PDK (ASU and ARM, Clark et al., Microelectronics Journal 2016), not a +foundry node, so the row is first stated at ASAP7 and then scaled by the foundry's published per-node power +reductions at the same speed: N7 to N5 x0.70 (TSMC: "30 percent lower power"), N5 to N3E x0.72 (TSMC: "25 to 30 +percent lower power", the midpoint), N3E to N2 x0.72 (TSMC: "25 to 30 percent lower power", the midpoint). So +N5 = 0.70, N3 = 0.50 and N2 = 0.36 of the ASAP7 figure. These are the foundry's claims for a whole design at a +fixed frequency, and a shadow core at a low clock could run at a lower voltage still; the N2 column is therefore +the chip's best case from this method, not its floor. Sources in section 7. + +### 2.4 What the method leaves out, on both sides + +On the chip side the figure omits instruction fetch and decode (a chip would run the program from a small SRAM +or a decoded instruction cache shared by many lanes, about 1 to 2 pJ per lane-instruction amortised across 32 +lanes, approximate), the clock tree beyond the block's own, and the result's move to a memory address unit. On +the card side the 15.1a figure is the whole card's marginal per counted op (the sleep floor subtracted), which +includes the card's own fetch, decode, operand collection and register file. So the `k` here is the chip's +datapath-and-window cost over the card's whole-lane cost: a FLOOR on the chip's cost and so a floor on `k`. + +## 3. The rows: per-unit floors (the minimal lane per family) + +Every row: ASAP7 routed, SPEF, OpenSTA `report_power` under the random-input gate-level VCD (every pin annotated, +0 unannotated), the TC corner (0.70 V). "pJ/op" is total power (internal + switching + leakage) times the clock +period over the ops per cycle. The propagated-0.5 cross-check is in `table.csv`; it agrees within 2x on the +logic lanes and overestimates the multiplier lanes 50x (OpenSTA's statistical propagation through a multiplier +is not a measurement), so the VCD row is the row. The GPU side is 15.1a: pJ per counted op on the RTX 5090 at +stock and at the 1,300 MHz lock, and 6.9 pJ per counted op on the Apple M5 Max (the class v4 shadow, measured). +`k` is absolute: the chip's own pJ per op at the node over the card's at its point. + +| Family (chip RTL) | Cells (with fill) | pJ/op ASAP7 | pJ/op N5 | pJ/op N3 | pJ/op N2 | 5090 pJ/op stock / lock | k, N3 chip vs 5090 stock / lock | k, N3 vs M5 Max 6.9 | k, ASAP7 unscaled vs lock | Label | +|---|---|---|---|---|---|---|---|---|---|---| +| ARX add | 11,631 | 2.16 | 1.51 | 1.09 | 0.78 | 11.3 / 6.2 | 0.096 / 0.18 | 0.16 | 0.35 | synthesised; N3 and N2 scaled (claimed) | +| ARX sub | | 2.14 | 1.50 | 1.08 | 0.78 | 11.3 / 6.2 | 0.095 / 0.17 | 0.16 | 0.35 | | +| ARX xor | | 2.09 | 1.46 | 1.05 | 0.76 | 11.3 / 6.2 | 0.093 / 0.17 | 0.15 | 0.34 | | +| ARX rotl (immediate) | | 2.15 | 1.50 | 1.08 | 0.78 | 11.3 / 6.2 | 0.095 / 0.17 | 0.16 | 0.35 | | +| ARX rotr (by register) | | 2.13 | 1.49 | 1.07 | 0.77 | 11.3 / 6.2 | 0.095 / 0.17 | 0.16 | 0.34 | | +| or (lossy; the values saturate to ones and the activity falls) | | 1.45 | 1.01 | 0.73 | 0.53 | 11.3 / 6.2 | 0.065 / 0.12 | 0.11 | 0.23 | | +| ARX random mix | | 2.24 | 1.57 | 1.13 | 0.81 | 11.3 / 6.2 | 0.10 / 0.18 | 0.16 | 0.36 | | +| mul (32 x 32, low word) | 19,603 | 1.35 | 0.94 | 0.68 | 0.49 | 13.9 / 8.3 | 0.049 / 0.082 | | 0.16 | | +| mulhi (high word; the same multiplier) | | 1.35 | 0.94 | 0.68 | 0.49 | 39.6 / 21.0 | 0.017 / 0.032 | | 0.064 | | +| mad (src x src2 + dst; three reads) | | 3.29 | 2.30 | 1.66 | 1.19 | 13.9 / 8.3 | 0.12 / 0.20 | | 0.40 | | +| mul random mix | | 1.53 | 1.07 | 0.77 | 0.56 | 13.9 / 8.3 | 0.055 / 0.093 | | 0.18 | | +| index fold (x M, rotl R, masks; era constants in registers) | 14,965 | 2.37 | 1.66 | 1.19 | 0.86 | 13.9 / 8.3 (read against int_mul) | 0.086 / 0.14 | | 0.29 | | +| prmt (byte permute) | 6,744 | 1.27 | 0.89 | 0.64 | 0.46 | 22.3 / 11.5 | 0.029 / 0.056 | | 0.11 | | +| lop3 (8-bit truth table) | 7,274 | 1.42 | 0.99 | 0.72 | 0.52 | 24.1 / 13.0 | 0.030 / 0.055 | | 0.11 | | +| 32-lane xor-mask shuffle (butterfly over the 1 KB window), per lane-op | 181,580 | 1.24 | 0.87 | 0.63 | 0.45 | 55.8 / 29.4 | 0.011 / 0.021 | | 0.042 | routed with SPEF; 15:1x UK | +| 32-lane general crossbar, per lane-op | ROW_XBAR | +| 8 KB scratch, one random read (flop array: the pessimistic form) | ROW_SCRATCH | +| int8 8x8x8 tile, per MAC | ROW_TILE | + +Reading the floors: a lane's add costs the chip about 2.2 pJ at ASAP7 and 1.1 at N3, against the 5090's 6.2 +at the lock and 11.3 at stock, so even the floor is not a tenth of the card's cost at the knee, and the claimed +"2 to 5 pJ for a SIMD array at N5" (15.1) was the right order for the unscaled lane with nothing around it. The +multiplier is the cheapest unit per op relative to the card (the 5090 pays 8.3 pJ for mul and 21 for mulhi, the +chip 0.68 for either, because the high word falls out of the same array), and mad is the dearest for the chip +(three register reads and two units). The fold costs a chip one multiply and a rotate: 1.2 pJ per address at N3, +or 0.15 nJ per hash over 128 loads, a third of a percent of the GDDR7 board's 0.466 microjoules. + +## 4. The headline row: the programmable sequencer core + +The per-unit lanes of section 3 are floors for a chip that cannot exist under class v6: layers 1 and 3 (per-era +op-mix, program-length and read-width draws, family epochs) kill any fixed lane, so the chip that competes is a +programmable core. `core_v6` (`tools/chip-model/rtl/rtl/core_v6.v`) is the minimal in-order SIMD sequencer that +executes the drawn class v6 shadow program: a 256 x 32-bit instruction memory (a flop array) sized to the drawn +program, a program counter wrapping at the era's drawn length, fetch into an instruction register, decode, the +era's parameter registers (M, R, WM, OFF, MASK, N), and per lane a 32 x 32-bit register file (flops) with three +read ports and one write port and every class unit: add, sub, xor, or, rotl, rotr, mul, mulhi, mad, prmt, lop3, +the xor-mask shuffle across the lanes (a butterfly), and the load (the index fold on the address path, the returned +word written next cycle). One instruction per cycle for every lane. The program is 256 instructions drawn with the +class v4 weights (`NONLOAD_WEIGHTS`), with a variant at one load per 16 instructions. Two builds: 8 lanes (the fast +row) and 32 lanes (the imem amortised over 32), and a 32-lane build with a 16-register file (the sensitivity). + +The activity: the gate-level random-input VCD of the synthesised netlist; two run lengths (150 and 800 cycles) +bracket the 260-cycle program-load phase, and solving the pair gives the steady-state run power (the load phase +draws 13.9 mW, the run phase 36.9 mW for 8 lanes at 1.5 ns). + +| Row | Stage | Cells | pJ per lane-op ASAP7 | N5 | N3 | N2 | 5090 stock / lock / M5 Max pJ per op | k absolute at N3 vs stock / lock / M5 Max | k at N2 | k unscaled ASAP7 vs lock | Label | +|---|---|---|---|---|---|---|---|---|---|---|---| +| core, 8 lanes, 32 registers | synthesis only (no wires, no clock tree) | 186,443 | 6.9 | 4.8 | 3.5 | 2.5 | 11.3 / 6.2 / 6.9 | 0.31 / 0.56 / 0.50 | 0.22 / 0.40 / 0.36 | 1.1 | synthesised; 14:0x UK | +| of which the sequential term (register file, imem and IR clock pins, no clock gating) | | | 2.4 | 1.7 | 1.2 | 0.9 | | | | | | +| of which the units, the read muxes and the butterfly | | | 4.5 | 3.1 | 2.3 | 1.6 | | | | | | +| core, 8 lanes, 32 registers | placed and routed, SPEF | ROW_CORE8_PLACED | +| core, 32 lanes, 32 registers | synthesis only (steady state from 150 and 400 run cycles) | 600,381 | 5.55 | 3.9 | 2.8 | 2.0 | 11.3 / 6.2 / 6.9 | 0.25 / 0.45 / 0.41 | 0.18 / 0.32 / 0.29 | 0.90 | synthesised; 15:2x UK | +| core, 32 lanes, 16 registers | synthesis only | 443,258 | 4.2 | 2.9 | 2.1 | 1.5 | 11.3 / 6.2 / 6.9 | 0.18 / 0.34 / 0.30 | 0.13 / 0.24 / 0.22 | 0.68 | synthesised; one run length, about plus or minus 10 percent; 15:0x UK | +| the bare ARX lane (section 3, the floor) | routed | 11,631 | 2.2 | 1.5 | 1.1 | 0.8 | 11.3 / 6.2 / 6.9 | 0.10 / 0.18 / 0.16 | 0.07 / 0.13 / 0.11 | 0.35 | the lower bound | + +Reading: fetch, decode, a 32-register file and the full unit set cost a chip 3.1x the bare lane, and the core's +k at the 1,300 lock is 0.56 at N3 (0.40 at N2), inside the record's claimed 0.3 to 0.8 band and at its centre, with +the bare lane's 0.18 as the lower bound. Two corrections pull opposite ways: placement adds wires and a clock tree +(+20 to +40 percent on a design like this, approximate; the placed rows read it) and a chip maker gates the +register-file clock (one of 32 registers is written per cycle; gating removes about 2.0 of the 2.4 pJ sequential +term, approximate), so the net figure for the re-fold is 7.0 / 4.9 / 3.5 / 2.5 pJ per lane-op at ASAP7 / N5 / N3 / +N2 (the synthesis-only figure within the rounding). The imem is amortised over 8 lanes in this row and over 32 in +the core32 row. + +### 4a. The design sweep: what a class could add to the core's cost (the coordinator's order, 14:3x UK; rows 14:5x) + +Synthesis-only (no wires, no clock tree), 8 lanes unless stated, the same corner and scaling; the steady-state run +power solved from two run lengths (150 and 600 cycles after the program load). The GPU side per knob is the hash +lane's: knob 3 measured on a rented 5090 and 4090 (RunPod, 14:28 to 14:39 UK), knobs 1 and 4 modelled until a +generator line exists (a 64-entry window is a new ISA: an init rule and a fold rule for the extra registers, about +half a day; the select tree is a new instruction kind), knob 2 has no GPU side (the warp's shuffle already spans +32 lanes). + +| Variant | Cells | pJ per lane-op ASAP7 | of which clocking (RF, imem, IR; no gating) | N5 | N3 | N2 | k at N3 vs 5090 stock / lock / M5 Max | GPU side | +|---|---|---|---|---|---|---|---|---| +| base: 32 registers, 256-entry imem (the headline row) | 186,443 | 6.9 | 2.4 | 4.8 | 3.5 | 2.5 | 0.31 / 0.56 / 0.50 | measured (class v4) | +| (1) 64-register file (a 40-bit instruction word) | 267,731 | 9.7 | 3.75 | 6.8 | 4.9 | 3.5 | 0.43 / 0.78 / 0.70 | new ISA; 64 live registers takes a 5090 or 4090 thread to about 110 of 255, occupancy to about half; under the latency-bound chain the rate is expected to hold and the energy to move little (the hash lane, modelled); the unmeasured term is the per-lane register traffic | +| (3) 1,024-entry imem, the program drawn at 1,024, as built (a flop array) | 324,543 | 12.0 | 6.0 | 8.4 | 6.1 | 4.4 | 0.53 / 0.97 / 0.87 | measured: the 1,024 block at 27 passes costs the 5090 1.58x the energy per hash (4.96 against 3.14 microjoules at stock, 111 against 141 MH/s at the 575 W cap) and the 4090 1.57x (7.01 against 4.48, the rate held at 62.5 MH/s), for 4x the shadow instructions: 0.40x per instruction | +| (3) the same with the imem as a 4 KB SRAM macro shared by the lanes (2 to 4 pJ per 32-bit read, approximate) | | about 7.2 | about 1.9 | 5.0 | 3.6 | 2.6 | about 0.32 / 0.58 / 0.52 | the same | +| (4) the drawn select tree (the era's 16-entry op permutation ahead of decode; every unit evaluated every cycle, as in the base) | 186,870 | 6.85 | 2.4 | 4.8 | 3.4 | 2.5 | 0.30 / 0.55 / 0.50 | the units' microbench sum (approximate) | +| (2) 32 lanes, 32 registers (the butterfly across 32; the imem amortised over 32) | 600,381 | 5.55 | 1.5 | 3.9 | 2.8 | 2.0 | 0.25 / 0.45 / 0.41 | no GPU knob | +| (2') 32 lanes, 16 registers (the register-file sensitivity the other way; one run length of 150 cycles, the load phase subtracted at the 8-lane ratio, about plus or minus 10 percent) | 443,258 | 4.2 | 2.3 | 2.9 | 2.1 | 1.5 | 0.18 / 0.34 / 0.30 | no GPU knob | +| (5) all four together (32 lanes, 64 registers, 1,024 imem, the select tree) | ROW_CORE32ALL | + +Reading, for the founder's "under 2x at the lock" (which needs k near 0.9 on the GDDR7 board): the 64-register +window is the one robust knob, because its cost is per lane and a chip cannot share it (+2.8 pJ per lane-op at +ASAP7, +0.22 of k at the lock). The long block's cost is instruction memory, which a chip shares across its lanes +as SRAM, so most of its 0.97 as built is the flop array's clock and the honest figure is about 0.6; the card +meanwhile pays 1.58x the energy per hash for it (measured), so on the GDDR7 board at stock the chip reads +4.96 / (0.466 + 408,400 x 3.6 pJ) = 2.6x (1.7x on the flop-array row, which is not a chip anyone builds). The +select tree costs the chip nothing because every unit already evaluates every cycle in the base core. (1) + (3) +together reach about 10 pJ per lane-op at ASAP7 with the SRAM imem (5.0 at N3, k about 0.81 at the lock, 0.45 at +stock), 14.8 as built (k about 1.2); so k 0.85 is reached only on the flop-array reading, and the DRAM board +under 2x at the lock needs the window, the long block and the knee together and holds only while the chip's imem +stays unamortised, which it does not. The GPU pays nothing for the window until occupancy binds, 1.58x per hash for +the long block (0.40x per forcing instruction), and nothing for the select tree. + +### 4b. The node column: which part of the edge is the node (the coordinator's order, 15:0x UK) + +The record compares a chip scaled to N3 or N2 against a TSMC 4N card (N5 class; the M5 Max is N3), so part of +the edge is the node. Factors as in 2.3, every one claimed. Absolute; the GDDR7 board at the lock = 2.33 / +(0.466 + 102,100 x e_chip). + +| Core | pJ per lane-op ASAP7 / N5 / N3 / N2 | k at the lock, N5 / N3 / N2 | GDDR7 board at the lock, N5 / N3 / N2 | +|---|---|---|---| +| base (32 registers, 256 imem) | 6.9 / 4.8 / 3.5 / 2.5 | 0.78 / 0.56 / 0.40 | 2.4x / 2.8x / 3.2x | +| 64-register file | 9.7 / 6.8 / 4.9 / 3.5 | 1.09 / 0.78 / 0.56 | 2.0x / 2.4x / 2.8x | +| all four together | ROW_CORE32ALL_NODE | + +One line: of the 2.8x at k 0.56, the N5-to-N3 node step is worth 0.4x (2.4x node-for-node, a factor of 1.17, +claimed); the rest is the memory system (3.6x at zero shadow at the lock, modelled) less what the class v4 shadow +takes back on the card's own node (3.6x to 2.4x), which is the design. Node-for-node the base core already sits at +k 0.78 and the 64-register core at 1.09, so "near 0.9" is reached node-for-node by the window alone; what it does +not survive is the node step a chip project would buy (an N3 core gives back 0.4x, an N2 core 0.8x). + +## 5. The chip edge at the measured k + +`E_chip = E_mem + N_ops x e_chip` (absolute: the chip's shadow cost is 102,100 x 3.5 pJ = 0.36 microjoules per hash +at N3, 0.26 at N2, whatever the card does); the record's convention `E_mem + k x F` beside it with `k` read at the +card's point (it coincides at the lock by construction and is the same at stock here because both are the same +arithmetic on the same ops; it differs when the chip's cost is held fixed while the card's point moves, which is +what the SRAM lane found flattered the die 1.6x). Cards: the 5090 at stock (3.36 microjoules, F 1.10) and at the +lock (2.33, F 0.652), the M5 Max (1.40 on class v4, 0.78 on class v3, 6.9 pJ per op). Chips: the record's GDDR7 +board and the hardware-future lane's rows (their `E_mem` at zero shadow). The record's rows that this recomputes +are hardware-future.md section 5: "2.6x to 2.8x at k = 1 and 4.4x to 5.1x at k = 0.5" for the strongest DRAM chips +against the stock 5090 (the GDDR7 board's own row there is 2.1x and 3.3x). + +Chip shadow per hash, absolute: N3 0.357 microjoules, N2 0.255. +| Chip (E_mem, microjoules) | Card row | Zero shadow | Absolute, N3 core | Absolute, N2 core | Record convention at the measured k (N3) | at k = 1 | at k = 0.5 | +|---|---|---|---|---|---|---|---| +| GDDR7 board, 28 nm controller (the record) (0.466) | 5090 stock (class v4) | 4.8x | 4.1x | 4.7x | 4.1x (k 0.32) | 2.1x | 3.3x | +| GDDR7 board, 28 nm controller (the record) (0.466) | 5090 at the 1,300 MHz lock | 3.6x | 2.8x | 3.2x | 2.8x (k 0.55) | 2.1x | 2.9x | +| GDDR7 board, 28 nm controller (the record) (0.466) | M5 Max (class v4 1.40 microjoules, class v3 0.78; 6.9 pJ per op; F 0.62) | 1.7x | 1.7x | 1.9x | 1.8x (k 0.51) | 1.3x | 1.8x | +| HBM3E one stack (0.321) | 5090 stock (class v4) | 7.0x | 5.0x | 5.8x | 5.0x (k 0.32) | 2.4x | 3.9x | +| HBM3E one stack (0.321) | 5090 at the 1,300 MHz lock | 5.2x | 3.4x | 4.0x | 3.4x (k 0.55) | 2.4x | 3.6x | +| HBM3E one stack (0.321) | M5 Max (class v4 1.40 microjoules, class v3 0.78; 6.9 pJ per op; F 0.62) | 2.4x | 2.1x | 2.4x | 2.2x (k 0.51) | 1.5x | 2.2x | +| HBM4 one stack, N12 base die (0.22) | 5090 stock (class v4) | 10.3x | 5.8x | 7.1x | 5.8x (k 0.32) | 2.5x | 4.4x | +| HBM4 one stack, N12 base die (0.22) | 5090 at the 1,300 MHz lock | 7.6x | 4.0x | 4.9x | 4.0x (k 0.55) | 2.7x | 4.3x | +| HBM4 one stack, N12 base die (0.22) | M5 Max (class v4 1.40 microjoules, class v3 0.78; 6.9 pJ per op; F 0.62) | 3.5x | 2.4x | 2.9x | 2.6x (k 0.51) | 1.7x | 2.6x | +| custom HBM4E base die, N3P (0.18) | 5090 stock (class v4) | 12.6x | 6.3x | 7.7x | 6.3x (k 0.32) | 2.6x | 4.6x | +| custom HBM4E base die, N3P (0.18) | 5090 at the 1,300 MHz lock | 9.3x | 4.3x | 5.4x | 4.3x (k 0.55) | 2.8x | 4.6x | +| custom HBM4E base die, N3P (0.18) | M5 Max (class v4 1.40 microjoules, class v3 0.78; 6.9 pJ per op; F 0.62) | 4.3x | 2.6x | 3.2x | 2.8x (k 0.51) | 1.7x | 2.9x | +| DRAM on logic, hybrid bonded (2029 to 2031) (0.15) | 5090 stock (class v4) | 15.1x | 6.6x | 8.3x | 6.6x (k 0.32) | 2.7x | 4.8x | +| DRAM on logic, hybrid bonded (2029 to 2031) (0.15) | 5090 at the 1,300 MHz lock | 11.2x | 4.6x | 5.7x | 4.6x (k 0.55) | 2.9x | 4.9x | +| DRAM on logic, hybrid bonded (2029 to 2031) (0.15) | M5 Max (class v4 1.40 microjoules, class v3 0.78; 6.9 pJ per op; F 0.62) | 5.2x | 2.8x | 3.5x | 3.0x (k 0.51) | 1.8x | 3.0x | +| SRAM full store, one N2 reticle (0.14) | 5090 stock (class v4) | 16.1x | 6.8x | 8.5x | 6.8x (k 0.32) | 2.7x | 4.9x | +| SRAM full store, one N2 reticle (0.14) | 5090 at the 1,300 MHz lock | 12.0x | 4.7x | 5.9x | 4.7x (k 0.55) | 2.9x | 5.0x | +| SRAM full store, one N2 reticle (0.14) | M5 Max (class v4 1.40 microjoules, class v3 0.78; 6.9 pJ per op; F 0.62) | 5.6x | 2.8x | 3.5x | 3.1x (k 0.51) | 1.8x | 3.1x | + +Floor lane 1 assumption (not in by 18:00 UK): the honest 5090 at 1.0 microjoules per hash at zero shadow. +| Chip | Premium kept at 0.652 (card 1.652): absolute N3 / N2 | Premium scaled with the card, 0.390 (card 1.390): absolute N3 / N2 | zero shadow | +|---|---|---|---| +| GDDR7 board, 28 nm controller (the record) | 2.0x / 2.3x | 1.7x / 1.9x | 2.1x | +| HBM3E one stack | 2.4x / 2.9x | 2.0x / 2.4x | 3.1x | +| HBM4 one stack, N12 base die | 2.9x / 3.5x | 2.4x / 2.9x | 4.5x | +| custom HBM4E base die, N3P | 3.1x / 3.8x | 2.6x / 3.2x | 5.6x | +| DRAM on logic, hybrid bonded (2029 to 2031) | 3.3x / 4.1x | 2.7x / 3.4x | 6.7x | +| SRAM full store, one N2 reticle | 3.3x / 4.2x | 2.8x / 3.5x | 7.1x | + +What the rows say. (1) At the knee the GDDR7 board keeps 2.8x with the shadow on (3.2x at an N2 core) against +3.6x at zero shadow: the class v4 shadow buys the honest card 0.8x of edge, not the 1.5x the k = 1 row served and +not the 0.1x the bare-lane floor would give. (2) The strongest DRAM chips read 4.3x to 4.7x at the knee and 6.3x to +6.8x at stock against the 5090 (the record's 2.6x to 2.8x at k = 1 and 4.4x to 5.1x at k = 0.5 bracket the stock +figure and understate the knee one, because the record's k was read at stock). (3) Against the M5 Max the whole +table compresses to 1.7x to 2.8x: the honest SoC's own operating point is most of the resistance. (4) Floor lane +1's honest 5090 at 1.0 microjoules (assumed; its rows were not in by 18:00 UK) would put the GDDR7 board at 1.7x +to 2.3x and the strongest chips at 2.6x to 4.2x. + +## 6. The ranking and the mixed draw + +### 6.1 Families ranked by k, highest first (the hardest for a chip), absolute at N3 against the 5090's lock + +| Rank | Family | Chip pJ/op N3 (unit floor) | 5090 pJ/op at the lock | k (floor) | In the class draw? | +|---|---|---|---|---|---| +| 1 | mad | 1.66 | 8.3 | 0.20 | yes (8 of 75) | +| 2 | add, sub, rotl, rotr, xor | 1.05 to 1.09 | 6.2 | 0.17 to 0.18 | yes (12, 6, 7, 6, 10) | +| 3 | the index fold | 1.19 | 8.3 | 0.14 | every load | +| 4 | or | 0.73 | 6.2 | 0.12 | yes (4) | +| 5 | mul | 0.68 | 8.3 | 0.082 | yes (8) | +| 6 | prmt, lop3 | 0.64, 0.72 | 11.5, 13.0 | 0.056, 0.055 | not drawn (RTL rows only) | +| 7 | mulhi | 0.68 | 21.0 | 0.032 | yes (6) | +| 8 | 32-lane shuffle (butterfly) | 0.63 | 29.4 | 0.021 | yes (8) | +| 9 | L1 scratch read (8 KB flop array) against the card's L2 hit | ROW_SCRATCH_K | 1,400 | pending | not drawn | +| 10 | int8 8x8x8 tile, per MAC | ROW_TILE_K | 2.2 | pending | not drawn (the tensor lever is dead on other grounds) | + +The order is set by the card's price, not the chip's: the chip pays 0.6 to 1.7 pJ for everything, and the card +pays 6.2 for an add, 8.3 for a multiply, 21 for a high word and 29 for a shuffle. So the families the card pays +MOST for (mulhi, shuffle) are the ones a chip undercuts most, and the forcing work is the plain ARX and mad the +card does cheapest. This is the measured form of 15.1b's hold on the shuffle-heavy re-weight. + +### 6.2 The mixed draw that maximises the expected k + +The objective at a FIXED GPU premium: `k_eff(w) = sum w_i e_chip_i / sum w_i e_gpu_i` (the chip's energy for the +drawn program over the card's for the same program; at a fixed `F` the chip pays `k_eff x F`). The band (layer +1, the research lane's 17:00 reading of lane D's cut): B = 4 points on the injecting families only (add, sub, +xor, mad, shfl, rotl, rotr), the lossy families (or, mul, mulhi) at or under their base, `or + mul + mulhi` at most +18 + B, the shuffle capped at its class v4 weight; the index fold on every address and the F8 uniformity floor are +not functions of the weights and do not move. `tools/chip-model/rtl/flow/mix.py` searches the band exhaustively +over the corners (the objective is a ratio of linear forms, so the optimum is at a corner of each family's range). + +On the unit floors (N3, the lock; the shuffle row in at 0.63 pJ, its k 0.021 the lowest of the drawn families): + +| Mix | add, xor, mul, mad, shfl, rotl, sub, mulhi, rotr, or | k_eff (floors) | The card's pJ per op at the lock | k_eff on the core (floor + the core's 2.4 pJ per op overhead at N3) | +|---|---|---|---|---| +| class v4 (the base) | 12, 10, 8, 8, 8, 7, 6, 6, 6, 4 | 0.097 (with the shuffle row in: 0.130 over the other 67 points) | 10.3 | 0.43 | +| the census lane's re-weight (d20eb04bd, PASS on 256 seeds, no era and eras 0 to 7, W = 4 and W = 16) | 13, 11, 6, 10, 8, 8, 7, 2, 6, 4 | 0.115 (+19 percent; 0.153 over the 67 points) | 9.0 | 0.49 (+14 percent) | +| the band's best corner (mix.py, exhaustive: every injecting family at +4, the shuffle at its floor, the lossy families at base minus 4) | 16, 14, 4, 12, 4, 11, 10, 2, 10, 0 | 0.137 (+42 percent) | 8.2 | about 0.51 (+19 percent) | +| the band's worst corner (shuffle and mulhi heavy) | 8, 6, 4, 4, 8, 3, 2, 6, 2, 0 | 0.074 | 12.3 | about 0.40 | + +Reading: the re-weight moves six points off the chip's two easiest families (mulhi, mul) onto the hardest (mad, +the ARX families) and lifts the expected k by about a seventh on the core; its acceptance is the census lane's +(0 exhausted, F8 0.999 to 1.008 of uniform, the verifier within 4 percent); its energy cost on the card is a +premium per instruction 11 percent LOWER (fewer mulhi), so at a fixed premium the program is 12 percent longer. +The shuffle's own row decides whether it goes to its floor: on the card it is the dearest op (29.4 pJ at the lock) +and on the chip a butterfly over the window, so its k is the lowest of the drawn families unless the synthesis says +otherwise. The recommendation: the band's best corner (add 16, xor 14, mad 12, rotl 11, sub 10, rotr 10, shfl 4, mul 4, mulhi 2, or 0) if its census passes (the census lane's harness, about five box-minutes per candidate); the census lane's draw as the passed fallback. Either way the shuffle goes to its floor: the card pays 29.4 pJ for a move the chip does for 0.6. + +## 7. Sources + +- The GPU side: `docs/analysis/counter-asic-4-research.md` 15.1a (the 5090 microbench, 8 October 2026, measured) + and 20.3 (the packs job: 10.8 / 6.4 pJ per counted op on the class v4 shadow); the M5 Max 6.9 pJ per counted + op from the coordinator's order of 8 October 2026 (the latency-shadow record). +- ASAP7: L. T. Clark et al., "ASAP7: A 7-nm finFET predictive process design kit", Microelectronics Journal 53 + (2016); the ORFS platform files (`flow/platforms/asap7`, the 7.5-track RVT library, TC corner 0.70 V / 0 C). +- The flow: OpenROAD-flow-scripts (docker image `openroad/orfs:latest`, Yosys 0.68, OpenROAD and OpenSTA with + `read_vcd`); iverilog 12 for the gate-level simulation. +- Node scaling (claimed): TSMC N5 "30 percent lower power at the same speed" against N7 (TSMC technology page, + https://www.tsmc.com/english/dedicatedFoundry/technology/logic/l_5nm); N3E "25 to 30 percent lower power" + against N5 (https://www.tsmc.com/english/dedicatedFoundry/technology/logic/l_3nm); N2 "25 to 30 percent lower + power" against N3E (https://www.tsmc.com/english/dedicatedFoundry/technology/logic/l_2nm). Read 8 October 2026. +- The chip model's memory rows: `docs/analysis/chip-model-v3.md` 5.5 (the GDDR7 board, 0.466 microjoules) and + `docs/analysis/class-v6/hardware-future.md` section 5 (the strongest DRAM and SRAM rows). +- The band: `docs/design/class-v6-rotating-family.md` section 2 (layer 1) and `igneum-pow/src/generator.rs` + `NONLOAD_WEIGHTS`. diff --git a/tools/chip-model/rtl/Makefile b/tools/chip-model/rtl/Makefile new file mode 100644 index 000000000..dea28cc08 --- /dev/null +++ b/tools/chip-model/rtl/Makefile @@ -0,0 +1,83 @@ +# Shadow-k floor lane: RTL -> ASAP7 (Yosys + OpenROAD, the ORFS docker image) -> pJ per op from a +# random-input gate-level simulation. Reproduces every row of docs/analysis/class-v6/floor/shadow-k.md. +# +# make row-arx synthesise, place, route, simulate and report one family +# make rows every family (serially; use -j for parallel on the box) +# make table collect every power log into table.md and table.csv +# +# Runs on build-4 or build-3 (docker, the build user in the docker group). Never on the Mac. +# Images: openroad/orfs:latest (yosys 0.68, OpenROAD) and orfs-sim:latest (the same plus iverilog, +# built once with: docker run --name b -u root openroad/orfs:latest bash -c 'apt-get update && apt-get install -y iverilog' && docker commit b orfs-sim:latest). + +WORK ?= $(abspath .) +ORFS_IMG ?= openroad/orfs:latest +SIM_IMG ?= orfs-sim:latest +UID_GID := $(shell id -u):$(shell id -g) +# Every run on a box goes through the lease tool (the coordinator's rule, 8 October 2026): THREADS cores from the +# bounded pool at nice 19; the docker container is pinned to the leased set and ORFS gets NUM_CORES=THREADS. +THREADS ?= 24 +LEASE ?= /srv/builds/_bin/lease +LEASE_ON ?= $(shell test -x $(LEASE) && echo 1) +LEASEPFX = $(if $(LEASE_ON),$(LEASE) pool $(THREADS) --label "floor-k shadow-k" --owner floor-k --nice 19 --,) +CPUSET = $(if $(LEASE_ON),--cpuset-cpus {cpuset},) +DOCKER := $(LEASEPFX) docker run --rm -u $(UID_GID) -e HOME=/tmp -e NUM_CORES=$(THREADS) $(CPUSET) -v $(WORK):/work +ORFS := $(DOCKER) -w /OpenROAD-flow-scripts/flow $(ORFS_IMG) +SIM := $(DOCKER) -w /work $(SIM_IMG) + +DESIGNS := arx mul prmt lop3 fold shfl xbar scratch tile core8 core32 core32r16 core8r64 core8i1k core8sel core32all +top = $(shell sed -n 's/^$(1) \([^ ]*\) .*/\1/p' flow/designs.txt) + +# per-family simulation tags (the op field fixed per row where the family has several ops) +SIMS_arx := mix add:+op=0 sub:+op=1 xor:+op=2 or:+op=3 rotl:+op=4 rotr:+op=5 +SIMS_mul := mix mul:+op=0 mulhi:+op=1 mad:+op=2 +SIMS_prmt := mix +SIMS_lop3 := mix +SIMS_fold := mix +SIMS_shfl := mix +SIMS_xbar := mix +SIMS_scratch := mix +SIMS_tile := mix +SIMS_core8 := mix mixld:+loads=1 +SIMS_core32 := mix mixld:+loads=1 +SIMS_core32r16 := mix mixld:+loads=1 +SIMS_core8r64 := mix +SIMS_core8i1k := mix +SIMS_core8sel := mix +SIMS_core32all := mix + +.PHONY: rows table clean + +rows: $(addprefix row-,$(DESIGNS)) + +row-%: power-% + @echo "row $* done: $(WORK)/out/$*/logs/asap7/$*/base/power.log" + +flow-%: + mkdir -p out/$* logs + $(ORFS) make DESIGN_CONFIG=/work/flow/$*.mk 2>&1 | tee logs/flow-$*.log + test -f out/$*/results/asap7/$*/base/6_final.v + +sim-%: flow-% + mkdir -p sim/$* + $(SIM) bash /work/flow/gl2sim.sh $* $(call top,$*) 2>&1 | tee logs/gl2sim-$*.log + for t in $(SIMS_$*); do tag=$${t%%:*}; args=$${t#*:}; [ "$$args" = "$$t" ] && args=""; \ + $(SIM) bash /work/flow/sim.sh $* $$tag $$args 2>&1 | tee logs/sim-$*-$$tag.log; done + +power-%: sim-% + $(ORFS) make DESIGN_CONFIG=/work/flow/$*.mk RUN_SCRIPT=/work/flow/power.tcl RUN_LOG_NAME_STEM=power run 2>&1 | tee logs/power-$*.log + +# synthesis-only row (no placement, no parasitics): the 14:45 fallback +synth-%: + mkdir -p out/$* logs sim/$*-synth + $(ORFS) make DESIGN_CONFIG=/work/flow/$*.mk synth 2>&1 | tee logs/synth-$*.log + NET=1_2_yosys.v $(SIM) bash -c 'NET=1_2_yosys.v bash /work/flow/gl2sim.sh $* $(call top,$*)' 2>&1 | tee logs/gl2sim-$*-synth.log + $(SIM) bash /work/flow/sim.sh $* s150 +cycles=150 2>&1 | tee logs/sim-$*-s150.log + $(SIM) bash /work/flow/sim.sh $* s600 +cycles=600 2>&1 | tee logs/sim-$*-s600.log + $(ORFS) make DESIGN_CONFIG=/work/flow/$*.mk RUN_SCRIPT=/work/flow/power.tcl RUN_LOG_NAME_STEM=power_synth FLOORK_ODB=1_synth.odb FLOORK_SDC=1_synth.sdc run 2>&1 | tee logs/power-$*-synth.log + +table: + python3 flow/collect.py $(WORK) > table.md + @echo wrote table.md and table.csv + +clean: + rm -rf out sim logs table.md table.csv diff --git a/tools/chip-model/rtl/flow/arx.mk b/tools/chip-model/rtl/flow/arx.mk new file mode 100644 index 000000000..ca1b61a86 --- /dev/null +++ b/tools/chip-model/rtl/flow/arx.mk @@ -0,0 +1,15 @@ +# ORFS design config for the arx shadow-core family (top lane_arx), ASAP7. +export PLATFORM = asap7 +export DESIGN_NAME = lane_arx +export DESIGN_NICKNAME = arx +export VERILOG_FILES = /work/rtl/lane_arx.v +export VERILOG_INCLUDE_DIRS = /work/rtl +export SDC_FILE = /work/flow/arx.sdc +export CORE_UTILIZATION = 40 +export CORE_ASPECT_RATIO = 1 +export CORE_MARGIN = 0.5 +export PLACE_DENSITY = 0.55 +export CORNER = TC +export SKIP_LAST_GASP = 1 +export WORK_HOME = /work/out/arx + diff --git a/tools/chip-model/rtl/flow/arx.sdc b/tools/chip-model/rtl/flow/arx.sdc new file mode 100644 index 000000000..81f424f33 --- /dev/null +++ b/tools/chip-model/rtl/flow/arx.sdc @@ -0,0 +1,10 @@ +current_design lane_arx +set clk_name core_clock +set clk_port_name clk +set clk_period 1000 +set clk_io_pct 0.2 +set clk_port [get_ports $clk_port_name] +create_clock -name $clk_name -period $clk_period $clk_port +set non_clock_inputs [all_inputs -no_clocks] +set_input_delay [expr $clk_period * $clk_io_pct] -clock $clk_name $non_clock_inputs +set_output_delay [expr $clk_period * $clk_io_pct] -clock $clk_name [all_outputs] diff --git a/tools/chip-model/rtl/flow/collect.py b/tools/chip-model/rtl/flow/collect.py new file mode 100644 index 000000000..842842019 --- /dev/null +++ b/tools/chip-model/rtl/flow/collect.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 +"""Collect the ORFS power logs into the shadow-k table: pJ per op per family at ASAP7 (TC corner), +scaled to N5, N3 and N2, against the 5090's measured pJ per counted op (counter-asic-4-research.md 15.1a). +Usage: collect.py (writes table.csv beside, prints table.md).""" +import re, sys, os, csv + +work = sys.argv[1] if len(sys.argv) > 1 else '.' +# ops per cycle per design (the per-op divisor) and the GPU row each family is read against +OPS = {'arx': 1, 'mul': 1, 'prmt': 1, 'lop3': 1, 'fold': 1, 'shfl': 32, 'xbar': 32, 'scratch': 1, 'tile': 512, 'core8': 8, 'core32': 32, 'core32r16': 32, 'core8r64': 8, 'core8i1k': 8, 'core8sel': 8, 'core32all': 32} +# 5090 measured pJ per counted op: (unlocked, at the 1,300 MHz lock); 15.1a +GPU = { + 'arx:mix': (11.3, 6.2), 'arx:add': (11.3, 6.2), 'arx:sub': (11.3, 6.2), 'arx:xor': (11.3, 6.2), 'arx:or': (11.3, 6.2), + 'arx:rotl': (11.3, 6.2), 'arx:rotr': (11.3, 6.2), + 'mul:mix': (13.9, 8.3), 'mul:mul': (13.9, 8.3), 'mul:mad': (13.9, 8.3), 'mul:mulhi': (39.6, 21.0), + 'prmt:mix': (22.3, 11.5), 'lop3:mix': (24.1, 13.0), + 'fold:mix': (13.9, 8.3), # the fold is a multiply plus a rotate and masks: read against int_mul + 'shfl:mix': (55.8, 29.4), 'xbar:mix': (55.8, 29.4), + 'scratch:mix': (2400.0, 1400.0), # the card's L2 hit (no shared-memory probe measured: owed) + 'tile:mix': (4.1, 2.2), + 'core8:mix': (11.3, 6.2), 'core8:mixld': (11.3, 6.2), 'core32:mix': (11.3, 6.2), 'core32:mixld': (11.3, 6.2), 'core32r16:mix': (11.3, 6.2), 'core8r64:mix': (11.3, 6.2), 'core8i1k:mix': (11.3, 6.2), 'core8sel:mix': (11.3, 6.2), 'core32all:mix': (11.3, 6.2), # the class v4 draw: read against int_arx (the packs job read the whole mix at 10.8 / 6.4) # dependent u8 m8n8k16 per MAC; the wide s8 tile reads 1.36 / 0.83 +} +# per-node energy scaling from ASAP7 (a 7 nm-class predictive PDK at 0.70 V), approximate and claimed: +# N7 -> N5 x0.70 (TSMC: "30 percent lower power at the same speed"), N5 -> N3E x0.72 (TSMC: 25 to 30 percent), +# N3E -> N2 x0.72 (TSMC: 25 to 30 percent). Sources in shadow-k.md section 3. +SCALE = {'ASAP7': 1.0, 'N5': 0.70, 'N3': 0.70 * 0.72, 'N2': 0.70 * 0.72 * 0.72} + +def parse_log(path): + txt = open(path, errors='replace').read() + period = None + m = re.search(r'FLOORK clock_period_ps ([\d.]+)', txt); period = float(m.group(1)) if m else None + cells = re.search(r'FLOORK cells (\d+)', txt); cells = int(cells.group(1)) if cells else None + rows = {} + for sec in re.split(r'FLOORK === ', txt)[1:]: + head, body = sec.split(' ===', 1) + tag = head.strip().replace('POWER_VCD ', 'vcd:').replace('POWER_PROPAGATED_0.5', 'prop') + # OpenSTA report_power: "Total 100.0%" in watts + tm = re.search(r'^Total\s+([\d.eE+-]+)\s+([\d.eE+-]+)\s+([\d.eE+-]+)\s+([\d.eE+-]+)', body, re.M) + ann = re.search(r'Annotated (\d+) pin activities', body) + unann = re.search(r'unannotated\s+(\d+)', body) + if tm: + rows[tag] = dict(internal=float(tm.group(1)), switching=float(tm.group(2)), leakage=float(tm.group(3)), + total=float(tm.group(4)), annotated=(f"{ann.group(1)} pins, {unann.group(1) if unann else '?'} unannotated") if ann else '') + return period, cells, rows + +LOAD_CYCLES = 260 # the core testbenches: 4 reset + 2 config + 256 program-load cycles before the run (1,028 at NPROG 1024) +def steady(rows, period, short, long_, cs, cl, loadc): + """Solve the run-phase power from two run lengths: P_i = (loadc x b + c_i x a) / (loadc + c_i).""" + out = {} + for key in ('internal', 'switching', 'leakage', 'total'): + Ps, Pl = rows[short][key], rows[long_][key] + A = Ps * (loadc + cs); B = Pl * (loadc + cl) + a = (B - A) / (cl - cs) + out[key] = a + out['annotated'] = rows[long_]['annotated'] + '; steady state solved from the two run lengths' + return out +out = [] +for d in OPS: + found = None + for stem in ('power_synth', 'power_synth_short', 'power'): + log = os.path.join(work, 'out', d, 'logs', 'asap7', d, 'base', stem + '.log') + if os.path.exists(log): + found = log + if stem == 'power': break + if not found: + continue + period, cells, rows = parse_log(found) + pairs = [('vcd:s150', 'vcd:s600', 150, 600), ('vcd:short', 'vcd:synth', 150, 800), ('vcd:s150', 'vcd:s400', 150, 400)] + for sh, lg, cs, cl in pairs: + if sh in rows and lg in rows: + rows['vcd:steady'] = steady(rows, period, sh, lg, cs, cl, 1028 if d in ('core8i1k', 'core32all') else LOAD_CYCLES) + for tag, r in rows.items(): + sub = tag.split(':')[1] if ':' in tag else 'prop' + key = f'{d}:{sub}' if sub in ('add','sub','xor','or','rotl','rotr','mul','mulhi','mad','mixld') else f'{d}:mix' + gpu = GPU.get(key, (None, None)) + pj = r['total'] * period * 1e-12 / OPS[d] * 1e12 # W * s / ops -> pJ + pj_dyn = (r['internal'] + r['switching']) * period / OPS[d] + row = dict(family=d, sim=tag, period_ps=period, cells=cells, ops_per_cycle=OPS[d], + total_W=r['total'], leak_W=r['leakage'], annotated_pct=r['annotated'], + pJ_asap7=pj, pJ_dyn_asap7=pj_dyn) + for node, s in SCALE.items(): + row[f'pJ_{node}'] = pj * s + row[f'k_{node}_lock'] = (pj * s / gpu[1]) if gpu[1] else None + row[f'k_{node}_unlocked'] = (pj * s / gpu[0]) if gpu[0] else None + row['gpu_pJ_unlocked'] = gpu[0]; row['gpu_pJ_lock'] = gpu[1] + # the Apple M5 Max: 6.9 pJ per counted op measured on the class v4 shadow (the honest tier's top); only the + # ARX-class and the core rows have a measured M5 Max figure + m5 = 6.9 if (d == 'arx' or d.startswith('core')) else None + row['m5_pJ'] = m5 + for node, sc in SCALE.items(): + row[f'k_{node}_m5'] = (pj * sc / m5) if m5 else None + out.append(row) + +if out: + with open(os.path.join(work, 'table.csv'), 'w', newline='') as f: + w = csv.DictWriter(f, fieldnames=list(out[0].keys())); w.writeheader(); w.writerows(out) +print('| Family | Sim | Period ps | Cells | Ops/cycle | Total W | Leak W | VCD annotated | pJ/op ASAP7 | pJ/op N5 | pJ/op N3 | pJ/op N2 | 5090 pJ/op unlocked / lock | k at N3 (lock) | k at N2 (lock) | k at N3 (unlocked) | k at N3 vs M5 Max (6.9) |') +print('|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|') +for r in out: + f = lambda v, n=3: ('' if v is None else (f'{v:.{n}g}' if isinstance(v, float) else str(v))) + print(f"| {r['family']} | {r['sim']} | {f(r['period_ps'])} | {r['cells']} | {r['ops_per_cycle']} | {f(r['total_W'])} | {f(r['leak_W'])} | {r['annotated_pct']} | {f(r['pJ_asap7'])} | {f(r['pJ_N5'])} | {f(r['pJ_N3'])} | {f(r['pJ_N2'])} | {f(r['gpu_pJ_unlocked'])} / {f(r['gpu_pJ_lock'])} | {f(r['k_N3_lock'])} | {f(r['k_N2_lock'])} | {f(r['k_N3_unlocked'])} | {f(r['k_N3_m5'])} |") diff --git a/tools/chip-model/rtl/flow/core32.mk b/tools/chip-model/rtl/flow/core32.mk new file mode 100644 index 000000000..9ec9f6dfa --- /dev/null +++ b/tools/chip-model/rtl/flow/core32.mk @@ -0,0 +1,15 @@ +# ORFS design config for the programmable shadow core (core32: core_v6_32), ASAP7. +export PLATFORM = asap7 +export DESIGN_NAME = core_v6_32 +export DESIGN_NICKNAME = core32 +export VERILOG_FILES = /work/rtl/core_v6_32.v +export VERILOG_INCLUDE_DIRS = /work/rtl +export SDC_FILE = /work/flow/core32.sdc +export CORE_UTILIZATION = 40 +export CORE_ASPECT_RATIO = 1 +export CORE_MARGIN = 0.5 +export PLACE_DENSITY = 0.55 +export CORNER = TC +export SKIP_LAST_GASP = 1 +export WORK_HOME = /work/out/core32 +export SYNTH_MEMORY_MAX_BITS = 2000000 diff --git a/tools/chip-model/rtl/flow/core32.sdc b/tools/chip-model/rtl/flow/core32.sdc new file mode 100644 index 000000000..118104793 --- /dev/null +++ b/tools/chip-model/rtl/flow/core32.sdc @@ -0,0 +1,10 @@ +current_design core_v6_32 +set clk_name core_clock +set clk_port_name clk +set clk_period 1500 +set clk_io_pct 0.2 +set clk_port [get_ports $clk_port_name] +create_clock -name $clk_name -period $clk_period $clk_port +set non_clock_inputs [all_inputs -no_clocks] +set_input_delay [expr $clk_period * $clk_io_pct] -clock $clk_name $non_clock_inputs +set_output_delay [expr $clk_period * $clk_io_pct] -clock $clk_name [all_outputs] diff --git a/tools/chip-model/rtl/flow/core32all.mk b/tools/chip-model/rtl/flow/core32all.mk new file mode 100644 index 000000000..c027bf40e --- /dev/null +++ b/tools/chip-model/rtl/flow/core32all.mk @@ -0,0 +1,15 @@ +# ORFS design config for the programmable shadow core (core8: core_v6_32all), ASAP7. +export PLATFORM = asap7 +export DESIGN_NAME = core_v6_32all +export DESIGN_NICKNAME = core32all +export VERILOG_FILES = /work/rtl/core_v6_32all.v +export VERILOG_INCLUDE_DIRS = /work/rtl +export SDC_FILE = /work/flow/core32all.sdc +export CORE_UTILIZATION = 40 +export CORE_ASPECT_RATIO = 1 +export CORE_MARGIN = 0.5 +export PLACE_DENSITY = 0.55 +export CORNER = TC +export SKIP_LAST_GASP = 1 +export WORK_HOME = /work/out/core32all +export SYNTH_MEMORY_MAX_BITS = 2000000 diff --git a/tools/chip-model/rtl/flow/core32all.sdc b/tools/chip-model/rtl/flow/core32all.sdc new file mode 100644 index 000000000..13993d335 --- /dev/null +++ b/tools/chip-model/rtl/flow/core32all.sdc @@ -0,0 +1,10 @@ +current_design core_v6_32all +set clk_name core_clock +set clk_port_name clk +set clk_period 1500 +set clk_io_pct 0.2 +set clk_port [get_ports $clk_port_name] +create_clock -name $clk_name -period $clk_period $clk_port +set non_clock_inputs [all_inputs -no_clocks] +set_input_delay [expr $clk_period * $clk_io_pct] -clock $clk_name $non_clock_inputs +set_output_delay [expr $clk_period * $clk_io_pct] -clock $clk_name [all_outputs] diff --git a/tools/chip-model/rtl/flow/core32r16.mk b/tools/chip-model/rtl/flow/core32r16.mk new file mode 100644 index 000000000..a1b779ece --- /dev/null +++ b/tools/chip-model/rtl/flow/core32r16.mk @@ -0,0 +1,15 @@ +# ORFS design config for the programmable shadow core (core32r16: core_v6_32r16), ASAP7. +export PLATFORM = asap7 +export DESIGN_NAME = core_v6_32r16 +export DESIGN_NICKNAME = core32r16 +export VERILOG_FILES = /work/rtl/core_v6_32r16.v +export VERILOG_INCLUDE_DIRS = /work/rtl +export SDC_FILE = /work/flow/core32r16.sdc +export CORE_UTILIZATION = 40 +export CORE_ASPECT_RATIO = 1 +export CORE_MARGIN = 0.5 +export PLACE_DENSITY = 0.55 +export CORNER = TC +export SKIP_LAST_GASP = 1 +export WORK_HOME = /work/out/core32r16 +export SYNTH_MEMORY_MAX_BITS = 2000000 diff --git a/tools/chip-model/rtl/flow/core32r16.sdc b/tools/chip-model/rtl/flow/core32r16.sdc new file mode 100644 index 000000000..d4aaf339a --- /dev/null +++ b/tools/chip-model/rtl/flow/core32r16.sdc @@ -0,0 +1,10 @@ +current_design core_v6_32r16 +set clk_name core_clock +set clk_port_name clk +set clk_period 1500 +set clk_io_pct 0.2 +set clk_port [get_ports $clk_port_name] +create_clock -name $clk_name -period $clk_period $clk_port +set non_clock_inputs [all_inputs -no_clocks] +set_input_delay [expr $clk_period * $clk_io_pct] -clock $clk_name $non_clock_inputs +set_output_delay [expr $clk_period * $clk_io_pct] -clock $clk_name [all_outputs] diff --git a/tools/chip-model/rtl/flow/core8.mk b/tools/chip-model/rtl/flow/core8.mk new file mode 100644 index 000000000..306449501 --- /dev/null +++ b/tools/chip-model/rtl/flow/core8.mk @@ -0,0 +1,15 @@ +# ORFS design config for the programmable shadow core (core8: core_v6_8), ASAP7. +export PLATFORM = asap7 +export DESIGN_NAME = core_v6_8 +export DESIGN_NICKNAME = core8 +export VERILOG_FILES = /work/rtl/core_v6_8.v +export VERILOG_INCLUDE_DIRS = /work/rtl +export SDC_FILE = /work/flow/core8.sdc +export CORE_UTILIZATION = 40 +export CORE_ASPECT_RATIO = 1 +export CORE_MARGIN = 0.5 +export PLACE_DENSITY = 0.55 +export CORNER = TC +export SKIP_LAST_GASP = 1 +export WORK_HOME = /work/out/core8 +export SYNTH_MEMORY_MAX_BITS = 2000000 diff --git a/tools/chip-model/rtl/flow/core8.sdc b/tools/chip-model/rtl/flow/core8.sdc new file mode 100644 index 000000000..10f3be2bd --- /dev/null +++ b/tools/chip-model/rtl/flow/core8.sdc @@ -0,0 +1,10 @@ +current_design core_v6_8 +set clk_name core_clock +set clk_port_name clk +set clk_period 1500 +set clk_io_pct 0.2 +set clk_port [get_ports $clk_port_name] +create_clock -name $clk_name -period $clk_period $clk_port +set non_clock_inputs [all_inputs -no_clocks] +set_input_delay [expr $clk_period * $clk_io_pct] -clock $clk_name $non_clock_inputs +set_output_delay [expr $clk_period * $clk_io_pct] -clock $clk_name [all_outputs] diff --git a/tools/chip-model/rtl/flow/core8i1k.mk b/tools/chip-model/rtl/flow/core8i1k.mk new file mode 100644 index 000000000..bbfe890c8 --- /dev/null +++ b/tools/chip-model/rtl/flow/core8i1k.mk @@ -0,0 +1,15 @@ +# ORFS design config for the programmable shadow core (core8: core_v6_8i1k), ASAP7. +export PLATFORM = asap7 +export DESIGN_NAME = core_v6_8i1k +export DESIGN_NICKNAME = core8i1k +export VERILOG_FILES = /work/rtl/core_v6_8i1k.v +export VERILOG_INCLUDE_DIRS = /work/rtl +export SDC_FILE = /work/flow/core8i1k.sdc +export CORE_UTILIZATION = 40 +export CORE_ASPECT_RATIO = 1 +export CORE_MARGIN = 0.5 +export PLACE_DENSITY = 0.55 +export CORNER = TC +export SKIP_LAST_GASP = 1 +export WORK_HOME = /work/out/core8i1k +export SYNTH_MEMORY_MAX_BITS = 2000000 diff --git a/tools/chip-model/rtl/flow/core8i1k.sdc b/tools/chip-model/rtl/flow/core8i1k.sdc new file mode 100644 index 000000000..fa6e34e57 --- /dev/null +++ b/tools/chip-model/rtl/flow/core8i1k.sdc @@ -0,0 +1,10 @@ +current_design core_v6_8i1k +set clk_name core_clock +set clk_port_name clk +set clk_period 1500 +set clk_io_pct 0.2 +set clk_port [get_ports $clk_port_name] +create_clock -name $clk_name -period $clk_period $clk_port +set non_clock_inputs [all_inputs -no_clocks] +set_input_delay [expr $clk_period * $clk_io_pct] -clock $clk_name $non_clock_inputs +set_output_delay [expr $clk_period * $clk_io_pct] -clock $clk_name [all_outputs] diff --git a/tools/chip-model/rtl/flow/core8r64.mk b/tools/chip-model/rtl/flow/core8r64.mk new file mode 100644 index 000000000..ff5d73224 --- /dev/null +++ b/tools/chip-model/rtl/flow/core8r64.mk @@ -0,0 +1,15 @@ +# ORFS design config for the programmable shadow core (core8: core_v6_8r64), ASAP7. +export PLATFORM = asap7 +export DESIGN_NAME = core_v6_8r64 +export DESIGN_NICKNAME = core8r64 +export VERILOG_FILES = /work/rtl/core_v6_8r64.v +export VERILOG_INCLUDE_DIRS = /work/rtl +export SDC_FILE = /work/flow/core8r64.sdc +export CORE_UTILIZATION = 40 +export CORE_ASPECT_RATIO = 1 +export CORE_MARGIN = 0.5 +export PLACE_DENSITY = 0.55 +export CORNER = TC +export SKIP_LAST_GASP = 1 +export WORK_HOME = /work/out/core8r64 +export SYNTH_MEMORY_MAX_BITS = 2000000 diff --git a/tools/chip-model/rtl/flow/core8r64.sdc b/tools/chip-model/rtl/flow/core8r64.sdc new file mode 100644 index 000000000..ffb92a9cd --- /dev/null +++ b/tools/chip-model/rtl/flow/core8r64.sdc @@ -0,0 +1,10 @@ +current_design core_v6_8r64 +set clk_name core_clock +set clk_port_name clk +set clk_period 1500 +set clk_io_pct 0.2 +set clk_port [get_ports $clk_port_name] +create_clock -name $clk_name -period $clk_period $clk_port +set non_clock_inputs [all_inputs -no_clocks] +set_input_delay [expr $clk_period * $clk_io_pct] -clock $clk_name $non_clock_inputs +set_output_delay [expr $clk_period * $clk_io_pct] -clock $clk_name [all_outputs] diff --git a/tools/chip-model/rtl/flow/core8sel.mk b/tools/chip-model/rtl/flow/core8sel.mk new file mode 100644 index 000000000..07f9f6ebb --- /dev/null +++ b/tools/chip-model/rtl/flow/core8sel.mk @@ -0,0 +1,15 @@ +# ORFS design config for the programmable shadow core (core8: core_v6_8sel), ASAP7. +export PLATFORM = asap7 +export DESIGN_NAME = core_v6_8sel +export DESIGN_NICKNAME = core8sel +export VERILOG_FILES = /work/rtl/core_v6_8sel.v +export VERILOG_INCLUDE_DIRS = /work/rtl +export SDC_FILE = /work/flow/core8sel.sdc +export CORE_UTILIZATION = 40 +export CORE_ASPECT_RATIO = 1 +export CORE_MARGIN = 0.5 +export PLACE_DENSITY = 0.55 +export CORNER = TC +export SKIP_LAST_GASP = 1 +export WORK_HOME = /work/out/core8sel +export SYNTH_MEMORY_MAX_BITS = 2000000 diff --git a/tools/chip-model/rtl/flow/core8sel.sdc b/tools/chip-model/rtl/flow/core8sel.sdc new file mode 100644 index 000000000..36f995570 --- /dev/null +++ b/tools/chip-model/rtl/flow/core8sel.sdc @@ -0,0 +1,10 @@ +current_design core_v6_8sel +set clk_name core_clock +set clk_port_name clk +set clk_period 1500 +set clk_io_pct 0.2 +set clk_port [get_ports $clk_port_name] +create_clock -name $clk_name -period $clk_period $clk_port +set non_clock_inputs [all_inputs -no_clocks] +set_input_delay [expr $clk_period * $clk_io_pct] -clock $clk_name $non_clock_inputs +set_output_delay [expr $clk_period * $clk_io_pct] -clock $clk_name [all_outputs] diff --git a/tools/chip-model/rtl/flow/designs.txt b/tools/chip-model/rtl/flow/designs.txt new file mode 100644 index 000000000..af4f9566e --- /dev/null +++ b/tools/chip-model/rtl/flow/designs.txt @@ -0,0 +1,16 @@ +arx lane_arx 1000 +mul lane_mul 1500 +prmt lane_prmt 1000 +lop3 lane_lop3 1000 +fold lane_fold 1500 +shfl shfl32 1000 +xbar xbar32 1000 +scratch scratch8k 1500 +tile tile8 2000 +core8 core_v6_8 1500 +core32 core_v6_32 1500 +core32r16 core_v6_32r16 1500 +core8r64 core_v6_8r64 1500 +core8i1k core_v6_8i1k 1500 +core8sel core_v6_8sel 1500 +core32all core_v6_32all 1500 diff --git a/tools/chip-model/rtl/flow/edge.py b/tools/chip-model/rtl/flow/edge.py new file mode 100644 index 000000000..1363f339d --- /dev/null +++ b/tools/chip-model/rtl/flow/edge.py @@ -0,0 +1,26 @@ +#!/usr/bin/env python3 +"""The chip edge rows at the synthesised chip cost. Usage: edge.py [] +Absolute convention: E_chip = E_mem + N_ops x e_chip (the chip's shadow cost does not depend on the card's point). +The record's convention beside it: E_chip = E_mem + k x F with k = e_chip / e_gpu at the card's point.""" +import sys +e3 = float(sys.argv[1]); e2 = float(sys.argv[2]) if len(sys.argv) > 2 else e3 * 0.72 +NOPS = 102100 # class v4 counted ops per hash (the packs job) +# the 5090 rows: (name, card microjoules per hash, premium F, pJ per counted op at that point) +cards = [('5090 stock (class v4)', 3.36, 1.10, 10.8), ('5090 at the 1,300 MHz lock', 2.33, 0.652, 6.4)] +honest = [('M5 Max (class v4 1.40 microjoules, class v3 0.78; 6.9 pJ per op; F 0.62)', 1.40, 0.62, 6.9)] +chips = [('GDDR7 board, 28 nm controller (the record)', 0.466), ('HBM3E one stack', 0.321), ('HBM4 one stack, N12 base die', 0.22), + ('custom HBM4E base die, N3P', 0.18), ('DRAM on logic, hybrid bonded (2029 to 2031)', 0.15), ('SRAM full store, one N2 reticle', 0.14)] +def edge_abs(card, mem, e): return card / (mem + NOPS * e * 1e-6) +def edge_rec(card, F, mem, k): return card / (mem + k * F) +print(f'Chip shadow per hash, absolute: N3 {NOPS*e3*1e-6:.3f} microjoules, N2 {NOPS*e2*1e-6:.3f}.') +print('| Chip (E_mem, microjoules) | Card row | Zero shadow | Absolute, N3 core | Absolute, N2 core | Record convention at the measured k (N3) | at k = 1 | at k = 0.5 |') +print('|---|---|---|---|---|---|---|---|') +for cn, mem in chips: + for rn, card, F, egpu in cards + honest: + k3 = e3 / egpu + print(f'| {cn} ({mem}) | {rn} | {(card-F)/mem:.1f}x | {edge_abs(card,mem,e3):.1f}x | {edge_abs(card,mem,e2):.1f}x | {edge_rec(card,F,mem,k3):.1f}x (k {k3:.2f}) | {edge_rec(card,F,mem,1):.1f}x | {edge_rec(card,F,mem,0.5):.1f}x |') +print('\nFloor lane 1 assumption (not in by 18:00 UK): the honest 5090 at 1.0 microjoules per hash at zero shadow.') +print('| Chip | Premium kept at 0.652 (card 1.652): absolute N3 / N2 | Premium scaled with the card, 0.390 (card 1.390): absolute N3 / N2 | zero shadow |') +print('|---|---|---|---|') +for cn, mem in chips: + print(f'| {cn} | {edge_abs(1.652,mem,e3):.1f}x / {edge_abs(1.652,mem,e2):.1f}x | {edge_abs(1.390,mem,e3):.1f}x / {edge_abs(1.390,mem,e2):.1f}x | {1.0/mem:.1f}x |') diff --git a/tools/chip-model/rtl/flow/fold.mk b/tools/chip-model/rtl/flow/fold.mk new file mode 100644 index 000000000..d5fe9d932 --- /dev/null +++ b/tools/chip-model/rtl/flow/fold.mk @@ -0,0 +1,15 @@ +# ORFS design config for the fold shadow-core family (top lane_fold), ASAP7. +export PLATFORM = asap7 +export DESIGN_NAME = lane_fold +export DESIGN_NICKNAME = fold +export VERILOG_FILES = /work/rtl/lane_fold.v +export VERILOG_INCLUDE_DIRS = /work/rtl +export SDC_FILE = /work/flow/fold.sdc +export CORE_UTILIZATION = 40 +export CORE_ASPECT_RATIO = 1 +export CORE_MARGIN = 0.5 +export PLACE_DENSITY = 0.55 +export CORNER = TC +export SKIP_LAST_GASP = 1 +export WORK_HOME = /work/out/fold + diff --git a/tools/chip-model/rtl/flow/fold.sdc b/tools/chip-model/rtl/flow/fold.sdc new file mode 100644 index 000000000..9dad665a3 --- /dev/null +++ b/tools/chip-model/rtl/flow/fold.sdc @@ -0,0 +1,10 @@ +current_design lane_fold +set clk_name core_clock +set clk_port_name clk +set clk_period 1500 +set clk_io_pct 0.2 +set clk_port [get_ports $clk_port_name] +create_clock -name $clk_name -period $clk_period $clk_port +set non_clock_inputs [all_inputs -no_clocks] +set_input_delay [expr $clk_period * $clk_io_pct] -clock $clk_name $non_clock_inputs +set_output_delay [expr $clk_period * $clk_io_pct] -clock $clk_name [all_outputs] diff --git a/tools/chip-model/rtl/flow/gl2sim.sh b/tools/chip-model/rtl/flow/gl2sim.sh new file mode 100755 index 000000000..a80c15115 --- /dev/null +++ b/tools/chip-model/rtl/flow/gl2sim.sh @@ -0,0 +1,21 @@ +#!/usr/bin/env bash +# gl2sim.sh : netlist -> sim netlist with cell bodies from the liberty. +set -euo pipefail +name=$1; top=$2 +net=/work/out/$name/results/asap7/$name/base/${NET:-6_final.v} +lib=/OpenROAD-flow-scripts/flow/platforms/asap7/lib/NLDM +out=/work/sim/$name; mkdir -p $out +cat > $out/gl2sim.ys < [node=N3] [state=lock|unlocked]""" +import csv, sys, itertools + +path = sys.argv[1]; node = 'N3'; state = 'lock' +for a in sys.argv[2:]: + k, v = a.split('='); + if k == 'node': node = v + if k == 'state': state = v +rows = {(r['family'], r['sim']): r for r in csv.DictReader(open(path))} +def chip(fam, sim='vcd:mix'): + r = rows.get((fam, sim)) or rows.get((fam, 'vcd:mix')) + return float(r[f'pJ_{node}']) if r else None +# the 5090 measured pJ per counted op (15.1a), per family of the generator's draw +gpu = {'add': (11.3, 6.2), 'sub': (11.3, 6.2), 'xor': (11.3, 6.2), 'or': (11.3, 6.2), 'rotl': (11.3, 6.2), 'rotr': (11.3, 6.2), + 'mul': (13.9, 8.3), 'mad': (13.9, 8.3), 'mulhi': (39.6, 21.0), 'shfl': (55.8, 29.4)} +si = 0 if state == 'unlocked' else 1 +e_gpu = {f: v[si] for f, v in gpu.items()} +e_chip = {'add': chip('arx', 'vcd:add'), 'sub': chip('arx', 'vcd:sub'), 'xor': chip('arx', 'vcd:xor'), 'or': chip('arx', 'vcd:or'), + 'rotl': chip('arx', 'vcd:rotl'), 'rotr': chip('arx', 'vcd:rotr'), + 'mul': chip('mul', 'vcd:mul'), 'mad': chip('mul', 'vcd:mad'), 'mulhi': chip('mul', 'vcd:mulhi'), + 'shfl': chip('shfl')} +base = {'add': 12, 'xor': 10, 'mul': 8, 'mad': 8, 'shfl': 8, 'rotl': 7, 'sub': 6, 'mulhi': 6, 'rotr': 6, 'or': 4} +B = 4 +inject = ['add', 'sub', 'xor', 'mad', 'shfl', 'rotl', 'rotr']; lossy = ['or', 'mul', 'mulhi'] +missing = [f for f, v in e_chip.items() if v is None] +if missing: + print('missing chip rows:', missing); sys.exit(1) +k_fam = {f: e_chip[f] / e_gpu[f] for f in base} +def keff(w): return sum(w[f] * e_chip[f] for f in w) / sum(w[f] * e_gpu[f] for f in w) +print(f'node {node}, 5090 state {state}') +print('| Family | Base weight | 5090 pJ/op | Chip pJ/op | k |') +print('|---|---|---|---|---|') +for f in sorted(base, key=lambda f: -k_fam[f]): + print(f'| {f} | {base[f]} | {e_gpu[f]} | {e_chip[f]:.3g} | {k_fam[f]:.3f} |') +print(f'\nclass v4 mix: k_eff = {keff(base):.4f}') +# exhaustive over the band: each family at one of the allowed values (steps of 1 point) +ranges = {} +for f in base: + lo = max(0, base[f] - B) + hi = base[f] + B if f in inject else base[f] + if f == 'shfl': hi = min(hi, 8) + ranges[f] = list(range(lo, hi + 1)) +best = None; worst = None +fams = list(base) +# reduce the search: injecting families only take their extremes and base (the objective is a ratio of linear forms, +# monotone in each weight), lossy families all values +cand = {f: ([ranges[f][0], base[f], ranges[f][-1]] if f in inject else ranges[f]) for f in fams} +for combo in itertools.product(*[cand[f] for f in fams]): + w = dict(zip(fams, combo)) + if w['or'] + w['mul'] + w['mulhi'] > 18 + B: continue + if sum(w.values()) == 0: continue + v = keff(w) + if best is None or v > best[0]: best = (v, dict(w)) + if worst is None or v < worst[0]: worst = (v, dict(w)) +for name, (v, w) in (('best', best), ('worst', worst)): + print(f'{name} mix in the band: k_eff = {v:.4f}: ' + ', '.join(f'{f} {w[f]}' for f in fams) + f' (sum {sum(w.values())})') diff --git a/tools/chip-model/rtl/flow/mul.mk b/tools/chip-model/rtl/flow/mul.mk new file mode 100644 index 000000000..bcafc64c3 --- /dev/null +++ b/tools/chip-model/rtl/flow/mul.mk @@ -0,0 +1,15 @@ +# ORFS design config for the mul shadow-core family (top lane_mul), ASAP7. +export PLATFORM = asap7 +export DESIGN_NAME = lane_mul +export DESIGN_NICKNAME = mul +export VERILOG_FILES = /work/rtl/lane_mul.v +export VERILOG_INCLUDE_DIRS = /work/rtl +export SDC_FILE = /work/flow/mul.sdc +export CORE_UTILIZATION = 40 +export CORE_ASPECT_RATIO = 1 +export CORE_MARGIN = 0.5 +export PLACE_DENSITY = 0.55 +export CORNER = TC +export SKIP_LAST_GASP = 1 +export WORK_HOME = /work/out/mul + diff --git a/tools/chip-model/rtl/flow/mul.sdc b/tools/chip-model/rtl/flow/mul.sdc new file mode 100644 index 000000000..4d393246f --- /dev/null +++ b/tools/chip-model/rtl/flow/mul.sdc @@ -0,0 +1,10 @@ +current_design lane_mul +set clk_name core_clock +set clk_port_name clk +set clk_period 1500 +set clk_io_pct 0.2 +set clk_port [get_ports $clk_port_name] +create_clock -name $clk_name -period $clk_period $clk_port +set non_clock_inputs [all_inputs -no_clocks] +set_input_delay [expr $clk_period * $clk_io_pct] -clock $clk_name $non_clock_inputs +set_output_delay [expr $clk_period * $clk_io_pct] -clock $clk_name [all_outputs] diff --git a/tools/chip-model/rtl/flow/power.tcl b/tools/chip-model/rtl/flow/power.tcl new file mode 100644 index 000000000..44c19072c --- /dev/null +++ b/tools/chip-model/rtl/flow/power.tcl @@ -0,0 +1,27 @@ +# Power of the routed design under (a) the VCD of the random-input gate-level simulation and +# (b) a propagated 0.5 input activity, both at the TC corner (0.70 V, 0 C on ASAP7). +# Run through ORFS: make DESIGN_CONFIG=/work/flow/.mk RUN_SCRIPT=/work/flow/power.tcl run +source $::env(SCRIPTS_DIR)/load.tcl +set odb [expr {[info exists ::env(FLOORK_ODB)] ? $::env(FLOORK_ODB) : "6_final.odb"}] +set sdc [expr {[info exists ::env(FLOORK_SDC)] ? $::env(FLOORK_SDC) : "6_final.sdc"}] +load_design $odb $sdc +puts "FLOORK stage $odb" +set spef $::env(RESULTS_DIR)/6_final.spef +if { $odb == "6_final.odb" && [file exists $spef] } { read_spef $spef } elseif { $odb == "6_final.odb" } { estimate_parasitics -global_routing } elseif { [string match "3_*" $odb] || [string match "4_*" $odb] } { estimate_parasitics -placement } else { puts "FLOORK no parasitics (synthesis only)" } +puts "FLOORK clock_period_ps [expr [get_property [lindex [all_clocks] 0] period]]" +puts "FLOORK cells [llength [get_cells *]]" +report_tns +report_wns +puts "FLOORK === POWER_PROPAGATED_0.5 ===" +set_power_activity -input -activity 0.5 -duty 0.5 +set_power_activity -input_port rst -activity 0 -duty 0 +report_power +foreach vcd [glob -nocomplain /work/sim/$::env(DESIGN_NICKNAME)/*.vcd] { + set tag [file rootname [file tail $vcd]] + if { $tag == "dump" } { continue } + puts "FLOORK === POWER_VCD $tag ===" + read_vcd -scope tb/dut $vcd + if { [info commands report_activity_annotation] != "" } { report_activity_annotation } + report_power +} +puts "FLOORK done" diff --git a/tools/chip-model/rtl/flow/prmt.mk b/tools/chip-model/rtl/flow/prmt.mk new file mode 100644 index 000000000..e17ad1bde --- /dev/null +++ b/tools/chip-model/rtl/flow/prmt.mk @@ -0,0 +1,15 @@ +# ORFS design config for the prmt shadow-core family (top lane_prmt), ASAP7. +export PLATFORM = asap7 +export DESIGN_NAME = lane_prmt +export DESIGN_NICKNAME = prmt +export VERILOG_FILES = /work/rtl/lane_prmt.v +export VERILOG_INCLUDE_DIRS = /work/rtl +export SDC_FILE = /work/flow/prmt.sdc +export CORE_UTILIZATION = 40 +export CORE_ASPECT_RATIO = 1 +export CORE_MARGIN = 0.5 +export PLACE_DENSITY = 0.55 +export CORNER = TC +export SKIP_LAST_GASP = 1 +export WORK_HOME = /work/out/prmt + diff --git a/tools/chip-model/rtl/flow/prmt.sdc b/tools/chip-model/rtl/flow/prmt.sdc new file mode 100644 index 000000000..dabb0c22f --- /dev/null +++ b/tools/chip-model/rtl/flow/prmt.sdc @@ -0,0 +1,10 @@ +current_design lane_prmt +set clk_name core_clock +set clk_port_name clk +set clk_period 1000 +set clk_io_pct 0.2 +set clk_port [get_ports $clk_port_name] +create_clock -name $clk_name -period $clk_period $clk_port +set non_clock_inputs [all_inputs -no_clocks] +set_input_delay [expr $clk_period * $clk_io_pct] -clock $clk_name $non_clock_inputs +set_output_delay [expr $clk_period * $clk_io_pct] -clock $clk_name [all_outputs] diff --git a/tools/chip-model/rtl/flow/scratch.mk b/tools/chip-model/rtl/flow/scratch.mk new file mode 100644 index 000000000..06d205dcd --- /dev/null +++ b/tools/chip-model/rtl/flow/scratch.mk @@ -0,0 +1,16 @@ +# ORFS design config for the scratch shadow-core family (top scratch8k), ASAP7. +export PLATFORM = asap7 +export DESIGN_NAME = scratch8k +export DESIGN_NICKNAME = scratch +export VERILOG_FILES = /work/rtl/scratch8k.v +export VERILOG_INCLUDE_DIRS = /work/rtl +export SDC_FILE = /work/flow/scratch.sdc +export CORE_UTILIZATION = 40 +export CORE_ASPECT_RATIO = 1 +export CORE_MARGIN = 0.5 +export PLACE_DENSITY = 0.55 +export CORNER = TC +export SKIP_LAST_GASP = 1 +export WORK_HOME = /work/out/scratch + +export SYNTH_MEMORY_MAX_BITS = 131072 diff --git a/tools/chip-model/rtl/flow/scratch.sdc b/tools/chip-model/rtl/flow/scratch.sdc new file mode 100644 index 000000000..822efb6cb --- /dev/null +++ b/tools/chip-model/rtl/flow/scratch.sdc @@ -0,0 +1,10 @@ +current_design scratch8k +set clk_name core_clock +set clk_port_name clk +set clk_period 1500 +set clk_io_pct 0.2 +set clk_port [get_ports $clk_port_name] +create_clock -name $clk_name -period $clk_period $clk_port +set non_clock_inputs [all_inputs -no_clocks] +set_input_delay [expr $clk_period * $clk_io_pct] -clock $clk_name $non_clock_inputs +set_output_delay [expr $clk_period * $clk_io_pct] -clock $clk_name [all_outputs] diff --git a/tools/chip-model/rtl/flow/shfl.mk b/tools/chip-model/rtl/flow/shfl.mk new file mode 100644 index 000000000..20b18916c --- /dev/null +++ b/tools/chip-model/rtl/flow/shfl.mk @@ -0,0 +1,16 @@ +# ORFS design config for the shfl shadow-core family (top shfl32), ASAP7. +export PLATFORM = asap7 +export DESIGN_NAME = shfl32 +export DESIGN_NICKNAME = shfl +export VERILOG_FILES = /work/rtl/shfl32.v +export VERILOG_INCLUDE_DIRS = /work/rtl +export SDC_FILE = /work/flow/shfl.sdc +export CORE_UTILIZATION = 40 +export CORE_ASPECT_RATIO = 1 +export CORE_MARGIN = 0.5 +export PLACE_DENSITY = 0.55 +export CORNER = TC +export SKIP_LAST_GASP = 1 +export WORK_HOME = /work/out/shfl + +export SYNTH_MEMORY_MAX_BITS = 131072 diff --git a/tools/chip-model/rtl/flow/shfl.sdc b/tools/chip-model/rtl/flow/shfl.sdc new file mode 100644 index 000000000..530484342 --- /dev/null +++ b/tools/chip-model/rtl/flow/shfl.sdc @@ -0,0 +1,10 @@ +current_design shfl32 +set clk_name core_clock +set clk_port_name clk +set clk_period 1000 +set clk_io_pct 0.2 +set clk_port [get_ports $clk_port_name] +create_clock -name $clk_name -period $clk_period $clk_port +set non_clock_inputs [all_inputs -no_clocks] +set_input_delay [expr $clk_period * $clk_io_pct] -clock $clk_name $non_clock_inputs +set_output_delay [expr $clk_period * $clk_io_pct] -clock $clk_name [all_outputs] diff --git a/tools/chip-model/rtl/flow/sim.sh b/tools/chip-model/rtl/flow/sim.sh new file mode 100755 index 000000000..88b05666d --- /dev/null +++ b/tools/chip-model/rtl/flow/sim.sh @@ -0,0 +1,9 @@ +#!/usr/bin/env bash +# sim.sh [plusargs...] : gate-level random-input simulation -> /work/sim//.vcd +set -euo pipefail +name=$1; tag=$2; shift 2 +out=/work/sim/$name +simcells=$(yosys-config --datdir)/simcells.v +iverilog -g2005 -I /work/tb -o $out/sim_$tag $out/sim_net.v /work/tb/tb_$(sed -n "s/^$name \([^ ]*\) .*/\1/p" /work/flow/designs.txt).v $simcells +( cd $out && vvp -n sim_$tag "$@" | tee sim_$tag.log && mv dump.vcd $tag.vcd ) +ls -la $out/$tag.vcd diff --git a/tools/chip-model/rtl/flow/synthrow.sh b/tools/chip-model/rtl/flow/synthrow.sh new file mode 100755 index 000000000..1b53a543f --- /dev/null +++ b/tools/chip-model/rtl/flow/synthrow.sh @@ -0,0 +1,14 @@ +#!/usr/bin/env bash +# synthrow.sh [cycles-short] [cycles-long] : the synthesis-only power row for a design whose +# ORFS synthesis has already run (1_2_yosys.v and 1_synth.odb present): gate-level sim at two run lengths, then +# OpenSTA power on the synthesised netlist (no wires). Every step under the lease tool when present. +set -euo pipefail +name=$1; cs=${2:-150}; cl=${3:-600} +cd "$(dirname "$0")/.." +WORK=$(pwd); top=$(sed -n "s/^$name \([^ ]*\) .*/\1/p" flow/designs.txt) +UG=$(id -u):$(id -g) +if [ -x /srv/builds/_bin/lease ]; then L="/srv/builds/_bin/lease pool 8 --label floor-k-synthrow-$name --owner floor-k --nice 19 --"; else L=""; fi +$L docker run --rm -u $UG -e HOME=/tmp -e NET=1_2_yosys.v -v $WORK:/work -w /work orfs-sim:latest bash /work/flow/gl2sim.sh $name $top +$L docker run --rm -u $UG -e HOME=/tmp -v $WORK:/work -w /work orfs-sim:latest bash /work/flow/sim.sh $name s$cs +cycles=$cs +$L docker run --rm -u $UG -e HOME=/tmp -v $WORK:/work -w /work orfs-sim:latest bash /work/flow/sim.sh $name s$cl +cycles=$cl +$L docker run --rm -u $UG -e HOME=/tmp -e NUM_CORES=8 -v $WORK:/work -w /OpenROAD-flow-scripts/flow openroad/orfs:latest make DESIGN_CONFIG=/work/flow/$name.mk RUN_SCRIPT=/work/flow/power.tcl RUN_LOG_NAME_STEM=power_synth FLOORK_ODB=1_synth.odb FLOORK_SDC=1_synth.sdc run diff --git a/tools/chip-model/rtl/flow/tile.mk b/tools/chip-model/rtl/flow/tile.mk new file mode 100644 index 000000000..69f85f03c --- /dev/null +++ b/tools/chip-model/rtl/flow/tile.mk @@ -0,0 +1,15 @@ +# ORFS design config for the tile shadow-core family (top tile8), ASAP7. +export PLATFORM = asap7 +export DESIGN_NAME = tile8 +export DESIGN_NICKNAME = tile +export VERILOG_FILES = /work/rtl/tile8.v +export VERILOG_INCLUDE_DIRS = /work/rtl +export SDC_FILE = /work/flow/tile.sdc +export CORE_UTILIZATION = 40 +export CORE_ASPECT_RATIO = 1 +export CORE_MARGIN = 0.5 +export PLACE_DENSITY = 0.55 +export CORNER = TC +export SKIP_LAST_GASP = 1 +export WORK_HOME = /work/out/tile + diff --git a/tools/chip-model/rtl/flow/tile.sdc b/tools/chip-model/rtl/flow/tile.sdc new file mode 100644 index 000000000..2bac92473 --- /dev/null +++ b/tools/chip-model/rtl/flow/tile.sdc @@ -0,0 +1,10 @@ +current_design tile8 +set clk_name core_clock +set clk_port_name clk +set clk_period 2000 +set clk_io_pct 0.2 +set clk_port [get_ports $clk_port_name] +create_clock -name $clk_name -period $clk_period $clk_port +set non_clock_inputs [all_inputs -no_clocks] +set_input_delay [expr $clk_period * $clk_io_pct] -clock $clk_name $non_clock_inputs +set_output_delay [expr $clk_period * $clk_io_pct] -clock $clk_name [all_outputs] diff --git a/tools/chip-model/rtl/flow/xbar.mk b/tools/chip-model/rtl/flow/xbar.mk new file mode 100644 index 000000000..612f4f154 --- /dev/null +++ b/tools/chip-model/rtl/flow/xbar.mk @@ -0,0 +1,16 @@ +# ORFS design config for the xbar shadow-core family (top xbar32), ASAP7. +export PLATFORM = asap7 +export DESIGN_NAME = xbar32 +export DESIGN_NICKNAME = xbar +export VERILOG_FILES = /work/rtl/xbar32.v +export VERILOG_INCLUDE_DIRS = /work/rtl +export SDC_FILE = /work/flow/xbar.sdc +export CORE_UTILIZATION = 40 +export CORE_ASPECT_RATIO = 1 +export CORE_MARGIN = 0.5 +export PLACE_DENSITY = 0.55 +export CORNER = TC +export SKIP_LAST_GASP = 1 +export WORK_HOME = /work/out/xbar + +export SYNTH_MEMORY_MAX_BITS = 131072 diff --git a/tools/chip-model/rtl/flow/xbar.sdc b/tools/chip-model/rtl/flow/xbar.sdc new file mode 100644 index 000000000..509dd19ff --- /dev/null +++ b/tools/chip-model/rtl/flow/xbar.sdc @@ -0,0 +1,10 @@ +current_design xbar32 +set clk_name core_clock +set clk_port_name clk +set clk_period 1000 +set clk_io_pct 0.2 +set clk_port [get_ports $clk_port_name] +create_clock -name $clk_name -period $clk_period $clk_port +set non_clock_inputs [all_inputs -no_clocks] +set_input_delay [expr $clk_period * $clk_io_pct] -clock $clk_name $non_clock_inputs +set_output_delay [expr $clk_period * $clk_io_pct] -clock $clk_name [all_outputs] diff --git a/tools/chip-model/rtl/rtl/core_v6.v b/tools/chip-model/rtl/rtl/core_v6.v new file mode 100644 index 000000000..da57559ca --- /dev/null +++ b/tools/chip-model/rtl/rtl/core_v6.v @@ -0,0 +1,120 @@ +// The programmable shadow core: the minimal in-order SIMD core that executes a class v6 shadow program as +// drawn. Per core: an instruction memory sized to the drawn program length (256 x 32-bit, a flop array), a +// program counter that wraps at the era's drawn length, fetch into an instruction register, decode, the era's +// parameter registers (fold constants M, R, WM, OFF, MASK; program length N). Per lane: a 32 x 32-bit register +// file (flops) with two or three read ports (mad reads three) and one write port, and the class's units: add, sub, +// xor, or, rotl by immediate, rotr by register, mul, mulhi, mad, prmt, lop3, the xor-mask shuffle across the +// lanes (a log2(LANES)-stage butterfly), and the load (the index fold on the address path, the returned word +// written on the next cycle). One instruction per cycle for every lane: LANES lane-ops per cycle. +// +// Instruction word: op[3:0] dst[8:4] src[13:9] src2[18:14] imm[23:19] aux[31:24]. +// 0 add 1 sub 2 xor 3 or 4 rotl(imm) 5 rotr(src) 6 mul 7 mulhi 8 mad 9 shfl(imm mask) 10 prmt(aux,aux) +// 11 lop3(aux lut) 12 load(fold(src) -> addr; dst <= returned word) 13 add 14 xor 15 sub +`include "lane_common.vh" +module core_v6 #(parameter LANES = 32, parameter LOG_LANES = 5, parameter REGS = 32, parameter LOG_REGS = 5, + parameter IW = 32, parameter IMEM_LOG = 8, parameter SELTREE = 0) ( + input clk, input rst, input run, + input prog_we, input [9:0] prog_addr, input [IW-1:0] prog_data, + input cfg_en, input [9:0] cfg_n, input [63:0] cfg_sel, input [31:0] cfg_m, input [4:0] cfg_r, input [31:0] cfg_wm, input [31:0] cfg_off, input [31:0] cfg_mask, + input [31:0] ld_val, + output [31:0] addr, output [31:0] out); + // instruction memory and the sequencer + localparam IMEM = 1 << IMEM_LOG; + reg [IW-1:0] imem [0:IMEM-1]; + reg [IMEM_LOG-1:0] pc; reg [IMEM_LOG-1:0] n_q; reg [IW-1:0] ir; + reg [63:0] sel_q; // the era's drawn op permutation (SELTREE = 1): 16 x 4-bit op codes + reg [31:0] m_q, wm_q, off_q, mask_q; reg [4:0] r_q; + integer i, l; + always @(posedge clk) begin + if (prog_we) imem[prog_addr[IMEM_LOG-1:0]] <= prog_data; + if (rst) begin pc <= 0; ir <= 0; n_q <= {IMEM_LOG{1'b1}}; sel_q <= 64'hfedcba9876543210; m_q <= 32'h9e3779b1; r_q <= 5'd13; wm_q <= 32'h0fffffc0; off_q <= 3; mask_q <= 32'h0fffffff; end + else begin + if (cfg_en) begin n_q <= cfg_n[IMEM_LOG-1:0]; sel_q <= cfg_sel; m_q <= cfg_m | 1; r_q <= cfg_r; wm_q <= cfg_wm; off_q <= cfg_off; mask_q <= cfg_mask; end + if (run) begin ir <= imem[pc]; pc <= (pc == n_q) ? {IMEM_LOG{1'b0}} : pc + 1'b1; end + end + end + // decode + // the drawn select tree (SELTREE = 1): the op code is remapped through the era's 16-entry permutation before + // decode, so the datapath's select structure is the era's draw, not a fixed table a chip could hard-wire + wire [3:0] op_raw = ir[3:0]; + wire [3:0] op = SELTREE ? sel_q[op_raw*4 +: 4] : op_raw; + wire [LOG_REGS-1:0] dst = ir[4 +: LOG_REGS]; wire [LOG_REGS-1:0] src = ir[4+LOG_REGS +: LOG_REGS]; wire [LOG_REGS-1:0] src2 = ir[4+2*LOG_REGS +: LOG_REGS]; + wire [4:0] imm = ir[4+3*LOG_REGS +: 5]; wire [7:0] aux = ir[9+3*LOG_REGS +: 8]; + wire is_load = (op == 4'd12); + wire [4:0] rn = (imm == 0) ? 5'd1 : imm; + wire [LOG_LANES-1:0] smask = imm[LOG_LANES-1:0]; + // lanes + reg [31:0] rf [0:LANES*REGS-1]; + wire [31:0] d_v [0:LANES-1]; wire [31:0] s_v [0:LANES-1]; wire [31:0] s2_v [0:LANES-1]; + wire [31:0] shin [0:LANES-1]; wire [31:0] shout [0:LANES-1]; + wire [31:0] res [0:LANES-1]; + genvar g, st; + generate for (g = 0; g < LANES; g = g + 1) begin : ln + assign d_v[g] = rf[g*REGS + dst]; + assign s_v[g] = rf[g*REGS + src]; + assign s2_v[g] = rf[g*REGS + src2]; + assign shin[g] = s_v[g]; + end endgenerate + // the butterfly shuffle network across the lanes + wire [31:0] bf [0:LOG_LANES][0:LANES-1]; + generate + for (g = 0; g < LANES; g = g + 1) begin : bf0 + assign bf[0][g] = shin[g]; + end + for (st = 0; st < LOG_LANES; st = st + 1) begin : bfs + for (g = 0; g < LANES; g = g + 1) begin : bfl + assign bf[st+1][g] = smask[st] ? bf[st][g ^ (1 << st)] : bf[st][g]; + end + end + for (g = 0; g < LANES; g = g + 1) begin : bfo + assign shout[g] = bf[LOG_LANES][g]; + end + endgenerate + // the units per lane + function [7:0] pick; input [63:0] b; input [3:0] k; reg [7:0] v; + begin v = b[8*k[2:0] +: 8]; pick = k[3] ? {8{v[7]}} : v; end + endfunction + generate for (g = 0; g < LANES; g = g + 1) begin : un + wire [31:0] d = d_v[g]; wire [31:0] s = s_v[g]; wire [31:0] s2 = s2_v[g]; + wire [4:0] sn = (s[4:0] == 0) ? 5'd1 : s[4:0]; + wire mad = (op == 4'd8); + wire [63:0] p = (mad ? s : d) * (mad ? s2 : s); + wire [63:0] bytes = {s, d}; wire [15:0] sel = {aux, aux}; + wire [31:0] prm = {pick(bytes, sel[15:12]), pick(bytes, sel[11:8]), pick(bytes, sel[7:4]), pick(bytes, sel[3:0])}; + reg [31:0] lp; integer b; + always @* for (b = 0; b < 32; b = b + 1) lp[b] = aux[{d[b], s[b], s2[b]}]; + reg [31:0] r; + always @* begin + case (op) + 4'd0, 4'd13: r = d + s; + 4'd1, 4'd15: r = d - s; + 4'd2, 4'd14: r = d ^ s; + 4'd3: r = d | s; + 4'd4: r = `ROTL32(d, rn); + 4'd5: r = `ROTR32(d, sn); + 4'd6: r = p[31:0]; + 4'd7: r = p[63:32]; + 4'd8: r = p[31:0] + d; + 4'd9: r = d ^ shout[g]; + 4'd10: r = prm; + 4'd11: r = lp; + default: r = ld_val ^ (32'h9e3779b9 * (g + 1)); // the returned word (lane-salted by the testbench's bus) + endcase + end + assign res[g] = r; + end endgenerate + // the address path: the fold of lane 0's source (one address per lane on a real part; lane 0 drives the port) + wire [31:0] fx = s_v[0] * m_q; + wire [4:0] frn = (r_q == 0) ? 5'd1 : r_q; + wire [31:0] fy = `ROTL32(fx, frn); + assign addr = is_load ? (((fy & wm_q) | off_q) & mask_q) : 32'd0; + // writeback + always @(posedge clk) begin + if (rst) begin for (i = 0; i < LANES*REGS; i = i + 1) rf[i] <= 32'h9e3779b9 * (i + 1); end + else if (run) for (l = 0; l < LANES; l = l + 1) rf[l*REGS + dst] <= res[l]; + end + // keep every lane alive: the xor over the lanes' results + reg [31:0] red; integer q; + always @* begin red = 0; for (q = 0; q < LANES; q = q + 1) red = red ^ res[q]; end + assign out = red; +endmodule diff --git a/tools/chip-model/rtl/rtl/core_v6_32.v b/tools/chip-model/rtl/rtl/core_v6_32.v new file mode 100644 index 000000000..a12f2b262 --- /dev/null +++ b/tools/chip-model/rtl/rtl/core_v6_32.v @@ -0,0 +1,7 @@ +`include "core_v6.v" +module core_v6_32(input clk, input rst, input run, input prog_we, input [9:0] prog_addr, input [32-1:0] prog_data, + input cfg_en, input [9:0] cfg_n, input [63:0] cfg_sel, input [31:0] cfg_m, input [4:0] cfg_r, input [31:0] cfg_wm, input [31:0] cfg_off, input [31:0] cfg_mask, + input [31:0] ld_val, output [31:0] addr, output [31:0] out); + core_v6 #(.LANES(32), .LOG_LANES(5)) c(.clk(clk), .rst(rst), .run(run), .prog_we(prog_we), .prog_addr(prog_addr), .prog_data(prog_data), + .cfg_en(cfg_en), .cfg_n(cfg_n), .cfg_sel(cfg_sel), .cfg_m(cfg_m), .cfg_r(cfg_r), .cfg_wm(cfg_wm), .cfg_off(cfg_off), .cfg_mask(cfg_mask), .ld_val(ld_val), .addr(addr), .out(out)); +endmodule diff --git a/tools/chip-model/rtl/rtl/core_v6_32all.v b/tools/chip-model/rtl/rtl/core_v6_32all.v new file mode 100644 index 000000000..918bdb392 --- /dev/null +++ b/tools/chip-model/rtl/rtl/core_v6_32all.v @@ -0,0 +1,7 @@ +`include "core_v6.v" +module core_v6_32all(input clk, input rst, input run, input prog_we, input [9:0] prog_addr, input [40-1:0] prog_data, + input cfg_en, input [9:0] cfg_n, input [63:0] cfg_sel, input [31:0] cfg_m, input [4:0] cfg_r, input [31:0] cfg_wm, input [31:0] cfg_off, input [31:0] cfg_mask, + input [31:0] ld_val, output [31:0] addr, output [31:0] out); + core_v6 #(.LANES(32), .LOG_LANES(5), .REGS(64), .LOG_REGS(6), .IW(40), .IMEM_LOG(10), .SELTREE(1)) c(.clk(clk), .rst(rst), .run(run), .prog_we(prog_we), .prog_addr(prog_addr), .prog_data(prog_data), + .cfg_en(cfg_en), .cfg_n(cfg_n), .cfg_sel(cfg_sel), .cfg_m(cfg_m), .cfg_r(cfg_r), .cfg_wm(cfg_wm), .cfg_off(cfg_off), .cfg_mask(cfg_mask), .ld_val(ld_val), .addr(addr), .out(out)); +endmodule diff --git a/tools/chip-model/rtl/rtl/core_v6_32r16.v b/tools/chip-model/rtl/rtl/core_v6_32r16.v new file mode 100644 index 000000000..43528b44c --- /dev/null +++ b/tools/chip-model/rtl/rtl/core_v6_32r16.v @@ -0,0 +1,7 @@ +`include "core_v6.v" +module core_v6_32r16(input clk, input rst, input run, input prog_we, input [9:0] prog_addr, input [32-1:0] prog_data, + input cfg_en, input [9:0] cfg_n, input [63:0] cfg_sel, input [31:0] cfg_m, input [4:0] cfg_r, input [31:0] cfg_wm, input [31:0] cfg_off, input [31:0] cfg_mask, + input [31:0] ld_val, output [31:0] addr, output [31:0] out); + core_v6 #(.LANES(32), .LOG_LANES(5), .REGS(16), .LOG_REGS(4)) c(.clk(clk), .rst(rst), .run(run), .prog_we(prog_we), .prog_addr(prog_addr), .prog_data(prog_data), + .cfg_en(cfg_en), .cfg_n(cfg_n), .cfg_sel(cfg_sel), .cfg_m(cfg_m), .cfg_r(cfg_r), .cfg_wm(cfg_wm), .cfg_off(cfg_off), .cfg_mask(cfg_mask), .ld_val(ld_val), .addr(addr), .out(out)); +endmodule diff --git a/tools/chip-model/rtl/rtl/core_v6_8.v b/tools/chip-model/rtl/rtl/core_v6_8.v new file mode 100644 index 000000000..77c025d50 --- /dev/null +++ b/tools/chip-model/rtl/rtl/core_v6_8.v @@ -0,0 +1,7 @@ +`include "core_v6.v" +module core_v6_8(input clk, input rst, input run, input prog_we, input [9:0] prog_addr, input [32-1:0] prog_data, + input cfg_en, input [9:0] cfg_n, input [63:0] cfg_sel, input [31:0] cfg_m, input [4:0] cfg_r, input [31:0] cfg_wm, input [31:0] cfg_off, input [31:0] cfg_mask, + input [31:0] ld_val, output [31:0] addr, output [31:0] out); + core_v6 #(.LANES(8), .LOG_LANES(3)) c(.clk(clk), .rst(rst), .run(run), .prog_we(prog_we), .prog_addr(prog_addr), .prog_data(prog_data), + .cfg_en(cfg_en), .cfg_n(cfg_n), .cfg_sel(cfg_sel), .cfg_m(cfg_m), .cfg_r(cfg_r), .cfg_wm(cfg_wm), .cfg_off(cfg_off), .cfg_mask(cfg_mask), .ld_val(ld_val), .addr(addr), .out(out)); +endmodule diff --git a/tools/chip-model/rtl/rtl/core_v6_8i1k.v b/tools/chip-model/rtl/rtl/core_v6_8i1k.v new file mode 100644 index 000000000..61758a7b9 --- /dev/null +++ b/tools/chip-model/rtl/rtl/core_v6_8i1k.v @@ -0,0 +1,7 @@ +`include "core_v6.v" +module core_v6_8i1k(input clk, input rst, input run, input prog_we, input [9:0] prog_addr, input [32-1:0] prog_data, + input cfg_en, input [9:0] cfg_n, input [63:0] cfg_sel, input [31:0] cfg_m, input [4:0] cfg_r, input [31:0] cfg_wm, input [31:0] cfg_off, input [31:0] cfg_mask, + input [31:0] ld_val, output [31:0] addr, output [31:0] out); + core_v6 #(.LANES(8), .LOG_LANES(3), .IMEM_LOG(10)) c(.clk(clk), .rst(rst), .run(run), .prog_we(prog_we), .prog_addr(prog_addr), .prog_data(prog_data), + .cfg_en(cfg_en), .cfg_n(cfg_n), .cfg_sel(cfg_sel), .cfg_m(cfg_m), .cfg_r(cfg_r), .cfg_wm(cfg_wm), .cfg_off(cfg_off), .cfg_mask(cfg_mask), .ld_val(ld_val), .addr(addr), .out(out)); +endmodule diff --git a/tools/chip-model/rtl/rtl/core_v6_8r64.v b/tools/chip-model/rtl/rtl/core_v6_8r64.v new file mode 100644 index 000000000..134032fce --- /dev/null +++ b/tools/chip-model/rtl/rtl/core_v6_8r64.v @@ -0,0 +1,7 @@ +`include "core_v6.v" +module core_v6_8r64(input clk, input rst, input run, input prog_we, input [9:0] prog_addr, input [40-1:0] prog_data, + input cfg_en, input [9:0] cfg_n, input [63:0] cfg_sel, input [31:0] cfg_m, input [4:0] cfg_r, input [31:0] cfg_wm, input [31:0] cfg_off, input [31:0] cfg_mask, + input [31:0] ld_val, output [31:0] addr, output [31:0] out); + core_v6 #(.LANES(8), .LOG_LANES(3), .REGS(64), .LOG_REGS(6), .IW(40)) c(.clk(clk), .rst(rst), .run(run), .prog_we(prog_we), .prog_addr(prog_addr), .prog_data(prog_data), + .cfg_en(cfg_en), .cfg_n(cfg_n), .cfg_sel(cfg_sel), .cfg_m(cfg_m), .cfg_r(cfg_r), .cfg_wm(cfg_wm), .cfg_off(cfg_off), .cfg_mask(cfg_mask), .ld_val(ld_val), .addr(addr), .out(out)); +endmodule diff --git a/tools/chip-model/rtl/rtl/core_v6_8sel.v b/tools/chip-model/rtl/rtl/core_v6_8sel.v new file mode 100644 index 000000000..b769156c7 --- /dev/null +++ b/tools/chip-model/rtl/rtl/core_v6_8sel.v @@ -0,0 +1,7 @@ +`include "core_v6.v" +module core_v6_8sel(input clk, input rst, input run, input prog_we, input [9:0] prog_addr, input [32-1:0] prog_data, + input cfg_en, input [9:0] cfg_n, input [63:0] cfg_sel, input [31:0] cfg_m, input [4:0] cfg_r, input [31:0] cfg_wm, input [31:0] cfg_off, input [31:0] cfg_mask, + input [31:0] ld_val, output [31:0] addr, output [31:0] out); + core_v6 #(.LANES(8), .LOG_LANES(3), .SELTREE(1)) c(.clk(clk), .rst(rst), .run(run), .prog_we(prog_we), .prog_addr(prog_addr), .prog_data(prog_data), + .cfg_en(cfg_en), .cfg_n(cfg_n), .cfg_sel(cfg_sel), .cfg_m(cfg_m), .cfg_r(cfg_r), .cfg_wm(cfg_wm), .cfg_off(cfg_off), .cfg_mask(cfg_mask), .ld_val(ld_val), .addr(addr), .out(out)); +endmodule diff --git a/tools/chip-model/rtl/rtl/lane_arx.v b/tools/chip-model/rtl/rtl/lane_arx.v new file mode 100644 index 000000000..be590f9a1 --- /dev/null +++ b/tools/chip-model/rtl/rtl/lane_arx.v @@ -0,0 +1,38 @@ +// 32-bit int ARX lane: add, sub, xor, or, rotl by immediate, rotr by register (the class v4/v6 families +// add, sub, xor, or, rotl, rotr). op: 0 add 1 sub 2 xor 3 or 4 rotl-imm 5 rotr-var 6 add 7 xor. +`include "lane_common.vh" +module lane_arx( + input clk, input rst, + input [2:0] op, input [2:0] dst, input [2:0] src, input [4:0] rot, + input ld_en, input [31:0] ld_val, + output [31:0] out); + reg [31:0] rf [0:7]; + reg [2:0] op_q, dst_q, src_q; reg [4:0] rot_q; reg ld_q; reg [31:0] ld_val_q; + integer i; + always @(posedge clk) begin + if (rst) begin op_q <= 0; dst_q <= 0; src_q <= 0; rot_q <= 0; ld_q <= 0; ld_val_q <= 0; end + else begin op_q <= op; dst_q <= dst; src_q <= src; rot_q <= rot; ld_q <= ld_en; ld_val_q <= ld_val; end + end + wire [31:0] d = rf[dst_q]; + wire [31:0] s = rf[src_q]; + wire [4:0] rn = (rot_q == 5'd0) ? 5'd1 : rot_q; // rotate by 1..31 + wire [4:0] sn = (s[4:0] == 5'd0) ? 5'd1 : s[4:0]; + reg [31:0] res; + always @* begin + case (op_q) + 3'd0: res = d + s; + 3'd1: res = d - s; + 3'd2: res = d ^ s; + 3'd3: res = d | s; + 3'd4: res = `ROTL32(d, rn); + 3'd5: res = `ROTR32(d, sn); + 3'd6: res = d + s; + default: res = d ^ s; + endcase + end + always @(posedge clk) begin + if (rst) begin for (i = 0; i < 8; i = i + 1) rf[i] <= 32'h9e3779b9 * (i + 1); end + else rf[dst_q] <= ld_q ? ld_val_q : res; + end + assign out = res; +endmodule diff --git a/tools/chip-model/rtl/rtl/lane_common.vh b/tools/chip-model/rtl/rtl/lane_common.vh new file mode 100644 index 000000000..c977a0d4c --- /dev/null +++ b/tools/chip-model/rtl/rtl/lane_common.vh @@ -0,0 +1,5 @@ +// Shared helpers for the shadow-core lanes (class v5/v6 mixer draw space). +// A lane = an instruction register, an 8 x 32-bit register window (flops), two or three +// read ports through muxes, one functional unit, one write port. One op per cycle. +`define ROTL32(x, n) (((x) << (n)) | ((x) >> (32 - (n)))) +`define ROTR32(x, n) (((x) >> (n)) | ((x) << (32 - (n)))) diff --git a/tools/chip-model/rtl/rtl/lane_fold.v b/tools/chip-model/rtl/rtl/lane_fold.v new file mode 100644 index 000000000..5b1fbc890 --- /dev/null +++ b/tools/chip-model/rtl/rtl/lane_fold.v @@ -0,0 +1,32 @@ +// The index fold (class v6 layer 1, the ring-A rule): idx = ((rotl(x * M, R) & WM) | OFF) & MASK. +// M, R, WM, OFF and MASK are the era's constants, held in registers (loaded by cfg_en, then static). +// One fold per cycle on a register read; the index is written back to the lane's address register. +`include "lane_common.vh" +module lane_fold( + input clk, input rst, + input [2:0] dst, input [2:0] src, + input cfg_en, input [31:0] cfg_m, input [4:0] cfg_r, input [31:0] cfg_wm, input [31:0] cfg_off, input [31:0] cfg_mask, + input ld_en, input [31:0] ld_val, + output [31:0] out); + reg [31:0] rf [0:7]; + reg [2:0] dst_q, src_q; reg ld_q; reg [31:0] ld_val_q; + reg [31:0] m_q, wm_q, off_q, mask_q; reg [4:0] r_q; + integer i; + always @(posedge clk) begin + if (rst) begin dst_q <= 0; src_q <= 0; ld_q <= 0; ld_val_q <= 0; m_q <= 32'h9e3779b1; r_q <= 5'd13; wm_q <= 32'h0fffffc0; off_q <= 32'h3; mask_q <= 32'h0fffffff; end + else begin + dst_q <= dst; src_q <= src; ld_q <= ld_en; ld_val_q <= ld_val; + if (cfg_en) begin m_q <= cfg_m | 32'h1; r_q <= cfg_r; wm_q <= cfg_wm; off_q <= cfg_off; mask_q <= cfg_mask; end + end + end + wire [31:0] x = rf[src_q]; + wire [31:0] prod = x * m_q; + wire [4:0] rn = (r_q == 5'd0) ? 5'd1 : r_q; + wire [31:0] y = `ROTL32(prod, rn); + wire [31:0] res = ((y & wm_q) | off_q) & mask_q; + always @(posedge clk) begin + if (rst) begin for (i = 0; i < 8; i = i + 1) rf[i] <= 32'h9e3779b9 * (i + 1) + 4; end + else rf[dst_q] <= ld_q ? ld_val_q : (res ^ rf[dst_q]); + end + assign out = res; +endmodule diff --git a/tools/chip-model/rtl/rtl/lane_lop3.v b/tools/chip-model/rtl/rtl/lane_lop3.v new file mode 100644 index 000000000..68c1c8690 --- /dev/null +++ b/tools/chip-model/rtl/rtl/lane_lop3.v @@ -0,0 +1,25 @@ +// Three-input logic lane (LOP3): an 8-bit truth table over (d, s, s2), bitwise. +`include "lane_common.vh" +module lane_lop3( + input clk, input rst, + input [2:0] dst, input [2:0] src, input [2:0] src2, input [7:0] lut, + input ld_en, input [31:0] ld_val, + output [31:0] out); + reg [31:0] rf [0:7]; + reg [2:0] dst_q, src_q, src2_q; reg [7:0] lut_q; reg ld_q; reg [31:0] ld_val_q; + integer i, b; + always @(posedge clk) begin + if (rst) begin dst_q <= 0; src_q <= 0; src2_q <= 0; lut_q <= 0; ld_q <= 0; ld_val_q <= 0; end + else begin dst_q <= dst; src_q <= src; src2_q <= src2; lut_q <= lut; ld_q <= ld_en; ld_val_q <= ld_val; end + end + wire [31:0] d = rf[dst_q]; + wire [31:0] s = rf[src_q]; + wire [31:0] s2 = rf[src2_q]; + reg [31:0] res; + always @* for (b = 0; b < 32; b = b + 1) res[b] = lut_q[{d[b], s[b], s2[b]}]; + always @(posedge clk) begin + if (rst) begin for (i = 0; i < 8; i = i + 1) rf[i] <= 32'h9e3779b9 * (i + 1) + 3; end + else rf[dst_q] <= ld_q ? ld_val_q : res; + end + assign out = res; +endmodule diff --git a/tools/chip-model/rtl/rtl/lane_mul.v b/tools/chip-model/rtl/rtl/lane_mul.v new file mode 100644 index 000000000..0b8b0736c --- /dev/null +++ b/tools/chip-model/rtl/rtl/lane_mul.v @@ -0,0 +1,36 @@ +// 32 x 32 multiplier lane: mul (low 32), mulhi (high 32 of the 64-bit product), mad (src*src2 + dst, low 32). +// op: 0 mul 1 mulhi 2 mad 3 mul. The mad reads three registers. +`include "lane_common.vh" +module lane_mul( + input clk, input rst, + input [1:0] op, input [2:0] dst, input [2:0] src, input [2:0] src2, + input ld_en, input [31:0] ld_val, + output [31:0] out); + reg [31:0] rf [0:7]; + reg [1:0] op_q; reg [2:0] dst_q, src_q, src2_q; reg ld_q; reg [31:0] ld_val_q; + integer i; + always @(posedge clk) begin + if (rst) begin op_q <= 0; dst_q <= 0; src_q <= 0; src2_q <= 0; ld_q <= 0; ld_val_q <= 0; end + else begin op_q <= op; dst_q <= dst; src_q <= src; src2_q <= src2; ld_q <= ld_en; ld_val_q <= ld_val; end + end + wire [31:0] d = rf[dst_q]; + wire [31:0] s = rf[src_q]; + wire [31:0] s2 = rf[src2_q]; + wire mad = (op_q == 2'd2); + wire [31:0] ma = mad ? s : d; + wire [31:0] mb = mad ? s2 : s; + wire [63:0] p = ma * mb; + reg [31:0] res; + always @* begin + case (op_q) + 2'd1: res = p[63:32]; + 2'd2: res = p[31:0] + d; + default: res = p[31:0]; + endcase + end + always @(posedge clk) begin + if (rst) begin for (i = 0; i < 8; i = i + 1) rf[i] <= 32'h9e3779b9 * (i + 1) + 1; end + else rf[dst_q] <= ld_q ? ld_val_q : res; + end + assign out = res; +endmodule diff --git a/tools/chip-model/rtl/rtl/lane_prmt.v b/tools/chip-model/rtl/rtl/lane_prmt.v new file mode 100644 index 000000000..d8d472f28 --- /dev/null +++ b/tools/chip-model/rtl/rtl/lane_prmt.v @@ -0,0 +1,29 @@ +// Byte-permute lane (PRMT): four output bytes, each one of the eight bytes of {s, d}, by a 16-bit selector +// (4 bits per output byte; bit 3 = sign-replicate as on the card). +`include "lane_common.vh" +module lane_prmt( + input clk, input rst, + input [2:0] dst, input [2:0] src, input [15:0] sel, + input ld_en, input [31:0] ld_val, + output [31:0] out); + reg [31:0] rf [0:7]; + reg [2:0] dst_q, src_q; reg [15:0] sel_q; reg ld_q; reg [31:0] ld_val_q; + integer i; + always @(posedge clk) begin + if (rst) begin dst_q <= 0; src_q <= 0; sel_q <= 0; ld_q <= 0; ld_val_q <= 0; end + else begin dst_q <= dst; src_q <= src; sel_q <= sel; ld_q <= ld_en; ld_val_q <= ld_val; end + end + wire [31:0] d = rf[dst_q]; + wire [31:0] s = rf[src_q]; + wire [63:0] bytes = {s, d}; + function [7:0] pick; input [63:0] b; input [3:0] k; + reg [7:0] v; + begin v = b[8*k[2:0] +: 8]; pick = k[3] ? {8{v[7]}} : v; end + endfunction + wire [31:0] res = {pick(bytes, sel_q[15:12]), pick(bytes, sel_q[11:8]), pick(bytes, sel_q[7:4]), pick(bytes, sel_q[3:0])}; + always @(posedge clk) begin + if (rst) begin for (i = 0; i < 8; i = i + 1) rf[i] <= 32'h9e3779b9 * (i + 1) + 2; end + else rf[dst_q] <= ld_q ? ld_val_q : res; + end + assign out = res; +endmodule diff --git a/tools/chip-model/rtl/rtl/scratch8k.v b/tools/chip-model/rtl/rtl/scratch8k.v new file mode 100644 index 000000000..8df1c1346 --- /dev/null +++ b/tools/chip-model/rtl/rtl/scratch8k.v @@ -0,0 +1,19 @@ +// 8 KB random-read scratch (2,048 x 32-bit) as a flop array: one random read per cycle (the op), a write on +// one cycle in eight. A flop array is the pessimistic form of the chip's L1; an SRAM macro reads lower. +module scratch8k( + input clk, input rst, + input [10:0] raddr, input we, input [10:0] waddr, input [31:0] wdata, + output reg [31:0] rdata); + reg [31:0] mem [0:2047]; + reg [10:0] raddr_q, waddr_q; reg we_q; reg [31:0] wdata_q; + integer i; + always @(posedge clk) begin + if (rst) begin raddr_q <= 0; waddr_q <= 0; we_q <= 0; wdata_q <= 0; end + else begin raddr_q <= raddr; waddr_q <= waddr; we_q <= we; wdata_q <= wdata; end + end + always @(posedge clk) begin + if (rst) begin for (i = 0; i < 2048; i = i + 1) mem[i] <= 32'h9e3779b9 * (i + 1); end + else if (we_q) mem[waddr_q] <= wdata_q; + end + always @(posedge clk) rdata <= rst ? 32'd0 : mem[raddr_q]; +endmodule diff --git a/tools/chip-model/rtl/rtl/shfl32.v b/tools/chip-model/rtl/rtl/shfl32.v new file mode 100644 index 000000000..1c88fbf45 --- /dev/null +++ b/tools/chip-model/rtl/rtl/shfl32.v @@ -0,0 +1,36 @@ +// 32-lane xor-mask shuffle over a 1 KB register window (32 lanes x 8 x 32 bits): +// r[dst][lane] ^= r[src][lane ^ m] for every lane, one instruction per cycle (32 lane-ops). +// The network is a 5-stage butterfly (one 2:1 mux per bit per stage). +`include "lane_common.vh" +module shfl32( + input clk, input rst, + input [2:0] dst, input [2:0] src, input [4:0] m, + input ld_en, input [4:0] ld_lane, input [31:0] ld_val, + output [31:0] out); + reg [31:0] rf [0:255]; // rf[lane*8 + reg] + reg [2:0] dst_q, src_q; reg [4:0] m_q, ld_lane_q; reg ld_q; reg [31:0] ld_val_q; + integer i, l; + always @(posedge clk) begin + if (rst) begin dst_q <= 0; src_q <= 0; m_q <= 0; ld_q <= 0; ld_lane_q <= 0; ld_val_q <= 0; end + else begin dst_q <= dst; src_q <= src; m_q <= m; ld_q <= ld_en; ld_lane_q <= ld_lane; ld_val_q <= ld_val; end + end + // explicit butterfly + wire [31:0] b0 [0:31]; wire [31:0] b1 [0:31]; wire [31:0] b2 [0:31]; wire [31:0] b3 [0:31]; wire [31:0] b4 [0:31]; wire [31:0] b5 [0:31]; + genvar g; + generate for (g = 0; g < 32; g = g + 1) begin : bf + assign b0[g] = rf[g*8 + src_q]; + assign b1[g] = m_q[0] ? b0[g ^ 1] : b0[g]; + assign b2[g] = m_q[1] ? b1[g ^ 2] : b1[g]; + assign b3[g] = m_q[2] ? b2[g ^ 4] : b2[g]; + assign b4[g] = m_q[3] ? b3[g ^ 8] : b3[g]; + assign b5[g] = m_q[4] ? b4[g ^ 16] : b4[g]; + end endgenerate + always @(posedge clk) begin + if (rst) begin for (i = 0; i < 256; i = i + 1) rf[i] <= 32'h9e3779b9 * (i + 1) + 5; end + else begin + for (l = 0; l < 32; l = l + 1) + rf[l*8 + dst_q] <= (ld_q && ld_lane_q == l) ? ld_val_q : (rf[l*8 + dst_q] ^ b5[l]); + end + end + assign out = b5[0] ^ b5[17]; +endmodule diff --git a/tools/chip-model/rtl/rtl/tile8.v b/tools/chip-model/rtl/rtl/tile8.v new file mode 100644 index 000000000..5c6b2e94a --- /dev/null +++ b/tools/chip-model/rtl/rtl/tile8.v @@ -0,0 +1,33 @@ +// int8 8x8x8 tile multiply: C[8][8] (int32) += A[8][8] (int8) x B[8][8] (int8): 512 MACs per cycle. +// A and B are loaded from the inputs every cycle; the accumulators clear on clr. +module tile8( + input clk, input rst, input clr, + input [511:0] a_in, input [511:0] b_in, + input [2:0] sel_r, input [2:0] sel_c, + output [31:0] out); + reg [511:0] a_q, b_q; reg clr_q; reg [2:0] sr_q, sc_q; + reg [31:0] c [0:63]; + integer i; + always @(posedge clk) begin + if (rst) begin a_q <= 0; b_q <= 0; clr_q <= 1; sr_q <= 0; sc_q <= 0; end + else begin a_q <= a_in; b_q <= b_in; clr_q <= clr; sr_q <= sel_r; sc_q <= sel_c; end + end + genvar r, cc, k; + generate for (r = 0; r < 8; r = r + 1) begin : row + for (cc = 0; cc < 8; cc = cc + 1) begin : col + wire signed [19:0] dot; + wire signed [15:0] p [0:7]; + for (k = 0; k < 8; k = k + 1) begin : mk + wire signed [7:0] av = a_q[(r*8 + k)*8 +: 8]; + wire signed [7:0] bv = b_q[(k*8 + cc)*8 +: 8]; + assign p[k] = av * bv; + end + assign dot = p[0] + p[1] + p[2] + p[3] + p[4] + p[5] + p[6] + p[7]; + always @(posedge clk) begin + if (rst || clr_q) c[r*8 + cc] <= 32'd0; + else c[r*8 + cc] <= c[r*8 + cc] + {{12{dot[19]}}, dot}; + end + end + end endgenerate + assign out = c[{sr_q, sc_q}]; +endmodule diff --git a/tools/chip-model/rtl/rtl/xbar32.v b/tools/chip-model/rtl/rtl/xbar32.v new file mode 100644 index 000000000..52b7f7398 --- /dev/null +++ b/tools/chip-model/rtl/rtl/xbar32.v @@ -0,0 +1,29 @@ +// 32-lane general crossbar over the same 1 KB window: each lane picks any source lane (a 5-bit select per +// lane), the upper bound on a shuffle's network cost. r[dst][lane] ^= r[src][sel[lane]]. +`include "lane_common.vh" +module xbar32( + input clk, input rst, + input [2:0] dst, input [2:0] src, input [159:0] sel, + input ld_en, input [4:0] ld_lane, input [31:0] ld_val, + output [31:0] out); + reg [31:0] rf [0:255]; + reg [2:0] dst_q, src_q; reg [159:0] sel_q; reg [4:0] ld_lane_q; reg ld_q; reg [31:0] ld_val_q; + integer i, l; + always @(posedge clk) begin + if (rst) begin dst_q <= 0; src_q <= 0; sel_q <= 0; ld_q <= 0; ld_lane_q <= 0; ld_val_q <= 0; end + else begin dst_q <= dst; src_q <= src; sel_q <= sel; ld_q <= ld_en; ld_lane_q <= ld_lane; ld_val_q <= ld_val; end + end + wire [31:0] s [0:31]; + wire [31:0] x [0:31]; + genvar g; + generate for (g = 0; g < 32; g = g + 1) begin : xb + assign s[g] = rf[g*8 + src_q]; + assign x[g] = s[sel_q[5*g +: 5]]; + end endgenerate + always @(posedge clk) begin + if (rst) begin for (i = 0; i < 256; i = i + 1) rf[i] <= 32'h9e3779b9 * (i + 1) + 6; end + else for (l = 0; l < 32; l = l + 1) + rf[l*8 + dst_q] <= (ld_q && ld_lane_q == l) ? ld_val_q : (rf[l*8 + dst_q] ^ x[l]); + end + assign out = x[0] ^ x[17]; +endmodule diff --git a/tools/chip-model/rtl/tb/tb_core_common.vh b/tools/chip-model/rtl/tb/tb_core_common.vh new file mode 100644 index 000000000..38e969897 --- /dev/null +++ b/tools/chip-model/rtl/tb/tb_core_common.vh @@ -0,0 +1,34 @@ +// shared body for the core testbenches: `TOP and `LANES set by the including file +`timescale 1ps/1ps +module tb; + reg clk = 0, rst = 1, run = 0, prog_we = 0, cfg_en = 0; reg [9:0] prog_addr = 0, cfg_n = 0; reg [`IW-1:0] prog_data = 0; reg [63:0] cfg_sel = 0; reg [31:0] cfg_m = 0, cfg_wm = 0, cfg_off = 0, cfg_mask = 0, ld_val = 0; reg [4:0] cfg_r = 0; + wire [31:0] addr, out; + `TOP dut(.clk(clk), .rst(rst), .run(run), .prog_we(prog_we), .prog_addr(prog_addr), .prog_data(prog_data), .cfg_en(cfg_en), .cfg_n(cfg_n), .cfg_sel(cfg_sel), .cfg_m(cfg_m), .cfg_r(cfg_r), .cfg_wm(cfg_wm), .cfg_off(cfg_off), .cfg_mask(cfg_mask), .ld_val(ld_val), .addr(addr), .out(out)); + integer n, cycles, loads, k, roll; reg [31:0] acc = 0; reg [3:0] opc; reg [63:0] w; + // the class v4 draw: add 12, xor 10, mul 8, mad 8, shfl 8, rotl 7, sub 6, mulhi 6, rotr 6, or 4 (sum 75) + function [3:0] draw_op; input integer r; integer x; + begin x = r % 75; if (x < 0) x = -x; + draw_op = (x < 12) ? 4'd0 : (x < 22) ? 4'd2 : (x < 30) ? 4'd6 : (x < 38) ? 4'd8 : (x < 46) ? 4'd9 : (x < 53) ? 4'd4 : (x < 59) ? 4'd1 : (x < 65) ? 4'd7 : (x < 71) ? 4'd5 : 4'd3; + end + endfunction + always #`HALF clk = ~clk; + initial begin + if (!$value$plusargs("cycles=%d", cycles)) cycles = 300; // the core VCDs run to gigabytes per thousand cycles + if (!$value$plusargs("loads=%d", loads)) loads = 0; // 1: one load in 16 instructions + $dumpfile("dump.vcd"); $dumpvars(0, tb.dut); + repeat (4) @(negedge clk); rst = 0; + // the era draw: constants and the program length + @(negedge clk); cfg_en = 1; cfg_n = `NPROG - 1; cfg_sel = `SEL; cfg_m = 32'h9e3779b1; cfg_r = 5'd13; cfg_wm = 32'h0fffffc0; cfg_off = 3; cfg_mask = 32'h0fffffff; + @(negedge clk); cfg_en = 0; + // the program: `NPROG instructions drawn with the class v4 weights + for (k = 0; k < `NPROG; k = k + 1) begin + @(negedge clk); w = {$random, $random}; opc = (loads && (k % 16 == 15)) ? 4'd12 : draw_op($random); + prog_we = 1; prog_addr = k; prog_data = {w[`IW-1:4], opc}; + end + @(negedge clk); prog_we = 0; run = 1; + for (n = 0; n < cycles; n = n + 1) begin + @(negedge clk); ld_val = $random; acc = acc ^ out ^ addr; + end + $display("CHECKSUM %08x", acc); $finish; + end +endmodule diff --git a/tools/chip-model/rtl/tb/tb_core_legacy.vh b/tools/chip-model/rtl/tb/tb_core_legacy.vh new file mode 100644 index 000000000..4fecabfee --- /dev/null +++ b/tools/chip-model/rtl/tb/tb_core_legacy.vh @@ -0,0 +1,34 @@ +// shared body for the core testbenches: `TOP and `LANES set by the including file +`timescale 1ps/1ps +module tb; + reg clk = 0, rst = 1, run = 0, prog_we = 0, cfg_en = 0; reg [7:0] prog_addr = 0, cfg_n = 0; reg [31:0] prog_data = 0, cfg_m = 0, cfg_wm = 0, cfg_off = 0, cfg_mask = 0, ld_val = 0; reg [4:0] cfg_r = 0; + wire [31:0] addr, out; + `TOP dut(.clk(clk), .rst(rst), .run(run), .prog_we(prog_we), .prog_addr(prog_addr), .prog_data(prog_data), .cfg_en(cfg_en), .cfg_n(cfg_n), .cfg_m(cfg_m), .cfg_r(cfg_r), .cfg_wm(cfg_wm), .cfg_off(cfg_off), .cfg_mask(cfg_mask), .ld_val(ld_val), .addr(addr), .out(out)); + integer n, cycles, loads, k, roll; reg [31:0] acc = 0; reg [3:0] opc; reg [31:0] w; + // the class v4 draw: add 12, xor 10, mul 8, mad 8, shfl 8, rotl 7, sub 6, mulhi 6, rotr 6, or 4 (sum 75) + function [3:0] draw_op; input integer r; integer x; + begin x = r % 75; if (x < 0) x = -x; + draw_op = (x < 12) ? 4'd0 : (x < 22) ? 4'd2 : (x < 30) ? 4'd6 : (x < 38) ? 4'd8 : (x < 46) ? 4'd9 : (x < 53) ? 4'd4 : (x < 59) ? 4'd1 : (x < 65) ? 4'd7 : (x < 71) ? 4'd5 : 4'd3; + end + endfunction + always #`HALF clk = ~clk; + initial begin + if (!$value$plusargs("cycles=%d", cycles)) cycles = 300; + if (!$value$plusargs("loads=%d", loads)) loads = 0; // 1: one load in 16 instructions + $dumpfile("dump.vcd"); $dumpvars(0, tb.dut); + repeat (4) @(negedge clk); rst = 0; + // the era draw: constants and the program length + @(negedge clk); cfg_en = 1; cfg_n = 8'd255; cfg_m = 32'h9e3779b1; cfg_r = 5'd13; cfg_wm = 32'h0fffffc0; cfg_off = 3; cfg_mask = 32'h0fffffff; + @(negedge clk); cfg_en = 0; + // the program: 256 instructions drawn with the class v4 weights + for (k = 0; k < 256; k = k + 1) begin + @(negedge clk); w = $random; opc = (loads && (k % 16 == 15)) ? 4'd12 : draw_op($random); + prog_we = 1; prog_addr = k; prog_data = {w[31:4], opc}; + end + @(negedge clk); prog_we = 0; run = 1; + for (n = 0; n < cycles; n = n + 1) begin + @(negedge clk); ld_val = $random; acc = acc ^ out ^ addr; + end + $display("CHECKSUM %08x", acc); $finish; + end +endmodule diff --git a/tools/chip-model/rtl/tb/tb_core_v6_32.v b/tools/chip-model/rtl/tb/tb_core_v6_32.v new file mode 100644 index 000000000..2c63f5096 --- /dev/null +++ b/tools/chip-model/rtl/tb/tb_core_v6_32.v @@ -0,0 +1,3 @@ +`define TOP core_v6_32 +`define HALF 750 +`include "tb_core_legacy.vh" diff --git a/tools/chip-model/rtl/tb/tb_core_v6_32all.v b/tools/chip-model/rtl/tb/tb_core_v6_32all.v new file mode 100644 index 000000000..9509ee182 --- /dev/null +++ b/tools/chip-model/rtl/tb/tb_core_v6_32all.v @@ -0,0 +1,6 @@ +`define TOP core_v6_32all +`define HALF 750 +`define IW 40 +`define NPROG 1024 +`define SEL 64'hc5e7092b4d6f81a3 +`include "tb_core_common.vh" diff --git a/tools/chip-model/rtl/tb/tb_core_v6_32r16.v b/tools/chip-model/rtl/tb/tb_core_v6_32r16.v new file mode 100644 index 000000000..696d54346 --- /dev/null +++ b/tools/chip-model/rtl/tb/tb_core_v6_32r16.v @@ -0,0 +1,6 @@ +`define TOP core_v6_32r16 +`define HALF 750 +`define IW 32 +`define NPROG 256 +`define SEL 64'hfedcba9876543210 +`include "tb_core_common.vh" diff --git a/tools/chip-model/rtl/tb/tb_core_v6_8.v b/tools/chip-model/rtl/tb/tb_core_v6_8.v new file mode 100644 index 000000000..b45dfef48 --- /dev/null +++ b/tools/chip-model/rtl/tb/tb_core_v6_8.v @@ -0,0 +1,6 @@ +`define TOP core_v6_8 +`define HALF 750 +`define IW 32 +`define NPROG 256 +`define SEL 64'hfedcba9876543210 +`include "tb_core_common.vh" diff --git a/tools/chip-model/rtl/tb/tb_core_v6_8i1k.v b/tools/chip-model/rtl/tb/tb_core_v6_8i1k.v new file mode 100644 index 000000000..738a4815b --- /dev/null +++ b/tools/chip-model/rtl/tb/tb_core_v6_8i1k.v @@ -0,0 +1,6 @@ +`define TOP core_v6_8i1k +`define HALF 750 +`define IW 32 +`define NPROG 1024 +`define SEL 64'hfedcba9876543210 +`include "tb_core_common.vh" diff --git a/tools/chip-model/rtl/tb/tb_core_v6_8r64.v b/tools/chip-model/rtl/tb/tb_core_v6_8r64.v new file mode 100644 index 000000000..5efbee6f6 --- /dev/null +++ b/tools/chip-model/rtl/tb/tb_core_v6_8r64.v @@ -0,0 +1,6 @@ +`define TOP core_v6_8r64 +`define HALF 750 +`define IW 40 +`define NPROG 256 +`define SEL 64'hfedcba9876543210 +`include "tb_core_common.vh" diff --git a/tools/chip-model/rtl/tb/tb_core_v6_8sel.v b/tools/chip-model/rtl/tb/tb_core_v6_8sel.v new file mode 100644 index 000000000..65094edc8 --- /dev/null +++ b/tools/chip-model/rtl/tb/tb_core_v6_8sel.v @@ -0,0 +1,6 @@ +`define TOP core_v6_8sel +`define HALF 750 +`define IW 32 +`define NPROG 256 +`define SEL 64'hc5e7092b4d6f81a3 +`include "tb_core_common.vh" diff --git a/tools/chip-model/rtl/tb/tb_lane_arx.v b/tools/chip-model/rtl/tb/tb_lane_arx.v new file mode 100644 index 000000000..b4b618ab7 --- /dev/null +++ b/tools/chip-model/rtl/tb/tb_lane_arx.v @@ -0,0 +1,20 @@ +`timescale 1ps/1ps +module tb; + reg clk = 0, rst = 1; reg [2:0] op = 0, dst = 0, src = 0; reg [4:0] rot = 0; reg ld_en = 0; reg [31:0] ld_val = 0; + wire [31:0] out; + lane_arx dut(.clk(clk), .rst(rst), .op(op), .dst(dst), .src(src), .rot(rot), .ld_en(ld_en), .ld_val(ld_val), .out(out)); + integer n, fixed_op, cycles; reg [31:0] acc = 0; + always #500 clk = ~clk; + initial begin + if (!$value$plusargs("op=%d", fixed_op)) fixed_op = -1; + if (!$value$plusargs("cycles=%d", cycles)) cycles = 4000; + $dumpfile("dump.vcd"); $dumpvars(0, tb.dut); + repeat (4) @(negedge clk); rst = 0; + for (n = 0; n < cycles; n = n + 1) begin + @(negedge clk); + op = (fixed_op < 0) ? $random : fixed_op; dst = $random; src = $random; rot = $random; + ld_en = (($random & 15) == 0); ld_val = $random; acc = acc ^ out; + end + $display("CHECKSUM %08x", acc); $finish; + end +endmodule diff --git a/tools/chip-model/rtl/tb/tb_lane_fold.v b/tools/chip-model/rtl/tb/tb_lane_fold.v new file mode 100644 index 000000000..644de5980 --- /dev/null +++ b/tools/chip-model/rtl/tb/tb_lane_fold.v @@ -0,0 +1,21 @@ +`timescale 1ps/1ps +module tb; + reg clk = 0, rst = 1; reg [2:0] dst = 0, src = 0; reg cfg_en = 0; reg [31:0] cfg_m = 0, cfg_wm = 0, cfg_off = 0, cfg_mask = 0; reg [4:0] cfg_r = 0; + reg ld_en = 0; reg [31:0] ld_val = 0; wire [31:0] out; + lane_fold dut(.clk(clk), .rst(rst), .dst(dst), .src(src), .cfg_en(cfg_en), .cfg_m(cfg_m), .cfg_r(cfg_r), .cfg_wm(cfg_wm), .cfg_off(cfg_off), .cfg_mask(cfg_mask), .ld_en(ld_en), .ld_val(ld_val), .out(out)); + integer n, cycles; reg [31:0] acc = 0; + always #750 clk = ~clk; + initial begin + if (!$value$plusargs("cycles=%d", cycles)) cycles = 4000; + $dumpfile("dump.vcd"); $dumpvars(0, tb.dut); + repeat (4) @(negedge clk); rst = 0; + // one era draw: the constants are written once and then static, as on the chip + @(negedge clk); cfg_en = 1; cfg_m = 32'h9e3779b1; cfg_r = 5'd13; cfg_wm = 32'h0fffffc0; cfg_off = 32'h3; cfg_mask = 32'h0fffffff; + @(negedge clk); cfg_en = 0; + for (n = 0; n < cycles; n = n + 1) begin + @(negedge clk); + dst = $random; src = $random; ld_en = (($random & 3) == 0); ld_val = $random; acc = acc ^ out; + end + $display("CHECKSUM %08x", acc); $finish; + end +endmodule diff --git a/tools/chip-model/rtl/tb/tb_lane_lop3.v b/tools/chip-model/rtl/tb/tb_lane_lop3.v new file mode 100644 index 000000000..3bccc1670 --- /dev/null +++ b/tools/chip-model/rtl/tb/tb_lane_lop3.v @@ -0,0 +1,18 @@ +`timescale 1ps/1ps +module tb; + reg clk = 0, rst = 1; reg [2:0] dst = 0, src = 0, src2 = 0; reg [7:0] lut = 0; reg ld_en = 0; reg [31:0] ld_val = 0; + wire [31:0] out; + lane_lop3 dut(.clk(clk), .rst(rst), .dst(dst), .src(src), .src2(src2), .lut(lut), .ld_en(ld_en), .ld_val(ld_val), .out(out)); + integer n, cycles; reg [31:0] acc = 0; + always #500 clk = ~clk; + initial begin + if (!$value$plusargs("cycles=%d", cycles)) cycles = 4000; + $dumpfile("dump.vcd"); $dumpvars(0, tb.dut); + repeat (4) @(negedge clk); rst = 0; + for (n = 0; n < cycles; n = n + 1) begin + @(negedge clk); + dst = $random; src = $random; src2 = $random; lut = $random; ld_en = (($random & 15) == 0); ld_val = $random; acc = acc ^ out; + end + $display("CHECKSUM %08x", acc); $finish; + end +endmodule diff --git a/tools/chip-model/rtl/tb/tb_lane_mul.v b/tools/chip-model/rtl/tb/tb_lane_mul.v new file mode 100644 index 000000000..f5f37f7a0 --- /dev/null +++ b/tools/chip-model/rtl/tb/tb_lane_mul.v @@ -0,0 +1,20 @@ +`timescale 1ps/1ps +module tb; + reg clk = 0, rst = 1; reg [1:0] op = 0; reg [2:0] dst = 0, src = 0, src2 = 0; reg ld_en = 0; reg [31:0] ld_val = 0; + wire [31:0] out; + lane_mul dut(.clk(clk), .rst(rst), .op(op), .dst(dst), .src(src), .src2(src2), .ld_en(ld_en), .ld_val(ld_val), .out(out)); + integer n, fixed_op, cycles; reg [31:0] acc = 0; + always #750 clk = ~clk; + initial begin + if (!$value$plusargs("op=%d", fixed_op)) fixed_op = -1; + if (!$value$plusargs("cycles=%d", cycles)) cycles = 4000; + $dumpfile("dump.vcd"); $dumpvars(0, tb.dut); + repeat (4) @(negedge clk); rst = 0; + for (n = 0; n < cycles; n = n + 1) begin + @(negedge clk); + op = (fixed_op < 0) ? $random : fixed_op; dst = $random; src = $random; src2 = $random; + ld_en = (($random & 15) == 0); ld_val = $random; acc = acc ^ out; + end + $display("CHECKSUM %08x", acc); $finish; + end +endmodule diff --git a/tools/chip-model/rtl/tb/tb_lane_prmt.v b/tools/chip-model/rtl/tb/tb_lane_prmt.v new file mode 100644 index 000000000..30764b26c --- /dev/null +++ b/tools/chip-model/rtl/tb/tb_lane_prmt.v @@ -0,0 +1,18 @@ +`timescale 1ps/1ps +module tb; + reg clk = 0, rst = 1; reg [2:0] dst = 0, src = 0; reg [15:0] sel = 0; reg ld_en = 0; reg [31:0] ld_val = 0; + wire [31:0] out; + lane_prmt dut(.clk(clk), .rst(rst), .dst(dst), .src(src), .sel(sel), .ld_en(ld_en), .ld_val(ld_val), .out(out)); + integer n, cycles; reg [31:0] acc = 0; + always #500 clk = ~clk; + initial begin + if (!$value$plusargs("cycles=%d", cycles)) cycles = 4000; + $dumpfile("dump.vcd"); $dumpvars(0, tb.dut); + repeat (4) @(negedge clk); rst = 0; + for (n = 0; n < cycles; n = n + 1) begin + @(negedge clk); + dst = $random; src = $random; sel = $random & 16'h7777; ld_en = (($random & 15) == 0); ld_val = $random; acc = acc ^ out; + end + $display("CHECKSUM %08x", acc); $finish; + end +endmodule diff --git a/tools/chip-model/rtl/tb/tb_scratch8k.v b/tools/chip-model/rtl/tb/tb_scratch8k.v new file mode 100644 index 000000000..ed118a378 --- /dev/null +++ b/tools/chip-model/rtl/tb/tb_scratch8k.v @@ -0,0 +1,17 @@ +`timescale 1ps/1ps +module tb; + reg clk = 0, rst = 1; reg [10:0] raddr = 0, waddr = 0; reg we = 0; reg [31:0] wdata = 0; wire [31:0] rdata; + scratch8k dut(.clk(clk), .rst(rst), .raddr(raddr), .we(we), .waddr(waddr), .wdata(wdata), .rdata(rdata)); + integer n, cycles; reg [31:0] acc = 0; + always #750 clk = ~clk; + initial begin + if (!$value$plusargs("cycles=%d", cycles)) cycles = 3000; + $dumpfile("dump.vcd"); $dumpvars(0, tb.dut); + repeat (4) @(negedge clk); rst = 0; + for (n = 0; n < cycles; n = n + 1) begin + @(negedge clk); + raddr = $random; waddr = $random; we = (($random & 7) == 0); wdata = $random; acc = acc ^ rdata; + end + $display("CHECKSUM %08x", acc); $finish; + end +endmodule diff --git a/tools/chip-model/rtl/tb/tb_shfl32.v b/tools/chip-model/rtl/tb/tb_shfl32.v new file mode 100644 index 000000000..2cf58a69f --- /dev/null +++ b/tools/chip-model/rtl/tb/tb_shfl32.v @@ -0,0 +1,18 @@ +`timescale 1ps/1ps +module tb; + reg clk = 0, rst = 1; reg [2:0] dst = 0, src = 0; reg [4:0] m = 0; reg ld_en = 0; reg [4:0] ld_lane = 0; reg [31:0] ld_val = 0; + wire [31:0] out; + shfl32 dut(.clk(clk), .rst(rst), .dst(dst), .src(src), .m(m), .ld_en(ld_en), .ld_lane(ld_lane), .ld_val(ld_val), .out(out)); + integer n, cycles; reg [31:0] acc = 0; + always #500 clk = ~clk; + initial begin + if (!$value$plusargs("cycles=%d", cycles)) cycles = 3000; + $dumpfile("dump.vcd"); $dumpvars(0, tb.dut); + repeat (4) @(negedge clk); rst = 0; + for (n = 0; n < cycles; n = n + 1) begin + @(negedge clk); + dst = $random; src = $random; m = $random; ld_en = 1; ld_lane = $random; ld_val = $random; acc = acc ^ out; + end + $display("CHECKSUM %08x", acc); $finish; + end +endmodule diff --git a/tools/chip-model/rtl/tb/tb_tile8.v b/tools/chip-model/rtl/tb/tb_tile8.v new file mode 100644 index 000000000..f9b66a2dc --- /dev/null +++ b/tools/chip-model/rtl/tb/tb_tile8.v @@ -0,0 +1,18 @@ +`timescale 1ps/1ps +module tb; + reg clk = 0, rst = 1, clr = 0; reg [511:0] a_in = 0, b_in = 0; reg [2:0] sel_r = 0, sel_c = 0; wire [31:0] out; + tile8 dut(.clk(clk), .rst(rst), .clr(clr), .a_in(a_in), .b_in(b_in), .sel_r(sel_r), .sel_c(sel_c), .out(out)); + integer n, j, cycles; reg [31:0] acc = 0; + always #1000 clk = ~clk; + initial begin + if (!$value$plusargs("cycles=%d", cycles)) cycles = 3000; + $dumpfile("dump.vcd"); $dumpvars(0, tb.dut); + repeat (4) @(negedge clk); rst = 0; + for (n = 0; n < cycles; n = n + 1) begin + @(negedge clk); + for (j = 0; j < 16; j = j + 1) begin a_in[j*32 +: 32] = $random; b_in[j*32 +: 32] = $random; end + clr = (($random & 63) == 0); sel_r = $random; sel_c = $random; acc = acc ^ out; + end + $display("CHECKSUM %08x", acc); $finish; + end +endmodule diff --git a/tools/chip-model/rtl/tb/tb_xbar32.v b/tools/chip-model/rtl/tb/tb_xbar32.v new file mode 100644 index 000000000..6440b2340 --- /dev/null +++ b/tools/chip-model/rtl/tb/tb_xbar32.v @@ -0,0 +1,18 @@ +`timescale 1ps/1ps +module tb; + reg clk = 0, rst = 1; reg [2:0] dst = 0, src = 0; reg [159:0] sel = 0; reg ld_en = 0; reg [4:0] ld_lane = 0; reg [31:0] ld_val = 0; + wire [31:0] out; + xbar32 dut(.clk(clk), .rst(rst), .dst(dst), .src(src), .sel(sel), .ld_en(ld_en), .ld_lane(ld_lane), .ld_val(ld_val), .out(out)); + integer n, cycles; reg [31:0] acc = 0; + always #500 clk = ~clk; + initial begin + if (!$value$plusargs("cycles=%d", cycles)) cycles = 3000; + $dumpfile("dump.vcd"); $dumpvars(0, tb.dut); + repeat (4) @(negedge clk); rst = 0; + for (n = 0; n < cycles; n = n + 1) begin + @(negedge clk); + dst = $random; src = $random; sel = {$random, $random, $random, $random, $random}; ld_en = 1; ld_lane = $random; ld_val = $random; acc = acc ^ out; + end + $display("CHECKSUM %08x", acc); $finish; + end +endmodule