diff --git a/docs/analysis/class-v6/floor/shadow-k.md b/docs/analysis/class-v6/floor/shadow-k.md index a5c561c52..1b8667439 100644 --- a/docs/analysis/class-v6/floor/shadow-k.md +++ b/docs/analysis/class-v6/floor/shadow-k.md @@ -104,7 +104,7 @@ stock and at the 1,300 MHz lock, and 6.9 pJ per counted op on the Apple M5 Max ( | lop3 (8-bit truth table) | 7,274 | 1.42 | 0.99 | 0.72 | 0.52 | 24.1 / 13.0 | 0.030 / 0.055 | | 0.11 | | | 32-lane xor-mask shuffle (butterfly over the 1 KB window), per lane-op | 181,580 | 1.24 | 0.87 | 0.63 | 0.45 | 55.8 / 29.4 | 0.011 / 0.021 | | 0.042 | routed with SPEF; 15:1x UK | | 32-lane general crossbar, per lane-op | ROW_XBAR | -| 8 KB scratch, one random read (flop array: the pessimistic form) | ROW_SCRATCH | +| 8 KB scratch, one random 32-bit read of a 2,048 x 32 flop array (the pessimistic form of a chip's L1; an SRAM macro reads lower) | 868,159 | 207 | 145 | 104 | 75 | 2,400 / 1,400 per L2 hit | 0.043 / 0.074 | | 0.15 | routed with SPEF, 19:4x UK; the card's shared-memory read is unmeasured (owed) | | int8 8x8x8 tile, per MAC | ROW_TILE | Reading the floors: a lane's add costs the chip about 2.2 pJ at ASAP7 and 1.1 at N3, against the 5090's 6.2 @@ -138,7 +138,7 @@ draws 13.9 mW, the run phase 36.9 mW for 8 lanes at 1.5 ns). | core, 8 lanes, 32 registers | synthesis only (no wires, no clock tree) | 186,443 | 6.9 | 4.8 | 3.5 | 2.5 | 11.3 / 6.2 / 6.9 | 0.31 / 0.56 / 0.50 | 0.22 / 0.40 / 0.36 | 1.1 | synthesised; 14:0x UK | | of which the sequential term (register file, imem and IR clock pins, no clock gating) | | | 2.4 | 1.7 | 1.2 | 0.9 | | | | | | | of which the units, the read muxes and the butterfly | | | 4.5 | 3.1 | 2.3 | 1.6 | | | | | | -| core, 8 lanes, 32 registers | placed and routed, SPEF | ROW_CORE8_PLACED | +| core, 8 lanes, 32 registers, ungated | placed and routed, SPEF, clock tree (one run length of 300 cycles, the load phase subtracted, about plus or minus 10 percent) | 444,478 | 11.3 | 7.9 | 5.6 | 4.1 | 11.3 / 6.2 / 6.9 | 0.50 / 0.91 / 0.82 | 0.36 / 0.66 / 0.59 | 1.8 | placed 16:0x UK on a rented pod; +64 percent over synthesis (wires, and a clock tree of 2.5 pJ per lane-op that gating removes) | | core, 32 lanes, 32 registers | synthesis only (steady state from 150 and 400 run cycles) | 600,381 | 5.55 | 3.9 | 2.8 | 2.0 | 11.3 / 6.2 / 6.9 | 0.25 / 0.45 / 0.41 | 0.18 / 0.32 / 0.29 | 0.90 | synthesised; 15:2x UK | | core, 32 lanes, 16 registers | synthesis only | 443,258 | 4.2 | 2.9 | 2.1 | 1.5 | 11.3 / 6.2 / 6.9 | 0.18 / 0.34 / 0.30 | 0.13 / 0.24 / 0.22 | 0.68 | synthesised; one run length, about plus or minus 10 percent; 15:0x UK | | the bare ARX lane (section 3, the floor) | routed | 11,631 | 2.2 | 1.5 | 1.1 | 0.8 | 11.3 / 6.2 / 6.9 | 0.10 / 0.18 / 0.16 | 0.07 / 0.13 / 0.11 | 0.35 | the lower bound | @@ -203,6 +203,128 @@ takes back on the card's own node (3.6x to 2.4x), which is the design. Node-for- k 0.78 and the 64-register core at 1.09, so "near 0.9" is reached node-for-node by the window alone; what it does not survive is the node step a chip project would buy (an N3 core gives back 0.4x, an N2 core 0.8x). +### 4c. The adversary's 64-register core: is the window a defence? (the coordinator's order, 15:3x UK) + +Every row in this section is a MODEL of a chip core, never a lower bound on what a chip maker can build; the +synthesis gives the cost of the circuit as drawn, and a better circuit is always possible. + +**The live-state analysis** (`tools/chip-model/rtl/flow/livestate.py`, run on build-3: programs drawn as the +core testbench draws them, the class v4 op weights, a load on one instruction in 16 as the dependent memory wait, +dst and src uniform over the window, the result fold reading every register at the end of the block; 64 drawn +programs, 1,024 waits per row): + +| Window R | Live values at a wait (mean, min to max) | Of which necessary (reach a later address or the result, transitively) | Dead writes per block | +|---|---|---|---| +| 8 (the class ISA) | 7.0 of 8 (6 to 7) | 6.9 | 2.9 percent | +| 32 | 30.5 of 32 (29 to 31) | 30.0 | 2.9 percent | +| 64 | 61.7 of 64 (59 to 63) | 61.0 | 2.8 percent | +| 64 at a 1,024-instruction block | 61.5 of 64 (58 to 63) | 60.6 | 3.1 percent | + +So under a fold that reads every register, 95 percent of the window is live AND necessary across every memory +wait: the adversary cannot shrink the state it keeps by liveness, and recomputing a value instead of keeping it +costs the dependent chain that produced it (every value feeds the result transitively). The window is a +defence ONLY because of the fold rule; a fold that read 8 of the 64 registers would let the chip drop the rest +(the dead fraction would rise toward the fraction never read before the fold), so the fold-reads-all rule is the +design rule that goes with the window. + +**What the adversary can do with the state it must keep** is make it cheaper per access, not smaller. The +GPU-shaped row (4a, 64 registers in flops, every flop clocked every cycle, three 64:1 read muxes) is 9.7 pJ per +lane-op at ASAP7. The forms a chip maker would use: + +| Form of the 64-register state (per lane, 256 bytes) | pJ per lane-op ASAP7 | N3 | k at the lock (N3) | Label | +|---|---|---|---|---| +| flops, no clock gating, 64:1 read muxes (the 4a row) | 9.7 | 4.9 | 0.78 | synthesised; a model | +| flops with the register-file clock gated (one of 64 registers written per cycle; the ICG cells allowed back in and inferred by Yosys, 320 gates) | 6.2 (225,441 cells; sequential 0.2) | 3.1 | 0.50 (0.70 node-for-node) | synthesised 16:0x UK; a model | +| the same gating on the 32-register base, for the penalty | 4.5 (156,833 cells; sequential 0.15) | 2.3 | 0.37 (0.51 node-for-node) | synthesised 16:0x UK; a model | +| the gated 32-register base PLACED AND ROUTED (SPEF, clock tree, 379,633 cells; steady state solved from 150 and 600 run cycles) | 6.7 (+49 percent over synthesis) | 3.4 | 0.54 (0.76 node-for-node) | placed 19:5x UK; the GDDR7 board at the lock 2.4x node-for-node, 2.8x a node ahead: the morning's headline figures to the digit | +| the gated 64-register window core PLACED AND ROUTED (563,339 cells; parasitics estimated from global routing, the SPEF lost to a full disk; steady state solved from 150 and 600 run cycles; plus or minus 15 percent) | 9.45 (+52 percent over synthesis; N5 6.6, N2 3.4) | 4.8 | 0.77 (1.07 node-for-node, 0.55 at N2) | placed 21:3x UK; the GDDR7 board at the lock 2.0x node-for-node, 2.45x a node ahead, 2.9x two ahead; the window's residual against the placed base +2.75 pJ per lane-op, +0.23 of k at the lock | +| latch-based register file (the clocked element halved; about 30 percent under the gated flop file, approximate) | about 0.7 x the gated row | | | modelled | +| SRAM-banked state shared across time-multiplexed lanes (one execution port serving many lanes' streams in turn, each lane's 256 bytes in a bank): modelled 8.5 to 10.5 from the access energies; BUILT as `core_tm` (8 lanes x 64 registers in banks, one port, round-robin, ungated): 8.8 pJ per lane-op synthesised (173,426 cells), against the SIMD ungated 9.7 | 8.8 (built) | 4.4 | 0.71 | synthesised 16:2x UK: port sharing saves 0.9 pJ of units and the bank-select muxes take most of it back; NOT the lever; the multi-family adversary lane's macro-window core reads 5.9 (its FakeRAM term modelled 2.0 to 7.0 pJ per access), within 5 percent of the gated flop row, so the file's form is not the lever either | +| values recomputed instead of kept | not available: 95 percent of the window is necessary (above) | | | measured on drawn programs | + +The defence, then, is the penalty that remains after the adversary's best form: the gated 64-register file +against the gated 32-register file (the two synthesised rows above when they land; measured: 6.2 against 4.5 pJ per lane-op at ASAP7 synthesised, a penalty of 1.7 pJ; PLACED 9.45 against 6.7, a +penalty of 2.75 pJ, 1.4 at N3, +0.23 of k at the lock, 0.54 to 0.77 at N3 and 0.76 to 1.07 node-for-node; the +analytic estimate had been 1.5 pJ). So on placed rows the window takes the GDDR7 board at the lock from 2.4x to +2.0x node-for-node and from 2.8x to 2.45x a node ahead, for at most 5 percent per load on the card. Two corrections this +forces: the honest adversary's BASE core is the gated one (k 0.37 at N3, 0.51 node-for-node), under the ungated +0.56 and 0.78 of section 4, which are the GPU-shaped core a maker would not build; and placement costs more than +the +20 to +40 percent estimated (the ungated placed base reads 11.3 against 6.9 pJ: wires plus a 2.5 pJ clock tree +that gating removes), so the placed gated rows (on a rented pod, 17:30 UK) are the figures to serve. On the 32-lane core the same +penalty applies per lane (the register file does not amortise), so the window moves the 32-lane core from k 0.45 +to about 0.57 at the lock at N3 (0.63 to about 0.80 node-for-node). + +**The GPU side** (the hash lane's hand): the compiled allocation of the 64-register measurement pack (ptxas +registers per thread, local-memory spill bytes, occupancy) and the rate beside the 8-register base, clock +18:00 UK; until then the modelled reading stands: about 110 of 255 registers per thread, occupancy about half, +the rate expected to hold under the latency-bound chain (the 5090 hides about 330,000 ops per hash before compute +binds) and the energy to move little, the per-lane register traffic the unmeasured term. + +The measured GPU side (the hash lane, 16:1x to 16:4x UK, RunPod secure pods, driver 580, the kit worker, 250 x +2^24, nvidia-smi 1 Hz; ptxas from nvcc 12.8 -Xptxas -v on the pack's kernel; pods destroyed, USD 1.22): + +| Card, pack | MH/s | W | microjoules per hash | registers per thread (ptxas) | spill | blocks per SM (occupancy) | per load | Label | +|---|---|---|---|---|---|---|---|---| +| 5090, the base (mx8-devnet-epoch0) | 141.74 | 303.1 | 2.139 | 30 | 0 B | 24 (4,080 warps) | 16.7 nJ | measured | +| 5090, the window, arithmetic-only (hl-reg64, twice the base's work by construction) | 80.38 | 308.6 | 3.839 | 96 | 0 B | 20 (83 percent) | 15.0 nJ | measured: level per unit of work | +| 5090, the window, full chain (hl-reg64c: every load's address mixes all 64 registers) | 70.96 | 320.3 | 4.513 | 88 | 0 B | 20 (83 percent) | 17.6 nJ (+5 percent) | measured | +| 4090, the base | 62.67 | 208.9 | 3.333 | 29 | 0 B | 24 | 26.0 nJ | measured | +| 4090, the window, arithmetic-only | 31.57 | 210.3 | 6.663 | 104 | 0 B | 16 (67 percent) | 26.0 nJ | measured: level | +| 4090, the window, full chain | 31.38 | 216.5 | 6.898 | 87 | 0 B | 20 (83 percent) | 27.0 nJ (+4 percent) | measured | + +So the card's side of the window defence is at most 5 percent per load: no spill on either card in either form, 88 +to 104 registers per thread, occupancy 67 to 83 percent, and the rate per unit of work held within 5 percent under +the latency-bound chain. The sound class form is the full chain (the arithmetic-only fold fails the liveness rule; +class string `+reg64c`, pack hl-v6-win with `check_window_liveness` in its suite). The chip's side (this section's +gated rows) therefore carries the whole defence. + +### 4d. The connected-state variant (cs64s27x16, the connected-state lane's structure) on the adversary's core + +The connected-state lane's program (a 64-register window; per step a load whose address register is the previous +block's last dst, the word landing in m_j, then a 27-instruction block whose first instruction reads m_j and every +later one draws its src from the block's last four dsts, the last instruction injecting; 16 steps of text, 448 +instructions, 16 passes per block; its liveness tool: 63 of 64 live at every address, about 11 registers in the +per-step dependent chain, about 20 touched per block) priced on the gated 64-register core with a 512-entry imem, +the program drawn by those rules in the testbench (`CS` mode of `tb_core_common.vh`), synthesis-only: + +| Row | Cells | pJ per lane-op ASAP7 | N5 | N3 | N2 | k at the lock N5 / N3 / N2 | k at stock N5 / N3 | +|---|---|---|---|---|---|---|---| +| cs64s27x16 on the gated 64-register core, 512 imem | 253,059 | 6.3 | 4.4 | 3.2 | 2.3 | 0.71 / 0.51 / 0.37 | 0.39 / 0.28 | +| the class v4 draw on the gated 64-register core, 256 imem (4c) | 225,441 | 6.2 | 4.3 | 3.1 | 2.2 | 0.70 / 0.50 / 0.36 | 0.38 / 0.27 | + +The chip's shadow for the program is 55,296 x 3.2 pJ = 0.18 microjoules per hash at N3 (0.24 node-for-node) +against 0.13 for the genesis window on the same core; the card pays +0.6 percent for the window on the 5090 (the +connected-state lane's measurement). The structure's other knobs do not reach the chip: the 16-pass loop and the +448 text cost the shared imem about 0.1 pJ per lane-op, hot-set banking is not needed (the gated file charges only +the written register), and the chain's width sets lane count, which is free. On the GDDR7 board at the lock the +window moves the chip's edge by about 1.1x (3.6x to 3.3x node-for-node), under the 1.25x gate that lane set. + +### 4e. The mixed-resource lane's FP32 units (class-v6-mixedfp) on the adversary's lane + +The mixed-resource lane's candidate adds four FP32 families (fadd, fmul, ffma, fcvt) to the shadow's draw, every +result injected by xor, with inputs masked to a 7-bit exponent range (never zero, denormal, NaN or Inf). The +adversary's simplified units (`rtl/fp32_units.v`): an FMA with the 24 x 24 mantissa multiplier, a 100-bit +alignment window, a full normaliser and RNE; a separate adder and multiplier; the int32 to float converter; the +exponent path narrowed to the range; no NaN, Inf, denormal or flag logic. The lane: an 8 x 32-bit window, the +four units, `d ^= bits(result)`. Routed with SPEF, random-input VCD, 42,936 cells, a 2 ns clock. + +| Op (every unit evaluating each cycle: an UPPER bound per op, no operand isolation) | pJ per op ASAP7 | N5 | N3 | N2 | 5090 fp32_fma stock / lock | k at N3 vs stock / lock | +|---|---|---|---|---|---|---| +| fadd | 6.6 | 4.6 | 3.3 | 2.4 | 9.2 / 5.2 | 0.36 / 0.64 | +| fmul | 7.0 | 4.9 | 3.5 | 2.5 | 9.2 / 5.2 | 0.38 / 0.68 | +| ffma | 6.7 | 4.7 | 3.4 | 2.4 | 9.2 / 5.2 | 0.37 / 0.65 | +| fcvt | 6.9 | 4.8 | 3.5 | 2.5 | 9.2 / 5.2 (cvt unmeasured on the card) | 0.38 / 0.67 | +| random mix | 7.1 | 4.9 | 3.6 | 2.6 | | 0.39 / 0.68 | + +Reading: the four read alike because all four units switch every cycle on the same operands, so each row is the +upper bound for its op (a chip isolates the idle units; by cell share about ffma 3.5 to 4, fmul 2.5, fadd 2, fcvt +1 pJ at ASAP7, approximate). Even on the upper bound the FP family is the chip's dearest per op relative to the +card: k 0.65 at N3 at the lock against 0.18 for the integer ARX lane floor, because the card does an FMA for 5.2 +pJ (under its own int add at 6.2) while the chip's multiply, alignment and normaliser cost about three int ops. +On the units' floors the shadow's k_eff rises from 0.097 (class v4) to about 0.14 at the fp12 mix and 0.17 at +fp24 (0.20 and 0.24 with isolation taken as half), the core's per-op overhead on top. The GPU-cost budget (10 +percent of energy per hash) is the binding side, and the vendor-rounding question is the class's, not the chip's. + ## 5. The chip edge at the measured k `E_chip = E_mem + N_ops x e_chip` (absolute: the chip's shadow cost is 102,100 x 3.5 pJ = 0.36 microjoules per hash @@ -270,7 +392,7 @@ to 2.3x and the strongest chips at 2.6x to 4.2x. | 6 | prmt, lop3 | 0.64, 0.72 | 11.5, 13.0 | 0.056, 0.055 | not drawn (RTL rows only) | | 7 | mulhi | 0.68 | 21.0 | 0.032 | yes (6) | | 8 | 32-lane shuffle (butterfly) | 0.63 | 29.4 | 0.021 | yes (8) | -| 9 | L1 scratch read (8 KB flop array) against the card's L2 hit | ROW_SCRATCH_K | 1,400 | pending | not drawn | +| 9 | L1 scratch read (8 KB flop array) against the card's L2 hit | 104 (pJ per read) | 1,400 | 0.074 | not drawn | | 10 | int8 8x8x8 tile, per MAC | ROW_TILE_K | 2.2 | pending | not drawn (the tensor lever is dead on other grounds) | The order is set by the card's price, not the chip's: the chip pays 0.6 to 1.7 pJ for everything, and the card diff --git a/docs/plans/igneum-2.0-test-harness-map.md b/docs/plans/igneum-2.0-test-harness-map.md index 26d5bf651..5711eb938 100644 --- a/docs/plans/igneum-2.0-test-harness-map.md +++ b/docs/plans/igneum-2.0-test-harness-map.md @@ -400,6 +400,14 @@ Rule: a case maps to a cell only where the cell's tests visibly answer it; cover - Cases: - INT-07 One v6 object agrees in node, pool, CPU verifier and each supported GPU host across activation.: partial: V6-12's clean-install half (one published object installed fresh, synced, mined, proved and paid on one host per artefact); the cross-host agreement half (node, pool, CPU verifier, each GPU host across activation) is harness:same-work's +### review:k-lane-shadow-k + +- Command: `floor lane 2's placed and routed cores on ASAP7 in tools/chip-model/rtl (the rows in docs/analysis/class-v6/floor/shadow-k.md); the register landing 50ff1611f recorded ADV-06 against it before the recorder existed` +- Box class: none +- Fixtures: none +- Cases: + - ADV-06 Separate process advantage from specialisation: partial: the placed gated cores and the adversary's forms are modelled in shadow-k.md; the independent review remains + ## Automated cases with no harness in the matrix (NOT RUN, the reason) - GOV-02 Approve thresholds before results: the approval is recorded in the registry's approval field; the automated half (thresholds frozen before any run_status) is the gate rule landing by 21:00 @@ -409,7 +417,6 @@ Rule: a case maps to a cell only where the cell's tests visibly answer it; cover - GPU-06 Measure accepted work under ordinary connectivity: accepted work under ordinary connectivity needs the fault network F4 - GPU-07 Survive sustained thermal and power operation: the sustained thermal and power soak has no harness tonight: the project's own rig mines nothing under the the earlier devnet off order - POW-05 Prevent amortised cheap winning attempts: amortised cheap winning attempts are the attack lanes' grind and era harnesses (tools/attack/f7-era, f9-grind), not in the release matrix; their rows come from those lanes -- ADV-06 Separate process advantage from specialisation: process-advantage separation is the adversary lanes' chip study - ROT-03 Test miner-voted bring-forward governance: miner-voted bring-forward needs a vote harness on the fault network F4 - ROT-04 Resist seed selection and faster evaluators: seed-selection resistance is the census harness (the class v6 invention lane), not yet in the matrix - EVM-01 Match the selected EVM semantics: no EVM conformance-vector harness is mapped tonight; the exec suite does not run the reference test vectors @@ -551,4 +558,4 @@ Rule: a case maps to a cell only where the cell's tests visibly answer it; cover ## Count -171 automated cases: 82 mapped to a cell, 146 NOT RUN with a reason. +171 automated cases: 83 mapped to a cell, 145 NOT RUN with a reason. diff --git a/docs/plans/igneum-2.0-test-registry.json b/docs/plans/igneum-2.0-test-registry.json index 014a6c01e..69b374ab3 100644 --- a/docs/plans/igneum-2.0-test-registry.json +++ b/docs/plans/igneum-2.0-test-registry.json @@ -952,7 +952,20 @@ "updated": "2026-10-08T21:00:32.846Z", "evidence_record": { "reason": "the memory-clock ladder has no harness tonight: the project's own rig mines nothing under the the earlier devnet off order", - "at": "2026-10-08T21:00:32.846Z" + "at": "2026-10-08T21:00:32.846Z", + "method": "static", + "requirement_id": "GPU-04", + "decision": "NOT RUN", + "reviewer": "", + "claim_impact": "", + "release_identity": { + "commit": "", + "lockfile": "", + "binary": "", + "network_object": "", + "activation": "", + "profile_hashes": "" + } }, "in_progress_since": "2026-10-08 18:3x UK", "approvals": { @@ -1187,7 +1200,20 @@ "updated": "2026-10-08T21:00:32.846Z", "evidence_record": { "reason": "the sustained thermal and power soak has no harness tonight: the project's own rig mines nothing under the the earlier devnet off order", - "at": "2026-10-08T21:00:32.846Z" + "at": "2026-10-08T21:00:32.846Z", + "method": "static", + "requirement_id": "GPU-07", + "decision": "NOT RUN", + "reviewer": "", + "claim_impact": "", + "release_identity": { + "commit": "", + "lockfile": "", + "binary": "", + "network_object": "", + "activation": "", + "profile_hashes": "" + } }, "in_progress_since": "2026-10-08 18:3x UK", "approvals": { @@ -2561,21 +2587,21 @@ "manual_page": 28, "owner_lane": "k lane (a3c9601a6d4686fe1)", "run_status": "NOT RUN", - "evidence_path": "docs/analysis/class-v6/multi-family-adversary.md; docs/analysis/class-v6/floor/shadow-k.md; docs/analysis/class-v6/multi-family-adversary.md", - "run_id": "adversary-20261008-placed-8lane", - "updated": "2026-10-08T20:41:20.327Z", + "evidence_path": "docs/analysis/class-v6/multi-family-adversary.md; docs/analysis/class-v6/floor/shadow-k.md; docs/analysis/class-v6/floor/shadow-k.md", + "run_id": "floor-k-20261008-rows-repeat", + "updated": "2026-10-08T21:09:38.300Z", "evidence_record": { "requirement_id": "ADV-05", "decision": "NOT RUN", "method": "model", "cell": "adversary:mf-placed", - "manifest_sha": "3a8874fef", - "run_id": "adversary-20261008-placed-8lane", - "evidence": "docs/analysis/class-v6/multi-family-adversary.md", + "manifest_sha": "86e5b0fb", + "run_id": "floor-k-20261008-rows-repeat", + "evidence": "docs/analysis/class-v6/floor/shadow-k.md", "in_progress": true, "coverage": "partial: the SRAM macros, ports, wiring, clocking and the complete-board terms are modelled in sections 2.3, 5 and 6 with the unmodelled items carried as uncertainty; the calibration against an existing hardware block is the k lane's bare-lane row beside it, owed as a named comparison", "release_identity": { - "commit": "3a8874fef", + "commit": "86e5b0fb", "lockfile": "", "binary": "", "network_object": "", @@ -2584,7 +2610,7 @@ }, "claim_impact": "", "reviewer": "", - "at": "2026-10-08T20:41:20.327Z" + "at": "2026-10-08T21:09:38.300Z" }, "in_progress_since": "2026-10-08 18:3x UK", "approvals": { @@ -2621,13 +2647,13 @@ "decision": "NOT RUN", "method": "model", "cell": "adversary:mf-placed", - "manifest_sha": "3a8874fef", - "run_id": "adversary-20261008-placed-8lane", - "evidence": "docs/analysis/class-v6/multi-family-adversary.md", + "manifest_sha": "86e5b0fb", + "run_id": "floor-k-20261008-rows-repeat", + "evidence": "docs/analysis/class-v6/floor/shadow-k.md", "in_progress": true, "coverage": "partial: the SRAM macros, ports, wiring, clocking and the complete-board terms are modelled in sections 2.3, 5 and 6 with the unmodelled items carried as uncertainty; the calibration against an existing hardware block is the k lane's bare-lane row beside it, owed as a named comparison", "release_identity": { - "commit": "3a8874fef", + "commit": "86e5b0fb", "lockfile": "", "binary": "", "network_object": "", @@ -2636,7 +2662,7 @@ }, "claim_impact": "", "reviewer": "", - "at": "2026-10-08T20:41:20.327Z" + "at": "2026-10-08T21:09:38.300Z" } } }, @@ -2669,24 +2695,29 @@ "owner_lane": "k lane (a3c9601a6d4686fe1)", "run_status": "NOT RUN", "evidence_path": "docs/analysis/class-v6/floor/shadow-k.md", - "run_id": "team-2026-10-08", - "updated": "2026-10-08T20:10:24.556Z", + "run_id": "floor-k-20261008-rows-repeat", + "updated": "2026-10-08T21:09:38.300Z", "evidence_record": { - "reason": "process-advantage separation is the adversary lanes' chip study", - "at": "2026-10-08T20:10:24.556Z", - "method": "static", "requirement_id": "ADV-06", "decision": "NOT RUN", - "reviewer": "", - "claim_impact": "", + "method": "model", + "cell": "review:k-lane-shadow-k", + "manifest_sha": "86e5b0fb", + "run_id": "floor-k-20261008-rows-repeat", + "evidence": "docs/analysis/class-v6/floor/shadow-k.md", + "in_progress": false, + "coverage": "partial: the placed gated cores and the adversary's forms are modelled in shadow-k.md; the independent review remains", "release_identity": { - "commit": "", + "commit": "86e5b0fb", "lockfile": "", "binary": "", "network_object": "", "activation": "", "profile_hashes": "" - } + }, + "claim_impact": "", + "reviewer": "", + "at": "2026-10-08T21:09:38.300Z" }, "in_progress_since": "2026-10-08 18:3x UK", "approvals": { @@ -2717,6 +2748,28 @@ "run_id": "team-2026-10-08", "evidence": "docs/analysis/class-v6/floor/shadow-k.md", "in_progress": true + }, + "review:k-lane-shadow-k": { + "requirement_id": "ADV-06", + "decision": "NOT RUN", + "method": "model", + "cell": "review:k-lane-shadow-k", + "manifest_sha": "86e5b0fb", + "run_id": "floor-k-20261008-rows-repeat", + "evidence": "docs/analysis/class-v6/floor/shadow-k.md", + "in_progress": false, + "coverage": "partial: the placed gated cores and the adversary's forms are modelled in shadow-k.md; the independent review remains", + "release_identity": { + "commit": "86e5b0fb", + "lockfile": "", + "binary": "", + "network_object": "", + "activation": "", + "profile_hashes": "" + }, + "claim_impact": "", + "reviewer": "", + "at": "2026-10-08T21:09:38.300Z" } } }, @@ -15225,6 +15278,8 @@ "owner_lane": "CI steward with the hash lane (a690540514aa453d7) and the worker lane (a9e87343f008e0edd)", "run_status": "BLOCKED", "master_status": "PROPOSED / NOT RUN", + "run_id": "canary-20261008-01", + "evidence_path": "tools/ci/canary-check.sh; packaging/ota/publish-manifest.sh; packaging/ota/publish-public.sh", "updated": "2026-10-08T21:00:32.846Z", "evidence_record": { "requirement_id": "INT-07", @@ -15295,9 +15350,7 @@ "implementation_complete": null, "evidence_reproduced": null, "claim_authorised": null - }, - "run_id": "canary-20261008-01", - "evidence_path": "tools/ci/canary-check.sh; packaging/ota/publish-manifest.sh; packaging/ota/publish-public.sh" + } }, { "id": "INT-08", diff --git a/tools/chip-model/rtl/Makefile b/tools/chip-model/rtl/Makefile index dea28cc08..cf77c1d4e 100644 --- a/tools/chip-model/rtl/Makefile +++ b/tools/chip-model/rtl/Makefile @@ -23,8 +23,14 @@ CPUSET = $(if $(LEASE_ON),--cpuset-cpus {cpuset},) DOCKER := $(LEASEPFX) docker run --rm -u $(UID_GID) -e HOME=/tmp -e NUM_CORES=$(THREADS) $(CPUSET) -v $(WORK):/work ORFS := $(DOCKER) -w /OpenROAD-flow-scripts/flow $(ORFS_IMG) SIM := $(DOCKER) -w /work $(SIM_IMG) +# NO_DOCKER=1: the box IS the ORFS image (a rented pod started from openroad/orfs:latest with iverilog installed +# and /work a symlink to this directory); the same targets run natively. +ifeq ($(NO_DOCKER),1) +ORFS := env NUM_CORES=$(THREADS) bash -c 'cd /OpenROAD-flow-scripts/flow && exec "$$@"' -- +SIM := env +endif -DESIGNS := arx mul prmt lop3 fold shfl xbar scratch tile core8 core32 core32r16 core8r64 core8i1k core8sel core32all +DESIGNS := arx mul prmt lop3 fold shfl xbar scratch tile core8 core32 core32r16 core8r64 core8i1k core8sel core32all core8g core8r64g coretm cs64 fp32 top = $(shell sed -n 's/^$(1) \([^ ]*\) .*/\1/p' flow/designs.txt) # per-family simulation tags (the op field fixed per row where the family has several ops) @@ -37,13 +43,18 @@ SIMS_shfl := mix SIMS_xbar := mix SIMS_scratch := mix SIMS_tile := mix -SIMS_core8 := mix mixld:+loads=1 -SIMS_core32 := mix mixld:+loads=1 -SIMS_core32r16 := mix mixld:+loads=1 -SIMS_core8r64 := mix +SIMS_core8 := s150:+cycles=150 s600:+cycles=600 +SIMS_core32 := s150:+cycles=150 s600:+cycles=600 +SIMS_core32r16 := s150:+cycles=150 s600:+cycles=600 +SIMS_core8r64 := s150:+cycles=150 s600:+cycles=600 SIMS_core8i1k := mix SIMS_core8sel := mix SIMS_core32all := mix +SIMS_core8g := s150:+cycles=150 s600:+cycles=600 +SIMS_core8r64g := s150:+cycles=150 s600:+cycles=600 +SIMS_coretm := mix +SIMS_cs64 := s150:+cycles=150 s600:+cycles=600 +SIMS_fp32 := mix fadd:+op=0 fmul:+op=1 ffma:+op=2 fcvt:+op=3 .PHONY: rows table clean @@ -65,6 +76,7 @@ sim-%: flow-% power-%: sim-% $(ORFS) make DESIGN_CONFIG=/work/flow/$*.mk RUN_SCRIPT=/work/flow/power.tcl RUN_LOG_NAME_STEM=power run 2>&1 | tee logs/power-$*.log + rm -f sim/$*/*.vcd # the traces run to gigabytes each; the power log keeps every number (a full disk killed four runs on 8 October 2026) # synthesis-only row (no placement, no parasitics): the 14:45 fallback synth-%: diff --git a/tools/chip-model/rtl/flow/asap7_icg_model.v b/tools/chip-model/rtl/flow/asap7_icg_model.v new file mode 100644 index 000000000..bd9c405a8 --- /dev/null +++ b/tools/chip-model/rtl/flow/asap7_icg_model.v @@ -0,0 +1,12 @@ +// Behavioural model of the ASAP7 integrated clock gate (latch_posedge_precontrol: the enable is latched while the +// clock is low, the gated clock is CLK AND the latched enable OR test). The liberty carries no function for it. +module ICGx1_ASAP7_75t_R(input CLK, input ENA, input SE, output GCLK); + reg en = 0; + always @(CLK or ENA or SE) if (!CLK) en = ENA | SE; + assign GCLK = CLK & en; +endmodule +module ICGx2_ASAP7_75t_R(input CLK, input ENA, input SE, output GCLK); + reg en = 0; + always @(CLK or ENA or SE) if (!CLK) en = ENA | SE; + assign GCLK = CLK & en; +endmodule diff --git a/tools/chip-model/rtl/flow/collect.py b/tools/chip-model/rtl/flow/collect.py index 842842019..79e01aff7 100644 --- a/tools/chip-model/rtl/flow/collect.py +++ b/tools/chip-model/rtl/flow/collect.py @@ -6,7 +6,7 @@ import re, sys, os, csv work = sys.argv[1] if len(sys.argv) > 1 else '.' # ops per cycle per design (the per-op divisor) and the GPU row each family is read against -OPS = {'arx': 1, 'mul': 1, 'prmt': 1, 'lop3': 1, 'fold': 1, 'shfl': 32, 'xbar': 32, 'scratch': 1, 'tile': 512, 'core8': 8, 'core32': 32, 'core32r16': 32, 'core8r64': 8, 'core8i1k': 8, 'core8sel': 8, 'core32all': 32} +OPS = {'arx': 1, 'mul': 1, 'prmt': 1, 'lop3': 1, 'fold': 1, 'shfl': 32, 'xbar': 32, 'scratch': 1, 'tile': 512, 'core8': 8, 'core32': 32, 'core32r16': 32, 'core8r64': 8, 'core8i1k': 8, 'core8sel': 8, 'core32all': 32, 'core8g': 8, 'core8r64g': 8, 'coretm': 1, 'cs64': 8, 'fp32': 1} # 5090 measured pJ per counted op: (unlocked, at the 1,300 MHz lock); 15.1a GPU = { 'arx:mix': (11.3, 6.2), 'arx:add': (11.3, 6.2), 'arx:sub': (11.3, 6.2), 'arx:xor': (11.3, 6.2), 'arx:or': (11.3, 6.2), @@ -17,7 +17,7 @@ GPU = { 'shfl:mix': (55.8, 29.4), 'xbar:mix': (55.8, 29.4), 'scratch:mix': (2400.0, 1400.0), # the card's L2 hit (no shared-memory probe measured: owed) 'tile:mix': (4.1, 2.2), - 'core8:mix': (11.3, 6.2), 'core8:mixld': (11.3, 6.2), 'core32:mix': (11.3, 6.2), 'core32:mixld': (11.3, 6.2), 'core32r16:mix': (11.3, 6.2), 'core8r64:mix': (11.3, 6.2), 'core8i1k:mix': (11.3, 6.2), 'core8sel:mix': (11.3, 6.2), 'core32all:mix': (11.3, 6.2), # the class v4 draw: read against int_arx (the packs job read the whole mix at 10.8 / 6.4) # dependent u8 m8n8k16 per MAC; the wide s8 tile reads 1.36 / 0.83 + 'core8:mix': (11.3, 6.2), 'core8:mixld': (11.3, 6.2), 'core32:mix': (11.3, 6.2), 'core32:mixld': (11.3, 6.2), 'core32r16:mix': (11.3, 6.2), 'core8r64:mix': (11.3, 6.2), 'core8i1k:mix': (11.3, 6.2), 'core8sel:mix': (11.3, 6.2), 'core32all:mix': (11.3, 6.2), 'core8g:mix': (11.3, 6.2), 'core8r64g:mix': (11.3, 6.2), 'coretm:mix': (11.3, 6.2), 'cs64:mix': (11.3, 6.2), 'fp32:mix': (9.2, 5.2), 'fp32:fadd': (9.2, 5.2), 'fp32:fmul': (9.2, 5.2), 'fp32:ffma': (9.2, 5.2), 'fp32:fcvt': (9.2, 5.2), # the class v4 draw: read against int_arx (the packs job read the whole mix at 10.8 / 6.4) # dependent u8 m8n8k16 per MAC; the wide s8 tile reads 1.36 / 0.83 } # per-node energy scaling from ASAP7 (a 7 nm-class predictive PDK at 0.70 V), approximate and claimed: # N7 -> N5 x0.70 (TSMC: "30 percent lower power at the same speed"), N5 -> N3E x0.72 (TSMC: 25 to 30 percent), @@ -67,10 +67,10 @@ for d in OPS: pairs = [('vcd:s150', 'vcd:s600', 150, 600), ('vcd:short', 'vcd:synth', 150, 800), ('vcd:s150', 'vcd:s400', 150, 400)] for sh, lg, cs, cl in pairs: if sh in rows and lg in rows: - rows['vcd:steady'] = steady(rows, period, sh, lg, cs, cl, 1028 if d in ('core8i1k', 'core32all') else LOAD_CYCLES) + rows['vcd:steady'] = steady(rows, period, sh, lg, cs, cl, 1028 if d in ('core8i1k', 'core32all') else (452 if d == 'cs64' else LOAD_CYCLES)) for tag, r in rows.items(): sub = tag.split(':')[1] if ':' in tag else 'prop' - key = f'{d}:{sub}' if sub in ('add','sub','xor','or','rotl','rotr','mul','mulhi','mad','mixld') else f'{d}:mix' + key = f'{d}:{sub}' if sub in ('add','sub','xor','or','rotl','rotr','mul','mulhi','mad','mixld','fadd','fmul','ffma','fcvt') else f'{d}:mix' gpu = GPU.get(key, (None, None)) pj = r['total'] * period * 1e-12 / OPS[d] * 1e12 # W * s / ops -> pJ pj_dyn = (r['internal'] + r['switching']) * period / OPS[d] diff --git a/tools/chip-model/rtl/flow/core8g.mk b/tools/chip-model/rtl/flow/core8g.mk new file mode 100644 index 000000000..c2601b8b8 --- /dev/null +++ b/tools/chip-model/rtl/flow/core8g.mk @@ -0,0 +1,18 @@ +# ORFS design config for the programmable shadow core (core8: core_v6_8), ASAP7. +export PLATFORM = asap7 +export DESIGN_NAME = core_v6_8 +export DESIGN_NICKNAME = core8g +export VERILOG_FILES = /work/rtl/core_v6_8.v +export VERILOG_INCLUDE_DIRS = /work/rtl +export SDC_FILE = /work/flow/core8g.sdc +export CORE_UTILIZATION = 40 +export CORE_ASPECT_RATIO = 1 +export CORE_MARGIN = 0.5 +export PLACE_DENSITY = 0.55 +export CORNER = TC +export SKIP_LAST_GASP = 1 +export WORK_HOME = /work/out/core8g +export SYNTH_MEMORY_MAX_BITS = 2000000 +# the adversary's register file: clock gating inferred (the ICG cells allowed back in) +export INFER_CLKGATES = 1 +export DONT_USE_CELLS = *x1p*_ASAP7* *xp*_ASAP7* SDF* diff --git a/tools/chip-model/rtl/flow/core8g.sdc b/tools/chip-model/rtl/flow/core8g.sdc new file mode 100644 index 000000000..10f3be2bd --- /dev/null +++ b/tools/chip-model/rtl/flow/core8g.sdc @@ -0,0 +1,10 @@ +current_design core_v6_8 +set clk_name core_clock +set clk_port_name clk +set clk_period 1500 +set clk_io_pct 0.2 +set clk_port [get_ports $clk_port_name] +create_clock -name $clk_name -period $clk_period $clk_port +set non_clock_inputs [all_inputs -no_clocks] +set_input_delay [expr $clk_period * $clk_io_pct] -clock $clk_name $non_clock_inputs +set_output_delay [expr $clk_period * $clk_io_pct] -clock $clk_name [all_outputs] diff --git a/tools/chip-model/rtl/flow/core8r64g.mk b/tools/chip-model/rtl/flow/core8r64g.mk new file mode 100644 index 000000000..e43857538 --- /dev/null +++ b/tools/chip-model/rtl/flow/core8r64g.mk @@ -0,0 +1,18 @@ +# ORFS design config for the programmable shadow core (core8: core_v6_8r64), ASAP7. +export PLATFORM = asap7 +export DESIGN_NAME = core_v6_8r64 +export DESIGN_NICKNAME = core8r64g +export VERILOG_FILES = /work/rtl/core_v6_8r64.v +export VERILOG_INCLUDE_DIRS = /work/rtl +export SDC_FILE = /work/flow/core8r64g.sdc +export CORE_UTILIZATION = 40 +export CORE_ASPECT_RATIO = 1 +export CORE_MARGIN = 0.5 +export PLACE_DENSITY = 0.55 +export CORNER = TC +export SKIP_LAST_GASP = 1 +export WORK_HOME = /work/out/core8r64g +export SYNTH_MEMORY_MAX_BITS = 2000000 +# the adversary's register file: clock gating inferred (the ICG cells allowed back in) +export INFER_CLKGATES = 1 +export DONT_USE_CELLS = *x1p*_ASAP7* *xp*_ASAP7* SDF* diff --git a/tools/chip-model/rtl/flow/core8r64g.sdc b/tools/chip-model/rtl/flow/core8r64g.sdc new file mode 100644 index 000000000..ffb92a9cd --- /dev/null +++ b/tools/chip-model/rtl/flow/core8r64g.sdc @@ -0,0 +1,10 @@ +current_design core_v6_8r64 +set clk_name core_clock +set clk_port_name clk +set clk_period 1500 +set clk_io_pct 0.2 +set clk_port [get_ports $clk_port_name] +create_clock -name $clk_name -period $clk_period $clk_port +set non_clock_inputs [all_inputs -no_clocks] +set_input_delay [expr $clk_period * $clk_io_pct] -clock $clk_name $non_clock_inputs +set_output_delay [expr $clk_period * $clk_io_pct] -clock $clk_name [all_outputs] diff --git a/tools/chip-model/rtl/flow/coretm.mk b/tools/chip-model/rtl/flow/coretm.mk new file mode 100644 index 000000000..ee4daff66 --- /dev/null +++ b/tools/chip-model/rtl/flow/coretm.mk @@ -0,0 +1,18 @@ +# ORFS design config for the programmable shadow core (core8: core_tm_8r64), ASAP7. +export PLATFORM = asap7 +export DESIGN_NAME = core_tm_8r64 +export DESIGN_NICKNAME = coretm +export VERILOG_FILES = /work/rtl/core_tm_8r64.v +export VERILOG_INCLUDE_DIRS = /work/rtl +export SDC_FILE = /work/flow/coretm.sdc +export CORE_UTILIZATION = 40 +export CORE_ASPECT_RATIO = 1 +export CORE_MARGIN = 0.5 +export PLACE_DENSITY = 0.55 +export CORNER = TC +export SKIP_LAST_GASP = 1 +export WORK_HOME = /work/out/coretm +export SYNTH_MEMORY_MAX_BITS = 2000000 +# the adversary's register file: clock gating inferred (the ICG cells allowed back in) +export INFER_CLKGATES = 1 +export DONT_USE_CELLS = *x1p*_ASAP7* *xp*_ASAP7* SDF* diff --git a/tools/chip-model/rtl/flow/coretm.sdc b/tools/chip-model/rtl/flow/coretm.sdc new file mode 100644 index 000000000..904853df5 --- /dev/null +++ b/tools/chip-model/rtl/flow/coretm.sdc @@ -0,0 +1,10 @@ +current_design core_tm_8r64 +set clk_name core_clock +set clk_port_name clk +set clk_period 1500 +set clk_io_pct 0.2 +set clk_port [get_ports $clk_port_name] +create_clock -name $clk_name -period $clk_period $clk_port +set non_clock_inputs [all_inputs -no_clocks] +set_input_delay [expr $clk_period * $clk_io_pct] -clock $clk_name $non_clock_inputs +set_output_delay [expr $clk_period * $clk_io_pct] -clock $clk_name [all_outputs] diff --git a/tools/chip-model/rtl/flow/cs64.mk b/tools/chip-model/rtl/flow/cs64.mk new file mode 100644 index 000000000..66f9c9f51 --- /dev/null +++ b/tools/chip-model/rtl/flow/cs64.mk @@ -0,0 +1,18 @@ +# ORFS design config for the programmable shadow core (core8: core_v6_8cs64), ASAP7. +export PLATFORM = asap7 +export DESIGN_NAME = core_v6_8cs64 +export DESIGN_NICKNAME = cs64 +export VERILOG_FILES = /work/rtl/core_v6_8cs64.v +export VERILOG_INCLUDE_DIRS = /work/rtl +export SDC_FILE = /work/flow/cs64.sdc +export CORE_UTILIZATION = 40 +export CORE_ASPECT_RATIO = 1 +export CORE_MARGIN = 0.5 +export PLACE_DENSITY = 0.55 +export CORNER = TC +export SKIP_LAST_GASP = 1 +export WORK_HOME = /work/out/cs64 +export SYNTH_MEMORY_MAX_BITS = 2000000 +# the adversary's register file: clock gating inferred (the ICG cells allowed back in) +export INFER_CLKGATES = 1 +export DONT_USE_CELLS = *x1p*_ASAP7* *xp*_ASAP7* SDF* diff --git a/tools/chip-model/rtl/flow/cs64.sdc b/tools/chip-model/rtl/flow/cs64.sdc new file mode 100644 index 000000000..8ab79b15b --- /dev/null +++ b/tools/chip-model/rtl/flow/cs64.sdc @@ -0,0 +1,10 @@ +current_design core_v6_8cs64 +set clk_name core_clock +set clk_port_name clk +set clk_period 1500 +set clk_io_pct 0.2 +set clk_port [get_ports $clk_port_name] +create_clock -name $clk_name -period $clk_period $clk_port +set non_clock_inputs [all_inputs -no_clocks] +set_input_delay [expr $clk_period * $clk_io_pct] -clock $clk_name $non_clock_inputs +set_output_delay [expr $clk_period * $clk_io_pct] -clock $clk_name [all_outputs] diff --git a/tools/chip-model/rtl/flow/designs.txt b/tools/chip-model/rtl/flow/designs.txt index af4f9566e..292d1983d 100644 --- a/tools/chip-model/rtl/flow/designs.txt +++ b/tools/chip-model/rtl/flow/designs.txt @@ -14,3 +14,8 @@ core8r64 core_v6_8r64 1500 core8i1k core_v6_8i1k 1500 core8sel core_v6_8sel 1500 core32all core_v6_32all 1500 +core8g core_v6_8 1500 +core8r64g core_v6_8r64 1500 +coretm core_tm_8r64 1500 +cs64 core_v6_8cs64 1500 +fp32 lane_fp32 2000 diff --git a/tools/chip-model/rtl/flow/fp32.mk b/tools/chip-model/rtl/flow/fp32.mk new file mode 100644 index 000000000..cb508f48a --- /dev/null +++ b/tools/chip-model/rtl/flow/fp32.mk @@ -0,0 +1,15 @@ +# ORFS design config for the mul shadow-core family (top lane_fp32), ASAP7. +export PLATFORM = asap7 +export DESIGN_NAME = lane_fp32 +export DESIGN_NICKNAME = fp32 +export VERILOG_FILES = /work/rtl/fp32_units.v +export VERILOG_INCLUDE_DIRS = /work/rtl +export SDC_FILE = /work/flow/fp32.sdc +export CORE_UTILIZATION = 40 +export CORE_ASPECT_RATIO = 1 +export CORE_MARGIN = 0.5 +export PLACE_DENSITY = 0.55 +export CORNER = TC +export SKIP_LAST_GASP = 1 +export WORK_HOME = /work/out/fp32 + diff --git a/tools/chip-model/rtl/flow/fp32.sdc b/tools/chip-model/rtl/flow/fp32.sdc new file mode 100644 index 000000000..b729bbae1 --- /dev/null +++ b/tools/chip-model/rtl/flow/fp32.sdc @@ -0,0 +1,10 @@ +current_design lane_fp32 +set clk_name core_clock +set clk_port_name clk +set clk_period 2000 +set clk_io_pct 0.2 +set clk_port [get_ports $clk_port_name] +create_clock -name $clk_name -period $clk_period $clk_port +set non_clock_inputs [all_inputs -no_clocks] +set_input_delay [expr $clk_period * $clk_io_pct] -clock $clk_name $non_clock_inputs +set_output_delay [expr $clk_period * $clk_io_pct] -clock $clk_name [all_outputs] diff --git a/tools/chip-model/rtl/flow/livestate.py b/tools/chip-model/rtl/flow/livestate.py new file mode 100644 index 000000000..f729ddd05 --- /dev/null +++ b/tools/chip-model/rtl/flow/livestate.py @@ -0,0 +1,79 @@ +#!/usr/bin/env python3 +"""Live-state analysis of a drawn shadow program on an R-register window (the coordinator's order, 15:3x UK). +The program is drawn as the core testbench draws it: NPROG instructions with the class v4 op weights, a load on +one instruction in 16 (the dependent memory wait), dst/src/src2 uniform over the R registers, and the result +fold reading every register at the end of the block. For every load (wait) the script reports the live set: +registers whose current value is read later (by an instruction, a later address, or the fold) before being +overwritten, split into those that feed a later ADDRESS or the RESULT and those that die inside an arithmetic +block. Dead writes (overwritten before any read) are counted too. Usage: livestate.py R [NPROG] [seeds]""" +import random, sys +R = int(sys.argv[1]) if len(sys.argv) > 1 else 64 +N = int(sys.argv[2]) if len(sys.argv) > 2 else 256 +SEEDS = int(sys.argv[3]) if len(sys.argv) > 3 else 64 +W = [('add',12),('xor',10),('mul',8),('mad',8),('shfl',8),('rotl',7),('sub',6),('mulhi',6),('rotr',6),('or',4)] +ops = [o for o,w in W for _ in range(w)] +def draw(rng): + prog = [] + for k in range(N): + op = 'load' if k % 16 == 15 else rng.choice(ops) + d, s, s2 = rng.randrange(R), rng.randrange(R), rng.randrange(R) + reads = [s] if op not in ('load',) else [s] # the load's address comes from src + if op in ('add','xor','mul','mad','sub','or','rotr','shfl','mulhi','rotl'): reads.append(d) # dst is read too (r[d] op= ...) + if op == 'mad': reads.append(s2) + if op == 'rotl': reads = [d] + prog.append((op, d, reads)) + return prog +tot_live = tot_addr = tot_dead = tot_waits = 0 +live_min, live_max = R, 0 +for seed in range(SEEDS): + rng = random.Random(seed) + prog = draw(rng) + # a value's "version" = (reg, write index); the fold at the end reads every register + # forward pass: for each instruction i and register r, next read of r's current value before its next write + n = len(prog) + # necessity: a version is NECESSARY if it reaches an address (a load's src) or the fold, transitively + # compute transitively by backward dataflow over versions + writes_at = {} # (i) -> reg written + # build version ids: version of reg r valid after instruction i + cur = {r: ('init', r) for r in range(R)} + uses = {} # version -> list of (consumer index, consumer version or 'addr'/'fold') + versions = set(cur.values()) + deps = {} # version -> set of versions it reads + for i, (op, d, reads) in enumerate(prog): + srcs = [cur[r] for r in reads] + if op == 'load': + v = ('load', i); deps[v] = set() # the returned word: its ADDRESS depends on srcs + for s in srcs: uses.setdefault(s, []).append(('addr', i)) + else: + v = (op, i); deps[v] = set(srcs) + for s in srcs: uses.setdefault(s, []).append(('op', i)) + cur[d] = v; versions.add(v) + fold = set(cur.values()) + # necessary = reaches an address or the fold + necessary = set(fold) + for v, us in uses.items(): + if any(k == 'addr' for k, _ in us): necessary.add(v) + changed = True + while changed: + changed = False + for v in list(versions): + if v in necessary: + for s in deps.get(v, ()): + if s not in necessary: necessary.add(s); changed = True + # per wait: the versions live at the load (written before it, read after it) + last_read = {} + for v, us in uses.items(): + last_read[v] = max(i for _, i in us) + for v in fold: last_read[v] = n + written_at = {v: (v[1] if v[0] != 'init' else -1) for v in versions} + for i, (op, d, reads) in enumerate(prog): + if op != 'load': continue + live = [v for v in versions if written_at[v] < i and last_read.get(v, -1) > i] + nec = [v for v in live if v in necessary] + tot_live += len(live); tot_addr += len(nec); tot_waits += 1 + live_min = min(live_min, len(live)); live_max = max(live_max, len(live)) + dead = sum(1 for v in versions if v[0] not in ('init',) and v not in uses and v not in fold) + tot_dead += dead +print(f'R = {R}, NPROG = {N}, {SEEDS} drawn programs, {tot_waits} waits') +print(f'live values at a wait: mean {tot_live/tot_waits:.1f} of {R} (min {live_min}, max {live_max}); of which necessary (reach a later address or the result): {tot_addr/tot_waits:.1f}') +print(f'dead writes (overwritten before any read): {tot_dead/SEEDS:.1f} per {N}-instruction block ({100*tot_dead/SEEDS/N:.1f} percent)') diff --git a/tools/chip-model/rtl/flow/pod.sh b/tools/chip-model/rtl/flow/pod.sh new file mode 100755 index 000000000..8344fc4ab --- /dev/null +++ b/tools/chip-model/rtl/flow/pod.sh @@ -0,0 +1,9 @@ +#!/usr/bin/env bash +# pod.sh: bootstrap a rented pod started from openroad/orfs:latest (RunPod, root): iverilog, /work -> this dir. +set -euo pipefail +cd "$(dirname "$0")/.." +export DEBIAN_FRONTEND=noninteractive +command -v iverilog >/dev/null || { apt-get update -qq >/dev/null 2>&1; apt-get install -y -qq iverilog rsync python3 >/dev/null 2>&1; } +[ -e /work ] || ln -s "$(pwd)" /work +export PATH=/OpenROAD-flow-scripts/tools/install/OpenROAD/bin:/OpenROAD-flow-scripts/tools/install/yosys/bin:$PATH +echo "pod ready: $(nproc) cores, $(free -g | awk '/Mem/{print $2}') GB, yosys $(yosys -V | cut -d' ' -f2), $(which iverilog)" diff --git a/tools/chip-model/rtl/flow/sim.sh b/tools/chip-model/rtl/flow/sim.sh index 88b05666d..f288d1c23 100755 --- a/tools/chip-model/rtl/flow/sim.sh +++ b/tools/chip-model/rtl/flow/sim.sh @@ -4,6 +4,7 @@ set -euo pipefail name=$1; tag=$2; shift 2 out=/work/sim/$name simcells=$(yosys-config --datdir)/simcells.v -iverilog -g2005 -I /work/tb -o $out/sim_$tag $out/sim_net.v /work/tb/tb_$(sed -n "s/^$name \([^ ]*\) .*/\1/p" /work/flow/designs.txt).v $simcells +tbfile=${TB:-/work/tb/tb_$(sed -n "s/^$name \([^ ]*\) .*/\1/p" /work/flow/designs.txt).v} +iverilog -g2005 -I /work/tb -o $out/sim_$tag $out/sim_net.v /work/flow/asap7_icg_model.v $tbfile $simcells ( cd $out && vvp -n sim_$tag "$@" | tee sim_$tag.log && mv dump.vcd $tag.vcd ) ls -la $out/$tag.vcd diff --git a/tools/chip-model/rtl/rtl/core_tm.v b/tools/chip-model/rtl/rtl/core_tm.v new file mode 100644 index 000000000..3e702511b --- /dev/null +++ b/tools/chip-model/rtl/rtl/core_tm.v @@ -0,0 +1,82 @@ +// The adversary's time-multiplexed core: ONE execution port (every class unit, once) serving LANES lanes' instruction +// streams round-robin, each lane's state (REGS x 32-bit) kept in its own bank; the imem and sequencer shared. +// One lane-op per cycle. Compared with core_v6 at the same LANES x REGS this removes LANES-1 copies of the units and +// keeps the register state and the imem: the energy per lane-op is the state's cost plus one unit set's. +// The butterfly shuffle across lanes needs every lane's source register in the same cycle, so the shuffle reads the +// bank-wide source column (as the SIMD core does) and the lane in turn takes its word. Loads return on ld_val. +`include "lane_common.vh" +module core_tm #(parameter LANES = 8, parameter LOG_LANES = 3, parameter REGS = 64, parameter LOG_REGS = 6, + parameter IW = 40, parameter IMEM_LOG = 8) ( + input clk, input rst, input run, + input prog_we, input [9:0] prog_addr, input [IW-1:0] prog_data, + input cfg_en, input [9:0] cfg_n, input [31:0] cfg_m, input [4:0] cfg_r, input [31:0] cfg_wm, input [31:0] cfg_off, input [31:0] cfg_mask, input [63:0] cfg_sel, + input [31:0] ld_val, + output [31:0] addr, output [31:0] out); + localparam IMEM = 1 << IMEM_LOG; + reg [IW-1:0] imem [0:IMEM-1]; + reg [IMEM_LOG-1:0] pc; reg [IMEM_LOG-1:0] n_q; reg [IW-1:0] ir; reg [LOG_LANES-1:0] lane; + reg [31:0] m_q, wm_q, off_q, mask_q; reg [4:0] r_q; + integer i; + // the sequencer: the same instruction is issued to each lane in turn (LANES cycles per instruction) + always @(posedge clk) begin + if (prog_we) imem[prog_addr[IMEM_LOG-1:0]] <= prog_data; + if (rst) begin pc <= 0; ir <= 0; lane <= 0; n_q <= {IMEM_LOG{1'b1}}; m_q <= 32'h9e3779b1; r_q <= 5'd13; wm_q <= 32'h0fffffc0; off_q <= 3; mask_q <= 32'h0fffffff; end + else begin + if (cfg_en) begin n_q <= cfg_n[IMEM_LOG-1:0]; m_q <= cfg_m | 1; r_q <= cfg_r; wm_q <= cfg_wm; off_q <= cfg_off; mask_q <= cfg_mask; end + if (run) begin + if (lane == {LOG_LANES{1'b1}}) begin ir <= imem[pc]; pc <= (pc == n_q) ? {IMEM_LOG{1'b0}} : pc + 1'b1; end + lane <= lane + 1'b1; + end + end + end + wire [3:0] op = ir[3:0]; + wire [LOG_REGS-1:0] dst = ir[4 +: LOG_REGS]; wire [LOG_REGS-1:0] src = ir[4+LOG_REGS +: LOG_REGS]; wire [LOG_REGS-1:0] src2 = ir[4+2*LOG_REGS +: LOG_REGS]; + wire [4:0] imm = ir[4+3*LOG_REGS +: 5]; wire [7:0] aux = ir[9+3*LOG_REGS +: 8]; + wire is_load = (op == 4'd12); + wire [4:0] rn = (imm == 0) ? 5'd1 : imm; + wire [LOG_LANES-1:0] smask = imm[LOG_LANES-1:0]; + // the banked state: one bank per lane, read through the lane select (a chip's SRAM bank select) + reg [31:0] rf [0:LANES*REGS-1]; + wire [31:0] d = rf[lane*REGS + dst]; + wire [31:0] s = rf[lane*REGS + src]; + wire [31:0] s2 = rf[lane*REGS + src2]; + wire [31:0] sx = rf[(lane ^ smask)*REGS + src]; // the shuffle partner's source word + // the one execution port + function [7:0] pick; input [63:0] b; input [3:0] k; reg [7:0] v; + begin v = b[8*k[2:0] +: 8]; pick = k[3] ? {8{v[7]}} : v; end + endfunction + wire [4:0] sn = (s[4:0] == 0) ? 5'd1 : s[4:0]; + wire mad = (op == 4'd8); + wire [63:0] p = (mad ? s : d) * (mad ? s2 : s); + wire [63:0] bytes = {s, d}; wire [15:0] sel = {aux, aux}; + wire [31:0] prm = {pick(bytes, sel[15:12]), pick(bytes, sel[11:8]), pick(bytes, sel[7:4]), pick(bytes, sel[3:0])}; + reg [31:0] lp; integer b; + always @* for (b = 0; b < 32; b = b + 1) lp[b] = aux[{d[b], s[b], s2[b]}]; + reg [31:0] r; + always @* begin + case (op) + 4'd0, 4'd13: r = d + s; + 4'd1, 4'd15: r = d - s; + 4'd2, 4'd14: r = d ^ s; + 4'd3: r = d | s; + 4'd4: r = `ROTL32(d, rn); + 4'd5: r = `ROTR32(d, sn); + 4'd6: r = p[31:0]; + 4'd7: r = p[63:32]; + 4'd8: r = p[31:0] + d; + 4'd9: r = d ^ sx; + 4'd10: r = prm; + 4'd11: r = lp; + default: r = ld_val ^ (32'h9e3779b9 * (lane + 1)); + endcase + end + wire [31:0] fx = s * m_q; + wire [4:0] frn = (r_q == 0) ? 5'd1 : r_q; + wire [31:0] fy = `ROTL32(fx, frn); + assign addr = is_load ? (((fy & wm_q) | off_q) & mask_q) : 32'd0; + always @(posedge clk) begin + if (rst) begin for (i = 0; i < LANES*REGS; i = i + 1) rf[i] <= 32'h9e3779b9 * (i + 1); end + else if (run) rf[lane*REGS + dst] <= r; + end + assign out = r; +endmodule diff --git a/tools/chip-model/rtl/rtl/core_tm_8r64.v b/tools/chip-model/rtl/rtl/core_tm_8r64.v new file mode 100644 index 000000000..5db91edc4 --- /dev/null +++ b/tools/chip-model/rtl/rtl/core_tm_8r64.v @@ -0,0 +1,7 @@ +`include "core_tm.v" +module core_tm_8r64(input clk, input rst, input run, input prog_we, input [9:0] prog_addr, input [39:0] prog_data, + input cfg_en, input [9:0] cfg_n, input [31:0] cfg_m, input [4:0] cfg_r, input [31:0] cfg_wm, input [31:0] cfg_off, input [31:0] cfg_mask, input [63:0] cfg_sel, + input [31:0] ld_val, output [31:0] addr, output [31:0] out); + core_tm #(.LANES(8), .LOG_LANES(3), .REGS(64), .LOG_REGS(6), .IW(40)) c(.clk(clk), .rst(rst), .run(run), .prog_we(prog_we), .prog_addr(prog_addr), .prog_data(prog_data), + .cfg_en(cfg_en), .cfg_n(cfg_n), .cfg_m(cfg_m), .cfg_r(cfg_r), .cfg_wm(cfg_wm), .cfg_off(cfg_off), .cfg_mask(cfg_mask), .cfg_sel(cfg_sel), .ld_val(ld_val), .addr(addr), .out(out)); +endmodule diff --git a/tools/chip-model/rtl/rtl/core_v6_8cs64.v b/tools/chip-model/rtl/rtl/core_v6_8cs64.v new file mode 100644 index 000000000..4d482e2c8 --- /dev/null +++ b/tools/chip-model/rtl/rtl/core_v6_8cs64.v @@ -0,0 +1,7 @@ +`include "core_v6.v" +module core_v6_8cs64(input clk, input rst, input run, input prog_we, input [9:0] prog_addr, input [40-1:0] prog_data, + input cfg_en, input [9:0] cfg_n, input [63:0] cfg_sel, input [31:0] cfg_m, input [4:0] cfg_r, input [31:0] cfg_wm, input [31:0] cfg_off, input [31:0] cfg_mask, + input [31:0] ld_val, output [31:0] addr, output [31:0] out); + core_v6 #(.LANES(8), .LOG_LANES(3), .REGS(64), .LOG_REGS(6), .IW(40), .IMEM_LOG(9)) c(.clk(clk), .rst(rst), .run(run), .prog_we(prog_we), .prog_addr(prog_addr), .prog_data(prog_data), + .cfg_en(cfg_en), .cfg_n(cfg_n), .cfg_sel(cfg_sel), .cfg_m(cfg_m), .cfg_r(cfg_r), .cfg_wm(cfg_wm), .cfg_off(cfg_off), .cfg_mask(cfg_mask), .ld_val(ld_val), .addr(addr), .out(out)); +endmodule diff --git a/tools/chip-model/rtl/rtl/fp32_units.v b/tools/chip-model/rtl/rtl/fp32_units.v new file mode 100644 index 000000000..73a76c0a4 --- /dev/null +++ b/tools/chip-model/rtl/rtl/fp32_units.v @@ -0,0 +1,123 @@ +// The adversary's simplified FP32 units for the mixed-resource lane's candidate (class-v6-mixedfp): inputs are +// f(x) = as_float((x & 0x807FFFFF) | ((96 + ((x >> 23) & 63)) << 23)): never zero, denormal, NaN or Inf; exponents +// in [96, 159]; results normal or +0 (exact cancellation). The units drop NaN/Inf/denormal handling and the flags, +// keep the full 24-bit mantissa path, a full alignment and a full normaliser (the mantissas are uniform), RNE. +// Each lane reads two or three registers of an 8 x 32-bit window, applies f(), computes, xors the bits into dst. +`include "lane_common.vh" +// ---- the shared pieces ---- +module fp_unpack(input [31:0] x, output s, output [8:0] e, output [23:0] m); + assign s = x[31]; + assign e = 9'd96 + {3'b0, x[28:23]}; // the masked exponent, 96..159 + assign m = {1'b1, x[22:0]}; +endmodule +module lzc48(input [47:0] v, output reg [5:0] n); // leading-zero count (v != 0) + integer i; always @* begin n = 6'd47; for (i = 47; i >= 0; i = i - 1) if (v[i]) begin n = 6'd47 - i; i = -1; end end +endmodule +module lzc32(input [31:0] v, output reg [5:0] n); + integer i; always @* begin n = 6'd31; for (i = 31; i >= 0; i = i - 1) if (v[i]) begin n = 6'd31 - i; i = -1; end end +endmodule +// ---- the FMA: fma(a, b, c) = a*b + c, one rounding (RNE), exponents in the lane's ranges ---- +module fp_fma(input [31:0] a, input [31:0] b, input [31:0] c, output [31:0] y); + wire sa, sb, sc; wire [8:0] ea, eb, ec; wire [23:0] ma, mb, mc; + fp_unpack ua(a, sa, ea, ma); fp_unpack ub(b, sb, eb, mb); fp_unpack uc(c, sc, ec, mc); + wire [47:0] prod = ma * mb; // 48-bit product, binary point after bit 46 + wire sp = sa ^ sb; + wire [9:0] ep = {1'b0, ea} + {1'b0, eb} - 10'd127; // product exponent (bias kept), 65..192 + // align the addend to the product: a 100-bit window keeps full precision for exponent gaps up to about 96 (the lane's bound) + wire [9:0] diff = (ep >= {1'b0, ec}) ? ep - {1'b0, ec} : {1'b0, ec} - ep; + wire prod_big = (ep >= {1'b0, ec}); + wire [99:0] pw = {2'b0, prod, 50'b0}; + wire [99:0] cw = {2'b0, mc, 24'b0, 50'b0}; // the addend at the product's scale when exponents equal + wire [6:0] sh = (diff > 10'd99) ? 7'd99 : diff[6:0]; + wire [99:0] smw = prod_big ? (cw >> sh) : (pw >> sh); + wire [99:0] bgw = prod_big ? pw : cw; + wire sbig = prod_big ? sp : sc; wire ssmall = prod_big ? sc : sp; + wire [9:0] ebig = prod_big ? ep : {1'b0, ec}; + wire [100:0] sum = (sbig == ssmall) ? ({1'b0, bgw} + {1'b0, smw}) : ({1'b0, bgw} - {1'b0, smw}); + wire [100:0] mag = sum[100] ? (~sum + 1'b1) : sum; // two's complement when the subtraction went negative + wire ssum = sum[100] ? ssmall : sbig; + // normalise: find the leading one in the 101-bit magnitude + reg [6:0] lz; integer i; + always @* begin lz = 7'd100; for (i = 100; i >= 0; i = i - 1) if (mag[i]) begin lz = 7'd100 - i; i = -1; end end + wire [100:0] norm = mag << lz; // leading one at bit 100 + wire [23:0] mant = norm[100:77]; + wire guard = norm[76]; wire sticky = |norm[75:0]; + wire round_up = guard & (sticky | mant[0]); + wire [24:0] mr = {1'b0, mant} + round_up; + wire carry = mr[24]; + wire [9:0] eres = ebig + 10'd2 - lz + carry; // the leading one of bgw sat at bit 98 (two headroom bits) + wire zero = (mag == 0); + wire [7:0] eout = eres[7:0]; + assign y = zero ? 32'h0 : {ssum, eout, carry ? mr[23:1] : mr[22:0]}; +endmodule +// ---- the adder and the multiplier as their own units ---- +module fp_add(input [31:0] a, input [31:0] b, output [31:0] y); + wire sa, sb; wire [8:0] ea, eb; wire [23:0] ma, mb; + fp_unpack ua(a, sa, ea, ma); fp_unpack ub(b, sb, eb, mb); + wire abig = (ea > eb) || (ea == eb && ma >= mb); + wire [8:0] ebig = abig ? ea : eb; wire [8:0] esm = abig ? eb : ea; + wire [23:0] mbig = abig ? ma : mb; wire [23:0] msm = abig ? mb : ma; + wire sbig = abig ? sa : sb; wire ssm = abig ? sb : sa; + wire [8:0] diff = ebig - esm; wire [6:0] sh = (diff > 9'd70) ? 7'd70 : diff[6:0]; + wire [73:0] bw = {1'b0, mbig, 49'b0}; wire [73:0] sw = {1'b0, msm, 49'b0} >> sh; + wire [74:0] sum = (sbig == ssm) ? ({1'b0, bw} + {1'b0, sw}) : ({1'b0, bw} - {1'b0, sw}); + reg [6:0] lz; integer i; + always @* begin lz = 7'd74; for (i = 74; i >= 0; i = i - 1) if (sum[i]) begin lz = 7'd74 - i; i = -1; end end + wire [74:0] norm = sum << lz; + wire [23:0] mant = norm[74:51]; wire guard = norm[50]; wire sticky = |norm[49:0]; + wire round_up = guard & (sticky | mant[0]); + wire [24:0] mr = {1'b0, mant} + round_up; wire carry = mr[24]; + wire [9:0] eres = {1'b0, ebig} + 10'd1 - lz + carry; + wire zero = (sum == 0); + assign y = zero ? 32'h0 : {sbig, eres[7:0], carry ? mr[23:1] : mr[22:0]}; +endmodule +module fp_mul(input [31:0] a, input [31:0] b, output [31:0] y); + wire sa, sb; wire [8:0] ea, eb; wire [23:0] ma, mb; + fp_unpack ua(a, sa, ea, ma); fp_unpack ub(b, sb, eb, mb); + wire [47:0] prod = ma * mb; + wire top = prod[47]; + wire [23:0] mant = top ? prod[47:24] : prod[46:23]; + wire guard = top ? prod[23] : prod[22]; wire sticky = top ? |prod[22:0] : |prod[21:0]; + wire round_up = guard & (sticky | mant[0]); + wire [24:0] mr = {1'b0, mant} + round_up; wire carry = mr[24]; + wire [9:0] eres = {1'b0, ea} + {1'b0, eb} - 10'd127 + top + carry; + assign y = {sa ^ sb, eres[7:0], carry ? mr[23:1] : mr[22:0]}; +endmodule +// ---- int32 to float, RNE ---- +module fp_cvt(input [31:0] a, output [31:0] y); + wire s = a[31]; wire [31:0] mag = s ? (~a + 1'b1) : a; + wire [5:0] lz; lzc32 l(mag, lz); + wire [31:0] norm = mag << lz; // leading one at bit 31 + wire [23:0] mant = norm[31:8]; wire guard = norm[7]; wire sticky = |norm[6:0]; + wire round_up = guard & (sticky | mant[0]); + wire [24:0] mr = {1'b0, mant} + round_up; wire carry = mr[24]; + wire [7:0] e = 8'd127 + 8'd31 - lz + carry; + assign y = (mag == 0) ? 32'h0 : {s, e, carry ? mr[23:1] : mr[22:0]}; +endmodule +// ---- the lane: op 0 fadd, 1 fmul, 2 ffma, 3 fcvt; d ^= bits(result) ---- +module lane_fp32( + input clk, input rst, + input [1:0] op, input [2:0] dst, input [2:0] src, input [2:0] src2, + input ld_en, input [31:0] ld_val, + output [31:0] out); + reg [31:0] rf [0:7]; + reg [1:0] op_q; reg [2:0] dst_q, src_q, src2_q; reg ld_q; reg [31:0] ld_val_q; + integer i; + always @(posedge clk) begin + if (rst) begin op_q <= 0; dst_q <= 0; src_q <= 0; src2_q <= 0; ld_q <= 0; ld_val_q <= 0; end + else begin op_q <= op; dst_q <= dst; src_q <= src; src2_q <= src2; ld_q <= ld_en; ld_val_q <= ld_val; end + end + wire [31:0] d = rf[dst_q]; wire [31:0] s = rf[src_q]; wire [31:0] s2 = rf[src2_q]; + wire [31:0] ya, ym, yf, yc; + fp_add A(d, s, ya); + fp_mul M(d, s, ym); + fp_fma F(s, s2, d, yf); + fp_cvt C(s, yc); + reg [31:0] res; + always @* case (op_q) 2'd0: res = d ^ ya; 2'd1: res = d ^ ym; 2'd2: res = d ^ yf; default: res = d ^ yc; endcase + always @(posedge clk) begin + if (rst) begin for (i = 0; i < 8; i = i + 1) rf[i] <= 32'h9e3779b9 * (i + 1) + 9; end + else rf[dst_q] <= ld_q ? ld_val_q : res; + end + assign out = res; +endmodule diff --git a/tools/chip-model/rtl/tb/tb_core_common.vh b/tools/chip-model/rtl/tb/tb_core_common.vh index 38e969897..5a6f88d70 100644 --- a/tools/chip-model/rtl/tb/tb_core_common.vh +++ b/tools/chip-model/rtl/tb/tb_core_common.vh @@ -21,10 +21,32 @@ module tb; @(negedge clk); cfg_en = 1; cfg_n = `NPROG - 1; cfg_sel = `SEL; cfg_m = 32'h9e3779b1; cfg_r = 5'd13; cfg_wm = 32'h0fffffc0; cfg_off = 3; cfg_mask = 32'h0fffffff; @(negedge clk); cfg_en = 0; // the program: `NPROG instructions drawn with the class v4 weights +`ifdef CS + // the connected-state draw: step = load + 27-instruction spine block; `NPROG = 16 x 28 = 448 + begin : cs + integer st, q, last0, last1, last2, last3, mreg, areg, dreg, sreg, s2reg; + areg = 0; mreg = 1; + for (st = 0; st < `NPROG / 28; st = st + 1) begin + @(negedge clk); w = {$random, $random}; mreg = $random & 63; + prog_we = 1; prog_addr = st*28; prog_data = {w[`IW-1:4], 4'd12}; prog_data[4 +: 6] = mreg; prog_data[10 +: 6] = areg; // load: dst m_j, src a_j + last0 = mreg; last1 = mreg; last2 = mreg; last3 = mreg; + for (q = 0; q < 27; q = q + 1) begin + @(negedge clk); w = {$random, $random}; opc = draw_op($random); dreg = $random & 63; + case ($random & 3) 0: sreg = last0; 1: sreg = last1; 2: sreg = last2; default: sreg = last3; endcase + s2reg = last0; + if (q == 26 && !(opc == 4'd0 || opc == 4'd1 || opc == 4'd2 || opc == 4'd8 || opc == 4'd9)) opc = 4'd0; // the last instruction injects + prog_we = 1; prog_addr = st*28 + 1 + q; prog_data = {w[`IW-1:4], opc}; prog_data[4 +: 6] = dreg; prog_data[10 +: 6] = sreg; prog_data[16 +: 6] = s2reg; + last3 = last2; last2 = last1; last1 = last0; last0 = dreg; + end + areg = last0; + end + end +`else for (k = 0; k < `NPROG; k = k + 1) begin @(negedge clk); w = {$random, $random}; opc = (loads && (k % 16 == 15)) ? 4'd12 : draw_op($random); prog_we = 1; prog_addr = k; prog_data = {w[`IW-1:4], opc}; end +`endif @(negedge clk); prog_we = 0; run = 1; for (n = 0; n < cycles; n = n + 1) begin @(negedge clk); ld_val = $random; acc = acc ^ out ^ addr; diff --git a/tools/chip-model/rtl/tb/tb_core_tm_8r64.v b/tools/chip-model/rtl/tb/tb_core_tm_8r64.v new file mode 100644 index 000000000..954ae207f --- /dev/null +++ b/tools/chip-model/rtl/tb/tb_core_tm_8r64.v @@ -0,0 +1,6 @@ +`define TOP core_tm_8r64 +`define HALF 750 +`define IW 40 +`define NPROG 256 +`define SEL 64'hfedcba9876543210 +`include "tb_core_common.vh" diff --git a/tools/chip-model/rtl/tb/tb_core_v6_8_legacy.v b/tools/chip-model/rtl/tb/tb_core_v6_8_legacy.v new file mode 100644 index 000000000..9607238cf --- /dev/null +++ b/tools/chip-model/rtl/tb/tb_core_v6_8_legacy.v @@ -0,0 +1,3 @@ +`define TOP core_v6_8 +`define HALF 750 +`include "tb_core_legacy.vh" diff --git a/tools/chip-model/rtl/tb/tb_core_v6_8cs64.v b/tools/chip-model/rtl/tb/tb_core_v6_8cs64.v new file mode 100644 index 000000000..83d71eb03 --- /dev/null +++ b/tools/chip-model/rtl/tb/tb_core_v6_8cs64.v @@ -0,0 +1,7 @@ +`define TOP core_v6_8cs64 +`define HALF 750 +`define IW 40 +`define NPROG 448 +`define CS 1 +`define SEL 64'hfedcba9876543210 +`include "tb_core_common.vh" diff --git a/tools/chip-model/rtl/tb/tb_lane_fp32.v b/tools/chip-model/rtl/tb/tb_lane_fp32.v new file mode 100644 index 000000000..721a54e69 --- /dev/null +++ b/tools/chip-model/rtl/tb/tb_lane_fp32.v @@ -0,0 +1,22 @@ +`timescale 1ps/1ps +module tb; + reg clk = 0, rst = 1; reg [1:0] op = 0; reg [2:0] dst = 0, src = 0, src2 = 0; reg ld_en = 0; reg [31:0] ld_val = 0; + wire [31:0] out; + lane_fp32 dut(.clk(clk), .rst(rst), .op(op), .dst(dst), .src(src), .src2(src2), .ld_en(ld_en), .ld_val(ld_val), .out(out)); + integer n, fixed_op, cycles; reg [31:0] acc = 0; + always #1000 clk = ~clk; + // a reference check of the units against the host's float arithmetic is the mixed lane's own (the ranges are its); + // this bench drives random registers and reports the checksum + initial begin + if (!$value$plusargs("op=%d", fixed_op)) fixed_op = -1; + if (!$value$plusargs("cycles=%d", cycles)) cycles = 3000; + $dumpfile("dump.vcd"); $dumpvars(0, tb.dut); + repeat (4) @(negedge clk); rst = 0; + for (n = 0; n < cycles; n = n + 1) begin + @(negedge clk); + op = (fixed_op < 0) ? $random : fixed_op; dst = $random; src = $random; src2 = $random; + ld_en = (($random & 7) == 0); ld_val = $random; acc = acc ^ out; + end + $display("CHECKSUM %08x", acc); $finish; + end +endmodule diff --git a/tools/ci/batches/floor-k-20261008-rows-repeat.json b/tools/ci/batches/floor-k-20261008-rows-repeat.json new file mode 100644 index 000000000..f525f5993 --- /dev/null +++ b/tools/ci/batches/floor-k-20261008-rows-repeat.json @@ -0,0 +1,26 @@ +{ + "run_id": "floor-k-20261008-rows-repeat", + "manifest_sha": "86e5b0fb", + "cut_tip": "class-v6-floor-k 86e5b0fb (the amendment carries shadow-k.md; floor lane 2's rows ADV-05 and ADV-06 move with it on its word, method model, no decision changed; the steward's rows no longer cite the directory since d2e1c192)", + "evidence_dir": "docs/analysis/class-v6/floor/shadow-k.md", + "boxes": [], + "cells": [ + { + "cell": "adversary:mf-placed", + "cases": [ + "ADV-05" + ], + "status": "RUNNING", + "method": "model", + "evidence": "docs/analysis/class-v6/floor/shadow-k.md", + "note": "repeat for ADV-05 at this manifest on floor lane 2's word (placed and routed RTL on ASAP7, a model); the rtl directory tools/chip-model/rtl/ is the flow, cited by its document since the recorder takes files only" + }, + { + "cell": "review:k-lane-shadow-k", + "status": "NOT RUN", + "method": "model", + "evidence": "docs/analysis/class-v6/floor/shadow-k.md", + "note": "repeat of team-2026-10-08 for ADV-06 at this manifest on floor lane 2's word; no decision changed" + } + ] +} diff --git a/tools/ci/test-map.json b/tools/ci/test-map.json index b1b508883..d303d539a 100644 --- a/tools/ci/test-map.json +++ b/tools/ci/test-map.json @@ -687,6 +687,17 @@ "coverage": { "INT-07": "partial: V6-12's clean-install half (one published object installed fresh, synced, mined, proved and paid on one host per artefact); the cross-host agreement half (node, pool, CPU verifier, each GPU host across activation) is harness:same-work's" } + }, + "review:k-lane-shadow-k": { + "command": "floor lane 2's placed and routed cores on ASAP7 in tools/chip-model/rtl (the rows in docs/analysis/class-v6/floor/shadow-k.md); the register landing 50ff1611f recorded ADV-06 against it before the recorder existed", + "box_class": "none", + "fixtures": [], + "cases": [ + "ADV-06" + ], + "coverage": { + "ADV-06": "partial: the placed gated cores and the adversary's forms are modelled in shadow-k.md; the independent review remains" + } } }, "not_run": { @@ -697,7 +708,6 @@ "GPU-06": "accepted work under ordinary connectivity needs the fault network F4", "GPU-07": "the sustained thermal and power soak has no harness tonight: the project's own rig mines nothing under the the earlier devnet off order", "POW-05": "amortised cheap winning attempts are the attack lanes' grind and era harnesses (tools/attack/f7-era, f9-grind), not in the release matrix; their rows come from those lanes", - "ADV-06": "process-advantage separation is the adversary lanes' chip study", "ROT-03": "miner-voted bring-forward needs a vote harness on the fault network F4", "ROT-04": "seed-selection resistance is the census harness (the class v6 invention lane), not yet in the matrix", "EVM-01": "no EVM conformance-vector harness is mapped tonight; the exec suite does not run the reference test vectors",