Merge class-v6-floor-k-docs-2 3200b22d into master (gate: green on 2fc7d5c4, recorded by tools/ci/pre-push.sh; landed on the box mirror under the exception declared by main: main's ruling, 7 Oct 2026 19:5x UK: the GitHub account is suspended, lanes land on the box mirror's master, the box gate stamp is the verdict; GitHub gets the fast-forward when it answers)
This commit is contained in:
commit
e85fad3d97
31 changed files with 795 additions and 43 deletions
|
|
@ -104,7 +104,7 @@ stock and at the 1,300 MHz lock, and 6.9 pJ per counted op on the Apple M5 Max (
|
|||
| lop3 (8-bit truth table) | 7,274 | 1.42 | 0.99 | 0.72 | 0.52 | 24.1 / 13.0 | 0.030 / 0.055 | | 0.11 | |
|
||||
| 32-lane xor-mask shuffle (butterfly over the 1 KB window), per lane-op | 181,580 | 1.24 | 0.87 | 0.63 | 0.45 | 55.8 / 29.4 | 0.011 / 0.021 | | 0.042 | routed with SPEF; 15:1x UK |
|
||||
| 32-lane general crossbar, per lane-op | ROW_XBAR |
|
||||
| 8 KB scratch, one random read (flop array: the pessimistic form) | ROW_SCRATCH |
|
||||
| 8 KB scratch, one random 32-bit read of a 2,048 x 32 flop array (the pessimistic form of a chip's L1; an SRAM macro reads lower) | 868,159 | 207 | 145 | 104 | 75 | 2,400 / 1,400 per L2 hit | 0.043 / 0.074 | | 0.15 | routed with SPEF, 19:4x UK; the card's shared-memory read is unmeasured (owed) |
|
||||
| int8 8x8x8 tile, per MAC | ROW_TILE |
|
||||
|
||||
Reading the floors: a lane's add costs the chip about 2.2 pJ at ASAP7 and 1.1 at N3, against the 5090's 6.2
|
||||
|
|
@ -138,7 +138,7 @@ draws 13.9 mW, the run phase 36.9 mW for 8 lanes at 1.5 ns).
|
|||
| core, 8 lanes, 32 registers | synthesis only (no wires, no clock tree) | 186,443 | 6.9 | 4.8 | 3.5 | 2.5 | 11.3 / 6.2 / 6.9 | 0.31 / 0.56 / 0.50 | 0.22 / 0.40 / 0.36 | 1.1 | synthesised; 14:0x UK |
|
||||
| of which the sequential term (register file, imem and IR clock pins, no clock gating) | | | 2.4 | 1.7 | 1.2 | 0.9 | | | | | |
|
||||
| of which the units, the read muxes and the butterfly | | | 4.5 | 3.1 | 2.3 | 1.6 | | | | | |
|
||||
| core, 8 lanes, 32 registers | placed and routed, SPEF | ROW_CORE8_PLACED |
|
||||
| core, 8 lanes, 32 registers, ungated | placed and routed, SPEF, clock tree (one run length of 300 cycles, the load phase subtracted, about plus or minus 10 percent) | 444,478 | 11.3 | 7.9 | 5.6 | 4.1 | 11.3 / 6.2 / 6.9 | 0.50 / 0.91 / 0.82 | 0.36 / 0.66 / 0.59 | 1.8 | placed 16:0x UK on a rented pod; +64 percent over synthesis (wires, and a clock tree of 2.5 pJ per lane-op that gating removes) |
|
||||
| core, 32 lanes, 32 registers | synthesis only (steady state from 150 and 400 run cycles) | 600,381 | 5.55 | 3.9 | 2.8 | 2.0 | 11.3 / 6.2 / 6.9 | 0.25 / 0.45 / 0.41 | 0.18 / 0.32 / 0.29 | 0.90 | synthesised; 15:2x UK |
|
||||
| core, 32 lanes, 16 registers | synthesis only | 443,258 | 4.2 | 2.9 | 2.1 | 1.5 | 11.3 / 6.2 / 6.9 | 0.18 / 0.34 / 0.30 | 0.13 / 0.24 / 0.22 | 0.68 | synthesised; one run length, about plus or minus 10 percent; 15:0x UK |
|
||||
| the bare ARX lane (section 3, the floor) | routed | 11,631 | 2.2 | 1.5 | 1.1 | 0.8 | 11.3 / 6.2 / 6.9 | 0.10 / 0.18 / 0.16 | 0.07 / 0.13 / 0.11 | 0.35 | the lower bound |
|
||||
|
|
@ -203,6 +203,128 @@ takes back on the card's own node (3.6x to 2.4x), which is the design. Node-for-
|
|||
k 0.78 and the 64-register core at 1.09, so "near 0.9" is reached node-for-node by the window alone; what it does
|
||||
not survive is the node step a chip project would buy (an N3 core gives back 0.4x, an N2 core 0.8x).
|
||||
|
||||
### 4c. The adversary's 64-register core: is the window a defence? (the coordinator's order, 15:3x UK)
|
||||
|
||||
Every row in this section is a MODEL of a chip core, never a lower bound on what a chip maker can build; the
|
||||
synthesis gives the cost of the circuit as drawn, and a better circuit is always possible.
|
||||
|
||||
**The live-state analysis** (`tools/chip-model/rtl/flow/livestate.py`, run on build-3: programs drawn as the
|
||||
core testbench draws them, the class v4 op weights, a load on one instruction in 16 as the dependent memory wait,
|
||||
dst and src uniform over the window, the result fold reading every register at the end of the block; 64 drawn
|
||||
programs, 1,024 waits per row):
|
||||
|
||||
| Window R | Live values at a wait (mean, min to max) | Of which necessary (reach a later address or the result, transitively) | Dead writes per block |
|
||||
|---|---|---|---|
|
||||
| 8 (the class ISA) | 7.0 of 8 (6 to 7) | 6.9 | 2.9 percent |
|
||||
| 32 | 30.5 of 32 (29 to 31) | 30.0 | 2.9 percent |
|
||||
| 64 | 61.7 of 64 (59 to 63) | 61.0 | 2.8 percent |
|
||||
| 64 at a 1,024-instruction block | 61.5 of 64 (58 to 63) | 60.6 | 3.1 percent |
|
||||
|
||||
So under a fold that reads every register, 95 percent of the window is live AND necessary across every memory
|
||||
wait: the adversary cannot shrink the state it keeps by liveness, and recomputing a value instead of keeping it
|
||||
costs the dependent chain that produced it (every value feeds the result transitively). The window is a
|
||||
defence ONLY because of the fold rule; a fold that read 8 of the 64 registers would let the chip drop the rest
|
||||
(the dead fraction would rise toward the fraction never read before the fold), so the fold-reads-all rule is the
|
||||
design rule that goes with the window.
|
||||
|
||||
**What the adversary can do with the state it must keep** is make it cheaper per access, not smaller. The
|
||||
GPU-shaped row (4a, 64 registers in flops, every flop clocked every cycle, three 64:1 read muxes) is 9.7 pJ per
|
||||
lane-op at ASAP7. The forms a chip maker would use:
|
||||
|
||||
| Form of the 64-register state (per lane, 256 bytes) | pJ per lane-op ASAP7 | N3 | k at the lock (N3) | Label |
|
||||
|---|---|---|---|---|
|
||||
| flops, no clock gating, 64:1 read muxes (the 4a row) | 9.7 | 4.9 | 0.78 | synthesised; a model |
|
||||
| flops with the register-file clock gated (one of 64 registers written per cycle; the ICG cells allowed back in and inferred by Yosys, 320 gates) | 6.2 (225,441 cells; sequential 0.2) | 3.1 | 0.50 (0.70 node-for-node) | synthesised 16:0x UK; a model |
|
||||
| the same gating on the 32-register base, for the penalty | 4.5 (156,833 cells; sequential 0.15) | 2.3 | 0.37 (0.51 node-for-node) | synthesised 16:0x UK; a model |
|
||||
| the gated 32-register base PLACED AND ROUTED (SPEF, clock tree, 379,633 cells; steady state solved from 150 and 600 run cycles) | 6.7 (+49 percent over synthesis) | 3.4 | 0.54 (0.76 node-for-node) | placed 19:5x UK; the GDDR7 board at the lock 2.4x node-for-node, 2.8x a node ahead: the morning's headline figures to the digit |
|
||||
| the gated 64-register window core PLACED AND ROUTED (563,339 cells; parasitics estimated from global routing, the SPEF lost to a full disk; steady state solved from 150 and 600 run cycles; plus or minus 15 percent) | 9.45 (+52 percent over synthesis; N5 6.6, N2 3.4) | 4.8 | 0.77 (1.07 node-for-node, 0.55 at N2) | placed 21:3x UK; the GDDR7 board at the lock 2.0x node-for-node, 2.45x a node ahead, 2.9x two ahead; the window's residual against the placed base +2.75 pJ per lane-op, +0.23 of k at the lock |
|
||||
| latch-based register file (the clocked element halved; about 30 percent under the gated flop file, approximate) | about 0.7 x the gated row | | | modelled |
|
||||
| SRAM-banked state shared across time-multiplexed lanes (one execution port serving many lanes' streams in turn, each lane's 256 bytes in a bank): modelled 8.5 to 10.5 from the access energies; BUILT as `core_tm` (8 lanes x 64 registers in banks, one port, round-robin, ungated): 8.8 pJ per lane-op synthesised (173,426 cells), against the SIMD ungated 9.7 | 8.8 (built) | 4.4 | 0.71 | synthesised 16:2x UK: port sharing saves 0.9 pJ of units and the bank-select muxes take most of it back; NOT the lever; the multi-family adversary lane's macro-window core reads 5.9 (its FakeRAM term modelled 2.0 to 7.0 pJ per access), within 5 percent of the gated flop row, so the file's form is not the lever either |
|
||||
| values recomputed instead of kept | not available: 95 percent of the window is necessary (above) | | | measured on drawn programs |
|
||||
|
||||
The defence, then, is the penalty that remains after the adversary's best form: the gated 64-register file
|
||||
against the gated 32-register file (the two synthesised rows above when they land; measured: 6.2 against 4.5 pJ per lane-op at ASAP7 synthesised, a penalty of 1.7 pJ; PLACED 9.45 against 6.7, a
|
||||
penalty of 2.75 pJ, 1.4 at N3, +0.23 of k at the lock, 0.54 to 0.77 at N3 and 0.76 to 1.07 node-for-node; the
|
||||
analytic estimate had been 1.5 pJ). So on placed rows the window takes the GDDR7 board at the lock from 2.4x to
|
||||
2.0x node-for-node and from 2.8x to 2.45x a node ahead, for at most 5 percent per load on the card. Two corrections this
|
||||
forces: the honest adversary's BASE core is the gated one (k 0.37 at N3, 0.51 node-for-node), under the ungated
|
||||
0.56 and 0.78 of section 4, which are the GPU-shaped core a maker would not build; and placement costs more than
|
||||
the +20 to +40 percent estimated (the ungated placed base reads 11.3 against 6.9 pJ: wires plus a 2.5 pJ clock tree
|
||||
that gating removes), so the placed gated rows (on a rented pod, 17:30 UK) are the figures to serve. On the 32-lane core the same
|
||||
penalty applies per lane (the register file does not amortise), so the window moves the 32-lane core from k 0.45
|
||||
to about 0.57 at the lock at N3 (0.63 to about 0.80 node-for-node).
|
||||
|
||||
**The GPU side** (the hash lane's hand): the compiled allocation of the 64-register measurement pack (ptxas
|
||||
registers per thread, local-memory spill bytes, occupancy) and the rate beside the 8-register base, clock
|
||||
18:00 UK; until then the modelled reading stands: about 110 of 255 registers per thread, occupancy about half,
|
||||
the rate expected to hold under the latency-bound chain (the 5090 hides about 330,000 ops per hash before compute
|
||||
binds) and the energy to move little, the per-lane register traffic the unmeasured term.
|
||||
|
||||
The measured GPU side (the hash lane, 16:1x to 16:4x UK, RunPod secure pods, driver 580, the kit worker, 250 x
|
||||
2^24, nvidia-smi 1 Hz; ptxas from nvcc 12.8 -Xptxas -v on the pack's kernel; pods destroyed, USD 1.22):
|
||||
|
||||
| Card, pack | MH/s | W | microjoules per hash | registers per thread (ptxas) | spill | blocks per SM (occupancy) | per load | Label |
|
||||
|---|---|---|---|---|---|---|---|---|
|
||||
| 5090, the base (mx8-devnet-epoch0) | 141.74 | 303.1 | 2.139 | 30 | 0 B | 24 (4,080 warps) | 16.7 nJ | measured |
|
||||
| 5090, the window, arithmetic-only (hl-reg64, twice the base's work by construction) | 80.38 | 308.6 | 3.839 | 96 | 0 B | 20 (83 percent) | 15.0 nJ | measured: level per unit of work |
|
||||
| 5090, the window, full chain (hl-reg64c: every load's address mixes all 64 registers) | 70.96 | 320.3 | 4.513 | 88 | 0 B | 20 (83 percent) | 17.6 nJ (+5 percent) | measured |
|
||||
| 4090, the base | 62.67 | 208.9 | 3.333 | 29 | 0 B | 24 | 26.0 nJ | measured |
|
||||
| 4090, the window, arithmetic-only | 31.57 | 210.3 | 6.663 | 104 | 0 B | 16 (67 percent) | 26.0 nJ | measured: level |
|
||||
| 4090, the window, full chain | 31.38 | 216.5 | 6.898 | 87 | 0 B | 20 (83 percent) | 27.0 nJ (+4 percent) | measured |
|
||||
|
||||
So the card's side of the window defence is at most 5 percent per load: no spill on either card in either form, 88
|
||||
to 104 registers per thread, occupancy 67 to 83 percent, and the rate per unit of work held within 5 percent under
|
||||
the latency-bound chain. The sound class form is the full chain (the arithmetic-only fold fails the liveness rule;
|
||||
class string `+reg64c`, pack hl-v6-win with `check_window_liveness` in its suite). The chip's side (this section's
|
||||
gated rows) therefore carries the whole defence.
|
||||
|
||||
### 4d. The connected-state variant (cs64s27x16, the connected-state lane's structure) on the adversary's core
|
||||
|
||||
The connected-state lane's program (a 64-register window; per step a load whose address register is the previous
|
||||
block's last dst, the word landing in m_j, then a 27-instruction block whose first instruction reads m_j and every
|
||||
later one draws its src from the block's last four dsts, the last instruction injecting; 16 steps of text, 448
|
||||
instructions, 16 passes per block; its liveness tool: 63 of 64 live at every address, about 11 registers in the
|
||||
per-step dependent chain, about 20 touched per block) priced on the gated 64-register core with a 512-entry imem,
|
||||
the program drawn by those rules in the testbench (`CS` mode of `tb_core_common.vh`), synthesis-only:
|
||||
|
||||
| Row | Cells | pJ per lane-op ASAP7 | N5 | N3 | N2 | k at the lock N5 / N3 / N2 | k at stock N5 / N3 |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| cs64s27x16 on the gated 64-register core, 512 imem | 253,059 | 6.3 | 4.4 | 3.2 | 2.3 | 0.71 / 0.51 / 0.37 | 0.39 / 0.28 |
|
||||
| the class v4 draw on the gated 64-register core, 256 imem (4c) | 225,441 | 6.2 | 4.3 | 3.1 | 2.2 | 0.70 / 0.50 / 0.36 | 0.38 / 0.27 |
|
||||
|
||||
The chip's shadow for the program is 55,296 x 3.2 pJ = 0.18 microjoules per hash at N3 (0.24 node-for-node)
|
||||
against 0.13 for the genesis window on the same core; the card pays +0.6 percent for the window on the 5090 (the
|
||||
connected-state lane's measurement). The structure's other knobs do not reach the chip: the 16-pass loop and the
|
||||
448 text cost the shared imem about 0.1 pJ per lane-op, hot-set banking is not needed (the gated file charges only
|
||||
the written register), and the chain's width sets lane count, which is free. On the GDDR7 board at the lock the
|
||||
window moves the chip's edge by about 1.1x (3.6x to 3.3x node-for-node), under the 1.25x gate that lane set.
|
||||
|
||||
### 4e. The mixed-resource lane's FP32 units (class-v6-mixedfp) on the adversary's lane
|
||||
|
||||
The mixed-resource lane's candidate adds four FP32 families (fadd, fmul, ffma, fcvt) to the shadow's draw, every
|
||||
result injected by xor, with inputs masked to a 7-bit exponent range (never zero, denormal, NaN or Inf). The
|
||||
adversary's simplified units (`rtl/fp32_units.v`): an FMA with the 24 x 24 mantissa multiplier, a 100-bit
|
||||
alignment window, a full normaliser and RNE; a separate adder and multiplier; the int32 to float converter; the
|
||||
exponent path narrowed to the range; no NaN, Inf, denormal or flag logic. The lane: an 8 x 32-bit window, the
|
||||
four units, `d ^= bits(result)`. Routed with SPEF, random-input VCD, 42,936 cells, a 2 ns clock.
|
||||
|
||||
| Op (every unit evaluating each cycle: an UPPER bound per op, no operand isolation) | pJ per op ASAP7 | N5 | N3 | N2 | 5090 fp32_fma stock / lock | k at N3 vs stock / lock |
|
||||
|---|---|---|---|---|---|---|
|
||||
| fadd | 6.6 | 4.6 | 3.3 | 2.4 | 9.2 / 5.2 | 0.36 / 0.64 |
|
||||
| fmul | 7.0 | 4.9 | 3.5 | 2.5 | 9.2 / 5.2 | 0.38 / 0.68 |
|
||||
| ffma | 6.7 | 4.7 | 3.4 | 2.4 | 9.2 / 5.2 | 0.37 / 0.65 |
|
||||
| fcvt | 6.9 | 4.8 | 3.5 | 2.5 | 9.2 / 5.2 (cvt unmeasured on the card) | 0.38 / 0.67 |
|
||||
| random mix | 7.1 | 4.9 | 3.6 | 2.6 | | 0.39 / 0.68 |
|
||||
|
||||
Reading: the four read alike because all four units switch every cycle on the same operands, so each row is the
|
||||
upper bound for its op (a chip isolates the idle units; by cell share about ffma 3.5 to 4, fmul 2.5, fadd 2, fcvt
|
||||
1 pJ at ASAP7, approximate). Even on the upper bound the FP family is the chip's dearest per op relative to the
|
||||
card: k 0.65 at N3 at the lock against 0.18 for the integer ARX lane floor, because the card does an FMA for 5.2
|
||||
pJ (under its own int add at 6.2) while the chip's multiply, alignment and normaliser cost about three int ops.
|
||||
On the units' floors the shadow's k_eff rises from 0.097 (class v4) to about 0.14 at the fp12 mix and 0.17 at
|
||||
fp24 (0.20 and 0.24 with isolation taken as half), the core's per-op overhead on top. The GPU-cost budget (10
|
||||
percent of energy per hash) is the binding side, and the vendor-rounding question is the class's, not the chip's.
|
||||
|
||||
## 5. The chip edge at the measured k
|
||||
|
||||
`E_chip = E_mem + N_ops x e_chip` (absolute: the chip's shadow cost is 102,100 x 3.5 pJ = 0.36 microjoules per hash
|
||||
|
|
@ -270,7 +392,7 @@ to 2.3x and the strongest chips at 2.6x to 4.2x.
|
|||
| 6 | prmt, lop3 | 0.64, 0.72 | 11.5, 13.0 | 0.056, 0.055 | not drawn (RTL rows only) |
|
||||
| 7 | mulhi | 0.68 | 21.0 | 0.032 | yes (6) |
|
||||
| 8 | 32-lane shuffle (butterfly) | 0.63 | 29.4 | 0.021 | yes (8) |
|
||||
| 9 | L1 scratch read (8 KB flop array) against the card's L2 hit | ROW_SCRATCH_K | 1,400 | pending | not drawn |
|
||||
| 9 | L1 scratch read (8 KB flop array) against the card's L2 hit | 104 (pJ per read) | 1,400 | 0.074 | not drawn |
|
||||
| 10 | int8 8x8x8 tile, per MAC | ROW_TILE_K | 2.2 | pending | not drawn (the tensor lever is dead on other grounds) |
|
||||
|
||||
The order is set by the card's price, not the chip's: the chip pays 0.6 to 1.7 pJ for everything, and the card
|
||||
|
|
|
|||
|
|
@ -400,6 +400,14 @@ Rule: a case maps to a cell only where the cell's tests visibly answer it; cover
|
|||
- Cases:
|
||||
- INT-07 One v6 object agrees in node, pool, CPU verifier and each supported GPU host across activation.: partial: V6-12's clean-install half (one published object installed fresh, synced, mined, proved and paid on one host per artefact); the cross-host agreement half (node, pool, CPU verifier, each GPU host across activation) is harness:same-work's
|
||||
|
||||
### review:k-lane-shadow-k
|
||||
|
||||
- Command: `floor lane 2's placed and routed cores on ASAP7 in tools/chip-model/rtl (the rows in docs/analysis/class-v6/floor/shadow-k.md); the register landing 50ff1611f recorded ADV-06 against it before the recorder existed`
|
||||
- Box class: none
|
||||
- Fixtures: none
|
||||
- Cases:
|
||||
- ADV-06 Separate process advantage from specialisation: partial: the placed gated cores and the adversary's forms are modelled in shadow-k.md; the independent review remains
|
||||
|
||||
## Automated cases with no harness in the matrix (NOT RUN, the reason)
|
||||
|
||||
- GOV-02 Approve thresholds before results: the approval is recorded in the registry's approval field; the automated half (thresholds frozen before any run_status) is the gate rule landing by 21:00
|
||||
|
|
@ -409,7 +417,6 @@ Rule: a case maps to a cell only where the cell's tests visibly answer it; cover
|
|||
- GPU-06 Measure accepted work under ordinary connectivity: accepted work under ordinary connectivity needs the fault network F4
|
||||
- GPU-07 Survive sustained thermal and power operation: the sustained thermal and power soak has no harness tonight: the project's own rig mines nothing under the the earlier devnet off order
|
||||
- POW-05 Prevent amortised cheap winning attempts: amortised cheap winning attempts are the attack lanes' grind and era harnesses (tools/attack/f7-era, f9-grind), not in the release matrix; their rows come from those lanes
|
||||
- ADV-06 Separate process advantage from specialisation: process-advantage separation is the adversary lanes' chip study
|
||||
- ROT-03 Test miner-voted bring-forward governance: miner-voted bring-forward needs a vote harness on the fault network F4
|
||||
- ROT-04 Resist seed selection and faster evaluators: seed-selection resistance is the census harness (the class v6 invention lane), not yet in the matrix
|
||||
- EVM-01 Match the selected EVM semantics: no EVM conformance-vector harness is mapped tonight; the exec suite does not run the reference test vectors
|
||||
|
|
@ -551,4 +558,4 @@ Rule: a case maps to a cell only where the cell's tests visibly answer it; cover
|
|||
|
||||
## Count
|
||||
|
||||
171 automated cases: 82 mapped to a cell, 146 NOT RUN with a reason.
|
||||
171 automated cases: 83 mapped to a cell, 145 NOT RUN with a reason.
|
||||
|
|
|
|||
|
|
@ -952,7 +952,20 @@
|
|||
"updated": "2026-10-08T21:00:32.846Z",
|
||||
"evidence_record": {
|
||||
"reason": "the memory-clock ladder has no harness tonight: the project's own rig mines nothing under the the earlier devnet off order",
|
||||
"at": "2026-10-08T21:00:32.846Z"
|
||||
"at": "2026-10-08T21:00:32.846Z",
|
||||
"method": "static",
|
||||
"requirement_id": "GPU-04",
|
||||
"decision": "NOT RUN",
|
||||
"reviewer": "",
|
||||
"claim_impact": "",
|
||||
"release_identity": {
|
||||
"commit": "",
|
||||
"lockfile": "",
|
||||
"binary": "",
|
||||
"network_object": "",
|
||||
"activation": "",
|
||||
"profile_hashes": ""
|
||||
}
|
||||
},
|
||||
"in_progress_since": "2026-10-08 18:3x UK",
|
||||
"approvals": {
|
||||
|
|
@ -1187,7 +1200,20 @@
|
|||
"updated": "2026-10-08T21:00:32.846Z",
|
||||
"evidence_record": {
|
||||
"reason": "the sustained thermal and power soak has no harness tonight: the project's own rig mines nothing under the the earlier devnet off order",
|
||||
"at": "2026-10-08T21:00:32.846Z"
|
||||
"at": "2026-10-08T21:00:32.846Z",
|
||||
"method": "static",
|
||||
"requirement_id": "GPU-07",
|
||||
"decision": "NOT RUN",
|
||||
"reviewer": "",
|
||||
"claim_impact": "",
|
||||
"release_identity": {
|
||||
"commit": "",
|
||||
"lockfile": "",
|
||||
"binary": "",
|
||||
"network_object": "",
|
||||
"activation": "",
|
||||
"profile_hashes": ""
|
||||
}
|
||||
},
|
||||
"in_progress_since": "2026-10-08 18:3x UK",
|
||||
"approvals": {
|
||||
|
|
@ -2561,21 +2587,21 @@
|
|||
"manual_page": 28,
|
||||
"owner_lane": "k lane (a3c9601a6d4686fe1)",
|
||||
"run_status": "NOT RUN",
|
||||
"evidence_path": "docs/analysis/class-v6/multi-family-adversary.md; docs/analysis/class-v6/floor/shadow-k.md; docs/analysis/class-v6/multi-family-adversary.md",
|
||||
"run_id": "adversary-20261008-placed-8lane",
|
||||
"updated": "2026-10-08T20:41:20.327Z",
|
||||
"evidence_path": "docs/analysis/class-v6/multi-family-adversary.md; docs/analysis/class-v6/floor/shadow-k.md; docs/analysis/class-v6/floor/shadow-k.md",
|
||||
"run_id": "floor-k-20261008-rows-repeat",
|
||||
"updated": "2026-10-08T21:09:38.300Z",
|
||||
"evidence_record": {
|
||||
"requirement_id": "ADV-05",
|
||||
"decision": "NOT RUN",
|
||||
"method": "model",
|
||||
"cell": "adversary:mf-placed",
|
||||
"manifest_sha": "3a8874fef",
|
||||
"run_id": "adversary-20261008-placed-8lane",
|
||||
"evidence": "docs/analysis/class-v6/multi-family-adversary.md",
|
||||
"manifest_sha": "86e5b0fb",
|
||||
"run_id": "floor-k-20261008-rows-repeat",
|
||||
"evidence": "docs/analysis/class-v6/floor/shadow-k.md",
|
||||
"in_progress": true,
|
||||
"coverage": "partial: the SRAM macros, ports, wiring, clocking and the complete-board terms are modelled in sections 2.3, 5 and 6 with the unmodelled items carried as uncertainty; the calibration against an existing hardware block is the k lane's bare-lane row beside it, owed as a named comparison",
|
||||
"release_identity": {
|
||||
"commit": "3a8874fef",
|
||||
"commit": "86e5b0fb",
|
||||
"lockfile": "",
|
||||
"binary": "",
|
||||
"network_object": "",
|
||||
|
|
@ -2584,7 +2610,7 @@
|
|||
},
|
||||
"claim_impact": "",
|
||||
"reviewer": "",
|
||||
"at": "2026-10-08T20:41:20.327Z"
|
||||
"at": "2026-10-08T21:09:38.300Z"
|
||||
},
|
||||
"in_progress_since": "2026-10-08 18:3x UK",
|
||||
"approvals": {
|
||||
|
|
@ -2621,13 +2647,13 @@
|
|||
"decision": "NOT RUN",
|
||||
"method": "model",
|
||||
"cell": "adversary:mf-placed",
|
||||
"manifest_sha": "3a8874fef",
|
||||
"run_id": "adversary-20261008-placed-8lane",
|
||||
"evidence": "docs/analysis/class-v6/multi-family-adversary.md",
|
||||
"manifest_sha": "86e5b0fb",
|
||||
"run_id": "floor-k-20261008-rows-repeat",
|
||||
"evidence": "docs/analysis/class-v6/floor/shadow-k.md",
|
||||
"in_progress": true,
|
||||
"coverage": "partial: the SRAM macros, ports, wiring, clocking and the complete-board terms are modelled in sections 2.3, 5 and 6 with the unmodelled items carried as uncertainty; the calibration against an existing hardware block is the k lane's bare-lane row beside it, owed as a named comparison",
|
||||
"release_identity": {
|
||||
"commit": "3a8874fef",
|
||||
"commit": "86e5b0fb",
|
||||
"lockfile": "",
|
||||
"binary": "",
|
||||
"network_object": "",
|
||||
|
|
@ -2636,7 +2662,7 @@
|
|||
},
|
||||
"claim_impact": "",
|
||||
"reviewer": "",
|
||||
"at": "2026-10-08T20:41:20.327Z"
|
||||
"at": "2026-10-08T21:09:38.300Z"
|
||||
}
|
||||
}
|
||||
},
|
||||
|
|
@ -2669,24 +2695,29 @@
|
|||
"owner_lane": "k lane (a3c9601a6d4686fe1)",
|
||||
"run_status": "NOT RUN",
|
||||
"evidence_path": "docs/analysis/class-v6/floor/shadow-k.md",
|
||||
"run_id": "team-2026-10-08",
|
||||
"updated": "2026-10-08T20:10:24.556Z",
|
||||
"run_id": "floor-k-20261008-rows-repeat",
|
||||
"updated": "2026-10-08T21:09:38.300Z",
|
||||
"evidence_record": {
|
||||
"reason": "process-advantage separation is the adversary lanes' chip study",
|
||||
"at": "2026-10-08T20:10:24.556Z",
|
||||
"method": "static",
|
||||
"requirement_id": "ADV-06",
|
||||
"decision": "NOT RUN",
|
||||
"reviewer": "",
|
||||
"claim_impact": "",
|
||||
"method": "model",
|
||||
"cell": "review:k-lane-shadow-k",
|
||||
"manifest_sha": "86e5b0fb",
|
||||
"run_id": "floor-k-20261008-rows-repeat",
|
||||
"evidence": "docs/analysis/class-v6/floor/shadow-k.md",
|
||||
"in_progress": false,
|
||||
"coverage": "partial: the placed gated cores and the adversary's forms are modelled in shadow-k.md; the independent review remains",
|
||||
"release_identity": {
|
||||
"commit": "",
|
||||
"commit": "86e5b0fb",
|
||||
"lockfile": "",
|
||||
"binary": "",
|
||||
"network_object": "",
|
||||
"activation": "",
|
||||
"profile_hashes": ""
|
||||
}
|
||||
},
|
||||
"claim_impact": "",
|
||||
"reviewer": "",
|
||||
"at": "2026-10-08T21:09:38.300Z"
|
||||
},
|
||||
"in_progress_since": "2026-10-08 18:3x UK",
|
||||
"approvals": {
|
||||
|
|
@ -2717,6 +2748,28 @@
|
|||
"run_id": "team-2026-10-08",
|
||||
"evidence": "docs/analysis/class-v6/floor/shadow-k.md",
|
||||
"in_progress": true
|
||||
},
|
||||
"review:k-lane-shadow-k": {
|
||||
"requirement_id": "ADV-06",
|
||||
"decision": "NOT RUN",
|
||||
"method": "model",
|
||||
"cell": "review:k-lane-shadow-k",
|
||||
"manifest_sha": "86e5b0fb",
|
||||
"run_id": "floor-k-20261008-rows-repeat",
|
||||
"evidence": "docs/analysis/class-v6/floor/shadow-k.md",
|
||||
"in_progress": false,
|
||||
"coverage": "partial: the placed gated cores and the adversary's forms are modelled in shadow-k.md; the independent review remains",
|
||||
"release_identity": {
|
||||
"commit": "86e5b0fb",
|
||||
"lockfile": "",
|
||||
"binary": "",
|
||||
"network_object": "",
|
||||
"activation": "",
|
||||
"profile_hashes": ""
|
||||
},
|
||||
"claim_impact": "",
|
||||
"reviewer": "",
|
||||
"at": "2026-10-08T21:09:38.300Z"
|
||||
}
|
||||
}
|
||||
},
|
||||
|
|
@ -15225,6 +15278,8 @@
|
|||
"owner_lane": "CI steward with the hash lane (a690540514aa453d7) and the worker lane (a9e87343f008e0edd)",
|
||||
"run_status": "BLOCKED",
|
||||
"master_status": "PROPOSED / NOT RUN",
|
||||
"run_id": "canary-20261008-01",
|
||||
"evidence_path": "tools/ci/canary-check.sh; packaging/ota/publish-manifest.sh; packaging/ota/publish-public.sh",
|
||||
"updated": "2026-10-08T21:00:32.846Z",
|
||||
"evidence_record": {
|
||||
"requirement_id": "INT-07",
|
||||
|
|
@ -15295,9 +15350,7 @@
|
|||
"implementation_complete": null,
|
||||
"evidence_reproduced": null,
|
||||
"claim_authorised": null
|
||||
},
|
||||
"run_id": "canary-20261008-01",
|
||||
"evidence_path": "tools/ci/canary-check.sh; packaging/ota/publish-manifest.sh; packaging/ota/publish-public.sh"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "INT-08",
|
||||
|
|
|
|||
|
|
@ -23,8 +23,14 @@ CPUSET = $(if $(LEASE_ON),--cpuset-cpus {cpuset},)
|
|||
DOCKER := $(LEASEPFX) docker run --rm -u $(UID_GID) -e HOME=/tmp -e NUM_CORES=$(THREADS) $(CPUSET) -v $(WORK):/work
|
||||
ORFS := $(DOCKER) -w /OpenROAD-flow-scripts/flow $(ORFS_IMG)
|
||||
SIM := $(DOCKER) -w /work $(SIM_IMG)
|
||||
# NO_DOCKER=1: the box IS the ORFS image (a rented pod started from openroad/orfs:latest with iverilog installed
|
||||
# and /work a symlink to this directory); the same targets run natively.
|
||||
ifeq ($(NO_DOCKER),1)
|
||||
ORFS := env NUM_CORES=$(THREADS) bash -c 'cd /OpenROAD-flow-scripts/flow && exec "$$@"' --
|
||||
SIM := env
|
||||
endif
|
||||
|
||||
DESIGNS := arx mul prmt lop3 fold shfl xbar scratch tile core8 core32 core32r16 core8r64 core8i1k core8sel core32all
|
||||
DESIGNS := arx mul prmt lop3 fold shfl xbar scratch tile core8 core32 core32r16 core8r64 core8i1k core8sel core32all core8g core8r64g coretm cs64 fp32
|
||||
top = $(shell sed -n 's/^$(1) \([^ ]*\) .*/\1/p' flow/designs.txt)
|
||||
|
||||
# per-family simulation tags (the op field fixed per row where the family has several ops)
|
||||
|
|
@ -37,13 +43,18 @@ SIMS_shfl := mix
|
|||
SIMS_xbar := mix
|
||||
SIMS_scratch := mix
|
||||
SIMS_tile := mix
|
||||
SIMS_core8 := mix mixld:+loads=1
|
||||
SIMS_core32 := mix mixld:+loads=1
|
||||
SIMS_core32r16 := mix mixld:+loads=1
|
||||
SIMS_core8r64 := mix
|
||||
SIMS_core8 := s150:+cycles=150 s600:+cycles=600
|
||||
SIMS_core32 := s150:+cycles=150 s600:+cycles=600
|
||||
SIMS_core32r16 := s150:+cycles=150 s600:+cycles=600
|
||||
SIMS_core8r64 := s150:+cycles=150 s600:+cycles=600
|
||||
SIMS_core8i1k := mix
|
||||
SIMS_core8sel := mix
|
||||
SIMS_core32all := mix
|
||||
SIMS_core8g := s150:+cycles=150 s600:+cycles=600
|
||||
SIMS_core8r64g := s150:+cycles=150 s600:+cycles=600
|
||||
SIMS_coretm := mix
|
||||
SIMS_cs64 := s150:+cycles=150 s600:+cycles=600
|
||||
SIMS_fp32 := mix fadd:+op=0 fmul:+op=1 ffma:+op=2 fcvt:+op=3
|
||||
|
||||
.PHONY: rows table clean
|
||||
|
||||
|
|
@ -65,6 +76,7 @@ sim-%: flow-%
|
|||
|
||||
power-%: sim-%
|
||||
$(ORFS) make DESIGN_CONFIG=/work/flow/$*.mk RUN_SCRIPT=/work/flow/power.tcl RUN_LOG_NAME_STEM=power run 2>&1 | tee logs/power-$*.log
|
||||
rm -f sim/$*/*.vcd # the traces run to gigabytes each; the power log keeps every number (a full disk killed four runs on 8 October 2026)
|
||||
|
||||
# synthesis-only row (no placement, no parasitics): the 14:45 fallback
|
||||
synth-%:
|
||||
|
|
|
|||
12
tools/chip-model/rtl/flow/asap7_icg_model.v
Normal file
12
tools/chip-model/rtl/flow/asap7_icg_model.v
Normal file
|
|
@ -0,0 +1,12 @@
|
|||
// Behavioural model of the ASAP7 integrated clock gate (latch_posedge_precontrol: the enable is latched while the
|
||||
// clock is low, the gated clock is CLK AND the latched enable OR test). The liberty carries no function for it.
|
||||
module ICGx1_ASAP7_75t_R(input CLK, input ENA, input SE, output GCLK);
|
||||
reg en = 0;
|
||||
always @(CLK or ENA or SE) if (!CLK) en = ENA | SE;
|
||||
assign GCLK = CLK & en;
|
||||
endmodule
|
||||
module ICGx2_ASAP7_75t_R(input CLK, input ENA, input SE, output GCLK);
|
||||
reg en = 0;
|
||||
always @(CLK or ENA or SE) if (!CLK) en = ENA | SE;
|
||||
assign GCLK = CLK & en;
|
||||
endmodule
|
||||
|
|
@ -6,7 +6,7 @@ import re, sys, os, csv
|
|||
|
||||
work = sys.argv[1] if len(sys.argv) > 1 else '.'
|
||||
# ops per cycle per design (the per-op divisor) and the GPU row each family is read against
|
||||
OPS = {'arx': 1, 'mul': 1, 'prmt': 1, 'lop3': 1, 'fold': 1, 'shfl': 32, 'xbar': 32, 'scratch': 1, 'tile': 512, 'core8': 8, 'core32': 32, 'core32r16': 32, 'core8r64': 8, 'core8i1k': 8, 'core8sel': 8, 'core32all': 32}
|
||||
OPS = {'arx': 1, 'mul': 1, 'prmt': 1, 'lop3': 1, 'fold': 1, 'shfl': 32, 'xbar': 32, 'scratch': 1, 'tile': 512, 'core8': 8, 'core32': 32, 'core32r16': 32, 'core8r64': 8, 'core8i1k': 8, 'core8sel': 8, 'core32all': 32, 'core8g': 8, 'core8r64g': 8, 'coretm': 1, 'cs64': 8, 'fp32': 1}
|
||||
# 5090 measured pJ per counted op: (unlocked, at the 1,300 MHz lock); 15.1a
|
||||
GPU = {
|
||||
'arx:mix': (11.3, 6.2), 'arx:add': (11.3, 6.2), 'arx:sub': (11.3, 6.2), 'arx:xor': (11.3, 6.2), 'arx:or': (11.3, 6.2),
|
||||
|
|
@ -17,7 +17,7 @@ GPU = {
|
|||
'shfl:mix': (55.8, 29.4), 'xbar:mix': (55.8, 29.4),
|
||||
'scratch:mix': (2400.0, 1400.0), # the card's L2 hit (no shared-memory probe measured: owed)
|
||||
'tile:mix': (4.1, 2.2),
|
||||
'core8:mix': (11.3, 6.2), 'core8:mixld': (11.3, 6.2), 'core32:mix': (11.3, 6.2), 'core32:mixld': (11.3, 6.2), 'core32r16:mix': (11.3, 6.2), 'core8r64:mix': (11.3, 6.2), 'core8i1k:mix': (11.3, 6.2), 'core8sel:mix': (11.3, 6.2), 'core32all:mix': (11.3, 6.2), # the class v4 draw: read against int_arx (the packs job read the whole mix at 10.8 / 6.4) # dependent u8 m8n8k16 per MAC; the wide s8 tile reads 1.36 / 0.83
|
||||
'core8:mix': (11.3, 6.2), 'core8:mixld': (11.3, 6.2), 'core32:mix': (11.3, 6.2), 'core32:mixld': (11.3, 6.2), 'core32r16:mix': (11.3, 6.2), 'core8r64:mix': (11.3, 6.2), 'core8i1k:mix': (11.3, 6.2), 'core8sel:mix': (11.3, 6.2), 'core32all:mix': (11.3, 6.2), 'core8g:mix': (11.3, 6.2), 'core8r64g:mix': (11.3, 6.2), 'coretm:mix': (11.3, 6.2), 'cs64:mix': (11.3, 6.2), 'fp32:mix': (9.2, 5.2), 'fp32:fadd': (9.2, 5.2), 'fp32:fmul': (9.2, 5.2), 'fp32:ffma': (9.2, 5.2), 'fp32:fcvt': (9.2, 5.2), # the class v4 draw: read against int_arx (the packs job read the whole mix at 10.8 / 6.4) # dependent u8 m8n8k16 per MAC; the wide s8 tile reads 1.36 / 0.83
|
||||
}
|
||||
# per-node energy scaling from ASAP7 (a 7 nm-class predictive PDK at 0.70 V), approximate and claimed:
|
||||
# N7 -> N5 x0.70 (TSMC: "30 percent lower power at the same speed"), N5 -> N3E x0.72 (TSMC: 25 to 30 percent),
|
||||
|
|
@ -67,10 +67,10 @@ for d in OPS:
|
|||
pairs = [('vcd:s150', 'vcd:s600', 150, 600), ('vcd:short', 'vcd:synth', 150, 800), ('vcd:s150', 'vcd:s400', 150, 400)]
|
||||
for sh, lg, cs, cl in pairs:
|
||||
if sh in rows and lg in rows:
|
||||
rows['vcd:steady'] = steady(rows, period, sh, lg, cs, cl, 1028 if d in ('core8i1k', 'core32all') else LOAD_CYCLES)
|
||||
rows['vcd:steady'] = steady(rows, period, sh, lg, cs, cl, 1028 if d in ('core8i1k', 'core32all') else (452 if d == 'cs64' else LOAD_CYCLES))
|
||||
for tag, r in rows.items():
|
||||
sub = tag.split(':')[1] if ':' in tag else 'prop'
|
||||
key = f'{d}:{sub}' if sub in ('add','sub','xor','or','rotl','rotr','mul','mulhi','mad','mixld') else f'{d}:mix'
|
||||
key = f'{d}:{sub}' if sub in ('add','sub','xor','or','rotl','rotr','mul','mulhi','mad','mixld','fadd','fmul','ffma','fcvt') else f'{d}:mix'
|
||||
gpu = GPU.get(key, (None, None))
|
||||
pj = r['total'] * period * 1e-12 / OPS[d] * 1e12 # W * s / ops -> pJ
|
||||
pj_dyn = (r['internal'] + r['switching']) * period / OPS[d]
|
||||
|
|
|
|||
18
tools/chip-model/rtl/flow/core8g.mk
Normal file
18
tools/chip-model/rtl/flow/core8g.mk
Normal file
|
|
@ -0,0 +1,18 @@
|
|||
# ORFS design config for the programmable shadow core (core8: core_v6_8), ASAP7.
|
||||
export PLATFORM = asap7
|
||||
export DESIGN_NAME = core_v6_8
|
||||
export DESIGN_NICKNAME = core8g
|
||||
export VERILOG_FILES = /work/rtl/core_v6_8.v
|
||||
export VERILOG_INCLUDE_DIRS = /work/rtl
|
||||
export SDC_FILE = /work/flow/core8g.sdc
|
||||
export CORE_UTILIZATION = 40
|
||||
export CORE_ASPECT_RATIO = 1
|
||||
export CORE_MARGIN = 0.5
|
||||
export PLACE_DENSITY = 0.55
|
||||
export CORNER = TC
|
||||
export SKIP_LAST_GASP = 1
|
||||
export WORK_HOME = /work/out/core8g
|
||||
export SYNTH_MEMORY_MAX_BITS = 2000000
|
||||
# the adversary's register file: clock gating inferred (the ICG cells allowed back in)
|
||||
export INFER_CLKGATES = 1
|
||||
export DONT_USE_CELLS = *x1p*_ASAP7* *xp*_ASAP7* SDF*
|
||||
10
tools/chip-model/rtl/flow/core8g.sdc
Normal file
10
tools/chip-model/rtl/flow/core8g.sdc
Normal file
|
|
@ -0,0 +1,10 @@
|
|||
current_design core_v6_8
|
||||
set clk_name core_clock
|
||||
set clk_port_name clk
|
||||
set clk_period 1500
|
||||
set clk_io_pct 0.2
|
||||
set clk_port [get_ports $clk_port_name]
|
||||
create_clock -name $clk_name -period $clk_period $clk_port
|
||||
set non_clock_inputs [all_inputs -no_clocks]
|
||||
set_input_delay [expr $clk_period * $clk_io_pct] -clock $clk_name $non_clock_inputs
|
||||
set_output_delay [expr $clk_period * $clk_io_pct] -clock $clk_name [all_outputs]
|
||||
18
tools/chip-model/rtl/flow/core8r64g.mk
Normal file
18
tools/chip-model/rtl/flow/core8r64g.mk
Normal file
|
|
@ -0,0 +1,18 @@
|
|||
# ORFS design config for the programmable shadow core (core8: core_v6_8r64), ASAP7.
|
||||
export PLATFORM = asap7
|
||||
export DESIGN_NAME = core_v6_8r64
|
||||
export DESIGN_NICKNAME = core8r64g
|
||||
export VERILOG_FILES = /work/rtl/core_v6_8r64.v
|
||||
export VERILOG_INCLUDE_DIRS = /work/rtl
|
||||
export SDC_FILE = /work/flow/core8r64g.sdc
|
||||
export CORE_UTILIZATION = 40
|
||||
export CORE_ASPECT_RATIO = 1
|
||||
export CORE_MARGIN = 0.5
|
||||
export PLACE_DENSITY = 0.55
|
||||
export CORNER = TC
|
||||
export SKIP_LAST_GASP = 1
|
||||
export WORK_HOME = /work/out/core8r64g
|
||||
export SYNTH_MEMORY_MAX_BITS = 2000000
|
||||
# the adversary's register file: clock gating inferred (the ICG cells allowed back in)
|
||||
export INFER_CLKGATES = 1
|
||||
export DONT_USE_CELLS = *x1p*_ASAP7* *xp*_ASAP7* SDF*
|
||||
10
tools/chip-model/rtl/flow/core8r64g.sdc
Normal file
10
tools/chip-model/rtl/flow/core8r64g.sdc
Normal file
|
|
@ -0,0 +1,10 @@
|
|||
current_design core_v6_8r64
|
||||
set clk_name core_clock
|
||||
set clk_port_name clk
|
||||
set clk_period 1500
|
||||
set clk_io_pct 0.2
|
||||
set clk_port [get_ports $clk_port_name]
|
||||
create_clock -name $clk_name -period $clk_period $clk_port
|
||||
set non_clock_inputs [all_inputs -no_clocks]
|
||||
set_input_delay [expr $clk_period * $clk_io_pct] -clock $clk_name $non_clock_inputs
|
||||
set_output_delay [expr $clk_period * $clk_io_pct] -clock $clk_name [all_outputs]
|
||||
18
tools/chip-model/rtl/flow/coretm.mk
Normal file
18
tools/chip-model/rtl/flow/coretm.mk
Normal file
|
|
@ -0,0 +1,18 @@
|
|||
# ORFS design config for the programmable shadow core (core8: core_tm_8r64), ASAP7.
|
||||
export PLATFORM = asap7
|
||||
export DESIGN_NAME = core_tm_8r64
|
||||
export DESIGN_NICKNAME = coretm
|
||||
export VERILOG_FILES = /work/rtl/core_tm_8r64.v
|
||||
export VERILOG_INCLUDE_DIRS = /work/rtl
|
||||
export SDC_FILE = /work/flow/coretm.sdc
|
||||
export CORE_UTILIZATION = 40
|
||||
export CORE_ASPECT_RATIO = 1
|
||||
export CORE_MARGIN = 0.5
|
||||
export PLACE_DENSITY = 0.55
|
||||
export CORNER = TC
|
||||
export SKIP_LAST_GASP = 1
|
||||
export WORK_HOME = /work/out/coretm
|
||||
export SYNTH_MEMORY_MAX_BITS = 2000000
|
||||
# the adversary's register file: clock gating inferred (the ICG cells allowed back in)
|
||||
export INFER_CLKGATES = 1
|
||||
export DONT_USE_CELLS = *x1p*_ASAP7* *xp*_ASAP7* SDF*
|
||||
10
tools/chip-model/rtl/flow/coretm.sdc
Normal file
10
tools/chip-model/rtl/flow/coretm.sdc
Normal file
|
|
@ -0,0 +1,10 @@
|
|||
current_design core_tm_8r64
|
||||
set clk_name core_clock
|
||||
set clk_port_name clk
|
||||
set clk_period 1500
|
||||
set clk_io_pct 0.2
|
||||
set clk_port [get_ports $clk_port_name]
|
||||
create_clock -name $clk_name -period $clk_period $clk_port
|
||||
set non_clock_inputs [all_inputs -no_clocks]
|
||||
set_input_delay [expr $clk_period * $clk_io_pct] -clock $clk_name $non_clock_inputs
|
||||
set_output_delay [expr $clk_period * $clk_io_pct] -clock $clk_name [all_outputs]
|
||||
18
tools/chip-model/rtl/flow/cs64.mk
Normal file
18
tools/chip-model/rtl/flow/cs64.mk
Normal file
|
|
@ -0,0 +1,18 @@
|
|||
# ORFS design config for the programmable shadow core (core8: core_v6_8cs64), ASAP7.
|
||||
export PLATFORM = asap7
|
||||
export DESIGN_NAME = core_v6_8cs64
|
||||
export DESIGN_NICKNAME = cs64
|
||||
export VERILOG_FILES = /work/rtl/core_v6_8cs64.v
|
||||
export VERILOG_INCLUDE_DIRS = /work/rtl
|
||||
export SDC_FILE = /work/flow/cs64.sdc
|
||||
export CORE_UTILIZATION = 40
|
||||
export CORE_ASPECT_RATIO = 1
|
||||
export CORE_MARGIN = 0.5
|
||||
export PLACE_DENSITY = 0.55
|
||||
export CORNER = TC
|
||||
export SKIP_LAST_GASP = 1
|
||||
export WORK_HOME = /work/out/cs64
|
||||
export SYNTH_MEMORY_MAX_BITS = 2000000
|
||||
# the adversary's register file: clock gating inferred (the ICG cells allowed back in)
|
||||
export INFER_CLKGATES = 1
|
||||
export DONT_USE_CELLS = *x1p*_ASAP7* *xp*_ASAP7* SDF*
|
||||
10
tools/chip-model/rtl/flow/cs64.sdc
Normal file
10
tools/chip-model/rtl/flow/cs64.sdc
Normal file
|
|
@ -0,0 +1,10 @@
|
|||
current_design core_v6_8cs64
|
||||
set clk_name core_clock
|
||||
set clk_port_name clk
|
||||
set clk_period 1500
|
||||
set clk_io_pct 0.2
|
||||
set clk_port [get_ports $clk_port_name]
|
||||
create_clock -name $clk_name -period $clk_period $clk_port
|
||||
set non_clock_inputs [all_inputs -no_clocks]
|
||||
set_input_delay [expr $clk_period * $clk_io_pct] -clock $clk_name $non_clock_inputs
|
||||
set_output_delay [expr $clk_period * $clk_io_pct] -clock $clk_name [all_outputs]
|
||||
|
|
@ -14,3 +14,8 @@ core8r64 core_v6_8r64 1500
|
|||
core8i1k core_v6_8i1k 1500
|
||||
core8sel core_v6_8sel 1500
|
||||
core32all core_v6_32all 1500
|
||||
core8g core_v6_8 1500
|
||||
core8r64g core_v6_8r64 1500
|
||||
coretm core_tm_8r64 1500
|
||||
cs64 core_v6_8cs64 1500
|
||||
fp32 lane_fp32 2000
|
||||
|
|
|
|||
15
tools/chip-model/rtl/flow/fp32.mk
Normal file
15
tools/chip-model/rtl/flow/fp32.mk
Normal file
|
|
@ -0,0 +1,15 @@
|
|||
# ORFS design config for the mul shadow-core family (top lane_fp32), ASAP7.
|
||||
export PLATFORM = asap7
|
||||
export DESIGN_NAME = lane_fp32
|
||||
export DESIGN_NICKNAME = fp32
|
||||
export VERILOG_FILES = /work/rtl/fp32_units.v
|
||||
export VERILOG_INCLUDE_DIRS = /work/rtl
|
||||
export SDC_FILE = /work/flow/fp32.sdc
|
||||
export CORE_UTILIZATION = 40
|
||||
export CORE_ASPECT_RATIO = 1
|
||||
export CORE_MARGIN = 0.5
|
||||
export PLACE_DENSITY = 0.55
|
||||
export CORNER = TC
|
||||
export SKIP_LAST_GASP = 1
|
||||
export WORK_HOME = /work/out/fp32
|
||||
|
||||
10
tools/chip-model/rtl/flow/fp32.sdc
Normal file
10
tools/chip-model/rtl/flow/fp32.sdc
Normal file
|
|
@ -0,0 +1,10 @@
|
|||
current_design lane_fp32
|
||||
set clk_name core_clock
|
||||
set clk_port_name clk
|
||||
set clk_period 2000
|
||||
set clk_io_pct 0.2
|
||||
set clk_port [get_ports $clk_port_name]
|
||||
create_clock -name $clk_name -period $clk_period $clk_port
|
||||
set non_clock_inputs [all_inputs -no_clocks]
|
||||
set_input_delay [expr $clk_period * $clk_io_pct] -clock $clk_name $non_clock_inputs
|
||||
set_output_delay [expr $clk_period * $clk_io_pct] -clock $clk_name [all_outputs]
|
||||
79
tools/chip-model/rtl/flow/livestate.py
Normal file
79
tools/chip-model/rtl/flow/livestate.py
Normal file
|
|
@ -0,0 +1,79 @@
|
|||
#!/usr/bin/env python3
|
||||
"""Live-state analysis of a drawn shadow program on an R-register window (the coordinator's order, 15:3x UK).
|
||||
The program is drawn as the core testbench draws it: NPROG instructions with the class v4 op weights, a load on
|
||||
one instruction in 16 (the dependent memory wait), dst/src/src2 uniform over the R registers, and the result
|
||||
fold reading every register at the end of the block. For every load (wait) the script reports the live set:
|
||||
registers whose current value is read later (by an instruction, a later address, or the fold) before being
|
||||
overwritten, split into those that feed a later ADDRESS or the RESULT and those that die inside an arithmetic
|
||||
block. Dead writes (overwritten before any read) are counted too. Usage: livestate.py R [NPROG] [seeds]"""
|
||||
import random, sys
|
||||
R = int(sys.argv[1]) if len(sys.argv) > 1 else 64
|
||||
N = int(sys.argv[2]) if len(sys.argv) > 2 else 256
|
||||
SEEDS = int(sys.argv[3]) if len(sys.argv) > 3 else 64
|
||||
W = [('add',12),('xor',10),('mul',8),('mad',8),('shfl',8),('rotl',7),('sub',6),('mulhi',6),('rotr',6),('or',4)]
|
||||
ops = [o for o,w in W for _ in range(w)]
|
||||
def draw(rng):
|
||||
prog = []
|
||||
for k in range(N):
|
||||
op = 'load' if k % 16 == 15 else rng.choice(ops)
|
||||
d, s, s2 = rng.randrange(R), rng.randrange(R), rng.randrange(R)
|
||||
reads = [s] if op not in ('load',) else [s] # the load's address comes from src
|
||||
if op in ('add','xor','mul','mad','sub','or','rotr','shfl','mulhi','rotl'): reads.append(d) # dst is read too (r[d] op= ...)
|
||||
if op == 'mad': reads.append(s2)
|
||||
if op == 'rotl': reads = [d]
|
||||
prog.append((op, d, reads))
|
||||
return prog
|
||||
tot_live = tot_addr = tot_dead = tot_waits = 0
|
||||
live_min, live_max = R, 0
|
||||
for seed in range(SEEDS):
|
||||
rng = random.Random(seed)
|
||||
prog = draw(rng)
|
||||
# a value's "version" = (reg, write index); the fold at the end reads every register
|
||||
# forward pass: for each instruction i and register r, next read of r's current value before its next write
|
||||
n = len(prog)
|
||||
# necessity: a version is NECESSARY if it reaches an address (a load's src) or the fold, transitively
|
||||
# compute transitively by backward dataflow over versions
|
||||
writes_at = {} # (i) -> reg written
|
||||
# build version ids: version of reg r valid after instruction i
|
||||
cur = {r: ('init', r) for r in range(R)}
|
||||
uses = {} # version -> list of (consumer index, consumer version or 'addr'/'fold')
|
||||
versions = set(cur.values())
|
||||
deps = {} # version -> set of versions it reads
|
||||
for i, (op, d, reads) in enumerate(prog):
|
||||
srcs = [cur[r] for r in reads]
|
||||
if op == 'load':
|
||||
v = ('load', i); deps[v] = set() # the returned word: its ADDRESS depends on srcs
|
||||
for s in srcs: uses.setdefault(s, []).append(('addr', i))
|
||||
else:
|
||||
v = (op, i); deps[v] = set(srcs)
|
||||
for s in srcs: uses.setdefault(s, []).append(('op', i))
|
||||
cur[d] = v; versions.add(v)
|
||||
fold = set(cur.values())
|
||||
# necessary = reaches an address or the fold
|
||||
necessary = set(fold)
|
||||
for v, us in uses.items():
|
||||
if any(k == 'addr' for k, _ in us): necessary.add(v)
|
||||
changed = True
|
||||
while changed:
|
||||
changed = False
|
||||
for v in list(versions):
|
||||
if v in necessary:
|
||||
for s in deps.get(v, ()):
|
||||
if s not in necessary: necessary.add(s); changed = True
|
||||
# per wait: the versions live at the load (written before it, read after it)
|
||||
last_read = {}
|
||||
for v, us in uses.items():
|
||||
last_read[v] = max(i for _, i in us)
|
||||
for v in fold: last_read[v] = n
|
||||
written_at = {v: (v[1] if v[0] != 'init' else -1) for v in versions}
|
||||
for i, (op, d, reads) in enumerate(prog):
|
||||
if op != 'load': continue
|
||||
live = [v for v in versions if written_at[v] < i and last_read.get(v, -1) > i]
|
||||
nec = [v for v in live if v in necessary]
|
||||
tot_live += len(live); tot_addr += len(nec); tot_waits += 1
|
||||
live_min = min(live_min, len(live)); live_max = max(live_max, len(live))
|
||||
dead = sum(1 for v in versions if v[0] not in ('init',) and v not in uses and v not in fold)
|
||||
tot_dead += dead
|
||||
print(f'R = {R}, NPROG = {N}, {SEEDS} drawn programs, {tot_waits} waits')
|
||||
print(f'live values at a wait: mean {tot_live/tot_waits:.1f} of {R} (min {live_min}, max {live_max}); of which necessary (reach a later address or the result): {tot_addr/tot_waits:.1f}')
|
||||
print(f'dead writes (overwritten before any read): {tot_dead/SEEDS:.1f} per {N}-instruction block ({100*tot_dead/SEEDS/N:.1f} percent)')
|
||||
9
tools/chip-model/rtl/flow/pod.sh
Executable file
9
tools/chip-model/rtl/flow/pod.sh
Executable file
|
|
@ -0,0 +1,9 @@
|
|||
#!/usr/bin/env bash
|
||||
# pod.sh: bootstrap a rented pod started from openroad/orfs:latest (RunPod, root): iverilog, /work -> this dir.
|
||||
set -euo pipefail
|
||||
cd "$(dirname "$0")/.."
|
||||
export DEBIAN_FRONTEND=noninteractive
|
||||
command -v iverilog >/dev/null || { apt-get update -qq >/dev/null 2>&1; apt-get install -y -qq iverilog rsync python3 >/dev/null 2>&1; }
|
||||
[ -e /work ] || ln -s "$(pwd)" /work
|
||||
export PATH=/OpenROAD-flow-scripts/tools/install/OpenROAD/bin:/OpenROAD-flow-scripts/tools/install/yosys/bin:$PATH
|
||||
echo "pod ready: $(nproc) cores, $(free -g | awk '/Mem/{print $2}') GB, yosys $(yosys -V | cut -d' ' -f2), $(which iverilog)"
|
||||
|
|
@ -4,6 +4,7 @@ set -euo pipefail
|
|||
name=$1; tag=$2; shift 2
|
||||
out=/work/sim/$name
|
||||
simcells=$(yosys-config --datdir)/simcells.v
|
||||
iverilog -g2005 -I /work/tb -o $out/sim_$tag $out/sim_net.v /work/tb/tb_$(sed -n "s/^$name \([^ ]*\) .*/\1/p" /work/flow/designs.txt).v $simcells
|
||||
tbfile=${TB:-/work/tb/tb_$(sed -n "s/^$name \([^ ]*\) .*/\1/p" /work/flow/designs.txt).v}
|
||||
iverilog -g2005 -I /work/tb -o $out/sim_$tag $out/sim_net.v /work/flow/asap7_icg_model.v $tbfile $simcells
|
||||
( cd $out && vvp -n sim_$tag "$@" | tee sim_$tag.log && mv dump.vcd $tag.vcd )
|
||||
ls -la $out/$tag.vcd
|
||||
|
|
|
|||
82
tools/chip-model/rtl/rtl/core_tm.v
Normal file
82
tools/chip-model/rtl/rtl/core_tm.v
Normal file
|
|
@ -0,0 +1,82 @@
|
|||
// The adversary's time-multiplexed core: ONE execution port (every class unit, once) serving LANES lanes' instruction
|
||||
// streams round-robin, each lane's state (REGS x 32-bit) kept in its own bank; the imem and sequencer shared.
|
||||
// One lane-op per cycle. Compared with core_v6 at the same LANES x REGS this removes LANES-1 copies of the units and
|
||||
// keeps the register state and the imem: the energy per lane-op is the state's cost plus one unit set's.
|
||||
// The butterfly shuffle across lanes needs every lane's source register in the same cycle, so the shuffle reads the
|
||||
// bank-wide source column (as the SIMD core does) and the lane in turn takes its word. Loads return on ld_val.
|
||||
`include "lane_common.vh"
|
||||
module core_tm #(parameter LANES = 8, parameter LOG_LANES = 3, parameter REGS = 64, parameter LOG_REGS = 6,
|
||||
parameter IW = 40, parameter IMEM_LOG = 8) (
|
||||
input clk, input rst, input run,
|
||||
input prog_we, input [9:0] prog_addr, input [IW-1:0] prog_data,
|
||||
input cfg_en, input [9:0] cfg_n, input [31:0] cfg_m, input [4:0] cfg_r, input [31:0] cfg_wm, input [31:0] cfg_off, input [31:0] cfg_mask, input [63:0] cfg_sel,
|
||||
input [31:0] ld_val,
|
||||
output [31:0] addr, output [31:0] out);
|
||||
localparam IMEM = 1 << IMEM_LOG;
|
||||
reg [IW-1:0] imem [0:IMEM-1];
|
||||
reg [IMEM_LOG-1:0] pc; reg [IMEM_LOG-1:0] n_q; reg [IW-1:0] ir; reg [LOG_LANES-1:0] lane;
|
||||
reg [31:0] m_q, wm_q, off_q, mask_q; reg [4:0] r_q;
|
||||
integer i;
|
||||
// the sequencer: the same instruction is issued to each lane in turn (LANES cycles per instruction)
|
||||
always @(posedge clk) begin
|
||||
if (prog_we) imem[prog_addr[IMEM_LOG-1:0]] <= prog_data;
|
||||
if (rst) begin pc <= 0; ir <= 0; lane <= 0; n_q <= {IMEM_LOG{1'b1}}; m_q <= 32'h9e3779b1; r_q <= 5'd13; wm_q <= 32'h0fffffc0; off_q <= 3; mask_q <= 32'h0fffffff; end
|
||||
else begin
|
||||
if (cfg_en) begin n_q <= cfg_n[IMEM_LOG-1:0]; m_q <= cfg_m | 1; r_q <= cfg_r; wm_q <= cfg_wm; off_q <= cfg_off; mask_q <= cfg_mask; end
|
||||
if (run) begin
|
||||
if (lane == {LOG_LANES{1'b1}}) begin ir <= imem[pc]; pc <= (pc == n_q) ? {IMEM_LOG{1'b0}} : pc + 1'b1; end
|
||||
lane <= lane + 1'b1;
|
||||
end
|
||||
end
|
||||
end
|
||||
wire [3:0] op = ir[3:0];
|
||||
wire [LOG_REGS-1:0] dst = ir[4 +: LOG_REGS]; wire [LOG_REGS-1:0] src = ir[4+LOG_REGS +: LOG_REGS]; wire [LOG_REGS-1:0] src2 = ir[4+2*LOG_REGS +: LOG_REGS];
|
||||
wire [4:0] imm = ir[4+3*LOG_REGS +: 5]; wire [7:0] aux = ir[9+3*LOG_REGS +: 8];
|
||||
wire is_load = (op == 4'd12);
|
||||
wire [4:0] rn = (imm == 0) ? 5'd1 : imm;
|
||||
wire [LOG_LANES-1:0] smask = imm[LOG_LANES-1:0];
|
||||
// the banked state: one bank per lane, read through the lane select (a chip's SRAM bank select)
|
||||
reg [31:0] rf [0:LANES*REGS-1];
|
||||
wire [31:0] d = rf[lane*REGS + dst];
|
||||
wire [31:0] s = rf[lane*REGS + src];
|
||||
wire [31:0] s2 = rf[lane*REGS + src2];
|
||||
wire [31:0] sx = rf[(lane ^ smask)*REGS + src]; // the shuffle partner's source word
|
||||
// the one execution port
|
||||
function [7:0] pick; input [63:0] b; input [3:0] k; reg [7:0] v;
|
||||
begin v = b[8*k[2:0] +: 8]; pick = k[3] ? {8{v[7]}} : v; end
|
||||
endfunction
|
||||
wire [4:0] sn = (s[4:0] == 0) ? 5'd1 : s[4:0];
|
||||
wire mad = (op == 4'd8);
|
||||
wire [63:0] p = (mad ? s : d) * (mad ? s2 : s);
|
||||
wire [63:0] bytes = {s, d}; wire [15:0] sel = {aux, aux};
|
||||
wire [31:0] prm = {pick(bytes, sel[15:12]), pick(bytes, sel[11:8]), pick(bytes, sel[7:4]), pick(bytes, sel[3:0])};
|
||||
reg [31:0] lp; integer b;
|
||||
always @* for (b = 0; b < 32; b = b + 1) lp[b] = aux[{d[b], s[b], s2[b]}];
|
||||
reg [31:0] r;
|
||||
always @* begin
|
||||
case (op)
|
||||
4'd0, 4'd13: r = d + s;
|
||||
4'd1, 4'd15: r = d - s;
|
||||
4'd2, 4'd14: r = d ^ s;
|
||||
4'd3: r = d | s;
|
||||
4'd4: r = `ROTL32(d, rn);
|
||||
4'd5: r = `ROTR32(d, sn);
|
||||
4'd6: r = p[31:0];
|
||||
4'd7: r = p[63:32];
|
||||
4'd8: r = p[31:0] + d;
|
||||
4'd9: r = d ^ sx;
|
||||
4'd10: r = prm;
|
||||
4'd11: r = lp;
|
||||
default: r = ld_val ^ (32'h9e3779b9 * (lane + 1));
|
||||
endcase
|
||||
end
|
||||
wire [31:0] fx = s * m_q;
|
||||
wire [4:0] frn = (r_q == 0) ? 5'd1 : r_q;
|
||||
wire [31:0] fy = `ROTL32(fx, frn);
|
||||
assign addr = is_load ? (((fy & wm_q) | off_q) & mask_q) : 32'd0;
|
||||
always @(posedge clk) begin
|
||||
if (rst) begin for (i = 0; i < LANES*REGS; i = i + 1) rf[i] <= 32'h9e3779b9 * (i + 1); end
|
||||
else if (run) rf[lane*REGS + dst] <= r;
|
||||
end
|
||||
assign out = r;
|
||||
endmodule
|
||||
7
tools/chip-model/rtl/rtl/core_tm_8r64.v
Normal file
7
tools/chip-model/rtl/rtl/core_tm_8r64.v
Normal file
|
|
@ -0,0 +1,7 @@
|
|||
`include "core_tm.v"
|
||||
module core_tm_8r64(input clk, input rst, input run, input prog_we, input [9:0] prog_addr, input [39:0] prog_data,
|
||||
input cfg_en, input [9:0] cfg_n, input [31:0] cfg_m, input [4:0] cfg_r, input [31:0] cfg_wm, input [31:0] cfg_off, input [31:0] cfg_mask, input [63:0] cfg_sel,
|
||||
input [31:0] ld_val, output [31:0] addr, output [31:0] out);
|
||||
core_tm #(.LANES(8), .LOG_LANES(3), .REGS(64), .LOG_REGS(6), .IW(40)) c(.clk(clk), .rst(rst), .run(run), .prog_we(prog_we), .prog_addr(prog_addr), .prog_data(prog_data),
|
||||
.cfg_en(cfg_en), .cfg_n(cfg_n), .cfg_m(cfg_m), .cfg_r(cfg_r), .cfg_wm(cfg_wm), .cfg_off(cfg_off), .cfg_mask(cfg_mask), .cfg_sel(cfg_sel), .ld_val(ld_val), .addr(addr), .out(out));
|
||||
endmodule
|
||||
7
tools/chip-model/rtl/rtl/core_v6_8cs64.v
Normal file
7
tools/chip-model/rtl/rtl/core_v6_8cs64.v
Normal file
|
|
@ -0,0 +1,7 @@
|
|||
`include "core_v6.v"
|
||||
module core_v6_8cs64(input clk, input rst, input run, input prog_we, input [9:0] prog_addr, input [40-1:0] prog_data,
|
||||
input cfg_en, input [9:0] cfg_n, input [63:0] cfg_sel, input [31:0] cfg_m, input [4:0] cfg_r, input [31:0] cfg_wm, input [31:0] cfg_off, input [31:0] cfg_mask,
|
||||
input [31:0] ld_val, output [31:0] addr, output [31:0] out);
|
||||
core_v6 #(.LANES(8), .LOG_LANES(3), .REGS(64), .LOG_REGS(6), .IW(40), .IMEM_LOG(9)) c(.clk(clk), .rst(rst), .run(run), .prog_we(prog_we), .prog_addr(prog_addr), .prog_data(prog_data),
|
||||
.cfg_en(cfg_en), .cfg_n(cfg_n), .cfg_sel(cfg_sel), .cfg_m(cfg_m), .cfg_r(cfg_r), .cfg_wm(cfg_wm), .cfg_off(cfg_off), .cfg_mask(cfg_mask), .ld_val(ld_val), .addr(addr), .out(out));
|
||||
endmodule
|
||||
123
tools/chip-model/rtl/rtl/fp32_units.v
Normal file
123
tools/chip-model/rtl/rtl/fp32_units.v
Normal file
|
|
@ -0,0 +1,123 @@
|
|||
// The adversary's simplified FP32 units for the mixed-resource lane's candidate (class-v6-mixedfp): inputs are
|
||||
// f(x) = as_float((x & 0x807FFFFF) | ((96 + ((x >> 23) & 63)) << 23)): never zero, denormal, NaN or Inf; exponents
|
||||
// in [96, 159]; results normal or +0 (exact cancellation). The units drop NaN/Inf/denormal handling and the flags,
|
||||
// keep the full 24-bit mantissa path, a full alignment and a full normaliser (the mantissas are uniform), RNE.
|
||||
// Each lane reads two or three registers of an 8 x 32-bit window, applies f(), computes, xors the bits into dst.
|
||||
`include "lane_common.vh"
|
||||
// ---- the shared pieces ----
|
||||
module fp_unpack(input [31:0] x, output s, output [8:0] e, output [23:0] m);
|
||||
assign s = x[31];
|
||||
assign e = 9'd96 + {3'b0, x[28:23]}; // the masked exponent, 96..159
|
||||
assign m = {1'b1, x[22:0]};
|
||||
endmodule
|
||||
module lzc48(input [47:0] v, output reg [5:0] n); // leading-zero count (v != 0)
|
||||
integer i; always @* begin n = 6'd47; for (i = 47; i >= 0; i = i - 1) if (v[i]) begin n = 6'd47 - i; i = -1; end end
|
||||
endmodule
|
||||
module lzc32(input [31:0] v, output reg [5:0] n);
|
||||
integer i; always @* begin n = 6'd31; for (i = 31; i >= 0; i = i - 1) if (v[i]) begin n = 6'd31 - i; i = -1; end end
|
||||
endmodule
|
||||
// ---- the FMA: fma(a, b, c) = a*b + c, one rounding (RNE), exponents in the lane's ranges ----
|
||||
module fp_fma(input [31:0] a, input [31:0] b, input [31:0] c, output [31:0] y);
|
||||
wire sa, sb, sc; wire [8:0] ea, eb, ec; wire [23:0] ma, mb, mc;
|
||||
fp_unpack ua(a, sa, ea, ma); fp_unpack ub(b, sb, eb, mb); fp_unpack uc(c, sc, ec, mc);
|
||||
wire [47:0] prod = ma * mb; // 48-bit product, binary point after bit 46
|
||||
wire sp = sa ^ sb;
|
||||
wire [9:0] ep = {1'b0, ea} + {1'b0, eb} - 10'd127; // product exponent (bias kept), 65..192
|
||||
// align the addend to the product: a 100-bit window keeps full precision for exponent gaps up to about 96 (the lane's bound)
|
||||
wire [9:0] diff = (ep >= {1'b0, ec}) ? ep - {1'b0, ec} : {1'b0, ec} - ep;
|
||||
wire prod_big = (ep >= {1'b0, ec});
|
||||
wire [99:0] pw = {2'b0, prod, 50'b0};
|
||||
wire [99:0] cw = {2'b0, mc, 24'b0, 50'b0}; // the addend at the product's scale when exponents equal
|
||||
wire [6:0] sh = (diff > 10'd99) ? 7'd99 : diff[6:0];
|
||||
wire [99:0] smw = prod_big ? (cw >> sh) : (pw >> sh);
|
||||
wire [99:0] bgw = prod_big ? pw : cw;
|
||||
wire sbig = prod_big ? sp : sc; wire ssmall = prod_big ? sc : sp;
|
||||
wire [9:0] ebig = prod_big ? ep : {1'b0, ec};
|
||||
wire [100:0] sum = (sbig == ssmall) ? ({1'b0, bgw} + {1'b0, smw}) : ({1'b0, bgw} - {1'b0, smw});
|
||||
wire [100:0] mag = sum[100] ? (~sum + 1'b1) : sum; // two's complement when the subtraction went negative
|
||||
wire ssum = sum[100] ? ssmall : sbig;
|
||||
// normalise: find the leading one in the 101-bit magnitude
|
||||
reg [6:0] lz; integer i;
|
||||
always @* begin lz = 7'd100; for (i = 100; i >= 0; i = i - 1) if (mag[i]) begin lz = 7'd100 - i; i = -1; end end
|
||||
wire [100:0] norm = mag << lz; // leading one at bit 100
|
||||
wire [23:0] mant = norm[100:77];
|
||||
wire guard = norm[76]; wire sticky = |norm[75:0];
|
||||
wire round_up = guard & (sticky | mant[0]);
|
||||
wire [24:0] mr = {1'b0, mant} + round_up;
|
||||
wire carry = mr[24];
|
||||
wire [9:0] eres = ebig + 10'd2 - lz + carry; // the leading one of bgw sat at bit 98 (two headroom bits)
|
||||
wire zero = (mag == 0);
|
||||
wire [7:0] eout = eres[7:0];
|
||||
assign y = zero ? 32'h0 : {ssum, eout, carry ? mr[23:1] : mr[22:0]};
|
||||
endmodule
|
||||
// ---- the adder and the multiplier as their own units ----
|
||||
module fp_add(input [31:0] a, input [31:0] b, output [31:0] y);
|
||||
wire sa, sb; wire [8:0] ea, eb; wire [23:0] ma, mb;
|
||||
fp_unpack ua(a, sa, ea, ma); fp_unpack ub(b, sb, eb, mb);
|
||||
wire abig = (ea > eb) || (ea == eb && ma >= mb);
|
||||
wire [8:0] ebig = abig ? ea : eb; wire [8:0] esm = abig ? eb : ea;
|
||||
wire [23:0] mbig = abig ? ma : mb; wire [23:0] msm = abig ? mb : ma;
|
||||
wire sbig = abig ? sa : sb; wire ssm = abig ? sb : sa;
|
||||
wire [8:0] diff = ebig - esm; wire [6:0] sh = (diff > 9'd70) ? 7'd70 : diff[6:0];
|
||||
wire [73:0] bw = {1'b0, mbig, 49'b0}; wire [73:0] sw = {1'b0, msm, 49'b0} >> sh;
|
||||
wire [74:0] sum = (sbig == ssm) ? ({1'b0, bw} + {1'b0, sw}) : ({1'b0, bw} - {1'b0, sw});
|
||||
reg [6:0] lz; integer i;
|
||||
always @* begin lz = 7'd74; for (i = 74; i >= 0; i = i - 1) if (sum[i]) begin lz = 7'd74 - i; i = -1; end end
|
||||
wire [74:0] norm = sum << lz;
|
||||
wire [23:0] mant = norm[74:51]; wire guard = norm[50]; wire sticky = |norm[49:0];
|
||||
wire round_up = guard & (sticky | mant[0]);
|
||||
wire [24:0] mr = {1'b0, mant} + round_up; wire carry = mr[24];
|
||||
wire [9:0] eres = {1'b0, ebig} + 10'd1 - lz + carry;
|
||||
wire zero = (sum == 0);
|
||||
assign y = zero ? 32'h0 : {sbig, eres[7:0], carry ? mr[23:1] : mr[22:0]};
|
||||
endmodule
|
||||
module fp_mul(input [31:0] a, input [31:0] b, output [31:0] y);
|
||||
wire sa, sb; wire [8:0] ea, eb; wire [23:0] ma, mb;
|
||||
fp_unpack ua(a, sa, ea, ma); fp_unpack ub(b, sb, eb, mb);
|
||||
wire [47:0] prod = ma * mb;
|
||||
wire top = prod[47];
|
||||
wire [23:0] mant = top ? prod[47:24] : prod[46:23];
|
||||
wire guard = top ? prod[23] : prod[22]; wire sticky = top ? |prod[22:0] : |prod[21:0];
|
||||
wire round_up = guard & (sticky | mant[0]);
|
||||
wire [24:0] mr = {1'b0, mant} + round_up; wire carry = mr[24];
|
||||
wire [9:0] eres = {1'b0, ea} + {1'b0, eb} - 10'd127 + top + carry;
|
||||
assign y = {sa ^ sb, eres[7:0], carry ? mr[23:1] : mr[22:0]};
|
||||
endmodule
|
||||
// ---- int32 to float, RNE ----
|
||||
module fp_cvt(input [31:0] a, output [31:0] y);
|
||||
wire s = a[31]; wire [31:0] mag = s ? (~a + 1'b1) : a;
|
||||
wire [5:0] lz; lzc32 l(mag, lz);
|
||||
wire [31:0] norm = mag << lz; // leading one at bit 31
|
||||
wire [23:0] mant = norm[31:8]; wire guard = norm[7]; wire sticky = |norm[6:0];
|
||||
wire round_up = guard & (sticky | mant[0]);
|
||||
wire [24:0] mr = {1'b0, mant} + round_up; wire carry = mr[24];
|
||||
wire [7:0] e = 8'd127 + 8'd31 - lz + carry;
|
||||
assign y = (mag == 0) ? 32'h0 : {s, e, carry ? mr[23:1] : mr[22:0]};
|
||||
endmodule
|
||||
// ---- the lane: op 0 fadd, 1 fmul, 2 ffma, 3 fcvt; d ^= bits(result) ----
|
||||
module lane_fp32(
|
||||
input clk, input rst,
|
||||
input [1:0] op, input [2:0] dst, input [2:0] src, input [2:0] src2,
|
||||
input ld_en, input [31:0] ld_val,
|
||||
output [31:0] out);
|
||||
reg [31:0] rf [0:7];
|
||||
reg [1:0] op_q; reg [2:0] dst_q, src_q, src2_q; reg ld_q; reg [31:0] ld_val_q;
|
||||
integer i;
|
||||
always @(posedge clk) begin
|
||||
if (rst) begin op_q <= 0; dst_q <= 0; src_q <= 0; src2_q <= 0; ld_q <= 0; ld_val_q <= 0; end
|
||||
else begin op_q <= op; dst_q <= dst; src_q <= src; src2_q <= src2; ld_q <= ld_en; ld_val_q <= ld_val; end
|
||||
end
|
||||
wire [31:0] d = rf[dst_q]; wire [31:0] s = rf[src_q]; wire [31:0] s2 = rf[src2_q];
|
||||
wire [31:0] ya, ym, yf, yc;
|
||||
fp_add A(d, s, ya);
|
||||
fp_mul M(d, s, ym);
|
||||
fp_fma F(s, s2, d, yf);
|
||||
fp_cvt C(s, yc);
|
||||
reg [31:0] res;
|
||||
always @* case (op_q) 2'd0: res = d ^ ya; 2'd1: res = d ^ ym; 2'd2: res = d ^ yf; default: res = d ^ yc; endcase
|
||||
always @(posedge clk) begin
|
||||
if (rst) begin for (i = 0; i < 8; i = i + 1) rf[i] <= 32'h9e3779b9 * (i + 1) + 9; end
|
||||
else rf[dst_q] <= ld_q ? ld_val_q : res;
|
||||
end
|
||||
assign out = res;
|
||||
endmodule
|
||||
|
|
@ -21,10 +21,32 @@ module tb;
|
|||
@(negedge clk); cfg_en = 1; cfg_n = `NPROG - 1; cfg_sel = `SEL; cfg_m = 32'h9e3779b1; cfg_r = 5'd13; cfg_wm = 32'h0fffffc0; cfg_off = 3; cfg_mask = 32'h0fffffff;
|
||||
@(negedge clk); cfg_en = 0;
|
||||
// the program: `NPROG instructions drawn with the class v4 weights
|
||||
`ifdef CS
|
||||
// the connected-state draw: step = load + 27-instruction spine block; `NPROG = 16 x 28 = 448
|
||||
begin : cs
|
||||
integer st, q, last0, last1, last2, last3, mreg, areg, dreg, sreg, s2reg;
|
||||
areg = 0; mreg = 1;
|
||||
for (st = 0; st < `NPROG / 28; st = st + 1) begin
|
||||
@(negedge clk); w = {$random, $random}; mreg = $random & 63;
|
||||
prog_we = 1; prog_addr = st*28; prog_data = {w[`IW-1:4], 4'd12}; prog_data[4 +: 6] = mreg; prog_data[10 +: 6] = areg; // load: dst m_j, src a_j
|
||||
last0 = mreg; last1 = mreg; last2 = mreg; last3 = mreg;
|
||||
for (q = 0; q < 27; q = q + 1) begin
|
||||
@(negedge clk); w = {$random, $random}; opc = draw_op($random); dreg = $random & 63;
|
||||
case ($random & 3) 0: sreg = last0; 1: sreg = last1; 2: sreg = last2; default: sreg = last3; endcase
|
||||
s2reg = last0;
|
||||
if (q == 26 && !(opc == 4'd0 || opc == 4'd1 || opc == 4'd2 || opc == 4'd8 || opc == 4'd9)) opc = 4'd0; // the last instruction injects
|
||||
prog_we = 1; prog_addr = st*28 + 1 + q; prog_data = {w[`IW-1:4], opc}; prog_data[4 +: 6] = dreg; prog_data[10 +: 6] = sreg; prog_data[16 +: 6] = s2reg;
|
||||
last3 = last2; last2 = last1; last1 = last0; last0 = dreg;
|
||||
end
|
||||
areg = last0;
|
||||
end
|
||||
end
|
||||
`else
|
||||
for (k = 0; k < `NPROG; k = k + 1) begin
|
||||
@(negedge clk); w = {$random, $random}; opc = (loads && (k % 16 == 15)) ? 4'd12 : draw_op($random);
|
||||
prog_we = 1; prog_addr = k; prog_data = {w[`IW-1:4], opc};
|
||||
end
|
||||
`endif
|
||||
@(negedge clk); prog_we = 0; run = 1;
|
||||
for (n = 0; n < cycles; n = n + 1) begin
|
||||
@(negedge clk); ld_val = $random; acc = acc ^ out ^ addr;
|
||||
|
|
|
|||
6
tools/chip-model/rtl/tb/tb_core_tm_8r64.v
Normal file
6
tools/chip-model/rtl/tb/tb_core_tm_8r64.v
Normal file
|
|
@ -0,0 +1,6 @@
|
|||
`define TOP core_tm_8r64
|
||||
`define HALF 750
|
||||
`define IW 40
|
||||
`define NPROG 256
|
||||
`define SEL 64'hfedcba9876543210
|
||||
`include "tb_core_common.vh"
|
||||
3
tools/chip-model/rtl/tb/tb_core_v6_8_legacy.v
Normal file
3
tools/chip-model/rtl/tb/tb_core_v6_8_legacy.v
Normal file
|
|
@ -0,0 +1,3 @@
|
|||
`define TOP core_v6_8
|
||||
`define HALF 750
|
||||
`include "tb_core_legacy.vh"
|
||||
7
tools/chip-model/rtl/tb/tb_core_v6_8cs64.v
Normal file
7
tools/chip-model/rtl/tb/tb_core_v6_8cs64.v
Normal file
|
|
@ -0,0 +1,7 @@
|
|||
`define TOP core_v6_8cs64
|
||||
`define HALF 750
|
||||
`define IW 40
|
||||
`define NPROG 448
|
||||
`define CS 1
|
||||
`define SEL 64'hfedcba9876543210
|
||||
`include "tb_core_common.vh"
|
||||
22
tools/chip-model/rtl/tb/tb_lane_fp32.v
Normal file
22
tools/chip-model/rtl/tb/tb_lane_fp32.v
Normal file
|
|
@ -0,0 +1,22 @@
|
|||
`timescale 1ps/1ps
|
||||
module tb;
|
||||
reg clk = 0, rst = 1; reg [1:0] op = 0; reg [2:0] dst = 0, src = 0, src2 = 0; reg ld_en = 0; reg [31:0] ld_val = 0;
|
||||
wire [31:0] out;
|
||||
lane_fp32 dut(.clk(clk), .rst(rst), .op(op), .dst(dst), .src(src), .src2(src2), .ld_en(ld_en), .ld_val(ld_val), .out(out));
|
||||
integer n, fixed_op, cycles; reg [31:0] acc = 0;
|
||||
always #1000 clk = ~clk;
|
||||
// a reference check of the units against the host's float arithmetic is the mixed lane's own (the ranges are its);
|
||||
// this bench drives random registers and reports the checksum
|
||||
initial begin
|
||||
if (!$value$plusargs("op=%d", fixed_op)) fixed_op = -1;
|
||||
if (!$value$plusargs("cycles=%d", cycles)) cycles = 3000;
|
||||
$dumpfile("dump.vcd"); $dumpvars(0, tb.dut);
|
||||
repeat (4) @(negedge clk); rst = 0;
|
||||
for (n = 0; n < cycles; n = n + 1) begin
|
||||
@(negedge clk);
|
||||
op = (fixed_op < 0) ? $random : fixed_op; dst = $random; src = $random; src2 = $random;
|
||||
ld_en = (($random & 7) == 0); ld_val = $random; acc = acc ^ out;
|
||||
end
|
||||
$display("CHECKSUM %08x", acc); $finish;
|
||||
end
|
||||
endmodule
|
||||
26
tools/ci/batches/floor-k-20261008-rows-repeat.json
Normal file
26
tools/ci/batches/floor-k-20261008-rows-repeat.json
Normal file
|
|
@ -0,0 +1,26 @@
|
|||
{
|
||||
"run_id": "floor-k-20261008-rows-repeat",
|
||||
"manifest_sha": "86e5b0fb",
|
||||
"cut_tip": "class-v6-floor-k 86e5b0fb (the amendment carries shadow-k.md; floor lane 2's rows ADV-05 and ADV-06 move with it on its word, method model, no decision changed; the steward's rows no longer cite the directory since d2e1c192)",
|
||||
"evidence_dir": "docs/analysis/class-v6/floor/shadow-k.md",
|
||||
"boxes": [],
|
||||
"cells": [
|
||||
{
|
||||
"cell": "adversary:mf-placed",
|
||||
"cases": [
|
||||
"ADV-05"
|
||||
],
|
||||
"status": "RUNNING",
|
||||
"method": "model",
|
||||
"evidence": "docs/analysis/class-v6/floor/shadow-k.md",
|
||||
"note": "repeat for ADV-05 at this manifest on floor lane 2's word (placed and routed RTL on ASAP7, a model); the rtl directory tools/chip-model/rtl/ is the flow, cited by its document since the recorder takes files only"
|
||||
},
|
||||
{
|
||||
"cell": "review:k-lane-shadow-k",
|
||||
"status": "NOT RUN",
|
||||
"method": "model",
|
||||
"evidence": "docs/analysis/class-v6/floor/shadow-k.md",
|
||||
"note": "repeat of team-2026-10-08 for ADV-06 at this manifest on floor lane 2's word; no decision changed"
|
||||
}
|
||||
]
|
||||
}
|
||||
|
|
@ -687,6 +687,17 @@
|
|||
"coverage": {
|
||||
"INT-07": "partial: V6-12's clean-install half (one published object installed fresh, synced, mined, proved and paid on one host per artefact); the cross-host agreement half (node, pool, CPU verifier, each GPU host across activation) is harness:same-work's"
|
||||
}
|
||||
},
|
||||
"review:k-lane-shadow-k": {
|
||||
"command": "floor lane 2's placed and routed cores on ASAP7 in tools/chip-model/rtl (the rows in docs/analysis/class-v6/floor/shadow-k.md); the register landing 50ff1611f recorded ADV-06 against it before the recorder existed",
|
||||
"box_class": "none",
|
||||
"fixtures": [],
|
||||
"cases": [
|
||||
"ADV-06"
|
||||
],
|
||||
"coverage": {
|
||||
"ADV-06": "partial: the placed gated cores and the adversary's forms are modelled in shadow-k.md; the independent review remains"
|
||||
}
|
||||
}
|
||||
},
|
||||
"not_run": {
|
||||
|
|
@ -697,7 +708,6 @@
|
|||
"GPU-06": "accepted work under ordinary connectivity needs the fault network F4",
|
||||
"GPU-07": "the sustained thermal and power soak has no harness tonight: the project's own rig mines nothing under the the earlier devnet off order",
|
||||
"POW-05": "amortised cheap winning attempts are the attack lanes' grind and era harnesses (tools/attack/f7-era, f9-grind), not in the release matrix; their rows come from those lanes",
|
||||
"ADV-06": "process-advantage separation is the adversary lanes' chip study",
|
||||
"ROT-03": "miner-voted bring-forward needs a vote harness on the fault network F4",
|
||||
"ROT-04": "seed-selection resistance is the census harness (the class v6 invention lane), not yet in the matrix",
|
||||
"EVM-01": "no EVM conformance-vector harness is mapped tonight; the exec suite does not run the reference test vectors",
|
||||
|
|
|
|||
Loading…
Reference in a new issue