Connected state: the chip rows (the k lane's re-optimised core: 4.4 pJ per lane-op at N5, k 0.71 at the lock), the score (3.3x node-for-node, 3.6x a node ahead; the window moves the chip's edge 1.10x against a 1.25x gate), verdict KILL as a class, what is kept and what v7 tests

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

Documents-only replay of 6ccd2a6fe (6ccd2a6fe) for the box mirror master
This commit is contained in:
igneum-labs 2026-10-08 15:23:51 +00:00
parent 0647e8976c
commit e6146368b5

View file

@ -87,11 +87,32 @@ Meaning: at stock the window costs the 5090 0.6 percent of energy per hash again
## 5. The chip side
ROW_CHIP
The k lane (floor lane 2) priced the re-optimised core on the drawn program at 17:2x BST (synthesis only, a model and never a lower bound; its placed row is due 21:00 as an amendment). The core: 8 lanes, a 64 x 32-bit window per lane in clock-gated flops (a macro file reads within 5 percent), a 512-entry imem holding the 448-instruction text, one in-order op per cycle per lane (the spine's ILP is met by lane count, which is free), every class unit, the load's fold on the address path; gate-level random-input VCD, every pin annotated; 253,059 cells.
| Row (the k lane's) | pJ per lane-op, ASAP7 | N5 (the card's node) | N3 | N2 | k at the 1,300 lock, N5 / N3 / N2 | k at stock, N5 / N3 |
|---|---|---|---|---|---|---|
| cs64s27x16, the re-optimised core (gated window, 512 imem) | 6.3 | 4.4 | 3.2 | 2.3 | 0.71 / 0.51 / 0.37 | 0.39 / 0.28 |
| the same window on the class v4 draw, 256 imem | 6.2 | 4.3 | 3.1 | 2.2 | 0.70 / 0.50 / 0.36 | 0.38 / 0.27 |
| the adversary's 32-register base, gated (the genesis window) | 4.5 | 3.2 | 2.3 | 1.6 | 0.51 / 0.37 / 0.26 | 0.28 / 0.20 |
| the GPU-shaped 64-register core, ungated (shadow-k.md, the earlier default) | 9.7 | 6.8 | 4.9 | 3.5 | 1.09 / 0.78 / 0.56 | 0.60 / 0.43 |
What the adversary's re-optimisation did to each part of the structure: the gated file charges only the register written, so the whole-window liveness costs it nothing beyond the write it would make anyway and the hot-20 banking of section 2 is not even needed; the 16-pass loop and the 448 text cost the shared imem 0.1 pJ per lane-op; the narrow per-step chain sets the lane count, which is free. The window itself is worth +1.2 pJ per lane-op at N5 over the genesis window (+0.14 of k at the lock), the same knob as the design document's 64-register row; the connected organisation around it adds about 0.1 pJ. The placed ungated core came in 64 percent over its synthesis, so the placed figure is expected near 8 to 10 pJ at ASAP7 (k node-for-node near 0.9 to 1.1, approximate); both rows move together and the ratio below holds.
## 6. The score and the verdict
ROW_SCORE
E_GPU over E_adversary, absolute convention, GDDR7 board (E_mem 0.466 microjoules per hash), E_chip = E_mem + 55,296 x e_chip, E_GPU = the 5090 at the lock (2.33 microjoules per hash on class v5) x 1.006 for the window (the stock rows of section 4; the lock row is owed):
| Core | Node-for-node (N5) | A node ahead (N3) | Two nodes (N2) |
|---|---|---|---|
| cs64s27x16, re-optimised | 2.344 / (0.466 + 0.243) = 3.3x | 2.344 / (0.466 + 0.177) = 3.6x | 2.344 / (0.466 + 0.127) = 4.0x |
| the genesis window (the control) | 2.33 / (0.466 + 0.177) = 3.6x | 2.33 / (0.466 + 0.127) = 3.9x | 2.33 / (0.466 + 0.088) = 4.2x |
| the window's effect on the chip's edge | 1.10x | 1.08x | 1.05x |
On the placed figures (both rows 64 percent higher) the pair reads about 2.2x and 2.4x node-for-node and the ratio stays near 1.1x. The gate was 1.25x node-for-node for the 1.5x one-node-ahead ambition; the row reads 1.10x node-for-node and 1.08x a node ahead.
**Verdict: KILL as a class.** The hypothesis was that a connected organisation of the same work, with the window independently necessary across the whole chain, would deny a specialist its separation of storage, arithmetic and scheduling. It does not: the liveness rows show the window is necessary (63 of 64 at every address) and the chip answers with a clock-gated file that pays per write, not per live register, so necessity costs it nothing; the only term that reaches the chip is the window's own width (+0.14 k at the lock), which the design document already holds as its one robust core knob, and the connected structure adds about 0.1 pJ around it. The GPU side passes its budget with room (+0.6 percent of energy per hash at stock on the 5090, -0.9 percent on the 4090, 80 to 87 registers per thread with no spill), and the census passes every instrument with fewer attempts than v5; neither moves the score. The founder's accepted review stands in a sharper form than before: a specialist's edge against this family is a per-op energy ratio on a known op mix, and reorganising the dependency graph of the same ops does not change what an op costs on either side.
What is kept: the generator variant and the liveness tool (research class, behind the flag) for the v7 tests below; the measured fact that a 64-register window costs a card under 1 percent at stock, which fixes the design document's modelled "about 0 rate" row; the kit worker and nvcc harness agreement on two cards. What is withdrawn: the "connected state" line as a resistance mechanism.
## 7. What a v7 variant would test next
@ -99,4 +120,5 @@ ROW_SCORE
- A wider per-step chain: a spine of depth 8 to 16 so `dep` rises from about 11 toward the window, at the cost of ILP on the card (measure the rate first; the chain per step is what a two-level file exploits).
- The window at 32 and 16 (`cs32s27x16`, `cs16s27x16`) for the k curve, and the block at 16 x 27 (`cs64s16x27`, text 272) if the imem matters to the re-optimised core.
- The op-mix re-weight of the census lane at this structure (the two closing instructions already take the injecting table).
- The kit worker row on both cards and the PC 1 lock row.
- The PC 1 lock row (informational now: the verdict does not turn on it; it lands as an amendment if PC 1 runs it).
- A knob that reaches a gated file: not more live state but more WRITES per op the chip cannot skip (every op writing two registers, or a window write per load), priced against the card's own write cost first; the k lane's placed row at 21:00 says whether even that moves k.