Connected state: the PC 1 lock row (the window +1.6 percent of energy per hash at the 1,300 lock, +5.9 unlocked; the score re-read at x 1.016: 1.08x node-for-node, 1.07x a node ahead; the verdict unchanged)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

Documents-only replay of 042f5ba98 (042f5ba98a) for the box mirror master
This commit is contained in:
igneum-labs 2026-10-08 16:56:07 +00:00
parent ecb70db5bb
commit 1c33395a75

View file

@ -70,7 +70,10 @@ Instrument: the class v5 nvcc harness (`proto-newpow/class-v5/bench.cu` on `box/
| RTX 5090, the same card | v5-genesis | 65.30 | 574.8 | 8.803 | 48 | 0 | 24 | ae74193ddad19e19 |
| RTX 4090 (RunPod 386yytbh4bkfnz, driver 580.159.04, 450 W cap), stock | cs64s27x16 | 62.44 | 268.8 | 4.305 | 87 | 0 | 20 | ad0cec2a42c84aff |
| RTX 4090, the same card | v5-genesis | 62.44 | 271.3 | 4.345 | 32 | 0 | 24 | ae74193ddad19e19 |
| RTX 5090 on PC 1 at the 1,300 MHz lock | both | OWED: PC 1 booked to 17:40 BST; the bound pack and the kit are at build-1:/srv/builds/igneum-wt-connected/cs-kit (sha c9aff54aaf10d79e) and the hash lane publishes the job when PC 1 frees | | | | | | |
| RTX 5090 on PC 1 at the 1,300 MHz core lock (the hash lane's job run-ca3-pc1-cs64-5090-20261008, 17:50 to 17:53 BST, the class v5 kit's CUDA worker, 250 x 2^24, nvidia-smi 1 Hz; PCIe gen 4 x16 since the eGPU swap) | cs64s27x16 (bound pack) | 132.4 | 316.4 | 2.390 | 80 | 0 | | ad0cec2a42c84aff |
| PC 1, the same lock | v5-genesis | 132.3 | 311.4 | 2.354 | 48 | 0 | | ae74193ddad19e19 |
| PC 1 unlocked (the same job) | cs64s27x16 | 139.9 | 492.6 | 3.521 | 80 | 0 | | ad0cec2a42c84aff |
| PC 1 unlocked | v5-genesis | 139.8 | 465.2 | 3.328 | 48 | 0 | | ae74193ddad19e19 |
The kit worker (the brief's instrument, `igneum-worker-cuda --bench --batch-log2 24 --batches 250`, the class v5 kit's NVRTC worker of 7 October loading the pack's `kernel_bound.cu`; check PASS on both packs):
@ -83,7 +86,7 @@ The kit worker (the brief's instrument, `igneum-worker-cuda --bench --batch-log2
Both instruments agree with each other on each card (the harness and the worker within 3 percent of rate) and agree on the comparison: the window moves energy per hash by +0.6 percent on the 5090 (harness and worker alike) and by -0.9 percent on the 4090 (harness and worker alike), inside the run-to-run noise of a power reading. The 4090 is not at its cap (269 W of 450) and both packs read the same rate to three figures, so there the hash is bound by the memory chain, not the ALU or the register file; the 5090 is at its 575 W cap and the window costs under 1 percent of rate. (The fingerprints are the same on both cards and both instruments: ad0cec2a42c84aff for cs64, ae74193ddad19e19 for v5-genesis.)
Meaning: at stock the window costs the 5090 0.6 percent of energy per hash against a 10 percent budget; the compiled allocation is 80 registers per thread with no spill, so the window is in registers, not local memory, and the occupancy under this harness is the same as v5's. The harness reads 65 MH/s where the NVRTC worker reads about twice that on a 5090 (one warp per block, 24 resident blocks); the ratio between two packs on the same harness is the measurement, the absolute rate is not. The lock row is where the energy comparison binds (the 5090 at the lock reads 2.33 microjoules per hash on v5); the stock rows say the card is bound by its power cap in both cases and the window moves the rate by under 1 percent.
Meaning: at stock the window costs the 5090 0.6 percent of energy per hash against a 10 percent budget; the compiled allocation is 80 registers per thread with no spill, so the window is in registers, not local memory, and the occupancy under this harness is the same as v5's. The harness reads 65 MH/s where the NVRTC worker reads about twice that on a 5090 (one warp per block, 24 resident blocks); the ratio between two packs on the same harness is the measurement, the absolute rate is not. At the lock, where the energy comparison binds, the window costs the 5090 +1.6 percent of energy per hash (2.39 against 2.35 microjoules, the rate equal); unlocked on PC 1 it costs +5.9 percent (3.52 against 3.33, the rate equal, 27 W more at the same clock): the window's register traffic is a power term the lock hides and the rented cards' caps hid. Both inside the 10 percent budget; the card side never decided this lane.
## 5. The chip side
@ -100,15 +103,15 @@ What the adversary's re-optimisation did to each part of the structure: the gate
## 6. The score and the verdict
E_GPU over E_adversary, absolute convention, GDDR7 board (E_mem 0.466 microjoules per hash), E_chip = E_mem + 55,296 x e_chip, E_GPU = the 5090 at the lock (2.33 microjoules per hash on class v5) x 1.006 for the window (the stock rows of section 4; the lock row is owed):
E_GPU over E_adversary, absolute convention, GDDR7 board (E_mem 0.466 microjoules per hash), E_chip = E_mem + 55,296 x e_chip, E_GPU = the 5090 at the lock (2.33 microjoules per hash on class v5) x 1.016 for the window (the PC 1 lock row of section 4; the table was first written at x 1.006 from the stock rows and the amendment moved the ratios from 1.10x and 1.08x to 1.08x and 1.07x):
| Core | Node-for-node (N5) | A node ahead (N3) | Two nodes (N2) |
|---|---|---|---|
| cs64s27x16, re-optimised | 2.344 / (0.466 + 0.243) = 3.3x | 2.344 / (0.466 + 0.177) = 3.6x | 2.344 / (0.466 + 0.127) = 4.0x |
| cs64s27x16, re-optimised | 2.367 / (0.466 + 0.243) = 3.3x | 2.367 / (0.466 + 0.177) = 3.7x | 2.367 / (0.466 + 0.127) = 4.0x |
| the genesis window (the control) | 2.33 / (0.466 + 0.177) = 3.6x | 2.33 / (0.466 + 0.127) = 3.9x | 2.33 / (0.466 + 0.088) = 4.2x |
| the window's effect on the chip's edge | 1.10x | 1.08x | 1.05x |
| the window's effect on the chip's edge | 1.08x | 1.07x | 1.05x |
On the placed figures (both rows 64 percent higher) the pair reads about 2.2x and 2.4x node-for-node and the ratio stays near 1.1x. The gate was 1.25x node-for-node for the 1.5x one-node-ahead ambition; the row reads 1.10x node-for-node and 1.08x a node ahead.
On the placed figures (both rows 64 percent higher) the pair reads about 2.2x and 2.4x node-for-node and the ratio stays near 1.1x. The gate was 1.25x node-for-node for the 1.5x one-node-ahead ambition; the row reads 1.08x node-for-node and 1.07x a node ahead.
**Verdict: KILL as a class.** The hypothesis was that a connected organisation of the same work, with the window independently necessary across the whole chain, would deny a specialist its separation of storage, arithmetic and scheduling. It does not: the liveness rows show the window is necessary (63 of 64 at every address) and the chip answers with a clock-gated file that pays per write, not per live register, so necessity costs it nothing; the only term that reaches the chip is the window's own width (+0.14 k at the lock), which the design document already holds as its one robust core knob, and the connected structure adds about 0.1 pJ around it. The GPU side passes its budget with room (+0.6 percent of energy per hash at stock on the 5090, -0.9 percent on the 4090, 80 to 87 registers per thread with no spill), and the census passes every instrument with fewer attempts than v5; neither moves the score. The founder's accepted review stands in a sharper form than before: a specialist's edge against this family is a per-op energy ratio on a known op mix, and reorganising the dependency graph of the same ops does not change what an op costs on either side.
@ -120,5 +123,4 @@ What is kept: the generator variant and the liveness tool (research class, behin
- A wider per-step chain: a spine of depth 8 to 16 so `dep` rises from about 11 toward the window, at the cost of ILP on the card (measure the rate first; the chain per step is what a two-level file exploits).
- The window at 32 and 16 (`cs32s27x16`, `cs16s27x16`) for the k curve, and the block at 16 x 27 (`cs64s16x27`, text 272) if the imem matters to the re-optimised core.
- The op-mix re-weight of the census lane at this structure (the two closing instructions already take the injecting table).
- The PC 1 lock row (informational now: the verdict does not turn on it; it lands as an amendment if PC 1 runs it).
- A knob that reaches a gated file: not more live state but more WRITES per op the chip cannot skip (every op writing two registers, or a window write per load), priced against the card's own write cost first; the k lane's placed row at 21:00 says whether even that moves k.