Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Documents-only replay of 6ccd2a6fe (6ccd2a6fe) for the box mirror master
17 KiB
Class v6: the connected-state variant (8 October 2026)
The experiment of the external review the founder accepted (the 15:4x BST rules): the remaining chip edge (about 2.0x to 2.2x node-for-node, 2.2x to 2.6x a node ahead) survives because a specialist can separate storage, arithmetic and memory scheduling. This lane reorganised the class v5 work, same dataset, same read width (4 bytes), about the same operation count, so that live state feeds each load address, the memory result feeds mixed arithmetic and cross-lane exchange, and that updates the live state for the next address, with a 64-register window per lane that stays necessary across the whole dependent chain. Research class only, behind --class cs<W>s<S>x<R> in igneum-pow (igneum-pow/src/connected.rs, branch class-v6-connected on the box mirror); never a chain class.
Every number below is measured unless marked modelled or owed. Times are UK (BST).
1. The structure
| Item | Class v5 (mx8+sh256x27, the control) |
Connected state (cs64s27x16) |
|---|---|---|
| Registers per lane | 8 | 64 (the window) |
| Per iteration | 64 base instructions with 16 loads, then a 256-instruction block run 27 times | 16 steps: a load, then a 27-instruction block run 16 times |
| Load address | a base register, x & MASK (the era stride and site window under an era) |
a window register, the same address rule |
| Memory result | xor into the load's destination | xor into r[m_j]; block instruction 0 reads it |
| Next address | whatever register the next load reads | written by the block's last instruction (add, sub, xor or shfl) from a fresh spine |
| Result | fold of 8 registers | fold of all 64 (lo over the first 32 at 7 i, hi over the second 32 at 9 i; the v5 fold at a window of 8) |
| Loads per hash | 128 | 128 |
| ALU instructions per hash | 55,680 (55,296 shadow + 384 base) | 55,296 (0.7 percent fewer) |
| Instruction text per iteration | 320 | 448 |
| Negative controls kept out | no long program (1,024), no select tree, no wide read (W = 16), no scratchpad |
The draw (deterministic from the seed words, attempt k re-seeded as every class): per step a_j then m_j != a_j (and the era window draws); per block instruction the op from the ten non-load families at the v5 weights, the source from the last four spine entries (the memory result first), the destination uniform over the window, the two immediates, the rotation, the selector bit and the shuffle mask. Four rules the first census pass forced, each a construction rather than a filter:
- The cover: the first 64 injecting destinations walk a drawn permutation of the window, so every register takes an injecting write (rule (b)); a uniform draw left one register without one in about 70 percent of candidates.
- The fresh spine: only destinations of add, sub, xor, mad and shfl enter the spine, so every address is fresh by dataflow (rule (a')); with a lossy spine every seed exhausted 256 attempts.
- Lossy ops feed the next injection: or, mul and mulhi write the register the next injecting instruction of the block writes, so their value enters the chain and no register accumulates a lossy op across the 16 passes (a 16-pass or saturated a register the block never re-injected, a 16-pass mul cleared its low bits: rule (c) refused every candidate).
- The two closing instructions of a block come from add, sub, xor and shfl (the census lane's rule for a load source's writer): no product on the address writer (a 16-pass mad on the address register read bit z 57.6 at one site).
The acceptance rule is the sub-version 3 rule over the window: (b), (a') (the fixpoint over the iteration's execution order), (c) over 64 units (constant bits per register, lane-constant sites, saturation at 1 percent of the window's final values, saturated sources, output bias, distinct addresses, the value-level index-bit read judged inside the site's era window), then (c'') at 0.98 and (c''') at 0.995 over 2^20 evaluations per site. (a) holds by construction.
2. Liveness (seed igneum-v6c/0, attempt 0, program id 9ad55de91485542b)
The tool (igneum-pow liveness --class cs64s27x16 --seed <s>) walks the unrolled trace of one hash (128 loads, 55,296 ALU instructions). live is the backward set at an address: registers whose value just before the load feeds this or any later address. dep is the forward set: registers at the previous address whose values feed this address.
| Measure | Value |
|---|---|
| live before each address | 63 of 64 at every address until the last iteration's tail; iteration 7: 63 62 60 60 60 57 57 54 53 51 45 36 25 20 9 1 |
| registers read by the result | 64 of 64 (the fold) |
| dep per step (the same every iteration, the text repeats) | 1 9 13 14 8 15 7 8 12 14 14 13 10 9 14 9; mean 11.2 of 64 |
| registers touched per block (read or written) | min 16, mean 20.4, max 24 of 64 |
| reads per register per hash | min 384, mean 1,724, max 3,208 |
| writes per register per hash | min 128, mean 866, max 1,664 |
| op mix of the 432 block instructions | add 82, mul 62, xor 51, shfl 51, sub 46, mulhi 39, rotl 35, mad 31, rotr 20, or 15 |
Meaning: the whole window is necessary over the hash (no register can be dropped or parked for long: the longest gap between writes to one register is under one iteration), and every value is read by a later address, so there is no side calculation a specialist can move to a separate engine. The dependent chain per step is narrow: about 11 of the 64 registers at one address feed the next address through 432 operations, and a block touches about 20. That is the shape a specialist will exploit: a two-level file, about 20 registers hot per step and 44 warm, the hot set moving along the text; the k lane prices exactly that (section 5). The program listing is liveness.txt and program.json of the pack.
3. Census (the census lane's sub-version 3 harness, tools/attack/v6-census/v6census.sh, (c''') 0.995 on, on a rented 5090 host with 40 threads, 16:01 to 16:08 BST)
| Run | Seeds | Accepted | Exhausted | Rejections per candidate | Mean attempt | Max attempt | Window-bit refusals | Over 6 sigma (reported) |
|---|---|---|---|---|---|---|---|---|
| cs64 no era | 256 | 256 | 0 | 0.283 | 0.39 | 5 | off | 146 of 256 (max z 128) |
control mx8+sh256x27 no era (census lane, build-3) |
256 | 256 | 0 | 0.668 | 2.01 | 18 | off | 143 of 256 (max z 97) |
| cs64 eras 0 to 7, window-bit refusal on | 8 x 32 | 256 | 0 | 0.635 | 1.74 | 8 | 306 | 0 (max z 5.9) |
| control eras 0 to 7, refusal on (census lane, build-4) | 8 x 32 | 256 | 0 | 0.830 | 4.87 | 22 | 218 | 0 (max z 5.9) |
| cs64 eras 0 to 7, refusal off | 8 x 32 | 256 | 0 | 0.304 | 0.44 | 2 | off | 94 of 256 (max z 71.5) |
The no-era rejections are all rule (c) (constant bits, saturation or bias in the 64-unit test); the control's are mostly (a') and (a), which this class satisfies by construction. The window-bit refusals sit on eras 0, 1, 4, 5 and 6 (55 to 64 each) and nearly vanish on eras 2 and 3 (0 and 1): the product-bit class (a product's low bits reaching an address bit through the era stride), the same class the control carries and the layer-1 index fold removes; this class does not change it either way.
F8 form (igneum-pow cs-uniform, 16 seeds x 2^20 nonces, no era): the top 0.1 percent item share reads 0.9987 to 1.0016 of a uniform control of the same size (the control class 0.9993 to 1.0016); the worst per-site distinct ratio 0.9941 (seed 9, site 9; the sequential-nonce sample of the F8 tool, the acceptance's own 2^20 sample passed 0.995 on every accepted program).
Meaning: the class censuses at least as well as v5 on every instrument of the harness and takes fewer attempts; it inherits v5's product-bit bias and the fix is the same fold. The acceptance costs about 3.7 s per candidate on one core (the (c'') pass dominates, 7 G lane-ops), against v5's 2.8 s.
4. GPU cost
Instrument: the class v5 nvcc harness (proto-newpow/class-v5/bench.cu on box/ds55-v5, the v5 design page's 4090 rows), compiled per pack with -Xptxas -v, 250 batches of 2^24, nvidia-smi at 1 Hz, both packs on the same card minutes apart; vectors 3 of 3 PASS and the dataset self-test PASS on every row. The cs64 pack is over the class v4 memory-hard dataset (no leaves); v5-genesis carries its 93 leaves. Stock means the card's own power limit and no clock lock.
| Card | Pack | MH/s | Mean W | Microjoules per hash | Registers per thread | Spills | Resident blocks per SM (1 warp per block) | Fingerprint of 2^24 at base 0 |
|---|---|---|---|---|---|---|---|---|
| RTX 5090 (Vast 54862507, driver 580.159.03, 575 W cap), stock | cs64s27x16 | 64.93 | 574.8 | 8.853 | 80 | 0 | 24 | ad0cec2a42c84aff |
| RTX 5090, the same card | v5-genesis | 65.30 | 574.8 | 8.803 | 48 | 0 | 24 | ae74193ddad19e19 |
| RTX 4090 (RunPod 386yytbh4bkfnz, driver 580.159.04, 450 W cap), stock | cs64s27x16 | 62.44 | 268.8 | 4.305 | 87 | 0 | 20 | ad0cec2a42c84aff |
| RTX 4090, the same card | v5-genesis | 62.44 | 271.3 | 4.345 | 32 | 0 | 24 | ae74193ddad19e19 |
| RTX 5090 on PC 1 at the 1,300 MHz lock | both | OWED: PC 1 booked to 17:40 BST; the bound pack and the kit are at build-1:/srv/builds/igneum-wt-connected/cs-kit (sha c9aff54aaf10d79e) and the hash lane publishes the job when PC 1 frees |
The kit worker (the brief's instrument, igneum-worker-cuda --bench --batch-log2 24 --batches 250, the class v5 kit's NVRTC worker of 7 October loading the pack's kernel_bound.cu; check PASS on both packs):
| Card | Pack | MH/s | Mean W | Microjoules per hash | Registers per thread | Fingerprint |
|---|---|---|---|---|---|---|
| RTX 5090 (the Vast card above), stock | cs64s27x16 (bound pack) | 62.88 | 574.2 | 9.131 | 80 | ad0cec2a42c84aff |
| RTX 5090, the same card | v5-genesis | 62.96 | 571.4 | 9.076 | 48 | ae74193ddad19e19 |
| RTX 4090 (the RunPod card above), stock | cs64s27x16 (bound pack) | 62.41 | 269.4 | 4.317 | 87 | ad0cec2a42c84aff |
| RTX 4090, the same card | v5-genesis | 62.39 | 271.9 | 4.358 | 32 | ae74193ddad19e19 |
Both instruments agree with each other on each card (the harness and the worker within 3 percent of rate) and agree on the comparison: the window moves energy per hash by +0.6 percent on the 5090 (harness and worker alike) and by -0.9 percent on the 4090 (harness and worker alike), inside the run-to-run noise of a power reading. The 4090 is not at its cap (269 W of 450) and both packs read the same rate to three figures, so there the hash is bound by the memory chain, not the ALU or the register file; the 5090 is at its 575 W cap and the window costs under 1 percent of rate. (The fingerprints are the same on both cards and both instruments: ad0cec2a42c84aff for cs64, ae74193ddad19e19 for v5-genesis.)
Meaning: at stock the window costs the 5090 0.6 percent of energy per hash against a 10 percent budget; the compiled allocation is 80 registers per thread with no spill, so the window is in registers, not local memory, and the occupancy under this harness is the same as v5's. The harness reads 65 MH/s where the NVRTC worker reads about twice that on a 5090 (one warp per block, 24 resident blocks); the ratio between two packs on the same harness is the measurement, the absolute rate is not. The lock row is where the energy comparison binds (the 5090 at the lock reads 2.33 microjoules per hash on v5); the stock rows say the card is bound by its power cap in both cases and the window moves the rate by under 1 percent.
5. The chip side
The k lane (floor lane 2) priced the re-optimised core on the drawn program at 17:2x BST (synthesis only, a model and never a lower bound; its placed row is due 21:00 as an amendment). The core: 8 lanes, a 64 x 32-bit window per lane in clock-gated flops (a macro file reads within 5 percent), a 512-entry imem holding the 448-instruction text, one in-order op per cycle per lane (the spine's ILP is met by lane count, which is free), every class unit, the load's fold on the address path; gate-level random-input VCD, every pin annotated; 253,059 cells.
| Row (the k lane's) | pJ per lane-op, ASAP7 | N5 (the card's node) | N3 | N2 | k at the 1,300 lock, N5 / N3 / N2 | k at stock, N5 / N3 |
|---|---|---|---|---|---|---|
| cs64s27x16, the re-optimised core (gated window, 512 imem) | 6.3 | 4.4 | 3.2 | 2.3 | 0.71 / 0.51 / 0.37 | 0.39 / 0.28 |
| the same window on the class v4 draw, 256 imem | 6.2 | 4.3 | 3.1 | 2.2 | 0.70 / 0.50 / 0.36 | 0.38 / 0.27 |
| the adversary's 32-register base, gated (the genesis window) | 4.5 | 3.2 | 2.3 | 1.6 | 0.51 / 0.37 / 0.26 | 0.28 / 0.20 |
| the GPU-shaped 64-register core, ungated (shadow-k.md, the earlier default) | 9.7 | 6.8 | 4.9 | 3.5 | 1.09 / 0.78 / 0.56 | 0.60 / 0.43 |
What the adversary's re-optimisation did to each part of the structure: the gated file charges only the register written, so the whole-window liveness costs it nothing beyond the write it would make anyway and the hot-20 banking of section 2 is not even needed; the 16-pass loop and the 448 text cost the shared imem 0.1 pJ per lane-op; the narrow per-step chain sets the lane count, which is free. The window itself is worth +1.2 pJ per lane-op at N5 over the genesis window (+0.14 of k at the lock), the same knob as the design document's 64-register row; the connected organisation around it adds about 0.1 pJ. The placed ungated core came in 64 percent over its synthesis, so the placed figure is expected near 8 to 10 pJ at ASAP7 (k node-for-node near 0.9 to 1.1, approximate); both rows move together and the ratio below holds.
6. The score and the verdict
E_GPU over E_adversary, absolute convention, GDDR7 board (E_mem 0.466 microjoules per hash), E_chip = E_mem + 55,296 x e_chip, E_GPU = the 5090 at the lock (2.33 microjoules per hash on class v5) x 1.006 for the window (the stock rows of section 4; the lock row is owed):
| Core | Node-for-node (N5) | A node ahead (N3) | Two nodes (N2) |
|---|---|---|---|
| cs64s27x16, re-optimised | 2.344 / (0.466 + 0.243) = 3.3x | 2.344 / (0.466 + 0.177) = 3.6x | 2.344 / (0.466 + 0.127) = 4.0x |
| the genesis window (the control) | 2.33 / (0.466 + 0.177) = 3.6x | 2.33 / (0.466 + 0.127) = 3.9x | 2.33 / (0.466 + 0.088) = 4.2x |
| the window's effect on the chip's edge | 1.10x | 1.08x | 1.05x |
On the placed figures (both rows 64 percent higher) the pair reads about 2.2x and 2.4x node-for-node and the ratio stays near 1.1x. The gate was 1.25x node-for-node for the 1.5x one-node-ahead ambition; the row reads 1.10x node-for-node and 1.08x a node ahead.
Verdict: KILL as a class. The hypothesis was that a connected organisation of the same work, with the window independently necessary across the whole chain, would deny a specialist its separation of storage, arithmetic and scheduling. It does not: the liveness rows show the window is necessary (63 of 64 at every address) and the chip answers with a clock-gated file that pays per write, not per live register, so necessity costs it nothing; the only term that reaches the chip is the window's own width (+0.14 k at the lock), which the design document already holds as its one robust core knob, and the connected structure adds about 0.1 pJ around it. The GPU side passes its budget with room (+0.6 percent of energy per hash at stock on the 5090, -0.9 percent on the 4090, 80 to 87 registers per thread with no spill), and the census passes every instrument with fewer attempts than v5; neither moves the score. The founder's accepted review stands in a sharper form than before: a specialist's edge against this family is a per-op energy ratio on a known op mix, and reorganising the dependency graph of the same ops does not change what an op costs on either side.
What is kept: the generator variant and the liveness tool (research class, behind the flag) for the v7 tests below; the measured fact that a 64-register window costs a card under 1 percent at stock, which fixes the design document's modelled "about 0 rate" row; the kit worker and nvcc harness agreement on two cards. What is withdrawn: the "connected state" line as a resistance mechanism.
7. What a v7 variant would test next
- The index fold on the address (the layer-1 row) in this class, which removes the window-bit refusals on eras 0, 1, 4, 5 and 6 for the control and this class alike.
- A wider per-step chain: a spine of depth 8 to 16 so
deprises from about 11 toward the window, at the cost of ILP on the card (measure the rate first; the chain per step is what a two-level file exploits). - The window at 32 and 16 (
cs32s27x16,cs16s27x16) for the k curve, and the block at 16 x 27 (cs64s16x27, text 272) if the imem matters to the re-optimised core. - The op-mix re-weight of the census lane at this structure (the two closing instructions already take the injecting table).
- The PC 1 lock row (informational now: the verdict does not turn on it; it lands as an amendment if PC 1 runs it).
- A knob that reaches a gated file: not more live state but more WRITES per op the chip cannot skip (every op writing two registers, or a window write per load), priced against the card's own write cost first; the k lane's placed row at 21:00 says whether even that moves k.