Connected state: the header-bound kernel in the pack (the NVRTC worker loads it: check PASS on a 5090), docs/analysis/class-v6/connected-state.md with the structure, the liveness rows, the census rows and the first GPU rows (the 4090, worker, lock and chip rows to follow)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Documents-only replay of ecf98c0dd (6ccd2a6fe) for the box mirror master; left on the branch: igneum-pow/src/connected.rs igneum-pow/src/main.rs
This commit is contained in:
parent
e0cfa64f17
commit
e9e88cf4f2
1 changed files with 93 additions and 0 deletions
93
docs/analysis/class-v6/connected-state.md
Normal file
93
docs/analysis/class-v6/connected-state.md
Normal file
|
|
@ -0,0 +1,93 @@
|
|||
# Class v6: the connected-state variant (8 October 2026)
|
||||
|
||||
The experiment of the external review the founder accepted (the 15:4x BST rules): the remaining chip edge (about 2.0x to 2.2x node-for-node, 2.2x to 2.6x a node ahead) survives because a specialist can separate storage, arithmetic and memory scheduling. This lane reorganised the class v5 work, same dataset, same read width (4 bytes), about the same operation count, so that live state feeds each load address, the memory result feeds mixed arithmetic and cross-lane exchange, and that updates the live state for the next address, with a 64-register window per lane that stays necessary across the whole dependent chain. Research class only, behind `--class cs<W>s<S>x<R>` in `igneum-pow` (`igneum-pow/src/connected.rs`, branch `class-v6-connected` on the box mirror); never a chain class.
|
||||
|
||||
Every number below is measured unless marked modelled or owed. Times are UK (BST).
|
||||
|
||||
## 1. The structure
|
||||
|
||||
| Item | Class v5 (`mx8+sh256x27`, the control) | Connected state (`cs64s27x16`) |
|
||||
|---|---|---|
|
||||
| Registers per lane | 8 | 64 (the window) |
|
||||
| Per iteration | 64 base instructions with 16 loads, then a 256-instruction block run 27 times | 16 steps: a load, then a 27-instruction block run 16 times |
|
||||
| Load address | a base register, `x & MASK` (the era stride and site window under an era) | a window register, the same address rule |
|
||||
| Memory result | xor into the load's destination | xor into `r[m_j]`; block instruction 0 reads it |
|
||||
| Next address | whatever register the next load reads | written by the block's last instruction (add, sub, xor or shfl) from a fresh spine |
|
||||
| Result | fold of 8 registers | fold of all 64 (`lo` over the first 32 at 7 i, `hi` over the second 32 at 9 i; the v5 fold at a window of 8) |
|
||||
| Loads per hash | 128 | 128 |
|
||||
| ALU instructions per hash | 55,680 (55,296 shadow + 384 base) | 55,296 (0.7 percent fewer) |
|
||||
| Instruction text per iteration | 320 | 448 |
|
||||
| Negative controls kept out | | no long program (1,024), no select tree, no wide read (W = 16), no scratchpad |
|
||||
|
||||
The draw (deterministic from the seed words, attempt k re-seeded as every class): per step `a_j` then `m_j != a_j` (and the era window draws); per block instruction the op from the ten non-load families at the v5 weights, the source from the last four spine entries (the memory result first), the destination uniform over the window, the two immediates, the rotation, the selector bit and the shuffle mask. Four rules the first census pass forced, each a construction rather than a filter:
|
||||
|
||||
1. The cover: the first 64 injecting destinations walk a drawn permutation of the window, so every register takes an injecting write (rule (b)); a uniform draw left one register without one in about 70 percent of candidates.
|
||||
2. The fresh spine: only destinations of add, sub, xor, mad and shfl enter the spine, so every address is fresh by dataflow (rule (a')); with a lossy spine every seed exhausted 256 attempts.
|
||||
3. Lossy ops feed the next injection: or, mul and mulhi write the register the next injecting instruction of the block writes, so their value enters the chain and no register accumulates a lossy op across the 16 passes (a 16-pass or saturated a register the block never re-injected, a 16-pass mul cleared its low bits: rule (c) refused every candidate).
|
||||
4. The two closing instructions of a block come from add, sub, xor and shfl (the census lane's rule for a load source's writer): no product on the address writer (a 16-pass mad on the address register read bit z 57.6 at one site).
|
||||
|
||||
The acceptance rule is the sub-version 3 rule over the window: (b), (a') (the fixpoint over the iteration's execution order), (c) over 64 units (constant bits per register, lane-constant sites, saturation at 1 percent of the window's final values, saturated sources, output bias, distinct addresses, the value-level index-bit read judged inside the site's era window), then (c'') at 0.98 and (c''') at 0.995 over 2^20 evaluations per site. (a) holds by construction.
|
||||
|
||||
## 2. Liveness (seed `igneum-v6c/0`, attempt 0, program id 9ad55de91485542b)
|
||||
|
||||
The tool (`igneum-pow liveness --class cs64s27x16 --seed <s>`) walks the unrolled trace of one hash (128 loads, 55,296 ALU instructions). `live` is the backward set at an address: registers whose value just before the load feeds this or any later address. `dep` is the forward set: registers at the previous address whose values feed this address.
|
||||
|
||||
| Measure | Value |
|
||||
|---|---|
|
||||
| live before each address | 63 of 64 at every address until the last iteration's tail; iteration 7: 63 62 60 60 60 57 57 54 53 51 45 36 25 20 9 1 |
|
||||
| registers read by the result | 64 of 64 (the fold) |
|
||||
| dep per step (the same every iteration, the text repeats) | 1 9 13 14 8 15 7 8 12 14 14 13 10 9 14 9; mean 11.2 of 64 |
|
||||
| registers touched per block (read or written) | min 16, mean 20.4, max 24 of 64 |
|
||||
| reads per register per hash | min 384, mean 1,724, max 3,208 |
|
||||
| writes per register per hash | min 128, mean 866, max 1,664 |
|
||||
| op mix of the 432 block instructions | add 82, mul 62, xor 51, shfl 51, sub 46, mulhi 39, rotl 35, mad 31, rotr 20, or 15 |
|
||||
|
||||
Meaning: the whole window is necessary over the hash (no register can be dropped or parked for long: the longest gap between writes to one register is under one iteration), and every value is read by a later address, so there is no side calculation a specialist can move to a separate engine. The dependent chain per step is narrow: about 11 of the 64 registers at one address feed the next address through 432 operations, and a block touches about 20. That is the shape a specialist will exploit: a two-level file, about 20 registers hot per step and 44 warm, the hot set moving along the text; the k lane prices exactly that (section 5). The program listing is `liveness.txt` and `program.json` of the pack.
|
||||
|
||||
## 3. Census (the census lane's sub-version 3 harness, `tools/attack/v6-census/v6census.sh`, (c''') 0.995 on, on a rented 5090 host with 40 threads, 16:01 to 16:08 BST)
|
||||
|
||||
| Run | Seeds | Accepted | Exhausted | Rejections per candidate | Mean attempt | Max attempt | Window-bit refusals | Over 6 sigma (reported) |
|
||||
|---|---|---|---|---|---|---|---|---|
|
||||
| cs64 no era | 256 | 256 | 0 | 0.283 | 0.39 | 5 | off | 146 of 256 (max z 128) |
|
||||
| control `mx8+sh256x27` no era (census lane, build-3) | 256 | 256 | 0 | 0.668 | 2.01 | 18 | off | 143 of 256 (max z 97) |
|
||||
| cs64 eras 0 to 7, window-bit refusal on | 8 x 32 | 256 | 0 | 0.635 | 1.74 | 8 | 306 | 0 (max z 5.9) |
|
||||
| control eras 0 to 7, refusal on (census lane, build-4) | 8 x 32 | 256 | 0 | 0.830 | 4.87 | 22 | 218 | 0 (max z 5.9) |
|
||||
| cs64 eras 0 to 7, refusal off | 8 x 32 | 256 | 0 | 0.304 | 0.44 | 2 | off | 94 of 256 (max z 71.5) |
|
||||
|
||||
The no-era rejections are all rule (c) (constant bits, saturation or bias in the 64-unit test); the control's are mostly (a') and (a), which this class satisfies by construction. The window-bit refusals sit on eras 0, 1, 4, 5 and 6 (55 to 64 each) and nearly vanish on eras 2 and 3 (0 and 1): the product-bit class (a product's low bits reaching an address bit through the era stride), the same class the control carries and the layer-1 index fold removes; this class does not change it either way.
|
||||
|
||||
F8 form (`igneum-pow cs-uniform`, 16 seeds x 2^20 nonces, no era): the top 0.1 percent item share reads 0.9987 to 1.0016 of a uniform control of the same size (the control class 0.9993 to 1.0016); the worst per-site distinct ratio 0.9941 (seed 9, site 9; the sequential-nonce sample of the F8 tool, the acceptance's own 2^20 sample passed 0.995 on every accepted program).
|
||||
|
||||
Meaning: the class censuses at least as well as v5 on every instrument of the harness and takes fewer attempts; it inherits v5's product-bit bias and the fix is the same fold. The acceptance costs about 3.7 s per candidate on one core (the (c'') pass dominates, 7 G lane-ops), against v5's 2.8 s.
|
||||
|
||||
## 4. GPU cost
|
||||
|
||||
Instrument: the class v5 nvcc harness (`proto-newpow/class-v5/bench.cu` on `box/ds55-v5`, the v5 design page's 4090 rows), compiled per pack with `-Xptxas -v`, 250 batches of 2^24, nvidia-smi at 1 Hz, both packs on the same card minutes apart; vectors 3 of 3 PASS and the dataset self-test PASS on every row. The cs64 pack is over the class v4 memory-hard dataset (no leaves); v5-genesis carries its 93 leaves. Stock means the card's own power limit and no clock lock.
|
||||
|
||||
| Card | Pack | MH/s | Mean W | Microjoules per hash | Registers per thread | Spills | Resident blocks per SM (1 warp per block) | Fingerprint of 2^24 at base 0 |
|
||||
|---|---|---|---|---|---|---|---|---|
|
||||
| RTX 5090 (Vast 54862507, driver 580.159.03, 575 W cap), stock | cs64s27x16 | 64.93 | 574.8 | 8.853 | 80 | 0 | 24 | ad0cec2a42c84aff |
|
||||
| RTX 5090, the same card | v5-genesis | 65.30 | 574.8 | 8.803 | 48 | 0 | 24 | ae74193ddad19e19 |
|
||||
| RTX 4090 (RunPod 386yytbh4bkfnz, driver 580.159.04, 450 W cap), stock | cs64s27x16 | ROW_4090_CS | | | | | | |
|
||||
| RTX 4090, the same card | v5-genesis | ROW_4090_V5 | | | | | | |
|
||||
| RTX 5090 on PC 1 at the 1,300 MHz lock | both | OWED: PC 1 booked to 17:40 BST; the bound pack and the kit are at build-1:/srv/builds/igneum-wt-connected/cs-kit (sha c9aff54aaf10d79e) and the hash lane publishes the job when PC 1 frees | | | | | | |
|
||||
|
||||
The kit worker (the brief's `worker --bench`): ROW_WORKER.
|
||||
|
||||
Meaning: at stock the window costs the 5090 0.6 percent of energy per hash against a 10 percent budget; the compiled allocation is 80 registers per thread with no spill, so the window is in registers, not local memory, and the occupancy under this harness is the same as v5's. The harness reads 65 MH/s where the NVRTC worker reads about twice that on a 5090 (one warp per block, 24 resident blocks); the ratio between two packs on the same harness is the measurement, the absolute rate is not. The lock row is where the energy comparison binds (the 5090 at the lock reads 2.33 microjoules per hash on v5); the stock rows say the card is bound by its power cap in both cases and the window moves the rate by under 1 percent.
|
||||
|
||||
## 5. The chip side
|
||||
|
||||
ROW_CHIP
|
||||
|
||||
## 6. The score and the verdict
|
||||
|
||||
ROW_SCORE
|
||||
|
||||
## 7. What a v7 variant would test next
|
||||
|
||||
- The index fold on the address (the layer-1 row) in this class, which removes the window-bit refusals on eras 0, 1, 4, 5 and 6 for the control and this class alike.
|
||||
- A wider per-step chain: a spine of depth 8 to 16 so `dep` rises from about 11 toward the window, at the cost of ILP on the card (measure the rate first; the chain per step is what a two-level file exploits).
|
||||
- The window at 32 and 16 (`cs32s27x16`, `cs16s27x16`) for the k curve, and the block at 16 x 27 (`cs64s16x27`, text 272) if the imem matters to the re-optimised core.
|
||||
- The op-mix re-weight of the census lane at this structure (the two closing instructions already take the injecting table).
|
||||
- The kit worker row on both cards and the PC 1 lock row.
|
||||
Loading…
Reference in a new issue