Squashed from six commits (tag ca2-era-pre-squash) for one merge. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
21 KiB
Read width of the lottery hash: 4, 16 and 64-byte loads, a per-load mix, and a written scratch (gate 1 experiment)
5 October 2026. Branch readwidth (worktree ../igneum-wt-readwidth), commits 019b014 and b970dda plus the measurement commit. Nothing here changes consensus, the live generator, the pinned vectors or a shipped binary: every class sits behind --class in igneum-pow and the default class is generator version 2 byte for byte (igneum-pow/tests/packs.rs still compares the four pinned packs against the emitters). Numbers and a recommendation; the decision is the project lead's.
1. The question
The bench-log entry "the 9070 XT on the eGPU" (5 October 2026) found the hash bound by dependent random 4-byte reads over the 1 GiB dataset, 128 per hash: the RX 9070 XT finishes 2.4 to 2.7 G such reads a second (18 MH/s), the RTX 5090 16 to 18 G (127 MH/s), the M5 Max 3.45 G (23 to 28 MH/s). AMD fetches a 64-byte line per 4-byte read, so 94 percent of its memory traffic is unused; NVIDIA fetches a 32-byte sector and its 96 MB L2 catches a share. the project lead's question: would wider reads keep the chip-resistance property (latency-bound, random access) while closing the vendor gap? Two additions from the coordinator: a per-load width drawn from an era-fixed mix so no chip is built for one width, and a written per-warp scratch so part of the memory work cannot be mirrored into read-only SRAM.
2. What was built (all behind the flag)
Class (--class) |
Loads per hash | What a load does | Dataset bytes per hash |
|---|---|---|---|
v2 (= w4, the lottery hash) |
128 | dst ^= dataset[src & MASK], one 4-byte word |
512 |
w16 |
128 | the 16-byte-aligned group of 4 words at src & MASK, every word folded into dst |
2,048 |
w64 |
128 | the 64-byte-aligned item (16 words), every word folded | 8,192 |
w64x4 |
32 (4 load slots) | as w64; the same bytes per hash as 512 loads of 4 bytes |
2,048 |
mix50-35-15 |
128 | per load, width 4, 16 or 64 bytes drawn from the program stream with probabilities 50/35/15 | 1,664 to 3,680 over the six programs measured (expected 2,202) |
mix25-50-25 |
128 | the same with 25/50/25 | 2,240 to 5,024 (expected 3,200) |
scr<k>k<kb> |
128 memory operations | k of the 16 slots are scratch read-modify-writes into a kb KiB per-warp scratch (16-byte slots, lane-major); the other 16 - k are 4-byte loads |
4 x (16 - k) x 8 reads plus 16 B read and 16 B written per scratch op |
The fold. A load of W words reads the W-word-aligned address b = (src AND MASK) AND NOT (W - 1) and sets x = dst XOR w[0]; for j in 1..W: x = (rotl(x, 11) * 0x9e3779b1) XOR w[j]; dst = x (verify::fold_words, mirrored in the three kernel dialects). The rotate-multiply between the words makes the fold state-dependent: two different lines give two different maps of dst (the multiply by an odd constant is not xor-linear), so no function of the line alone can stand in for it, and a dataset of pre-folded lines cannot replace the dataset. The dependent chain is unchanged: the next load's address comes from a register that the fold wrote. Width 1 is the lottery hash's xor of one word, so w4 is the pinned program bcc1248b10cc90f2 exactly.
The per-load draw. Every class other than v2 takes one extra draw per instruction (below(100), the width roll, consumed on every slot so the stream stays uniform) after the nine draws of spec 01 section 1.4.3; on a load slot the width is the first entry of the mix whose cumulative weight exceeds the roll. The program id of a class is FNV-1a-64("igneum-program-rw/" || 2 || seed words || attempt || mix[3] || slots [|| "scratch/" k kb]), so no class program can pass for a version 2 program. The acceptance rule of 1.4.6 runs unchanged on the aligned addresses (lane-constant sites and the distinct-address bound, scaled to the dataset loads per hash).
The scratch (variant 5, measurement only). The kernel runs N persistent warps (one per block or work-group of 32); warp w owns scratch w and runs units w, w + N, ... of the launch. A scratch op reads the lane's 16-byte slot src AND (slots - 1): three data words behind a per-unit tag; a slot whose tag is not this unit's reads as its fill splitmix32(((base + lane) XOR seed[j]) + slot x 0x9e3779b1 + (j + 1) x 0x85ebca77), the three words are folded into dst as above, and the slot is rewritten (tag, x XOR w1, rotl(x, 7) XOR w2, x + w0). The CPU verifier holds the touched slots of one unit (at most 32 x k x 8) and nothing else. The scratch is per lane (a 32 KiB warp scratch is 64 slots per lane, 128 KiB is 256), so two lanes never race on a slot and the result is a function of (program, day, unit) alone; the GPU's tags make the lazy fill exact as long as a tag is not reused within the arena's history (the salt advances per unit; it wraps after 2^32 units, a measurement caveat, not a design).
3. Method
| Step | Command (every figure in the bench log carries its command) |
|---|---|
| Packs | igneum-pow export --seed <s> --class <c> --out proto-cuda/packs-readwidth/<name> (23 packs; the six mix seeds per mix are igneum-readwidth/A/<k> and /B/<k> with attempt 0 accepted, A/4 skipped: rejected at attempt 0) |
| CPU verifier | igneum-pow bench --seed igneum-genesis --class <c> --warps 50 (M5 Max, one core; load average 4 to 9 from other agents' builds during the run) |
| Metal | proto-metal/packbench --pack <dir> --batches 5 --batch-log2 24 --group 256 [--warps N] (new harness: runs the pack's own text; vectors, cache FNV, dataset words, 2^24 fingerprint, MH/s by GPU time) under the measure lock |
| Apple OpenCL | proto-opencl/igneum-bench-cl-rw --bench-pack --pack <dir> --batches 5 --batch-log2 24 and --memprobe (Apple's OpenCL, a correctness check and an approximate rate) |
| Emulators | proto-cuda/emu/emu.sh ../packs-readwidth/<p> --batch-log2 13 --batches 1 --block-warps 2 (the CUDA text as C++); proto-opencl/emu/emu.sh ../packs-readwidth/<p> 0 32 --sg 32 and 1 64 --sg 64 (the OpenCL text, wave32 and wave64) |
| RTX 5090 | job run-readwidth-5090-20261005 on PC 2 (relay/playbooks/readwidth-5090.ps1): the NVIDIA card switched off in the app through POST app.url/api/cards and restored after; igneum-worker-cuda --memprobe, then --bench --pack <dir> --batches 5 --batch-log2 24 per pack (NVRTC, the pack's own text, vectors through the bound kernel, 2^24 fingerprint) |
| RX 9070 XT | job run-readwidth-9070-20261005 on PC 1 (relay/playbooks/readwidth-9070.ps1): only the gfx1201 card switched off; igneum-worker-opencl --device D --memprobe, then --bench-pack --pack <dir> --batches 5 --batch-log2 24 per pack |
Latency-bound share = measured MH/s x loads per hash / the card's dependent-read ceiling for that width from its own probe at 1024 MiB (for a mix, the harmonic combination of the widths' ceilings weighted by the program's width counts). A share near 1 means the hash runs at the card's random-access limit, the property the design wants; a share well under 1 means something else bounds it (bandwidth, ALU, occupancy).
4. Results (full tables with commands in docs/bench-log.md, "read width of the lottery hash")
Probe ceilings at 1024 MiB (G dependent reads/s): RTX 5090 4 B 17.5, 16 B 18.0, 64 B 9.1 (584 GB/s), stream 1,579 GB/s; RX 9070 XT 4 B 2.42, 16 B 2.43, 64 B 2.47 (158 GB/s), stream 636; M5 Max (Apple OpenCL, approximate) 3.50 / 3.51 / 3.51, stream 522.
| Class | dataset B/hash | RTX 5090 MH/s (latency-bound share) | RX 9070 XT (share) | M5 Max Metal (share) | 5090 / 9070 | DRAM bytes moved per hash, NVIDIA 32 B sector / AMD 64 B line | CPU verify ms per unit |
|---|---|---|---|---|---|---|---|
| v2 = w4 (today) | 512 | 136.1 (0.96) | 18.15 (0.87) | 27.74 (1.01) | 7.5x | 4,096 / 8,192 | 0.604 |
| w16 | 2,048 | 139.8 (0.90) | 17.90 (0.84) | 28.26 (1.03) | 7.8x | 4,096 / 8,192 | 0.610 |
| w64 | 8,192 | 71.9 (0.58) | 17.59 (0.78) | 28.27 (1.03) | 4.1x | 8,192 / 8,192 | 0.630 |
| w64x4 (32 loads) | 2,048 | 275.3 (0.56) | 75.19 (0.84) | 109.7 (1.00) | 3.7x | 2,048 / 2,048 | 0.160 |
| mix50-35-15 (6 programs, min / median / max) | 1,664 to 3,680 | 99.5 / 114.2 / 121.0, spread 18.8% | 17.45 / 18.76 / 18.83, 7.4% | 25.36 / 27.26 / 28.43, 11.3% | 6.1x | 5,939 / 8,192 expected | 0.620 |
| mix25-50-25 (6 programs) | 2,240 to 5,024 | 95.9 / 107.3 / 119.8, 22.3% | 17.84 / 18.45 / 18.85, 5.5% | 23.21 / 24.68 / 25.21, 8.1% | 5.8x | 7,168 / 8,192 expected | 0.614 |
4.1 Per watt and per pound (consequences review C11)
The runs carried no power sampling; the watts are the telemetry entry's (docs/bench-log.md, opencl-rdna4-telemetry, 5 October 2026: the RTX 5090 at 307.6 W under its 450 W cap for 122.3 MH/s, the RX 9070 XT at 199 W of its 304 W rating for about 17.8 MH/s, both on v2 with the shader clock at its top and the die waiting on memory), held constant across classes because every class is memory-bound on both cards (approximate: a class that moves more bytes per hash draws somewhat more at the memory controller, unmeasured). The Mac's GPU power is not measurable without root (powermetrics) and is taken as about 50 W (approximate, from memory). Prices are UK list, approximate, from memory.
| Class | RTX 5090 MH/W (at 307.6 W) | RX 9070 XT MH/W (at 199 W) | 5090 / 9070 per watt | M5 Max MH/W (at about 50 W GPU, approximate) | 5090 MH per pound (at about 1,900, approximate) | 9070 XT MH per pound (at about 570, approximate) |
|---|---|---|---|---|---|---|
| v2 (w4) | 0.442 | 0.091 | 4.9x | 0.55 | 0.072 | 0.032 |
| w16 | 0.454 | 0.090 | 5.1x | 0.57 | 0.074 | 0.031 |
| w64 | 0.234 | 0.088 | 2.6x | 0.57 | 0.038 | 0.031 |
| w64x4 | 0.895 | 0.378 | 2.4x | 2.19 | 0.145 | 0.132 |
| mix50-35-15 (median) | 0.371 | 0.094 | 3.9x | 0.55 | 0.060 | 0.033 |
| mix25-50-25 (median) | 0.349 | 0.093 | 3.8x | 0.49 | 0.056 | 0.032 |
| scr8k32 | 0.397 | 0.071 | 5.6x | 0.98 | 0.064 | 0.025 |
| scr2k32 | 0.372 | 0.074 | 5.1x | 0.52 | 0.060 | 0.026 |
Reading: whatever width is chosen, an AMD home miner keeps about a seventh of a 5090's rate and pays about 4.5x the electricity per hash, because every width costs the 9070 XT the same 2.4 G line fetches a second; per pound of card the 5090 is 2.2x the 9070 XT at v2 and w16 (0.072 against 0.032 MH/s per pound) and 4.9x per watt; only w64x4 narrows the per-pound gap (0.145 against 0.132), and that class fails the width rule. The consequence for the decision (D6, the project lead's): AMD's line width is not a read-width question at all; it is the card's random-access rate, and the levers that act on it (the 64 MB Infinity Cache against the dataset size, the memory path) are v3-or-3.0 questions outside this experiment.
Scratch, variant 5 (N persistent warps; GPU cost against the persistent control scr0k32; working set = 1 GiB + 256 MiB + 128 MiB output + N x size):
| Class | RMW share | dataset B/hash | scratch B/hash read + written | RTX 5090 MH/s, 2,048 warps of 4,080 resident (vs control, share) | M5 Max Metal (vs control) | RX 9070 XT, 4,096 warps | working set 5090 / 9070 / Mac |
|---|---|---|---|---|---|---|---|
| scr0k32 | 0 | 512 | 0 | 139.1 (control, 0.98) | 28.25 (control) | 17.88 (control, 0.86) | 1.4 GiB / 1.5 GiB / 1.5 GiB |
| scr2k32 | 12.5% | 448 | 256 + 256 | 114.4 (-18%, 0.80) | 26.14 (-7%) | 14.65 (-18%) | same |
| scr4k32 | 25% | 384 | 512 + 512 | 109.8 (-21%, 0.76) | 31.74 (+12%) | 14.00 (-22%) | same |
| scr8k32 | 50% | 256 | 1,024 + 1,024 | 122.1 (-12%, 0.82) | 49.08 (+74%) | 14.17 (-21%) | same |
| scr2k128 | 12.5% | 448 | 256 + 256 | 110.1 (-21%, 0.77) | 26.24 (-7%) | 14.07 (-21%) | 1.6 GiB / 1.9 GiB / 1.9 GiB |
| scr4k128 | 25% | 384 | 512 + 512 | 98.0 (-30%, 0.68) | 28.08 (-1%) | 13.14 (-27%) | same |
| scr8k128 | 50% | 256 | 1,024 + 1,024 | 72.8 (-48%, 0.49) | 35.44 (+25%) | 12.03 (-33%) | same |
Resident warps and the cap: the 5090 holds 4,080 warps at one warp per block (24 blocks per SM x 170 SMs; 8,160 at 8 warps per block), so 128 KiB each is 510 MiB and the whole working set 1.9 GiB; a 1 MB scratch would have been 4.0 GiB at this geometry and 10.6 GiB at the 64-warp-per-SM figure, which is why the cap moved the size to the tens of kilobytes. The occupancy query returned 24 blocks per SM before and after the arena allocation: the allocation did not change it. The 9070 XT's OpenCL runtime has no occupancy query; 4,096 persistent warps were launched (64 per compute unit over 64 CUs, approximate) and the arena is 128 MiB at 32 KiB, 512 MiB at 128 KiB. The Mac's residency is not reported; 2,048 to 16,384 warps were swept and the best row kept.
Chip model, re-run with the measured widths (the M16 arithmetic of docs/analysis/m16-recompute-attacker-2026-10-05.md; the on-die-cache recompute chip's row per scratch variant is the ca2-soundness branch's, as agreed with the Counter ASIC 2.0 coordinator):
| Class | what a chip with its own DRAM controller gains over the GPU's memory system | what a chip with on-die SRAM gains |
|---|---|---|
| v2 | the GPU fetches 8 to 16x the bytes it uses (AMD 64 B, NVIDIA 32 B per 4 B); a chip fetching 32 B bursts moves 4,096 B per hash, the 5090's figure, so nothing over NVIDIA and 2x over AMD in traffic, none in latency (the chain is 128 dependent DRAM latencies on either) | the recompute attacker of M16: 150,000 integer ops per hash against the 256 MiB cache; 2.4x at equal silicon before a fixed-function factor (unchanged by the width) |
| w16 | the same: 4,096 / 8,192 bytes moved, 2,048 used; traffic efficiency 50 percent on NVIDIA, 25 on AMD | unchanged: the fold uses every byte, so the chip recomputes 128 items per hash exactly as before; the SRAM mirror of the read-only dataset (1 GiB) stays out of reach |
| w64 | every byte moved is used on both vendors (8,192 moved, 8,192 used); the 5090 is bandwidth-bound at 589 GB/s, so a chip with HBM3 class bandwidth (several TB/s, approximate) is bandwidth-advantaged: the Ethash shape | unchanged in op count; but the chain of 128 loads now moves 8 KB, so a chip's advantage shifts from latency to bandwidth per dollar, which is the wrong direction for the design's 2x target |
| w64x4 | 2,048 moved and used; 32 latencies per hash; every card 4x faster; a bandwidth-rich chip gains as above | 32 items per hash: the recompute attacker's op count falls 4x (37,500 per hash), so the M16 gain rises 4x: fails the 2x target by arithmetic |
| mixes | between v2 and w64 per program; the chip cannot be built for one width, but the GPU pays the 64-byte hours (the 5090 loses up to 27 percent in a heavy hour) | as v2 per item; the recompute attacker is indifferent to the width |
| scratch | a chip must provide writable memory for N units in flight: 32 KiB x N at the GPU's geometry (128 MiB at 4,080), against the 256 MiB read-only cache it could mirror into SRAM (54 to 83 mm^2 at a leading node, the coordinator's figure, approximate); but a unit touches at most k x 8 x 32 slots (4 KiB at 50 percent), the fill is a function and the tags are per unit, so a chip need only hold the touched set per unit in flight (the soundness caveat below) | the dataset reads replaced by scratch ops are reads the chip no longer has to serve from the 1 GiB; at 50 percent the recompute attacker computes 64 items instead of 128 |
Soundness (variant 5, measurement only, as instructed; the chip row is the ca2-soundness branch's, a465881: the on-die-cache recompute chip's gain is 2.4x at 0, 12.5, 25 and 50 percent, replaced or added, 32 or 128 KB, so the scratch does not move it): the per-unit scratch starts from a fill that any implementation can compute, and a unit writes at most k x 8 slots per lane; an implementation that keeps only the touched slots of each unit in flight (the CPU verifier does exactly this) needs 16 B x touched slots, not the nominal arena, so the "real memory a chip must provide" is bounded by units in flight x touched slots, not by N x 32 KiB. The variant forces memory that is written, which SRAM can hold as well as DRAM; it does not force memory that is large. A written region that outlives the unit (state carried across units) would, and the CPU verifier could not replay it. This is the finding, not a recommendation.
5. Recommendation (the decision is the project lead's)
the project lead's rules, as passed by the coordinator: width = the widest read that keeps every card latency-bound (achieved within 90 percent of the probe ceiling at that width) with margin on the 5090 (bytes per hash x rate under a third of the 1,579 GB/s stream); the mix is in only if the six-program spread is under 5 percent per card; the scratch share is the smallest at which the chip model's gain falls under 1.5x at the lowest GPU cost within the 6 GB cap.
| Variant | Verdict under the rules | Numbers |
|---|---|---|
| w16 (16-byte loads, 128 per hash) | PASSES the rules: shares 0.90 / 0.84 / 1.03 (the 9070 XT's 0.84 equals its v2 share of 0.87 within noise: the card is at its ceiling in both), 286 GB/s on the 5090 = 18 percent of the stream. It does NOT close the vendor gap (7.8x against 7.5x), because the memory systems already move a sector or a line per load; it changes what the fold consumes, nothing the DRAM does | the only width row that passes; a no-cost change in rate (+2.7 percent 5090, -1.4 percent 9070 XT, +1.9 percent M5 Max) |
| w64 | FAILS: 5090 share 0.58, 37 percent of the stream; closes the gap to 4.1x only by making the 5090 bandwidth-bound | |
| w64x4 | FAILS: shares 0.56 / 0.84 / 1.00, the recompute gain rises 4x | |
| mix 50/35/15 and 25/50/25 | OUT: spreads 18.8 and 22.3 percent on the 5090, 7.4 and 5.5 on the 9070 XT, 11.3 and 8.1 on the M5 Max, all over 5 percent; a chip is not built for a width anyway (see the model: the width does not change the recompute attacker) | |
| scratch | OUT: every share costs the 5090 12 to 48 percent and the 9070 XT 18 to 33 percent, and raises the M5 Max's rate (the arena is cached there); the soundness branch's chip row (ca2-soundness a465881, the on-die-cache recompute chip) stays at 2.4x at every share, 32 or 128 KB, because the verifier resets the scratch per unit and the live state is the hash's own read-modify-writes, which a chip keeps in 80 to 320 B per lane; under the rule the share is 0. The rows stay as the measurement that decided it |
Recommendation: keep 128 loads per hash and 4 bytes per load (v2) for the devnet; if a width change is wanted for the fold's sake (every byte of the sector consumed, which removes the "94 percent waste" statement from the AMD entry without changing what the card does), w16 is the one that passes every rule and costs nothing measurable, and it is the only width worth a vector re-cut. The AMD gap is a random-access gap (2.4 G against 17.5 G dependent reads per second at 1 GiB on the cards we own), and no read width closes it without turning the 5090 bandwidth-bound; the levers that act on the gap are the ones outside this experiment (the AMD card's memory path, and the dataset size against the 5090's 96 MB L2 share, which the probe's 64 MiB rows show at 9 G reads/s against 2.4 at 1 GiB). The per-load mix is out on stability; the scratch is out on GPU cost and on the soundness caveat.
6. What w16 would change if adopted (not done; the decision is the project lead's)
| Where | Change |
|---|---|
docs/spec/01-lottery-hash.md 1.4.1 |
load: dst = fold(dst, dataset[b .. b + 4)), b = (src AND MASK) AND NOT 3, with the fold written out; 1.4.3 unchanged (no width draw for a fixed width); 1.4.6 unchanged (the aligned address is the address the rule sees) |
| 1.5 | "A load reads one 4-byte word" becomes 16 bytes aligned; the single-form text-search rule of 1.14 item 2 becomes the wide form; the item size (64 B) and dataset[w] = item(w >> 4)[w AND 15] unchanged |
| 1.11 | unchanged in count (4,096 items per unit; the verifier derives the same items) |
| 1.15, 1.17 | every vector re-cut (new program ids: the class enters the id or the generator version steps to 3); the four pinned packs replaced; the conformance fuzz re-run on Metal, CUDA and OpenCL (this branch's 23 packs and the three emulators are the template) |
| Litepaper, Mining ("random reads over a multi-gigabyte dataset") and the vs-RandomX "128 dataset addresses" rows | "128 reads of 16 bytes"; site/bench.html sector arithmetic (32 B per 4 B) becomes 32 B per 16 B |
| Workers | no host change: the kernel text carries the loads; proto-cuda/host.cu's static mask check (TESTS.md section 5) learns the wide form |
| Cost on the 5090 | none measured (+2.7 percent); on the 9070 XT -1.4 percent; verifier +1 percent |
7. Files
igneum-pow/src/{generator,verify,accept,emit,memhard,main}.rs (the classes, behind --class), proto-cuda/packs-readwidth/ (23 packs), proto-metal/packbench.swift (Metal from a pack's files), proto-opencl/host.c (--bench-pack, --warps, the 16-byte probe row, the scratch arguments), proto-cuda/nvrtc/worker.cpp (--bench, --memprobe, the scratch arena), proto-cuda/nvrtc/packfile.h (class fields; string-seed packs), proto-cuda/emu/cuda_runtime.h and proto-opencl/emu/{emu_opencl.h,emu_main.cpp} (vector types, the persistent launch), relay/playbooks/readwidth-*.ps1 (the PC jobs: the card under test off in the app and restored, never the other card).