diff --git a/docs/analysis/chip-model-v3.md b/docs/analysis/chip-model-v3.md
index 8a48ca7af..0510d7ce0 100644
--- a/docs/analysis/chip-model-v3.md
+++ b/docs/analysis/chip-model-v3.md
@@ -23,6 +23,7 @@ is a measurement of a chip; every GPU figure says where it was measured. "Approx
| Cache mirror plus a 96 MB hot table, N5 headline | 175 mm^2, $68 | same, so a hot table costs 0.49 mm^2 and $0.23 per MB (linear, approximate) |
| 512 MiB and 1 GiB mirrors, N5 headline | 255 mm^2 and 510 mm^2; $111 to $306 | same, section 4 (the growth rule's cache at years 4 and 12, priced at today's node) |
| GPU-class die | 750 mm^2 (the equal-silicon comparison) | M16 section 3 |
+| CPU verifier, one M5 Max core (loaded, load average 5.6; ratios are the measurement) | v2 1.31 to 1.36 ms per unit, x4 1.92 to 1.96 (1.45x), x8 2.79 (2.1x); worst cold 1.58 / 2.04 / 2.94 ms | `docs/plans/mixer-x4.md` section 6.4, 5 October 2026 21:40 UTC |
## 2. The rows
@@ -35,11 +36,12 @@ silicon" = bare x (750 - SRAM) / 750 x 3: the SRAM takes die area the logic does
| v2 as shipped (the M16 and scratch-soundness row) | x1 | 149,760 | 334 MH/s | 256 MiB | 128 / $46 | 2.45x (2.39x against 139.7) | 7.4x | 6.1x |
| v2 at w16 (not adopted; the chip's cost is items, not bytes: unchanged) | x1 | 149,760 | 334 | 256 MiB | 128 / $46 | 2.39x against 139.8 | 7.2x | 5.9x |
| v3: mixer x4 | x4 | 599,040 | 83.5 MH/s | 256 MiB | 128 / $46 | 0.61x | 1.84x | 1.53x |
-| v3 plus a 32 MiB hot table, added form (16 dataset loads and k hot loads) | x4 | 599,040 (a hot load is one SRAM read, no item) | 83.5 | 288 MiB | 144 / $53 | 0.61x or below (owed: the 5090's added-form rate; the hot loads cost it something, the chip nothing) | 1.84x or below | 1.49x |
-| v3 plus a 64 MiB hot table, added form | x4 | 599,040 | 83.5 | 320 MiB | 160 / $61 | 0.61x or below (owed) | 1.84x or below | 1.45x |
-| v3 at year 4 (cache 512 MiB under option C, dataset 4 GiB), 64 MiB hot table | x4 | 599,040 | 83.5 | 576 MiB | 287 / $130 | 0.61x | 1.84x | 1.14x |
-| v3 at year 12 (cache 1 GiB, dataset 8 GiB), 64 MiB hot table | x4 | 599,040 | 83.5 | 1,088 MiB | 542 / $330 | 0.61x | 1.84x | 0.51x |
-| v3 with the mixer at x8 instead (the next lever, not adopted) | x8 | 1,198,080 | 41.7 | 256 MiB | 128 / $46 | 0.31x | 0.92x | 0.76x |
+| MEASURED, NOT ADOPTED (layer 5 decided out of v3 on the PC rows, coordinator 21:40 UTC): v3 plus a 32 MiB hot table, added form (16 dataset loads and k hot loads): the honest card pays the hot loads, this chip pays SRAM only | x4 | 599,040 (a hot load is one SRAM read, no item) | 83.5 | 288 MiB | 144 / $53 | 0.66x at the Mac's g = 0.93 (126.6 MH/s); 0.70x at the 5090's g = 0.87 (118.4); the 9070 XT's g 0.84 | 1.98x (Mac g), 2.11x (5090 g) | 1.60x, 1.71x |
+| MEASURED, NOT ADOPTED: v3 plus a 64 MiB hot table, added form | x4 | 599,040 | 83.5 | 320 MiB | 160 / $61 | 0.71x at the Mac's g = 0.87 (118.4 MH/s); 0.73x at the 5090's g = 0.84 (114.3); the 9070 XT's g 0.80 | 2.12x (Mac g), 2.19x (5090 g) | 1.67x, 1.73x |
+| v3 at year 4 (cache 512 MiB under option C, dataset 4 GiB), no hot table | x4 | 599,040 | 83.5 | 512 MiB | 255 / $111 | 0.61x | 1.84x | 1.21x |
+| v3 at year 12 (cache 1 GiB, dataset 8 GiB) | x4 | 599,040 | 83.5 | 1 GiB | 510 / $306 | 0.61x | 1.84x | 0.59x |
+| v3 with the mixer at x8 (the candidate under the coordinator's rule of 21:30 UTC: in if the verify stays under 10 ms per warp on one Mac core and the daily 1 GiB build under 1 s on every discrete card) | x8 | 1,198,080 | 41.7 | 256 MiB | 128 / $46 | 0.31x | 0.92x | 0.76x |
+| x8 at year 4 | x8 | 1,198,080 | 41.7 | 512 MiB | 255 / $111 | 0.31x | 0.92x | 0.61x |
The era draws of spec 1.13.1 cost the chip nothing in this model: the mixer round count is not drawn, the op
weights and fold rotations change the program, not the item derivation, so the chip's ops per hash stand. The
@@ -49,15 +51,25 @@ chip's cost not at all.
Arithmetic, row v3: 36 x 130 = 4,680 ops per item; x 128 = 599,040 per hash; 50 x 10^12 / 599,040 = 83.5 x 10^6
hashes per second; 83.5 / 136.1 = 0.613; x 3 = 1.84; equal silicon (750 - 128) / 750 = 0.829, x 1.84 = 1.53.
Hot table rows: 32 MiB x 0.49 mm^2 per MB = 16 mm^2, 64 MiB = 32 mm^2 (the 96 MB column of `sram-mirror.md`
-scaled linearly); (750 - 144) / 750 = 0.808 and (750 - 160) / 750 = 0.787. Year 4 and 12 rows: the mirror of
+scaled linearly); (750 - 144) / 750 = 0.808 and (750 - 160) / 750 = 0.787. The honest denominator in the added
+form is the v2 rate times `g`, the card's measured ratio with the hot loads added: on the M5 Max tonight
+`g = 0.93 / 0.87 / 0.83` at 32 / 64 / 96 MiB (the cache agent, relayed by the coordinator at 21:23 UTC;
+`docs/plans/hot-table.md` carries the runs); the 5090's and the 9070 XT's `g` are the PC rows, owed, and until they
+land the row carries the Mac's `g` against the 5090's rate, which is a mixed figure and is marked so. Year 4 and 12 rows: the mirror of
`sram-mirror.md` section 4 at N5 for 512 MiB and 1 GiB plus the 64 MiB table, at today's density (the node of
those years is denser by about 1.8x at year 10 on the trend the same file cites; the row is a floor on the area,
not a forecast).
## 3. The margin, plainly
-The combined headline row reads 1.84x with the 3x factor at an equal integer budget, 1.5x with the SRAM area
-deducted. The claim is "under 2x", and the margin is thin:
+The combined headline row is the mixer row alone (layer 5 is out: the added form costs the 5090 13 to 16 percent
+and the 9070 XT 16 to 20 percent against the 0.97 bar, coordinator 21:40 UTC; the width stays 4 bytes; the era
+draws and the cache growth cost this chip nothing at year 0). At x4 it reads 1.84x with the 3x factor at an equal
+integer budget, 1.53x with the SRAM area deducted; at x8 0.92x and 0.76x. The hot-table rows above are kept as
+measured, not adopted: against THIS chip an added hot table is a cost to the honest card and none to the chip, so
+it would have moved the row the wrong way by the card's own `g` (2.11x to 2.19x at the 5090's g). The claim is
+"under 2x" at x4 on both conventions, with the margin thin on the equal-budget one; at x8 the chip is under 1x on
+both:
- the 3x fixed-function factor is approximate and from memory; at 3.3x the equal-budget row reads 2.0x;
- the denominator is one card's measured rate on one night (136.1 against 139.7 the night before: 2.6% apart);
@@ -76,8 +88,9 @@ order:
at the edge of the 10 ms gate; the measured v3 row of `docs/plans/mixer-x4.md` section 6 is what to scale from
now, and whether a 2019-class laptop core (unmeasured, O-1.14) passes 10 ms is what decides it.
2. The hot table: adopted or not on the PC rows (`docs/plans/hot-table.md`); in the added form it costs the GPU
- nothing it was not already paying in cache misses and the chip die area only, so it is the second lever for the
- chip only through area and the first against a DRAM-only chip.
+ 7 to 17 percent on the Mac and the chip die area only, so against this chip it is a lever in the wrong
+ direction and against a DRAM-only chip the first lever; if it is adopted, the mixer must carry the extra `1/g`
+ (x8 at g = 0.87 reads 1.06x at the equal budget, 0.84x with the SRAM deducted).
## 4. What this does not settle
diff --git a/docs/bench-log.md b/docs/bench-log.md
index 8a82a73b2..8550d873c 100644
--- a/docs/bench-log.md
+++ b/docs/bench-log.md
@@ -1645,3 +1645,22 @@ Rates (5 dispatches of 2^24 after a warm-up; 5090 `--bench --block-warps 1`, 907
| hot96k4a | 114.4 | 0.84 | 14.56 | 0.80 | 1 |
Reading: the probe promises a full hit rate on the 5090 (every S inside the 96 MiB L2 at one ceiling, 6.4x DRAM) and the hash gets 2 to 8% at k = 4 and 20% at k = 8; the 9070 XT the same shape. The dataset's random lines evict the table from the shared cache on every card. The added form costs 13 to 20% of the rate. Recommendation in `docs/plans/hot-table.md` section 6.4: do not adopt layer 5 in either form on these measurements.
+## 5 October 2026 (night), mixer x4 and the cache growth rule: the class v3 dataset construction, with the x8 candidate (Counter ASIC 2.0; branch ca2-mixer on ca2-v3 6c75dad; cryptographer's lane)
+
+Machine: Apple M5 Max, 64 GiB, Darwin 25.6.0. Write-up `docs/plans/mixer-x4.md`; chip model `docs/analysis/chip-model-v3.md`; code `igneum-pow` (LoadClass mixer_mult and growth, memhard::Shape, the schedule, the three emitters), packs `proto-cuda/packs-ca2-mixer/`, tests `igneum-pow/tests/mixer.rs` and `tests/packs.rs`. Commits 0fc0ad1, 66eeba3, e4c04a7, 7ce8d1e, 504cae4, fe4e193 and this entry's.
+
+What changed. Under program class v3 (`V3_CLASS = LoadClass::MX4`) every mixer application of the item derivation is `m = 4` applications with round keys `(r m + j + 1) x 0x9E3779B9`, the 8 dependent cache reads per item unchanged; the cache doubles when the dataset doubles (`growth_doublings(d) = floor(log2(1 + d / 1460))`: 2^26 words to day 1,459, 2^27 from day 1,460, 2^28 from day 4,380). Version 2 is byte-identical: fresh exports of igneum-genesis-mh and igneum-devnet-v4-epoch0 `diff -r` IDENTICAL against the checked-in packs, and the crate tests regenerate every pinned file. A v3 program of a seed is the v2 program of that seed instruction for instruction (v2 loads take no width roll); only the dataset words and the hashes change.
+
+Bit-exactness, `with-lock.sh run`, 22:05 and 21:45 UTC: the two pinned v3 packs (mx4-genesis, mx4-devnet-epoch0: dataset words 0..15 `61ff2180 0d4c7e6c ...` and `afe80d67 b9fbd029 ...`, word MASK `5020180e` and `e6a99c7a`, unit at base 0 lane 0 `63acd2d273f475ba` and `212c6442b51e87ae`) and the two x8 candidate packs on Metal (`packbench`, built from this branch) and Apple OpenCL (`igneum-bench-cl --bench-pack`): 3/3 standalone and 3/3 in batch, 96 of 96 lanes, cache FNV-1a 64 unchanged from v2 (48c4f5bf24166b2e, 448274a57f508cbc), dataset head, word MASK and 64 samples PASS, one 2^24 fingerprint per pack across both harnesses (mx4 6f48d5a2aa0dbe5f and 73caaebb28e808fe; mx8 7c28cfb06c5c65a9 and bbb183f72692f840); hash rate the v2 rate (27.5 to 27.7 MH/s GPU time, the hash kernel is unchanged). Fuzz: 200 class v3 programs (4 units each across the 32-bit range, one in the top 256 nonces) interpreted twice on the CPU, 800 of 800; the same 200 packs on Metal 200 of 200 (`--batch-log2 9 --batch-base 4294967040`, the wrapping unit inside the window), every tenth on Apple OpenCL 20 of 20; x8: 50 of 50 on Metal, 5 of 5 on OpenCL. Stats (8,192 outputs per seed, two seeds): v3 avalanche 49.97 to 49.99 percent, worst bit z 1.92 to 3.09, 0 duplicates (v2 beside it 49.87 to 49.98, z 2.25 to 2.30). Edges: items 0, 1, 2^28 - 1, 2^32 - 1 by hand at m = 1, 2, 4, 8; words 0, 15, 16, 17, MASK - 1, MASK through the fetch path. Determinism: two epochs, every vector and file equal and equal to the pinned pack. The scratch soundness tests of ca2-soundness (cherry-pick 0d8f745) 7 of 7 on this tree. Crate: 44 lib + 12 packs + 4 mixer + 7 scratch tests pass. A first Metal fuzz run reported 200 of 200 FAIL on an empty RESULT line (a packbench built before the `--batch-base` cherry-pick); it was read as a failure, the harness rebuilt, the run repeated.
+
+Timings, `with-lock.sh measure`, one session 21:40:12 to 21:40:23 UTC, one core, two rounds; the box carried a load average of 5.6 (one minute) and 26 (fifteen minutes) from unlocked processes, so the absolute figures are about 2.2x the quiet readwidth night's 0.604 ms v2 row and the ratios are the measurement:
+
+| Construction | Verifier ms per 32-lane unit, avg of 50 (two rounds) | Worst cold unit | Against v2 | 256 MiB fill, one core | Metal 1 GiB build, GPU ms |
+|---|---|---|---|---|---|
+| v2 | 1.361 / 1.310 | 1.579 | 1 | 172 to 173 ms | 29.7 (first touch) / 21.0 |
+| x4 (class v3) | 1.956 / 1.923 | 2.043 | 1.45x | 172 to 175 ms | 20.9 / 21.0 |
+| x8 (candidate) | 2.785 / 2.790 | 2.942 | 2.09x | 172 ms | 21.9 / 21.9 |
+
+Reading: the mixer multiplies the verifier's ALU part only (the 8 dependent misses per item are unchanged), hence 1.45x and 2.1x and not 4x and 8x; the Mac's GPU build is latency-bound and does not move with the mixer, so the "under 1 s on every discrete card" half of the x8 rule is the PC job (five packs, `relay/playbooks/mixer-x4-pc1-bench.ps1`, waiting for the go). Verification throughput (C19): a quiet 2026 core serves about 1,100 shares per second at x4 and 800 at x8 (1,660 at v2, re-cutting spec 09's 2,270), a 22,000-member pool at one share per 10 s needs 2 cores at x4 and 3 at x8, IBD over 108,000 headers is 1.6 min at x4 and 2.3 at x8 on that core; the 10 ms gate keeps 8.0 ms (x4) and 7.1 ms (x8) of margin on the loaded core, 6 to 7 ms on a 2019-class laptop core (approximate, unmeasured, O-1.14).
+
+Chip model (`docs/analysis/chip-model-v3.md`): the on-die-cache recompute chip at 50 T op/s against the 5090's measured 136.1 MH/s: v2 334 MH/s, 2.45x bare, 7.4x with the 3x fixed-function factor; x4 83.5 MH/s, 0.61x bare, 1.84x with the factor, 1.53x with the 128 mm^2 N5 mirror deducted at equal silicon; x8 41.7 MH/s, 0.31x, 0.92x, 0.76x. The claim at x4 is "under 2x" with the margin thin on the equal-budget convention (a 3.3x factor or a 10 percent larger budget reads 2.0x); the hot table in the added form would have raised it to 2.1x to 2.2x at the 5090's g (kept as measured, not adopted). Nothing here is a measurement of a chip.
diff --git a/docs/plans/mixer-x4.md b/docs/plans/mixer-x4.md
index 2635da058..ce91b34d5 100644
--- a/docs/plans/mixer-x4.md
+++ b/docs/plans/mixer-x4.md
@@ -125,8 +125,57 @@ node agent).
## 4. Vectors (class v3, `proto-cuda/packs-ca2-mixer/`)
-Filled in section 6 from the exported packs: `mx4-genesis` (seed `igneum-genesis`, day `2026-10-03`, 2^28 words,
-2^26-word cache) and `mx4-devnet-epoch0` (the devnet genesis hash as the epoch seed, day bytes of 2026-10-04).
+Produced by `igneum-pow export --program-class v3` (the Rust CPU interpreter, 5 October 2026, commit 66eeba3) and
+checked on the GPUs in section 6. The program of each pack is the version 2 program of the same seed instruction
+for instruction (`tests/packs.rs`, `v3_packs_are_the_v2_seeds_under_mixer_x4`); the cache is the version 2 cache
+(day 0 of the growth rule); the dataset words and the hashes are new. The 64 sampled indices are those of every
+pack (`emit::sample_indices`).
+
+Pack `mx4-genesis` (seed igneum-genesis, day 2026-10-03, generator 3, class mx4, program id e323b9dcaf283a6f, 2^28 words, 2^26-word cache, cache FNV-1a 64 `48c4f5bf24166b2e` as under v2):
+
+```
+dataset words 0..15 (item 0):
+ 61ff2180 0d4c7e6c 2177d443 60df9025 cf8b2e10 63675bfb 25289e58 9c45dc42
+ 2d271c54 9652369b 2dd77508 5921392c 3afa60ee c640ad68 f2bb56ff cfa46438
+dataset[0x0fffffff] = 5020180e
+sampled words (the first 8 of the 64 in vectors.json):
+ dataset[59471966] = de85726d dataset[217795994] = 7cfc31c7 dataset[208353206] = d3cd5289 dataset[42483309] = 6ebbeef1
+ dataset[172547758] = 9858413b dataset[148076330] = e786f141 dataset[183853158] = 64f13833 dataset[214389424] = ca229d04
+unit at base nonce 0, lanes 0..31:
+ 63acd2d273f475ba e929c78b34b80d4b 0b1011cb19982558 1457a0df5497aa11
+ 957d0f3bb71d98fb ac16901e6e6f6057 8ea1c6279f4b177a f28146e60bd08ba9
+ fe5b8cfe87f8e65b 49f87240566ace62 6ef6d6b7bdea8e41 46d9c0dc29a97b9c
+ 111fe30128db9398 66dc39084f0946d4 8ee11bdfd35fecf2 2861fcfc75db6677
+ 31c7667d4bde8556 c5989c48858b4ce0 276395e734a9d30d 84217b41e91368ff
+ 3604861e34d9f697 9f51d8ee16bf3639 c89e47bafa84401c 7ae78c1f10b70e19
+ 0b8c947157a29a48 d67192e8cfb43842 05a4c6d182c8c675 188e2661f3263f2e
+ a2df24238f7fea2e ed69ea7e13ad3a48 2d44ae509bab91b8 adad61931ea4fb70
+unit at base nonce 4096: lane 0 edd508ac57e5699a, lane 1 4eefd56d526cdaeb, lane 31 8892f8604733b1e0
+unit at base nonce 1000000: lane 0 8b3183778a49f59c, lane 1 1831b72a8797e895, lane 31 75eae55eba53a506
+```
+
+Pack `mx4-devnet-epoch0` (seed the devnet genesis hash, day bytes of 2026-10-04, generator 3, class mx4, program id 73bcbfe8ccf988f1, 2^28 words, 2^26-word cache, cache FNV-1a 64 `448274a57f508cbc` as under v2):
+
+```
+dataset words 0..15 (item 0):
+ afe80d67 b9fbd029 6c79f193 95139ad9 96310aff 4609f8b1 75279e63 28235be1
+ 47b17dcb 718e0ef2 a52588c8 a8bf49d5 19cf243e 5ec8905e a4851f66 af9cd9f3
+dataset[0x0fffffff] = e6a99c7a
+sampled words (the first 8 of the 64 in vectors.json):
+ dataset[59471966] = 57642b58 dataset[217795994] = c279badd dataset[208353206] = cbccbaad dataset[42483309] = 32cce392
+ dataset[172547758] = 71fdb4c6 dataset[148076330] = f5a268ce dataset[183853158] = 3ca1d676 dataset[214389424] = 7977b03d
+unit at base nonce 0, lanes 0..31:
+ 212c6442b51e87ae c374795c00839331 b6036a220a98f4b3 eb8b8013e637367b
+ db5866e9b73930fd f3f3d01f46e90333 9d913991ab8ed428 7ccb1d8fa100a800
+ 3cf45ba44f09a0fe 91acf48ef1a63082 6ea46c69fb082f99 581f0218977a9d72
+ 9a4623a5c62ddf2d ab6eb5e768f0feb4 07b70bdccf8aca12 d666311ae5e4311e
+ 53114757d669f0a4 bd5d6ace87ce2ce4 fb712015e8189192 a32cec81103e134b
+ 83f3d18c3289c124 fe29f1984b132b3d c9ffcf4e3774497a ac99c9243dc63809
+ d78a7e8217a32f3c ab81ad63d242fc31 0e4c30b7e00024af ce014289fff6778d
+ 64a1292e2a8b4a91 d5b8c90e681d7e3a 06078117673030fd 51bf77b280173930
+unit at base nonce 4096: lane 0 3c797978566b5950, lane 1 7c759e60185b6411, lane 31 96a903eb9a0ca390
+unit at base nonce 1000000: lane 0 d5a8da0568df8ee7, lane 1 68cb69c04208285a, lane 31 f8ca84a1a5d78cf5
+```
## 5. The v2 path is byte-identical
@@ -136,7 +185,136 @@ section 6 also records a fresh `igneum-pow export` of both packs diffed against
## 6. Measurements
-Filled as they land. Every row names the machine, the date, the command and the lock mode.
+Machine: Apple M5 Max, 64 GiB, Darwin 25.6.0, 5 October 2026 (night), other agents' builds running beside every
+run; a timing row says which lock it ran under (`measure` is exclusive; `run` and `build` are not timings).
+
+### 6.1 The v2 path, byte for byte (no lock needed)
+
+`igneum-pow export --seed igneum-genesis --day 2026-10-03` and `igneum-pow export --epoch-hex edc4fa84...fb07
+--day-hex 69676e65756d2d6461792ffa50000000000000` on commit 66eeba3, `diff -r` against
+`proto-cuda/packs/igneum-genesis-mh` and `igneum-devnet-v4-epoch0`: IDENTICAL, both (twelve files each). The
+crate tests regenerate the same files and compare them on every run (`tests/packs.rs`, 12 of 12 pass).
+
+### 6.2 Bit-exactness of the class v3 construction on the GPUs (`with-lock.sh run`, 22:05 UTC)
+
+`packbench --pack
--batches 1 --batch-log2 24 --group 256` (Metal, built from this branch) and
+`igneum-bench-cl-igneum-genesis-mh --bench-pack --pack --batches 1 --batch-log2 24` (Apple OpenCL, built from
+this branch's host.c). Vectors are the Rust interpreter's; the fingerprint is FNV-1a 64 over the 2^24 outputs at
+base nonce 0.
+
+| Pack | Harness | Cache FNV-1a 64 | Dataset head, word MASK, 64 samples | Vectors standalone / in batch | Fingerprint 2^24 | MH/s (GPU time; not a measurement, the run lock) |
+|---|---|---|---|---|---|---|
+| mx4-genesis | Metal | 48c4f5bf24166b2e PASS | PASS | 3/3, 3/3 | 6f48d5a2aa0dbe5f | 27.6 |
+| mx4-genesis | Apple OpenCL | PASS | PASS | 96 of 96 lanes | 6f48d5a2aa0dbe5f | 27.7 (wall) |
+| mx4-devnet-epoch0 | Metal | 448274a57f508cbc PASS | PASS | 3/3, 3/3 | 73caaebb28e808fe | 27.5 |
+| mx4-devnet-epoch0 | Apple OpenCL | PASS | PASS | 96 of 96 lanes | 73caaebb28e808fe | 27.6 (wall) |
+| mx8-genesis (the x8 candidate, 21:45 UTC) | Metal | 48c4f5bf24166b2e PASS | PASS | 3/3, 3/3 | 7c28cfb06c5c65a9 | 27.7 |
+| mx8-genesis | Apple OpenCL | PASS | PASS | 96 of 96 lanes | 7c28cfb06c5c65a9 | 27.6 (wall) |
+| mx8-devnet-epoch0 | Metal | 448274a57f508cbc PASS | PASS | 3/3, 3/3 | bbb183f72692f840 | 27.7 |
+| mx8-devnet-epoch0 | Apple OpenCL | PASS | PASS | 96 of 96 lanes | bbb183f72692f840 | 27.7 (wall) |
+
+Reading: the Rust interpreter, Metal and Apple OpenCL agree on the v3 dataset (head, MASK word, 64 samples), on
+every vector lane and on the 2^24-output fingerprint of each pack; the hash rate is the v2 rate (27.7 MH/s on this
+card tonight, readwidth table), as it must be: the hash kernel only loads, the mixer is paid in the build.
+
+Dataset build under the run lock (indicative only; the measured rows are in 6.4): Metal 30.2 ms GPU (mx4-genesis)
+and 21.7 ms GPU (mx4-devnet-epoch0) for 1 GiB; Apple OpenCL 55 and 56 ms wall. The readwidth entry's v2 figure on
+this card is 0.6 to 2.1 ms cache fill and a 1 GiB build the bench-log's memory-hard entry puts at 13 to 30 ms; the
+x4 build on the GPU is the row the measure lock will settle.
+
+### 6.3 Soundness suite on the v3 construction (commit 66eeba3 and after; `with-lock.sh build` for cargo, `run` for the GPU)
+
+| Suite | Command | Result |
+|---|---|---|
+| Crate lib tests (memhard schedule table, the by-hand multiplied mixer, the seam, the generator) | `cargo test -j4 --release` | 44 of 44 pass |
+| Pinned packs, v2 and v3 (`tests/packs.rs`: programs, ids, dataset words, 96 vectors per pack, every emitted file byte for byte, 16 masked loads per kernel, the v3 packs as the v2 seeds under mixer x4) | same | 12 of 12 pass |
+| Scratch soundness tests of branch ca2-soundness (cherry-pick 0d8f745, one conflict in packbench.swift's RESULT line resolved by hand, `era_bytes: None` added to the edge program literal) | `cargo test -j4 --release --test scratch` | 7 of 7 pass: bijections, re-hit rates, 56 of 56 edge units, 42 of 42 emitted scr kernels, 200 scratch programs and 800 units on the CPU |
+| v3 fuzz on the CPU (`tests/mixer.rs`): 200 programs through the seam, the contract on every instruction (each is the v2 program of its seed), 4 units each across the 32-bit range with one unit in the top 256 nonces, interpreted twice | `IGNEUM_MIXER_PACKS_OUT= cargo test -j4 --release --test mixer` | 200 of 200, 800 of 800 units; 200 packs written for the GPU runs (82 s with the three cache fills) |
+| v3 stats beside v2 (`tests/mixer.rs`): 8,192 outputs per seed, bit balance, single-bit avalanche within the unit and across units, duplicates | same | igneum-genesis v3: avalanche 49.99 percent, worst bit z 1.92, 0 duplicates (v2: 49.87, z 2.25); igneum-genesis/stats1 v3: 49.97, z 3.09 (v2: 49.98, z 2.30) |
+| v3 edge (`tests/mixer.rs`): items 0, 1, 2^28 - 1 and 2^32 - 1 by hand at m = 1, 2, 4, 8 on a 2^14-word cache; words 0, 15, 16, 17, MASK - 1, MASK through the interpreter's fetch path; the index wrap at MASK + 1 | same | pass |
+| v3 determinism (`tests/mixer.rs`): two independent epochs, every vector and every emitted file equal, and equal to the pinned pack | same | pass |
+| Metal and Apple OpenCL on the two pinned v3 packs | section 6.2 | 3/3 standalone, 3/3 in batch, 96 of 96 lanes, dataset head, MASK word and 64 samples, one fingerprint per pack across both harnesses |
+| Metal fuzz: the 200 packs, 4 units each standalone and the top-256 unit inside a 512-nonce batch at base 4,294,967,040; every tenth pack on Apple OpenCL as well | `packbench --pack --batches 1 --batch-log2 9 --batch-base 4294967040` (Metal), `igneum-bench-cl-igneum-genesis-mh --bench-pack --pack --batches 1 --batch-log2 10` (Apple OpenCL), `with-lock.sh run`, 21:22 to 21:24 UTC | Metal 200 of 200 packs PASS (800 of 800 standalone units, 200 of 200 inside the wrapping window, cache and dataset self-tests on every pack); Apple OpenCL 20 of 20 packs PASS (the three vectors.h units, the self-tests); a first run with a packbench built before the `--batch-base` cherry-pick reported 200 of 200 FAIL on an empty RESULT line and was read as such (the watcher rule), the harness rebuilt and the run repeated |
+
+### 6.4 Timings (`with-lock.sh measure`, one session, 21:40:12 to 21:40:23 UTC, commit 504cae4)
+
+Script `measure-v3.sh` (session scratchpad): `igneum-pow bench --seed igneum-genesis --day 2026-10-03 --warps 50`
+(v2), `... --program-class v3` (x4), `... --class mx8` (x8), two rounds each, then the devnet seeds, then
+`packbench --pack --batches 2 --batch-log2 22 --group 256` on igneum-genesis-mh, mx4-genesis and mx8-genesis,
+two rounds. The lock was exclusive among the agents' builds and measurements, but the box was not quiet: load
+average 5.6 (one minute) and 26 (fifteen minutes) at the start, from unlocked processes (the devnet node, other
+agents' editors); the v2 row reads 1.31 to 1.36 ms where the quiet readwidth night read 0.604 to 0.626. So the
+absolute numbers below are a loaded-core figure, about 2.2x the quiet one, and the ratios between the rows are the
+measurement (two rounds within 4 percent). A quiet-box re-run is owed (section 9).
+
+| Construction | Verifier, ms per 32-lane unit, avg of 50 (round 1 / round 2) | Worst cold unit of three | Against v2 | 256 MiB cache fill, one core | Metal 1 GiB dataset build, GPU ms (round 1 / round 2) |
+|---|---|---|---|---|---|
+| v2 (igneum-genesis) | 1.361 / 1.310 | 1.579 | 1 | 172.1 / 172.6 ms | 29.7 / 21.0 |
+| x4 (mx4, class v3) | 1.956 / 1.923 | 2.043 | 1.45x | 175.3 / 172.3 ms | 20.9 / 21.0 |
+| x8 (mx8) | 2.785 / 2.790 | 2.942 | 2.09x | 172.3 / 172.3 ms | 21.9 / 21.9 |
+| x4, the devnet seeds (mx4-devnet-epoch0) | 1.923 | 2.012 | | 173.9 ms | |
+| x8, the devnet seeds | 2.972 | 2.885 | | 173.6 ms | |
+
+Reading. The verifier's ALU part is what grows: x4 adds 0.6 ms per unit for 27 more mixer applications on each of
+4,096 items (110,592 applications, about 5.5 ns each on this core, the lanes' chains interleaved), x8 another
+0.85 ms for 36 more; the latency part (8 dependent misses per item) is the same in every row, which is why the
+measured ratios are 1.45x and 2.1x and not the 4x and 8x of the M16 table's scaling. Shape B (32 rounds of one
+read, section 1) would have multiplied the latency part too; x4 is under the 4.8 ms bar even on the loaded core, so
+B stays unimplemented. The cache fill does not depend on the mixer (it is the ChaCha chain): 172 to 175 ms, the
+spec's 175 to 181 ms of 1.8.3. The Metal 1 GiB build does not move with the mixer at all (21 ms at v2, x4 and x8
+once warm; the 29.7 ms first v2 run is the first-touch cost the hosts fill twice for): on this card the build is
+bound by the 8 dependent cache-line reads per item, not by the arithmetic, so the Mac says nothing about whether
+the 5090's or the 9070 XT's build is arithmetic-bound; that is the PC job (section 8).
+
+Against the x4 / x8 rule (section 6.5): the verifier half passes for x8 with 7.1 ms of the 10 ms gate to spare on
+this loaded core (worst cold 2.94 ms; the quiet-core figure would be about 1.3 ms, scaled by the 2.2x of the v2
+row, approximate); x4 leaves 8.0 ms. The build half waits on the PC rows.
+
+### 6.5 Verification throughput per tier (consequences row C19), from the loaded-core figures above
+
+Warps verified per second on one core = 1,000 / (ms per warp); a pool core verifying members' shares handles that
+many shares per second; a node verifies a block with one unit (plus the header path, under 0.1 ms, not measured
+here); IBD over the 108,000-header pruning window (spec 02) on one core = 108,000 x ms per warp.
+
+| Figure | v2 | x4 | x8 | Note |
+|---|---|---|---|---|
+| ms per warp, steady (this session, loaded core) | 1.33 | 1.94 | 2.79 | avg of the two rounds |
+| ms per warp, quiet M5 Max core (scaled by 0.604 / 1.33 = 0.45, approximate) | 0.60 | 0.88 | 1.26 | the readwidth night's v2 figure is measured; x4 and x8 scaled |
+| ms per warp, 2019-class laptop core (approximate: 2.5x the quiet M5 Max figure, the ratio the design document assumes for the gate; unmeasured, O-1.14) | 1.5 | 2.2 | 3.2 | the figure that fixes the gate is a measurement, not this row |
+| Shares per second per core (loaded / quiet, approximate) | 750 / 1,660 | 515 / 1,140 | 358 / 790 | spec 09 section 9.8 item 5 carried 2,270 at v2; re-cut from the quiet row: 1,660 |
+| Cores for a 22,000-member pool at one share per member per 10 s (2,200 shares per second), loaded / quiet | 2.9 / 1.3 | 4.3 / 1.9 | 6.1 / 2.8 | |
+| Node: worst cold single unit (loaded core) | 1.58 ms | 2.04 ms | 2.94 ms | per block |
+| IBD over 108,000 headers on one core, loaded / quiet, minutes | 2.4 / 1.1 | 3.5 / 1.6 | 5.0 / 2.3 | laptop (approximate): 2.7 / 4.0 / 5.8 min; a seed VM core (unmeasured) sits between the laptop and the quiet M5 Max |
+| Margin left under the 10 ms gate for Counter ASIC 3.0 (worst cold, loaded core) | 8.4 ms | 8.0 ms | 7.1 ms | on the 2019-class laptop row (approximate) 7.5 / 6.8 / 5.9 ms steady |
+
+Reading: at x4 a pool core serves about 1,100 shares per second on a quiet 2026 core (a 22,000-member pool needs
+two cores); at x8 about 800 (three cores). A node's block verification stays a few milliseconds. The gate's
+remaining margin is what Counter ASIC 3.0 has to spend, and on the unmeasured laptop core it is 6 to 7 ms at x4 and
+about 6 at x8, which is the number the 2019-class measurement (O-1.14) must confirm before x8 is final.
+
+### 6.5 The daily build per tier, and the x4 / x8 rule
+
+The coordinator's rule (21:30 UTC): x8 enters v3 if the per-warp verify stays under 10 ms on one Mac core AND the
+daily 1 GiB build stays under 1 s on every discrete card we own; else x4 with the thin margin stated and x8 named
+as the next lever. The integrated tier is decided beside it (consequences row C23): its build is per prepare, not
+per day, so its consequence is per-day dataset reuse in the workers or a restart per epoch.
+
+| Card | Build at x1 | x4 | x8 | Source |
+|---|---|---|---|---|
+| RTX 5090 (PC 2) | 13.4 ms (1 GiB) | owed (PC job, section 8) | owed | bench-log, 3 October, memory-hard entry |
+| RX 9070 XT (PC 1, eGPU) | owed | owed | owed | PC job, section 8 |
+| M5 Max, Metal | 13 to 30 ms (the two runs of the memory-hard entry; tonight's run-lock figures 21.7 to 30.2 ms at x4 and 22.0 to 30.0 at x8 say the Mac's build is latency-bound, not mixer-bound) | section 6.4 | section 6.4 | this file |
+| Radeon integrated gfx1036 (PC 2), OpenCL, per prepare | 6.9 / 9.4 / 11.7 s prepare total with the 1 GiB build inside | about 28 to 47 s (approximate: scaled x4; the iGPU's build is arithmetic-bound at x1 already) | about 55 to 94 s (approximate) | `docs/plans/epoch-length.md` section 7 (branch ca2-epoch), M11 table |
+| gfx1036 beside WSL build jobs (PC 1) | 55 / 116 / 124 s | about 4 to 8 min (approximate) | about 7 to 17 min (approximate) | same |
+| 8 GB-class discrete card (not owned; about a tenth of the 5090's rate, approximate) | about 0.13 s | about 0.5 s | about 1 s, on the edge of the rule | scaled from the 5090 row, approximate |
+
+Reading, before the measured rows: on the discrete cards the rule is about the 5090 and 9070 XT rows (owed, the PC
+job). The integrated tier misses the rule at x4 already: a per-prepare build of 28 to 47 s is a tenth to a quarter
+of the 600-DAA-second lead the devnet gives the next program (spec 1.12), and under load it is the whole lead; so
+if x4 or x8 goes in, the iGPU tier needs the workers to build the day's dataset once a day and keep it across
+epochs (today a prepare rebuilds it: `proto-opencl/host.c` prepareTask builds the pair's cache and dataset per
+prepare), or to restart per epoch. The node agent is asked whether the per-day reuse is bounded tonight
+(coordinator, 21:45 UTC); until then the iGPU consequence stands as written.
## 7. The chip model
@@ -144,8 +322,30 @@ Filled as they land. Every row names the machine, the date, the command and the
## 8. What is unverified
-Filled at the end.
+1. The 5090's and the 9070 XT's dataset build at x4 and x8 (the "under 1 s on every discrete card" half of the x8
+ rule): the PC 1 job (`relay/playbooks/mixer-x4-pc1-bench.ps1`, package `mixer-x4-pcjob.zip` sha256
+ ec3be97e...bdf87, five packs, the worker's `cache ... dataset ... ms` line per pack) waits for the coordinator's
+ go after the era job; by the M16 arithmetic the 5090 is 54 ms at x4 and 107 ms at x8 if its build is
+ arithmetic-bound and 13.4 ms if it is latency-bound like the Mac, both far under 1 s; the gfx1036 is the tier
+ that fails (section 6.5 of the build table), and its consequence (per-day dataset reuse in the workers) is with
+ the node agent.
+2. The absolute verifier figures were taken on a loaded core (load average 5.6); the ratios are the measurement
+ and the quiet-core figures are scaled. A 2019-class laptop core has not run any construction (O-1.14).
+3. The mixer has had no cryptanalysis (spec 1.8.4); `m` applications with distinct round keys is `m` times the
+ work only if no shortcut composes them, which is the same open question as for one application.
+4. The x8 packs are generator 2 with the class in the id (`--class mx8`); if x8 is chosen, the pinned v3 packs are
+ re-cut through the seam (`V3_CLASS = MX8`, generator 3) and the tests re-pinned, one commit.
+5. The 5090's rate for the v3 program is the v2 rate by construction (the hash kernel is unchanged, the Mac shows
+ 27.7 MH/s at v2, x4 and x8); the chip row's denominator stays the readwidth table's 136.1 MH/s until a v3 pack
+ runs on the card, which the PC job also gives.
## 9. Owed
-Filled at the end.
+| Item | Owner | When |
+|---|---|---|
+| PC 1 run of the five packs (the build-time rows, the v3 fingerprints on NVIDIA and AMD) | ca2-mixer, on the coordinator's go | after the era job, about 22:05 UTC |
+| Quiet-box re-run of the verifier session for absolute numbers | ca2-mixer | when the Mac is quiet |
+| The x4 / x8 choice recorded from the rule, then the vectors re-cut once through the seam | coordinator, then ca2-mixer | after the PC rows |
+| `Epoch::chain_dataset_day` wired to the genesis day index in the node (`days_since_genesis(day_index(header), day_index(genesis))`) and `pow_genesis_dataset_log2` in the override | ca2-node | the integration |
+| The spec text of section 2 into `docs/spec/01-lottery-hash.md` 1.8.5 and 1.13.3 (with the v3 vectors into 1.17) | the integration | after the choice |
+| The 2019-class laptop core measurement that fixes the gate (O-1.14) | cryptographer | gate 1 |
diff --git a/igneum-pow/src/generator.rs b/igneum-pow/src/generator.rs
index 0c22b7841..00a05124b 100644
--- a/igneum-pow/src/generator.rs
+++ b/igneum-pow/src/generator.rs
@@ -400,6 +400,12 @@ impl LoadClass {
self.era.map(|e| e.layout()).unwrap_or(crate::memhard::Layout::LINEAR)
}
+ /// The x8 candidate beside [`LoadClass::MX4`] (coordinator's rule of 5 October 2026, 21:30 UTC: x8 enters v3 if
+ /// the per-warp verify stays under 10 ms on one Mac core and the daily 1 GiB build under 1 s on every card):
+ /// the same loads and growth rule, the mixer applied 8 times per round. Name "mx8".
+ pub const MX8: LoadClass =
+ LoadClass { mix: [100, 0, 0], load_slots: LOAD_SLOTS as u8, scratch: None, scratch_kb: 0, mixer_mult: 8, growth: true };
+
/// A fixed width (1, 4 or 16 words) with `load_slots` loads per program.
pub fn fixed(width_words: u8, load_slots: u8) -> LoadClass {
let mut mix = [0u8; 3];
@@ -496,6 +502,9 @@ impl LoadClass {
if s == "mx4" {
return Some(LoadClass::MX4);
}
+ if s == "mx8" {
+ return Some(LoadClass::MX8);
+ }
// the mixer suffix: "...m" then an optional "g"
let (s, growth) = match s.strip_suffix('g') {
Some(base) if base.rsplit_once('m').map(|(_, d)| !d.is_empty() && d.bytes().all(|b| b.is_ascii_digit())).unwrap_or(false) => (base, true),
@@ -589,7 +598,7 @@ impl LoadClass {
*self == LoadClass::V2
}
- /// "v2", "w4", "w16", "w64", "w64x4", "mix50-35-15", "mix25-50-25x8", "scr4k32"; "mx4" for the v3 construction;
+ /// "v2", "w4", "w16", "w64", "w64x4", "mix50-35-15", "mix25-50-25x8", "scr4k32"; "mx4" for the v3 construction, "mx8" for its x8 candidate;
/// any other mixer setting appends "m" and, with the growth rule, "g" ("v2m2", "w16m4g").
/// An era class is the base name with "-era" appended ("w4-era401998a5", "mx4-era...").
/// A hot class appends "hotk[a]" ("hot64k4", "scr4k32+hot64k4a"; measured and not adopted).
@@ -615,6 +624,9 @@ impl LoadClass {
if *self == LoadClass::MX4 {
return "mx4".to_string();
}
+ if *self == LoadClass::MX8 {
+ return "mx8".to_string();
+ }
let loads = LoadClass { mixer_mult: 1, growth: false, ..*self };
let base = if loads.is_v2() {
"v2".to_string()
diff --git a/igneum-pow/src/verify.rs b/igneum-pow/src/verify.rs
index ba7338b7c..91f9ee31f 100644
--- a/igneum-pow/src/verify.rs
+++ b/igneum-pow/src/verify.rs
@@ -77,6 +77,22 @@ pub struct ScratchModel {
data: Vec<[u32; 3]>,
pub reads: usize,
pub writes: usize,
+ /// Soundness tests (`tests/scratch.rs`, `docs/analysis/scratch-soundness.md`): when `Some`, every
+ /// read-modify-write is appended as it happened. `None` on every verification path.
+ pub trace: Option>,
+}
+
+/// One scratch read-modify-write as the interpreter saw it (variant 5 soundness tests).
+#[derive(Clone, Copy, Debug, PartialEq, Eq)]
+pub struct ScratchEvent {
+ pub lane: u8,
+ pub slot: u32,
+ /// The slot had been written earlier in this unit (a re-hit): the words read were a rewrite, not the fill.
+ pub hit: bool,
+ pub read: [u32; 3],
+ /// The fold result, the new value of `dst`.
+ pub x: u32,
+ pub written: [u32; 3],
}
impl ScratchModel {
@@ -87,6 +103,7 @@ impl ScratchModel {
data: vec![[0; 3]; LANES * slots_per_lane],
reads: 0,
writes: 0,
+ trace: None,
}
}
/// Read slot `slot` of `lane`, then rewrite it from the fold result `x`. Returns the three words read.
@@ -103,7 +120,11 @@ impl ScratchModel {
]
};
let x = fold_words(dst, &w);
- self.data[i] = scratch_rewrite(x, &w);
+ let out = scratch_rewrite(x, &w);
+ if let Some(t) = self.trace.as_mut() {
+ t.push(ScratchEvent { lane: lane as u8, slot, hit: self.written[i], read: w, x, written: out });
+ }
+ self.data[i] = out;
self.written[i] = true;
self.reads += 1;
self.writes += 1;
@@ -305,6 +326,19 @@ pub fn interpret_warp(program: &Program, base_nonce: u32, ds: &DatasetSource) ->
/// [`interpret_warp`] with explicit init words `I` (section 1.6 of the spec). The packs use `I = program.seed`;
/// a block uses `I = bind::block_init_words(H, nonce)`.
pub fn interpret_warp_init(program: &Program, seed: &[u32; 8], base_nonce: u32, ds: &DatasetSource) -> WarpResult {
+ interpret_warp_scratch(program, seed, base_nonce, ds, false).0
+}
+
+/// [`interpret_warp_init`] that also returns every scratch read-modify-write of the unit in execution order
+/// (lane-minor within an instruction, as the interpreter runs them) when `trace` is set; empty otherwise and for
+/// a class without a scratch. For the soundness tests of variant 5 only.
+pub fn interpret_warp_scratch(
+ program: &Program,
+ seed: &[u32; 8],
+ base_nonce: u32,
+ ds: &DatasetSource,
+ trace: bool,
+) -> (WarpResult, Vec) {
let mask = ds.mask;
let log2 = ds.log2_words;
let era = program.class.era;
@@ -323,6 +357,11 @@ pub fn interpret_warp_init(program: &Program, seed: &[u32; 8], base_nonce: u32,
let mut idx = [0u32; LANES];
let mut val = [0u32; LANES];
let mut scratch = if program.has_scratch() { Some(ScratchModel::new(program.class.scratch_slots_per_lane())) } else { None };
+ if trace {
+ if let Some(m) = scratch.as_mut() {
+ m.trace = Some(Vec::new());
+ }
+ }
let slot_mask = program.class.scratch_slot_mask();
if program.has_hot() {
let h = ds.hot.as_ref().expect("a hot-table program needs the epoch's hot table on the dataset source");
@@ -348,7 +387,8 @@ pub fn interpret_warp_init(program: &Program, seed: &[u32; 8], base_nonce: u32,
let hi = r[4][lane] ^ r[5][lane].rotate_left(9) ^ r[6][lane].rotate_left(18) ^ r[7][lane].rotate_left(27);
hashes[lane] = ((hi as u64) << 32) | lo as u64;
}
- WarpResult { hashes, items_derived }
+ let events = scratch.and_then(|m| m.trace).unwrap_or_default();
+ (WarpResult { hashes, items_derived }, events)
}
#[inline(always)]
diff --git a/igneum-pow/tests/mixer.rs b/igneum-pow/tests/mixer.rs
new file mode 100644
index 000000000..78de4768f
--- /dev/null
+++ b/igneum-pow/tests/mixer.rs
@@ -0,0 +1,247 @@
+//! The class v3 dataset construction (mixer x4, cache growth option C; `docs/plans/mixer-x4.md`): the soundness
+//! runs the brief asks for, on the CPU, with the packs for the GPU runs written on request.
+//!
+//! 1. Fuzz: `IGNEUM_MIXER_FUZZ` (default 200) programs through the seam (`ProgramClass::V3`), the contract on every
+//! instruction (the v2 program of the seed, instruction for instruction), 4 units each across the 32-bit range
+//! including the wrap, interpreted twice on the CPU; with `IGNEUM_MIXER_PACKS_OUT=` every program is written
+//! as a pack with its 4 bases in vectors.json for `packbench` and the OpenCL host (the Metal fuzz).
+//! 2. Stats: bit balance and single-bit-flip avalanche of the v3 hash against v2 on the same programs and nonces.
+//! 3. Edge: the dataset at word 0, word MASK and the item boundary, derived through the interpreter's fetch path and
+//! by hand at every multiplier 1, 2, 4, 8, on a small cache.
+//! 4. Determinism: two independent epochs of the same seed and day agree on every vector and every emitted file.
+
+use igneum_pow::emit::{export_pack, vectors_json};
+use igneum_pow::generator::{generate_from_seed_bytes, generate_from_seed_bytes_class, generate_from_seed_bytes_program_class, LoadClass, Op, Program, ProgramClass, GENERATOR_VERSION_V3, INSTR_COUNT, V3_CLASS};
+use igneum_pow::memhard::{derive_item, mixer, round_key, Cache, MixParams, Shape};
+use igneum_pow::seed::{day_key, SplitMix64};
+use igneum_pow::verify::{DatasetMode, DatasetSource, Epoch};
+use std::collections::HashMap;
+use std::path::PathBuf;
+
+const DAY: &str = "2026-10-03";
+
+fn contract(p: &Program, seed: &str, class: LoadClass) {
+ if class == V3_CLASS {
+ assert_eq!(p.generator, GENERATOR_VERSION_V3);
+ }
+ assert_eq!(p.class, class);
+ assert_eq!(p.instrs.len(), INSTR_COUNT);
+ assert_eq!(p.instrs.iter().filter(|i| i.op == Op::Load).count(), 16);
+ for (k, i) in p.instrs.iter().enumerate() {
+ assert!(i.src != i.dst, "#{k}: src == dst");
+ assert!((1..=31).contains(&i.rot), "#{k}: rot {}", i.rot);
+ assert!([1u8, 2, 4, 8, 16].contains(&i.mask), "#{k}: mask {}", i.mask);
+ assert!(i.dst < 8 && i.src < 8 && i.src2 < 8);
+ assert_eq!(i.width, 1, "#{k}: a v3 load reads one word");
+ }
+ assert!(igneum_pow::accept::check(p).is_ok(), "an accepted program");
+ let v2 = generate_from_seed_bytes(seed, seed.as_bytes());
+ assert_eq!(p.instrs, v2.instrs, "the v2 program of the seed under the v3 construction");
+ assert_eq!(p.attempt, v2.attempt);
+}
+
+/// Write a pack whose vectors.json carries `bases` instead of the three standard bases (packbench and the OpenCL
+/// host check every unit standalone and the ones inside the batch window).
+fn write_pack_with_bases(dir: &PathBuf, e: &Epoch, day: &str, bases: &[u32], source: &str) {
+ let mut pack = export_pack(e, day, source);
+ let outs: Vec<[u64; 32]> = bases.iter().map(|&b| e.hash_warp(b)).collect();
+ let vj = vectors_json(&e.program, day, e.dataset.log2_words, bases, &outs, &pack.vectors, e.dataset.mask, source, true);
+ for f in pack.files.iter_mut() {
+ if f.0 == "vectors.json" {
+ f.1 = vj.clone();
+ }
+ }
+ pack.write_to(dir).unwrap();
+}
+
+#[test]
+fn fuzz_v3_programs_cpu() {
+ let n: usize = std::env::var("IGNEUM_MIXER_FUZZ").ok().and_then(|s| s.parse().ok()).unwrap_or(200);
+ let out = std::env::var("IGNEUM_MIXER_PACKS_OUT").ok().map(PathBuf::from);
+ // IGNEUM_MIXER_CLASS=mx8 fuzzes the x8 candidate as a load class (generator 2 with the class in the id); the
+ // default is V3_CLASS through the seam
+ let class = std::env::var("IGNEUM_MIXER_CLASS").ok().map(|s| LoadClass::parse(&s).expect("a load class")).unwrap_or(V3_CLASS);
+ let mut rng = SplitMix64::new(0x6967_6e65_756d_2d6d); // "igneum-m"
+ let shape = Shape::for_class(&class);
+ assert_eq!(shape.cache_log2_words, 26);
+ assert!(shape.mixer_mult > 1);
+ // one memory-hard source per dataset size (the 256 MiB cache fill is 0.2 s each)
+ let mut mh: HashMap = HashMap::new();
+ let mut manifest = String::from("pack\tlog2\tprogram_id\tbases\n");
+ let mut units = 0usize;
+ let mut wraps = 0usize;
+ for i in 0..n {
+ let seed = format!("igneum-mixer-fuzz/{i}");
+ let p = if class == V3_CLASS {
+ generate_from_seed_bytes_program_class(&seed, seed.as_bytes(), ProgramClass::V3, None)
+ } else {
+ generate_from_seed_bytes_class(&seed, seed.as_bytes(), class)
+ };
+ contract(&p, &seed, class);
+ let b0 = (rng.below(8) as u32) * 32;
+ let b1 = 0x8000_0000u32.wrapping_sub(256).wrapping_add((rng.below(16) as u32) * 32);
+ let b2 = 0xffff_ff00u32.wrapping_add((rng.below(8) as u32) * 32);
+ let b3 = (rng.next() as u32) & !31;
+ let bases = [b0, b1, b2, b3];
+ wraps += bases.iter().filter(|&&b| b >= 0xffff_ff00).count();
+ let log2 = [24u32, 26, 28][rng.below(3) as usize];
+ let ds = mh.remove(&log2).unwrap_or_else(|| DatasetSource::new_shape(DAY, DatasetMode::MemoryHard, log2, shape));
+ let e = Epoch { program: p, dataset: ds };
+ for &b in &bases {
+ let r1 = e.interpret_warp(b);
+ let r2 = e.interpret_warp(b);
+ assert_eq!(r1.hashes, r2.hashes);
+ assert!(r1.items_derived >= 120 * 32 / 32 && r1.items_derived <= 4_096, "{seed}: {} items", r1.items_derived);
+ units += 1;
+ }
+ if let Some(dir) = &out {
+ let pack_name = format!("fuzz-{i:03}-{}-l{log2}", class.name());
+ write_pack_with_bases(&dir.join(&pack_name), &e, DAY, &bases, "igneum-pow tests/mixer.rs fuzz");
+ manifest.push_str(&format!(
+ "{pack_name}\t{log2}\t{:016x}\t{}\n",
+ e.program.program_id(),
+ bases.iter().map(|b| format!("{b}")).collect::>().join(",")
+ ));
+ }
+ mh.insert(log2, e.dataset);
+ }
+ println!("fuzz: {n} {} programs, {units} units on the CPU, {wraps} units in the top 256 nonces", class.name());
+ assert_eq!(units, 4 * n);
+ assert_eq!(wraps, n);
+ if let Some(dir) = &out {
+ std::fs::create_dir_all(dir).unwrap();
+ std::fs::write(dir.join("manifest.tsv"), manifest).unwrap();
+ println!("packs written to {}", dir.display());
+ }
+}
+
+/// Bit balance and avalanche of the v3 hash beside v2 on the same program (the TESTS.md section 3 shape, on the
+/// CPU, 2^13 nonces per seed): every output bit within 5 sigma of half ones; a single nonce-bit flip moves 50 percent
+/// of the output bits within 2 points; no duplicate among the outputs.
+#[test]
+fn stats_v3_against_v2() {
+ let n_warps = 256usize; // 8,192 nonces
+ for seed in ["igneum-genesis", "igneum-genesis/stats1"] {
+ let v3 = Epoch {
+ program: generate_from_seed_bytes_program_class(seed, seed.as_bytes(), ProgramClass::V3, None),
+ dataset: DatasetSource::new_shape(DAY, DatasetMode::MemoryHard, 24, Shape::for_class(&V3_CLASS)),
+ };
+ let v2 = Epoch::new(seed, DAY, DatasetMode::MemoryHard, 24);
+ for (name, e) in [("v3", &v3), ("v2", &v2)] {
+ let mut ones = [0u64; 64];
+ let mut outs = Vec::with_capacity(n_warps * 32);
+ for w in 0..n_warps {
+ let h = e.hash_warp(w as u32 * 32);
+ for &x in &h {
+ outs.push(x);
+ for b in 0..64 {
+ ones[b] += (x >> b) & 1;
+ }
+ }
+ }
+ let total = (n_warps * 32) as f64;
+ let sigma = (total / 4.0).sqrt();
+ for (b, &c) in ones.iter().enumerate() {
+ let z = (c as f64 - total / 2.0).abs() / sigma;
+ assert!(z < 5.0, "{seed} {name}: bit {b} ones {c} of {total}, z {z:.2}");
+ }
+ // avalanche: flip one bit of the nonce within the unit (lanes 0..31 differ in the low 5 bits) and across
+ // units (bit 5 and up): compare lane l of unit u with lane l ^ (1 << k) and with unit u ^ (1 << k)
+ let mut flips = 0u64;
+ let mut moved = 0u64;
+ for w in 0..64usize {
+ let h = e.hash_warp(w as u32 * 32);
+ for k in 0..5 {
+ for l in 0..32usize {
+ moved += (h[l] ^ h[l ^ (1 << k)]).count_ones() as u64;
+ flips += 1;
+ }
+ }
+ let h2 = e.hash_warp((w ^ 1) as u32 * 32);
+ for l in 0..32usize {
+ moved += (h[l] ^ h2[l]).count_ones() as u64;
+ flips += 1;
+ }
+ }
+ let avg = moved as f64 / flips as f64 / 64.0 * 100.0;
+ assert!((avg - 50.0).abs() < 2.0, "{seed} {name}: avalanche {avg:.2} percent");
+ outs.sort_unstable();
+ let dups = outs.windows(2).filter(|p| p[0] == p[1]).count();
+ assert_eq!(dups, 0, "{seed} {name}: duplicate outputs");
+ println!("{seed} {name}: {} outputs, avalanche {avg:.2} percent, worst bit z {:.2}", outs.len(), ones.iter().map(|&c| (c as f64 - total / 2.0).abs() / sigma).fold(0.0, f64::max));
+ }
+ assert_ne!(v3.hash_warp(0), v2.hash_warp(0));
+ }
+}
+
+/// The dataset edges under every multiplier on a small cache: word 0, word MASK, the last word of item 0 and the
+/// first of item 1, through `DatasetSource::word` and by hand.
+#[test]
+fn edge_items_every_multiplier() {
+ let key = day_key(DAY);
+ let cache = Cache::fill_log2(key, 14);
+ for m in [1u32, 2, 4, 8] {
+ let mp = MixParams::with_shape(key, Shape { mixer_mult: m, cache_log2_words: 14 });
+ let by_hand = |t: u32| -> [u32; 16] {
+ let mut s = [0u32; 16];
+ s[..8].copy_from_slice(&key);
+ for i in 0..8 {
+ s[8 + i] = t.wrapping_mul(mp.mul[i]).wrapping_add(mp.rc[i]);
+ }
+ for r in 0..8usize {
+ for j in 0..m as usize {
+ mixer(&mut s, round_key(r * m as usize + j), &mp);
+ }
+ let line = cache.line(s[0]);
+ for i in 0..16 {
+ s[i] ^= line[i];
+ }
+ }
+ for j in 0..m as usize {
+ mixer(&mut s, round_key(8 * m as usize + j), &mp);
+ }
+ s
+ };
+ for t in [0u32, 1, 0x0fff_ffff, 0xffff_ffff] {
+ assert_eq!(derive_item(t, &mp, &cache), by_hand(t), "m {m} item {t}");
+ }
+ }
+ // the interpreter's fetch path at the genesis cache: words 0, 15, 16 and MASK of a 2^20-word dataset agree with
+ // the item derivation, under v3
+ let ds = DatasetSource::new_shape(DAY, DatasetMode::MemoryHard, 20, Shape::for_class(&V3_CLASS));
+ let m = ds.memhard().unwrap();
+ for w in [0u32, 15, 16, 17, ds.mask - 1, ds.mask] {
+ assert_eq!(ds.word(w), derive_item(w >> 4, &m.params, &m.cache)[(w & 15) as usize]);
+ assert_eq!(ds.word(w), m.word(w));
+ }
+ // a load at an out-of-range register masks to the dataset: the word at mask + 1 is the word at 0
+ assert_eq!(ds.word(ds.mask.wrapping_add(1)), ds.word(0));
+}
+
+/// Two independent epochs of the same seed and day: every vector and every emitted file identical; the pinned v3
+/// pack is what a third export writes.
+#[test]
+fn determinism_v3() {
+ let build = || Epoch {
+ program: generate_from_seed_bytes_program_class("igneum-genesis", b"igneum-genesis", ProgramClass::V3, None),
+ dataset: DatasetSource::new_shape(DAY, DatasetMode::MemoryHard, 28, Shape::for_class(&V3_CLASS)),
+ };
+ let a = build();
+ let b = build();
+ let pa = export_pack(&a, DAY, "a");
+ let pb = export_pack(&b, DAY, "a");
+ assert_eq!(pa.outs, pb.outs);
+ assert_eq!(pa.vectors, pb.vectors);
+ assert_eq!(pa.files, pb.files);
+ for (w, warp) in [(0u32, 0usize), (4096, 1), (1_000_000, 2)] {
+ assert_eq!(a.hash_warp(w), pa.outs[warp]);
+ }
+ let dir = PathBuf::from(env!("CARGO_MANIFEST_DIR")).join("../proto-cuda/packs-ca2-mixer/mx4-genesis");
+ for (name, text) in &pa.files {
+ if name == "vectors.json" || name == "vectors.h" {
+ continue; // the source string differs ("a" here)
+ }
+ let on_disk = std::fs::read_to_string(dir.join(name)).unwrap();
+ assert_eq!(&on_disk, text, "{name}");
+ }
+}
diff --git a/igneum-pow/tests/scratch.rs b/igneum-pow/tests/scratch.rs
new file mode 100644
index 000000000..75e48f583
--- /dev/null
+++ b/igneum-pow/tests/scratch.rs
@@ -0,0 +1,766 @@
+//! Soundness tests of layer 3 of `docs/plans/counter-asic-2.md`: the per-warp scratch with read-modify-writes
+//! (variant 5 of the read-width experiment, `LoadClass::scratch(k, kb)`). Analysis and results:
+//! `docs/analysis/scratch-soundness.md`. Every test is parametric over the class's slot count
+//! (`scratch_slots_per_lane()`), so the 32 and 128 KiB geometries and any later one run the same checks.
+//!
+//! What runs under plain `cargo test`:
+//! 1. `rewrite_is_a_bijection_of_the_fold_value`, `fill_is_a_bijection_of_the_nonce`: the written words as
+//! functions (question 1).
+//! 2. `written_words_unbiased_and_rehit_rates`: bit bias of every written word over 2^11 units x 3 seeds per class
+//! (the TESTS.md section 3 shape), and the measured slot re-hit rate against the birthday formula (question 2).
+//! 3. `edge_programs_match_the_hand_model`: hand-built programs that drive every read-modify-write of a hash to
+//! slot 0, slot MASK, through out-of-range registers, to one slot per lane, alternating two slots, and 16
+//! read-modify-writes per iteration on one slot; the interpreter against an independent hand model, and the
+//! hand model shown to have teeth (question 3, CPU half).
+//! 4. `scr_packs_regenerate_and_pass_the_static_scratch_check`: every emitted kernel of every scr pack under
+//! `proto-cuda/packs-readwidth` regenerates from its program.json and passes the static scratch-mask check;
+//! the check is shown to fail on four deliberate breaks (question 4).
+//! 5. `fuzz_scr_programs_cpu`: 200 generated scratch programs over the six classes, generator contract on every
+//! instruction, 4 units each at base nonces across the 32-bit range including the wrap; with
+//! `IGNEUM_SCRATCH_PACKS_OUT=` it also writes the packs (and the edge packs) for the Metal runs of
+//! `proto-metal/packbench` (question 3 GPU half, question 4, `TESTS.md` section 9 shape).
+
+use igneum_pow::emit::{
+ cuda_kernel, cuda_kernel_bound, export_pack, metal_program, metal_program_bound, opencl_kernel,
+ opencl_kernel_bound, vectors_json, LoadSource,
+};
+use igneum_pow::generator::{
+ generate_class, generate_from_seed_bytes_class, Instr, LoadClass, Op, Program, GENERATOR_VERSION, INSTR_COUNT,
+ ITERATIONS, LANES,
+};
+use igneum_pow::seed::{seed_words_from_bytes, SplitMix64};
+use igneum_pow::verify::{
+ fold_words, interpret_warp_scratch, scratch_fill, scratch_rewrite, splitmix32, DatasetMode, DatasetSource,
+ Epoch, ScratchEvent, FOLD_MUL, FOLD_ROT,
+};
+use serde_json::Value;
+use std::collections::HashMap;
+use std::path::PathBuf;
+
+/// The classes under study: the two capped geometries (32 and 128 KiB per warp: 64 and 256 slots per lane) at the
+/// RMW shares the readwidth branch measures.
+const CLASSES: [&str; 6] = ["scr2k32", "scr4k32", "scr8k32", "scr2k128", "scr4k128", "scr8k128"];
+
+fn class(name: &str) -> LoadClass {
+ LoadClass::parse(name).unwrap_or_else(|| panic!("class {name}"))
+}
+
+// ---------------------------------------------------------------------------------------------------------------
+// 1. The written words as functions (question 1)
+// ---------------------------------------------------------------------------------------------------------------
+
+/// For a fixed slot content `w`, each of the three rewritten words is a bijection of the fold value `x`
+/// (`x ^ w1`, `rotl(x, 7) ^ w2`, `x + w0`), so the rewrite is injective in `x` and a uniform `x` gives a uniform
+/// word in every position. Checked over 2^16 consecutive `x` for 16 random `w`.
+#[test]
+fn rewrite_is_a_bijection_of_the_fold_value() {
+ let mut rng = SplitMix64::new(0x7363_7261_7463_6801);
+ for _ in 0..16 {
+ let w = [rng.next() as u32, rng.next() as u32, rng.next() as u32];
+ let x0 = rng.next() as u32;
+ let mut seen = [vec![false; 1 << 16], vec![false; 1 << 16], vec![false; 1 << 16]];
+ for i in 0..(1u32 << 16) {
+ let x = x0.wrapping_add(i);
+ let out = scratch_rewrite(x, &w);
+ for j in 0..3 {
+ // a bijection of x maps 2^16 consecutive x to 2^16 distinct words; the low 16 bits alone are
+ // distinct for the xor words (x ^ c) and for the add word (x + c), since both act on the low 16
+ // bits as bijections of the low 16 bits of x; the rotl word is checked on its rotated-back bits
+ let key = if j == 1 { out[j].rotate_right(7) & 0xffff } else { out[j] & 0xffff };
+ assert!(!seen[j][key as usize], "word {j} repeats inside 2^16 consecutive x");
+ seen[j][key as usize] = true;
+ }
+ }
+ }
+ // The rewrite inverts: from the old content and any ONE written word the fold value is recovered, so a
+ // rewritten slot carries exactly 32 bits of new state (the point of question 2's arithmetic).
+ let w = [0x1234_5678, 0x9abc_def0, 0x0fed_cba9];
+ let x = 0xdead_beef;
+ let out = scratch_rewrite(x, &w);
+ assert_eq!(out[0] ^ w[1], x);
+ assert_eq!((out[1] ^ w[2]).rotate_right(7), x);
+ assert_eq!(out[2].wrapping_sub(w[0]), x);
+}
+
+/// For a fixed (seed, slot, j) the fill is a bijection of the lane nonce: `splitmix32` is a bijection of its
+/// 32-bit input and the input `((base + lane) ^ s) + c` is a bijection of `base + lane`. Over 2^16 consecutive
+/// nonces no fill word repeats, for 8 slots x 3 words.
+#[test]
+fn fill_is_a_bijection_of_the_nonce() {
+ let seed = seed_words_from_bytes(b"igneum-genesis");
+ for slot in [0u32, 1, 63, 64, 255, 1023, 2047] {
+ for j in 0..3u32 {
+ let mut words: Vec = (0..(1u32 << 16)).map(|n| scratch_fill(&seed, n, 0, slot, j)).collect();
+ words.sort_unstable();
+ words.dedup();
+ assert_eq!(words.len(), 1 << 16, "slot {slot} word {j}: fill words of 2^16 consecutive nonces are distinct");
+ }
+ }
+ // base + lane is the lane nonce: the fill of lane l at base b is the fill of lane 0 at base b + l
+ assert_eq!(scratch_fill(&seed, 0x1000, 7, 5, 2), scratch_fill(&seed, 0x1007, 0, 5, 2));
+ // and it wraps with the nonce: base 0xffffffe0, lane 31 is nonce 0xffffffff; lane 32 would be nonce 0
+ assert_eq!(scratch_fill(&seed, 0xffff_ffe0, 32, 5, 2), scratch_fill(&seed, 0, 0, 5, 2));
+ // the three word positions of one slot and nonce are three different permutation outputs
+ let f: Vec = (0..3).map(|j| scratch_fill(&seed, 12345, 7, 17, j)).collect();
+ assert!(f[0] != f[1] && f[1] != f[2] && f[0] != f[2]);
+}
+
+// ---------------------------------------------------------------------------------------------------------------
+// 2. Uniformity of the written words and the slot re-hit rate (questions 1 and 2)
+// ---------------------------------------------------------------------------------------------------------------
+
+/// Birthday arithmetic: the expected number of distinct slots after `n` uniform draws from `s` slots.
+fn expected_distinct(s: usize, n: usize) -> f64 {
+ let s = s as f64;
+ s * (1.0 - (1.0 - 1.0 / s).powi(n as i32))
+}
+
+struct ClassStats {
+ units: usize,
+ events: usize,
+ hits: usize,
+ /// ones count per bit of the written words, 3 x 32
+ ones: [[u64; 32]; 3],
+ /// ones count per bit of written XOR read (the change the rewrite makes to the slot)
+ delta_ones: [[u64; 32]; 3],
+ /// re-hit depth histogram: how many earlier RMWs the slot had seen in this unit (0 = first touch)
+ depth: Vec,
+ max_depth: usize,
+ /// how often each slot index was addressed (the slot comes from a register's low bits)
+ slot_hist: Vec,
+}
+
+fn class_stats(name: &str, seeds: &[&str], units_per_seed: usize) -> ClassStats {
+ let c = class(name);
+ let mut st = ClassStats {
+ units: 0,
+ events: 0,
+ hits: 0,
+ ones: [[0; 32]; 3],
+ delta_ones: [[0; 32]; 3],
+ depth: vec![0; 256],
+ max_depth: 0,
+ slot_hist: vec![0; c.scratch_slots_per_lane()],
+ };
+ let ds = DatasetSource::new("2026-10-03", DatasetMode::ClosedForm, 28);
+ for seed in seeds {
+ let p = generate_class(seed, c);
+ assert_eq!(p.scratch_ops_per_hash(), c.scratch_slots() * ITERATIONS);
+ for u in 0..units_per_seed {
+ let base = (u as u32).wrapping_mul(32).wrapping_add(0x4000_0000);
+ let (_, ev) = interpret_warp_scratch(&p, &p.seed, base, &ds, true);
+ assert_eq!(ev.len(), p.scratch_ops_per_hash() * LANES);
+ let mut count: HashMap<(u8, u32), usize> = HashMap::new();
+ for e in &ev {
+ assert!(e.slot < c.scratch_slots_per_lane() as u32, "slot inside the lane's scratch");
+ let d = count.entry((e.lane, e.slot)).or_insert(0);
+ assert_eq!(e.hit, *d > 0, "hit flag agrees with the unit's own history");
+ assert_eq!(e.written, scratch_rewrite(e.x, &e.read));
+ if !e.hit {
+ let fill = [
+ scratch_fill(&p.seed, base, e.lane as u32, e.slot, 0),
+ scratch_fill(&p.seed, base, e.lane as u32, e.slot, 1),
+ scratch_fill(&p.seed, base, e.lane as u32, e.slot, 2),
+ ];
+ assert_eq!(e.read, fill, "a first touch reads the fill");
+ }
+ st.depth[(*d).min(255)] += 1;
+ st.max_depth = st.max_depth.max(*d);
+ st.slot_hist[e.slot as usize] += 1;
+ *d += 1;
+ st.events += 1;
+ st.hits += e.hit as usize;
+ for j in 0..3 {
+ for b in 0..32 {
+ st.ones[j][b] += ((e.written[j] >> b) & 1) as u64;
+ st.delta_ones[j][b] += (((e.written[j] ^ e.read[j]) >> b) & 1) as u64;
+ }
+ }
+ }
+ st.units += 1;
+ }
+ }
+ st
+}
+
+/// Bit bias of every written word (and of the change each rewrite makes) within 6 sigma of a fair coin, over
+/// 3 seeds x 2^11 units per class (131,072 hashes per seed set); the slot re-hit rate against the birthday
+/// formula within 3 percent relative. The table printed here is the one in the analysis.
+#[test]
+fn written_words_unbiased_and_rehit_rates() {
+ let seeds = ["igneum-genesis", "igneum-genesis/stats1", "igneum-genesis/stats2"];
+ let units = 1usize << 11;
+ println!("class | slots/lane | RMW/hash | events | re-hits | re-hit % | birthday % | slot chi2 z (spread) | max depth | max bias sigma | max delta bias sigma");
+ for name in CLASSES {
+ let c = class(name);
+ let st = class_stats(name, &seeds, units);
+ let n = st.events as f64;
+ let sigma = (n / 4.0).sqrt();
+ let mut worst = 0.0f64;
+ let mut worst_delta = 0.0f64;
+ for j in 0..3 {
+ for b in 0..32 {
+ let z = (st.ones[j][b] as f64 - n / 2.0).abs() / sigma;
+ let zd = (st.delta_ones[j][b] as f64 - n / 2.0).abs() / sigma;
+ assert!(z <= 6.0, "{name}: written word {j} bit {b} biased: {z:.2} sigma");
+ assert!(zd <= 6.0, "{name}: rewrite delta word {j} bit {b} biased: {zd:.2} sigma");
+ worst = worst.max(z);
+ worst_delta = worst_delta.max(zd);
+ }
+ }
+ let per_lane_hash = c.scratch_slots() * ITERATIONS;
+ let s = c.scratch_slots_per_lane();
+ let exp_hits = per_lane_hash as f64 - expected_distinct(s, per_lane_hash);
+ let exp_pct = 100.0 * exp_hits / per_lane_hash as f64;
+ let got_pct = 100.0 * st.hits as f64 / st.events as f64;
+ // chi-square of the slot histogram against uniform (df = s - 1): the slot is a register's low bits, and
+ // the measured re-hit rate runs above the uniform birthday rate (the finding of the analysis, question 2)
+ let expect_per_slot = n / s as f64;
+ let chi2: f64 = st.slot_hist.iter().map(|&h| (h as f64 - expect_per_slot).powi(2) / expect_per_slot).sum();
+ let chi2_z = (chi2 - (s as f64 - 1.0)) / (2.0 * (s as f64 - 1.0)).sqrt();
+ let hot = *st.slot_hist.iter().max().unwrap() as f64 / expect_per_slot;
+ let cold = *st.slot_hist.iter().min().unwrap() as f64 / expect_per_slot;
+ println!(
+ "{name} | {s} | {per_lane_hash} | {} | {} | {got_pct:.2} | {exp_pct:.2} | {chi2_z:.1} (hottest slot {hot:.2}x, coldest {cold:.2}x) | {} | {worst:.2} | {worst_delta:.2}",
+ st.events, st.hits, st.max_depth
+ );
+ // a regression band, not a uniformity claim: the rate sits between the uniform birthday rate and twice it
+ assert!(
+ got_pct >= 0.9 * exp_pct && got_pct <= 2.0 * exp_pct,
+ "{name}: re-hit rate {got_pct:.2}% against birthday {exp_pct:.2}%"
+ );
+ // depth histogram: the number of earlier RMWs a re-hit slot had seen in the unit
+ let shown: Vec = st.depth.iter().take(st.max_depth + 1).enumerate().map(|(d, n)| format!("{d}:{n}")).collect();
+ println!(" depth histogram {}", shown.join(" "));
+ }
+}
+
+// ---------------------------------------------------------------------------------------------------------------
+// 3. Hand-built edge programs against an independent hand model (question 3, CPU half)
+// ---------------------------------------------------------------------------------------------------------------
+
+fn ins(op: Op, dst: u8, src: u8) -> Instr {
+ Instr { op, dst, src, src2: 0, imm: 0, imm2: 0, rot: 1, bit: 0, mask: 1, width: 1 }
+}
+fn add_imm(dst: u8, src: u8, imm: u32) -> Instr {
+ Instr { op: Op::Add, dst, src, src2: 0, imm, imm2: imm, rot: 1, bit: 0, mask: 1, width: 1 }
+}
+
+/// A hand-built program of class `c` named `name` (its seed is the name, so its fill words and init words are
+/// its own). These bypass the generator and the acceptance rule, like `TESTS.md` section 2; `sub r, r` zeroes a
+/// register as the Swift edge set does.
+fn edge(name: &str, c: LoadClass, instrs: Vec) -> Program {
+ let seed_string = format!("igneum-scratch-edge/{name}");
+ let seed_bytes = seed_string.as_bytes().to_vec();
+ let k = instrs.iter().filter(|i| i.op == Op::Scratch).count();
+ assert_eq!(k, c.scratch_slots(), "{name}: the class carries the program's scratch count");
+ Program {
+ seed: seed_words_from_bytes(&seed_bytes),
+ seed_string,
+ seed_bytes,
+ generator: GENERATOR_VERSION,
+ attempt: 0,
+ class: c,
+ era_bytes: None,
+ instrs,
+ }
+}
+
+/// The edge set for a scratch of `kb` KiB per warp. Each entry: (name, what it drives, program).
+fn edge_programs(kb: u8) -> Vec<(String, &'static str, Program)> {
+ let m = LoadClass::scratch(1, kb).scratch_slot_mask();
+ let dsts = [2u8, 3, 4, 5, 6, 7, 0, 2, 3, 4, 5, 6, 7, 0, 2, 3];
+ let scr = |n: usize, src: u8| -> Vec { (0..n).map(|i| ins(Op::Scratch, dsts[i], src)).collect() };
+ let mut v = Vec::new();
+ // every RMW of the hash to slot 0 through a zero register: 64 dependent RMWs on one slot per lane
+ let mut p = vec![ins(Op::Sub, 1, 1)];
+ p.extend(scr(8, 1));
+ v.push(("slot0".to_string(), "r1 = 0: every RMW to slot 0", edge(&format!("slot0/k{kb}"), LoadClass::scratch(8, kb), p)));
+ // slot MASK through the in-range register MASK
+ let mut p = vec![ins(Op::Sub, 1, 1), ins(Op::Sub, 2, 2), add_imm(1, 2, m)];
+ p.extend(scr(8, 1));
+ v.push(("slotmask".to_string(), "r1 = MASK: every RMW to the last slot", edge(&format!("slotmask/k{kb}"), LoadClass::scratch(8, kb), p)));
+ // slot MASK through the out-of-range register 0xffffffff
+ let mut p = vec![ins(Op::Sub, 1, 1), ins(Op::Sub, 2, 2), add_imm(2, 1, 1), ins(Op::Sub, 1, 2)];
+ p.extend(scr(8, 1));
+ v.push(("ones".to_string(), "r1 = 0xffffffff: masked to the last slot", edge(&format!("ones/k{kb}"), LoadClass::scratch(8, kb), p)));
+ // slot 0 through the out-of-range register MASK + 1
+ let mut p = vec![ins(Op::Sub, 1, 1), ins(Op::Sub, 2, 2), add_imm(1, 2, m.wrapping_add(1))];
+ p.extend(scr(8, 1));
+ v.push(("maskplus1".to_string(), "r1 = MASK + 1: masked to slot 0", edge(&format!("maskplus1/k{kb}"), LoadClass::scratch(8, kb), p)));
+ // 16 RMWs per iteration on slot 0: 128 dependent RMWs on one slot per lane per hash
+ let mut p = vec![ins(Op::Sub, 1, 1)];
+ p.extend(scr(16, 1));
+ v.push(("sixteen".to_string(), "16 RMWs per iteration on slot 0", edge(&format!("sixteen/k{kb}"), LoadClass::scratch(16, kb), p)));
+ // one slot per lane from the init words: lanes with equal slots would show any cross-lane aliasing
+ // (r5 is the slot register and is never a destination here)
+ let p: Vec = [0u8, 1, 2, 3, 4, 6, 7, 0].iter().map(|&d| ins(Op::Scratch, d, 5)).collect();
+ v.push(("lanevar".to_string(), "r5 never written: one init-dependent slot per lane", edge(&format!("lanevar/k{kb}"), LoadClass::scratch(8, kb), p)));
+ // alternating slot 0 and slot MASK inside one iteration
+ let mut p = vec![ins(Op::Sub, 1, 1), ins(Op::Sub, 2, 2), add_imm(2, 1, m)];
+ for (i, &d) in [3u8, 4, 5, 6, 7, 0, 3, 4].iter().enumerate() {
+ // r1 and r2 hold the two slots and are never destinations
+ p.push(ins(Op::Scratch, d, if i % 2 == 0 { 1 } else { 2 }));
+ }
+ v.push(("twoslots".to_string(), "slot 0 and slot MASK alternating", edge(&format!("twoslots/k{kb}"), LoadClass::scratch(8, kb), p)));
+ v
+}
+
+/// The hand model: a second, minimal interpreter for the ops the edge programs use (sub, add, scratch), with its
+/// own slot store keyed by (lane, slot). `mutate` swaps the rewrite's words to show the comparison has teeth.
+fn hand_model(p: &Program, base: u32, mutate: bool) -> [u64; 32] {
+ let seed = &p.seed;
+ let m = p.class.scratch_slot_mask();
+ let mut r = [[0u32; LANES]; 8];
+ for lane in 0..LANES {
+ let nonce = base.wrapping_add(lane as u32);
+ for i in 0..8 {
+ let mut x = nonce ^ seed[i];
+ x = x.wrapping_add(0x9e3779b9u32.wrapping_mul(i as u32 + 1));
+ x = splitmix32(x);
+ r[i][lane] = x ^ seed[(i + 1) & 7];
+ }
+ }
+ let mut store: HashMap<(usize, u32), [u32; 3]> = HashMap::new();
+ for _ in 0..ITERATIONS {
+ let sel = r[0];
+ for ins in &p.instrs {
+ let (d, a) = (ins.dst as usize, ins.src as usize);
+ match ins.op {
+ Op::Sub => {
+ for lane in 0..LANES {
+ r[d][lane] = r[d][lane].wrapping_sub(r[a][lane]);
+ }
+ }
+ Op::Add => {
+ for lane in 0..LANES {
+ let c = if (sel[lane] >> ins.bit) & 1 != 0 { ins.imm2 } else { ins.imm };
+ r[d][lane] = r[d][lane].wrapping_add(r[a][lane]).wrapping_add(c);
+ }
+ }
+ Op::Scratch => {
+ for lane in 0..LANES {
+ let slot = r[a][lane] & m;
+ let w = *store.entry((lane, slot)).or_insert_with(|| {
+ let mut f = [0u32; 3];
+ for j in 0..3u32 {
+ // the fill, written out in full rather than through verify::scratch_fill
+ let n = base.wrapping_add(lane as u32);
+ f[j as usize] = splitmix32(
+ (n ^ seed[j as usize])
+ .wrapping_add(slot.wrapping_mul(0x9E37_79B1))
+ .wrapping_add((j + 1).wrapping_mul(0x85EB_CA77)),
+ );
+ }
+ f
+ });
+ let mut x = r[d][lane] ^ w[0];
+ x = x.rotate_left(FOLD_ROT).wrapping_mul(FOLD_MUL) ^ w[1];
+ x = x.rotate_left(FOLD_ROT).wrapping_mul(FOLD_MUL) ^ w[2];
+ r[d][lane] = x;
+ let out = if mutate {
+ [x.rotate_left(7) ^ w[2], x ^ w[1], x.wrapping_add(w[0])]
+ } else {
+ [x ^ w[1], x.rotate_left(7) ^ w[2], x.wrapping_add(w[0])]
+ };
+ store.insert((lane, slot), out);
+ }
+ }
+ other => panic!("the hand model does not implement {other:?}"),
+ }
+ }
+ }
+ let mut out = [0u64; 32];
+ for lane in 0..LANES {
+ let lo = r[0][lane] ^ r[1][lane].rotate_left(7) ^ r[2][lane].rotate_left(14) ^ r[3][lane].rotate_left(21);
+ let hi = r[4][lane] ^ r[5][lane].rotate_left(9) ^ r[6][lane].rotate_left(18) ^ r[7][lane].rotate_left(27);
+ out[lane] = ((hi as u64) << 32) | lo as u64;
+ }
+ out
+}
+
+/// The four unit bases of every edge vector: 0 and 32 (two consecutive units, the pair a one-warp persistent
+/// launch runs on one arena), a unit straddling 2^31, and the unit that wraps past 2^32.
+const EDGE_BASES: [u32; 4] = [0, 32, 0x7fff_fff0, 0xffff_ffe0];
+
+#[test]
+fn edge_programs_match_the_hand_model() {
+ let ds = DatasetSource::new("2026-10-03", DatasetMode::ClosedForm, 24);
+ let mut cases = 0;
+ for kb in [32u8, 128] {
+ for (name, what, p) in edge_programs(kb) {
+ let slots = p.class.scratch_slots_per_lane();
+ for base in EDGE_BASES {
+ let (res, ev) = interpret_warp_scratch(&p, &p.seed, base, &ds, true);
+ let hand = hand_model(&p, base, false);
+ assert_eq!(res.hashes, hand, "{name} k{kb} base {base:#x}: interpreter against the hand model ({what})");
+ assert_ne!(res.hashes, hand_model(&p, base, true), "{name} k{kb}: the comparison has teeth");
+ // the slots the trace saw are the ones the program was built to drive
+ let slot_set: std::collections::BTreeSet = ev.iter().map(|e| e.slot).collect();
+ let m = (slots - 1) as u32;
+ match name.as_str() {
+ "slot0" | "maskplus1" | "sixteen" => assert_eq!(slot_set.into_iter().collect::>(), vec![0]),
+ "slotmask" | "ones" => assert_eq!(slot_set.into_iter().collect::>(), vec![m]),
+ "twoslots" => assert_eq!(slot_set.into_iter().collect::>(), vec![0, m]),
+ "lanevar" => {
+ for e in &ev {
+ assert!(e.slot <= m);
+ }
+ }
+ _ => unreachable!(),
+ }
+ // the chain depth on the driven slot: every RMW after the first per lane is a re-hit
+ let per_lane = p.scratch_ops_per_hash();
+ let hits = ev.iter().filter(|e| e.hit).count();
+ let expected_hits = match name.as_str() {
+ "twoslots" => (per_lane - 2) * LANES,
+ _ => (per_lane - 1) * LANES,
+ };
+ assert_eq!(hits, expected_hits, "{name} k{kb}: re-hits");
+ cases += 1;
+ }
+ }
+ }
+ assert_eq!(cases, 2 * 7 * 4);
+}
+
+// ---------------------------------------------------------------------------------------------------------------
+// 4. The static scratch check over every emitted kernel of every scr pack (question 4)
+// ---------------------------------------------------------------------------------------------------------------
+
+#[derive(Clone, Copy, Debug, PartialEq, Eq)]
+pub enum Dialect {
+ Metal,
+ Cuda,
+ OpenCl,
+}
+
+/// The static scratch check: every scratch read-modify-write in an emitted kernel has the one masked form the
+/// emitter writes, the arena is the lane's own `slots x 4` words, the tag is `salt + unit`, and nothing else
+/// touches the scratch. Like the dataset mask check of `TESTS.md` section 5 and `tests/packs.rs`, a text check:
+/// the guarantee is that the emitter has one template and it masks.
+pub fn scratch_text_check(text: &str, dialect: Dialect, k: usize, slots: usize, kernels: usize) -> Result<(), String> {
+ assert!(kernels >= 1);
+ // every count below is per hash kernel; an OpenCL bound file carries igneum_hash and igneum_hash_bound
+ let k = k * kernels;
+ assert!(slots.is_power_of_two() && slots >= 1);
+ let mask = (slots - 1) as u32;
+ let wpl = slots * 4;
+ let (u, load, store, ptr) = match dialect {
+ Dialect::Metal => ("uint", "uint4 v_ = *(device const uint4*)(arena + s_ * 4u);", "*(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); }", "device uint* arena"),
+ Dialect::Cuda => ("uint32_t", "uint4 v_ = *(const uint4*)(arena + s_ * 4u);", "*(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); }", "uint32_t* arena"),
+ Dialect::OpenCl => ("uint", "uint4 v_ = vload4(s_, arena);", "vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); }", "__global uint* arena"),
+ };
+ let count = |needle: &str| text.matches(needle).count();
+ let mut errs = Vec::new();
+ let mut expect = |what: &str, got: usize, want: usize| {
+ if got != want {
+ errs.push(format!("{what}: {got}, expected {want}"));
+ }
+ };
+ // k slot computations, each masked with exactly the class's mask and immediately followed by the one load form
+ expect("slot definitions `{ u s_ = r`", count(&format!("{{ {u} s_ = r")), k);
+ expect("masked slot followed by the load", count(&format!(" & {mask}u; {load}")), k);
+ expect("stores of the tagged slot", count(store), k);
+ expect("tag compares", count("(v_.x == tag)"), k);
+ expect("fill calls (three per RMW)", count("scr_fill(gbase, lane, s_, "), 3 * k);
+ // the arena: one definition with the class's words per lane, and 2k uses (one load, one store per RMW)
+ expect("arena definition", count(&format!("{ptr} = scratch + ((size_t)warp_ * 32u + lane) * {wpl}u;")), kernels);
+ expect("arena mentions (definition + load + store per RMW)", count("arena"), kernels + 2 * k);
+ expect("tag definition `tag = salt + g_`", count(&format!("{u} tag = salt + g_;")), kernels);
+ expect("direct scratch indexing", count("scratch["), 0);
+ expect("scratch pointer arithmetic outside the arena definition", count("scratch +"), kernels);
+ // no other mask value on a slot: every `s_ = r` line carries the class mask and nothing else carries ` & Nu; uint4 v_`
+ let any_mask_load = count(&format!("u; {load}"));
+ expect("loads preceded by some mask (must all be the class mask)", any_mask_load, k);
+ if errs.is_empty() {
+ Ok(())
+ } else {
+ Err(errs.join("; "))
+ }
+}
+
+fn packs_rw_dir() -> PathBuf {
+ PathBuf::from(env!("CARGO_MANIFEST_DIR")).join("../proto-cuda/packs-readwidth")
+}
+
+fn scr_packs() -> Vec {
+ let mut v: Vec = std::fs::read_dir(packs_rw_dir())
+ .unwrap()
+ .map(|d| d.unwrap().file_name().to_string_lossy().to_string())
+ .filter(|n| n.starts_with("scr"))
+ .collect();
+ v.sort();
+ v
+}
+
+fn read_pack(pack: &str, file: &str) -> String {
+ let p = packs_rw_dir().join(pack).join(file);
+ std::fs::read_to_string(&p).unwrap_or_else(|e| panic!("read {}: {e}", p.display()))
+}
+
+/// Every scr pack regenerates from its program.json (seed bytes, class, day bytes, size) to the same six kernel
+/// texts, byte for byte, and every one of those texts passes the static scratch check for the class's k and slot
+/// count; the check fails on four deliberate breaks of a copy of the Metal text (mask dropped, mask changed, arena
+/// stride changed, a stray scratch access) and on the OpenCL and CUDA twins of the first.
+#[test]
+fn scr_packs_regenerate_and_pass_the_static_scratch_check() {
+ let packs = scr_packs();
+ assert!(packs.len() >= 6, "the scr packs: {packs:?}");
+ let mut checked = 0;
+ let mut sample_metal = String::new();
+ let mut sample_cl = String::new();
+ let mut sample_cu = String::new();
+ let mut sample_k = 0;
+ let mut sample_slots = 0;
+ for pack in &packs {
+ let j: Value = serde_json::from_str(&read_pack(pack, "program.json")).unwrap();
+ let name = j["load_class"].as_str().unwrap();
+ let c = class(name);
+ assert_eq!(&format!("{name}"), pack, "pack directory named after its class");
+ let seed = j["seed"].as_str().unwrap();
+ let seed_bytes = igneum_pow::bind::unhex(j["seed_bytes"].as_str().unwrap()).unwrap();
+ let day_bytes = igneum_pow::bind::unhex(j["dataset"]["day_bytes"].as_str().unwrap()).unwrap();
+ let log2 = j["dataset"]["log2_words"].as_u64().unwrap() as u32;
+ assert_eq!(j["dataset_mode"].as_str().unwrap(), "memory-hard");
+ let program = generate_from_seed_bytes_class(seed, &seed_bytes, c);
+ assert_eq!(program.class, c);
+ assert_eq!(program.program_id(), u64::from_str_radix(j["program_id"].as_str().unwrap().trim_start_matches("0x"), 16).unwrap());
+ let mut dataset = DatasetSource::from_key(seed_words_from_bytes(&day_bytes), DatasetMode::MemoryHard, log2);
+ dataset.key_bytes = day_bytes;
+ let e = Epoch { program, dataset };
+ let p = &e.program;
+ let mp = e.dataset.memhard().map(|m| &m.params);
+ let k = c.scratch_slots();
+ let slots = c.scratch_slots_per_lane();
+ assert_eq!(p.scratch_ops_per_hash(), k * ITERATIONS);
+ for (file, text, dialect, kernels) in [
+ ("program.metal", metal_program(p, log2, LoadSource::Stored), Dialect::Metal, 1),
+ ("program_bound.metal", metal_program_bound(p, log2), Dialect::Metal, 1),
+ ("kernel.cu", cuda_kernel(p, mp), Dialect::Cuda, 1),
+ ("kernel_bound.cu", cuda_kernel_bound(p, mp), Dialect::Cuda, 1),
+ ("kernel.cl", opencl_kernel(p, mp), Dialect::OpenCl, 1),
+ // the OpenCL bound file carries igneum_hash and igneum_hash_bound
+ ("kernel_bound.cl", opencl_kernel_bound(p, mp), Dialect::OpenCl, 2),
+ ] {
+ let on_disk = read_pack(pack, file);
+ assert_eq!(on_disk, text, "{pack}/{file}: the pack is the emitter's text");
+ // scr0 is the persistent control: an arena and a tag, no read-modify-write; the check holds with k = 0
+ scratch_text_check(&on_disk, dialect, k, slots, kernels).unwrap_or_else(|e| panic!("{pack}/{file}: {e}"));
+ checked += 1;
+ }
+ // the vectors of the pack are the CPU's
+ let v: Value = serde_json::from_str(&read_pack(pack, "vectors.json")).unwrap();
+ for w in v["warps"].as_array().unwrap() {
+ let base = w["base_nonce"].as_u64().unwrap() as u32;
+ let got = e.hash_warp(base);
+ for (lane, x) in w["expected"].as_array().unwrap().iter().enumerate() {
+ let want = u64::from_str_radix(x.as_str().unwrap().trim_start_matches("0x"), 16).unwrap();
+ assert_eq!(got[lane], want, "{pack}: base {base} lane {lane}");
+ }
+ }
+ if k == 4 && slots == 64 {
+ sample_metal = read_pack(pack, "program.metal");
+ sample_cl = read_pack(pack, "kernel.cl");
+ sample_cu = read_pack(pack, "kernel.cu");
+ sample_k = k;
+ sample_slots = slots;
+ }
+ }
+ assert_eq!(checked, packs.len() * 6);
+ println!("static scratch check: {checked} kernels over {} scr packs", packs.len());
+
+ // The deliberate breaks (the watcher rule of CLAUDE.md: a check is trusted once it fails on a known-broken
+ // case). Each must be caught; the message names what.
+ assert!(sample_k == 4 && sample_slots == 64, "scr4k32 is in the pack set");
+ let mask = format!(" & {}u; uint4 v_", sample_slots - 1);
+ let broken_mask = sample_metal.replacen(&mask, "; uint4 v_", 1);
+ assert_ne!(broken_mask, sample_metal);
+ let e = scratch_text_check(&broken_mask, Dialect::Metal, 4, 64, 1).unwrap_err();
+ assert!(e.contains("masked slot followed by the load: 3, expected 4"), "{e}");
+ println!("break 1 (one mask dropped, Metal): {e}");
+ let wrong_mask = sample_metal.replace(" & 63u;", " & 127u;");
+ let e = scratch_text_check(&wrong_mask, Dialect::Metal, 4, 64, 1).unwrap_err();
+ assert!(e.contains("masked slot followed by the load: 0, expected 4"), "{e}");
+ println!("break 2 (mask 63 -> 127 on every RMW, Metal): {e}");
+ let wrong_stride = sample_metal.replace("* 256u;", "* 128u;");
+ let e = scratch_text_check(&wrong_stride, Dialect::Metal, 4, 64, 1).unwrap_err();
+ assert!(e.contains("arena definition: 0, expected 1"), "{e}");
+ println!("break 3 (arena stride 256 -> 128 words, Metal): {e}");
+ let stray = format!("{sample_metal}\n// stray\n// arena[0] = 0u; scratch[1] = 1u;\n");
+ let e = scratch_text_check(&stray, Dialect::Metal, 4, 64, 1).unwrap_err();
+ assert!(e.contains("arena mentions") && e.contains("direct scratch indexing: 1, expected 0"), "{e}");
+ println!("break 4 (a stray arena and scratch access, Metal): {e}");
+ let e = scratch_text_check(&sample_cl.replacen(" & 63u; uint4 v_ = vload4", "; uint4 v_ = vload4", 1), Dialect::OpenCl, 4, 64, 1).unwrap_err();
+ assert!(e.contains("masked slot followed by the load: 3, expected 4"), "{e}");
+ println!("break 5 (one mask dropped, OpenCL): {e}");
+ let e = scratch_text_check(&sample_cu.replacen(" & 63u; uint4 v_ = *(const uint4*)", "; uint4 v_ = *(const uint4*)", 1), Dialect::Cuda, 4, 64, 1).unwrap_err();
+ assert!(e.contains("masked slot followed by the load: 3, expected 4"), "{e}");
+ println!("break 6 (one mask dropped, CUDA): {e}");
+ // and the unbroken texts pass under the same calls
+ scratch_text_check(&sample_metal, Dialect::Metal, 4, 64, 1).unwrap();
+ scratch_text_check(&sample_cl, Dialect::OpenCl, 4, 64, 1).unwrap();
+ scratch_text_check(&sample_cu, Dialect::Cuda, 4, 64, 1).unwrap();
+ // a wrong slot count, RMW count or kernel count against a right text fails too (the check is tied to the class)
+ assert!(scratch_text_check(&sample_metal, Dialect::Metal, 4, 256, 1).is_err());
+ assert!(scratch_text_check(&sample_metal, Dialect::Metal, 3, 64, 1).is_err());
+ assert!(scratch_text_check(&sample_metal, Dialect::Metal, 4, 64, 2).is_err());
+}
+
+// ---------------------------------------------------------------------------------------------------------------
+// 5. The fuzz: 200 generated scratch programs, contract on every instruction, 4 units each across the 32-bit
+// range including the wrap; with IGNEUM_SCRATCH_PACKS_OUT the packs for the Metal runs (question 3, 4)
+// ---------------------------------------------------------------------------------------------------------------
+
+/// Write a pack whose vectors.json carries `bases` (any number of units) instead of the three standard bases.
+fn write_pack_with_bases(dir: &PathBuf, e: &Epoch, day: &str, bases: &[u32], source: &str) -> Vec<[u64; 32]> {
+ let mut pack = export_pack(e, day, source);
+ let outs: Vec<[u64; 32]> = bases.iter().map(|&b| e.hash_warp(b)).collect();
+ let vj = vectors_json(&e.program, day, e.dataset.log2_words, bases, &outs, &pack.vectors, e.dataset.mask, source, true);
+ for f in pack.files.iter_mut() {
+ if f.0 == "vectors.json" {
+ f.1 = vj.clone();
+ }
+ }
+ pack.write_to(dir).unwrap();
+ outs
+}
+
+fn contract(p: &Program) {
+ assert_eq!(p.instrs.len(), INSTR_COUNT);
+ assert_eq!(p.instrs.iter().filter(|i| i.op == Op::Load).count() + p.instrs.iter().filter(|i| i.op == Op::Scratch).count(), 16);
+ assert_eq!(p.instrs.iter().filter(|i| i.op == Op::Scratch).count(), p.class.scratch_slots());
+ assert!(p.instrs[0].op != Op::Load && p.instrs[0].op != Op::Scratch, "instruction 0 is never a memory op");
+ for (k, i) in p.instrs.iter().enumerate() {
+ assert!(i.src != i.dst, "#{k}: src == dst");
+ assert!((1..=31).contains(&i.rot), "#{k}: rot {}", i.rot);
+ assert!([1u8, 2, 4, 8, 16].contains(&i.mask), "#{k}: mask {}", i.mask);
+ assert!(i.dst < 8 && i.src < 8 && i.src2 < 8);
+ assert_eq!(i.width, 1, "#{k}: a scratch class reads one-word loads");
+ }
+ assert!(igneum_pow::accept::check(p).is_ok(), "an accepted program");
+}
+
+#[test]
+fn fuzz_scr_programs_cpu() {
+ let n: usize = std::env::var("IGNEUM_SCRATCH_FUZZ").ok().and_then(|s| s.parse().ok()).unwrap_or(200);
+ let out = std::env::var("IGNEUM_SCRATCH_PACKS_OUT").ok().map(PathBuf::from);
+ let mut rng = SplitMix64::new(0x6967_6e65_756d_2d73); // "igneum-s"
+ let day = "2026-10-03";
+ let closed = DatasetSource::new(day, DatasetMode::ClosedForm, 28);
+ // memory-hard sources per size, built once each (the cache fill is 0.2 s); only when packs are written
+ let mut mh: HashMap = HashMap::new();
+ let mut manifest = String::from("pack\tclass\tlog2\tprogram_id\tscratch_ops_per_hash\tbases\n");
+ let mut per_class: HashMap = HashMap::new();
+ let mut units = 0usize;
+ let mut wraps = 0usize;
+ if let Some(dir) = &out {
+ std::fs::create_dir_all(dir).unwrap();
+ // the edge packs first: 64 MiB datasets (no dataset load in them), the four edge bases
+ for kb in [32u8, 128] {
+ for (name, _what, p) in edge_programs(kb) {
+ let log2 = 24;
+ let ds = mh.remove(&log2).unwrap_or_else(|| DatasetSource::new(day, DatasetMode::MemoryHard, log2));
+ let e = Epoch { program: p, dataset: ds };
+ let pack_name = format!("edge-{name}-k{kb}");
+ write_pack_with_bases(&dir.join(&pack_name), &e, day, &EDGE_BASES, "igneum-pow tests/scratch.rs edge");
+ manifest.push_str(&format!(
+ "{pack_name}\t{}\t{log2}\t{:016x}\t{}\t{}\n",
+ e.program.class.name(),
+ e.program.program_id(),
+ e.program.scratch_ops_per_hash(),
+ EDGE_BASES.iter().map(|b| format!("{b}")).collect::>().join(",")
+ ));
+ mh.insert(log2, e.dataset);
+ }
+ }
+ }
+ for i in 0..n {
+ let name = CLASSES[rng.below(CLASSES.len() as u64) as usize];
+ let c = class(name);
+ let seed = format!("igneum-scratch-fuzz/{i}");
+ let p = generate_class(&seed, c);
+ contract(&p);
+ *per_class.entry(name.to_string()).or_insert(0) += 1;
+ // four bases: one inside a 256-nonce batch (in-batch check on the GPU), one straddling 2^31, one in
+ // the last 256 nonces (the unit wraps past 2^32 or ends on it), one uniform
+ let b0 = (rng.below(8) as u32) * 32;
+ let b1 = 0x8000_0000u32.wrapping_sub(256).wrapping_add((rng.below(16) as u32) * 32);
+ let b2 = 0xffff_ff00u32.wrapping_add((rng.below(8) as u32) * 32);
+ let b3 = (rng.next() as u32) & !31;
+ let bases = [b0, b1, b2, b3];
+ // an aligned unit never straddles 2^32 (spec 1.9); the top unit ends on 0xffffffff and the persistent
+ // kernel's unit sequence wraps inside a launch, which the Metal run checks with packbench --batch-base
+ wraps += bases.iter().filter(|&&b| b >= 0xffff_ff00).count();
+ // the CPU: the interpreter is deterministic and every scratch event is inside the lane's slots
+ for &b in &bases {
+ let (r1, ev) = interpret_warp_scratch(&p, &p.seed, b, &closed, true);
+ let r2 = interpret_warp_scratch(&p, &p.seed, b, &closed, false).0;
+ assert_eq!(r1.hashes, r2.hashes);
+ assert_eq!(ev.len(), p.scratch_ops_per_hash() * LANES);
+ assert!(ev.iter().all(|e: &ScratchEvent| e.slot < c.scratch_slots_per_lane() as u32));
+ units += 1;
+ }
+ if let Some(dir) = &out {
+ let log2 = [24u32, 26, 28][rng.below(3) as usize];
+ let ds = mh.remove(&log2).unwrap_or_else(|| DatasetSource::new(day, DatasetMode::MemoryHard, log2));
+ let e = Epoch { program: p, dataset: ds };
+ let pack_name = format!("fuzz-{i:03}-{name}-l{log2}");
+ write_pack_with_bases(&dir.join(&pack_name), &e, day, &bases, "igneum-pow tests/scratch.rs fuzz");
+ manifest.push_str(&format!(
+ "{pack_name}\t{name}\t{log2}\t{:016x}\t{}\t{}\n",
+ e.program.program_id(),
+ e.program.scratch_ops_per_hash(),
+ bases.iter().map(|b| format!("{b}")).collect::>().join(",")
+ ));
+ mh.insert(log2, e.dataset);
+ } else {
+ let _ = rng.below(3);
+ }
+ }
+ let mut classes: Vec<_> = per_class.iter().collect();
+ classes.sort();
+ println!("fuzz: {n} programs, {units} units on the CPU, {wraps} units in the top 256 nonces, classes {classes:?}");
+ assert_eq!(units, 4 * n);
+ assert_eq!(wraps, n, "every program has a unit in the top 256 nonces");
+ if let Some(dir) = &out {
+ std::fs::write(dir.join("manifest.tsv"), manifest).unwrap();
+ println!("packs written to {}", dir.display());
+ }
+}
+
+/// The fold and rewrite, restated: a slot after `d` dependent RMWs holds 96 bits that are a function of the fill
+/// (3 words, a pure function of nonce, slot and seed) and the `d` fold values; a chip that keeps the `d` fold
+/// values (32 bits each) instead of the 96-bit slot recomputes the slot in `d` rewrites. This test pins the
+/// arithmetic the analysis uses (question 2): the replay from the fold values reproduces the slot.
+#[test]
+fn slot_is_replayable_from_its_fold_values() {
+ let seed = seed_words_from_bytes(b"igneum-genesis");
+ let (base, lane, slot) = (0x1234_5600u32, 5u32, 17u32);
+ let fill = [scratch_fill(&seed, base, lane, slot, 0), scratch_fill(&seed, base, lane, slot, 1), scratch_fill(&seed, base, lane, slot, 2)];
+ let mut rng = SplitMix64::new(99);
+ let dsts: Vec = (0..64).map(|_| rng.next() as u32).collect();
+ // the honest sequence: read, fold, rewrite, 64 times
+ let mut w = fill;
+ let mut xs = Vec::new();
+ for &d in &dsts {
+ let x = fold_words(d, &w);
+ xs.push(x);
+ w = scratch_rewrite(x, &w);
+ }
+ // the replay: from the fill and the stored fold values alone
+ let mut w2 = fill;
+ for &x in &xs {
+ w2 = scratch_rewrite(x, &w2);
+ }
+ assert_eq!(w, w2);
+ // and nothing shorter: the fold value at step d depends on the slot content at step d, which depends on
+ // every earlier fold value (drop one and the chain diverges)
+ let mut w3 = fill;
+ for (i, &x) in xs.iter().enumerate() {
+ if i != 10 {
+ w3 = scratch_rewrite(x, &w3);
+ }
+ }
+ assert_ne!(w, w3);
+}
diff --git a/proto-cuda/packs-ca2-mixer/mx8-devnet-epoch0/kernel.cl b/proto-cuda/packs-ca2-mixer/mx8-devnet-epoch0/kernel.cl
new file mode 100644
index 000000000..e4f3bb63d
--- /dev/null
+++ b/proto-cuda/packs-ca2-mixer/mx8-devnet-epoch0/kernel.cl
@@ -0,0 +1,279 @@
+// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
+// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal).
+// Built from source at runtime by proto-opencl/host.c, which passes these defines:
+// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit)
+// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default)
+// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32
+// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition
+// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units;
+// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32.
+#ifndef IGNEUM_GROUP
+#define IGNEUM_GROUP 32
+#endif
+#ifndef IGNEUM_EXCHANGE
+#define IGNEUM_EXCHANGE 0
+#endif
+#ifdef __OPENCL_VERSION__
+#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1)))
+#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n]
+#if IGNEUM_EXCHANGE == 1
+#ifdef cl_khr_subgroups
+#pragma OPENCL EXTENSION cl_khr_subgroups : enable
+#endif
+#ifdef cl_khr_subgroup_shuffle
+#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable
+#endif
+#elif IGNEUM_EXCHANGE == 2
+#pragma OPENCL EXTENSION cl_intel_subgroups : enable
+#endif
+#else
+// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros.
+#include "emu_opencl.h"
+#endif
+
+#if IGNEUM_EXCHANGE == 1
+#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m))
+#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
+#elif IGNEUM_EXCHANGE == 2
+#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m))
+#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
+#else
+// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per
+// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane
+// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's
+// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier.
+#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; }
+#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; }
+#endif
+
+static inline uint splitmix32(uint x) {
+ x ^= x >> 16; x *= 0x7feb352du;
+ x ^= x >> 15; x *= 0x846ca68bu;
+ x ^= x >> 16;
+ return x;
+}
+// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32.
+static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); }
+// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x.
+static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); }
+static inline uint ds_elem(uint i, uint d0, uint d1) {
+ uint x = i ^ d0;
+ x *= 0x9E3779B1u; x ^= x >> 15;
+ x += d1;
+ x *= 0x85EBCA77u; x ^= x >> 13;
+ x *= 0xC2B2AE3Du; x ^= x >> 16;
+ return x;
+}
+
+// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
+// Item: 8 rounds of 8 x seed-parameterised mixer + one 64-byte cache read, then 8 x final mixer (class v3, mixer multiplier 8,
+// docs/plans/mixer-x4.md: the round key of application j of round r is 0x9E3779B9 * (r * m + j + 1)). All parameters are literals.
+#define MH_CACHE_LINE_MASK 0x003fffffu
+#define MH_SEGMENT_LINES 64u
+#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
+static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
+
+// y = ChaCha12 core(x) + x
+static inline void mh_chacha_block(const uint* x, uint* y) {
+ for (uint i = 0u; i < 16u; ++i) y[i] = x[i];
+ for (uint r = 0u; r < 6u; ++r) {
+ MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
+ MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
+ MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
+ MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
+ }
+ for (uint i = 0u; i < 16u; ++i) y[i] += x[i];
+}
+
+// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
+static inline void mh_cache_segment(__global uint* cache, uint seg) {
+ uint prev[16]; uint x[16]; uint y[16];
+ for (uint i = 0u; i < 16u; ++i) prev[i] = 0u;
+ for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) {
+ x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
+ x[4] = 0xceed56d7u ^ prev[4];
+ x[5] = 0x9ba270d2u ^ prev[5];
+ x[6] = 0x82caab2du ^ prev[6];
+ x[7] = 0x81ebce0eu ^ prev[7];
+ x[8] = 0x12b6ecf1u ^ prev[8];
+ x[9] = 0xd0f3fd7cu ^ prev[9];
+ x[10] = 0xd872eefeu ^ prev[10];
+ x[11] = 0xc158c7bdu ^ prev[11];
+ x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
+ mh_chacha_block(x, y);
+ __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
+ for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
+ }
+}
+
+// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
+static inline void mh_mixer(uint* s, uint rk) {
+ s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
+ s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
+ s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
+ s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
+ s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
+ s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
+ s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
+ s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
+ s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
+ s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
+ s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
+ s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
+ s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
+ s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
+ s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
+ s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
+ MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
+ MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
+ MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
+ MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
+}
+
+// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of 8 x mixer + cache line s[0] & mask; 8 x final mixer.
+static inline void mh_item(__global const uint* cache, uint t, uint* s) {
+ s[0] = 0xceed56d7u;
+ s[1] = 0x9ba270d2u;
+ s[2] = 0x82caab2du;
+ s[3] = 0x81ebce0eu;
+ s[4] = 0x12b6ecf1u;
+ s[5] = 0xd0f3fd7cu;
+ s[6] = 0xd872eefeu;
+ s[7] = 0xc158c7bdu;
+ s[8] = t * 0xf351d601u + 0xc6892460u;
+ s[9] = t * 0xa3bb398fu + 0x25b7228au;
+ s[10] = t * 0xb5a09e35u + 0xcd515004u;
+ s[11] = t * 0x7509c9c1u + 0x2846527au;
+ s[12] = t * 0x6bbf31e9u + 0xa6324241u;
+ s[13] = t * 0xfc849a79u + 0x36e3ec53u;
+ s[14] = t * 0xded91851u + 0x82961bacu;
+ s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
+ for (uint r = 0u; r < 8u; ++r) {
+ for (uint j = 0u; j < 8u; ++j) mh_mixer(s, 0x9E3779B9u * (r * 8u + j + 1u));
+ __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
+ for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i];
+ }
+ for (uint j = 0u; j < 8u; ++j) mh_mixer(s, 0x9E3779B9u * (64u + j + 1u));
+}
+// dataset[w] without the dataset: derive item w >> 4 and take word w & 15.
+static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; }
+
+// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item.
+// The same constants as memhard.h in this pack (one emitter, three dialects).
+__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) {
+ uint seg = (uint)get_global_id(0);
+ if (seg < nSegments) mh_cache_segment(cache, seg);
+}
+__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) {
+ uint t = (uint)get_global_id(0);
+ if (t < nItems) {
+ uint s[16];
+ mh_item(cache, t, s);
+ __global uint* d = ds + ((ulong)t * 16u);
+ for (uint i = 0u; i < 16u; ++i) d[i] = s[i];
+ }
+}
+
+// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the
+// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and
+// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all).
+IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) {
+ uint gid = (uint)get_global_id(0);
+ uint lid = (uint)get_local_id(0);
+ uint nonce = baseNonce + gid;
+ uint r0, r1, r2, r3, r4, r5, r6, r7;
+#if IGNEUM_EXCHANGE == 0
+ IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
+ uint xk = 0u;
+#else
+ (void)lid;
+#endif
+ { uint x = nonce ^ 0x667d0fbdu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x7b8e5963u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
+ { uint x = nonce ^ 0x7b8e5963u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x31c67e5eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
+ { uint x = nonce ^ 0x31c67e5eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4529ddc6u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
+ { uint x = nonce ^ 0x4529ddc6u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef19d6d8u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
+ { uint x = nonce ^ 0xef19d6d8u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xaccf6211u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
+ { uint x = nonce ^ 0xaccf6211u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xda0aed32u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
+ { uint x = nonce ^ 0xda0aed32u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xabc6df31u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
+ { uint x = nonce ^ 0xabc6df31u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x667d0fbdu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
+
+ for (uint it = 0u; it < 8u; ++it) {
+ uint sel = r0;
+ r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
+ { uint t_; IGNEUM_SHFL_XOR(t_, r0, 4u); r2 = r2 ^ t_; } // 1 shfl
+ r3 = r3 + r2 + ((((sel >> 10u) & 1u) != 0u) ? 0x642e66dbu : 0x2cccb6cau); // 2 add
+ r0 = rotl_imm(r0, 19u); // 3 rotl
+ r7 = rotr_var(r7, r6); // 4 rotr
+ r7 = r7 + r4 + ((((sel >> 21u) & 1u) != 0u) ? 0xc1535555u : 0xee02465fu); // 5 add
+ r1 = mul_hi(r1, r7); // 6 mulhi
+ r4 = r4 ^ ds[r2 & mask]; // 7 load
+ r7 = r7 ^ ds[r4 & mask]; // 8 load
+ r0 = r0 ^ ds[r3 & mask]; // 9 load
+ r5 = r5 ^ ds[r1 & mask]; // 10 load
+ r1 = r1 ^ ds[r5 & mask]; // 11 load
+ r3 = mul_hi(r3, r5); // 12 mulhi
+ r1 = r1 ^ ds[r3 & mask]; // 13 load
+ r0 = r0 - r3; // 14 sub
+ r5 = r1 * r3 + r5; // 15 mad
+ r6 = mul_hi(r6, r1); // 16 mulhi
+ r5 = r5 + r2 + ((((sel >> 28u) & 1u) != 0u) ? 0x8b965b57u : 0x697b3d00u); // 17 add
+ r0 = mul_hi(r0, r6); // 18 mulhi
+ r5 = rotr_var(r5, r3); // 19 rotr
+ r5 = mul_hi(r5, r2); // 20 mulhi
+ r1 = r1 + r0 + ((((sel >> 1u) & 1u) != 0u) ? 0x6d7e8d05u : 0xebcf247au); // 21 add
+ r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 22 add
+ r1 = mul_hi(r1, r5); // 23 mulhi
+ r2 = r2 - r5; // 24 sub
+ r7 = r7 + r4 + ((((sel >> 2u) & 1u) != 0u) ? 0x699ef1bbu : 0x08ffa6c7u); // 25 add
+ { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 26 shfl
+ r7 = r7 + r1 + ((((sel >> 14u) & 1u) != 0u) ? 0xb4ead2fbu : 0xe60fea84u); // 27 add
+ r3 = r3 + r1 + ((((sel >> 6u) & 1u) != 0u) ? 0x8f30d21du : 0x65c76dabu); // 28 add
+ r2 = r2 ^ ds[r1 & mask]; // 29 load
+ r5 = r5 ^ ds[r7 & mask]; // 30 load
+ r2 = r2 ^ ds[r5 & mask]; // 31 load
+ { uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r1 = r1 ^ t_; } // 32 shfl
+ r4 = r5 * r7 + r4; // 33 mad
+ r4 = r4 + r2 + ((((sel >> 21u) & 1u) != 0u) ? 0xc7ce690cu : 0x0480debeu); // 34 add
+ { uint t_; IGNEUM_SHFL_XOR(t_, r7, 8u); r3 = r3 ^ t_; } // 35 shfl
+ r7 = r7 + r1 + ((((sel >> 2u) & 1u) != 0u) ? 0xe10c2c95u : 0xc53b542eu); // 36 add
+ r5 = r5 ^ r7; // 37 xor
+ r2 = r2 | r1; // 38 or
+ r1 = mul_hi(r1, r0); // 39 mulhi
+ r6 = rotl_imm(r6, 19u); // 40 rotl
+ r4 = mul_hi(r4, r6); // 41 mulhi
+ r6 = r6 - r0; // 42 sub
+ { uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r6 = r6 ^ t_; } // 43 shfl
+ r4 = r4 ^ ds[r2 & mask]; // 44 load
+ r1 = r1 ^ r3; // 45 xor
+ r7 = r7 ^ ds[r0 & mask]; // 46 load
+ r3 = r3 ^ ds[r1 & mask]; // 47 load
+ r5 = r5 * r3; // 48 mul
+ r1 = r1 - r5; // 49 sub
+ r2 = rotl_imm(r2, 8u); // 50 rotl
+ r1 = r1 + r5 + ((((sel >> 23u) & 1u) != 0u) ? 0x77b9bd43u : 0xa900fec4u); // 51 add
+ r4 = r4 ^ ds[r7 & mask]; // 52 load
+ r2 = r2 - r7; // 53 sub
+ r4 = r4 ^ r0; // 54 xor
+ r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 55 add
+ r2 = r2 ^ ds[r4 & mask]; // 56 load
+ r0 = r1 * r4 + r0; // 57 mad
+ r3 = r3 ^ ds[r5 & mask]; // 58 load
+ r5 = r5 | r6; // 59 or
+ r6 = r5 * r7 + r6; // 60 mad
+ r4 = rotl_imm(r4, 28u); // 61 rotl
+ r5 = mul_hi(r5, r0); // 62 mulhi
+ r3 = r3 ^ ds[r6 & mask]; // 63 load
+ }
+ uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
+ uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
+ out[gid] = ((ulong)hi << 32) | (ulong)lo;
+}
+
+#if IGNEUM_EXCHANGE != 0
+// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the
+// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a
+// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md.
+IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) {
+ if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); }
+}
+#endif
diff --git a/proto-cuda/packs-ca2-mixer/mx8-devnet-epoch0/kernel.cu b/proto-cuda/packs-ca2-mixer/mx8-devnet-epoch0/kernel.cu
new file mode 100644
index 000000000..c129edeec
--- /dev/null
+++ b/proto-cuda/packs-ca2-mixer/mx8-devnet-epoch0/kernel.cu
@@ -0,0 +1,164 @@
+// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
+// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal).
+// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC.
+#include
+#include
+#include "program.h"
+#include "memhard.h"
+
+__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
+ x ^= x >> 16; x *= 0x7feb352du;
+ x ^= x >> 15; x *= 0x846ca68bu;
+ x ^= x >> 16;
+ return x;
+}
+// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31.
+__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
+// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x.
+__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
+__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) {
+ uint32_t x = i ^ d0;
+ x *= 0x9E3779B1u; x ^= x >> 15;
+ x += d1;
+ x *= 0x85EBCA77u; x ^= x >> 13;
+ x *= 0xC2B2AE3Du; x ^= x >> 16;
+ return x;
+}
+
+// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item.
+// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host.
+__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) {
+ uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x;
+ if (seg < nSegments) mh_cache_segment(cache, seg);
+}
+__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
+ uint32_t t = blockIdx.x * blockDim.x + threadIdx.x;
+ if (t < nItems) {
+ uint32_t s[16];
+ mh_item(cache, t, s);
+ uint32_t* d = ds + (size_t)t * 16u;
+ for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i];
+ }
+}
+
+// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every
+// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a
+// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid.
+__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) {
+ uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
+ uint32_t nonce = baseNonce + gid;
+ uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
+ { uint32_t x = nonce ^ 0x667d0fbdu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x7b8e5963u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
+ { uint32_t x = nonce ^ 0x7b8e5963u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x31c67e5eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
+ { uint32_t x = nonce ^ 0x31c67e5eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4529ddc6u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
+ { uint32_t x = nonce ^ 0x4529ddc6u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef19d6d8u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
+ { uint32_t x = nonce ^ 0xef19d6d8u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xaccf6211u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
+ { uint32_t x = nonce ^ 0xaccf6211u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xda0aed32u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
+ { uint32_t x = nonce ^ 0xda0aed32u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xabc6df31u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
+ { uint32_t x = nonce ^ 0xabc6df31u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x667d0fbdu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
+
+ for (uint32_t it = 0u; it < 8u; ++it) {
+ uint32_t sel = r0;
+ r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
+ r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r0, 4); // 1 shfl
+ r3 = r3 + r2 + ((((sel >> 10u) & 1u) != 0u) ? 0x642e66dbu : 0x2cccb6cau); // 2 add
+ r0 = rotl_imm(r0, 19u); // 3 rotl
+ r7 = rotr_var(r7, r6); // 4 rotr
+ r7 = r7 + r4 + ((((sel >> 21u) & 1u) != 0u) ? 0xc1535555u : 0xee02465fu); // 5 add
+ r1 = __umulhi(r1, r7); // 6 mulhi
+ r4 = r4 ^ ds[r2 & mask]; // 7 load
+ r7 = r7 ^ ds[r4 & mask]; // 8 load
+ r0 = r0 ^ ds[r3 & mask]; // 9 load
+ r5 = r5 ^ ds[r1 & mask]; // 10 load
+ r1 = r1 ^ ds[r5 & mask]; // 11 load
+ r3 = __umulhi(r3, r5); // 12 mulhi
+ r1 = r1 ^ ds[r3 & mask]; // 13 load
+ r0 = r0 - r3; // 14 sub
+ r5 = r1 * r3 + r5; // 15 mad
+ r6 = __umulhi(r6, r1); // 16 mulhi
+ r5 = r5 + r2 + ((((sel >> 28u) & 1u) != 0u) ? 0x8b965b57u : 0x697b3d00u); // 17 add
+ r0 = __umulhi(r0, r6); // 18 mulhi
+ r5 = rotr_var(r5, r3); // 19 rotr
+ r5 = __umulhi(r5, r2); // 20 mulhi
+ r1 = r1 + r0 + ((((sel >> 1u) & 1u) != 0u) ? 0x6d7e8d05u : 0xebcf247au); // 21 add
+ r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 22 add
+ r1 = __umulhi(r1, r5); // 23 mulhi
+ r2 = r2 - r5; // 24 sub
+ r7 = r7 + r4 + ((((sel >> 2u) & 1u) != 0u) ? 0x699ef1bbu : 0x08ffa6c7u); // 25 add
+ r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 26 shfl
+ r7 = r7 + r1 + ((((sel >> 14u) & 1u) != 0u) ? 0xb4ead2fbu : 0xe60fea84u); // 27 add
+ r3 = r3 + r1 + ((((sel >> 6u) & 1u) != 0u) ? 0x8f30d21du : 0x65c76dabu); // 28 add
+ r2 = r2 ^ ds[r1 & mask]; // 29 load
+ r5 = r5 ^ ds[r7 & mask]; // 30 load
+ r2 = r2 ^ ds[r5 & mask]; // 31 load
+ r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r7, 4); // 32 shfl
+ r4 = r5 * r7 + r4; // 33 mad
+ r4 = r4 + r2 + ((((sel >> 21u) & 1u) != 0u) ? 0xc7ce690cu : 0x0480debeu); // 34 add
+ r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r7, 8); // 35 shfl
+ r7 = r7 + r1 + ((((sel >> 2u) & 1u) != 0u) ? 0xe10c2c95u : 0xc53b542eu); // 36 add
+ r5 = r5 ^ r7; // 37 xor
+ r2 = r2 | r1; // 38 or
+ r1 = __umulhi(r1, r0); // 39 mulhi
+ r6 = rotl_imm(r6, 19u); // 40 rotl
+ r4 = __umulhi(r4, r6); // 41 mulhi
+ r6 = r6 - r0; // 42 sub
+ r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 43 shfl
+ r4 = r4 ^ ds[r2 & mask]; // 44 load
+ r1 = r1 ^ r3; // 45 xor
+ r7 = r7 ^ ds[r0 & mask]; // 46 load
+ r3 = r3 ^ ds[r1 & mask]; // 47 load
+ r5 = r5 * r3; // 48 mul
+ r1 = r1 - r5; // 49 sub
+ r2 = rotl_imm(r2, 8u); // 50 rotl
+ r1 = r1 + r5 + ((((sel >> 23u) & 1u) != 0u) ? 0x77b9bd43u : 0xa900fec4u); // 51 add
+ r4 = r4 ^ ds[r7 & mask]; // 52 load
+ r2 = r2 - r7; // 53 sub
+ r4 = r4 ^ r0; // 54 xor
+ r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 55 add
+ r2 = r2 ^ ds[r4 & mask]; // 56 load
+ r0 = r1 * r4 + r0; // 57 mad
+ r3 = r3 ^ ds[r5 & mask]; // 58 load
+ r5 = r5 | r6; // 59 or
+ r6 = r5 * r7 + r6; // 60 mad
+ r4 = rotl_imm(r4, 28u); // 61 rotl
+ r5 = __umulhi(r5, r0); // 62 mulhi
+ r3 = r3 ^ ds[r6 & mask]; // 63 load
+ }
+ uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
+ uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
+ out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
+}
+
+// Host-side launch wrappers. Declared in program.h, called from host.cu.
+cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) {
+ if (nSegments == 0u) return cudaErrorInvalidValue;
+ uint32_t block = 256u;
+ uint32_t grid = (nSegments + block - 1u) / block;
+ igneum_cache_fill<<>>(cache, nSegments);
+ return cudaGetLastError();
+}
+
+cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
+ if (nItems == 0u) return cudaErrorInvalidValue;
+ uint32_t block = 256u;
+ uint32_t grid = (nItems + block - 1u) / block;
+ igneum_build<<>>(ds, cache, nItems);
+ return cudaGetLastError();
+}
+
+cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
+ uint32_t nonces, uint32_t blockWarps) {
+ if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
+ uint32_t block = 32u * blockWarps;
+ if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
+ igneum_hash<<>>(ds, out, baseNonce, mask);
+ return cudaGetLastError();
+}
+
+cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
+ cudaFuncAttributes attr;
+ cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash);
+ if (e != cudaSuccess) return e;
+ *numRegs = attr.numRegs;
+ return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0);
+}
diff --git a/proto-cuda/packs-ca2-mixer/mx8-devnet-epoch0/kernel_bound.cl b/proto-cuda/packs-ca2-mixer/mx8-devnet-epoch0/kernel_bound.cl
new file mode 100644
index 000000000..51b5d85d6
--- /dev/null
+++ b/proto-cuda/packs-ca2-mixer/mx8-devnet-epoch0/kernel_bound.cl
@@ -0,0 +1,373 @@
+// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
+// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal).
+// Built from source at runtime by proto-opencl/host.c, which passes these defines:
+// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit)
+// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default)
+// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32
+// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition
+// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units;
+// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32.
+#ifndef IGNEUM_GROUP
+#define IGNEUM_GROUP 32
+#endif
+#ifndef IGNEUM_EXCHANGE
+#define IGNEUM_EXCHANGE 0
+#endif
+#ifdef __OPENCL_VERSION__
+#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1)))
+#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n]
+#if IGNEUM_EXCHANGE == 1
+#ifdef cl_khr_subgroups
+#pragma OPENCL EXTENSION cl_khr_subgroups : enable
+#endif
+#ifdef cl_khr_subgroup_shuffle
+#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable
+#endif
+#elif IGNEUM_EXCHANGE == 2
+#pragma OPENCL EXTENSION cl_intel_subgroups : enable
+#endif
+#else
+// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros.
+#include "emu_opencl.h"
+#endif
+
+#if IGNEUM_EXCHANGE == 1
+#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m))
+#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
+#elif IGNEUM_EXCHANGE == 2
+#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m))
+#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
+#else
+// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per
+// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane
+// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's
+// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier.
+#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; }
+#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; }
+#endif
+
+static inline uint splitmix32(uint x) {
+ x ^= x >> 16; x *= 0x7feb352du;
+ x ^= x >> 15; x *= 0x846ca68bu;
+ x ^= x >> 16;
+ return x;
+}
+// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32.
+static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); }
+// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x.
+static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); }
+static inline uint ds_elem(uint i, uint d0, uint d1) {
+ uint x = i ^ d0;
+ x *= 0x9E3779B1u; x ^= x >> 15;
+ x += d1;
+ x *= 0x85EBCA77u; x ^= x >> 13;
+ x *= 0xC2B2AE3Du; x ^= x >> 16;
+ return x;
+}
+
+// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
+// Item: 8 rounds of 8 x seed-parameterised mixer + one 64-byte cache read, then 8 x final mixer (class v3, mixer multiplier 8,
+// docs/plans/mixer-x4.md: the round key of application j of round r is 0x9E3779B9 * (r * m + j + 1)). All parameters are literals.
+#define MH_CACHE_LINE_MASK 0x003fffffu
+#define MH_SEGMENT_LINES 64u
+#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
+static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
+
+// y = ChaCha12 core(x) + x
+static inline void mh_chacha_block(const uint* x, uint* y) {
+ for (uint i = 0u; i < 16u; ++i) y[i] = x[i];
+ for (uint r = 0u; r < 6u; ++r) {
+ MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
+ MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
+ MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
+ MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
+ }
+ for (uint i = 0u; i < 16u; ++i) y[i] += x[i];
+}
+
+// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
+static inline void mh_cache_segment(__global uint* cache, uint seg) {
+ uint prev[16]; uint x[16]; uint y[16];
+ for (uint i = 0u; i < 16u; ++i) prev[i] = 0u;
+ for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) {
+ x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
+ x[4] = 0xceed56d7u ^ prev[4];
+ x[5] = 0x9ba270d2u ^ prev[5];
+ x[6] = 0x82caab2du ^ prev[6];
+ x[7] = 0x81ebce0eu ^ prev[7];
+ x[8] = 0x12b6ecf1u ^ prev[8];
+ x[9] = 0xd0f3fd7cu ^ prev[9];
+ x[10] = 0xd872eefeu ^ prev[10];
+ x[11] = 0xc158c7bdu ^ prev[11];
+ x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
+ mh_chacha_block(x, y);
+ __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
+ for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
+ }
+}
+
+// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
+static inline void mh_mixer(uint* s, uint rk) {
+ s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
+ s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
+ s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
+ s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
+ s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
+ s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
+ s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
+ s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
+ s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
+ s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
+ s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
+ s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
+ s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
+ s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
+ s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
+ s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
+ MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
+ MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
+ MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
+ MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
+}
+
+// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of 8 x mixer + cache line s[0] & mask; 8 x final mixer.
+static inline void mh_item(__global const uint* cache, uint t, uint* s) {
+ s[0] = 0xceed56d7u;
+ s[1] = 0x9ba270d2u;
+ s[2] = 0x82caab2du;
+ s[3] = 0x81ebce0eu;
+ s[4] = 0x12b6ecf1u;
+ s[5] = 0xd0f3fd7cu;
+ s[6] = 0xd872eefeu;
+ s[7] = 0xc158c7bdu;
+ s[8] = t * 0xf351d601u + 0xc6892460u;
+ s[9] = t * 0xa3bb398fu + 0x25b7228au;
+ s[10] = t * 0xb5a09e35u + 0xcd515004u;
+ s[11] = t * 0x7509c9c1u + 0x2846527au;
+ s[12] = t * 0x6bbf31e9u + 0xa6324241u;
+ s[13] = t * 0xfc849a79u + 0x36e3ec53u;
+ s[14] = t * 0xded91851u + 0x82961bacu;
+ s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
+ for (uint r = 0u; r < 8u; ++r) {
+ for (uint j = 0u; j < 8u; ++j) mh_mixer(s, 0x9E3779B9u * (r * 8u + j + 1u));
+ __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
+ for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i];
+ }
+ for (uint j = 0u; j < 8u; ++j) mh_mixer(s, 0x9E3779B9u * (64u + j + 1u));
+}
+// dataset[w] without the dataset: derive item w >> 4 and take word w & 15.
+static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; }
+
+// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item.
+// The same constants as memhard.h in this pack (one emitter, three dialects).
+__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) {
+ uint seg = (uint)get_global_id(0);
+ if (seg < nSegments) mh_cache_segment(cache, seg);
+}
+__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) {
+ uint t = (uint)get_global_id(0);
+ if (t < nItems) {
+ uint s[16];
+ mh_item(cache, t, s);
+ __global uint* d = ds + ((ulong)t * 16u);
+ for (uint i = 0u; i < 16u; ++i) d[i] = s[i];
+ }
+}
+
+// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the
+// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and
+// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all).
+IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) {
+ uint gid = (uint)get_global_id(0);
+ uint lid = (uint)get_local_id(0);
+ uint nonce = baseNonce + gid;
+ uint r0, r1, r2, r3, r4, r5, r6, r7;
+#if IGNEUM_EXCHANGE == 0
+ IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
+ uint xk = 0u;
+#else
+ (void)lid;
+#endif
+ { uint x = nonce ^ 0x667d0fbdu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x7b8e5963u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
+ { uint x = nonce ^ 0x7b8e5963u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x31c67e5eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
+ { uint x = nonce ^ 0x31c67e5eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4529ddc6u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
+ { uint x = nonce ^ 0x4529ddc6u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef19d6d8u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
+ { uint x = nonce ^ 0xef19d6d8u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xaccf6211u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
+ { uint x = nonce ^ 0xaccf6211u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xda0aed32u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
+ { uint x = nonce ^ 0xda0aed32u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xabc6df31u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
+ { uint x = nonce ^ 0xabc6df31u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x667d0fbdu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
+
+ for (uint it = 0u; it < 8u; ++it) {
+ uint sel = r0;
+ r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
+ { uint t_; IGNEUM_SHFL_XOR(t_, r0, 4u); r2 = r2 ^ t_; } // 1 shfl
+ r3 = r3 + r2 + ((((sel >> 10u) & 1u) != 0u) ? 0x642e66dbu : 0x2cccb6cau); // 2 add
+ r0 = rotl_imm(r0, 19u); // 3 rotl
+ r7 = rotr_var(r7, r6); // 4 rotr
+ r7 = r7 + r4 + ((((sel >> 21u) & 1u) != 0u) ? 0xc1535555u : 0xee02465fu); // 5 add
+ r1 = mul_hi(r1, r7); // 6 mulhi
+ r4 = r4 ^ ds[r2 & mask]; // 7 load
+ r7 = r7 ^ ds[r4 & mask]; // 8 load
+ r0 = r0 ^ ds[r3 & mask]; // 9 load
+ r5 = r5 ^ ds[r1 & mask]; // 10 load
+ r1 = r1 ^ ds[r5 & mask]; // 11 load
+ r3 = mul_hi(r3, r5); // 12 mulhi
+ r1 = r1 ^ ds[r3 & mask]; // 13 load
+ r0 = r0 - r3; // 14 sub
+ r5 = r1 * r3 + r5; // 15 mad
+ r6 = mul_hi(r6, r1); // 16 mulhi
+ r5 = r5 + r2 + ((((sel >> 28u) & 1u) != 0u) ? 0x8b965b57u : 0x697b3d00u); // 17 add
+ r0 = mul_hi(r0, r6); // 18 mulhi
+ r5 = rotr_var(r5, r3); // 19 rotr
+ r5 = mul_hi(r5, r2); // 20 mulhi
+ r1 = r1 + r0 + ((((sel >> 1u) & 1u) != 0u) ? 0x6d7e8d05u : 0xebcf247au); // 21 add
+ r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 22 add
+ r1 = mul_hi(r1, r5); // 23 mulhi
+ r2 = r2 - r5; // 24 sub
+ r7 = r7 + r4 + ((((sel >> 2u) & 1u) != 0u) ? 0x699ef1bbu : 0x08ffa6c7u); // 25 add
+ { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 26 shfl
+ r7 = r7 + r1 + ((((sel >> 14u) & 1u) != 0u) ? 0xb4ead2fbu : 0xe60fea84u); // 27 add
+ r3 = r3 + r1 + ((((sel >> 6u) & 1u) != 0u) ? 0x8f30d21du : 0x65c76dabu); // 28 add
+ r2 = r2 ^ ds[r1 & mask]; // 29 load
+ r5 = r5 ^ ds[r7 & mask]; // 30 load
+ r2 = r2 ^ ds[r5 & mask]; // 31 load
+ { uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r1 = r1 ^ t_; } // 32 shfl
+ r4 = r5 * r7 + r4; // 33 mad
+ r4 = r4 + r2 + ((((sel >> 21u) & 1u) != 0u) ? 0xc7ce690cu : 0x0480debeu); // 34 add
+ { uint t_; IGNEUM_SHFL_XOR(t_, r7, 8u); r3 = r3 ^ t_; } // 35 shfl
+ r7 = r7 + r1 + ((((sel >> 2u) & 1u) != 0u) ? 0xe10c2c95u : 0xc53b542eu); // 36 add
+ r5 = r5 ^ r7; // 37 xor
+ r2 = r2 | r1; // 38 or
+ r1 = mul_hi(r1, r0); // 39 mulhi
+ r6 = rotl_imm(r6, 19u); // 40 rotl
+ r4 = mul_hi(r4, r6); // 41 mulhi
+ r6 = r6 - r0; // 42 sub
+ { uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r6 = r6 ^ t_; } // 43 shfl
+ r4 = r4 ^ ds[r2 & mask]; // 44 load
+ r1 = r1 ^ r3; // 45 xor
+ r7 = r7 ^ ds[r0 & mask]; // 46 load
+ r3 = r3 ^ ds[r1 & mask]; // 47 load
+ r5 = r5 * r3; // 48 mul
+ r1 = r1 - r5; // 49 sub
+ r2 = rotl_imm(r2, 8u); // 50 rotl
+ r1 = r1 + r5 + ((((sel >> 23u) & 1u) != 0u) ? 0x77b9bd43u : 0xa900fec4u); // 51 add
+ r4 = r4 ^ ds[r7 & mask]; // 52 load
+ r2 = r2 - r7; // 53 sub
+ r4 = r4 ^ r0; // 54 xor
+ r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 55 add
+ r2 = r2 ^ ds[r4 & mask]; // 56 load
+ r0 = r1 * r4 + r0; // 57 mad
+ r3 = r3 ^ ds[r5 & mask]; // 58 load
+ r5 = r5 | r6; // 59 or
+ r6 = r5 * r7 + r6; // 60 mad
+ r4 = rotl_imm(r4, 28u); // 61 rotl
+ r5 = mul_hi(r5, r0); // 62 mulhi
+ r3 = r3 ^ ds[r6 & mask]; // 63 load
+ }
+ uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
+ uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
+ out[gid] = ((ulong)hi << 32) | (ulong)lo;
+}
+
+#if IGNEUM_EXCHANGE != 0
+// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the
+// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a
+// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md.
+IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) {
+ if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); }
+}
+#endif
+
+// Header-bound variant (bind.rs): the init words come from initw, not SEEDW. Same body as igneum_hash.
+IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw) {
+ uint gid = (uint)get_global_id(0);
+ uint lid = (uint)get_local_id(0);
+ uint nonce = baseNonce + gid;
+ uint r0, r1, r2, r3, r4, r5, r6, r7;
+ uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7];
+#if IGNEUM_EXCHANGE == 0
+ IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
+ uint xk = 0u;
+#else
+ (void)lid;
+#endif
+ { uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; }
+ { uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; }
+ { uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; }
+ { uint x = nonce ^ iw3; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw4; }
+ { uint x = nonce ^ iw4; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw5; }
+ { uint x = nonce ^ iw5; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw6; }
+ { uint x = nonce ^ iw6; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw7; }
+ { uint x = nonce ^ iw7; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw0; }
+
+ for (uint it = 0u; it < 8u; ++it) {
+ uint sel = r0;
+ r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
+ { uint t_; IGNEUM_SHFL_XOR(t_, r0, 4u); r2 = r2 ^ t_; } // 1 shfl
+ r3 = r3 + r2 + ((((sel >> 10u) & 1u) != 0u) ? 0x642e66dbu : 0x2cccb6cau); // 2 add
+ r0 = rotl_imm(r0, 19u); // 3 rotl
+ r7 = rotr_var(r7, r6); // 4 rotr
+ r7 = r7 + r4 + ((((sel >> 21u) & 1u) != 0u) ? 0xc1535555u : 0xee02465fu); // 5 add
+ r1 = mul_hi(r1, r7); // 6 mulhi
+ r4 = r4 ^ ds[r2 & mask]; // 7 load
+ r7 = r7 ^ ds[r4 & mask]; // 8 load
+ r0 = r0 ^ ds[r3 & mask]; // 9 load
+ r5 = r5 ^ ds[r1 & mask]; // 10 load
+ r1 = r1 ^ ds[r5 & mask]; // 11 load
+ r3 = mul_hi(r3, r5); // 12 mulhi
+ r1 = r1 ^ ds[r3 & mask]; // 13 load
+ r0 = r0 - r3; // 14 sub
+ r5 = r1 * r3 + r5; // 15 mad
+ r6 = mul_hi(r6, r1); // 16 mulhi
+ r5 = r5 + r2 + ((((sel >> 28u) & 1u) != 0u) ? 0x8b965b57u : 0x697b3d00u); // 17 add
+ r0 = mul_hi(r0, r6); // 18 mulhi
+ r5 = rotr_var(r5, r3); // 19 rotr
+ r5 = mul_hi(r5, r2); // 20 mulhi
+ r1 = r1 + r0 + ((((sel >> 1u) & 1u) != 0u) ? 0x6d7e8d05u : 0xebcf247au); // 21 add
+ r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 22 add
+ r1 = mul_hi(r1, r5); // 23 mulhi
+ r2 = r2 - r5; // 24 sub
+ r7 = r7 + r4 + ((((sel >> 2u) & 1u) != 0u) ? 0x699ef1bbu : 0x08ffa6c7u); // 25 add
+ { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 26 shfl
+ r7 = r7 + r1 + ((((sel >> 14u) & 1u) != 0u) ? 0xb4ead2fbu : 0xe60fea84u); // 27 add
+ r3 = r3 + r1 + ((((sel >> 6u) & 1u) != 0u) ? 0x8f30d21du : 0x65c76dabu); // 28 add
+ r2 = r2 ^ ds[r1 & mask]; // 29 load
+ r5 = r5 ^ ds[r7 & mask]; // 30 load
+ r2 = r2 ^ ds[r5 & mask]; // 31 load
+ { uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r1 = r1 ^ t_; } // 32 shfl
+ r4 = r5 * r7 + r4; // 33 mad
+ r4 = r4 + r2 + ((((sel >> 21u) & 1u) != 0u) ? 0xc7ce690cu : 0x0480debeu); // 34 add
+ { uint t_; IGNEUM_SHFL_XOR(t_, r7, 8u); r3 = r3 ^ t_; } // 35 shfl
+ r7 = r7 + r1 + ((((sel >> 2u) & 1u) != 0u) ? 0xe10c2c95u : 0xc53b542eu); // 36 add
+ r5 = r5 ^ r7; // 37 xor
+ r2 = r2 | r1; // 38 or
+ r1 = mul_hi(r1, r0); // 39 mulhi
+ r6 = rotl_imm(r6, 19u); // 40 rotl
+ r4 = mul_hi(r4, r6); // 41 mulhi
+ r6 = r6 - r0; // 42 sub
+ { uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r6 = r6 ^ t_; } // 43 shfl
+ r4 = r4 ^ ds[r2 & mask]; // 44 load
+ r1 = r1 ^ r3; // 45 xor
+ r7 = r7 ^ ds[r0 & mask]; // 46 load
+ r3 = r3 ^ ds[r1 & mask]; // 47 load
+ r5 = r5 * r3; // 48 mul
+ r1 = r1 - r5; // 49 sub
+ r2 = rotl_imm(r2, 8u); // 50 rotl
+ r1 = r1 + r5 + ((((sel >> 23u) & 1u) != 0u) ? 0x77b9bd43u : 0xa900fec4u); // 51 add
+ r4 = r4 ^ ds[r7 & mask]; // 52 load
+ r2 = r2 - r7; // 53 sub
+ r4 = r4 ^ r0; // 54 xor
+ r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 55 add
+ r2 = r2 ^ ds[r4 & mask]; // 56 load
+ r0 = r1 * r4 + r0; // 57 mad
+ r3 = r3 ^ ds[r5 & mask]; // 58 load
+ r5 = r5 | r6; // 59 or
+ r6 = r5 * r7 + r6; // 60 mad
+ r4 = rotl_imm(r4, 28u); // 61 rotl
+ r5 = mul_hi(r5, r0); // 62 mulhi
+ r3 = r3 ^ ds[r6 & mask]; // 63 load
+ }
+ uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
+ uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
+ out[gid] = ((ulong)hi << 32) | (ulong)lo;
+}
diff --git a/proto-cuda/packs-ca2-mixer/mx8-devnet-epoch0/kernel_bound.cu b/proto-cuda/packs-ca2-mixer/mx8-devnet-epoch0/kernel_bound.cu
new file mode 100644
index 000000000..9a46b7480
--- /dev/null
+++ b/proto-cuda/packs-ca2-mixer/mx8-devnet-epoch0/kernel_bound.cu
@@ -0,0 +1,123 @@
+// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
+// Header-bound twin of igneum_hash in kernel.cu: the init words come from a kernel argument, not SEEDW.
+// Host declarations (also in program_bound.h if present):
+// struct IgneumInitWords { uint32_t w[8]; };
+// cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
+// IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps);
+// cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
+#include
+#include
+#include "program.h"
+
+struct IgneumInitWords { uint32_t w[8]; };
+
+__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
+ x ^= x >> 16; x *= 0x7feb352du;
+ x ^= x >> 15; x *= 0x846ca68bu;
+ x ^= x >> 16;
+ return x;
+}
+__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
+__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
+
+__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw) {
+ uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
+ uint32_t nonce = baseNonce + gid;
+ uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
+ { uint32_t x = nonce ^ iw.w[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw.w[1]; }
+ { uint32_t x = nonce ^ iw.w[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw.w[2]; }
+ { uint32_t x = nonce ^ iw.w[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw.w[3]; }
+ { uint32_t x = nonce ^ iw.w[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw.w[4]; }
+ { uint32_t x = nonce ^ iw.w[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw.w[5]; }
+ { uint32_t x = nonce ^ iw.w[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw.w[6]; }
+ { uint32_t x = nonce ^ iw.w[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw.w[7]; }
+ { uint32_t x = nonce ^ iw.w[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw.w[0]; }
+
+ for (uint32_t it = 0u; it < 8u; ++it) {
+ uint32_t sel = r0;
+ r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
+ r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r0, 4); // 1 shfl
+ r3 = r3 + r2 + ((((sel >> 10u) & 1u) != 0u) ? 0x642e66dbu : 0x2cccb6cau); // 2 add
+ r0 = rotl_imm(r0, 19u); // 3 rotl
+ r7 = rotr_var(r7, r6); // 4 rotr
+ r7 = r7 + r4 + ((((sel >> 21u) & 1u) != 0u) ? 0xc1535555u : 0xee02465fu); // 5 add
+ r1 = __umulhi(r1, r7); // 6 mulhi
+ r4 = r4 ^ ds[r2 & mask]; // 7 load
+ r7 = r7 ^ ds[r4 & mask]; // 8 load
+ r0 = r0 ^ ds[r3 & mask]; // 9 load
+ r5 = r5 ^ ds[r1 & mask]; // 10 load
+ r1 = r1 ^ ds[r5 & mask]; // 11 load
+ r3 = __umulhi(r3, r5); // 12 mulhi
+ r1 = r1 ^ ds[r3 & mask]; // 13 load
+ r0 = r0 - r3; // 14 sub
+ r5 = r1 * r3 + r5; // 15 mad
+ r6 = __umulhi(r6, r1); // 16 mulhi
+ r5 = r5 + r2 + ((((sel >> 28u) & 1u) != 0u) ? 0x8b965b57u : 0x697b3d00u); // 17 add
+ r0 = __umulhi(r0, r6); // 18 mulhi
+ r5 = rotr_var(r5, r3); // 19 rotr
+ r5 = __umulhi(r5, r2); // 20 mulhi
+ r1 = r1 + r0 + ((((sel >> 1u) & 1u) != 0u) ? 0x6d7e8d05u : 0xebcf247au); // 21 add
+ r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 22 add
+ r1 = __umulhi(r1, r5); // 23 mulhi
+ r2 = r2 - r5; // 24 sub
+ r7 = r7 + r4 + ((((sel >> 2u) & 1u) != 0u) ? 0x699ef1bbu : 0x08ffa6c7u); // 25 add
+ r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 26 shfl
+ r7 = r7 + r1 + ((((sel >> 14u) & 1u) != 0u) ? 0xb4ead2fbu : 0xe60fea84u); // 27 add
+ r3 = r3 + r1 + ((((sel >> 6u) & 1u) != 0u) ? 0x8f30d21du : 0x65c76dabu); // 28 add
+ r2 = r2 ^ ds[r1 & mask]; // 29 load
+ r5 = r5 ^ ds[r7 & mask]; // 30 load
+ r2 = r2 ^ ds[r5 & mask]; // 31 load
+ r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r7, 4); // 32 shfl
+ r4 = r5 * r7 + r4; // 33 mad
+ r4 = r4 + r2 + ((((sel >> 21u) & 1u) != 0u) ? 0xc7ce690cu : 0x0480debeu); // 34 add
+ r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r7, 8); // 35 shfl
+ r7 = r7 + r1 + ((((sel >> 2u) & 1u) != 0u) ? 0xe10c2c95u : 0xc53b542eu); // 36 add
+ r5 = r5 ^ r7; // 37 xor
+ r2 = r2 | r1; // 38 or
+ r1 = __umulhi(r1, r0); // 39 mulhi
+ r6 = rotl_imm(r6, 19u); // 40 rotl
+ r4 = __umulhi(r4, r6); // 41 mulhi
+ r6 = r6 - r0; // 42 sub
+ r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 43 shfl
+ r4 = r4 ^ ds[r2 & mask]; // 44 load
+ r1 = r1 ^ r3; // 45 xor
+ r7 = r7 ^ ds[r0 & mask]; // 46 load
+ r3 = r3 ^ ds[r1 & mask]; // 47 load
+ r5 = r5 * r3; // 48 mul
+ r1 = r1 - r5; // 49 sub
+ r2 = rotl_imm(r2, 8u); // 50 rotl
+ r1 = r1 + r5 + ((((sel >> 23u) & 1u) != 0u) ? 0x77b9bd43u : 0xa900fec4u); // 51 add
+ r4 = r4 ^ ds[r7 & mask]; // 52 load
+ r2 = r2 - r7; // 53 sub
+ r4 = r4 ^ r0; // 54 xor
+ r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 55 add
+ r2 = r2 ^ ds[r4 & mask]; // 56 load
+ r0 = r1 * r4 + r0; // 57 mad
+ r3 = r3 ^ ds[r5 & mask]; // 58 load
+ r5 = r5 | r6; // 59 or
+ r6 = r5 * r7 + r6; // 60 mad
+ r4 = rotl_imm(r4, 28u); // 61 rotl
+ r5 = __umulhi(r5, r0); // 62 mulhi
+ r3 = r3 ^ ds[r6 & mask]; // 63 load
+ }
+ uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
+ uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
+ out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
+}
+
+cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
+ IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps) {
+ if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
+ uint32_t block = 32u * blockWarps;
+ if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
+ igneum_hash_bound<<>>(ds, out, baseNonce, mask, iw);
+ return cudaGetLastError();
+}
+
+cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
+ cudaFuncAttributes attr;
+ cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash_bound);
+ if (e != cudaSuccess) return e;
+ *numRegs = attr.numRegs;
+ return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash_bound, (int)(32u * blockWarps), 0);
+}
diff --git a/proto-cuda/packs-ca2-mixer/mx8-devnet-epoch0/memhard.h b/proto-cuda/packs-ca2-mixer/mx8-devnet-epoch0/memhard.h
new file mode 100644
index 000000000..b16f7a4b7
--- /dev/null
+++ b/proto-cuda/packs-ca2-mixer/mx8-devnet-epoch0/memhard.h
@@ -0,0 +1,109 @@
+// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
+// Memory-hard dataset core, the same text that the Mac's Metal kernels and CPU verifier were checked against.
+// Included by kernel.cu (device), host.cu (host reference) and proto-opencl/host.c (C99 host reference).
+// See proto-metal/MEMHARD.md for the construction. kernel.cl carries the same text in OpenCL C.
+#pragma once
+#ifdef __cplusplus
+#include
+#else
+#include
+#endif
+#if defined(__CUDACC__)
+#define IGNEUM_HD __host__ __device__ __forceinline__
+#elif defined(_MSC_VER) && !defined(__cplusplus)
+#define IGNEUM_HD static __inline
+#else
+#define IGNEUM_HD static inline
+#endif
+// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
+// Item: 8 rounds of 8 x seed-parameterised mixer + one 64-byte cache read, then 8 x final mixer (class v3, mixer multiplier 8,
+// docs/plans/mixer-x4.md: the round key of application j of round r is 0x9E3779B9 * (r * m + j + 1)). All parameters are literals.
+#define MH_CACHE_LINE_MASK 0x003fffffu
+#define MH_SEGMENT_LINES 64u
+#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
+IGNEUM_HD uint32_t mh_rotl(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
+
+// y = ChaCha12 core(x) + x
+IGNEUM_HD void mh_chacha_block(const uint32_t* x, uint32_t* y) {
+ for (uint32_t i = 0u; i < 16u; ++i) y[i] = x[i];
+ for (uint32_t r = 0u; r < 6u; ++r) {
+ MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
+ MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
+ MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
+ MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
+ }
+ for (uint32_t i = 0u; i < 16u; ++i) y[i] += x[i];
+}
+
+// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
+IGNEUM_HD void mh_cache_segment(uint32_t* cache, uint32_t seg) {
+ uint32_t prev[16]; uint32_t x[16]; uint32_t y[16];
+ for (uint32_t i = 0u; i < 16u; ++i) prev[i] = 0u;
+ for (uint32_t j = 0u; j < MH_SEGMENT_LINES; ++j) {
+ x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
+ x[4] = 0xceed56d7u ^ prev[4];
+ x[5] = 0x9ba270d2u ^ prev[5];
+ x[6] = 0x82caab2du ^ prev[6];
+ x[7] = 0x81ebce0eu ^ prev[7];
+ x[8] = 0x12b6ecf1u ^ prev[8];
+ x[9] = 0xd0f3fd7cu ^ prev[9];
+ x[10] = 0xd872eefeu ^ prev[10];
+ x[11] = 0xc158c7bdu ^ prev[11];
+ x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
+ mh_chacha_block(x, y);
+ uint32_t* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
+ for (uint32_t i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
+ }
+}
+
+// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
+IGNEUM_HD void mh_mixer(uint32_t* s, uint32_t rk) {
+ s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
+ s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
+ s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
+ s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
+ s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
+ s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
+ s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
+ s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
+ s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
+ s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
+ s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
+ s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
+ s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
+ s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
+ s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
+ s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
+ MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
+ MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
+ MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
+ MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
+}
+
+// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of 8 x mixer + cache line s[0] & mask; 8 x final mixer.
+IGNEUM_HD void mh_item(const uint32_t* cache, uint32_t t, uint32_t* s) {
+ s[0] = 0xceed56d7u;
+ s[1] = 0x9ba270d2u;
+ s[2] = 0x82caab2du;
+ s[3] = 0x81ebce0eu;
+ s[4] = 0x12b6ecf1u;
+ s[5] = 0xd0f3fd7cu;
+ s[6] = 0xd872eefeu;
+ s[7] = 0xc158c7bdu;
+ s[8] = t * 0xf351d601u + 0xc6892460u;
+ s[9] = t * 0xa3bb398fu + 0x25b7228au;
+ s[10] = t * 0xb5a09e35u + 0xcd515004u;
+ s[11] = t * 0x7509c9c1u + 0x2846527au;
+ s[12] = t * 0x6bbf31e9u + 0xa6324241u;
+ s[13] = t * 0xfc849a79u + 0x36e3ec53u;
+ s[14] = t * 0xded91851u + 0x82961bacu;
+ s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
+ for (uint32_t r = 0u; r < 8u; ++r) {
+ for (uint32_t j = 0u; j < 8u; ++j) mh_mixer(s, 0x9E3779B9u * (r * 8u + j + 1u));
+ const uint32_t* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
+ for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= line[i];
+ }
+ for (uint32_t j = 0u; j < 8u; ++j) mh_mixer(s, 0x9E3779B9u * (64u + j + 1u));
+}
+// dataset[w] without the dataset: derive item w >> 4 and take word w & 15.
+IGNEUM_HD uint32_t mh_word(const uint32_t* cache, uint32_t w) { uint32_t s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; }
diff --git a/proto-cuda/packs-ca2-mixer/mx8-devnet-epoch0/memhard.metal b/proto-cuda/packs-ca2-mixer/mx8-devnet-epoch0/memhard.metal
new file mode 100644
index 000000000..d9fed7f50
--- /dev/null
+++ b/proto-cuda/packs-ca2-mixer/mx8-devnet-epoch0/memhard.metal
@@ -0,0 +1,107 @@
+#include
+using namespace metal;
+// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
+// Item: 8 rounds of 8 x seed-parameterised mixer + one 64-byte cache read, then 8 x final mixer (class v3, mixer multiplier 8,
+// docs/plans/mixer-x4.md: the round key of application j of round r is 0x9E3779B9 * (r * m + j + 1)). All parameters are literals.
+#define MH_CACHE_LINE_MASK 0x003fffffu
+#define MH_SEGMENT_LINES 64u
+#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
+inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
+
+// y = ChaCha12 core(x) + x
+inline void mh_chacha_block(const thread uint* x, thread uint* y) {
+ for (uint i = 0u; i < 16u; ++i) y[i] = x[i];
+ for (uint r = 0u; r < 6u; ++r) {
+ MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
+ MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
+ MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
+ MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
+ }
+ for (uint i = 0u; i < 16u; ++i) y[i] += x[i];
+}
+
+// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
+inline void mh_cache_segment(device uint* cache, uint seg) {
+ uint prev[16]; uint x[16]; uint y[16];
+ for (uint i = 0u; i < 16u; ++i) prev[i] = 0u;
+ for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) {
+ x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
+ x[4] = 0xceed56d7u ^ prev[4];
+ x[5] = 0x9ba270d2u ^ prev[5];
+ x[6] = 0x82caab2du ^ prev[6];
+ x[7] = 0x81ebce0eu ^ prev[7];
+ x[8] = 0x12b6ecf1u ^ prev[8];
+ x[9] = 0xd0f3fd7cu ^ prev[9];
+ x[10] = 0xd872eefeu ^ prev[10];
+ x[11] = 0xc158c7bdu ^ prev[11];
+ x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
+ mh_chacha_block(x, y);
+ device uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
+ for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
+ }
+}
+
+// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
+inline void mh_mixer(thread uint* s, uint rk) {
+ s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
+ s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
+ s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
+ s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
+ s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
+ s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
+ s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
+ s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
+ s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
+ s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
+ s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
+ s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
+ s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
+ s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
+ s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
+ s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
+ MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
+ MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
+ MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
+ MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
+}
+
+// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of 8 x mixer + cache line s[0] & mask; 8 x final mixer.
+inline void mh_item(device const uint* cache, uint t, thread uint* s) {
+ s[0] = 0xceed56d7u;
+ s[1] = 0x9ba270d2u;
+ s[2] = 0x82caab2du;
+ s[3] = 0x81ebce0eu;
+ s[4] = 0x12b6ecf1u;
+ s[5] = 0xd0f3fd7cu;
+ s[6] = 0xd872eefeu;
+ s[7] = 0xc158c7bdu;
+ s[8] = t * 0xf351d601u + 0xc6892460u;
+ s[9] = t * 0xa3bb398fu + 0x25b7228au;
+ s[10] = t * 0xb5a09e35u + 0xcd515004u;
+ s[11] = t * 0x7509c9c1u + 0x2846527au;
+ s[12] = t * 0x6bbf31e9u + 0xa6324241u;
+ s[13] = t * 0xfc849a79u + 0x36e3ec53u;
+ s[14] = t * 0xded91851u + 0x82961bacu;
+ s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
+ for (uint r = 0u; r < 8u; ++r) {
+ for (uint j = 0u; j < 8u; ++j) mh_mixer(s, 0x9E3779B9u * (r * 8u + j + 1u));
+ device const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
+ for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i];
+ }
+ for (uint j = 0u; j < 8u; ++j) mh_mixer(s, 0x9E3779B9u * (64u + j + 1u));
+}
+// dataset[w] without the dataset: derive item w >> 4 and take word w & 15.
+inline uint mh_word(device const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; }
+
+// One thread per segment (2^16 threads).
+kernel void igneum_cache_fill(device uint* cache [[buffer(0)]], uint gid [[thread_position_in_grid]]) {
+ mh_cache_segment(cache, gid);
+}
+// One thread per 64-byte item (dataset words / 16 threads).
+kernel void igneum_build(device const uint* cache [[buffer(0)]], device uint* dataset [[buffer(1)]],
+ uint gid [[thread_position_in_grid]]) {
+ uint s[16];
+ mh_item(cache, gid, s);
+ device uint* d = dataset + gid * 16u;
+ for (uint i = 0u; i < 16u; ++i) d[i] = s[i];
+}
diff --git a/proto-cuda/packs-ca2-mixer/mx8-devnet-epoch0/program.h b/proto-cuda/packs-ca2-mixer/mx8-devnet-epoch0/program.h
new file mode 100644
index 000000000..3eacb5594
--- /dev/null
+++ b/proto-cuda/packs-ca2-mixer/mx8-devnet-epoch0/program.h
@@ -0,0 +1,63 @@
+// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
+// Program metadata for host.cu plus the launch wrappers defined in kernel.cu.
+// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros.
+#pragma once
+#ifdef __cplusplus
+#include
+#else
+#include
+#endif
+#ifndef IGNEUM_NO_CUDA
+#include
+#endif
+
+#define IGNEUM_SEED_STRING "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000"
+#define IGNEUM_SEED_BYTES_HEX "edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07"
+#define IGNEUM_GENERATOR 2
+#define IGNEUM_PROGRAM_ATTEMPT 0
+#define IGNEUM_PROGRAM_ID 0xa696c464d9e9649bull
+#define IGNEUM_DAY_STRING "bytes:69676e65756d2d6461792ffa50000000000000"
+#define IGNEUM_DAY_BYTES_HEX "69676e65756d2d6461792ffa50000000000000"
+#define IGNEUM_DAY0 0xceed56d7u
+#define IGNEUM_DAY1 0x9ba270d2u
+#define IGNEUM_DATASET_LOG2 28
+#define IGNEUM_MASK 0x0fffffffu
+#define IGNEUM_LANES 32
+#define IGNEUM_ITERATIONS 8
+#define IGNEUM_INSTR_COUNT 64
+#define IGNEUM_LOADS_PER_HASH 128
+#define IGNEUM_WIDE_LOADS_PER_HASH 0
+#define IGNEUM_OP_MIX "load=16 add=13 mulhi=9 shfl=5 sub=5 mad=4 rotl=4 xor=3 or=2 rotr=2 mul=1"
+// Class v3 construction (Counter ASIC 2.0, 5 October 2026, docs/plans/mixer-x4.md): version 2 loads; the dataset item
+// derivation applies the mixer IGNEUM_MIXER_MULT times per round (memhard.h), and the cache follows the growth rule.
+#define IGNEUM_LOAD_CLASS "mx8"
+#define IGNEUM_CLASS_MIXER_MULT 8
+#define IGNEUM_CACHE_GROWTH 1 // 1: cache words = 2^(26 + doublings(day)), doublings = floor(log2(1 + day / 1460))
+#define IGNEUM_LOAD_SLOTS 16
+#define IGNEUM_LOAD_MIX { 100, 0, 0 }
+#define IGNEUM_LOAD_WIDTH_COUNTS { 16, 0, 0 } // loads of 4, 16, 64 bytes per program
+#define IGNEUM_BYTES_PER_HASH 512
+#define IGNEUM_FOLD_ROT 11
+#define IGNEUM_FOLD_MUL 0x9e3779b1u
+// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h)
+#define IGNEUM_DATASET_MODE 1
+
+#define IGNEUM_SEEDW_INIT { 0x667d0fbdu, 0x7b8e5963u, 0x31c67e5eu, 0x4529ddc6u, 0xef19d6d8u, 0xaccf6211u, 0xda0aed32u, 0xabc6df31u }
+#define IGNEUM_KEY_INIT { 0xceed56d7u, 0x9ba270d2u, 0x82caab2du, 0x81ebce0eu, 0x12b6ecf1u, 0xd0f3fd7cu, 0xd872eefeu, 0xc158c7bdu }
+#define IGNEUM_CACHE_LOG2_WORDS 26
+#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6
+#define IGNEUM_CACHE_SEGMENTS 65536u
+#define IGNEUM_ITEM_ROUNDS 8
+#define IGNEUM_MIXER_MULT 8 // mixer applications per round and after the last read (class v3, docs/plans/mixer-x4.md)
+#define IGNEUM_MIX_ROT_INIT { 17u, 12u, 20u, 23u, 7u, 3u, 27u, 16u }
+#define IGNEUM_MIX_MUL_INIT { 0xf351d601u, 0xa3bb398fu, 0xb5a09e35u, 0x7509c9c1u, 0x6bbf31e9u, 0xfc849a79u, 0xded91851u, 0x8d9113d1u, 0x0ff15225u, 0x3a5bdd41u, 0xab533435u, 0xe1c55ad5u, 0xe6d3bd0du, 0x9d9ffbbdu, 0xbb2a3cf3u, 0x50a7c08du }
+#define IGNEUM_MIX_RC_INIT { 0xc6892460u, 0x25b7228au, 0xcd515004u, 0x2846527au, 0xa6324241u, 0x36e3ec53u, 0x82961bacu, 0x0f97ba7du, 0xb6f921a9u, 0x3ada24e5u, 0xde20ab91u, 0x5378eeb2u, 0x7d161662u, 0x89353cc1u, 0xb1aa03a2u, 0x788acae6u }
+
+#ifndef IGNEUM_NO_CUDA
+// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError().
+cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments);
+cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems);
+cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
+ uint32_t nonces, uint32_t blockWarps);
+cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
+#endif
diff --git a/proto-cuda/packs-ca2-mixer/mx8-devnet-epoch0/program.json b/proto-cuda/packs-ca2-mixer/mx8-devnet-epoch0/program.json
new file mode 100644
index 000000000..3f423788f
--- /dev/null
+++ b/proto-cuda/packs-ca2-mixer/mx8-devnet-epoch0/program.json
@@ -0,0 +1,131 @@
+{
+ "format": "igneum-program-pack-3",
+ "generator": 2,
+ "attempt": 0,
+ "program_id": "0xa696c464d9e9649b",
+ "program_id_derivation": "FNV-1a 64 over 'igneum-program/' || generator_le32 || seed_words as little-endian bytes || attempt_le32",
+ "dataset_mode": "memory-hard",
+ "seed": "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000",
+ "seed_bytes": "edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07",
+ "seed_words": ["0x667d0fbd", "0x7b8e5963", "0x31c67e5e", "0x4529ddc6", "0xef19d6d8", "0xaccf6211", "0xda0aed32", "0xabc6df31"],
+ "seed_derivation": "seed_words = FNV-1a 64 over seed_bytes (attempt 0) or seed_bytes || attempt_le32 (attempt k >= 1), basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32",
+ "generator_rule": "version 2: exactly 16 load slots drawn first from instructions 1..63 (partial Fisher-Yates), the other 48 ops from the ten non-load weights (sum 75); a load's source is drawn from the registers other than dst written by an earlier instruction and not read by a load since; the candidate must pass the acceptance rule of spec 01 section 1.4.6 (static: no cyclically stale load source, every register has an injecting write; dynamic: 64 units on the seed-keyed closed-form dataset with no constant register bit, no lane-constant load site, under 164 saturated final values, every output bit within 136 of 1024, distinct addresses above 245760), else the next attempt of the seed is tried",
+ "lanes": 32,
+ "registers": 8,
+ "iterations": 8,
+ "instruction_count": 64,
+ "loads_per_hash": 128,
+ "load_class": "mx8",
+ "mixer_mult": 8,
+ "cache_growth": true,
+ "mixer": "class v3 (Counter ASIC 2.0, 5 October 2026, docs/plans/mixer-x4.md): every mixer application of the item derivation is 8 applications with round keys (r * 8 + j + 1) * 0x9E3779B9, the 8 dependent cache reads per item unchanged; cache growth rule option C: cache words = 2^(26 + doublings(day)), dataset words = 2^(genesis_log2 + doublings(day)), doublings(day) = floor(log2(1 + day / 1460)) for day = days since genesis",
+ "load_slots": 16,
+ "load_mix_percent_4_16_64": [100, 0, 0],
+ "load_width_counts_4_16_64": [16, 0, 0],
+ "bytes_per_hash": 512,
+ "wide_load": "read-width experiment (5 October 2026, docs/plans/read-width.md), NOT the lottery hash: a load of W words (width field, 4 or 16) reads dataset[b .. b + W) with b = (src & mask) & ~(W - 1) and folds every word into dst: x = dst ^ w[0]; for j in 1..W: x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x; width 1 is the plain load; the width is drawn per instruction from the class mix with one extra below(100) draw after the nine of version 2, and the program id is FNV-1a 64 over 'igneum-program-rw/' || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots",
+ "op_mix": {"load": 16, "add": 13, "mulhi": 9, "shfl": 5, "sub": 5, "mad": 4, "rotl": 4, "xor": 3, "or": 2, "rotr": 2, "mul": 1},
+ "register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]",
+ "splitmix32": "x ^= x>>16; x *= 0x7feb352d; x ^= x>>15; x *= 0x846ca68b; x ^= x>>16",
+ "iteration": "sel = r0 sampled once at the top of each iteration, then all instructions in order",
+ "output": "lo = r0 ^ rotl(r1,7) ^ rotl(r2,14) ^ rotl(r3,21); hi = r4 ^ rotl(r5,9) ^ rotl(r6,18) ^ rotl(r7,27); out = (hi << 32) | lo",
+ "op_semantics": {
+ "add": "dst = dst + src + (bit `bit` of sel ? imm2 : imm)",
+ "sub": "dst = dst - src",
+ "mul": "dst = dst * src (low 32)",
+ "mulhi": "dst = high 32 bits of dst * src",
+ "xor": "dst = dst ^ src",
+ "or": "dst = dst | src",
+ "rotl": "dst = rotl(dst, rot), rot in 1..31",
+ "rotr": "dst = rotr(dst, src & 31)",
+ "mad": "dst = src * src2 + dst",
+ "shfl": "dst = dst ^ (src of lane (lane ^ mask)), mask in {1,2,4,8,16}, within the 32-lane warp",
+ "load": "dst = dst ^ dataset[src & dataset.mask]",
+ "wload": "base = (src of lane 0 & dataset.mask) & ~31; dst = dst ^ dataset[base + lane] (warp-coalesced 128-byte load, lever b, only when --wide-frac > 0)"
+ },
+ "dataset": {
+ "log2_words": 28,
+ "bytes": 1073741824,
+ "mask": "0x0fffffff",
+ "day": "bytes:69676e65756d2d6461792ffa50000000000000",
+ "day_bytes": "69676e65756d2d6461792ffa50000000000000",
+ "day_words_from": "seed_words_from_bytes(day_bytes)",
+ "d0": "0xceed56d7",
+ "d1": "0x9ba270d2",
+ "mode": "memory-hard",
+ "spec": "proto-metal/MEMHARD.md",
+ "key": ["0xceed56d7", "0x9ba270d2", "0x82caab2d", "0x81ebce0e", "0x12b6ecf1", "0xd0f3fd7c", "0xd872eefe", "0xc158c7bd"],
+ "key_derivation": "the 8 words of seed_words_from_bytes(day_bytes); d0, d1 are key[0], key[1]",
+ "cache": {"log2_words": 26, "bytes": 268435456, "line_words": 16, "segment_lines": 64, "segments": 65536, "block": "ChaCha12 core + feed-forward, rotations 16 12 8 7", "sigma": ["0x61707865", "0x3320646e", "0x79622d32", "0x6b206574"], "tag": ["0x49676e65", "0x756d4d48"], "chain": "in_j = prev_line ^ (sigma[0..3] || key[0..7] || seg || j || tag[0..1]); line_j = block(in_j); prev_0 = 0"},
+ "mixer": {"draw": "SplitMix64 seeded with key[0] | key[1] << 32: rot[0..7] = 1 + next() % 31, mul[0..15] = low32(next()) | 1, rc[0..15] = low32(next())", "rot": [17, 12, 20, 23, 7, 3, 27, 16], "mul": ["0xf351d601", "0xa3bb398f", "0xb5a09e35", "0x7509c9c1", "0x6bbf31e9", "0xfc849a79", "0xded91851", "0x8d9113d1", "0x0ff15225", "0x3a5bdd41", "0xab533435", "0xe1c55ad5", "0xe6d3bd0d", "0x9d9ffbbd", "0xbb2a3cf3", "0x50a7c08d"], "rc": ["0xc6892460", "0x25b7228a", "0xcd515004", "0x2846527a", "0xa6324241", "0x36e3ec53", "0x82961bac", "0x0f97ba7d", "0xb6f921a9", "0x3ada24e5", "0xde20ab91", "0x5378eeb2", "0x7d161662", "0x89353cc1", "0xb1aa03a2", "0x788acae6"], "round": "for i in 0..15: s[i] = (s[i] ^ (rc[i] + (r+1) * 0x9E3779B9)) * mul[i]; then quarter rounds on columns (0,4,8,12) (1,5,9,13) (2,6,10,14) (3,7,11,15) with rot[0..3] and diagonals (0,5,10,15) (1,6,11,12) (2,7,8,13) (3,4,9,14) with rot[4..7]", "quarter_round": "a += b; d ^= a; d = rotl(d, r1); c += d; b ^= c; b = rotl(b, r2); a += b; d ^= a; d = rotl(d, r3); c += d; b ^= c; b = rotl(b, r4)"},
+ "mixer_mult": 8,
+ "item": "s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..7: for j in 0..7: s = M(s, rk = (r * 8 + j + 1) * 0x9E3779B9); line = s[0] & 0x003fffff; s[i] ^= cache[line * 16 + i]; then for j in 0..7: s = M(s, rk = (64 + j + 1) * 0x9E3779B9); item(t) = s",
+ "word": "dataset[w] = item(w >> 4)[w & 15]"
+ },
+ "instructions": [
+ {"i": 0, "op": "add", "dst": 4, "src": 5, "src2": 7, "imm": "0xea86e152", "imm2": "0x5810667a", "rot": 27, "bit": 13, "mask": 2, "width": 1},
+ {"i": 1, "op": "shfl", "dst": 2, "src": 0, "src2": 7, "imm": "0xe3c2f9cb", "imm2": "0xde0bea4e", "rot": 6, "bit": 9, "mask": 4, "width": 1},
+ {"i": 2, "op": "add", "dst": 3, "src": 2, "src2": 0, "imm": "0x2cccb6ca", "imm2": "0x642e66db", "rot": 27, "bit": 10, "mask": 2, "width": 1},
+ {"i": 3, "op": "rotl", "dst": 0, "src": 2, "src2": 1, "imm": "0xb59e83b2", "imm2": "0x19c26fb9", "rot": 19, "bit": 20, "mask": 4, "width": 1},
+ {"i": 4, "op": "rotr", "dst": 7, "src": 6, "src2": 5, "imm": "0x6f055f55", "imm2": "0x550e4ea1", "rot": 31, "bit": 20, "mask": 4, "width": 1},
+ {"i": 5, "op": "add", "dst": 7, "src": 4, "src2": 7, "imm": "0xee02465f", "imm2": "0xc1535555", "rot": 31, "bit": 21, "mask": 4, "width": 1},
+ {"i": 6, "op": "mulhi", "dst": 1, "src": 7, "src2": 5, "imm": "0x7826a6a7", "imm2": "0x946f7818", "rot": 18, "bit": 21, "mask": 8, "width": 1},
+ {"i": 7, "op": "load", "dst": 4, "src": 2, "src2": 0, "imm": "0x5d080878", "imm2": "0xdf885578", "rot": 5, "bit": 18, "mask": 4, "width": 1},
+ {"i": 8, "op": "load", "dst": 7, "src": 4, "src2": 1, "imm": "0x875bbb36", "imm2": "0x594a838f", "rot": 24, "bit": 4, "mask": 2, "width": 1},
+ {"i": 9, "op": "load", "dst": 0, "src": 3, "src2": 1, "imm": "0xe5e607c2", "imm2": "0xa5cd9f75", "rot": 26, "bit": 28, "mask": 4, "width": 1},
+ {"i": 10, "op": "load", "dst": 5, "src": 1, "src2": 3, "imm": "0xea3f7b43", "imm2": "0x10c8d4e7", "rot": 5, "bit": 29, "mask": 8, "width": 1},
+ {"i": 11, "op": "load", "dst": 1, "src": 5, "src2": 4, "imm": "0x4454980f", "imm2": "0xebd31581", "rot": 10, "bit": 28, "mask": 16, "width": 1},
+ {"i": 12, "op": "mulhi", "dst": 3, "src": 5, "src2": 7, "imm": "0xbfd5615c", "imm2": "0xd8224ae2", "rot": 21, "bit": 12, "mask": 2, "width": 1},
+ {"i": 13, "op": "load", "dst": 1, "src": 3, "src2": 5, "imm": "0x35c07cc5", "imm2": "0xe82db54f", "rot": 7, "bit": 6, "mask": 8, "width": 1},
+ {"i": 14, "op": "sub", "dst": 0, "src": 3, "src2": 0, "imm": "0x587ee0f1", "imm2": "0xfd23eefd", "rot": 16, "bit": 21, "mask": 16, "width": 1},
+ {"i": 15, "op": "mad", "dst": 5, "src": 1, "src2": 3, "imm": "0x6974dd29", "imm2": "0xc8148960", "rot": 3, "bit": 11, "mask": 2, "width": 1},
+ {"i": 16, "op": "mulhi", "dst": 6, "src": 1, "src2": 7, "imm": "0x3072c3c6", "imm2": "0x55ee21f8", "rot": 26, "bit": 1, "mask": 1, "width": 1},
+ {"i": 17, "op": "add", "dst": 5, "src": 2, "src2": 2, "imm": "0x697b3d00", "imm2": "0x8b965b57", "rot": 9, "bit": 28, "mask": 1, "width": 1},
+ {"i": 18, "op": "mulhi", "dst": 0, "src": 6, "src2": 3, "imm": "0x2910cacb", "imm2": "0x6ac79431", "rot": 7, "bit": 6, "mask": 4, "width": 1},
+ {"i": 19, "op": "rotr", "dst": 5, "src": 3, "src2": 0, "imm": "0xed96a94a", "imm2": "0x4c988c10", "rot": 24, "bit": 21, "mask": 8, "width": 1},
+ {"i": 20, "op": "mulhi", "dst": 5, "src": 2, "src2": 7, "imm": "0x40d2fc76", "imm2": "0x2f7c7eca", "rot": 13, "bit": 8, "mask": 4, "width": 1},
+ {"i": 21, "op": "add", "dst": 1, "src": 0, "src2": 5, "imm": "0xebcf247a", "imm2": "0x6d7e8d05", "rot": 31, "bit": 1, "mask": 16, "width": 1},
+ {"i": 22, "op": "add", "dst": 7, "src": 5, "src2": 7, "imm": "0xf66e7017", "imm2": "0xb9e3577e", "rot": 9, "bit": 12, "mask": 16, "width": 1},
+ {"i": 23, "op": "mulhi", "dst": 1, "src": 5, "src2": 7, "imm": "0xfdfe72fd", "imm2": "0x735eeb8d", "rot": 30, "bit": 25, "mask": 8, "width": 1},
+ {"i": 24, "op": "sub", "dst": 2, "src": 5, "src2": 0, "imm": "0x32cf1258", "imm2": "0xd813deb6", "rot": 30, "bit": 6, "mask": 16, "width": 1},
+ {"i": 25, "op": "add", "dst": 7, "src": 4, "src2": 2, "imm": "0x08ffa6c7", "imm2": "0x699ef1bb", "rot": 7, "bit": 2, "mask": 16, "width": 1},
+ {"i": 26, "op": "shfl", "dst": 3, "src": 4, "src2": 5, "imm": "0x6f53c70d", "imm2": "0x3357513f", "rot": 3, "bit": 26, "mask": 2, "width": 1},
+ {"i": 27, "op": "add", "dst": 7, "src": 1, "src2": 4, "imm": "0xe60fea84", "imm2": "0xb4ead2fb", "rot": 14, "bit": 14, "mask": 1, "width": 1},
+ {"i": 28, "op": "add", "dst": 3, "src": 1, "src2": 0, "imm": "0x65c76dab", "imm2": "0x8f30d21d", "rot": 24, "bit": 6, "mask": 1, "width": 1},
+ {"i": 29, "op": "load", "dst": 2, "src": 1, "src2": 2, "imm": "0x82fad9a6", "imm2": "0x8c6358db", "rot": 7, "bit": 31, "mask": 2, "width": 1},
+ {"i": 30, "op": "load", "dst": 5, "src": 7, "src2": 3, "imm": "0x6e947ee0", "imm2": "0xaf9a2dda", "rot": 2, "bit": 30, "mask": 4, "width": 1},
+ {"i": 31, "op": "load", "dst": 2, "src": 5, "src2": 1, "imm": "0x608bb7ce", "imm2": "0x4be663db", "rot": 19, "bit": 1, "mask": 16, "width": 1},
+ {"i": 32, "op": "shfl", "dst": 1, "src": 7, "src2": 6, "imm": "0x88e52e20", "imm2": "0x77647269", "rot": 20, "bit": 25, "mask": 4, "width": 1},
+ {"i": 33, "op": "mad", "dst": 4, "src": 5, "src2": 7, "imm": "0xb48420ae", "imm2": "0x3d1f2485", "rot": 14, "bit": 28, "mask": 1, "width": 1},
+ {"i": 34, "op": "add", "dst": 4, "src": 2, "src2": 3, "imm": "0x0480debe", "imm2": "0xc7ce690c", "rot": 12, "bit": 21, "mask": 16, "width": 1},
+ {"i": 35, "op": "shfl", "dst": 3, "src": 7, "src2": 5, "imm": "0xa73f59de", "imm2": "0x84b9e329", "rot": 21, "bit": 27, "mask": 8, "width": 1},
+ {"i": 36, "op": "add", "dst": 7, "src": 1, "src2": 7, "imm": "0xc53b542e", "imm2": "0xe10c2c95", "rot": 21, "bit": 2, "mask": 4, "width": 1},
+ {"i": 37, "op": "xor", "dst": 5, "src": 7, "src2": 4, "imm": "0x81cd7b0e", "imm2": "0x21a51823", "rot": 12, "bit": 10, "mask": 2, "width": 1},
+ {"i": 38, "op": "or", "dst": 2, "src": 1, "src2": 3, "imm": "0x7894e657", "imm2": "0xf8e4b972", "rot": 18, "bit": 9, "mask": 4, "width": 1},
+ {"i": 39, "op": "mulhi", "dst": 1, "src": 0, "src2": 4, "imm": "0xbee8421f", "imm2": "0x070888a8", "rot": 20, "bit": 28, "mask": 4, "width": 1},
+ {"i": 40, "op": "rotl", "dst": 6, "src": 1, "src2": 4, "imm": "0x4609857a", "imm2": "0xaeecb156", "rot": 19, "bit": 21, "mask": 1, "width": 1},
+ {"i": 41, "op": "mulhi", "dst": 4, "src": 6, "src2": 0, "imm": "0x1a84e1e9", "imm2": "0x9b26bb72", "rot": 28, "bit": 19, "mask": 1, "width": 1},
+ {"i": 42, "op": "sub", "dst": 6, "src": 0, "src2": 6, "imm": "0xbf62908e", "imm2": "0xdf03ea88", "rot": 11, "bit": 27, "mask": 16, "width": 1},
+ {"i": 43, "op": "shfl", "dst": 6, "src": 3, "src2": 4, "imm": "0xa966241c", "imm2": "0x9c639efa", "rot": 12, "bit": 4, "mask": 4, "width": 1},
+ {"i": 44, "op": "load", "dst": 4, "src": 2, "src2": 3, "imm": "0x7b5b5474", "imm2": "0x45cfc5dd", "rot": 17, "bit": 18, "mask": 8, "width": 1},
+ {"i": 45, "op": "xor", "dst": 1, "src": 3, "src2": 6, "imm": "0x006193d0", "imm2": "0xfc2acc3f", "rot": 25, "bit": 1, "mask": 1, "width": 1},
+ {"i": 46, "op": "load", "dst": 7, "src": 0, "src2": 6, "imm": "0x4777c4f8", "imm2": "0x3cf0a02f", "rot": 11, "bit": 14, "mask": 4, "width": 1},
+ {"i": 47, "op": "load", "dst": 3, "src": 1, "src2": 3, "imm": "0x10692532", "imm2": "0x1a292ea5", "rot": 20, "bit": 23, "mask": 1, "width": 1},
+ {"i": 48, "op": "mul", "dst": 5, "src": 3, "src2": 7, "imm": "0xadf5bd13", "imm2": "0xb999de2e", "rot": 23, "bit": 10, "mask": 4, "width": 1},
+ {"i": 49, "op": "sub", "dst": 1, "src": 5, "src2": 4, "imm": "0x68ff101e", "imm2": "0xbdaaf46a", "rot": 25, "bit": 23, "mask": 16, "width": 1},
+ {"i": 50, "op": "rotl", "dst": 2, "src": 6, "src2": 4, "imm": "0x92d9a412", "imm2": "0x0daf96ea", "rot": 8, "bit": 4, "mask": 4, "width": 1},
+ {"i": 51, "op": "add", "dst": 1, "src": 5, "src2": 0, "imm": "0xa900fec4", "imm2": "0x77b9bd43", "rot": 6, "bit": 23, "mask": 2, "width": 1},
+ {"i": 52, "op": "load", "dst": 4, "src": 7, "src2": 1, "imm": "0x51392a72", "imm2": "0x99e8bb36", "rot": 11, "bit": 9, "mask": 1, "width": 1},
+ {"i": 53, "op": "sub", "dst": 2, "src": 7, "src2": 2, "imm": "0x0ffe2ac7", "imm2": "0x030743df", "rot": 9, "bit": 30, "mask": 16, "width": 1},
+ {"i": 54, "op": "xor", "dst": 4, "src": 0, "src2": 4, "imm": "0x9123ff15", "imm2": "0x10c329a7", "rot": 28, "bit": 15, "mask": 2, "width": 1},
+ {"i": 55, "op": "add", "dst": 1, "src": 6, "src2": 5, "imm": "0xe09f54e9", "imm2": "0x83e825bf", "rot": 23, "bit": 14, "mask": 8, "width": 1},
+ {"i": 56, "op": "load", "dst": 2, "src": 4, "src2": 3, "imm": "0x239de52c", "imm2": "0xf80bae18", "rot": 15, "bit": 24, "mask": 1, "width": 1},
+ {"i": 57, "op": "mad", "dst": 0, "src": 1, "src2": 4, "imm": "0x31dede8e", "imm2": "0xd3f619e6", "rot": 29, "bit": 7, "mask": 2, "width": 1},
+ {"i": 58, "op": "load", "dst": 3, "src": 5, "src2": 3, "imm": "0x206437d6", "imm2": "0x28d1c290", "rot": 17, "bit": 28, "mask": 4, "width": 1},
+ {"i": 59, "op": "or", "dst": 5, "src": 6, "src2": 3, "imm": "0x8fffd674", "imm2": "0x0507903a", "rot": 26, "bit": 27, "mask": 2, "width": 1},
+ {"i": 60, "op": "mad", "dst": 6, "src": 5, "src2": 7, "imm": "0xf572bdb9", "imm2": "0xeda2af31", "rot": 21, "bit": 8, "mask": 2, "width": 1},
+ {"i": 61, "op": "rotl", "dst": 4, "src": 2, "src2": 7, "imm": "0x84f12ddf", "imm2": "0x81ef22e1", "rot": 28, "bit": 30, "mask": 1, "width": 1},
+ {"i": 62, "op": "mulhi", "dst": 5, "src": 0, "src2": 6, "imm": "0xf6bb45ee", "imm2": "0x6bfb632d", "rot": 22, "bit": 0, "mask": 4, "width": 1},
+ {"i": 63, "op": "load", "dst": 3, "src": 6, "src2": 6, "imm": "0x50ec702a", "imm2": "0xae6ee96e", "rot": 20, "bit": 25, "mask": 2, "width": 1}
+ ]
+}
diff --git a/proto-cuda/packs-ca2-mixer/mx8-devnet-epoch0/program.metal b/proto-cuda/packs-ca2-mixer/mx8-devnet-epoch0/program.metal
new file mode 100644
index 000000000..b899adecb
--- /dev/null
+++ b/proto-cuda/packs-ca2-mixer/mx8-devnet-epoch0/program.metal
@@ -0,0 +1,109 @@
+#include
+using namespace metal;
+
+#define MASK 0x0fffffffu
+constant uint SEEDW[8] = { 0x667d0fbdu, 0x7b8e5963u, 0x31c67e5eu, 0x4529ddc6u, 0xef19d6d8u, 0xaccf6211u, 0xda0aed32u, 0xabc6df31u };
+
+inline uint splitmix32(uint x) {
+ x ^= x >> 16; x *= 0x7feb352du;
+ x ^= x >> 15; x *= 0x846ca68bu;
+ x ^= x >> 16;
+ return x;
+}
+inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31
+inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
+inline uint ds_elem(uint i, uint d0, uint d1) {
+ uint x = i ^ d0;
+ x *= 0x9E3779B1u; x ^= x >> 15;
+ x += d1;
+ x *= 0x85EBCA77u; x ^= x >> 13;
+ x *= 0xC2B2AE3Du; x ^= x >> 16;
+ return x;
+}
+
+kernel void igneum_hash(device const uint* dataset [[buffer(0)]],
+ device ulong* out [[buffer(1)]],
+ constant uint& baseNonce [[buffer(2)]],
+ uint gid [[thread_position_in_grid]]) {
+ uint nonce = baseNonce + gid;
+ uint r0, r1, r2, r3, r4, r5, r6, r7;
+ { uint x = nonce ^ SEEDW[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ SEEDW[1]; }
+ { uint x = nonce ^ SEEDW[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ SEEDW[2]; }
+ { uint x = nonce ^ SEEDW[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ SEEDW[3]; }
+ { uint x = nonce ^ SEEDW[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ SEEDW[4]; }
+ { uint x = nonce ^ SEEDW[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ SEEDW[5]; }
+ { uint x = nonce ^ SEEDW[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ SEEDW[6]; }
+ { uint x = nonce ^ SEEDW[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ SEEDW[7]; }
+ { uint x = nonce ^ SEEDW[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ SEEDW[0]; }
+
+ for (uint it = 0u; it < 8u; ++it) {
+ uint sel = r0;
+ r4 = r4 + r5 + select(0xea86e152u, 0x5810667au, ((sel >> 13u) & 1u) != 0u); // 0
+ r2 = r2 ^ simd_shuffle_xor(r0, (ushort)4); // 1
+ r3 = r3 + r2 + select(0x2cccb6cau, 0x642e66dbu, ((sel >> 10u) & 1u) != 0u); // 2
+ r0 = rotl_imm(r0, 19u); // 3
+ r7 = rotr_var(r7, r6); // 4
+ r7 = r7 + r4 + select(0xee02465fu, 0xc1535555u, ((sel >> 21u) & 1u) != 0u); // 5
+ r1 = mulhi(r1, r7); // 6
+ r4 = r4 ^ dataset[r2 & MASK]; // 7
+ r7 = r7 ^ dataset[r4 & MASK]; // 8
+ r0 = r0 ^ dataset[r3 & MASK]; // 9
+ r5 = r5 ^ dataset[r1 & MASK]; // 10
+ r1 = r1 ^ dataset[r5 & MASK]; // 11
+ r3 = mulhi(r3, r5); // 12
+ r1 = r1 ^ dataset[r3 & MASK]; // 13
+ r0 = r0 - r3; // 14
+ r5 = r1 * r3 + r5; // 15
+ r6 = mulhi(r6, r1); // 16
+ r5 = r5 + r2 + select(0x697b3d00u, 0x8b965b57u, ((sel >> 28u) & 1u) != 0u); // 17
+ r0 = mulhi(r0, r6); // 18
+ r5 = rotr_var(r5, r3); // 19
+ r5 = mulhi(r5, r2); // 20
+ r1 = r1 + r0 + select(0xebcf247au, 0x6d7e8d05u, ((sel >> 1u) & 1u) != 0u); // 21
+ r7 = r7 + r5 + select(0xf66e7017u, 0xb9e3577eu, ((sel >> 12u) & 1u) != 0u); // 22
+ r1 = mulhi(r1, r5); // 23
+ r2 = r2 - r5; // 24
+ r7 = r7 + r4 + select(0x08ffa6c7u, 0x699ef1bbu, ((sel >> 2u) & 1u) != 0u); // 25
+ r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 26
+ r7 = r7 + r1 + select(0xe60fea84u, 0xb4ead2fbu, ((sel >> 14u) & 1u) != 0u); // 27
+ r3 = r3 + r1 + select(0x65c76dabu, 0x8f30d21du, ((sel >> 6u) & 1u) != 0u); // 28
+ r2 = r2 ^ dataset[r1 & MASK]; // 29
+ r5 = r5 ^ dataset[r7 & MASK]; // 30
+ r2 = r2 ^ dataset[r5 & MASK]; // 31
+ r1 = r1 ^ simd_shuffle_xor(r7, (ushort)4); // 32
+ r4 = r5 * r7 + r4; // 33
+ r4 = r4 + r2 + select(0x0480debeu, 0xc7ce690cu, ((sel >> 21u) & 1u) != 0u); // 34
+ r3 = r3 ^ simd_shuffle_xor(r7, (ushort)8); // 35
+ r7 = r7 + r1 + select(0xc53b542eu, 0xe10c2c95u, ((sel >> 2u) & 1u) != 0u); // 36
+ r5 = r5 ^ r7; // 37
+ r2 = r2 | r1; // 38
+ r1 = mulhi(r1, r0); // 39
+ r6 = rotl_imm(r6, 19u); // 40
+ r4 = mulhi(r4, r6); // 41
+ r6 = r6 - r0; // 42
+ r6 = r6 ^ simd_shuffle_xor(r3, (ushort)4); // 43
+ r4 = r4 ^ dataset[r2 & MASK]; // 44
+ r1 = r1 ^ r3; // 45
+ r7 = r7 ^ dataset[r0 & MASK]; // 46
+ r3 = r3 ^ dataset[r1 & MASK]; // 47
+ r5 = r5 * r3; // 48
+ r1 = r1 - r5; // 49
+ r2 = rotl_imm(r2, 8u); // 50
+ r1 = r1 + r5 + select(0xa900fec4u, 0x77b9bd43u, ((sel >> 23u) & 1u) != 0u); // 51
+ r4 = r4 ^ dataset[r7 & MASK]; // 52
+ r2 = r2 - r7; // 53
+ r4 = r4 ^ r0; // 54
+ r1 = r1 + r6 + select(0xe09f54e9u, 0x83e825bfu, ((sel >> 14u) & 1u) != 0u); // 55
+ r2 = r2 ^ dataset[r4 & MASK]; // 56
+ r0 = r1 * r4 + r0; // 57
+ r3 = r3 ^ dataset[r5 & MASK]; // 58
+ r5 = r5 | r6; // 59
+ r6 = r5 * r7 + r6; // 60
+ r4 = rotl_imm(r4, 28u); // 61
+ r5 = mulhi(r5, r0); // 62
+ r3 = r3 ^ dataset[r6 & MASK]; // 63
+ }
+ uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
+ uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
+ out[gid] = ((ulong)hi << 32) | (ulong)lo;
+}
diff --git a/proto-cuda/packs-ca2-mixer/mx8-devnet-epoch0/program_bound.metal b/proto-cuda/packs-ca2-mixer/mx8-devnet-epoch0/program_bound.metal
new file mode 100644
index 000000000..2f0218149
--- /dev/null
+++ b/proto-cuda/packs-ca2-mixer/mx8-devnet-epoch0/program_bound.metal
@@ -0,0 +1,111 @@
+#include
+using namespace metal;
+
+#define MASK 0x0fffffffu
+constant uint SEEDW[8] = { 0x667d0fbdu, 0x7b8e5963u, 0x31c67e5eu, 0x4529ddc6u, 0xef19d6d8u, 0xaccf6211u, 0xda0aed32u, 0xabc6df31u };
+
+inline uint splitmix32(uint x) {
+ x ^= x >> 16; x *= 0x7feb352du;
+ x ^= x >> 15; x *= 0x846ca68bu;
+ x ^= x >> 16;
+ return x;
+}
+inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31
+inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
+inline uint ds_elem(uint i, uint d0, uint d1) {
+ uint x = i ^ d0;
+ x *= 0x9E3779B1u; x ^= x >> 15;
+ x += d1;
+ x *= 0x85EBCA77u; x ^= x >> 13;
+ x *= 0xC2B2AE3Du; x ^= x >> 16;
+ return x;
+}
+
+// Header-bound variant: the init words come from buffer 3 (bind.rs), not from SEEDW.
+kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]],
+ device ulong* out [[buffer(1)]],
+ constant uint& baseNonce [[buffer(2)]],
+ constant uint* initw [[buffer(3)]],
+ uint gid [[thread_position_in_grid]]) {
+ uint nonce = baseNonce + gid;
+ uint r0, r1, r2, r3, r4, r5, r6, r7;
+ { uint x = nonce ^ initw[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ initw[1]; }
+ { uint x = nonce ^ initw[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ initw[2]; }
+ { uint x = nonce ^ initw[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ initw[3]; }
+ { uint x = nonce ^ initw[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ initw[4]; }
+ { uint x = nonce ^ initw[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ initw[5]; }
+ { uint x = nonce ^ initw[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ initw[6]; }
+ { uint x = nonce ^ initw[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ initw[7]; }
+ { uint x = nonce ^ initw[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ initw[0]; }
+
+ for (uint it = 0u; it < 8u; ++it) {
+ uint sel = r0;
+ r4 = r4 + r5 + select(0xea86e152u, 0x5810667au, ((sel >> 13u) & 1u) != 0u); // 0
+ r2 = r2 ^ simd_shuffle_xor(r0, (ushort)4); // 1
+ r3 = r3 + r2 + select(0x2cccb6cau, 0x642e66dbu, ((sel >> 10u) & 1u) != 0u); // 2
+ r0 = rotl_imm(r0, 19u); // 3
+ r7 = rotr_var(r7, r6); // 4
+ r7 = r7 + r4 + select(0xee02465fu, 0xc1535555u, ((sel >> 21u) & 1u) != 0u); // 5
+ r1 = mulhi(r1, r7); // 6
+ r4 = r4 ^ dataset[r2 & MASK]; // 7
+ r7 = r7 ^ dataset[r4 & MASK]; // 8
+ r0 = r0 ^ dataset[r3 & MASK]; // 9
+ r5 = r5 ^ dataset[r1 & MASK]; // 10
+ r1 = r1 ^ dataset[r5 & MASK]; // 11
+ r3 = mulhi(r3, r5); // 12
+ r1 = r1 ^ dataset[r3 & MASK]; // 13
+ r0 = r0 - r3; // 14
+ r5 = r1 * r3 + r5; // 15
+ r6 = mulhi(r6, r1); // 16
+ r5 = r5 + r2 + select(0x697b3d00u, 0x8b965b57u, ((sel >> 28u) & 1u) != 0u); // 17
+ r0 = mulhi(r0, r6); // 18
+ r5 = rotr_var(r5, r3); // 19
+ r5 = mulhi(r5, r2); // 20
+ r1 = r1 + r0 + select(0xebcf247au, 0x6d7e8d05u, ((sel >> 1u) & 1u) != 0u); // 21
+ r7 = r7 + r5 + select(0xf66e7017u, 0xb9e3577eu, ((sel >> 12u) & 1u) != 0u); // 22
+ r1 = mulhi(r1, r5); // 23
+ r2 = r2 - r5; // 24
+ r7 = r7 + r4 + select(0x08ffa6c7u, 0x699ef1bbu, ((sel >> 2u) & 1u) != 0u); // 25
+ r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 26
+ r7 = r7 + r1 + select(0xe60fea84u, 0xb4ead2fbu, ((sel >> 14u) & 1u) != 0u); // 27
+ r3 = r3 + r1 + select(0x65c76dabu, 0x8f30d21du, ((sel >> 6u) & 1u) != 0u); // 28
+ r2 = r2 ^ dataset[r1 & MASK]; // 29
+ r5 = r5 ^ dataset[r7 & MASK]; // 30
+ r2 = r2 ^ dataset[r5 & MASK]; // 31
+ r1 = r1 ^ simd_shuffle_xor(r7, (ushort)4); // 32
+ r4 = r5 * r7 + r4; // 33
+ r4 = r4 + r2 + select(0x0480debeu, 0xc7ce690cu, ((sel >> 21u) & 1u) != 0u); // 34
+ r3 = r3 ^ simd_shuffle_xor(r7, (ushort)8); // 35
+ r7 = r7 + r1 + select(0xc53b542eu, 0xe10c2c95u, ((sel >> 2u) & 1u) != 0u); // 36
+ r5 = r5 ^ r7; // 37
+ r2 = r2 | r1; // 38
+ r1 = mulhi(r1, r0); // 39
+ r6 = rotl_imm(r6, 19u); // 40
+ r4 = mulhi(r4, r6); // 41
+ r6 = r6 - r0; // 42
+ r6 = r6 ^ simd_shuffle_xor(r3, (ushort)4); // 43
+ r4 = r4 ^ dataset[r2 & MASK]; // 44
+ r1 = r1 ^ r3; // 45
+ r7 = r7 ^ dataset[r0 & MASK]; // 46
+ r3 = r3 ^ dataset[r1 & MASK]; // 47
+ r5 = r5 * r3; // 48
+ r1 = r1 - r5; // 49
+ r2 = rotl_imm(r2, 8u); // 50
+ r1 = r1 + r5 + select(0xa900fec4u, 0x77b9bd43u, ((sel >> 23u) & 1u) != 0u); // 51
+ r4 = r4 ^ dataset[r7 & MASK]; // 52
+ r2 = r2 - r7; // 53
+ r4 = r4 ^ r0; // 54
+ r1 = r1 + r6 + select(0xe09f54e9u, 0x83e825bfu, ((sel >> 14u) & 1u) != 0u); // 55
+ r2 = r2 ^ dataset[r4 & MASK]; // 56
+ r0 = r1 * r4 + r0; // 57
+ r3 = r3 ^ dataset[r5 & MASK]; // 58
+ r5 = r5 | r6; // 59
+ r6 = r5 * r7 + r6; // 60
+ r4 = rotl_imm(r4, 28u); // 61
+ r5 = mulhi(r5, r0); // 62
+ r3 = r3 ^ dataset[r6 & MASK]; // 63
+ }
+ uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
+ uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
+ out[gid] = ((ulong)hi << 32) | (ulong)lo;
+}
diff --git a/proto-cuda/packs-ca2-mixer/mx8-devnet-epoch0/vectors.h b/proto-cuda/packs-ca2-mixer/mx8-devnet-epoch0/vectors.h
new file mode 100644
index 000000000..f78994362
--- /dev/null
+++ b/proto-cuda/packs-ca2-mixer/mx8-devnet-epoch0/vectors.h
@@ -0,0 +1,57 @@
+// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
+// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset
+#pragma once
+#ifdef __cplusplus
+#include
+#else
+#include
+#endif
+
+#define IGNEUM_VEC_WARPS 3
+static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u };
+static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = {
+ { // base nonce 0
+ 0x731250f1008c6f8eull, 0xb89b396aa5e33597ull, 0xbc96ebcdb47c61e4ull, 0xac1414cd0fb91152ull, 0x1cc2d1847875555eull, 0x071f703c82efff0cull, 0x21147b8c58bc96e4ull, 0xf412422862c643e6ull,
+ 0xb8a5c330668e9283ull, 0x994558a2b24a6557ull, 0xb2fb321e884a1fedull, 0xa8f928a38622b10eull, 0x9ea4cfcfc2c96d8bull, 0xc082abc94d5dfb18ull, 0x695239a4a852e0f9ull, 0xef804c526240ce15ull,
+ 0x624303d31812df0dull, 0xcfd82ed4955f140aull, 0x6569ae6a4a436a9aull, 0xc17f46cf926a7a0cull, 0x721fe9c93267d13bull, 0xf09fe97d5ce6bfbbull, 0x455d853d3913d6a3ull, 0x4fbb5f43bd1af7eaull,
+ 0x864d689410bb8ed0ull, 0xe8c93188d7f02b04ull, 0xebaafb6120560360ull, 0xde81676bff855560ull, 0xd0b3ae8111aab7c3ull, 0xacd97035767867caull, 0x1970c0b030b99c03ull, 0x09fed95e63215205ull
+ },
+ { // base nonce 4096
+ 0x3d4f7e76fa55a96aull, 0xd54574ccb1a27843ull, 0xeb1635ddd5f1e181ull, 0x1e49a6cf24085bd9ull, 0x9f929d79ad731c0eull, 0x5caca654a4bd4f0full, 0x33eb9defb51424b9ull, 0x04f1f3c93663c27eull,
+ 0xad1c775687a46553ull, 0xc57c600167bc7571ull, 0xaf0b8e2f9a316b22ull, 0xa01088974b376d8eull, 0x14555ccee9288fddull, 0x0d1d576ea396fbe7ull, 0x8cf63b15a58299bfull, 0x49f530db99fd5f01ull,
+ 0x30324c49b12da61bull, 0xeb6fc871f103ad7bull, 0x8ac3271ce7edfc29ull, 0x2804099c65265affull, 0xc4d9377a427c3337ull, 0x566c1e580b940448ull, 0x790b966beb7bebb3ull, 0xba5da14cc619efe0ull,
+ 0x5da22ad9ae88fcd7ull, 0x24221e506877ef84ull, 0xfcda89486e6e78caull, 0xfd9a09b3587ba8fbull, 0x4fdb52f5b31ca54aull, 0x9cf4a2730edbf331ull, 0xb28717ae699944bbull, 0x0889c8c6adecb410ull
+ },
+ { // base nonce 1000000
+ 0x377dfe4959f008b5ull, 0xd1bc9ed622919f80ull, 0x7c8c576aebc618bbull, 0x7930d463eceb4becull, 0xd7c8081c372c0e96ull, 0xff093d43f46685c2ull, 0x64d59d5e4abdef5cull, 0xd0b6c0699e0d51f8ull,
+ 0xbf73d4c2e7d7bb5cull, 0xb461b9770521b1b8ull, 0x224eaf296658b9c0ull, 0x81233334fb928726ull, 0xa10c5ee0aecc8251ull, 0xcd1a7e1b3cb970e7ull, 0x573c61ddaaa8fcccull, 0x67dd7ae97a699881ull,
+ 0x6df554eae8e0482aull, 0x61f528d7eba4cf89ull, 0xc1f241f1f41fc34dull, 0x25e1b51094413da0ull, 0x9be1d90ff6f51065ull, 0x21fc77cf9089f049ull, 0x1fadecae2996b343ull, 0xbd170b4f84703b24ull,
+ 0x0a7493c2ebd8bfb5ull, 0xd1d6e43170ced136ull, 0xe87fadf2a2c02a52ull, 0x2f873ddc2cf9f022ull, 0xa3963afdf94f3307ull, 0x3dfff37956161da1ull, 0xa60a7a56137b4231ull, 0xb8abcc39f8cb021aull
+ }
+};
+
+// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455).
+static const uint32_t IGNEUM_DS_HEAD[16] = {
+ 0x19c6837fu, 0xdf3adfb8u, 0xc5f0629fu, 0x03e38ae8u, 0xc5de2525u, 0xa34a4906u, 0x898367cbu, 0x35b56808u,
+ 0xb886e673u, 0x1ee29eceu, 0x0cc36c88u, 0xc984181bu, 0x3d9d15f6u, 0x27c7d52au, 0xc268cf69u, 0x99b3c5c2u
+};
+static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u;
+static const uint32_t IGNEUM_DS_LAST = 0x8d66c390u;
+// 64 sampled dataset words (index, value) computed on the Mac.
+#define IGNEUM_DS_SAMPLES 64
+static const uint32_t IGNEUM_DS_SAMPLE_INDEX[IGNEUM_DS_SAMPLES] = {
+ 59471966u, 217795994u, 208353206u, 42483309u, 172547758u, 148076330u, 183853158u, 214389424u, 267488061u, 169781097u, 184093494u, 153880993u, 84977930u, 46426879u, 3093825u, 225364072u, 44593546u, 260713159u, 168250303u, 52384140u, 223401610u, 45554030u, 95410555u, 175039924u, 79171087u, 267580473u, 24168642u, 37981670u, 171551130u, 195559979u, 204611762u, 140997658u, 138925853u, 86637313u, 20736778u, 219665210u, 160430336u, 264654675u, 8013395u, 228945585u, 213884386u, 104419827u, 44185464u, 142737231u, 99284897u, 132475900u, 61861762u, 132056166u, 262388043u, 91878046u, 117353561u, 124768597u, 71352993u, 190698941u, 46055428u, 55281366u, 165145231u, 106810753u, 171985651u, 232085256u, 159510492u, 40072060u, 209107596u, 39023794u
+};
+static const uint32_t IGNEUM_DS_SAMPLE_VALUE[IGNEUM_DS_SAMPLES] = {
+ 0xe53754ceu, 0xc6792663u, 0x0ca9fdceu, 0x96159926u, 0xfca71449u, 0x54522663u, 0xd0160df7u, 0x301be5a9u, 0xb0de07c5u, 0xb8f1c964u, 0xcef19286u, 0xc9c7985du, 0x9db14da9u, 0xda9f746du, 0x8747e5a1u, 0x545158e0u, 0xdf31afbbu, 0x4daa9b4cu, 0x1f0762ffu, 0x80a2d158u, 0xf762ad72u, 0x7ee04a69u, 0xd07df0e9u, 0x87cfc4abu, 0x9de7bb05u, 0xcd214af2u, 0xbd15dc85u, 0x89f48a42u, 0xd49c8a27u, 0xbc60a7e1u, 0xafc36693u, 0x4876871fu, 0x703246c4u, 0xbb823e84u, 0xa13f71a1u, 0xf2dd8540u, 0x74b039d1u, 0x84e8a275u, 0x750744bbu, 0x3f690b5fu, 0x5e75ecbdu, 0x36a11cc7u, 0x65e888cfu, 0x2e8538c8u, 0x14f5fec9u, 0x7329f625u, 0x42ad9261u, 0x120aea25u, 0x5374b09fu, 0x4fa85aa3u, 0xcb856f4eu, 0x526707bau, 0x82d847dau, 0xaed963e5u, 0x1516404bu, 0x2e565dddu, 0x46d1dacfu, 0xf7bb9e4cu, 0x61ae3b87u, 0x9daeed43u, 0xcf6bc86fu, 0x24f111d5u, 0x4356a558u, 0x0890de16u
+};
+// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^26 words.
+static const uint32_t IGNEUM_CACHE_HEAD[16] = {
+ 0xebd9055cu, 0x7eaf6b21u, 0x9b610855u, 0x134a8562u, 0xb10059bau, 0x7f459e58u, 0x8f3c7873u, 0x3523e609u,
+ 0x75b8f4cau, 0x750d31b4u, 0x0dae0782u, 0x26f10015u, 0x7ba87a41u, 0x0efd8543u, 0x691d6368u, 0x8929d967u
+};
+static const uint32_t IGNEUM_CACHE_LAST[16] = {
+ 0x51474edfu, 0xc3cfba93u, 0xf21454cfu, 0x9b79baacu, 0x4d4fce23u, 0xd701dfd5u, 0x37357ab3u, 0x1be693fau,
+ 0xa701cb3bu, 0x7467c620u, 0x428184e8u, 0xf4010df0u, 0xe33aa88fu, 0xc78d6d62u, 0xae9ba9c0u, 0xf4bba97eu
+};
+static const uint64_t IGNEUM_CACHE_FNV64 = 0x448274a57f508cbcull;
diff --git a/proto-cuda/packs-ca2-mixer/mx8-devnet-epoch0/vectors.json b/proto-cuda/packs-ca2-mixer/mx8-devnet-epoch0/vectors.json
new file mode 100644
index 000000000..4e8720b3b
--- /dev/null
+++ b/proto-cuda/packs-ca2-mixer/mx8-devnet-epoch0/vectors.json
@@ -0,0 +1,36 @@
+{
+ "seed": "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000",
+ "day": "bytes:69676e65756d2d6461792ffa50000000000000",
+ "dataset_mode": "memory-hard",
+ "dataset_log2_words": 28,
+ "mask": "0x0fffffff",
+ "lanes": 32,
+ "source": "igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset",
+ "warps": [
+ {"base_nonce": 0, "expected": [
+ "0x731250f1008c6f8e", "0xb89b396aa5e33597", "0xbc96ebcdb47c61e4", "0xac1414cd0fb91152", "0x1cc2d1847875555e", "0x071f703c82efff0c", "0x21147b8c58bc96e4", "0xf412422862c643e6",
+ "0xb8a5c330668e9283", "0x994558a2b24a6557", "0xb2fb321e884a1fed", "0xa8f928a38622b10e", "0x9ea4cfcfc2c96d8b", "0xc082abc94d5dfb18", "0x695239a4a852e0f9", "0xef804c526240ce15",
+ "0x624303d31812df0d", "0xcfd82ed4955f140a", "0x6569ae6a4a436a9a", "0xc17f46cf926a7a0c", "0x721fe9c93267d13b", "0xf09fe97d5ce6bfbb", "0x455d853d3913d6a3", "0x4fbb5f43bd1af7ea",
+ "0x864d689410bb8ed0", "0xe8c93188d7f02b04", "0xebaafb6120560360", "0xde81676bff855560", "0xd0b3ae8111aab7c3", "0xacd97035767867ca", "0x1970c0b030b99c03", "0x09fed95e63215205"
+ ]},
+ {"base_nonce": 4096, "expected": [
+ "0x3d4f7e76fa55a96a", "0xd54574ccb1a27843", "0xeb1635ddd5f1e181", "0x1e49a6cf24085bd9", "0x9f929d79ad731c0e", "0x5caca654a4bd4f0f", "0x33eb9defb51424b9", "0x04f1f3c93663c27e",
+ "0xad1c775687a46553", "0xc57c600167bc7571", "0xaf0b8e2f9a316b22", "0xa01088974b376d8e", "0x14555ccee9288fdd", "0x0d1d576ea396fbe7", "0x8cf63b15a58299bf", "0x49f530db99fd5f01",
+ "0x30324c49b12da61b", "0xeb6fc871f103ad7b", "0x8ac3271ce7edfc29", "0x2804099c65265aff", "0xc4d9377a427c3337", "0x566c1e580b940448", "0x790b966beb7bebb3", "0xba5da14cc619efe0",
+ "0x5da22ad9ae88fcd7", "0x24221e506877ef84", "0xfcda89486e6e78ca", "0xfd9a09b3587ba8fb", "0x4fdb52f5b31ca54a", "0x9cf4a2730edbf331", "0xb28717ae699944bb", "0x0889c8c6adecb410"
+ ]},
+ {"base_nonce": 1000000, "expected": [
+ "0x377dfe4959f008b5", "0xd1bc9ed622919f80", "0x7c8c576aebc618bb", "0x7930d463eceb4bec", "0xd7c8081c372c0e96", "0xff093d43f46685c2", "0x64d59d5e4abdef5c", "0xd0b6c0699e0d51f8",
+ "0xbf73d4c2e7d7bb5c", "0xb461b9770521b1b8", "0x224eaf296658b9c0", "0x81233334fb928726", "0xa10c5ee0aecc8251", "0xcd1a7e1b3cb970e7", "0x573c61ddaaa8fccc", "0x67dd7ae97a699881",
+ "0x6df554eae8e0482a", "0x61f528d7eba4cf89", "0xc1f241f1f41fc34d", "0x25e1b51094413da0", "0x9be1d90ff6f51065", "0x21fc77cf9089f049", "0x1fadecae2996b343", "0xbd170b4f84703b24",
+ "0x0a7493c2ebd8bfb5", "0xd1d6e43170ced136", "0xe87fadf2a2c02a52", "0x2f873ddc2cf9f022", "0xa3963afdf94f3307", "0x3dfff37956161da1", "0xa60a7a56137b4231", "0xb8abcc39f8cb021a"
+ ]}
+ ],
+ "dataset_head": ["0x19c6837f", "0xdf3adfb8", "0xc5f0629f", "0x03e38ae8", "0xc5de2525", "0xa34a4906", "0x898367cb", "0x35b56808", "0xb886e673", "0x1ee29ece", "0x0cc36c88", "0xc984181b", "0x3d9d15f6", "0x27c7d52a", "0xc268cf69", "0x99b3c5c2"],
+ "dataset_last_index": 268435455,
+ "dataset_last": "0x8d66c390",
+ "dataset_samples": [{"index": 59471966, "value": "0xe53754ce"}, {"index": 217795994, "value": "0xc6792663"}, {"index": 208353206, "value": "0x0ca9fdce"}, {"index": 42483309, "value": "0x96159926"}, {"index": 172547758, "value": "0xfca71449"}, {"index": 148076330, "value": "0x54522663"}, {"index": 183853158, "value": "0xd0160df7"}, {"index": 214389424, "value": "0x301be5a9"}, {"index": 267488061, "value": "0xb0de07c5"}, {"index": 169781097, "value": "0xb8f1c964"}, {"index": 184093494, "value": "0xcef19286"}, {"index": 153880993, "value": "0xc9c7985d"}, {"index": 84977930, "value": "0x9db14da9"}, {"index": 46426879, "value": "0xda9f746d"}, {"index": 3093825, "value": "0x8747e5a1"}, {"index": 225364072, "value": "0x545158e0"}, {"index": 44593546, "value": "0xdf31afbb"}, {"index": 260713159, "value": "0x4daa9b4c"}, {"index": 168250303, "value": "0x1f0762ff"}, {"index": 52384140, "value": "0x80a2d158"}, {"index": 223401610, "value": "0xf762ad72"}, {"index": 45554030, "value": "0x7ee04a69"}, {"index": 95410555, "value": "0xd07df0e9"}, {"index": 175039924, "value": "0x87cfc4ab"}, {"index": 79171087, "value": "0x9de7bb05"}, {"index": 267580473, "value": "0xcd214af2"}, {"index": 24168642, "value": "0xbd15dc85"}, {"index": 37981670, "value": "0x89f48a42"}, {"index": 171551130, "value": "0xd49c8a27"}, {"index": 195559979, "value": "0xbc60a7e1"}, {"index": 204611762, "value": "0xafc36693"}, {"index": 140997658, "value": "0x4876871f"}, {"index": 138925853, "value": "0x703246c4"}, {"index": 86637313, "value": "0xbb823e84"}, {"index": 20736778, "value": "0xa13f71a1"}, {"index": 219665210, "value": "0xf2dd8540"}, {"index": 160430336, "value": "0x74b039d1"}, {"index": 264654675, "value": "0x84e8a275"}, {"index": 8013395, "value": "0x750744bb"}, {"index": 228945585, "value": "0x3f690b5f"}, {"index": 213884386, "value": "0x5e75ecbd"}, {"index": 104419827, "value": "0x36a11cc7"}, {"index": 44185464, "value": "0x65e888cf"}, {"index": 142737231, "value": "0x2e8538c8"}, {"index": 99284897, "value": "0x14f5fec9"}, {"index": 132475900, "value": "0x7329f625"}, {"index": 61861762, "value": "0x42ad9261"}, {"index": 132056166, "value": "0x120aea25"}, {"index": 262388043, "value": "0x5374b09f"}, {"index": 91878046, "value": "0x4fa85aa3"}, {"index": 117353561, "value": "0xcb856f4e"}, {"index": 124768597, "value": "0x526707ba"}, {"index": 71352993, "value": "0x82d847da"}, {"index": 190698941, "value": "0xaed963e5"}, {"index": 46055428, "value": "0x1516404b"}, {"index": 55281366, "value": "0x2e565ddd"}, {"index": 165145231, "value": "0x46d1dacf"}, {"index": 106810753, "value": "0xf7bb9e4c"}, {"index": 171985651, "value": "0x61ae3b87"}, {"index": 232085256, "value": "0x9daeed43"}, {"index": 159510492, "value": "0xcf6bc86f"}, {"index": 40072060, "value": "0x24f111d5"}, {"index": 209107596, "value": "0x4356a558"}, {"index": 39023794, "value": "0x0890de16"}],
+ "cache_head": ["0xebd9055c", "0x7eaf6b21", "0x9b610855", "0x134a8562", "0xb10059ba", "0x7f459e58", "0x8f3c7873", "0x3523e609", "0x75b8f4ca", "0x750d31b4", "0x0dae0782", "0x26f10015", "0x7ba87a41", "0x0efd8543", "0x691d6368", "0x8929d967"],
+ "cache_last_line": ["0x51474edf", "0xc3cfba93", "0xf21454cf", "0x9b79baac", "0x4d4fce23", "0xd701dfd5", "0x37357ab3", "0x1be693fa", "0xa701cb3b", "0x7467c620", "0x428184e8", "0xf4010df0", "0xe33aa88f", "0xc78d6d62", "0xae9ba9c0", "0xf4bba97e"],
+ "cache_fnv1a64": "0x448274a57f508cbc"
+}
diff --git a/proto-cuda/packs-ca2-mixer/mx8-genesis/kernel.cl b/proto-cuda/packs-ca2-mixer/mx8-genesis/kernel.cl
new file mode 100644
index 000000000..a8ccd9f8f
--- /dev/null
+++ b/proto-cuda/packs-ca2-mixer/mx8-genesis/kernel.cl
@@ -0,0 +1,279 @@
+// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
+// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal).
+// Built from source at runtime by proto-opencl/host.c, which passes these defines:
+// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit)
+// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default)
+// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32
+// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition
+// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units;
+// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32.
+#ifndef IGNEUM_GROUP
+#define IGNEUM_GROUP 32
+#endif
+#ifndef IGNEUM_EXCHANGE
+#define IGNEUM_EXCHANGE 0
+#endif
+#ifdef __OPENCL_VERSION__
+#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1)))
+#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n]
+#if IGNEUM_EXCHANGE == 1
+#ifdef cl_khr_subgroups
+#pragma OPENCL EXTENSION cl_khr_subgroups : enable
+#endif
+#ifdef cl_khr_subgroup_shuffle
+#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable
+#endif
+#elif IGNEUM_EXCHANGE == 2
+#pragma OPENCL EXTENSION cl_intel_subgroups : enable
+#endif
+#else
+// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros.
+#include "emu_opencl.h"
+#endif
+
+#if IGNEUM_EXCHANGE == 1
+#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m))
+#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
+#elif IGNEUM_EXCHANGE == 2
+#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m))
+#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
+#else
+// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per
+// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane
+// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's
+// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier.
+#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; }
+#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; }
+#endif
+
+static inline uint splitmix32(uint x) {
+ x ^= x >> 16; x *= 0x7feb352du;
+ x ^= x >> 15; x *= 0x846ca68bu;
+ x ^= x >> 16;
+ return x;
+}
+// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32.
+static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); }
+// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x.
+static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); }
+static inline uint ds_elem(uint i, uint d0, uint d1) {
+ uint x = i ^ d0;
+ x *= 0x9E3779B1u; x ^= x >> 15;
+ x += d1;
+ x *= 0x85EBCA77u; x ^= x >> 13;
+ x *= 0xC2B2AE3Du; x ^= x >> 16;
+ return x;
+}
+
+// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
+// Item: 8 rounds of 8 x seed-parameterised mixer + one 64-byte cache read, then 8 x final mixer (class v3, mixer multiplier 8,
+// docs/plans/mixer-x4.md: the round key of application j of round r is 0x9E3779B9 * (r * m + j + 1)). All parameters are literals.
+#define MH_CACHE_LINE_MASK 0x003fffffu
+#define MH_SEGMENT_LINES 64u
+#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
+static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
+
+// y = ChaCha12 core(x) + x
+static inline void mh_chacha_block(const uint* x, uint* y) {
+ for (uint i = 0u; i < 16u; ++i) y[i] = x[i];
+ for (uint r = 0u; r < 6u; ++r) {
+ MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
+ MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
+ MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
+ MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
+ }
+ for (uint i = 0u; i < 16u; ++i) y[i] += x[i];
+}
+
+// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
+static inline void mh_cache_segment(__global uint* cache, uint seg) {
+ uint prev[16]; uint x[16]; uint y[16];
+ for (uint i = 0u; i < 16u; ++i) prev[i] = 0u;
+ for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) {
+ x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
+ x[4] = 0x3067619fu ^ prev[4];
+ x[5] = 0x3c269176u ^ prev[5];
+ x[6] = 0x84a03b03u ^ prev[6];
+ x[7] = 0xf8c63294u ^ prev[7];
+ x[8] = 0xff977c5bu ^ prev[8];
+ x[9] = 0xe60def3eu ^ prev[9];
+ x[10] = 0x63630141u ^ prev[10];
+ x[11] = 0xb8fbcb58u ^ prev[11];
+ x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
+ mh_chacha_block(x, y);
+ __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
+ for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
+ }
+}
+
+// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
+static inline void mh_mixer(uint* s, uint rk) {
+ s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u;
+ s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu;
+ s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u;
+ s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu;
+ s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u;
+ s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u;
+ s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u;
+ s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u;
+ s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du;
+ s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u;
+ s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du;
+ s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu;
+ s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du;
+ s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu;
+ s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u;
+ s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u;
+ MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u)
+ MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u)
+ MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u)
+ MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u)
+}
+
+// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of 8 x mixer + cache line s[0] & mask; 8 x final mixer.
+static inline void mh_item(__global const uint* cache, uint t, uint* s) {
+ s[0] = 0x3067619fu;
+ s[1] = 0x3c269176u;
+ s[2] = 0x84a03b03u;
+ s[3] = 0xf8c63294u;
+ s[4] = 0xff977c5bu;
+ s[5] = 0xe60def3eu;
+ s[6] = 0x63630141u;
+ s[7] = 0xb8fbcb58u;
+ s[8] = t * 0x42146205u + 0xbab68293u;
+ s[9] = t * 0x52cbe0fbu + 0xcc162340u;
+ s[10] = t * 0x7ecf4a03u + 0x6ce151ccu;
+ s[11] = t * 0x6728907fu + 0xe62b8997u;
+ s[12] = t * 0xd81d9751u + 0xc9c80297u;
+ s[13] = t * 0x132952c3u + 0xf74a1654u;
+ s[14] = t * 0xf60de277u + 0x3d704af5u;
+ s[15] = t * 0x05358035u + 0x3cf522b7u;
+ for (uint r = 0u; r < 8u; ++r) {
+ for (uint j = 0u; j < 8u; ++j) mh_mixer(s, 0x9E3779B9u * (r * 8u + j + 1u));
+ __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
+ for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i];
+ }
+ for (uint j = 0u; j < 8u; ++j) mh_mixer(s, 0x9E3779B9u * (64u + j + 1u));
+}
+// dataset[w] without the dataset: derive item w >> 4 and take word w & 15.
+static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; }
+
+// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item.
+// The same constants as memhard.h in this pack (one emitter, three dialects).
+__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) {
+ uint seg = (uint)get_global_id(0);
+ if (seg < nSegments) mh_cache_segment(cache, seg);
+}
+__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) {
+ uint t = (uint)get_global_id(0);
+ if (t < nItems) {
+ uint s[16];
+ mh_item(cache, t, s);
+ __global uint* d = ds + ((ulong)t * 16u);
+ for (uint i = 0u; i < 16u; ++i) d[i] = s[i];
+ }
+}
+
+// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the
+// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and
+// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all).
+IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) {
+ uint gid = (uint)get_global_id(0);
+ uint lid = (uint)get_local_id(0);
+ uint nonce = baseNonce + gid;
+ uint r0, r1, r2, r3, r4, r5, r6, r7;
+#if IGNEUM_EXCHANGE == 0
+ IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
+ uint xk = 0u;
+#else
+ (void)lid;
+#endif
+ { uint x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
+ { uint x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
+ { uint x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
+ { uint x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
+ { uint x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
+ { uint x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
+ { uint x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
+ { uint x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
+
+ for (uint it = 0u; it < 8u; ++it) {
+ uint sel = r0;
+ r2 = r3 * r4 + r2; // 0 mad
+ r2 = r1 * r1 + r2; // 1 mad
+ r2 = r3 * r2 + r2; // 2 mad
+ r3 = r3 ^ r5; // 3 xor
+ r7 = r7 ^ ds[r2 & mask]; // 4 load
+ r5 = r5 ^ ds[r7 & mask]; // 5 load
+ { uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r1 = r1 ^ t_; } // 6 shfl
+ { uint t_; IGNEUM_SHFL_XOR(t_, r3, 8u); r7 = r7 ^ t_; } // 7 shfl
+ r1 = mul_hi(r1, r5); // 8 mulhi
+ r6 = rotr_var(r6, r3); // 9 rotr
+ r3 = r3 | r4; // 10 or
+ r4 = r4 ^ ds[r3 & mask]; // 11 load
+ r0 = mul_hi(r0, r4); // 12 mulhi
+ r5 = r5 + r1 + ((((sel >> 30u) & 1u) != 0u) ? 0xd3177981u : 0xc7934706u); // 13 add
+ r0 = r0 ^ ds[r4 & mask]; // 14 load
+ r2 = r2 - r4; // 15 sub
+ r2 = r2 ^ ds[r0 & mask]; // 16 load
+ r7 = r7 ^ ds[r2 & mask]; // 17 load
+ { uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 18 shfl
+ r5 = r5 * r0; // 19 mul
+ { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 20 shfl
+ { uint t_; IGNEUM_SHFL_XOR(t_, r4, 16u); r2 = r2 ^ t_; } // 21 shfl
+ r6 = mul_hi(r6, r2); // 22 mulhi
+ r6 = r6 ^ ds[r1 & mask]; // 23 load
+ r5 = r5 * r0; // 24 mul
+ r5 = rotl_imm(r5, 19u); // 25 rotl
+ { uint t_; IGNEUM_SHFL_XOR(t_, r6, 2u); r7 = r7 ^ t_; } // 26 shfl
+ r0 = r0 ^ r5; // 27 xor
+ r0 = r0 ^ r4; // 28 xor
+ r3 = r3 - r0; // 29 sub
+ r5 = r5 * r1; // 30 mul
+ r7 = r7 ^ ds[r2 & mask]; // 31 load
+ r1 = r1 ^ ds[r0 & mask]; // 32 load
+ r5 = r5 ^ r6; // 33 xor
+ r5 = r5 ^ ds[r1 & mask]; // 34 load
+ r0 = mul_hi(r0, r5); // 35 mulhi
+ { uint t_; IGNEUM_SHFL_XOR(t_, r2, 4u); r5 = r5 ^ t_; } // 36 shfl
+ r7 = r7 ^ ds[r0 & mask]; // 37 load
+ r3 = r3 + r1 + ((((sel >> 27u) & 1u) != 0u) ? 0x230c005cu : 0x75ba2fadu); // 38 add
+ { uint t_; IGNEUM_SHFL_XOR(t_, r5, 4u); r1 = r1 ^ t_; } // 39 shfl
+ r2 = r2 ^ r5; // 40 xor
+ r3 = r6 * r3 + r3; // 41 mad
+ r6 = r6 - r7; // 42 sub
+ r7 = r7 ^ r0; // 43 xor
+ r1 = r1 ^ ds[r7 & mask]; // 44 load
+ r2 = r2 * r3; // 45 mul
+ r1 = mul_hi(r1, r5); // 46 mulhi
+ r4 = r4 - r3; // 47 sub
+ r2 = rotr_var(r2, r6); // 48 rotr
+ r3 = r3 ^ ds[r5 & mask]; // 49 load
+ r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 50 add
+ r0 = r0 * r2; // 51 mul
+ r0 = r0 + r2 + ((((sel >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); // 52 add
+ r1 = r1 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0x77b1520du : 0x2bb965afu); // 53 add
+ r7 = rotl_imm(r7, 14u); // 54 rotl
+ r3 = r3 + r7 + ((((sel >> 1u) & 1u) != 0u) ? 0xa54c55a0u : 0x7b0fe07au); // 55 add
+ r6 = r6 ^ ds[r7 & mask]; // 56 load
+ r1 = rotr_var(r1, r5); // 57 rotr
+ r5 = r5 ^ ds[r4 & mask]; // 58 load
+ r6 = r6 ^ ds[r2 & mask]; // 59 load
+ r3 = r5 * r0 + r3; // 60 mad
+ r5 = r5 + r7 + ((((sel >> 31u) & 1u) != 0u) ? 0xad7493e7u : 0xaf9dd72du); // 61 add
+ r4 = r4 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); // 62 add
+ r5 = rotl_imm(r5, 19u); // 63 rotl
+ }
+ uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
+ uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
+ out[gid] = ((ulong)hi << 32) | (ulong)lo;
+}
+
+#if IGNEUM_EXCHANGE != 0
+// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the
+// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a
+// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md.
+IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) {
+ if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); }
+}
+#endif
diff --git a/proto-cuda/packs-ca2-mixer/mx8-genesis/kernel.cu b/proto-cuda/packs-ca2-mixer/mx8-genesis/kernel.cu
new file mode 100644
index 000000000..e22472314
--- /dev/null
+++ b/proto-cuda/packs-ca2-mixer/mx8-genesis/kernel.cu
@@ -0,0 +1,164 @@
+// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
+// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal).
+// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC.
+#include
+#include
+#include "program.h"
+#include "memhard.h"
+
+__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
+ x ^= x >> 16; x *= 0x7feb352du;
+ x ^= x >> 15; x *= 0x846ca68bu;
+ x ^= x >> 16;
+ return x;
+}
+// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31.
+__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
+// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x.
+__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
+__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) {
+ uint32_t x = i ^ d0;
+ x *= 0x9E3779B1u; x ^= x >> 15;
+ x += d1;
+ x *= 0x85EBCA77u; x ^= x >> 13;
+ x *= 0xC2B2AE3Du; x ^= x >> 16;
+ return x;
+}
+
+// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item.
+// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host.
+__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) {
+ uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x;
+ if (seg < nSegments) mh_cache_segment(cache, seg);
+}
+__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
+ uint32_t t = blockIdx.x * blockDim.x + threadIdx.x;
+ if (t < nItems) {
+ uint32_t s[16];
+ mh_item(cache, t, s);
+ uint32_t* d = ds + (size_t)t * 16u;
+ for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i];
+ }
+}
+
+// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every
+// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a
+// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid.
+__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) {
+ uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
+ uint32_t nonce = baseNonce + gid;
+ uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
+ { uint32_t x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
+ { uint32_t x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
+ { uint32_t x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
+ { uint32_t x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
+ { uint32_t x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
+ { uint32_t x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
+ { uint32_t x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
+ { uint32_t x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
+
+ for (uint32_t it = 0u; it < 8u; ++it) {
+ uint32_t sel = r0;
+ r2 = r3 * r4 + r2; // 0 mad
+ r2 = r1 * r1 + r2; // 1 mad
+ r2 = r3 * r2 + r2; // 2 mad
+ r3 = r3 ^ r5; // 3 xor
+ r7 = r7 ^ ds[r2 & mask]; // 4 load
+ r5 = r5 ^ ds[r7 & mask]; // 5 load
+ r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 6 shfl
+ r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 8); // 7 shfl
+ r1 = __umulhi(r1, r5); // 8 mulhi
+ r6 = rotr_var(r6, r3); // 9 rotr
+ r3 = r3 | r4; // 10 or
+ r4 = r4 ^ ds[r3 & mask]; // 11 load
+ r0 = __umulhi(r0, r4); // 12 mulhi
+ r5 = r5 + r1 + ((((sel >> 30u) & 1u) != 0u) ? 0xd3177981u : 0xc7934706u); // 13 add
+ r0 = r0 ^ ds[r4 & mask]; // 14 load
+ r2 = r2 - r4; // 15 sub
+ r2 = r2 ^ ds[r0 & mask]; // 16 load
+ r7 = r7 ^ ds[r2 & mask]; // 17 load
+ r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 18 shfl
+ r5 = r5 * r0; // 19 mul
+ r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 20 shfl
+ r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r4, 16); // 21 shfl
+ r6 = __umulhi(r6, r2); // 22 mulhi
+ r6 = r6 ^ ds[r1 & mask]; // 23 load
+ r5 = r5 * r0; // 24 mul
+ r5 = rotl_imm(r5, 19u); // 25 rotl
+ r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r6, 2); // 26 shfl
+ r0 = r0 ^ r5; // 27 xor
+ r0 = r0 ^ r4; // 28 xor
+ r3 = r3 - r0; // 29 sub
+ r5 = r5 * r1; // 30 mul
+ r7 = r7 ^ ds[r2 & mask]; // 31 load
+ r1 = r1 ^ ds[r0 & mask]; // 32 load
+ r5 = r5 ^ r6; // 33 xor
+ r5 = r5 ^ ds[r1 & mask]; // 34 load
+ r0 = __umulhi(r0, r5); // 35 mulhi
+ r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r2, 4); // 36 shfl
+ r7 = r7 ^ ds[r0 & mask]; // 37 load
+ r3 = r3 + r1 + ((((sel >> 27u) & 1u) != 0u) ? 0x230c005cu : 0x75ba2fadu); // 38 add
+ r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r5, 4); // 39 shfl
+ r2 = r2 ^ r5; // 40 xor
+ r3 = r6 * r3 + r3; // 41 mad
+ r6 = r6 - r7; // 42 sub
+ r7 = r7 ^ r0; // 43 xor
+ r1 = r1 ^ ds[r7 & mask]; // 44 load
+ r2 = r2 * r3; // 45 mul
+ r1 = __umulhi(r1, r5); // 46 mulhi
+ r4 = r4 - r3; // 47 sub
+ r2 = rotr_var(r2, r6); // 48 rotr
+ r3 = r3 ^ ds[r5 & mask]; // 49 load
+ r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 50 add
+ r0 = r0 * r2; // 51 mul
+ r0 = r0 + r2 + ((((sel >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); // 52 add
+ r1 = r1 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0x77b1520du : 0x2bb965afu); // 53 add
+ r7 = rotl_imm(r7, 14u); // 54 rotl
+ r3 = r3 + r7 + ((((sel >> 1u) & 1u) != 0u) ? 0xa54c55a0u : 0x7b0fe07au); // 55 add
+ r6 = r6 ^ ds[r7 & mask]; // 56 load
+ r1 = rotr_var(r1, r5); // 57 rotr
+ r5 = r5 ^ ds[r4 & mask]; // 58 load
+ r6 = r6 ^ ds[r2 & mask]; // 59 load
+ r3 = r5 * r0 + r3; // 60 mad
+ r5 = r5 + r7 + ((((sel >> 31u) & 1u) != 0u) ? 0xad7493e7u : 0xaf9dd72du); // 61 add
+ r4 = r4 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); // 62 add
+ r5 = rotl_imm(r5, 19u); // 63 rotl
+ }
+ uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
+ uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
+ out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
+}
+
+// Host-side launch wrappers. Declared in program.h, called from host.cu.
+cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) {
+ if (nSegments == 0u) return cudaErrorInvalidValue;
+ uint32_t block = 256u;
+ uint32_t grid = (nSegments + block - 1u) / block;
+ igneum_cache_fill<<>>(cache, nSegments);
+ return cudaGetLastError();
+}
+
+cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
+ if (nItems == 0u) return cudaErrorInvalidValue;
+ uint32_t block = 256u;
+ uint32_t grid = (nItems + block - 1u) / block;
+ igneum_build<<>>(ds, cache, nItems);
+ return cudaGetLastError();
+}
+
+cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
+ uint32_t nonces, uint32_t blockWarps) {
+ if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
+ uint32_t block = 32u * blockWarps;
+ if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
+ igneum_hash<<>>(ds, out, baseNonce, mask);
+ return cudaGetLastError();
+}
+
+cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
+ cudaFuncAttributes attr;
+ cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash);
+ if (e != cudaSuccess) return e;
+ *numRegs = attr.numRegs;
+ return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0);
+}
diff --git a/proto-cuda/packs-ca2-mixer/mx8-genesis/kernel_bound.cl b/proto-cuda/packs-ca2-mixer/mx8-genesis/kernel_bound.cl
new file mode 100644
index 000000000..a8164a82a
--- /dev/null
+++ b/proto-cuda/packs-ca2-mixer/mx8-genesis/kernel_bound.cl
@@ -0,0 +1,373 @@
+// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
+// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal).
+// Built from source at runtime by proto-opencl/host.c, which passes these defines:
+// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit)
+// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default)
+// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32
+// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition
+// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units;
+// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32.
+#ifndef IGNEUM_GROUP
+#define IGNEUM_GROUP 32
+#endif
+#ifndef IGNEUM_EXCHANGE
+#define IGNEUM_EXCHANGE 0
+#endif
+#ifdef __OPENCL_VERSION__
+#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1)))
+#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n]
+#if IGNEUM_EXCHANGE == 1
+#ifdef cl_khr_subgroups
+#pragma OPENCL EXTENSION cl_khr_subgroups : enable
+#endif
+#ifdef cl_khr_subgroup_shuffle
+#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable
+#endif
+#elif IGNEUM_EXCHANGE == 2
+#pragma OPENCL EXTENSION cl_intel_subgroups : enable
+#endif
+#else
+// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros.
+#include "emu_opencl.h"
+#endif
+
+#if IGNEUM_EXCHANGE == 1
+#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m))
+#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
+#elif IGNEUM_EXCHANGE == 2
+#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m))
+#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
+#else
+// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per
+// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane
+// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's
+// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier.
+#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; }
+#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; }
+#endif
+
+static inline uint splitmix32(uint x) {
+ x ^= x >> 16; x *= 0x7feb352du;
+ x ^= x >> 15; x *= 0x846ca68bu;
+ x ^= x >> 16;
+ return x;
+}
+// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32.
+static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); }
+// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x.
+static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); }
+static inline uint ds_elem(uint i, uint d0, uint d1) {
+ uint x = i ^ d0;
+ x *= 0x9E3779B1u; x ^= x >> 15;
+ x += d1;
+ x *= 0x85EBCA77u; x ^= x >> 13;
+ x *= 0xC2B2AE3Du; x ^= x >> 16;
+ return x;
+}
+
+// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
+// Item: 8 rounds of 8 x seed-parameterised mixer + one 64-byte cache read, then 8 x final mixer (class v3, mixer multiplier 8,
+// docs/plans/mixer-x4.md: the round key of application j of round r is 0x9E3779B9 * (r * m + j + 1)). All parameters are literals.
+#define MH_CACHE_LINE_MASK 0x003fffffu
+#define MH_SEGMENT_LINES 64u
+#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
+static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
+
+// y = ChaCha12 core(x) + x
+static inline void mh_chacha_block(const uint* x, uint* y) {
+ for (uint i = 0u; i < 16u; ++i) y[i] = x[i];
+ for (uint r = 0u; r < 6u; ++r) {
+ MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
+ MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
+ MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
+ MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
+ }
+ for (uint i = 0u; i < 16u; ++i) y[i] += x[i];
+}
+
+// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
+static inline void mh_cache_segment(__global uint* cache, uint seg) {
+ uint prev[16]; uint x[16]; uint y[16];
+ for (uint i = 0u; i < 16u; ++i) prev[i] = 0u;
+ for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) {
+ x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
+ x[4] = 0x3067619fu ^ prev[4];
+ x[5] = 0x3c269176u ^ prev[5];
+ x[6] = 0x84a03b03u ^ prev[6];
+ x[7] = 0xf8c63294u ^ prev[7];
+ x[8] = 0xff977c5bu ^ prev[8];
+ x[9] = 0xe60def3eu ^ prev[9];
+ x[10] = 0x63630141u ^ prev[10];
+ x[11] = 0xb8fbcb58u ^ prev[11];
+ x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
+ mh_chacha_block(x, y);
+ __global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
+ for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
+ }
+}
+
+// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
+static inline void mh_mixer(uint* s, uint rk) {
+ s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u;
+ s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu;
+ s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u;
+ s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu;
+ s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u;
+ s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u;
+ s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u;
+ s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u;
+ s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du;
+ s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u;
+ s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du;
+ s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu;
+ s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du;
+ s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu;
+ s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u;
+ s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u;
+ MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u)
+ MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u)
+ MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u)
+ MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u)
+}
+
+// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of 8 x mixer + cache line s[0] & mask; 8 x final mixer.
+static inline void mh_item(__global const uint* cache, uint t, uint* s) {
+ s[0] = 0x3067619fu;
+ s[1] = 0x3c269176u;
+ s[2] = 0x84a03b03u;
+ s[3] = 0xf8c63294u;
+ s[4] = 0xff977c5bu;
+ s[5] = 0xe60def3eu;
+ s[6] = 0x63630141u;
+ s[7] = 0xb8fbcb58u;
+ s[8] = t * 0x42146205u + 0xbab68293u;
+ s[9] = t * 0x52cbe0fbu + 0xcc162340u;
+ s[10] = t * 0x7ecf4a03u + 0x6ce151ccu;
+ s[11] = t * 0x6728907fu + 0xe62b8997u;
+ s[12] = t * 0xd81d9751u + 0xc9c80297u;
+ s[13] = t * 0x132952c3u + 0xf74a1654u;
+ s[14] = t * 0xf60de277u + 0x3d704af5u;
+ s[15] = t * 0x05358035u + 0x3cf522b7u;
+ for (uint r = 0u; r < 8u; ++r) {
+ for (uint j = 0u; j < 8u; ++j) mh_mixer(s, 0x9E3779B9u * (r * 8u + j + 1u));
+ __global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
+ for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i];
+ }
+ for (uint j = 0u; j < 8u; ++j) mh_mixer(s, 0x9E3779B9u * (64u + j + 1u));
+}
+// dataset[w] without the dataset: derive item w >> 4 and take word w & 15.
+static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; }
+
+// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item.
+// The same constants as memhard.h in this pack (one emitter, three dialects).
+__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) {
+ uint seg = (uint)get_global_id(0);
+ if (seg < nSegments) mh_cache_segment(cache, seg);
+}
+__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) {
+ uint t = (uint)get_global_id(0);
+ if (t < nItems) {
+ uint s[16];
+ mh_item(cache, t, s);
+ __global uint* d = ds + ((ulong)t * 16u);
+ for (uint i = 0u; i < 16u; ++i) d[i] = s[i];
+ }
+}
+
+// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the
+// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and
+// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all).
+IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) {
+ uint gid = (uint)get_global_id(0);
+ uint lid = (uint)get_local_id(0);
+ uint nonce = baseNonce + gid;
+ uint r0, r1, r2, r3, r4, r5, r6, r7;
+#if IGNEUM_EXCHANGE == 0
+ IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
+ uint xk = 0u;
+#else
+ (void)lid;
+#endif
+ { uint x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
+ { uint x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
+ { uint x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
+ { uint x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
+ { uint x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
+ { uint x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
+ { uint x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
+ { uint x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
+
+ for (uint it = 0u; it < 8u; ++it) {
+ uint sel = r0;
+ r2 = r3 * r4 + r2; // 0 mad
+ r2 = r1 * r1 + r2; // 1 mad
+ r2 = r3 * r2 + r2; // 2 mad
+ r3 = r3 ^ r5; // 3 xor
+ r7 = r7 ^ ds[r2 & mask]; // 4 load
+ r5 = r5 ^ ds[r7 & mask]; // 5 load
+ { uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r1 = r1 ^ t_; } // 6 shfl
+ { uint t_; IGNEUM_SHFL_XOR(t_, r3, 8u); r7 = r7 ^ t_; } // 7 shfl
+ r1 = mul_hi(r1, r5); // 8 mulhi
+ r6 = rotr_var(r6, r3); // 9 rotr
+ r3 = r3 | r4; // 10 or
+ r4 = r4 ^ ds[r3 & mask]; // 11 load
+ r0 = mul_hi(r0, r4); // 12 mulhi
+ r5 = r5 + r1 + ((((sel >> 30u) & 1u) != 0u) ? 0xd3177981u : 0xc7934706u); // 13 add
+ r0 = r0 ^ ds[r4 & mask]; // 14 load
+ r2 = r2 - r4; // 15 sub
+ r2 = r2 ^ ds[r0 & mask]; // 16 load
+ r7 = r7 ^ ds[r2 & mask]; // 17 load
+ { uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 18 shfl
+ r5 = r5 * r0; // 19 mul
+ { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 20 shfl
+ { uint t_; IGNEUM_SHFL_XOR(t_, r4, 16u); r2 = r2 ^ t_; } // 21 shfl
+ r6 = mul_hi(r6, r2); // 22 mulhi
+ r6 = r6 ^ ds[r1 & mask]; // 23 load
+ r5 = r5 * r0; // 24 mul
+ r5 = rotl_imm(r5, 19u); // 25 rotl
+ { uint t_; IGNEUM_SHFL_XOR(t_, r6, 2u); r7 = r7 ^ t_; } // 26 shfl
+ r0 = r0 ^ r5; // 27 xor
+ r0 = r0 ^ r4; // 28 xor
+ r3 = r3 - r0; // 29 sub
+ r5 = r5 * r1; // 30 mul
+ r7 = r7 ^ ds[r2 & mask]; // 31 load
+ r1 = r1 ^ ds[r0 & mask]; // 32 load
+ r5 = r5 ^ r6; // 33 xor
+ r5 = r5 ^ ds[r1 & mask]; // 34 load
+ r0 = mul_hi(r0, r5); // 35 mulhi
+ { uint t_; IGNEUM_SHFL_XOR(t_, r2, 4u); r5 = r5 ^ t_; } // 36 shfl
+ r7 = r7 ^ ds[r0 & mask]; // 37 load
+ r3 = r3 + r1 + ((((sel >> 27u) & 1u) != 0u) ? 0x230c005cu : 0x75ba2fadu); // 38 add
+ { uint t_; IGNEUM_SHFL_XOR(t_, r5, 4u); r1 = r1 ^ t_; } // 39 shfl
+ r2 = r2 ^ r5; // 40 xor
+ r3 = r6 * r3 + r3; // 41 mad
+ r6 = r6 - r7; // 42 sub
+ r7 = r7 ^ r0; // 43 xor
+ r1 = r1 ^ ds[r7 & mask]; // 44 load
+ r2 = r2 * r3; // 45 mul
+ r1 = mul_hi(r1, r5); // 46 mulhi
+ r4 = r4 - r3; // 47 sub
+ r2 = rotr_var(r2, r6); // 48 rotr
+ r3 = r3 ^ ds[r5 & mask]; // 49 load
+ r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 50 add
+ r0 = r0 * r2; // 51 mul
+ r0 = r0 + r2 + ((((sel >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); // 52 add
+ r1 = r1 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0x77b1520du : 0x2bb965afu); // 53 add
+ r7 = rotl_imm(r7, 14u); // 54 rotl
+ r3 = r3 + r7 + ((((sel >> 1u) & 1u) != 0u) ? 0xa54c55a0u : 0x7b0fe07au); // 55 add
+ r6 = r6 ^ ds[r7 & mask]; // 56 load
+ r1 = rotr_var(r1, r5); // 57 rotr
+ r5 = r5 ^ ds[r4 & mask]; // 58 load
+ r6 = r6 ^ ds[r2 & mask]; // 59 load
+ r3 = r5 * r0 + r3; // 60 mad
+ r5 = r5 + r7 + ((((sel >> 31u) & 1u) != 0u) ? 0xad7493e7u : 0xaf9dd72du); // 61 add
+ r4 = r4 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); // 62 add
+ r5 = rotl_imm(r5, 19u); // 63 rotl
+ }
+ uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
+ uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
+ out[gid] = ((ulong)hi << 32) | (ulong)lo;
+}
+
+#if IGNEUM_EXCHANGE != 0
+// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the
+// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a
+// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md.
+IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) {
+ if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); }
+}
+#endif
+
+// Header-bound variant (bind.rs): the init words come from initw, not SEEDW. Same body as igneum_hash.
+IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw) {
+ uint gid = (uint)get_global_id(0);
+ uint lid = (uint)get_local_id(0);
+ uint nonce = baseNonce + gid;
+ uint r0, r1, r2, r3, r4, r5, r6, r7;
+ uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7];
+#if IGNEUM_EXCHANGE == 0
+ IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
+ uint xk = 0u;
+#else
+ (void)lid;
+#endif
+ { uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; }
+ { uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; }
+ { uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; }
+ { uint x = nonce ^ iw3; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw4; }
+ { uint x = nonce ^ iw4; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw5; }
+ { uint x = nonce ^ iw5; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw6; }
+ { uint x = nonce ^ iw6; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw7; }
+ { uint x = nonce ^ iw7; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw0; }
+
+ for (uint it = 0u; it < 8u; ++it) {
+ uint sel = r0;
+ r2 = r3 * r4 + r2; // 0 mad
+ r2 = r1 * r1 + r2; // 1 mad
+ r2 = r3 * r2 + r2; // 2 mad
+ r3 = r3 ^ r5; // 3 xor
+ r7 = r7 ^ ds[r2 & mask]; // 4 load
+ r5 = r5 ^ ds[r7 & mask]; // 5 load
+ { uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r1 = r1 ^ t_; } // 6 shfl
+ { uint t_; IGNEUM_SHFL_XOR(t_, r3, 8u); r7 = r7 ^ t_; } // 7 shfl
+ r1 = mul_hi(r1, r5); // 8 mulhi
+ r6 = rotr_var(r6, r3); // 9 rotr
+ r3 = r3 | r4; // 10 or
+ r4 = r4 ^ ds[r3 & mask]; // 11 load
+ r0 = mul_hi(r0, r4); // 12 mulhi
+ r5 = r5 + r1 + ((((sel >> 30u) & 1u) != 0u) ? 0xd3177981u : 0xc7934706u); // 13 add
+ r0 = r0 ^ ds[r4 & mask]; // 14 load
+ r2 = r2 - r4; // 15 sub
+ r2 = r2 ^ ds[r0 & mask]; // 16 load
+ r7 = r7 ^ ds[r2 & mask]; // 17 load
+ { uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 18 shfl
+ r5 = r5 * r0; // 19 mul
+ { uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 20 shfl
+ { uint t_; IGNEUM_SHFL_XOR(t_, r4, 16u); r2 = r2 ^ t_; } // 21 shfl
+ r6 = mul_hi(r6, r2); // 22 mulhi
+ r6 = r6 ^ ds[r1 & mask]; // 23 load
+ r5 = r5 * r0; // 24 mul
+ r5 = rotl_imm(r5, 19u); // 25 rotl
+ { uint t_; IGNEUM_SHFL_XOR(t_, r6, 2u); r7 = r7 ^ t_; } // 26 shfl
+ r0 = r0 ^ r5; // 27 xor
+ r0 = r0 ^ r4; // 28 xor
+ r3 = r3 - r0; // 29 sub
+ r5 = r5 * r1; // 30 mul
+ r7 = r7 ^ ds[r2 & mask]; // 31 load
+ r1 = r1 ^ ds[r0 & mask]; // 32 load
+ r5 = r5 ^ r6; // 33 xor
+ r5 = r5 ^ ds[r1 & mask]; // 34 load
+ r0 = mul_hi(r0, r5); // 35 mulhi
+ { uint t_; IGNEUM_SHFL_XOR(t_, r2, 4u); r5 = r5 ^ t_; } // 36 shfl
+ r7 = r7 ^ ds[r0 & mask]; // 37 load
+ r3 = r3 + r1 + ((((sel >> 27u) & 1u) != 0u) ? 0x230c005cu : 0x75ba2fadu); // 38 add
+ { uint t_; IGNEUM_SHFL_XOR(t_, r5, 4u); r1 = r1 ^ t_; } // 39 shfl
+ r2 = r2 ^ r5; // 40 xor
+ r3 = r6 * r3 + r3; // 41 mad
+ r6 = r6 - r7; // 42 sub
+ r7 = r7 ^ r0; // 43 xor
+ r1 = r1 ^ ds[r7 & mask]; // 44 load
+ r2 = r2 * r3; // 45 mul
+ r1 = mul_hi(r1, r5); // 46 mulhi
+ r4 = r4 - r3; // 47 sub
+ r2 = rotr_var(r2, r6); // 48 rotr
+ r3 = r3 ^ ds[r5 & mask]; // 49 load
+ r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 50 add
+ r0 = r0 * r2; // 51 mul
+ r0 = r0 + r2 + ((((sel >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); // 52 add
+ r1 = r1 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0x77b1520du : 0x2bb965afu); // 53 add
+ r7 = rotl_imm(r7, 14u); // 54 rotl
+ r3 = r3 + r7 + ((((sel >> 1u) & 1u) != 0u) ? 0xa54c55a0u : 0x7b0fe07au); // 55 add
+ r6 = r6 ^ ds[r7 & mask]; // 56 load
+ r1 = rotr_var(r1, r5); // 57 rotr
+ r5 = r5 ^ ds[r4 & mask]; // 58 load
+ r6 = r6 ^ ds[r2 & mask]; // 59 load
+ r3 = r5 * r0 + r3; // 60 mad
+ r5 = r5 + r7 + ((((sel >> 31u) & 1u) != 0u) ? 0xad7493e7u : 0xaf9dd72du); // 61 add
+ r4 = r4 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); // 62 add
+ r5 = rotl_imm(r5, 19u); // 63 rotl
+ }
+ uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
+ uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
+ out[gid] = ((ulong)hi << 32) | (ulong)lo;
+}
diff --git a/proto-cuda/packs-ca2-mixer/mx8-genesis/kernel_bound.cu b/proto-cuda/packs-ca2-mixer/mx8-genesis/kernel_bound.cu
new file mode 100644
index 000000000..e8f1a9a79
--- /dev/null
+++ b/proto-cuda/packs-ca2-mixer/mx8-genesis/kernel_bound.cu
@@ -0,0 +1,123 @@
+// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
+// Header-bound twin of igneum_hash in kernel.cu: the init words come from a kernel argument, not SEEDW.
+// Host declarations (also in program_bound.h if present):
+// struct IgneumInitWords { uint32_t w[8]; };
+// cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
+// IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps);
+// cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
+#include
+#include
+#include "program.h"
+
+struct IgneumInitWords { uint32_t w[8]; };
+
+__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
+ x ^= x >> 16; x *= 0x7feb352du;
+ x ^= x >> 15; x *= 0x846ca68bu;
+ x ^= x >> 16;
+ return x;
+}
+__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
+__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
+
+__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw) {
+ uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
+ uint32_t nonce = baseNonce + gid;
+ uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
+ { uint32_t x = nonce ^ iw.w[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw.w[1]; }
+ { uint32_t x = nonce ^ iw.w[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw.w[2]; }
+ { uint32_t x = nonce ^ iw.w[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw.w[3]; }
+ { uint32_t x = nonce ^ iw.w[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw.w[4]; }
+ { uint32_t x = nonce ^ iw.w[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw.w[5]; }
+ { uint32_t x = nonce ^ iw.w[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw.w[6]; }
+ { uint32_t x = nonce ^ iw.w[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw.w[7]; }
+ { uint32_t x = nonce ^ iw.w[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw.w[0]; }
+
+ for (uint32_t it = 0u; it < 8u; ++it) {
+ uint32_t sel = r0;
+ r2 = r3 * r4 + r2; // 0 mad
+ r2 = r1 * r1 + r2; // 1 mad
+ r2 = r3 * r2 + r2; // 2 mad
+ r3 = r3 ^ r5; // 3 xor
+ r7 = r7 ^ ds[r2 & mask]; // 4 load
+ r5 = r5 ^ ds[r7 & mask]; // 5 load
+ r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 6 shfl
+ r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 8); // 7 shfl
+ r1 = __umulhi(r1, r5); // 8 mulhi
+ r6 = rotr_var(r6, r3); // 9 rotr
+ r3 = r3 | r4; // 10 or
+ r4 = r4 ^ ds[r3 & mask]; // 11 load
+ r0 = __umulhi(r0, r4); // 12 mulhi
+ r5 = r5 + r1 + ((((sel >> 30u) & 1u) != 0u) ? 0xd3177981u : 0xc7934706u); // 13 add
+ r0 = r0 ^ ds[r4 & mask]; // 14 load
+ r2 = r2 - r4; // 15 sub
+ r2 = r2 ^ ds[r0 & mask]; // 16 load
+ r7 = r7 ^ ds[r2 & mask]; // 17 load
+ r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 18 shfl
+ r5 = r5 * r0; // 19 mul
+ r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 20 shfl
+ r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r4, 16); // 21 shfl
+ r6 = __umulhi(r6, r2); // 22 mulhi
+ r6 = r6 ^ ds[r1 & mask]; // 23 load
+ r5 = r5 * r0; // 24 mul
+ r5 = rotl_imm(r5, 19u); // 25 rotl
+ r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r6, 2); // 26 shfl
+ r0 = r0 ^ r5; // 27 xor
+ r0 = r0 ^ r4; // 28 xor
+ r3 = r3 - r0; // 29 sub
+ r5 = r5 * r1; // 30 mul
+ r7 = r7 ^ ds[r2 & mask]; // 31 load
+ r1 = r1 ^ ds[r0 & mask]; // 32 load
+ r5 = r5 ^ r6; // 33 xor
+ r5 = r5 ^ ds[r1 & mask]; // 34 load
+ r0 = __umulhi(r0, r5); // 35 mulhi
+ r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r2, 4); // 36 shfl
+ r7 = r7 ^ ds[r0 & mask]; // 37 load
+ r3 = r3 + r1 + ((((sel >> 27u) & 1u) != 0u) ? 0x230c005cu : 0x75ba2fadu); // 38 add
+ r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r5, 4); // 39 shfl
+ r2 = r2 ^ r5; // 40 xor
+ r3 = r6 * r3 + r3; // 41 mad
+ r6 = r6 - r7; // 42 sub
+ r7 = r7 ^ r0; // 43 xor
+ r1 = r1 ^ ds[r7 & mask]; // 44 load
+ r2 = r2 * r3; // 45 mul
+ r1 = __umulhi(r1, r5); // 46 mulhi
+ r4 = r4 - r3; // 47 sub
+ r2 = rotr_var(r2, r6); // 48 rotr
+ r3 = r3 ^ ds[r5 & mask]; // 49 load
+ r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 50 add
+ r0 = r0 * r2; // 51 mul
+ r0 = r0 + r2 + ((((sel >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); // 52 add
+ r1 = r1 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0x77b1520du : 0x2bb965afu); // 53 add
+ r7 = rotl_imm(r7, 14u); // 54 rotl
+ r3 = r3 + r7 + ((((sel >> 1u) & 1u) != 0u) ? 0xa54c55a0u : 0x7b0fe07au); // 55 add
+ r6 = r6 ^ ds[r7 & mask]; // 56 load
+ r1 = rotr_var(r1, r5); // 57 rotr
+ r5 = r5 ^ ds[r4 & mask]; // 58 load
+ r6 = r6 ^ ds[r2 & mask]; // 59 load
+ r3 = r5 * r0 + r3; // 60 mad
+ r5 = r5 + r7 + ((((sel >> 31u) & 1u) != 0u) ? 0xad7493e7u : 0xaf9dd72du); // 61 add
+ r4 = r4 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); // 62 add
+ r5 = rotl_imm(r5, 19u); // 63 rotl
+ }
+ uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
+ uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
+ out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
+}
+
+cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
+ IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps) {
+ if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
+ uint32_t block = 32u * blockWarps;
+ if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
+ igneum_hash_bound<<>>(ds, out, baseNonce, mask, iw);
+ return cudaGetLastError();
+}
+
+cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
+ cudaFuncAttributes attr;
+ cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash_bound);
+ if (e != cudaSuccess) return e;
+ *numRegs = attr.numRegs;
+ return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash_bound, (int)(32u * blockWarps), 0);
+}
diff --git a/proto-cuda/packs-ca2-mixer/mx8-genesis/memhard.h b/proto-cuda/packs-ca2-mixer/mx8-genesis/memhard.h
new file mode 100644
index 000000000..f7f34c732
--- /dev/null
+++ b/proto-cuda/packs-ca2-mixer/mx8-genesis/memhard.h
@@ -0,0 +1,109 @@
+// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
+// Memory-hard dataset core, the same text that the Mac's Metal kernels and CPU verifier were checked against.
+// Included by kernel.cu (device), host.cu (host reference) and proto-opencl/host.c (C99 host reference).
+// See proto-metal/MEMHARD.md for the construction. kernel.cl carries the same text in OpenCL C.
+#pragma once
+#ifdef __cplusplus
+#include
+#else
+#include
+#endif
+#if defined(__CUDACC__)
+#define IGNEUM_HD __host__ __device__ __forceinline__
+#elif defined(_MSC_VER) && !defined(__cplusplus)
+#define IGNEUM_HD static __inline
+#else
+#define IGNEUM_HD static inline
+#endif
+// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
+// Item: 8 rounds of 8 x seed-parameterised mixer + one 64-byte cache read, then 8 x final mixer (class v3, mixer multiplier 8,
+// docs/plans/mixer-x4.md: the round key of application j of round r is 0x9E3779B9 * (r * m + j + 1)). All parameters are literals.
+#define MH_CACHE_LINE_MASK 0x003fffffu
+#define MH_SEGMENT_LINES 64u
+#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
+IGNEUM_HD uint32_t mh_rotl(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
+
+// y = ChaCha12 core(x) + x
+IGNEUM_HD void mh_chacha_block(const uint32_t* x, uint32_t* y) {
+ for (uint32_t i = 0u; i < 16u; ++i) y[i] = x[i];
+ for (uint32_t r = 0u; r < 6u; ++r) {
+ MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
+ MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
+ MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
+ MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
+ }
+ for (uint32_t i = 0u; i < 16u; ++i) y[i] += x[i];
+}
+
+// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
+IGNEUM_HD void mh_cache_segment(uint32_t* cache, uint32_t seg) {
+ uint32_t prev[16]; uint32_t x[16]; uint32_t y[16];
+ for (uint32_t i = 0u; i < 16u; ++i) prev[i] = 0u;
+ for (uint32_t j = 0u; j < MH_SEGMENT_LINES; ++j) {
+ x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
+ x[4] = 0x3067619fu ^ prev[4];
+ x[5] = 0x3c269176u ^ prev[5];
+ x[6] = 0x84a03b03u ^ prev[6];
+ x[7] = 0xf8c63294u ^ prev[7];
+ x[8] = 0xff977c5bu ^ prev[8];
+ x[9] = 0xe60def3eu ^ prev[9];
+ x[10] = 0x63630141u ^ prev[10];
+ x[11] = 0xb8fbcb58u ^ prev[11];
+ x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
+ mh_chacha_block(x, y);
+ uint32_t* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
+ for (uint32_t i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
+ }
+}
+
+// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
+IGNEUM_HD void mh_mixer(uint32_t* s, uint32_t rk) {
+ s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u;
+ s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu;
+ s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u;
+ s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu;
+ s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u;
+ s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u;
+ s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u;
+ s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u;
+ s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du;
+ s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u;
+ s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du;
+ s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu;
+ s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du;
+ s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu;
+ s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u;
+ s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u;
+ MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u)
+ MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u)
+ MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u)
+ MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u)
+}
+
+// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of 8 x mixer + cache line s[0] & mask; 8 x final mixer.
+IGNEUM_HD void mh_item(const uint32_t* cache, uint32_t t, uint32_t* s) {
+ s[0] = 0x3067619fu;
+ s[1] = 0x3c269176u;
+ s[2] = 0x84a03b03u;
+ s[3] = 0xf8c63294u;
+ s[4] = 0xff977c5bu;
+ s[5] = 0xe60def3eu;
+ s[6] = 0x63630141u;
+ s[7] = 0xb8fbcb58u;
+ s[8] = t * 0x42146205u + 0xbab68293u;
+ s[9] = t * 0x52cbe0fbu + 0xcc162340u;
+ s[10] = t * 0x7ecf4a03u + 0x6ce151ccu;
+ s[11] = t * 0x6728907fu + 0xe62b8997u;
+ s[12] = t * 0xd81d9751u + 0xc9c80297u;
+ s[13] = t * 0x132952c3u + 0xf74a1654u;
+ s[14] = t * 0xf60de277u + 0x3d704af5u;
+ s[15] = t * 0x05358035u + 0x3cf522b7u;
+ for (uint32_t r = 0u; r < 8u; ++r) {
+ for (uint32_t j = 0u; j < 8u; ++j) mh_mixer(s, 0x9E3779B9u * (r * 8u + j + 1u));
+ const uint32_t* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
+ for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= line[i];
+ }
+ for (uint32_t j = 0u; j < 8u; ++j) mh_mixer(s, 0x9E3779B9u * (64u + j + 1u));
+}
+// dataset[w] without the dataset: derive item w >> 4 and take word w & 15.
+IGNEUM_HD uint32_t mh_word(const uint32_t* cache, uint32_t w) { uint32_t s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; }
diff --git a/proto-cuda/packs-ca2-mixer/mx8-genesis/memhard.metal b/proto-cuda/packs-ca2-mixer/mx8-genesis/memhard.metal
new file mode 100644
index 000000000..01b263d6d
--- /dev/null
+++ b/proto-cuda/packs-ca2-mixer/mx8-genesis/memhard.metal
@@ -0,0 +1,107 @@
+#include
+using namespace metal;
+// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
+// Item: 8 rounds of 8 x seed-parameterised mixer + one 64-byte cache read, then 8 x final mixer (class v3, mixer multiplier 8,
+// docs/plans/mixer-x4.md: the round key of application j of round r is 0x9E3779B9 * (r * m + j + 1)). All parameters are literals.
+#define MH_CACHE_LINE_MASK 0x003fffffu
+#define MH_SEGMENT_LINES 64u
+#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
+inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
+
+// y = ChaCha12 core(x) + x
+inline void mh_chacha_block(const thread uint* x, thread uint* y) {
+ for (uint i = 0u; i < 16u; ++i) y[i] = x[i];
+ for (uint r = 0u; r < 6u; ++r) {
+ MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
+ MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
+ MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
+ MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
+ }
+ for (uint i = 0u; i < 16u; ++i) y[i] += x[i];
+}
+
+// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
+inline void mh_cache_segment(device uint* cache, uint seg) {
+ uint prev[16]; uint x[16]; uint y[16];
+ for (uint i = 0u; i < 16u; ++i) prev[i] = 0u;
+ for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) {
+ x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
+ x[4] = 0x3067619fu ^ prev[4];
+ x[5] = 0x3c269176u ^ prev[5];
+ x[6] = 0x84a03b03u ^ prev[6];
+ x[7] = 0xf8c63294u ^ prev[7];
+ x[8] = 0xff977c5bu ^ prev[8];
+ x[9] = 0xe60def3eu ^ prev[9];
+ x[10] = 0x63630141u ^ prev[10];
+ x[11] = 0xb8fbcb58u ^ prev[11];
+ x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
+ mh_chacha_block(x, y);
+ device uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
+ for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
+ }
+}
+
+// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
+inline void mh_mixer(thread uint* s, uint rk) {
+ s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u;
+ s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu;
+ s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u;
+ s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu;
+ s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u;
+ s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u;
+ s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u;
+ s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u;
+ s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du;
+ s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u;
+ s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du;
+ s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu;
+ s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du;
+ s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu;
+ s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u;
+ s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u;
+ MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u)
+ MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u)
+ MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u)
+ MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u)
+}
+
+// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of 8 x mixer + cache line s[0] & mask; 8 x final mixer.
+inline void mh_item(device const uint* cache, uint t, thread uint* s) {
+ s[0] = 0x3067619fu;
+ s[1] = 0x3c269176u;
+ s[2] = 0x84a03b03u;
+ s[3] = 0xf8c63294u;
+ s[4] = 0xff977c5bu;
+ s[5] = 0xe60def3eu;
+ s[6] = 0x63630141u;
+ s[7] = 0xb8fbcb58u;
+ s[8] = t * 0x42146205u + 0xbab68293u;
+ s[9] = t * 0x52cbe0fbu + 0xcc162340u;
+ s[10] = t * 0x7ecf4a03u + 0x6ce151ccu;
+ s[11] = t * 0x6728907fu + 0xe62b8997u;
+ s[12] = t * 0xd81d9751u + 0xc9c80297u;
+ s[13] = t * 0x132952c3u + 0xf74a1654u;
+ s[14] = t * 0xf60de277u + 0x3d704af5u;
+ s[15] = t * 0x05358035u + 0x3cf522b7u;
+ for (uint r = 0u; r < 8u; ++r) {
+ for (uint j = 0u; j < 8u; ++j) mh_mixer(s, 0x9E3779B9u * (r * 8u + j + 1u));
+ device const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
+ for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i];
+ }
+ for (uint j = 0u; j < 8u; ++j) mh_mixer(s, 0x9E3779B9u * (64u + j + 1u));
+}
+// dataset[w] without the dataset: derive item w >> 4 and take word w & 15.
+inline uint mh_word(device const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; }
+
+// One thread per segment (2^16 threads).
+kernel void igneum_cache_fill(device uint* cache [[buffer(0)]], uint gid [[thread_position_in_grid]]) {
+ mh_cache_segment(cache, gid);
+}
+// One thread per 64-byte item (dataset words / 16 threads).
+kernel void igneum_build(device const uint* cache [[buffer(0)]], device uint* dataset [[buffer(1)]],
+ uint gid [[thread_position_in_grid]]) {
+ uint s[16];
+ mh_item(cache, gid, s);
+ device uint* d = dataset + gid * 16u;
+ for (uint i = 0u; i < 16u; ++i) d[i] = s[i];
+}
diff --git a/proto-cuda/packs-ca2-mixer/mx8-genesis/program.h b/proto-cuda/packs-ca2-mixer/mx8-genesis/program.h
new file mode 100644
index 000000000..9c99a1697
--- /dev/null
+++ b/proto-cuda/packs-ca2-mixer/mx8-genesis/program.h
@@ -0,0 +1,63 @@
+// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
+// Program metadata for host.cu plus the launch wrappers defined in kernel.cu.
+// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros.
+#pragma once
+#ifdef __cplusplus
+#include
+#else
+#include
+#endif
+#ifndef IGNEUM_NO_CUDA
+#include
+#endif
+
+#define IGNEUM_SEED_STRING "igneum-genesis"
+#define IGNEUM_SEED_BYTES_HEX "69676e65756d2d67656e65736973"
+#define IGNEUM_GENERATOR 2
+#define IGNEUM_PROGRAM_ATTEMPT 0
+#define IGNEUM_PROGRAM_ID 0x9544515847736361ull
+#define IGNEUM_DAY_STRING "2026-10-03"
+#define IGNEUM_DAY_BYTES_HEX "6461792f323032362d31302d3033"
+#define IGNEUM_DAY0 0x3067619fu
+#define IGNEUM_DAY1 0x3c269176u
+#define IGNEUM_DATASET_LOG2 28
+#define IGNEUM_MASK 0x0fffffffu
+#define IGNEUM_LANES 32
+#define IGNEUM_ITERATIONS 8
+#define IGNEUM_INSTR_COUNT 64
+#define IGNEUM_LOADS_PER_HASH 128
+#define IGNEUM_WIDE_LOADS_PER_HASH 0
+#define IGNEUM_OP_MIX "load=16 add=8 shfl=8 xor=6 mad=5 mul=5 mulhi=5 sub=4 rotl=3 rotr=3 or=1"
+// Class v3 construction (Counter ASIC 2.0, 5 October 2026, docs/plans/mixer-x4.md): version 2 loads; the dataset item
+// derivation applies the mixer IGNEUM_MIXER_MULT times per round (memhard.h), and the cache follows the growth rule.
+#define IGNEUM_LOAD_CLASS "mx8"
+#define IGNEUM_CLASS_MIXER_MULT 8
+#define IGNEUM_CACHE_GROWTH 1 // 1: cache words = 2^(26 + doublings(day)), doublings = floor(log2(1 + day / 1460))
+#define IGNEUM_LOAD_SLOTS 16
+#define IGNEUM_LOAD_MIX { 100, 0, 0 }
+#define IGNEUM_LOAD_WIDTH_COUNTS { 16, 0, 0 } // loads of 4, 16, 64 bytes per program
+#define IGNEUM_BYTES_PER_HASH 512
+#define IGNEUM_FOLD_ROT 11
+#define IGNEUM_FOLD_MUL 0x9e3779b1u
+// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h)
+#define IGNEUM_DATASET_MODE 1
+
+#define IGNEUM_SEEDW_INIT { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u }
+#define IGNEUM_KEY_INIT { 0x3067619fu, 0x3c269176u, 0x84a03b03u, 0xf8c63294u, 0xff977c5bu, 0xe60def3eu, 0x63630141u, 0xb8fbcb58u }
+#define IGNEUM_CACHE_LOG2_WORDS 26
+#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6
+#define IGNEUM_CACHE_SEGMENTS 65536u
+#define IGNEUM_ITEM_ROUNDS 8
+#define IGNEUM_MIXER_MULT 8 // mixer applications per round and after the last read (class v3, docs/plans/mixer-x4.md)
+#define IGNEUM_MIX_ROT_INIT { 20u, 20u, 19u, 4u, 26u, 3u, 3u, 27u }
+#define IGNEUM_MIX_MUL_INIT { 0x42146205u, 0x52cbe0fbu, 0x7ecf4a03u, 0x6728907fu, 0xd81d9751u, 0x132952c3u, 0xf60de277u, 0x05358035u, 0xbaf6499du, 0xe4db9667u, 0x3e98f45du, 0xd0004eddu, 0x2691630du, 0x9beb3bcfu, 0xab310379u, 0x99cfb423u }
+#define IGNEUM_MIX_RC_INIT { 0xbab68293u, 0xcc162340u, 0x6ce151ccu, 0xe62b8997u, 0xc9c80297u, 0xf74a1654u, 0x3d704af5u, 0x3cf522b7u, 0x2b9cac04u, 0xa880ac10u, 0x13e5dd1du, 0x6fc3e233u, 0x2d83eeacu, 0x9006e8bfu, 0x2c4b5362u, 0x31b49ee2u }
+
+#ifndef IGNEUM_NO_CUDA
+// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError().
+cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments);
+cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems);
+cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
+ uint32_t nonces, uint32_t blockWarps);
+cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
+#endif
diff --git a/proto-cuda/packs-ca2-mixer/mx8-genesis/program.json b/proto-cuda/packs-ca2-mixer/mx8-genesis/program.json
new file mode 100644
index 000000000..d4e1c079a
--- /dev/null
+++ b/proto-cuda/packs-ca2-mixer/mx8-genesis/program.json
@@ -0,0 +1,131 @@
+{
+ "format": "igneum-program-pack-3",
+ "generator": 2,
+ "attempt": 0,
+ "program_id": "0x9544515847736361",
+ "program_id_derivation": "FNV-1a 64 over 'igneum-program/' || generator_le32 || seed_words as little-endian bytes || attempt_le32",
+ "dataset_mode": "memory-hard",
+ "seed": "igneum-genesis",
+ "seed_bytes": "69676e65756d2d67656e65736973",
+ "seed_words": ["0x67a9a7be", "0x1a155b25", "0xfddfb732", "0x4b5af2e8", "0xc55caf33", "0xa27c13b7", "0x06628a48", "0x03852469"],
+ "seed_derivation": "seed_words = FNV-1a 64 over seed_bytes (attempt 0) or seed_bytes || attempt_le32 (attempt k >= 1), basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32",
+ "generator_rule": "version 2: exactly 16 load slots drawn first from instructions 1..63 (partial Fisher-Yates), the other 48 ops from the ten non-load weights (sum 75); a load's source is drawn from the registers other than dst written by an earlier instruction and not read by a load since; the candidate must pass the acceptance rule of spec 01 section 1.4.6 (static: no cyclically stale load source, every register has an injecting write; dynamic: 64 units on the seed-keyed closed-form dataset with no constant register bit, no lane-constant load site, under 164 saturated final values, every output bit within 136 of 1024, distinct addresses above 245760), else the next attempt of the seed is tried",
+ "lanes": 32,
+ "registers": 8,
+ "iterations": 8,
+ "instruction_count": 64,
+ "loads_per_hash": 128,
+ "load_class": "mx8",
+ "mixer_mult": 8,
+ "cache_growth": true,
+ "mixer": "class v3 (Counter ASIC 2.0, 5 October 2026, docs/plans/mixer-x4.md): every mixer application of the item derivation is 8 applications with round keys (r * 8 + j + 1) * 0x9E3779B9, the 8 dependent cache reads per item unchanged; cache growth rule option C: cache words = 2^(26 + doublings(day)), dataset words = 2^(genesis_log2 + doublings(day)), doublings(day) = floor(log2(1 + day / 1460)) for day = days since genesis",
+ "load_slots": 16,
+ "load_mix_percent_4_16_64": [100, 0, 0],
+ "load_width_counts_4_16_64": [16, 0, 0],
+ "bytes_per_hash": 512,
+ "wide_load": "read-width experiment (5 October 2026, docs/plans/read-width.md), NOT the lottery hash: a load of W words (width field, 4 or 16) reads dataset[b .. b + W) with b = (src & mask) & ~(W - 1) and folds every word into dst: x = dst ^ w[0]; for j in 1..W: x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x; width 1 is the plain load; the width is drawn per instruction from the class mix with one extra below(100) draw after the nine of version 2, and the program id is FNV-1a 64 over 'igneum-program-rw/' || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots",
+ "op_mix": {"load": 16, "add": 8, "shfl": 8, "xor": 6, "mad": 5, "mul": 5, "mulhi": 5, "sub": 4, "rotl": 3, "rotr": 3, "or": 1},
+ "register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]",
+ "splitmix32": "x ^= x>>16; x *= 0x7feb352d; x ^= x>>15; x *= 0x846ca68b; x ^= x>>16",
+ "iteration": "sel = r0 sampled once at the top of each iteration, then all instructions in order",
+ "output": "lo = r0 ^ rotl(r1,7) ^ rotl(r2,14) ^ rotl(r3,21); hi = r4 ^ rotl(r5,9) ^ rotl(r6,18) ^ rotl(r7,27); out = (hi << 32) | lo",
+ "op_semantics": {
+ "add": "dst = dst + src + (bit `bit` of sel ? imm2 : imm)",
+ "sub": "dst = dst - src",
+ "mul": "dst = dst * src (low 32)",
+ "mulhi": "dst = high 32 bits of dst * src",
+ "xor": "dst = dst ^ src",
+ "or": "dst = dst | src",
+ "rotl": "dst = rotl(dst, rot), rot in 1..31",
+ "rotr": "dst = rotr(dst, src & 31)",
+ "mad": "dst = src * src2 + dst",
+ "shfl": "dst = dst ^ (src of lane (lane ^ mask)), mask in {1,2,4,8,16}, within the 32-lane warp",
+ "load": "dst = dst ^ dataset[src & dataset.mask]",
+ "wload": "base = (src of lane 0 & dataset.mask) & ~31; dst = dst ^ dataset[base + lane] (warp-coalesced 128-byte load, lever b, only when --wide-frac > 0)"
+ },
+ "dataset": {
+ "log2_words": 28,
+ "bytes": 1073741824,
+ "mask": "0x0fffffff",
+ "day": "2026-10-03",
+ "day_bytes": "6461792f323032362d31302d3033",
+ "day_words_from": "seed_words_from_bytes(day_bytes)",
+ "d0": "0x3067619f",
+ "d1": "0x3c269176",
+ "mode": "memory-hard",
+ "spec": "proto-metal/MEMHARD.md",
+ "key": ["0x3067619f", "0x3c269176", "0x84a03b03", "0xf8c63294", "0xff977c5b", "0xe60def3e", "0x63630141", "0xb8fbcb58"],
+ "key_derivation": "the 8 words of seed_words_from_bytes(day_bytes); d0, d1 are key[0], key[1]",
+ "cache": {"log2_words": 26, "bytes": 268435456, "line_words": 16, "segment_lines": 64, "segments": 65536, "block": "ChaCha12 core + feed-forward, rotations 16 12 8 7", "sigma": ["0x61707865", "0x3320646e", "0x79622d32", "0x6b206574"], "tag": ["0x49676e65", "0x756d4d48"], "chain": "in_j = prev_line ^ (sigma[0..3] || key[0..7] || seg || j || tag[0..1]); line_j = block(in_j); prev_0 = 0"},
+ "mixer": {"draw": "SplitMix64 seeded with key[0] | key[1] << 32: rot[0..7] = 1 + next() % 31, mul[0..15] = low32(next()) | 1, rc[0..15] = low32(next())", "rot": [20, 20, 19, 4, 26, 3, 3, 27], "mul": ["0x42146205", "0x52cbe0fb", "0x7ecf4a03", "0x6728907f", "0xd81d9751", "0x132952c3", "0xf60de277", "0x05358035", "0xbaf6499d", "0xe4db9667", "0x3e98f45d", "0xd0004edd", "0x2691630d", "0x9beb3bcf", "0xab310379", "0x99cfb423"], "rc": ["0xbab68293", "0xcc162340", "0x6ce151cc", "0xe62b8997", "0xc9c80297", "0xf74a1654", "0x3d704af5", "0x3cf522b7", "0x2b9cac04", "0xa880ac10", "0x13e5dd1d", "0x6fc3e233", "0x2d83eeac", "0x9006e8bf", "0x2c4b5362", "0x31b49ee2"], "round": "for i in 0..15: s[i] = (s[i] ^ (rc[i] + (r+1) * 0x9E3779B9)) * mul[i]; then quarter rounds on columns (0,4,8,12) (1,5,9,13) (2,6,10,14) (3,7,11,15) with rot[0..3] and diagonals (0,5,10,15) (1,6,11,12) (2,7,8,13) (3,4,9,14) with rot[4..7]", "quarter_round": "a += b; d ^= a; d = rotl(d, r1); c += d; b ^= c; b = rotl(b, r2); a += b; d ^= a; d = rotl(d, r3); c += d; b ^= c; b = rotl(b, r4)"},
+ "mixer_mult": 8,
+ "item": "s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..7: for j in 0..7: s = M(s, rk = (r * 8 + j + 1) * 0x9E3779B9); line = s[0] & 0x003fffff; s[i] ^= cache[line * 16 + i]; then for j in 0..7: s = M(s, rk = (64 + j + 1) * 0x9E3779B9); item(t) = s",
+ "word": "dataset[w] = item(w >> 4)[w & 15]"
+ },
+ "instructions": [
+ {"i": 0, "op": "mad", "dst": 2, "src": 3, "src2": 4, "imm": "0xbf7b174d", "imm2": "0x337b762e", "rot": 17, "bit": 2, "mask": 2, "width": 1},
+ {"i": 1, "op": "mad", "dst": 2, "src": 1, "src2": 1, "imm": "0xdd04a5da", "imm2": "0x42da7657", "rot": 15, "bit": 30, "mask": 16, "width": 1},
+ {"i": 2, "op": "mad", "dst": 2, "src": 3, "src2": 2, "imm": "0x734003fa", "imm2": "0x5bb67700", "rot": 3, "bit": 20, "mask": 1, "width": 1},
+ {"i": 3, "op": "xor", "dst": 3, "src": 5, "src2": 5, "imm": "0xc55a1b1c", "imm2": "0xa19720f3", "rot": 7, "bit": 8, "mask": 1, "width": 1},
+ {"i": 4, "op": "load", "dst": 7, "src": 2, "src2": 5, "imm": "0xad572dd7", "imm2": "0x9ceb3ea7", "rot": 18, "bit": 30, "mask": 2, "width": 1},
+ {"i": 5, "op": "load", "dst": 5, "src": 7, "src2": 3, "imm": "0x769a53be", "imm2": "0x80f9067e", "rot": 12, "bit": 22, "mask": 1, "width": 1},
+ {"i": 6, "op": "shfl", "dst": 1, "src": 4, "src2": 0, "imm": "0xd3613d88", "imm2": "0x262fb219", "rot": 10, "bit": 30, "mask": 8, "width": 1},
+ {"i": 7, "op": "shfl", "dst": 7, "src": 3, "src2": 4, "imm": "0xce38e42f", "imm2": "0xb868b818", "rot": 11, "bit": 8, "mask": 8, "width": 1},
+ {"i": 8, "op": "mulhi", "dst": 1, "src": 5, "src2": 2, "imm": "0xa5eebca5", "imm2": "0x5703a72b", "rot": 13, "bit": 13, "mask": 16, "width": 1},
+ {"i": 9, "op": "rotr", "dst": 6, "src": 3, "src2": 4, "imm": "0x17a5a9c7", "imm2": "0xdcfb93a1", "rot": 20, "bit": 27, "mask": 2, "width": 1},
+ {"i": 10, "op": "or", "dst": 3, "src": 4, "src2": 1, "imm": "0xccb7d785", "imm2": "0xc335364c", "rot": 14, "bit": 12, "mask": 4, "width": 1},
+ {"i": 11, "op": "load", "dst": 4, "src": 3, "src2": 2, "imm": "0x88cb9af3", "imm2": "0x4e7dc10d", "rot": 24, "bit": 17, "mask": 4, "width": 1},
+ {"i": 12, "op": "mulhi", "dst": 0, "src": 4, "src2": 4, "imm": "0x45374321", "imm2": "0x3cd91989", "rot": 11, "bit": 4, "mask": 2, "width": 1},
+ {"i": 13, "op": "add", "dst": 5, "src": 1, "src2": 2, "imm": "0xc7934706", "imm2": "0xd3177981", "rot": 16, "bit": 30, "mask": 2, "width": 1},
+ {"i": 14, "op": "load", "dst": 0, "src": 4, "src2": 3, "imm": "0xfd7f56bb", "imm2": "0x65e14f52", "rot": 13, "bit": 22, "mask": 2, "width": 1},
+ {"i": 15, "op": "sub", "dst": 2, "src": 4, "src2": 4, "imm": "0x35a80b49", "imm2": "0x060f2d13", "rot": 16, "bit": 20, "mask": 16, "width": 1},
+ {"i": 16, "op": "load", "dst": 2, "src": 0, "src2": 2, "imm": "0xae0a32c2", "imm2": "0x4c2a4cfe", "rot": 8, "bit": 31, "mask": 16, "width": 1},
+ {"i": 17, "op": "load", "dst": 7, "src": 2, "src2": 6, "imm": "0x82a84cc3", "imm2": "0x21a38d68", "rot": 15, "bit": 21, "mask": 2, "width": 1},
+ {"i": 18, "op": "shfl", "dst": 7, "src": 3, "src2": 3, "imm": "0xa3818806", "imm2": "0x8f66b5c8", "rot": 14, "bit": 6, "mask": 4, "width": 1},
+ {"i": 19, "op": "mul", "dst": 5, "src": 0, "src2": 1, "imm": "0xa00de107", "imm2": "0x77bfcaa5", "rot": 3, "bit": 10, "mask": 2, "width": 1},
+ {"i": 20, "op": "shfl", "dst": 3, "src": 4, "src2": 7, "imm": "0x1d2b8cab", "imm2": "0x80b4f9a2", "rot": 14, "bit": 25, "mask": 2, "width": 1},
+ {"i": 21, "op": "shfl", "dst": 2, "src": 4, "src2": 5, "imm": "0x3ac915d2", "imm2": "0x5fba7bc2", "rot": 16, "bit": 1, "mask": 16, "width": 1},
+ {"i": 22, "op": "mulhi", "dst": 6, "src": 2, "src2": 0, "imm": "0xdc3ec8fd", "imm2": "0x599e2fa3", "rot": 22, "bit": 3, "mask": 2, "width": 1},
+ {"i": 23, "op": "load", "dst": 6, "src": 1, "src2": 5, "imm": "0x2a6b16d5", "imm2": "0xd73e396f", "rot": 28, "bit": 29, "mask": 2, "width": 1},
+ {"i": 24, "op": "mul", "dst": 5, "src": 0, "src2": 7, "imm": "0x376d0223", "imm2": "0xe1c2169a", "rot": 4, "bit": 16, "mask": 16, "width": 1},
+ {"i": 25, "op": "rotl", "dst": 5, "src": 7, "src2": 3, "imm": "0x78ad8c60", "imm2": "0x6f5b77d5", "rot": 19, "bit": 11, "mask": 16, "width": 1},
+ {"i": 26, "op": "shfl", "dst": 7, "src": 6, "src2": 5, "imm": "0x93915b9f", "imm2": "0x1e61fb6b", "rot": 28, "bit": 23, "mask": 2, "width": 1},
+ {"i": 27, "op": "xor", "dst": 0, "src": 5, "src2": 7, "imm": "0x6378fe15", "imm2": "0x66c78f42", "rot": 12, "bit": 31, "mask": 8, "width": 1},
+ {"i": 28, "op": "xor", "dst": 0, "src": 4, "src2": 7, "imm": "0x20a57fda", "imm2": "0x088c848e", "rot": 16, "bit": 13, "mask": 4, "width": 1},
+ {"i": 29, "op": "sub", "dst": 3, "src": 0, "src2": 2, "imm": "0x49d95fd5", "imm2": "0x1a5f946a", "rot": 6, "bit": 12, "mask": 1, "width": 1},
+ {"i": 30, "op": "mul", "dst": 5, "src": 1, "src2": 7, "imm": "0x0a816217", "imm2": "0x405c4f73", "rot": 13, "bit": 27, "mask": 4, "width": 1},
+ {"i": 31, "op": "load", "dst": 7, "src": 2, "src2": 2, "imm": "0x09ed045e", "imm2": "0xd69c4715", "rot": 5, "bit": 9, "mask": 2, "width": 1},
+ {"i": 32, "op": "load", "dst": 1, "src": 0, "src2": 6, "imm": "0xeb79ea49", "imm2": "0xcc587f5a", "rot": 6, "bit": 8, "mask": 16, "width": 1},
+ {"i": 33, "op": "xor", "dst": 5, "src": 6, "src2": 1, "imm": "0x3027401e", "imm2": "0x5f20c27e", "rot": 18, "bit": 9, "mask": 2, "width": 1},
+ {"i": 34, "op": "load", "dst": 5, "src": 1, "src2": 3, "imm": "0x0e1cab07", "imm2": "0x09356c5b", "rot": 19, "bit": 31, "mask": 1, "width": 1},
+ {"i": 35, "op": "mulhi", "dst": 0, "src": 5, "src2": 2, "imm": "0x90e31357", "imm2": "0xabd32484", "rot": 26, "bit": 5, "mask": 8, "width": 1},
+ {"i": 36, "op": "shfl", "dst": 5, "src": 2, "src2": 5, "imm": "0xee9a955f", "imm2": "0x31b3faed", "rot": 8, "bit": 24, "mask": 4, "width": 1},
+ {"i": 37, "op": "load", "dst": 7, "src": 0, "src2": 7, "imm": "0x3ba2f832", "imm2": "0x1160dcd3", "rot": 4, "bit": 29, "mask": 1, "width": 1},
+ {"i": 38, "op": "add", "dst": 3, "src": 1, "src2": 7, "imm": "0x75ba2fad", "imm2": "0x230c005c", "rot": 4, "bit": 27, "mask": 1, "width": 1},
+ {"i": 39, "op": "shfl", "dst": 1, "src": 5, "src2": 3, "imm": "0xdbf37e75", "imm2": "0xb5ac1969", "rot": 30, "bit": 13, "mask": 4, "width": 1},
+ {"i": 40, "op": "xor", "dst": 2, "src": 5, "src2": 5, "imm": "0x47f136c5", "imm2": "0x06ce9153", "rot": 19, "bit": 10, "mask": 2, "width": 1},
+ {"i": 41, "op": "mad", "dst": 3, "src": 6, "src2": 3, "imm": "0xce13eff8", "imm2": "0x04cc1d55", "rot": 3, "bit": 1, "mask": 4, "width": 1},
+ {"i": 42, "op": "sub", "dst": 6, "src": 7, "src2": 1, "imm": "0x6a65ab71", "imm2": "0x8fbc1bcd", "rot": 4, "bit": 1, "mask": 8, "width": 1},
+ {"i": 43, "op": "xor", "dst": 7, "src": 0, "src2": 7, "imm": "0xdaeb4928", "imm2": "0xc0423027", "rot": 24, "bit": 11, "mask": 8, "width": 1},
+ {"i": 44, "op": "load", "dst": 1, "src": 7, "src2": 7, "imm": "0x778f01c9", "imm2": "0x28cedcea", "rot": 12, "bit": 4, "mask": 16, "width": 1},
+ {"i": 45, "op": "mul", "dst": 2, "src": 3, "src2": 2, "imm": "0xf4264f1b", "imm2": "0x0f627d56", "rot": 5, "bit": 28, "mask": 8, "width": 1},
+ {"i": 46, "op": "mulhi", "dst": 1, "src": 5, "src2": 1, "imm": "0xffb2147a", "imm2": "0xccde9b05", "rot": 13, "bit": 9, "mask": 2, "width": 1},
+ {"i": 47, "op": "sub", "dst": 4, "src": 3, "src2": 5, "imm": "0x73b36234", "imm2": "0x3f5d5997", "rot": 7, "bit": 18, "mask": 2, "width": 1},
+ {"i": 48, "op": "rotr", "dst": 2, "src": 6, "src2": 3, "imm": "0x3a4d9aa9", "imm2": "0x212bec7b", "rot": 4, "bit": 29, "mask": 16, "width": 1},
+ {"i": 49, "op": "load", "dst": 3, "src": 5, "src2": 2, "imm": "0x626f11df", "imm2": "0x56cd5bfd", "rot": 7, "bit": 1, "mask": 1, "width": 1},
+ {"i": 50, "op": "add", "dst": 1, "src": 5, "src2": 6, "imm": "0x81b8bc2c", "imm2": "0x1907970c", "rot": 28, "bit": 7, "mask": 4, "width": 1},
+ {"i": 51, "op": "mul", "dst": 0, "src": 2, "src2": 2, "imm": "0xa8848b30", "imm2": "0xef6ac348", "rot": 9, "bit": 15, "mask": 8, "width": 1},
+ {"i": 52, "op": "add", "dst": 0, "src": 2, "src2": 0, "imm": "0x4f92b968", "imm2": "0x699fd448", "rot": 22, "bit": 6, "mask": 4, "width": 1},
+ {"i": 53, "op": "add", "dst": 1, "src": 0, "src2": 2, "imm": "0x2bb965af", "imm2": "0x77b1520d", "rot": 2, "bit": 12, "mask": 8, "width": 1},
+ {"i": 54, "op": "rotl", "dst": 7, "src": 1, "src2": 0, "imm": "0x553e678b", "imm2": "0x3cc8eae0", "rot": 14, "bit": 20, "mask": 2, "width": 1},
+ {"i": 55, "op": "add", "dst": 3, "src": 7, "src2": 2, "imm": "0x7b0fe07a", "imm2": "0xa54c55a0", "rot": 10, "bit": 1, "mask": 1, "width": 1},
+ {"i": 56, "op": "load", "dst": 6, "src": 7, "src2": 4, "imm": "0x01eba9aa", "imm2": "0x2758c0f7", "rot": 14, "bit": 15, "mask": 4, "width": 1},
+ {"i": 57, "op": "rotr", "dst": 1, "src": 5, "src2": 2, "imm": "0x1f5267b3", "imm2": "0x236f5a27", "rot": 2, "bit": 31, "mask": 16, "width": 1},
+ {"i": 58, "op": "load", "dst": 5, "src": 4, "src2": 3, "imm": "0xa9954a9b", "imm2": "0x6a54d4e8", "rot": 11, "bit": 10, "mask": 16, "width": 1},
+ {"i": 59, "op": "load", "dst": 6, "src": 2, "src2": 4, "imm": "0x9923ff88", "imm2": "0x9357254e", "rot": 16, "bit": 1, "mask": 16, "width": 1},
+ {"i": 60, "op": "mad", "dst": 3, "src": 5, "src2": 0, "imm": "0xc0cc51a6", "imm2": "0x3fd7701b", "rot": 20, "bit": 1, "mask": 4, "width": 1},
+ {"i": 61, "op": "add", "dst": 5, "src": 7, "src2": 7, "imm": "0xaf9dd72d", "imm2": "0xad7493e7", "rot": 7, "bit": 31, "mask": 16, "width": 1},
+ {"i": 62, "op": "add", "dst": 4, "src": 6, "src2": 2, "imm": "0x89841d87", "imm2": "0x1e07c3d9", "rot": 6, "bit": 27, "mask": 1, "width": 1},
+ {"i": 63, "op": "rotl", "dst": 5, "src": 4, "src2": 2, "imm": "0xae210f8d", "imm2": "0x8e499ba4", "rot": 19, "bit": 9, "mask": 1, "width": 1}
+ ]
+}
diff --git a/proto-cuda/packs-ca2-mixer/mx8-genesis/program.metal b/proto-cuda/packs-ca2-mixer/mx8-genesis/program.metal
new file mode 100644
index 000000000..d90be2f77
--- /dev/null
+++ b/proto-cuda/packs-ca2-mixer/mx8-genesis/program.metal
@@ -0,0 +1,109 @@
+#include
+using namespace metal;
+
+#define MASK 0x0fffffffu
+constant uint SEEDW[8] = { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u };
+
+inline uint splitmix32(uint x) {
+ x ^= x >> 16; x *= 0x7feb352du;
+ x ^= x >> 15; x *= 0x846ca68bu;
+ x ^= x >> 16;
+ return x;
+}
+inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31
+inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
+inline uint ds_elem(uint i, uint d0, uint d1) {
+ uint x = i ^ d0;
+ x *= 0x9E3779B1u; x ^= x >> 15;
+ x += d1;
+ x *= 0x85EBCA77u; x ^= x >> 13;
+ x *= 0xC2B2AE3Du; x ^= x >> 16;
+ return x;
+}
+
+kernel void igneum_hash(device const uint* dataset [[buffer(0)]],
+ device ulong* out [[buffer(1)]],
+ constant uint& baseNonce [[buffer(2)]],
+ uint gid [[thread_position_in_grid]]) {
+ uint nonce = baseNonce + gid;
+ uint r0, r1, r2, r3, r4, r5, r6, r7;
+ { uint x = nonce ^ SEEDW[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ SEEDW[1]; }
+ { uint x = nonce ^ SEEDW[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ SEEDW[2]; }
+ { uint x = nonce ^ SEEDW[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ SEEDW[3]; }
+ { uint x = nonce ^ SEEDW[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ SEEDW[4]; }
+ { uint x = nonce ^ SEEDW[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ SEEDW[5]; }
+ { uint x = nonce ^ SEEDW[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ SEEDW[6]; }
+ { uint x = nonce ^ SEEDW[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ SEEDW[7]; }
+ { uint x = nonce ^ SEEDW[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ SEEDW[0]; }
+
+ for (uint it = 0u; it < 8u; ++it) {
+ uint sel = r0;
+ r2 = r3 * r4 + r2; // 0
+ r2 = r1 * r1 + r2; // 1
+ r2 = r3 * r2 + r2; // 2
+ r3 = r3 ^ r5; // 3
+ r7 = r7 ^ dataset[r2 & MASK]; // 4
+ r5 = r5 ^ dataset[r7 & MASK]; // 5
+ r1 = r1 ^ simd_shuffle_xor(r4, (ushort)8); // 6
+ r7 = r7 ^ simd_shuffle_xor(r3, (ushort)8); // 7
+ r1 = mulhi(r1, r5); // 8
+ r6 = rotr_var(r6, r3); // 9
+ r3 = r3 | r4; // 10
+ r4 = r4 ^ dataset[r3 & MASK]; // 11
+ r0 = mulhi(r0, r4); // 12
+ r5 = r5 + r1 + select(0xc7934706u, 0xd3177981u, ((sel >> 30u) & 1u) != 0u); // 13
+ r0 = r0 ^ dataset[r4 & MASK]; // 14
+ r2 = r2 - r4; // 15
+ r2 = r2 ^ dataset[r0 & MASK]; // 16
+ r7 = r7 ^ dataset[r2 & MASK]; // 17
+ r7 = r7 ^ simd_shuffle_xor(r3, (ushort)4); // 18
+ r5 = r5 * r0; // 19
+ r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 20
+ r2 = r2 ^ simd_shuffle_xor(r4, (ushort)16); // 21
+ r6 = mulhi(r6, r2); // 22
+ r6 = r6 ^ dataset[r1 & MASK]; // 23
+ r5 = r5 * r0; // 24
+ r5 = rotl_imm(r5, 19u); // 25
+ r7 = r7 ^ simd_shuffle_xor(r6, (ushort)2); // 26
+ r0 = r0 ^ r5; // 27
+ r0 = r0 ^ r4; // 28
+ r3 = r3 - r0; // 29
+ r5 = r5 * r1; // 30
+ r7 = r7 ^ dataset[r2 & MASK]; // 31
+ r1 = r1 ^ dataset[r0 & MASK]; // 32
+ r5 = r5 ^ r6; // 33
+ r5 = r5 ^ dataset[r1 & MASK]; // 34
+ r0 = mulhi(r0, r5); // 35
+ r5 = r5 ^ simd_shuffle_xor(r2, (ushort)4); // 36
+ r7 = r7 ^ dataset[r0 & MASK]; // 37
+ r3 = r3 + r1 + select(0x75ba2fadu, 0x230c005cu, ((sel >> 27u) & 1u) != 0u); // 38
+ r1 = r1 ^ simd_shuffle_xor(r5, (ushort)4); // 39
+ r2 = r2 ^ r5; // 40
+ r3 = r6 * r3 + r3; // 41
+ r6 = r6 - r7; // 42
+ r7 = r7 ^ r0; // 43
+ r1 = r1 ^ dataset[r7 & MASK]; // 44
+ r2 = r2 * r3; // 45
+ r1 = mulhi(r1, r5); // 46
+ r4 = r4 - r3; // 47
+ r2 = rotr_var(r2, r6); // 48
+ r3 = r3 ^ dataset[r5 & MASK]; // 49
+ r1 = r1 + r5 + select(0x81b8bc2cu, 0x1907970cu, ((sel >> 7u) & 1u) != 0u); // 50
+ r0 = r0 * r2; // 51
+ r0 = r0 + r2 + select(0x4f92b968u, 0x699fd448u, ((sel >> 6u) & 1u) != 0u); // 52
+ r1 = r1 + r0 + select(0x2bb965afu, 0x77b1520du, ((sel >> 12u) & 1u) != 0u); // 53
+ r7 = rotl_imm(r7, 14u); // 54
+ r3 = r3 + r7 + select(0x7b0fe07au, 0xa54c55a0u, ((sel >> 1u) & 1u) != 0u); // 55
+ r6 = r6 ^ dataset[r7 & MASK]; // 56
+ r1 = rotr_var(r1, r5); // 57
+ r5 = r5 ^ dataset[r4 & MASK]; // 58
+ r6 = r6 ^ dataset[r2 & MASK]; // 59
+ r3 = r5 * r0 + r3; // 60
+ r5 = r5 + r7 + select(0xaf9dd72du, 0xad7493e7u, ((sel >> 31u) & 1u) != 0u); // 61
+ r4 = r4 + r6 + select(0x89841d87u, 0x1e07c3d9u, ((sel >> 27u) & 1u) != 0u); // 62
+ r5 = rotl_imm(r5, 19u); // 63
+ }
+ uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
+ uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
+ out[gid] = ((ulong)hi << 32) | (ulong)lo;
+}
diff --git a/proto-cuda/packs-ca2-mixer/mx8-genesis/program_bound.metal b/proto-cuda/packs-ca2-mixer/mx8-genesis/program_bound.metal
new file mode 100644
index 000000000..8fafce718
--- /dev/null
+++ b/proto-cuda/packs-ca2-mixer/mx8-genesis/program_bound.metal
@@ -0,0 +1,111 @@
+#include
+using namespace metal;
+
+#define MASK 0x0fffffffu
+constant uint SEEDW[8] = { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u };
+
+inline uint splitmix32(uint x) {
+ x ^= x >> 16; x *= 0x7feb352du;
+ x ^= x >> 15; x *= 0x846ca68bu;
+ x ^= x >> 16;
+ return x;
+}
+inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31
+inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
+inline uint ds_elem(uint i, uint d0, uint d1) {
+ uint x = i ^ d0;
+ x *= 0x9E3779B1u; x ^= x >> 15;
+ x += d1;
+ x *= 0x85EBCA77u; x ^= x >> 13;
+ x *= 0xC2B2AE3Du; x ^= x >> 16;
+ return x;
+}
+
+// Header-bound variant: the init words come from buffer 3 (bind.rs), not from SEEDW.
+kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]],
+ device ulong* out [[buffer(1)]],
+ constant uint& baseNonce [[buffer(2)]],
+ constant uint* initw [[buffer(3)]],
+ uint gid [[thread_position_in_grid]]) {
+ uint nonce = baseNonce + gid;
+ uint r0, r1, r2, r3, r4, r5, r6, r7;
+ { uint x = nonce ^ initw[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ initw[1]; }
+ { uint x = nonce ^ initw[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ initw[2]; }
+ { uint x = nonce ^ initw[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ initw[3]; }
+ { uint x = nonce ^ initw[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ initw[4]; }
+ { uint x = nonce ^ initw[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ initw[5]; }
+ { uint x = nonce ^ initw[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ initw[6]; }
+ { uint x = nonce ^ initw[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ initw[7]; }
+ { uint x = nonce ^ initw[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ initw[0]; }
+
+ for (uint it = 0u; it < 8u; ++it) {
+ uint sel = r0;
+ r2 = r3 * r4 + r2; // 0
+ r2 = r1 * r1 + r2; // 1
+ r2 = r3 * r2 + r2; // 2
+ r3 = r3 ^ r5; // 3
+ r7 = r7 ^ dataset[r2 & MASK]; // 4
+ r5 = r5 ^ dataset[r7 & MASK]; // 5
+ r1 = r1 ^ simd_shuffle_xor(r4, (ushort)8); // 6
+ r7 = r7 ^ simd_shuffle_xor(r3, (ushort)8); // 7
+ r1 = mulhi(r1, r5); // 8
+ r6 = rotr_var(r6, r3); // 9
+ r3 = r3 | r4; // 10
+ r4 = r4 ^ dataset[r3 & MASK]; // 11
+ r0 = mulhi(r0, r4); // 12
+ r5 = r5 + r1 + select(0xc7934706u, 0xd3177981u, ((sel >> 30u) & 1u) != 0u); // 13
+ r0 = r0 ^ dataset[r4 & MASK]; // 14
+ r2 = r2 - r4; // 15
+ r2 = r2 ^ dataset[r0 & MASK]; // 16
+ r7 = r7 ^ dataset[r2 & MASK]; // 17
+ r7 = r7 ^ simd_shuffle_xor(r3, (ushort)4); // 18
+ r5 = r5 * r0; // 19
+ r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 20
+ r2 = r2 ^ simd_shuffle_xor(r4, (ushort)16); // 21
+ r6 = mulhi(r6, r2); // 22
+ r6 = r6 ^ dataset[r1 & MASK]; // 23
+ r5 = r5 * r0; // 24
+ r5 = rotl_imm(r5, 19u); // 25
+ r7 = r7 ^ simd_shuffle_xor(r6, (ushort)2); // 26
+ r0 = r0 ^ r5; // 27
+ r0 = r0 ^ r4; // 28
+ r3 = r3 - r0; // 29
+ r5 = r5 * r1; // 30
+ r7 = r7 ^ dataset[r2 & MASK]; // 31
+ r1 = r1 ^ dataset[r0 & MASK]; // 32
+ r5 = r5 ^ r6; // 33
+ r5 = r5 ^ dataset[r1 & MASK]; // 34
+ r0 = mulhi(r0, r5); // 35
+ r5 = r5 ^ simd_shuffle_xor(r2, (ushort)4); // 36
+ r7 = r7 ^ dataset[r0 & MASK]; // 37
+ r3 = r3 + r1 + select(0x75ba2fadu, 0x230c005cu, ((sel >> 27u) & 1u) != 0u); // 38
+ r1 = r1 ^ simd_shuffle_xor(r5, (ushort)4); // 39
+ r2 = r2 ^ r5; // 40
+ r3 = r6 * r3 + r3; // 41
+ r6 = r6 - r7; // 42
+ r7 = r7 ^ r0; // 43
+ r1 = r1 ^ dataset[r7 & MASK]; // 44
+ r2 = r2 * r3; // 45
+ r1 = mulhi(r1, r5); // 46
+ r4 = r4 - r3; // 47
+ r2 = rotr_var(r2, r6); // 48
+ r3 = r3 ^ dataset[r5 & MASK]; // 49
+ r1 = r1 + r5 + select(0x81b8bc2cu, 0x1907970cu, ((sel >> 7u) & 1u) != 0u); // 50
+ r0 = r0 * r2; // 51
+ r0 = r0 + r2 + select(0x4f92b968u, 0x699fd448u, ((sel >> 6u) & 1u) != 0u); // 52
+ r1 = r1 + r0 + select(0x2bb965afu, 0x77b1520du, ((sel >> 12u) & 1u) != 0u); // 53
+ r7 = rotl_imm(r7, 14u); // 54
+ r3 = r3 + r7 + select(0x7b0fe07au, 0xa54c55a0u, ((sel >> 1u) & 1u) != 0u); // 55
+ r6 = r6 ^ dataset[r7 & MASK]; // 56
+ r1 = rotr_var(r1, r5); // 57
+ r5 = r5 ^ dataset[r4 & MASK]; // 58
+ r6 = r6 ^ dataset[r2 & MASK]; // 59
+ r3 = r5 * r0 + r3; // 60
+ r5 = r5 + r7 + select(0xaf9dd72du, 0xad7493e7u, ((sel >> 31u) & 1u) != 0u); // 61
+ r4 = r4 + r6 + select(0x89841d87u, 0x1e07c3d9u, ((sel >> 27u) & 1u) != 0u); // 62
+ r5 = rotl_imm(r5, 19u); // 63
+ }
+ uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
+ uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
+ out[gid] = ((ulong)hi << 32) | (ulong)lo;
+}
diff --git a/proto-cuda/packs-ca2-mixer/mx8-genesis/vectors.h b/proto-cuda/packs-ca2-mixer/mx8-genesis/vectors.h
new file mode 100644
index 000000000..141affeeb
--- /dev/null
+++ b/proto-cuda/packs-ca2-mixer/mx8-genesis/vectors.h
@@ -0,0 +1,57 @@
+// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
+// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset
+#pragma once
+#ifdef __cplusplus
+#include
+#else
+#include
+#endif
+
+#define IGNEUM_VEC_WARPS 3
+static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u };
+static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = {
+ { // base nonce 0
+ 0x19b56348bc85304dull, 0xb08a9cfb44aa720full, 0xe1f8f627f780eff7ull, 0x2ff3e86ff1696161ull, 0xb65e0578c257e8acull, 0xf37b5705c2bebaaeull, 0xb19fef670982389eull, 0x2374331b28827a11ull,
+ 0xf492cdd05dda7f88ull, 0x37700f1385e19885ull, 0xa27524f3010b3e87ull, 0x6c8a24de8b938c43ull, 0x3dc017c8820cfd16ull, 0x5c80d37146657b68ull, 0x214082cac03f734aull, 0x1d67c665145d72f3ull,
+ 0x9fb0ae736eb87ff5ull, 0xcd7d416ae3a6be61ull, 0x0069cda85c6de58full, 0x0b7d82e717224baaull, 0x600960f6d7be30f0ull, 0xeb29032663b0c4d2ull, 0xf29583a5766408f1ull, 0x6c9532a99ad7314dull,
+ 0x9a949fd0959cc80full, 0xa20258520e7f6c25ull, 0xf2ab11f9bb032e38ull, 0xcc967bcd0c8d07c1ull, 0x37745267bb3231f2ull, 0x35a046048c2b69b3ull, 0xaa51834cd3f364f3ull, 0x359192708e4f754aull
+ },
+ { // base nonce 4096
+ 0x62fb132a9943127aull, 0x0b703e577e7f4ecaull, 0xf9f24f5522ce7593ull, 0x3cf5c516abc4332aull, 0xd25523f5f6d7a127ull, 0xd2081a002f983682ull, 0xbf46c54e9b3c4254ull, 0xca362e291e5e5f4dull,
+ 0x6039712f10f457a3ull, 0x8a34b7cabf97c23bull, 0xa473c6a2e0bf59bcull, 0x6cf3926513a4b069ull, 0x297ec2998376a40dull, 0x8efd7f601a8f28dbull, 0x8e72532dfdc1e544ull, 0x917c2b2ebe2a7e00ull,
+ 0x923fbb2d2f635c25ull, 0xce864ea5c0dedad9ull, 0x4b8ec7e874e446efull, 0x1b69b69465449196ull, 0x5ef3a8a6edb369cfull, 0x06c263ef9ce63fc4ull, 0x9c2048fd9d9e2639ull, 0x457fdd96ca4a138eull,
+ 0xd1904018b8d7b6e3ull, 0x8682312fb2e96ab8ull, 0xdc3257e0d0f979a5ull, 0xa51b0a8519d87db5ull, 0x334f08ec056e618bull, 0x3464ce71dc65119dull, 0x6a4d6df066332e04ull, 0x7d7866cb9cfca8ffull
+ },
+ { // base nonce 1000000
+ 0x86b6cb0e13d89b03ull, 0x96299a3f19d7ef15ull, 0x67d2c55100d2f876ull, 0x0a4dfe97d671b728ull, 0x41e4489014d42595ull, 0xf11cb1958c0c0e82ull, 0xf8b70b0c0a03175full, 0x632299df87d5063eull,
+ 0xe198417776130492ull, 0x8ffc5449290d7be2ull, 0x5f2e264eb1311f1bull, 0x988376463ac88586ull, 0x83969eadda489c26ull, 0xbed0a2c3f255d306ull, 0x1a949d271961a819ull, 0x5bce06eb6984725cull,
+ 0x94d5d6a1b0ffd4e9ull, 0xf3c78bae6c2182b4ull, 0xb97e9fe1bbfcdd55ull, 0x70262d1d4c0eccb2ull, 0x1fc93b427dba28d9ull, 0x02b2e3c4317f2a2dull, 0x54d3d42a588edcb9ull, 0x79998677846e7cceull,
+ 0x486522a5425f821aull, 0x95fa88e933360e52ull, 0xc8bae2da2b883f6cull, 0xbe3eb610ad33614full, 0x20efb3c4de82907full, 0xd6b650cfedfb26b7ull, 0x8c24447a646dba26ull, 0x9c004678515e44ecull
+ }
+};
+
+// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455).
+static const uint32_t IGNEUM_DS_HEAD[16] = {
+ 0xfdad4319u, 0x1a7b68e1u, 0xde6db608u, 0x13d73892u, 0xd17f447au, 0xb2221ccfu, 0x9db004bdu, 0x57d7d367u,
+ 0xdbc4cf34u, 0x697c009au, 0xc43af1d4u, 0x97f12b2eu, 0x74c37cd0u, 0xc651ea15u, 0x665a6d29u, 0x22330a2du
+};
+static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u;
+static const uint32_t IGNEUM_DS_LAST = 0xa83e7aa6u;
+// 64 sampled dataset words (index, value) computed on the Mac.
+#define IGNEUM_DS_SAMPLES 64
+static const uint32_t IGNEUM_DS_SAMPLE_INDEX[IGNEUM_DS_SAMPLES] = {
+ 59471966u, 217795994u, 208353206u, 42483309u, 172547758u, 148076330u, 183853158u, 214389424u, 267488061u, 169781097u, 184093494u, 153880993u, 84977930u, 46426879u, 3093825u, 225364072u, 44593546u, 260713159u, 168250303u, 52384140u, 223401610u, 45554030u, 95410555u, 175039924u, 79171087u, 267580473u, 24168642u, 37981670u, 171551130u, 195559979u, 204611762u, 140997658u, 138925853u, 86637313u, 20736778u, 219665210u, 160430336u, 264654675u, 8013395u, 228945585u, 213884386u, 104419827u, 44185464u, 142737231u, 99284897u, 132475900u, 61861762u, 132056166u, 262388043u, 91878046u, 117353561u, 124768597u, 71352993u, 190698941u, 46055428u, 55281366u, 165145231u, 106810753u, 171985651u, 232085256u, 159510492u, 40072060u, 209107596u, 39023794u
+};
+static const uint32_t IGNEUM_DS_SAMPLE_VALUE[IGNEUM_DS_SAMPLES] = {
+ 0x3230bc7bu, 0x7fbfe2c9u, 0xb2690991u, 0x1745c7c5u, 0x0ab0ccafu, 0x1bf87d6bu, 0x160139fdu, 0x719817acu, 0x0155df4bu, 0xbe1e86c3u, 0x680bcd6cu, 0x79c3dc6cu, 0x181e7e5fu, 0x0713a109u, 0xc705dd9fu, 0x3933b7a8u, 0xdd1c0431u, 0x50522b30u, 0xa0020b38u, 0xbff39e96u, 0x21b67e18u, 0x740f8db3u, 0x2baba568u, 0x2c9bef83u, 0x0ad9b671u, 0xc4327869u, 0x7b4fd7d0u, 0x2c29965fu, 0xec56f15fu, 0x61111746u, 0x303a1d6eu, 0xbddcfd1au, 0xf829a355u, 0x6d5df2a9u, 0x01ab8e44u, 0x06d13507u, 0xda8dcfc6u, 0x01a703e1u, 0xafe7d2c1u, 0xc091c3a2u, 0xac1814feu, 0x6e6ff62au, 0x8fdf01bau, 0xdd3f7159u, 0xdfa0d75cu, 0x26684c35u, 0x7f441e63u, 0x88df2570u, 0x8aa4d5ebu, 0xcc816c05u, 0x434df890u, 0xcd392ad6u, 0x1ab4cb63u, 0x595926fau, 0x7cd76b41u, 0x20cb95c4u, 0x13cf823fu, 0xf9daf901u, 0xff9af40au, 0x2c7dfa51u, 0x871206dbu, 0x938c116cu, 0xb64bf199u, 0x5751f874u
+};
+// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^26 words.
+static const uint32_t IGNEUM_CACHE_HEAD[16] = {
+ 0x355a86d2u, 0x7957db1cu, 0xd21772afu, 0x6fc1e09bu, 0xd55ce61du, 0x6e6a278bu, 0xd3f543ceu, 0x223d8e82u,
+ 0x143ab337u, 0x2e9f05bdu, 0x2eb389bfu, 0x0c6e449eu, 0x5cfa4222u, 0xba6560feu, 0x8e3e1aa4u, 0xdbcc1d53u
+};
+static const uint32_t IGNEUM_CACHE_LAST[16] = {
+ 0x41190d91u, 0xbd277957u, 0x22ddbb49u, 0x6986f207u, 0xdf69a4d6u, 0x26401a3au, 0x818230fbu, 0xc417122du,
+ 0x3597b211u, 0xb553ce55u, 0xcf39cc0du, 0x3b7fc43au, 0x3fd43b00u, 0x67e1c80eu, 0xffa7ea7du, 0xca2960abu
+};
+static const uint64_t IGNEUM_CACHE_FNV64 = 0x48c4f5bf24166b2eull;
diff --git a/proto-cuda/packs-ca2-mixer/mx8-genesis/vectors.json b/proto-cuda/packs-ca2-mixer/mx8-genesis/vectors.json
new file mode 100644
index 000000000..916b155f2
--- /dev/null
+++ b/proto-cuda/packs-ca2-mixer/mx8-genesis/vectors.json
@@ -0,0 +1,36 @@
+{
+ "seed": "igneum-genesis",
+ "day": "2026-10-03",
+ "dataset_mode": "memory-hard",
+ "dataset_log2_words": 28,
+ "mask": "0x0fffffff",
+ "lanes": 32,
+ "source": "igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset",
+ "warps": [
+ {"base_nonce": 0, "expected": [
+ "0x19b56348bc85304d", "0xb08a9cfb44aa720f", "0xe1f8f627f780eff7", "0x2ff3e86ff1696161", "0xb65e0578c257e8ac", "0xf37b5705c2bebaae", "0xb19fef670982389e", "0x2374331b28827a11",
+ "0xf492cdd05dda7f88", "0x37700f1385e19885", "0xa27524f3010b3e87", "0x6c8a24de8b938c43", "0x3dc017c8820cfd16", "0x5c80d37146657b68", "0x214082cac03f734a", "0x1d67c665145d72f3",
+ "0x9fb0ae736eb87ff5", "0xcd7d416ae3a6be61", "0x0069cda85c6de58f", "0x0b7d82e717224baa", "0x600960f6d7be30f0", "0xeb29032663b0c4d2", "0xf29583a5766408f1", "0x6c9532a99ad7314d",
+ "0x9a949fd0959cc80f", "0xa20258520e7f6c25", "0xf2ab11f9bb032e38", "0xcc967bcd0c8d07c1", "0x37745267bb3231f2", "0x35a046048c2b69b3", "0xaa51834cd3f364f3", "0x359192708e4f754a"
+ ]},
+ {"base_nonce": 4096, "expected": [
+ "0x62fb132a9943127a", "0x0b703e577e7f4eca", "0xf9f24f5522ce7593", "0x3cf5c516abc4332a", "0xd25523f5f6d7a127", "0xd2081a002f983682", "0xbf46c54e9b3c4254", "0xca362e291e5e5f4d",
+ "0x6039712f10f457a3", "0x8a34b7cabf97c23b", "0xa473c6a2e0bf59bc", "0x6cf3926513a4b069", "0x297ec2998376a40d", "0x8efd7f601a8f28db", "0x8e72532dfdc1e544", "0x917c2b2ebe2a7e00",
+ "0x923fbb2d2f635c25", "0xce864ea5c0dedad9", "0x4b8ec7e874e446ef", "0x1b69b69465449196", "0x5ef3a8a6edb369cf", "0x06c263ef9ce63fc4", "0x9c2048fd9d9e2639", "0x457fdd96ca4a138e",
+ "0xd1904018b8d7b6e3", "0x8682312fb2e96ab8", "0xdc3257e0d0f979a5", "0xa51b0a8519d87db5", "0x334f08ec056e618b", "0x3464ce71dc65119d", "0x6a4d6df066332e04", "0x7d7866cb9cfca8ff"
+ ]},
+ {"base_nonce": 1000000, "expected": [
+ "0x86b6cb0e13d89b03", "0x96299a3f19d7ef15", "0x67d2c55100d2f876", "0x0a4dfe97d671b728", "0x41e4489014d42595", "0xf11cb1958c0c0e82", "0xf8b70b0c0a03175f", "0x632299df87d5063e",
+ "0xe198417776130492", "0x8ffc5449290d7be2", "0x5f2e264eb1311f1b", "0x988376463ac88586", "0x83969eadda489c26", "0xbed0a2c3f255d306", "0x1a949d271961a819", "0x5bce06eb6984725c",
+ "0x94d5d6a1b0ffd4e9", "0xf3c78bae6c2182b4", "0xb97e9fe1bbfcdd55", "0x70262d1d4c0eccb2", "0x1fc93b427dba28d9", "0x02b2e3c4317f2a2d", "0x54d3d42a588edcb9", "0x79998677846e7cce",
+ "0x486522a5425f821a", "0x95fa88e933360e52", "0xc8bae2da2b883f6c", "0xbe3eb610ad33614f", "0x20efb3c4de82907f", "0xd6b650cfedfb26b7", "0x8c24447a646dba26", "0x9c004678515e44ec"
+ ]}
+ ],
+ "dataset_head": ["0xfdad4319", "0x1a7b68e1", "0xde6db608", "0x13d73892", "0xd17f447a", "0xb2221ccf", "0x9db004bd", "0x57d7d367", "0xdbc4cf34", "0x697c009a", "0xc43af1d4", "0x97f12b2e", "0x74c37cd0", "0xc651ea15", "0x665a6d29", "0x22330a2d"],
+ "dataset_last_index": 268435455,
+ "dataset_last": "0xa83e7aa6",
+ "dataset_samples": [{"index": 59471966, "value": "0x3230bc7b"}, {"index": 217795994, "value": "0x7fbfe2c9"}, {"index": 208353206, "value": "0xb2690991"}, {"index": 42483309, "value": "0x1745c7c5"}, {"index": 172547758, "value": "0x0ab0ccaf"}, {"index": 148076330, "value": "0x1bf87d6b"}, {"index": 183853158, "value": "0x160139fd"}, {"index": 214389424, "value": "0x719817ac"}, {"index": 267488061, "value": "0x0155df4b"}, {"index": 169781097, "value": "0xbe1e86c3"}, {"index": 184093494, "value": "0x680bcd6c"}, {"index": 153880993, "value": "0x79c3dc6c"}, {"index": 84977930, "value": "0x181e7e5f"}, {"index": 46426879, "value": "0x0713a109"}, {"index": 3093825, "value": "0xc705dd9f"}, {"index": 225364072, "value": "0x3933b7a8"}, {"index": 44593546, "value": "0xdd1c0431"}, {"index": 260713159, "value": "0x50522b30"}, {"index": 168250303, "value": "0xa0020b38"}, {"index": 52384140, "value": "0xbff39e96"}, {"index": 223401610, "value": "0x21b67e18"}, {"index": 45554030, "value": "0x740f8db3"}, {"index": 95410555, "value": "0x2baba568"}, {"index": 175039924, "value": "0x2c9bef83"}, {"index": 79171087, "value": "0x0ad9b671"}, {"index": 267580473, "value": "0xc4327869"}, {"index": 24168642, "value": "0x7b4fd7d0"}, {"index": 37981670, "value": "0x2c29965f"}, {"index": 171551130, "value": "0xec56f15f"}, {"index": 195559979, "value": "0x61111746"}, {"index": 204611762, "value": "0x303a1d6e"}, {"index": 140997658, "value": "0xbddcfd1a"}, {"index": 138925853, "value": "0xf829a355"}, {"index": 86637313, "value": "0x6d5df2a9"}, {"index": 20736778, "value": "0x01ab8e44"}, {"index": 219665210, "value": "0x06d13507"}, {"index": 160430336, "value": "0xda8dcfc6"}, {"index": 264654675, "value": "0x01a703e1"}, {"index": 8013395, "value": "0xafe7d2c1"}, {"index": 228945585, "value": "0xc091c3a2"}, {"index": 213884386, "value": "0xac1814fe"}, {"index": 104419827, "value": "0x6e6ff62a"}, {"index": 44185464, "value": "0x8fdf01ba"}, {"index": 142737231, "value": "0xdd3f7159"}, {"index": 99284897, "value": "0xdfa0d75c"}, {"index": 132475900, "value": "0x26684c35"}, {"index": 61861762, "value": "0x7f441e63"}, {"index": 132056166, "value": "0x88df2570"}, {"index": 262388043, "value": "0x8aa4d5eb"}, {"index": 91878046, "value": "0xcc816c05"}, {"index": 117353561, "value": "0x434df890"}, {"index": 124768597, "value": "0xcd392ad6"}, {"index": 71352993, "value": "0x1ab4cb63"}, {"index": 190698941, "value": "0x595926fa"}, {"index": 46055428, "value": "0x7cd76b41"}, {"index": 55281366, "value": "0x20cb95c4"}, {"index": 165145231, "value": "0x13cf823f"}, {"index": 106810753, "value": "0xf9daf901"}, {"index": 171985651, "value": "0xff9af40a"}, {"index": 232085256, "value": "0x2c7dfa51"}, {"index": 159510492, "value": "0x871206db"}, {"index": 40072060, "value": "0x938c116c"}, {"index": 209107596, "value": "0xb64bf199"}, {"index": 39023794, "value": "0x5751f874"}],
+ "cache_head": ["0x355a86d2", "0x7957db1c", "0xd21772af", "0x6fc1e09b", "0xd55ce61d", "0x6e6a278b", "0xd3f543ce", "0x223d8e82", "0x143ab337", "0x2e9f05bd", "0x2eb389bf", "0x0c6e449e", "0x5cfa4222", "0xba6560fe", "0x8e3e1aa4", "0xdbcc1d53"],
+ "cache_last_line": ["0x41190d91", "0xbd277957", "0x22ddbb49", "0x6986f207", "0xdf69a4d6", "0x26401a3a", "0x818230fb", "0xc417122d", "0x3597b211", "0xb553ce55", "0xcf39cc0d", "0x3b7fc43a", "0x3fd43b00", "0x67e1c80e", "0xffa7ea7d", "0xca2960ab"],
+ "cache_fnv1a64": "0x48c4f5bf24166b2e"
+}
diff --git a/proto-metal/packbench.swift b/proto-metal/packbench.swift
index a38012db8..a378e736a 100644
--- a/proto-metal/packbench.swift
+++ b/proto-metal/packbench.swift
@@ -5,7 +5,7 @@
// against the Rust CPU reference and timed without a Swift mirror of the generator. One file, no packages.
//
// swiftc -O -target arm64-apple-macos11 -o packbench packbench.swift -framework Metal
-// ./packbench --pack [--batches 5] [--batch-log2 24] [--group 256] [--warps 2048]
+// ./packbench --pack [--batches 5] [--batch-log2 24] [--group 256] [--warps 2048] [--batch-base 0]
//
// Prints one RESULT line per run: vectors, cache and dataset checks, the batch fingerprint (FNV-1a 64 over the 2^B
// outputs at base nonce 0) and MH/s by wall and by GPU time. Variant 5 packs (IGNEUM_PERSISTENT_WARPS) are launched
@@ -16,7 +16,7 @@ import Metal
func nowMs() -> Double { return Double(DispatchTime.now().uptimeNanoseconds) / 1e6 }
func fail(_ m: String) -> Never { print("FAIL: \(m)"); exit(1) }
-struct Opts { var pack = ""; var batches = 5; var batchLog2 = 24; var group = 256; var warps = 2048 }
+struct Opts { var pack = ""; var batches = 5; var batchLog2 = 24; var group = 256; var warps = 2048; var batchBase: UInt32 = 0 }
var opts = Opts()
var args = Array(CommandLine.arguments.dropFirst())
while !args.isEmpty {
@@ -28,6 +28,7 @@ while !args.isEmpty {
case "--batch-log2": opts.batchLog2 = Int(next())!
case "--group": opts.group = Int(next())!
case "--warps": opts.warps = Int(next())!
+ case "--batch-base": opts.batchBase = UInt32(next())! // base nonce of the fingerprint batch (default 0; a base near 2^32 makes the persistent unit sequence wrap inside the launch)
default: fail("unknown argument \(a)")
}
}
@@ -216,12 +217,13 @@ for (i, base) in vecBases.enumerated() {
var nonces = 1 << opts.batchLog2
if persistent { let unit = 32 * min(warpsN, nonces / 32); nonces = (nonces / unit) * unit }
let out = device.makeBuffer(length: nonces * 8, options: .storageModeShared)!
-let (warmWall, warmGpu) = run { enc in encodeHash(enc, out: out, base: 0, nonces: nonces, group: opts.group) }
+let (warmWall, warmGpu) = run { enc in encodeHash(enc, out: out, base: opts.batchBase, nonces: nonces, group: opts.group) }
let outPtr = out.contents().bindMemory(to: UInt64.self, capacity: nonces)
var batchVecPass = 0, batchVecN = 0
-for (i, base) in vecBases.enumerated() where Int(base) + 32 <= nonces {
+for (i, base) in vecBases.enumerated() where Int(base &- opts.batchBase) + 32 <= nonces { // the batch window, wrapping past 2^32
+ let off = Int(base &- opts.batchBase)
batchVecN += 1
- if (0..<32).allSatisfy({ outPtr[Int(base) + $0] == vecOuts[i][$0] }) { batchVecPass += 1 }
+ if (0..<32).allSatisfy({ outPtr[off + $0] == vecOuts[i][$0] }) { batchVecPass += 1 }
}
let fingerprint = fnv1a64(out.contents(), nonces * 8)
// Timed batches
@@ -239,5 +241,5 @@ print("cache FNV-1a 64 \(String(format: "%016llx", cacheFnv)) \(cacheOk ? "PASS"
if hotMb > 0 { print("hot table \(hotMb) MiB (\(hotSlots) of 16 load slots): fill \(String(format: "%.2f", hotGpu)) ms GPU (\(String(format: "%.2f", hotWall)) wall), FNV-1a 64 \(String(format: "%016llx", hotFnv)) \(hotFnvWant == nil ? "(not in vectors.json)" : (hotOk ? "PASS" : "FAIL"))") }
print("warm-up batch \(nonces) hashes: \(String(format: "%.1f", warmGpu)) ms GPU, \(String(format: "%.1f", warmWall)) ms wall")
let overall = cacheOk && dsOk && hotOk && vecPass == vecBases.count && batchVecPass == batchVecN
-print("RESULT pack=\(packName) class=\(className) device=\(device.name.replacingOccurrences(of: " ", with: "_")) group=\(opts.group) warps=\(warpsN) arena_mib=\(persistent ? warpsN * 32 * scratchWordsPerLane * 4 / 1048576 : 0) hot_mib=\(hotMb) hot_slots=\(hotSlots) hot_fill_ms=\(String(format: "%.2f", hotGpu)) hot=\(hotMb > 0 ? (hotOk ? "PASS" : "FAIL") : "none") nonces=\(nonces) batches=\(opts.batches) vectors=\(vecPass)/\(vecBases.count) batch_vectors=\(batchVecPass)/\(batchVecN) cache=\(cacheOk ? "PASS" : "FAIL") dataset=\(dsOk ? "PASS" : "FAIL") fingerprint=\(String(format: "%016llx", fingerprint)) mhs_gpu=\(String(format: "%.3f", mhsGpu)) mhs_wall=\(String(format: "%.3f", mhsWall)) loads=\(loadsPerHash) bytes=\(bytesPerHash) scratch_ops=\(scratchOps * 8) overall=\(overall ? "PASS" : "FAIL")")
+print("RESULT pack=\(packName) class=\(className) device=\(device.name.replacingOccurrences(of: " ", with: "_")) group=\(opts.group) warps=\(warpsN) batch_base=\(opts.batchBase) arena_mib=\(persistent ? warpsN * 32 * scratchWordsPerLane * 4 / 1048576 : 0) hot_mib=\(hotMb) hot_slots=\(hotSlots) hot_fill_ms=\(String(format: "%.2f", hotGpu)) hot=\(hotMb > 0 ? (hotOk ? "PASS" : "FAIL") : "none") nonces=\(nonces) batches=\(opts.batches) vectors=\(vecPass)/\(vecBases.count) batch_vectors=\(batchVecPass)/\(batchVecN) cache=\(cacheOk ? "PASS" : "FAIL") dataset=\(dsOk ? "PASS" : "FAIL") fingerprint=\(String(format: "%016llx", fingerprint)) mhs_gpu=\(String(format: "%.3f", mhsGpu)) mhs_wall=\(String(format: "%.3f", mhsWall)) loads=\(loadsPerHash) bytes=\(bytesPerHash) scratch_ops=\(scratchOps * 8) overall=\(overall ? "PASS" : "FAIL")")
exit(overall ? 0 : 1)
diff --git a/relay/playbooks/mixer-x4-5090-bench.ps1 b/relay/playbooks/mixer-x4-5090-bench.ps1
index 74f4a184a..c0aaa4f40 100644
--- a/relay/playbooks/mixer-x4-5090-bench.ps1
+++ b/relay/playbooks/mixer-x4-5090-bench.ps1
@@ -4,7 +4,7 @@
# igneum-worker-cuda.exe on the v2 pack and the two v3 packs of the fetched packs folder (the dataset build time per pack is the
# number this job is for: the worker's own `cache ... dataset ... ms` line, the same line the readwidth round printed),
# and switches the card back on with the settings it had. Every result line starts with RESULT so `node tools/jobs.mjs ` shows them.
-# The packs folder of the fetch job: proto-cuda/packs-ca2-mixer/mx4-genesis and mx4-devnet-epoch0 from branch ca2-mixer, plus
+# The packs folder of the fetch job: proto-cuda/packs-ca2-mixer/mx4-genesis, mx4-devnet-epoch0, mx8-genesis and mx8-devnet-epoch0 (the x8 candidate) from branch ca2-mixer, plus
# proto-cuda/packs/igneum-genesis-mh copied in as v2-genesis-mh (the version 2 control, same card, same run).
$ErrorActionPreference = 'Continue'
function Say([string] $m) { Write-Host ("[" + (Get-Date -Format 'HH:mm:ss') + "] " + $m) }
@@ -48,7 +48,7 @@ if ($card) {
& nvidia-smi --query-gpu=name,driver_version,power.limit,clocks.sm,clocks.mem,memory.used,temperature.gpu --format=csv,noheader 2>&1 | ForEach-Object { "RESULT gpu-before $_" }
# the probe ran in run-readwidth-5090-20261005 (bench-log, 5 October 2026); this second run is the bench only
-foreach ($pk in @('v2-genesis-mh', 'mx4-genesis', 'mx4-devnet-epoch0')) {
+foreach ($pk in @('v2-genesis-mh', 'mx4-genesis', 'mx4-devnet-epoch0', 'mx8-genesis', 'mx8-devnet-epoch0')) {
$d = Join-Path $packs $pk
Write-Output "RESULT bench $pk start $(Get-Date -Format HH:mm:ss)"
& $exe --bench --pack $d --batches 5 --batch-log2 24 --block-warps 1 2>&1 | ForEach-Object { "RESULT $_" }
diff --git a/relay/playbooks/mixer-x4-9070-bench.ps1 b/relay/playbooks/mixer-x4-9070-bench.ps1
index 50df5b654..4d808e14a 100644
--- a/relay/playbooks/mixer-x4-9070-bench.ps1
+++ b/relay/playbooks/mixer-x4-9070-bench.ps1
@@ -4,7 +4,7 @@
# igneum-worker-opencl.exe on the v2 pack and the two v3 packs of the fetched packs folder (the dataset build time per pack is the
# number this job is for: the worker's own `cache ... dataset ... ms` line, the same line the readwidth round printed),
# and switches the card back on with the settings it had. Every result line starts with RESULT so `node tools/jobs.mjs ` shows them.
-# The packs folder of the fetch job: proto-cuda/packs-ca2-mixer/mx4-genesis and mx4-devnet-epoch0 from branch ca2-mixer, plus
+# The packs folder of the fetch job: proto-cuda/packs-ca2-mixer/mx4-genesis, mx4-devnet-epoch0, mx8-genesis and mx8-devnet-epoch0 (the x8 candidate) from branch ca2-mixer, plus
# proto-cuda/packs/igneum-genesis-mh copied in as v2-genesis-mh (the version 2 control, same card, same run).
$ErrorActionPreference = 'Continue'
function Say([string] $m) { Write-Host ("[" + (Get-Date -Format 'HH:mm:ss') + "] " + $m) }
@@ -20,7 +20,9 @@ $list = & $exe --list 2>&1
$list | ForEach-Object { "RESULT list $_" }
$dev = $null
foreach ($l in $list) { if ($l -match '^\s*\[(\d+)\].*gfx1201' -and $l -notmatch 'dup') { $dev = [int]$Matches[1]; break } }
-if ($null -eq $dev) { Write-Output 'RESULT error no gfx1201 device in --list'; exit 2 }
+# the 9070 XT when it is on the bus, else the integrated gfx1036 (the coordinator's rule for the AMD build-time row)
+if ($null -eq $dev) { foreach ($l in $list) { if ($l -match '^\s*\[(\d+)\].*gfx1036' -and $l -notmatch 'dup') { $dev = [int]$Matches[1]; Write-Output 'RESULT device-fallback gfx1036 (no gfx1201 on the bus)'; break } } }
+if ($null -eq $dev) { Write-Output 'RESULT error no gfx1201 or gfx1036 device in --list'; exit 2 }
Write-Output "RESULT device $dev"
# the app: switch off the NVIDIA card only, remember its settings
@@ -51,7 +53,7 @@ if ($card) {
} else { Write-Output 'RESULT card none-found (the app is not running or has no NVIDIA card); measuring with whatever else runs on the GPU' }
# the probe ran in run-readwidth-9070-20261005 (bench-log, 5 October 2026); this second run is the bench only
-foreach ($pk in @('v2-genesis-mh', 'mx4-genesis', 'mx4-devnet-epoch0')) {
+foreach ($pk in @('v2-genesis-mh', 'mx4-genesis', 'mx4-devnet-epoch0', 'mx8-genesis', 'mx8-devnet-epoch0')) {
$d = Join-Path $packs $pk
Write-Output "RESULT bench $pk start $(Get-Date -Format HH:mm:ss)"
& $exe --bench-pack --pack $d --batches 5 --batch-log2 24 --device $dev 2>&1 | ForEach-Object { "RESULT $_" }
diff --git a/relay/playbooks/mixer-x4-pc1-bench.ps1 b/relay/playbooks/mixer-x4-pc1-bench.ps1
new file mode 100644
index 000000000..6a92de28b
--- /dev/null
+++ b/relay/playbooks/mixer-x4-pc1-bench.ps1
@@ -0,0 +1,95 @@
+# Igneum run job: the class v3 dataset construction, mixer x4 and the x8 candidate (docs/plans/mixer-x4.md), on PC 1
+# (machine ae432dc7): the RTX 5090 through igneum-worker-cuda.exe --bench, then the AMD card through
+# igneum-worker-opencl.exe --bench-pack, each card switched off in the app ONLY while it is under test and restored
+# after with the settings it had (the ca2-era-pc1.ps1 shape). 5 October 2026. The AMD card: the RX 9070 XT (gfx1201)
+# when its eGPU box is on the bus, else the integrated gfx1036. The number this job is for: the worker's own
+# `cache ... dataset ... ms` line per pack, the daily 1 GiB dataset build time at x1, x4 and x8 (the coordinator's rule:
+# x8 enters v3 only if that build stays under 1 s on every discrete card we own), plus the vectors and the 2^24
+# fingerprint per pack (bit-exactness against the Mac: mx4-genesis 6f48d5a2aa0dbe5f, mx4-devnet-epoch0 73caaebb28e808fe,
+# mx8-genesis 7c28cfb06c5c65a9, mx8-devnet-epoch0 bbb183f72692f840, v2-genesis-mh 25f96e7dce90bd4e).
+# Published as a plain `run` job (NOT --stop-miners): the installed app keeps every other card mining. Packs:
+# proto-cuda/packs-ca2-mixer/{mx4-genesis, mx4-devnet-epoch0, mx8-genesis, mx8-devnet-epoch0} from branch ca2-mixer and
+# proto-cuda/packs/igneum-genesis-mh as v2-genesis-mh (the control). Every result line starts with RESULT.
+$ErrorActionPreference = 'Continue'
+function Say([string] $m) { Write-Host ("[" + (Get-Date -Format 'HH:mm:ss') + "] " + $m) }
+$jobs = Split-Path $env:IGNEUM_JOB_DIR
+$fetched = Join-Path $jobs 'fetch-mixer-x4-20261005'
+$cuda = Join-Path $fetched 'igneum-worker-cuda.exe'
+$ocl = Join-Path $fetched 'igneum-worker-opencl.exe'
+$packs = Join-Path $fetched 'packs-ca2-mixer'
+$packList = @('v2-genesis-mh', 'mx4-genesis', 'mx4-devnet-epoch0', 'mx8-genesis', 'mx8-devnet-epoch0')
+if (-not (Test-Path $cuda)) { Write-Output "RESULT error cuda worker missing at $cuda (the fetch job runs first)"; exit 2 }
+if (-not (Test-Path $ocl)) { Write-Output "RESULT error opencl worker missing at $ocl"; exit 2 }
+if (-not (Test-Path $packs)) { Write-Output "RESULT error packs missing at $packs"; exit 2 }
+foreach ($pk in $packList) { if (-not (Test-Path (Join-Path $packs $pk))) { Write-Output "RESULT error pack $pk missing"; exit 2 } }
+$inst = @("$env:LOCALAPPDATA\Programs\Igneum Miner", "$env:ProgramFiles\Igneum Miner") | Where-Object { Test-Path (Join-Path $_ 'igneum-worker-cuda.exe') } | Select-Object -First 1
+if (-not $inst) { Write-Output 'RESULT error no installed igneum-worker-cuda.exe (the NVRTC DLLs come from there)'; exit 2 }
+Get-ChildItem $inst -Filter 'nvrtc*.dll' | Copy-Item -Destination $fetched -Force
+Write-Output "RESULT worker-cuda $cuda sha256 $((Get-FileHash -Algorithm SHA256 $cuda).Hash.ToLower()) with $((Get-ChildItem $fetched -Filter 'nvrtc*.dll').Count) NVRTC DLL(s) from $inst"
+Write-Output "RESULT worker-opencl $ocl sha256 $((Get-FileHash -Algorithm SHA256 $ocl).Hash.ToLower())"
+
+# the app
+$appDir = $env:IGNEUM_APP_DIR
+if (-not $appDir) { $appDir = Join-Path $env:LOCALAPPDATA 'igneum\app' }
+$urlFile = Join-Path $appDir 'app.url'
+$url = $null
+if (Test-Path $urlFile) { $url = (Get-Content -LiteralPath $urlFile -Raw).Trim() }
+function Find-Card([string] $vendor, [string] $keyMatch) {
+ if (-not $url) { return $null }
+ try {
+ $st = Invoke-RestMethod -Uri ($url + 'api/state') -Method GET -TimeoutSec 10
+ $c = $st.mining.cards | Where-Object { $_.vendor -eq $vendor -and $_.key -match $keyMatch } | Select-Object -First 1
+ if (-not $c) { $c = $st.cards | Where-Object { $_.vendor -eq $vendor -and $_.key -match $keyMatch } | Select-Object -First 1 }
+ return $c
+ } catch { Say ("api/state: " + $_.Exception.Message); return $null }
+}
+function Card-Off($card) {
+ Write-Output ("RESULT card " + $card.key + " enabled=" + $card.enabled + " identities=" + $card.identities + " power_pct=" + $card.power_pct + " state=" + $card.state)
+ $body = @{ cards = @(@{ key = $card.key; enabled = $false; identities = [int]$card.identities; power_pct = [int]$card.power_pct }) } | ConvertTo-Json -Depth 5
+ try { Invoke-RestMethod -Uri ($url + 'api/cards') -Method POST -Body $body -ContentType 'application/json' -TimeoutSec 10 | Out-Null; Say "card off requested" } catch { Say ("api/cards off: " + $_.Exception.Message) }
+ $t = 0
+ while ($t -lt 90) {
+ Start-Sleep -Seconds 5; $t += 5
+ try { $st = Invoke-RestMethod -Uri ($url + 'api/state') -Method GET -TimeoutSec 10; $c2 = $st.mining.cards | Where-Object { $_.key -eq $card.key }; if (-not $c2) { $c2 = $st.cards | Where-Object { $_.key -eq $card.key } }; if ($c2 -and $c2.state -eq 'off' -and $c2.pid -eq 0) { break } } catch { }
+ }
+ Write-Output ("RESULT card-off " + $card.key + " after " + $t + " s")
+ Start-Sleep -Seconds 5
+}
+function Card-Restore($card) {
+ $body = @{ cards = @(@{ key = $card.key; enabled = [bool]$card.enabled; identities = [int]$card.identities; power_pct = [int]$card.power_pct }) } | ConvertTo-Json -Depth 5
+ try { Invoke-RestMethod -Uri ($url + 'api/cards') -Method POST -Body $body -ContentType 'application/json' -TimeoutSec 10 | Out-Null; Write-Output ("RESULT card restored " + $card.key + " enabled=" + $card.enabled) } catch { Write-Output ("RESULT error card restore " + $card.key + ": " + $_.Exception.Message) }
+}
+
+# ---- the RTX 5090 (CUDA, NVRTC) ----
+$nv = Find-Card 'nvidia' '.'
+if ($nv) { Card-Off $nv } else { Write-Output 'RESULT card none-found nvidia (the app is not running or has no NVIDIA card); measuring with whatever else runs on the GPU' }
+& nvidia-smi --query-gpu=name,driver_version,power.limit,clocks.sm,clocks.mem,memory.used,temperature.gpu --format=csv,noheader 2>&1 | ForEach-Object { "RESULT gpu-before $_" }
+foreach ($pk in $packList) {
+ $d = Join-Path $packs $pk
+ Write-Output "RESULT bench-5090 $pk start $(Get-Date -Format HH:mm:ss)"
+ & $cuda --bench --pack $d --batches 5 --batch-log2 24 --block-warps 1 2>&1 | ForEach-Object { "RESULT $_" }
+ & $cuda --bench --pack $d --batches 5 --batch-log2 24 --block-warps 8 2>&1 | Where-Object { $_ -match '^RESULT|error|FAIL|dataset' } | ForEach-Object { "RESULT $_" }
+}
+& nvidia-smi --query-gpu=power.draw,clocks.sm,clocks.mem,memory.used,temperature.gpu --format=csv,noheader 2>&1 | ForEach-Object { "RESULT gpu-after $_" }
+if ($nv) { Card-Restore $nv }
+
+# ---- the AMD card (OpenCL): the gfx1201 on the eGPU when present, else the integrated gfx1036 ----
+$list = & $ocl --list 2>&1
+$list | ForEach-Object { "RESULT list $_" }
+$dev = $null; $gfx = $null
+foreach ($l in $list) { if ($l -match '^\s*\[(\d+)\].*gfx1201' -and $l -notmatch 'dup') { $dev = [int]$Matches[1]; $gfx = 'gfx1201'; break } }
+if ($null -eq $dev) { foreach ($l in $list) { if ($l -match '^\s*\[(\d+)\].*gfx1036' -and $l -notmatch 'dup') { $dev = [int]$Matches[1]; $gfx = 'gfx1036'; break } } }
+if ($null -eq $dev) { Write-Output 'RESULT error no gfx1201 and no gfx1036 device in --list'; exit 2 }
+Write-Output "RESULT device $dev $gfx"
+$amd = Find-Card 'amd' $gfx
+if ($amd) { Card-Off $amd } else { Write-Output "RESULT card none-found $gfx (the app is not running or has no such card); measuring with whatever else runs on the GPU" }
+# the 9070 XT: the full shape; the gfx1036 (about 3 MH/s): the same 2^24 fingerprint range, 2 timed dispatches, one shape
+$batches = 5; if ($gfx -eq 'gfx1036') { $batches = 2 }
+foreach ($pk in $packList) {
+ $d = Join-Path $packs $pk
+ Write-Output "RESULT bench-$gfx $pk start $(Get-Date -Format HH:mm:ss)"
+ & $ocl --bench-pack --pack $d --batches $batches --batch-log2 24 --device $dev 2>&1 | ForEach-Object { "RESULT $_" }
+ if ($gfx -eq 'gfx1201') { & $ocl --bench-pack --pack $d --batches 5 --batch-log2 24 --device $dev --group-warps 8 2>&1 | Where-Object { $_ -match '^RESULT|error|FAIL|dataset' } | ForEach-Object { "RESULT $_" } }
+}
+if ($amd) { Card-Restore $amd }
+exit 0