Merge branch 'ca2-v3' into release-0.3.11

This commit is contained in:
igneum-labs 2026-10-05 22:38:59 +00:00
commit 74515d5c39
550 changed files with 80997 additions and 322 deletions

View file

@ -0,0 +1,100 @@
# The on-die-cache recompute chip against the RTX 5090, class v2 and class v3, everything combined
5 October 2026 (night), Counter ASIC 2.0, worker ca2-mixer. The model is M16's
(`docs/analysis/m16-recompute-attacker-2026-10-05.md`): the strongest chip the plan has priced holds the whole
cache in SRAM and derives every dataset item instead of reading it, so its cost per hash is item derivations,
and its rate at a 50 T op/s integer budget (an RTX 5090's, approximate) is `50 T / (ops per hash)`. Nothing here
is a measurement of a chip; every GPU figure says where it was measured. "Approximate" marks a figure from memory.
## 1. Inputs
| Input | Value | Source |
|---|---|---|
| Items per hash | 128 (one item per load, 128 loads per hash, median 128.00 distinct) | spec 01 sections 1.4.2 and 1.8.5; the 20,000-program census |
| Integer operations per mixer application | about 130 | spec 01 section 1.8.4 |
| Mixer applications per item | 9 under v2; 36 under v3 (`m = 4`, `docs/plans/mixer-x4.md`) | `memhard::Shape::mixers_per_item` |
| Integer operations per item | 1,170 (v2); 4,680 (v3) | 9 x 130; 36 x 130 |
| Integer operations per hash | 149,760 (v2, "150,000"); 599,040 (v3, "600,000") | 128 x the above |
| Chip integer budget | 50 T op/s (approximate: 21,760 ALUs at about 2.4 GHz, one 32-bit operation each per clock) | M16 section 3 |
| Fixed-function factor | 3x (approximate, from memory: 2x to 5x is the usual credit for a pipeline with no scheduling or divergence) | M16 section 3 |
| RTX 5090, version 2 programs, measured | 136.1 MH/s (readwidth, tonight, `docs/plans/read-width.md`, pack w4 on PC 2); 139.7 MH/s (M11, 4 October, `docs/bench-log.md`) | this analysis uses tonight's 136.1 as the denominator and quotes both |
| RTX 5090 at w16 (16-byte loads), measured | 139.8 MH/s | readwidth table, tonight (the width stays 4 B: w16 closes nothing) |
| Cache mirror, 256 MiB, N5 headline density | 128 mm^2, $46 per good die (64 mm^2, $21 at the bit-cell lower bound) | `docs/analysis/sram-mirror.md` revision 2, sections 4 and 5 (`ca2-analysis` e6085c6) |
| Cache mirror plus a 96 MB hot table, N5 headline | 175 mm^2, $68 | same, so a hot table costs 0.49 mm^2 and $0.23 per MB (linear, approximate) |
| 512 MiB and 1 GiB mirrors, N5 headline | 255 mm^2 and 510 mm^2; $111 to $306 | same, section 4 (the growth rule's cache at years 4 and 12, priced at today's node) |
| GPU-class die | 750 mm^2 (the equal-silicon comparison) | M16 section 3 |
| CPU verifier, one M5 Max core (loaded, load average 5.6; ratios are the measurement) | v2 1.31 to 1.36 ms per unit, x4 1.92 to 1.96 (1.45x), x8 2.79 (2.1x); worst cold 1.58 / 2.04 / 2.94 ms | `docs/plans/mixer-x4.md` section 6.4, 5 October 2026 21:40 UTC |
## 2. The rows
Chip rate = 50 T op/s / ops per hash. "Bare" = chip rate / 136.1 MH/s. "With the factor" = bare x 3. "Equal
silicon" = bare x (750 - SRAM) / 750 x 3: the SRAM takes die area the logic does not get, the M16 convention
("minus the area the SRAM takes"). SRAM in mm^2 and dollars at the N5 headline density.
| Row | Mixer | Ops per hash | Chip rate at 50 T op/s | SRAM the chip holds | mm^2 / $ (N5 headline) | Bare gain against 136.1 MH/s | With the 3x factor | Equal silicon, SRAM deducted, with the factor |
|---|---|---|---|---|---|---|---|---|
| v2 as shipped (the M16 and scratch-soundness row) | x1 | 149,760 | 334 MH/s | 256 MiB | 128 / $46 | 2.45x (2.39x against 139.7) | 7.4x | 6.1x |
| v2 at w16 (not adopted; the chip's cost is items, not bytes: unchanged) | x1 | 149,760 | 334 | 256 MiB | 128 / $46 | 2.39x against 139.8 | 7.2x | 5.9x |
| x4 (the candidate measured beside v3; not v3) | x4 | 599,040 | 83.5 MH/s | 256 MiB | 128 / $46 | 0.61x | 1.84x | 1.53x |
| MEASURED, NOT ADOPTED (layer 5 decided out of v3 on the PC rows, coordinator 21:40 UTC): v3 plus a 32 MiB hot table, added form (16 dataset loads and k hot loads): the honest card pays the hot loads, this chip pays SRAM only | x4 | 599,040 (a hot load is one SRAM read, no item) | 83.5 | 288 MiB | 144 / $53 | 0.66x at the Mac's g = 0.93 (126.6 MH/s); 0.70x at the 5090's g = 0.87 (118.4); the 9070 XT's g 0.84 | 1.98x (Mac g), 2.11x (5090 g) | 1.60x, 1.71x |
| MEASURED, NOT ADOPTED: v3 plus a 64 MiB hot table, added form | x4 | 599,040 | 83.5 | 320 MiB | 160 / $61 | 0.71x at the Mac's g = 0.87 (118.4 MH/s); 0.73x at the 5090's g = 0.84 (114.3); the 9070 XT's g 0.80 | 2.12x (Mac g), 2.19x (5090 g) | 1.67x, 1.73x |
| v3 at year 4 (cache 512 MiB under option C, dataset 4 GiB), no hot table | x4 | 599,040 | 83.5 | 512 MiB | 255 / $111 | 0.61x | 1.84x | 1.21x |
| v3 at year 12 (cache 1 GiB, dataset 8 GiB) | x4 | 599,040 | 83.5 | 1 GiB | 510 / $306 | 0.61x | 1.84x | 0.59x |
| **v3: mixer x8** (decided 22:05 UTC under the delegated rule: verify 2.1 ms per unit on one Mac core against the 10 ms gate, the daily 1 GiB build 23 to 77 ms on the 5090 and the 9070 XT) | x8 | 1,198,080 | 41.7 | 256 MiB | 128 / $46 | 0.31x | 0.92x | 0.76x |
| x8 at year 4 | x8 | 1,198,080 | 41.7 | 512 MiB | 255 / $111 | 0.31x | 0.92x | 0.61x |
The era draws of spec 1.13.1 cost the chip nothing in this model: the mixer round count is not drawn, the op
weights and fold rotations change the program, not the item derivation, so the chip's ops per hash stand. The
width rule (4-byte loads kept) changes nothing either: w16 would have moved the honest denominator by 2.7% and the
chip's cost not at all.
Arithmetic, row v3: 36 x 130 = 4,680 ops per item; x 128 = 599,040 per hash; 50 x 10^12 / 599,040 = 83.5 x 10^6
hashes per second; 83.5 / 136.1 = 0.613; x 3 = 1.84; equal silicon (750 - 128) / 750 = 0.829, x 1.84 = 1.53.
Hot table rows: 32 MiB x 0.49 mm^2 per MB = 16 mm^2, 64 MiB = 32 mm^2 (the 96 MB column of `sram-mirror.md`
scaled linearly); (750 - 144) / 750 = 0.808 and (750 - 160) / 750 = 0.787. The honest denominator in the added
form is the v2 rate times `g`, the card's measured ratio with the hot loads added: on the M5 Max tonight
`g = 0.93 / 0.87 / 0.83` at 32 / 64 / 96 MiB (the cache agent, relayed by the coordinator at 21:23 UTC;
`docs/plans/hot-table.md` carries the runs); the 5090's and the 9070 XT's `g` are the PC rows, owed, and until they
land the row carries the Mac's `g` against the 5090's rate, which is a mixed figure and is marked so. Year 4 and 12 rows: the mirror of
`sram-mirror.md` section 4 at N5 for 512 MiB and 1 GiB plus the 64 MiB table, at today's density (the node of
those years is denser by about 1.8x at year 10 on the trend the same file cites; the row is a floor on the area,
not a forecast).
## 3. The margin, plainly
The combined headline row is the mixer row alone (layer 5 is out: the added form costs the 5090 13 to 16 percent
and the 9070 XT 16 to 20 percent against the 0.97 bar, coordinator 21:40 UTC; the width stays 4 bytes; the era
draws and the cache growth cost this chip nothing at year 0), and class v3 is x8 (decided 22:05 UTC). The headline:
**the on-die-cache recompute chip at 50 T op/s reaches 41.7 MH/s against the 5090's 136.1, 0.31x bare, 0.92x with
the 3x fixed-function factor, 0.76x with the mirror's area deducted: under 1x with the factor, 0.92x, a margin of 8
percent on the factor (a 3.3x factor reads 1.0x) and of 9 percent on the budget (55 T op/s reads 1.0x).** The x4
candidate, measured beside it, read 1.84x and 1.53x. The hot-table rows above are kept as measured, not adopted:
against THIS chip an added hot table is a cost to the honest card and none to the chip, so it would have moved the
row the wrong way by the card's own `g`. The margin, plainly:
- the 3x fixed-function factor is approximate and from memory; at 3.3x the equal-budget row reads 2.0x;
- the denominator is one card's measured rate on one night (136.1 against 139.7 the night before: 2.6% apart);
- the 50 T op/s budget is approximate; a chip at 55 T op/s reads 2.0x;
- the hot table in the added form lowers the honest denominator by whatever the hot loads cost the GPU (owed from
the PC rows), which raises the chip's gain by the same share, 1.84x or more if the hot loads are free, higher if
not; the hot table's only cost to this chip is 16 to 32 mm^2 of die.
What keeps it under 1x is the mixer, and nothing else in Counter ASIC 2.0 moves this chip (the scratch at any share
gave 2.4x, `docs/analysis/scratch-soundness.md` section 3.4; the hot table taxes the DRAM-only chip, not this one;
the cache growth taxes it only in die area, which is cheap at year 0 and real at year 12). The next levers, in
order:
1. Mixer x16 (the next step of the same lever): 0.16x bare and 0.46x with the factor against 136.1; the verifier
by the measured increments (+0.63 ms at x4, +1.46 at x8 on the M5 Max core: about +3.1 ms at x16, 3.7 ms per
unit, 9 ms on a 2.5x slower laptop core, approximate) is at the edge of the 10 ms gate, so a 2019-class laptop
core measurement (O-1.14) decides it, not this model.
2. The hot table: adopted or not on the PC rows (`docs/plans/hot-table.md`); in the added form it costs the GPU
7 to 17 percent on the Mac and the chip die area only, so against this chip it is a lever in the wrong
direction and against a DRAM-only chip the first lever; if it is adopted, the mixer must carry the extra `1/g`
(x8 at g = 0.87 reads 1.06x at the equal budget, 0.84x with the SRAM deducted).
## 4. What this does not settle
The items of M16 section 5 stand: the inline kernel on NVIDIA with a 64 MiB cache inside L2 (a measured point
under the "50 T op/s" row) is a PC job not yet run; the time-memory curve (O-1.6) is not drawn; the mixer has had
no cryptanalysis, and a shortcut inside it cuts the 4,680 directly; no chip has been priced beyond its SRAM.

View file

@ -1551,6 +1551,148 @@ What is measured: one BLS12-381 aggregate signature over 16 summed G1 keys plus
Reading (the NEW finding, ledger C4). With the module off GHOSTDAG alone converges on the heavier chain and the losing side's records re-determine (F24 works when the chain moves). With the module on the overlay holds during the split (A, with 30% of the frozen table, locks nothing; B locks 7 and 8) and then fails at the heal in the shipped node: B's certificates for blocks off n0's chain are "kept pending until the chain decides (no lock at this index)", n0's chain never decides because GHOSTDAG keeps its heavier tip and nothing turns the certificate into a fork-choice constraint, and once n0's last lock (index 7, DAA 209) is one window old (DAA 329) the frozen table stops applying on A's chain ("no frozen table (no lock on this chain inside the window)"), A's two keys are 100% of A's own window (B's post-cut blocks are red there) and n0 locks 10, 11, 12 alone; B's certificates for 10 and 11 then log CONFLICTING on n0 (n0 log, 17:27:04 to 17:29:54 BST). A finality fork from a 96-s honest partition, no attacker, table intact at the heal; the 150-s run and the v2 control end the same way. The spec's fork choice ("GHOSTDAG among tips through all certified checkpoints", 3.5) is therefore implemented only for certificates over blocks already on the node's chain. Fix named in the ledger entry: verify an off-chain certificate against the table at its own block and let it constrain fork choice (a certificate-driven reorg), then re-determine. Raw: `scratchpad fud-a/c4-results-*.md`, node logs `c4-on90-tmp/`, `c4-v2-control-tmp/`.
## 5 October 2026 (night), read width of the lottery hash: 4, 16 and 64-byte loads, a per-load mix, a written scratch; three cards (gate 1 experiment, cryptographer)
Branch `readwidth` (commits 019b014, b970dda, 4badcee, a9e002c, d0018cf and the entry commit); plan and recommendation in `docs/plans/read-width.md`. Nothing here changes consensus: every class sits behind `igneum-pow --class` and the default class is generator version 2 byte for byte (`igneum-pow/tests/packs.rs` passes on the four pinned packs after every commit). Question (the project lead, after "the 9070 XT on the eGPU" above): would wider reads keep the latency-bound random-access property while closing the vendor gap. Additions from the coordinator: a per-load width drawn from an era-fixed mix, and a written per-warp scratch (measurement only, no soundness claim).
**What a class does** (`igneum-pow/src/generator.rs` `LoadClass`, `verify::fold_words`, the three emitters): a load of W words reads the W-word-aligned address `(src AND MASK) AND NOT (W - 1)` and folds every word into `dst` (`x = dst ^ w0; x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]`); W = 1 is the lottery hash exactly (`w4` = pack `bcc1248b10cc90f2`). A mix class draws W per load with one extra `below(100)` roll per instruction. A scratch class `scr<k>k<kb>` turns `k` of the 16 memory slots into read-modify-writes of a 16-byte slot of the lane's share of a `kb` KiB per-warp scratch (kernels run persistent warps, one per block or work-group; a slot reads as a seed-and-base fill until the unit writes it, behind a per-unit tag). Program ids carry the class. Dependent chain and 32-lane unit unchanged.
**Correctness**: 23 packs (`proto-cuda/packs-readwidth/`, Rust CPU reference vectors). Every pack passed its three vector units and the cache and dataset checks on Metal (M5 Max, `proto-metal/packbench`), Apple OpenCL (`--bench-pack`), the RTX 5090 (NVRTC, `igneum-worker-cuda --bench`) and, the 16 width and mix packs, the RX 9070 XT (`igneum-worker-opencl --bench-pack`); the 2^24 batch fingerprints agree across all four runtimes on every pack (for example w16 `e7c890445b47af60`, w64 `836e56e7d496e980`, mixB-2 `a18ac73098c76007`). The clang CUDA emulation (w16, w64, w64x4, mixA-0, mixB-0: 3 of 3 units standalone and 2 of 2 in batch at 2 warps per block) and the clang OpenCL emulation (the same five plus scr2k32 and scr8k128, sub-group 32 and, width packs, wave64 with sub-group shuffles) pass with equal fingerprints per configuration. Acceptance rule on the classes: 60 candidates per class, rejection 0 to 14 of 60 (w16 and w64 as v2; the mixes the same; the scratch classes' distinct-address bound now covers dataset loads only, since a 64-slot lane scratch repeats slots by design). CPU verifier (M5 Max, one core, avg of 50 units, `igneum-pow bench --class`): v2 0.604 ms, w16 0.610, w64 0.630, w64x4 0.160, mix50-35-15 0.620, mix25-50-25 0.614, scr0k32 0.600 (1.004 on a loaded re-run), scr2k32 0.657, scr4k32 0.458, scr8k32 0.317, scr2k128 0.535, scr4k128 0.458, scr8k128 0.311; per hash divide by 32. The wide reads cost the verifier nothing (a lane's words lie in one item); scratch ops replace item derivations and make it cheaper.
**Probes** (`--memprobe`, dependent random reads at 1024 MiB, G reads/s, best over lanes in flight; 4 B = the hash's pattern; the 5090 and 9070 XT with the card off in the app, the Mac through Apple OpenCL under a load average of 5 to 10):
| Card | 4 B chase | 16 B | 64 B | 64 B as GB/s | coalesced stream GB/s (rated) | integer chain |
|---|---|---|---|---|---|---|
| RTX 5090 (PC 2, CUDA) | 17.5 to 18.2 | 18.0 to 19.9 | 9.1 to 15.7 (9.1 at 4 M lanes) | 584 | 1,579 (1,792) | 39.0 T op/s |
| RX 9070 XT (PC 1, eGPU, OpenCL) | 2.42 to 2.66 | 2.43 to 2.73 | 2.47 to 2.87 | 158 | 636 (640) | 6.2 T op/s |
| Apple M5 Max (Apple OpenCL, approximate) | 3.50 | 3.51 | 3.51 | 225 | 522 | |
Reading: on the 9070 XT and the M5 Max a 64-byte dependent read costs exactly what a 4-byte one costs (the line is fetched either way); on the 5090 a 64-byte read costs about two 4-byte reads (two 32-byte sectors) and the 64 B chase at full occupancy sits at 584 GB/s, a third of the stream.
**Hash rates** (5 timed dispatches of 2^24 nonces after a warm-up; Metal and the 9070 XT by device time, the 5090 by wall time around the stream sync; the PC cards switched off in the app for the run and restored, PC 1's 5090 and the integrated chip kept mining; the Mac under other agents' builds, load 4 to 9, so its absolute numbers carry that; the share = measured / (the card's probe ceiling at the class's widths / loads per hash)):
| Class | dataset B/hash | RTX 5090 MH/s (share) | RX 9070 XT MH/s (share) | M5 Max Metal MH/s (share) | 5090 / 9070 |
|---|---|---|---|---|---|
| v2 (w4, the lottery hash) | 512 | 136.1 (0.96) | 18.15 (0.87) | 27.74 (1.01) | 7.5x |
| w16 | 2,048 | 139.8 (0.90) | 17.90 (0.84) | 28.26 (1.03) | 7.8x |
| w64 | 8,192 | 71.9 (0.58) | 17.59 (0.78) | 28.27 (1.03) | 4.1x |
| w64x4 (32 loads) | 2,048 | 275.3 (0.56) | 75.19 (0.84) | 109.7 (1.00) | 3.7x |
| mix50-35-15, 6 programs: min / median / max (spread of median) | 1,664 to 3,680 | 99.5 / 114.2 / 121.0 (18.8%) | 17.45 / 18.76 / 18.83 (7.4%) | 25.36 / 27.26 / 28.43 (11.3%) | 6.1x |
| mix25-50-25, 6 programs | 2,240 to 5,024 | 95.9 / 107.3 / 119.8 (22.3%) | 17.84 / 18.45 / 18.85 (5.5%) | 23.21 / 24.68 / 25.21 (8.1%) | 5.8x |
Scratch (variant 5; N persistent warps; 5090: 2,048 warps launched against a resident capacity of 4,080 = 24 blocks/SM x 1 warp/block x 170 SMs at `--block-warps 1`, the occupancy query unchanged by the allocation (24 before and after); Metal: 2,048 to 16,384 warps swept, best shown; arena = N x per-warp size; the whole working set = 1 GiB dataset + 256 MiB cache + 128 MiB output + arena, under 2 GB on every row):
| Class (k of 16 slots, KiB per warp) | scratch ops/hash | dataset B/hash | RTX 5090 MH/s (vs scr0, share) | M5 Max Metal MH/s (vs scr0) | RX 9070 XT MH/s | 5090 arena / working set |
|---|---|---|---|---|---|---|
| scr0k32 (control, persistent loop, no RMW) | 0 | 512 | 139.1 (0, 0.98) | 28.25 (0) | 17.88 (control, 0.86) | 64 MiB / 1.4 GiB |
| scr2k32 (12.5%) | 16 | 448 | 114.4 (-18%, 0.80) | 26.14 (-7%) | 14.65 (-18%) | 64 MiB / 1.4 GiB |
| scr4k32 (25%) | 32 | 384 | 109.8 (-21%, 0.76) | 31.74 (+12%) | 14.00 (-22%) | 64 MiB / 1.4 GiB |
| scr8k32 (50%) | 64 | 256 | 122.1 (-12%, 0.82) | 49.08 (+74%) | 14.17 (-21%) | 64 MiB / 1.4 GiB |
| scr2k128 (12.5%) | 16 | 448 | 110.1 (-21%, 0.77) | 26.24 (-7%) | 14.07 (-21%) | 256 MiB / 1.6 GiB |
| scr4k128 (25%) | 32 | 384 | 98.0 (-30%, 0.68) | 28.08 (-1%) | 13.14 (-27%) | 256 MiB / 1.6 GiB |
| scr8k128 (50%) | 64 | 256 | 72.8 (-48%, 0.49) | 35.44 (+25%) | 12.03 (-33%) | 256 MiB / 1.6 GiB |
The 9070 XT rows are 2,048 persistent warps (4,096 within 1 percent), arena 64 MiB at 32 KiB and 256 MiB at 128 KiB, working set 1.4 and 1.6 GiB; its control (17.88, the persistent loop) equals its v2 rate (18.15) within 2 percent, and every RMW share costs it 18 to 33 percent: on AMD a scratch op is a dependent 16-byte read plus a write into a region the 64 MB Infinity Cache does not hold for 2,048 warps, so it is memory work there as on the 5090, not the cached op it is on Apple. Apple OpenCL on the same scratch packs (wall time, `--bench-pack --warps 2048`): scr0k32 27.85, scr2k32 28.58, scr4k32 32.43, scr8k32 47.93, scr2k128 25.67, scr4k128 27.24, scr8k128 32.93 MH/s, the Metal shape within 4 percent, fingerprints equal. Bytes moved per scratch op: 16 read + 16 written (the tag word included); per hash at 50 percent, 1,024 read + 1,024 written beside 256 of dataset reads. The 5090 at 4,096 launched warps (above its 4,080 resident) lost 2 to 26 percent (scr8k32 90.0 MH/s), so the rows above are the in-capacity launch.
**Readings.** (1) Same count, wider: the vendor gap does not move at 16 B (7.8x) because on the 9070 XT a 4-byte read already costs a 64-byte line and on the 5090 a 16-byte read costs one 32-byte sector, the same as 4 bytes: the memory systems do identical work, only the fold's input grows. At 64 B the gap closes to 4.1x, entirely by the 5090 losing half its rate (its share falls to 0.58 and its DRAM traffic reaches 589 GB/s, 37 percent of the stream: bandwidth, not latency, bounds it), while the 9070 XT and the M5 Max do not move. (2) Fewer, wider (w64x4): 3.7x, but every card runs 4x faster because the dependent chain is 32 loads long instead of 128; the 5090 sits at a 0.56 share (bandwidth), so a chip with more bandwidth per dollar than a GPU gains, which is the Ethash shape the design avoids. (3) The mix: the hour-to-hour spread is 7 to 22 percent of the median per card (the 5090 the widest, because its 64-byte loads are the expensive ones and their count per program runs 2 to 8 of 16); the programs with many 64-byte loads (mixA-3, mixA-5, mixB-2) are the slow hours on the 5090 and the fast ones nowhere. (4) The scratch: on the 5090 every RMW share costs 12 to 48 percent against the persistent control, the 32 KiB arena less than the 128 KiB one (the smaller arena, 64 MiB over 2,048 warps, sits inside the 96 MB L2); on the M5 Max the 32 KiB rows are FASTER than the control (+12 and +74 percent at 25 and 50 percent), because the arena (128 MiB over 4,096 warps) lives in the chip's caches and a scratch op is cheaper than a dataset read, so replacing dataset loads raises the rate: the scratch at these sizes is not memory work on Apple and is partly cached on NVIDIA. The chip row for these variants comes from the ca2-soundness branch; what this entry gives is the GPU cost and the share. (5) Latency-bound shares: v2 0.87 to 1.01 on the three cards, w16 0.84 to 1.03, w64 0.58 (5090) and 0.78 (9070 XT); the Mac's shares above 1 are an Apple OpenCL probe under load against a Metal rate.
Jobs: `run-readwidth-5090-20261005` and `run-readwidth-9070-20261005` (probes; the packs refused for their string seeds, fixed in a9e002c), `run-readwidth-5090-20261005c`, `run-readwidth-9070-20261005c` (benches), `run-readwidth-9070-scratch-20261005d` (the scratch packs after the `__local` fix d0018cf, AMD's compiler requires the exchange buffer at the kernel's outermost scope); read back with `node tools/jobs.mjs <id> --all`. Mac commands and logs: `docs/plans/read-width.md` section 3. The worker exes for the jobs: `proto-cuda/nvrtc/build-windows.sh` on this branch (mingw), sha256 of the CUDA one `6f46336f...defe1`.
## 5 October 2026 (night), the hot table on the M5 Max: a second table sized to GPU cache beside the 1 GiB dataset (Counter ASIC 2.0 layer 5)
Branch `ca2-cache` (on readwidth 1ea7a52), `docs/plans/hot-table.md`. Apple M5 Max, measure lock held, the Mac's load average 14 to 27 throughout (other agents' CPU work; the lock serialises builds and measurements, not every process), so the ratios inside one session are the result and the absolute rates are not quiet numbers. Packs `proto-cuda/packs-ca2-hot/hot{32,64,96}k4`, `hot64k2`, `hot64k8` from `igneum-pow export --seed igneum-genesis --day 2026-10-03 --class hot<S>k<k>` (the version 2 genesis program with k of its 16 loads redirected to an S MiB table H keyed by `seed_words("igneum-hot/" || seed bytes)`, read at `H[mulhi(src, words)]`).
**Probe** (`proto-opencl/igneum-bench-cl-igneum-genesis-mh --memprobe --probe-mib S`, Apple OpenCL, wall time, best of 3, 256 dependent steps per lane, work-group 256; the ceiling row is 4,194,304 lanes):
| MiB | chase at 4,096 lanes | ns per dependent load | chase ceiling, G loads/s | indep x8 ceiling | stream |
|---|---|---|---|---|---|
| 32 | 3.51 G/s | 1,168 | 21.7 | 21.8 | 138.7 GB/s |
| 64 | 3.63 | 1,129 | 12.8 | 13.0 | 199.1 |
| 96 | 3.30 | 1,242 | 12.3 | 12.7 | 242.6 |
| 1024 | 2.22 | 1,844 | 3.50 | 3.50 | 521.5 |
**Hash rate and bit-exactness** (Metal `proto-metal/packbench --pack <dir> --batches 5 --batch-log2 24`, GPU time; Apple OpenCL `--bench-pack --pack <dir> --batches 5 --batch-log2 24`, wall; both fill H on the device from the pack's `igneum_hot_fill` and check it; fingerprint = FNV-1a 64 over 2^24 outputs at base 0):
| Pack | Metal Mhash/s | Apple OpenCL Mhash/s | fingerprint (equal on both) | vectors | hot table head, last line, FNV | g against v2 (Metal) | probe-predicted g | ideal g |
|---|---|---|---|---|---|---|---|---|
| igneum-genesis-mh (v2) | 27.68 | 27.61 | 25f96e7dce90bd4e | 96/96 both | none | 1 | 1 | 1 |
| hot32k4 | 33.90 | 33.92 | d2e6cf3b61d0b9fe | 96/96 both | PASS both | 1.22 | 1.27 | 1.33 |
| hot64k4 | 30.93 | 30.89 | e4c5263ac650cc0d | 96/96 both | PASS both | 1.12 | 1.22 | 1.33 |
| hot96k4 | 29.06 | 28.97 | 5d63439b6e394521 | 96/96 both | PASS both | 1.05 | 1.22 | 1.33 |
| hot64k2 | 27.67 | 27.27 | 352633bdbbb0d2b6 | 96/96 both | PASS both | 1.00 | 1.10 | 1.14 |
| hot64k8 | 47.42 | 46.73 | da54630d7dfaaf85 | 96/96 both | PASS both | 1.71 | 1.57 | 2.0 |
Hot table fill, Metal GPU time: 0.07 ms (32 MiB), 0.15 (64), 0.22 (96). Hot table FNV-1a 64 of the genesis epoch: c1767ba3ef02719f (32 MiB), 77ca4b9527104530 (64), 79bcf436c4e5bc47 (96); cache 48c4f5bf24166b2e unchanged.
**CPU verifier** (`igneum-pow bench --seed igneum-genesis --day 2026-10-03 --class <c> --warps 50`, one core, release):
| Class | hot fill, one core | items per warp | ms per warp |
|---|---|---|---|
| v2 | none | 4,096 | 0.626 |
| hot32k4 | 24.0 ms | 3,072 | 0.489 |
| hot64k4 | 46.4 ms | 3,072 | 0.504 |
| hot96k4 | 73.0 ms | 3,072 | 0.488 |
| hot64k2 | 45.5 ms | 3,584 | 0.560 |
| hot64k8 | 47.7 ms | 2,048 | 0.344 |
Reading: bit-exact across Metal, Apple OpenCL and the Rust reference on every hot pack, hot table included. On this card the 32 MiB table delivers 92% of the probe's predicted gain with the dataset streaming beside it, 64 MiB about half, 96 MiB a quarter; k = 8 at 64 MiB gives 1.71x against an ideal 2.0x. The verifier gets cheaper with k (a hot load is one table read, a dataset load is an item derivation) and pays 24 to 73 ms per epoch for the fill. Chip model with these g in the plan, section 6.4. The RTX 5090 and RX 9070 XT rows are a prepared PC job (`relay/playbooks/ca2-hot-{5090,9070}-bench.ps1`, zip `~/Desktop/igneum-ca2-hot.zip`), not run.
Crate: `cargo test --release` 52 pass (39 unit, 13 pack tests: the four pinned v2 packs byte-identical, the five hot packs pinned with their load-form count: exactly 16 - k masked dataset loads and k hot loads per hash kernel).
**Addendum, the added form** (coordinator's form of 5 October 2026: 16 + k load slots, the k hot ones drawn among them, the 16 dataset loads and the 4,096-item verifier bound unchanged; packs `hot32k4a`, `hot64k4a`, `hot96k4a`; second Mac session 21:03 to 21:19 UTC, load average 7 to 14; same harnesses and commands, branch `ca2-cache` on ca2-v3 464d6e1, the hosts rebuilt on the merged packfile.h):
| Pack | Metal Mhash/s | Apple OpenCL Mhash/s | fingerprint (equal on both) | vectors | hot table | g against v2 (Metal, v2 27.63 in this session) | probe-predicted g | CPU verify ms/warp (v2 0.602) | hot fill, one core |
|---|---|---|---|---|---|---|---|---|---|
| hot32k4a | 25.76 | 25.72 | 8a3414735db4523c | 96/96 both | PASS both | 0.93 | 0.96 | 0.631 | 21.7 ms |
| hot64k4a | 23.92 | 23.87 | 45668f34105f6307 | 96/96 both | PASS both | 0.87 | 0.94 | 0.609 | 43.3 ms |
| hot96k4a | 22.92 | 22.88 | af763997dfee4c82 | 96/96 both | PASS both | 0.83 | 0.93 | 0.614 | 64.9 ms |
Reading: the added form costs this card 7, 13 and 17% of its rate at 32, 64 and 96 MiB for four extra loads per iteration, more than the probe predicts as the table grows; the verifier is unchanged (4,096 items, plus 32 table reads) and pays the fill per epoch. Chip arithmetic in the plan, section 6.4. All eight packs load and self-test through the rebuilt OpenCL host (the Windows exe's host.c) on the Mac.
**Addendum, the PCs** (5 October 2026, 21:29 to 21:35 UTC, PC 1 ae432dc7, app 0.3.9 before and after; fetch `fetch-ca2-hot-20261005` (zip sha256 bd49faa1c9d48024f49c615481faff5c68a4c09f0889dbaf009c208674d67b3f), jobs `run-ca2-hot-5090-20261005` (126 s) and `run-ca2-hot-9070-20261005` (247 s), both exit 0, the card under test switched off in the app through `api/cards` and restored; workers `igneum-worker-cuda.exe` sha256 956c4ab34f42cbcd1d2c1c6fb1a58fd9b3a8c70166df771cafcd0296ca6a27d4 and `igneum-worker-opencl.exe` sha256 32d3d34390aad70485c3524424c354223387137d383b5c5daf01f40073c12703, built from ca2-cache 196db96 on ca2-v3's merged packfile.h d2cd6e1; read back with `node tools/jobs.mjs <id> --all`):
Probe (`--memprobe --probe-mib S`, dependent 4 B chase ceiling at 4,194,304 lanes, G loads/s; ns per dependent load at 4,096 lanes in brackets):
| Card | 32 MiB | 64 | 96 | 1024 | stream at 1024 MiB |
|---|---|---|---|---|---|
| RTX 5090 (CUDA, wall) | 112.6 (320) | 112.6 (340) | 112.6 (336) | 17.6 (610) | 1,563 GB/s |
| RX 9070 XT (OpenCL, event) | 9.88 (396) | 9.47 (457) | 8.18 (454) | 2.43 (1,579) | 633 GB/s |
Rates (5 dispatches of 2^24 after a warm-up; 5090 `--bench --block-warps 1`, 9070 XT `--bench-pack --device 1` work-group 256; every row check=PASS with the Mac's fingerprint; v2 references from the readwidth entry, same night, same workers: 136.1 and 18.15 MH/s):
| Pack | 5090 MH/s | g | 9070 XT MH/s | g | ideal g |
|---|---|---|---|---|---|
| hot32k4 | 146.6 | 1.08 | 19.79 | 1.09 | 1.33 |
| hot64k4 | 140.8 | 1.03 | 18.73 | 1.03 | 1.33 |
| hot96k4 | 138.5 | 1.02 | 18.33 | 1.01 | 1.33 |
| hot64k2 | 137.5 | 1.01 | 18.17 | 1.00 | 1.14 |
| hot64k8 | 163.6 | 1.20 | 22.32 | 1.23 | 2.0 |
| hot32k4a | 118.7 | 0.87 | 15.27 | 0.84 | 1 |
| hot64k4a | 115.4 | 0.85 | 14.62 | 0.81 | 1 |
| hot96k4a | 114.4 | 0.84 | 14.56 | 0.80 | 1 |
Reading: the probe promises a full hit rate on the 5090 (every S inside the 96 MiB L2 at one ceiling, 6.4x DRAM) and the hash gets 2 to 8% at k = 4 and 20% at k = 8; the 9070 XT the same shape. The dataset's random lines evict the table from the shared cache on every card. The added form costs 13 to 20% of the rate. Recommendation in `docs/plans/hot-table.md` section 6.4: do not adopt layer 5 in either form on these measurements.
## 5 October 2026 (night), mixer x4 and the cache growth rule: the class v3 dataset construction, with the x8 candidate (Counter ASIC 2.0; branch ca2-mixer on ca2-v3 6c75dad; cryptographer's lane)
Machine: Apple M5 Max, 64 GiB, Darwin 25.6.0. Write-up `docs/plans/mixer-x4.md`; chip model `docs/analysis/chip-model-v3.md`; code `igneum-pow` (LoadClass mixer_mult and growth, memhard::Shape, the schedule, the three emitters), packs `proto-cuda/packs-ca2-mixer/`, tests `igneum-pow/tests/mixer.rs` and `tests/packs.rs`. Commits 0fc0ad1, 66eeba3, e4c04a7, 7ce8d1e, 504cae4, fe4e193 and this entry's.
What changed. Under program class v3 (`V3_CLASS = LoadClass::MX4`) every mixer application of the item derivation is `m = 4` applications with round keys `(r m + j + 1) x 0x9E3779B9`, the 8 dependent cache reads per item unchanged; the cache doubles when the dataset doubles (`growth_doublings(d) = floor(log2(1 + d / 1460))`: 2^26 words to day 1,459, 2^27 from day 1,460, 2^28 from day 4,380). Version 2 is byte-identical: fresh exports of igneum-genesis-mh and igneum-devnet-v4-epoch0 `diff -r` IDENTICAL against the checked-in packs, and the crate tests regenerate every pinned file. A v3 program of a seed is the v2 program of that seed instruction for instruction (v2 loads take no width roll); only the dataset words and the hashes change.
Bit-exactness, `with-lock.sh run`, 22:05 and 21:45 UTC: the two pinned v3 packs (mx4-genesis, mx4-devnet-epoch0: dataset words 0..15 `61ff2180 0d4c7e6c ...` and `afe80d67 b9fbd029 ...`, word MASK `5020180e` and `e6a99c7a`, unit at base 0 lane 0 `63acd2d273f475ba` and `212c6442b51e87ae`) and the two x8 candidate packs on Metal (`packbench`, built from this branch) and Apple OpenCL (`igneum-bench-cl --bench-pack`): 3/3 standalone and 3/3 in batch, 96 of 96 lanes, cache FNV-1a 64 unchanged from v2 (48c4f5bf24166b2e, 448274a57f508cbc), dataset head, word MASK and 64 samples PASS, one 2^24 fingerprint per pack across both harnesses (mx4 6f48d5a2aa0dbe5f and 73caaebb28e808fe; mx8 7c28cfb06c5c65a9 and bbb183f72692f840); hash rate the v2 rate (27.5 to 27.7 MH/s GPU time, the hash kernel is unchanged). Fuzz: 200 class v3 programs (4 units each across the 32-bit range, one in the top 256 nonces) interpreted twice on the CPU, 800 of 800; the same 200 packs on Metal 200 of 200 (`--batch-log2 9 --batch-base 4294967040`, the wrapping unit inside the window), every tenth on Apple OpenCL 20 of 20; x8: 50 of 50 on Metal, 5 of 5 on OpenCL. Stats (8,192 outputs per seed, two seeds): v3 avalanche 49.97 to 49.99 percent, worst bit z 1.92 to 3.09, 0 duplicates (v2 beside it 49.87 to 49.98, z 2.25 to 2.30). Edges: items 0, 1, 2^28 - 1, 2^32 - 1 by hand at m = 1, 2, 4, 8; words 0, 15, 16, 17, MASK - 1, MASK through the fetch path. Determinism: two epochs, every vector and file equal and equal to the pinned pack. The scratch soundness tests of ca2-soundness (cherry-pick 0d8f745) 7 of 7 on this tree. Crate: 44 lib + 12 packs + 4 mixer + 7 scratch tests pass. A first Metal fuzz run reported 200 of 200 FAIL on an empty RESULT line (a packbench built before the `--batch-base` cherry-pick); it was read as a failure, the harness rebuilt, the run repeated.
Timings, `with-lock.sh measure`, one session 21:40:12 to 21:40:23 UTC, one core, two rounds; the box carried a load average of 5.6 (one minute) and 26 (fifteen minutes) from unlocked processes, so the absolute figures are about 2.2x the quiet readwidth night's 0.604 ms v2 row and the ratios are the measurement:
| Construction | Verifier ms per 32-lane unit, avg of 50 (two rounds) | Worst cold unit | Against v2 | 256 MiB fill, one core | Metal 1 GiB build, GPU ms |
|---|---|---|---|---|---|
| v2 | 1.361 / 1.310 | 1.579 | 1 | 172 to 173 ms | 29.7 (first touch) / 21.0 |
| x4 (class v3) | 1.956 / 1.923 | 2.043 | 1.45x | 172 to 175 ms | 20.9 / 21.0 |
| x8 (candidate) | 2.785 / 2.790 | 2.942 | 2.09x | 172 ms | 21.9 / 21.9 |
Reading: the mixer multiplies the verifier's ALU part only (the 8 dependent misses per item are unchanged), hence 1.45x and 2.1x and not 4x and 8x; the Mac's GPU build is latency-bound and does not move with the mixer, so the "under 1 s on every discrete card" half of the x8 rule is the PC job (five packs, `relay/playbooks/mixer-x4-pc1-bench.ps1`, waiting for the go). Verification throughput (C19): a quiet 2026 core serves about 1,100 shares per second at x4 and 800 at x8 (1,660 at v2, re-cutting spec 09's 2,270), a 22,000-member pool at one share per 10 s needs 2 cores at x4 and 3 at x8, IBD over 108,000 headers is 1.6 min at x4 and 2.3 at x8 on that core; the 10 ms gate keeps 8.0 ms (x4) and 7.1 ms (x8) of margin on the loaded core, 6 to 7 ms on a 2019-class laptop core (approximate, unmeasured, O-1.14).
Chip model (`docs/analysis/chip-model-v3.md`): the on-die-cache recompute chip at 50 T op/s against the 5090's measured 136.1 MH/s: v2 334 MH/s, 2.45x bare, 7.4x with the 3x fixed-function factor; x4 83.5 MH/s, 0.61x bare, 1.84x with the factor, 1.53x with the 128 mm^2 N5 mirror deducted at equal silicon; x8 41.7 MH/s, 0.31x, 0.92x, 0.76x. The claim at x4 is "under 2x" with the margin thin on the equal-budget convention (a 3.3x factor or a 10 percent larger budget reads 2.0x); the hot table in the added form would have raised it to 2.1x to 2.2x at the 5090's g (kept as measured, not adopted). Nothing here is a measurement of a chip.
**Addendum, 22:15 UTC: the verifier regression, the PC 1 build rows, and x8 into v3.** The era agent measured the same v2 input with readwidth's binary (0.604 ms) and ca2-v3 HEAD's (1.33) in one minute; bisected under the measure lock to this branch's 0fc0ad1 (seam 6c75dad 0.610, 0fc0ad1 1.332; the "loaded box" reading above was wrong by that factor, the load was real but the 2x was the code). Cause: the item loop (`derive_items`) inlined into `MemhardCpu::fetch`; the mask hoisted, the mask constant, and the constant-mask loop inlined all stayed at 1.33, the same loop `#[inline(never)]` read 0.60 to 0.62. Fix: `derive_items_mask`, out of line, one instance per cache size with the line mask a constant. Measured the era agent's way (readwidth's binary beside the fixed one, same input, same minute, 22:07 UTC): v2 0.607 / 0.610 against 0.609 / 0.611; on the fixed binary x4 1.238 / 1.237 (2.0x), x8 2.077 / 2.058 (3.4x), worst cold 2.15 ms; the increments (+0.63, +1.46 ms per unit) equal the slow binary's. Lesson, the class: an inlined item loop costs 2.2x and nothing in the suite sees it; a verifier benchmark with a pinned bound in the crate's CI is filed for the next cut, and until then every change to the item loop is measured against the previous binary on the same input in the same minute. PC 1 (job run-mixer-x4-pc1-20261005, 22:00 to 22:04 UTC, the worker's `cache ... dataset ... ms` wall line): RTX 5090 dataset 23 to 25 ms at v2, x4 and x8; RX 9070 XT (gfx1201) 72 to 77 ms at all three; every fingerprint equal to the Mac's; rates the v2 rate (136.5 to 137.4 and 18.0 to 18.2 MH/s). Decision under the delegated rule (coordinator, 22:05 UTC): x8 enters class v3 (`V3_CLASS = MX8`); pinned packs mx8-genesis (7c28cfb06c5c65a9) and mx8-devnet-epoch0 through the chain path with the era inside (90f794dd556f7a3b, Metal and Apple OpenCL, 22:12 UTC); the x4 packs kept as the candidate's record. Chip headline at x8: 41.7 MH/s, 0.31x bare, 0.92x with the 3x factor, 0.76x at equal silicon (`docs/analysis/chip-model-v3.md`).
## 5 October 2026 (evening), EVM transaction relay: three nodes in a chain, every transaction sent to one end included by the other two miners (execution and networking engineer)
Until this change the node did not relay EVM transactions to its peers, so a transaction sent to one node was only ever included by that node's own templates (this file, "5 October 2026 (afternoon), live devnet: real transactions": 3,794 transfers, all in the Mac's blocks; execution-layer ledger item 9). Fork branch `tx-gossip` (worktree `vendor/igneum-node-txgossip`, from release-0.3.6 a24ab01a, commit e242acd0), main repo branch `tx-gossip`. Design in `docs/design/execution-layer.md` 1.4 "Relay"; the hand-out cooldown of its 10.2 table is gone with it (row "Mempool hold").

View file

@ -0,0 +1,381 @@
{
"pass": true,
"checks": {
"switch_line_on_every_node": true,
"switch_line_names_the_rounded_epoch": true,
"template_switched_at_the_first_v3_epoch": true,
"blocks_before_the_boundary": true,
"blocks_after_the_boundary": true,
"v2_and_v3_programs_seen": true,
"program_ids_differ_across_the_switch": true,
"miners_agree_on_every_program": true,
"zero_rejected_by_miners": true,
"zero_rejected_by_nodes": true,
"sinks_agree": true,
"block_counts_agree": true
},
"activation": 150,
"epoch_blocks": 60,
"first_v3_epoch": 3,
"boundary_daa": 180,
"secs": 480,
"threads": 1,
"node": "/Users/joshm/Projects/igneum-wt-ca2-v3/vendor/igneum-node-ca2/target-ca2/release/igneumd",
"miner": "/Users/joshm/Projects/igneum-wt-ca2-v3/vendor/igneum-node-ca2/target-ca2/release/igneum-miner",
"genesis_bits": "0x1f010000",
"template_switch": {
"epoch": 3,
"daa": 180,
"at": 168
},
"run_ended_at_s": 298.3,
"final_daa": 300,
"blocks": {
"total": 305,
"before_boundary": 181,
"after_boundary": 124,
"chain_before": 176,
"chain_after": 123
},
"programs": [
{
"epoch": 0,
"class": "v2",
"program_id": "8f8806638d59850f",
"seed": "234e082d653dc69d",
"miners": 3,
"disagree": false,
"ready_ms": [
219,
235,
232
]
},
{
"epoch": 1,
"class": "v2",
"program_id": "fd9562df32a68313",
"seed": "9b36731951ad7fb3",
"miners": 3,
"disagree": false,
"ready_ms": [
2,
2,
2
]
},
{
"epoch": 2,
"class": "v2",
"program_id": "1ae6d90ab299154c",
"seed": "ce88239dfe686691",
"miners": 3,
"disagree": false,
"ready_ms": [
2,
2,
2
]
},
{
"epoch": 3,
"class": "v3",
"program_id": "5d0dedd9fd9e29a1",
"seed": "f177ab855facbeb4",
"miners": 3,
"disagree": false,
"ready_ms": [
185,
179,
183,
177,
183,
177
]
},
{
"epoch": 4,
"class": "v3",
"program_id": "e81808dcdb02ce05",
"seed": "bcbc22513d3b73a2",
"miners": 3,
"disagree": false,
"ready_ms": [
2,
2,
2
]
},
{
"epoch": 5,
"class": "v3",
"program_id": "06aff9c1d33e7a13",
"seed": "8d6755bc0eb15465",
"miners": 3,
"disagree": false,
"ready_ms": [
2,
2,
2
]
}
],
"accepted_per_miner": [
97,
103,
104
],
"rejected_by_miners": [
0,
0,
0
],
"rejected_by_nodes": [
0,
0,
0
],
"rejected_lines": [],
"sinks": [
"082fd39ba65df2ff",
"082fd39ba65df2ff",
"082fd39ba65df2ff"
],
"block_counts": [
304,
304,
304
],
"tips_per_node": [
1,
1,
1
],
"switch_lines": [
"Program class v3 from the override file: active from epoch 3 (DAA score 150 rounded up to the epoch boundary at 180, epochs of 60 DAA)",
"Program class v3 from the override file: active from epoch 3 (DAA score 150 rounded up to the epoch boundary at 180, epochs of 60 DAA)",
"Program class v3 from the override file: active from epoch 3 (DAA score 150 rounded up to the epoch boundary at 180, epochs of 60 DAA)"
],
"samples": [
{
"t": 4.7,
"daa": 0,
"epoch": 0,
"class": 2,
"nodes": [
"0/234e082d",
"0/234e082d",
"0/234e082d"
]
},
{
"t": 19.7,
"daa": 11,
"epoch": 0,
"class": 2,
"nodes": [
"11/c036063c",
"11/c036063c",
"11/c036063c"
]
},
{
"t": 34.7,
"daa": 34,
"epoch": 0,
"class": 2,
"nodes": [
"34/62ae0e8e",
"34/62ae0e8e",
"34/62ae0e8e"
]
},
{
"t": 49.8,
"daa": 53,
"epoch": 0,
"class": 2,
"nodes": [
"53/c21d7e86",
"53/c21d7e86",
"53/c21d7e86"
]
},
{
"t": 64.8,
"daa": 70,
"epoch": 1,
"class": 2,
"nodes": [
"70/9cb940da",
"70/9cb940da",
"70/9cb940da"
]
},
{
"t": 79.8,
"daa": 88,
"epoch": 1,
"class": 2,
"nodes": [
"88/c84b8514",
"88/c84b8514",
"88/c84b8514"
]
},
{
"t": 94.9,
"daa": 103,
"epoch": 1,
"class": 2,
"nodes": [
"103/1e2dfed3",
"103/1e2dfed3",
"103/1e2dfed3"
]
},
{
"t": 109.9,
"daa": 111,
"epoch": 1,
"class": 2,
"nodes": [
"111/aa7ea4cd",
"111/aa7ea4cd",
"111/aa7ea4cd"
]
},
{
"t": 124.9,
"daa": 139,
"epoch": 2,
"class": 2,
"nodes": [
"139/d52be92d",
"139/d52be92d",
"139/d52be92d"
]
},
{
"t": 139.9,
"daa": 153,
"epoch": 2,
"class": 2,
"nodes": [
"153/02a4323f",
"153/02a4323f",
"153/02a4323f"
]
},
{
"t": 155,
"daa": 165,
"epoch": 2,
"class": 2,
"nodes": [
"165/f2e38b3d",
"165/f2e38b3d",
"165/f2e38b3d"
]
},
{
"t": 170,
"daa": 183,
"epoch": 3,
"class": 3,
"nodes": [
"183/98f2ca9d",
"183/98f2ca9d",
"183/98f2ca9d"
]
},
{
"t": 185.1,
"daa": 195,
"epoch": 3,
"class": 3,
"nodes": [
"195/eecbf07e",
"195/eecbf07e",
"195/eecbf07e"
]
},
{
"t": 200.1,
"daa": 205,
"epoch": 3,
"class": 3,
"nodes": [
"205/ad4a2646",
"205/ad4a2646",
"205/ad4a2646"
]
},
{
"t": 215.1,
"daa": 221,
"epoch": 3,
"class": 3,
"nodes": [
"221/f8626ffd",
"221/f8626ffd",
"221/f8626ffd"
]
},
{
"t": 230.2,
"daa": 233,
"epoch": 3,
"class": 3,
"nodes": [
"233/d94e1815",
"233/d94e1815",
"233/d94e1815"
]
},
{
"t": 245.2,
"daa": 251,
"epoch": 4,
"class": 3,
"nodes": [
"251/f724fd9e",
"251/f724fd9e",
"251/f724fd9e"
]
},
{
"t": 260.2,
"daa": 266,
"epoch": 4,
"class": 3,
"nodes": [
"266/b91936fd",
"266/b91936fd",
"266/b91936fd"
]
},
{
"t": 275.2,
"daa": 276,
"epoch": 4,
"class": 3,
"nodes": [
"276/7c8f720d",
"276/7c8f720d",
"276/7c8f720d"
]
},
{
"t": 290.3,
"daa": 297,
"epoch": 4,
"class": 3,
"nodes": [
"297/63eb2574",
"297/63eb2574",
"297/63eb2574"
]
}
]
}

View file

@ -0,0 +1,378 @@
{
"pass": true,
"checks": {
"switch_line_on_every_node": true,
"switch_line_names_the_rounded_epoch": true,
"template_switched_at_the_first_v3_epoch": true,
"blocks_before_the_boundary": true,
"blocks_after_the_boundary": true,
"v2_and_v3_programs_seen": true,
"program_ids_differ_across_the_switch": true,
"miners_agree_on_every_program": true,
"zero_rejected_by_miners": true,
"zero_rejected_by_nodes": true,
"sinks_agree": true,
"block_counts_agree": true
},
"activation": 150,
"epoch_blocks": 60,
"first_v3_epoch": 3,
"boundary_daa": 180,
"secs": 480,
"threads": 1,
"node": "/Users/joshm/Projects/igneum-wt-ca2-v3/vendor/igneum-node-ca2/target-ca2/release/igneumd",
"miner": "/Users/joshm/Projects/igneum-wt-ca2-v3/vendor/igneum-node-ca2/target-ca2/release/igneum-miner",
"genesis_bits": "0x1f010000",
"template_switch": {
"epoch": 3,
"daa": 180,
"at": 171
},
"run_ended_at_s": 292.3,
"final_daa": 300,
"blocks": {
"total": 305,
"before_boundary": 181,
"after_boundary": 124,
"chain_before": 180,
"chain_after": 122
},
"programs": [
{
"epoch": 0,
"class": "v2",
"program_id": "8f8806638d59850f",
"seed": "234e082d653dc69d",
"miners": 3,
"disagree": false,
"ready_ms": [
177,
177,
178
]
},
{
"epoch": 1,
"class": "v2",
"program_id": "bb8dd9ddbf9eb63f",
"seed": "2f66af56da44ca92",
"miners": 3,
"disagree": false,
"ready_ms": [
2,
2,
2
]
},
{
"epoch": 2,
"class": "v2",
"program_id": "8ee7a9f33d418e48",
"seed": "5194dfa8a53259bc",
"miners": 3,
"disagree": false,
"ready_ms": [
2,
2,
2
]
},
{
"epoch": 3,
"class": "v3",
"program_id": "2d278041ba482dba",
"seed": "54c353de1d8609d7",
"miners": 3,
"disagree": false,
"ready_ms": [
181,
179,
179
]
},
{
"epoch": 4,
"class": "v3",
"program_id": "2ae786d294a8a59d",
"seed": "8529c69223a2194c",
"miners": 3,
"disagree": false,
"ready_ms": [
2,
2,
2
]
},
{
"epoch": 5,
"class": "v3",
"program_id": "bc36813df2f41b5f",
"seed": "fdc233b69c42ffea",
"miners": 3,
"disagree": false,
"ready_ms": [
2,
2,
2
]
}
],
"accepted_per_miner": [
102,
102,
100
],
"rejected_by_miners": [
0,
0,
0
],
"rejected_by_nodes": [
0,
0,
0
],
"rejected_lines": [],
"sinks": [
"712c1b212091dcdc",
"712c1b212091dcdc",
"712c1b212091dcdc"
],
"block_counts": [
303,
303,
303
],
"tips_per_node": [
1,
1,
1
],
"switch_lines": [
"Program class v3 from the override file: active from epoch 3 (DAA score 150 rounded up to the epoch boundary at 180, epochs of 60 DAA)",
"Program class v3 from the override file: active from epoch 3 (DAA score 150 rounded up to the epoch boundary at 180, epochs of 60 DAA)",
"Program class v3 from the override file: active from epoch 3 (DAA score 150 rounded up to the epoch boundary at 180, epochs of 60 DAA)"
],
"samples": [
{
"t": 4.6,
"daa": 0,
"epoch": 0,
"class": 2,
"nodes": [
"0/234e082d",
"0/234e082d",
"0/234e082d"
]
},
{
"t": 19.7,
"daa": 18,
"epoch": 0,
"class": 2,
"nodes": [
"18/5b5f2d36",
"18/5b5f2d36",
"18/5b5f2d36"
]
},
{
"t": 34.7,
"daa": 51,
"epoch": 0,
"class": 2,
"nodes": [
"51/333c61cc",
"51/333c61cc",
"51/333c61cc"
]
},
{
"t": 49.8,
"daa": 67,
"epoch": 1,
"class": 2,
"nodes": [
"67/f1fcbc43",
"67/f1fcbc43",
"67/f1fcbc43"
]
},
{
"t": 64.8,
"daa": 89,
"epoch": 1,
"class": 2,
"nodes": [
"89/91180594",
"89/91180594",
"89/91180594"
]
},
{
"t": 79.8,
"daa": 101,
"epoch": 1,
"class": 2,
"nodes": [
"101/49dd0422",
"101/49dd0422",
"101/49dd0422"
]
},
{
"t": 94.9,
"daa": 111,
"epoch": 1,
"class": 2,
"nodes": [
"111/a65a10b3",
"111/a65a10b3",
"111/a65a10b3"
]
},
{
"t": 109.9,
"daa": 122,
"epoch": 2,
"class": 2,
"nodes": [
"122/f06b1eb3",
"122/f06b1eb3",
"122/f06b1eb3"
]
},
{
"t": 124.9,
"daa": 140,
"epoch": 2,
"class": 2,
"nodes": [
"140/bbfe6710",
"140/bbfe6710",
"140/bbfe6710"
]
},
{
"t": 140,
"daa": 155,
"epoch": 2,
"class": 2,
"nodes": [
"155/a4e736dd",
"155/a4e736dd",
"155/a4e736dd"
]
},
{
"t": 155,
"daa": 171,
"epoch": 2,
"class": 2,
"nodes": [
"171/c754adeb",
"171/c754adeb",
"171/c754adeb"
]
},
{
"t": 170,
"daa": 179,
"epoch": 2,
"class": 2,
"nodes": [
"179/da89ab4d",
"179/da89ab4d",
"179/da89ab4d"
]
},
{
"t": 185,
"daa": 187,
"epoch": 3,
"class": 3,
"nodes": [
"187/e0f61041",
"187/e0f61041",
"187/e0f61041"
]
},
{
"t": 200.1,
"daa": 204,
"epoch": 3,
"class": 3,
"nodes": [
"204/5d37e885",
"204/5d37e885",
"204/5d37e885"
]
},
{
"t": 215.1,
"daa": 222,
"epoch": 3,
"class": 3,
"nodes": [
"222/46643165",
"222/46643165",
"222/46643165"
]
},
{
"t": 230.1,
"daa": 236,
"epoch": 3,
"class": 3,
"nodes": [
"236/17eb5a95",
"236/17eb5a95",
"236/17eb5a95"
]
},
{
"t": 245.2,
"daa": 247,
"epoch": 4,
"class": 3,
"nodes": [
"247/7aada19d",
"247/7aada19d",
"247/7aada19d"
]
},
{
"t": 260.2,
"daa": 261,
"epoch": 4,
"class": 3,
"nodes": [
"261/15d8f83e",
"261/15d8f83e",
"261/15d8f83e"
]
},
{
"t": 275.3,
"daa": 274,
"epoch": 4,
"class": 3,
"nodes": [
"274/9f3d19c5",
"274/9f3d19c5",
"274/9f3d19c5"
]
},
{
"t": 290.3,
"daa": 297,
"epoch": 4,
"class": 3,
"nodes": [
"297/fa7975bd",
"297/fa7975bd",
"297/fa7975bd"
]
}
]
}

View file

@ -0,0 +1,367 @@
{
"pass": true,
"checks": {
"switch_line_on_every_node": true,
"switch_line_names_the_rounded_epoch": true,
"template_switched_at_the_first_v3_epoch": true,
"blocks_before_the_boundary": true,
"blocks_after_the_boundary": true,
"v2_and_v3_programs_seen": true,
"program_ids_differ_across_the_switch": true,
"miners_agree_on_every_program": true,
"zero_rejected_by_miners": true,
"zero_rejected_by_nodes": true,
"sinks_agree": true,
"block_counts_agree": true
},
"activation": 150,
"epoch_blocks": 60,
"first_v3_epoch": 3,
"boundary_daa": 180,
"secs": 480,
"threads": 1,
"node": "/Users/joshm/Projects/igneum-wt-ca2-v3/vendor/igneum-node-ca2/target-ca2/release/igneumd",
"miner": "/Users/joshm/Projects/igneum-wt-ca2-v3/vendor/igneum-node-ca2/target-ca2/release/igneum-miner",
"genesis_bits": "0x1f010000",
"template_switch": {
"epoch": 3,
"daa": 181,
"at": 127.9
},
"run_ended_at_s": 280.2,
"final_daa": 300,
"blocks": {
"total": 304,
"before_boundary": 182,
"after_boundary": 122,
"chain_before": 180,
"chain_after": 120
},
"programs": [
{
"epoch": 0,
"class": "v2",
"program_id": "8f8806638d59850f",
"seed": "234e082d653dc69d",
"miners": 3,
"disagree": false,
"ready_ms": [
179,
178,
179
]
},
{
"epoch": 1,
"class": "v2",
"program_id": "e145305446a6b4ce",
"seed": "09952ab515cc5a10",
"miners": 3,
"disagree": false,
"ready_ms": [
2,
2,
2
]
},
{
"epoch": 2,
"class": "v2",
"program_id": "743ad2a3cab0518a",
"seed": "1812a8eabb5a452c",
"miners": 3,
"disagree": false,
"ready_ms": [
2,
2,
2
]
},
{
"epoch": 3,
"class": "v3",
"program_id": "a6523b90cff501e3",
"seed": "6296e3d38df15872",
"miners": 3,
"disagree": false,
"ready_ms": [
191,
191,
191
]
},
{
"epoch": 4,
"class": "v3",
"program_id": "bc811b3c4b8b1ced",
"seed": "a59152f149a517e0",
"miners": 3,
"disagree": false,
"ready_ms": [
2,
2,
2
]
},
{
"epoch": 5,
"class": "v3",
"program_id": "e784541f19cdebe5",
"seed": "87ec7849fce0019f",
"miners": 3,
"disagree": false,
"ready_ms": [
2,
2,
2
]
}
],
"accepted_per_miner": [
101,
89,
113
],
"rejected_by_miners": [
0,
0,
0
],
"rejected_by_nodes": [
0,
0,
0
],
"rejected_lines": [],
"sinks": [
"a9ce45df8beeaf13",
"a9ce45df8beeaf13",
"a9ce45df8beeaf13"
],
"block_counts": [
303,
303,
303
],
"tips_per_node": [
1,
1,
1
],
"switch_lines": [
"Program class v3 from the override file: active from epoch 3 (DAA score 150 rounded up to the epoch boundary at 180, epochs of 60 DAA)",
"Program class v3 from the override file: active from epoch 3 (DAA score 150 rounded up to the epoch boundary at 180, epochs of 60 DAA)",
"Program class v3 from the override file: active from epoch 3 (DAA score 150 rounded up to the epoch boundary at 180, epochs of 60 DAA)"
],
"samples": [
{
"t": 4.6,
"daa": 0,
"epoch": 0,
"class": 2,
"nodes": [
"0/234e082d",
"0/234e082d",
"0/234e082d"
]
},
{
"t": 19.7,
"daa": 48,
"epoch": 0,
"class": 2,
"nodes": [
"48/78df0834",
"48/78df0834",
"48/78df0834"
]
},
{
"t": 34.7,
"daa": 67,
"epoch": 1,
"class": 2,
"nodes": [
"67/0b934b49",
"67/0b934b49",
"67/0b934b49"
]
},
{
"t": 49.7,
"daa": 84,
"epoch": 1,
"class": 2,
"nodes": [
"84/2ba5b427",
"84/2ba5b427",
"84/2ba5b427"
]
},
{
"t": 64.8,
"daa": 101,
"epoch": 1,
"class": 2,
"nodes": [
"101/6574b577",
"101/6574b577",
"101/6574b577"
]
},
{
"t": 79.8,
"daa": 123,
"epoch": 2,
"class": 2,
"nodes": [
"123/21e2b6ce",
"123/21e2b6ce",
"123/21e2b6ce"
]
},
{
"t": 94.8,
"daa": 134,
"epoch": 2,
"class": 2,
"nodes": [
"134/74b83d55",
"134/74b83d55",
"134/74b83d55"
]
},
{
"t": 109.9,
"daa": 153,
"epoch": 2,
"class": 2,
"nodes": [
"153/71b251a6",
"153/71b251a6",
"153/71b251a6"
]
},
{
"t": 124.9,
"daa": 175,
"epoch": 2,
"class": 2,
"nodes": [
"175/b1500fb5",
"175/b1500fb5",
"175/b1500fb5"
]
},
{
"t": 139.9,
"daa": 185,
"epoch": 3,
"class": 3,
"nodes": [
"185/dd8ff66c",
"185/dd8ff66c",
"185/dd8ff66c"
]
},
{
"t": 155,
"daa": 191,
"epoch": 3,
"class": 3,
"nodes": [
"191/2ef77979",
"191/2ef77979",
"191/2ef77979"
]
},
{
"t": 170,
"daa": 194,
"epoch": 3,
"class": 3,
"nodes": [
"194/1871cd31",
"194/1871cd31",
"194/1871cd31"
]
},
{
"t": 185,
"daa": 206,
"epoch": 3,
"class": 3,
"nodes": [
"206/808070af",
"206/808070af",
"206/808070af"
]
},
{
"t": 200.1,
"daa": 221,
"epoch": 3,
"class": 3,
"nodes": [
"221/1bccc51c",
"221/1bccc51c",
"221/1bccc51c"
]
},
{
"t": 215.1,
"daa": 237,
"epoch": 3,
"class": 3,
"nodes": [
"237/d0bfaef8",
"237/d0bfaef8",
"237/d0bfaef8"
]
},
{
"t": 230.1,
"daa": 253,
"epoch": 4,
"class": 3,
"nodes": [
"253/a4136a32",
"253/a4136a32",
"253/a4136a32"
]
},
{
"t": 245.2,
"daa": 268,
"epoch": 4,
"class": 3,
"nodes": [
"268/51ef3166",
"268/51ef3166",
"268/51ef3166"
]
},
{
"t": 260.2,
"daa": 277,
"epoch": 4,
"class": 3,
"nodes": [
"277/869d0dd3",
"277/869d0dd3",
"277/869d0dd3"
]
},
{
"t": 275.2,
"daa": 292,
"epoch": 4,
"class": 3,
"nodes": [
"292/fe54acb8",
"292/fe54acb8",
"292/fe54acb8"
]
}
]
}

View file

@ -0,0 +1,474 @@
{
"pass": false,
"checks": {
"switch_line_on_every_node": true,
"switch_line_names_the_rounded_epoch": true,
"template_switched_at_the_first_v3_epoch": true,
"blocks_before_the_boundary": true,
"blocks_after_the_boundary": true,
"v2_and_v3_programs_seen": true,
"program_ids_differ_across_the_switch": true,
"miners_agree_on_every_program": true,
"zero_rejected_by_miners": true,
"zero_rejected_by_nodes": true,
"sinks_agree": true,
"block_counts_agree": true,
"metal_prepare_sent_for_v3": true,
"metal_worker_prepared_v3_pack": true,
"metal_no_need_or_mismatch": false,
"metal_accepted_blocks_after_switch": true,
"metal_cpu_recheck_clean": true,
"metal_swapped_without_pause": true
},
"activation": 150,
"epoch_blocks": 60,
"first_v3_epoch": 3,
"boundary_daa": 180,
"secs": 480,
"threads": 1,
"node": "/Users/joshm/Projects/igneum-wt-ca2-v3/vendor/igneum-node-ca2/target-ca2/release/igneumd",
"miner": "/Users/joshm/Projects/igneum-wt-ca2-v3/vendor/igneum-node-ca2/target-ca2/release/igneum-miner",
"genesis_bits": "0x1e010000",
"template_switch": {
"epoch": 3,
"daa": 180,
"at": 264.5
},
"run_ended_at_s": 366.8,
"final_daa": 301,
"blocks": {
"total": 305,
"before_boundary": 182,
"after_boundary": 123,
"chain_before": 91,
"chain_after": 67
},
"programs": [
{
"epoch": 0,
"class": "v2",
"program_id": "3e974094b5ce543f",
"seed": "130e3b7d568db07a",
"miners": 2,
"disagree": false,
"ready_ms": [
192,
192
]
},
{
"epoch": 1,
"class": "v2",
"program_id": "1d5b129e89ec2995",
"seed": "bc056261dafe0ec2",
"miners": 2,
"disagree": false,
"ready_ms": [
4,
4
]
},
{
"epoch": 2,
"class": "v2",
"program_id": "835e953cdc2c11be",
"seed": "1908705502833bdf",
"miners": 2,
"disagree": false,
"ready_ms": [
9,
7
]
},
{
"epoch": 3,
"class": "v3",
"program_id": "082b7ee882aaf418",
"seed": "4ae011065b84f679",
"miners": 2,
"disagree": false,
"ready_ms": [
1069,
1085
]
},
{
"epoch": 4,
"class": "v3",
"program_id": "7ef81e08e4611665",
"seed": "d5964a745dbe2ce2",
"miners": 2,
"disagree": false,
"ready_ms": [
6,
6
]
},
{
"epoch": 5,
"class": "v3",
"program_id": "ecee6312983c15ee",
"seed": "f9d5de9eea7330ae",
"miners": 2,
"disagree": false,
"ready_ms": [
5,
5
]
}
],
"accepted_per_miner": [
301,
2,
1
],
"rejected_by_miners": [
0,
0,
0
],
"rejected_by_nodes": [
0,
0,
0
],
"rejected_lines": [],
"sinks": [
"66a33eca15d33ba2",
"66a33eca15d33ba2",
"66a33eca15d33ba2"
],
"block_counts": [
304,
304,
304
],
"tips_per_node": [
2,
2,
2
],
"switch_lines": [
"Program class v3 from the override file: active from epoch 3 (DAA score 150 rounded up to the epoch boundary at 180, epochs of 60 DAA)",
"Program class v3 from the override file: active from epoch 3 (DAA score 150 rounded up to the epoch boundary at 180, epochs of 60 DAA)",
"Program class v3 from the override file: active from epoch 3 (DAA score 150 rounded up to the epoch boundary at 180, epochs of 60 DAA)"
],
"samples": [
{
"t": 4.6,
"daa": 1,
"epoch": 0,
"class": 2,
"nodes": [
"1/97755438",
"0/130e3b7d",
"0/130e3b7d"
]
},
{
"t": 19.7,
"daa": 3,
"epoch": 0,
"class": 2,
"nodes": [
"3/52e2c440",
"3/52e2c440",
"3/52e2c440"
]
},
{
"t": 34.7,
"daa": 3,
"epoch": 0,
"class": 2,
"nodes": [
"3/52e2c440",
"3/52e2c440",
"3/52e2c440"
]
},
{
"t": 49.7,
"daa": 13,
"epoch": 0,
"class": 2,
"nodes": [
"13/fa702087",
"13/fa702087",
"13/fa702087"
]
},
{
"t": 64.8,
"daa": 38,
"epoch": 0,
"class": 2,
"nodes": [
"38/2a687376",
"38/2a687376",
"38/2a687376"
]
},
{
"t": 79.8,
"daa": 55,
"epoch": 0,
"class": 2,
"nodes": [
"55/2f1f0ea8",
"55/2f1f0ea8",
"55/2f1f0ea8"
]
},
{
"t": 94.8,
"daa": 57,
"epoch": 0,
"class": 2,
"nodes": [
"57/e1b2abc1",
"57/e1b2abc1",
"57/e1b2abc1"
]
},
{
"t": 109.9,
"daa": 61,
"epoch": 1,
"class": 2,
"nodes": [
"61/86fa7300",
"61/86fa7300",
"61/86fa7300"
]
},
{
"t": 124.9,
"daa": 61,
"epoch": 1,
"class": 2,
"nodes": [
"61/86fa7300",
"61/86fa7300",
"61/86fa7300"
]
},
{
"t": 140,
"daa": 61,
"epoch": 1,
"class": 2,
"nodes": [
"61/86fa7300",
"61/86fa7300",
"61/86fa7300"
]
},
{
"t": 155,
"daa": 73,
"epoch": 1,
"class": 2,
"nodes": [
"73/e7f6e171",
"73/e7f6e171",
"73/e7f6e171"
]
},
{
"t": 170.1,
"daa": 94,
"epoch": 1,
"class": 2,
"nodes": [
"94/0b1c93b1",
"94/0b1c93b1",
"94/0b1c93b1"
]
},
{
"t": 185.2,
"daa": 114,
"epoch": 1,
"class": 2,
"nodes": [
"114/16034702",
"114/16034702",
"114/16034702"
]
},
{
"t": 200.2,
"daa": 117,
"epoch": 1,
"class": 2,
"nodes": [
"117/d018a345",
"117/d018a345",
"117/d018a345"
]
},
{
"t": 215.4,
"daa": 120,
"epoch": 2,
"class": 2,
"nodes": [
"120/e022d64e",
"120/e022d64e",
"120/e022d64e"
]
},
{
"t": 230.4,
"daa": 134,
"epoch": 2,
"class": 2,
"nodes": [
"134/e7d07203",
"134/e7d07203",
"134/e7d07203"
]
},
{
"t": 245.4,
"daa": 157,
"epoch": 2,
"class": 2,
"nodes": [
"157/d23b3a1b",
"157/d23b3a1b",
"157/d23b3a1b"
]
},
{
"t": 260.5,
"daa": 175,
"epoch": 2,
"class": 2,
"nodes": [
"175/4bf5acbe",
"175/4bf5acbe",
"175/4bf5acbe"
]
},
{
"t": 275.5,
"daa": 198,
"epoch": 3,
"class": 3,
"nodes": [
"198/f337dbc1",
"198/f337dbc1",
"198/f337dbc1"
]
},
{
"t": 290.6,
"daa": 217,
"epoch": 3,
"class": 3,
"nodes": [
"217/dc28325d",
"217/dc28325d",
"217/dc28325d"
]
},
{
"t": 305.6,
"daa": 234,
"epoch": 3,
"class": 3,
"nodes": [
"234/a0079166",
"234/a0079166",
"234/a0079166"
]
},
{
"t": 320.7,
"daa": 249,
"epoch": 4,
"class": 3,
"nodes": [
"249/eb206cb2",
"249/eb206cb2",
"249/eb206cb2"
]
},
{
"t": 335.7,
"daa": 268,
"epoch": 4,
"class": 3,
"nodes": [
"268/9b9866b1",
"268/9b9866b1",
"268/9b9866b1"
]
},
{
"t": 350.7,
"daa": 285,
"epoch": 4,
"class": 3,
"nodes": [
"285/3263ff7a",
"285/3263ff7a",
"285/3263ff7a"
]
},
{
"t": 365.8,
"daa": 299,
"epoch": 4,
"class": 3,
"nodes": [
"299/31eff7d0",
"299/31eff7d0",
"299/31eff7d0"
]
}
],
"metal": {
"worker": "/Users/joshm/Projects/igneum-wt-ca2-v3/proto-metal/igneum-bench-ca2",
"prepares": [
"PREPARE sent for epoch seed b7eda892cff02b39c043f5fa683a7c1629c96d02fa480522603be595372e371f day 1243916 class v2 (6 DAA blocks before the boundary at 60; CPU side ready in 960 ms)",
"PREPARE sent for epoch seed bc056261dafe0ec266f2e8f508e8fe7349e4d0c064853c6fd37ffa6a5e1fab3c day 1243916 class v2 (the current pair, after seed mismatches; CPU side ready in 1248 ms)",
"PREPARE sent for epoch seed 1908705502833bdf6c5d08b6d6ae3a364e3fe2bf0956504b20bc5071eecbdfc3 day 1243916 class v2 (6 DAA blocks before the boundary at 120; CPU side ready in 1768 ms)",
"PREPARE sent for epoch seed 4ae011065b84f67924d0c42ff63669e6124b6bddc4a7f71a3f1075b5891efd8c day 1243916 class v3 (6 DAA blocks before the boundary at 180; CPU side ready in 1740 ms)",
"PREPARE sent for epoch seed d5964a745dbe2ce2a88d6a4198a244996f1f4aba6cc1bc20841bb532e469fdd6 day 1243916 class v3 (6 DAA blocks before the boundary at 240; CPU side ready in 1289 ms)",
"PREPARE sent for epoch seed f9d5de9eea7330aeec83c0994b8929e61a2ac2f13dbae6462cceace3b1d6f24e day 1243916 class v3 (7 DAA blocks before the boundary at 300; CPU side ready in 1166 ms)"
],
"prepared": [
"worker: prepared b7eda892cff02b39c043f5fa683a7c1629c96d02fa480522603be595372e371f 69676e65756d2d6461792f0cfb120000000000 35663.4 program 64.5 dataset 0.0 race 35598.9 variant base class v2 loads/hash 128 cache-fill 2.0 build 30.5 resident 2 programs 1 datasets (prepare answered 1.3 s after it was sent)",
"worker: prepared bc056261dafe0ec266f2e8f508e8fe7349e4d0c064853c6fd37ffa6a5e1fab3c 69676e65756d2d6461792f0cfb120000000000 33632.6 program 108.1 dataset 0.0 race 33524.5 variant base class v2 loads/hash 128 cache-fill 2.0 build 30.5 resident 2 programs 1 datasets (prepare answered 35.2 s after it was sent)",
"worker: prepared 1908705502833bdf6c5d08b6d6ae3a364e3fe2bf0956504b20bc5071eecbdfc3 69676e65756d2d6461792f0cfb120000000000 34728.4 program 77.3 dataset 0.0 race 34651.0 variant base class v2 loads/hash 128 cache-fill 2.0 build 30.5 resident 2 programs 1 datasets (prepare answered 0.0 s after it was sent)",
"worker: prepared 4ae011065b84f67924d0c42ff63669e6124b6bddc4a7f71a3f1075b5891efd8c 69676e65756d2d6461792f0cfb120000000000 252.5 program 55.5 dataset 196.9 race 0.0 variant base class v3 loads/hash 128 cache-fill 0.9 build 41.1 resident 2 programs 2 datasets (prepare answered 0.3 s after it was sent)",
"worker: prepared d5964a745dbe2ce2a88d6a4198a244996f1f4aba6cc1bc20841bb532e469fdd6 69676e65756d2d6461792f0cfb120000000000 51.1 program 51.1 dataset 0.0 race 0.0 variant base class v3 loads/hash 128 cache-fill 0.9 build 41.1 resident 2 programs 2 datasets (prepare answered 0.1 s after it was sent)",
"worker: prepared f9d5de9eea7330aeec83c0994b8929e61a2ac2f13dbae6462cceace3b1d6f24e 69676e65756d2d6461792f0cfb120000000000 46.0 program 46.0 dataset 0.0 race 0.0 variant base class v3 loads/hash 128 cache-fill 0.9 build 41.1 resident 2 programs 2 datasets (prepare answered 0.0 s after it was sent)"
],
"prepared_v3": [
"worker: prepared 4ae011065b84f67924d0c42ff63669e6124b6bddc4a7f71a3f1075b5891efd8c 69676e65756d2d6461792f0cfb120000000000 252.5 program 55.5 dataset 196.9 race 0.0 variant base class v3 loads/hash 128 cache-fill 0.9 build 41.1 resident 2 programs 2 datasets (prepare answered 0.3 s after it was sent)",
"worker: prepared d5964a745dbe2ce2a88d6a4198a244996f1f4aba6cc1bc20841bb532e469fdd6 69676e65756d2d6461792f0cfb120000000000 51.1 program 51.1 dataset 0.0 race 0.0 variant base class v3 loads/hash 128 cache-fill 0.9 build 41.1 resident 2 programs 2 datasets (prepare answered 0.1 s after it was sent)",
"worker: prepared f9d5de9eea7330aeec83c0994b8929e61a2ac2f13dbae6462cceace3b1d6f24e 69676e65756d2d6461792f0cfb120000000000 46.0 program 46.0 dataset 0.0 race 0.0 variant base class v3 loads/hash 128 cache-fill 0.9 build 41.1 resident 2 programs 2 datasets (prepare answered 0.0 s after it was sent)"
],
"need": 0,
"mismatch_lines": [
"1791239226.967 PACK OUT OF DATE: the prepared pair b7eda892cff02b39 is not the node's epoch bc056261dafe0ec2; rebuilding the program pack for the worker"
],
"refused": [],
"swaps": [
"SEED CHANGE at daa 60: epoch seed 130e3b7d568db07aad638f67b2ac26e7cf93c51e4f564d45c6724fc4b2d3da65 -> bc056261dafe0ec266f2e8f508e8fe7349e4d0c064853c6fd37ffa6a5e1fab3c, day 1243916 -> 1243916 (epoch 1): the pair was not prepared (unexpected seeds); the worker compiles inline",
"SEED CHANGE at daa 120: epoch seed bc056261dafe0ec266f2e8f508e8fe7349e4d0c064853c6fd37ffa6a5e1fab3c -> 1908705502833bdf6c5d08b6d6ae3a364e3fe2bf0956504b20bc5071eecbdfc3, day 1243916 -> 1243916 (epoch 2): prepare was sent but the worker has not answered prepared yet; it compiles inline",
"SEED CHANGE at daa 180: epoch seed 1908705502833bdf6c5d08b6d6ae3a364e3fe2bf0956504b20bc5071eecbdfc3 -> 4ae011065b84f67924d0c42ff63669e6124b6bddc4a7f71a3f1075b5891efd8c, day 1243916 -> 1243916 (epoch 3): swapped with no pause (prepared 4 s ago, prepare took 252 ms)",
"SEED CHANGE at daa 240: epoch seed 4ae011065b84f67924d0c42ff63669e6124b6bddc4a7f71a3f1075b5891efd8c -> d5964a745dbe2ce2a88d6a4198a244996f1f4aba6cc1bc20841bb532e469fdd6, day 1243916 -> 1243916 (epoch 4): swapped with no pause (prepared 4 s ago, prepare took 51 ms)",
"SEED CHANGE at daa 300: epoch seed d5964a745dbe2ce2a88d6a4198a244996f1f4aba6cc1bc20841bb532e469fdd6 -> f9d5de9eea7330aeec83c0994b8929e61a2ac2f13dbae6462cceace3b1d6f24e, day 1243916 -> 1243916 (epoch 5): swapped with no pause (prepared 6 s ago, prepare took 46 ms)"
],
"accepted_total": 301,
"accepted_after_switch": 124,
"found_lines": 0,
"cpu_recheck_mismatched": 0,
"last_status": ""
}
}

View file

@ -0,0 +1,193 @@
# Counter ASIC 2.0: the node side (program class v3 as a height switch)
5 October 2026, night, worker "ca2-node". Branches: `ca2-v3` (main repository: the igneum-pow seam, the workers, the fast-time gate) and `ca2-v3-node` (the fork, from the 0.3.10 tip 21d4c73c plus pack-loop 05ef0fa3). Plan: `docs/plans/counter-asic-2-rollout.md`; status: `docs/plans/counter-asic-2-status.md`. The shape follows finality v3 (`docs/plans/finality-v3-rollout-devnet.md` section 6): one height switch read from the override file, every node carries the same object before the height.
Everything below is the SEAM. The class itself (`igneum_pow::V3_CLASS`) is a placeholder, w16 (`LoadClass::fixed(4, 16)`), that the integration branch replaces with the decided width, mix and scratch share; the ca2-era draw and the ca2-mixer item construction fill what `Epoch::from_chain_seeds` calls. Nothing here changes a v2 program, a pinned pack or any live node.
## 1. What changed, where
| Piece | What |
|---|---|
| igneum-pow `generator.rs` | `ProgramClass { V2, V3 }`, `V3_CLASS` (placeholder), `GENERATOR_VERSION_V3 = 3`, `generate_from_seed_bytes_program_class(label, seed, class, era)`; `Program::era_bytes`; a v3 program's id is `program_id(3, seed, attempt)` (spec 01 section 1.4.6) |
| igneum-pow `verify.rs` | `Epoch::from_chain_seeds(epoch, day, era, class, label)`, `Epoch::chain_program` (no cache fill), `Epoch::chain_dataset(day, class)` (the one entry the node's day cache goes through; today both classes build the same cache) |
| igneum-pow `emit.rs` | program.h: `IGNEUM_PROGRAM_CLASS "v3"` and `IGNEUM_ERA_SEED_HEX` beside `IGNEUM_GENERATOR 3`; program.json: `program_class`, `era_seed_bytes`; nothing on a v2 pack (`tests/packs.rs` diffs the pinned packs `igneum-genesis-mh` and `igneum-devnet-v4-epoch0`: identical) |
| igneum-pow `packcheck.rs` | `verify_pack_texts_chain` / `verify_pack_dir_chain(dir, epoch, day, want_class, want_era)`: `PackFault::WrongClass` for the wrong class, the wrong era, or a generator other than 2 or 3; `PackIdentity` carries `generator`, `class`, `era_hex` |
| `proto-cuda/nvrtc/packfile.h` | `pf_load` refuses a generator other than 2 or 3 (spec 1.4.5), reads the class (must match the generator) and the era; `pf_pack_class_ok`, `pf_class_token` (the one rule for the `class=` / `era=` tokens) |
| `proto-cuda/nvrtc/worker.cpp`, `proto-opencl/host.c` | a pair's identity includes its class and era when the line names them; right seeds and the wrong class answer `need <e> <d>` and `error <id> pack <dir>: program class mismatch ...`, so the miner prepares the pair from a pack of the right class; a prepare on a wrong-class pack fails in plain words |
| `proto-metal/main.swift` | refuses every `class=v3` line (the Swift generator is version 2; the integration adds 3) |
| fork `consensus/core/src/igneum.rs` | `ProgramClass`, `program_class_for_epoch_at(e, N4, L)`, `program_class_v3_first_epoch_at`, the process-wide activation (`install_program_class_v3_activation`, `program_class_for_epoch`, `program_class_at`), `POW_ERA_BLOCKS`, `POW_ERA_LEAD`, `pow_era_index`, `pow_era_seed_score`; `PowEpochInfo` + `program_class`, `next_program_class`, `program_class_v3_activation_daa`, `era_index`, `era_seed` (serde defaults: v2, never, none) |
| fork `consensus/core/src/config/params.rs` | `program_class_v3_activation_daa` in `Params` and `OverrideParams` (default never on devnet, simnet, mainnet; 0 on the testnet like every other switch), `override_params`, `From<Params>`, the digest (unconditionally, right after `finality_v3_activation_daa`), `Params::install_program_class_v3_activation`, `Params::program_class_v3_first_epoch`; tests `override_params_carry_the_program_class_v3_activation`, the digest test's 11th edit, `fast_time_60x_file_is_the_devnet_at_60x` (every field present) |
| fork `kaspad/src/daemon.rs` | `Program class v3 from the override file: active from epoch E (DAA score N4 rounded up to the epoch boundary at 3600*E, epochs of 3600 DAA)`; the activation installed next to the PoW schedule |
| fork `consensus/pow/src/igneum.rs` | `EpochSeeds { epoch, day, class, era }` (+ `EpochSeeds::v2`), day caches keyed on `(day, class)`, the program through `Epoch::chain_program`, the cache through `Epoch::chain_dataset`, `standalone_epoch` through `Epoch::from_chain_seeds`; test `program_class_v3_seeds_hash_their_own_program_over_their_own_cache` |
| fork `consensus/src/pipeline/header_processor/{processor,pre_ghostdag_validation}.rs` | the class of the header's epoch (`program_class_at(daa)`), `HeaderProcessor::era_seed` (the stand-in, memoised per era), `RuleError::EraSeedUnavailable` |
| fork `consensus/src/processes/pruning_proof/igneum_pow.rs` | `seeds_for`: the class of the epoch; era 0 = genesis; a later era is `PruningImportError::MissingEraSeed` (no era witness in the proof format yet) |
| fork `consensus/src/consensus/mod.rs` | `get_pow_epoch_info`: the class of this and the next epoch, the activation, the era index and seed (the era walk memoised once per era per process) |
| fork `rpc/core/src/model/message.rs`, `rpc/grpc/core/proto/rpc.proto` (fields 12 to 16), `rpc/grpc/core/src/convert/message.rs` | `RpcPowEpochInfo` + `program_class` (generator number), `next_program_class`, `program_class_v3_activation_daa`, `era_index`, `era_seed`; an old node's absent fields read as v2, never, none |
| fork `igneum/miner/src/main.rs` | seeds from the template (`seeds_from_info`), the activation installed from the template, the legacy seed walk keys the class on it; `next_pair` takes the next epoch's class; the job and prepare lines end with `class=v3 era=<hex>` for a v3 epoch; `seeds.txt` carries `program_class` and `era_seed_hex`; `write_pack_checked` checks class and era (`verify_pack_dir_chain`); the "program and 256 MiB cache ready" lines and `export-pack` print the class and the program id |
| `infra/fast-time/override-60x.json`, `tools/finality-attacks/redteam/override-60x-v3.json`, `infra/fast-time/README.md` | the field at never, the README row |
| `infra/fast-time/class-v3.mjs` | the G4 gate (section 5) |
## 2. The epoch-boundary rule
One epoch has one program (spec 01 section 1.12), so the switch keys on the EPOCH: epoch `e` is class v3 when `L * e >= N4` with `L` the live epoch length (`pow_epoch_blocks()`, 3,600 on the devnet, 60 on the fast-time profile). The first v3 epoch is `ceil(N4 / L)`; a height inside an epoch rounds UP to the next boundary and never splits an epoch between two programs. A block's class is a function of its DAA score alone (`program_class_at(daa)`), as its epoch seed is.
| N4 | L | first v3 epoch | first v3 DAA | the epoch before |
|---|---|---|---|---|
| 150 | 60 | 3 | 180 | epoch 2 (DAA 120 to 179) is v2 to its last block |
| 180 | 60 | 3 | 180 | |
| 181 | 60 | 4 | 240 | |
| 136,000 | 3,600 | 38 | 136,800 | epoch 37 (133,200 to 136,799) is v2 |
| 136,800 | 3,600 | 38 | 136,800 | |
| 0 | any | 0 | 0 | (the testnet) |
| never | any | none | | |
The unit tests `program_class_switch_rounds_up_to_the_epoch_boundary` (consensus-core) and `override_params_carry_the_program_class_v3_activation` (params) pin these rows and sweep every activation 0..399 at L = 60.
The digest: the field (and `pow_genesis_dataset_log2`) enters `consensus_digest` unconditionally, so the digest flips the moment a binary carrying the field (at never) runs, exactly as the finality v3 field did. This is intended: every node must carry the object before any node reaches the height, and a node without the field is refused at the handshake (rollout section 1, order step 1). Measured: the devnet digest with no override file moves from `9409dedac4bf9f0f...` (0.3.10) to `c562d70e1428c9789823cc40067623b4767f7c555ce7ff4ea11c1498f013ef6c` (0.3.11; the pinned value of `consensus_digest_keeps_the_0_3_5_value_until_the_fee_switch_is_set`), which is the expected digest of a scratch node at the publish.
## 3. The era stand-in
`E_n` of spec 04 section 4.4, until the 1-hour VDF is in the node (`docs/plans/era-layout.md` section 2):
| Era | `E_n` |
|---|---|
| 0 | the genesis block hash |
| n >= 1 | the hash of the last selected-chain block whose DAA score is below `15,552,000 n - 7,200` |
One function per path: `HeaderProcessor::era_seed` (the header's era from its DAA score, the walk down the selected parents from the header's selected parent, memoised per era in `era_seed_memo`), `Consensus::get_pow_epoch_info` (the same walk from the sink for the template, memoised once per era per process), `ProofSeeds::seeds_for` (era 0 only). The seed block is at least one era lead (7,200 DAA) below any header that uses it, past the merge depth (3,600), so one walk per era per process is sound; the walk itself is up to an era long on the first header of era `n >= 1` (about 15.5 million selected parents), which is why it is memoised and why the VDF should land before era 1 (180 days after genesis). A v2 program never reads the era; the placeholder v3 class does not either (the ca2-era draw will); the era is carried and recorded in the pack so a worker of the wrong era is refused from the first v3 build.
## 4. The job line, the prepare line, the pack
| Surface | Class v2 (today) | Class v3 |
|---|---|---|
| job line | `job <id> <prehash> <target> <start> <count> <epoch> <day>` | the same with ` class=v3 era=<64 hex>` at the end |
| prepare line | `prepare <epoch> <day> [<dir>]` | the same with ` class=v3 era=<64 hex>` at the end |
| `program.h` | `IGNEUM_GENERATOR 2` | `IGNEUM_GENERATOR 3`, `IGNEUM_PROGRAM_CLASS "v3"`, `IGNEUM_ERA_SEED_HEX "<64 hex>"` |
| `program.json` | `"generator": 2` | `"generator": 3`, `"program_class": "v3"`, `"era_seed_bytes": "<hex>"` |
| `seeds.txt` | `epoch_seed_hex`, `day_seed_hex`, `day_index` | plus `program_class v3`, `era_seed_hex <hex>` |
| template `pow_epoch` | `programClass 2` | `programClass 3`, `nextProgramClass`, `programClassV3ActivationDaa`, `eraIndex`, `eraSeed` |
A v3 program's identity is the pair (program id, era seed): the id covers the generator, the seed words and the attempt (spec 01 section 1.4.6, unchanged), so every era of one epoch seed shares one id, and the era seed, carried by the pack (`IGNEUM_ERA_SEED_HEX`) and the job line (`era=`), tells them apart. The workers and `packcheck` compare both.
A v2 line and a v2 pack are byte for byte what the workers read before this branch (the tokens are sent only for a v3 epoch), so a 0.3.10 worker on a 0.3.11 miner mines v2 epochs unchanged and refuses nothing until the switch; by the switch every worker is 0.3.11 (rollout order).
Refusals: `pf_load` refuses a generator that is not 2 or 3 (`error 0 pack <dir>: program pack generator N is not a generator version this worker runs (2 or 3)`, the exit-44 path of 05ef0fa3: the miner rebuilds the pack before the restart). A job of class v3 against a resident v2 pack of the same seeds answers `need <e> <d>` and `error <id> pack <dir>: program class mismatch: this pack is class v2, the job names class v3 (export the pack again)`; the miner's `need` handling prepares the pair again, `write_pack_checked` writes a v3 pack (checked with `verify_pack_dir_chain` before the worker hears of it), and the prepared v3 pair wins over the resident v2 pair because the pair identity now carries the class. The Metal worker answers `error <id> program class v3 is not implemented by this worker` until the Swift generator carries version 3.
## 5. The fast-time gate (rollout G4)
`node infra/fast-time/class-v3.mjs [--secs 420] [--activation 150] [--epochs-after 2]` under `tools/lock/with-lock.sh run`: three nodes on `override-60x.json` merged with `genesis_bits` 0x1f010000 (2^16 hashes per block, the CPU difficulty of `sim/difficulty/testnet_v2.py`) and `program_class_v3_activation_daa` 150 (inside epoch 2, so the rounding rule is exercised: the first v3 epoch is 3 at DAA 180); one real CPU miner per node (`--engine igneum-pow`, 1 thread, real lottery-hash solutions, every node verifying the other two); the run ends two epochs after the boundary.
### Result, run 1 (5 October 2026, 21:33:04Z to 21:38:02Z, Apple M5 Max shared with other agents' builds)
Binaries: fork ca2-v3-node 79bd8e10 and igneum-pow at ca2-v3 66eeba3 (the mixer-x4 class, `V3_CLASS` = MX4, before the era and hot-table fields), `target-ca2/release`, built on the Mac under the build lock. Summary: `docs/plans/counter-asic-2-gate/class-v3-20261005-2133Z-mx4.json`. PASS: every check true.
| Check | Measured |
|---|---|
| Switch line on every node | 3 of 3: `Program class v3 from the override file: active from epoch 3 (DAA score 150 rounded up to the epoch boundary at 180, epochs of 60 DAA)` |
| Digest, all three nodes | `0186df7d0834d054...` (the 60x profile with the CPU bits and the switch) |
| Template class per epoch | epochs 0 to 2 class 2 (`nextProgramClass` 3 from epoch 2), epochs 3 to 5 class 3; the switch seen at DAA 180, 168.0 s wall |
| Blocks before / after the boundary (node 0's DAG) | 181 / 124 (selected chain 176 / 123), 305 in all |
| Program id per epoch, all three miners agreeing | e0 v2 `8f8806638d59850f`, e1 v2 `fd9562df32a68313`, e2 v2 `1ae6d90ab299154c`, e3 v3 `5d0dedd9fd9e29a1`, e4 v3 `e81808dcdb02ce05`, e5 v3 `06aff9c1d33e7a13`; no v2 id reappears under v3 |
| Rejected blocks | miners 0 / 0 / 0 (97, 103, 104 accepted); nodes 0 / 0 / 0 `PoW rejected` lines |
| Forks | sinks `082fd39ba65df2ff` on all three nodes, block counts 304 / 304 / 304, one tip each |
| Program and cache ready, one CPU core | a new `(day, class)` cache: v2 epoch 0 219 to 235 ms, the first v3 epoch 177 to 185 ms (two per miner: the 24-minute day rolled at DAA 190); a program swap inside a day 2 ms |
The mixer x4 build-time number the rollout asks for is not visible here: the CPU miner derives dataset words on demand from the cache (no dataset build), so the x4 cost lands on the GPU workers' dataset build, measured by the ca2-mixer playbooks on the PCs.
### Result, run 2 (5 October 2026, 21:46:36Z to 21:51:31Z): the composed class
Binaries rebuilt on ca2-v3 b105a55 (era layout merged on the mixer: `V3_CLASS` = MX4 with the era drawn inside `generate_from_seed_bytes_program_class`, hot `None`), fork 79bd8e10 unchanged. Summary: `docs/plans/counter-asic-2-gate/class-v3-20261005-2146Z-era-mx4.json`. PASS: every check true.
| Check | Measured |
|---|---|
| Switch lines, digest | 3 of 3, the same line as run 1; digest `0186df7d0834d054...` |
| Template class per epoch | epochs 0 to 2 class 2, 3 to 5 class 3; the switch at DAA 180, 171.0 s wall |
| Blocks before / after the boundary | 181 / 124 (selected chain 180 / 122), 305 in all; 102, 102, 100 accepted per miner |
| Program id per epoch, all three miners agreeing | e0 v2 `8f8806638d59850f`, e1 v2 `bb8dd9ddbf9eb63f`, e2 v2 `8ee7a9f33d418e48`, e3 v3 `2d278041ba482dba`, e4 v3 `2ae786d294a8a59d`, e5 v3 `bc36813df2f41b5f` |
| Rejected blocks | miners 0 / 0 / 0, nodes 0 / 0 / 0 |
| Forks | sinks `712c1b212091dcdc` on all three nodes, block counts 303 / 303 / 303, one tip each |
| Program and cache ready, one CPU core | v2 epoch 0 178 ms, the first v3 epoch 181 ms, an in-day swap 2 ms |
CPU hash rate across the switch, run 2, miner cpu0 (one thread, the Mac shared with other agents' builds, so approximate): the cumulative rate read 0.024 MH/s through the v2 epochs (30 to 151 s), then fell to 0.021 MH/s cumulative by 271 s (100 s under v3), which puts the v3 interval rate near 0.017 MH/s, about 30 percent under v2 on the CPU interpreter (the era's strided windowed loads and the mixer path). The node has no per-block verify timing line; the CPU verifier is measured in `igneum-pow` (the mixer agent, one M5 Max core, ms per 32-lane unit, same minute, cited from ca2-mixer 1ab8b21's message of 5 October 2026 22:10Z):
| Path | readwidth | ca2-v3 88dafbc (before the fix) | ca2-v3 d233fa1 (after) |
|---|---|---|---|
| v2 (the live devnet) | 0.607 / 0.610 | 1.332 | 0.609 / 0.611 |
| v3 at x4 | | | 1.238 |
| v3 at x8 (the class) | | | 2.077 (worst cold 2.15) |
The regression was `memhard::derive_items` at m = 1 (2.2x); the fix dispatches to an out-of-line `derive_items_mask`, one instance per cache size with the line mask a constant. A v3 block costs the node about 3.4x a v2 block to verify (2.08 against 0.61 ms per unit).
Epoch 0's v2 id is the same in both runs (`8f8806638d59850f`: a v2 program is untouched by the era code, on the chain as in the packs); the v3 ids differ from run 1 because the era draw is now inside the class.
### Result, run 3 (5 October 2026, 22:15:04Z to 22:19:44Z): the final class, x8 with the era
Binaries rebuilt on ca2-v3 d233fa1 (mixer x8, the era drawn inside the class, the verifier fix), fork 89dfcb95. Summary: `docs/plans/counter-asic-2-gate/class-v3-20261005-2215Z-era-mx8.json`. PASS: every check true.
| Check | Measured |
|---|---|
| Switch lines, digest | 3 of 3, the same line; digest `0186df7d0834d054...` |
| Template class per epoch | epochs 0 to 2 class 2, 3 to 5 class 3; the switch seen at DAA 181, 127.9 s wall |
| Blocks before / after the boundary | 182 / 122 (selected chain 180 / 120), 304 in all; 101, 89, 113 accepted per miner |
| Program id per epoch, all three miners agreeing | e0 v2 `8f8806638d59850f`, e1 v2 `e145305446a6b4ce`, e2 v2 `743ad2a3cab0518a`, e3 v3 `a6523b90cff501e3`, e4 v3 `bc811b3c4b8b1ced`, e5 v3 `e784541f19cdebe5` |
| Rejected blocks | miners 0 / 0 / 0, nodes 0 / 0 / 0 |
| Forks | sinks `a9ce45df8beeaf13` on all three nodes, block counts 303 / 303 / 303, one tip each |
| Program and cache ready, one CPU core | v2 epoch 0 179 ms, the first v3 epoch 191 ms (x8 construction: the cache fill is unchanged, the items are derived on demand), an in-day swap 2 ms |
Epoch 0's v2 id `8f8806638d59850f` is the same in all three runs. Three runs, three v3 classes (mixer x4; x4 with the era; x8 with the era), the same chain behaviour each time: the switch rounds up to epoch 3, no block rejected, one chain.
### Result, run 4 (5 October 2026, 22:25:22Z to 22:31:28Z): a real Metal miner across the boundary (gate G4b)
The fleet-outage case: the three runs above used CPU miners, so the Metal worker's class v3 path had not mined. Run 4 puts node 0's miner on the Metal worker the way the app does (`igneum-miner --worker igneum-bench --prepare-packs <dir> --exit-on-seed-change`, `igneum-bench` built from ca2-v3 00c55aa: a class v3 program comes from the pack's `program_bound.metal`, its day from the pack's `memhard.metal` (the x8 construction, the era layout), keyed by (day, class, era)), genesis bits 0x1e010000 (2^24 hashes per block), CPU miners on nodes 1 and 2. Summary: `docs/plans/counter-asic-2-gate/class-v3-20261005-2225Z-metal-mx8.json`.
| Check | Measured |
|---|---|
| The chain | 182 / 123 blocks across DAA 180, 0 rejected (miners and nodes), sinks `66a33eca15d33ba2` on all three, 304 / 304 / 304, 3 of 3 switch lines, the switch at DAA 180, 264.5 s wall |
| PREPARE lines | 6 (three class v2, three class v3), each 6 to 7 DAA before its boundary, the v3 ones with the pack directory and `class=v3 era=<hex>` |
| The worker's `prepared` for the v3 packs | three: `prepared <seed> <day> 252.5 program 55.5 dataset 196.9 race 0.0 variant base class v3 loads/hash 128 cache-fill 0.9 build 41.1` (the first, with the day built from the pack: 0.9 ms cache fill, 41.1 ms build on the GPU), then 51.1 and 46.0 ms (the day resident) |
| The swap at the boundary | `SEED CHANGE at daa 180 ... (epoch 3): swapped with no pause (prepared 4 s ago, prepare took 252 ms)`; the same at 240 and 300 |
| Blocks on v3 | 124 accepted after the switch (301 in the run), every one re-checked on the CPU: `mismatched 0`; `need` 0; no class or era mismatch line; no refusal, no exit 42 or 44 |
One line before the switch, on the v2 path: `PACK OUT OF DATE: the prepared pair b7eda892... is not the node's epoch bc056261...; rebuilding the program pack` at DAA 60: the epoch-1 seed reported at `boundary - lead` flipped (the quarter-lead confirm is 3 DAA on the 60x profile, 150 on the devnet) and the miner's rebuild path of 05ef0fa3 wrote the right pack. The v2 prepares answered 33 to 35 s after they were sent (the Metal variant race runs inside the prepare, longer than the 10-DAA lead of the profile; 600 s on the devnet), so epochs 1 and 2 compiled inline; a pack program races nothing and answered in 252 ms. The v3 path is clean; the script now judges the Metal checks from the first v3 prepare on (the run's own report flagged the v2 line and read FAIL on that one check).
The PC 2 suite jobs for the same fork (79bd8e10), the G6 evidence (the coordinator's status file carries the SUMMARY lines):
| Job | Tree | What | Result |
|---|---|---|---|
| `build-20261005-215219` | main b105a55 | the Linux node, the six node suites, the app tests | Linux build and app tests passed; `kaspa-consensus` failed on the known flake (`ban_is_decided_by_the_carrying_block`, `UnexpectedDifficulty` in `mine_on_all`, the 0.3.10 cut's section 11 case) |
| `build-20261005-215712` | main b105a55 | `kaspa-consensus` alone | 97 passed, 0 failed, 3 ignored, `ban_is_decided ...` ok, 21:59:45Z |
| `build-20261005-220351` | main 8ea6740 | the other five crates (`kaspa-consensus-core igneum-exec kaspa-pow igneum-miner kaspa-p2p-flows`) and the app tests | 22:04 to 22:06:55Z: igneum-exec 17, igneum-miner 18, consensus-core and p2p-flows green, the app 78 + 26 + 8; `kaspa-pow` FAILED on `program_class_v3_seeds_hash_their_own_program_over_their_own_cache` (the fork test compared a chain v3 program's class to `V3_CLASS` with `era: None`; on the era-merged crate the class carries the drawn era inside). Fixed on the fork at 89dfcb95 (the class minus its era is `V3_CLASS`, the era is drawn, another era seed keeps the program id and hashes another program over the same day cache). Re-run: job 4 |
| `build-20261005-221237` | main d233fa1, fork 89dfcb95 | the five crates and the app | 22:12:37 to 22:15:44Z (185 s): every stage ok; kaspa-consensus-core 108 (2 ignored), igneum-exec 17, kaspa-pow 33 + 14 (the v3 engine test, with the igneum-pow feature), igneum-miner 18, kaspa-p2p-flows 7, igneum-app 78 + 26 + 8; 0 failed. With job 2, G6 is green |
The PC's test stage builds the fork's test binaries with the plan's feature set, so the v3 engine test DID run on PC 2 (the earlier reading that it was Mac-only is withdrawn): the PC job is the G6 evidence.
## 6. Tests
| Where | What | State |
|---|---|---|
| igneum-pow `cargo test --release` | 42 unit + 11 pack tests, including `program_classes`, `program_class_and_era_are_checked`, the pinned-pack diffs | pass (Mac, 5 Oct 2026) |
| `proto-cuda/nvrtc/emu/packfile-test.sh` | 13 checks: the attempt rule, the generator rule, a v3 pack with its era, the token matcher | pass (Mac) |
| `host.c`, `worker.cpp` (emulation), `main.swift` | syntax / compile | pass (Mac) |
| fork `cargo test -p kaspa-pow --features igneum-pow` | 14 engine tests including `program_class_v3_seeds_hash_their_own_program_over_their_own_cache` (rewritten at 89dfcb95 for the era-in-class rule) | pass (Mac, 89dfcb95 against d233fa1; PC 2 job 4: 33 + 14) |
| fork `cargo test -p kaspa-consensus-core` | 108 + 7: the switch rounding, the era clock, the params, digest and fast-time file tests | pass (Mac, fork 79bd8e10) |
| PC 2 `build-job.mjs` suites | `kaspa-consensus-core igneum-exec kaspa-pow kaspa-consensus igneum-miner` + `igneum-app` | the coordinator publishes on its go |
## 6a. The integration merges for the ship (5 October 2026, 22:35Z to 22:45Z)
| Merge | Commit | Conflicts, and how they were kept |
|---|---|---|
| readwidth 30ff674 (the OpenCL `__local` declaration rule, the read-width decision record, the per-watt rows) | 3566afd | `emit.rs` (2 hunks: the hot-table kernel arguments, HEAD's superset), `packbench.swift` (readwidth's resident-footprint lines added to HEAD's hot-aware RESULT line), `bench-log.md` (both entries), `read-width.md` add/add (readwidth's version, a pure superset of HEAD's) |
| origin/master 1f0d62c (the 0.3.10 merge) | 49c7e78 | `host.c` (one struct hunk: both fields, `dupOf` and `pci[32]`; master's topology, duplicate-platform and read-back code and ca2-v3's class/era tokens, `mh_word`, hot and mixer fields all auto-merged; `cc -fsyntax-only` clean), `bench-log.md` (both entries) |
Checks after the merges (49c7e78, 22:50Z): the igneum-pow suite 53 + 4 + 19 + 7 pass; the packfile test 0 failures; the CUDA emulation `emu/test.sh` PASS (ready + prepare 1, 64 + 64 + 32 found on pack A, pack B prepared with its self-test, the swap, the self-heal rebuild of pack A, 17 sampled hashes equal to `igneum-pow hash-bound`, the NVRTC source check PASS for 6 files); `proto-opencl/test-generic.sh` on Apple OpenCL PASS (the same protocol, job 5 refused after the swap, 15 sampled hashes equal); the fork's `kaspa-pow`, `igneum-miner`, `kaspad` check clean. The emulation needs `IGNEUM_CUDA_INC` pointed at a checkout's `proto-cuda/nvrtc/redist/include` (the worktree has none).
## 7. Unverified, and what is owed
- The era walk for era >= 1 has never run (the devnet is 180 days from era 1); a pruning-proof sync past era 0 fails with `MissingEraSeed` until an era witness exists in the proof format.
- The Metal worker takes class v3 from the pack (program and day) and mined across the boundary in run 4; the app must pass `--prepare-packs` to its Metal worker (the app branch fixes `engine.rs`, which passed it to the non-Metal workers only).
- The GPU workers' class refusal was checked by the C test of `pf_pack_class_ok` and the syntax of both hosts, not by a live worker on a v3 pack: the integration's bit-exact gate (G1) is where a real worker first builds a v3 pack.
- `next_pair` keeps the current era seed for the next epoch; an epoch boundary that is also an era boundary (once per 180 days) would prepare the wrong era, and the job line then names the right one, so the worker refuses the prepared pair and the miner prepares again (one wasted compile, no wrong block). The VDF era seed replaces the stand-in before this matters.
- Done 22:05Z: ca2-cache rebased as 1950661 fast-forwarded (the first attempt, 2de19e5 on 464d6e1, conflicted in 8 files and was aborted); igneum-pow 53 + 19 tests and the packfile test pass; the fork's kaspa-pow, miner and kaspad check clean against the merged crate (seam unchanged). `V3_CLASS = { era: None, hot: None, ..LoadClass::MX4 }`.
- Done 22:12Z: ca2-mixer 16dfd1e (MX8 candidate, tests) merged at 4e733bb with two one-line fixes the merge needed (3a7fba7 `MX8` gets `era: None, hot: None`; 795472e tests/scratch.rs's `Instr` literals get `win: 0, off: 0`); then ca2-mixer 1ab8b21 (the verifier fix, `V3_CLASS = { era: None, hot: None, ..LoadClass::MX8 }`: class v3 is x8 by the coordinator's decision of 22:05Z) merged at d233fa1. Checks on d233fa1: igneum-pow 53 + 4 + 19 + 7, packfile 0 failures, the fork's `cargo test -p kaspa-pow --features igneum-pow` 14 pass, the release rebuild 1 min 57 s.
- Owed (0.3.12, coordinator's ask of 5 October 2026 22:20Z): per-day dataset reuse in the CUDA and OpenCL workers (a `Day` object shared by consecutive pairs, the cache freed after the build); the Metal worker already keys datasets by day. Until then the iGPU tier mines v3 with a dataset rebuild per epoch on those two workers.
- Wire compatibility: `RpcPowEpochInfo` gained five fields in its Borsh form (wRPC) and five proto fields (gRPC); the gRPC side reads an old node's zeros as v2 / never / none; the Borsh form is versioned by `GetBlockTemplateResponse` (version 2 carries the whole struct), so a 0.3.11 wRPC client against a 0.3.10 node reads short: the miner uses gRPC, the console reads JSON (serde defaults).

177
docs/plans/era-layout.md Normal file
View file

@ -0,0 +1,177 @@
# Era layout: table layout and working set drawn per era and per program (Counter ASIC 2.0, layers 4 and 8)
5 October 2026, branch `ca2-era`, worker "ca2-era". Status: Designed and Implemented behind the class flag (`LoadClass::era`, not the lottery hash); nothing here changes the default generator, the pinned packs or any live program. Measured sections are marked as such; everything else is design.
Plan: `docs/plans/counter-asic-2.md`, layers 4 ("table layout drawn per era: item size, stride, interleave") and 8 ("working-set size drawn per program"). The era seed is `E_n` of spec 04 section 4.4. Confirmed on 5 October 2026 by grep over `igneum-pow/src` on branches master, readwidth and opencl-rdna4-telemetry: no era draw existed in code before this branch (`igneum-era`, `EraParams`, `era_seed`: no match).
## 1. What is drawn, and from what
Two streams, nothing else:
| Stream | Seeded from | Draws | Sets |
|---|---|---|---|
| Era stream | `seed_words_from_bytes("igneum-era/" \|\| E_n)` words 0 and 1 (spec 01 section 1.13.1 wrote `n_le64 \|\| E_n`; the index is dropped here because `E_n` already commits to `n` through the VDF input of section 4.4 step 2, and the node's seam hands the generator the era bytes alone: `Epoch::from_chain_seeds(epoch, day, era, class, label)`, branch ca2-v3) | 7 per era, fixed | the load width `W`, the stride `(M, R)`, the interleave `pos[0..3]` |
| Program stream | the epoch seed words as today (spec 01 section 1.3.3) | 11 per instruction instead of 9: the 9 of version 2, then 2 window draws (a class whose loads are not version 2's takes the read-width width roll between them, ca2-v3's rule) | per load site: the window shrink `k_off` and its offset `o` |
The era parameters change the class of every program of the era. The program stream does not see the era parameters (two eras with the same epoch seed draw the same instruction list and the same windows, and differ in width, stride and layout); this is what makes the six era packs below a controlled comparison.
### 1.1 Era draw (proposed spec text for section 1.13.1, replacing its parameter table)
One SplitMix64 stream `S` seeded with `lo = words[0] | (words[1] << 32)` of `seed_words_from_bytes("igneum-era/" || E_n)`. Seven draws, in this order, whether or not a value is used:
1. `W = allowed[below(|allowed|)]`: the width in words of every dataset load of the era, drawn from the genesis-fixed ascending set `allowed`, a subset of {1, 4, 16} (4, 16 or 64 bytes). A set of one element pins the width; the draw is still consumed. The set is `{1}` (4 bytes, v2's load): the read-width decision of 5 October 2026 (`docs/plans/read-width.md`, "keep v2; w16 the only width that passes the rules and closes nothing") and the adoption rule of the same evening (a draw that changes the bytes per hash changes the rate; the six-era hash-rate spread must stay under 5 percent per card). The set is recorded in every pack (`IGNEUM_ERA_ALLOWED_WIDTHS`, `program.json` `era.allowed_widths`) and enters the program id. The code keeps the draw general so the set can be widened at genesis without a new derivation.
2. `M = low32(next()) OR 1`: the stride multiplier, odd, so `x -> x * M` is a bijection on 32-bit words.
3. `R = 1 + below(31)`: the stride rotation, in 1..31 (never 0: spec 01 section 1.14 item 3).
4. to 7. `r_i = next()` for `i` in 0..3: the interleave draws. Let `b = log2(W)` (0, 2 or 4) and `free = 4 - b`. Let `c = [b, b + 1, ..., 15]` (16 - b candidates). For `i` in `0..free`: `j = i + (r_i mod (16 - b - i))`, swap `c[i]` and `c[j]`. The interleave is `pos = [0, ..., b - 1] ++ sort(c[0..free])`, four ascending bit positions in 0..15. Draws `r_free..r_3` are consumed and ignored.
The era parameters are `(W, M, R, pos)`. Era 0 of the devnet packs is listed in section 5.
### 1.2 Dataset mapping with the interleave (replaces the last sentence of section 1.8.5)
Word `w` of the dataset holds word `j(w)` of item `t(w)`, where `j(w)` is the 4-bit number whose bit `i` is bit `pos[i]` of `w`, and `t(w)` is `w` with bits `pos[0..3]` removed (the remaining bits in order). With `pos = [0, 1, 2, 3]` this is today's `dataset[w] = item(w >> 4)[w AND 15]`, byte for byte.
Properties kept:
- An item has the same value at every dataset size (item derivation is untouched), and because every `pos[i] < 16`, `dataset[w]` is the same at every dataset size of at least 2^16 words. The 1 GiB vectors of an era remain valid for words below 2^28 at any larger size, as today.
- The low `b = log2(W)` positions are 0..b-1, so the `W` words of one aligned load lie in one item (`t` is the same for all of them and `j` runs `j0 .. j0 + W - 1`). The verifier derives one item per lane per load, as today: the 4,096-item bound of section 1.11 holds (16 loads x 8 iterations x 32 lanes, whatever the width).
- The dataset build writes 16 words of one item to 16 addresses `w(t, j)` (a scatter of 4-byte writes instead of one 64-byte line when `pos != [0, 1, 2, 3]`). This is the only GPU cost of the interleave and is paid once per day; section 6 measures it.
What the interleave does and does not buy. A chip that hard-wires today's layout (64-byte items, a 64-byte line per item) reads the wrong 15 words with every word once the era draws another layout; the layout changes every 180 days inside rules fixed at genesis. A chip whose address decoder can permute 28 address lines under firmware control pays nothing for it. The honest claim is the first sentence only. The stride below is the same kind of lever: two integer operations per load on a chip, nothing on a GPU.
### 1.3 Load address (replaces "a load reads one 4-byte word at `src AND MASK`" in section 1.5 for era programs)
For a load site with window draws `(k_off, o)` (section 1.4), a dataset of `2^D` words (`MASK = 2^D - 1`), and register value `x`:
```
k = min(k_off, D - 26) (0 when D <= 26)
y = rotl(x * M, R)
idx = ((y AND (MASK >> k)) OR ((o AND (2^k - 1)) << (D - k))) AND MASK
base = idx AND NOT (W - 1) (W words from base are folded as verify::fold_words)
```
Uniformity: `x * M` with `M` odd and `rotl` are bijections of the 32-bit word, so `y` is uniform when `x` is; `y AND (MASK >> k)` is uniform on the window; the offset picks which of the `2^k` aligned windows. Branch-free, integer only, three operations before the mask (multiply, rotate, and-or) against one today. The emitted text has one form per dialect, checkable by text search (section 1.14 item 2): CUDA and OpenCL `ds[((rotl_imm(rN * 0x........u, Ru) & 0x........u) | 0x........u) & mask]`, Metal the same with `dataset[` and `& MASK]`.
### 1.4 Window draw per load site (layer 8; proposed text for section 1.4.3)
After the nine draws of version 2 (and the width roll, for a class whose loads are not version 2's), every instruction takes two more draws, used only on a load slot:
```
k_off = below(3) window = the dataset, a half or a quarter of it (2^(D - k_off) words)
o = low32(next()) AND (2^k_off - 1) which aligned window
```
Bounds: the window never goes below `2^26` words (256 MiB; `k = min(k_off, D - 26)` in 1.3), which exceeds the largest on-chip cache of any card in the benchmark (the RTX 5090's 96 MiB L2, the RX 9070 XT's 64 MB Infinity Cache, vendor figures), and never above the dataset. At the prototype dataset (2^28) the windows are 1 GiB, 512 MiB and 256 MiB; at the genesis dataset (2^29) 2 GiB, 1 GiB and 512 MiB. The dataset grows by the step schedule recommended to the project lead (spec 01 section 1.13.3 option (b), `docs/analysis/card-lifetime-2026-10-05.md`: power-of-two steps, 4 GiB at year 4, 8 GiB at year 12, 16 GiB at year 28, 32 GiB at year 60, every index `AND MASK`), so the window ceiling follows the steps and the floor stays the genesis constant 2^26 words; nothing in the address of 1.3 needs a range reduction. A program has 16 load sites and so up to 16 windows; the set a program reads is their union (section 7 computes its distribution). The verifier bound is unchanged (1.2).
Why per load site and not per program: a per-program window of a quarter of the dataset would hand a 256 MiB SRAM mirror a third of the hours at the prototype size. Sixteen sites with drawn offsets cover the dataset with high probability (section 7), so the mirror a chip would need is the whole dataset in every hour, and the hour-to-hour variation lands on the memory design (which quarter, which half, how many distinct windows), not on its size.
### 1.5 Program stream (replaces "592 draws per program" in section 1.4.3 for era programs)
16 slot draws, then 64 x 11 = 704: 720 draws per program (768 and 784 for a class with the width roll). On the chain an era program is a class v3 program (branch ca2-v3's seam: `ProgramClass::V3`, generator version 3): its id is `program_id(3, seed words, attempt)` as that branch defines it, and the era it was drawn under is identified beside the id by `IGNEUM_ERA_SEED_HEX` (`packcheck::verify_pack_dir_chain` refuses a pack whose era is not the job's), so the pair (id, era seed) names the program. The experiment classes that are not class v3 (`--class <other> --era ...`) carry the era inside the class id instead: `FNV-1a-64("igneum-program-rw/" || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots || "era/" || allowed[3] || W || M_le32 || R || pos[4])`. The class name is the base name with `-era<first stream word as hex>` (`w4-era401998a5`).
### 1.5.1 The seam (branch ca2-v3, kept as its signatures stand)
`V3_CLASS` is the base class of class v3 (16 loads of 4 bytes, the lottery hash's load; the era rides inside it), `V3_ALLOWED = [1]` its width set. `generate_from_seed_bytes_program_class(label, seed, V3, Some(era))` draws `LoadClass::era(V3_CLASS, era, &V3_ALLOWED)` and stamps generator 3; without era bytes (a template before the era is known) the bare `V3_CLASS` stands. `Epoch::chain_dataset(day, class)` is unchanged: the layout of 1.2 is a property of the program (`program.class.layout()`), applied by the interpreter and by `Epoch::dataset_word`, so the day's cache is shared by every era of a day and the engine keys its caches on `(day, class)` as before.
### 1.6 Acceptance
The rule of section 1.4.6 is unchanged in its tests. Its interpreter mirrors 1.3 at the rule's constant `D = 28` (`idx` as above with `MASK = 0x0fffffff`), as `igneum-pow/src/accept.rs` does. The distinct-address bound counts dataset loads as the read-width branch defines it.
## 2. The era seed on the devnet (stand-in for `E_n`)
Until the 1-hour VDF of section 4.4 is in the node, in the shape of the epoch seed's stand-in (`docs/fork-divergence.md` "Epoch seed"):
| Era | `E_n` |
|---|---|
| 0 | the genesis block hash (32 bytes) |
| n >= 1 | the hash of the last selected-chain block whose DAA score is below `15,552,000 n - 7,200` (the 2-hour lead of section 4.4 step 1) |
The era of a block is `floor(DAA score / 15,552,000)`, a function of the header alone. `E_n` for `n >= 1` is known 7,200 DAA seconds before the era starts, which covers the 1-hour VDF when it arrives and the kernel compile and dataset rebuild now. Test seeds for packs and tests: `E_n` = the 32 bytes (little-endian words) of `seed_words_from_bytes("igneum-era-test/<n>")` (`igneum-pow ... --era igneum-era-test/<n>`, the number only names the pack); raw bytes with `--era <n>:<64 hex>`.
## 3. Memory budget (the project lead, 5 October 2026: under 6 GB on an 8 GB card)
The era layout adds no resident memory: a window is a mask and an offset in the kernel text, the interleave is address arithmetic, the stride is two operations. The whole working set on a card, every item from this branch and the others:
| Item | Bytes | Who |
|---|---|---|
| Dataset (prototype) | 1 GiB | existing |
| Cache (256 MiB, resident only while the day's dataset is built, then free; the decided reading, confirmed by the coordinator on 5 October 2026, `docs/analysis/card-lifetime-2026-10-05.md` carries the per-tier working set with the cache freed as the best case) | 256 MiB peak | existing |
| Layer 5 hot table | the ca2-cache worker's figure | not this branch |
| Scratch per resident warp (read-width variant 5, 32 or 128 KiB per warp) | 2,048 warps x 128 KiB = 256 MiB at most | readwidth branch |
| Output and read-back buffers | 2^24 nonces x 8 B = 128 MiB per dispatch | existing harness |
| Era windows, stride, interleave | 0 | this branch |
The era window never exceeds the dataset, so it never grows the footprint.
## 4. Implementation (behind the flag)
| Piece | Where | What |
|---|---|---|
| `EraParams`, `era_draw`, `LoadClass::era` | `igneum-pow/src/generator.rs` | the 7-draw era stream of 1.1; the class carries `era: Option<EraParams>` beside `mix`, `load_slots`, `scratch`, `scratch_kb`; the two window draws per instruction (`Instr::win`, `Instr::off`); the program id of 1.5 |
| `Layout` | `igneum-pow/src/memhard.rs` | `split(w) -> (t, j)`, `join(t, j) -> w`, `LINEAR = [0, 1, 2, 3]`; `MemhardCpu` and `DatasetSource` carry it |
| `load_index` | `igneum-pow/src/verify.rs` | the address of 1.3, shared by the interpreter and the acceptance mirror |
| Emitters | `igneum-pow/src/emit.rs` | the one load form of 1.3 in Metal, CUDA and OpenCL C; `mh_word`, the three `igneum_build` kernels and the Metal build kernel with `mh_t`, `mh_j`, `mh_addr` when the layout is not linear; `IGNEUM_ERA_*` in `program.h`, an `"era"` object in `program.json` |
| CLI | `igneum-pow/src/main.rs` | `--era <igneum-era-test/n \| n:hex>` and `--era-widths 4,16,64` (one width pins) on every command |
| Tests | `igneum-pow/src/*.rs`, `igneum-pow/tests/packs.rs` | the draw is deterministic and within bounds; `split`/`join` are inverse and the dataset is a prefix at every size; six era programs pass the generator contract and the acceptance rule; vectors round-trip; the pinned packs are byte-identical; the six era packs match the emitters and every load has the form of 1.3 |
The default class is untouched: `LoadClass::V2` has `era: None`, every emitter branch on `era` keeps today's text, and `tests/packs.rs` diffs the pinned packs (`igneum-genesis-mh`, `igneum-devnet-v4-epoch0`) against the emitters as before. Section 6 records the diff of a fresh export against the checked-in files.
## 5. The six era packs
`proto-cuda/packs-ca2-era/era-<n>`, `n` in 0..5: the devnet's 32-byte epoch seed `edc4fa84...fb07` and day bytes `igneum-day/20730` (the seeds of the pinned pack `igneum-devnet-v4-epoch0`, which is the v2 baseline with the same program seed), dataset 2^28 words, era seed `igneum-era-test/<n>`, width pinned at 4 bytes (`--era-widths 4`, the default). The program seed is held fixed so that the six packs differ in the era parameters only (section 1); each carries `seeds.txt` for the one-click workers. Every pack is attempt 1 (attempt 0 of this seed is rejected under the era class: 12 draws per instruction give a different stream from v2's). The drawn parameters (`igneum-pow show --epoch-hex edc4... --era igneum-era-test/<n>`, 5 October 2026):
| Pack | Class | Era seed `E_n` (first 16 hex) | W (bytes) | M | R | pos |
|---|---|---|---|---|---|---|
| era-0 | w4-erab2ed8a89 | 5e0587f455a86e91 | 4 | 0x625e5ab3 | 19 | 0, 2, 10, 15 |
| era-1 | w4-era676a17fc | df57136f2ad5f410 | 4 | 0xb2a9d70d | 6 | 1, 3, 8, 13 |
| era-2 | w4-era843155d7 | 7f450623297a954f | 4 | 0x2b4a5b97 | 28 | 1, 3, 4, 8 |
| era-3 | w4-erad6367bfe | 8bffdd3366b9c3ff | 4 | 0x27ea7eff | 30 | 2, 3, 8, 13 |
| era-4 | w4-era4488f3ed | e593fc1d48475c88 | 4 | 0x4d38603d | 10 | 2, 9, 13, 15 |
| era-5 | w4-eraf897c84e | ff87ad96a1b53f36 | 4 | 0x03ac37ad | 22 | 0, 2, 10, 13 |
All six are class v3 packs (generator 3, `IGNEUM_PROGRAM_CLASS "v3"`, `IGNEUM_ERA_SEED_HEX`), attempt 0, program id `73bcbfe8ccf988f1` in every pack (the seam's `program_id(3, seed, attempt)`; the era seed beside it names the program), the era layout over version 2's item construction (mixer x1, the genesis cache) so that the v2 baseline pack is the same dataset; the chain's class v3 composes the same draw over `LoadClass::MX4` (mixer x4, growth), and the integration re-exports these packs on it after the PC rows. The era-seed-to-pack assignment above is from `program.h` of each pack; the test-seed numbering is only the pack name.
The windows are a property of the program, so they are the same in all six packs (site:shrink:offset): `1:2:3 4:2:2 6:0:0 10:1:0 12:1:1 20:2:0 27:1:0 30:0:0 35:0:0 40:0:0 41:0:0 43:2:3 45:1:0 52:0:0 53:2:2 54:0:0`: 7 sites read the whole dataset, 5 a half, 4 a quarter; the union is the whole dataset.
## 6. Measurements
Pending at the time of this commit; each table below says the machine, the date, the harness and the command when filled.
### 6.1 Byte-identical default path
### 6.2 Bit-exactness on the Mac (Metal, Apple OpenCL, CUDA emulation)
### 6.3 Hash rate per era on the M5 Max (Metal) and the CPU verifier
### 6.4 PCs (RTX 5090 CUDA, RX 9070 XT OpenCL): prepared, waiting for the go
## 7. The chip-model line per draw
From `docs/analysis/m16-recompute-attacker-2026-10-05.md` and the random-read ceilings of `docs/bench-log.md` ("the 9070 XT on the eGPU": 9070 XT 2.42 to 2.68 G loads/s at 1 GiB, 5090 16.4 to 18.0, M5 Max 3.41 to 3.49; every random 4-byte read costs AMD a 64-byte line):
| Quantity | Formula |
|---|---|
| Bytes read per hash | 128 loads x W bytes |
| Distinct 64-byte lines per hash | 128 (one line per load at every W up to 64 bytes; the census's 120 to 128 distinct addresses per hash) |
| SRAM a chip needs to mirror what the hash reads | the union of the program's 16 windows (a distribution over programs; section 7.1) |
| Latency-bound share | measured rate / (the card's 4-byte random-read ceiling / 128) |
The latency-bound share uses the 4-byte ceiling for every width because the 9070 XT line probe showed the same count per second for 4-byte and 64-byte random reads; the 5090's 16-byte and 64-byte ceilings are the read-width branch's measurement, cited when they land.
### 7.1 Union of windows per program
Filled from a CPU census over programs (section 6).
## 7.2 Found on the way (harness defects, both fixed on this branch)
| Where | Defect | Fix | Checked |
|---|---|---|---|
| `proto-cuda/nvrtc/packfile.h` (the one-click workers) | the seed words were re-derived from the bare epoch seed, so every pack of attempt 1 or higher was refused ("the epoch seed bytes do not give the pack's IGNEUM_SEEDW_INIT"); 5.14 percent of epochs under v2, all six era packs, and the epoch 34 fleet outage of 18:23Z on 5 October 2026 (branch pack-loop af983a7, which this branch takes: `pf_program_words`) | the pack-loop derivation merged over the readwidth packfile (class fields and string seeds kept) | the devnet pack (attempt 0) loads, the six era packs (attempt 1) load, a copy of era-0 with the attempt tampered to 0 is refused on the re-derivation (section 6) |
| `proto-cuda/host.cu`, `proto-opencl/host.c` | the host-side dataset word was `mh_item(w >> 4)[w AND 15]`, the harness's own copy of the linear layout; under an interleaved layout the "64 random points vs host derivation" check failed while the Mac samples and the vectors passed | `host_ds_word` calls the pack's `mh_word` (memhard.h), which carries the layout | `proto-cuda/emu/test-layout.sh`: the CUDA emulation on era-1 (interleaved) and the devnet pack (linear) must pass the random-point check; it failed on era-1, era-3 and era-5 before the fix (section 6) |
## 8. What is unverified
- Everything in section 6 marked pending.
- The 1-hour VDF does not exist; the devnet stand-in of section 2 is a proposal.
- The interleave's value against a chip with a programmable address decoder is nil (1.2); the claim is limited to hard-wired layouts.
- The window floor of 2^26 words is set by the 5090's L2 (96 MiB) and the 9070 XT's Infinity Cache (64 MB, vendor figures); a future card with a larger cache moves the floor, which is a genesis constant.
- No cryptanalysis of the stride (a multiply and a rotate before the mask); it is a bijection, so the address distribution is that of the register value, as today.

266
docs/plans/hot-table.md Normal file
View file

@ -0,0 +1,266 @@
# Hot table: a second table sized to GPU cache, read beside the 1 GiB dataset
Counter ASIC 2.0, layer 5 (`docs/plans/counter-asic-2.md`). Experiment branch `ca2-cache`, 5 October 2026 (night), on top of the read-width branch (`readwidth` b970dda: `LoadClass`, `verify::fold_words`, the scratch variant, `proto-metal/packbench.swift`, `proto-opencl/host.c --bench-pack`). Nothing here is the lottery hash: every hot class sits behind the generator flag and the default v2 path is byte-identical (the pinned packs `igneum-genesis-mh` and `igneum-devnet-v4-epoch0` are diffed by `igneum-pow/tests/packs.rs`).
## 1. The idea
The honest hash is 128 dependent random 4-byte reads over 1 GiB (spec 01 section 1.5). A chip that wants a gain on that must beat a GPU at DRAM random reads, which is the same physics for both (the plan's last section). What a chip can do that a GPU cannot is choose its memory: a chip can mirror read-only data into SRAM and serve it at SRAM latency, if the data fits.
The hot table turns that around. A second table `H` of `S` MiB (32, 64 or 96 in this experiment) is derived from the epoch seed and read by `k` of the 16 load slots (2, 4 or 8). `S` is chosen to fit the caches of the cards that mine: 96 MiB L2 on the RTX 5090, 64 MB Infinity Cache plus 8 MB L2 on the RX 9070 XT, the system level cache on the M5 Max (figures approximate, from memory; the probe of section 6 measures what each card does at each size). A GPU gets the hot loads as cache hits for free. A chip must carry `S` MiB of SRAM for the same hits, beside the DRAM path it still needs for the other `16 - k` slots, or serve `H` from DRAM and fall behind the GPU by the hot share. Either way the chip pays for something the GPU already has.
What this does not change: the cold loads (the 1 GiB dataset) stay a dependent chain at DRAM latency, so the latency-bound property holds for them; the hot loads are interleaved in the same chain (every load's address is a fresh register of the same iteration, G2), so a hot hit shortens the chain by one DRAM latency and nothing else.
## 2. The specification text (proposed; prototype values, to be fixed at gate 1)
### 2.1 Hot key and fill
For an epoch whose program seed bytes are `e` (spec 01 section 1.12: the UTF-8 of a seed string in the packs, the 32-byte VDF output on the chain; the bytes before the attempt suffix of 1.4.6, so every attempt of one epoch shares one table):
```
KH = seed_words_from_bytes("igneum-hot/" || e) 8 words
```
`H` has `N_H = S x 2^18` words (`S` MiB) in `N_seg = S x 256` segments of 64 chained lines of 16 words, filled exactly as the cache of section 1.8.3 with `KH` in place of `K` and the tag `("Igne", "umHT") = (0x49676e65, 0x756d4854)` in place of `("Igne", "umMH")`:
```
prev = 0^16
for j in 0..63:
in = prev XOR (sigma[0..3] || KH[0..7] || s || j || tag[0..1])
line = B(in) ChaCha12 with feed-forward, section 1.8.2
H[segment s, line j] = line
prev = line
```
One GPU thread per segment, as the cache fill. The fill is a once-per-epoch cost: `S / 256` of the 256 MiB cache fill (section 1.8.3 table: 0.6 to 2.1 ms on the M5 Max GPU, 0.67 ms on the RTX 5090, 175 to 190 ms on one CPU core for 256 MiB), so under 1 ms on a GPU and 22 to 71 ms on one core, measured below.
### 2.2 The hot load
A hot load slot reads `H` instead of the dataset with the same fold (width 1: a plain XOR):
```
hot: dst = dst XOR H[mulhi(src, N_H)]
```
`mulhi(a, b)` is the high 32 bits of the 64-bit product (the `mulhi` family of section 1.4.1, bit-exact on Metal, CUDA and OpenCL). The index lies in `[0, N_H)` for any `N_H`, which is what lets `S = 96` exist: 96 MiB is not a power of two, so `src AND MASK` cannot address it. This is the multiply-shift range reduction section 1.13.3 proposes for the growing dataset, so the hot table is also its first measured use. For `S = 32` and `64` the mapping takes the top 23 or 24 bits of `src` where the dataset load takes the low 28; a fresh source is uniform, so neither choice costs uniformity (section 6, hot-load uniformity test).
The emitted text has one form per dialect, checkable by text search as the mask check of section 1.14 item 2: `hot[mulhi(rN, HOT_WORDS)]` (Metal), `hot[__umulhi(rN, HOT_WORDS)]` (CUDA), `hot[mul_hi(rN, HOT_WORDS)]` (OpenCL), with `HOT_WORDS` a literal of the pack. A hot pack's hash kernels carry exactly `16 - k` masked dataset loads and exactly `k` hot loads (`igneum-pow/tests/packs.rs`).
### 2.3 Which slots are hot
The 16 load slots are drawn first by the partial Fisher-Yates of section 1.4.3, which emits them in a uniformly random order. The first `k` slots in that draw order are the hot slots. Drawn, not fixed, because a fixed pattern (every fourth load, say) would let a pipeline schedule its SRAM reads statically for every hour; drawn costs no extra draw, so a `hot(S, k)` program takes the version 2 stream exactly and is the version 2 program with `k` of its loads redirected (the width roll of the read-width classes is not taken: `LoadClass::takes_width_roll`). A hot slot keeps every rule of a load: the fresh-source draw (G2), the injecting-write count (acceptance (b)), the lane-constant test and the distinct-address count (acceptance (c)); its address is tagged apart from dataset addresses in the count so a hot word and a dataset word at the same index are two addresses.
The class composes with the read-width fields of `LoadClass` (width mix, scratch `k` and `kb`): the hot slots are taken from the drawn slots after the scratch slots, so a class may carry a width mix, a scratch and a hot table at once. The experiment packs below use the version 2 widths and no scratch.
Two forms (coordinator's decision, 5 October 2026, after the first Mac measurement):
| Form | Load slots drawn | Dataset loads per hash | Hot loads per hash | Items per warp (verifier bound) | Class name | What a chip pays |
|---|---|---|---|---|---|---|
| replaced (`hot(S, k)`) | 16, the first `k` in draw order hot | `(16 - k) x 8` | `8k` | `(16 - k) x 256` | `hot64k4` | the on-die-cache recompute chip of `docs/analysis/scratch-soundness.md` 3.4 (the 256 MiB cache in SRAM, items derived on the fly, 333 MH/s at 50 T op/s, 2.4x the 5090) GAINS: a hot load replaces an item derivation (1,170 ops) with an SRAM read, so at `k = 4` its rate rises 1.33x against the GPU's measured 1.05 to 1.22x |
| added (`hot(S, k, added)`) | `16 + k`, the first `k` in draw order hot | `16 x 8 = 128` | `8k` | 4,096, unchanged | `hot64k4a` | `S` MiB of SRAM and `k` reads per iteration for nothing: the 16 item derivations stay; the GPU pays `k` cache hits |
The added form is the one that taxes the named chip; the replaced form stays as the measured record (section 6). In the added form the slot draw is a partial Fisher-Yates of `16 + k` slots over 1..63, so the program stream differs from version 2 (another slot count), and the width roll is still not taken (`takes_width_roll`: the slot count is 16 plus the class's own added hot slots). `loads_per_hash` is `128 + 8k`; the acceptance rule's distinct-address bound scales with it as for the read-width classes.
Program id: `FNV-1a 64 over "igneum-program-rw/" || generator || seed words || attempt || mix || load_slots || "hot/" || S || k [|| "added"]` (the read-width id with the hot fields appended), so no hot pack can be mistaken for a version 2 one, for another hot class or for the other form.
### 2.4 Acceptance (section 1.4.6)
The dynamic test stays a pure function of the program. A hot load reads `dataset_elem(mulhi(src, N_H), SW[2], SW[3])` (the six-operation closed form keyed by seed words 2 and 3, where the dataset stand-in is keyed by words 0 and 1), so the two stand-ins are two tables without any cache or day. Every other test is unchanged; the distinct-address bound is the read-width rule's (dataset and hot loads counted, scratch read-modify-writes not).
### 2.5 The verifier (section 1.11)
A verifier holds, per day, the mixer parameters and the 256 MiB cache, and, per epoch, `H` (`S` MiB, filled on one core in the times of section 6). A hot load is one table read per lane; a dataset load is the item derivation of section 1.11 as before. The verifier never holds the dataset. Per epoch the verifier's memory is `256 MiB + S MiB`.
## 3. Memory budget (the project lead's cap, 5 October 2026: the working set on a card stays under 6 GB on an 8 GB card)
Decided by the coordinator on 5 October 2026 after the card-lifetime review (`docs/analysis/card-lifetime-2026-10-05.md`, branch `card-lifetime` 1fecfe2, merged into `ca2-coord`): the GPU frees the 256 MiB cache after the daily dataset build (the hash never reads it; the rebuild costs 0.67 ms of fill and 13.4 ms of build on the 5090, bench-log 3 October 2026), so the cache is not resident and the working set is dataset + hot table + scratch (0 in v3) + buffers. The table below is that review's per-tier table with the hot table at its largest (96 MiB), buffers 128 MiB (the harnesses' 2^24-nonce output), resident warps = SMs x 48 on NVIDIA Ampere and later (SM counts approximate, from memory; the GTX 1650 is Turing at 32 warps per SM), 2,048 launched warps on Apple (the Metal harness). The scratch columns are the read-width variant's 32 and 128 KiB per resident warp, kept for the record; v3 carries no scratch.
| Tier | Card assumed (SMs, approximate) | Resident warps | Scratch at 32 KiB | Scratch at 128 KiB | Hot table | Dataset at genesis | Buffers | Total, no scratch | Total at 32 KiB | Total at 128 KiB | Years the dataset leaves under mapping (b), cache freed |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 4 GB | GTX 1650 (14 SMs x 32) | 448 | 14 MiB | 56 MiB | 96 MiB | 2,048 MiB | 128 MiB | 2,272 MiB | 2,286 MiB | 2,328 MiB | to year 4 (the 4 GiB step) |
| 8 GB | RTX 3050 (20) | 960 | 30 | 120 | 96 | 2,048 | 128 | 2,272 | 2,302 | 2,392 | to year 12 (the 8 GiB step) |
| 12 GB | RTX 3060 (28) | 1,344 | 42 | 168 | 96 | 2,048 | 128 | 2,272 | 2,314 | 2,440 | to year 28 (the 16 GiB step; year 12 if the cache were resident) |
| 16 GB | RTX 5060 Ti (36) | 1,728 | 54 | 216 | 96 | 2,048 | 128 | 2,272 | 2,326 | 2,488 | to year 28 |
| 24 GB | RTX 4090 (128) | 6,144 | 192 | 768 | 96 | 2,048 | 128 | 2,272 | 2,464 | 3,040 | to year 60 (the 32 GiB step) |
| 32 GB | RTX 5090 (170) | 8,160 | 255 | 1,020 | 96 | 2,048 | 128 | 2,272 | 2,527 | 3,292 | to year 60 |
| Apple 8 to 64 GB | M-series, 2,048 launched | 2,048 | 64 | 256 | 96 | 2,048 | 128 | 2,272 | 2,336 | 2,528 | 8 GB to year 4, 16 GB to year 12, 32 GB to year 28, 64 GB to year 60 (50% of unified memory usable, the review's assumption) |
The prototype packs here use the 1 GiB dataset of spec 1.5 (1,024 MiB less in every total). Every total is under 6 GB with the hot table at its largest, so `S` is not what the cap binds: the dataset's growth is, and the hot table takes 96 MiB of the room at every tier (about 2 months of the 0.5 GiB-a-year schedule). The verifier's memory is `256 MiB + S MiB` (section 2.5); the GPU's is the table above.
## 4. The chip model with the SRAM it would need
Figures: SRAM area per bit from `docs/analysis/m16-recompute-attacker-2026-10-05.md` section 3 (256 MiB in about 100 to 300 mm^2 at a current node: the low end from a 0.02 um^2 bit cell with array overhead, the high end from wafer-scale parts at about 1 MB per mm^2; approximate, from memory) and from `docs/plans/counter-asic-2.md` (256 MB in about 45 mm^2 at a leading node, approximate). That is 0.18, 0.39 and 1.17 mm^2 per MiB. Die areas: a 750 mm^2 GPU-class die (the M16 model's equal-silicon comparison) and a 100 mm^2 memory-chip die (an assumption for a latency-bound chip whose die holds memory controllers and little else; labelled as such).
| S (MiB) | SRAM at 0.18 mm^2/MiB | at 0.39 | at 1.17 | Share of a 100 mm^2 die (0.39) | Share of a 750 mm^2 die (0.39) |
|---|---|---|---|---|---|
| 32 | 6 mm^2 | 12 mm^2 | 37 mm^2 | 11% | 1.6% |
| 64 | 12 mm^2 | 25 mm^2 | 75 mm^2 | 20% | 3.2% |
| 96 | 17 mm^2 | 37 mm^2 | 112 mm^2 | 27% | 4.8% |
The gain arithmetic. Let `G0` be a chip's gain over a GPU on the version 2 hash (the plan's public claim: under 2x; the plan's last section says the bound is DRAM latency, the same physics on both). Let `g` be the GPU's own measured speed-up from the hot class over version 2 on the same card (section 6: `hash rate hot(S, k) / hash rate v2`). The ideal `g` with every hot load a hit at zero cost is `16 / (16 - k)`: 1.14x at `k = 2`, 1.33x at `k = 4`, 2.0x at `k = 8`; the measured `g` says what share of that a real cache delivers while the 1 GiB dataset streams through the same cache.
| Chip | Hot loads served from | Gain after the hot table | Arithmetic |
|---|---|---|---|
| A, no SRAM for H | DRAM | `G0 / g` | the chip's rate is what it was; the honest GPU gained `g` |
| B, S MiB of SRAM for H | SRAM | `G0 x A_die / (A_die + A_S)` per unit of silicon | the chip regains `g` and pays `A_S` on top of its die |
| B on a 100 mm^2 die, S = 64, 0.39 mm^2/MiB | SRAM | `0.80 x G0` | 100 / 125 |
| B on a 750 mm^2 die, S = 96, 0.39 mm^2/MiB | SRAM | `0.95 x G0` | 750 / 787 |
What the big-table loads still cost the chip: `(16 - k) x 8` dependent DRAM reads per hash at the card's loaded latency (the 9070 XT entry's probe: 451 ns on the 5090 at 4,096 lanes, 276 ns unloaded on the 9070 XT, 1,949 ns on the M5 Max at 4,096 lanes; section 6 repeats the probe here). At `k = 4` that is 96 reads per hash; to match one RTX 5090 at its honest 229 Mhash/s (the M16 model's reference) a chip must keep `229 M x 96 x 300 ns = about 6,600` DRAM reads in flight at a 300 ns latency, and 11,000 at 500 ns, whatever its arithmetic. That queue depth is a memory-controller property, which is the latency-bound argument of the plan restated for the cold share.
What the hot table does to the recompute attacker of M16: nothing good. `H` is a plain ChaCha12 chain, so a word of it can be recomputed from `KH` at `j + 1` block evaluations (32.5 on average, about 40,000 integer operations per hot load, approximate, against 1,170 per dataset item), which is 35x the cost of recomputing a dataset word. A chip stores `H` or reads it from DRAM; it does not recompute it.
Reading, before the measurements: the hot table costs a chip `A_S` of die area or `g` of rate. Which one binds depends on the measured `g`, which is why the measurement comes first. If a GPU's cache delivers most of the ideal `g` with the dataset streaming beside it, `k = 8` at `S = 64` doubles the honest rate and halves chip A. If the cache delivers little, the hot table is a cost to the verifier (`S` MiB per epoch) with no gain, and the layer is dropped.
## 5. Implementation (branch `ca2-cache`)
| Where | What |
|---|---|
| `igneum-pow/src/generator.rs` | `LoadClass { hot: Option<HotClass> }` beside `mix`, `load_slots`, `scratch`, `scratch_kb`; `HotClass { mb, k }`; names `hot32k4`, `hot64k2`; `Op::Hot`; the first `k` drawn load slots after the scratch slots are hot; no width roll for a hot class with version 2 widths (`takes_width_roll`); the program id carries `hot/S/k` |
| `igneum-pow/src/memhard.rs` | `HotTable` (key, S, words), `hot_key(seed_bytes)`, `hot_words(mb)`, `hot_segments(mb)`, `hot_index(src, words)`, the tagged segment fill shared with the cache, `HOT_TAG` |
| `igneum-pow/src/verify.rs` | `DatasetSource::hot: Option<HotTable>`; `Op::Hot` in `step`; `Epoch::new_class` and `from_seed_bytes_class` fill `H` from the program's seed bytes when the class has a hot table |
| `igneum-pow/src/accept.rs` | `Op::Hot` with the closed-form stand-in of 2.4 |
| `igneum-pow/src/emit.rs` | the hot load statement in the three dialects; `hot` as the buffer after the init words (Metal buffer 3, or 4 when bound; CUDA and OpenCL argument after `mask`, or after the init words when bound) and before the scratch triple; `HOT_WORDS` literal; `ht_cache_segment` and `igneum_hot_fill` kernels in memhard.h, kernel.cu, kernel.cl and memhard.metal; `program.h` `IGNEUM_HOT_MB`, `IGNEUM_HOT_WORDS`, `IGNEUM_HOT_SEGMENTS`, `IGNEUM_HOT_SLOTS`, `IGNEUM_HOT_KEY_INIT`; `vectors.h` and `vectors.json` the hot head, last line and FNV-1a 64 |
| `igneum-pow/src/main.rs` | `--class hot<S>k<k>` on every command (`bench` reports the fill time and the verifier ms per warp) |
| `igneum-pow/tests/packs.rs` | the five hot packs pinned (program, vectors, every emitted file, the load-form count: `16 - k` masked loads and `k` hot loads) |
| `proto-cuda/packs-ca2-hot/` | replaced: `hot32k4`, `hot64k4`, `hot96k4`, `hot64k2`, `hot64k8`; added: `hot32k4a`, `hot64k4a`, `hot96k4a`; all from seed `igneum-genesis`, day `2026-10-03` |
| `proto-cuda/nvrtc/packfile.h` | `hotMb`, `hotWords`, `hotSegments`, `hotSlots`; the hot self-test values; `pf_selftest` checks them when the pack carries them |
| `proto-opencl/host.c` | `--bench-pack` and the self-test allocate and fill `H` from the pack's `igneum_hot_fill`, check its head, last line and FNV-1a 64, pass it as the argument after the init words |
| `proto-metal/packbench.swift` | the same on Metal from `memhard.metal` |
| `proto-cuda/nvrtc/worker.cpp` | `--check` and `--bench` (new: the base kernel timed over `--batches` dispatches, one RESULT line) allocate, fill and self-test `H`; the launch passes it |
## 6. Measurements
Every row names the machine, the harness and the command; the bench-log entry of 5 October 2026 ("the hot table on the M5 Max") carries the raw lines. The Mac rows were taken on 5 October 2026, 20:19 to 20:21 UTC, under the measure lock, with the Mac's load average at 14 to 27 from other agents' CPU work (the lock serialises builds and measurements, not every process), so they are ordered, repeatable to within a few percent against each other, and not the Mac's quiet numbers. The PC rows wait for the coordinator's go.
### 6.1 Random-read probe at the hot sizes
`igneum-bench-cl-igneum-genesis-mh --memprobe --probe-mib S` (`proto-opencl/host.c`; dependent random 4-byte loads, best of 3, 256 steps per lane, work-group 256; the ceiling is the 4,194,304-lane row; Apple OpenCL reports wall time).
| Card | 32 MiB ceiling | 64 MiB | 96 MiB | 1024 MiB | Ratio 32 / 1024 | 64 / 1024 | 96 / 1024 | ns per dependent load at 4,096 lanes (32, 64, 96, 1024 MiB) |
|---|---|---|---|---|---|---|---|---|
| M5 Max (Apple OpenCL) | 21.7 G loads/s | 12.8 | 12.3 | 3.50 | 6.2 | 3.7 | 3.5 | 1,168; 1,129; 1,242; 1,844 |
| RTX 5090 (CUDA worker `--memprobe`, PC 1, job run-ca2-hot-5090-20261005, card off in the app) | 112.6 | 112.6 | 112.6 | 17.6 | 6.4 | 6.4 | 6.4 | 320; 340; 336; 610 |
| RX 9070 XT (OpenCL worker `--memprobe --device 1`, PC 1, job run-ca2-hot-9070-20261005, card off in the app) | 9.88 (10.8 at 262,144 lanes) | 9.47 | 8.18 | 2.43 | 4.1 | 3.9 | 3.4 | 396; 457; 454; 1,579 |
Reading, RTX 5090: all three sizes sit inside the 96 MiB L2 at one ceiling (112.6 G loads/s, 6.4x the DRAM figure), so the probe alone promises a full hit rate for every S. RX 9070 XT: 32 and 64 MiB inside the Infinity Cache at 9.5 to 10.8 G loads/s (3.9 to 4.1x), 96 MiB at 8.2 (3.4x), as the 9070 XT entry's 64 MiB row said. Streams: 5090 1,563 GB/s at 1024 MiB, 9070 XT 633 GB/s (both at their rated figures, the cards were not parked).
Reading, M5 Max: the step from 32 to 64 MiB halves the ceiling (21.7 to 12.8 G loads/s) and 96 MiB sits with 64, so 32 MiB is inside a cache level that 64 MiB is not (the M5 Max's system level cache size is not published; approximate reading: the 32 MiB table fits, the two larger ones mostly do not and run at a 3.5 to 3.7x advantage over DRAM from whatever hits they get). The coalesced stream grows with the buffer (139, 199, 243, 522 GB/s) because the small buffers are read once from cold.
Predicted `g` from the probe alone, with a hot load costing `1 / ceiling_S` and a cold load `1 / ceiling_1024`: `g = 1 / ((16 - k) / 16 + (k / 16) x ceiling_1024 / ceiling_S)`.
### 6.2 Bit-exactness and hash rate per hot pack
Metal: `proto-metal/packbench --pack <dir> --batches 5 --batch-log2 24` (GPU time). Apple OpenCL: `igneum-bench-cl-igneum-genesis-mh --bench-pack --pack <dir> --batches 5 --batch-log2 24` (wall time). Both harnesses fill the hot table on the device from the pack's `igneum_hot_fill` and check its head, last line and FNV-1a 64 against vectors.json or vectors.h; vectors are the 96 lanes of the Rust reference; the fingerprint is FNV-1a 64 over the 2^24 outputs at base nonce 0.
| Pack | Vectors (Metal, OpenCL) | Hot table FNV (Metal, OpenCL) | Fingerprint 2^24 (both harnesses equal) | Metal Mhash/s | Apple OpenCL Mhash/s | g against v2 (Metal) | Predicted g from the probe | Ideal g |
|---|---|---|---|---|---|---|---|---|
| v2 (igneum-genesis-mh) | 3/3, 96/96 | none | 25f96e7dce90bd4e | 27.68 | 27.61 | 1 | 1 | 1 |
| hot32k4 | 3/3, 96/96 | PASS, PASS | d2e6cf3b61d0b9fe | 33.90 | 33.92 | 1.22 | 1.27 | 1.33 |
| hot64k4 | 3/3, 96/96 | PASS, PASS | e4c5263ac650cc0d | 30.93 | 30.89 | 1.12 | 1.22 | 1.33 |
| hot96k4 | 3/3, 96/96 | PASS, PASS | 5d63439b6e394521 | 29.06 | 28.97 | 1.05 | 1.22 | 1.33 |
| hot64k2 | 3/3, 96/96 | PASS, PASS | 352633bdbbb0d2b6 | 27.67 | 27.27 | 1.00 | 1.10 | 1.14 |
| hot64k8 | 3/3, 96/96 | PASS, PASS | da54630d7dfaaf85 | 47.42 | 46.73 | 1.71 | 1.57 | 2.0 |
| hot32k4a (added) | 3/3, 96/96 | PASS, PASS | 8a3414735db4523c | 25.76 | 25.72 | 0.93 | 0.96 | 1 |
| hot64k4a (added) | 3/3, 96/96 | PASS, PASS | 45668f34105f6307 | 23.92 | 23.87 | 0.87 | 0.94 | 1 |
| hot96k4a (added) | 3/3, 96/96 | PASS, PASS | af763997dfee4c82 | 22.92 | 22.88 | 0.83 | 0.93 | 1 |
For the added form the ideal `g` is 1 (the 16 dataset loads stay) and the probe predicts `g = 1 / (1 + (k / 16) x ceiling_1024 / ceiling_S)`: 0.96 at 32 MiB, 0.94 at 64 and 96 MiB for `k = 4` on the M5 Max; what matters is how far below 1 the GPU lands (its cost of the layer) against the chip's `S` MiB of SRAM and `k` reads.
The PCs (5 October 2026, 21:29 to 21:35 UTC, PC 1 ae432dc7 on app 0.3.9 before and after, the card under test switched off in the app through `api/cards` and restored; CUDA worker `--bench --batches 5 --batch-log2 24 --block-warps 1` on the RTX 5090, wall time around the stream sync; OpenCL worker `--bench-pack --batches 5 --batch-log2 24 --device 1` on the RX 9070 XT, device event time, work-group 256; the v2 references are the readwidth entry's same-night, same-worker numbers: 136.1 and 18.15 MH/s). Every pack bit-exact with the Mac's fingerprint, hot table head, last line and FNV PASS, 96/96 lanes, on both cards.
| Pack | RTX 5090 MH/s | g (v2 136.1) | probe-predicted g (5090) | RX 9070 XT MH/s | g (v2 18.15) | probe-predicted g (9070 XT) | ideal g |
|---|---|---|---|---|---|---|---|
| hot32k4 | 146.6 | 1.08 | 1.27 | 19.79 | 1.09 | 1.23 | 1.33 |
| hot64k4 | 140.8 | 1.03 | 1.27 | 18.73 | 1.03 | 1.23 | 1.33 |
| hot96k4 | 138.5 | 1.02 | 1.27 | 18.33 | 1.01 | 1.21 | 1.33 |
| hot64k2 | 137.5 | 1.01 | 1.13 | 18.17 | 1.00 | 1.11 | 1.14 |
| hot64k8 | 163.6 | 1.20 | 1.73 | 22.32 | 1.23 | 1.59 | 2.0 |
| hot32k4a (added) | 118.7 | 0.87 | 0.96 | 15.27 | 0.84 | 0.94 | 1 |
| hot64k4a (added) | 115.4 | 0.85 | 0.96 | 14.62 | 0.81 | 0.94 | 1 |
| hot96k4a (added) | 114.4 | 0.84 | 0.96 | 14.56 | 0.80 | 0.93 | 1 |
The 5090 at `--block-warps 8` (6 blocks per SM, 8,160 resident warps) is within 0.7% of every row above; the 9070 XT at work-group 32 within 0.5%. Hot fill on the 5090: 1.2 to 2.5 ms (wall, driver API); on the 9070 XT 1.6 to 5.3 ms.
Reading, the PCs. The probe promised a full hit rate for every S on the 5090 (all three tables inside the 96 MiB L2 at one ceiling) and 3.4 to 4.1x on the 9070 XT, and the hash delivered a fraction of it: 1.02 to 1.08x at `k = 4` against the probe's 1.27x and the ideal 1.33x on the 5090, 1.01 to 1.09x on the 9070 XT, 1.20 to 1.23x at `k = 8` against 1.73 and 2.0x. The table that stands alone in the probe does not stand up with the 1 GiB dataset streaming through the same cache: the dataset's random lines evict it (the 5090's L2 and the 9070 XT's Infinity Cache are shared by every load; neither card partitions them). The added form costs the 5090 13 to 16% and the 9070 XT 16 to 20% of its rate for four extra loads per iteration, three to five times the probe's 4 to 7%. The three cards agree on the shape; the 5090's larger cache buys it nothing over the Mac at 32 MiB (1.08 against 1.22).
Hot table fill on the M5 Max GPU (Metal, GPU time): 0.07 ms at 32 MiB, 0.15 ms at 64 MiB, 0.22 ms at 96 MiB.
Reading, M5 Max. Bit-exactness holds: Metal and Apple OpenCL give one fingerprint per pack and every vector lane against the Rust reference, hot table included. On the rate, the hot loads are worth 92% of the probe's prediction at 32 MiB (1.22 against 1.27), 50% at 64 MiB (1.12 against 1.22 is 0.12 of 0.22) and 23% at 96 MiB; at `k = 8` the measured 1.71 is above the prediction (1.57), which says the cold loads also go faster when half of them are gone (fewer dependent DRAM reads in the chain per hash). `hot64k2` gained nothing on this Mac at load (27.67 against 27.68). So on Apple silicon, with the 1 GiB dataset streaming through the same cache, the table that fits (32 MiB) delivers most of its ideal and the larger ones lose most of theirs. The 5090 and 9070 XT, whose caches are the plan's targets, are the PC job.
### 6.3 CPU verifier, one core, avg of 50 warps
`igneum-pow bench --seed igneum-genesis --day 2026-10-03 --class <class> --warps 50` (release build, one M5 Max core, the Mac at load as above; the v2 row from the same session, the 0.604 ms of the brief was an earlier quiet run).
| Class | Hot fill, one core | Items derived per warp | ms per warp | Against v2 |
|---|---|---|---|---|
| v2 | none | 4,096 | 0.626 | 1 |
| hot32k4 | 24.0 ms | 3,072 | 0.489 | 0.78 |
| hot64k4 | 46.4 ms | 3,072 | 0.504 | 0.81 |
| hot96k4 | 73.0 ms | 3,072 | 0.488 | 0.78 |
| hot64k2 | 45.5 ms | 3,584 | 0.560 | 0.89 |
| hot64k8 | 47.7 ms | 2,048 | 0.344 | 0.55 |
| hot32k4a (added) | 21.7 ms | 4,096 | 0.631 | 1.05 (v2 in the same session 0.602) |
| hot64k4a (added) | 43.3 ms | 4,096 | 0.609 | 1.01 |
| hot96k4a (added) | 64.9 ms | 4,096 | 0.614 | 1.02 |
Reading: the verifier gets cheaper with `k`, because a hot load is one table read where a dataset load is an item derivation (9 mixers and 8 cache reads); the per-epoch cost is the fill, 24 to 73 ms on one core, against a 3,600 s epoch and the 20-minute seed lead. The 10 ms gate (spec 1.16) is unaffected.
### 6.4 The chip model with the measured g (M5 Max; the PC rows will replace it)
Chip A (no SRAM for `H`): gain after = `G0 / g`. Chip B (`H` in SRAM): `G0 x A_die / (A_die + A_S)` at 0.39 mm^2 per MiB, 100 mm^2 die.
| Class | g (M5 Max) | Chip A, gain after as a share of G0 | Chip B, share of G0 at 100 mm^2 | Chip B at 750 mm^2 |
|---|---|---|---|---|
| hot32k4 | 1.22 | 0.82 | 0.89 | 0.98 |
| hot64k4 | 1.12 | 0.89 | 0.80 | 0.97 |
| hot96k4 | 1.05 | 0.95 | 0.73 | 0.95 |
| hot64k8 | 1.71 | 0.58 | 0.80 | 0.97 |
The added form (second Mac session, 21:03 to 21:19 UTC, load average 7 to 14; v2 in that session 27.63 Metal, 27.59 OpenCL): the GPU keeps 0.93 / 0.87 / 0.83 of its rate at 32 / 64 / 96 MiB (`k = 4`), against the probe's 0.96 / 0.94 / 0.93, so the hot hits cost this card more than the probe says as the table grows, and 96 MiB is not resident. The named chip's position under the added form, same assumptions:
| Chip | Hot loads served from | Rate against its own version 2 rate | Gain after, as a share of G0 (S = 32 / 64 / 96) | Arithmetic |
|---|---|---|---|---|
| A, no SRAM for H | DRAM | at most 16 / 20 = 0.80 (20 dependent DRAM loads per iteration in place of 16) | 0.86 / 0.92 / 0.96 | `0.80 / g` |
| B, S MiB of SRAM for H, 100 mm^2 die, 0.39 mm^2/MiB | SRAM (the k reads near free) | 1.00 | 0.96 / 0.92 / 0.88 | `(1 / g) x A_die / (A_die + A_S)` |
| B on a 750 mm^2 die | SRAM | 1.00 | 1.06 / 1.12 / 1.15 | the SRAM is 1.6 to 4.8% of the die, the GPU's loss is larger |
With the PCs' `g` (added form, `k = 4`): RTX 5090 0.87 / 0.85 / 0.84, RX 9070 XT 0.84 / 0.81 / 0.80 at 32 / 64 / 96 MiB.
| Chip, added form, S = 32 / 64 / 96 MiB | Gain after as a share of G0, 5090 g | 9070 XT g |
|---|---|---|
| A, H from DRAM (rate 0.80 of its own) | 0.92 / 0.94 / 0.95 | 0.95 / 0.99 / 1.00 |
| B, S MiB of SRAM, 100 mm^2 die, 0.39 mm^2/MiB | 1.03 / 0.94 / 0.87 | 1.06 / 0.99 / 0.91 |
| B, 750 mm^2 die | 1.13 / 1.14 / 1.17 | 1.17 / 1.20 / 1.22 |
Reading: on the cards the plan named, the added form taxes no chip. A chip that serves H from its DRAM loses at most 8% of its gain on the 5090's figures and nothing on the 9070 XT's; a chip with the SRAM comes out ahead on every row except the smallest die at 64 and 96 MiB. The replaced form helps the recompute chip outright (section 2.3). The hot hits are not free on a GPU while the 1 GiB dataset streams through the same cache, and the layer's premise (section 1) needed them to be.
**Recommendation (owner: the project lead, gate 1):** do not adopt layer 5 in either form on these measurements. What could change it: a GPU-side way to keep H resident (cache partitioning or persisting-access controls exist on NVIDIA, approximate, from memory, and are a driver setting, not a consensus rule), or a dataset access pattern that bypasses the cache; both are outside the hash and were not measured. The experiment stays behind the flag with its packs and vectors as the record.
Reading, replaced form: on the M5 Max figures the chip's cheaper way out of `hot64k8` is the SRAM (0.80 of `G0` at a 100 mm^2 die) rather than serving `H` from DRAM (0.58). The layer's value is then the SRAM area, which is small on a large die (0.97 at 750 mm^2). The strongest configuration on this card is the one whose table fits the cache and whose `k` is large; whether 64 MiB fits the 5090's L2 and the 9070 XT's Infinity Cache with the dataset streaming beside it is the PC measurement.
## 7. Decision and what would make it pay (the Counter ASIC 3.0 note)
Decided by the coordinator on 5 October 2026 after the PC rows: layer 5 is out of v3. The rule was a GPU cost under 3% (`g` above 0.97) for a chip cost worth having; measured `g` 0.84 to 0.87 on the RTX 5090 and 0.80 to 0.84 on the RX 9070 XT for the added form, and the replaced form helps the recompute chip. No re-run.
What would make a hot table pay, for a later round:
| Condition | What the measurements say | What would have to be shown |
|---|---|---|
| A table that stays resident beside a streaming 1 GiB | The probe's ceiling at every S inside the cache (5090: 112.6 G loads/s at 32, 64 and 96 MiB) and the hash's 1.02 to 1.08x say the dataset's random lines evict it; 32 MiB on the M5 Max kept 92% of its probe gain, the 5090 kept 29% | A size small enough to survive: the probe cannot say which (it has no competing stream); a sweep of `S` down from 32 MiB (16, 8, 4, 2) with the hash itself, on the 5090 first, would. A 2 to 4 MiB table costs a chip under 2 mm^2 of SRAM, so the layer would then tax nothing; the point of the layer was a table a chip cannot afford, and a table a GPU keeps is one a chip affords |
| The `k` and `S` the probe says | `g` moved with `k` (1.20x at `k = 8`, 1.02 to 1.08 at `k = 4`, 1.00 at `k = 2`) and barely with `S`; the probe predicted 1.27 to 1.73 | Only a resident table makes `k` worth raising; at `k = 8` with a resident table the replaced form reaches 2x (ideal) and the added form costs 0 by the probe. Without residency, `k` buys rate on the replaced form (which helps the chip) and costs rate on the added form |
| A different access shape | Both forms read H at one random word per hot load, the dataset's own pattern, and share the cache with 128 dataset reads per hash | A shape that touches the table in cache-line units and few lines per hash (one 64-byte line per iteration, say) would need a smaller resident set per hash and could be measured with the read-width emitters (`wide_load_stmt`) pointed at H. Or a hot table whose lines are read in a fixed order per epoch (a stream, not a random read), which a GPU prefetches and a chip must still hold or fetch; not designed here |
| Cache partitioning on the GPU | NVIDIA exposes persisting L2 access controls to CUDA programs (approximate, from memory); AMD and Apple do not expose an equivalent to OpenCL or Metal | A consensus rule cannot depend on a driver feature of one vendor; it could be a miner-side optimisation if the layer were in, which it is not |
What the round leaves in place: the code behind the flag (`LoadClass::hot`, both forms, the hosts filling H on the device from the epoch seed, the eight packs and their vectors), the probe at the hot sizes on three cards, and the measured rule that a read-only table a GPU cache could hold is not one it does hold while the dataset streams.
## 8. What is unverified
Listed here until measured, and carried into the bench-log entry.
1. Measured on all three cards (6.2): no card keeps `H` resident enough to deliver the probe's hit rate while the dataset streams. Not measured: whether a driver-side cache partition (NVIDIA's persisting L2 access controls, approximate, from memory) would; it is not something a consensus rule can rely on.
2. The PC rows are one run each (jobs run-ca2-hot-5090-20261005 and run-ca2-hot-9070-20261005, 5 October 2026); the two launch shapes per card agree within 1%, and no re-run was taken (PC 1 was handed on).
3. The resident warp counts of section 3 are approximate; the occupancy the hosts reach is printed by each harness (`kernel:` lines) and should replace them.
4. The SRAM area figures are approximate (section 4 cites their sources); no chip was priced.
5. `S = 96` uses the multiply-shift mapping; the spec's dataset still uses `AND MASK`. If the layer goes in, gate 1 decides whether the dataset mapping follows (section 1.13.3) or `S` stays a power of two.
6. The hot table on the chain needs the epoch seed about 1 ms (GPU) or up to 71 ms (CPU) before the epoch starts; the 20-minute lead of section 4.3 covers it. Not exercised on a node.
7. The Mac numbers were taken at load average 14 to 27 (other agents' CPU work); the ratios between packs of one session are the result, the absolute Mhash/s are not quiet numbers.

418
docs/plans/mixer-x4.md Normal file
View file

@ -0,0 +1,418 @@
# Mixer x4 and the cache growth rule: the class v3 dataset construction
5 October 2026 (night). Counter ASIC 2.0, layer 6 (option C) and ledger M16's lever, decided by the coordinator
under the project lead's delegation at 22:00 UTC (`docs/plans/counter-asic-2-status.md`, "22:00 decided"; the project lead confirms for
the public testnet genesis). Branch `ca2-mixer`. Worker: ca2-mixer (cryptographer's lane).
What this changes, in one line: under program class v3 every mixer application of the dataset item derivation
becomes four applications with distinct round keys, the eight dependent cache reads per item stay eight, and the
256 MiB cache doubles on the days the dataset doubles (years 4 and 12). Version 2 is byte-identical: the two
pinned packs re-export without a changed byte (section 5).
## 1. Why this form
The recompute attacker of `docs/analysis/m16-recompute-attacker-2026-10-05.md` holds the 256 MiB cache on a die
and derives every dataset word instead of reading it: 128 items per hash at about 1,170 integer operations and 8
dependent cache reads each. Its cost is linear in operations per item; the honest miner pays the mixer once a day
in the dataset build and never per hash; the verifier pays it per item it checks. The multiplier `m` is the one
parameter that moves the attacker and leaves the honest hash rate untouched.
Two shapes give the attacker 4x the operations:
| Shape | Mixer applications per item | Dependent cache reads per item | What grows for the verifier | What grows for the chip | What grows for the honest build |
|---|---|---|---|---|---|
| A, chosen: `m = 4` applications per round, 8 rounds | 36 | 8 | the ALU part only; the latency part (8 dependent misses per item) unchanged | integer operations 4x; SRAM bandwidth unchanged (1,024 reads per hash) | 4x the mixer arithmetic, same reads |
| B, alternative: 32 rounds of one application and one read | 33 | 32 | both parts: 4x the dependent misses per item, so about 4x the latency-bound time (spec 1.11: about 8 x 100 ns per item in series without interleaving) | integer operations 3.7x and SRAM bandwidth 4x (4,096 reads per hash, 60 TB/s to match one 5090 at the M16 rate) | 4x the reads too; the GPU build becomes latency-bound at 4x the dependent line fetches |
Shape A is chosen because the verifier's latency part is the part the 10 ms gate protects (section 1.11: the
distinct items of a unit are derived with their chains interleaved so the 8 misses of each item overlap across up
to 32 items; shape B would make that 32 misses deep). Shape B is implemented nowhere; the rule in the brief: if
shape A's measured verifier time exceeds 4.8 ms per warp on one M5 Max core, measure both and recommend. Section 6
has the measurement; it is under that bound, so B stays unimplemented.
## 2. Spec text (replaces 1.8.5 and 1.13.3 under class v3; v2 text unchanged)
### 1.8.5 Item derivation and dataset mapping
Item `t` (16 words) under mixer multiplier `m` (`m = 1` for program class v2, `m = 4` for class v3; a class
parameter, `LoadClass::mixer_mult`):
```
s[i] = K[i] for i in 0..7
s[8 + i] = t * MUL[i] + RC[i] for i in 0..7
for r in 0..7:
for j in 0..m-1:
s = M(s, rk = (r * m + j + 1) * 0x9E3779B9)
a = s[0] AND (2^(C - 4) - 1) cache line index, 2^(C - 4) lines of a 2^C-word cache
s[i] = s[i] XOR cache[line a][i] for i in 0..15
for j in 0..m-1:
s = M(s, rk = (8 * m + j + 1) * 0x9E3779B9)
item(t) = s
```
`M(s, rk)` is the mixer of 1.8.4 with round key `rk`; under `m = 1` the keys are `(r + 1) * 0x9E3779B9` and
`9 * 0x9E3779B9`, the version 2 text exactly. The round keys of the `9 m` applications are the first `9 m` values
of the version 2 key sequence, all distinct (the sequence is `k * 0x9E3779B9` for `k = 1 .. 9 m`, and
`0x9E3779B9` is odd, so no two of the first 2^32 keys coincide). Eight dependent cache reads per item at every
`m` (`ITEM_ROUNDS = 8`, prototype value): the address of read `r` depends on every earlier read. `9 m` mixer
applications, 36 under class v3, about 4,700 integer operations per item (130 per application, 1.8.4).
`dataset[w] = item(w >> 4)[w AND 15]`. A dataset of 2^D words is the prefix of items `0 .. 2^(D-4) - 1`, so an
item has the same value at every dataset size; and a cache of 2^C words is the prefix of segments of every
larger cache (1.8.3 fills segments independently of the cache size), but an item's value depends on `C` through
the line mask, so the item changes on the day the cache doubles.
Source: `igneum-pow/src/memhard.rs` (`derive_items`, `round_key_mult`, `Shape`), the emitted `mh_item` of
`memhard.h`, `memhard.metal` and `kernel.cl` (`emit.rs`, `emit_memhard_core`: the `m` loop is emitted only for
`m > 1`, so every version 2 pack keeps its text).
### 1.13.3 Dataset growth (class v3: option (b) with the cache tied to it, "option C")
Designed: 2 GiB at genesis plus 0.5 GiB per year. The linear schedule in bytes, `G x (1 + 86,400 d /
(4 x 31,536,000)) = G x (1 + d / 1,460)` for the genesis size `G` and the chain day `d` (DAA days since genesis,
section 1.12), doubles at day 1,460 (year 4), quadruples at day 4,380 (year 12), reaches 8x at day 10,220
(year 28). Rule (Designed, decided 5 October 2026 for class v3):
```
doublings(d) = floor(log2(1 + d / 1460)) integer division, then integer log2
dataset_words(d) = 2^(D_0 + doublings(d)) D_0 = 29 designed (2 GiB), 28 on the devnet (1 GiB); capped at 32
cache_words(d) = 2^(26 + doublings(d)) 256 MiB, 512 MiB from year 4, 1 GiB from year 12
```
Power-of-two sizes only (option (b)), so every load keeps the `src AND MASK` form of 1.14 item 2 and the cache
line index keeps `s[0] AND mask`. The cache doubles exactly when the dataset doubles ("option C",
`docs/analysis/sram-mirror.md` section 7): the cache's job is to stay above any GPU's last-level cache and that
needs growth; the recompute attacker is priced by the mixer, not by the cache (section 7 below).
`d` is `day_index(header.timestamp) - day_index(genesis.timestamp)` with `day_index = timestamp_ms / 86,400,000`
(`bind::day_index`, the interim day rule), clamped at 0 (`memhard::days_since_genesis`). Under class v2 nothing
grows: the cache is 2^26 words and the dataset the genesis size on every day.
| Chain day `d` | Years | `doublings` | Cache words | Cache | Dataset words (devnet `D_0 = 28`) | Dataset (designed `D_0 = 29`) | Verifier cache fill, one M5 Max core (measured at 256 MiB, section 6, scaled linearly) |
|---|---|---|---|---|---|---|---|
| 0 to 1,459 | 0 to 4 | 0 | 2^26 | 256 MiB | 2^28 (1 GiB) | 2 GiB | 0.18 s |
| 1,460 to 4,379 | 4 to 12 | 1 | 2^27 | 512 MiB | 2^29 (2 GiB) | 4 GiB | 0.36 s |
| 4,380 to 10,219 | 12 to 28 | 2 | 2^28 | 1 GiB | 2^30 (4 GiB) | 8 GiB | 0.72 s |
| 10,220 to 21,899 | 28 to 60 | 3 | 2^29 | 2 GiB | 2^31 (8 GiB) | 16 GiB | 1.4 s |
| 21,900 and on | 60 and on | 4 | 2^30 | 4 GiB | 2^32 (16 GiB, the index cap) | 2^32 words, the cap | 2.9 s |
Test: `memhard::tests::growth_schedule_table` pins every row and the day before each step. The devnet pack
`igneum-devnet-v4-epoch0` is day 20,730 of the Unix count against genesis day 20,729, `d = 1`, so every existing
size and vector stands.
Consequences for the tiers (the rule of 5 October): a verifier (any node, any pool core) holds 512 MiB from year 4
and 1 GiB from year 12, and fills it once a day in under a second on one 2026 core (the table); a miner's card
holds the dataset, 4 GiB from year 4 and 8 GiB from year 12 on the designed schedule, so an 8 GB card mines until
year 12 and a 16 GB card until year 28 (the cache is not in the card's working set at hash time: it is built,
the dataset built from it, and dropped). Those dates are the design document's own schedule restated as steps;
option (a) would have faded a 4 GiB card out in year 4 instead of year 4.
## 3. Interfaces
| Item | Where | Note |
|---|---|---|
| `LoadClass { mixer_mult: u8, growth: bool }`, `LoadClass::MX4` ("mx4"), `with_mixer(m, growth)`, `v2_loads()`, `takes_width_roll()` | `generator.rs` | a class with v2 loads takes no width roll: its program stream is version 2's draw for draw, so the v3 program of a seed is the v2 program of that seed, only the dataset differs |
| `Shape { mixer_mult, cache_log2_words }`, `Shape::for_class_day(class, d)`, `MixParams.shape`, `Cache::fill_log2(key, log2)`, `round_key_mult(r, j, m)` | `memhard.rs` | `Shape::V2` is version 2 |
| `growth_doublings(d)`, `cache_log2_words(d)`, `dataset_log2_words(D_0, d)`, `days_since_genesis(day, genesis_day)` | `memhard.rs` | the schedule, one function and its two sizes |
| `DatasetSource::{new_shape, from_key_shape, shape}`, `Epoch::new_class_day`, `Epoch::from_seed_bytes_day(epoch, day, label, class, d, D_0)` | `verify.rs` | the day-sized entries; the v2 entries are unchanged and build the v2 shape |
| `IGNEUM_MIXER_MULT`, `IGNEUM_CLASS_MIXER_MULT`, `IGNEUM_CACHE_GROWTH` in program.h; `"mixer_mult"`, `"cache_growth"`, the `"item"` string in program.json | `emit.rs` | written only for a class with `m != 1` or growth, so v2 packs do not change |
| `packfile.h` `mixerMult`; `packbench` and the OpenCL host print the multiplier and the cache size | the three hosts | the kernels carry the construction in their text (one emitter, three dialects); the hosts size the cache from `IGNEUM_CACHE_LOG2_WORDS` already (packbench.swift line 56, host.cu line 65, host.c line 1966) |
| `igneum-pow --class mx4 [--days d]` on every command | `main.rs` | `--days` sizes the cache for a growth class |
Under the ca2-v3 seam (`ProgramClass::V3`, `V3_CLASS`), the integration sets `V3_CLASS = LoadClass::MX4`; the
chain's day-sized dataset needs the day index, so `Epoch::chain_dataset(day, class)` builds the genesis-size
cache and a `chain_dataset_day(day_bytes, class, d, D_0)` beside it is the growth entry (section 9, owed to the
node agent).
## 4. Vectors (class v3, `proto-cuda/packs-ca2-mixer/`)
Produced by `igneum-pow export --program-class v3` (the Rust CPU interpreter, 5 October 2026, commit 66eeba3) and
checked on the GPUs in section 6. The program of each pack is the version 2 program of the same seed instruction
for instruction (`tests/packs.rs`, `v3_packs_are_the_v2_seeds_under_mixer_x4`); the cache is the version 2 cache
(day 0 of the growth rule); the dataset words and the hashes are new. The 64 sampled indices are those of every
pack (`emit::sample_indices`).
Pack `mx4-genesis` (seed igneum-genesis, day 2026-10-03, generator 3, class mx4, program id e323b9dcaf283a6f, 2^28 words, 2^26-word cache, cache FNV-1a 64 `48c4f5bf24166b2e` as under v2):
```
dataset words 0..15 (item 0):
61ff2180 0d4c7e6c 2177d443 60df9025 cf8b2e10 63675bfb 25289e58 9c45dc42
2d271c54 9652369b 2dd77508 5921392c 3afa60ee c640ad68 f2bb56ff cfa46438
dataset[0x0fffffff] = 5020180e
sampled words (the first 8 of the 64 in vectors.json):
dataset[59471966] = de85726d dataset[217795994] = 7cfc31c7 dataset[208353206] = d3cd5289 dataset[42483309] = 6ebbeef1
dataset[172547758] = 9858413b dataset[148076330] = e786f141 dataset[183853158] = 64f13833 dataset[214389424] = ca229d04
unit at base nonce 0, lanes 0..31:
63acd2d273f475ba e929c78b34b80d4b 0b1011cb19982558 1457a0df5497aa11
957d0f3bb71d98fb ac16901e6e6f6057 8ea1c6279f4b177a f28146e60bd08ba9
fe5b8cfe87f8e65b 49f87240566ace62 6ef6d6b7bdea8e41 46d9c0dc29a97b9c
111fe30128db9398 66dc39084f0946d4 8ee11bdfd35fecf2 2861fcfc75db6677
31c7667d4bde8556 c5989c48858b4ce0 276395e734a9d30d 84217b41e91368ff
3604861e34d9f697 9f51d8ee16bf3639 c89e47bafa84401c 7ae78c1f10b70e19
0b8c947157a29a48 d67192e8cfb43842 05a4c6d182c8c675 188e2661f3263f2e
a2df24238f7fea2e ed69ea7e13ad3a48 2d44ae509bab91b8 adad61931ea4fb70
unit at base nonce 4096: lane 0 edd508ac57e5699a, lane 1 4eefd56d526cdaeb, lane 31 8892f8604733b1e0
unit at base nonce 1000000: lane 0 8b3183778a49f59c, lane 1 1831b72a8797e895, lane 31 75eae55eba53a506
```
Pack `mx4-devnet-epoch0` (seed the devnet genesis hash, day bytes of 2026-10-04, generator 3, class mx4, program id 73bcbfe8ccf988f1, 2^28 words, 2^26-word cache, cache FNV-1a 64 `448274a57f508cbc` as under v2):
```
dataset words 0..15 (item 0):
afe80d67 b9fbd029 6c79f193 95139ad9 96310aff 4609f8b1 75279e63 28235be1
47b17dcb 718e0ef2 a52588c8 a8bf49d5 19cf243e 5ec8905e a4851f66 af9cd9f3
dataset[0x0fffffff] = e6a99c7a
sampled words (the first 8 of the 64 in vectors.json):
dataset[59471966] = 57642b58 dataset[217795994] = c279badd dataset[208353206] = cbccbaad dataset[42483309] = 32cce392
dataset[172547758] = 71fdb4c6 dataset[148076330] = f5a268ce dataset[183853158] = 3ca1d676 dataset[214389424] = 7977b03d
unit at base nonce 0, lanes 0..31:
212c6442b51e87ae c374795c00839331 b6036a220a98f4b3 eb8b8013e637367b
db5866e9b73930fd f3f3d01f46e90333 9d913991ab8ed428 7ccb1d8fa100a800
3cf45ba44f09a0fe 91acf48ef1a63082 6ea46c69fb082f99 581f0218977a9d72
9a4623a5c62ddf2d ab6eb5e768f0feb4 07b70bdccf8aca12 d666311ae5e4311e
53114757d669f0a4 bd5d6ace87ce2ce4 fb712015e8189192 a32cec81103e134b
83f3d18c3289c124 fe29f1984b132b3d c9ffcf4e3774497a ac99c9243dc63809
d78a7e8217a32f3c ab81ad63d242fc31 0e4c30b7e00024af ce014289fff6778d
64a1292e2a8b4a91 d5b8c90e681d7e3a 06078117673030fd 51bf77b280173930
unit at base nonce 4096: lane 0 3c797978566b5950, lane 1 7c759e60185b6411, lane 31 96a903eb9a0ca390
unit at base nonce 1000000: lane 0 d5a8da0568df8ee7, lane 1 68cb69c04208285a, lane 31 f8ca84a1a5d78cf5
```
## 5. The v2 path is byte-identical
`cargo test --test packs` regenerates every file of `igneum-genesis-mh` and `igneum-devnet-v4-epoch0` from
`program.json` and compares byte for byte (`emitted_sources_match_all_packs`, `export_pack_matches_all_packs`);
section 6 also records a fresh `igneum-pow export` of both packs diffed against the checked-in directories.
## 6. Measurements
Machine: Apple M5 Max, 64 GiB, Darwin 25.6.0, 5 October 2026 (night), other agents' builds running beside every
run; a timing row says which lock it ran under (`measure` is exclusive; `run` and `build` are not timings).
### 6.1 The v2 path, byte for byte (no lock needed)
`igneum-pow export --seed igneum-genesis --day 2026-10-03` and `igneum-pow export --epoch-hex edc4fa84...fb07
--day-hex 69676e65756d2d6461792ffa50000000000000` on commit 66eeba3, `diff -r` against
`proto-cuda/packs/igneum-genesis-mh` and `igneum-devnet-v4-epoch0`: IDENTICAL, both (twelve files each). The
crate tests regenerate the same files and compare them on every run (`tests/packs.rs`, 12 of 12 pass).
### 6.2 Bit-exactness of the class v3 construction on the GPUs (`with-lock.sh run`, 22:05 UTC)
`packbench --pack <dir> --batches 1 --batch-log2 24 --group 256` (Metal, built from this branch) and
`igneum-bench-cl-igneum-genesis-mh --bench-pack --pack <dir> --batches 1 --batch-log2 24` (Apple OpenCL, built from
this branch's host.c). Vectors are the Rust interpreter's; the fingerprint is FNV-1a 64 over the 2^24 outputs at
base nonce 0.
| Pack | Harness | Cache FNV-1a 64 | Dataset head, word MASK, 64 samples | Vectors standalone / in batch | Fingerprint 2^24 | MH/s (GPU time; not a measurement, the run lock) |
|---|---|---|---|---|---|---|
| mx4-genesis | Metal | 48c4f5bf24166b2e PASS | PASS | 3/3, 3/3 | 6f48d5a2aa0dbe5f | 27.6 |
| mx4-genesis | Apple OpenCL | PASS | PASS | 96 of 96 lanes | 6f48d5a2aa0dbe5f | 27.7 (wall) |
| mx4-devnet-epoch0 | Metal | 448274a57f508cbc PASS | PASS | 3/3, 3/3 | 73caaebb28e808fe | 27.5 |
| mx4-devnet-epoch0 | Apple OpenCL | PASS | PASS | 96 of 96 lanes | 73caaebb28e808fe | 27.6 (wall) |
| mx8-genesis (the x8 candidate, 21:45 UTC) | Metal | 48c4f5bf24166b2e PASS | PASS | 3/3, 3/3 | 7c28cfb06c5c65a9 | 27.7 |
| mx8-genesis | Apple OpenCL | PASS | PASS | 96 of 96 lanes | 7c28cfb06c5c65a9 | 27.6 (wall) |
| mx8-devnet-epoch0 | Metal | 448274a57f508cbc PASS | PASS | 3/3, 3/3 | bbb183f72692f840 | 27.7 |
| mx8-devnet-epoch0 | Apple OpenCL | PASS | PASS | 96 of 96 lanes | bbb183f72692f840 | 27.7 (wall) |
Reading: the Rust interpreter, Metal and Apple OpenCL agree on the v3 dataset (head, MASK word, 64 samples), on
every vector lane and on the 2^24-output fingerprint of each pack; the hash rate is the v2 rate (27.7 MH/s on this
card tonight, readwidth table), as it must be: the hash kernel only loads, the mixer is paid in the build.
Dataset build under the run lock (indicative only; the measured rows are in 6.4): Metal 30.2 ms GPU (mx4-genesis)
and 21.7 ms GPU (mx4-devnet-epoch0) for 1 GiB; Apple OpenCL 55 and 56 ms wall. The readwidth entry's v2 figure on
this card is 0.6 to 2.1 ms cache fill and a 1 GiB build the bench-log's memory-hard entry puts at 13 to 30 ms; the
x4 build on the GPU is the row the measure lock will settle.
### 6.2a The packs as pinned after the x8 decision (`with-lock.sh run`, 22:12 UTC, the fixed binary's exports)
| Pack | How exported | Program id | Unit 0 lane 0 | Fingerprint 2^24 (Metal = Apple OpenCL) | Vectors, self-tests |
|---|---|---|---|---|---|
| mx8-genesis | `--program-class v3` (generator 3, V3_CLASS = mx8) | e323b9dcaf283a6f | 19b56348bc85304d | 7c28cfb06c5c65a9 | 3/3 + 3/3, 96 of 96, PASS |
| mx8-devnet-epoch0 | the chain path, `--program-class v3 --era-hex <genesis>` (the era drawn inside the class, load class `mx8-erad810f22d`) | 73bcbfe8ccf988f1 | d424577fce4a7a60 | 90f794dd556f7a3b | 3/3 + 3/3, 96 of 96, PASS |
| mx4-genesis (the x4 record) | `--class mx4` (generator 2, the class in the id) | 951b89584750bd75 | 63acd2d273f475ba | 6f48d5a2aa0dbe5f | 3/3 + 3/3, 96 of 96, PASS |
| mx4-devnet-epoch0 (the x4 record) | `--class mx4` on the chain seeds | (program.json) | 212c6442b51e87ae | 73caaebb28e808fe | 3/3 + 3/3, 96 of 96, PASS |
The PC 1 rows of 6.5 ran the earlier mx8-devnet-epoch0 (the load-class export, fingerprint bbb183f72692f840, era
not inside); the composed pack's PC fingerprints are owed with the next PC round (section 9).
### 6.3 Soundness suite on the v3 construction (commit 66eeba3 and after; `with-lock.sh build` for cargo, `run` for the GPU)
| Suite | Command | Result |
|---|---|---|
| Crate lib tests (memhard schedule table, the by-hand multiplied mixer, the seam, the generator) | `cargo test -j4 --release` | 44 of 44 pass |
| Pinned packs, v2 and v3 (`tests/packs.rs`: programs, ids, dataset words, 96 vectors per pack, every emitted file byte for byte, 16 masked loads per kernel, the v3 packs as the v2 seeds under mixer x4) | same | 12 of 12 pass |
| Scratch soundness tests of branch ca2-soundness (cherry-pick 0d8f745, one conflict in packbench.swift's RESULT line resolved by hand, `era_bytes: None` added to the edge program literal) | `cargo test -j4 --release --test scratch` | 7 of 7 pass: bijections, re-hit rates, 56 of 56 edge units, 42 of 42 emitted scr kernels, 200 scratch programs and 800 units on the CPU |
| v3 fuzz on the CPU (`tests/mixer.rs`): 200 programs through the seam, the contract on every instruction (each is the v2 program of its seed), 4 units each across the 32-bit range with one unit in the top 256 nonces, interpreted twice | `IGNEUM_MIXER_PACKS_OUT=<dir> cargo test -j4 --release --test mixer` | 200 of 200, 800 of 800 units; 200 packs written for the GPU runs (82 s with the three cache fills) |
| v3 stats beside v2 (`tests/mixer.rs`): 8,192 outputs per seed, bit balance, single-bit avalanche within the unit and across units, duplicates | same | igneum-genesis v3: avalanche 49.99 percent, worst bit z 1.92, 0 duplicates (v2: 49.87, z 2.25); igneum-genesis/stats1 v3: 49.97, z 3.09 (v2: 49.98, z 2.30) |
| v3 edge (`tests/mixer.rs`): items 0, 1, 2^28 - 1 and 2^32 - 1 by hand at m = 1, 2, 4, 8 on a 2^14-word cache; words 0, 15, 16, 17, MASK - 1, MASK through the interpreter's fetch path; the index wrap at MASK + 1 | same | pass |
| v3 determinism (`tests/mixer.rs`): two independent epochs, every vector and every emitted file equal, and equal to the pinned pack | same | pass |
| Metal and Apple OpenCL on the two pinned v3 packs | section 6.2 | 3/3 standalone, 3/3 in batch, 96 of 96 lanes, dataset head, MASK word and 64 samples, one fingerprint per pack across both harnesses |
| Metal fuzz: the 200 packs, 4 units each standalone and the top-256 unit inside a 512-nonce batch at base 4,294,967,040; every tenth pack on Apple OpenCL as well | `packbench --pack <dir> --batches 1 --batch-log2 9 --batch-base 4294967040` (Metal), `igneum-bench-cl-igneum-genesis-mh --bench-pack --pack <dir> --batches 1 --batch-log2 10` (Apple OpenCL), `with-lock.sh run`, 21:22 to 21:24 UTC | Metal 200 of 200 packs PASS (800 of 800 standalone units, 200 of 200 inside the wrapping window, cache and dataset self-tests on every pack); Apple OpenCL 20 of 20 packs PASS (the three vectors.h units, the self-tests); a first run with a packbench built before the `--batch-base` cherry-pick reported 200 of 200 FAIL on an empty RESULT line and was read as such (the watcher rule), the harness rebuilt and the run repeated |
### 6.4 Timings (`with-lock.sh measure`, one session, 21:40:12 to 21:40:23 UTC, commit 504cae4)
Script `measure-v3.sh` (session scratchpad): `igneum-pow bench --seed igneum-genesis --day 2026-10-03 --warps 50`
(v2), `... --program-class v3` (x4), `... --class mx8` (x8), two rounds each, then the devnet seeds, then
`packbench --pack <dir> --batches 2 --batch-log2 22 --group 256` on igneum-genesis-mh, mx4-genesis and mx8-genesis,
two rounds. The lock was exclusive among the agents' builds and measurements, but the box was not quiet: load
average 5.6 (one minute) and 26 (fifteen minutes) at the start, from unlocked processes (the devnet node, other
agents' editors); the v2 row reads 1.31 to 1.36 ms where the quiet readwidth night read 0.604 to 0.626. So the
absolute numbers below are a loaded-core figure, about 2.2x the quiet one, and the ratios between the rows are the
measurement (two rounds within 4 percent). A quiet-box re-run is owed (section 9).
| Construction | Verifier, ms per 32-lane unit, avg of 50 (round 1 / round 2) | Worst cold unit of three | Against v2 | 256 MiB cache fill, one core | Metal 1 GiB dataset build, GPU ms (round 1 / round 2) |
|---|---|---|---|---|---|
| v2 (igneum-genesis) | 1.361 / 1.310 | 1.579 | 1 | 172.1 / 172.6 ms | 29.7 / 21.0 |
| x4 (mx4, class v3) | 1.956 / 1.923 | 2.043 | 1.45x | 175.3 / 172.3 ms | 20.9 / 21.0 |
| x8 (mx8) | 2.785 / 2.790 | 2.942 | 2.09x | 172.3 / 172.3 ms | 21.9 / 21.9 |
| x4, the devnet seeds (mx4-devnet-epoch0) | 1.923 | 2.012 | | 173.9 ms | |
| x8, the devnet seeds | 2.972 | 2.885 | | 173.6 ms | |
Reading. The verifier's ALU part is what grows: x4 adds 0.6 ms per unit for 27 more mixer applications on each of
4,096 items (110,592 applications, about 5.5 ns each on this core, the lanes' chains interleaved), x8 another
0.85 ms for 36 more; the latency part (8 dependent misses per item) is the same in every row, which is why the
measured ratios are 1.45x and 2.1x and not the 4x and 8x of the M16 table's scaling. Shape B (32 rounds of one
read, section 1) would have multiplied the latency part too; x4 is under the 4.8 ms bar even on the loaded core, so
B stays unimplemented. The cache fill does not depend on the mixer (it is the ChaCha chain): 172 to 175 ms, the
spec's 175 to 181 ms of 1.8.3. The Metal 1 GiB build does not move with the mixer at all (21 ms at v2, x4 and x8
once warm; the 29.7 ms first v2 run is the first-touch cost the hosts fill twice for): on this card the build is
bound by the 8 dependent cache-line reads per item, not by the arithmetic, so the Mac says nothing about whether
the 5090's or the 9070 XT's build is arithmetic-bound; that is the PC job (section 8).
Against the x4 / x8 rule (section 6.5): the verifier half passes for x8 with 7.1 ms of the 10 ms gate to spare on
this loaded core (worst cold 2.94 ms; the quiet-core figure would be about 1.3 ms, scaled by the 2.2x of the v2
row, approximate); x4 leaves 8.0 ms. The build half waits on the PC rows.
### 6.5 Verification throughput per tier (consequences row C19), from the loaded-core figures above
Warps verified per second on one core = 1,000 / (ms per warp); a pool core verifying members' shares handles that
many shares per second; a node verifies a block with one unit (plus the header path, under 0.1 ms, not measured
here); IBD over the 108,000-header pruning window (spec 02) on one core = 108,000 x ms per warp.
| Figure | v2 | x4 | x8 | Note |
|---|---|---|---|---|
| ms per warp, steady (this session, loaded core) | 1.33 | 1.94 | 2.79 | avg of the two rounds |
| ms per warp, quiet M5 Max core (scaled by 0.604 / 1.33 = 0.45, approximate) | 0.60 | 0.88 | 1.26 | the readwidth night's v2 figure is measured; x4 and x8 scaled |
| ms per warp, 2019-class laptop core (approximate: 2.5x the quiet M5 Max figure, the ratio the design document assumes for the gate; unmeasured, O-1.14) | 1.5 | 2.2 | 3.2 | the figure that fixes the gate is a measurement, not this row |
| Shares per second per core (loaded / quiet, approximate) | 750 / 1,660 | 515 / 1,140 | 358 / 790 | spec 09 section 9.8 item 5 carried 2,270 at v2; re-cut from the quiet row: 1,660 |
| Cores for a 22,000-member pool at one share per member per 10 s (2,200 shares per second), loaded / quiet | 2.9 / 1.3 | 4.3 / 1.9 | 6.1 / 2.8 | |
| Node: worst cold single unit (loaded core) | 1.58 ms | 2.04 ms | 2.94 ms | per block |
| IBD over 108,000 headers on one core, loaded / quiet, minutes | 2.4 / 1.1 | 3.5 / 1.6 | 5.0 / 2.3 | laptop (approximate): 2.7 / 4.0 / 5.8 min; a seed VM core (unmeasured) sits between the laptop and the quiet M5 Max |
| Margin left under the 10 ms gate for Counter ASIC 3.0 (worst cold, loaded core) | 8.4 ms | 8.0 ms | 7.1 ms | on the 2019-class laptop row (approximate) 7.5 / 6.8 / 5.9 ms steady |
Reading: at x4 a pool core serves about 1,100 shares per second on a quiet 2026 core (a 22,000-member pool needs
two cores); at x8 about 800 (three cores). A node's block verification stays a few milliseconds. The gate's
remaining margin is what Counter ASIC 3.0 has to spend, and on the unmeasured laptop core it is 6 to 7 ms at x4 and
about 6 at x8, which is the number the 2019-class measurement (O-1.14) must confirm before x8 is final.
### 6.4a The same session on the fixed binary (section 6.6), `with-lock.sh measure`, 22:06:59 to 22:07:06 UTC
The verifier of section 6.4 was measured on a binary that carried the inlining regression of section 6.6; after the
fix, readwidth's binary and the fixed one on the same v2 input in the same minute (checksum 19297e99c7b9a55e), then
x4 and x8 on the fixed binary; load average 5.5 (the same box, so the ratios of 6.4 stand and the absolute row is
now the measured one):
| Construction | Verifier, ms per unit, avg of 50 (round 1 / round 2) | Worst cold unit of three | Against v2 |
|---|---|---|---|
| v2, readwidth e752fc7's binary | 0.607 / 0.610 | 0.666 | 1 |
| v2, the fixed binary | 0.609 / 0.611 | 0.657 | 1.00 |
| x4 (mx4) | 1.238 / 1.237 | 1.396 | 2.03x |
| x8 (mx8, class v3) | 2.077 / 2.058 | 2.145 | 3.4x |
| x4, the devnet seeds | 1.240 | 1.289 | |
| x8, the devnet seeds | 2.058 | 2.144 | |
The added cost per unit is the same as on the slow binary (x4 + 0.63 ms, x8 + 1.46 ms: the regression was a
constant 0.72 ms per unit in the shared item loop), so the ALU reading of 6.4 holds; the ratios against v2 are
2.0x and 3.4x once v2 is back at 0.61. The 10 ms gate keeps 7.9 ms at x8 (worst cold 2.15 ms) on this core.
### 6.5 The daily build per tier, and the x4 / x8 rule
The coordinator's rule (21:30 UTC): x8 enters v3 if the per-warp verify stays under 10 ms on one Mac core AND the
daily 1 GiB build stays under 1 s on every discrete card we own; else x4 with the thin margin stated and x8 named
as the next lever. The integrated tier is decided beside it (consequences row C23): its build is per prepare, not
per day, so its consequence is per-day dataset reuse in the workers or a restart per epoch.
| Card | Build at x1 | x4 | x8 | Source |
|---|---|---|---|---|
| RTX 5090 (PC 1, job run-mixer-x4-pc1-20261005, 22:00 to 22:04 UTC, the worker's `cache ... dataset ... ms` wall line, two dispatches per pack) | 23 to 25 ms | 23 to 25 ms | 23 ms | the PC 1 job (13.4 ms GPU time on 3 October: the wall line carries the launch) |
| RX 9070 XT (PC 1, gfx1201 on the eGPU, the same job) | 74 ms | 73 to 77 ms | 72 to 76 ms | the PC 1 job |
| M5 Max, Metal | 13 to 30 ms (the two runs of the memory-hard entry; tonight's run-lock figures 21.7 to 30.2 ms at x4 and 22.0 to 30.0 at x8 say the Mac's build is latency-bound, not mixer-bound) | section 6.4 | section 6.4 | this file |
| Radeon integrated gfx1036 (PC 2), OpenCL, per prepare | 6.9 / 9.4 / 11.7 s prepare total with the 1 GiB build inside | about 28 to 47 s (approximate: scaled x4; the iGPU's build is arithmetic-bound at x1 already) | about 55 to 94 s (approximate) | `docs/plans/epoch-length.md` section 7 (branch ca2-epoch), M11 table |
| gfx1036 beside WSL build jobs (PC 1) | 55 / 116 / 124 s | about 4 to 8 min (approximate) | about 7 to 17 min (approximate) | same |
| 8 GB-class discrete card (not owned; about a tenth of the 5090's rate, approximate) | about 0.13 s | about 0.5 s | about 1 s, on the edge of the rule | scaled from the 5090 row, approximate |
Decision (coordinator under the delegated rule, 22:05 UTC, on these rows): x8 enters class v3. Both halves pass:
the verifier at x8 is 2.1 ms per unit on one M5 Max core (6.4a) against the 10 ms gate, and the daily 1 GiB build
does not move with the mixer on any discrete card we own (5090 23 to 25 ms, 9070 XT 72 to 77 ms, M5 Max 21 ms at
x1, x4 and x8: latency-bound), 13x to 40x under the 1 s bar. `V3_CLASS = { era: None, hot: None, ..LoadClass::MX8 }`;
the pinned v3 packs are mx8-genesis and mx8-devnet-epoch0 (the latter through the chain path with the era inside the
class); the x4 packs stay pinned as the candidate's record (generator 2, the class in the id).
Reading of the integrated tier: on the discrete cards the rule is settled by the 5090 and 9070 XT rows above. The integrated tier misses the rule at x4 already: a per-prepare build of 28 to 47 s is a tenth to a quarter
of the 600-DAA-second lead the devnet gives the next program (spec 1.12), and under load it is the whole lead; so
if x4 or x8 goes in, the iGPU tier needs the workers to build the day's dataset once a day and keep it across
epochs (today a prepare rebuilds it: `proto-opencl/host.c` prepareTask builds the pair's cache and dataset per
prepare), or to restart per epoch. The node agent is asked whether the per-day reuse is bounded tonight
(coordinator, 21:45 UTC); until then the iGPU consequence stands as written.
### 6.6 The verifier regression of 0fc0ad1, found and fixed (5 October 2026, 21:59 to 22:07 UTC)
The era agent measured the same v2 input with two binaries in one minute: readwidth's 0.604 ms per unit, ca2-v3
HEAD's 1.33. Bisected under the measure lock (one session, four binaries, two rounds, checksum 19297e99c7b9a55e):
readwidth e752fc7 0.607 / 0.609; the ca2-v3 seam 6c75dad (before this branch) 0.610 / 0.609; this branch's 0fc0ad1
1.332 / 1.316; ca2-v3 88dafbc 1.325 / 1.347. So the 2.2x was in 0fc0ad1's `derive_items`, on the version 2 path the
devnet verifies with, and section 6.4's "loaded box" reading was wrong: the load was real (the same session shows
it) but the 2x was the code.
Variants, each a one-change copy measured against readwidth's binary in the same session:
| Variant | ms per unit | Reading |
|---|---|---|
| 0fc0ad1 as written (the line mask read from the cache at run time, the loop inlined into `MemhardCpu::fetch`) | 1.33 | the regression |
| the mask hoisted into a local before the item loop | 1.32 to 1.37 | not the reload |
| `Cache::line` with the constant mask, on the 0fc0ad1 tree | 0.60 to 0.65 | fixed there |
| the same constant mask on the merged ca2-v3 tree | 1.33 to 1.46 | not the mask either |
| one instance per cache size with the mask a constant, `#[inline(always)]` | 1.32 to 1.39 | not the mask |
| the same instances `#[inline(never)]` | 0.604 / 0.617 / 0.618 / 0.624 | the fix |
So it is inlining: the item loop inlined into its callers (`fetch`, `fetch_wide`, `word_at`) runs at 2.2x the
out-of-line loop, and which small change tips LLVM's decision depends on the rest of the tree (the constant mask
tipped it on one tree and not on the other). The fix (`memhard::derive_items`): the loop is `derive_items_mask`,
`#[inline(never)]`, one instance per cache size the growth rule reaches (2^26 to 2^30 words) with the line mask a
constant, and a run-time-mask instance for every other size (tests). Measured in 6.4a: 0.609 / 0.611 against
readwidth's 0.607 / 0.610.
The class, not the instance: a verifier benchmark with a pinned bound in the crate's CI (the v2 unit at a known
input against a stored ms-per-unit on a named core, failing on a 1.3x drift) would have caught this at the first
commit; filed for the next cut (section 9). Until then the era agent's two-binary check (same input, same minute)
is the rule for every change that touches the item loop.
## 7. The chip model
`docs/analysis/chip-model-v3.md`.
## 8. What is unverified
1. The 5090's and the 9070 XT's dataset build at x4 and x8 are measured (6.5, the PC 1 job): both latency-bound,
under 0.1 s. The gfx1036 is the tier that fails the per-prepare build (6.5), and its consequence (per-day dataset
reuse in the workers) is with the node agent.
2. The absolute verifier figures were taken on a loaded core (load average 5.6); the ratios are the measurement
and the quiet-core figures are scaled. A 2019-class laptop core has not run any construction (O-1.14).
3. The mixer has had no cryptanalysis (spec 1.8.4); `m` applications with distinct round keys is `m` times the
work only if no shortcut composes them, which is the same open question as for one application.
4. The x8 packs are generator 2 with the class in the id (`--class mx8`); if x8 is chosen, the pinned v3 packs are
re-cut through the seam (`V3_CLASS = MX8`, generator 3) and the tests re-pinned, one commit.
5. The 5090's rate for the v3 program is the v2 rate by construction (the hash kernel is unchanged, the Mac shows
27.7 MH/s at v2, x4 and x8); the chip row's denominator stays the readwidth table's 136.1 MH/s until a v3 pack
runs on the card, which the PC job also gives.
## 9. Owed
| Item | Owner | When |
|---|---|---|
| PC 1 run of the five packs | done 22:04 UTC (run-mixer-x4-pc1-20261005; 6.5 and 6.2a) | |
| A verifier benchmark with a pinned bound in the crate's CI (section 6.6, the class rule) | ca2-mixer | the next cut |
| GPU bit-exactness of the re-exported mx8-devnet-epoch0 (the composed class, the era inside) and the mx4 record packs on the PCs (the Mac rows are in 6.2a) | ca2-mixer | the next PC round |
| The x4 / x8 choice recorded from the rule, then the vectors re-cut once through the seam | done 22:05 UTC: x8 (6.5), V3_CLASS = MX8, mx8 packs re-exported | |
| `Epoch::chain_dataset_day` wired to the genesis day index in the node (`days_since_genesis(day_index(header), day_index(genesis))`) and `pow_genesis_dataset_log2` in the override | ca2-node | the integration |
| The spec text of section 2 into `docs/spec/01-lottery-hash.md` 1.8.5 and 1.13.3 (with the v3 vectors into 1.17) | the integration | after the choice |
| The 2019-class laptop core measurement that fixes the gate (O-1.14) | cryptographer | gate 1 |

126
docs/plans/read-width.md Normal file
View file

@ -0,0 +1,126 @@
# Read width of the lottery hash: 4, 16 and 64-byte loads, a per-load mix, and a written scratch (gate 1 experiment)
5 October 2026. Branch `readwidth` (worktree `../igneum-wt-readwidth`), commits 019b014 and b970dda plus the measurement commit. Nothing here changes consensus, the live generator, the pinned vectors or a shipped binary: every class sits behind `--class` in `igneum-pow` and the default class is generator version 2 byte for byte (`igneum-pow/tests/packs.rs` still compares the four pinned packs against the emitters). Numbers and a recommendation; the decision is the project lead's.
## 1. The question
The bench-log entry "the 9070 XT on the eGPU" (5 October 2026) found the hash bound by dependent random 4-byte reads over the 1 GiB dataset, 128 per hash: the RX 9070 XT finishes 2.4 to 2.7 G such reads a second (18 MH/s), the RTX 5090 16 to 18 G (127 MH/s), the M5 Max 3.45 G (23 to 28 MH/s). AMD fetches a 64-byte line per 4-byte read, so 94 percent of its memory traffic is unused; NVIDIA fetches a 32-byte sector and its 96 MB L2 catches a share. the project lead's question: would wider reads keep the chip-resistance property (latency-bound, random access) while closing the vendor gap? Two additions from the coordinator: a per-load width drawn from an era-fixed mix so no chip is built for one width, and a written per-warp scratch so part of the memory work cannot be mirrored into read-only SRAM.
## 2. What was built (all behind the flag)
| Class (`--class`) | Loads per hash | What a load does | Dataset bytes per hash |
|---|---|---|---|
| `v2` (= `w4`, the lottery hash) | 128 | `dst ^= dataset[src & MASK]`, one 4-byte word | 512 |
| `w16` | 128 | the 16-byte-aligned group of 4 words at `src & MASK`, every word folded into `dst` | 2,048 |
| `w64` | 128 | the 64-byte-aligned item (16 words), every word folded | 8,192 |
| `w64x4` | 32 (4 load slots) | as `w64`; the same bytes per hash as 512 loads of 4 bytes | 2,048 |
| `mix50-35-15` | 128 | per load, width 4, 16 or 64 bytes drawn from the program stream with probabilities 50/35/15 | 1,664 to 3,680 over the six programs measured (expected 2,202) |
| `mix25-50-25` | 128 | the same with 25/50/25 | 2,240 to 5,024 (expected 3,200) |
| `scr<k>k<kb>` | 128 memory operations | `k` of the 16 slots are scratch read-modify-writes into a `kb` KiB per-warp scratch (16-byte slots, lane-major); the other `16 - k` are 4-byte loads | 4 x (16 - k) x 8 reads plus 16 B read and 16 B written per scratch op |
The fold. A load of W words reads the W-word-aligned address `b = (src AND MASK) AND NOT (W - 1)` and sets `x = dst XOR w[0]; for j in 1..W: x = (rotl(x, 11) * 0x9e3779b1) XOR w[j]; dst = x` (`verify::fold_words`, mirrored in the three kernel dialects). The rotate-multiply between the words makes the fold state-dependent: two different lines give two different maps of `dst` (the multiply by an odd constant is not xor-linear), so no function of the line alone can stand in for it, and a dataset of pre-folded lines cannot replace the dataset. The dependent chain is unchanged: the next load's address comes from a register that the fold wrote. Width 1 is the lottery hash's xor of one word, so `w4` is the pinned program `bcc1248b10cc90f2` exactly.
The per-load draw. Every class other than `v2` takes one extra draw per instruction (`below(100)`, the width roll, consumed on every slot so the stream stays uniform) after the nine draws of spec 01 section 1.4.3; on a load slot the width is the first entry of the mix whose cumulative weight exceeds the roll. The program id of a class is `FNV-1a-64("igneum-program-rw/" || 2 || seed words || attempt || mix[3] || slots [|| "scratch/" k kb])`, so no class program can pass for a version 2 program. The acceptance rule of 1.4.6 runs unchanged on the aligned addresses (lane-constant sites and the distinct-address bound, scaled to the dataset loads per hash).
The scratch (variant 5, measurement only). The kernel runs N persistent warps (one per block or work-group of 32); warp `w` owns scratch `w` and runs units `w, w + N, ...` of the launch. A scratch op reads the lane's 16-byte slot `src AND (slots - 1)`: three data words behind a per-unit tag; a slot whose tag is not this unit's reads as its fill `splitmix32(((base + lane) XOR seed[j]) + slot x 0x9e3779b1 + (j + 1) x 0x85ebca77)`, the three words are folded into `dst` as above, and the slot is rewritten `(tag, x XOR w1, rotl(x, 7) XOR w2, x + w0)`. The CPU verifier holds the touched slots of one unit (at most 32 x k x 8) and nothing else. The scratch is per lane (a 32 KiB warp scratch is 64 slots per lane, 128 KiB is 256), so two lanes never race on a slot and the result is a function of (program, day, unit) alone; the GPU's tags make the lazy fill exact as long as a tag is not reused within the arena's history (the salt advances per unit; it wraps after 2^32 units, a measurement caveat, not a design).
## 3. Method
| Step | Command (every figure in the bench log carries its command) |
|---|---|
| Packs | `igneum-pow export --seed <s> --class <c> --out proto-cuda/packs-readwidth/<name>` (23 packs; the six mix seeds per mix are `igneum-readwidth/A/<k>` and `/B/<k>` with attempt 0 accepted, `A/4` skipped: rejected at attempt 0) |
| CPU verifier | `igneum-pow bench --seed igneum-genesis --class <c> --warps 50` (M5 Max, one core; load average 4 to 9 from other agents' builds during the run) |
| Metal | `proto-metal/packbench --pack <dir> --batches 5 --batch-log2 24 --group 256 [--warps N]` (new harness: runs the pack's own text; vectors, cache FNV, dataset words, 2^24 fingerprint, MH/s by GPU time) under the measure lock |
| Apple OpenCL | `proto-opencl/igneum-bench-cl-rw --bench-pack --pack <dir> --batches 5 --batch-log2 24` and `--memprobe` (Apple's OpenCL, a correctness check and an approximate rate) |
| Emulators | `proto-cuda/emu/emu.sh ../packs-readwidth/<p> --batch-log2 13 --batches 1 --block-warps 2` (the CUDA text as C++); `proto-opencl/emu/emu.sh ../packs-readwidth/<p> 0 32 --sg 32` and `1 64 --sg 64` (the OpenCL text, wave32 and wave64) |
| RTX 5090 | job `run-readwidth-5090-20261005` on PC 2 (`relay/playbooks/readwidth-5090.ps1`): the NVIDIA card switched off in the app through `POST app.url/api/cards` and restored after; `igneum-worker-cuda --memprobe`, then `--bench --pack <dir> --batches 5 --batch-log2 24` per pack (NVRTC, the pack's own text, vectors through the bound kernel, 2^24 fingerprint) |
| RX 9070 XT | job `run-readwidth-9070-20261005` on PC 1 (`relay/playbooks/readwidth-9070.ps1`): only the gfx1201 card switched off; `igneum-worker-opencl --device D --memprobe`, then `--bench-pack --pack <dir> --batches 5 --batch-log2 24` per pack |
Latency-bound share = measured MH/s x loads per hash / the card's dependent-read ceiling for that width from its own probe at 1024 MiB (for a mix, the harmonic combination of the widths' ceilings weighted by the program's width counts). A share near 1 means the hash runs at the card's random-access limit, the property the design wants; a share well under 1 means something else bounds it (bandwidth, ALU, occupancy).
## 4. Results (full tables with commands in `docs/bench-log.md`, "read width of the lottery hash")
Probe ceilings at 1024 MiB (G dependent reads/s): RTX 5090 4 B 17.5, 16 B 18.0, 64 B 9.1 (584 GB/s), stream 1,579 GB/s; RX 9070 XT 4 B 2.42, 16 B 2.43, 64 B 2.47 (158 GB/s), stream 636; M5 Max (Apple OpenCL, approximate) 3.50 / 3.51 / 3.51, stream 522.
| Class | dataset B/hash | RTX 5090 MH/s (latency-bound share) | RX 9070 XT (share) | M5 Max Metal (share) | 5090 / 9070 | DRAM bytes moved per hash, NVIDIA 32 B sector / AMD 64 B line | CPU verify ms per unit |
|---|---|---|---|---|---|---|---|
| v2 = w4 (today) | 512 | 136.1 (0.96) | 18.15 (0.87) | 27.74 (1.01) | 7.5x | 4,096 / 8,192 | 0.604 |
| w16 | 2,048 | 139.8 (0.90) | 17.90 (0.84) | 28.26 (1.03) | 7.8x | 4,096 / 8,192 | 0.610 |
| w64 | 8,192 | 71.9 (0.58) | 17.59 (0.78) | 28.27 (1.03) | 4.1x | 8,192 / 8,192 | 0.630 |
| w64x4 (32 loads) | 2,048 | 275.3 (0.56) | 75.19 (0.84) | 109.7 (1.00) | 3.7x | 2,048 / 2,048 | 0.160 |
| mix50-35-15 (6 programs, min / median / max) | 1,664 to 3,680 | 99.5 / 114.2 / 121.0, spread 18.8% | 17.45 / 18.76 / 18.83, 7.4% | 25.36 / 27.26 / 28.43, 11.3% | 6.1x | 5,939 / 8,192 expected | 0.620 |
| mix25-50-25 (6 programs) | 2,240 to 5,024 | 95.9 / 107.3 / 119.8, 22.3% | 17.84 / 18.45 / 18.85, 5.5% | 23.21 / 24.68 / 25.21, 8.1% | 5.8x | 7,168 / 8,192 expected | 0.614 |
### 4.1 Per watt and per pound (consequences review C11)
The runs carried no power sampling; the watts are the telemetry entry's (`docs/bench-log.md`, opencl-rdna4-telemetry, 5 October 2026: the RTX 5090 at 307.6 W under its 450 W cap for 122.3 MH/s, the RX 9070 XT at 199 W of its 304 W rating for about 17.8 MH/s, both on v2 with the shader clock at its top and the die waiting on memory), held constant across classes because every class is memory-bound on both cards (approximate: a class that moves more bytes per hash draws somewhat more at the memory controller, unmeasured). The Mac's GPU power is not measurable without root (`powermetrics`) and is taken as about 50 W (approximate, from memory). Prices are UK list, approximate, from memory.
| Class | RTX 5090 MH/W (at 307.6 W) | RX 9070 XT MH/W (at 199 W) | 5090 / 9070 per watt | M5 Max MH/W (at about 50 W GPU, approximate) | 5090 MH per pound (at about 1,900, approximate) | 9070 XT MH per pound (at about 570, approximate) |
|---|---|---|---|---|---|---|
| v2 (w4) | 0.442 | 0.091 | 4.9x | 0.55 | 0.072 | 0.032 |
| w16 | 0.454 | 0.090 | 5.1x | 0.57 | 0.074 | 0.031 |
| w64 | 0.234 | 0.088 | 2.6x | 0.57 | 0.038 | 0.031 |
| w64x4 | 0.895 | 0.378 | 2.4x | 2.19 | 0.145 | 0.132 |
| mix50-35-15 (median) | 0.371 | 0.094 | 3.9x | 0.55 | 0.060 | 0.033 |
| mix25-50-25 (median) | 0.349 | 0.093 | 3.8x | 0.49 | 0.056 | 0.032 |
| scr8k32 | 0.397 | 0.071 | 5.6x | 0.98 | 0.064 | 0.025 |
| scr2k32 | 0.372 | 0.074 | 5.1x | 0.52 | 0.060 | 0.026 |
Reading: whatever width is chosen, an AMD home miner keeps about a seventh of a 5090's rate and pays about 4.5x the electricity per hash, because every width costs the 9070 XT the same 2.4 G line fetches a second; per pound of card the 5090 is 2.2x the 9070 XT at v2 and w16 (0.072 against 0.032 MH/s per pound) and 4.9x per watt; only w64x4 narrows the per-pound gap (0.145 against 0.132), and that class fails the width rule. The consequence for the decision (D6, the project lead's): AMD's line width is not a read-width question at all; it is the card's random-access rate, and the levers that act on it (the 64 MB Infinity Cache against the dataset size, the memory path) are v3-or-3.0 questions outside this experiment.
Scratch, variant 5 (N persistent warps; GPU cost against the persistent control scr0k32; working set = 1 GiB + 256 MiB + 128 MiB output + N x size):
| Class | RMW share | dataset B/hash | scratch B/hash read + written | RTX 5090 MH/s, 2,048 warps of 4,080 resident (vs control, share) | M5 Max Metal (vs control) | RX 9070 XT, 4,096 warps | working set 5090 / 9070 / Mac |
|---|---|---|---|---|---|---|---|
| scr0k32 | 0 | 512 | 0 | 139.1 (control, 0.98) | 28.25 (control) | 17.88 (control, 0.86) | 1.4 GiB / 1.5 GiB / 1.5 GiB |
| scr2k32 | 12.5% | 448 | 256 + 256 | 114.4 (-18%, 0.80) | 26.14 (-7%) | 14.65 (-18%) | same |
| scr4k32 | 25% | 384 | 512 + 512 | 109.8 (-21%, 0.76) | 31.74 (+12%) | 14.00 (-22%) | same |
| scr8k32 | 50% | 256 | 1,024 + 1,024 | 122.1 (-12%, 0.82) | 49.08 (+74%) | 14.17 (-21%) | same |
| scr2k128 | 12.5% | 448 | 256 + 256 | 110.1 (-21%, 0.77) | 26.24 (-7%) | 14.07 (-21%) | 1.6 GiB / 1.9 GiB / 1.9 GiB |
| scr4k128 | 25% | 384 | 512 + 512 | 98.0 (-30%, 0.68) | 28.08 (-1%) | 13.14 (-27%) | same |
| scr8k128 | 50% | 256 | 1,024 + 1,024 | 72.8 (-48%, 0.49) | 35.44 (+25%) | 12.03 (-33%) | same |
Resident warps and the cap: the 5090 holds 4,080 warps at one warp per block (24 blocks per SM x 170 SMs; 8,160 at 8 warps per block), so 128 KiB each is 510 MiB and the whole working set 1.9 GiB; a 1 MB scratch would have been 4.0 GiB at this geometry and 10.6 GiB at the 64-warp-per-SM figure, which is why the cap moved the size to the tens of kilobytes. The occupancy query returned 24 blocks per SM before and after the arena allocation: the allocation did not change it. The 9070 XT's OpenCL runtime has no occupancy query; 4,096 persistent warps were launched (64 per compute unit over 64 CUs, approximate) and the arena is 128 MiB at 32 KiB, 512 MiB at 128 KiB. The Mac's residency is not reported; 2,048 to 16,384 warps were swept and the best row kept.
Chip model, re-run with the measured widths (the M16 arithmetic of `docs/analysis/m16-recompute-attacker-2026-10-05.md`; the on-die-cache recompute chip's row per scratch variant is the ca2-soundness branch's, as agreed with the Counter ASIC 2.0 coordinator):
| Class | what a chip with its own DRAM controller gains over the GPU's memory system | what a chip with on-die SRAM gains |
|---|---|---|
| v2 | the GPU fetches 8 to 16x the bytes it uses (AMD 64 B, NVIDIA 32 B per 4 B); a chip fetching 32 B bursts moves 4,096 B per hash, the 5090's figure, so nothing over NVIDIA and 2x over AMD in traffic, none in latency (the chain is 128 dependent DRAM latencies on either) | the recompute attacker of M16: 150,000 integer ops per hash against the 256 MiB cache; 2.4x at equal silicon before a fixed-function factor (unchanged by the width) |
| w16 | the same: 4,096 / 8,192 bytes moved, 2,048 used; traffic efficiency 50 percent on NVIDIA, 25 on AMD | unchanged: the fold uses every byte, so the chip recomputes 128 items per hash exactly as before; the SRAM mirror of the read-only dataset (1 GiB) stays out of reach |
| w64 | every byte moved is used on both vendors (8,192 moved, 8,192 used); the 5090 is bandwidth-bound at 589 GB/s, so a chip with HBM3 class bandwidth (several TB/s, approximate) is bandwidth-advantaged: the Ethash shape | unchanged in op count; but the chain of 128 loads now moves 8 KB, so a chip's advantage shifts from latency to bandwidth per dollar, which is the wrong direction for the design's 2x target |
| w64x4 | 2,048 moved and used; 32 latencies per hash; every card 4x faster; a bandwidth-rich chip gains as above | 32 items per hash: the recompute attacker's op count falls 4x (37,500 per hash), so the M16 gain rises 4x: fails the 2x target by arithmetic |
| mixes | between v2 and w64 per program; the chip cannot be built for one width, but the GPU pays the 64-byte hours (the 5090 loses up to 27 percent in a heavy hour) | as v2 per item; the recompute attacker is indifferent to the width |
| scratch | a chip must provide writable memory for N units in flight: 32 KiB x N at the GPU's geometry (128 MiB at 4,080), against the 256 MiB read-only cache it could mirror into SRAM (54 to 83 mm^2 at a leading node, the coordinator's figure, approximate); but a unit touches at most k x 8 x 32 slots (4 KiB at 50 percent), the fill is a function and the tags are per unit, so a chip need only hold the touched set per unit in flight (the soundness caveat below) | the dataset reads replaced by scratch ops are reads the chip no longer has to serve from the 1 GiB; at 50 percent the recompute attacker computes 64 items instead of 128 |
Soundness (variant 5, measurement only, as instructed; the chip row is the ca2-soundness branch's, a465881: the on-die-cache recompute chip's gain is 2.4x at 0, 12.5, 25 and 50 percent, replaced or added, 32 or 128 KB, so the scratch does not move it): the per-unit scratch starts from a fill that any implementation can compute, and a unit writes at most `k x 8` slots per lane; an implementation that keeps only the touched slots of each unit in flight (the CPU verifier does exactly this) needs 16 B x touched slots, not the nominal arena, so the "real memory a chip must provide" is bounded by units in flight x touched slots, not by N x 32 KiB. The variant forces memory that is written, which SRAM can hold as well as DRAM; it does not force memory that is large. A written region that outlives the unit (state carried across units) would, and the CPU verifier could not replay it. This is the finding, not a recommendation.
## 5. Recommendation (the decision is the project lead's)
the project lead's rules, as passed by the coordinator: width = the widest read that keeps every card latency-bound (achieved within 90 percent of the probe ceiling at that width) with margin on the 5090 (bytes per hash x rate under a third of the 1,579 GB/s stream); the mix is in only if the six-program spread is under 5 percent per card; the scratch share is the smallest at which the chip model's gain falls under 1.5x at the lowest GPU cost within the 6 GB cap.
| Variant | Verdict under the rules | Numbers |
|---|---|---|
| w16 (16-byte loads, 128 per hash) | PASSES the rules: shares 0.90 / 0.84 / 1.03 (the 9070 XT's 0.84 equals its v2 share of 0.87 within noise: the card is at its ceiling in both), 286 GB/s on the 5090 = 18 percent of the stream. It does NOT close the vendor gap (7.8x against 7.5x), because the memory systems already move a sector or a line per load; it changes what the fold consumes, nothing the DRAM does | the only width row that passes; a no-cost change in rate (+2.7 percent 5090, -1.4 percent 9070 XT, +1.9 percent M5 Max) |
| w64 | FAILS: 5090 share 0.58, 37 percent of the stream; closes the gap to 4.1x only by making the 5090 bandwidth-bound | |
| w64x4 | FAILS: shares 0.56 / 0.84 / 1.00, the recompute gain rises 4x | |
| mix 50/35/15 and 25/50/25 | OUT: spreads 18.8 and 22.3 percent on the 5090, 7.4 and 5.5 on the 9070 XT, 11.3 and 8.1 on the M5 Max, all over 5 percent; a chip is not built for a width anyway (see the model: the width does not change the recompute attacker) | |
| scratch | OUT: every share costs the 5090 12 to 48 percent and the 9070 XT 18 to 33 percent, and raises the M5 Max's rate (the arena is cached there); the soundness branch's chip row (ca2-soundness a465881, the on-die-cache recompute chip) stays at 2.4x at every share, 32 or 128 KB, because the verifier resets the scratch per unit and the live state is the hash's own read-modify-writes, which a chip keeps in 80 to 320 B per lane; under the rule the share is 0. The rows stay as the measurement that decided it | |
Recommendation: keep 128 loads per hash and 4 bytes per load (v2) for the devnet; if a width change is wanted for the fold's sake (every byte of the sector consumed, which removes the "94 percent waste" statement from the AMD entry without changing what the card does), w16 is the one that passes every rule and costs nothing measurable, and it is the only width worth a vector re-cut. The AMD gap is a random-access gap (2.4 G against 17.5 G dependent reads per second at 1 GiB on the cards we own), and no read width closes it without turning the 5090 bandwidth-bound; the levers that act on the gap are the ones outside this experiment (the AMD card's memory path, and the dataset size against the 5090's 96 MB L2 share, which the probe's 64 MiB rows show at 9 G reads/s against 2.4 at 1 GiB). The per-load mix is out on stability; the scratch is out on GPU cost and on the soundness caveat.
## 6. What w16 would change if adopted (not done; the decision is the project lead's)
| Where | Change |
|---|---|
| `docs/spec/01-lottery-hash.md` 1.4.1 | `load`: `dst = fold(dst, dataset[b .. b + 4))`, `b = (src AND MASK) AND NOT 3`, with the fold written out; 1.4.3 unchanged (no width draw for a fixed width); 1.4.6 unchanged (the aligned address is the address the rule sees) |
| 1.5 | "A load reads one 4-byte word" becomes 16 bytes aligned; the single-form text-search rule of 1.14 item 2 becomes the wide form; the item size (64 B) and `dataset[w] = item(w >> 4)[w AND 15]` unchanged |
| 1.11 | unchanged in count (4,096 items per unit; the verifier derives the same items) |
| 1.15, 1.17 | every vector re-cut (new program ids: the class enters the id or the generator version steps to 3); the four pinned packs replaced; the conformance fuzz re-run on Metal, CUDA and OpenCL (this branch's 23 packs and the three emulators are the template) |
| Litepaper, Mining ("random reads over a multi-gigabyte dataset") and the vs-RandomX "128 dataset addresses" rows | "128 reads of 16 bytes"; `site/bench.html` sector arithmetic (32 B per 4 B) becomes 32 B per 16 B |
| Workers | no host change: the kernel text carries the loads; `proto-cuda/host.cu`'s static mask check (`TESTS.md` section 5) learns the wide form |
| Cost on the 5090 | none measured (+2.7 percent); on the 9070 XT -1.4 percent; verifier +1 percent |
## 7. Files
`igneum-pow/src/{generator,verify,accept,emit,memhard,main}.rs` (the classes, behind `--class`), `proto-cuda/packs-readwidth/` (23 packs), `proto-metal/packbench.swift` (Metal from a pack's files), `proto-opencl/host.c` (`--bench-pack`, `--warps`, the 16-byte probe row, the scratch arguments), `proto-cuda/nvrtc/worker.cpp` (`--bench`, `--memprobe`, the scratch arena), `proto-cuda/nvrtc/packfile.h` (class fields; string-seed packs), `proto-cuda/emu/cuda_runtime.h` and `proto-opencl/emu/{emu_opencl.h,emu_main.cpp}` (vector types, the persistent launch), `relay/playbooks/readwidth-*.ps1` (the PC jobs: the card under test off in the app and restored, never the other card).

View file

@ -11,12 +11,14 @@
//! | (c) dynamic | the program is interpreted for [`ACCEPT_UNITS`] (64) units of 32 lanes at base nonces drawn from SplitMix64 seeded with `FNV-1a-64("igneum-accept/" \|\| seed words as little-endian bytes)`, each `low32(next()) AND NOT 31`, with init words equal to the seed words and the closed-form dataset `dataset_elem(idx, S[0], S[1])` at [`ACCEPT_DATASET_LOG2`] (2^28 words) in place of the memory-hard dataset. Over the 2,048 evaluations: no register has a bit equal in every final value; no load site (iteration, instruction) reads one address in all 32 lanes of any unit; fewer than [`MAX_SATURATED`] (164, 1 percent of 16,384) final register values are 0 or 2^32 - 1; every output bit's ones count is within [`BIAS_TOLERANCE`] (136, 6 sigma) of 1,024; the distinct masked addresses read by one lane in one evaluation, summed over the 2,048 evaluations, exceed [`MIN_DISTINCT_SUM`] (245,760, a mean above 120 of the 128 loads) |
//!
//! The dynamic test uses the closed form so that it is a pure function of the program (no cache, no day) and
//! costs about a millisecond on one core. The census (section 7.3) checked on 100,000 programs that the
//! costs about a millisecond on one core. A hot-table load (`docs/plans/hot-table.md`) reads the closed form keyed by
//! seed words 2 and 3 at its multiply-shift index, a second pure table beside the dataset stand-in (words 0 and 1). The census (section 7.3) checked on 100,000 programs that the
//! closed-form verdict agrees with the memory-hard one on all but 39 threshold-edge cases.
use crate::generator::{Instr, Op, Program, INSTR_COUNT, ITERATIONS, LANES};
use crate::seed::{fnv1a64, SplitMix64};
use crate::verify::{dataset_elem, splitmix32};
use crate::memhard::hot_index;
use crate::verify::{dataset_elem, fold_words, load_index, splitmix32, ScratchModel};
/// Units (32-lane warps) the dynamic test interprets.
pub const ACCEPT_UNITS: usize = 64;
@ -33,6 +35,14 @@ pub const BIAS_TOLERANCE: u32 = 136;
/// Distinct addresses per lane per evaluation, summed over 2,048 evaluations, must exceed this (mean above 120).
pub const MIN_DISTINCT_SUM: u64 = 245_760;
/// The distinct-address bound for a program with `loads` dataset loads per hash: the same 120 of 128 ratio, so
/// [`MIN_DISTINCT_SUM`] for the lottery hash and `loads x 1,920` for the read-width classes with other counts.
/// Variant 5's scratch read-modify-writes are not dataset loads: their slots repeat by design (a later
/// read-modify-write sees an earlier write), so they are neither counted nor bounded here.
pub fn min_distinct_sum(loads: usize) -> u64 {
loads as u64 * ACCEPT_HASHES as u64 * 120 / 128
}
/// Why a candidate was rejected. The verdict (accept or reject) is what consensus depends on; the reason is the
/// first failing test in the order of the module table.
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
@ -67,7 +77,7 @@ impl std::fmt::Display for Reject {
Reject::Saturated { count } => write!(f, "(c) {count} of 16384 final register values saturated (limit 163)"),
Reject::OutputBias { bit, ones } => write!(f, "(c) output bit {bit} set in {ones} of 2048 hashes"),
Reject::DistinctAddresses { sum } => {
write!(f, "(c) distinct addresses {sum} over 2048 hashes (mean {:.2}, needs above 120)", *sum as f64 / 2048.0)
write!(f, "(c) distinct dataset addresses {sum} over 2048 hashes (mean {:.2}, needs above 120 of 128 of the dataset loads)", *sum as f64 / 2048.0)
}
}
}
@ -170,6 +180,8 @@ fn run_unit(p: &Program, unit: usize, base: u32, acc: &mut Acc, lane_addrs: &mut
let seed = &p.seed;
let mask: u32 = (1u32 << ACCEPT_DATASET_LOG2) - 1;
let (d0, d1) = (seed[0], seed[1]);
let (h0, h1) = (seed[2], seed[3]);
let hot_words = p.hot_words();
let loads = p.loads_per_hash();
let mut r = [[0u32; LANES]; 8];
for lane in 0..LANES {
@ -183,12 +195,30 @@ fn run_unit(p: &Program, unit: usize, base: u32, acc: &mut Acc, lane_addrs: &mut
}
let mut idx = [0u32; LANES];
let mut nload = 0usize;
let mut scratch = if p.has_scratch() { Some(ScratchModel::new(p.class.scratch_slots_per_lane())) } else { None };
let slot_mask = p.class.scratch_slot_mask();
let era = p.class.era;
for it in 0..ITERATIONS {
let sel = r[0];
for (k, ins) in p.instrs.iter().enumerate() {
let d = ins.dst as usize;
let a = ins.src as usize;
match ins.op {
Op::Scratch => {
// Variant 5: the slot stands in for the address (bit 31 set so it never aliases a dataset word).
let m = scratch.as_mut().expect("a scratch op needs a scratch class");
for lane in 0..LANES {
idx[lane] = r[a][lane] & slot_mask;
}
if idx.iter().all(|&x| x == idx[0]) {
return Err(Reject::LaneConstantSite { iteration: it as u8, instr: k as u8, unit: unit as u8 });
}
for lane in 0..LANES {
r[d][lane] = m.rmw(&p.seed, base, lane, idx[lane], r[d][lane]);
lane_addrs[lane * loads + nload] = 0x8000_0000 | idx[lane];
}
nload += 1;
}
Op::Add => {
let (imm, imm2, bit) = (ins.imm, ins.imm2, ins.bit as u32);
let src = r[a];
@ -255,18 +285,45 @@ fn run_unit(p: &Program, unit: usize, base: u32, acc: &mut Acc, lane_addrs: &mut
}
}
Op::Load => {
// Read-width experiment: a load of `width` words reads from the aligned address and folds every
// word (verify::fold_words); width 1 is the lottery hash's xor of one word.
let width = ins.width as usize;
let align = !(ins.width as u32 - 1);
for lane in 0..LANES {
idx[lane] = r[a][lane] & mask;
idx[lane] = load_index(era.as_ref(), ins, r[a][lane], mask, ACCEPT_DATASET_LOG2) & align;
}
if idx.iter().all(|&x| x == idx[0]) {
return Err(Reject::LaneConstantSite { iteration: it as u8, instr: k as u8, unit: unit as u8 });
}
for lane in 0..LANES {
r[d][lane] ^= dataset_elem(idx[lane], d0, d1);
if width == 1 {
r[d][lane] ^= dataset_elem(idx[lane], d0, d1);
} else {
let mut w = [0u32; 16];
for j in 0..width {
w[j] = dataset_elem(idx[lane] + j as u32, d0, d1);
}
r[d][lane] = fold_words(r[d][lane], &w[..width]);
}
lane_addrs[lane * loads + nload] = idx[lane];
}
nload += 1;
}
Op::Hot => {
// Hot table: the stand-in is dataset_elem keyed by seed words 2 and 3; the address is tagged with
// bit 30 so a hot word and a dataset word at one index count as two addresses.
for lane in 0..LANES {
idx[lane] = hot_index(r[a][lane], hot_words);
}
if idx.iter().all(|&x| x == idx[0]) {
return Err(Reject::LaneConstantSite { iteration: it as u8, instr: k as u8, unit: unit as u8 });
}
for lane in 0..LANES {
r[d][lane] ^= dataset_elem(idx[lane], h0, h1);
lane_addrs[lane * loads + nload] = 0x4000_0000 | idx[lane];
}
nload += 1;
}
Op::WLoad => {
let b = (r[a][0] & mask) & !31;
for lane in 0..LANES {
@ -298,7 +355,8 @@ fn run_unit(p: &Program, unit: usize, base: u32, acc: &mut Acc, lane_addrs: &mut
sl.sort_unstable();
let mut distinct = 0u64;
for k in 0..loads {
if k == 0 || sl[k] != sl[k - 1] {
// scratch slots carry bit 31 (variant 5) and are not dataset addresses
if sl[k] & 0x8000_0000 == 0 && (k == 0 || sl[k] != sl[k - 1]) {
distinct += 1;
}
}
@ -333,7 +391,7 @@ pub fn check_dynamic(p: &Program) -> Result<AcceptReport, Reject> {
}
bias_max = bias_max.max(d);
}
if acc.distinct_sum <= MIN_DISTINCT_SUM {
if acc.distinct_sum <= min_distinct_sum(loads - p.scratch_ops_per_hash()) {
return Err(Reject::DistinctAddresses { sum: acc.distinct_sum });
}
Ok(AcceptReport { distinct_sum: acc.distinct_sum, saturated: acc.saturated, bias_max })
@ -348,9 +406,50 @@ pub fn check(p: &Program) -> Result<AcceptReport, Reject> {
#[cfg(test)]
mod tests {
use super::*;
use crate::generator::{candidate, generate, GeneratorConfig, generate_v1};
use crate::generator::{candidate, candidate_class, generate, generate_class, GeneratorConfig, generate_v1, LoadClass};
use crate::verify::{DatasetMode, DatasetSource};
#[test]
fn distinct_bound_scales_with_the_load_count() {
assert_eq!(min_distinct_sum(128), MIN_DISTINCT_SUM);
assert_eq!(min_distinct_sum(32), 61_440);
}
/// The read-width classes pass the rule at about the version 2 rate, and the instrumented interpreter agrees
/// with `verify.rs` on every class (the fold is shared, the addresses are aligned the same way).
#[test]
fn classes_pass_and_match_verify() {
for name in ["w16", "w64", "w64x4", "50,35,15", "25,50,25", "scr2k32", "scr8k128"] {
let c = LoadClass::parse(name).unwrap();
let p = generate_class("igneum-genesis", c);
assert!(check(&p).is_ok(), "{name}");
let mut rejected = 0;
for i in 0..60u32 {
let s = format!("igneum-rw-accept/{i}");
let q = candidate_class(&s, s.as_bytes(), 0, c);
if check(&q).is_err() {
rejected += 1;
}
}
assert!(rejected < 15, "{name}: {rejected} of 60 rejected");
let ds = DatasetSource::from_key(p.seed, DatasetMode::ClosedForm, ACCEPT_DATASET_LOG2);
let bases = accept_base_nonces(&p.seed);
let loads = p.loads_per_hash();
let mut acc = Acc { and_acc: [u32::MAX; 8], or_acc: [0; 8], saturated: 0, bit_ones: [0; 64], distinct_sum: 0 };
let mut la = vec![0u32; LANES * loads];
let mut ones = [0u32; 64];
for (u, &b) in bases.iter().enumerate() {
run_unit(&p, u, b, &mut acc, &mut la).unwrap();
for h in crate::verify::hash_warp(&p, b, &ds) {
for j in 0..64 {
ones[j] += ((h >> j) & 1) as u32;
}
}
}
assert_eq!(acc.bit_ones, ones, "{name}: bit counts match the reference interpreter");
}
}
/// The instrumented interpreter agrees with `verify.rs` on the closed-form dataset keyed by the seed words.
#[test]
fn instrumented_interpreter_matches_verify() {
@ -381,6 +480,52 @@ mod tests {
}
}
/// Hot-table experiment: the hot classes pass the rule at about the version 2 rate, and the hot addresses are
/// uniform over the table (16 buckets of the index over 64 units x 32 lanes x 32 hot loads).
#[test]
fn hot_classes_pass_and_hot_loads_are_uniform() {
for name in ["hot32k4", "hot64k4", "hot96k4", "hot64k2", "hot64k8", "scr4k32+hot64k4", "hot32k4a", "hot64k4a", "hot96k4a"] {
let c = LoadClass::parse(name).unwrap();
let p = generate_class("igneum-genesis", c);
assert!(check(&p).is_ok(), "{name}");
let mut rejected = 0;
for i in 0..60u32 {
let s = format!("igneum-hot-accept/{i}");
let q = candidate_class(&s, s.as_bytes(), 0, c);
if check(&q).is_err() {
rejected += 1;
}
}
assert!(rejected < 15, "{name}: {rejected} of 60 rejected");
}
let p = generate_class("igneum-genesis", LoadClass::hot(96, 4));
let words = p.hot_words();
let loads = p.loads_per_hash();
let mut acc = Acc { and_acc: [u32::MAX; 8], or_acc: [0; 8], saturated: 0, bit_ones: [0; 64], distinct_sum: 0 };
let mut la = vec![0u32; LANES * loads];
let mut buckets = [0u64; 16];
let mut hot_count = 0u64;
for (u, &b) in accept_base_nonces(&p.seed).iter().enumerate() {
run_unit(&p, u, b, &mut acc, &mut la).unwrap();
for &a in &la {
if a & 0xC000_0000 == 0x4000_0000 {
let idx = a & 0x3FFF_FFFF;
assert!(idx < words);
buckets[(idx as u64 * 16 / words as u64) as usize] += 1;
hot_count += 1;
}
}
}
assert_eq!(hot_count, 64 * 32 * 32, "32 hot loads per hash over 2,048 hashes");
let mean = hot_count as f64 / 16.0;
for (i, &b) in buckets.iter().enumerate() {
assert!((b as f64 - mean).abs() < 0.15 * mean, "bucket {i}: {b} against a mean of {mean}");
}
// the dataset distinct count still holds for the dataset loads alone
let r = check(&p).unwrap();
assert!(r.distinct_mean() > 120.0);
}
#[test]
fn base_nonces_are_aligned_and_seed_dependent() {
let a = accept_base_nonces(&[1, 2, 3, 4, 5, 6, 7, 8]);

File diff suppressed because it is too large Load diff

File diff suppressed because it is too large Load diff

View file

@ -32,7 +32,7 @@ pub mod verify;
pub use bind::{block_init_words, day_bytes, pow256_from_lane, target64_from_le256};
pub use accept::{check as accept_program, AcceptReport, Reject};
pub use generator::{generate, generate_from_seed_bytes, Instr, Op, Program, GENERATOR_VERSION};
pub use memhard::{Cache, MemhardCpu, MixParams};
pub use generator::{generate, generate_from_seed_bytes, generate_from_seed_bytes_program_class, Instr, LoadClass, Op, Program, ProgramClass, GENERATOR_VERSION, GENERATOR_VERSION_V3, V3_CLASS};
pub use memhard::{cache_log2_words, dataset_log2_words, days_since_genesis, growth_doublings, Cache, MemhardCpu, MixParams, Shape};
pub use seed::{fnv1a64, seed_words, SplitMix64};
pub use verify::{hash_warp, interpret_warp_init, verify_block, DatasetMode, DatasetSource, Epoch};

View file

@ -7,11 +7,22 @@
//! [--epoch-hex <64 hex> --day-hex <hex>] byte seeds instead of strings (Epoch::from_seed_bytes)
//! igneum-pow accept --seed <s> [--epoch-hex <64 hex>] every candidate of the seed with its verdict (spec 01 section 1.4.6)
//! igneum-pow show --seed <s> [--epoch-hex <64 hex>] the accepted program, one instruction per line
//!
//! Read-width experiment (5 October 2026, docs/plans/read-width.md): `--class v2|w4|w16|w64|w64x4|p4,p16,p64[xN]`
//! on every command selects the load class (default v2, the lottery hash). Nothing in a v2 run changes.
//!
//! Era layout (5 October 2026, docs/plans/era-layout.md): `--era igneum-era-test/<n>` (a test era seed: the 32 bytes
//! of seed_words_from_bytes of the string, era index n) or `--era <n>:<64 hex>` (the chain's 32-byte era seed E_n)
//! turns the chosen class into its era class; `--era-widths 4` (default: the read-width decision of 5 October 2026
//! keeps v2's 4-byte load) is the allowed width set the era draws from; `4,16,64` lets the era draw the width.
use igneum_pow::generator::{EraParams, GENERATOR_VERSION_V3};
use igneum_pow::emit::export_pack;
use igneum_pow::memhard::Cache;
use igneum_pow::generator::{LoadClass, ProgramClass};
use igneum_pow::memhard::{Cache, Shape};
use igneum_pow::seed::day_key;
use igneum_pow::verify::{DatasetMode, Epoch, DEFAULT_DATASET_LOG2};
use igneum_pow::verify::{DatasetMode, DatasetSource, Epoch, DEFAULT_DATASET_LOG2};
use std::time::Instant;
struct Args {
@ -26,6 +37,51 @@ struct Args {
prehash: String,
epoch_hex: Option<String>,
day_hex: Option<String>,
class: LoadClass,
/// Days since genesis for the cache growth rule of a class with `growth` (0: the genesis cache).
days: u64,
/// The program class (Counter ASIC 2.0 seam): v2 (default) or v3, which draws from V3_CLASS with generator 3.
program_class: Option<ProgramClass>,
/// The era seed bytes a class v3 chain program records (`--era-hex`).
era_hex: Option<String>,
/// Era layout: `--era igneum-era-test/<n>` or `--era <n>:<64 hex>` composes the era class over `--class` with
/// generator 3 and the era bytes recorded (the measurement packs: v2's mixer under the era layout).
era: Option<(u64, Vec<u8>, String)>,
era_widths: Vec<u8>,
}
/// `igneum-era-test/<n>` or `<n>:<64 hex>` -> (index, 32 era bytes, label).
fn parse_era(s: &str) -> Option<(u64, Vec<u8>, String)> {
if let Some(n) = s.strip_prefix("igneum-era-test/") {
let index: u64 = n.parse().ok()?;
return Some((index, EraParams::test_era_bytes(s).to_vec(), s.to_string()));
}
let (n, hex) = s.split_once(':')?;
let index: u64 = n.parse().ok()?;
let bytes = igneum_pow::bind::unhex(hex)?;
if bytes.len() != 32 {
return None;
}
Some((index, bytes, format!("igneum-era/{index}/{hex}")))
}
/// "4,16,64" (bytes) -> ascending words.
fn parse_widths(s: &str) -> Option<Vec<u8>> {
let mut v: Vec<u8> = s
.split(',')
.map(|x| match x.trim() {
"4" => Some(1u8),
"16" => Some(4),
"64" => Some(16),
_ => None,
})
.collect::<Option<Vec<_>>>()?;
v.sort_unstable();
v.dedup();
if v.is_empty() {
return None;
}
Some(v)
}
fn usage() -> ! {
@ -36,7 +92,12 @@ fn usage() -> ! {
\x20 hash --nonce <n> print the 64-bit hash of one nonce (pack form, init words = seed words)\n\
\x20 hash-bound --prehash <64 hex> --nonce <u64> print the header-bound hash (bind.rs) of one 64-bit nonce\n\
\x20 accept every candidate of the seed (or --epoch-hex) with its acceptance verdict\n\
\x20 show the accepted program, one instruction per line"
\x20 show the accepted program, one instruction per line\n\
\x20 --class C load class: v2 (default), mx4 (class v3: mixer x4, cache growth), w4, w16, w64, w64x4, p4,p16,p64[xN], <class>m<mult>[g]\n\
\x20 --days N days since genesis for the cache growth rule of a class with it (default 0: the 2^26-word cache)\n\
\x20 --program-class v2|v3 the program class of the seam (v3 = generator 3 on V3_CLASS, the chain's own derivation; --era-hex records the era seed)\n\
\x20 --era E era layout over --class: igneum-era-test/<n> or <n>:<64 hex> (the 32-byte era seed E_n)\n\
\x20 --era-widths 4[,16,64] the width set the era draws from, in bytes (default 4: pinned; more lets the era draw it)"
);
std::process::exit(2)
}
@ -54,6 +115,12 @@ fn parse() -> Args {
prehash: "00".repeat(32),
epoch_hex: None,
day_hex: None,
class: LoadClass::V2,
days: 0,
program_class: None,
era_hex: None,
era: None,
era_widths: vec![1],
};
let mut it = std::env::args().skip(1);
a.cmd = it.next().unwrap_or_else(|| usage());
@ -70,12 +137,30 @@ fn parse() -> Args {
"--prehash" => a.prehash = val(),
"--epoch-hex" => a.epoch_hex = Some(val()),
"--day-hex" => a.day_hex = Some(val()),
"--class" => a.class = LoadClass::parse(&val()).unwrap_or_else(|| usage()),
"--days" => a.days = val().parse().unwrap_or_else(|_| usage()),
"--program-class" => a.program_class = Some(ProgramClass::parse(&val()).unwrap_or_else(|| usage())),
"--era-hex" => a.era_hex = Some(val()),
"--era" => a.era = Some(parse_era(&val()).unwrap_or_else(|| usage())),
"--era-widths" => a.era_widths = parse_widths(&val()).unwrap_or_else(|| usage()),
_ => usage(),
}
}
if let Some((_, bytes, _)) = &a.era {
a.class = LoadClass::era(a.class, bytes, &a.era_widths);
}
a
}
/// An era program is a class v3 program: generator 3 and the era bytes recorded (what the chain's
/// `Epoch::from_chain_seeds` does); the pack then carries IGNEUM_PROGRAM_CLASS "v3" and IGNEUM_ERA_SEED_HEX.
fn stamp_era(e: &mut Epoch, a: &Args) {
if let Some((_, bytes, _)) = &a.era {
e.program.generator = GENERATOR_VERSION_V3;
e.program.era_bytes = Some(bytes.clone());
}
}
fn main() {
let a = parse();
let mode = if a.closed_form { DatasetMode::ClosedForm } else { DatasetMode::MemoryHard };
@ -85,21 +170,14 @@ fn main() {
"accept" => accept(&a),
"show" => show(&a),
"hash" => {
let e = Epoch::new(&a.seed, &a.day, mode, a.dataset_log2);
let (e, _) = epoch_of(&a, mode);
println!("{:016x}", e.hash(a.nonce as u32));
}
"hash-bound" => {
let bytes = igneum_pow::bind::unhex(&a.prehash).unwrap_or_else(|| usage());
let prehash: [u8; 32] = bytes.as_slice().try_into().unwrap_or_else(|_| usage());
// --epoch-hex / --day-hex: the chain's byte seeds (Epoch::from_seed_bytes), as the worker protocol carries them
let e = match (&a.epoch_hex, &a.day_hex) {
(Some(eh), Some(dh)) => {
let eb = igneum_pow::bind::unhex(eh).unwrap_or_else(|| usage());
let db = igneum_pow::bind::unhex(dh).unwrap_or_else(|| usage());
Epoch::from_seed_bytes(&eb, &db, "cli")
}
_ => Epoch::new(&a.seed, &a.day, mode, a.dataset_log2),
};
let (e, _) = epoch_of(&a, mode);
let init = igneum_pow::bind::block_init_words(&prehash, a.nonce);
println!("init words {}", init.iter().map(|w| format!("{w:08x}")).collect::<Vec<_>>().join(" "));
println!("{:016x}", e.hash_bound(&prehash, a.nonce));
@ -108,6 +186,49 @@ fn main() {
}
}
/// The epoch every command works on, and the day label for packs. `--epoch-hex`/`--day-hex` give the chain's byte
/// seeds (the day label then names the day bytes); else the string seed and day. `--program-class v3` draws the
/// program through the seam (generator 3 on `V3_CLASS`, the era bytes of `--era-hex` recorded) and sizes the
/// dataset for `--days` through `Epoch::chain_dataset_day`; `--class` is ignored under a program class (the class
/// names the load class). Closed-form mode is only for string seeds under the default class.
fn epoch_of(a: &Args, mode: DatasetMode) -> (Epoch, String) {
let (mut e, label) = epoch_of_class(a, mode);
stamp_era(&mut e, a);
(e, label)
}
fn epoch_of_class(a: &Args, mode: DatasetMode) -> (Epoch, String) {
let era = a.era_hex.as_ref().map(|h| igneum_pow::bind::unhex(h).unwrap_or_else(|| usage()));
match (&a.epoch_hex, &a.day_hex) {
(Some(eh), Some(dh)) => {
let eb = igneum_pow::bind::unhex(eh).unwrap_or_else(|| usage());
let db = igneum_pow::bind::unhex(dh).unwrap_or_else(|| usage());
let label = format!("igneum-epoch/{eh}/day/{dh}");
let e = match a.program_class {
Some(pc) => Epoch {
program: Epoch::chain_program(&eb, era.as_deref(), pc, &label),
dataset: Epoch::chain_dataset_day(&db, pc, a.days, a.dataset_log2),
},
None => Epoch::from_seed_bytes_day(&eb, &db, &label, a.class, a.days, a.dataset_log2),
};
(e, format!("bytes:{dh}"))
}
_ => {
let e = match a.program_class {
Some(pc) => {
let program = igneum_pow::generator::generate_from_seed_bytes_program_class(&a.seed, a.seed.as_bytes(), pc, era.as_deref());
let lc = pc.load_class();
let shape = Shape::for_class_day(&lc, a.days);
let log2 = if lc.growth { igneum_pow::memhard::dataset_log2_words(a.dataset_log2, a.days) } else { a.dataset_log2 };
Epoch { program, dataset: DatasetSource::new_shape(&a.day, mode, log2, shape) }
}
None => Epoch::new_class_day(&a.seed, &a.day, mode, a.dataset_log2, a.class, a.days),
};
(e, a.day.clone())
}
}
}
fn bench(a: &Args, mode: DatasetMode) {
println!(
"igneum-pow bench: seed \"{}\", day \"{}\", dataset 2^{} words ({})",
@ -116,21 +237,41 @@ fn bench(a: &Args, mode: DatasetMode) {
a.dataset_log2,
mode.name()
);
let shape = Shape::for_class_day(&a.program_class.map(|pc| pc.load_class()).unwrap_or(a.class), a.days);
if mode == DatasetMode::MemoryHard {
// Time the cache fill on its own first (one core), then build the epoch (which fills it again).
let t0 = Instant::now();
let c = Cache::fill(day_key(&a.day));
let c = Cache::fill_log2(day_key(&a.day), shape.cache_log2_words);
let fill_ms = t0.elapsed().as_secs_f64() * 1e3;
println!("cache: fill {fill_ms:.1} ms on one core (2^26 words, 65536 chains of 64 ChaCha12 blocks), FNV-1a 64 {:016x}", c.fnv1a64());
println!(
"cache: fill {fill_ms:.1} ms on one core (2^{} words, {} MiB, {} chains of 64 ChaCha12 blocks), FNV-1a 64 {:016x}",
shape.cache_log2_words,
shape.cache_words() * 4 / (1 << 20),
c.segments(),
c.fnv1a64()
);
drop(c);
}
if let Some(h) = a.class.hot {
// the hot table of the epoch on its own first (one core), then the epoch (which fills it again)
let t0 = Instant::now();
let t = igneum_pow::memhard::HotTable::for_seed_bytes(a.seed.as_bytes(), h.mb as u32);
let fill_ms = t0.elapsed().as_secs_f64() * 1e3;
println!("hot table: {} MiB filled in {fill_ms:.1} ms on one core ({} chains of 64 ChaCha12 blocks), FNV-1a 64 {:016x}", h.mb, igneum_pow::memhard::hot_segments(h.mb as u32), t.fnv1a64());
}
let t0 = Instant::now();
let e = Epoch::new(&a.seed, &a.day, mode, a.dataset_log2);
let (e, _) = epoch_of(a, mode);
let build_ms = t0.elapsed().as_secs_f64() * 1e3;
println!(
"program: {} loads/hash, {} items/warp, op mix {}; epoch built in {build_ms:.1} ms",
"program: class {}, {} loads/hash, {} bytes/hash, widths (1,4,16 words) {:?}, {} items/warp, mixer x{} ({} mixers/item), cache 2^{} words, op mix {}; epoch built in {build_ms:.1} ms",
e.program.class.name(),
e.program.loads_per_hash(),
e.program.bytes_per_hash(),
e.program.width_counts(),
e.program.items_per_warp(),
shape.mixer_mult,
shape.mixers_per_item(),
shape.cache_log2_words,
e.program.op_mix()
);
let bases = [0u32, 4096, 1_000_000];
@ -158,14 +299,7 @@ fn export(a: &Args, mode: DatasetMode) {
let out = a.out.clone().unwrap_or_else(|| usage());
let t0 = Instant::now();
// --epoch-hex / --day-hex: the chain's byte seeds; the day label then names the day bytes
let (e, day_label) = match (&a.epoch_hex, &a.day_hex) {
(Some(eh), Some(dh)) => {
let eb = igneum_pow::bind::unhex(eh).unwrap_or_else(|| usage());
let db = igneum_pow::bind::unhex(dh).unwrap_or_else(|| usage());
(Epoch::from_seed_bytes(&eb, &db, &format!("igneum-epoch/{eh}/day/{dh}")), format!("bytes:{dh}"))
}
_ => (Epoch::new(&a.seed, &a.day, mode, a.dataset_log2), a.day.clone()),
};
let (e, day_label) = epoch_of(a, mode);
let build_ms = t0.elapsed().as_secs_f64() * 1e3;
println!("igneum-pow export {out}");
println!(
@ -179,7 +313,7 @@ fn export(a: &Args, mode: DatasetMode) {
e.program.program_id(),
e.program.loads_per_hash()
);
println!("op mix: {}", e.program.op_mix());
println!("op mix: {}; class {}, {} bytes/hash, widths (1,4,16 words) {:?}", e.program.op_mix(), e.program.class.name(), e.program.bytes_per_hash(), e.program.width_counts());
let source = format!("igneum-pow (Rust) CPU interpreter, generator v{}, {} dataset", e.program.generator, e.dataset.mode().name());
let pack = export_pack(&e, &day_label, &source);
let dir = std::path::Path::new(&out);
@ -209,7 +343,7 @@ fn seed_bytes_of(a: &Args) -> (String, Vec<u8>) {
fn accept(a: &Args) {
let (label, bytes) = seed_bytes_of(a);
let t0 = Instant::now();
let tries = igneum_pow::generator::attempts(&label, &bytes);
let tries = igneum_pow::generator::attempts_class(&label, &bytes, a.class);
let ms = t0.elapsed().as_secs_f64() * 1e3;
for (p, verdict) in &tries {
match verdict {
@ -233,19 +367,42 @@ fn accept(a: &Args) {
fn show(a: &Args) {
let (label, bytes) = seed_bytes_of(a);
let p = igneum_pow::generator::generate_from_seed_bytes(&label, &bytes);
let mut p = igneum_pow::generator::generate_from_seed_bytes_class(&label, &bytes, a.class);
if let Some((_, eb, _)) = &a.era {
p.generator = GENERATOR_VERSION_V3;
p.era_bytes = Some(eb.clone());
}
println!(
"seed \"{}\" generator v{} attempt {} program id {:016x} seed words {}",
"seed \"{}\" generator v{} class {} attempt {} program id {:016x} seed words {}",
p.seed_string,
p.generator,
p.class.name(),
p.attempt,
p.program_id(),
p.seed.iter().map(|w| format!("{w:08x}")).collect::<Vec<_>>().join(" ")
);
println!("op mix {} loads/hash {}", p.op_mix(), p.loads_per_hash());
println!("op mix {} loads/hash {} bytes/hash {}", p.op_mix(), p.loads_per_hash(), p.bytes_per_hash());
if let Some(e) = p.class.era {
println!(
"era {} ({}): width {} B, stride mul {:#010x} rot {}, interleave {:?}, windows (site:shrink:offset) {}",
e.label(),
a.era.as_ref().map(|x| x.2.as_str()).unwrap_or("?"),
e.width_words as u32 * 4,
e.stride_mul,
e.stride_rot,
e.pos,
p.instrs
.iter()
.enumerate()
.filter(|(_, i)| i.op == igneum_pow::generator::Op::Load)
.map(|(k, i)| format!("{k}:{}:{}", i.win, i.off))
.collect::<Vec<_>>()
.join(" ")
);
}
for (k, i) in p.instrs.iter().enumerate() {
println!(
"{k:2}: {:5} dst={} src={} src2={} imm={:#010x} imm2={:#010x} rot={} bit={} mask={}",
"{k:2}: {:5} dst={} src={} src2={} imm={:#010x} imm2={:#010x} rot={} bit={} mask={}{}",
i.op.name(),
i.dst,
i.src,
@ -254,7 +411,8 @@ fn show(a: &Args) {
i.imm2,
i.rot,
i.bit,
i.mask
i.mask,
if i.op == igneum_pow::generator::Op::Load && i.width > 1 { format!(" width={}B", i.width as u32 * 4) } else { String::new() }
);
}
}

View file

@ -3,12 +3,18 @@
//! ARX-multiply mixer. The verifier holds the cache and never the dataset.
//!
//! All arithmetic is on u32 modulo 2^32. Rotations are by 1..31 at every call site.
//!
//! Counter ASIC 2.0 (5 October 2026, `docs/plans/mixer-x4.md`, behind the program class): the construction has a
//! [`Shape`], the mixer multiplier `m` and the cache size. Under `m` every mixer application of an item becomes
//! `m` applications with distinct round keys, the 8 dependent cache reads unchanged; the cache doubles when the
//! dataset doubles ([`growth_doublings`]). [`Shape::V2`] (`m = 1`, 2^26 words) is version 2 bit for bit.
use crate::generator::LoadClass;
use crate::seed::{day_key, fnv1a64_words, SplitMix64};
pub const CACHE_LOG2_WORDS: usize = 26;
pub const CACHE_SEGMENT_LOG2_LINES: usize = 6;
/// 2^26 words = 256 MiB.
/// 2^26 words = 256 MiB (the version 2 cache, and the v3 cache until the first dataset doubling).
pub const CACHE_WORDS: usize = 1 << CACHE_LOG2_WORDS;
/// 2^22 lines of 16 words.
pub const CACHE_LINES: usize = CACHE_WORDS >> 4;
@ -23,6 +29,102 @@ pub const CHACHA_ROUNDS: usize = 12;
pub const CHACHA_SIGMA: [u32; 4] = [0x61707865, 0x3320646e, 0x79622d32, 0x6b206574];
/// "Igne", "umMH".
pub const CACHE_TAG: [u32; 2] = [0x49676e65, 0x756d4d48];
/// "Igne", "umHT": the chain tag of the hot table (hot-table experiment, `docs/plans/hot-table.md`).
pub const HOT_TAG: [u32; 2] = [0x49676e65, 0x756d4854];
/// Domain tag of the hot key: `KH = seed_words_from_bytes("igneum-hot/" || epoch seed bytes)`.
pub const HOT_KEY_TAG: &[u8] = b"igneum-hot/";
/// Words per MiB of hot table.
pub const HOT_WORDS_PER_MIB: u32 = 1 << 18;
/// Segments (64 chained lines of 16 words, 4 KiB) per MiB of hot table.
pub const HOT_SEGMENTS_PER_MIB: u32 = 256;
/// The shape of the item derivation and of the cache: the mixer multiplier and the cache size.
#[derive(Clone, Copy, Debug, PartialEq, Eq, Hash)]
pub struct Shape {
/// Mixer applications per round (and after the last read): 1 under version 2, 4 under class v3.
pub mixer_mult: u32,
/// The cache is 2^cache_log2_words words (26 at genesis; 27 and 28 after the dataset doublings of 1.13.3).
pub cache_log2_words: u32,
}
impl Shape {
/// Version 2: one mixer application per round, a 2^26-word cache.
pub const V2: Shape = Shape { mixer_mult: 1, cache_log2_words: CACHE_LOG2_WORDS as u32 };
/// The shape of a load class on day 0 of the chain (and on every day for a class without the growth rule).
pub fn for_class(class: &LoadClass) -> Shape {
Shape::for_class_day(class, 0)
}
/// The shape of a load class on day `days_since_genesis` of the chain: the class's multiplier, and the cache
/// of [`cache_log2_words`] when the class has the growth rule, else 2^26 words.
pub fn for_class_day(class: &LoadClass, days_since_genesis: u64) -> Shape {
Shape {
mixer_mult: class.mixer_mult(),
cache_log2_words: if class.growth { cache_log2_words(days_since_genesis) } else { CACHE_LOG2_WORDS as u32 },
}
}
pub fn is_v2(&self) -> bool {
*self == Shape::V2
}
pub fn cache_words(&self) -> usize {
1usize << self.cache_log2_words
}
pub fn cache_lines(&self) -> usize {
self.cache_words() >> 4
}
pub fn cache_line_mask(&self) -> u32 {
(self.cache_lines() - 1) as u32
}
pub fn cache_segments(&self) -> usize {
self.cache_lines() >> CACHE_SEGMENT_LOG2_LINES
}
pub fn log2_segments(&self) -> u32 {
self.cache_log2_words - 4 - CACHE_SEGMENT_LOG2_LINES as u32
}
/// Mixer applications per item: `(ITEM_ROUNDS + 1) x m`.
pub fn mixers_per_item(&self) -> u32 {
(ITEM_ROUNDS as u32 + 1) * self.mixer_mult
}
}
// --------------------------------------------------------------------------------------------------------------
// Dataset growth, option C (spec 01 section 1.13.3 option (b) with the cache tied to the dataset's doublings)
// --------------------------------------------------------------------------------------------------------------
/// Days per year of the growth schedule: one year = 31,536,000 DAA seconds of 86,400 (spec 01 section 1.13.3).
pub const GROWTH_DAYS_PER_YEAR: u64 = 365;
/// The linear schedule of 1.13.3, 2 GiB at genesis plus 0.5 GiB per year, is `G x (1 + d / 1460)` for the genesis
/// size `G` and the day `d`: it doubles at day 1,460 (year 4), quadruples at day 4,380 (year 12), reaches 8x at
/// day 10,220 (year 28) and 16x at day 21,900 (year 60).
pub const GROWTH_DOUBLING_DAYS: u64 = 4 * GROWTH_DAYS_PER_YEAR;
/// The number of dataset doublings reached by day `days_since_genesis` of the chain: `floor(log2(1 + d / 1460))`,
/// in integers (`1 + d / 1460` rounded down, then its integer log2, which equals the real log2's floor because a
/// power of two is an integer). 0 until day 1,459; 1 from day 1,460 (year 4); 2 from day 4,380 (year 12).
pub fn growth_doublings(days_since_genesis: u64) -> u32 {
(1 + days_since_genesis / GROWTH_DOUBLING_DAYS).ilog2()
}
/// The cache size on day `d` under option C: 2^26 words doubled once per dataset doubling (256 MiB, 512 MiB from
/// year 4, 1 GiB from year 12).
pub fn cache_log2_words(days_since_genesis: u64) -> u32 {
CACHE_LOG2_WORDS as u32 + growth_doublings(days_since_genesis)
}
/// The dataset size on day `d` under option (b) of 1.13.3: the genesis size (2^`genesis_log2_words` words: 28 for
/// the 1 GiB packs and the devnet, 29 for the designed 2 GiB) doubled once per doubling of the linear schedule. The
/// result is capped at 32 (the item index is 32 bits, spec 1.13.3).
pub fn dataset_log2_words(genesis_log2_words: u32, days_since_genesis: u64) -> u32 {
(genesis_log2_words + growth_doublings(days_since_genesis)).min(32)
}
/// Days since genesis from two day indices of `bind::day_index` (the header's `timestamp_ms / 86,400,000`): the
/// day of the block and the day of the genesis header. A block before the genesis day (clock skew) is day 0.
pub fn days_since_genesis(day_index: u64, genesis_day_index: u64) -> u64 {
day_index.saturating_sub(genesis_day_index)
}
#[inline(always)]
fn rotl(x: u32, n: u32) -> u32 {
@ -66,17 +168,23 @@ pub fn chacha_block(x: &[u32; 16]) -> [u32; 16] {
y
}
/// Mixer parameters drawn from the day key. Draw order: ROT[0..7] (1..31), MUL[0..15] (odd), RC[0..15].
/// Mixer parameters drawn from the day key, plus the [`Shape`] the mixer is applied under. Draw order:
/// ROT[0..7] (1..31), MUL[0..15] (odd), RC[0..15]. The shape is not drawn: it is the class's.
#[derive(Clone, Debug, PartialEq, Eq)]
pub struct MixParams {
pub key: [u32; 8],
pub rot: [u32; 8],
pub mul: [u32; 16],
pub rc: [u32; 16],
pub shape: Shape,
}
impl MixParams {
/// Version 2 shape.
pub fn new(key: [u32; 8]) -> Self {
Self::with_shape(key, Shape::V2)
}
pub fn with_shape(key: [u32; 8], shape: Shape) -> Self {
let mut rng = SplitMix64::new(key[0] as u64 | ((key[1] as u64) << 32));
let mut rot = [0u32; 8];
let mut mul = [0u32; 16];
@ -90,7 +198,7 @@ impl MixParams {
for c in rc.iter_mut() {
*c = rng.next() as u32;
}
Self { key, rot, mul, rc }
Self { key, rot, mul, rc, shape }
}
/// Parameters for a day string: the key is `seed_words("day/" + day)`.
pub fn for_day(day: &str) -> Self {
@ -98,12 +206,84 @@ impl MixParams {
}
}
/// The dataset layout (era layout, `docs/plans/era-layout.md` section 1.2): word `w` of the dataset holds word
/// `j(w)` of item `t(w)`, where `j(w)` gathers the four bits of `w` at the ascending positions `pos` and `t(w)` is
/// `w` with those bits removed. [`Layout::LINEAR`] (`pos = [0, 1, 2, 3]`) is `dataset[w] = item(w >> 4)[w & 15]`,
/// the lottery hash's mapping. Every position is below 16, so the mapping is the same at every dataset size of
/// at least 2^16 words and an item keeps its value at every size.
#[derive(Clone, Copy, Debug, PartialEq, Eq, Hash)]
pub struct Layout {
pub pos: [u8; 4],
}
impl Layout {
pub const LINEAR: Layout = Layout { pos: [0, 1, 2, 3] };
pub fn is_linear(&self) -> bool {
self.pos == [0, 1, 2, 3]
}
/// Positions ascending, distinct, below 16.
pub fn is_valid(&self) -> bool {
self.pos.iter().all(|&p| p < 16) && (1..4).all(|i| self.pos[i] > self.pos[i - 1])
}
/// `(t, j)` of word index `w`.
#[inline(always)]
pub fn split(&self, w: u32) -> (u32, u32) {
if self.is_linear() {
return (w >> 4, w & 15);
}
let mut j = 0u32;
for (i, &p) in self.pos.iter().enumerate() {
j |= ((w >> p) & 1) << i;
}
// remove the highest position first so the lower ones stay where they are
let mut t = w;
for &p in self.pos.iter().rev() {
let p = p as u32;
let low = (1u32 << p) - 1;
t = (t & low) | ((t >> (p + 1)) << p);
}
(t, j)
}
/// The word index of word `j` of item `t`: the inverse of [`Layout::split`].
#[inline(always)]
pub fn join(&self, t: u32, j: u32) -> u32 {
if self.is_linear() {
return (t << 4) | (j & 15);
}
// insert the lowest position first: every later position counts the bit just inserted
let mut w = t;
for (i, &p) in self.pos.iter().enumerate() {
let p = p as u32;
let low = (1u32 << p) - 1;
w = ((w >> p) << (p + 1)) | (w & low) | (((j >> i) & 1) << p);
}
w
}
}
impl Default for Layout {
fn default() -> Self {
Layout::LINEAR
}
}
/// Round key `(r + 1) * 0x9E3779B9` mod 2^32.
#[inline(always)]
pub fn round_key(r: usize) -> u32 {
((r + 1) as u32).wrapping_mul(0x9E3779B9)
}
/// The round key of application `j` (0 <= j < m) of round `r` under multiplier `m`: `round_key(r * m + j)`. For
/// `m = 1` this is `round_key(r)`, version 2's key.
#[inline(always)]
pub fn round_key_mult(r: usize, j: usize, m: usize) -> u32 {
round_key(r * m + j)
}
/// `M_r` on 16 words in place: per word `(s ^ (RC + rk)) * MUL`, then one ChaCha-shaped double round with
/// the four column rotations `ROT[0..3]` and the four diagonal rotations `ROT[4..7]`.
#[inline(always)]
@ -122,9 +302,11 @@ pub fn mixer(s: &mut [u32; 16], rk: u32, mp: &MixParams) {
qr(s, 3, 4, 9, 14, r[4], r[5], r[6], r[7]);
}
/// The 256 MiB cache for one day key.
/// The cache for one day key: 2^log2_words words (256 MiB under version 2).
pub struct Cache {
pub key: [u32; 8],
pub log2_words: u32,
line_mask: u32,
words: Vec<u32>,
}
@ -132,6 +314,12 @@ impl Cache {
/// One segment: 64 chained lines written at `cache[seg * 1024 ..]`.
/// `in_j = prev XOR (sigma || K || seg || j || tag)`, `line_j = B(in_j)`, `prev_0 = 0`.
pub fn fill_segment(words: &mut [u32], seg: usize, key: &[u32; 8]) {
Self::fill_segment_tagged(words, seg, key, &CACHE_TAG)
}
/// [`Cache::fill_segment`] with an explicit chain tag: [`CACHE_TAG`] for the cache, [`HOT_TAG`] for the hot
/// table of `docs/plans/hot-table.md` (the same chain, another key and tag).
pub fn fill_segment_tagged(words: &mut [u32], seg: usize, key: &[u32; 8], tag: &[u32; 2]) {
let base = (seg << CACHE_SEGMENT_LOG2_LINES) * 16;
let seg_words = &mut words[base..base + CACHE_LINES_PER_SEGMENT * 16];
let mut prev = [0u32; 16];
@ -141,8 +329,8 @@ impl Cache {
x[4..12].copy_from_slice(key);
x[12] = seg as u32;
x[13] = j as u32;
x[14] = CACHE_TAG[0];
x[15] = CACHE_TAG[1];
x[14] = tag[0];
x[15] = tag[1];
for i in 0..16 {
x[i] ^= prev[i];
}
@ -152,13 +340,22 @@ impl Cache {
}
}
/// The whole cache on the calling thread: 65,536 chains of 64 ChaCha12 blocks, in segment order.
/// The version 2 cache on the calling thread: 65,536 chains of 64 ChaCha12 blocks, in segment order.
pub fn fill(key: [u32; 8]) -> Cache {
let mut words = vec![0u32; CACHE_WORDS];
for seg in 0..CACHE_SEGMENTS {
Self::fill_log2(key, CACHE_LOG2_WORDS as u32)
}
/// A cache of 2^`log2_words` words (26, 27 or 28 under the growth rule; smaller sizes for tests): 2^(log2 - 10)
/// independent chains of 64 lines, the same chain function at every size, so a larger cache's first segments
/// are the smaller cache's segments word for word.
pub fn fill_log2(key: [u32; 8], log2_words: u32) -> Cache {
assert!((10..=30).contains(&log2_words), "cache log2 words must be in 10..=30");
let shape = Shape { mixer_mult: 1, cache_log2_words: log2_words };
let mut words = vec![0u32; shape.cache_words()];
for seg in 0..shape.cache_segments() {
Self::fill_segment(&mut words, seg, &key);
}
Cache { key, words }
Cache { key, log2_words, line_mask: shape.cache_line_mask(), words }
}
pub fn for_day(day: &str) -> Cache {
@ -170,10 +367,27 @@ impl Cache {
&self.words
}
/// Cache line `a` (0 <= a < 2^22) as 16 words.
pub fn lines(&self) -> usize {
self.words.len() >> 4
}
pub fn line_mask(&self) -> u32 {
self.line_mask
}
pub fn segments(&self) -> usize {
self.lines() >> CACHE_SEGMENT_LOG2_LINES
}
/// Cache line `a` (masked to the cache's lines) as 16 words.
#[inline(always)]
pub fn line(&self, a: u32) -> &[u32] {
let o = (a & CACHE_LINE_MASK) as usize * 16;
let o = (a & self.line_mask) as usize * 16;
&self.words[o..o + 16]
}
/// [`Cache::line`] with the mask as a constant (the verifier's hot path, see [`derive_items`]).
#[inline(always)]
pub fn line_const<const LINE_MASK: u32>(&self, a: u32) -> &[u32] {
let o = (a & LINE_MASK) as usize * 16;
&self.words[o..o + 16]
}
@ -183,11 +397,114 @@ impl Cache {
}
}
/// The hot key of an epoch: `seed_words_from_bytes("igneum-hot/" || seed_bytes)`, `seed_bytes` the program seed
/// bytes before any attempt suffix, so every attempt of one epoch shares one table.
pub fn hot_key(seed_bytes: &[u8]) -> [u32; 8] {
let mut b = Vec::with_capacity(HOT_KEY_TAG.len() + seed_bytes.len());
b.extend_from_slice(HOT_KEY_TAG);
b.extend_from_slice(seed_bytes);
crate::seed::seed_words_from_bytes(&b)
}
/// Words of a hot table of `mb` MiB.
pub fn hot_words(mb: u32) -> u32 {
mb * HOT_WORDS_PER_MIB
}
/// Segments of a hot table of `mb` MiB.
pub fn hot_segments(mb: u32) -> u32 {
mb * HOT_SEGMENTS_PER_MIB
}
/// The hot index of a source word: `mulhi(src, words)`, the high 32 bits of the 64-bit product, in `[0, words)`
/// for any table size (the multiply-shift range reduction of spec 01 section 1.13.3).
#[inline(always)]
pub fn hot_index(src: u32, words: u32) -> u32 {
((src as u64 * words as u64) >> 32) as u32
}
/// The hot table `H` of one epoch (hot-table experiment): `mb` MiB of chained ChaCha12 lines under the hot key,
/// read by the hot load slots as `dst ^= H[hot_index(src, words)]`. The verifier holds it beside the cache.
pub struct HotTable {
pub key: [u32; 8],
pub mb: u32,
words: Vec<u32>,
}
impl HotTable {
/// Fill `mb` MiB under `key` on the calling thread.
pub fn fill(key: [u32; 8], mb: u32) -> HotTable {
assert!(mb >= 1 && mb <= 4096, "hot table size in MiB out of range");
let n = hot_words(mb) as usize;
let mut words = vec![0u32; n];
for seg in 0..hot_segments(mb) as usize {
Cache::fill_segment_tagged(&mut words, seg, &key, &HOT_TAG);
}
HotTable { key, mb, words }
}
/// The table of the epoch whose program seed bytes are `seed_bytes`.
pub fn for_seed_bytes(seed_bytes: &[u8], mb: u32) -> HotTable {
Self::fill(hot_key(seed_bytes), mb)
}
#[inline(always)]
pub fn n_words(&self) -> u32 {
self.words.len() as u32
}
/// `H[i]`.
#[inline(always)]
pub fn at(&self, i: u32) -> u32 {
self.words[i as usize]
}
/// `H[hot_index(src, words)]`: what a hot load reads for source word `src`.
#[inline(always)]
pub fn word(&self, src: u32) -> u32 {
self.words[hot_index(src, self.n_words()) as usize]
}
#[inline(always)]
pub fn words(&self) -> &[u32] {
&self.words
}
/// FNV-1a 64 over the table as little-endian bytes (what `vectors.h` carries as `IGNEUM_HOT_FNV64`).
pub fn fnv1a64(&self) -> u64 {
fnv1a64_words(&self.words)
}
}
/// Derive `ts.len()` items into `out`, all chains interleaved round by round so the cache-line misses of
/// independent items overlap in the memory system (`deriveItems` in the Swift).
/// independent items overlap in the memory system (`deriveItems` in the Swift). Under multiplier `m`
/// (`mp.shape.mixer_mult`) round `r` applies `M` with keys `round_key(r m + j)` for `j = 0 .. m - 1` before its
/// one cache read; the final mixer applies `M` with keys `round_key(8 m + j)`. `m = 1` is version 2.
pub fn derive_items(ts: &[u32], mp: &MixParams, cache: &Cache, out: &mut [[u32; 16]]) {
// The item loop lives in its own function, one instance per cache size the growth rule can reach with the line
// mask a constant, never inlined into the callers. Inlined into `MemhardCpu::fetch` it ran at 1.33 ms per unit
// against 0.61 out of line (the version 2 verifier, bisected on one core under the measure lock, 5 October 2026,
// `docs/plans/mixer-x4.md` section 6.6: the constant mask alone, or the mask hoisted into a local, or the
// constant with the loop still inlined, all stayed at 1.33; the out-of-line instances read 0.60 to 0.62). Any
// other cache size (tests) takes the instance with the run-time mask.
match cache.log2_words {
26 => derive_items_mask::<{ (1u32 << 22) - 1 }>(ts, mp, cache, out),
27 => derive_items_mask::<{ (1u32 << 23) - 1 }>(ts, mp, cache, out),
28 => derive_items_mask::<{ (1u32 << 24) - 1 }>(ts, mp, cache, out),
29 => derive_items_mask::<{ (1u32 << 25) - 1 }>(ts, mp, cache, out),
30 => derive_items_mask::<{ (1u32 << 26) - 1 }>(ts, mp, cache, out),
_ => derive_items_mask::<0>(ts, mp, cache, out),
}
}
/// [`derive_items`] with the cache line mask as a constant (`LINE_MASK = 0`: the cache's own run-time mask). Kept
/// out of line on purpose (see [`derive_items`]).
#[inline(never)]
fn derive_items_mask<const LINE_MASK: u32>(ts: &[u32], mp: &MixParams, cache: &Cache, out: &mut [[u32; 16]]) {
let n = ts.len();
debug_assert!(out.len() >= n);
debug_assert!(LINE_MASK == 0 || LINE_MASK == cache.line_mask);
let m = mp.shape.mixer_mult as usize;
for k in 0..n {
let s = &mut out[k];
let t = ts[k];
@ -197,20 +514,24 @@ pub fn derive_items(ts: &[u32], mp: &MixParams, cache: &Cache, out: &mut [[u32;
}
}
for r in 0..ITEM_ROUNDS {
let rk = round_key(r);
for s in out[..n].iter_mut() {
mixer(s, rk, mp);
for j in 0..m {
let rk = round_key_mult(r, j, m);
for s in out[..n].iter_mut() {
mixer(s, rk, mp);
}
}
for s in out[..n].iter_mut() {
let line = cache.line(s[0]);
let line = if LINE_MASK != 0 { cache.line_const::<LINE_MASK>(s[0]) } else { cache.line(s[0]) };
for i in 0..16 {
s[i] ^= line[i];
}
}
}
let rk = round_key(ITEM_ROUNDS);
for s in out[..n].iter_mut() {
mixer(s, rk, mp);
for j in 0..m {
let rk = round_key_mult(ITEM_ROUNDS, j, m);
for s in out[..n].iter_mut() {
mixer(s, rk, mp);
}
}
}
@ -221,7 +542,7 @@ pub fn derive_item(t: u32, mp: &MixParams, cache: &Cache) -> [u32; 16] {
out[0]
}
/// The CPU verifier's view of the memory-hard dataset: the mixer parameters and the 256 MiB cache.
/// The CPU verifier's view of the memory-hard dataset: the mixer parameters (with the shape) and the cache.
pub struct MemhardCpu {
pub params: MixParams,
pub cache: Cache,
@ -231,26 +552,41 @@ pub struct MemhardCpu {
pub const FETCH_MAX: usize = 64;
impl MemhardCpu {
/// Version 2 shape.
pub fn new(key: [u32; 8]) -> Self {
Self { params: MixParams::new(key), cache: Cache::fill(key) }
Self::with_shape(key, Shape::V2)
}
pub fn with_shape(key: [u32; 8], shape: Shape) -> Self {
Self { params: MixParams::with_shape(key, shape), cache: Cache::fill_log2(key, shape.cache_log2_words) }
}
pub fn for_day(day: &str) -> Self {
Self::new(day_key(day))
}
/// `dataset[w] = item(w >> 4)[w & 15]`.
pub fn shape(&self) -> Shape {
self.params.shape
}
/// `dataset[w] = item(w >> 4)[w & 15]` (the linear layout).
pub fn word(&self, w: u32) -> u32 {
derive_item(w >> 4, &self.params, &self.cache)[(w & 15) as usize]
self.word_at(Layout::LINEAR, w)
}
/// `dataset[w] = item(t(w))[j(w)]` under `layout` (era layout; the layout is the program's, the cache the
/// day's, so one cache serves every era of a day).
pub fn word_at(&self, layout: Layout, w: u32) -> u32 {
let (t, j) = layout.split(w);
derive_item(t, &self.params, &self.cache)[j as usize]
}
/// `out[k] = dataset[idx[k]]` for every k, `idx.len() <= FETCH_MAX`. Equal items are derived once.
/// Returns the number of distinct items derived.
pub fn fetch(&self, idx: &[u32], out: &mut [u32]) -> usize {
pub fn fetch(&self, idx: &[u32], out: &mut [u32], layout: Layout) -> usize {
let n = idx.len();
assert!(n <= FETCH_MAX && out.len() >= n);
let mut uniq = [0u32; FETCH_MAX];
let mut slot = [0u8; FETCH_MAX];
let mut word = [0u8; FETCH_MAX];
let mut u = 0usize;
for k in 0..n {
let t = idx[k] >> 4;
let (t, j) = layout.split(idx[k]);
word[k] = j as u8;
let found = uniq[..u].iter().position(|&x| x == t);
let j = match found {
Some(j) => j,
@ -265,7 +601,39 @@ impl MemhardCpu {
let mut items = [[0u32; 16]; FETCH_MAX];
derive_items(&uniq[..u], &self.params, &self.cache, &mut items);
for k in 0..n {
out[k] = items[slot[k] as usize][(idx[k] & 15) as usize];
out[k] = items[slot[k] as usize][word[k] as usize];
}
u
}
/// `out[k][j] = dataset[base[k] + j]` for `j < width` (read-width experiment): `base[k]` is aligned to `width`
/// words and the layout's low `log2(width)` positions are the identity, so every lane's words lie in one item
/// at consecutive word offsets, derived once per distinct item. Returns the distinct items.
pub fn fetch_wide(&self, base: &[u32], width: usize, out: &mut [[u32; 16]], layout: Layout) -> usize {
let n = base.len();
assert!(n <= FETCH_MAX && out.len() >= n && width <= 16);
debug_assert!((0..width.trailing_zeros() as usize).all(|i| layout.pos[i] == i as u8), "a wide load needs the identity on its low positions");
let mut uniq = [0u32; FETCH_MAX];
let mut slot = [0u8; FETCH_MAX];
let mut word = [0u8; FETCH_MAX];
let mut u = 0usize;
for k in 0..n {
let (t, j0) = layout.split(base[k]);
word[k] = j0 as u8;
let j = match uniq[..u].iter().position(|&x| x == t) {
Some(j) => j,
None => {
uniq[u] = t;
u += 1;
u - 1
}
};
slot[k] = j as u8;
}
let mut items = [[0u32; 16]; FETCH_MAX];
derive_items(&uniq[..u], &self.params, &self.cache, &mut items);
for k in 0..n {
let o = word[k] as usize;
out[k][..width].copy_from_slice(&items[slot[k] as usize][o..o + width]);
}
u
}
@ -285,6 +653,40 @@ mod tests {
assert_eq!(mp.rc[0], 0xbab68293);
assert_eq!(mp.rc[15], 0x31b49ee2);
assert!(mp.mul.iter().all(|m| m & 1 == 1));
assert_eq!(mp.shape, Shape::V2);
}
/// Era layout: split and join are inverse, the linear layout is today's mapping, and an interleaved layout
/// keeps every position below 16 so the mapping is the same at every size of at least 2^16 words.
#[test]
fn layout_split_join() {
let lin = Layout::LINEAR;
assert!(lin.is_linear() && lin.is_valid());
for w in [0u32, 1, 15, 16, 17, 0x0fff_ffff, 0xffff_ffff] {
assert_eq!(lin.split(w), (w >> 4, w & 15));
assert_eq!(lin.join(w >> 4, w & 15), w);
}
let l = Layout { pos: [0, 1, 7, 12] };
assert!(!l.is_linear() && l.is_valid());
for w in [0u32, 1, 2, 3, 4, 127, 128, 129, 4095, 4096, 0x0fff_ffff, 0x1234_5678, 0xffff_ffff] {
let (t, j) = l.split(w);
assert!(j < 16);
assert_eq!(l.join(t, j), w, "w {w:#x}");
}
// bits: j0 = bit 0, j1 = bit 1, j2 = bit 7, j3 = bit 12; t = the other 28 bits in order
assert_eq!(l.split(0b1_0000_0000_0000), (0, 8));
assert_eq!(l.split(1 << 7), (0, 4));
assert_eq!(l.split(0b100), (1, 0));
// every t in 0..2^(D-4) appears exactly once among w < 2^D (D = 16), with every j
let mut seen = vec![0u32; 1 << 12];
for w in 0..(1u32 << 16) {
let (t, j) = l.split(w);
seen[t as usize] |= 1 << j;
}
assert!(seen.iter().all(|&s| s == 0xffff));
assert!(!Layout { pos: [0, 1, 1, 5] }.is_valid());
assert!(!Layout { pos: [0, 1, 2, 16] }.is_valid());
assert!(!Layout { pos: [1, 0, 2, 3] }.is_valid());
}
#[test]
@ -296,6 +698,40 @@ mod tests {
assert_eq!(y, z);
}
/// Hot-table experiment: the genesis epoch's table (seed bytes "igneum-genesis") as the hot packs carry it
/// (`proto-cuda/packs-ca2-hot/hot32k4/vectors.json`: hot_head, hot_fnv1a64; the head is the same at every size,
/// a larger table is more segments). The index mapping stays inside the table for any size.
#[test]
fn hot_table_fill_vector_and_index() {
assert_ne!(HOT_TAG, CACHE_TAG);
let h = HotTable::for_seed_bytes(b"igneum-genesis", 32);
assert_eq!(h.n_words(), 1 << 23);
assert_eq!(hot_segments(32), 8192);
assert_eq!(
&h.words()[..16],
&[
0x8068cc73, 0x6036ebf9, 0xb604cd25, 0x8ffb840e, 0xc54074a2, 0x285c0695, 0x77512425, 0xc26a58a7,
0x72c88757, 0xc10fca78, 0x513825dd, 0x30d6ccc8, 0x9a05e7cf, 0xb9533f50, 0x4bac3ba0, 0xa5c19528
]
);
assert_eq!(h.fnv1a64(), 0xc1767ba3ef02719f, "hot32k4 pack, hot_fnv1a64");
assert_eq!(h.key, hot_key(b"igneum-genesis"));
assert_ne!(h.key, day_key("2026-10-03"));
// a different seed, a different table; the same seed under the cache tag is not the hot table
assert_ne!(HotTable::for_seed_bytes(b"igneum-genesis\x01\x00\x00\x00", 1).words()[..16], h.words()[..16]);
let mut under_cache_tag = vec![0u32; 1024];
Cache::fill_segment(&mut under_cache_tag, 0, &h.key);
assert_ne!(&under_cache_tag[..16], &h.words()[..16]);
for words in [hot_words(32), hot_words(64), hot_words(96)] {
assert_eq!(hot_index(0, words), 0);
assert!(hot_index(u32::MAX, words) < words);
assert_eq!(hot_index(u32::MAX, words), words - 1);
assert!(hot_index(0x8000_0000, words) == words / 2);
}
assert_eq!(hot_index(0x1234_5678, 1 << 24), 0x1234_5678 >> 8);
assert_eq!(h.word(0x8000_0000), h.at(1 << 22));
}
#[test]
fn first_cache_line_matches_pack() {
// vectors.json cache_head for day 2026-10-03: segment 0, line 0, with prev = 0.
@ -310,4 +746,96 @@ mod tests {
]
);
}
/// Option C: the schedule table of `docs/plans/mixer-x4.md` (day -> doublings, cache words, dataset words at a
/// 2^28 genesis). The doublings fall at years 4 and 12 exactly, never a day early.
#[test]
fn growth_schedule_table() {
let table: [(u64, u32, u32, u32); 12] = [
(0, 0, 26, 28),
(1, 0, 26, 28),
(365, 0, 26, 28),
(1_459, 0, 26, 28),
(1_460, 1, 27, 29),
(2_920, 1, 27, 29),
(4_379, 1, 27, 29),
(4_380, 2, 28, 30),
(10_219, 2, 28, 30),
(10_220, 3, 29, 31),
(21_900, 4, 30, 32),
(100_000, 6, 32, 32),
];
for (d, k, c, s) in table {
assert_eq!(growth_doublings(d), k, "day {d}");
assert_eq!(cache_log2_words(d), c, "day {d}");
assert_eq!(dataset_log2_words(28, d), s, "day {d}");
}
// the designed 2 GiB genesis: 2^29 words, 2^30 at year 4, 2^31 at year 12
assert_eq!(dataset_log2_words(29, 0), 29);
assert_eq!(dataset_log2_words(29, 1_460), 30);
assert_eq!(dataset_log2_words(29, 4_380), 31);
// the linear schedule itself: 2 GiB x (1 + d / 1460) crosses 4 GiB at day 1,460 and 8 GiB at day 4,380
for d in [1_459u64, 1_460, 4_379, 4_380] {
let bytes = 2u64 * (1 << 30) + (1u64 << 29) * d / 365;
let k = (bytes / (2u64 << 30)).ilog2();
assert_eq!(growth_doublings(d), k, "day {d}: linear {bytes} bytes");
}
assert_eq!(days_since_genesis(20_730, 20_729), 1);
assert_eq!(days_since_genesis(20_729, 20_729), 0);
assert_eq!(days_since_genesis(20_000, 20_729), 0);
let v2 = Shape::for_class_day(&LoadClass::V2, 100_000);
assert_eq!(v2, Shape::V2);
let v3 = Shape::for_class_day(&LoadClass::MX4, 0);
assert_eq!(v3, Shape { mixer_mult: 4, cache_log2_words: 26 });
assert_eq!(Shape::for_class_day(&LoadClass::MX4, 1_460).cache_log2_words, 27);
assert_eq!(v3.mixers_per_item(), 36);
assert_eq!(Shape::V2.mixers_per_item(), 9);
assert_eq!(Shape::V2.cache_segments(), CACHE_SEGMENTS);
assert_eq!(Shape::V2.cache_line_mask(), CACHE_LINE_MASK);
assert_eq!(Shape::V2.log2_segments(), 16);
}
/// The multiplied mixer, restated by hand on a small cache: `m` applications with keys `round_key(r m + j)`
/// before every read, the same 8 reads; `m = 1` is `derive_item` of version 2 word for word; a larger cache's
/// first segments equal the smaller cache's.
#[test]
fn mixer_mult_by_hand() {
let key = day_key("2026-10-03");
let small = Cache::fill_log2(key, 16);
let big = Cache::fill_log2(key, 18);
assert_eq!(&big.words()[..small.words().len()], small.words());
assert_eq!(small.segments(), 64);
assert_eq!(small.line_mask(), 4095);
for m in [1u32, 2, 4] {
let mp = MixParams::with_shape(key, Shape { mixer_mult: m, cache_log2_words: 16 });
for t in [0u32, 1, 12_345, u32::MAX] {
let got = derive_item(t, &mp, &small);
let mut s = [0u32; 16];
s[..8].copy_from_slice(&key);
for i in 0..8 {
s[8 + i] = t.wrapping_mul(mp.mul[i]).wrapping_add(mp.rc[i]);
}
for r in 0..8usize {
for j in 0..m as usize {
mixer(&mut s, round_key(r * m as usize + j), &mp);
}
let line = small.line(s[0]);
for i in 0..16 {
s[i] ^= line[i];
}
}
for j in 0..m as usize {
mixer(&mut s, round_key(8 * m as usize + j), &mp);
}
assert_eq!(got, s, "m {m} t {t}");
}
}
let v2 = MixParams::with_shape(key, Shape { mixer_mult: 1, cache_log2_words: 16 });
let v3 = MixParams::with_shape(key, Shape { mixer_mult: 4, cache_log2_words: 16 });
assert_ne!(derive_item(0, &v2, &small), derive_item(0, &v3, &small));
assert_eq!(round_key_mult(0, 0, 1), round_key(0));
assert_eq!(round_key_mult(8, 0, 1), round_key(8));
assert_eq!(round_key_mult(2, 3, 4), round_key(11));
assert_eq!(round_key_mult(8, 3, 4), round_key(35));
}
}

View file

@ -10,7 +10,7 @@
//! a later attempt. This module is the one place that rule is written in Rust; the miner checks every pack it
//! writes with it before a worker sees the pack, and the tests pin the attempt vectors the C side also pins.
use crate::generator::attempt_words;
use crate::generator::{attempt_words, ProgramClass};
use crate::seed::seed_words_from_bytes;
use std::fmt;
use std::path::Path;
@ -23,6 +23,12 @@ pub struct PackIdentity {
pub attempt: u32,
pub seedw: [u32; 8],
pub keyw: [u32; 8],
/// `IGNEUM_GENERATOR` (2 or 3; a pack without the line is generator 1, which no worker runs).
pub generator: u32,
/// The program class the generator version names (Counter ASIC 2.0).
pub class: ProgramClass,
/// `IGNEUM_ERA_SEED_HEX` when the pack carries one (class v3 chain packs).
pub era_hex: Option<String>,
}
/// Why a pack is not the one a worker should mine with. `Display` is the plain-words line the logs carry.
@ -35,6 +41,9 @@ pub enum PackFault {
/// The files of one pack contradict each other (seeds.txt against program.h, or the init words against the
/// seeds and the attempt): a half-written or hand-edited pack, or a worker and an exporter on different rules.
Disagree(String),
/// The pack is of another program class than the one the chain is on (spec 01 section 1.4.5: an implementation
/// refuses a pack whose generator version is not its own), or its era seed is not the era the job names.
WrongClass(String),
}
impl fmt::Display for PackFault {
@ -50,6 +59,7 @@ impl fmt::Display for PackFault {
day_label(want_day)
),
PackFault::Disagree(w) => write!(f, "program pack and its seeds disagree: {w}"),
PackFault::WrongClass(w) => write!(f, "program pack of the wrong class: {w}"),
}
}
}
@ -151,6 +161,44 @@ fn seeds_line(text: &str, key: &str) -> Option<String> {
/// Checks the texts of a pack (program.h, and seeds.txt when it exists) against the seeds a worker will be asked
/// to mine with. Pure: the miner and the tests call it with file contents.
pub fn verify_pack_texts(program_h: &str, seeds_txt: Option<&str>, want_epoch: &[u8], want_day: &[u8]) -> Result<PackIdentity, PackFault> {
verify_pack_texts_chain(program_h, seeds_txt, want_epoch, want_day, None, None)
}
/// [`verify_pack_texts`] that also demands a program class and, for class v3, the era seed the chain is on
/// (Counter ASIC 2.0, 5 October 2026). `want_class` `None` accepts either class; `want_era` `None` skips the era.
/// A pack whose `IGNEUM_GENERATOR` is neither 2 nor 3 is refused whatever is wanted.
pub fn verify_pack_texts_chain(
program_h: &str,
seeds_txt: Option<&str>,
want_epoch: &[u8],
want_day: &[u8],
want_class: Option<ProgramClass>,
want_era: Option<&[u8]>,
) -> Result<PackIdentity, PackFault> {
let generator = define_u32(program_h, "IGNEUM_GENERATOR").unwrap_or(1);
let Some(class) = ProgramClass::from_generator(generator) else {
return Err(PackFault::WrongClass(format!("IGNEUM_GENERATOR {generator} is not a generator version this software runs (2 or 3)")));
};
// IGNEUM_PROGRAM_CLASS, when present, must name the class the generator version names
if let Some(named) = define_str(program_h, "IGNEUM_PROGRAM_CLASS") {
if ProgramClass::parse(&named) != Some(class) {
return Err(PackFault::Disagree(format!("IGNEUM_PROGRAM_CLASS {named:?} does not match IGNEUM_GENERATOR {generator}")));
}
}
let era_hex = define_str(program_h, "IGNEUM_ERA_SEED_HEX").map(|h| h.to_ascii_lowercase());
if let Some(want) = want_class {
if want != class {
return Err(PackFault::WrongClass(format!("the pack is program class {} (generator {generator}), the chain is on class {}", class.name(), want.name())));
}
}
if let (Some(want), ProgramClass::V3) = (want_era, class) {
let want_hex = hex(want);
match &era_hex {
Some(h) if *h == want_hex => {}
Some(h) => return Err(PackFault::WrongClass(format!("the pack's era seed {} is not the era seed {} the job names", short(h), short(&want_hex)))),
None => return Err(PackFault::WrongClass("a class v3 pack without IGNEUM_ERA_SEED_HEX; the job names an era seed".into())),
}
}
let seedw = define_words(program_h, "IGNEUM_SEEDW_INIT").ok_or_else(|| PackFault::Unreadable("program.h has no IGNEUM_SEEDW_INIT with 8 words".into()))?;
let keyw = define_words(program_h, "IGNEUM_KEY_INIT").ok_or_else(|| PackFault::Unreadable("program.h has no IGNEUM_KEY_INIT with 8 words".into()))?;
let attempt = define_u32(program_h, "IGNEUM_PROGRAM_ATTEMPT").unwrap_or(0);
@ -194,14 +242,19 @@ pub fn verify_pack_texts(program_h: &str, seeds_txt: Option<&str>, want_epoch: &
if epoch_hex != want_epoch_hex || day_hex != want_day_hex {
return Err(PackFault::OutOfDate { pack_epoch: epoch_hex, pack_day: day_hex, want_epoch: want_epoch_hex, want_day: want_day_hex });
}
Ok(PackIdentity { epoch_hex, day_hex, attempt, seedw, keyw })
Ok(PackIdentity { epoch_hex, day_hex, attempt, seedw, keyw, generator, class, era_hex })
}
/// [`verify_pack_texts`] over a pack directory.
pub fn verify_pack_dir(dir: &Path, want_epoch: &[u8], want_day: &[u8]) -> Result<PackIdentity, PackFault> {
verify_pack_dir_chain(dir, want_epoch, want_day, None, None)
}
/// [`verify_pack_texts_chain`] over a pack directory.
pub fn verify_pack_dir_chain(dir: &Path, want_epoch: &[u8], want_day: &[u8], want_class: Option<ProgramClass>, want_era: Option<&[u8]>) -> Result<PackIdentity, PackFault> {
let program_h = std::fs::read_to_string(dir.join("program.h")).map_err(|e| PackFault::Unreadable(format!("cannot read {}/program.h: {e}", dir.display())))?;
let seeds = std::fs::read_to_string(dir.join("seeds.txt")).ok();
verify_pack_texts(&program_h, seeds.as_deref(), want_epoch, want_day)
verify_pack_texts_chain(&program_h, seeds.as_deref(), want_epoch, want_day, want_class, want_era)
}
#[cfg(test)]
@ -301,4 +354,44 @@ mod tests {
assert_eq!(day_label(DAY_20731), "20731");
assert_eq!(day_label("abcd"), "abcd");
}
/// Counter ASIC 2.0: a class v3 chain pack carries generator 3, the class line and the era seed; it is refused
/// when the chain wants class v2, when the era differs, and a v2 pack is refused when the chain wants v3; a
/// generator this software does not run is refused whatever is wanted.
#[test]
fn program_class_and_era_are_checked() {
let e = bytes(EPOCH_34);
let d = bytes(DAY_20731);
let era = [0x5au8; 32];
let v3 = Epoch::from_chain_seeds(&e, &d, Some(&era), ProgramClass::V3, "class test");
let h3 = program_header(&v3.program, "test day", &v3.dataset);
assert!(h3.contains("#define IGNEUM_GENERATOR 3\n"));
assert!(h3.contains("#define IGNEUM_PROGRAM_CLASS \"v3\"\n"));
assert!(h3.contains(&format!("#define IGNEUM_ERA_SEED_HEX \"{}\"\n", hex(&era))));
let id = verify_pack_texts_chain(&h3, None, &e, &d, Some(ProgramClass::V3), Some(&era)).unwrap();
assert_eq!((id.generator, id.class, id.era_hex.as_deref()), (3, ProgramClass::V3, Some(hex(&era).as_str())));
assert_eq!(id.attempt, v3.program.attempt);
assert!(verify_pack_texts(&h3, None, &e, &d).is_ok(), "no class wanted: either class passes");
let err = verify_pack_texts_chain(&h3, None, &e, &d, Some(ProgramClass::V2), None).unwrap_err();
assert!(matches!(err, PackFault::WrongClass(_)), "{err}");
assert!(err.to_string().contains("program pack of the wrong class"), "{err}");
let err = verify_pack_texts_chain(&h3, None, &e, &d, Some(ProgramClass::V3), Some(&[1u8; 32])).unwrap_err();
assert!(err.to_string().contains("era seed"), "{err}");
// the v2 pack of the same seeds: generator 2, no class line, no era line, refused when v3 is wanted
let v2 = Epoch::from_chain_seeds(&e, &d, Some(&era), ProgramClass::V2, "class test");
let h2 = program_header(&v2.program, "test day", &v2.dataset);
assert!(h2.contains("#define IGNEUM_GENERATOR 2\n"));
assert!(!h2.contains("IGNEUM_PROGRAM_CLASS") && !h2.contains("IGNEUM_ERA_SEED_HEX"));
let plain = Epoch::from_seed_bytes(&e, &d, "class test");
assert_eq!(program_header(&plain.program, "test day", &plain.dataset), h2, "class v2 from the chain is the v2 export byte for byte");
let id = verify_pack_texts_chain(&h2, None, &e, &d, Some(ProgramClass::V2), Some(&era)).unwrap();
assert_eq!((id.generator, id.class, id.era_hex), (2, ProgramClass::V2, None));
assert!(matches!(verify_pack_texts_chain(&h2, None, &e, &d, Some(ProgramClass::V3), None), Err(PackFault::WrongClass(_))));
// a generator nobody runs
let h9 = h2.replace("#define IGNEUM_GENERATOR 2\n", "#define IGNEUM_GENERATOR 9\n");
assert!(matches!(verify_pack_texts(&h9, None, &e, &d), Err(PackFault::WrongClass(_))));
// a class line that contradicts the generator
let bad = h3.replace("#define IGNEUM_PROGRAM_CLASS \"v3\"\n", "#define IGNEUM_PROGRAM_CLASS \"v2\"\n");
assert!(matches!(verify_pack_texts(&bad, None, &e, &d), Err(PackFault::Disagree(_))));
}
}

View file

@ -1,10 +1,137 @@
//! The CPU reference interpreter for one 32-lane warp (`cpuWarpTraced` in the Swift) and the API the node
//! calls. Dataset words come from the memory-hard cache (default) or from the closed form (old packs).
use crate::generator::{generate, Instr, Op, Program, ITERATIONS, LANES};
use crate::memhard::MemhardCpu;
use crate::generator::{generate, generate_class, EraParams, Instr, LoadClass, Op, Program, ProgramClass, ITERATIONS, LANES};
use crate::memhard::{hot_index, HotTable, Layout, MemhardCpu, Shape};
use crate::seed::day_key;
/// The load address of an era program (`docs/plans/era-layout.md` section 1.3): `y = rotl(x * M, R)`, then the
/// window of the load site, `k = min(win, D - 26)` (0 when `D <= 26`), `idx = ((y & (MASK >> k)) | ((off &
/// (2^k - 1)) << (D - k))) & MASK`. For every other class `idx = x & MASK`, the lottery hash's address. `mask` is
/// `2^D - 1`. The acceptance mirror calls this at the rule's constant `D = 28`.
#[inline(always)]
pub fn load_index(era: Option<&EraParams>, ins: &Instr, x: u32, mask: u32, log2: u32) -> u32 {
match era {
None => x & mask,
Some(e) => {
let (wm, off) = window(ins, mask, log2);
let y = x.wrapping_mul(e.stride_mul).rotate_left(e.stride_rot);
((y & wm) | off) & mask
}
}
}
/// The window of a load site at a dataset of `2^log2` words: `(window mask, offset)` such that
/// `idx = (y & window mask) | offset` lies in the site's aligned window of `2^(log2 - k)` words.
#[inline(always)]
pub fn window(ins: &Instr, mask: u32, log2: u32) -> (u32, u32) {
let k = (ins.win as u32).min(log2.saturating_sub(26));
let wm = mask >> k;
let off = ((ins.off as u32) & ((1u32 << k) - 1)) << (log2 - k);
(wm, off)
}
/// Read-width experiment (5 October 2026): a `load` of `W` words folds every word into `dst`:
/// `x = dst XOR w[0]; for j in 1..W: x = (rotl(x, FOLD_ROT) * FOLD_MUL) XOR w[j]; dst = x`. For `W = 1` this is the
/// lottery hash's `dst XOR dataset[...]`. The fold is state-dependent (the rotate-multiply sits between the words),
/// so no function of the line alone replaces it: two different lines give two different maps of `dst`, and a
/// dataset of folded lines cannot be stored in place of the dataset (see `docs/plans/read-width.md`).
pub const FOLD_ROT: u32 = 11;
pub const FOLD_MUL: u32 = 0x9E3779B1;
/// The fold of `words` into `dst` (at least one word).
#[inline(always)]
pub fn fold_words(dst: u32, words: &[u32]) -> u32 {
let mut x = dst ^ words[0];
for &w in &words[1..] {
x = x.rotate_left(FOLD_ROT).wrapping_mul(FOLD_MUL) ^ w;
}
x
}
/// Variant 5 (scratch): the fill value of word `j` (0..2) of slot `slot` of lane `lane` of the unit at base nonce
/// `base`, under program seed words `seed`. The scratch of a unit starts as these values; a slot written during
/// the unit's hash holds what was written. Mirrored as `scr_fill` in every emitted kernel.
#[inline(always)]
pub fn scratch_fill(seed: &[u32; 8], base: u32, lane: u32, slot: u32, j: u32) -> u32 {
splitmix32(
(base.wrapping_add(lane) ^ seed[j as usize])
.wrapping_add(slot.wrapping_mul(0x9E3779B1))
.wrapping_add((j + 1).wrapping_mul(0x85EBCA77)),
)
}
/// Variant 5: the 16-byte slot after a read-modify-write that read `w` and folded to `x`: `(x ^ w1, rotl(x, 7) ^ w2,
/// x + w0)` behind the slot's tag.
#[inline(always)]
pub fn scratch_rewrite(x: u32, w: &[u32; 3]) -> [u32; 3] {
[x ^ w[1], x.rotate_left(7) ^ w[2], x.wrapping_add(w[0])]
}
/// The CPU model of one unit's scratch (variant 5): per lane, the written slots and their words. Unwritten slots
/// read as [`scratch_fill`]. A unit touches at most `scratch ops x 32` slots; a GPU keeps the real scratch per
/// resident warp with a per-unit tag per slot.
pub struct ScratchModel {
slots: usize,
written: Vec<bool>,
data: Vec<[u32; 3]>,
pub reads: usize,
pub writes: usize,
/// Soundness tests (`tests/scratch.rs`, `docs/analysis/scratch-soundness.md`): when `Some`, every
/// read-modify-write is appended as it happened. `None` on every verification path.
pub trace: Option<Vec<ScratchEvent>>,
}
/// One scratch read-modify-write as the interpreter saw it (variant 5 soundness tests).
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
pub struct ScratchEvent {
pub lane: u8,
pub slot: u32,
/// The slot had been written earlier in this unit (a re-hit): the words read were a rewrite, not the fill.
pub hit: bool,
pub read: [u32; 3],
/// The fold result, the new value of `dst`.
pub x: u32,
pub written: [u32; 3],
}
impl ScratchModel {
pub fn new(slots_per_lane: usize) -> Self {
Self {
slots: slots_per_lane,
written: vec![false; LANES * slots_per_lane],
data: vec![[0; 3]; LANES * slots_per_lane],
reads: 0,
writes: 0,
trace: None,
}
}
/// Read slot `slot` of `lane`, then rewrite it from the fold result `x`. Returns the three words read.
#[inline]
pub fn rmw(&mut self, seed: &[u32; 8], base: u32, lane: usize, slot: u32, dst: u32) -> u32 {
let i = lane * self.slots + slot as usize;
let w = if self.written[i] {
self.data[i]
} else {
[
scratch_fill(seed, base, lane as u32, slot, 0),
scratch_fill(seed, base, lane as u32, slot, 1),
scratch_fill(seed, base, lane as u32, slot, 2),
]
};
let x = fold_words(dst, &w);
let out = scratch_rewrite(x, &w);
if let Some(t) = self.trace.as_mut() {
t.push(ScratchEvent { lane: lane as u8, slot, hit: self.written[i], read: w, x, written: out });
}
self.data[i] = out;
self.written[i] = true;
self.reads += 1;
self.writes += 1;
x
}
}
/// Dataset element, closed form of (day words, index). The original prototype's six-operation element.
#[inline(always)]
pub fn dataset_elem(i: u32, d0: u32, d1: u32) -> u32 {
@ -65,24 +192,55 @@ pub struct DatasetSource {
/// in packs so any implementation can rebuild the key. Empty when the key was given directly.
pub key_bytes: Vec<u8>,
pub dataset: Dataset,
/// The hot table of the epoch (hot-table experiment, `docs/plans/hot-table.md`): `Some` when the program's
/// class has one; filled by [`Epoch::new_class`] and [`Epoch::from_seed_bytes_class`] from the program's seed
/// bytes. A hot load reads `hot[hot_index(src, words)]`.
pub hot: Option<HotTable>,
}
impl DatasetSource {
/// Build the source for a day. Memory-hard mode fills the 256 MiB cache on the calling thread.
pub fn new(day: &str, mode: DatasetMode, log2_words: u32) -> Self {
let mut ds = Self::from_key(day_key(day), mode, log2_words);
Self::new_shape(day, mode, log2_words, Shape::V2)
}
/// [`DatasetSource::new`] with the construction's shape (mixer multiplier, cache size; Counter ASIC 2.0).
pub fn new_shape(day: &str, mode: DatasetMode, log2_words: u32, shape: Shape) -> Self {
let mut ds = Self::from_key_shape(day_key(day), mode, log2_words, shape);
ds.key_bytes = format!("day/{day}").into_bytes();
ds
}
pub fn from_key(key: [u32; 8], mode: DatasetMode, log2_words: u32) -> Self {
Self::from_key_shape(key, mode, log2_words, Shape::V2)
}
/// [`DatasetSource::from_key`] with the construction's shape. Memory-hard mode fills a cache of
/// `2^shape.cache_log2_words` words on the calling thread.
pub fn from_key_shape(key: [u32; 8], mode: DatasetMode, log2_words: u32, shape: Shape) -> Self {
assert!((4..=32).contains(&log2_words), "dataset log2 must be in 4..=32");
let mask = if log2_words == 32 { u32::MAX } else { (1u32 << log2_words) - 1 };
let dataset = match mode {
DatasetMode::ClosedForm => Dataset::ClosedForm { d0: key[0], d1: key[1] },
DatasetMode::MemoryHard => Dataset::MemoryHard(MemhardCpu::new(key)),
DatasetMode::MemoryHard => Dataset::MemoryHard(MemhardCpu::with_shape(key, shape)),
};
Self { log2_words, mask, key, key_bytes: Vec::new(), dataset }
Self { log2_words, mask, key, key_bytes: Vec::new(), dataset, hot: None }
}
/// This source with the hot table of the epoch whose program seed bytes are `seed_bytes` (`mb` MiB).
pub fn with_hot(mut self, seed_bytes: &[u8], mb: u32) -> Self {
self.hot = Some(HotTable::for_seed_bytes(seed_bytes, mb));
self
}
/// The hot table of a program's class, filled from its seed bytes (none for a class without one).
pub fn attach_hot_for(&mut self, program: &Program) {
self.hot = program.class.hot.map(|h| HotTable::for_seed_bytes(&program.seed_bytes, h.mb as u32));
}
/// The shape of the memory-hard construction ([`Shape::V2`] for the closed form, which has none).
pub fn shape(&self) -> Shape {
self.memhard().map(|m| m.shape()).unwrap_or(Shape::V2)
}
pub fn mode(&self) -> DatasetMode {
@ -99,18 +257,23 @@ impl DatasetSource {
}
}
/// `dataset[w & mask]`.
/// `dataset[w & mask]` under the linear layout (the lottery hash).
pub fn word(&self, w: u32) -> u32 {
self.word_at(Layout::LINEAR, w)
}
/// `dataset[w & mask]` under a program's layout (era layout). The closed form has no items and ignores it.
pub fn word_at(&self, layout: Layout, w: u32) -> u32 {
let w = w & self.mask;
match &self.dataset {
Dataset::ClosedForm { d0, d1 } => dataset_elem(w, *d0, *d1),
Dataset::MemoryHard(m) => m.word(w),
Dataset::MemoryHard(m) => m.word_at(layout, w),
}
}
/// `out[k] = dataset[idx[k]]`; indices are already masked. Returns items derived (0 for the closed form).
#[inline]
fn fetch(&self, idx: &[u32; LANES], out: &mut [u32; LANES]) -> usize {
fn fetch(&self, idx: &[u32; LANES], out: &mut [u32; LANES], layout: Layout) -> usize {
match &self.dataset {
Dataset::ClosedForm { d0, d1 } => {
for k in 0..LANES {
@ -118,7 +281,24 @@ impl DatasetSource {
}
0
}
Dataset::MemoryHard(m) => m.fetch(idx, out),
Dataset::MemoryHard(m) => m.fetch(idx, out, layout),
}
}
/// `out[k][j] = dataset[base[k] + j]` for `j < width`; bases are masked and aligned to `width` words
/// (`width` 4 or 16, so a lane's words lie in one item). Returns items derived (0 for the closed form).
#[inline]
fn fetch_wide(&self, base: &[u32; LANES], width: usize, out: &mut [[u32; 16]; LANES], layout: Layout) -> usize {
match &self.dataset {
Dataset::ClosedForm { d0, d1 } => {
for k in 0..LANES {
for j in 0..width {
out[k][j] = dataset_elem(base[k] + j as u32, *d0, *d1);
}
}
0
}
Dataset::MemoryHard(m) => m.fetch_wide(base, width, out, layout),
}
}
}
@ -146,7 +326,23 @@ pub fn interpret_warp(program: &Program, base_nonce: u32, ds: &DatasetSource) ->
/// [`interpret_warp`] with explicit init words `I` (section 1.6 of the spec). The packs use `I = program.seed`;
/// a block uses `I = bind::block_init_words(H, nonce)`.
pub fn interpret_warp_init(program: &Program, seed: &[u32; 8], base_nonce: u32, ds: &DatasetSource) -> WarpResult {
interpret_warp_scratch(program, seed, base_nonce, ds, false).0
}
/// [`interpret_warp_init`] that also returns every scratch read-modify-write of the unit in execution order
/// (lane-minor within an instruction, as the interpreter runs them) when `trace` is set; empty otherwise and for
/// a class without a scratch. For the soundness tests of variant 5 only.
pub fn interpret_warp_scratch(
program: &Program,
seed: &[u32; 8],
base_nonce: u32,
ds: &DatasetSource,
trace: bool,
) -> (WarpResult, Vec<ScratchEvent>) {
let mask = ds.mask;
let log2 = ds.log2_words;
let era = program.class.era;
let layout = program.class.layout();
let mut r = [[0u32; LANES]; 8];
for lane in 0..LANES {
let nonce = base_nonce.wrapping_add(lane as u32);
@ -160,10 +356,29 @@ pub fn interpret_warp_init(program: &Program, seed: &[u32; 8], base_nonce: u32,
let mut items_derived = 0usize;
let mut idx = [0u32; LANES];
let mut val = [0u32; LANES];
let mut scratch = if program.has_scratch() { Some(ScratchModel::new(program.class.scratch_slots_per_lane())) } else { None };
if trace {
if let Some(m) = scratch.as_mut() {
m.trace = Some(Vec::new());
}
}
let slot_mask = program.class.scratch_slot_mask();
if program.has_hot() {
let h = ds.hot.as_ref().expect("a hot-table program needs the epoch's hot table on the dataset source");
assert_eq!(h.n_words(), program.hot_words(), "the hot table's size is the class's");
}
for _ in 0..ITERATIONS {
let sel = r[0];
for ins in &program.instrs {
step(ins, &mut r, &sel, mask, ds, &mut idx, &mut val, &mut items_derived);
step(ins, &mut r, &sel, mask, log2, era.as_ref(), layout, ds, &mut idx, &mut val, &mut items_derived);
if ins.op == Op::Scratch {
let m = scratch.as_mut().expect("a scratch op needs a scratch class");
let (d, a) = (ins.dst as usize, ins.src as usize);
for lane in 0..LANES {
let slot = r[a][lane] & slot_mask;
r[d][lane] = m.rmw(&program.seed, base_nonce, lane, slot, r[d][lane]);
}
}
}
}
let mut hashes = [0u64; LANES];
@ -172,7 +387,8 @@ pub fn interpret_warp_init(program: &Program, seed: &[u32; 8], base_nonce: u32,
let hi = r[4][lane] ^ r[5][lane].rotate_left(9) ^ r[6][lane].rotate_left(18) ^ r[7][lane].rotate_left(27);
hashes[lane] = ((hi as u64) << 32) | lo as u64;
}
WarpResult { hashes, items_derived }
let events = scratch.and_then(|m| m.trace).unwrap_or_default();
(WarpResult { hashes, items_derived }, events)
}
#[inline(always)]
@ -182,6 +398,9 @@ fn step(
r: &mut [[u32; LANES]; 8],
sel: &[u32; LANES],
mask: u32,
log2: u32,
era: Option<&EraParams>,
layout: Layout,
ds: &DatasetSource,
idx: &mut [u32; LANES],
val: &mut [u32; LANES],
@ -255,22 +474,46 @@ fn step(
r[d][lane] ^= src[lane ^ m];
}
}
Op::Load => {
Op::Load if ins.width == 1 => {
for lane in 0..LANES {
idx[lane] = r[a][lane] & mask;
idx[lane] = load_index(era, ins, r[a][lane], mask, log2);
}
*items_derived += ds.fetch(idx, val);
*items_derived += ds.fetch(idx, val, layout);
for lane in 0..LANES {
r[d][lane] ^= val[lane];
}
}
Op::Load => {
// Read-width experiment: `width` words from the aligned address, every word folded into dst.
let width = ins.width as usize;
let align = !(ins.width as u32 - 1);
for lane in 0..LANES {
idx[lane] = load_index(era, ins, r[a][lane], mask, log2) & align;
}
let mut vals = [[0u32; 16]; LANES];
*items_derived += ds.fetch_wide(idx, width, &mut vals, layout);
for lane in 0..LANES {
r[d][lane] = fold_words(r[d][lane], &vals[lane][..width]);
}
}
Op::Scratch => {
// handled by the caller (interpret_warp_init), which owns the unit's scratch model
}
Op::Hot => {
// Hot-table experiment: one word of the epoch table at the multiply-shift index, plain xor fold.
let h = ds.hot.as_ref().expect("a hot load needs the hot table");
let n = h.n_words();
for lane in 0..LANES {
r[d][lane] ^= h.at(hot_index(r[a][lane], n));
}
}
Op::WLoad => {
// Lane 0's register, masked, aligned down to 32 words; lane l reads word base + l.
let base = (r[a][0] & mask) & !31;
for lane in 0..LANES {
idx[lane] = base + lane as u32;
}
*items_derived += ds.fetch(idx, val);
*items_derived += ds.fetch(idx, val, Layout::LINEAR);
for lane in 0..LANES {
r[d][lane] ^= val[lane];
}
@ -294,11 +537,39 @@ pub struct Epoch {
/// Default dataset size: 2^28 words = 1 GiB.
pub const DEFAULT_DATASET_LOG2: u32 = 28;
/// Days a day index lies after the network's genesis day (0 for the genesis day and any day before it). The node's
/// entry; the same function as `memhard::days_since_genesis`.
pub fn days_since_genesis(day_index: u64, genesis_day_index: u64) -> u64 {
crate::memhard::days_since_genesis(day_index, genesis_day_index)
}
impl Epoch {
pub fn new(seed: &str, day: &str, mode: DatasetMode, dataset_log2: u32) -> Self {
Self { program: generate(seed), dataset: DatasetSource::new(day, mode, dataset_log2) }
}
/// [`Epoch::new`] with a load class (read-width experiment; Counter ASIC 2.0: the class's mixer multiplier
/// shapes the dataset, the cache is the genesis size since a string day has no day index).
pub fn new_class(seed: &str, day: &str, mode: DatasetMode, dataset_log2: u32, class: LoadClass) -> Self {
Self::new_class_day(seed, day, mode, dataset_log2, class, 0)
}
/// [`Epoch::new_class`] on day `days_since_genesis` of the growth schedule (the cache of
/// `memhard::cache_log2_words` for a class with the growth rule; the dataset size is the caller's).
pub fn new_class_day(seed: &str, day: &str, mode: DatasetMode, dataset_log2: u32, class: LoadClass, days_since_genesis: u64) -> Self {
let shape = Shape::for_class_day(&class, days_since_genesis);
let program = generate_class(seed, class);
let mut dataset = DatasetSource::new_shape(day, mode, dataset_log2, shape);
// hot-table experiment: a hot class fills its table from the seed bytes
dataset.attach_hot_for(&program);
Self { program, dataset }
}
/// `dataset[w]` as this epoch's program reads it: under the program's layout (era layout; linear for v2).
pub fn dataset_word(&self, w: u32) -> u32 {
self.dataset.word_at(self.program.class.layout(), w)
}
/// The production shape: memory-hard, 1 GiB dataset.
pub fn memory_hard(seed: &str, day: &str) -> Self {
Self::new(seed, day, DatasetMode::MemoryHard, DEFAULT_DATASET_LOG2)
@ -309,13 +580,71 @@ impl Epoch {
/// `seed_words_from_bytes(day_bytes)` (`bind::day_bytes`). Memory-hard, 1 GiB dataset. `label` is only
/// recorded in emitted packs.
pub fn from_seed_bytes(epoch_seed: &[u8], day_bytes: &[u8], label: &str) -> Self {
let program = crate::generator::generate_from_seed_bytes(label, epoch_seed);
Self::from_seed_bytes_class(epoch_seed, day_bytes, label, LoadClass::V2)
}
/// [`Epoch::from_seed_bytes`] with a load class (read-width experiment; Counter ASIC 2.0: the class's mixer
/// multiplier shapes the dataset). Day 0 of the growth schedule: the 2^26-word cache and the 2^28-word dataset,
/// which is every devnet pack and vector. A node past the first doubling calls [`Epoch::from_seed_bytes_day`].
pub fn from_seed_bytes_class(epoch_seed: &[u8], day_bytes: &[u8], label: &str, class: LoadClass) -> Self {
Self::from_seed_bytes_day(epoch_seed, day_bytes, label, class, 0, DEFAULT_DATASET_LOG2)
}
/// The chain's shape on day `days_since_genesis` (`memhard::days_since_genesis(day_index(header), day_index(genesis))`,
/// the node's two day indices): the program of the class, and under the class's growth rule the cache of
/// `memhard::cache_log2_words(d)` and the dataset of `memhard::dataset_log2_words(genesis_dataset_log2, d)`
/// (the genesis size is 28 for the 1 GiB devnet, 29 for the designed 2 GiB). Without the growth rule the cache
/// is 2^26 words and the dataset `2^genesis_dataset_log2` on every day.
pub fn from_seed_bytes_day(epoch_seed: &[u8], day_bytes: &[u8], label: &str, class: LoadClass, days_since_genesis: u64, genesis_dataset_log2: u32) -> Self {
let program = crate::generator::generate_from_seed_bytes_class(label, epoch_seed, class);
let key = crate::seed::seed_words_from_bytes(day_bytes);
let mut dataset = DatasetSource::from_key(key, DatasetMode::MemoryHard, DEFAULT_DATASET_LOG2);
let shape = Shape::for_class_day(&class, days_since_genesis);
let dataset_log2 = if class.growth { crate::memhard::dataset_log2_words(genesis_dataset_log2, days_since_genesis) } else { genesis_dataset_log2 };
let mut dataset = DatasetSource::from_key_shape(key, DatasetMode::MemoryHard, dataset_log2, shape);
dataset.key_bytes = day_bytes.to_vec();
dataset.attach_hot_for(&program);
Self { program, dataset }
}
/// The chain's shape with the program class (Counter ASIC 2.0, 5 October 2026): what the node's engine and the
/// miner's pack export build from the seeds a block template carries. Class v2 is [`Epoch::from_seed_bytes`]
/// exactly (the era bytes are ignored and not recorded); class v3 draws from [`crate::generator::V3_CLASS`]
/// with generator version 3 and records the era seed bytes (`E_n`) in the program for the pack.
pub fn from_chain_seeds(epoch_seed: &[u8], day_bytes: &[u8], era_bytes: Option<&[u8]>, class: ProgramClass, label: &str) -> Self {
Self { program: Self::chain_program(epoch_seed, era_bytes, class, label), dataset: Self::chain_dataset(day_bytes, class) }
}
/// The program alone of [`Epoch::from_chain_seeds`] (no cache fill): for an engine that shares the day's cache.
pub fn chain_program(epoch_seed: &[u8], era_bytes: Option<&[u8]>, class: ProgramClass, label: &str) -> Program {
crate::generator::generate_from_seed_bytes_program_class(label, epoch_seed, class, era_bytes)
}
/// The day's cache and dataset of [`Epoch::from_chain_seeds`], the one entry the node's engine builds a day
/// cache through. The class is an argument because the Counter ASIC 2.0 integration gives class v3 its own item
/// construction (the mixer multiplier) and cache size schedule (ca2-mixer); today both classes build the day of
/// [`Epoch::from_seed_bytes`], and the engine keys its day caches on `(day, class)` so the two never share one.
pub fn chain_dataset(day_bytes: &[u8], class: ProgramClass) -> DatasetSource {
Self::chain_dataset_day(day_bytes, class, 0, DEFAULT_DATASET_LOG2)
}
/// [`Epoch::chain_dataset`] with the day's position since genesis and the network's genesis dataset size: the
/// entry the node's engine and the miner's export build every day cache through, so the cache growth schedule
/// of spec 01 section 1.13.3 has one place to act (ca2-mixer, 5 October 2026, `docs/plans/mixer-x4.md`): the
/// class's load class gives the mixer multiplier and whether the growth rule applies (`Shape::for_class_day`);
/// under the rule the cache is `2^memhard::cache_log2_words(d)` words and the dataset
/// `2^memhard::dataset_log2_words(genesis_dataset_log2, d)`; without it (class v2) the cache is 2^26 words and
/// the dataset the genesis size on every day. `days_since_genesis` is [`days_since_genesis`] of the block's and
/// the genesis header's day indices.
pub fn chain_dataset_day(day_bytes: &[u8], class: ProgramClass, days_since_genesis: u64, genesis_dataset_log2: u32) -> DatasetSource {
let lc = class.load_class();
let shape = Shape::for_class_day(&lc, days_since_genesis);
let dataset_log2 = if lc.growth { crate::memhard::dataset_log2_words(genesis_dataset_log2, days_since_genesis) } else { genesis_dataset_log2 };
let key = crate::seed::seed_words_from_bytes(day_bytes);
let mut dataset = DatasetSource::from_key_shape(key, DatasetMode::MemoryHard, dataset_log2, shape);
dataset.key_bytes = day_bytes.to_vec();
dataset
}
/// The 32 hashes of the warp starting at `base_nonce`.
pub fn hash_warp(&self, base_nonce: u32) -> [u64; LANES] {
hash_warp(&self.program, base_nonce, &self.dataset)
@ -356,6 +685,190 @@ mod tests {
assert_eq!(ds.word(0x0fffffff), 0xf78c84a4);
}
/// Read-width experiment: the fold for one word is a plain xor; a wide fetch hands each lane the words the
/// scalar path would; two distinct lines give two distinct maps of dst (one point suffices as a smoke check).
#[test]
fn fold_and_wide_fetch() {
assert_eq!(fold_words(0x1234_5678, &[0xdead_beef]), 0x1234_5678 ^ 0xdead_beef);
let w = [1u32, 2, 3, 4];
let x = fold_words(7, &w);
let mut y: u32 = 7 ^ 1;
for &v in &w[1..] {
y = y.rotate_left(FOLD_ROT).wrapping_mul(FOLD_MUL) ^ v;
}
assert_eq!(x, y);
assert_ne!(fold_words(7, &[1, 2, 3, 4]), fold_words(7, &[1, 2, 3, 5]));
let ds = DatasetSource::new("2026-10-03", DatasetMode::MemoryHard, 20);
let mut base = [0u32; LANES];
for (k, b) in base.iter_mut().enumerate() {
*b = ((k as u32).wrapping_mul(0x9E37_79B1) & ds.mask) & !15;
}
let mut out = [[0u32; 16]; LANES];
let items = ds.fetch_wide(&base, 16, &mut out, Layout::LINEAR);
assert!(items >= 1 && items <= LANES);
for k in 0..LANES {
for j in 0..16 {
assert_eq!(out[k][j], ds.word(base[k] + j as u32), "lane {k} word {j}");
}
}
let mut base4 = base;
for b in base4.iter_mut() {
*b += 8;
}
let items4 = ds.fetch_wide(&base4, 4, &mut out, Layout::LINEAR);
assert_eq!(items4, items);
for k in 0..LANES {
for j in 0..4 {
assert_eq!(out[k][j], ds.word(base4[k] + j as u32));
}
}
}
/// A wide-load program interprets identically on the closed form and through the memory-hard path's fold
/// (the same fold code), and a mixed-class epoch builds and hashes.
/// Variant 5: a fill word is deterministic, a rewrite changes the slot, and a second read of a written slot
/// returns the rewrite, not the fill.
#[test]
fn scratch_model() {
let seed = [1u32, 2, 3, 4, 5, 6, 7, 8];
assert_eq!(scratch_fill(&seed, 32, 3, 100, 1), scratch_fill(&seed, 32, 3, 100, 1));
assert_ne!(scratch_fill(&seed, 32, 3, 100, 1), scratch_fill(&seed, 32, 3, 100, 2));
assert_ne!(scratch_fill(&seed, 32, 3, 100, 1), scratch_fill(&seed, 64, 3, 100, 1));
let mut m = ScratchModel::new(256);
let w = [scratch_fill(&seed, 32, 3, 100, 0), scratch_fill(&seed, 32, 3, 100, 1), scratch_fill(&seed, 32, 3, 100, 2)];
let x = m.rmw(&seed, 32, 3, 100, 0xabcd);
assert_eq!(x, fold_words(0xabcd, &w));
let x2 = m.rmw(&seed, 32, 3, 100, 0xabcd);
assert_eq!(x2, fold_words(0xabcd, &scratch_rewrite(x, &w)));
assert_eq!(m.reads, 2);
let e = Epoch::new_class("igneum-genesis", "2026-10-03", DatasetMode::ClosedForm, 20, LoadClass::scratch(4, 128));
assert_eq!(e.program.scratch_ops_per_hash(), 32);
assert_eq!(e.hash_warp(0), e.hash_warp(0));
}
/// Era layout: the load address stays inside the site's window and below the mask at every dataset size (the
/// window floor of 2^26 words clamps the shrink), the interleaved memory-hard dataset reads item(t(w))[j(w)] and
/// is the same prefix at 2^20 and 2^22 words, the wide fetch agrees word for word, and an era epoch hashes
/// deterministically through the interpreter and the single-nonce API.
#[test]
fn era_windows_layout_and_epochs() {
let eb = EraParams::test_era_bytes("igneum-era-test/1");
let c = LoadClass::era(LoadClass::V2, &eb, &[1]);
let e = c.era.unwrap();
let mut ins = Instr { op: Op::Load, dst: 0, src: 1, src2: 0, imm: 0, imm2: 0, rot: 1, bit: 0, mask: 1, width: 1, win: 2, off: 3 };
let mut s = crate::seed::SplitMix64::new(7);
for log2 in [20u32, 26, 27, 28, 29] {
let mask = (1u64 << log2) as u32 - 1;
let k = ins.win.min(log2.saturating_sub(26) as u8) as u32;
for _ in 0..1000 {
let x = s.next() as u32;
let idx = load_index(Some(&e), &ins, x, mask, log2);
assert!(idx <= mask);
let (wm, off) = window(&ins, mask, log2);
assert_eq!(idx & !wm, off, "log2 {log2}");
assert_eq!(wm, mask >> k);
assert_eq!(idx, ((x.wrapping_mul(e.stride_mul).rotate_left(e.stride_rot) & wm) | off) & mask);
}
}
ins.win = 0;
assert_eq!(load_index(None, &ins, 0xdead_beef, 0x0fff_ffff, 28), 0xdead_beef & 0x0fff_ffff);
// the interleaved dataset: one day cache, the layout per program
let l = e.layout();
assert_eq!(l.pos, [1, 3, 8, 13]);
let small = DatasetSource::new("2026-10-03", DatasetMode::MemoryHard, 20);
let big = DatasetSource::new("2026-10-03", DatasetMode::MemoryHard, 22);
let m = small.memhard().unwrap();
for w in [0u32, 1, 4, 5, 255, 256, 4095, 8192, 0x0f_ffff] {
let (t, j) = l.split(w);
assert_eq!(small.word_at(l, w), crate::memhard::derive_item(t, &m.params, &m.cache)[j as usize], "w {w}");
assert_eq!(small.word_at(l, w), big.word_at(l, w), "prefix at w {w}");
assert_eq!(small.word(w), small.word_at(Layout::LINEAR, w));
}
let mut idx = [0u32; LANES];
for (k, i) in idx.iter_mut().enumerate() {
*i = (k as u32).wrapping_mul(0x9E37_79B1) & small.mask;
}
let mut out = [0u32; LANES];
small.fetch(&idx, &mut out, l);
for k in 0..LANES {
assert_eq!(out[k], small.word_at(l, idx[k]));
}
// a 16-byte era: the wide fetch keeps a lane's four words in one item
let eb3 = EraParams::test_era_bytes("igneum-era-test/3");
let c3 = LoadClass::era(LoadClass::fixed(4, 16), &eb3, &[4]);
let l3 = c3.layout();
assert_eq!(l3.pos[..2], [0, 1]);
let mut base = [0u32; LANES];
for (k, b) in base.iter_mut().enumerate() {
*b = ((k as u32).wrapping_mul(0x9E37_79B1) & small.mask) & !3;
}
let mut wide = [[0u32; 16]; LANES];
small.fetch_wide(&base, 4, &mut wide, l3);
for k in 0..LANES {
for j in 0..4 {
assert_eq!(wide[k][j], small.word_at(l3, base[k] + j as u32), "lane {k} word {j}");
}
}
// era epochs hash deterministically, differ per era, and the single-nonce API agrees with the warp
let mut seen = std::collections::HashSet::new();
for (n, c) in [(1u64, c), (3, c3)] {
let ep = Epoch::new_class("igneum-genesis", "2026-10-03", DatasetMode::MemoryHard, 20, c);
assert_eq!(ep.dataset_word(5), ep.dataset.word_at(c.layout(), 5));
let a = ep.hash_warp(64);
assert_eq!(a, ep.hash_warp(64));
assert_eq!(ep.hash(64 + 5), a[5]);
assert!(seen.insert(a[0]), "era {n}");
let closed = Epoch::new_class("igneum-genesis", "2026-10-03", DatasetMode::ClosedForm, 20, c);
assert_ne!(closed.hash_warp(64), a);
}
}
/// Hot-table experiment: an epoch of a hot class carries the table, hashes deterministically and differs from
/// version 2; the reference interpreter agrees with a hand-stepped hot load; a hot program without its table is
/// refused.
#[test]
fn hot_epochs_hash() {
let e = Epoch::new_class("igneum-genesis", "2026-10-03", DatasetMode::ClosedForm, 20, LoadClass::hot(32, 4));
let h = e.dataset.hot.as_ref().expect("the epoch fills the hot table");
assert_eq!(h.n_words(), 1 << 23);
assert_eq!(h.key, crate::memhard::hot_key(b"igneum-genesis"));
let v2 = Epoch::new_class("igneum-genesis", "2026-10-03", DatasetMode::ClosedForm, 20, LoadClass::V2);
let a = e.hash_warp(0);
assert_eq!(a, e.hash_warp(0));
assert_ne!(a, v2.hash_warp(0));
assert_ne!(a[0], a[1]);
// the same program under a 64 MiB table reads other words
let e64 = Epoch::new_class("igneum-genesis", "2026-10-03", DatasetMode::ClosedForm, 20, LoadClass::hot(64, 4));
assert_eq!(e64.program.instrs, e.program.instrs);
assert_ne!(e64.hash_warp(0), a);
// from seed bytes, the chain's shape, with a hot class
let genesis = crate::bind::unhex("edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07").unwrap();
let ec = Epoch::from_seed_bytes_class(&genesis, &crate::bind::day_bytes(20_730), "devnet", LoadClass::hot(32, 2));
assert_eq!(ec.dataset.hot.as_ref().unwrap().key, crate::memhard::hot_key(&genesis));
assert_eq!(ec.hash_warp(0), ec.hash_warp(0));
}
#[test]
#[should_panic(expected = "needs the epoch's hot table")]
fn hot_program_without_a_table_is_refused() {
let p = generate_class("igneum-genesis", LoadClass::hot(32, 4));
let ds = DatasetSource::new("2026-10-03", DatasetMode::ClosedForm, 20);
let _ = hash_warp(&p, 0, &ds);
}
#[test]
fn wide_class_epochs_hash() {
for name in ["w16", "w64x4", "50,35,15"] {
let c = LoadClass::parse(name).unwrap();
let e = Epoch::new_class("igneum-genesis", "2026-10-03", DatasetMode::ClosedForm, 20, c);
assert_eq!(e.program.class, c);
let a = e.hash_warp(0);
let b = e.hash_warp(0);
assert_eq!(a, b);
assert_ne!(a[0], a[1]);
}
}
#[test]
fn closed_form_genesis_vector_lane0() {
// Generator v2 vectors (4 October 2026), proto-cuda/packs/igneum-genesis/vectors.json.

247
igneum-pow/tests/mixer.rs Normal file
View file

@ -0,0 +1,247 @@
//! The class v3 dataset construction (mixer x4, cache growth option C; `docs/plans/mixer-x4.md`): the soundness
//! runs the brief asks for, on the CPU, with the packs for the GPU runs written on request.
//!
//! 1. Fuzz: `IGNEUM_MIXER_FUZZ` (default 200) programs through the seam (`ProgramClass::V3`), the contract on every
//! instruction (the v2 program of the seed, instruction for instruction), 4 units each across the 32-bit range
//! including the wrap, interpreted twice on the CPU; with `IGNEUM_MIXER_PACKS_OUT=<dir>` every program is written
//! as a pack with its 4 bases in vectors.json for `packbench` and the OpenCL host (the Metal fuzz).
//! 2. Stats: bit balance and single-bit-flip avalanche of the v3 hash against v2 on the same programs and nonces.
//! 3. Edge: the dataset at word 0, word MASK and the item boundary, derived through the interpreter's fetch path and
//! by hand at every multiplier 1, 2, 4, 8, on a small cache.
//! 4. Determinism: two independent epochs of the same seed and day agree on every vector and every emitted file.
use igneum_pow::emit::{export_pack, vectors_json};
use igneum_pow::generator::{generate_from_seed_bytes, generate_from_seed_bytes_class, generate_from_seed_bytes_program_class, LoadClass, Op, Program, ProgramClass, GENERATOR_VERSION_V3, INSTR_COUNT, V3_CLASS};
use igneum_pow::memhard::{derive_item, mixer, round_key, Cache, MixParams, Shape};
use igneum_pow::seed::{day_key, SplitMix64};
use igneum_pow::verify::{DatasetMode, DatasetSource, Epoch};
use std::collections::HashMap;
use std::path::PathBuf;
const DAY: &str = "2026-10-03";
fn contract(p: &Program, seed: &str, class: LoadClass) {
if class == V3_CLASS {
assert_eq!(p.generator, GENERATOR_VERSION_V3);
}
assert_eq!(p.class, class);
assert_eq!(p.instrs.len(), INSTR_COUNT);
assert_eq!(p.instrs.iter().filter(|i| i.op == Op::Load).count(), 16);
for (k, i) in p.instrs.iter().enumerate() {
assert!(i.src != i.dst, "#{k}: src == dst");
assert!((1..=31).contains(&i.rot), "#{k}: rot {}", i.rot);
assert!([1u8, 2, 4, 8, 16].contains(&i.mask), "#{k}: mask {}", i.mask);
assert!(i.dst < 8 && i.src < 8 && i.src2 < 8);
assert_eq!(i.width, 1, "#{k}: a v3 load reads one word");
}
assert!(igneum_pow::accept::check(p).is_ok(), "an accepted program");
let v2 = generate_from_seed_bytes(seed, seed.as_bytes());
assert_eq!(p.instrs, v2.instrs, "the v2 program of the seed under the v3 construction");
assert_eq!(p.attempt, v2.attempt);
}
/// Write a pack whose vectors.json carries `bases` instead of the three standard bases (packbench and the OpenCL
/// host check every unit standalone and the ones inside the batch window).
fn write_pack_with_bases(dir: &PathBuf, e: &Epoch, day: &str, bases: &[u32], source: &str) {
let mut pack = export_pack(e, day, source);
let outs: Vec<[u64; 32]> = bases.iter().map(|&b| e.hash_warp(b)).collect();
let vj = vectors_json(&e.program, day, e.dataset.log2_words, bases, &outs, &pack.vectors, e.dataset.mask, source, true);
for f in pack.files.iter_mut() {
if f.0 == "vectors.json" {
f.1 = vj.clone();
}
}
pack.write_to(dir).unwrap();
}
#[test]
fn fuzz_v3_programs_cpu() {
let n: usize = std::env::var("IGNEUM_MIXER_FUZZ").ok().and_then(|s| s.parse().ok()).unwrap_or(200);
let out = std::env::var("IGNEUM_MIXER_PACKS_OUT").ok().map(PathBuf::from);
// IGNEUM_MIXER_CLASS=mx8 fuzzes the x8 candidate as a load class (generator 2 with the class in the id); the
// default is V3_CLASS through the seam
let class = std::env::var("IGNEUM_MIXER_CLASS").ok().map(|s| LoadClass::parse(&s).expect("a load class")).unwrap_or(V3_CLASS);
let mut rng = SplitMix64::new(0x6967_6e65_756d_2d6d); // "igneum-m"
let shape = Shape::for_class(&class);
assert_eq!(shape.cache_log2_words, 26);
assert!(shape.mixer_mult > 1);
// one memory-hard source per dataset size (the 256 MiB cache fill is 0.2 s each)
let mut mh: HashMap<u32, DatasetSource> = HashMap::new();
let mut manifest = String::from("pack\tlog2\tprogram_id\tbases\n");
let mut units = 0usize;
let mut wraps = 0usize;
for i in 0..n {
let seed = format!("igneum-mixer-fuzz/{i}");
let p = if class == V3_CLASS {
generate_from_seed_bytes_program_class(&seed, seed.as_bytes(), ProgramClass::V3, None)
} else {
generate_from_seed_bytes_class(&seed, seed.as_bytes(), class)
};
contract(&p, &seed, class);
let b0 = (rng.below(8) as u32) * 32;
let b1 = 0x8000_0000u32.wrapping_sub(256).wrapping_add((rng.below(16) as u32) * 32);
let b2 = 0xffff_ff00u32.wrapping_add((rng.below(8) as u32) * 32);
let b3 = (rng.next() as u32) & !31;
let bases = [b0, b1, b2, b3];
wraps += bases.iter().filter(|&&b| b >= 0xffff_ff00).count();
let log2 = [24u32, 26, 28][rng.below(3) as usize];
let ds = mh.remove(&log2).unwrap_or_else(|| DatasetSource::new_shape(DAY, DatasetMode::MemoryHard, log2, shape));
let e = Epoch { program: p, dataset: ds };
for &b in &bases {
let r1 = e.interpret_warp(b);
let r2 = e.interpret_warp(b);
assert_eq!(r1.hashes, r2.hashes);
assert!(r1.items_derived >= 120 * 32 / 32 && r1.items_derived <= 4_096, "{seed}: {} items", r1.items_derived);
units += 1;
}
if let Some(dir) = &out {
let pack_name = format!("fuzz-{i:03}-{}-l{log2}", class.name());
write_pack_with_bases(&dir.join(&pack_name), &e, DAY, &bases, "igneum-pow tests/mixer.rs fuzz");
manifest.push_str(&format!(
"{pack_name}\t{log2}\t{:016x}\t{}\n",
e.program.program_id(),
bases.iter().map(|b| format!("{b}")).collect::<Vec<_>>().join(",")
));
}
mh.insert(log2, e.dataset);
}
println!("fuzz: {n} {} programs, {units} units on the CPU, {wraps} units in the top 256 nonces", class.name());
assert_eq!(units, 4 * n);
assert_eq!(wraps, n);
if let Some(dir) = &out {
std::fs::create_dir_all(dir).unwrap();
std::fs::write(dir.join("manifest.tsv"), manifest).unwrap();
println!("packs written to {}", dir.display());
}
}
/// Bit balance and avalanche of the v3 hash beside v2 on the same program (the TESTS.md section 3 shape, on the
/// CPU, 2^13 nonces per seed): every output bit within 5 sigma of half ones; a single nonce-bit flip moves 50 percent
/// of the output bits within 2 points; no duplicate among the outputs.
#[test]
fn stats_v3_against_v2() {
let n_warps = 256usize; // 8,192 nonces
for seed in ["igneum-genesis", "igneum-genesis/stats1"] {
let v3 = Epoch {
program: generate_from_seed_bytes_program_class(seed, seed.as_bytes(), ProgramClass::V3, None),
dataset: DatasetSource::new_shape(DAY, DatasetMode::MemoryHard, 24, Shape::for_class(&V3_CLASS)),
};
let v2 = Epoch::new(seed, DAY, DatasetMode::MemoryHard, 24);
for (name, e) in [("v3", &v3), ("v2", &v2)] {
let mut ones = [0u64; 64];
let mut outs = Vec::with_capacity(n_warps * 32);
for w in 0..n_warps {
let h = e.hash_warp(w as u32 * 32);
for &x in &h {
outs.push(x);
for b in 0..64 {
ones[b] += (x >> b) & 1;
}
}
}
let total = (n_warps * 32) as f64;
let sigma = (total / 4.0).sqrt();
for (b, &c) in ones.iter().enumerate() {
let z = (c as f64 - total / 2.0).abs() / sigma;
assert!(z < 5.0, "{seed} {name}: bit {b} ones {c} of {total}, z {z:.2}");
}
// avalanche: flip one bit of the nonce within the unit (lanes 0..31 differ in the low 5 bits) and across
// units (bit 5 and up): compare lane l of unit u with lane l ^ (1 << k) and with unit u ^ (1 << k)
let mut flips = 0u64;
let mut moved = 0u64;
for w in 0..64usize {
let h = e.hash_warp(w as u32 * 32);
for k in 0..5 {
for l in 0..32usize {
moved += (h[l] ^ h[l ^ (1 << k)]).count_ones() as u64;
flips += 1;
}
}
let h2 = e.hash_warp((w ^ 1) as u32 * 32);
for l in 0..32usize {
moved += (h[l] ^ h2[l]).count_ones() as u64;
flips += 1;
}
}
let avg = moved as f64 / flips as f64 / 64.0 * 100.0;
assert!((avg - 50.0).abs() < 2.0, "{seed} {name}: avalanche {avg:.2} percent");
outs.sort_unstable();
let dups = outs.windows(2).filter(|p| p[0] == p[1]).count();
assert_eq!(dups, 0, "{seed} {name}: duplicate outputs");
println!("{seed} {name}: {} outputs, avalanche {avg:.2} percent, worst bit z {:.2}", outs.len(), ones.iter().map(|&c| (c as f64 - total / 2.0).abs() / sigma).fold(0.0, f64::max));
}
assert_ne!(v3.hash_warp(0), v2.hash_warp(0));
}
}
/// The dataset edges under every multiplier on a small cache: word 0, word MASK, the last word of item 0 and the
/// first of item 1, through `DatasetSource::word` and by hand.
#[test]
fn edge_items_every_multiplier() {
let key = day_key(DAY);
let cache = Cache::fill_log2(key, 14);
for m in [1u32, 2, 4, 8] {
let mp = MixParams::with_shape(key, Shape { mixer_mult: m, cache_log2_words: 14 });
let by_hand = |t: u32| -> [u32; 16] {
let mut s = [0u32; 16];
s[..8].copy_from_slice(&key);
for i in 0..8 {
s[8 + i] = t.wrapping_mul(mp.mul[i]).wrapping_add(mp.rc[i]);
}
for r in 0..8usize {
for j in 0..m as usize {
mixer(&mut s, round_key(r * m as usize + j), &mp);
}
let line = cache.line(s[0]);
for i in 0..16 {
s[i] ^= line[i];
}
}
for j in 0..m as usize {
mixer(&mut s, round_key(8 * m as usize + j), &mp);
}
s
};
for t in [0u32, 1, 0x0fff_ffff, 0xffff_ffff] {
assert_eq!(derive_item(t, &mp, &cache), by_hand(t), "m {m} item {t}");
}
}
// the interpreter's fetch path at the genesis cache: words 0, 15, 16 and MASK of a 2^20-word dataset agree with
// the item derivation, under v3
let ds = DatasetSource::new_shape(DAY, DatasetMode::MemoryHard, 20, Shape::for_class(&V3_CLASS));
let m = ds.memhard().unwrap();
for w in [0u32, 15, 16, 17, ds.mask - 1, ds.mask] {
assert_eq!(ds.word(w), derive_item(w >> 4, &m.params, &m.cache)[(w & 15) as usize]);
assert_eq!(ds.word(w), m.word(w));
}
// a load at an out-of-range register masks to the dataset: the word at mask + 1 is the word at 0
assert_eq!(ds.word(ds.mask.wrapping_add(1)), ds.word(0));
}
/// Two independent epochs of the same seed and day: every vector and every emitted file identical; the pinned v3
/// pack is what a third export writes.
#[test]
fn determinism_v3() {
let build = || Epoch {
program: generate_from_seed_bytes_program_class("igneum-genesis", b"igneum-genesis", ProgramClass::V3, None),
dataset: DatasetSource::new_shape(DAY, DatasetMode::MemoryHard, 28, Shape::for_class(&V3_CLASS)),
};
let a = build();
let b = build();
let pa = export_pack(&a, DAY, "a");
let pb = export_pack(&b, DAY, "a");
assert_eq!(pa.outs, pb.outs);
assert_eq!(pa.vectors, pb.vectors);
assert_eq!(pa.files, pb.files);
for (w, warp) in [(0u32, 0usize), (4096, 1), (1_000_000, 2)] {
assert_eq!(a.hash_warp(w), pa.outs[warp]);
}
let dir = PathBuf::from(env!("CARGO_MANIFEST_DIR")).join("../proto-cuda/packs-ca2-mixer/mx8-genesis");
for (name, text) in &pa.files {
if name == "vectors.json" || name == "vectors.h" {
continue; // the source string differs ("a" here)
}
let on_disk = std::fs::read_to_string(dir.join(name)).unwrap();
assert_eq!(&on_disk, text, "{name}");
}
}

View file

@ -5,28 +5,52 @@
//!
//! Packs: igneum-genesis-mh and igneum-devnet-v4-epoch0 (memory-hard; the latter from the devnet genesis hash as
//! the epoch seed and the day bytes of 2026-10-04), igneum-genesis and igneum-hourly (closed-form dataset,
//! interpreter regression only).
//! interpreter regression only); and, under `proto-cuda/packs-ca2-mixer/`, the class v3 packs mx8-genesis and
//! mx8-devnet-epoch0 (Counter ASIC 2.0, 5 October 2026: generator 3 on `V3_CLASS` = mixer x8 with the cache growth
//! rule, decided 22:05 UTC under the delegated rule; the same seeds and days as the two memory-hard v2 packs, so the
//! v2 program and cache carry over and only the dataset words and the hashes change) and the x4 candidate's packs
//! mx4-genesis and mx4-devnet-epoch0 (generator 2 with the load class in the id, the record of the x4 rows).
use igneum_pow::accept;
use igneum_pow::emit::{
cuda_kernel, cuda_kernel_bound, cuda_memhard_header, export_pack, metal_memhard, metal_program,
cuda_kernel, cuda_kernel_bound, cuda_memhard_header, export_pack, metal_memhard, metal_memhard_for, metal_program,
metal_program_bound, opencl_kernel, opencl_kernel_bound, program_header, program_json, LoadSource,
};
use igneum_pow::generator::{generate_from_seed_bytes, Op, GENERATOR_VERSION, LOAD_SLOTS};
use igneum_pow::memhard::CACHE_WORDS;
use igneum_pow::generator::{generate_from_seed_bytes, generate_from_seed_bytes_class, generate_from_seed_bytes_program_class, LoadClass, Op, ProgramClass, GENERATOR_VERSION, GENERATOR_VERSION_V3, LOAD_SLOTS, V3_CLASS};
use igneum_pow::memhard::{Shape, CACHE_WORDS};
use igneum_pow::verify::{DatasetMode, DatasetSource, Epoch};
use serde_json::Value;
use std::path::PathBuf;
use std::sync::OnceLock;
const PACKS: [&str; 4] = ["igneum-genesis-mh", "igneum-devnet-v4-epoch0", "igneum-genesis", "igneum-hourly"];
const PACKS: [&str; 8] = [
"igneum-genesis-mh",
"igneum-devnet-v4-epoch0",
"igneum-genesis",
"igneum-hourly",
"mx8-genesis",
"mx8-devnet-epoch0",
"mx4-genesis",
"mx4-devnet-epoch0",
];
/// The class v3 packs (generator 3 through the seam).
const PACKS_V3: [&str; 2] = ["mx8-genesis", "mx8-devnet-epoch0"];
fn packs_dir() -> PathBuf {
PathBuf::from(env!("CARGO_MANIFEST_DIR")).join("../proto-cuda/packs")
}
/// The directory a pack lives in: the class v3 packs under packs-ca2-mixer, the rest under packs.
fn pack_dir(pack: &str) -> PathBuf {
if pack.starts_with("mx4-") || pack.starts_with("mx8-") {
PathBuf::from(env!("CARGO_MANIFEST_DIR")).join("../proto-cuda/packs-ca2-mixer").join(pack)
} else {
packs_dir().join(pack)
}
}
fn read(pack: &str, file: &str) -> String {
let p = packs_dir().join(pack).join(file);
let p = pack_dir(pack).join(file);
std::fs::read_to_string(&p).unwrap_or_else(|e| panic!("read {}: {e}", p.display()))
}
@ -62,9 +86,22 @@ fn epoch(pack: &str) -> &'static Epoch {
_ => DatasetMode::ClosedForm,
};
let log2 = j["dataset"]["log2_words"].as_u64().unwrap() as u32;
let program = generate_from_seed_bytes(seed, &seed_bytes);
// a class v3 pack: generator 3 on V3_CLASS through the seam, the era bytes it records, the dataset
// in the class's shape on day 0 (the growth rule's genesis cache: every pinned pack is a day-0 size)
let program = match j["generator"].as_u64().unwrap() as u32 {
GENERATOR_VERSION_V3 => {
let era = j.get("era_seed_bytes").map(unhex);
generate_from_seed_bytes_program_class(seed, &seed_bytes, ProgramClass::V3, era.as_deref())
}
// a generator 2 pack with a load class (the x4 candidate's packs): the class from program.json
_ => match j.get("load_class").and_then(|c| c.as_str()) {
Some(c) => generate_from_seed_bytes_class(seed, &seed_bytes, LoadClass::parse(c).expect("a load class name")),
None => generate_from_seed_bytes(seed, &seed_bytes),
},
};
let shape = Shape::for_class(&program.class);
let mut dataset =
DatasetSource::from_key(igneum_pow::seed::seed_words_from_bytes(&day_bytes), mode, log2);
DatasetSource::from_key_shape(igneum_pow::seed::seed_words_from_bytes(&day_bytes), mode, log2, shape);
dataset.key_bytes = day_bytes;
(p.to_string(), Epoch { program, dataset })
})
@ -81,7 +118,8 @@ fn check_program_json(pack: &str) {
let j = json(pack, "program.json");
let p = &epoch(pack).program;
assert_eq!(j["format"].as_str().unwrap(), "igneum-program-pack-3");
assert_eq!(j["generator"].as_u64().unwrap() as u32, GENERATOR_VERSION, "{pack}: generator version");
assert_eq!(j["generator"].as_u64().unwrap() as u32, p.generator, "{pack}: generator version");
assert_eq!(p.generator, if PACKS_V3.contains(&pack) { GENERATOR_VERSION_V3 } else { GENERATOR_VERSION });
assert_eq!(j["attempt"].as_u64().unwrap() as u32, p.attempt, "{pack}: attempt");
assert_eq!(hex64(&j["program_id"]), p.program_id(), "{pack}: program id");
let sw: Vec<u32> = j["seed_words"].as_array().unwrap().iter().map(hex32).collect();
@ -129,7 +167,7 @@ fn genesis_program_shape() {
#[test]
fn mixer_params_match_pack() {
for pack in ["igneum-genesis-mh", "igneum-devnet-v4-epoch0"] {
for pack in ["igneum-genesis-mh", "igneum-devnet-v4-epoch0", "mx8-genesis", "mx8-devnet-epoch0", "mx4-genesis", "mx4-devnet-epoch0"] {
let j = json(pack, "program.json");
let mp = &epoch(pack).dataset.memhard().unwrap().params;
let key: Vec<u32> = j["dataset"]["key"].as_array().unwrap().iter().map(hex32).collect();
@ -166,19 +204,21 @@ fn cache_matches_vectors() {
fn check_dataset_words(pack: &str) {
let v = json(pack, "vectors.json");
let ds = &epoch(pack).dataset;
let e = epoch(pack);
let ds = &e.dataset;
// a pack's self-test words are read under its program's layout (the era layout; linear for every v2 pack)
let head: Vec<u32> = v["dataset_head"].as_array().unwrap().iter().map(hex32).collect();
for (i, h) in head.iter().enumerate() {
assert_eq!(ds.word(i as u32), *h, "{pack}: dataset[{i}]");
assert_eq!(e.dataset_word(i as u32), *h, "{pack}: dataset[{i}]");
}
let last_index = v["dataset_last_index"].as_u64().unwrap() as u32;
assert_eq!(last_index, ds.mask);
assert_eq!(ds.word(last_index), hex32(&v["dataset_last"]), "{pack}: dataset[MASK]");
assert_eq!(e.dataset_word(last_index), hex32(&v["dataset_last"]), "{pack}: dataset[MASK]");
let samples = v["dataset_samples"].as_array().unwrap();
assert_eq!(samples.len(), 64);
for s in samples {
let idx = s["index"].as_u64().unwrap() as u32;
assert_eq!(ds.word(idx), hex32(&s["value"]), "{pack}: dataset[{idx}]");
assert_eq!(e.dataset_word(idx), hex32(&s["value"]), "{pack}: dataset[{idx}]");
}
}
@ -261,12 +301,15 @@ fn check_sources(pack: &str) {
assert_same_text(pack, "program.h", &program_header(p, &day, &e.dataset));
if let Some(mp) = mp {
assert_same_text(pack, "memhard.h", &cuda_memhard_header(p, mp));
assert_same_text(pack, "memhard.metal", &metal_memhard(mp));
assert_same_text(pack, "memhard.metal", &igneum_pow::emit::metal_memhard_layout(mp, p.class.layout()));
}
let got = program_json(p, &day, &e.dataset);
assert_same_text(pack, "program.json", &got);
let _: Value = serde_json::from_str(&got).expect("program.json is valid JSON");
// Every load in every emitted hash kernel has the masked form, and there are exactly 16 of them.
// Every load in every emitted hash kernel has the masked form, and there are exactly 16 of them (a class with
// the era layout inside has the era form instead: `((rotl_imm(rN * M, R) & WM) | OFF) & mask`, checked by
// era_emitted_sources_match_and_loads_have_the_era_form over the era packs, and here by the same count).
let era_load = if p.class.era.is_some() { "((rotl_imm(r" } else { "" };
for (file, load, masked) in [
("kernel.cu", "ds[r", " & mask]"),
("kernel_bound.cu", "ds[r", " & mask]"),
@ -274,6 +317,7 @@ fn check_sources(pack: &str) {
("program_bound.metal", "dataset[r", " & MASK]"),
] {
let text = read(pack, file);
let load = if era_load.is_empty() { load } else { era_load };
assert_eq!(text.matches(load).count(), LOAD_SLOTS, "{pack}/{file}: 16 loads");
assert_eq!(text.matches(masked).count(), LOAD_SLOTS, "{pack}/{file}: 16 masked loads");
}
@ -287,6 +331,85 @@ fn emitted_sources_match_all_packs() {
}
/// The whole pack as `export` writes it: vectors.json and vectors.h match, and the file list is the full set.
/// The class v3 packs (Counter ASIC 2.0, `docs/plans/mixer-x4.md`): generator 3 on V3_CLASS = mx8; the program of
/// each is the v2 program of the same seed instruction for instruction (v2 loads take no width roll); the cache is
/// the v2 cache (day 0 of the growth rule: 2^26 words, the same FNV-1a 64); the dataset words differ from v2's;
/// program.json, program.h and the emitted memhard core say so; the id carries generator 3.
#[test]
fn v3_packs_are_the_v2_seeds_under_mixer_x8() {
assert_eq!(V3_CLASS.name(), "mx8");
assert_eq!(V3_CLASS.mixer_mult, 8);
assert!(V3_CLASS.growth);
assert_eq!(LoadClass { mixer_mult: 4, ..V3_CLASS }, LoadClass::MX4, "the x4 candidate differs from v3 in the multiplier alone");
for (v3, v2) in [("mx8-genesis", "igneum-genesis-mh"), ("mx8-devnet-epoch0", "igneum-devnet-v4-epoch0")] {
let e3 = epoch(v3);
let e2 = epoch(v2);
let j = json(v3, "program.json");
assert_eq!(j["program_class"].as_str().unwrap(), "v3");
assert!(j["load_class"].as_str().unwrap().starts_with("mx8"), "{v3}: mx8, or mx8 with the era inside");
assert_eq!(j["mixer_mult"].as_u64().unwrap(), 8);
assert_eq!(j["cache_growth"].as_bool().unwrap(), true);
assert_eq!(j["dataset"]["mixer_mult"].as_u64().unwrap(), 8);
assert_eq!(j["dataset"]["cache"]["log2_words"].as_u64().unwrap(), 26);
assert_eq!(e3.program.generator, GENERATOR_VERSION_V3);
if e3.program.era_bytes.is_some() {
// the era layout composed into class v3 (docs/plans/era-layout.md, 5 October 2026): a chain pack carries an
// era, so its class is V3_CLASS with the era drawn inside and its stream takes two window draws per
// instruction; the v2 program carries over only in the seed, the attempt and the day
assert_eq!(LoadClass { era: None, ..e3.program.class }, V3_CLASS, "{v3}: the composed class");
assert!(e3.program.class.era.is_some());
assert_ne!(e3.program.instrs, e2.program.instrs, "{v3}: the era windows change the stream");
} else {
assert_eq!(e3.program.class, V3_CLASS);
assert_eq!(e3.program.instrs, e2.program.instrs, "{v3}: the v2 program under the v3 construction");
}
assert_eq!(e3.program.seed, e2.program.seed);
assert_eq!(e3.program.attempt, e2.program.attempt);
assert_ne!(e3.program.program_id(), e2.program.program_id());
assert_eq!(e3.program.program_id(), igneum_pow::generator::program_id(GENERATOR_VERSION_V3, &e3.program.seed, e3.program.attempt));
let m3 = e3.dataset.memhard().unwrap();
let m2 = e2.dataset.memhard().unwrap();
assert_eq!(m3.shape(), Shape { mixer_mult: 8, cache_log2_words: 26 });
assert_eq!(m3.cache.fnv1a64(), m2.cache.fnv1a64(), "{v3}: the same cache as v2 on day 0");
assert_eq!(m3.params.rot, m2.params.rot);
assert_eq!(e3.dataset.log2_words, 28);
assert_ne!(e3.dataset.word(0), e2.dataset.word(0), "{v3}: the dataset words differ");
assert_ne!(e3.hash_warp(0), e2.hash_warp(0));
let h = read(v3, "program.h");
assert!(h.contains("#define IGNEUM_GENERATOR 3\n"));
assert!(h.contains("#define IGNEUM_PROGRAM_CLASS \"v3\"\n"));
assert!(h.contains("#define IGNEUM_MIXER_MULT 8"));
assert!(h.contains("#define IGNEUM_CACHE_GROWTH 1"));
assert!(h.contains("#define IGNEUM_CACHE_LOG2_WORDS 26\n"));
assert!(h.contains("#define IGNEUM_LOAD_CLASS \"mx8"), "mx8, or mx8 with the era inside");
for file in ["memhard.h", "memhard.metal", "kernel.cl"] {
let text = read(v3, file);
assert_eq!(text.matches("j < 8u; ++j) mh_mixer(s, 0x9E3779B9u * (r * 8u + j + 1u))").count(), 1, "{v3}/{file}");
assert_eq!(text.matches("j < 8u; ++j) mh_mixer(s, 0x9E3779B9u * (64u + j + 1u))").count(), 1, "{v3}/{file}");
}
for file in ["memhard.h", "memhard.metal", "kernel.cl"] {
let text = read(v2, file);
assert_eq!(text.matches("j < 8u").count(), 0, "{v2}/{file}: the v2 text has no multiplier loop");
}
}
// the devnet v3 pack records era 0's stand-in, the devnet genesis hash
let j = json("mx8-devnet-epoch0", "program.json");
assert_eq!(unhex(&j["era_seed_bytes"]), unhex(&j["seed_bytes"]));
assert!(read("mx8-devnet-epoch0", "program.h").contains("#define IGNEUM_ERA_SEED_HEX \"edc4fa844da9dc98"));
assert!(json("mx8-genesis", "program.json").get("era_seed_bytes").is_none());
// the x4 candidate's packs: generator 2, the class in the id, the v2 program of the seed, mixer x4
for (x4, v2) in [("mx4-genesis", "igneum-genesis-mh"), ("mx4-devnet-epoch0", "igneum-devnet-v4-epoch0")] {
let e4 = epoch(x4);
let j = json(x4, "program.json");
assert_eq!(e4.program.generator, GENERATOR_VERSION);
assert_eq!(j["load_class"].as_str().unwrap(), "mx4");
assert_eq!(e4.program.class, LoadClass::MX4);
assert_eq!(e4.program.instrs, epoch(v2).program.instrs);
assert_eq!(e4.dataset.memhard().unwrap().shape(), Shape { mixer_mult: 4, cache_log2_words: 26 });
assert!(read(x4, "memhard.h").contains("j < 4u; ++j) mh_mixer(s, 0x9E3779B9u * (r * 4u + j + 1u))"));
}
}
fn check_export(pack: &str) {
let e = epoch(pack);
let v = json(pack, "vectors.json");
@ -311,7 +434,7 @@ fn check_export(pack: &str) {
expected.extend(["memhard.h", "memhard.metal"]);
}
assert_eq!(out.files.iter().map(|(n, _)| n.as_str()).collect::<Vec<_>>(), expected);
let mut on_disk: Vec<String> = std::fs::read_dir(packs_dir().join(pack))
let mut on_disk: Vec<String> = std::fs::read_dir(pack_dir(pack))
.unwrap()
.map(|d| d.unwrap().file_name().to_string_lossy().to_string())
.filter(|n| !n.starts_with('.'))
@ -350,3 +473,429 @@ fn devnet_pack_is_the_chain_derivation() {
assert_eq!(e.program.instrs, epoch("igneum-devnet-v4-epoch0").program.instrs);
assert_eq!(e.hash_warp(0), epoch("igneum-devnet-v4-epoch0").hash_warp(0));
}
// ---------------------------------------------------------------------------------------------------------
// Era layout packs (5 October 2026, docs/plans/era-layout.md): proto-cuda/packs-ca2-era/era-<n>, n in 0..5, the
// devnet epoch seed and day bytes under the era class of test era seed igneum-era-test/<n>, the width pinned at
// 4 bytes (allowed_widths in program.json). Checked like the pinned packs, plus the one load form of 1.3 by text search.
// ---------------------------------------------------------------------------------------------------------
use igneum_pow::generator::{generate_era, EraParams, V3_ALLOWED};
const ERA_PACKS: [&str; 6] = ["era-0", "era-1", "era-2", "era-3", "era-4", "era-5"];
fn era_packs_dir() -> PathBuf {
PathBuf::from(env!("CARGO_MANIFEST_DIR")).join("../proto-cuda/packs-ca2-era")
}
fn era_read(pack: &str, file: &str) -> String {
let p = era_packs_dir().join(pack).join(file);
std::fs::read_to_string(&p).unwrap_or_else(|e| panic!("read {}: {e}", p.display()))
}
fn era_json(pack: &str, file: &str) -> Value {
serde_json::from_str(&era_read(pack, file)).unwrap_or_else(|e| panic!("{pack}/{file}: {e}"))
}
/// The era class of a pack: the era bytes from program.json (`era_seed_bytes`, the chain's `E_n`; the test seed
/// `igneum-era-test/<n>` of the pack's number gives the same bytes), the allowed set from `era.allowed_widths`
/// (class v3's `V3_ALLOWED`); the pack's recorded stream words and draw must be the class's.
fn era_class(pack: &str) -> LoadClass {
let j = era_json(pack, "program.json");
let n: u64 = pack.trim_start_matches("era-").parse().unwrap();
let eb = unhex(&j["era_seed_bytes"]);
assert_eq!(eb, EraParams::test_era_bytes(&format!("igneum-era-test/{n}")).to_vec(), "{pack}: the era bytes of test seed {n}");
let allowed: Vec<u8> = j["era"]["allowed_widths"].as_array().unwrap().iter().map(|v| v.as_u64().unwrap() as u8).collect();
assert_eq!(allowed, V3_ALLOWED.to_vec(), "{pack}: class v3's width set");
// the measurement packs of 5 October 2026: the era layout over version 2's construction (mixer x1, the genesis
// cache), generator 3 and the era bytes recorded; the chain's class v3 composes the same draw over LoadClass::MX4
// (generator tests era_programs_are_accepted and program_classes), and the integration re-exports these packs
let c = LoadClass::era(LoadClass::V2, &eb, &allowed);
let e = c.era.unwrap();
let words: Vec<u32> = j["era"]["seed_words"].as_array().unwrap().iter().map(hex32).collect();
assert_eq!(e.words.to_vec(), words, "{pack}: era seed words");
assert_eq!(e.width_words as u64, j["era"]["width_words"].as_u64().unwrap(), "{pack}: width");
assert_eq!(e.stride_mul, hex32(&j["era"]["stride_mul"]), "{pack}: stride mul");
assert_eq!(e.stride_rot as u64, j["era"]["stride_rot"].as_u64().unwrap(), "{pack}: stride rot");
let pos: Vec<u8> = j["era"]["interleave"].as_array().unwrap().iter().map(|v| v.as_u64().unwrap() as u8).collect();
assert_eq!(e.pos.to_vec(), pos, "{pack}: interleave");
c
}
fn era_epoch(pack: &str) -> &'static Epoch {
static E: OnceLock<Vec<(String, Epoch)>> = OnceLock::new();
let all = E.get_or_init(|| {
ERA_PACKS
.iter()
.map(|p| {
let j = era_json(p, "program.json");
let seed = j["seed"].as_str().unwrap();
let seed_bytes = unhex(&j["seed_bytes"]);
let day_bytes = unhex(&j["dataset"]["day_bytes"]);
assert_eq!(j["dataset_mode"].as_str().unwrap(), "memory-hard");
let log2 = j["dataset"]["log2_words"].as_u64().unwrap() as u32;
let class = era_class(p);
assert_eq!(log2, igneum_pow::verify::DEFAULT_DATASET_LOG2);
let eb = unhex(&j["era_seed_bytes"]);
let program = generate_era(seed, &seed_bytes, LoadClass::V2, &eb, &V3_ALLOWED);
assert_eq!(program.class, class, "{p}: the pack's era class");
assert_eq!(program.generator, GENERATOR_VERSION_V3);
assert_eq!(program.era_bytes.as_deref(), Some(&eb[..]));
let mut dataset = DatasetSource::from_key(igneum_pow::seed::seed_words_from_bytes(&day_bytes), DatasetMode::MemoryHard, log2);
dataset.key_bytes = day_bytes;
(p.to_string(), Epoch { program, dataset })
})
.collect()
});
&all.iter().find(|(n, _)| n == pack).unwrap().1
}
/// program.json of an era pack: generator, attempt, id, class, every instruction with width, win and off, and the
/// program passes the acceptance rule.
#[test]
fn era_program_json_matches() {
for pack in ERA_PACKS {
let j = era_json(pack, "program.json");
let p = &era_epoch(pack).program;
assert_eq!(j["generator"].as_u64().unwrap() as u32, GENERATOR_VERSION_V3, "{pack}: a class v3 pack");
assert_eq!(j["program_class"].as_str().unwrap(), "v3");
assert_eq!(j["attempt"].as_u64().unwrap() as u32, p.attempt, "{pack}: attempt");
assert_eq!(hex64(&j["program_id"]), p.program_id(), "{pack}: program id");
assert_eq!(j["load_class"].as_str().unwrap(), p.class.name(), "{pack}: class");
assert_eq!(j["bytes_per_hash"].as_u64().unwrap() as usize, p.bytes_per_hash());
assert_eq!(p.loads_per_hash(), 8 * LOAD_SLOTS);
assert!(accept::check(p).is_ok(), "{pack}: acceptance");
let instrs = j["instructions"].as_array().unwrap();
assert_eq!(instrs.len(), p.instrs.len());
for (k, (ins, ji)) in p.instrs.iter().zip(instrs).enumerate() {
assert_eq!(Op::from_name(ji["op"].as_str().unwrap()).unwrap(), ins.op, "{pack} #{k} op");
assert_eq!(ji["dst"].as_u64().unwrap(), ins.dst as u64);
assert_eq!(ji["src"].as_u64().unwrap(), ins.src as u64);
assert_eq!(hex32(&ji["imm"]), ins.imm);
assert_eq!(ji["width"].as_u64().unwrap(), ins.width as u64, "{pack} #{k} width");
assert_eq!(ji["win"].as_u64().unwrap(), ins.win as u64, "{pack} #{k} win");
assert_eq!(ji["off"].as_u64().unwrap(), ins.off as u64, "{pack} #{k} off");
if ins.op == Op::Load {
assert_eq!(ins.width, p.class.era.unwrap().width_words);
assert!(ins.win <= 2 && (ins.off as u32) < (1u32 << ins.win));
}
}
}
}
/// The six era packs are the same program seed under six draws: the instruction lists agree, the widths and layouts
/// follow the draw, and the dataset words differ from the linear layout exactly when the interleave is not linear.
#[test]
fn era_packs_share_the_program_and_differ_in_layout() {
let linear = &epoch("igneum-devnet-v4-epoch0").dataset;
for pack in ERA_PACKS {
let e = era_epoch(pack);
assert_eq!(e.program.seed_bytes, epoch("igneum-devnet-v4-epoch0").program.seed_bytes, "{pack}: the devnet seed");
let strip = |p: &igneum_pow::generator::Program| {
p.instrs.iter().map(|i| (i.op, i.dst, i.src, i.src2, i.imm, i.imm2, i.rot, i.bit, i.mask, i.win, i.off)).collect::<Vec<_>>()
};
assert_eq!(strip(&e.program), strip(&era_epoch("era-0").program), "{pack}: same stream as era-0");
let l = e.program.class.layout();
let same_at_1 = (0..64u32).all(|w| e.dataset_word(w * 977 + 1) == linear.word(w * 977 + 1));
assert_eq!(same_at_1, l.is_linear(), "{pack}: layout {:?}", l.pos);
assert_eq!(e.dataset_word(0), linear.word(0), "{pack}: word 0 is item 0 word 0 in every layout");
// the chain's shared day cache serves every era: the day's dataset source is the pinned pack's, bit for bit
assert_eq!(e.dataset.key, linear.key);
assert_eq!(e.dataset.memhard().unwrap().cache.fnv1a64(), linear.memhard().unwrap().cache.fnv1a64());
}
}
#[test]
fn era_dataset_words_and_vectors_match() {
for pack in ERA_PACKS {
let v = era_json(pack, "vectors.json");
let e = era_epoch(pack);
let ds = &e.dataset;
let head: Vec<u32> = v["dataset_head"].as_array().unwrap().iter().map(hex32).collect();
for (i, h) in head.iter().enumerate() {
assert_eq!(e.dataset_word(i as u32), *h, "{pack}: dataset[{i}]");
}
assert_eq!(e.dataset_word(ds.mask), hex32(&v["dataset_last"]), "{pack}: dataset[MASK]");
for s in v["dataset_samples"].as_array().unwrap() {
let idx = s["index"].as_u64().unwrap() as u32;
assert_eq!(e.dataset_word(idx), hex32(&s["value"]), "{pack}: dataset[{idx}]");
}
assert_eq!(ds.memhard().unwrap().cache.fnv1a64(), hex64(&v["cache_fnv1a64"]));
let mut n = 0;
for w in v["warps"].as_array().unwrap() {
let base = w["base_nonce"].as_u64().unwrap() as u32;
let expected: Vec<u64> = w["expected"].as_array().unwrap().iter().map(hex64).collect();
let got = e.hash_warp(base);
for lane in 0..32 {
assert_eq!(got[lane], expected[lane], "{pack}: base {base} lane {lane}");
n += 1;
}
assert_eq!(e.hash(base + 7), expected[7]);
}
assert_eq!(n, 96, "{pack}");
}
}
/// Every emitted file of every era pack matches the emitters byte for byte, the export reproduces vectors.json and
/// vectors.h, and every dataset load in every hash kernel has the one era form (no plain `ds[rN & mask]` remains).
#[test]
fn era_emitted_sources_match_and_loads_have_the_era_form() {
for pack in ERA_PACKS {
let e = era_epoch(pack);
let day = era_json(pack, "program.json")["dataset"]["day"].as_str().unwrap().to_string();
let source = era_json(pack, "vectors.json")["source"].as_str().unwrap().to_string();
let out = export_pack(e, &day, &source);
for (name, text) in &out.files {
let want = era_read(pack, name);
assert!(text == &want, "{pack}/{name} differs from the emitter");
}
let mut on_disk: Vec<String> = std::fs::read_dir(era_packs_dir().join(pack))
.unwrap()
.map(|d| d.unwrap().file_name().to_string_lossy().to_string())
.filter(|n| !n.starts_with('.') && n != "seeds.txt")
.collect();
on_disk.sort();
let mut want: Vec<String> = out.files.iter().map(|(n, _)| n.clone()).collect();
want.sort();
assert_eq!(on_disk, want, "{pack}: the pack holds the export's files and seeds.txt only");
let era = e.program.class.era.unwrap();
let mul = format!("0x{:08x}u", era.stride_mul);
for (file, mask) in [
("kernel.cu", "mask"),
("kernel_bound.cu", "mask"),
("kernel.cl", "mask"),
("kernel_bound.cl", "mask"),
("program.metal", "MASK"),
("program_bound.metal", "MASK"),
] {
let text = era_read(pack, file);
let kernels = if file.starts_with("kernel_bound") || file == "kernel.cu" || file == "kernel.cl" || file.starts_with("program") { 1 } else { 1 };
// kernel_bound.cl carries igneum_hash and igneum_hash_bound: two kernels
let kernels = if file == "kernel_bound.cl" { 2 } else { kernels };
let era_form: usize = text
.lines()
.filter(|l| l.contains("rotl_imm(r") && l.contains(&format!(" * {mul}, {}u) & ", era.stride_rot)) && l.contains(&format!(") & {mask}")))
.filter(|l| l.contains("ds[") || l.contains("dataset[") || l.contains("b_ = "))
.count();
assert_eq!(era_form, LOAD_SLOTS * kernels, "{pack}/{file}: {} loads of the era form", LOAD_SLOTS * kernels);
let plain = text.lines().filter(|l| l.contains("ds[r") || l.contains("dataset[r")).count();
assert_eq!(plain, 0, "{pack}/{file}: a load without the era form");
}
// the layout helpers appear exactly when the layout is not linear
let mh = era_read(pack, "memhard.h");
assert_eq!(mh.contains("mh_addr("), !era.layout().is_linear(), "{pack}: memhard.h layout helpers");
}
}
/// An era pack's dataset is a prefix at every size of at least 2^16 words: the 2^20-word source gives the pack's
/// words below 2^20.
#[test]
fn era_dataset_is_a_prefix_at_smaller_sizes() {
for pack in ["era-1", "era-3"] {
let e = era_epoch(pack);
let small = DatasetSource::from_key(e.dataset.key, DatasetMode::MemoryHard, 20);
let l = e.program.class.layout();
for w in [0u32, 1, 2, 3, 16, 255, 4096, 65_535, 65_536, 0x000f_ffff] {
assert_eq!(small.word_at(l, w), e.dataset_word(w), "{pack}: w {w}");
}
}
}
// ---------------------------------------------------------------------------------------------------------
// Hot-table experiment (5 October 2026, docs/plans/hot-table.md): the five packs under proto-cuda/packs-ca2-hot/ are
// pinned the same way (program, vectors, every emitted file byte for byte), plus the hot table's fingerprint and the
// one-form load check: exactly 16 - k masked dataset loads and exactly k hot loads in every hash kernel.
// ---------------------------------------------------------------------------------------------------------
const HOT_PACKS: [&str; 8] = ["hot32k4", "hot64k4", "hot96k4", "hot64k2", "hot64k8", "hot32k4a", "hot64k4a", "hot96k4a"];
fn hot_packs_dir() -> PathBuf {
PathBuf::from(env!("CARGO_MANIFEST_DIR")).join("../proto-cuda/packs-ca2-hot")
}
fn hread(pack: &str, file: &str) -> String {
let p = hot_packs_dir().join(pack).join(file);
std::fs::read_to_string(&p).unwrap_or_else(|e| panic!("read {}: {e}", p.display()))
}
fn hjson(pack: &str, file: &str) -> Value {
serde_json::from_str(&hread(pack, file)).unwrap_or_else(|e| panic!("{pack}/{file}: {e}"))
}
/// The epoch of a hot pack from program.json alone: the class from `load_class`, the program from the seed bytes,
/// the dataset from the day bytes, the hot table from the seed bytes (what `Epoch::from_seed_bytes_class` does).
fn hepoch(pack: &str) -> &'static Epoch {
static E: OnceLock<Vec<(String, Epoch)>> = OnceLock::new();
let all = E.get_or_init(|| {
HOT_PACKS
.iter()
.map(|p| {
let j = hjson(p, "program.json");
let seed = j["seed"].as_str().unwrap();
let seed_bytes = unhex(&j["seed_bytes"]);
let day_bytes = unhex(&j["dataset"]["day_bytes"]);
assert_eq!(j["dataset_mode"].as_str().unwrap(), "memory-hard");
let class = LoadClass::parse(j["load_class"].as_str().unwrap()).unwrap();
assert_eq!(class.name(), *p, "the pack directory is the class name");
let log2 = j["dataset"]["log2_words"].as_u64().unwrap() as u32;
let program = generate_from_seed_bytes_class(seed, &seed_bytes, class);
let mut dataset =
DatasetSource::from_key(igneum_pow::seed::seed_words_from_bytes(&day_bytes), DatasetMode::MemoryHard, log2);
dataset.key_bytes = day_bytes;
dataset.attach_hot_for(&program);
(p.to_string(), Epoch { program, dataset })
})
.collect()
});
&all.iter().find(|(n, _)| n == pack).unwrap().1
}
fn hassert_same_text(pack: &str, file: &str, got: &str) {
let want = hread(pack, file);
if got != want {
let (gl, wl): (Vec<&str>, Vec<&str>) = (got.lines().collect(), want.lines().collect());
for i in 0..gl.len().max(wl.len()) {
let g = gl.get(i).copied().unwrap_or("<eof>");
let w = wl.get(i).copied().unwrap_or("<eof>");
if g != w {
panic!("{pack}/{file} differs at line {}:\n pack: {w}\n rust: {g}", i + 1);
}
}
panic!("{pack}/{file} differs only in trailing bytes (len {} vs {})", got.len(), want.len());
}
}
#[test]
fn hot_packs_program_and_vectors() {
for pack in HOT_PACKS {
let e = hepoch(pack);
let p = &e.program;
let j = hjson(pack, "program.json");
let h = p.class.hot.unwrap();
assert_eq!(j["generator"].as_u64().unwrap() as u32, GENERATOR_VERSION);
assert_eq!(j["attempt"].as_u64().unwrap() as u32, p.attempt);
assert_eq!(hex64(&j["program_id"]), p.program_id(), "{pack}: program id");
let dataset_slots = if h.added { 16 } else { 16 - h.k as usize };
assert_eq!(j["loads_per_hash"].as_u64().unwrap() as usize, (dataset_slots + h.k as usize) * 8);
assert_eq!(j["hot_table"]["mb"].as_u64().unwrap(), h.mb as u64);
assert_eq!(j["hot_table"]["slots"].as_u64().unwrap(), h.k as u64);
assert_eq!(j["hot_table"]["dataset_slots"].as_u64().unwrap() as usize, dataset_slots);
assert_eq!(j["hot_table"]["words"].as_u64().unwrap() as u32, p.hot_words());
assert_eq!(j["op_mix"]["hot"].as_u64().unwrap(), h.k as u64, "{pack}: k hot instructions");
assert_eq!(j["op_mix"]["load"].as_u64().unwrap() as usize, dataset_slots);
assert_eq!(p.items_per_warp(), dataset_slots * 8 * 32);
assert!(accept::check(p).is_ok(), "{pack}: passes the acceptance rule");
let v2 = &epoch("igneum-genesis-mh").program;
if !h.added {
// replaced form: the version 2 genesis program with k loads redirected (attempt 0 on both)
assert_eq!(p.attempt, v2.attempt);
for (a, b) in p.instrs.iter().zip(v2.instrs.iter()) {
if a.op == Op::Hot {
assert_eq!(b.op, Op::Load);
} else {
assert_eq!(a, b);
}
}
} else {
// added form: 16 + k load slots, so another slot draw and another program; 16 dataset loads stay
assert_ne!(p.instrs, v2.instrs);
assert_eq!(p.instrs.iter().filter(|i| i.op == Op::Load).count(), 16);
assert!(p.instrs.iter().all(|i| i.width == 1));
}
// the hot table: the pack's head, last line and fingerprint
let v = hjson(pack, "vectors.json");
let t = e.dataset.hot.as_ref().unwrap();
assert_eq!(t.n_words(), p.hot_words());
let head: Vec<u32> = v["hot_head"].as_array().unwrap().iter().map(hex32).collect();
assert_eq!(&t.words()[..16], &head[..]);
let last: Vec<u32> = v["hot_last_line"].as_array().unwrap().iter().map(hex32).collect();
assert_eq!(&t.words()[t.words().len() - 16..], &last[..]);
assert_eq!(t.fnv1a64(), hex64(&v["hot_fnv1a64"]), "{pack}: hot_fnv1a64");
assert_eq!(t.key, igneum_pow::memhard::hot_key(&p.seed_bytes));
// the cache is the day's, unchanged by the class
assert_eq!(e.dataset.memhard().unwrap().cache.fnv1a64(), 0x48c4f5bf24166b2e);
// 96 vectors
let warps = v["warps"].as_array().unwrap();
assert_eq!(warps.len(), 3);
for w in warps {
let base = w["base_nonce"].as_u64().unwrap() as u32;
let expected: Vec<u64> = w["expected"].as_array().unwrap().iter().map(hex64).collect();
let got = e.hash_warp(base);
for lane in 0..32 {
assert_eq!(got[lane], expected[lane], "{pack}: base {base} lane {lane}");
}
assert_eq!(e.hash(base + 5), expected[5]);
}
// the dataset words are the day's
let head: Vec<u32> = v["dataset_head"].as_array().unwrap().iter().map(hex32).collect();
for (i, hd) in head.iter().enumerate() {
assert_eq!(e.dataset.word(i as u32), *hd);
}
}
// the same k at three sizes: identical programs, three fingerprints, three vector sets
let a = hepoch("hot32k4");
let b = hepoch("hot64k4");
let c = hepoch("hot96k4");
assert_eq!(a.program.instrs, b.program.instrs);
assert_eq!(b.program.instrs, c.program.instrs);
assert_ne!(a.hash_warp(0), b.hash_warp(0));
assert_ne!(b.hash_warp(0), c.hash_warp(0));
}
#[test]
fn hot_packs_emitted_sources_and_load_forms() {
for pack in HOT_PACKS {
let e = hepoch(pack);
let p = &e.program;
let k = p.class.hot.unwrap().k as usize;
let dataset_loads = p.class.dataset_slots();
let day = hjson(pack, "program.json")["dataset"]["day"].as_str().unwrap().to_string();
let mp = &e.dataset.memhard().unwrap().params;
hassert_same_text(pack, "kernel.cu", &cuda_kernel(p, Some(mp)));
hassert_same_text(pack, "kernel_bound.cu", &cuda_kernel_bound(p, Some(mp)));
hassert_same_text(pack, "program.metal", &metal_program(p, e.dataset.log2_words, LoadSource::Stored));
hassert_same_text(pack, "program_bound.metal", &metal_program_bound(p, e.dataset.log2_words));
hassert_same_text(pack, "kernel.cl", &opencl_kernel(p, Some(mp)));
hassert_same_text(pack, "kernel_bound.cl", &opencl_kernel_bound(p, Some(mp)));
hassert_same_text(pack, "program.h", &program_header(p, &day, &e.dataset));
hassert_same_text(pack, "memhard.h", &cuda_memhard_header(p, mp));
hassert_same_text(pack, "memhard.metal", &metal_memhard_for(p, mp));
assert_ne!(metal_memhard_for(p, mp), metal_memhard(mp), "{pack}: the hot fill kernel is in memhard.metal");
let got = program_json(p, &day, &e.dataset);
hassert_same_text(pack, "program.json", &got);
let _: Value = serde_json::from_str(&got).expect("program.json is valid JSON");
let v = hjson(pack, "vectors.json");
let out = export_pack(e, &day, v["source"].as_str().unwrap());
let file = |name: &str| -> &str { &out.files.iter().find(|(n, _)| n == name).unwrap().1 };
hassert_same_text(pack, "vectors.json", file("vectors.json"));
hassert_same_text(pack, "vectors.h", file("vectors.h"));
assert_eq!(out.files.len(), 12);
// One form per dialect, exactly 16 - k masked dataset loads and k hot loads in every hash kernel; the fill
// kernel is present once per source that builds the table.
for (file, load, masked, hot) in [
("kernel.cu", "ds[r", " & mask]", "hot[__umulhi(r"),
("kernel_bound.cu", "ds[r", " & mask]", "hot[__umulhi(r"),
("program.metal", "dataset[r", " & MASK]", "hot[mulhi(r"),
("program_bound.metal", "dataset[r", " & MASK]", "hot[mulhi(r"),
("kernel.cl", "ds[r", " & mask]", "hot[mul_hi(r"),
] {
let text = hread(pack, file);
assert_eq!(text.matches(load).count(), dataset_loads, "{pack}/{file}: {dataset_loads} dataset loads");
assert_eq!(text.matches(masked).count(), dataset_loads, "{pack}/{file}: masked loads");
assert_eq!(text.matches(hot).count(), k, "{pack}/{file}: {k} hot loads");
assert!(text.contains(&format!("#define HOT_WORDS 0x{:08x}u", p.hot_words())), "{pack}/{file}: HOT_WORDS literal");
}
// kernel_bound.cl carries both kernels
let text = hread(pack, "kernel_bound.cl");
assert_eq!(text.matches("hot[mul_hi(r").count(), 2 * k);
assert_eq!(text.matches("ds[r").count(), 2 * dataset_loads);
for file in ["kernel.cu", "kernel.cl", "kernel_bound.cl", "memhard.metal"] {
assert_eq!(hread(pack, file).matches("igneum_hot_fill(").count(), 1, "{pack}/{file}: one hot fill kernel");
}
assert_eq!(hread(pack, "memhard.h").matches("void ht_segment(").count(), 1);
let ph = hread(pack, "program.h");
assert!(ph.contains(&format!("#define IGNEUM_HOT_MB {}", p.class.hot.unwrap().mb)));
assert!(ph.contains(&format!("#define IGNEUM_HOT_SLOTS {k}")));
assert!(ph.contains("igneum_launch_hot_fill("));
}
}

766
igneum-pow/tests/scratch.rs Normal file
View file

@ -0,0 +1,766 @@
//! Soundness tests of layer 3 of `docs/plans/counter-asic-2.md`: the per-warp scratch with read-modify-writes
//! (variant 5 of the read-width experiment, `LoadClass::scratch(k, kb)`). Analysis and results:
//! `docs/analysis/scratch-soundness.md`. Every test is parametric over the class's slot count
//! (`scratch_slots_per_lane()`), so the 32 and 128 KiB geometries and any later one run the same checks.
//!
//! What runs under plain `cargo test`:
//! 1. `rewrite_is_a_bijection_of_the_fold_value`, `fill_is_a_bijection_of_the_nonce`: the written words as
//! functions (question 1).
//! 2. `written_words_unbiased_and_rehit_rates`: bit bias of every written word over 2^11 units x 3 seeds per class
//! (the TESTS.md section 3 shape), and the measured slot re-hit rate against the birthday formula (question 2).
//! 3. `edge_programs_match_the_hand_model`: hand-built programs that drive every read-modify-write of a hash to
//! slot 0, slot MASK, through out-of-range registers, to one slot per lane, alternating two slots, and 16
//! read-modify-writes per iteration on one slot; the interpreter against an independent hand model, and the
//! hand model shown to have teeth (question 3, CPU half).
//! 4. `scr_packs_regenerate_and_pass_the_static_scratch_check`: every emitted kernel of every scr pack under
//! `proto-cuda/packs-readwidth` regenerates from its program.json and passes the static scratch-mask check;
//! the check is shown to fail on four deliberate breaks (question 4).
//! 5. `fuzz_scr_programs_cpu`: 200 generated scratch programs over the six classes, generator contract on every
//! instruction, 4 units each at base nonces across the 32-bit range including the wrap; with
//! `IGNEUM_SCRATCH_PACKS_OUT=<dir>` it also writes the packs (and the edge packs) for the Metal runs of
//! `proto-metal/packbench` (question 3 GPU half, question 4, `TESTS.md` section 9 shape).
use igneum_pow::emit::{
cuda_kernel, cuda_kernel_bound, export_pack, metal_program, metal_program_bound, opencl_kernel,
opencl_kernel_bound, vectors_json, LoadSource,
};
use igneum_pow::generator::{
generate_class, generate_from_seed_bytes_class, Instr, LoadClass, Op, Program, GENERATOR_VERSION, INSTR_COUNT,
ITERATIONS, LANES,
};
use igneum_pow::seed::{seed_words_from_bytes, SplitMix64};
use igneum_pow::verify::{
fold_words, interpret_warp_scratch, scratch_fill, scratch_rewrite, splitmix32, DatasetMode, DatasetSource,
Epoch, ScratchEvent, FOLD_MUL, FOLD_ROT,
};
use serde_json::Value;
use std::collections::HashMap;
use std::path::PathBuf;
/// The classes under study: the two capped geometries (32 and 128 KiB per warp: 64 and 256 slots per lane) at the
/// RMW shares the readwidth branch measures.
const CLASSES: [&str; 6] = ["scr2k32", "scr4k32", "scr8k32", "scr2k128", "scr4k128", "scr8k128"];
fn class(name: &str) -> LoadClass {
LoadClass::parse(name).unwrap_or_else(|| panic!("class {name}"))
}
// ---------------------------------------------------------------------------------------------------------------
// 1. The written words as functions (question 1)
// ---------------------------------------------------------------------------------------------------------------
/// For a fixed slot content `w`, each of the three rewritten words is a bijection of the fold value `x`
/// (`x ^ w1`, `rotl(x, 7) ^ w2`, `x + w0`), so the rewrite is injective in `x` and a uniform `x` gives a uniform
/// word in every position. Checked over 2^16 consecutive `x` for 16 random `w`.
#[test]
fn rewrite_is_a_bijection_of_the_fold_value() {
let mut rng = SplitMix64::new(0x7363_7261_7463_6801);
for _ in 0..16 {
let w = [rng.next() as u32, rng.next() as u32, rng.next() as u32];
let x0 = rng.next() as u32;
let mut seen = [vec![false; 1 << 16], vec![false; 1 << 16], vec![false; 1 << 16]];
for i in 0..(1u32 << 16) {
let x = x0.wrapping_add(i);
let out = scratch_rewrite(x, &w);
for j in 0..3 {
// a bijection of x maps 2^16 consecutive x to 2^16 distinct words; the low 16 bits alone are
// distinct for the xor words (x ^ c) and for the add word (x + c), since both act on the low 16
// bits as bijections of the low 16 bits of x; the rotl word is checked on its rotated-back bits
let key = if j == 1 { out[j].rotate_right(7) & 0xffff } else { out[j] & 0xffff };
assert!(!seen[j][key as usize], "word {j} repeats inside 2^16 consecutive x");
seen[j][key as usize] = true;
}
}
}
// The rewrite inverts: from the old content and any ONE written word the fold value is recovered, so a
// rewritten slot carries exactly 32 bits of new state (the point of question 2's arithmetic).
let w = [0x1234_5678, 0x9abc_def0, 0x0fed_cba9];
let x = 0xdead_beef;
let out = scratch_rewrite(x, &w);
assert_eq!(out[0] ^ w[1], x);
assert_eq!((out[1] ^ w[2]).rotate_right(7), x);
assert_eq!(out[2].wrapping_sub(w[0]), x);
}
/// For a fixed (seed, slot, j) the fill is a bijection of the lane nonce: `splitmix32` is a bijection of its
/// 32-bit input and the input `((base + lane) ^ s) + c` is a bijection of `base + lane`. Over 2^16 consecutive
/// nonces no fill word repeats, for 8 slots x 3 words.
#[test]
fn fill_is_a_bijection_of_the_nonce() {
let seed = seed_words_from_bytes(b"igneum-genesis");
for slot in [0u32, 1, 63, 64, 255, 1023, 2047] {
for j in 0..3u32 {
let mut words: Vec<u32> = (0..(1u32 << 16)).map(|n| scratch_fill(&seed, n, 0, slot, j)).collect();
words.sort_unstable();
words.dedup();
assert_eq!(words.len(), 1 << 16, "slot {slot} word {j}: fill words of 2^16 consecutive nonces are distinct");
}
}
// base + lane is the lane nonce: the fill of lane l at base b is the fill of lane 0 at base b + l
assert_eq!(scratch_fill(&seed, 0x1000, 7, 5, 2), scratch_fill(&seed, 0x1007, 0, 5, 2));
// and it wraps with the nonce: base 0xffffffe0, lane 31 is nonce 0xffffffff; lane 32 would be nonce 0
assert_eq!(scratch_fill(&seed, 0xffff_ffe0, 32, 5, 2), scratch_fill(&seed, 0, 0, 5, 2));
// the three word positions of one slot and nonce are three different permutation outputs
let f: Vec<u32> = (0..3).map(|j| scratch_fill(&seed, 12345, 7, 17, j)).collect();
assert!(f[0] != f[1] && f[1] != f[2] && f[0] != f[2]);
}
// ---------------------------------------------------------------------------------------------------------------
// 2. Uniformity of the written words and the slot re-hit rate (questions 1 and 2)
// ---------------------------------------------------------------------------------------------------------------
/// Birthday arithmetic: the expected number of distinct slots after `n` uniform draws from `s` slots.
fn expected_distinct(s: usize, n: usize) -> f64 {
let s = s as f64;
s * (1.0 - (1.0 - 1.0 / s).powi(n as i32))
}
struct ClassStats {
units: usize,
events: usize,
hits: usize,
/// ones count per bit of the written words, 3 x 32
ones: [[u64; 32]; 3],
/// ones count per bit of written XOR read (the change the rewrite makes to the slot)
delta_ones: [[u64; 32]; 3],
/// re-hit depth histogram: how many earlier RMWs the slot had seen in this unit (0 = first touch)
depth: Vec<usize>,
max_depth: usize,
/// how often each slot index was addressed (the slot comes from a register's low bits)
slot_hist: Vec<u64>,
}
fn class_stats(name: &str, seeds: &[&str], units_per_seed: usize) -> ClassStats {
let c = class(name);
let mut st = ClassStats {
units: 0,
events: 0,
hits: 0,
ones: [[0; 32]; 3],
delta_ones: [[0; 32]; 3],
depth: vec![0; 256],
max_depth: 0,
slot_hist: vec![0; c.scratch_slots_per_lane()],
};
let ds = DatasetSource::new("2026-10-03", DatasetMode::ClosedForm, 28);
for seed in seeds {
let p = generate_class(seed, c);
assert_eq!(p.scratch_ops_per_hash(), c.scratch_slots() * ITERATIONS);
for u in 0..units_per_seed {
let base = (u as u32).wrapping_mul(32).wrapping_add(0x4000_0000);
let (_, ev) = interpret_warp_scratch(&p, &p.seed, base, &ds, true);
assert_eq!(ev.len(), p.scratch_ops_per_hash() * LANES);
let mut count: HashMap<(u8, u32), usize> = HashMap::new();
for e in &ev {
assert!(e.slot < c.scratch_slots_per_lane() as u32, "slot inside the lane's scratch");
let d = count.entry((e.lane, e.slot)).or_insert(0);
assert_eq!(e.hit, *d > 0, "hit flag agrees with the unit's own history");
assert_eq!(e.written, scratch_rewrite(e.x, &e.read));
if !e.hit {
let fill = [
scratch_fill(&p.seed, base, e.lane as u32, e.slot, 0),
scratch_fill(&p.seed, base, e.lane as u32, e.slot, 1),
scratch_fill(&p.seed, base, e.lane as u32, e.slot, 2),
];
assert_eq!(e.read, fill, "a first touch reads the fill");
}
st.depth[(*d).min(255)] += 1;
st.max_depth = st.max_depth.max(*d);
st.slot_hist[e.slot as usize] += 1;
*d += 1;
st.events += 1;
st.hits += e.hit as usize;
for j in 0..3 {
for b in 0..32 {
st.ones[j][b] += ((e.written[j] >> b) & 1) as u64;
st.delta_ones[j][b] += (((e.written[j] ^ e.read[j]) >> b) & 1) as u64;
}
}
}
st.units += 1;
}
}
st
}
/// Bit bias of every written word (and of the change each rewrite makes) within 6 sigma of a fair coin, over
/// 3 seeds x 2^11 units per class (131,072 hashes per seed set); the slot re-hit rate against the birthday
/// formula within 3 percent relative. The table printed here is the one in the analysis.
#[test]
fn written_words_unbiased_and_rehit_rates() {
let seeds = ["igneum-genesis", "igneum-genesis/stats1", "igneum-genesis/stats2"];
let units = 1usize << 11;
println!("class | slots/lane | RMW/hash | events | re-hits | re-hit % | birthday % | slot chi2 z (spread) | max depth | max bias sigma | max delta bias sigma");
for name in CLASSES {
let c = class(name);
let st = class_stats(name, &seeds, units);
let n = st.events as f64;
let sigma = (n / 4.0).sqrt();
let mut worst = 0.0f64;
let mut worst_delta = 0.0f64;
for j in 0..3 {
for b in 0..32 {
let z = (st.ones[j][b] as f64 - n / 2.0).abs() / sigma;
let zd = (st.delta_ones[j][b] as f64 - n / 2.0).abs() / sigma;
assert!(z <= 6.0, "{name}: written word {j} bit {b} biased: {z:.2} sigma");
assert!(zd <= 6.0, "{name}: rewrite delta word {j} bit {b} biased: {zd:.2} sigma");
worst = worst.max(z);
worst_delta = worst_delta.max(zd);
}
}
let per_lane_hash = c.scratch_slots() * ITERATIONS;
let s = c.scratch_slots_per_lane();
let exp_hits = per_lane_hash as f64 - expected_distinct(s, per_lane_hash);
let exp_pct = 100.0 * exp_hits / per_lane_hash as f64;
let got_pct = 100.0 * st.hits as f64 / st.events as f64;
// chi-square of the slot histogram against uniform (df = s - 1): the slot is a register's low bits, and
// the measured re-hit rate runs above the uniform birthday rate (the finding of the analysis, question 2)
let expect_per_slot = n / s as f64;
let chi2: f64 = st.slot_hist.iter().map(|&h| (h as f64 - expect_per_slot).powi(2) / expect_per_slot).sum();
let chi2_z = (chi2 - (s as f64 - 1.0)) / (2.0 * (s as f64 - 1.0)).sqrt();
let hot = *st.slot_hist.iter().max().unwrap() as f64 / expect_per_slot;
let cold = *st.slot_hist.iter().min().unwrap() as f64 / expect_per_slot;
println!(
"{name} | {s} | {per_lane_hash} | {} | {} | {got_pct:.2} | {exp_pct:.2} | {chi2_z:.1} (hottest slot {hot:.2}x, coldest {cold:.2}x) | {} | {worst:.2} | {worst_delta:.2}",
st.events, st.hits, st.max_depth
);
// a regression band, not a uniformity claim: the rate sits between the uniform birthday rate and twice it
assert!(
got_pct >= 0.9 * exp_pct && got_pct <= 2.0 * exp_pct,
"{name}: re-hit rate {got_pct:.2}% against birthday {exp_pct:.2}%"
);
// depth histogram: the number of earlier RMWs a re-hit slot had seen in the unit
let shown: Vec<String> = st.depth.iter().take(st.max_depth + 1).enumerate().map(|(d, n)| format!("{d}:{n}")).collect();
println!(" depth histogram {}", shown.join(" "));
}
}
// ---------------------------------------------------------------------------------------------------------------
// 3. Hand-built edge programs against an independent hand model (question 3, CPU half)
// ---------------------------------------------------------------------------------------------------------------
fn ins(op: Op, dst: u8, src: u8) -> Instr {
Instr { op, dst, src, src2: 0, imm: 0, imm2: 0, rot: 1, bit: 0, mask: 1, width: 1, win: 0, off: 0 }
}
fn add_imm(dst: u8, src: u8, imm: u32) -> Instr {
Instr { op: Op::Add, dst, src, src2: 0, imm, imm2: imm, rot: 1, bit: 0, mask: 1, width: 1, win: 0, off: 0 }
}
/// A hand-built program of class `c` named `name` (its seed is the name, so its fill words and init words are
/// its own). These bypass the generator and the acceptance rule, like `TESTS.md` section 2; `sub r, r` zeroes a
/// register as the Swift edge set does.
fn edge(name: &str, c: LoadClass, instrs: Vec<Instr>) -> Program {
let seed_string = format!("igneum-scratch-edge/{name}");
let seed_bytes = seed_string.as_bytes().to_vec();
let k = instrs.iter().filter(|i| i.op == Op::Scratch).count();
assert_eq!(k, c.scratch_slots(), "{name}: the class carries the program's scratch count");
Program {
seed: seed_words_from_bytes(&seed_bytes),
seed_string,
seed_bytes,
generator: GENERATOR_VERSION,
attempt: 0,
class: c,
era_bytes: None,
instrs,
}
}
/// The edge set for a scratch of `kb` KiB per warp. Each entry: (name, what it drives, program).
fn edge_programs(kb: u8) -> Vec<(String, &'static str, Program)> {
let m = LoadClass::scratch(1, kb).scratch_slot_mask();
let dsts = [2u8, 3, 4, 5, 6, 7, 0, 2, 3, 4, 5, 6, 7, 0, 2, 3];
let scr = |n: usize, src: u8| -> Vec<Instr> { (0..n).map(|i| ins(Op::Scratch, dsts[i], src)).collect() };
let mut v = Vec::new();
// every RMW of the hash to slot 0 through a zero register: 64 dependent RMWs on one slot per lane
let mut p = vec![ins(Op::Sub, 1, 1)];
p.extend(scr(8, 1));
v.push(("slot0".to_string(), "r1 = 0: every RMW to slot 0", edge(&format!("slot0/k{kb}"), LoadClass::scratch(8, kb), p)));
// slot MASK through the in-range register MASK
let mut p = vec![ins(Op::Sub, 1, 1), ins(Op::Sub, 2, 2), add_imm(1, 2, m)];
p.extend(scr(8, 1));
v.push(("slotmask".to_string(), "r1 = MASK: every RMW to the last slot", edge(&format!("slotmask/k{kb}"), LoadClass::scratch(8, kb), p)));
// slot MASK through the out-of-range register 0xffffffff
let mut p = vec![ins(Op::Sub, 1, 1), ins(Op::Sub, 2, 2), add_imm(2, 1, 1), ins(Op::Sub, 1, 2)];
p.extend(scr(8, 1));
v.push(("ones".to_string(), "r1 = 0xffffffff: masked to the last slot", edge(&format!("ones/k{kb}"), LoadClass::scratch(8, kb), p)));
// slot 0 through the out-of-range register MASK + 1
let mut p = vec![ins(Op::Sub, 1, 1), ins(Op::Sub, 2, 2), add_imm(1, 2, m.wrapping_add(1))];
p.extend(scr(8, 1));
v.push(("maskplus1".to_string(), "r1 = MASK + 1: masked to slot 0", edge(&format!("maskplus1/k{kb}"), LoadClass::scratch(8, kb), p)));
// 16 RMWs per iteration on slot 0: 128 dependent RMWs on one slot per lane per hash
let mut p = vec![ins(Op::Sub, 1, 1)];
p.extend(scr(16, 1));
v.push(("sixteen".to_string(), "16 RMWs per iteration on slot 0", edge(&format!("sixteen/k{kb}"), LoadClass::scratch(16, kb), p)));
// one slot per lane from the init words: lanes with equal slots would show any cross-lane aliasing
// (r5 is the slot register and is never a destination here)
let p: Vec<Instr> = [0u8, 1, 2, 3, 4, 6, 7, 0].iter().map(|&d| ins(Op::Scratch, d, 5)).collect();
v.push(("lanevar".to_string(), "r5 never written: one init-dependent slot per lane", edge(&format!("lanevar/k{kb}"), LoadClass::scratch(8, kb), p)));
// alternating slot 0 and slot MASK inside one iteration
let mut p = vec![ins(Op::Sub, 1, 1), ins(Op::Sub, 2, 2), add_imm(2, 1, m)];
for (i, &d) in [3u8, 4, 5, 6, 7, 0, 3, 4].iter().enumerate() {
// r1 and r2 hold the two slots and are never destinations
p.push(ins(Op::Scratch, d, if i % 2 == 0 { 1 } else { 2 }));
}
v.push(("twoslots".to_string(), "slot 0 and slot MASK alternating", edge(&format!("twoslots/k{kb}"), LoadClass::scratch(8, kb), p)));
v
}
/// The hand model: a second, minimal interpreter for the ops the edge programs use (sub, add, scratch), with its
/// own slot store keyed by (lane, slot). `mutate` swaps the rewrite's words to show the comparison has teeth.
fn hand_model(p: &Program, base: u32, mutate: bool) -> [u64; 32] {
let seed = &p.seed;
let m = p.class.scratch_slot_mask();
let mut r = [[0u32; LANES]; 8];
for lane in 0..LANES {
let nonce = base.wrapping_add(lane as u32);
for i in 0..8 {
let mut x = nonce ^ seed[i];
x = x.wrapping_add(0x9e3779b9u32.wrapping_mul(i as u32 + 1));
x = splitmix32(x);
r[i][lane] = x ^ seed[(i + 1) & 7];
}
}
let mut store: HashMap<(usize, u32), [u32; 3]> = HashMap::new();
for _ in 0..ITERATIONS {
let sel = r[0];
for ins in &p.instrs {
let (d, a) = (ins.dst as usize, ins.src as usize);
match ins.op {
Op::Sub => {
for lane in 0..LANES {
r[d][lane] = r[d][lane].wrapping_sub(r[a][lane]);
}
}
Op::Add => {
for lane in 0..LANES {
let c = if (sel[lane] >> ins.bit) & 1 != 0 { ins.imm2 } else { ins.imm };
r[d][lane] = r[d][lane].wrapping_add(r[a][lane]).wrapping_add(c);
}
}
Op::Scratch => {
for lane in 0..LANES {
let slot = r[a][lane] & m;
let w = *store.entry((lane, slot)).or_insert_with(|| {
let mut f = [0u32; 3];
for j in 0..3u32 {
// the fill, written out in full rather than through verify::scratch_fill
let n = base.wrapping_add(lane as u32);
f[j as usize] = splitmix32(
(n ^ seed[j as usize])
.wrapping_add(slot.wrapping_mul(0x9E37_79B1))
.wrapping_add((j + 1).wrapping_mul(0x85EB_CA77)),
);
}
f
});
let mut x = r[d][lane] ^ w[0];
x = x.rotate_left(FOLD_ROT).wrapping_mul(FOLD_MUL) ^ w[1];
x = x.rotate_left(FOLD_ROT).wrapping_mul(FOLD_MUL) ^ w[2];
r[d][lane] = x;
let out = if mutate {
[x.rotate_left(7) ^ w[2], x ^ w[1], x.wrapping_add(w[0])]
} else {
[x ^ w[1], x.rotate_left(7) ^ w[2], x.wrapping_add(w[0])]
};
store.insert((lane, slot), out);
}
}
other => panic!("the hand model does not implement {other:?}"),
}
}
}
let mut out = [0u64; 32];
for lane in 0..LANES {
let lo = r[0][lane] ^ r[1][lane].rotate_left(7) ^ r[2][lane].rotate_left(14) ^ r[3][lane].rotate_left(21);
let hi = r[4][lane] ^ r[5][lane].rotate_left(9) ^ r[6][lane].rotate_left(18) ^ r[7][lane].rotate_left(27);
out[lane] = ((hi as u64) << 32) | lo as u64;
}
out
}
/// The four unit bases of every edge vector: 0 and 32 (two consecutive units, the pair a one-warp persistent
/// launch runs on one arena), a unit straddling 2^31, and the unit that wraps past 2^32.
const EDGE_BASES: [u32; 4] = [0, 32, 0x7fff_fff0, 0xffff_ffe0];
#[test]
fn edge_programs_match_the_hand_model() {
let ds = DatasetSource::new("2026-10-03", DatasetMode::ClosedForm, 24);
let mut cases = 0;
for kb in [32u8, 128] {
for (name, what, p) in edge_programs(kb) {
let slots = p.class.scratch_slots_per_lane();
for base in EDGE_BASES {
let (res, ev) = interpret_warp_scratch(&p, &p.seed, base, &ds, true);
let hand = hand_model(&p, base, false);
assert_eq!(res.hashes, hand, "{name} k{kb} base {base:#x}: interpreter against the hand model ({what})");
assert_ne!(res.hashes, hand_model(&p, base, true), "{name} k{kb}: the comparison has teeth");
// the slots the trace saw are the ones the program was built to drive
let slot_set: std::collections::BTreeSet<u32> = ev.iter().map(|e| e.slot).collect();
let m = (slots - 1) as u32;
match name.as_str() {
"slot0" | "maskplus1" | "sixteen" => assert_eq!(slot_set.into_iter().collect::<Vec<_>>(), vec![0]),
"slotmask" | "ones" => assert_eq!(slot_set.into_iter().collect::<Vec<_>>(), vec![m]),
"twoslots" => assert_eq!(slot_set.into_iter().collect::<Vec<_>>(), vec![0, m]),
"lanevar" => {
for e in &ev {
assert!(e.slot <= m);
}
}
_ => unreachable!(),
}
// the chain depth on the driven slot: every RMW after the first per lane is a re-hit
let per_lane = p.scratch_ops_per_hash();
let hits = ev.iter().filter(|e| e.hit).count();
let expected_hits = match name.as_str() {
"twoslots" => (per_lane - 2) * LANES,
_ => (per_lane - 1) * LANES,
};
assert_eq!(hits, expected_hits, "{name} k{kb}: re-hits");
cases += 1;
}
}
}
assert_eq!(cases, 2 * 7 * 4);
}
// ---------------------------------------------------------------------------------------------------------------
// 4. The static scratch check over every emitted kernel of every scr pack (question 4)
// ---------------------------------------------------------------------------------------------------------------
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
pub enum Dialect {
Metal,
Cuda,
OpenCl,
}
/// The static scratch check: every scratch read-modify-write in an emitted kernel has the one masked form the
/// emitter writes, the arena is the lane's own `slots x 4` words, the tag is `salt + unit`, and nothing else
/// touches the scratch. Like the dataset mask check of `TESTS.md` section 5 and `tests/packs.rs`, a text check:
/// the guarantee is that the emitter has one template and it masks.
pub fn scratch_text_check(text: &str, dialect: Dialect, k: usize, slots: usize, kernels: usize) -> Result<(), String> {
assert!(kernels >= 1);
// every count below is per hash kernel; an OpenCL bound file carries igneum_hash and igneum_hash_bound
let k = k * kernels;
assert!(slots.is_power_of_two() && slots >= 1);
let mask = (slots - 1) as u32;
let wpl = slots * 4;
let (u, load, store, ptr) = match dialect {
Dialect::Metal => ("uint", "uint4 v_ = *(device const uint4*)(arena + s_ * 4u);", "*(device uint4*)(arena + s_ * 4u) = uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); }", "device uint* arena"),
Dialect::Cuda => ("uint32_t", "uint4 v_ = *(const uint4*)(arena + s_ * 4u);", "*(uint4*)(arena + s_ * 4u) = make_uint4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_); }", "uint32_t* arena"),
Dialect::OpenCl => ("uint", "uint4 v_ = vload4(s_, arena);", "vstore4(IGNEUM_U4(tag, x_ ^ w1_, rotl_imm(x_, 7u) ^ w2_, x_ + w0_), s_, arena); }", "__global uint* arena"),
};
let count = |needle: &str| text.matches(needle).count();
let mut errs = Vec::new();
let mut expect = |what: &str, got: usize, want: usize| {
if got != want {
errs.push(format!("{what}: {got}, expected {want}"));
}
};
// k slot computations, each masked with exactly the class's mask and immediately followed by the one load form
expect("slot definitions `{ u s_ = r`", count(&format!("{{ {u} s_ = r")), k);
expect("masked slot followed by the load", count(&format!(" & {mask}u; {load}")), k);
expect("stores of the tagged slot", count(store), k);
expect("tag compares", count("(v_.x == tag)"), k);
expect("fill calls (three per RMW)", count("scr_fill(gbase, lane, s_, "), 3 * k);
// the arena: one definition with the class's words per lane, and 2k uses (one load, one store per RMW)
expect("arena definition", count(&format!("{ptr} = scratch + ((size_t)warp_ * 32u + lane) * {wpl}u;")), kernels);
expect("arena mentions (definition + load + store per RMW)", count("arena"), kernels + 2 * k);
expect("tag definition `tag = salt + g_`", count(&format!("{u} tag = salt + g_;")), kernels);
expect("direct scratch indexing", count("scratch["), 0);
expect("scratch pointer arithmetic outside the arena definition", count("scratch +"), kernels);
// no other mask value on a slot: every `s_ = r` line carries the class mask and nothing else carries ` & Nu; uint4 v_`
let any_mask_load = count(&format!("u; {load}"));
expect("loads preceded by some mask (must all be the class mask)", any_mask_load, k);
if errs.is_empty() {
Ok(())
} else {
Err(errs.join("; "))
}
}
fn packs_rw_dir() -> PathBuf {
PathBuf::from(env!("CARGO_MANIFEST_DIR")).join("../proto-cuda/packs-readwidth")
}
fn scr_packs() -> Vec<String> {
let mut v: Vec<String> = std::fs::read_dir(packs_rw_dir())
.unwrap()
.map(|d| d.unwrap().file_name().to_string_lossy().to_string())
.filter(|n| n.starts_with("scr"))
.collect();
v.sort();
v
}
fn read_pack(pack: &str, file: &str) -> String {
let p = packs_rw_dir().join(pack).join(file);
std::fs::read_to_string(&p).unwrap_or_else(|e| panic!("read {}: {e}", p.display()))
}
/// Every scr pack regenerates from its program.json (seed bytes, class, day bytes, size) to the same six kernel
/// texts, byte for byte, and every one of those texts passes the static scratch check for the class's k and slot
/// count; the check fails on four deliberate breaks of a copy of the Metal text (mask dropped, mask changed, arena
/// stride changed, a stray scratch access) and on the OpenCL and CUDA twins of the first.
#[test]
fn scr_packs_regenerate_and_pass_the_static_scratch_check() {
let packs = scr_packs();
assert!(packs.len() >= 6, "the scr packs: {packs:?}");
let mut checked = 0;
let mut sample_metal = String::new();
let mut sample_cl = String::new();
let mut sample_cu = String::new();
let mut sample_k = 0;
let mut sample_slots = 0;
for pack in &packs {
let j: Value = serde_json::from_str(&read_pack(pack, "program.json")).unwrap();
let name = j["load_class"].as_str().unwrap();
let c = class(name);
assert_eq!(&format!("{name}"), pack, "pack directory named after its class");
let seed = j["seed"].as_str().unwrap();
let seed_bytes = igneum_pow::bind::unhex(j["seed_bytes"].as_str().unwrap()).unwrap();
let day_bytes = igneum_pow::bind::unhex(j["dataset"]["day_bytes"].as_str().unwrap()).unwrap();
let log2 = j["dataset"]["log2_words"].as_u64().unwrap() as u32;
assert_eq!(j["dataset_mode"].as_str().unwrap(), "memory-hard");
let program = generate_from_seed_bytes_class(seed, &seed_bytes, c);
assert_eq!(program.class, c);
assert_eq!(program.program_id(), u64::from_str_radix(j["program_id"].as_str().unwrap().trim_start_matches("0x"), 16).unwrap());
let mut dataset = DatasetSource::from_key(seed_words_from_bytes(&day_bytes), DatasetMode::MemoryHard, log2);
dataset.key_bytes = day_bytes;
let e = Epoch { program, dataset };
let p = &e.program;
let mp = e.dataset.memhard().map(|m| &m.params);
let k = c.scratch_slots();
let slots = c.scratch_slots_per_lane();
assert_eq!(p.scratch_ops_per_hash(), k * ITERATIONS);
for (file, text, dialect, kernels) in [
("program.metal", metal_program(p, log2, LoadSource::Stored), Dialect::Metal, 1),
("program_bound.metal", metal_program_bound(p, log2), Dialect::Metal, 1),
("kernel.cu", cuda_kernel(p, mp), Dialect::Cuda, 1),
("kernel_bound.cu", cuda_kernel_bound(p, mp), Dialect::Cuda, 1),
("kernel.cl", opencl_kernel(p, mp), Dialect::OpenCl, 1),
// the OpenCL bound file carries igneum_hash and igneum_hash_bound
("kernel_bound.cl", opencl_kernel_bound(p, mp), Dialect::OpenCl, 2),
] {
let on_disk = read_pack(pack, file);
assert_eq!(on_disk, text, "{pack}/{file}: the pack is the emitter's text");
// scr0 is the persistent control: an arena and a tag, no read-modify-write; the check holds with k = 0
scratch_text_check(&on_disk, dialect, k, slots, kernels).unwrap_or_else(|e| panic!("{pack}/{file}: {e}"));
checked += 1;
}
// the vectors of the pack are the CPU's
let v: Value = serde_json::from_str(&read_pack(pack, "vectors.json")).unwrap();
for w in v["warps"].as_array().unwrap() {
let base = w["base_nonce"].as_u64().unwrap() as u32;
let got = e.hash_warp(base);
for (lane, x) in w["expected"].as_array().unwrap().iter().enumerate() {
let want = u64::from_str_radix(x.as_str().unwrap().trim_start_matches("0x"), 16).unwrap();
assert_eq!(got[lane], want, "{pack}: base {base} lane {lane}");
}
}
if k == 4 && slots == 64 {
sample_metal = read_pack(pack, "program.metal");
sample_cl = read_pack(pack, "kernel.cl");
sample_cu = read_pack(pack, "kernel.cu");
sample_k = k;
sample_slots = slots;
}
}
assert_eq!(checked, packs.len() * 6);
println!("static scratch check: {checked} kernels over {} scr packs", packs.len());
// The deliberate breaks (the watcher rule of CLAUDE.md: a check is trusted once it fails on a known-broken
// case). Each must be caught; the message names what.
assert!(sample_k == 4 && sample_slots == 64, "scr4k32 is in the pack set");
let mask = format!(" & {}u; uint4 v_", sample_slots - 1);
let broken_mask = sample_metal.replacen(&mask, "; uint4 v_", 1);
assert_ne!(broken_mask, sample_metal);
let e = scratch_text_check(&broken_mask, Dialect::Metal, 4, 64, 1).unwrap_err();
assert!(e.contains("masked slot followed by the load: 3, expected 4"), "{e}");
println!("break 1 (one mask dropped, Metal): {e}");
let wrong_mask = sample_metal.replace(" & 63u;", " & 127u;");
let e = scratch_text_check(&wrong_mask, Dialect::Metal, 4, 64, 1).unwrap_err();
assert!(e.contains("masked slot followed by the load: 0, expected 4"), "{e}");
println!("break 2 (mask 63 -> 127 on every RMW, Metal): {e}");
let wrong_stride = sample_metal.replace("* 256u;", "* 128u;");
let e = scratch_text_check(&wrong_stride, Dialect::Metal, 4, 64, 1).unwrap_err();
assert!(e.contains("arena definition: 0, expected 1"), "{e}");
println!("break 3 (arena stride 256 -> 128 words, Metal): {e}");
let stray = format!("{sample_metal}\n// stray\n// arena[0] = 0u; scratch[1] = 1u;\n");
let e = scratch_text_check(&stray, Dialect::Metal, 4, 64, 1).unwrap_err();
assert!(e.contains("arena mentions") && e.contains("direct scratch indexing: 1, expected 0"), "{e}");
println!("break 4 (a stray arena and scratch access, Metal): {e}");
let e = scratch_text_check(&sample_cl.replacen(" & 63u; uint4 v_ = vload4", "; uint4 v_ = vload4", 1), Dialect::OpenCl, 4, 64, 1).unwrap_err();
assert!(e.contains("masked slot followed by the load: 3, expected 4"), "{e}");
println!("break 5 (one mask dropped, OpenCL): {e}");
let e = scratch_text_check(&sample_cu.replacen(" & 63u; uint4 v_ = *(const uint4*)", "; uint4 v_ = *(const uint4*)", 1), Dialect::Cuda, 4, 64, 1).unwrap_err();
assert!(e.contains("masked slot followed by the load: 3, expected 4"), "{e}");
println!("break 6 (one mask dropped, CUDA): {e}");
// and the unbroken texts pass under the same calls
scratch_text_check(&sample_metal, Dialect::Metal, 4, 64, 1).unwrap();
scratch_text_check(&sample_cl, Dialect::OpenCl, 4, 64, 1).unwrap();
scratch_text_check(&sample_cu, Dialect::Cuda, 4, 64, 1).unwrap();
// a wrong slot count, RMW count or kernel count against a right text fails too (the check is tied to the class)
assert!(scratch_text_check(&sample_metal, Dialect::Metal, 4, 256, 1).is_err());
assert!(scratch_text_check(&sample_metal, Dialect::Metal, 3, 64, 1).is_err());
assert!(scratch_text_check(&sample_metal, Dialect::Metal, 4, 64, 2).is_err());
}
// ---------------------------------------------------------------------------------------------------------------
// 5. The fuzz: 200 generated scratch programs, contract on every instruction, 4 units each across the 32-bit
// range including the wrap; with IGNEUM_SCRATCH_PACKS_OUT the packs for the Metal runs (question 3, 4)
// ---------------------------------------------------------------------------------------------------------------
/// Write a pack whose vectors.json carries `bases` (any number of units) instead of the three standard bases.
fn write_pack_with_bases(dir: &PathBuf, e: &Epoch, day: &str, bases: &[u32], source: &str) -> Vec<[u64; 32]> {
let mut pack = export_pack(e, day, source);
let outs: Vec<[u64; 32]> = bases.iter().map(|&b| e.hash_warp(b)).collect();
let vj = vectors_json(&e.program, day, e.dataset.log2_words, bases, &outs, &pack.vectors, e.dataset.mask, source, true);
for f in pack.files.iter_mut() {
if f.0 == "vectors.json" {
f.1 = vj.clone();
}
}
pack.write_to(dir).unwrap();
outs
}
fn contract(p: &Program) {
assert_eq!(p.instrs.len(), INSTR_COUNT);
assert_eq!(p.instrs.iter().filter(|i| i.op == Op::Load).count() + p.instrs.iter().filter(|i| i.op == Op::Scratch).count(), 16);
assert_eq!(p.instrs.iter().filter(|i| i.op == Op::Scratch).count(), p.class.scratch_slots());
assert!(p.instrs[0].op != Op::Load && p.instrs[0].op != Op::Scratch, "instruction 0 is never a memory op");
for (k, i) in p.instrs.iter().enumerate() {
assert!(i.src != i.dst, "#{k}: src == dst");
assert!((1..=31).contains(&i.rot), "#{k}: rot {}", i.rot);
assert!([1u8, 2, 4, 8, 16].contains(&i.mask), "#{k}: mask {}", i.mask);
assert!(i.dst < 8 && i.src < 8 && i.src2 < 8);
assert_eq!(i.width, 1, "#{k}: a scratch class reads one-word loads");
}
assert!(igneum_pow::accept::check(p).is_ok(), "an accepted program");
}
#[test]
fn fuzz_scr_programs_cpu() {
let n: usize = std::env::var("IGNEUM_SCRATCH_FUZZ").ok().and_then(|s| s.parse().ok()).unwrap_or(200);
let out = std::env::var("IGNEUM_SCRATCH_PACKS_OUT").ok().map(PathBuf::from);
let mut rng = SplitMix64::new(0x6967_6e65_756d_2d73); // "igneum-s"
let day = "2026-10-03";
let closed = DatasetSource::new(day, DatasetMode::ClosedForm, 28);
// memory-hard sources per size, built once each (the cache fill is 0.2 s); only when packs are written
let mut mh: HashMap<u32, DatasetSource> = HashMap::new();
let mut manifest = String::from("pack\tclass\tlog2\tprogram_id\tscratch_ops_per_hash\tbases\n");
let mut per_class: HashMap<String, usize> = HashMap::new();
let mut units = 0usize;
let mut wraps = 0usize;
if let Some(dir) = &out {
std::fs::create_dir_all(dir).unwrap();
// the edge packs first: 64 MiB datasets (no dataset load in them), the four edge bases
for kb in [32u8, 128] {
for (name, _what, p) in edge_programs(kb) {
let log2 = 24;
let ds = mh.remove(&log2).unwrap_or_else(|| DatasetSource::new(day, DatasetMode::MemoryHard, log2));
let e = Epoch { program: p, dataset: ds };
let pack_name = format!("edge-{name}-k{kb}");
write_pack_with_bases(&dir.join(&pack_name), &e, day, &EDGE_BASES, "igneum-pow tests/scratch.rs edge");
manifest.push_str(&format!(
"{pack_name}\t{}\t{log2}\t{:016x}\t{}\t{}\n",
e.program.class.name(),
e.program.program_id(),
e.program.scratch_ops_per_hash(),
EDGE_BASES.iter().map(|b| format!("{b}")).collect::<Vec<_>>().join(",")
));
mh.insert(log2, e.dataset);
}
}
}
for i in 0..n {
let name = CLASSES[rng.below(CLASSES.len() as u64) as usize];
let c = class(name);
let seed = format!("igneum-scratch-fuzz/{i}");
let p = generate_class(&seed, c);
contract(&p);
*per_class.entry(name.to_string()).or_insert(0) += 1;
// four bases: one inside a 256-nonce batch (in-batch check on the GPU), one straddling 2^31, one in
// the last 256 nonces (the unit wraps past 2^32 or ends on it), one uniform
let b0 = (rng.below(8) as u32) * 32;
let b1 = 0x8000_0000u32.wrapping_sub(256).wrapping_add((rng.below(16) as u32) * 32);
let b2 = 0xffff_ff00u32.wrapping_add((rng.below(8) as u32) * 32);
let b3 = (rng.next() as u32) & !31;
let bases = [b0, b1, b2, b3];
// an aligned unit never straddles 2^32 (spec 1.9); the top unit ends on 0xffffffff and the persistent
// kernel's unit sequence wraps inside a launch, which the Metal run checks with packbench --batch-base
wraps += bases.iter().filter(|&&b| b >= 0xffff_ff00).count();
// the CPU: the interpreter is deterministic and every scratch event is inside the lane's slots
for &b in &bases {
let (r1, ev) = interpret_warp_scratch(&p, &p.seed, b, &closed, true);
let r2 = interpret_warp_scratch(&p, &p.seed, b, &closed, false).0;
assert_eq!(r1.hashes, r2.hashes);
assert_eq!(ev.len(), p.scratch_ops_per_hash() * LANES);
assert!(ev.iter().all(|e: &ScratchEvent| e.slot < c.scratch_slots_per_lane() as u32));
units += 1;
}
if let Some(dir) = &out {
let log2 = [24u32, 26, 28][rng.below(3) as usize];
let ds = mh.remove(&log2).unwrap_or_else(|| DatasetSource::new(day, DatasetMode::MemoryHard, log2));
let e = Epoch { program: p, dataset: ds };
let pack_name = format!("fuzz-{i:03}-{name}-l{log2}");
write_pack_with_bases(&dir.join(&pack_name), &e, day, &bases, "igneum-pow tests/scratch.rs fuzz");
manifest.push_str(&format!(
"{pack_name}\t{name}\t{log2}\t{:016x}\t{}\t{}\n",
e.program.program_id(),
e.program.scratch_ops_per_hash(),
bases.iter().map(|b| format!("{b}")).collect::<Vec<_>>().join(",")
));
mh.insert(log2, e.dataset);
} else {
let _ = rng.below(3);
}
}
let mut classes: Vec<_> = per_class.iter().collect();
classes.sort();
println!("fuzz: {n} programs, {units} units on the CPU, {wraps} units in the top 256 nonces, classes {classes:?}");
assert_eq!(units, 4 * n);
assert_eq!(wraps, n, "every program has a unit in the top 256 nonces");
if let Some(dir) = &out {
std::fs::write(dir.join("manifest.tsv"), manifest).unwrap();
println!("packs written to {}", dir.display());
}
}
/// The fold and rewrite, restated: a slot after `d` dependent RMWs holds 96 bits that are a function of the fill
/// (3 words, a pure function of nonce, slot and seed) and the `d` fold values; a chip that keeps the `d` fold
/// values (32 bits each) instead of the 96-bit slot recomputes the slot in `d` rewrites. This test pins the
/// arithmetic the analysis uses (question 2): the replay from the fold values reproduces the slot.
#[test]
fn slot_is_replayable_from_its_fold_values() {
let seed = seed_words_from_bytes(b"igneum-genesis");
let (base, lane, slot) = (0x1234_5600u32, 5u32, 17u32);
let fill = [scratch_fill(&seed, base, lane, slot, 0), scratch_fill(&seed, base, lane, slot, 1), scratch_fill(&seed, base, lane, slot, 2)];
let mut rng = SplitMix64::new(99);
let dsts: Vec<u32> = (0..64).map(|_| rng.next() as u32).collect();
// the honest sequence: read, fold, rewrite, 64 times
let mut w = fill;
let mut xs = Vec::new();
for &d in &dsts {
let x = fold_words(d, &w);
xs.push(x);
w = scratch_rewrite(x, &w);
}
// the replay: from the fill and the stored fold values alone
let mut w2 = fill;
for &x in &xs {
w2 = scratch_rewrite(x, &w2);
}
assert_eq!(w, w2);
// and nothing shorter: the fold value at step d depends on the slot content at step d, which depends on
// every earlier fold value (drop one and the chain diverges)
let mut w3 = fill;
for (i, &x) in xs.iter().enumerate() {
if i != 10 {
w3 = scratch_rewrite(x, &w3);
}
}
assert_ne!(w, w3);
}

View file

@ -24,6 +24,12 @@ simulator takes the same file: `igneum-harness-sim --override-params-file infra/
Proof script: `node infra/fast-time/simnet.mjs` (3 nodes on ports 29500 and up, `igneum-devnet-950`, data
`/tmp/igneum-fast-time`; records the first lock and the first program swap; see "Measured" below).
Class switch gate (Counter ASIC 2.0 rollout G4): `node infra/fast-time/class-v3.mjs` (3 nodes on ports 29600 and up,
`igneum-devnet-960`, data `/tmp/igneum-fast-time-v3`, CPU difficulty `genesis_bits` 0x1f010000, real CPU mining on
every node, `program_class_v3_activation_daa` a few epochs ahead; reports blocks on both sides of the boundary, the
program id and class before and after, rejected blocks, forks, and every node's switch line; see
`docs/plans/counter-asic-2-node.md`).
Miners: `igneum-miner` follows the epoch length and lead its node reports in every template (`pow_epoch`), no flag.
The dataset day is not in the template, so a real-hash miner on a fast-time network takes `IGNEUM_POW_DAY_MS=1440000`
in its environment (the node reads it from the file). The environment variables `IGNEUM_POW_EPOCH_BLOCKS`,
@ -51,6 +57,11 @@ Time parameters, divided by 60 (devnet value, 60x value):
| `difficulty_v2_activation_daa` | never (`18446744073709551615`) | never | a height, not a clock: difficulty rule v2 (4 Oct 2026, `docs/analysis/difficulty-2026-10-04-oscillation.md`) switches on at this DAA score; a test network sets it in its merged file (`sim/difficulty/testnet_v2.py` uses 900) |
| `proving_v0_activation_daa` | never | never | a height: proving v0 payouts (spec 7.7) switch on at this DAA score; `tools/proving-v0/run.mjs` sets 60 in its merged file |
| `finality_v3_activation_daa` | never | never | a height, not a clock: finality rule v3 (the frozen weight table of ledger F21 and the certificate fold of F22, 4 Oct 2026 evening) applies to checkpoints at or above this DAA score; `tools/finality-attacks/v3.mjs` sets 0 in its merged file |
| `pow_genesis_dataset_log2` | 28 | 28 | a size, not a clock: the dataset at genesis (2^28 words, 1 GiB); the cache growth rule (ca2-mixer) doubles it with the days since the genesis day |
| `proving_v1_activation_daa` | never | never | a height: proving v1 (the fork's proving-v1 branch, 5 Oct 2026) switches on at this DAA score |
| `proving_v1_segment_blocks`, `proving_v1_aggregator_share_bps` | 8, 1,000 | 8, 1,000 | a block count and a share |
| `proving_v1_unproven_daa` | 600 | 10 | a DAA clock (the unproven allowance), divided by 60; the proving agent confirms the value |
| `program_class_v3_activation_daa` | never | never | a height, not a clock: the lottery hash draws programs from class v3 (Counter ASIC 2.0, 5 Oct 2026) from the first EPOCH whose start is at or above this DAA score (rounded up to an epoch boundary: at 60 DAA per epoch, 150 means epoch 3 at DAA 180); `infra/fast-time/class-v3.mjs` sets it a few epochs ahead in its merged file |
Unchanged, and why:

View file

@ -0,0 +1,282 @@
#!/usr/bin/env node
// Counter ASIC 2.0 rollout gate G4 (docs/plans/counter-asic-2-rollout.md section 7): the fast-time 3-node network
// mining across a program class v3 activation. A private network on 127.0.0.1 ports 29600 and up, data under
// /tmp/igneum-fast-time-v3, network id igneum-devnet-960, every node on infra/fast-time/override-60x.json merged with
// a CPU genesis difficulty (genesis_bits 0x1f010000, 2^16 hashes per block, as sim/difficulty/testnet_v2.py) and
// `program_class_v3_activation_daa` a few epochs ahead (default 150: inside epoch 2 at 60 DAA per epoch, so the
// switch rounds UP to epoch 3 at DAA 180, which is the boundary rule under test). One real CPU miner per node
// (igneum-miner --engine igneum-pow, 1 thread) follows its node's templates, so every block of the run is a real
// lottery-hash solution and every node verifies every block of the other two under the class of its epoch.
//
// Reports, from the nodes' RPC and the logs: blocks on each side of the boundary, the program class and id of every
// epoch (before and after), rejected blocks (the miners' submit answers and the nodes' "PoW rejected" lines), forks
// (every node's sink, selected tip and block count at the end), and every node's switch line. Exit 0 when every
// check passes. Never touches the live devnet (26610/26611, 26640/26641, 28640) or the simnet.mjs ports.
//
// node infra/fast-time/class-v3.mjs [--secs 420] [--activation 150] [--epochs-after 2] [--threads 1]
// [--metal <igneum-bench>] [--genesis-bits 0x1f010000]
// IGNEUMD and IGNEUM_MINER name the binaries (default: the ca2 fork worktree's target-ca2/release).
// --metal: node 0's miner is a real Metal GPU worker (igneum-miner --worker <igneum-bench> --prepare-packs <dir>
// --exit-on-seed-change, the app's own shape), so the Metal worker's class v3 path (the prepare line with the pack
// directory and the class and era tokens, the pack-built program and day) mines across the boundary (gate G4b);
// the report then adds the miner's PREPARE lines, the worker's prepared lines, any need / mismatch / refusal line,
// the accepted blocks on each side and the swap time. --genesis-bits raises the CPU difficulty for a GPU miner
// (0x1e010000 = 2^24 hashes per block: under a second on an M5 Max; the CPU miners on nodes 1 and 2 then verify
// and rarely find).
import { spawn } from 'node:child_process';
import { mkdirSync, rmSync, writeFileSync, readFileSync, openSync, existsSync } from 'node:fs';
import { connectRpc } from '../../tools/finality-attacks/lib/rpc.mjs';
import { devAddress } from '../../tools/harness/lib/address.mjs';
const ROOT = new URL('../../', import.meta.url).pathname;
const FILE = `${ROOT}infra/fast-time/override-60x.json`;
const BIN = process.env.IGNEUM_CA2_BIN || `${ROOT}vendor/igneum-node-ca2/target-ca2/release`;
const IGNEUMD = process.env.IGNEUMD || `${BIN}/igneumd`;
const CPU_MINER = process.env.IGNEUM_MINER || `${BIN}/igneum-miner`;
const TMP = '/tmp/igneum-fast-time-v3';
const BASE = 29600, SUFFIX = 960;
const args = process.argv.slice(2);
const flag = (name, dflt) => { const i = args.indexOf(`--${name}`); return i >= 0 ? +args[i + 1] : dflt; };
const sflag = (name) => { const i = args.indexOf(`--${name}`); return i >= 0 ? args[i + 1] : null; };
const GENESIS_BITS = flag('genesis-bits', 0x1f010000);
const METAL = sflag('metal');
const SECS = flag('secs', 420);
const ACTIVATION = flag('activation', 150);
const EPOCHS_AFTER = flag('epochs-after', 2);
const THREADS = flag('threads', 1);
const started = [];
const log = (...a) => console.log(new Date().toISOString().slice(11, 23), ...a);
const sleep = (ms) => new Promise(r => setTimeout(r, ms));
for (const b of [IGNEUMD, CPU_MINER, ...(METAL ? [METAL] : [])]) if (!existsSync(b)) { console.error(`missing ${b}`); process.exit(2); }
rmSync(TMP, { recursive: true, force: true }); mkdirSync(TMP, { recursive: true });
// The file is merged as TEXT, never through JSON.parse: a `never` height is 18446744073709551615, which JavaScript
// rounds to 1.8446744073709552e+19 and the node refuses ("expected u64"; the first gate run, 5 October 2026 21:30Z).
// The merged fields are appended after the file's last field; a field already in the file is removed first.
const baseText = readFileSync(FILE, 'utf8');
const field = (name) => { const m = new RegExp(`"${name}":\\s*([0-9]+)`).exec(baseText); return m ? +m[1] : undefined; };
const EPOCH = field('pow_epoch_blocks');
const DAY_MS = field('pow_day_ms');
const FIRST_V3_EPOCH = Math.ceil(ACTIVATION / EPOCH);
const BOUNDARY = FIRST_V3_EPOCH * EPOCH;
const override = `${TMP}/override.json`;
export function mergeOverrideText(text, fields) {
let out = text;
for (const k of Object.keys(fields)) out = out.replace(new RegExp(`\\s*"${k}":\\s*[^,}\\n]+,?`), '');
const extra = Object.entries(fields).map(([k, v]) => `"${k}": ${typeof v === 'string' ? JSON.stringify(v) : v}`).join(', ');
return out.replace(/,?\s*}\s*$/, `,\n ${extra}\n}\n`);
}
writeFileSync(override, mergeOverrideText(baseText, { genesis_bits: GENESIS_BITS, skip_proof_of_work: false, program_class_v3_activation_daa: ACTIVATION }));
log(`activation ${ACTIVATION} at ${EPOCH} DAA per epoch: the first v3 epoch is ${FIRST_V3_EPOCH} (DAA ${BOUNDARY}); run ${SECS} s, CPU genesis bits 0x${GENESIS_BITS.toString(16)}`);
class Node {
constructor(i, connect = []) {
this.i = i; this.grpcPort = BASE + i * 10; this.p2pPort = BASE + i * 10 + 1; this.jsonPort = BASE + i * 10 + 2;
this.connect = connect; this.dir = `${TMP}/n${i}`; this.logFile = `${this.dir}/node.log`;
}
get grpc() { return `grpc://127.0.0.1:${this.grpcPort}`; }
async start() {
mkdirSync(this.dir, { recursive: true });
const a = ['--devnet', `--devnet-suffix=${SUFFIX}`, '--nodnsseed', '--disable-upnp', '--nologfiles', '--enable-unsynced-mining', '--utxoindex',
`--appdir=${this.dir}`, `--rpclisten=127.0.0.1:${this.grpcPort}`, `--rpclisten-json=127.0.0.1:${this.jsonPort}`,
`--listen=127.0.0.1:${this.p2pPort}`, `--override-params-file=${override}`, '--loglevel=info', '--yes'];
if (this.connect.length) a.push(`--connect=${this.connect.join(',')}`); else a.push('--outpeers=0');
const out = openSync(this.logFile, 'a');
this.proc = spawn(IGNEUMD, a, { stdio: ['ignore', out, out] });
started.push(this.proc);
await sleep(1200);
this.rpc = await connectRpc(`ws://127.0.0.1:${this.jsonPort}`);
log(`n${this.i} up pid ${this.proc.pid} json ${this.jsonPort} p2p ${this.p2pPort}`);
return this;
}
grepLog(re) { try { return readFileSync(this.logFile, 'utf8').split('\n').filter(l => re.test(l)); } catch { return []; } }
}
function miner(bin, argv, name, env = {}) {
const out = openSync(`${TMP}/${name}.log`, 'a');
const p = spawn(bin, argv, { stdio: ['ignore', out, out], env: { ...process.env, ...env } });
started.push(p);
return p;
}
async function stopAll() {
for (const p of started.reverse()) { try { p.kill('SIGINT'); } catch { } }
await sleep(1500);
for (const p of started) { try { p.kill('SIGKILL'); } catch { } }
}
process.on('SIGINT', async () => { await stopAll(); process.exit(130); });
process.on('unhandledRejection', async (e) => { log(`FAILED: ${e?.stack || e}`); await stopAll(); process.exit(3); }); // a thrown start leaves no node behind
const minerName = (i) => (METAL && i === 0) ? 'metal0' : `cpu${i}`;
const minerLog = (i) => { try { return readFileSync(`${TMP}/${minerName(i)}.log`, 'utf8').split('\n'); } catch { return []; } };
const t0 = Date.now();
const since = () => ((Date.now() - t0) / 1000).toFixed(1);
const n0 = await new Node(0).start();
const n1 = await new Node(1, [`127.0.0.1:${n0.p2pPort}`]).start();
const n2 = await new Node(2, [`127.0.0.1:${n0.p2pPort}`]).start(); // --connect takes one address
const nodes = [n0, n1, n2];
for (const n of nodes) {
const line = n.grepLog(/Program class v3 from the override file/)[0];
log(`n${n.i} switch line: ${line ? line.replace(/^.*?(Program class v3)/, '$1') : '(none)'}`);
}
log(`n0 PoW schedule: ${n0.grepLog(/PoW schedule from the override file/).map(l => l.replace(/^.*?(PoW schedule)/, '$1')).join(' | ') || '(no line)'}`);
log(`n0 digest: ${n0.grepLog(/Consensus params digest/).map(l => l.replace(/^.*?digest: /, '').slice(0, 16)).join(' ')}`);
// one real CPU miner per node, 1 thread, 2^16 expected hashes per block at genesis; with --metal, node 0's miner drives
// the Metal worker the way the app does (the pack for every prepared pair under packs/prepare, exit 42/44 on a refusal)
const PACKS = `${TMP}/packs/prepare`;
mkdirSync(PACKS, { recursive: true });
nodes.forEach((n, i) => {
if (METAL && i === 0) {
miner(CPU_MINER, ['mine', n.grpc, '1', String(SECS), 'metal0', '--engine', 'igneum-pow', '--worker', METAL, '--prepare-packs', PACKS, '--exit-on-seed-change', '--payout-label', 'metal0', '--status-secs', '30', '--no-vote'], 'metal0', { IGNEUM_POW_DAY_MS: String(DAY_MS) });
log(`n0 miner: Metal worker ${METAL}, packs under ${PACKS}`);
} else {
miner(CPU_MINER, ['mine', n.grpc, String(THREADS), String(SECS), `cpu${i}`, '--engine', 'igneum-pow', '--payout-label', `cpu${i}`, '--status-secs', '30', '--no-vote'], `cpu${i}`, { IGNEUM_POW_DAY_MS: String(DAY_MS) });
}
});
const pay = devAddress('fast-time-v3');
const epochs = new Map(); // epoch index -> { class, firstSeenDaa, at }
let firstV3 = null, lastEpoch = -1, lastReport = 0, lastDaa = 0, switchEnd = null;
const samples = [];
while (Date.now() - t0 < SECS * 1000) {
await sleep(1000);
let daa = null, epoch = null, cls = null, nextCls = null, era = null, eraSeed = null, act = null;
try {
const t = await n0.rpc.call('getBlockTemplate', { payAddress: pay, extraData: [] });
const pe = t.powEpoch || t.pow_epoch || {};
daa = pe.virtualDaaScore ?? t.block?.header?.daaScore; epoch = pe.epochIndex; cls = pe.programClass; nextCls = pe.nextProgramClass;
era = pe.eraIndex; eraSeed = pe.eraSeed; act = pe.programClassV3ActivationDaa;
} catch (e) { log(`template: ${e.message}`); }
if (epoch != null && epoch !== lastEpoch) {
epochs.set(epoch, { class: cls, firstSeenDaa: daa, at: +since() });
log(`epoch ${lastEpoch} -> ${epoch} at daa ${daa}, ${since()} s: template class ${cls} (generator), next epoch class ${nextCls}, era ${era} seed ${String(eraSeed).slice(0, 16)}, activation ${act}`);
if (firstV3 == null && cls === 3) { firstV3 = { epoch, daa, at: +since() }; log(`CLASS SWITCH: the template is class v3 from epoch ${epoch} (daa ${daa}) at ${since()} s wall`); }
lastEpoch = epoch;
}
lastDaa = daa ?? lastDaa;
if (Date.now() - lastReport > 15000) {
lastReport = Date.now();
const counts = await Promise.all(nodes.map(async n => { try { const d = await n.rpc.call('getBlockDagInfo'); return `${d.blockCount}/${String(d.sink).slice(0, 8)}`; } catch { return '?'; } }));
log(`t=${since()} s daa ${daa} epoch ${epoch} class ${cls} blocks/sink per node ${counts.join(' ')}`);
samples.push({ t: +since(), daa, epoch, class: cls, nodes: counts });
}
if (firstV3 != null && daa != null && daa >= BOUNDARY + EPOCHS_AFTER * EPOCH) { switchEnd = +since(); break; }
}
await sleep(3000); // let the last blocks relay before the end-of-run reads
// ---- the end-of-run reads -----------------------------------------------------------------------------------------
const dag = await Promise.all(nodes.map(async n => { try { return await n.rpc.call('getBlockDagInfo'); } catch (e) { return { error: e.message }; } }));
const genesis = dag[0].pruningPointHash;
// every block the first node holds, by DAA score, from genesis
async function allBlocks(n) {
const out = []; let low = genesis; const seen = new Set();
for (let round = 0; round < 500; round++) {
const r = await n.rpc.call('getBlocks', { lowHash: low, includeBlocks: true, includeTransactions: false });
const blocks = r.blocks || [];
let added = 0;
for (const b of blocks) { const h = b.verboseData?.hash || b.header?.hash; if (seen.has(h)) continue; seen.add(h); out.push({ hash: h, daa: +b.header.daaScore, chain: !!b.verboseData?.isChainBlock, blue: +(b.verboseData?.blueScore ?? 0) }); added++; }
if (!blocks.length || added === 0) break;
low = (r.blockHashes || []).at(-1) || blocks.at(-1).verboseData?.hash; if (!low) break;
}
return out;
}
let blocks = [];
try { blocks = await allBlocks(n0); } catch (e) { log(`getBlocks: ${e.message}`); }
const before = blocks.filter(b => b.daa < BOUNDARY), after = blocks.filter(b => b.daa >= BOUNDARY);
const chainBefore = before.filter(b => b.chain).length, chainAfter = after.filter(b => b.chain).length;
// program ids and classes per epoch from the miners' "program and 256 MiB cache ready" lines
const programs = new Map(); // epoch seed -> { class, id, miners: Set }
for (const i of [0, 1, 2]) for (const l of minerLog(i)) {
const m = /epoch seed ([0-9a-f]{64}) day (\d+) \(daa (\d+)\): program and 256 MiB cache ready in ([\d.]+) ms; class (v\d) program id ([0-9a-f]{16})/.exec(l);
if (!m) continue;
const k = m[1]; const e = programs.get(k) || { seed: k.slice(0, 16), epoch: Math.floor(+m[3] / EPOCH), class: m[5], id: m[6], miners: new Set(), ready_ms: [] };
if (e.id !== m[6] || e.class !== m[5]) e.disagree = true;
e.miners.add(i); e.ready_ms.push(+m[4]); programs.set(k, e);
}
// ready_ms: the program generation plus the day's cache and dataset build on one CPU core (the first epoch of a day
// pays the cache; under the mixer x4 construction the first v3 epoch's build is the number the rollout asks for)
const programRows = [...programs.values()].sort((a, b) => a.epoch - b.epoch).map(p => ({ epoch: p.epoch, class: p.class, program_id: p.id, seed: p.seed, miners: p.miners.size, disagree: !!p.disagree, ready_ms: p.ready_ms }));
const v2Ids = programRows.filter(p => p.class === 'v2').map(p => p.program_id), v3Ids = programRows.filter(p => p.class === 'v3').map(p => p.program_id);
// rejections: the miners' submit answers and the nodes' PoW lines
const accepted = [0, 1, 2].map(i => minerLog(i).filter(l => /ACCEPTED block/.test(l)).length);
const rejectedMiner = [0, 1, 2].map(i => minerLog(i).filter(l => /rejected nonce=|submit error/.test(l)));
const rejectedNode = nodes.map(n => n.grepLog(/PoW rejected|Rejected block|rejected block/i));
const switchLines = nodes.map(n => n.grepLog(/Program class v3 from the override file/).map(l => l.replace(/^.*?(Program class v3)/, '$1'))[0] || null);
const sinks = dag.map(d => String(d.sink || '?').slice(0, 16));
const counts = dag.map(d => d.blockCount ?? '?');
const tips = dag.map(d => (d.tipHashes || []).length);
// the Metal miner's protocol lines (gate G4b): PREPARE sent (the miner), prepared / need / error (the worker), the swap
let metal = null;
if (METAL) {
const L = minerLog(0);
const t = (l) => { const m = /^(\d+\.\d+) /.exec(l); return m ? +m[1] * 1000 : null; };
const switchAt = firstV3 ? t0 + firstV3.at * 1000 : null;
const prepares = L.filter(l => /PREPARE sent for epoch seed/.test(l)).map(l => l.replace(/^\S+ /, ''));
const prepared = L.filter(l => /worker: prepared /.test(l)).map(l => l.replace(/^\S+ /, ''));
const preparedV3 = prepared.filter(l => / class v3 /.test(l));
const need = L.filter(l => /worker: need |^\S+ worker error:.*need /.test(l) || /\bneed [0-9a-f]{64}/.test(l));
// a PACK OUT OF DATE before the first class v3 prepare is the v2 rebuild path at work (a seed that flipped inside the
// quarter-lead confirm window, 3 DAA on the 60x profile); the gate judges the v3 path from its first prepare on
const firstV3Prepare = L.findIndex(l => /PREPARE sent for epoch seed .* class v3/.test(l));
const sinceV3 = (l, i) => firstV3Prepare < 0 || i >= firstV3Prepare;
const mismatch = L.filter((l, i) => sinceV3(l, i) && /program class mismatch|era seed mismatch|PACK OUT OF DATE|pack .*: program pack/.test(l));
const mismatchBefore = L.filter((l, i) => !sinceV3(l, i) && /program class mismatch|era seed mismatch|PACK OUT OF DATE/.test(l));
const refused = L.filter(l => /exiting with code 4[24]|refused the program pack|prepare-failed/.test(l));
const swaps = L.filter(l => /swapped with no pause|compiles inline|worker without prepare support/.test(l)).map(l => l.replace(/^\S+ /, ''));
const accepted = L.filter(l => /ACCEPTED block/.test(l));
const acceptedAfter = switchAt ? accepted.filter(l => (t(l) || 0) >= switchAt).length : 0;
const found = L.filter(l => /worker: found |^\S+ found /.test(l)).length;
const status = L.filter(l => /miner 'metal0' \[igneum-pow\]/.test(l)).map(l => l.replace(/^\S+ /, ''));
const lastStatus = status.at(-1) || '';
const mismatched = +(/mismatched=(\d+)/.exec(lastStatus)?.[1] ?? 0);
const rate = /hash=([\d.]+) MH\/s/.exec(lastStatus)?.[1];
metal = { worker: METAL, prepares, prepared, prepared_v3: preparedV3, need: need.length, mismatch_lines: mismatch, v2_rebuild_lines_before_v3: mismatchBefore, refused, swaps, accepted_total: accepted.length, accepted_after_switch: acceptedAfter, found_lines: found, cpu_recheck_mismatched: mismatched, rate_mh_s: rate, last_status: lastStatus };
}
const checks = {
switch_line_on_every_node: switchLines.every(Boolean),
switch_line_names_the_rounded_epoch: switchLines.every(l => l && l.includes(`active from epoch ${FIRST_V3_EPOCH} `)),
template_switched_at_the_first_v3_epoch: firstV3 != null && firstV3.epoch === FIRST_V3_EPOCH,
blocks_before_the_boundary: before.length > 0,
blocks_after_the_boundary: after.length > 0,
v2_and_v3_programs_seen: v2Ids.length > 0 && v3Ids.length > 0,
program_ids_differ_across_the_switch: v2Ids.length > 0 && v3Ids.length > 0 && !v2Ids.some(id => v3Ids.includes(id)),
miners_agree_on_every_program: programRows.every(p => !p.disagree),
zero_rejected_by_miners: rejectedMiner.every(r => r.length === 0),
zero_rejected_by_nodes: rejectedNode.every(r => r.length === 0),
sinks_agree: new Set(sinks).size === 1,
block_counts_agree: new Set(counts.map(String)).size === 1,
...(METAL ? {
metal_prepare_sent_for_v3: metal.prepares.some(l => / class v3 /.test(l)),
metal_worker_prepared_v3_pack: metal.prepared_v3.length > 0,
metal_no_need_or_mismatch: metal.need === 0 && metal.mismatch_lines.length === 0 && metal.refused.length === 0,
metal_accepted_blocks_after_switch: metal.accepted_after_switch > 0,
metal_cpu_recheck_clean: metal.cpu_recheck_mismatched === 0,
metal_swapped_without_pause: metal.swaps.some(l => /swapped with no pause/.test(l)),
} : {}),
};
const pass = Object.values(checks).every(Boolean);
const summary = {
pass, checks, activation: ACTIVATION, epoch_blocks: EPOCH, first_v3_epoch: FIRST_V3_EPOCH, boundary_daa: BOUNDARY, secs: SECS, threads: THREADS,
node: IGNEUMD, miner: CPU_MINER, genesis_bits: `0x${GENESIS_BITS.toString(16)}`,
template_switch: firstV3, run_ended_at_s: switchEnd, final_daa: lastDaa,
blocks: { total: blocks.length, before_boundary: before.length, after_boundary: after.length, chain_before: chainBefore, chain_after: chainAfter },
programs: programRows, accepted_per_miner: accepted,
rejected_by_miners: rejectedMiner.map(r => r.length), rejected_by_nodes: rejectedNode.map(r => r.length),
rejected_lines: [...rejectedMiner.flat(), ...rejectedNode.flat()].slice(0, 20),
sinks, block_counts: counts, tips_per_node: tips, switch_lines: switchLines, samples, metal,
};
writeFileSync(`${TMP}/summary.json`, JSON.stringify(summary, null, 2));
log(`SUMMARY ${pass ? 'PASS' : 'FAIL'}: blocks ${before.length} before / ${after.length} after the boundary at DAA ${BOUNDARY} (chain ${chainBefore} / ${chainAfter}); programs ${programRows.map(p => `e${p.epoch}:${p.class}:${p.program_id}:${Math.max(...p.ready_ms)}ms`).join(' ')}; rejected miners ${rejectedMiner.map(r => r.length).join('/')} nodes ${rejectedNode.map(r => r.length).join('/')}; sinks ${sinks.join(' ')} (${checks.sinks_agree ? 'agree' : 'DIFFER'}); block counts ${counts.join('/')}; switch lines ${switchLines.filter(Boolean).length}/3`);
if (metal) {
log(`METAL: ${metal.prepares.length} PREPARE lines (${metal.prepares.filter(l => / class v3 /.test(l)).length} class v3), ${metal.prepared.length} prepared (${metal.prepared_v3.length} v3), need ${metal.need}, mismatch/refusal lines ${metal.mismatch_lines.length + metal.refused.length}, accepted ${metal.accepted_total} (${metal.accepted_after_switch} after the switch), cpu re-check mismatched ${metal.cpu_recheck_mismatched}, rate ${metal.rate_mh_s} MH/s`);
for (const l of [...metal.prepares, ...metal.prepared, ...metal.swaps]) log(` ${l.slice(0, 260)}`);
for (const l of [...metal.mismatch_lines, ...metal.refused].slice(0, 10)) log(` BAD ${l.slice(0, 260)}`);
}
for (const [k, v] of Object.entries(checks)) if (!v) log(`FAILED CHECK ${k}`);
log(`summary: ${TMP}/summary.json`);
await stopAll();
process.exit(pass ? 0 : 1);

View file

@ -54,6 +54,12 @@
"difficulty_v2_activation_daa": 18446744073709551615,
"proving_v0_activation_daa": 18446744073709551615,
"finality_v3_activation_daa": 18446744073709551615,
"program_class_v3_activation_daa": 18446744073709551615,
"pow_genesis_dataset_log2": 28,
"proving_v1_activation_daa": 18446744073709551615,
"proving_v1_segment_blocks": 8,
"proving_v1_unproven_daa": 10,
"proving_v1_aggregator_share_bps": 1000,
"fees_v1_activation_daa": 0,
"fees": {"pgas": {"version": 1, "cycles_per_pgas": 1000, "intrinsic_pgas_per_tx": 300, "modexp_base": 10, "modexp_per_byte_numer": 1, "modexp_per_byte_denom": 10}, "block_proving_gas_limit": 120000, "shard_proving_gas_budget": 30000, "min_execution_base_fee_wei": 100000000000, "min_proving_base_fee_wei": 10000000000000, "initial_execution_base_fee_wei": 100000000000, "initial_proving_base_fee_wei": 10000000000000, "base_fee_change_denominator": 8}
}

View file

@ -35,7 +35,9 @@ for (const b of [IGNEUMD, VMINE]) if (!existsSync(b)) { console.error(`missing $
rmSync(TMP, { recursive: true, force: true }); mkdirSync(TMP, { recursive: true });
const override = `${TMP}/override.json`;
writeFileSync(override, JSON.stringify({ ...JSON.parse(readFileSync(FILE, 'utf8')), skip_proof_of_work: true }));
// merged as text, not through JSON.parse: a `never` height (18446744073709551615) does not survive a JavaScript number
// (the class-v3.mjs gate run of 5 October 2026 21:30Z: "invalid type: floating point 1.8446744073709552e+19")
writeFileSync(override, readFileSync(FILE, 'utf8').replace(/\s*"skip_proof_of_work":\s*[^,}\n]+,?/, '').replace(/,?\s*}\s*$/, ',\n "skip_proof_of_work": true\n}\n'));
class Node {
constructor(i, connect = []) {

View file

@ -20,6 +20,9 @@
#define __forceinline__ inline
struct uint3 { unsigned x, y, z; };
// uint4 and make_uint4 (read-width experiment, 5 October 2026): the wide loads and the scratch slots are 16-byte vectors.
struct uint4 { unsigned x, y, z, w; };
static inline uint4 make_uint4(unsigned x, unsigned y, unsigned z, unsigned w) { uint4 v; v.x = x; v.y = y; v.z = z; v.w = w; return v; }
struct dim3 { unsigned x, y, z; dim3(unsigned x_ = 1, unsigned y_ = 1, unsigned z_ = 1) : x(x_), y(y_), z(z_) {} };
extern thread_local uint3 threadIdx;
extern thread_local uint3 blockIdx;

28
proto-cuda/emu/test-layout.sh Executable file
View file

@ -0,0 +1,28 @@
#!/usr/bin/env bash
# The host-side dataset derivation of the harnesses under a non-linear item-to-word layout (era layout,
# docs/plans/era-layout.md 1.2): proto-cuda/host.cu (and proto-opencl/host.c, the same function) derive dataset words
# on the host through the pack's own mh_word, which carries the layout. Until 5 October 2026 both derived
# mh_item(w >> 4)[w & 15] and the "64 random points vs host derivation" check failed on every interleaved pack while
# the Mac's samples and the vectors passed (the known-failed case, docs/bench-log.md, era layout entry). This script
# runs the CUDA CPU emulation on an interleaved era pack and a linear one and requires OVERALL: PASS on both; it is
# the test that fails on the old derivation. Usage: emu/test-layout.sh [pack dir ...]
set -euo pipefail
HERE="$(cd "$(dirname "$0")" && pwd)"
CUDA_DIR="$(cd "$HERE/.." && pwd)"
PACKS=("$@")
[ ${#PACKS[@]} -gt 0 ] || PACKS=("$CUDA_DIR/packs-ca2-era/era-1" "$CUDA_DIR/packs/igneum-devnet-v4-epoch0")
CXX="${CXX:-c++}"
for P in "${PACKS[@]}"; do
grep -q "mh_addr(" "$P/memhard.h" && kind=interleaved || kind=linear
BUILD="$HERE/build-layout-$(basename "$P")"
mkdir -p "$BUILD"
cp "$CUDA_DIR/host.cu" "$BUILD/host_emu.cpp"
sed -E 's/([A-Za-z_0-9]+)<<<([^,]+), ([^>]+)>>>\(/emu_launch(\1, \2, \3, /' "$P/kernel.cu" > "$BUILD/kernel_emu.cpp"
sed -E 's/([A-Za-z_0-9]+)<<<([^,]+), ([^>]+)>>>\(/emu_launch(\1, \2, \3, /' "$P/kernel_bound.cu" > "$BUILD/kernel_bound_emu.cpp"
"$CXX" -std=c++17 -O2 -w -I "$HERE" -I "$P" -o "$BUILD/igneum-emu" "$BUILD/host_emu.cpp" "$BUILD/kernel_emu.cpp" -DIGNEUM_BOUND "$BUILD/kernel_bound_emu.cpp" "$HERE/shim.cpp" -pthread
out="$("$BUILD/igneum-emu" --batch-log2 13 --batches 1 2>&1)"
echo "$out" | grep -E "dataset self-test|OVERALL" | sed "s|^|$(basename "$P") ($kind): |"
echo "$out" | grep -q "^OVERALL: PASS" || { echo "FAIL: $P ($kind layout) did not pass the emulated harness"; exit 1; }
echo "$out" | grep -q "64 random points vs host derivation PASS" || { echo "FAIL: $P ($kind layout): the host derivation disagrees with the pack's layout"; exit 1; }
done
echo "PASS: the host derivation follows the pack's layout on ${#PACKS[@]} pack(s)"

View file

@ -115,11 +115,11 @@ static bool setupCache() {
return gCachePass;
}
// dataset[w] derived on the host from the host cache, exactly as proto-metal's verifier does it.
// dataset[w] derived on the host from the host cache through the pack's own mh_word (memhard.h), which carries the
// pack's item-to-word layout (era layout, 5 October 2026: the harness's former w >> 4 / w & 15 failed the random
// points of every interleaved pack while the Mac samples and vectors passed).
static uint32_t host_ds_word(uint32_t w) {
uint32_t s[16];
mh_item(hCache.data(), w >> 4u, s);
return s[w & 15u];
return mh_word(hCache.data(), w);
}
#endif

View file

@ -103,11 +103,41 @@ int main(int argc, char** argv) {
{
char text[2000], sw[200], kw[200];
words_hex(bare, sw); words_hex(keyw, kw);
snprintf(text, sizeof(text), "#define IGNEUM_SEED_BYTES_HEX \"%s\"\n#define IGNEUM_DAY_BYTES_HEX \"%s\"\n#define IGNEUM_DATASET_LOG2 28\n#define IGNEUM_DATASET_MODE 1\n#define IGNEUM_SEEDW_INIT { %s }\n#define IGNEUM_KEY_INIT { %s }\n#define IGNEUM_CACHE_LOG2_WORDS 26\n#define IGNEUM_CACHE_SEGMENTS 4096u\n", EPOCH_34, DAY_20731, sw, kw);
snprintf(text, sizeof(text), "#define IGNEUM_SEED_BYTES_HEX \"%s\"\n#define IGNEUM_DAY_BYTES_HEX \"%s\"\n#define IGNEUM_GENERATOR 2\n#define IGNEUM_DATASET_LOG2 28\n#define IGNEUM_DATASET_MODE 1\n#define IGNEUM_SEEDW_INIT { %s }\n#define IGNEUM_KEY_INIT { %s }\n#define IGNEUM_CACHE_LOG2_WORDS 26\n#define IGNEUM_CACHE_SEGMENTS 4096u\n", EPOCH_34, DAY_20731, sw, kw);
write_file(dir, "program.h", text);
write_file(dir, "seeds.txt", "epoch_seed_hex " EPOCH_34 "\nday_seed_hex " DAY_20731 "\n");
err[0] = 0;
CHECK(pf_load(dir, &pk, err, sizeof(err)) == 1 && pk.attempt == 0, "no attempt line reads as attempt 0 and the bare words load");
CHECK(strcmp(pk.programClass, "v2") == 0 && pk.eraHex[0] == 0, "a generator 2 pack without a class line is class v2 with no era");
}
// 4. Program classes (Counter ASIC 2.0, 5 October 2026, spec 01 section 1.4.5): a generator this worker does not
// run is refused; a generator 3 pack is class v3 and carries its era seed; a class line that contradicts the
// generator is refused; a pack with no generator line at all (generator 1, the retired lever generator) is refused.
{
char text[2400], sw[200], kw[200], why[256];
words_hex(bare, sw); words_hex(keyw, kw);
snprintf(text, sizeof(text), "#define IGNEUM_SEED_BYTES_HEX \"%s\"\n#define IGNEUM_DAY_BYTES_HEX \"%s\"\n#define IGNEUM_GENERATOR 9\n#define IGNEUM_DATASET_LOG2 28\n#define IGNEUM_DATASET_MODE 1\n#define IGNEUM_SEEDW_INIT { %s }\n#define IGNEUM_KEY_INIT { %s }\n#define IGNEUM_CACHE_LOG2_WORDS 26\n#define IGNEUM_CACHE_SEGMENTS 4096u\n", EPOCH_34, DAY_20731, sw, kw);
write_file(dir, "program.h", text);
err[0] = 0;
CHECK(pf_load(dir, &pk, err, sizeof(err)) == 0 && strstr(err, "generator 9 is not a generator version this worker runs") != NULL, "generator 9 is refused in plain words");
snprintf(text, sizeof(text), "#define IGNEUM_SEED_BYTES_HEX \"%s\"\n#define IGNEUM_DAY_BYTES_HEX \"%s\"\n#define IGNEUM_GENERATOR 3\n#define IGNEUM_PROGRAM_CLASS \"v3\"\n#define IGNEUM_ERA_SEED_HEX \"%s\"\n#define IGNEUM_DATASET_LOG2 28\n#define IGNEUM_DATASET_MODE 1\n#define IGNEUM_SEEDW_INIT { %s }\n#define IGNEUM_KEY_INIT { %s }\n#define IGNEUM_CACHE_LOG2_WORDS 26\n#define IGNEUM_CACHE_SEGMENTS 4096u\n", EPOCH_34, DAY_20731, EPOCH_33, sw, kw);
write_file(dir, "program.h", text);
err[0] = 0;
CHECK(pf_load(dir, &pk, err, sizeof(err)) == 1 && strcmp(pk.programClass, "v3") == 0 && strcmp(pk.eraHex, EPOCH_33) == 0, "a generator 3 pack loads as class v3 with its era seed");
CHECK(pf_pack_class_ok(pk.programClass, pk.eraHex, "v3", EPOCH_33, why, sizeof(why)) == 1, "the v3 pack matches a job naming class v3 and its era");
CHECK(pf_pack_class_ok(pk.programClass, pk.eraHex, "", "", why, sizeof(why)) == 1, "a job naming no class accepts the pack");
CHECK(pf_pack_class_ok(pk.programClass, pk.eraHex, "v2", "", why, sizeof(why)) == 0 && strstr(why, "program class mismatch") == why, "a job naming class v2 refuses the v3 pack");
CHECK(pf_pack_class_ok(pk.programClass, pk.eraHex, "v3", EPOCH_34, why, sizeof(why)) == 0 && strstr(why, "era seed mismatch") == why, "a job naming another era refuses the v3 pack");
CHECK(pf_pack_class_ok("v2", "", "v3", EPOCH_33, why, sizeof(why)) == 0, "a job naming class v3 refuses a v2 pack");
CHECK(pf_pack_class_ok("v2", "", "v2", EPOCH_33, why, sizeof(why)) == 1, "an era named on a v2 job is ignored");
snprintf(text, sizeof(text), "#define IGNEUM_SEED_BYTES_HEX \"%s\"\n#define IGNEUM_DAY_BYTES_HEX \"%s\"\n#define IGNEUM_GENERATOR 3\n#define IGNEUM_PROGRAM_CLASS \"v2\"\n#define IGNEUM_DATASET_LOG2 28\n#define IGNEUM_DATASET_MODE 1\n#define IGNEUM_SEEDW_INIT { %s }\n#define IGNEUM_KEY_INIT { %s }\n#define IGNEUM_CACHE_LOG2_WORDS 26\n#define IGNEUM_CACHE_SEGMENTS 4096u\n", EPOCH_34, DAY_20731, sw, kw);
write_file(dir, "program.h", text);
err[0] = 0;
CHECK(pf_load(dir, &pk, err, sizeof(err)) == 0 && strstr(err, "does not match IGNEUM_GENERATOR 3") != NULL, "a class line that contradicts the generator is refused");
snprintf(text, sizeof(text), "#define IGNEUM_SEED_BYTES_HEX \"%s\"\n#define IGNEUM_DAY_BYTES_HEX \"%s\"\n#define IGNEUM_DATASET_LOG2 28\n#define IGNEUM_DATASET_MODE 1\n#define IGNEUM_SEEDW_INIT { %s }\n#define IGNEUM_KEY_INIT { %s }\n#define IGNEUM_CACHE_LOG2_WORDS 26\n#define IGNEUM_CACHE_SEGMENTS 4096u\n", EPOCH_34, DAY_20731, sw, kw);
write_file(dir, "program.h", text);
err[0] = 0;
CHECK(pf_load(dir, &pk, err, sizeof(err)) == 0 && strstr(err, "generator 1 is not") != NULL, "a pack with no generator line (generator 1) is refused");
}
printf("%s: %d failure(s)\n", argv[0], failures);

View file

@ -26,6 +26,21 @@ typedef struct {
uint32_t attempt; // IGNEUM_PROGRAM_ATTEMPT: seedw are the words of this attempt of the epoch seed (0 = bare seed)
uint32_t seedw[8], keyw[8];
char seedString[600];
// read-width experiment (5 October 2026): the load class (0 when absent), bytes per hash, variant 5's scratch
uint32_t loadsPerHash, bytesPerHash, scratchOps, persistent, scratchWordsPerLane;
char loadClass[64];
char programClass[8]; /* IGNEUM_PROGRAM_CLASS: "v2" or "v3" (Counter ASIC 2.0); absent = the generator's class */
char eraHex[65]; /* IGNEUM_ERA_SEED_HEX of a class v3 chain pack; empty otherwise */
// Counter ASIC 2.0 (5 October 2026): the mixer multiplier of the item derivation (IGNEUM_MIXER_MULT, 1 when absent:
// version 2; 4 under class v3). The emitted memhard.h / kernel.cl carry it in their text; this is for the log lines.
uint32_t mixerMult;
// hot-table experiment (5 October 2026, docs/plans/hot-table.md): hotMb 0 when the pack has no hot table; the
// table is filled on the device from the pack's igneum_hot_fill (never shipped), its self-test values from vectors.h
uint32_t hotMb, hotWords, hotSegments, hotSlots;
uint32_t hotKey[8];
int haveHot;
uint32_t hotHead[16], hotLast[16];
uint64_t hotFnv;
// seeds.txt (or program.h): the seeds as the worker protocol carries them
char epochHex[65];
char dayHex[PF_HEX_CAP];
@ -268,8 +283,36 @@ static int pf_load(const char* dir, PfPack* pk, char* err, size_t cap) {
if (!pf_define_u32(prog, "IGNEUM_CACHE_LOG2_WORDS", &pk->cacheLog2Words)) { free(prog); return pf_fail(err, cap, "program.h has no IGNEUM_CACHE_LOG2_WORDS"); }
if (!pf_define_u32(prog, "IGNEUM_CACHE_SEGMENTS", &pk->cacheSegments)) { free(prog); return pf_fail(err, cap, "program.h has no IGNEUM_CACHE_SEGMENTS"); }
if (!pf_define_u32(prog, "IGNEUM_GENERATOR", &pk->generator)) pk->generator = 1;
/* Spec 01 section 1.4.5: a pack whose generator version is not one this worker runs is refused. Generator 2 is
* program class v2 (the lottery hash of 4 October 2026), generator 3 is class v3 (Counter ASIC 2.0). */
if (pk->generator != 2 && pk->generator != 3) {
char m[200]; snprintf(m, sizeof(m), "program pack generator %u is not a generator version this worker runs (2 or 3)", (unsigned)pk->generator);
free(prog); return pf_fail(err, cap, m);
}
strcpy(pk->programClass, pk->generator == 3 ? "v3" : "v2");
{
char named[8] = {0};
if (pf_define_str(prog, "IGNEUM_PROGRAM_CLASS", named, sizeof(named)) && strcmp(named, pk->programClass) != 0) {
char m[200]; snprintf(m, sizeof(m), "program pack IGNEUM_PROGRAM_CLASS \"%.7s\" does not match IGNEUM_GENERATOR %u", named, (unsigned)pk->generator);
free(prog); return pf_fail(err, cap, m);
}
}
pk->eraHex[0] = 0; pf_define_str(prog, "IGNEUM_ERA_SEED_HEX", pk->eraHex, sizeof(pk->eraHex));
if (!pf_define_u32(prog, "IGNEUM_PROGRAM_ATTEMPT", &pk->attempt)) pk->attempt = 0;
if (pf_define_words(prog, "IGNEUM_SEEDW_INIT", pk->seedw, 8) != 8) { free(prog); return pf_fail(err, cap, "program.h has no IGNEUM_SEEDW_INIT with 8 words"); }
pk->loadsPerHash = 128; pf_define_u32(prog, "IGNEUM_LOADS_PER_HASH", &pk->loadsPerHash);
pk->bytesPerHash = pk->loadsPerHash * 4u; pf_define_u32(prog, "IGNEUM_BYTES_PER_HASH", &pk->bytesPerHash);
pk->scratchOps = 0; pf_define_u32(prog, "IGNEUM_SCRATCH_OPS", &pk->scratchOps);
pk->persistent = 0; pf_define_u32(prog, "IGNEUM_PERSISTENT_WARPS", &pk->persistent);
pk->scratchWordsPerLane = 8192; pf_define_u32(prog, "IGNEUM_SCRATCH_WORDS_PER_LANE", &pk->scratchWordsPerLane);
strcpy(pk->loadClass, "v2"); pf_define_str(prog, "IGNEUM_LOAD_CLASS", pk->loadClass, sizeof(pk->loadClass));
pk->mixerMult = 1; pf_define_u32(prog, "IGNEUM_MIXER_MULT", &pk->mixerMult);
pk->hotMb = 0; pf_define_u32(prog, "IGNEUM_HOT_MB", &pk->hotMb);
if (pk->hotMb) {
if (!pf_define_u32(prog, "IGNEUM_HOT_WORDS", &pk->hotWords) || !pf_define_u32(prog, "IGNEUM_HOT_SEGMENTS", &pk->hotSegments) ||
!pf_define_u32(prog, "IGNEUM_HOT_SLOTS", &pk->hotSlots) || pf_define_words(prog, "IGNEUM_HOT_KEY_INIT", pk->hotKey, 8) != 8) { free(prog); return pf_fail(err, cap, "program.h has IGNEUM_HOT_MB but not IGNEUM_HOT_WORDS, IGNEUM_HOT_SEGMENTS, IGNEUM_HOT_SLOTS and IGNEUM_HOT_KEY_INIT"); }
if (pk->hotWords != pk->hotMb * 262144u || pk->hotSegments != pk->hotMb * 256u || pk->hotMb > 4096u) { free(prog); return pf_fail(err, cap, "program.h hot table sizes disagree (words must be MiB x 2^18, segments MiB x 256)"); }
}
if (pf_define_words(prog, "IGNEUM_KEY_INIT", pk->keyw, 8) != 8) { free(prog); return pf_fail(err, cap, "program.h has no IGNEUM_KEY_INIT with 8 words"); }
if (!pf_define_str(prog, "IGNEUM_SEED_STRING", pk->seedString, sizeof(pk->seedString))) strncpy(pk->seedString, "(no IGNEUM_SEED_STRING)", sizeof(pk->seedString) - 1);
pf_define_str(prog, "IGNEUM_SEED_BYTES_HEX", ehex, sizeof(ehex));
@ -290,15 +333,19 @@ static int pf_load(const char* dir, PfPack* pk, char* err, size_t cap) {
strcpy(ehex, e2); strcpy(dhex, d2);
}
if (!ehex[0] || !dhex[0]) return pf_fail(err, cap, "no seeds: neither seeds.txt nor IGNEUM_SEED_BYTES_HEX / IGNEUM_DAY_BYTES_HEX in program.h (a pack from igneum-pow export --seed <name> has no byte seeds)");
if (strlen(ehex) != 64) return pf_fail(err, cap, "epoch seed is not 64 hex characters");
// The chain's epoch seed is 32 bytes (64 hex characters). A pack exported from a seed STRING (igneum-pow export
// --seed <name>, the read-width experiment's packs of 5 October 2026) carries the string's bytes instead; the
// seed-word re-derivation below checks either form, so any even-length hex seed is accepted here. The serve
// protocol still carries 64-hex seeds; a string-seed pack can only be benched (--bench, --bench-pack, --check).
if (strlen(ehex) < 2 || strlen(ehex) % 2 != 0) return pf_fail(err, cap, "epoch seed is not an even-length hex string");
strcpy(pk->epochHex, ehex); strcpy(pk->dayHex, dhex);
// The seed words derived from the bytes AND the attempt must be the pack's own words: otherwise the pack and
// its seeds disagree (a half rewritten directory, or an exporter on another rule)
{
uint32_t w[8];
char m[400];
if (!pf_unhex(ehex, bytes, 32, &blen) || blen != 32) return pf_fail(err, cap, "epoch seed hex is malformed");
pf_program_words(bytes, 32, pk->attempt, w);
if (!pf_unhex(ehex, bytes, sizeof(bytes), &blen) || blen == 0) return pf_fail(err, cap, "epoch seed hex is malformed");
pf_program_words(bytes, blen, pk->attempt, w);
if (memcmp(w, pk->seedw, 32) != 0) {
snprintf(m, sizeof(m), "program pack and its seeds disagree: IGNEUM_SEEDW_INIT is not attempt %u of the epoch seed %.16s (attempt %u gives %08x %08x ..., the pack has %08x %08x ...); run igneum-miner export-pack again",
(unsigned)pk->attempt, ehex, (unsigned)pk->attempt, w[0], w[1], pk->seedw[0], pk->seedw[1]);
@ -330,6 +377,9 @@ static int pf_load(const char* dir, PfPack* pk, char* err, size_t cap) {
pf_symbol_numbers(vec, "IGNEUM_CACHE_FNV64", &pk->cacheFnv, 1) == 1 && pk->vecWarps > 0) {
uint32_t ns = 0;
pk->haveVectors = 1;
if (pk->hotMb && pf_symbol_u32s(vec, "IGNEUM_HOT_HEAD", pk->hotHead, 16) == 16 &&
pf_symbol_u32s(vec, "IGNEUM_HOT_LAST", pk->hotLast, 16) == 16 &&
pf_symbol_numbers(vec, "IGNEUM_HOT_FNV64", &pk->hotFnv, 1) == 1) pk->haveHot = 1;
if (pf_define_u32(vec, "IGNEUM_DS_SAMPLES", &ns) && ns > 0 && ns <= PF_MAX_SAMPLES) {
int a = pf_symbol_u32s(vec, "IGNEUM_DS_SAMPLE_INDEX", pk->sampleIdx, (int)ns);
int b = pf_symbol_u32s(vec, "IGNEUM_DS_SAMPLE_VALUE", pk->sampleVal, (int)ns);
@ -343,23 +393,32 @@ static int pf_load(const char* dir, PfPack* pk, char* err, size_t cap) {
// The self-test verdict from values the host read back from the device. `vec` holds vecWarps x 32 outputs of the
// bound kernel run with the pack's own seed words as init words (that is igneum_hash of kernel.cu). Writes one line.
// A hot-table pack (pk->hotMb) also hands the hot table's head, last line and FNV-1a 64 (NULL and 0 otherwise); a
// hot pack whose vectors.h carries no hot values is not checked on the table (the vectors cover it) and says so.
static int pf_selftest(const PfPack* pk, const uint32_t* cacheHead, const uint32_t* cacheLast, uint64_t cacheFnv,
const uint32_t* dsHead, uint32_t dsLast, const uint32_t* sampleVals, const uint64_t* vec,
const uint32_t* hotHead, const uint32_t* hotLast, uint64_t hotFnv,
char* out, size_t cap) {
int okCH = memcmp(cacheHead, pk->cacheHead, 64) == 0, okCL = memcmp(cacheLast, pk->cacheLast, 64) == 0;
int okFnv = (cacheFnv == pk->cacheFnv);
int okDH = memcmp(dsHead, pk->dsHead, 64) == 0, okDL = (dsLast == pk->dsLast);
int hotChecked = (pk->hotMb && pk->haveHot && hotHead && hotLast);
int okHot = !hotChecked || (memcmp(hotHead, pk->hotHead, 64) == 0 && memcmp(hotLast, pk->hotLast, 64) == 0 && hotFnv == pk->hotFnv);
int badS = 0, badV = 0, i, w, l, firstBadWarp = -1, firstBadLane = -1;
for (i = 0; i < pk->nSamples; ++i) if (sampleVals[i] != pk->sampleVal[i]) ++badS;
for (w = 0; w < pk->vecWarps; ++w) for (l = 0; l < 32; ++l) if (vec[w * 32 + l] != pk->vecOut[w][l]) { if (firstBadWarp < 0) { firstBadWarp = w; firstBadLane = l; } ++badV; }
if (okCH && okCL && okFnv && okDH && okDL && badS == 0 && badV == 0) {
snprintf(out, cap, "self-test PASS (cache head, last line and FNV-1a 64 %016llx; dataset head, word [%u] and %d samples; %d of %d vector lanes)",
(unsigned long long)cacheFnv, pk->dsLastIndex, pk->nSamples, pk->vecWarps * 32, pk->vecWarps * 32);
if (okCH && okCL && okFnv && okDH && okDL && okHot && badS == 0 && badV == 0) {
if (pk->hotMb)
snprintf(out, cap, "self-test PASS (cache head, last line and FNV-1a 64 %016llx; dataset head, word [%u] and %d samples; hot table %u MiB %s; %d of %d vector lanes)",
(unsigned long long)cacheFnv, pk->dsLastIndex, pk->nSamples, pk->hotMb, hotChecked ? "head, last line and FNV-1a 64 ok" : "not in vectors.h (the vectors cover it)", pk->vecWarps * 32, pk->vecWarps * 32);
else
snprintf(out, cap, "self-test PASS (cache head, last line and FNV-1a 64 %016llx; dataset head, word [%u] and %d samples; %d of %d vector lanes)",
(unsigned long long)cacheFnv, pk->dsLastIndex, pk->nSamples, pk->vecWarps * 32, pk->vecWarps * 32);
return 1;
}
snprintf(out, cap, "self-test FAIL (cache head %s, cache last %s, cache FNV %016llx vs pack %016llx %s, dataset head %s, dataset last %s, samples %d bad of %d, vector lanes %d bad of %d%s)",
snprintf(out, cap, "self-test FAIL (cache head %s, cache last %s, cache FNV %016llx vs pack %016llx %s, dataset head %s, dataset last %s, samples %d bad of %d, hot table %s, vector lanes %d bad of %d%s)",
okCH ? "ok" : "BAD", okCL ? "ok" : "BAD", (unsigned long long)cacheFnv, (unsigned long long)pk->cacheFnv, okFnv ? "ok" : "BAD",
okDH ? "ok" : "BAD", okDL ? "ok" : "BAD", badS, pk->nSamples, badV, pk->vecWarps * 32,
okDH ? "ok" : "BAD", okDL ? "ok" : "BAD", badS, pk->nSamples, hotChecked ? (okHot ? "ok" : "BAD") : "none", badV, pk->vecWarps * 32,
firstBadWarp >= 0 ? " (first bad lane in the warp at base nonce" : "");
if (firstBadWarp >= 0) {
size_t n = strlen(out);
@ -369,4 +428,29 @@ static int pf_selftest(const PfPack* pk, const uint32_t* cacheHead, const uint32
return 0;
}
/* Counter ASIC 2.0 (5 October 2026): a job or prepare line may end with `class=<v2|v3>` and `era=<hex>` tokens (sent
* only when the chain is on class v3, so every v2 line is the line of before). A pack matches the line when its class
* is the named class and, when an era is named, its era seed is that era. Empty wanted strings accept any pack.
* Returns 1 on a match, else 0 with the reason in `why`. */
static int pf_pack_class_ok(const char* packClass, const char* packEra, const char* wantClass, const char* wantEra, char* why, size_t cap) {
if (wantClass && wantClass[0] && strcmp(wantClass, packClass) != 0) {
snprintf(why, cap, "program class mismatch: this pack is class %s, the job names class %s (export the pack again)", packClass, wantClass);
return 0;
}
if (wantEra && wantEra[0] && strcmp(packClass, "v3") == 0) {
size_t i; int eq = strlen(packEra) == strlen(wantEra);
for (i = 0; eq && packEra[i]; ++i) if (tolower((unsigned char)packEra[i]) != tolower((unsigned char)wantEra[i])) eq = 0;
if (!eq) { snprintf(why, cap, "era seed mismatch: this pack was drawn under era %.16s, the job names era %.16s (export the pack again)", packEra[0] ? packEra : "(none)", wantEra); return 0; }
}
return 1;
}
/* Reads a `class=` or `era=` token into `cls` / `era` (small fixed buffers). Returns 1 when the token was one of them. */
static int pf_class_token(const char* tok, char* cls, size_t clsCap, char* era, size_t eraCap) {
if (strncmp(tok, "class=", 6) == 0) { strncpy(cls, tok + 6, clsCap - 1); cls[clsCap - 1] = 0; return 1; }
if (strncmp(tok, "era=", 4) == 0) { strncpy(era, tok + 4, eraCap - 1); era[eraCap - 1] = 0; return 1; }
return 0;
}
#endif

View file

@ -241,6 +241,10 @@ static bool loadNvrtc(Rtc& r, std::string& err, std::string& libName) {
struct Ctx {
Drv drv;
Rtc rtc;
// read-width experiment, variant 5 (5 October 2026): persistent warps and their scratch; --warps caps the launch
int warps = 0; // 0 = the resident capacity from the occupancy query, rounded down to a power of two
uint32_t salt = 1; // the running per-unit tag salt (+= units per launch)
int batches = 5; // --bench: timed dispatches
CUdevice dev = 0;
CUcontext ctx = nullptr;
std::string name;
@ -421,12 +425,45 @@ struct Pair {
std::string variant = "base"; // the bound kernel in service: a variant name (see allVariants)
std::string raceLine; // the race's one-line report, emitted by the main thread with "prepared"
double raceMs = 0;
// read-width experiment (5 October 2026): the pack's load class and, for variant 5, the persistent-warp scratch
std::string loadClass = "v2";
std::string programClass = "v2", eraHex; // Counter ASIC 2.0: the pack's class and era seed (packfile.h)
uint32_t loadsPerHash = 128, bytesPerHash = 512, scratchOps = 0;
bool persistent = false;
CUdeviceptr scratch = 0;
int warps = 0; // persistent warps launched (the arena holds this many)
int residentWarps = 0; // the occupancy query's capacity: blocks/SM x warps/block x SMs
size_t scratchBytes = 0;
// hot-table experiment (5 October 2026, docs/plans/hot-table.md): the epoch's hot table, filled on the device by the
// pack's igneum_hot_fill, the argument after the init words
uint32_t hotMb = 0, hotWords = 0, hotSegments = 0, hotSlots = 0;
CUfunction fHotFill = nullptr;
CUdeviceptr hot = 0;
double hotMs = 0;
};
static bool pairIs(const Pair* p, const std::string& epochHex, const std::string& dayHex) {
return p && hexEq(p->epochHex, epochHex) && hexEq(p->dayHex, dayHex);
}
// Counter ASIC 2.0: a job that names a class (and an era) belongs to a pair of that class (and era) only, so a pack
// of the old class for the same seeds is not this job's pair and the prepared pack of the right class wins.
static bool pairIsClass(const Pair* p, const std::string& epochHex, const std::string& dayHex, const std::string& cls, const std::string& era) {
char why[256];
return pairIs(p, epochHex, dayHex) && pf_pack_class_ok(p->programClass.c_str(), p->eraHex.c_str(), cls.c_str(), era.c_str(), why, sizeof(why));
}
// The trailing `class=` and `era=` tokens of a job or prepare line (absent on every class v2 line), removed from `f`.
static void takeClassTokens(std::vector<std::string>& f, std::string& cls, std::string& era) {
while (!f.empty()) {
char c[8] = {0}, e[65] = {0};
if (!pf_class_token(f.back().c_str(), c, sizeof(c), e, sizeof(e))) break;
if (c[0]) cls = c;
if (e[0]) era = e;
f.pop_back();
}
}
// The job loop and a race take turns on the card: a variant is timed with no job running (exclusive numbers), and
// mining resumes between variants. Held per chunk by the job loop, per variant by the race.
static std::mutex gpuMutex;
@ -435,6 +472,8 @@ static void releasePair(Ctx& c, Pair* p) {
if (!p) return;
if (p->ds) c.drv.memFree(p->ds);
if (p->cache) c.drv.memFree(p->cache);
if (p->scratch) c.drv.memFree(p->scratch);
if (p->hot) c.drv.memFree(p->hot);
if (p->modBound) c.drv.moduleUnload(p->modBound);
if (p->modKernel) c.drv.moduleUnload(p->modKernel);
delete p;
@ -446,7 +485,20 @@ struct IgneumInitWordsArg { uint32_t w[8]; };
static bool launchHash(Ctx& c, Pair* p, CUdeviceptr out, uint32_t baseNonce, const uint32_t iw[8], uint32_t nonces, uint32_t block, CUstream s, std::string& err) {
uint32_t mask = p->words - 1u;
IgneumInitWordsArg a; std::memcpy(a.w, iw, 32);
void* args[5] = { &p->ds, &out, &baseNonce, &mask, &a };
if (p->persistent) {
// Variant 5: N persistent warps over nonces / 32 units; the arena was sized for p->warps warps in buildPair.
uint32_t units = nonces / 32u, warps = (uint32_t)p->warps;
if (warps > units) warps = units;
while (warps > 1u && units % warps != 0u) warps >>= 1;
if (block != 32u) { err = "a variant-5 pack runs one warp per block (--block-warps 1)"; return false; }
uint32_t salt = c.salt; c.salt += units;
// the hot table (when the pack has one) sits between the init words and the scratch triple
void* args[9] = { &p->ds, &out, &baseNonce, &mask, &a, &p->scratch, &units, &salt, nullptr };
if (p->hot) { args[5] = &p->hot; args[6] = &p->scratch; args[7] = &units; args[8] = &salt; }
DRV_CHECK(c, c.drv.launchKernel(p->fHashBound, warps, 1, 1, 32, 1, 1, 0, s, args, nullptr), "cuLaunchKernel igneum_hash_bound (persistent)");
return true;
}
void* args[6] = { &p->ds, &out, &baseNonce, &mask, &a, &p->hot };
DRV_CHECK(c, c.drv.launchKernel(p->fHashBound, nonces / block, 1, 1, block, 1, 1, 0, s, args, nullptr), "cuLaunchKernel igneum_hash_bound");
return true;
}
@ -769,7 +821,9 @@ static Pair* buildPair(Ctx& c, const std::string& dir, CUstream s, std::string&
p->datasetLog2 = pk.datasetLog2; p->words = 1u << pk.datasetLog2; p->cacheWords = 1u << pk.cacheLog2Words; p->cacheSegments = pk.cacheSegments;
// Compile
Compiled ck, cb;
if (!rtcCompile(c, kernelDev, "kernel.cu", programH, memhardH, { "igneum_cache_fill", "igneum_build" }, ck, err)) { releasePair(c, p); return nullptr; }
std::vector<std::string> kernelNames = { "igneum_cache_fill", "igneum_build" };
if (pk.hotMb) kernelNames.push_back("igneum_hot_fill");
if (!rtcCompile(c, kernelDev, "kernel.cu", programH, memhardH, kernelNames, ck, err)) { releasePair(c, p); return nullptr; }
if (!rtcCompile(c, boundDev, "kernel_bound.cu", programH, memhardH, { "igneum_hash_bound" }, cb, err)) { releasePair(c, p); return nullptr; }
p->compileMs = ck.ms + cb.ms;
// Load
@ -781,19 +835,45 @@ static Pair* buildPair(Ctx& c, const std::string& dir, CUstream s, std::string&
if (c.drv.moduleGetFunction(&p->fCacheFill, p->modKernel, ck.lowered[0].c_str()) != CUDA_SUCCESS) { err = "igneum_cache_fill (" + ck.lowered[0] + ") not in the module"; releasePair(c, p); return nullptr; }
if (c.drv.moduleGetFunction(&p->fBuild, p->modKernel, ck.lowered[1].c_str()) != CUDA_SUCCESS) { err = "igneum_build (" + ck.lowered[1] + ") not in the module"; releasePair(c, p); return nullptr; }
if (c.drv.moduleGetFunction(&p->fHashBound, p->modBound, cb.lowered[0].c_str()) != CUDA_SUCCESS) { err = "igneum_hash_bound (" + cb.lowered[0] + ") not in the module"; releasePair(c, p); return nullptr; }
if (pk.hotMb && c.drv.moduleGetFunction(&p->fHotFill, p->modKernel, ck.lowered[2].c_str()) != CUDA_SUCCESS) { err = "igneum_hot_fill (" + ck.lowered[2] + ") not in the module"; releasePair(c, p); return nullptr; }
c.drv.funcGetAttribute(&p->regs, CU_FUNC_ATTRIBUTE_NUM_REGS, p->fHashBound);
c.drv.occupancy(&p->blocksPerSM, p->fHashBound, 32 * c.blockWarps, 0);
}
p->loadClass = pk.loadClass; p->loadsPerHash = pk.loadsPerHash; p->bytesPerHash = pk.bytesPerHash; p->scratchOps = pk.scratchOps;
p->programClass = pk.programClass; p->eraHex = pk.eraHex;
p->persistent = pk.persistent != 0;
p->hotMb = pk.hotMb; p->hotWords = pk.hotWords; p->hotSegments = pk.hotSegments; p->hotSlots = pk.hotSlots;
p->residentWarps = p->blocksPerSM * c.blockWarps * c.sms;
size_t scratchBytes = 0, hotBytes = (size_t)pk.hotWords * 4u;
if (p->persistent) {
// Variant 5: one arena per launched warp. The launch is the resident capacity (the occupancy query), rounded
// down to a power of two so it divides every batch, or --warps; the allocation cannot change the occupancy
// (registers and shared memory decide it), and the number is re-queried after the allocation below to show it.
if (c.blockWarps != 1) { err = "a variant-5 pack runs one warp per block: use --block-warps 1"; releasePair(c, p); return nullptr; }
int w = c.warps > 0 ? c.warps : p->residentWarps;
int pw = 1; while (pw * 2 <= w) pw *= 2;
p->warps = pw;
scratchBytes = (size_t)p->warps * 32u * (size_t)pk.scratchWordsPerLane * 4u;
}
// Cache
double t0 = wallMs();
size_t cacheBytes = (size_t)p->cacheWords * 4u, dsBytes = (size_t)p->words * 4u;
{
size_t freeB = 0, totalB = 0;
if (c.drv.memGetInfo(&freeB, &totalB) == CUDA_SUCCESS && freeB < cacheBytes + dsBytes + (64u << 20)) {
err = fmt("%llu MiB free on the device, this pack needs %llu MiB (cache %llu + dataset %llu)", (unsigned long long)(freeB >> 20), (unsigned long long)((cacheBytes + dsBytes) >> 20), (unsigned long long)(cacheBytes >> 20), (unsigned long long)(dsBytes >> 20));
if (c.drv.memGetInfo(&freeB, &totalB) == CUDA_SUCCESS && freeB < cacheBytes + dsBytes + scratchBytes + hotBytes + (64u << 20)) {
err = fmt("%llu MiB free on the device, this pack needs %llu MiB (cache %llu + dataset %llu + scratch %llu + hot %llu)", (unsigned long long)(freeB >> 20), (unsigned long long)((cacheBytes + dsBytes + scratchBytes + hotBytes) >> 20), (unsigned long long)(cacheBytes >> 20), (unsigned long long)(dsBytes >> 20), (unsigned long long)(scratchBytes >> 20), (unsigned long long)(hotBytes >> 20));
releasePair(c, p); return nullptr;
}
}
if (p->persistent) {
CUresult r = c.drv.memAlloc(&p->scratch, scratchBytes);
if (r != CUDA_SUCCESS) { err = "cuMemAlloc scratch: " + c.err(r); p->scratch = 0; releasePair(c, p); return nullptr; }
p->scratchBytes = scratchBytes;
int after = 0;
c.drv.occupancy(&after, p->fHashBound, 32 * c.blockWarps, 0);
info(fmt("variant 5: %d persistent warps (resident capacity %d = %d blocks/SM x %d warps/block x %d SMs; occupancy query after the allocation %d blocks/SM), scratch %llu MiB (%u KiB per warp)",
p->warps, p->residentWarps, p->blocksPerSM, c.blockWarps, c.sms, after, (unsigned long long)(scratchBytes >> 20), pk.scratchWordsPerLane * 4u * 32u / 1024u));
}
{
CUresult r = c.drv.memAlloc(&p->cache, cacheBytes);
if (r != CUDA_SUCCESS) { err = "cuMemAlloc cache: " + c.err(r); p->cache = 0; releasePair(c, p); return nullptr; }
@ -816,6 +896,18 @@ static Pair* buildPair(Ctx& c, const std::string& dir, CUstream s, std::string&
if (r != CUDA_SUCCESS) { err = "dataset build: " + c.err(r); releasePair(c, p); return nullptr; }
}
p->dsMs = wallMs() - t0;
// Hot table (hot-table experiment): filled from the epoch seed by the pack's own kernel, never shipped
if (pk.hotMb) {
t0 = wallMs();
CUresult r = c.drv.memAlloc(&p->hot, hotBytes);
if (r != CUDA_SUCCESS) { err = "cuMemAlloc hot table: " + c.err(r); p->hot = 0; releasePair(c, p); return nullptr; }
uint32_t nSeg = pk.hotSegments, block = 256u, grid = (nSeg + block - 1u) / block;
void* args[2] = { &p->hot, &nSeg };
r = c.drv.launchKernel(p->fHotFill, grid, 1, 1, block, 1, 1, 0, s, args, nullptr);
if (r == CUDA_SUCCESS) r = c.drv.streamSynchronize(s);
if (r != CUDA_SUCCESS) { err = "hot table fill: " + c.err(r); releasePair(c, p); return nullptr; }
p->hotMs = wallMs() - t0;
}
// Self-test against vectors.h
t0 = wallMs();
if (!pk.haveVectors) {
@ -844,8 +936,17 @@ static Pair* buildPair(Ctx& c, const std::string& dir, CUstream s, std::string&
if ((r = c.drv.streamSynchronize(s)) != CUDA_SUCCESS || (r = c.drv.memcpyDtoH(&vec[(size_t)w * 32u], out, 32u * 8u)) != CUDA_SUCCESS) { err = "vector warp: " + c.err(r); c.drv.memFree(out); releasePair(c, p); return nullptr; }
}
c.drv.memFree(out);
uint32_t hotHead[16] = {0}, hotLast[16] = {0};
uint64_t hotFnv = 0;
if (p->hot) {
std::vector<uint32_t> hw(p->hotWords);
if ((r = c.drv.memcpyDtoH(hw.data(), p->hot, hotBytes)) != CUDA_SUCCESS) { err = "cuMemcpyDtoH hot table: " + c.err(r); releasePair(c, p); return nullptr; }
std::memcpy(hotHead, hw.data(), 64);
std::memcpy(hotLast, hw.data() + p->hotWords - 16u, 64);
hotFnv = pf_fnv1a64(hw.data(), hotBytes);
}
char line[1024];
p->checkPass = pf_selftest(&pk, cacheHead, cacheLast, fnv, dsHead, dsLast, samples.data(), vec.data(), line, sizeof(line)) != 0;
p->checkPass = pf_selftest(&pk, cacheHead, cacheLast, fnv, dsHead, dsLast, samples.data(), vec.data(), p->hot ? hotHead : nullptr, p->hot ? hotLast : nullptr, hotFnv, line, sizeof(line)) != 0;
p->checked = true;
p->check = line;
}
@ -857,7 +958,7 @@ static Pair* buildPair(Ctx& c, const std::string& dir, CUstream s, std::string&
}
static std::string pairSummary(const Pair* p) {
return fmt("nvrtc %.0f cache %.0f dataset %.0f check %.0f race %.0f ms variant %s; %s", p->compileMs, p->cacheMs, p->dsMs, p->checkMs, p->raceMs, p->variant.c_str(), p->check.c_str());
return fmt("nvrtc %.0f cache %.0f dataset %.0f hot %.0f check %.0f race %.0f ms variant %s; %s", p->compileMs, p->cacheMs, p->dsMs, p->hotMs, p->checkMs, p->raceMs, p->variant.c_str(), p->check.c_str());
}
// ---------------------------------------------------------------------------------------------
@ -865,6 +966,7 @@ static std::string pairSummary(const Pair* p) {
struct PrepareTask {
std::string epochHex, dayHex, dir, error;
std::string wantClass, wantEra; // the class and era the prepare line named (empty: any)
std::atomic<bool> done{false};
Pair* result = nullptr;
double t0 = 0;
@ -882,6 +984,11 @@ static void prepareRun(Ctx* c, PrepareTask* t) {
err = "the pack in " + t->dir + " is for epoch " + p->epochHex.substr(0, 16) + " day " + p->dayHex + ", not the prepared seeds";
releasePair(*c, p); p = nullptr;
}
char why[256];
if (p && !pf_pack_class_ok(p->programClass.c_str(), p->eraHex.c_str(), t->wantClass.c_str(), t->wantEra.c_str(), why, sizeof(why))) {
err = "pack " + t->dir + ": " + why;
releasePair(*c, p); p = nullptr;
}
t->result = p;
t->error = err;
t->done = true;
@ -892,6 +999,8 @@ static void prepareRun(Ctx* c, PrepareTask* t) {
struct Options {
bool serve = false, check = false, raceOnly = false;
bool bench = false, memprobe = false; // read-width experiment (5 October 2026)
int batches = 5, warps = 0, probeMib = 0;
int device = 0, batchLog2 = 22, blockWarps = 1;
std::string pack, arch = "auto";
std::string race = "on", pinned, tuningPath;
@ -906,6 +1015,12 @@ static void usage() {
" --batch-log2 B nonces per dispatch = 2^B (default 22)\n"
" --block-warps W warps per thread block (default 1)\n"
" --arch sm_XY|compute_XY|auto NVRTC target (default auto: the device's architecture)\n"
" --bench --pack <dir> read-width experiment: build and self-test the pack, time --batches dispatches of 2^B nonces,\n"
" print the 2^B fingerprint at base nonce 0 (one RESULT line); a variant-5 pack runs --warps persistent warps\n"
" --memprobe [--probe-mib N] no pack: dependent random 4, 16 and 64-byte reads, independent reads, a coalesced stream and an\n"
" integer chain at 4, 64 and 1024 MiB (the same table as igneum-worker-opencl --memprobe)\n"
" --batches N --bench: timed dispatches (default 5)\n"
" --warps N --bench on a variant-5 pack: persistent warps (default: the occupancy capacity, rounded down to a power of two)\n"
" --race --pack <dir> the variant race alone (3 rounds): one line per variant, the race line, exit 0 or 1\n"
" --race on|off|a,b,c in --serve: race every variant (default), none, or these names\n"
" --race-bench-ms N timed window per variant (default 2000)\n"
@ -922,6 +1037,11 @@ static Options parseArgs(int argc, char** argv) {
auto next = [&]() -> std::string { if (i + 1 >= argc) { usage(); std::exit(2); } return argv[++i]; };
if (a == "--serve") o.serve = true;
else if (a == "--check") o.check = true;
else if (a == "--bench") o.bench = true;
else if (a == "--memprobe") o.memprobe = true;
else if (a == "--batches") o.batches = std::atoi(next().c_str());
else if (a == "--warps") o.warps = std::atoi(next().c_str());
else if (a == "--probe-mib") o.probeMib = std::atoi(next().c_str());
else if (a == "--race" && (i + 1 >= argc || std::string(argv[i + 1]).rfind("--", 0) == 0)) o.raceOnly = true;
else if (a == "--race") o.race = next();
else if (a == "--race-bench-ms") o.raceBenchMs = std::atoi(next().c_str());
@ -940,12 +1060,12 @@ static Options parseArgs(int argc, char** argv) {
}
if (o.batchLog2 < 10 || o.batchLog2 > 28) { std::printf("--batch-log2 must be between 10 and 28\n"); std::exit(2); }
if (o.blockWarps < 1 || o.blockWarps > 32) { std::printf("--block-warps must be between 1 and 32\n"); std::exit(2); }
if (!o.serve && !o.check && !o.raceOnly) { usage(); std::exit(2); }
if (!o.serve && !o.check && !o.raceOnly && !o.bench && !o.memprobe) { usage(); std::exit(2); }
if (o.raceBenchMs < 200 || o.raceBenchMs > 20000) { std::printf("--race-bench-ms must be between 200 and 20000\n"); std::exit(2); }
if (o.raceBudgetS < 5 || o.raceBudgetS > 540) { std::printf("--race-budget-s must be between 5 and 540 (the prepare lead is 600 DAA)\n"); std::exit(2); }
if (o.raceRounds == 0) o.raceRounds = o.raceOnly ? 3 : 1;
if (o.tuningPath.empty()) if (const char* t = std::getenv("IGNEUM_TUNING_FILE")) o.tuningPath = t;
if (o.pack.empty()) { std::printf("--pack <dir> is required (igneum-miner export-pack <node> <dir> writes one)\n"); std::exit(2); }
if (o.pack.empty() && !o.memprobe) { std::printf("--pack <dir> is required (igneum-miner export-pack <node> <dir> writes one)\n"); std::exit(2); }
while (o.pack.size() > 1 && (o.pack.back() == '/' || o.pack.back() == '\\')) o.pack.pop_back();
return o;
}
@ -1023,6 +1143,8 @@ static int runServe(Ctx& c, const Options& o, Pair* cur) {
}
delete task; task = nullptr;
}
std::string wantClass, wantEra;
takeClassTokens(f, wantClass, wantEra);
if (f[0] == "prepare") {
if (f.size() < 4) { emit(fmt("prepare-failed %s %s a pack directory is needed as the third field (igneum-miner --prepare-packs <dir>)", f.size() > 1 ? f[1].c_str() : "0", f.size() > 2 ? f[2].c_str() : "0")); continue; }
if (f[1].size() != 64) { emit(fmt("prepare-failed %s %s bad field (epoch_seed 64 hex, day_seed hex)", f[1].c_str(), f[2].c_str())); continue; }
@ -1034,6 +1156,7 @@ static int runServe(Ctx& c, const Options& o, Pair* cur) {
prepareRoot = parentDir(dir);
task = new PrepareTask();
task->epochHex = f[1]; task->dayHex = f[2]; task->dir = dir; task->t0 = wallMs();
task->wantClass = wantClass; task->wantEra = wantEra;
task->thread = std::thread(prepareRun, &c, task);
info(fmt("prepare started for epoch %.16s day %s from %s (NVRTC %s in the background)", f[1].c_str(), f[2].c_str(), dir.c_str(), c.archOpt.c_str()));
continue;
@ -1055,7 +1178,7 @@ static int runServe(Ctx& c, const Options& o, Pair* cur) {
pf_seed_words_from_bytes(daySeed, dl, kw);
double t0 = wallMs();
bool switched = false;
if (!pairIs(cur, f[6], f[7]) && !task && !pairIs(prepared, f[6], f[7])) {
if (!pairIsClass(cur, f[6], f[7], wantClass, wantEra) && !task && !pairIsClass(prepared, f[6], f[7], wantClass, wantEra)) {
// Self-heal: a job on seeds this worker has no pair for and no prepare in flight (a prepare failed, or
// the miner never sent one). The miner writes a pack per pair under its --prepare-packs root; find it by
// seeds.txt and build it now, in the foreground. The miner only re-sends prepare for the pair after this one.
@ -1071,11 +1194,18 @@ static int runServe(Ctx& c, const Options& o, Pair* cur) {
else emit("error " + jobId + " could not build " + dir + ": " + berr);
}
}
if (!pairIs(cur, f[6], f[7])) {
if (pairIs(prepared, f[6], f[7])) {
if (!pairIsClass(cur, f[6], f[7], wantClass, wantEra)) {
char why[256];
if (pairIsClass(prepared, f[6], f[7], wantClass, wantEra)) {
if (old) releasePair(c, old);
old = cur; cur = prepared; prepared = nullptr; switched = true;
info(fmt("switched to the prepared pair epoch %.16s day %s in %.2f ms", cur->epochHex.c_str(), cur->dayHex.c_str(), wallMs() - t0));
info(fmt("switched to the prepared pair epoch %.16s day %s (class %s) in %.2f ms", cur->epochHex.c_str(), cur->dayHex.c_str(), cur->programClass.c_str(), wallMs() - t0));
} else if (pairIs(cur, f[6], f[7]) && !pf_pack_class_ok(cur->programClass.c_str(), cur->eraHex.c_str(), wantClass.c_str(), wantEra.c_str(), why, sizeof(why))) {
// Counter ASIC 2.0: right seeds, wrong class or era. The miner prepares the pair again from a pack of
// the class the chain is on (the `need` line), and the pack of the other class is never mined.
emit(fmt("need %s %s", f[6].c_str(), f[7].c_str()));
emit(fmt("error %s pack %s: %s", jobId.c_str(), cur->dir.c_str(), why));
continue;
} else if (!hexEq(cur->epochHex, f[6])) {
emit(fmt("need %s %s", f[6].c_str(), f[7].c_str())); // the miner prepares this pair (4 October 2026)
emit(fmt("error %s epoch seed mismatch: this worker holds epoch %.16s (program words %08x %08x ...)%s, the job is for epoch %.16s (bare seed words %08x %08x ...); send prepare with a pack directory",
@ -1142,11 +1272,201 @@ static int runServe(Ctx& c, const Options& o, Pair* cur) {
// ---------------------------------------------------------------------------------------------
// Main
// ---------------------------------------------------------------------------------------------
// Read-width experiment (5 October 2026, docs/plans/read-width.md): --bench and --memprobe
static uint64_t fnv1a64Bytes(const void* p, size_t n) {
const uint8_t* b = (const uint8_t*)p;
uint64_t h = 0xcbf29ce484222325ull;
for (size_t i = 0; i < n; ++i) { h ^= b[i]; h *= 0x100000001b3ull; }
return h;
}
// --bench: the pair is built and self-tested (vectors through the bound kernel with the seed words); then a warm-up
// dispatch at base nonce 0 (fingerprinted) and --batches timed dispatches of 2^B nonces, wall time around
// cuStreamSynchronize (the driver API path loads no event symbols; a 2^24 dispatch is 60 to 900 ms on the cards here,
// so the launch overhead is under 1 percent).
static int runBench(Ctx& c, const Options& o, Pair* p) {
uint32_t nonces = 1u << o.batchLog2, block = 32u * (uint32_t)c.blockWarps;
if (p->persistent) { uint32_t unit = 32u * (uint32_t)p->warps; nonces = (nonces / unit) * unit; if (nonces == 0) nonces = unit; }
CUdeviceptr dOut = 0;
std::string err;
if (c.drv.memAlloc(&dOut, (size_t)nonces * 8u) != CUDA_SUCCESS) { std::printf("FAIL: cuMemAlloc out\n"); return 2; }
std::vector<uint64_t> hOut(nonces);
double sum = 0, warm = 0;
uint64_t fp = 0;
for (int b = -1; b < o.batches; ++b) {
double t0 = wallMs();
if (!launchHash(c, p, dOut, (uint32_t)(b + 1) * nonces, p->sw, nonces, block, nullptr, err)) { std::printf("FAIL: %s\n", err.c_str()); return 2; }
CUresult r = c.drv.streamSynchronize(nullptr);
if (r != CUDA_SUCCESS) { std::printf("FAIL: dispatch %d: %s\n", b, c.err(r).c_str()); return 2; }
double ms = wallMs() - t0;
if (b < 0) {
warm = ms;
if (c.drv.memcpyDtoH(hOut.data(), dOut, (size_t)nonces * 8u) != CUDA_SUCCESS) { std::printf("FAIL: read-back\n"); return 2; }
fp = fnv1a64Bytes(hOut.data(), (size_t)nonces * 8u);
} else sum += ms;
}
c.drv.memFree(dOut);
std::string dev = c.name; for (char& ch : dev) if (ch == ' ') ch = '_';
std::printf("warm-up dispatch (base 0): %.2f ms; %d timed dispatches of %u nonces: mean %.2f ms\n", warm, o.batches, nonces, sum / o.batches);
std::printf("RESULT pack=%s class=%s device=%s arch=%s regs=%d blocks_per_sm=%d warps=%d resident=%d arena_mib=%llu hot_mib=%u hot_slots=%u hot_fill_ms=%.2f nonces=%u batches=%d check=%s fingerprint=%016llx mhs=%.3f loads=%u bytes=%u scratch_ops=%u time=wall\n",
p->dir.c_str(), p->loadClass.c_str(), dev.c_str(), c.archOpt.c_str(), p->regs, p->blocksPerSM, p->warps, p->residentWarps, (unsigned long long)(p->scratchBytes >> 20), p->hotMb, p->hotSlots, p->hotMs, nonces, o.batches,
p->checked ? (p->checkPass ? "PASS" : "FAIL") : "skipped", (unsigned long long)fp, (double)nonces * (double)o.batches / (sum / 1000.0) / 1e6,
p->loadsPerHash, p->bytesPerHash, p->scratchOps * 8u);
return 0;
}
// --memprobe: the OpenCL worker's table (proto-opencl/host.c, 5 October 2026) in CUDA C through NVRTC, so the two
// vendors are probed with the same access patterns: a dependent chain of random 4-byte reads (the hash's pattern),
// eight independent chains per lane, dependent random 16-byte and 64-byte reads, a coalesced stream and an integer
// chain, at 4, 64 and 1024 MiB. Wall time around cuStreamSynchronize, best of 3, a fresh seed per repetition.
static const char* PROBE_CUDA =
"#include <cstdint>\n"
"__device__ __forceinline__ uint32_t pm_mix(uint32_t x) { x ^= x >> 16; x *= 0x7feb352du; x ^= x >> 15; x *= 0x846ca68bu; x ^= x >> 16; return x; }\n"
"extern \"C\" __global__ void probe_fill(uint32_t* ds, uint32_t n) { uint32_t i = blockIdx.x * blockDim.x + threadIdx.x; if (i < n) ds[i] = pm_mix(i ^ 0x9E3779B9u); }\n"
"extern \"C\" __global__ void probe_chase(const uint32_t* ds, uint32_t mask, uint32_t steps, uint32_t seed, uint32_t* out) {\n"
" uint32_t g = blockIdx.x * blockDim.x + threadIdx.x; uint32_t x = pm_mix(g ^ seed);\n"
" for (uint32_t s = 0u; s < steps; ++s) x = ds[x & mask] ^ (x * 0x9E3779B1u + s);\n"
" out[g] = x;\n"
"}\n"
"extern \"C\" __global__ void probe_indep(const uint32_t* ds, uint32_t mask, uint32_t steps, uint32_t seed, uint32_t* out) {\n"
" uint32_t g = blockIdx.x * blockDim.x + threadIdx.x;\n"
" uint32_t x0 = pm_mix(g * 8u ^ seed), x1 = pm_mix((g * 8u + 1u) ^ seed), x2 = pm_mix((g * 8u + 2u) ^ seed), x3 = pm_mix((g * 8u + 3u) ^ seed);\n"
" uint32_t x4 = pm_mix((g * 8u + 4u) ^ seed), x5 = pm_mix((g * 8u + 5u) ^ seed), x6 = pm_mix((g * 8u + 6u) ^ seed), x7 = pm_mix((g * 8u + 7u) ^ seed);\n"
" for (uint32_t s = 0u; s < steps; ++s) {\n"
" x0 = ds[x0 & mask] ^ (x0 * 0x9E3779B1u + s); x1 = ds[x1 & mask] ^ (x1 * 0x9E3779B1u + s);\n"
" x2 = ds[x2 & mask] ^ (x2 * 0x9E3779B1u + s); x3 = ds[x3 & mask] ^ (x3 * 0x9E3779B1u + s);\n"
" x4 = ds[x4 & mask] ^ (x4 * 0x9E3779B1u + s); x5 = ds[x5 & mask] ^ (x5 * 0x9E3779B1u + s);\n"
" x6 = ds[x6 & mask] ^ (x6 * 0x9E3779B1u + s); x7 = ds[x7 & mask] ^ (x7 * 0x9E3779B1u + s);\n"
" }\n"
" out[g] = x0 ^ x1 ^ x2 ^ x3 ^ x4 ^ x5 ^ x6 ^ x7;\n"
"}\n"
"extern \"C\" __global__ void probe_line16(const uint4* ds, uint32_t vecMask, uint32_t steps, uint32_t seed, uint32_t* out) {\n"
" uint32_t g = blockIdx.x * blockDim.x + threadIdx.x; uint32_t x = pm_mix(g ^ seed);\n"
" for (uint32_t s = 0u; s < steps; ++s) { uint4 a = ds[x & vecMask]; x = (a.x ^ a.y ^ a.z ^ a.w) ^ (x * 0x9E3779B1u + s); }\n"
" out[g] = x;\n"
"}\n"
"extern \"C\" __global__ void probe_line(const uint4* ds, uint32_t lineMask, uint32_t steps, uint32_t seed, uint32_t* out) {\n"
" uint32_t g = blockIdx.x * blockDim.x + threadIdx.x; uint32_t x = pm_mix(g ^ seed);\n"
" for (uint32_t s = 0u; s < steps; ++s) { uint32_t l = (x & lineMask) * 4u; uint4 a = ds[l], b = ds[l + 1u], c = ds[l + 2u], d = ds[l + 3u]; x = (a.x ^ b.y ^ c.z ^ d.w) ^ (x * 0x9E3779B1u + s); }\n"
" out[g] = x;\n"
"}\n"
"extern \"C\" __global__ void probe_stream(const uint4* ds, uint32_t perLane, uint32_t* out) {\n"
" uint32_t g = blockIdx.x * blockDim.x + threadIdx.x, n = gridDim.x * blockDim.x; uint4 acc = make_uint4(0u, 0u, 0u, 0u);\n"
" for (uint32_t s = 0u; s < perLane; ++s) { uint4 v = ds[s * n + g]; acc.x ^= v.x; acc.y ^= v.y; acc.z ^= v.z; acc.w ^= v.w; }\n"
" out[g] = acc.x ^ acc.y ^ acc.z ^ acc.w;\n"
"}\n"
"extern \"C\" __global__ void probe_alu(uint32_t steps, uint32_t seed, uint32_t* out) {\n"
" uint32_t g = blockIdx.x * blockDim.x + threadIdx.x; uint32_t x = pm_mix(g ^ seed), y = x ^ 0x5bd1e995u;\n"
" for (uint32_t s = 0u; s < steps; ++s) { x = x * 0x9E3779B1u + ((y << 7u) | (y >> 25u)); y = (y ^ x) + s; }\n"
" out[g] = x ^ y;\n"
"}\n";
static double probeLaunch(Ctx& c, CUfunction f, size_t lanes, size_t local, int reps, int seedArg, uint32_t seed, void** args) {
double best = -1;
for (int r = 0; r < reps; ++r) {
uint32_t s = seed + (uint32_t)r * 0x9E3779B9u;
if (seedArg >= 0) args[seedArg] = &s;
double t0 = wallMs();
if (c.drv.launchKernel(f, (unsigned)(lanes / local), 1, 1, (unsigned)local, 1, 1, 0, nullptr, args, nullptr) != CUDA_SUCCESS) return -1;
if (c.drv.streamSynchronize(nullptr) != CUDA_SUCCESS) return -1;
double ms = wallMs() - t0;
if (best < 0 || ms < best) best = ms;
}
return best;
}
static int runMemprobe(Ctx& c, const Options& o) {
Compiled cp;
std::string err;
if (!rtcCompile(c, PROBE_CUDA, "probe.cu", "", "", {}, cp, err)) { std::printf("memprobe: build FAILED: %s\n", err.c_str()); return 2; }
CUmodule mod = nullptr;
if (c.drv.moduleLoadData(&mod, cp.image.data()) != CUDA_SUCCESS) { std::printf("memprobe: cuModuleLoadData failed\n"); return 2; }
CUfunction kFill, kChase, kIndep, kLine16, kLine, kStream, kAlu;
const char* names[7] = { "probe_fill", "probe_chase", "probe_indep", "probe_line16", "probe_line", "probe_stream", "probe_alu" };
CUfunction* fns[7] = { &kFill, &kChase, &kIndep, &kLine16, &kLine, &kStream, &kAlu };
for (int i = 0; i < 7; ++i) if (c.drv.moduleGetFunction(fns[i], mod, names[i]) != CUDA_SUCCESS) { std::printf("memprobe: %s not in the module\n", names[i]); return 2; }
int sizes[3] = { 4, 64, 1024 }, nSizes = 3;
if (o.probeMib > 0) { sizes[0] = o.probeMib; nSizes = 1; }
const size_t lanesList[8] = { 256, 1024, 1u << 12, 1u << 14, 1u << 16, 1u << 18, 1u << 20, 1u << 22 };
const size_t groups[2] = { 32, 256 };
const uint32_t STEPS = 256u, ALU_STEPS = 4096u;
const size_t maxLanes = 1u << 22;
CUdeviceptr dOut = 0;
if (c.drv.memAlloc(&dOut, maxLanes * 4u) != CUDA_SUCCESS) { std::printf("memprobe: cuMemAlloc out\n"); return 2; }
std::printf("memprobe on %s (sm_%d%d, %d SMs, driver %d.%d, NVRTC %d.%d), wall time around cuStreamSynchronize\n", c.name.c_str(), c.major, c.minor, c.sms, c.driverVersion / 1000, (c.driverVersion % 100) / 10, c.rtcMajor, c.rtcMinor);
std::printf("| probe | MiB | block | lanes in flight | steps per lane | best ms | G loads/s | ns per dependent load |\n|---|---|---|---|---|---|---|---|\n");
for (int si = 0; si < nSizes; ++si) {
int mib = sizes[si];
uint64_t bytes = (uint64_t)mib << 20;
uint32_t words = (uint32_t)(bytes / 4ull), mask = words - 1u, n = words;
CUdeviceptr dDs = 0;
if (c.drv.memAlloc(&dDs, (size_t)bytes) != CUDA_SUCCESS) { std::printf("| chase | %d | skipped: cuMemAlloc failed | | | | | |\n", mib); continue; }
{ void* a[2] = { &dDs, &n }; probeLaunch(c, kFill, ((size_t)words + 255) / 256 * 256, 256, 1, -1, 0, a); }
for (int gi = 0; gi < 2; ++gi) {
size_t local = groups[gi];
for (int li = 0; li < 8; ++li) {
size_t lanes = lanesList[li];
if (lanes < local) continue;
uint32_t seed = 0x1234567u + (uint32_t)li * 977u, steps = STEPS;
void* a[5] = { &dDs, &mask, &steps, &seed, &dOut };
double ms = probeLaunch(c, kChase, lanes, local, 3, 3, seed, a);
std::printf("| chase | %d | %zu | %zu | %u | %.3f | %.3f | %.0f |\n", mib, local, lanes, STEPS, ms, (double)lanes * STEPS / (ms / 1000.0) / 1e9, ms * 1e6 / STEPS);
std::fflush(stdout);
}
}
for (size_t lanes = 1u << 16; lanes <= maxLanes; lanes <<= 2) {
uint32_t seed = 0x7654321u, steps = STEPS;
void* a[5] = { &dDs, &mask, &steps, &seed, &dOut };
double ms = probeLaunch(c, kIndep, lanes, 256, 3, 3, seed, a);
std::printf("| indep x8 | %d | 256 | %zu | %u | %.3f | %.3f | (8 loads in flight per lane) |\n", mib, lanes, STEPS, ms, (double)lanes * 8.0 * STEPS / (ms / 1000.0) / 1e9);
}
for (size_t lanes = 1u << 14; lanes <= maxLanes; lanes <<= 2) {
uint32_t vecMask = (words / 4u) - 1u, seed = 0x2718281u, steps = STEPS;
void* a[5] = { &dDs, &vecMask, &steps, &seed, &dOut };
double ms = probeLaunch(c, kLine16, lanes, 256, 3, 3, seed, a);
std::printf("| line 16 B | %d | 256 | %zu | %u | %.3f | %.3f G reads/s | %.1f GB/s in 16 B reads |\n", mib, lanes, STEPS, ms, (double)lanes * STEPS / (ms / 1000.0) / 1e9, (double)lanes * STEPS * 16.0 / (ms / 1000.0) / 1e9);
}
for (size_t lanes = 1u << 14; lanes <= maxLanes; lanes <<= 2) {
uint32_t lineMask = (words / 16u) - 1u, seed = 0x3141592u, steps = STEPS;
void* a[5] = { &dDs, &lineMask, &steps, &seed, &dOut };
double ms = probeLaunch(c, kLine, lanes, 256, 3, 3, seed, a);
std::printf("| line 64 B | %d | 256 | %zu | %u | %.3f | %.3f G lines/s | %.1f GB/s in lines |\n", mib, lanes, STEPS, ms, (double)lanes * STEPS / (ms / 1000.0) / 1e9, (double)lanes * STEPS * 64.0 / (ms / 1000.0) / 1e9);
}
{
size_t lanes = 1u << 20;
uint32_t perLane = (uint32_t)((uint64_t)words / 4ull / (uint64_t)lanes);
if (perLane == 0) { perLane = 1; lanes = (size_t)words / 4u; }
double bytesRead = (double)perLane * (double)lanes * 16.0;
void* a[3] = { &dDs, &perLane, &dOut };
double ms = probeLaunch(c, kStream, lanes, 256, 3, -1, 0, a);
std::printf("| stream | %d | 256 | %zu | %u | %.3f | %.1f GB/s coalesced | (%.0f MiB read once) |\n", mib, lanes, perLane, ms, bytesRead / (ms / 1000.0) / 1e9, bytesRead / 1048576.0);
}
c.drv.memFree(dDs);
std::fflush(stdout);
}
{
size_t lanes = 1u << 20;
uint32_t seed = 0x2468aceu, steps = ALU_STEPS;
void* a[3] = { &steps, &seed, &dOut };
double ms = probeLaunch(c, kAlu, lanes, 256, 3, 1, seed, a);
double ops = (double)lanes * ALU_STEPS * 5.0;
std::printf("| alu | 0 | 256 | %zu | %u | %.3f | %.1f G int ops/s | %.3f G steps/s per SM (approximate: 5 ops per step counted) |\n", lanes, ALU_STEPS, ms, ops / (ms / 1000.0) / 1e9, (double)lanes * ALU_STEPS / (ms / 1000.0) / 1e9 / (c.sms ? c.sms : 1));
}
c.drv.memFree(dOut);
c.drv.moduleUnload(mod);
std::printf("memprobe: done\n");
return 0;
}
int main(int argc, char** argv) {
Options o = parseArgs(argc, argv);
Ctx c;
c.blockWarps = o.blockWarps;
c.race = o.race; c.raceBenchMs = o.raceBenchMs; c.raceBudgetS = o.raceBudgetS; c.raceRounds = o.raceRounds; c.batchLog2 = o.batchLog2; c.pinned = o.pinned;
c.warps = o.warps; c.batches = o.batches;
if (o.bench || o.memprobe) c.race = "off";
if (!o.tuningPath.empty()) { bool ok = false; c.tuning = readText(o.tuningPath, ok); if (!ok) c.tuning.clear(); }
std::string err, drvLib, rtcLib;
if (!loadDriver(c.drv, err, drvLib)) { emit("error 0 " + err); return 2; }
@ -1155,9 +1475,17 @@ int main(int argc, char** argv) {
info(fmt("igneum-worker-cuda %s: device %d %s (sm_%d%d, %d SMs), driver %d.%d from %s, NVRTC %d.%d from %s, target %s (%s)",
WORKER_VERSION, o.device, c.name.c_str(), c.major, c.minor, c.sms, c.driverVersion / 1000, (c.driverVersion % 100) / 10, drvLib.c_str(), c.rtcMajor, c.rtcMinor, rtcLib.c_str(), c.archOpt.c_str(), c.why.c_str()));
if (!c.tuning.empty()) info(fmt("tuning file %s (%zu bytes): %s", o.tuningPath.c_str(), c.tuning.size(), readTuning(c.tuning, c.name).found ? "has an entry for this card" : "no entry for this card"));
if (o.memprobe) { int rc = runMemprobe(c, o); c.drv.primaryCtxRelease(c.dev); return rc; }
double t0 = wallMs();
Pair* cur = buildPair(c, o.pack, nullptr, err, !o.check);
Pair* cur = buildPair(c, o.pack, nullptr, err, !o.check && !o.bench);
if (!cur) { emit("error 0 " + err); return 1; }
if (o.bench) {
std::printf("pack %s on %s: %s\n", o.pack.c_str(), c.name.c_str(), pairSummary(cur).c_str());
int rc = runBench(c, o, cur);
releasePair(c, cur);
c.drv.primaryCtxRelease(c.dev);
return rc;
}
if (o.raceOnly) {
std::printf("race %s on %s (%s, %d SMs, driver %d.%d, NVRTC %d.%d, %s): %s\n", o.pack.c_str(), c.name.c_str(), c.archOpt.c_str(), c.sms, c.driverVersion / 1000, (c.driverVersion % 100) / 10, c.rtcMajor, c.rtcMinor, c.why.c_str(), pairSummary(cur).c_str());
std::printf("%s\n", cur->raceLine.c_str());

View file

@ -0,0 +1,281 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal).
// Built from source at runtime by proto-opencl/host.c, which passes these defines:
// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit)
// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default)
// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32
// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition
// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units;
// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32.
#ifndef IGNEUM_GROUP
#define IGNEUM_GROUP 32
#endif
#ifndef IGNEUM_EXCHANGE
#define IGNEUM_EXCHANGE 0
#endif
#ifdef __OPENCL_VERSION__
#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1)))
#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n]
#if IGNEUM_EXCHANGE == 1
#ifdef cl_khr_subgroups
#pragma OPENCL EXTENSION cl_khr_subgroups : enable
#endif
#ifdef cl_khr_subgroup_shuffle
#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable
#endif
#elif IGNEUM_EXCHANGE == 2
#pragma OPENCL EXTENSION cl_intel_subgroups : enable
#endif
#else
// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros.
#include "emu_opencl.h"
#endif
#if IGNEUM_EXCHANGE == 1
#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m))
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
#elif IGNEUM_EXCHANGE == 2
#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m))
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
#else
// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per
// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane
// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's
// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier.
#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; }
#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; }
#endif
static inline uint splitmix32(uint x) {
x ^= x >> 16; x *= 0x7feb352du;
x ^= x >> 15; x *= 0x846ca68bu;
x ^= x >> 16;
return x;
}
// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32.
static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); }
// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x.
static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); }
static inline uint ds_elem(uint i, uint d0, uint d1) {
uint x = i ^ d0;
x *= 0x9E3779B1u; x ^= x >> 15;
x += d1;
x *= 0x85EBCA77u; x ^= x >> 13;
x *= 0xC2B2AE3Du; x ^= x >> 16;
return x;
}
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
#define MH_CACHE_LINE_MASK 0x003fffffu
#define MH_SEGMENT_LINES 64u
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
// y = ChaCha12 core(x) + x
static inline void mh_chacha_block(const uint* x, uint* y) {
for (uint i = 0u; i < 16u; ++i) y[i] = x[i];
for (uint r = 0u; r < 6u; ++r) {
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
}
for (uint i = 0u; i < 16u; ++i) y[i] += x[i];
}
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
static inline void mh_cache_segment(__global uint* cache, uint seg) {
uint prev[16]; uint x[16]; uint y[16];
for (uint i = 0u; i < 16u; ++i) prev[i] = 0u;
for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) {
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
x[4] = 0xceed56d7u ^ prev[4];
x[5] = 0x9ba270d2u ^ prev[5];
x[6] = 0x82caab2du ^ prev[6];
x[7] = 0x81ebce0eu ^ prev[7];
x[8] = 0x12b6ecf1u ^ prev[8];
x[9] = 0xd0f3fd7cu ^ prev[9];
x[10] = 0xd872eefeu ^ prev[10];
x[11] = 0xc158c7bdu ^ prev[11];
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
mh_chacha_block(x, y);
__global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
}
}
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
static inline void mh_mixer(uint* s, uint rk) {
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
}
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
static inline void mh_item(__global const uint* cache, uint t, uint* s) {
s[0] = 0xceed56d7u;
s[1] = 0x9ba270d2u;
s[2] = 0x82caab2du;
s[3] = 0x81ebce0eu;
s[4] = 0x12b6ecf1u;
s[5] = 0xd0f3fd7cu;
s[6] = 0xd872eefeu;
s[7] = 0xc158c7bdu;
s[8] = t * 0xf351d601u + 0xc6892460u;
s[9] = t * 0xa3bb398fu + 0x25b7228au;
s[10] = t * 0xb5a09e35u + 0xcd515004u;
s[11] = t * 0x7509c9c1u + 0x2846527au;
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
s[14] = t * 0xded91851u + 0x82961bacu;
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
for (uint r = 0u; r < 8u; ++r) {
mh_mixer(s, 0x9E3779B9u * (r + 1u));
__global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i];
}
mh_mixer(s, 0x9E3779B9u * 9u);
}
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 0 2 10 15 of w.
static inline uint mh_j(uint w) { return ((w >> 0u) & 1u) | (((w >> 2u) & 1u) << 1) | (((w >> 10u) & 1u) << 2) | (((w >> 15u) & 1u) << 3); }
static inline uint mh_t(uint w) { w = (w & 0x00007fffu) | ((w >> 16u) << 15u); w = (w & 0x000003ffu) | ((w >> 11u) << 10u); w = (w & 0x00000003u) | ((w >> 3u) << 2u); w = (w & 0x00000000u) | ((w >> 1u) << 0u); return w; }
static inline uint mh_addr(uint t, uint j) { uint w = t; w = ((w >> 0u) << 1u) | (w & 0x00000000u) | (((j >> 0u) & 1u) << 0u); w = ((w >> 2u) << 3u) | (w & 0x00000003u) | (((j >> 1u) & 1u) << 2u); w = ((w >> 10u) << 11u) | (w & 0x000003ffu) | (((j >> 2u) & 1u) << 10u); w = ((w >> 15u) << 16u) | (w & 0x00007fffu) | (((j >> 3u) & 1u) << 15u); return w; }
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }
// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item.
// The same constants as memhard.h in this pack (one emitter, three dialects).
__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) {
uint seg = (uint)get_global_id(0);
if (seg < nSegments) mh_cache_segment(cache, seg);
}
__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) {
uint t = (uint)get_global_id(0);
if (t < nItems) {
uint s[16];
mh_item(cache, t, s);
for (uint i = 0u; i < 16u; ++i) ds[(ulong)mh_addr(t, i)] = s[i];
}
}
// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the
// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and
// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all).
IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) {
uint gid = (uint)get_global_id(0);
uint lid = (uint)get_local_id(0);
uint nonce = baseNonce + gid;
uint r0, r1, r2, r3, r4, r5, r6, r7;
#if IGNEUM_EXCHANGE == 0
IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
uint xk = 0u;
#else
(void)lid;
#endif
{ uint x = nonce ^ 0x667d0fbdu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x7b8e5963u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
{ uint x = nonce ^ 0x7b8e5963u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x31c67e5eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
{ uint x = nonce ^ 0x31c67e5eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4529ddc6u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
{ uint x = nonce ^ 0x4529ddc6u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef19d6d8u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
{ uint x = nonce ^ 0xef19d6d8u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xaccf6211u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
{ uint x = nonce ^ 0xaccf6211u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xda0aed32u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
{ uint x = nonce ^ 0xda0aed32u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xabc6df31u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
{ uint x = nonce ^ 0xabc6df31u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x667d0fbdu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
for (uint it = 0u; it < 8u; ++it) {
uint sel = r0;
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
r7 = r7 ^ r0; // 1 xor
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 1u); r3 = r3 ^ t_; } // 2 shfl
r4 = r4 - r1; // 3 sub
r2 = r0 * r4 + r2; // 4 mad
r4 = rotl_imm(r4, 9u); // 5 rotl
r0 = rotr_var(r0, r2); // 6 rotr
r6 = r6 ^ ds[((rotl_imm(r7 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
r1 = r1 ^ ds[((rotl_imm(r4 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
r1 = r1 ^ ds[((rotl_imm(r2 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
r7 = r7 ^ ds[((rotl_imm(r0 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
r7 = r7 ^ ds[((rotl_imm(r1 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
r5 = r5 * r4; // 12 mul
r7 = r7 ^ ds[((rotl_imm(r6 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 16u); r6 = r6 ^ t_; } // 14 shfl
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r3 = r3 ^ t_; } // 15 shfl
r2 = r2 | r7; // 16 or
r0 = rotr_var(r0, r6); // 17 rotr
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
r7 = rotl_imm(r7, 24u); // 19 rotl
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
r2 = r2 | r6; // 21 or
r6 = r5 * r4 + r6; // 22 mad
r1 = r1 | r0; // 23 or
r6 = r6 ^ r1; // 24 xor
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
r7 = rotr_var(r7, r0); // 26 rotr
r4 = r5 * r7 + r4; // 27 mad
r3 = r2 * r4 + r3; // 28 mad
r1 = r1 ^ ds[((rotl_imm(r4 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
r2 = r2 ^ ds[((rotl_imm(r3 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
r1 = r1 ^ ds[((rotl_imm(r5 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
r2 = mul_hi(r2, r5); // 35 mulhi
r4 = r4 ^ r2; // 36 xor
r6 = r6 * r5; // 37 mul
r7 = r7 ^ r0; // 38 xor
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
r2 = rotr_var(r2, r3); // 40 rotr
r7 = r7 - r0; // 41 sub
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 43 shfl
r0 = r0 ^ ds[((rotl_imm(r7 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
r3 = r3 ^ ds[((rotl_imm(r1 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
r6 = r6 ^ ds[((rotl_imm(r3 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
r4 = mul_hi(r4, r2); // 48 mulhi
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r0 = r0 ^ t_; } // 50 shfl
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
r5 = r5 ^ ds[((rotl_imm(r2 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
r6 = rotl_imm(r6, 12u); // 53 rotl
r3 = rotl_imm(r3, 12u); // 54 rotl
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
r1 = r1 ^ ds[((rotl_imm(r4 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
r5 = r5 ^ ds[((rotl_imm(r0 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
r3 = mul_hi(r3, r2); // 60 mulhi
r6 = r6 | r4; // 61 or
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r5 = r5 ^ t_; } // 62 shfl
r3 = r3 ^ ds[((rotl_imm(r6 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
}
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((ulong)hi << 32) | (ulong)lo;
}
#if IGNEUM_EXCHANGE != 0
// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the
// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a
// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md.
IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) {
if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); }
}
#endif

View file

@ -0,0 +1,163 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal).
// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC.
#include <cuda_runtime.h>
#include <cstdint>
#include "program.h"
#include "memhard.h"
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
x ^= x >> 16; x *= 0x7feb352du;
x ^= x >> 15; x *= 0x846ca68bu;
x ^= x >> 16;
return x;
}
// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31.
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x.
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) {
uint32_t x = i ^ d0;
x *= 0x9E3779B1u; x ^= x >> 15;
x += d1;
x *= 0x85EBCA77u; x ^= x >> 13;
x *= 0xC2B2AE3Du; x ^= x >> 16;
return x;
}
// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item.
// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host.
__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) {
uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x;
if (seg < nSegments) mh_cache_segment(cache, seg);
}
__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
uint32_t t = blockIdx.x * blockDim.x + threadIdx.x;
if (t < nItems) {
uint32_t s[16];
mh_item(cache, t, s);
for (uint32_t i = 0u; i < 16u; ++i) ds[(size_t)mh_addr(t, i)] = s[i];
}
}
// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every
// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a
// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid.
__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) {
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
uint32_t nonce = baseNonce + gid;
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
{ uint32_t x = nonce ^ 0x667d0fbdu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x7b8e5963u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
{ uint32_t x = nonce ^ 0x7b8e5963u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x31c67e5eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
{ uint32_t x = nonce ^ 0x31c67e5eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4529ddc6u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
{ uint32_t x = nonce ^ 0x4529ddc6u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef19d6d8u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
{ uint32_t x = nonce ^ 0xef19d6d8u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xaccf6211u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
{ uint32_t x = nonce ^ 0xaccf6211u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xda0aed32u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
{ uint32_t x = nonce ^ 0xda0aed32u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xabc6df31u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
{ uint32_t x = nonce ^ 0xabc6df31u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x667d0fbdu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
for (uint32_t it = 0u; it < 8u; ++it) {
uint32_t sel = r0;
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
r7 = r7 ^ r0; // 1 xor
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r6, 1); // 2 shfl
r4 = r4 - r1; // 3 sub
r2 = r0 * r4 + r2; // 4 mad
r4 = rotl_imm(r4, 9u); // 5 rotl
r0 = rotr_var(r0, r2); // 6 rotr
r6 = r6 ^ ds[((rotl_imm(r7 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
r1 = r1 ^ ds[((rotl_imm(r4 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
r1 = r1 ^ ds[((rotl_imm(r2 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
r7 = r7 ^ ds[((rotl_imm(r0 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
r7 = r7 ^ ds[((rotl_imm(r1 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
r5 = r5 * r4; // 12 mul
r7 = r7 ^ ds[((rotl_imm(r6 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r5, 16); // 14 shfl
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r5, 1); // 15 shfl
r2 = r2 | r7; // 16 or
r0 = rotr_var(r0, r6); // 17 rotr
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
r7 = rotl_imm(r7, 24u); // 19 rotl
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
r2 = r2 | r6; // 21 or
r6 = r5 * r4 + r6; // 22 mad
r1 = r1 | r0; // 23 or
r6 = r6 ^ r1; // 24 xor
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
r7 = rotr_var(r7, r0); // 26 rotr
r4 = r5 * r7 + r4; // 27 mad
r3 = r2 * r4 + r3; // 28 mad
r1 = r1 ^ ds[((rotl_imm(r4 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
r2 = r2 ^ ds[((rotl_imm(r3 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
r1 = r1 ^ ds[((rotl_imm(r5 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
r2 = __umulhi(r2, r5); // 35 mulhi
r4 = r4 ^ r2; // 36 xor
r6 = r6 * r5; // 37 mul
r7 = r7 ^ r0; // 38 xor
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
r2 = rotr_var(r2, r3); // 40 rotr
r7 = r7 - r0; // 41 sub
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 43 shfl
r0 = r0 ^ ds[((rotl_imm(r7 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
r3 = r3 ^ ds[((rotl_imm(r1 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
r6 = r6 ^ ds[((rotl_imm(r3 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
r4 = __umulhi(r4, r2); // 48 mulhi
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r7, 4); // 50 shfl
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
r5 = r5 ^ ds[((rotl_imm(r2 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
r6 = rotl_imm(r6, 12u); // 53 rotl
r3 = rotl_imm(r3, 12u); // 54 rotl
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
r1 = r1 ^ ds[((rotl_imm(r4 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
r5 = r5 ^ ds[((rotl_imm(r0 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
r3 = __umulhi(r3, r2); // 60 mulhi
r6 = r6 | r4; // 61 or
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 62 shfl
r3 = r3 ^ ds[((rotl_imm(r6 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
}
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
}
// Host-side launch wrappers. Declared in program.h, called from host.cu.
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) {
if (nSegments == 0u) return cudaErrorInvalidValue;
uint32_t block = 256u;
uint32_t grid = (nSegments + block - 1u) / block;
igneum_cache_fill<<<grid, block>>>(cache, nSegments);
return cudaGetLastError();
}
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
if (nItems == 0u) return cudaErrorInvalidValue;
uint32_t block = 256u;
uint32_t grid = (nItems + block - 1u) / block;
igneum_build<<<grid, block>>>(ds, cache, nItems);
return cudaGetLastError();
}
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
uint32_t nonces, uint32_t blockWarps) {
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
uint32_t block = 32u * blockWarps;
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
igneum_hash<<<nonces / block, block>>>(ds, out, baseNonce, mask);
return cudaGetLastError();
}
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
cudaFuncAttributes attr;
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash);
if (e != cudaSuccess) return e;
*numRegs = attr.numRegs;
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0);
}

View file

@ -0,0 +1,375 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal).
// Built from source at runtime by proto-opencl/host.c, which passes these defines:
// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit)
// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default)
// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32
// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition
// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units;
// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32.
#ifndef IGNEUM_GROUP
#define IGNEUM_GROUP 32
#endif
#ifndef IGNEUM_EXCHANGE
#define IGNEUM_EXCHANGE 0
#endif
#ifdef __OPENCL_VERSION__
#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1)))
#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n]
#if IGNEUM_EXCHANGE == 1
#ifdef cl_khr_subgroups
#pragma OPENCL EXTENSION cl_khr_subgroups : enable
#endif
#ifdef cl_khr_subgroup_shuffle
#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable
#endif
#elif IGNEUM_EXCHANGE == 2
#pragma OPENCL EXTENSION cl_intel_subgroups : enable
#endif
#else
// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros.
#include "emu_opencl.h"
#endif
#if IGNEUM_EXCHANGE == 1
#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m))
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
#elif IGNEUM_EXCHANGE == 2
#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m))
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
#else
// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per
// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane
// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's
// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier.
#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; }
#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; }
#endif
static inline uint splitmix32(uint x) {
x ^= x >> 16; x *= 0x7feb352du;
x ^= x >> 15; x *= 0x846ca68bu;
x ^= x >> 16;
return x;
}
// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32.
static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); }
// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x.
static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); }
static inline uint ds_elem(uint i, uint d0, uint d1) {
uint x = i ^ d0;
x *= 0x9E3779B1u; x ^= x >> 15;
x += d1;
x *= 0x85EBCA77u; x ^= x >> 13;
x *= 0xC2B2AE3Du; x ^= x >> 16;
return x;
}
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
#define MH_CACHE_LINE_MASK 0x003fffffu
#define MH_SEGMENT_LINES 64u
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
// y = ChaCha12 core(x) + x
static inline void mh_chacha_block(const uint* x, uint* y) {
for (uint i = 0u; i < 16u; ++i) y[i] = x[i];
for (uint r = 0u; r < 6u; ++r) {
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
}
for (uint i = 0u; i < 16u; ++i) y[i] += x[i];
}
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
static inline void mh_cache_segment(__global uint* cache, uint seg) {
uint prev[16]; uint x[16]; uint y[16];
for (uint i = 0u; i < 16u; ++i) prev[i] = 0u;
for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) {
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
x[4] = 0xceed56d7u ^ prev[4];
x[5] = 0x9ba270d2u ^ prev[5];
x[6] = 0x82caab2du ^ prev[6];
x[7] = 0x81ebce0eu ^ prev[7];
x[8] = 0x12b6ecf1u ^ prev[8];
x[9] = 0xd0f3fd7cu ^ prev[9];
x[10] = 0xd872eefeu ^ prev[10];
x[11] = 0xc158c7bdu ^ prev[11];
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
mh_chacha_block(x, y);
__global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
}
}
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
static inline void mh_mixer(uint* s, uint rk) {
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
}
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
static inline void mh_item(__global const uint* cache, uint t, uint* s) {
s[0] = 0xceed56d7u;
s[1] = 0x9ba270d2u;
s[2] = 0x82caab2du;
s[3] = 0x81ebce0eu;
s[4] = 0x12b6ecf1u;
s[5] = 0xd0f3fd7cu;
s[6] = 0xd872eefeu;
s[7] = 0xc158c7bdu;
s[8] = t * 0xf351d601u + 0xc6892460u;
s[9] = t * 0xa3bb398fu + 0x25b7228au;
s[10] = t * 0xb5a09e35u + 0xcd515004u;
s[11] = t * 0x7509c9c1u + 0x2846527au;
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
s[14] = t * 0xded91851u + 0x82961bacu;
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
for (uint r = 0u; r < 8u; ++r) {
mh_mixer(s, 0x9E3779B9u * (r + 1u));
__global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i];
}
mh_mixer(s, 0x9E3779B9u * 9u);
}
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 0 2 10 15 of w.
static inline uint mh_j(uint w) { return ((w >> 0u) & 1u) | (((w >> 2u) & 1u) << 1) | (((w >> 10u) & 1u) << 2) | (((w >> 15u) & 1u) << 3); }
static inline uint mh_t(uint w) { w = (w & 0x00007fffu) | ((w >> 16u) << 15u); w = (w & 0x000003ffu) | ((w >> 11u) << 10u); w = (w & 0x00000003u) | ((w >> 3u) << 2u); w = (w & 0x00000000u) | ((w >> 1u) << 0u); return w; }
static inline uint mh_addr(uint t, uint j) { uint w = t; w = ((w >> 0u) << 1u) | (w & 0x00000000u) | (((j >> 0u) & 1u) << 0u); w = ((w >> 2u) << 3u) | (w & 0x00000003u) | (((j >> 1u) & 1u) << 2u); w = ((w >> 10u) << 11u) | (w & 0x000003ffu) | (((j >> 2u) & 1u) << 10u); w = ((w >> 15u) << 16u) | (w & 0x00007fffu) | (((j >> 3u) & 1u) << 15u); return w; }
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }
// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item.
// The same constants as memhard.h in this pack (one emitter, three dialects).
__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) {
uint seg = (uint)get_global_id(0);
if (seg < nSegments) mh_cache_segment(cache, seg);
}
__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) {
uint t = (uint)get_global_id(0);
if (t < nItems) {
uint s[16];
mh_item(cache, t, s);
for (uint i = 0u; i < 16u; ++i) ds[(ulong)mh_addr(t, i)] = s[i];
}
}
// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the
// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and
// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all).
IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) {
uint gid = (uint)get_global_id(0);
uint lid = (uint)get_local_id(0);
uint nonce = baseNonce + gid;
uint r0, r1, r2, r3, r4, r5, r6, r7;
#if IGNEUM_EXCHANGE == 0
IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
uint xk = 0u;
#else
(void)lid;
#endif
{ uint x = nonce ^ 0x667d0fbdu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x7b8e5963u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
{ uint x = nonce ^ 0x7b8e5963u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x31c67e5eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
{ uint x = nonce ^ 0x31c67e5eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4529ddc6u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
{ uint x = nonce ^ 0x4529ddc6u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef19d6d8u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
{ uint x = nonce ^ 0xef19d6d8u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xaccf6211u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
{ uint x = nonce ^ 0xaccf6211u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xda0aed32u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
{ uint x = nonce ^ 0xda0aed32u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xabc6df31u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
{ uint x = nonce ^ 0xabc6df31u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x667d0fbdu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
for (uint it = 0u; it < 8u; ++it) {
uint sel = r0;
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
r7 = r7 ^ r0; // 1 xor
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 1u); r3 = r3 ^ t_; } // 2 shfl
r4 = r4 - r1; // 3 sub
r2 = r0 * r4 + r2; // 4 mad
r4 = rotl_imm(r4, 9u); // 5 rotl
r0 = rotr_var(r0, r2); // 6 rotr
r6 = r6 ^ ds[((rotl_imm(r7 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
r1 = r1 ^ ds[((rotl_imm(r4 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
r1 = r1 ^ ds[((rotl_imm(r2 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
r7 = r7 ^ ds[((rotl_imm(r0 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
r7 = r7 ^ ds[((rotl_imm(r1 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
r5 = r5 * r4; // 12 mul
r7 = r7 ^ ds[((rotl_imm(r6 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 16u); r6 = r6 ^ t_; } // 14 shfl
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r3 = r3 ^ t_; } // 15 shfl
r2 = r2 | r7; // 16 or
r0 = rotr_var(r0, r6); // 17 rotr
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
r7 = rotl_imm(r7, 24u); // 19 rotl
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
r2 = r2 | r6; // 21 or
r6 = r5 * r4 + r6; // 22 mad
r1 = r1 | r0; // 23 or
r6 = r6 ^ r1; // 24 xor
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
r7 = rotr_var(r7, r0); // 26 rotr
r4 = r5 * r7 + r4; // 27 mad
r3 = r2 * r4 + r3; // 28 mad
r1 = r1 ^ ds[((rotl_imm(r4 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
r2 = r2 ^ ds[((rotl_imm(r3 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
r1 = r1 ^ ds[((rotl_imm(r5 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
r2 = mul_hi(r2, r5); // 35 mulhi
r4 = r4 ^ r2; // 36 xor
r6 = r6 * r5; // 37 mul
r7 = r7 ^ r0; // 38 xor
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
r2 = rotr_var(r2, r3); // 40 rotr
r7 = r7 - r0; // 41 sub
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 43 shfl
r0 = r0 ^ ds[((rotl_imm(r7 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
r3 = r3 ^ ds[((rotl_imm(r1 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
r6 = r6 ^ ds[((rotl_imm(r3 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
r4 = mul_hi(r4, r2); // 48 mulhi
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r0 = r0 ^ t_; } // 50 shfl
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
r5 = r5 ^ ds[((rotl_imm(r2 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
r6 = rotl_imm(r6, 12u); // 53 rotl
r3 = rotl_imm(r3, 12u); // 54 rotl
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
r1 = r1 ^ ds[((rotl_imm(r4 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
r5 = r5 ^ ds[((rotl_imm(r0 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
r3 = mul_hi(r3, r2); // 60 mulhi
r6 = r6 | r4; // 61 or
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r5 = r5 ^ t_; } // 62 shfl
r3 = r3 ^ ds[((rotl_imm(r6 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
}
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((ulong)hi << 32) | (ulong)lo;
}
#if IGNEUM_EXCHANGE != 0
// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the
// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a
// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md.
IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) {
if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); }
}
#endif
// Header-bound variant (bind.rs): the init words come from initw, not SEEDW. Same body as igneum_hash.
IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw) {
uint gid = (uint)get_global_id(0);
uint lid = (uint)get_local_id(0);
uint nonce = baseNonce + gid;
uint r0, r1, r2, r3, r4, r5, r6, r7;
uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7];
#if IGNEUM_EXCHANGE == 0
IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
uint xk = 0u;
#else
(void)lid;
#endif
{ uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; }
{ uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; }
{ uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; }
{ uint x = nonce ^ iw3; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw4; }
{ uint x = nonce ^ iw4; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw5; }
{ uint x = nonce ^ iw5; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw6; }
{ uint x = nonce ^ iw6; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw7; }
{ uint x = nonce ^ iw7; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw0; }
for (uint it = 0u; it < 8u; ++it) {
uint sel = r0;
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
r7 = r7 ^ r0; // 1 xor
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 1u); r3 = r3 ^ t_; } // 2 shfl
r4 = r4 - r1; // 3 sub
r2 = r0 * r4 + r2; // 4 mad
r4 = rotl_imm(r4, 9u); // 5 rotl
r0 = rotr_var(r0, r2); // 6 rotr
r6 = r6 ^ ds[((rotl_imm(r7 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
r1 = r1 ^ ds[((rotl_imm(r4 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
r1 = r1 ^ ds[((rotl_imm(r2 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
r7 = r7 ^ ds[((rotl_imm(r0 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
r7 = r7 ^ ds[((rotl_imm(r1 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
r5 = r5 * r4; // 12 mul
r7 = r7 ^ ds[((rotl_imm(r6 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 16u); r6 = r6 ^ t_; } // 14 shfl
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r3 = r3 ^ t_; } // 15 shfl
r2 = r2 | r7; // 16 or
r0 = rotr_var(r0, r6); // 17 rotr
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
r7 = rotl_imm(r7, 24u); // 19 rotl
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
r2 = r2 | r6; // 21 or
r6 = r5 * r4 + r6; // 22 mad
r1 = r1 | r0; // 23 or
r6 = r6 ^ r1; // 24 xor
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
r7 = rotr_var(r7, r0); // 26 rotr
r4 = r5 * r7 + r4; // 27 mad
r3 = r2 * r4 + r3; // 28 mad
r1 = r1 ^ ds[((rotl_imm(r4 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
r2 = r2 ^ ds[((rotl_imm(r3 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
r1 = r1 ^ ds[((rotl_imm(r5 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
r2 = mul_hi(r2, r5); // 35 mulhi
r4 = r4 ^ r2; // 36 xor
r6 = r6 * r5; // 37 mul
r7 = r7 ^ r0; // 38 xor
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
r2 = rotr_var(r2, r3); // 40 rotr
r7 = r7 - r0; // 41 sub
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 43 shfl
r0 = r0 ^ ds[((rotl_imm(r7 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
r3 = r3 ^ ds[((rotl_imm(r1 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
r6 = r6 ^ ds[((rotl_imm(r3 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
r4 = mul_hi(r4, r2); // 48 mulhi
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r0 = r0 ^ t_; } // 50 shfl
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
r5 = r5 ^ ds[((rotl_imm(r2 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
r6 = rotl_imm(r6, 12u); // 53 rotl
r3 = rotl_imm(r3, 12u); // 54 rotl
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
r1 = r1 ^ ds[((rotl_imm(r4 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
r5 = r5 ^ ds[((rotl_imm(r0 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
r3 = mul_hi(r3, r2); // 60 mulhi
r6 = r6 | r4; // 61 or
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r5 = r5 ^ t_; } // 62 shfl
r3 = r3 ^ ds[((rotl_imm(r6 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
}
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((ulong)hi << 32) | (ulong)lo;
}

View file

@ -0,0 +1,123 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
// Header-bound twin of igneum_hash in kernel.cu: the init words come from a kernel argument, not SEEDW.
// Host declarations (also in program_bound.h if present):
// struct IgneumInitWords { uint32_t w[8]; };
// cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
// IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps);
// cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
#include <cuda_runtime.h>
#include <cstdint>
#include "program.h"
struct IgneumInitWords { uint32_t w[8]; };
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
x ^= x >> 16; x *= 0x7feb352du;
x ^= x >> 15; x *= 0x846ca68bu;
x ^= x >> 16;
return x;
}
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw) {
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
uint32_t nonce = baseNonce + gid;
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
{ uint32_t x = nonce ^ iw.w[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw.w[1]; }
{ uint32_t x = nonce ^ iw.w[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw.w[2]; }
{ uint32_t x = nonce ^ iw.w[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw.w[3]; }
{ uint32_t x = nonce ^ iw.w[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw.w[4]; }
{ uint32_t x = nonce ^ iw.w[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw.w[5]; }
{ uint32_t x = nonce ^ iw.w[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw.w[6]; }
{ uint32_t x = nonce ^ iw.w[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw.w[7]; }
{ uint32_t x = nonce ^ iw.w[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw.w[0]; }
for (uint32_t it = 0u; it < 8u; ++it) {
uint32_t sel = r0;
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
r7 = r7 ^ r0; // 1 xor
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r6, 1); // 2 shfl
r4 = r4 - r1; // 3 sub
r2 = r0 * r4 + r2; // 4 mad
r4 = rotl_imm(r4, 9u); // 5 rotl
r0 = rotr_var(r0, r2); // 6 rotr
r6 = r6 ^ ds[((rotl_imm(r7 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
r1 = r1 ^ ds[((rotl_imm(r4 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
r1 = r1 ^ ds[((rotl_imm(r2 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
r7 = r7 ^ ds[((rotl_imm(r0 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
r7 = r7 ^ ds[((rotl_imm(r1 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
r5 = r5 * r4; // 12 mul
r7 = r7 ^ ds[((rotl_imm(r6 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r5, 16); // 14 shfl
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r5, 1); // 15 shfl
r2 = r2 | r7; // 16 or
r0 = rotr_var(r0, r6); // 17 rotr
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
r7 = rotl_imm(r7, 24u); // 19 rotl
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
r2 = r2 | r6; // 21 or
r6 = r5 * r4 + r6; // 22 mad
r1 = r1 | r0; // 23 or
r6 = r6 ^ r1; // 24 xor
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
r7 = rotr_var(r7, r0); // 26 rotr
r4 = r5 * r7 + r4; // 27 mad
r3 = r2 * r4 + r3; // 28 mad
r1 = r1 ^ ds[((rotl_imm(r4 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
r2 = r2 ^ ds[((rotl_imm(r3 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
r1 = r1 ^ ds[((rotl_imm(r5 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
r2 = __umulhi(r2, r5); // 35 mulhi
r4 = r4 ^ r2; // 36 xor
r6 = r6 * r5; // 37 mul
r7 = r7 ^ r0; // 38 xor
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
r2 = rotr_var(r2, r3); // 40 rotr
r7 = r7 - r0; // 41 sub
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 43 shfl
r0 = r0 ^ ds[((rotl_imm(r7 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
r3 = r3 ^ ds[((rotl_imm(r1 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
r6 = r6 ^ ds[((rotl_imm(r3 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
r4 = __umulhi(r4, r2); // 48 mulhi
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r7, 4); // 50 shfl
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
r5 = r5 ^ ds[((rotl_imm(r2 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
r6 = rotl_imm(r6, 12u); // 53 rotl
r3 = rotl_imm(r3, 12u); // 54 rotl
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
r1 = r1 ^ ds[((rotl_imm(r4 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
r5 = r5 ^ ds[((rotl_imm(r0 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
r3 = __umulhi(r3, r2); // 60 mulhi
r6 = r6 | r4; // 61 or
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 62 shfl
r3 = r3 ^ ds[((rotl_imm(r6 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
}
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
}
cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps) {
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
uint32_t block = 32u * blockWarps;
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
igneum_hash_bound<<<nonces / block, block>>>(ds, out, baseNonce, mask, iw);
return cudaGetLastError();
}
cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
cudaFuncAttributes attr;
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash_bound);
if (e != cudaSuccess) return e;
*numRegs = attr.numRegs;
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash_bound, (int)(32u * blockWarps), 0);
}

View file

@ -0,0 +1,112 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
// Memory-hard dataset core, the same text that the Mac's Metal kernels and CPU verifier were checked against.
// Included by kernel.cu (device), host.cu (host reference) and proto-opencl/host.c (C99 host reference).
// See proto-metal/MEMHARD.md for the construction. kernel.cl carries the same text in OpenCL C.
#pragma once
#ifdef __cplusplus
#include <cstdint>
#else
#include <stdint.h>
#endif
#if defined(__CUDACC__)
#define IGNEUM_HD __host__ __device__ __forceinline__
#elif defined(_MSC_VER) && !defined(__cplusplus)
#define IGNEUM_HD static __inline
#else
#define IGNEUM_HD static inline
#endif
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
#define MH_CACHE_LINE_MASK 0x003fffffu
#define MH_SEGMENT_LINES 64u
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
IGNEUM_HD uint32_t mh_rotl(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
// y = ChaCha12 core(x) + x
IGNEUM_HD void mh_chacha_block(const uint32_t* x, uint32_t* y) {
for (uint32_t i = 0u; i < 16u; ++i) y[i] = x[i];
for (uint32_t r = 0u; r < 6u; ++r) {
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
}
for (uint32_t i = 0u; i < 16u; ++i) y[i] += x[i];
}
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
IGNEUM_HD void mh_cache_segment(uint32_t* cache, uint32_t seg) {
uint32_t prev[16]; uint32_t x[16]; uint32_t y[16];
for (uint32_t i = 0u; i < 16u; ++i) prev[i] = 0u;
for (uint32_t j = 0u; j < MH_SEGMENT_LINES; ++j) {
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
x[4] = 0xceed56d7u ^ prev[4];
x[5] = 0x9ba270d2u ^ prev[5];
x[6] = 0x82caab2du ^ prev[6];
x[7] = 0x81ebce0eu ^ prev[7];
x[8] = 0x12b6ecf1u ^ prev[8];
x[9] = 0xd0f3fd7cu ^ prev[9];
x[10] = 0xd872eefeu ^ prev[10];
x[11] = 0xc158c7bdu ^ prev[11];
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
mh_chacha_block(x, y);
uint32_t* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
for (uint32_t i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
}
}
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
IGNEUM_HD void mh_mixer(uint32_t* s, uint32_t rk) {
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
}
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
IGNEUM_HD void mh_item(const uint32_t* cache, uint32_t t, uint32_t* s) {
s[0] = 0xceed56d7u;
s[1] = 0x9ba270d2u;
s[2] = 0x82caab2du;
s[3] = 0x81ebce0eu;
s[4] = 0x12b6ecf1u;
s[5] = 0xd0f3fd7cu;
s[6] = 0xd872eefeu;
s[7] = 0xc158c7bdu;
s[8] = t * 0xf351d601u + 0xc6892460u;
s[9] = t * 0xa3bb398fu + 0x25b7228au;
s[10] = t * 0xb5a09e35u + 0xcd515004u;
s[11] = t * 0x7509c9c1u + 0x2846527au;
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
s[14] = t * 0xded91851u + 0x82961bacu;
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
for (uint32_t r = 0u; r < 8u; ++r) {
mh_mixer(s, 0x9E3779B9u * (r + 1u));
const uint32_t* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= line[i];
}
mh_mixer(s, 0x9E3779B9u * 9u);
}
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 0 2 10 15 of w.
IGNEUM_HD uint32_t mh_j(uint32_t w) { return ((w >> 0u) & 1u) | (((w >> 2u) & 1u) << 1) | (((w >> 10u) & 1u) << 2) | (((w >> 15u) & 1u) << 3); }
IGNEUM_HD uint32_t mh_t(uint32_t w) { w = (w & 0x00007fffu) | ((w >> 16u) << 15u); w = (w & 0x000003ffu) | ((w >> 11u) << 10u); w = (w & 0x00000003u) | ((w >> 3u) << 2u); w = (w & 0x00000000u) | ((w >> 1u) << 0u); return w; }
IGNEUM_HD uint32_t mh_addr(uint32_t t, uint32_t j) { uint32_t w = t; w = ((w >> 0u) << 1u) | (w & 0x00000000u) | (((j >> 0u) & 1u) << 0u); w = ((w >> 2u) << 3u) | (w & 0x00000003u) | (((j >> 1u) & 1u) << 2u); w = ((w >> 10u) << 11u) | (w & 0x000003ffu) | (((j >> 2u) & 1u) << 10u); w = ((w >> 15u) << 16u) | (w & 0x00007fffu) | (((j >> 3u) & 1u) << 15u); return w; }
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
IGNEUM_HD uint32_t mh_word(const uint32_t* cache, uint32_t w) { uint32_t s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }

View file

@ -0,0 +1,109 @@
#include <metal_stdlib>
using namespace metal;
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
#define MH_CACHE_LINE_MASK 0x003fffffu
#define MH_SEGMENT_LINES 64u
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
// y = ChaCha12 core(x) + x
inline void mh_chacha_block(const thread uint* x, thread uint* y) {
for (uint i = 0u; i < 16u; ++i) y[i] = x[i];
for (uint r = 0u; r < 6u; ++r) {
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
}
for (uint i = 0u; i < 16u; ++i) y[i] += x[i];
}
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
inline void mh_cache_segment(device uint* cache, uint seg) {
uint prev[16]; uint x[16]; uint y[16];
for (uint i = 0u; i < 16u; ++i) prev[i] = 0u;
for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) {
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
x[4] = 0xceed56d7u ^ prev[4];
x[5] = 0x9ba270d2u ^ prev[5];
x[6] = 0x82caab2du ^ prev[6];
x[7] = 0x81ebce0eu ^ prev[7];
x[8] = 0x12b6ecf1u ^ prev[8];
x[9] = 0xd0f3fd7cu ^ prev[9];
x[10] = 0xd872eefeu ^ prev[10];
x[11] = 0xc158c7bdu ^ prev[11];
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
mh_chacha_block(x, y);
device uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
}
}
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
inline void mh_mixer(thread uint* s, uint rk) {
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
}
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
inline void mh_item(device const uint* cache, uint t, thread uint* s) {
s[0] = 0xceed56d7u;
s[1] = 0x9ba270d2u;
s[2] = 0x82caab2du;
s[3] = 0x81ebce0eu;
s[4] = 0x12b6ecf1u;
s[5] = 0xd0f3fd7cu;
s[6] = 0xd872eefeu;
s[7] = 0xc158c7bdu;
s[8] = t * 0xf351d601u + 0xc6892460u;
s[9] = t * 0xa3bb398fu + 0x25b7228au;
s[10] = t * 0xb5a09e35u + 0xcd515004u;
s[11] = t * 0x7509c9c1u + 0x2846527au;
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
s[14] = t * 0xded91851u + 0x82961bacu;
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
for (uint r = 0u; r < 8u; ++r) {
mh_mixer(s, 0x9E3779B9u * (r + 1u));
device const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i];
}
mh_mixer(s, 0x9E3779B9u * 9u);
}
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 0 2 10 15 of w.
inline uint mh_j(uint w) { return ((w >> 0u) & 1u) | (((w >> 2u) & 1u) << 1) | (((w >> 10u) & 1u) << 2) | (((w >> 15u) & 1u) << 3); }
inline uint mh_t(uint w) { w = (w & 0x00007fffu) | ((w >> 16u) << 15u); w = (w & 0x000003ffu) | ((w >> 11u) << 10u); w = (w & 0x00000003u) | ((w >> 3u) << 2u); w = (w & 0x00000000u) | ((w >> 1u) << 0u); return w; }
inline uint mh_addr(uint t, uint j) { uint w = t; w = ((w >> 0u) << 1u) | (w & 0x00000000u) | (((j >> 0u) & 1u) << 0u); w = ((w >> 2u) << 3u) | (w & 0x00000003u) | (((j >> 1u) & 1u) << 2u); w = ((w >> 10u) << 11u) | (w & 0x000003ffu) | (((j >> 2u) & 1u) << 10u); w = ((w >> 15u) << 16u) | (w & 0x00007fffu) | (((j >> 3u) & 1u) << 15u); return w; }
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
inline uint mh_word(device const uint* cache, uint w) { uint s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }
// One thread per segment (2^16 threads).
kernel void igneum_cache_fill(device uint* cache [[buffer(0)]], uint gid [[thread_position_in_grid]]) {
mh_cache_segment(cache, gid);
}
// One thread per 64-byte item (dataset words / 16 threads).
kernel void igneum_build(device const uint* cache [[buffer(0)]], device uint* dataset [[buffer(1)]],
uint gid [[thread_position_in_grid]]) {
uint s[16];
mh_item(cache, gid, s);
for (uint i = 0u; i < 16u; ++i) dataset[mh_addr(gid, i)] = s[i];
}

View file

@ -0,0 +1,76 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
// Program metadata for host.cu plus the launch wrappers defined in kernel.cu.
// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros.
#pragma once
#ifdef __cplusplus
#include <cstdint>
#else
#include <stdint.h>
#endif
#ifndef IGNEUM_NO_CUDA
#include <cuda_runtime.h>
#endif
#define IGNEUM_SEED_STRING "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000"
#define IGNEUM_SEED_BYTES_HEX "edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07"
#define IGNEUM_GENERATOR 3
#define IGNEUM_PROGRAM_ATTEMPT 0
#define IGNEUM_PROGRAM_ID 0x73bcbfe8ccf988f1ull
#define IGNEUM_DAY_STRING "bytes:69676e65756d2d6461792ffa50000000000000"
#define IGNEUM_DAY_BYTES_HEX "69676e65756d2d6461792ffa50000000000000"
#define IGNEUM_DAY0 0xceed56d7u
#define IGNEUM_DAY1 0x9ba270d2u
#define IGNEUM_DATASET_LOG2 28
#define IGNEUM_MASK 0x0fffffffu
#define IGNEUM_LANES 32
#define IGNEUM_ITERATIONS 8
#define IGNEUM_INSTR_COUNT 64
#define IGNEUM_LOADS_PER_HASH 128
#define IGNEUM_WIDE_LOADS_PER_HASH 0
#define IGNEUM_OP_MIX "load=16 add=15 shfl=6 mad=4 or=4 rotl=4 rotr=4 xor=4 mulhi=3 mul=2 sub=2"
// Program class v3 (Counter ASIC 2.0, docs/plans/counter-asic-2-rollout.md): generator version 3; a worker that
// runs another class refuses this pack, and a job line names the class it wants (class=v3 era=<hex>).
#define IGNEUM_PROGRAM_CLASS "v3"
#define IGNEUM_ERA_SEED_HEX "8bffdd3366b9c3ffe89c1231e91538d531cfa717307df5c48ccba43096e17212"
// Class v3 construction (Counter ASIC 2.0, 5 October 2026, docs/plans/mixer-x4.md): version 2 loads; the dataset item
// derivation applies the mixer IGNEUM_MIXER_MULT times per round (memhard.h), and the cache follows the growth rule.
#define IGNEUM_LOAD_CLASS "w4-erab2ed8a89"
#define IGNEUM_LOAD_SLOTS 16
#define IGNEUM_LOAD_MIX { 100, 0, 0 }
#define IGNEUM_LOAD_WIDTH_COUNTS { 16, 0, 0 } // loads of 4, 16, 64 bytes per program
#define IGNEUM_BYTES_PER_HASH 512
#define IGNEUM_FOLD_ROT 11
#define IGNEUM_FOLD_MUL 0x9e3779b1u
// Era layout (5 October 2026, docs/plans/era-layout.md): NOT the lottery hash. Every dataset load reads
// idx = ((rotl(src * STRIDE_MUL, STRIDE_ROT) & window mask) | window offset) & MASK; the window of a load site is the
// dataset, a half or a quarter of it (IGNEUM_ERA_WINDOWS: site:shrink:offset); dataset word w holds word j(w) of item
// t(w) with j's bits at the INTERLEAVE positions (memhard.h: mh_t, mh_j, mh_addr).
#define IGNEUM_ERA_LABEL "b2ed8a89"
#define IGNEUM_ERA_SEED_WORDS { 0xb2ed8a89u, 0xb023f2bau, 0x6bbf405eu, 0x98153cddu, 0x49428e54u, 0xfbfa65eeu, 0xbcb76d0du, 0x2f509891u }
#define IGNEUM_ERA_ALLOWED_WIDTHS { 1, 0, 0 } // words, ascending, 0 = unused; one entry pins the width
#define IGNEUM_ERA_WIDTH_WORDS 1
#define IGNEUM_ERA_STRIDE_MUL 0x625e5ab3u
#define IGNEUM_ERA_STRIDE_ROT 19
#define IGNEUM_ERA_INTERLEAVE { 0, 2, 10, 15 }
#define IGNEUM_ERA_WINDOWS "7:2:1 8:1:1 9:1:1 10:1:1 11:0:0 13:1:1 29:0:0 30:2:2 31:1:1 44:1:1 46:2:0 47:0:0 52:0:0 56:0:0 58:2:0 63:1:1"
// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h)
#define IGNEUM_DATASET_MODE 1
#define IGNEUM_SEEDW_INIT { 0x667d0fbdu, 0x7b8e5963u, 0x31c67e5eu, 0x4529ddc6u, 0xef19d6d8u, 0xaccf6211u, 0xda0aed32u, 0xabc6df31u }
#define IGNEUM_KEY_INIT { 0xceed56d7u, 0x9ba270d2u, 0x82caab2du, 0x81ebce0eu, 0x12b6ecf1u, 0xd0f3fd7cu, 0xd872eefeu, 0xc158c7bdu }
#define IGNEUM_CACHE_LOG2_WORDS 26
#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6
#define IGNEUM_CACHE_SEGMENTS 65536u
#define IGNEUM_ITEM_ROUNDS 8
#define IGNEUM_MIX_ROT_INIT { 17u, 12u, 20u, 23u, 7u, 3u, 27u, 16u }
#define IGNEUM_MIX_MUL_INIT { 0xf351d601u, 0xa3bb398fu, 0xb5a09e35u, 0x7509c9c1u, 0x6bbf31e9u, 0xfc849a79u, 0xded91851u, 0x8d9113d1u, 0x0ff15225u, 0x3a5bdd41u, 0xab533435u, 0xe1c55ad5u, 0xe6d3bd0du, 0x9d9ffbbdu, 0xbb2a3cf3u, 0x50a7c08du }
#define IGNEUM_MIX_RC_INIT { 0xc6892460u, 0x25b7228au, 0xcd515004u, 0x2846527au, 0xa6324241u, 0x36e3ec53u, 0x82961bacu, 0x0f97ba7du, 0xb6f921a9u, 0x3ada24e5u, 0xde20ab91u, 0x5378eeb2u, 0x7d161662u, 0x89353cc1u, 0xb1aa03a2u, 0x788acae6u }
#ifndef IGNEUM_NO_CUDA
// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError().
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments);
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems);
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
uint32_t nonces, uint32_t blockWarps);
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
#endif

View file

@ -0,0 +1,143 @@
{
"format": "igneum-program-pack-3",
"generator": 3,
"attempt": 0,
"program_id": "0x73bcbfe8ccf988f1",
"program_id_derivation": "FNV-1a 64 over 'igneum-program/' || generator_le32 || seed_words as little-endian bytes || attempt_le32",
"dataset_mode": "memory-hard",
"seed": "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000",
"seed_bytes": "edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07",
"seed_words": ["0x667d0fbd", "0x7b8e5963", "0x31c67e5e", "0x4529ddc6", "0xef19d6d8", "0xaccf6211", "0xda0aed32", "0xabc6df31"],
"seed_derivation": "seed_words = FNV-1a 64 over seed_bytes (attempt 0) or seed_bytes || attempt_le32 (attempt k >= 1), basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32",
"generator_rule": "version 2: exactly 16 load slots drawn first from instructions 1..63 (partial Fisher-Yates), the other 48 ops from the ten non-load weights (sum 75); a load's source is drawn from the registers other than dst written by an earlier instruction and not read by a load since; the candidate must pass the acceptance rule of spec 01 section 1.4.6 (static: no cyclically stale load source, every register has an injecting write; dynamic: 64 units on the seed-keyed closed-form dataset with no constant register bit, no lane-constant load site, under 164 saturated final values, every output bit within 136 of 1024, distinct addresses above 245760), else the next attempt of the seed is tried",
"lanes": 32,
"registers": 8,
"iterations": 8,
"instruction_count": 64,
"loads_per_hash": 128,
"program_class": "v3",
"era_seed_bytes": "8bffdd3366b9c3ffe89c1231e91538d531cfa717307df5c48ccba43096e17212",
"load_class": "w4-erab2ed8a89",
"load_slots": 16,
"load_mix_percent_4_16_64": [100, 0, 0],
"load_width_counts_4_16_64": [16, 0, 0],
"bytes_per_hash": 512,
"wide_load": "read-width experiment (5 October 2026, docs/plans/read-width.md), NOT the lottery hash: a load of W words (width field, 4 or 16) reads dataset[b .. b + W) with b = (src & mask) & ~(W - 1) and folds every word into dst: x = dst ^ w[0]; for j in 1..W: x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x; width 1 is the plain load; the width is drawn per instruction from the class mix with one extra below(100) draw after the nine of version 2, and the program id is FNV-1a 64 over 'igneum-program-rw/' || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots",
"era": {
"label": "b2ed8a89",
"seed_words": ["0xb2ed8a89", "0xb023f2ba", "0x6bbf405e", "0x98153cdd", "0x49428e54", "0xfbfa65ee", "0xbcb76d0d", "0x2f509891"],
"draw": "docs/plans/era-layout.md 1.1: SplitMix64 seeded with seed_words[0] | seed_words[1] << 32 of seed_words_from_bytes('igneum-era/' || n_le64 || E_n); width = allowed[below(|allowed|)], stride_mul = low32(next()) | 1, stride_rot = 1 + below(31), then four next() draws for a partial Fisher-Yates over positions log2(W)..15 of which 4 - log2(W) are used",
"allowed_widths": [1],
"width_words": 1,
"stride_mul": "0x625e5ab3",
"stride_rot": 19,
"interleave": [0, 2, 10, 15],
"address": "y = rotl(src * stride_mul, stride_rot); k = min(win, D - 26); idx = ((y & (mask >> k)) | ((off & (2^k - 1)) << (D - k))) & mask; a wide load aligns idx down to W words",
"windows": "per instruction, after the width roll: win = below(3), off = low32(next()) & (2^win - 1); used on a load slot (the instruction's win and off fields)",
"dataset_word": "dataset[w] = item(t(w))[j(w)]: j(w) gathers the bits of w at the interleave positions, t(w) is w with those bits removed",
"program_id_suffix": "'era/' || allowed[3] || width_words || stride_mul_le32 || stride_rot_le32 || interleave[4]"
},
"op_mix": {"load": 16, "add": 15, "shfl": 6, "mad": 4, "or": 4, "rotl": 4, "rotr": 4, "xor": 4, "mulhi": 3, "mul": 2, "sub": 2},
"register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]",
"splitmix32": "x ^= x>>16; x *= 0x7feb352d; x ^= x>>15; x *= 0x846ca68b; x ^= x>>16",
"iteration": "sel = r0 sampled once at the top of each iteration, then all instructions in order",
"output": "lo = r0 ^ rotl(r1,7) ^ rotl(r2,14) ^ rotl(r3,21); hi = r4 ^ rotl(r5,9) ^ rotl(r6,18) ^ rotl(r7,27); out = (hi << 32) | lo",
"op_semantics": {
"add": "dst = dst + src + (bit `bit` of sel ? imm2 : imm)",
"sub": "dst = dst - src",
"mul": "dst = dst * src (low 32)",
"mulhi": "dst = high 32 bits of dst * src",
"xor": "dst = dst ^ src",
"or": "dst = dst | src",
"rotl": "dst = rotl(dst, rot), rot in 1..31",
"rotr": "dst = rotr(dst, src & 31)",
"mad": "dst = src * src2 + dst",
"shfl": "dst = dst ^ (src of lane (lane ^ mask)), mask in {1,2,4,8,16}, within the 32-lane warp",
"load": "dst = dst ^ dataset[src & dataset.mask]",
"wload": "base = (src of lane 0 & dataset.mask) & ~31; dst = dst ^ dataset[base + lane] (warp-coalesced 128-byte load, lever b, only when --wide-frac > 0)"
},
"dataset": {
"log2_words": 28,
"bytes": 1073741824,
"mask": "0x0fffffff",
"day": "bytes:69676e65756d2d6461792ffa50000000000000",
"day_bytes": "69676e65756d2d6461792ffa50000000000000",
"day_words_from": "seed_words_from_bytes(day_bytes)",
"d0": "0xceed56d7",
"d1": "0x9ba270d2",
"mode": "memory-hard",
"spec": "proto-metal/MEMHARD.md",
"key": ["0xceed56d7", "0x9ba270d2", "0x82caab2d", "0x81ebce0e", "0x12b6ecf1", "0xd0f3fd7c", "0xd872eefe", "0xc158c7bd"],
"key_derivation": "the 8 words of seed_words_from_bytes(day_bytes); d0, d1 are key[0], key[1]",
"cache": {"log2_words": 26, "bytes": 268435456, "line_words": 16, "segment_lines": 64, "segments": 65536, "block": "ChaCha12 core + feed-forward, rotations 16 12 8 7", "sigma": ["0x61707865", "0x3320646e", "0x79622d32", "0x6b206574"], "tag": ["0x49676e65", "0x756d4d48"], "chain": "in_j = prev_line ^ (sigma[0..3] || key[0..7] || seg || j || tag[0..1]); line_j = block(in_j); prev_0 = 0"},
"mixer": {"draw": "SplitMix64 seeded with key[0] | key[1] << 32: rot[0..7] = 1 + next() % 31, mul[0..15] = low32(next()) | 1, rc[0..15] = low32(next())", "rot": [17, 12, 20, 23, 7, 3, 27, 16], "mul": ["0xf351d601", "0xa3bb398f", "0xb5a09e35", "0x7509c9c1", "0x6bbf31e9", "0xfc849a79", "0xded91851", "0x8d9113d1", "0x0ff15225", "0x3a5bdd41", "0xab533435", "0xe1c55ad5", "0xe6d3bd0d", "0x9d9ffbbd", "0xbb2a3cf3", "0x50a7c08d"], "rc": ["0xc6892460", "0x25b7228a", "0xcd515004", "0x2846527a", "0xa6324241", "0x36e3ec53", "0x82961bac", "0x0f97ba7d", "0xb6f921a9", "0x3ada24e5", "0xde20ab91", "0x5378eeb2", "0x7d161662", "0x89353cc1", "0xb1aa03a2", "0x788acae6"], "round": "for i in 0..15: s[i] = (s[i] ^ (rc[i] + (r+1) * 0x9E3779B9)) * mul[i]; then quarter rounds on columns (0,4,8,12) (1,5,9,13) (2,6,10,14) (3,7,11,15) with rot[0..3] and diagonals (0,5,10,15) (1,6,11,12) (2,7,8,13) (3,4,9,14) with rot[4..7]", "quarter_round": "a += b; d ^= a; d = rotl(d, r1); c += d; b ^= c; b = rotl(b, r2); a += b; d ^= a; d = rotl(d, r3); c += d; b ^= c; b = rotl(b, r4)"},
"item": "s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..7: s = M_r(s); line = s[0] & 0x003fffff; s[i] ^= cache[line * 16 + i]; then s = M_8(s); item(t) = s",
"word": "dataset[w] = item(w >> 4)[w & 15]"
},
"instructions": [
{"i": 0, "op": "add", "dst": 4, "src": 5, "src2": 7, "imm": "0xea86e152", "imm2": "0x5810667a", "rot": 27, "bit": 13, "mask": 2, "width": 1, "win": 0, "off": 0},
{"i": 1, "op": "xor", "dst": 7, "src": 0, "src2": 6, "imm": "0xbaab6229", "imm2": "0xed861989", "rot": 26, "bit": 22, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 2, "op": "shfl", "dst": 3, "src": 6, "src2": 2, "imm": "0x5b623116", "imm2": "0xff12e5b2", "rot": 12, "bit": 24, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 3, "op": "sub", "dst": 4, "src": 1, "src2": 1, "imm": "0xe99741c7", "imm2": "0xf5fa5009", "rot": 1, "bit": 21, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 4, "op": "mad", "dst": 2, "src": 0, "src2": 4, "imm": "0x673c2157", "imm2": "0xee02465f", "rot": 20, "bit": 22, "mask": 2, "width": 1, "win": 0, "off": 0},
{"i": 5, "op": "rotl", "dst": 4, "src": 7, "src2": 7, "imm": "0x946f7818", "imm2": "0x45d3399e", "rot": 9, "bit": 2, "mask": 16, "width": 1, "win": 0, "off": 0},
{"i": 6, "op": "rotr", "dst": 0, "src": 2, "src2": 4, "imm": "0x5f6a0ed2", "imm2": "0x7043a636", "rot": 19, "bit": 31, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 7, "op": "load", "dst": 6, "src": 7, "src2": 1, "imm": "0x5c61dcf7", "imm2": "0x7466aa40", "rot": 19, "bit": 9, "mask": 2, "width": 1, "win": 2, "off": 1},
{"i": 8, "op": "load", "dst": 1, "src": 4, "src2": 5, "imm": "0x85668475", "imm2": "0xdb8cc483", "rot": 29, "bit": 7, "mask": 4, "width": 1, "win": 1, "off": 1},
{"i": 9, "op": "load", "dst": 1, "src": 2, "src2": 4, "imm": "0x4454980f", "imm2": "0xebd31581", "rot": 10, "bit": 28, "mask": 16, "width": 1, "win": 1, "off": 1},
{"i": 10, "op": "load", "dst": 7, "src": 0, "src2": 2, "imm": "0xe075297c", "imm2": "0x5779c44c", "rot": 10, "bit": 22, "mask": 2, "width": 1, "win": 1, "off": 1},
{"i": 11, "op": "load", "dst": 7, "src": 1, "src2": 6, "imm": "0x65aa4311", "imm2": "0x4fe48ea9", "rot": 15, "bit": 9, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 12, "op": "mul", "dst": 5, "src": 4, "src2": 2, "imm": "0x1383d3ad", "imm2": "0xf3094b29", "rot": 8, "bit": 9, "mask": 2, "width": 1, "win": 0, "off": 0},
{"i": 13, "op": "load", "dst": 7, "src": 6, "src2": 2, "imm": "0xed8a496f", "imm2": "0x3072c3c6", "rot": 28, "bit": 19, "mask": 8, "width": 1, "win": 1, "off": 1},
{"i": 14, "op": "shfl", "dst": 6, "src": 5, "src2": 0, "imm": "0x8b965b57", "imm2": "0xcfeca6c1", "rot": 12, "bit": 27, "mask": 16, "width": 1, "win": 0, "off": 0},
{"i": 15, "op": "shfl", "dst": 3, "src": 5, "src2": 3, "imm": "0x877c7586", "imm2": "0xa9cb2a03", "rot": 2, "bit": 29, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 16, "op": "or", "dst": 2, "src": 7, "src2": 3, "imm": "0xb740221a", "imm2": "0x89d38d6d", "rot": 6, "bit": 31, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 17, "op": "rotr", "dst": 0, "src": 6, "src2": 1, "imm": "0x26f3ad8a", "imm2": "0x27256f15", "rot": 18, "bit": 5, "mask": 2, "width": 1, "win": 0, "off": 0},
{"i": 18, "op": "add", "dst": 7, "src": 5, "src2": 7, "imm": "0xf66e7017", "imm2": "0xb9e3577e", "rot": 9, "bit": 12, "mask": 16, "width": 1, "win": 0, "off": 0},
{"i": 19, "op": "rotl", "dst": 7, "src": 0, "src2": 5, "imm": "0x849ae6ee", "imm2": "0x02b358f9", "rot": 24, "bit": 18, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 20, "op": "add", "dst": 6, "src": 7, "src2": 6, "imm": "0x52334d12", "imm2": "0x8c9f0ef8", "rot": 11, "bit": 23, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 21, "op": "or", "dst": 2, "src": 6, "src2": 5, "imm": "0xb1871e63", "imm2": "0xb2e40191", "rot": 5, "bit": 13, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 22, "op": "mad", "dst": 6, "src": 5, "src2": 4, "imm": "0x97df29e4", "imm2": "0xe60fea84", "rot": 11, "bit": 16, "mask": 2, "width": 1, "win": 0, "off": 0},
{"i": 23, "op": "or", "dst": 1, "src": 0, "src2": 3, "imm": "0x8f30d21d", "imm2": "0x2df685a0", "rot": 31, "bit": 0, "mask": 16, "width": 1, "win": 0, "off": 0},
{"i": 24, "op": "xor", "dst": 6, "src": 1, "src2": 2, "imm": "0xd2c4025f", "imm2": "0x5269eb4d", "rot": 31, "bit": 21, "mask": 16, "width": 1, "win": 0, "off": 0},
{"i": 25, "op": "add", "dst": 2, "src": 6, "src2": 5, "imm": "0x659fc3d3", "imm2": "0x9cec0e12", "rot": 6, "bit": 17, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 26, "op": "rotr", "dst": 7, "src": 0, "src2": 1, "imm": "0xd89ef484", "imm2": "0x20be3846", "rot": 12, "bit": 9, "mask": 16, "width": 1, "win": 0, "off": 0},
{"i": 27, "op": "mad", "dst": 4, "src": 5, "src2": 7, "imm": "0xb48420ae", "imm2": "0x3d1f2485", "rot": 14, "bit": 28, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 28, "op": "mad", "dst": 3, "src": 2, "src2": 4, "imm": "0xc5c46d76", "imm2": "0x700044b5", "rot": 22, "bit": 10, "mask": 2, "width": 1, "win": 0, "off": 0},
{"i": 29, "op": "load", "dst": 1, "src": 4, "src2": 3, "imm": "0x0fbaf177", "imm2": "0xfff4f2ed", "rot": 20, "bit": 31, "mask": 2, "width": 1, "win": 0, "off": 0},
{"i": 30, "op": "load", "dst": 2, "src": 3, "src2": 0, "imm": "0xe90eb2e5", "imm2": "0xb0f9eb79", "rot": 27, "bit": 14, "mask": 4, "width": 1, "win": 2, "off": 2},
{"i": 31, "op": "load", "dst": 1, "src": 5, "src2": 2, "imm": "0x97ba3fc3", "imm2": "0x7894e657", "rot": 3, "bit": 30, "mask": 16, "width": 1, "win": 1, "off": 1},
{"i": 32, "op": "add", "dst": 7, "src": 2, "src2": 7, "imm": "0x070888a8", "imm2": "0xe403240e", "rot": 2, "bit": 16, "mask": 2, "width": 1, "win": 0, "off": 0},
{"i": 33, "op": "add", "dst": 2, "src": 0, "src2": 5, "imm": "0xf2e46d55", "imm2": "0x29701828", "rot": 31, "bit": 28, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 34, "op": "add", "dst": 2, "src": 3, "src2": 0, "imm": "0x58f75b87", "imm2": "0x343b7aee", "rot": 12, "bit": 14, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 35, "op": "mulhi", "dst": 2, "src": 5, "src2": 6, "imm": "0x5892a9e6", "imm2": "0xc9824c94", "rot": 19, "bit": 26, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 36, "op": "xor", "dst": 4, "src": 2, "src2": 3, "imm": "0x7b5b5474", "imm2": "0x45cfc5dd", "rot": 17, "bit": 18, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 37, "op": "mul", "dst": 6, "src": 5, "src2": 7, "imm": "0xccf564a5", "imm2": "0x873ad101", "rot": 7, "bit": 11, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 38, "op": "xor", "dst": 7, "src": 0, "src2": 6, "imm": "0xcac8f06d", "imm2": "0x6b97c683", "rot": 18, "bit": 28, "mask": 2, "width": 1, "win": 0, "off": 0},
{"i": 39, "op": "add", "dst": 7, "src": 2, "src2": 4, "imm": "0xb8180e9d", "imm2": "0x32bbd117", "rot": 23, "bit": 19, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 40, "op": "rotr", "dst": 2, "src": 3, "src2": 0, "imm": "0x2d6070bc", "imm2": "0x68ff101e", "rot": 13, "bit": 18, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 41, "op": "sub", "dst": 7, "src": 0, "src2": 2, "imm": "0x0daf96ea", "imm2": "0x36f37be1", "rot": 5, "bit": 0, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 42, "op": "add", "dst": 4, "src": 3, "src2": 6, "imm": "0x6ced15b7", "imm2": "0x6df7aed4", "rot": 19, "bit": 4, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 43, "op": "shfl", "dst": 7, "src": 3, "src2": 5, "imm": "0x8ace05f3", "imm2": "0xd378ec12", "rot": 23, "bit": 10, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 44, "op": "load", "dst": 0, "src": 7, "src2": 4, "imm": "0xb0607786", "imm2": "0xc4acabbc", "rot": 13, "bit": 7, "mask": 16, "width": 1, "win": 1, "off": 1},
{"i": 45, "op": "add", "dst": 1, "src": 6, "src2": 5, "imm": "0xe09f54e9", "imm2": "0x83e825bf", "rot": 23, "bit": 14, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 46, "op": "load", "dst": 3, "src": 1, "src2": 0, "imm": "0x63cc1e4e", "imm2": "0xa1be8118", "rot": 12, "bit": 6, "mask": 2, "width": 1, "win": 2, "off": 0},
{"i": 47, "op": "load", "dst": 6, "src": 3, "src2": 7, "imm": "0x353f1d79", "imm2": "0x3b2e7456", "rot": 18, "bit": 18, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 48, "op": "mulhi", "dst": 4, "src": 2, "src2": 7, "imm": "0x00d8a3cd", "imm2": "0x231866d2", "rot": 21, "bit": 20, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 49, "op": "add", "dst": 5, "src": 0, "src2": 2, "imm": "0xa8bae6df", "imm2": "0xf572bdb9", "rot": 14, "bit": 7, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 50, "op": "shfl", "dst": 0, "src": 7, "src2": 7, "imm": "0x81ef22e1", "imm2": "0x74438fc5", "rot": 28, "bit": 18, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 51, "op": "add", "dst": 6, "src": 0, "src2": 6, "imm": "0x383b9260", "imm2": "0x11e17c61", "rot": 12, "bit": 19, "mask": 16, "width": 1, "win": 0, "off": 0},
{"i": 52, "op": "load", "dst": 5, "src": 2, "src2": 2, "imm": "0xfb84f451", "imm2": "0x11cd863e", "rot": 21, "bit": 20, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 53, "op": "rotl", "dst": 6, "src": 5, "src2": 4, "imm": "0xb1a7db6b", "imm2": "0x76686b9b", "rot": 12, "bit": 4, "mask": 16, "width": 1, "win": 0, "off": 0},
{"i": 54, "op": "rotl", "dst": 3, "src": 6, "src2": 3, "imm": "0x6f981f52", "imm2": "0xd99aeba2", "rot": 12, "bit": 27, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 55, "op": "add", "dst": 2, "src": 1, "src2": 2, "imm": "0xac6be8e3", "imm2": "0x18d67dbb", "rot": 26, "bit": 10, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 56, "op": "load", "dst": 1, "src": 4, "src2": 0, "imm": "0x7e7f6a00", "imm2": "0x6f0747da", "rot": 25, "bit": 28, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 57, "op": "add", "dst": 5, "src": 0, "src2": 4, "imm": "0xf03673fe", "imm2": "0xa75cd60d", "rot": 16, "bit": 12, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 58, "op": "load", "dst": 5, "src": 0, "src2": 2, "imm": "0x227f94a6", "imm2": "0x0e8344f9", "rot": 20, "bit": 10, "mask": 2, "width": 1, "win": 2, "off": 0},
{"i": 59, "op": "add", "dst": 1, "src": 4, "src2": 3, "imm": "0xdecd4794", "imm2": "0x8dfb96bb", "rot": 21, "bit": 7, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 60, "op": "mulhi", "dst": 3, "src": 2, "src2": 2, "imm": "0x0dd268e0", "imm2": "0x53034ca9", "rot": 1, "bit": 8, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 61, "op": "or", "dst": 6, "src": 4, "src2": 7, "imm": "0x3a45a321", "imm2": "0x9bc59a5f", "rot": 25, "bit": 11, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 62, "op": "shfl", "dst": 5, "src": 4, "src2": 2, "imm": "0x8f229cc1", "imm2": "0xcaac64a2", "rot": 17, "bit": 13, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 63, "op": "load", "dst": 3, "src": 6, "src2": 6, "imm": "0xb2574178", "imm2": "0xcbcc798d", "rot": 28, "bit": 0, "mask": 16, "width": 1, "win": 1, "off": 1}
]
}

View file

@ -0,0 +1,109 @@
#include <metal_stdlib>
using namespace metal;
#define MASK 0x0fffffffu
constant uint SEEDW[8] = { 0x667d0fbdu, 0x7b8e5963u, 0x31c67e5eu, 0x4529ddc6u, 0xef19d6d8u, 0xaccf6211u, 0xda0aed32u, 0xabc6df31u };
inline uint splitmix32(uint x) {
x ^= x >> 16; x *= 0x7feb352du;
x ^= x >> 15; x *= 0x846ca68bu;
x ^= x >> 16;
return x;
}
inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31
inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
inline uint ds_elem(uint i, uint d0, uint d1) {
uint x = i ^ d0;
x *= 0x9E3779B1u; x ^= x >> 15;
x += d1;
x *= 0x85EBCA77u; x ^= x >> 13;
x *= 0xC2B2AE3Du; x ^= x >> 16;
return x;
}
kernel void igneum_hash(device const uint* dataset [[buffer(0)]],
device ulong* out [[buffer(1)]],
constant uint& baseNonce [[buffer(2)]],
uint gid [[thread_position_in_grid]]) {
uint nonce = baseNonce + gid;
uint r0, r1, r2, r3, r4, r5, r6, r7;
{ uint x = nonce ^ SEEDW[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ SEEDW[1]; }
{ uint x = nonce ^ SEEDW[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ SEEDW[2]; }
{ uint x = nonce ^ SEEDW[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ SEEDW[3]; }
{ uint x = nonce ^ SEEDW[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ SEEDW[4]; }
{ uint x = nonce ^ SEEDW[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ SEEDW[5]; }
{ uint x = nonce ^ SEEDW[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ SEEDW[6]; }
{ uint x = nonce ^ SEEDW[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ SEEDW[7]; }
{ uint x = nonce ^ SEEDW[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ SEEDW[0]; }
for (uint it = 0u; it < 8u; ++it) {
uint sel = r0;
r4 = r4 + r5 + select(0xea86e152u, 0x5810667au, ((sel >> 13u) & 1u) != 0u); // 0
r7 = r7 ^ r0; // 1
r3 = r3 ^ simd_shuffle_xor(r6, (ushort)1); // 2
r4 = r4 - r1; // 3
r2 = r0 * r4 + r2; // 4
r4 = rotl_imm(r4, 9u); // 5
r0 = rotr_var(r0, r2); // 6
r6 = r6 ^ dataset[((rotl_imm(r7 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x04000000u) & MASK]; // 7
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 8
r1 = r1 ^ dataset[((rotl_imm(r2 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 9
r7 = r7 ^ dataset[((rotl_imm(r0 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 10
r7 = r7 ^ dataset[((rotl_imm(r1 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 11
r5 = r5 * r4; // 12
r7 = r7 ^ dataset[((rotl_imm(r6 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 13
r6 = r6 ^ simd_shuffle_xor(r5, (ushort)16); // 14
r3 = r3 ^ simd_shuffle_xor(r5, (ushort)1); // 15
r2 = r2 | r7; // 16
r0 = rotr_var(r0, r6); // 17
r7 = r7 + r5 + select(0xf66e7017u, 0xb9e3577eu, ((sel >> 12u) & 1u) != 0u); // 18
r7 = rotl_imm(r7, 24u); // 19
r6 = r6 + r7 + select(0x52334d12u, 0x8c9f0ef8u, ((sel >> 23u) & 1u) != 0u); // 20
r2 = r2 | r6; // 21
r6 = r5 * r4 + r6; // 22
r1 = r1 | r0; // 23
r6 = r6 ^ r1; // 24
r2 = r2 + r6 + select(0x659fc3d3u, 0x9cec0e12u, ((sel >> 17u) & 1u) != 0u); // 25
r7 = rotr_var(r7, r0); // 26
r4 = r5 * r7 + r4; // 27
r3 = r2 * r4 + r3; // 28
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 29
r2 = r2 ^ dataset[((rotl_imm(r3 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x08000000u) & MASK]; // 30
r1 = r1 ^ dataset[((rotl_imm(r5 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 31
r7 = r7 + r2 + select(0x070888a8u, 0xe403240eu, ((sel >> 16u) & 1u) != 0u); // 32
r2 = r2 + r0 + select(0xf2e46d55u, 0x29701828u, ((sel >> 28u) & 1u) != 0u); // 33
r2 = r2 + r3 + select(0x58f75b87u, 0x343b7aeeu, ((sel >> 14u) & 1u) != 0u); // 34
r2 = mulhi(r2, r5); // 35
r4 = r4 ^ r2; // 36
r6 = r6 * r5; // 37
r7 = r7 ^ r0; // 38
r7 = r7 + r2 + select(0xb8180e9du, 0x32bbd117u, ((sel >> 19u) & 1u) != 0u); // 39
r2 = rotr_var(r2, r3); // 40
r7 = r7 - r0; // 41
r4 = r4 + r3 + select(0x6ced15b7u, 0x6df7aed4u, ((sel >> 4u) & 1u) != 0u); // 42
r7 = r7 ^ simd_shuffle_xor(r3, (ushort)4); // 43
r0 = r0 ^ dataset[((rotl_imm(r7 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 44
r1 = r1 + r6 + select(0xe09f54e9u, 0x83e825bfu, ((sel >> 14u) & 1u) != 0u); // 45
r3 = r3 ^ dataset[((rotl_imm(r1 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 46
r6 = r6 ^ dataset[((rotl_imm(r3 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 47
r4 = mulhi(r4, r2); // 48
r5 = r5 + r0 + select(0xa8bae6dfu, 0xf572bdb9u, ((sel >> 7u) & 1u) != 0u); // 49
r0 = r0 ^ simd_shuffle_xor(r7, (ushort)4); // 50
r6 = r6 + r0 + select(0x383b9260u, 0x11e17c61u, ((sel >> 19u) & 1u) != 0u); // 51
r5 = r5 ^ dataset[((rotl_imm(r2 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 52
r6 = rotl_imm(r6, 12u); // 53
r3 = rotl_imm(r3, 12u); // 54
r2 = r2 + r1 + select(0xac6be8e3u, 0x18d67dbbu, ((sel >> 10u) & 1u) != 0u); // 55
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 56
r5 = r5 + r0 + select(0xf03673feu, 0xa75cd60du, ((sel >> 12u) & 1u) != 0u); // 57
r5 = r5 ^ dataset[((rotl_imm(r0 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 58
r1 = r1 + r4 + select(0xdecd4794u, 0x8dfb96bbu, ((sel >> 7u) & 1u) != 0u); // 59
r3 = mulhi(r3, r2); // 60
r6 = r6 | r4; // 61
r5 = r5 ^ simd_shuffle_xor(r4, (ushort)8); // 62
r3 = r3 ^ dataset[((rotl_imm(r6 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 63
}
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((ulong)hi << 32) | (ulong)lo;
}

View file

@ -0,0 +1,111 @@
#include <metal_stdlib>
using namespace metal;
#define MASK 0x0fffffffu
constant uint SEEDW[8] = { 0x667d0fbdu, 0x7b8e5963u, 0x31c67e5eu, 0x4529ddc6u, 0xef19d6d8u, 0xaccf6211u, 0xda0aed32u, 0xabc6df31u };
inline uint splitmix32(uint x) {
x ^= x >> 16; x *= 0x7feb352du;
x ^= x >> 15; x *= 0x846ca68bu;
x ^= x >> 16;
return x;
}
inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31
inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
inline uint ds_elem(uint i, uint d0, uint d1) {
uint x = i ^ d0;
x *= 0x9E3779B1u; x ^= x >> 15;
x += d1;
x *= 0x85EBCA77u; x ^= x >> 13;
x *= 0xC2B2AE3Du; x ^= x >> 16;
return x;
}
// Header-bound variant: the init words come from buffer 3 (bind.rs), not from SEEDW.
kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]],
device ulong* out [[buffer(1)]],
constant uint& baseNonce [[buffer(2)]],
constant uint* initw [[buffer(3)]],
uint gid [[thread_position_in_grid]]) {
uint nonce = baseNonce + gid;
uint r0, r1, r2, r3, r4, r5, r6, r7;
{ uint x = nonce ^ initw[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ initw[1]; }
{ uint x = nonce ^ initw[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ initw[2]; }
{ uint x = nonce ^ initw[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ initw[3]; }
{ uint x = nonce ^ initw[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ initw[4]; }
{ uint x = nonce ^ initw[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ initw[5]; }
{ uint x = nonce ^ initw[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ initw[6]; }
{ uint x = nonce ^ initw[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ initw[7]; }
{ uint x = nonce ^ initw[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ initw[0]; }
for (uint it = 0u; it < 8u; ++it) {
uint sel = r0;
r4 = r4 + r5 + select(0xea86e152u, 0x5810667au, ((sel >> 13u) & 1u) != 0u); // 0
r7 = r7 ^ r0; // 1
r3 = r3 ^ simd_shuffle_xor(r6, (ushort)1); // 2
r4 = r4 - r1; // 3
r2 = r0 * r4 + r2; // 4
r4 = rotl_imm(r4, 9u); // 5
r0 = rotr_var(r0, r2); // 6
r6 = r6 ^ dataset[((rotl_imm(r7 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x04000000u) & MASK]; // 7
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 8
r1 = r1 ^ dataset[((rotl_imm(r2 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 9
r7 = r7 ^ dataset[((rotl_imm(r0 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 10
r7 = r7 ^ dataset[((rotl_imm(r1 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 11
r5 = r5 * r4; // 12
r7 = r7 ^ dataset[((rotl_imm(r6 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 13
r6 = r6 ^ simd_shuffle_xor(r5, (ushort)16); // 14
r3 = r3 ^ simd_shuffle_xor(r5, (ushort)1); // 15
r2 = r2 | r7; // 16
r0 = rotr_var(r0, r6); // 17
r7 = r7 + r5 + select(0xf66e7017u, 0xb9e3577eu, ((sel >> 12u) & 1u) != 0u); // 18
r7 = rotl_imm(r7, 24u); // 19
r6 = r6 + r7 + select(0x52334d12u, 0x8c9f0ef8u, ((sel >> 23u) & 1u) != 0u); // 20
r2 = r2 | r6; // 21
r6 = r5 * r4 + r6; // 22
r1 = r1 | r0; // 23
r6 = r6 ^ r1; // 24
r2 = r2 + r6 + select(0x659fc3d3u, 0x9cec0e12u, ((sel >> 17u) & 1u) != 0u); // 25
r7 = rotr_var(r7, r0); // 26
r4 = r5 * r7 + r4; // 27
r3 = r2 * r4 + r3; // 28
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 29
r2 = r2 ^ dataset[((rotl_imm(r3 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x08000000u) & MASK]; // 30
r1 = r1 ^ dataset[((rotl_imm(r5 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 31
r7 = r7 + r2 + select(0x070888a8u, 0xe403240eu, ((sel >> 16u) & 1u) != 0u); // 32
r2 = r2 + r0 + select(0xf2e46d55u, 0x29701828u, ((sel >> 28u) & 1u) != 0u); // 33
r2 = r2 + r3 + select(0x58f75b87u, 0x343b7aeeu, ((sel >> 14u) & 1u) != 0u); // 34
r2 = mulhi(r2, r5); // 35
r4 = r4 ^ r2; // 36
r6 = r6 * r5; // 37
r7 = r7 ^ r0; // 38
r7 = r7 + r2 + select(0xb8180e9du, 0x32bbd117u, ((sel >> 19u) & 1u) != 0u); // 39
r2 = rotr_var(r2, r3); // 40
r7 = r7 - r0; // 41
r4 = r4 + r3 + select(0x6ced15b7u, 0x6df7aed4u, ((sel >> 4u) & 1u) != 0u); // 42
r7 = r7 ^ simd_shuffle_xor(r3, (ushort)4); // 43
r0 = r0 ^ dataset[((rotl_imm(r7 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 44
r1 = r1 + r6 + select(0xe09f54e9u, 0x83e825bfu, ((sel >> 14u) & 1u) != 0u); // 45
r3 = r3 ^ dataset[((rotl_imm(r1 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 46
r6 = r6 ^ dataset[((rotl_imm(r3 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 47
r4 = mulhi(r4, r2); // 48
r5 = r5 + r0 + select(0xa8bae6dfu, 0xf572bdb9u, ((sel >> 7u) & 1u) != 0u); // 49
r0 = r0 ^ simd_shuffle_xor(r7, (ushort)4); // 50
r6 = r6 + r0 + select(0x383b9260u, 0x11e17c61u, ((sel >> 19u) & 1u) != 0u); // 51
r5 = r5 ^ dataset[((rotl_imm(r2 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 52
r6 = rotl_imm(r6, 12u); // 53
r3 = rotl_imm(r3, 12u); // 54
r2 = r2 + r1 + select(0xac6be8e3u, 0x18d67dbbu, ((sel >> 10u) & 1u) != 0u); // 55
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 56
r5 = r5 + r0 + select(0xf03673feu, 0xa75cd60du, ((sel >> 12u) & 1u) != 0u); // 57
r5 = r5 ^ dataset[((rotl_imm(r0 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 58
r1 = r1 + r4 + select(0xdecd4794u, 0x8dfb96bbu, ((sel >> 7u) & 1u) != 0u); // 59
r3 = mulhi(r3, r2); // 60
r6 = r6 | r4; // 61
r5 = r5 ^ simd_shuffle_xor(r4, (ushort)8); // 62
r3 = r3 ^ dataset[((rotl_imm(r6 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 63
}
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((ulong)hi << 32) | (ulong)lo;
}

View file

@ -0,0 +1,5 @@
epoch_seed_hex edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07
day_seed_hex 69676e65756d2d6461792ffa50000000000000
epoch_index 0
day_index 20730
daa_score 0

View file

@ -0,0 +1,57 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v3, memory-hard dataset
#pragma once
#ifdef __cplusplus
#include <cstdint>
#else
#include <stdint.h>
#endif
#define IGNEUM_VEC_WARPS 3
static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u };
static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = {
{ // base nonce 0
0xa32ad2dd06264089ull, 0x854fea26361763aaull, 0xd958908d36ee63b5ull, 0x33641fe988c44ad9ull, 0xb5290b7892dbc7baull, 0xd815f329e6f806c0ull, 0x9fb28ff3fc29e39cull, 0xdbc28a5d8bf42ab4ull,
0x06f955808ff71bb9ull, 0x1ec10fff7bcdeaddull, 0x1acbc289cddd83cdull, 0x395e1564cda10137ull, 0x8a2aac2aa5b190d2ull, 0x499b8401e1f3e412ull, 0x9e2320319c85fc43ull, 0x66ba9661bc421ad6ull,
0x3ba69971f1c9d8bdull, 0x00d7a64d460bfefeull, 0xc0f6487a1fbc5978ull, 0x1a261ac8673e1146ull, 0xf5a95a7caaaafa0aull, 0x7bbdabd107842744ull, 0x49dd3bf9b77a16cfull, 0xf7026cc441495a05ull,
0x49fc48c1cba2be91ull, 0x84c2fc11de378cf9ull, 0x96de9e40c9da551cull, 0x3d20f01e363aaf61ull, 0xef36922a5fcf96caull, 0x31f1b0c4b0b42aafull, 0x33c6406d516d5992ull, 0x847af3a4248a972full
},
{ // base nonce 4096
0x2039f40a61003341ull, 0x8238174df3761142ull, 0x32a72a04c4ca4a3cull, 0x16aa0fd3f65da6b3ull, 0x502bf9db9bd98782ull, 0x109f60bc6167ff30ull, 0x3873c59792d6b552ull, 0xea98fba68fc32149ull,
0x56a8440169cb0f83ull, 0x7362edc3a7c98252ull, 0x8f393a7ad9f4ff40ull, 0x6b3cb3a1e0b02453ull, 0xdcbdc40aedfa06c0ull, 0x2bc381227147a2f3ull, 0x1570ef95b0c8e31full, 0xf3ca64d0f9ca8c91ull,
0x1c767a63cd6a0bf8ull, 0x8e1a13a1237a948bull, 0xff2241b0b53b3efdull, 0x95949c528978b1fbull, 0xce3519499b78db7eull, 0x56a0fcecaaa62096ull, 0x0a62c2bddc7041b2ull, 0x5ec463510d31d7c0ull,
0x2fc74273ed47b17full, 0xedc5740b0c5c191eull, 0xe62f737c106216cbull, 0x185df496808afd70ull, 0x29070790e64a0bc7ull, 0xc0ddaeac8b22b14bull, 0x302982837b6d96b4ull, 0x523972727aa906b5ull
},
{ // base nonce 1000000
0x3527292a5f4afb4eull, 0x8c4128a03d946030ull, 0xb5b598dd18eb207dull, 0x9f1b21b5bf9dce8full, 0xa791a53fa1a4733eull, 0x12d2f9629aaca3b8ull, 0x400e844a00ba254cull, 0x1e9f98bee0272859ull,
0x9983064542aabb54ull, 0x017873e17fab719cull, 0xf3e747e327968d2bull, 0x562380f8016b9ca7ull, 0x3f5007c1ec21b4a8ull, 0x5612489676e34c2bull, 0x97e13b52bc125dabull, 0xe15cf6d8d12f9b00ull,
0xd8ba508db7b12919ull, 0x2e16d66db0bd1837ull, 0x0026e228553d7b18ull, 0x0c6d22f5018958ddull, 0xe76da566fc36580dull, 0x7824813a77faff44ull, 0x663811c5504ff04dull, 0x75efd5493353d65aull,
0x926b72ebe963caa2ull, 0x392723945438b76dull, 0x4df61663eb2bfe05ull, 0x9a305821ba9ceb59ull, 0x317f4a6ac3f65a88ull, 0xdbae9b884f6829a9ull, 0x97cc8c8866b1a0cbull, 0x5db142b6d7777df1ull
}
};
// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455).
static const uint32_t IGNEUM_DS_HEAD[16] = {
0x3dd50b1fu, 0x48edec90u, 0x93123701u, 0x3a7d2407u, 0x5119540eu, 0x657ac748u, 0x3358ef2du, 0x09e2105cu,
0xcdd88b45u, 0xb03a0ac0u, 0x2c37c6b0u, 0x10e2c6a5u, 0x25e2929cu, 0x4317cabdu, 0x5fa0cdccu, 0x6178ee52u
};
static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u;
static const uint32_t IGNEUM_DS_LAST = 0xf7b7180eu;
// 64 sampled dataset words (index, value) computed on the Mac.
#define IGNEUM_DS_SAMPLES 64
static const uint32_t IGNEUM_DS_SAMPLE_INDEX[IGNEUM_DS_SAMPLES] = {
59471966u, 217795994u, 208353206u, 42483309u, 172547758u, 148076330u, 183853158u, 214389424u, 267488061u, 169781097u, 184093494u, 153880993u, 84977930u, 46426879u, 3093825u, 225364072u, 44593546u, 260713159u, 168250303u, 52384140u, 223401610u, 45554030u, 95410555u, 175039924u, 79171087u, 267580473u, 24168642u, 37981670u, 171551130u, 195559979u, 204611762u, 140997658u, 138925853u, 86637313u, 20736778u, 219665210u, 160430336u, 264654675u, 8013395u, 228945585u, 213884386u, 104419827u, 44185464u, 142737231u, 99284897u, 132475900u, 61861762u, 132056166u, 262388043u, 91878046u, 117353561u, 124768597u, 71352993u, 190698941u, 46055428u, 55281366u, 165145231u, 106810753u, 171985651u, 232085256u, 159510492u, 40072060u, 209107596u, 39023794u
};
static const uint32_t IGNEUM_DS_SAMPLE_VALUE[IGNEUM_DS_SAMPLES] = {
0x1faef64au, 0x068e4a54u, 0x24551c62u, 0x1234f539u, 0xc15100f5u, 0x2f53b6b0u, 0xbab5fb76u, 0x06edf897u, 0xb0880fc6u, 0xccfddc4du, 0xf640c0f4u, 0xf145a054u, 0xa2425e8fu, 0x0a218118u, 0xd310512au, 0xb612d7ffu, 0x9a290b4du, 0xaf0227edu, 0x81b9c33cu, 0x5f352daeu, 0x46889c3du, 0x8ddc5d17u, 0x368b4a6cu, 0x16acb57au, 0x103f86c1u, 0x7144efafu, 0x87cd0dc1u, 0x1d9132b1u, 0x28b0d87cu, 0x01c44610u, 0x908b6bacu, 0x9e127820u, 0x9e695a6fu, 0xfbb80109u, 0x39ad9ba2u, 0x0c867f51u, 0xac1e69e5u, 0xf16f763du, 0x2de60ff0u, 0xfabaf858u, 0x53a35511u, 0x249e753fu, 0x88f84300u, 0xd79db283u, 0x7ac395cau, 0x5df23d8bu, 0xe3f98318u, 0x33ad1217u, 0x2dd19e3fu, 0xf1a97b54u, 0xb1fabd1eu, 0x1bc27b43u, 0xb208c2e5u, 0x1f109813u, 0xccb6a56bu, 0x64b0398eu, 0x9ac3815eu, 0x14970614u, 0x6139a686u, 0x6a5eeda3u, 0xbb567051u, 0xd2cd3e4du, 0x08c46214u, 0xc15f2c97u
};
// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^26 words.
static const uint32_t IGNEUM_CACHE_HEAD[16] = {
0xebd9055cu, 0x7eaf6b21u, 0x9b610855u, 0x134a8562u, 0xb10059bau, 0x7f459e58u, 0x8f3c7873u, 0x3523e609u,
0x75b8f4cau, 0x750d31b4u, 0x0dae0782u, 0x26f10015u, 0x7ba87a41u, 0x0efd8543u, 0x691d6368u, 0x8929d967u
};
static const uint32_t IGNEUM_CACHE_LAST[16] = {
0x51474edfu, 0xc3cfba93u, 0xf21454cfu, 0x9b79baacu, 0x4d4fce23u, 0xd701dfd5u, 0x37357ab3u, 0x1be693fau,
0xa701cb3bu, 0x7467c620u, 0x428184e8u, 0xf4010df0u, 0xe33aa88fu, 0xc78d6d62u, 0xae9ba9c0u, 0xf4bba97eu
};
static const uint64_t IGNEUM_CACHE_FNV64 = 0x448274a57f508cbcull;

View file

@ -0,0 +1,36 @@
{
"seed": "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000",
"day": "bytes:69676e65756d2d6461792ffa50000000000000",
"dataset_mode": "memory-hard",
"dataset_log2_words": 28,
"mask": "0x0fffffff",
"lanes": 32,
"source": "igneum-pow (Rust) CPU interpreter, generator v3, memory-hard dataset",
"warps": [
{"base_nonce": 0, "expected": [
"0xa32ad2dd06264089", "0x854fea26361763aa", "0xd958908d36ee63b5", "0x33641fe988c44ad9", "0xb5290b7892dbc7ba", "0xd815f329e6f806c0", "0x9fb28ff3fc29e39c", "0xdbc28a5d8bf42ab4",
"0x06f955808ff71bb9", "0x1ec10fff7bcdeadd", "0x1acbc289cddd83cd", "0x395e1564cda10137", "0x8a2aac2aa5b190d2", "0x499b8401e1f3e412", "0x9e2320319c85fc43", "0x66ba9661bc421ad6",
"0x3ba69971f1c9d8bd", "0x00d7a64d460bfefe", "0xc0f6487a1fbc5978", "0x1a261ac8673e1146", "0xf5a95a7caaaafa0a", "0x7bbdabd107842744", "0x49dd3bf9b77a16cf", "0xf7026cc441495a05",
"0x49fc48c1cba2be91", "0x84c2fc11de378cf9", "0x96de9e40c9da551c", "0x3d20f01e363aaf61", "0xef36922a5fcf96ca", "0x31f1b0c4b0b42aaf", "0x33c6406d516d5992", "0x847af3a4248a972f"
]},
{"base_nonce": 4096, "expected": [
"0x2039f40a61003341", "0x8238174df3761142", "0x32a72a04c4ca4a3c", "0x16aa0fd3f65da6b3", "0x502bf9db9bd98782", "0x109f60bc6167ff30", "0x3873c59792d6b552", "0xea98fba68fc32149",
"0x56a8440169cb0f83", "0x7362edc3a7c98252", "0x8f393a7ad9f4ff40", "0x6b3cb3a1e0b02453", "0xdcbdc40aedfa06c0", "0x2bc381227147a2f3", "0x1570ef95b0c8e31f", "0xf3ca64d0f9ca8c91",
"0x1c767a63cd6a0bf8", "0x8e1a13a1237a948b", "0xff2241b0b53b3efd", "0x95949c528978b1fb", "0xce3519499b78db7e", "0x56a0fcecaaa62096", "0x0a62c2bddc7041b2", "0x5ec463510d31d7c0",
"0x2fc74273ed47b17f", "0xedc5740b0c5c191e", "0xe62f737c106216cb", "0x185df496808afd70", "0x29070790e64a0bc7", "0xc0ddaeac8b22b14b", "0x302982837b6d96b4", "0x523972727aa906b5"
]},
{"base_nonce": 1000000, "expected": [
"0x3527292a5f4afb4e", "0x8c4128a03d946030", "0xb5b598dd18eb207d", "0x9f1b21b5bf9dce8f", "0xa791a53fa1a4733e", "0x12d2f9629aaca3b8", "0x400e844a00ba254c", "0x1e9f98bee0272859",
"0x9983064542aabb54", "0x017873e17fab719c", "0xf3e747e327968d2b", "0x562380f8016b9ca7", "0x3f5007c1ec21b4a8", "0x5612489676e34c2b", "0x97e13b52bc125dab", "0xe15cf6d8d12f9b00",
"0xd8ba508db7b12919", "0x2e16d66db0bd1837", "0x0026e228553d7b18", "0x0c6d22f5018958dd", "0xe76da566fc36580d", "0x7824813a77faff44", "0x663811c5504ff04d", "0x75efd5493353d65a",
"0x926b72ebe963caa2", "0x392723945438b76d", "0x4df61663eb2bfe05", "0x9a305821ba9ceb59", "0x317f4a6ac3f65a88", "0xdbae9b884f6829a9", "0x97cc8c8866b1a0cb", "0x5db142b6d7777df1"
]}
],
"dataset_head": ["0x3dd50b1f", "0x48edec90", "0x93123701", "0x3a7d2407", "0x5119540e", "0x657ac748", "0x3358ef2d", "0x09e2105c", "0xcdd88b45", "0xb03a0ac0", "0x2c37c6b0", "0x10e2c6a5", "0x25e2929c", "0x4317cabd", "0x5fa0cdcc", "0x6178ee52"],
"dataset_last_index": 268435455,
"dataset_last": "0xf7b7180e",
"dataset_samples": [{"index": 59471966, "value": "0x1faef64a"}, {"index": 217795994, "value": "0x068e4a54"}, {"index": 208353206, "value": "0x24551c62"}, {"index": 42483309, "value": "0x1234f539"}, {"index": 172547758, "value": "0xc15100f5"}, {"index": 148076330, "value": "0x2f53b6b0"}, {"index": 183853158, "value": "0xbab5fb76"}, {"index": 214389424, "value": "0x06edf897"}, {"index": 267488061, "value": "0xb0880fc6"}, {"index": 169781097, "value": "0xccfddc4d"}, {"index": 184093494, "value": "0xf640c0f4"}, {"index": 153880993, "value": "0xf145a054"}, {"index": 84977930, "value": "0xa2425e8f"}, {"index": 46426879, "value": "0x0a218118"}, {"index": 3093825, "value": "0xd310512a"}, {"index": 225364072, "value": "0xb612d7ff"}, {"index": 44593546, "value": "0x9a290b4d"}, {"index": 260713159, "value": "0xaf0227ed"}, {"index": 168250303, "value": "0x81b9c33c"}, {"index": 52384140, "value": "0x5f352dae"}, {"index": 223401610, "value": "0x46889c3d"}, {"index": 45554030, "value": "0x8ddc5d17"}, {"index": 95410555, "value": "0x368b4a6c"}, {"index": 175039924, "value": "0x16acb57a"}, {"index": 79171087, "value": "0x103f86c1"}, {"index": 267580473, "value": "0x7144efaf"}, {"index": 24168642, "value": "0x87cd0dc1"}, {"index": 37981670, "value": "0x1d9132b1"}, {"index": 171551130, "value": "0x28b0d87c"}, {"index": 195559979, "value": "0x01c44610"}, {"index": 204611762, "value": "0x908b6bac"}, {"index": 140997658, "value": "0x9e127820"}, {"index": 138925853, "value": "0x9e695a6f"}, {"index": 86637313, "value": "0xfbb80109"}, {"index": 20736778, "value": "0x39ad9ba2"}, {"index": 219665210, "value": "0x0c867f51"}, {"index": 160430336, "value": "0xac1e69e5"}, {"index": 264654675, "value": "0xf16f763d"}, {"index": 8013395, "value": "0x2de60ff0"}, {"index": 228945585, "value": "0xfabaf858"}, {"index": 213884386, "value": "0x53a35511"}, {"index": 104419827, "value": "0x249e753f"}, {"index": 44185464, "value": "0x88f84300"}, {"index": 142737231, "value": "0xd79db283"}, {"index": 99284897, "value": "0x7ac395ca"}, {"index": 132475900, "value": "0x5df23d8b"}, {"index": 61861762, "value": "0xe3f98318"}, {"index": 132056166, "value": "0x33ad1217"}, {"index": 262388043, "value": "0x2dd19e3f"}, {"index": 91878046, "value": "0xf1a97b54"}, {"index": 117353561, "value": "0xb1fabd1e"}, {"index": 124768597, "value": "0x1bc27b43"}, {"index": 71352993, "value": "0xb208c2e5"}, {"index": 190698941, "value": "0x1f109813"}, {"index": 46055428, "value": "0xccb6a56b"}, {"index": 55281366, "value": "0x64b0398e"}, {"index": 165145231, "value": "0x9ac3815e"}, {"index": 106810753, "value": "0x14970614"}, {"index": 171985651, "value": "0x6139a686"}, {"index": 232085256, "value": "0x6a5eeda3"}, {"index": 159510492, "value": "0xbb567051"}, {"index": 40072060, "value": "0xd2cd3e4d"}, {"index": 209107596, "value": "0x08c46214"}, {"index": 39023794, "value": "0xc15f2c97"}],
"cache_head": ["0xebd9055c", "0x7eaf6b21", "0x9b610855", "0x134a8562", "0xb10059ba", "0x7f459e58", "0x8f3c7873", "0x3523e609", "0x75b8f4ca", "0x750d31b4", "0x0dae0782", "0x26f10015", "0x7ba87a41", "0x0efd8543", "0x691d6368", "0x8929d967"],
"cache_last_line": ["0x51474edf", "0xc3cfba93", "0xf21454cf", "0x9b79baac", "0x4d4fce23", "0xd701dfd5", "0x37357ab3", "0x1be693fa", "0xa701cb3b", "0x7467c620", "0x428184e8", "0xf4010df0", "0xe33aa88f", "0xc78d6d62", "0xae9ba9c0", "0xf4bba97e"],
"cache_fnv1a64": "0x448274a57f508cbc"
}

View file

@ -0,0 +1,281 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal).
// Built from source at runtime by proto-opencl/host.c, which passes these defines:
// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit)
// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default)
// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32
// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition
// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units;
// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32.
#ifndef IGNEUM_GROUP
#define IGNEUM_GROUP 32
#endif
#ifndef IGNEUM_EXCHANGE
#define IGNEUM_EXCHANGE 0
#endif
#ifdef __OPENCL_VERSION__
#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1)))
#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n]
#if IGNEUM_EXCHANGE == 1
#ifdef cl_khr_subgroups
#pragma OPENCL EXTENSION cl_khr_subgroups : enable
#endif
#ifdef cl_khr_subgroup_shuffle
#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable
#endif
#elif IGNEUM_EXCHANGE == 2
#pragma OPENCL EXTENSION cl_intel_subgroups : enable
#endif
#else
// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros.
#include "emu_opencl.h"
#endif
#if IGNEUM_EXCHANGE == 1
#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m))
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
#elif IGNEUM_EXCHANGE == 2
#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m))
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
#else
// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per
// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane
// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's
// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier.
#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; }
#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; }
#endif
static inline uint splitmix32(uint x) {
x ^= x >> 16; x *= 0x7feb352du;
x ^= x >> 15; x *= 0x846ca68bu;
x ^= x >> 16;
return x;
}
// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32.
static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); }
// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x.
static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); }
static inline uint ds_elem(uint i, uint d0, uint d1) {
uint x = i ^ d0;
x *= 0x9E3779B1u; x ^= x >> 15;
x += d1;
x *= 0x85EBCA77u; x ^= x >> 13;
x *= 0xC2B2AE3Du; x ^= x >> 16;
return x;
}
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
#define MH_CACHE_LINE_MASK 0x003fffffu
#define MH_SEGMENT_LINES 64u
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
// y = ChaCha12 core(x) + x
static inline void mh_chacha_block(const uint* x, uint* y) {
for (uint i = 0u; i < 16u; ++i) y[i] = x[i];
for (uint r = 0u; r < 6u; ++r) {
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
}
for (uint i = 0u; i < 16u; ++i) y[i] += x[i];
}
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
static inline void mh_cache_segment(__global uint* cache, uint seg) {
uint prev[16]; uint x[16]; uint y[16];
for (uint i = 0u; i < 16u; ++i) prev[i] = 0u;
for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) {
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
x[4] = 0xceed56d7u ^ prev[4];
x[5] = 0x9ba270d2u ^ prev[5];
x[6] = 0x82caab2du ^ prev[6];
x[7] = 0x81ebce0eu ^ prev[7];
x[8] = 0x12b6ecf1u ^ prev[8];
x[9] = 0xd0f3fd7cu ^ prev[9];
x[10] = 0xd872eefeu ^ prev[10];
x[11] = 0xc158c7bdu ^ prev[11];
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
mh_chacha_block(x, y);
__global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
}
}
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
static inline void mh_mixer(uint* s, uint rk) {
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
}
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
static inline void mh_item(__global const uint* cache, uint t, uint* s) {
s[0] = 0xceed56d7u;
s[1] = 0x9ba270d2u;
s[2] = 0x82caab2du;
s[3] = 0x81ebce0eu;
s[4] = 0x12b6ecf1u;
s[5] = 0xd0f3fd7cu;
s[6] = 0xd872eefeu;
s[7] = 0xc158c7bdu;
s[8] = t * 0xf351d601u + 0xc6892460u;
s[9] = t * 0xa3bb398fu + 0x25b7228au;
s[10] = t * 0xb5a09e35u + 0xcd515004u;
s[11] = t * 0x7509c9c1u + 0x2846527au;
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
s[14] = t * 0xded91851u + 0x82961bacu;
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
for (uint r = 0u; r < 8u; ++r) {
mh_mixer(s, 0x9E3779B9u * (r + 1u));
__global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i];
}
mh_mixer(s, 0x9E3779B9u * 9u);
}
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 1 3 8 13 of w.
static inline uint mh_j(uint w) { return ((w >> 1u) & 1u) | (((w >> 3u) & 1u) << 1) | (((w >> 8u) & 1u) << 2) | (((w >> 13u) & 1u) << 3); }
static inline uint mh_t(uint w) { w = (w & 0x00001fffu) | ((w >> 14u) << 13u); w = (w & 0x000000ffu) | ((w >> 9u) << 8u); w = (w & 0x00000007u) | ((w >> 4u) << 3u); w = (w & 0x00000001u) | ((w >> 2u) << 1u); return w; }
static inline uint mh_addr(uint t, uint j) { uint w = t; w = ((w >> 1u) << 2u) | (w & 0x00000001u) | (((j >> 0u) & 1u) << 1u); w = ((w >> 3u) << 4u) | (w & 0x00000007u) | (((j >> 1u) & 1u) << 3u); w = ((w >> 8u) << 9u) | (w & 0x000000ffu) | (((j >> 2u) & 1u) << 8u); w = ((w >> 13u) << 14u) | (w & 0x00001fffu) | (((j >> 3u) & 1u) << 13u); return w; }
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }
// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item.
// The same constants as memhard.h in this pack (one emitter, three dialects).
__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) {
uint seg = (uint)get_global_id(0);
if (seg < nSegments) mh_cache_segment(cache, seg);
}
__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) {
uint t = (uint)get_global_id(0);
if (t < nItems) {
uint s[16];
mh_item(cache, t, s);
for (uint i = 0u; i < 16u; ++i) ds[(ulong)mh_addr(t, i)] = s[i];
}
}
// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the
// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and
// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all).
IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) {
uint gid = (uint)get_global_id(0);
uint lid = (uint)get_local_id(0);
uint nonce = baseNonce + gid;
uint r0, r1, r2, r3, r4, r5, r6, r7;
#if IGNEUM_EXCHANGE == 0
IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
uint xk = 0u;
#else
(void)lid;
#endif
{ uint x = nonce ^ 0x667d0fbdu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x7b8e5963u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
{ uint x = nonce ^ 0x7b8e5963u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x31c67e5eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
{ uint x = nonce ^ 0x31c67e5eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4529ddc6u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
{ uint x = nonce ^ 0x4529ddc6u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef19d6d8u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
{ uint x = nonce ^ 0xef19d6d8u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xaccf6211u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
{ uint x = nonce ^ 0xaccf6211u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xda0aed32u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
{ uint x = nonce ^ 0xda0aed32u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xabc6df31u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
{ uint x = nonce ^ 0xabc6df31u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x667d0fbdu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
for (uint it = 0u; it < 8u; ++it) {
uint sel = r0;
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
r7 = r7 ^ r0; // 1 xor
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 1u); r3 = r3 ^ t_; } // 2 shfl
r4 = r4 - r1; // 3 sub
r2 = r0 * r4 + r2; // 4 mad
r4 = rotl_imm(r4, 9u); // 5 rotl
r0 = rotr_var(r0, r2); // 6 rotr
r6 = r6 ^ ds[((rotl_imm(r7 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
r1 = r1 ^ ds[((rotl_imm(r4 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
r1 = r1 ^ ds[((rotl_imm(r2 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
r7 = r7 ^ ds[((rotl_imm(r0 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
r7 = r7 ^ ds[((rotl_imm(r1 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
r5 = r5 * r4; // 12 mul
r7 = r7 ^ ds[((rotl_imm(r6 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 16u); r6 = r6 ^ t_; } // 14 shfl
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r3 = r3 ^ t_; } // 15 shfl
r2 = r2 | r7; // 16 or
r0 = rotr_var(r0, r6); // 17 rotr
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
r7 = rotl_imm(r7, 24u); // 19 rotl
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
r2 = r2 | r6; // 21 or
r6 = r5 * r4 + r6; // 22 mad
r1 = r1 | r0; // 23 or
r6 = r6 ^ r1; // 24 xor
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
r7 = rotr_var(r7, r0); // 26 rotr
r4 = r5 * r7 + r4; // 27 mad
r3 = r2 * r4 + r3; // 28 mad
r1 = r1 ^ ds[((rotl_imm(r4 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
r2 = r2 ^ ds[((rotl_imm(r3 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
r1 = r1 ^ ds[((rotl_imm(r5 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
r2 = mul_hi(r2, r5); // 35 mulhi
r4 = r4 ^ r2; // 36 xor
r6 = r6 * r5; // 37 mul
r7 = r7 ^ r0; // 38 xor
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
r2 = rotr_var(r2, r3); // 40 rotr
r7 = r7 - r0; // 41 sub
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 43 shfl
r0 = r0 ^ ds[((rotl_imm(r7 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
r3 = r3 ^ ds[((rotl_imm(r1 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
r6 = r6 ^ ds[((rotl_imm(r3 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
r4 = mul_hi(r4, r2); // 48 mulhi
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r0 = r0 ^ t_; } // 50 shfl
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
r5 = r5 ^ ds[((rotl_imm(r2 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
r6 = rotl_imm(r6, 12u); // 53 rotl
r3 = rotl_imm(r3, 12u); // 54 rotl
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
r1 = r1 ^ ds[((rotl_imm(r4 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
r5 = r5 ^ ds[((rotl_imm(r0 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
r3 = mul_hi(r3, r2); // 60 mulhi
r6 = r6 | r4; // 61 or
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r5 = r5 ^ t_; } // 62 shfl
r3 = r3 ^ ds[((rotl_imm(r6 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
}
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((ulong)hi << 32) | (ulong)lo;
}
#if IGNEUM_EXCHANGE != 0
// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the
// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a
// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md.
IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) {
if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); }
}
#endif

View file

@ -0,0 +1,163 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal).
// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC.
#include <cuda_runtime.h>
#include <cstdint>
#include "program.h"
#include "memhard.h"
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
x ^= x >> 16; x *= 0x7feb352du;
x ^= x >> 15; x *= 0x846ca68bu;
x ^= x >> 16;
return x;
}
// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31.
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x.
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) {
uint32_t x = i ^ d0;
x *= 0x9E3779B1u; x ^= x >> 15;
x += d1;
x *= 0x85EBCA77u; x ^= x >> 13;
x *= 0xC2B2AE3Du; x ^= x >> 16;
return x;
}
// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item.
// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host.
__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) {
uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x;
if (seg < nSegments) mh_cache_segment(cache, seg);
}
__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
uint32_t t = blockIdx.x * blockDim.x + threadIdx.x;
if (t < nItems) {
uint32_t s[16];
mh_item(cache, t, s);
for (uint32_t i = 0u; i < 16u; ++i) ds[(size_t)mh_addr(t, i)] = s[i];
}
}
// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every
// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a
// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid.
__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) {
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
uint32_t nonce = baseNonce + gid;
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
{ uint32_t x = nonce ^ 0x667d0fbdu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x7b8e5963u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
{ uint32_t x = nonce ^ 0x7b8e5963u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x31c67e5eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
{ uint32_t x = nonce ^ 0x31c67e5eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4529ddc6u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
{ uint32_t x = nonce ^ 0x4529ddc6u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef19d6d8u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
{ uint32_t x = nonce ^ 0xef19d6d8u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xaccf6211u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
{ uint32_t x = nonce ^ 0xaccf6211u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xda0aed32u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
{ uint32_t x = nonce ^ 0xda0aed32u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xabc6df31u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
{ uint32_t x = nonce ^ 0xabc6df31u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x667d0fbdu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
for (uint32_t it = 0u; it < 8u; ++it) {
uint32_t sel = r0;
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
r7 = r7 ^ r0; // 1 xor
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r6, 1); // 2 shfl
r4 = r4 - r1; // 3 sub
r2 = r0 * r4 + r2; // 4 mad
r4 = rotl_imm(r4, 9u); // 5 rotl
r0 = rotr_var(r0, r2); // 6 rotr
r6 = r6 ^ ds[((rotl_imm(r7 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
r1 = r1 ^ ds[((rotl_imm(r4 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
r1 = r1 ^ ds[((rotl_imm(r2 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
r7 = r7 ^ ds[((rotl_imm(r0 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
r7 = r7 ^ ds[((rotl_imm(r1 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
r5 = r5 * r4; // 12 mul
r7 = r7 ^ ds[((rotl_imm(r6 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r5, 16); // 14 shfl
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r5, 1); // 15 shfl
r2 = r2 | r7; // 16 or
r0 = rotr_var(r0, r6); // 17 rotr
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
r7 = rotl_imm(r7, 24u); // 19 rotl
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
r2 = r2 | r6; // 21 or
r6 = r5 * r4 + r6; // 22 mad
r1 = r1 | r0; // 23 or
r6 = r6 ^ r1; // 24 xor
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
r7 = rotr_var(r7, r0); // 26 rotr
r4 = r5 * r7 + r4; // 27 mad
r3 = r2 * r4 + r3; // 28 mad
r1 = r1 ^ ds[((rotl_imm(r4 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
r2 = r2 ^ ds[((rotl_imm(r3 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
r1 = r1 ^ ds[((rotl_imm(r5 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
r2 = __umulhi(r2, r5); // 35 mulhi
r4 = r4 ^ r2; // 36 xor
r6 = r6 * r5; // 37 mul
r7 = r7 ^ r0; // 38 xor
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
r2 = rotr_var(r2, r3); // 40 rotr
r7 = r7 - r0; // 41 sub
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 43 shfl
r0 = r0 ^ ds[((rotl_imm(r7 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
r3 = r3 ^ ds[((rotl_imm(r1 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
r6 = r6 ^ ds[((rotl_imm(r3 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
r4 = __umulhi(r4, r2); // 48 mulhi
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r7, 4); // 50 shfl
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
r5 = r5 ^ ds[((rotl_imm(r2 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
r6 = rotl_imm(r6, 12u); // 53 rotl
r3 = rotl_imm(r3, 12u); // 54 rotl
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
r1 = r1 ^ ds[((rotl_imm(r4 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
r5 = r5 ^ ds[((rotl_imm(r0 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
r3 = __umulhi(r3, r2); // 60 mulhi
r6 = r6 | r4; // 61 or
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 62 shfl
r3 = r3 ^ ds[((rotl_imm(r6 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
}
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
}
// Host-side launch wrappers. Declared in program.h, called from host.cu.
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) {
if (nSegments == 0u) return cudaErrorInvalidValue;
uint32_t block = 256u;
uint32_t grid = (nSegments + block - 1u) / block;
igneum_cache_fill<<<grid, block>>>(cache, nSegments);
return cudaGetLastError();
}
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
if (nItems == 0u) return cudaErrorInvalidValue;
uint32_t block = 256u;
uint32_t grid = (nItems + block - 1u) / block;
igneum_build<<<grid, block>>>(ds, cache, nItems);
return cudaGetLastError();
}
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
uint32_t nonces, uint32_t blockWarps) {
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
uint32_t block = 32u * blockWarps;
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
igneum_hash<<<nonces / block, block>>>(ds, out, baseNonce, mask);
return cudaGetLastError();
}
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
cudaFuncAttributes attr;
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash);
if (e != cudaSuccess) return e;
*numRegs = attr.numRegs;
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0);
}

View file

@ -0,0 +1,375 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal).
// Built from source at runtime by proto-opencl/host.c, which passes these defines:
// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit)
// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default)
// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32
// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition
// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units;
// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32.
#ifndef IGNEUM_GROUP
#define IGNEUM_GROUP 32
#endif
#ifndef IGNEUM_EXCHANGE
#define IGNEUM_EXCHANGE 0
#endif
#ifdef __OPENCL_VERSION__
#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1)))
#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n]
#if IGNEUM_EXCHANGE == 1
#ifdef cl_khr_subgroups
#pragma OPENCL EXTENSION cl_khr_subgroups : enable
#endif
#ifdef cl_khr_subgroup_shuffle
#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable
#endif
#elif IGNEUM_EXCHANGE == 2
#pragma OPENCL EXTENSION cl_intel_subgroups : enable
#endif
#else
// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros.
#include "emu_opencl.h"
#endif
#if IGNEUM_EXCHANGE == 1
#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m))
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
#elif IGNEUM_EXCHANGE == 2
#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m))
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
#else
// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per
// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane
// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's
// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier.
#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; }
#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; }
#endif
static inline uint splitmix32(uint x) {
x ^= x >> 16; x *= 0x7feb352du;
x ^= x >> 15; x *= 0x846ca68bu;
x ^= x >> 16;
return x;
}
// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32.
static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); }
// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x.
static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); }
static inline uint ds_elem(uint i, uint d0, uint d1) {
uint x = i ^ d0;
x *= 0x9E3779B1u; x ^= x >> 15;
x += d1;
x *= 0x85EBCA77u; x ^= x >> 13;
x *= 0xC2B2AE3Du; x ^= x >> 16;
return x;
}
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
#define MH_CACHE_LINE_MASK 0x003fffffu
#define MH_SEGMENT_LINES 64u
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
// y = ChaCha12 core(x) + x
static inline void mh_chacha_block(const uint* x, uint* y) {
for (uint i = 0u; i < 16u; ++i) y[i] = x[i];
for (uint r = 0u; r < 6u; ++r) {
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
}
for (uint i = 0u; i < 16u; ++i) y[i] += x[i];
}
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
static inline void mh_cache_segment(__global uint* cache, uint seg) {
uint prev[16]; uint x[16]; uint y[16];
for (uint i = 0u; i < 16u; ++i) prev[i] = 0u;
for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) {
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
x[4] = 0xceed56d7u ^ prev[4];
x[5] = 0x9ba270d2u ^ prev[5];
x[6] = 0x82caab2du ^ prev[6];
x[7] = 0x81ebce0eu ^ prev[7];
x[8] = 0x12b6ecf1u ^ prev[8];
x[9] = 0xd0f3fd7cu ^ prev[9];
x[10] = 0xd872eefeu ^ prev[10];
x[11] = 0xc158c7bdu ^ prev[11];
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
mh_chacha_block(x, y);
__global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
}
}
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
static inline void mh_mixer(uint* s, uint rk) {
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
}
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
static inline void mh_item(__global const uint* cache, uint t, uint* s) {
s[0] = 0xceed56d7u;
s[1] = 0x9ba270d2u;
s[2] = 0x82caab2du;
s[3] = 0x81ebce0eu;
s[4] = 0x12b6ecf1u;
s[5] = 0xd0f3fd7cu;
s[6] = 0xd872eefeu;
s[7] = 0xc158c7bdu;
s[8] = t * 0xf351d601u + 0xc6892460u;
s[9] = t * 0xa3bb398fu + 0x25b7228au;
s[10] = t * 0xb5a09e35u + 0xcd515004u;
s[11] = t * 0x7509c9c1u + 0x2846527au;
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
s[14] = t * 0xded91851u + 0x82961bacu;
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
for (uint r = 0u; r < 8u; ++r) {
mh_mixer(s, 0x9E3779B9u * (r + 1u));
__global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i];
}
mh_mixer(s, 0x9E3779B9u * 9u);
}
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 1 3 8 13 of w.
static inline uint mh_j(uint w) { return ((w >> 1u) & 1u) | (((w >> 3u) & 1u) << 1) | (((w >> 8u) & 1u) << 2) | (((w >> 13u) & 1u) << 3); }
static inline uint mh_t(uint w) { w = (w & 0x00001fffu) | ((w >> 14u) << 13u); w = (w & 0x000000ffu) | ((w >> 9u) << 8u); w = (w & 0x00000007u) | ((w >> 4u) << 3u); w = (w & 0x00000001u) | ((w >> 2u) << 1u); return w; }
static inline uint mh_addr(uint t, uint j) { uint w = t; w = ((w >> 1u) << 2u) | (w & 0x00000001u) | (((j >> 0u) & 1u) << 1u); w = ((w >> 3u) << 4u) | (w & 0x00000007u) | (((j >> 1u) & 1u) << 3u); w = ((w >> 8u) << 9u) | (w & 0x000000ffu) | (((j >> 2u) & 1u) << 8u); w = ((w >> 13u) << 14u) | (w & 0x00001fffu) | (((j >> 3u) & 1u) << 13u); return w; }
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }
// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item.
// The same constants as memhard.h in this pack (one emitter, three dialects).
__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) {
uint seg = (uint)get_global_id(0);
if (seg < nSegments) mh_cache_segment(cache, seg);
}
__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) {
uint t = (uint)get_global_id(0);
if (t < nItems) {
uint s[16];
mh_item(cache, t, s);
for (uint i = 0u; i < 16u; ++i) ds[(ulong)mh_addr(t, i)] = s[i];
}
}
// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the
// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and
// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all).
IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) {
uint gid = (uint)get_global_id(0);
uint lid = (uint)get_local_id(0);
uint nonce = baseNonce + gid;
uint r0, r1, r2, r3, r4, r5, r6, r7;
#if IGNEUM_EXCHANGE == 0
IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
uint xk = 0u;
#else
(void)lid;
#endif
{ uint x = nonce ^ 0x667d0fbdu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x7b8e5963u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
{ uint x = nonce ^ 0x7b8e5963u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x31c67e5eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
{ uint x = nonce ^ 0x31c67e5eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4529ddc6u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
{ uint x = nonce ^ 0x4529ddc6u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef19d6d8u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
{ uint x = nonce ^ 0xef19d6d8u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xaccf6211u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
{ uint x = nonce ^ 0xaccf6211u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xda0aed32u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
{ uint x = nonce ^ 0xda0aed32u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xabc6df31u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
{ uint x = nonce ^ 0xabc6df31u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x667d0fbdu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
for (uint it = 0u; it < 8u; ++it) {
uint sel = r0;
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
r7 = r7 ^ r0; // 1 xor
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 1u); r3 = r3 ^ t_; } // 2 shfl
r4 = r4 - r1; // 3 sub
r2 = r0 * r4 + r2; // 4 mad
r4 = rotl_imm(r4, 9u); // 5 rotl
r0 = rotr_var(r0, r2); // 6 rotr
r6 = r6 ^ ds[((rotl_imm(r7 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
r1 = r1 ^ ds[((rotl_imm(r4 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
r1 = r1 ^ ds[((rotl_imm(r2 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
r7 = r7 ^ ds[((rotl_imm(r0 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
r7 = r7 ^ ds[((rotl_imm(r1 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
r5 = r5 * r4; // 12 mul
r7 = r7 ^ ds[((rotl_imm(r6 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 16u); r6 = r6 ^ t_; } // 14 shfl
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r3 = r3 ^ t_; } // 15 shfl
r2 = r2 | r7; // 16 or
r0 = rotr_var(r0, r6); // 17 rotr
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
r7 = rotl_imm(r7, 24u); // 19 rotl
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
r2 = r2 | r6; // 21 or
r6 = r5 * r4 + r6; // 22 mad
r1 = r1 | r0; // 23 or
r6 = r6 ^ r1; // 24 xor
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
r7 = rotr_var(r7, r0); // 26 rotr
r4 = r5 * r7 + r4; // 27 mad
r3 = r2 * r4 + r3; // 28 mad
r1 = r1 ^ ds[((rotl_imm(r4 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
r2 = r2 ^ ds[((rotl_imm(r3 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
r1 = r1 ^ ds[((rotl_imm(r5 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
r2 = mul_hi(r2, r5); // 35 mulhi
r4 = r4 ^ r2; // 36 xor
r6 = r6 * r5; // 37 mul
r7 = r7 ^ r0; // 38 xor
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
r2 = rotr_var(r2, r3); // 40 rotr
r7 = r7 - r0; // 41 sub
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 43 shfl
r0 = r0 ^ ds[((rotl_imm(r7 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
r3 = r3 ^ ds[((rotl_imm(r1 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
r6 = r6 ^ ds[((rotl_imm(r3 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
r4 = mul_hi(r4, r2); // 48 mulhi
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r0 = r0 ^ t_; } // 50 shfl
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
r5 = r5 ^ ds[((rotl_imm(r2 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
r6 = rotl_imm(r6, 12u); // 53 rotl
r3 = rotl_imm(r3, 12u); // 54 rotl
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
r1 = r1 ^ ds[((rotl_imm(r4 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
r5 = r5 ^ ds[((rotl_imm(r0 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
r3 = mul_hi(r3, r2); // 60 mulhi
r6 = r6 | r4; // 61 or
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r5 = r5 ^ t_; } // 62 shfl
r3 = r3 ^ ds[((rotl_imm(r6 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
}
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((ulong)hi << 32) | (ulong)lo;
}
#if IGNEUM_EXCHANGE != 0
// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the
// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a
// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md.
IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) {
if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); }
}
#endif
// Header-bound variant (bind.rs): the init words come from initw, not SEEDW. Same body as igneum_hash.
IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw) {
uint gid = (uint)get_global_id(0);
uint lid = (uint)get_local_id(0);
uint nonce = baseNonce + gid;
uint r0, r1, r2, r3, r4, r5, r6, r7;
uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7];
#if IGNEUM_EXCHANGE == 0
IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
uint xk = 0u;
#else
(void)lid;
#endif
{ uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; }
{ uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; }
{ uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; }
{ uint x = nonce ^ iw3; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw4; }
{ uint x = nonce ^ iw4; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw5; }
{ uint x = nonce ^ iw5; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw6; }
{ uint x = nonce ^ iw6; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw7; }
{ uint x = nonce ^ iw7; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw0; }
for (uint it = 0u; it < 8u; ++it) {
uint sel = r0;
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
r7 = r7 ^ r0; // 1 xor
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 1u); r3 = r3 ^ t_; } // 2 shfl
r4 = r4 - r1; // 3 sub
r2 = r0 * r4 + r2; // 4 mad
r4 = rotl_imm(r4, 9u); // 5 rotl
r0 = rotr_var(r0, r2); // 6 rotr
r6 = r6 ^ ds[((rotl_imm(r7 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
r1 = r1 ^ ds[((rotl_imm(r4 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
r1 = r1 ^ ds[((rotl_imm(r2 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
r7 = r7 ^ ds[((rotl_imm(r0 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
r7 = r7 ^ ds[((rotl_imm(r1 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
r5 = r5 * r4; // 12 mul
r7 = r7 ^ ds[((rotl_imm(r6 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 16u); r6 = r6 ^ t_; } // 14 shfl
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r3 = r3 ^ t_; } // 15 shfl
r2 = r2 | r7; // 16 or
r0 = rotr_var(r0, r6); // 17 rotr
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
r7 = rotl_imm(r7, 24u); // 19 rotl
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
r2 = r2 | r6; // 21 or
r6 = r5 * r4 + r6; // 22 mad
r1 = r1 | r0; // 23 or
r6 = r6 ^ r1; // 24 xor
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
r7 = rotr_var(r7, r0); // 26 rotr
r4 = r5 * r7 + r4; // 27 mad
r3 = r2 * r4 + r3; // 28 mad
r1 = r1 ^ ds[((rotl_imm(r4 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
r2 = r2 ^ ds[((rotl_imm(r3 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
r1 = r1 ^ ds[((rotl_imm(r5 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
r2 = mul_hi(r2, r5); // 35 mulhi
r4 = r4 ^ r2; // 36 xor
r6 = r6 * r5; // 37 mul
r7 = r7 ^ r0; // 38 xor
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
r2 = rotr_var(r2, r3); // 40 rotr
r7 = r7 - r0; // 41 sub
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 43 shfl
r0 = r0 ^ ds[((rotl_imm(r7 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
r3 = r3 ^ ds[((rotl_imm(r1 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
r6 = r6 ^ ds[((rotl_imm(r3 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
r4 = mul_hi(r4, r2); // 48 mulhi
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r0 = r0 ^ t_; } // 50 shfl
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
r5 = r5 ^ ds[((rotl_imm(r2 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
r6 = rotl_imm(r6, 12u); // 53 rotl
r3 = rotl_imm(r3, 12u); // 54 rotl
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
r1 = r1 ^ ds[((rotl_imm(r4 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
r5 = r5 ^ ds[((rotl_imm(r0 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
r3 = mul_hi(r3, r2); // 60 mulhi
r6 = r6 | r4; // 61 or
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r5 = r5 ^ t_; } // 62 shfl
r3 = r3 ^ ds[((rotl_imm(r6 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
}
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((ulong)hi << 32) | (ulong)lo;
}

View file

@ -0,0 +1,123 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
// Header-bound twin of igneum_hash in kernel.cu: the init words come from a kernel argument, not SEEDW.
// Host declarations (also in program_bound.h if present):
// struct IgneumInitWords { uint32_t w[8]; };
// cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
// IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps);
// cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
#include <cuda_runtime.h>
#include <cstdint>
#include "program.h"
struct IgneumInitWords { uint32_t w[8]; };
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
x ^= x >> 16; x *= 0x7feb352du;
x ^= x >> 15; x *= 0x846ca68bu;
x ^= x >> 16;
return x;
}
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw) {
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
uint32_t nonce = baseNonce + gid;
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
{ uint32_t x = nonce ^ iw.w[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw.w[1]; }
{ uint32_t x = nonce ^ iw.w[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw.w[2]; }
{ uint32_t x = nonce ^ iw.w[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw.w[3]; }
{ uint32_t x = nonce ^ iw.w[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw.w[4]; }
{ uint32_t x = nonce ^ iw.w[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw.w[5]; }
{ uint32_t x = nonce ^ iw.w[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw.w[6]; }
{ uint32_t x = nonce ^ iw.w[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw.w[7]; }
{ uint32_t x = nonce ^ iw.w[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw.w[0]; }
for (uint32_t it = 0u; it < 8u; ++it) {
uint32_t sel = r0;
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
r7 = r7 ^ r0; // 1 xor
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r6, 1); // 2 shfl
r4 = r4 - r1; // 3 sub
r2 = r0 * r4 + r2; // 4 mad
r4 = rotl_imm(r4, 9u); // 5 rotl
r0 = rotr_var(r0, r2); // 6 rotr
r6 = r6 ^ ds[((rotl_imm(r7 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
r1 = r1 ^ ds[((rotl_imm(r4 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
r1 = r1 ^ ds[((rotl_imm(r2 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
r7 = r7 ^ ds[((rotl_imm(r0 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
r7 = r7 ^ ds[((rotl_imm(r1 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
r5 = r5 * r4; // 12 mul
r7 = r7 ^ ds[((rotl_imm(r6 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r5, 16); // 14 shfl
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r5, 1); // 15 shfl
r2 = r2 | r7; // 16 or
r0 = rotr_var(r0, r6); // 17 rotr
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
r7 = rotl_imm(r7, 24u); // 19 rotl
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
r2 = r2 | r6; // 21 or
r6 = r5 * r4 + r6; // 22 mad
r1 = r1 | r0; // 23 or
r6 = r6 ^ r1; // 24 xor
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
r7 = rotr_var(r7, r0); // 26 rotr
r4 = r5 * r7 + r4; // 27 mad
r3 = r2 * r4 + r3; // 28 mad
r1 = r1 ^ ds[((rotl_imm(r4 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
r2 = r2 ^ ds[((rotl_imm(r3 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
r1 = r1 ^ ds[((rotl_imm(r5 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
r2 = __umulhi(r2, r5); // 35 mulhi
r4 = r4 ^ r2; // 36 xor
r6 = r6 * r5; // 37 mul
r7 = r7 ^ r0; // 38 xor
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
r2 = rotr_var(r2, r3); // 40 rotr
r7 = r7 - r0; // 41 sub
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 43 shfl
r0 = r0 ^ ds[((rotl_imm(r7 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
r3 = r3 ^ ds[((rotl_imm(r1 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
r6 = r6 ^ ds[((rotl_imm(r3 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
r4 = __umulhi(r4, r2); // 48 mulhi
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r7, 4); // 50 shfl
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
r5 = r5 ^ ds[((rotl_imm(r2 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
r6 = rotl_imm(r6, 12u); // 53 rotl
r3 = rotl_imm(r3, 12u); // 54 rotl
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
r1 = r1 ^ ds[((rotl_imm(r4 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
r5 = r5 ^ ds[((rotl_imm(r0 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
r3 = __umulhi(r3, r2); // 60 mulhi
r6 = r6 | r4; // 61 or
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 62 shfl
r3 = r3 ^ ds[((rotl_imm(r6 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
}
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
}
cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps) {
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
uint32_t block = 32u * blockWarps;
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
igneum_hash_bound<<<nonces / block, block>>>(ds, out, baseNonce, mask, iw);
return cudaGetLastError();
}
cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
cudaFuncAttributes attr;
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash_bound);
if (e != cudaSuccess) return e;
*numRegs = attr.numRegs;
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash_bound, (int)(32u * blockWarps), 0);
}

View file

@ -0,0 +1,112 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
// Memory-hard dataset core, the same text that the Mac's Metal kernels and CPU verifier were checked against.
// Included by kernel.cu (device), host.cu (host reference) and proto-opencl/host.c (C99 host reference).
// See proto-metal/MEMHARD.md for the construction. kernel.cl carries the same text in OpenCL C.
#pragma once
#ifdef __cplusplus
#include <cstdint>
#else
#include <stdint.h>
#endif
#if defined(__CUDACC__)
#define IGNEUM_HD __host__ __device__ __forceinline__
#elif defined(_MSC_VER) && !defined(__cplusplus)
#define IGNEUM_HD static __inline
#else
#define IGNEUM_HD static inline
#endif
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
#define MH_CACHE_LINE_MASK 0x003fffffu
#define MH_SEGMENT_LINES 64u
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
IGNEUM_HD uint32_t mh_rotl(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
// y = ChaCha12 core(x) + x
IGNEUM_HD void mh_chacha_block(const uint32_t* x, uint32_t* y) {
for (uint32_t i = 0u; i < 16u; ++i) y[i] = x[i];
for (uint32_t r = 0u; r < 6u; ++r) {
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
}
for (uint32_t i = 0u; i < 16u; ++i) y[i] += x[i];
}
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
IGNEUM_HD void mh_cache_segment(uint32_t* cache, uint32_t seg) {
uint32_t prev[16]; uint32_t x[16]; uint32_t y[16];
for (uint32_t i = 0u; i < 16u; ++i) prev[i] = 0u;
for (uint32_t j = 0u; j < MH_SEGMENT_LINES; ++j) {
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
x[4] = 0xceed56d7u ^ prev[4];
x[5] = 0x9ba270d2u ^ prev[5];
x[6] = 0x82caab2du ^ prev[6];
x[7] = 0x81ebce0eu ^ prev[7];
x[8] = 0x12b6ecf1u ^ prev[8];
x[9] = 0xd0f3fd7cu ^ prev[9];
x[10] = 0xd872eefeu ^ prev[10];
x[11] = 0xc158c7bdu ^ prev[11];
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
mh_chacha_block(x, y);
uint32_t* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
for (uint32_t i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
}
}
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
IGNEUM_HD void mh_mixer(uint32_t* s, uint32_t rk) {
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
}
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
IGNEUM_HD void mh_item(const uint32_t* cache, uint32_t t, uint32_t* s) {
s[0] = 0xceed56d7u;
s[1] = 0x9ba270d2u;
s[2] = 0x82caab2du;
s[3] = 0x81ebce0eu;
s[4] = 0x12b6ecf1u;
s[5] = 0xd0f3fd7cu;
s[6] = 0xd872eefeu;
s[7] = 0xc158c7bdu;
s[8] = t * 0xf351d601u + 0xc6892460u;
s[9] = t * 0xa3bb398fu + 0x25b7228au;
s[10] = t * 0xb5a09e35u + 0xcd515004u;
s[11] = t * 0x7509c9c1u + 0x2846527au;
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
s[14] = t * 0xded91851u + 0x82961bacu;
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
for (uint32_t r = 0u; r < 8u; ++r) {
mh_mixer(s, 0x9E3779B9u * (r + 1u));
const uint32_t* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= line[i];
}
mh_mixer(s, 0x9E3779B9u * 9u);
}
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 1 3 8 13 of w.
IGNEUM_HD uint32_t mh_j(uint32_t w) { return ((w >> 1u) & 1u) | (((w >> 3u) & 1u) << 1) | (((w >> 8u) & 1u) << 2) | (((w >> 13u) & 1u) << 3); }
IGNEUM_HD uint32_t mh_t(uint32_t w) { w = (w & 0x00001fffu) | ((w >> 14u) << 13u); w = (w & 0x000000ffu) | ((w >> 9u) << 8u); w = (w & 0x00000007u) | ((w >> 4u) << 3u); w = (w & 0x00000001u) | ((w >> 2u) << 1u); return w; }
IGNEUM_HD uint32_t mh_addr(uint32_t t, uint32_t j) { uint32_t w = t; w = ((w >> 1u) << 2u) | (w & 0x00000001u) | (((j >> 0u) & 1u) << 1u); w = ((w >> 3u) << 4u) | (w & 0x00000007u) | (((j >> 1u) & 1u) << 3u); w = ((w >> 8u) << 9u) | (w & 0x000000ffu) | (((j >> 2u) & 1u) << 8u); w = ((w >> 13u) << 14u) | (w & 0x00001fffu) | (((j >> 3u) & 1u) << 13u); return w; }
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
IGNEUM_HD uint32_t mh_word(const uint32_t* cache, uint32_t w) { uint32_t s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }

View file

@ -0,0 +1,109 @@
#include <metal_stdlib>
using namespace metal;
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
#define MH_CACHE_LINE_MASK 0x003fffffu
#define MH_SEGMENT_LINES 64u
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
// y = ChaCha12 core(x) + x
inline void mh_chacha_block(const thread uint* x, thread uint* y) {
for (uint i = 0u; i < 16u; ++i) y[i] = x[i];
for (uint r = 0u; r < 6u; ++r) {
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
}
for (uint i = 0u; i < 16u; ++i) y[i] += x[i];
}
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
inline void mh_cache_segment(device uint* cache, uint seg) {
uint prev[16]; uint x[16]; uint y[16];
for (uint i = 0u; i < 16u; ++i) prev[i] = 0u;
for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) {
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
x[4] = 0xceed56d7u ^ prev[4];
x[5] = 0x9ba270d2u ^ prev[5];
x[6] = 0x82caab2du ^ prev[6];
x[7] = 0x81ebce0eu ^ prev[7];
x[8] = 0x12b6ecf1u ^ prev[8];
x[9] = 0xd0f3fd7cu ^ prev[9];
x[10] = 0xd872eefeu ^ prev[10];
x[11] = 0xc158c7bdu ^ prev[11];
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
mh_chacha_block(x, y);
device uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
}
}
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
inline void mh_mixer(thread uint* s, uint rk) {
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
}
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
inline void mh_item(device const uint* cache, uint t, thread uint* s) {
s[0] = 0xceed56d7u;
s[1] = 0x9ba270d2u;
s[2] = 0x82caab2du;
s[3] = 0x81ebce0eu;
s[4] = 0x12b6ecf1u;
s[5] = 0xd0f3fd7cu;
s[6] = 0xd872eefeu;
s[7] = 0xc158c7bdu;
s[8] = t * 0xf351d601u + 0xc6892460u;
s[9] = t * 0xa3bb398fu + 0x25b7228au;
s[10] = t * 0xb5a09e35u + 0xcd515004u;
s[11] = t * 0x7509c9c1u + 0x2846527au;
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
s[14] = t * 0xded91851u + 0x82961bacu;
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
for (uint r = 0u; r < 8u; ++r) {
mh_mixer(s, 0x9E3779B9u * (r + 1u));
device const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i];
}
mh_mixer(s, 0x9E3779B9u * 9u);
}
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 1 3 8 13 of w.
inline uint mh_j(uint w) { return ((w >> 1u) & 1u) | (((w >> 3u) & 1u) << 1) | (((w >> 8u) & 1u) << 2) | (((w >> 13u) & 1u) << 3); }
inline uint mh_t(uint w) { w = (w & 0x00001fffu) | ((w >> 14u) << 13u); w = (w & 0x000000ffu) | ((w >> 9u) << 8u); w = (w & 0x00000007u) | ((w >> 4u) << 3u); w = (w & 0x00000001u) | ((w >> 2u) << 1u); return w; }
inline uint mh_addr(uint t, uint j) { uint w = t; w = ((w >> 1u) << 2u) | (w & 0x00000001u) | (((j >> 0u) & 1u) << 1u); w = ((w >> 3u) << 4u) | (w & 0x00000007u) | (((j >> 1u) & 1u) << 3u); w = ((w >> 8u) << 9u) | (w & 0x000000ffu) | (((j >> 2u) & 1u) << 8u); w = ((w >> 13u) << 14u) | (w & 0x00001fffu) | (((j >> 3u) & 1u) << 13u); return w; }
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
inline uint mh_word(device const uint* cache, uint w) { uint s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }
// One thread per segment (2^16 threads).
kernel void igneum_cache_fill(device uint* cache [[buffer(0)]], uint gid [[thread_position_in_grid]]) {
mh_cache_segment(cache, gid);
}
// One thread per 64-byte item (dataset words / 16 threads).
kernel void igneum_build(device const uint* cache [[buffer(0)]], device uint* dataset [[buffer(1)]],
uint gid [[thread_position_in_grid]]) {
uint s[16];
mh_item(cache, gid, s);
for (uint i = 0u; i < 16u; ++i) dataset[mh_addr(gid, i)] = s[i];
}

View file

@ -0,0 +1,76 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
// Program metadata for host.cu plus the launch wrappers defined in kernel.cu.
// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros.
#pragma once
#ifdef __cplusplus
#include <cstdint>
#else
#include <stdint.h>
#endif
#ifndef IGNEUM_NO_CUDA
#include <cuda_runtime.h>
#endif
#define IGNEUM_SEED_STRING "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000"
#define IGNEUM_SEED_BYTES_HEX "edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07"
#define IGNEUM_GENERATOR 3
#define IGNEUM_PROGRAM_ATTEMPT 0
#define IGNEUM_PROGRAM_ID 0x73bcbfe8ccf988f1ull
#define IGNEUM_DAY_STRING "bytes:69676e65756d2d6461792ffa50000000000000"
#define IGNEUM_DAY_BYTES_HEX "69676e65756d2d6461792ffa50000000000000"
#define IGNEUM_DAY0 0xceed56d7u
#define IGNEUM_DAY1 0x9ba270d2u
#define IGNEUM_DATASET_LOG2 28
#define IGNEUM_MASK 0x0fffffffu
#define IGNEUM_LANES 32
#define IGNEUM_ITERATIONS 8
#define IGNEUM_INSTR_COUNT 64
#define IGNEUM_LOADS_PER_HASH 128
#define IGNEUM_WIDE_LOADS_PER_HASH 0
#define IGNEUM_OP_MIX "load=16 add=15 shfl=6 mad=4 or=4 rotl=4 rotr=4 xor=4 mulhi=3 mul=2 sub=2"
// Program class v3 (Counter ASIC 2.0, docs/plans/counter-asic-2-rollout.md): generator version 3; a worker that
// runs another class refuses this pack, and a job line names the class it wants (class=v3 era=<hex>).
#define IGNEUM_PROGRAM_CLASS "v3"
#define IGNEUM_ERA_SEED_HEX "ff87ad96a1b53f367a95d5ca123bab64211bd9aa57fc7ad3c5e2d496c20138d4"
// Class v3 construction (Counter ASIC 2.0, 5 October 2026, docs/plans/mixer-x4.md): version 2 loads; the dataset item
// derivation applies the mixer IGNEUM_MIXER_MULT times per round (memhard.h), and the cache follows the growth rule.
#define IGNEUM_LOAD_CLASS "w4-era676a17fc"
#define IGNEUM_LOAD_SLOTS 16
#define IGNEUM_LOAD_MIX { 100, 0, 0 }
#define IGNEUM_LOAD_WIDTH_COUNTS { 16, 0, 0 } // loads of 4, 16, 64 bytes per program
#define IGNEUM_BYTES_PER_HASH 512
#define IGNEUM_FOLD_ROT 11
#define IGNEUM_FOLD_MUL 0x9e3779b1u
// Era layout (5 October 2026, docs/plans/era-layout.md): NOT the lottery hash. Every dataset load reads
// idx = ((rotl(src * STRIDE_MUL, STRIDE_ROT) & window mask) | window offset) & MASK; the window of a load site is the
// dataset, a half or a quarter of it (IGNEUM_ERA_WINDOWS: site:shrink:offset); dataset word w holds word j(w) of item
// t(w) with j's bits at the INTERLEAVE positions (memhard.h: mh_t, mh_j, mh_addr).
#define IGNEUM_ERA_LABEL "676a17fc"
#define IGNEUM_ERA_SEED_WORDS { 0x676a17fcu, 0x60bc956eu, 0x3e9865f4u, 0x68ae9e61u, 0xc0b3f442u, 0xaf406bebu, 0xb6126b9au, 0xaa30959bu }
#define IGNEUM_ERA_ALLOWED_WIDTHS { 1, 0, 0 } // words, ascending, 0 = unused; one entry pins the width
#define IGNEUM_ERA_WIDTH_WORDS 1
#define IGNEUM_ERA_STRIDE_MUL 0xb2a9d70du
#define IGNEUM_ERA_STRIDE_ROT 6
#define IGNEUM_ERA_INTERLEAVE { 1, 3, 8, 13 }
#define IGNEUM_ERA_WINDOWS "7:2:1 8:1:1 9:1:1 10:1:1 11:0:0 13:1:1 29:0:0 30:2:2 31:1:1 44:1:1 46:2:0 47:0:0 52:0:0 56:0:0 58:2:0 63:1:1"
// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h)
#define IGNEUM_DATASET_MODE 1
#define IGNEUM_SEEDW_INIT { 0x667d0fbdu, 0x7b8e5963u, 0x31c67e5eu, 0x4529ddc6u, 0xef19d6d8u, 0xaccf6211u, 0xda0aed32u, 0xabc6df31u }
#define IGNEUM_KEY_INIT { 0xceed56d7u, 0x9ba270d2u, 0x82caab2du, 0x81ebce0eu, 0x12b6ecf1u, 0xd0f3fd7cu, 0xd872eefeu, 0xc158c7bdu }
#define IGNEUM_CACHE_LOG2_WORDS 26
#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6
#define IGNEUM_CACHE_SEGMENTS 65536u
#define IGNEUM_ITEM_ROUNDS 8
#define IGNEUM_MIX_ROT_INIT { 17u, 12u, 20u, 23u, 7u, 3u, 27u, 16u }
#define IGNEUM_MIX_MUL_INIT { 0xf351d601u, 0xa3bb398fu, 0xb5a09e35u, 0x7509c9c1u, 0x6bbf31e9u, 0xfc849a79u, 0xded91851u, 0x8d9113d1u, 0x0ff15225u, 0x3a5bdd41u, 0xab533435u, 0xe1c55ad5u, 0xe6d3bd0du, 0x9d9ffbbdu, 0xbb2a3cf3u, 0x50a7c08du }
#define IGNEUM_MIX_RC_INIT { 0xc6892460u, 0x25b7228au, 0xcd515004u, 0x2846527au, 0xa6324241u, 0x36e3ec53u, 0x82961bacu, 0x0f97ba7du, 0xb6f921a9u, 0x3ada24e5u, 0xde20ab91u, 0x5378eeb2u, 0x7d161662u, 0x89353cc1u, 0xb1aa03a2u, 0x788acae6u }
#ifndef IGNEUM_NO_CUDA
// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError().
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments);
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems);
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
uint32_t nonces, uint32_t blockWarps);
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
#endif

View file

@ -0,0 +1,143 @@
{
"format": "igneum-program-pack-3",
"generator": 3,
"attempt": 0,
"program_id": "0x73bcbfe8ccf988f1",
"program_id_derivation": "FNV-1a 64 over 'igneum-program/' || generator_le32 || seed_words as little-endian bytes || attempt_le32",
"dataset_mode": "memory-hard",
"seed": "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000",
"seed_bytes": "edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07",
"seed_words": ["0x667d0fbd", "0x7b8e5963", "0x31c67e5e", "0x4529ddc6", "0xef19d6d8", "0xaccf6211", "0xda0aed32", "0xabc6df31"],
"seed_derivation": "seed_words = FNV-1a 64 over seed_bytes (attempt 0) or seed_bytes || attempt_le32 (attempt k >= 1), basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32",
"generator_rule": "version 2: exactly 16 load slots drawn first from instructions 1..63 (partial Fisher-Yates), the other 48 ops from the ten non-load weights (sum 75); a load's source is drawn from the registers other than dst written by an earlier instruction and not read by a load since; the candidate must pass the acceptance rule of spec 01 section 1.4.6 (static: no cyclically stale load source, every register has an injecting write; dynamic: 64 units on the seed-keyed closed-form dataset with no constant register bit, no lane-constant load site, under 164 saturated final values, every output bit within 136 of 1024, distinct addresses above 245760), else the next attempt of the seed is tried",
"lanes": 32,
"registers": 8,
"iterations": 8,
"instruction_count": 64,
"loads_per_hash": 128,
"program_class": "v3",
"era_seed_bytes": "ff87ad96a1b53f367a95d5ca123bab64211bd9aa57fc7ad3c5e2d496c20138d4",
"load_class": "w4-era676a17fc",
"load_slots": 16,
"load_mix_percent_4_16_64": [100, 0, 0],
"load_width_counts_4_16_64": [16, 0, 0],
"bytes_per_hash": 512,
"wide_load": "read-width experiment (5 October 2026, docs/plans/read-width.md), NOT the lottery hash: a load of W words (width field, 4 or 16) reads dataset[b .. b + W) with b = (src & mask) & ~(W - 1) and folds every word into dst: x = dst ^ w[0]; for j in 1..W: x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x; width 1 is the plain load; the width is drawn per instruction from the class mix with one extra below(100) draw after the nine of version 2, and the program id is FNV-1a 64 over 'igneum-program-rw/' || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots",
"era": {
"label": "676a17fc",
"seed_words": ["0x676a17fc", "0x60bc956e", "0x3e9865f4", "0x68ae9e61", "0xc0b3f442", "0xaf406beb", "0xb6126b9a", "0xaa30959b"],
"draw": "docs/plans/era-layout.md 1.1: SplitMix64 seeded with seed_words[0] | seed_words[1] << 32 of seed_words_from_bytes('igneum-era/' || n_le64 || E_n); width = allowed[below(|allowed|)], stride_mul = low32(next()) | 1, stride_rot = 1 + below(31), then four next() draws for a partial Fisher-Yates over positions log2(W)..15 of which 4 - log2(W) are used",
"allowed_widths": [1],
"width_words": 1,
"stride_mul": "0xb2a9d70d",
"stride_rot": 6,
"interleave": [1, 3, 8, 13],
"address": "y = rotl(src * stride_mul, stride_rot); k = min(win, D - 26); idx = ((y & (mask >> k)) | ((off & (2^k - 1)) << (D - k))) & mask; a wide load aligns idx down to W words",
"windows": "per instruction, after the width roll: win = below(3), off = low32(next()) & (2^win - 1); used on a load slot (the instruction's win and off fields)",
"dataset_word": "dataset[w] = item(t(w))[j(w)]: j(w) gathers the bits of w at the interleave positions, t(w) is w with those bits removed",
"program_id_suffix": "'era/' || allowed[3] || width_words || stride_mul_le32 || stride_rot_le32 || interleave[4]"
},
"op_mix": {"load": 16, "add": 15, "shfl": 6, "mad": 4, "or": 4, "rotl": 4, "rotr": 4, "xor": 4, "mulhi": 3, "mul": 2, "sub": 2},
"register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]",
"splitmix32": "x ^= x>>16; x *= 0x7feb352d; x ^= x>>15; x *= 0x846ca68b; x ^= x>>16",
"iteration": "sel = r0 sampled once at the top of each iteration, then all instructions in order",
"output": "lo = r0 ^ rotl(r1,7) ^ rotl(r2,14) ^ rotl(r3,21); hi = r4 ^ rotl(r5,9) ^ rotl(r6,18) ^ rotl(r7,27); out = (hi << 32) | lo",
"op_semantics": {
"add": "dst = dst + src + (bit `bit` of sel ? imm2 : imm)",
"sub": "dst = dst - src",
"mul": "dst = dst * src (low 32)",
"mulhi": "dst = high 32 bits of dst * src",
"xor": "dst = dst ^ src",
"or": "dst = dst | src",
"rotl": "dst = rotl(dst, rot), rot in 1..31",
"rotr": "dst = rotr(dst, src & 31)",
"mad": "dst = src * src2 + dst",
"shfl": "dst = dst ^ (src of lane (lane ^ mask)), mask in {1,2,4,8,16}, within the 32-lane warp",
"load": "dst = dst ^ dataset[src & dataset.mask]",
"wload": "base = (src of lane 0 & dataset.mask) & ~31; dst = dst ^ dataset[base + lane] (warp-coalesced 128-byte load, lever b, only when --wide-frac > 0)"
},
"dataset": {
"log2_words": 28,
"bytes": 1073741824,
"mask": "0x0fffffff",
"day": "bytes:69676e65756d2d6461792ffa50000000000000",
"day_bytes": "69676e65756d2d6461792ffa50000000000000",
"day_words_from": "seed_words_from_bytes(day_bytes)",
"d0": "0xceed56d7",
"d1": "0x9ba270d2",
"mode": "memory-hard",
"spec": "proto-metal/MEMHARD.md",
"key": ["0xceed56d7", "0x9ba270d2", "0x82caab2d", "0x81ebce0e", "0x12b6ecf1", "0xd0f3fd7c", "0xd872eefe", "0xc158c7bd"],
"key_derivation": "the 8 words of seed_words_from_bytes(day_bytes); d0, d1 are key[0], key[1]",
"cache": {"log2_words": 26, "bytes": 268435456, "line_words": 16, "segment_lines": 64, "segments": 65536, "block": "ChaCha12 core + feed-forward, rotations 16 12 8 7", "sigma": ["0x61707865", "0x3320646e", "0x79622d32", "0x6b206574"], "tag": ["0x49676e65", "0x756d4d48"], "chain": "in_j = prev_line ^ (sigma[0..3] || key[0..7] || seg || j || tag[0..1]); line_j = block(in_j); prev_0 = 0"},
"mixer": {"draw": "SplitMix64 seeded with key[0] | key[1] << 32: rot[0..7] = 1 + next() % 31, mul[0..15] = low32(next()) | 1, rc[0..15] = low32(next())", "rot": [17, 12, 20, 23, 7, 3, 27, 16], "mul": ["0xf351d601", "0xa3bb398f", "0xb5a09e35", "0x7509c9c1", "0x6bbf31e9", "0xfc849a79", "0xded91851", "0x8d9113d1", "0x0ff15225", "0x3a5bdd41", "0xab533435", "0xe1c55ad5", "0xe6d3bd0d", "0x9d9ffbbd", "0xbb2a3cf3", "0x50a7c08d"], "rc": ["0xc6892460", "0x25b7228a", "0xcd515004", "0x2846527a", "0xa6324241", "0x36e3ec53", "0x82961bac", "0x0f97ba7d", "0xb6f921a9", "0x3ada24e5", "0xde20ab91", "0x5378eeb2", "0x7d161662", "0x89353cc1", "0xb1aa03a2", "0x788acae6"], "round": "for i in 0..15: s[i] = (s[i] ^ (rc[i] + (r+1) * 0x9E3779B9)) * mul[i]; then quarter rounds on columns (0,4,8,12) (1,5,9,13) (2,6,10,14) (3,7,11,15) with rot[0..3] and diagonals (0,5,10,15) (1,6,11,12) (2,7,8,13) (3,4,9,14) with rot[4..7]", "quarter_round": "a += b; d ^= a; d = rotl(d, r1); c += d; b ^= c; b = rotl(b, r2); a += b; d ^= a; d = rotl(d, r3); c += d; b ^= c; b = rotl(b, r4)"},
"item": "s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..7: s = M_r(s); line = s[0] & 0x003fffff; s[i] ^= cache[line * 16 + i]; then s = M_8(s); item(t) = s",
"word": "dataset[w] = item(w >> 4)[w & 15]"
},
"instructions": [
{"i": 0, "op": "add", "dst": 4, "src": 5, "src2": 7, "imm": "0xea86e152", "imm2": "0x5810667a", "rot": 27, "bit": 13, "mask": 2, "width": 1, "win": 0, "off": 0},
{"i": 1, "op": "xor", "dst": 7, "src": 0, "src2": 6, "imm": "0xbaab6229", "imm2": "0xed861989", "rot": 26, "bit": 22, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 2, "op": "shfl", "dst": 3, "src": 6, "src2": 2, "imm": "0x5b623116", "imm2": "0xff12e5b2", "rot": 12, "bit": 24, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 3, "op": "sub", "dst": 4, "src": 1, "src2": 1, "imm": "0xe99741c7", "imm2": "0xf5fa5009", "rot": 1, "bit": 21, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 4, "op": "mad", "dst": 2, "src": 0, "src2": 4, "imm": "0x673c2157", "imm2": "0xee02465f", "rot": 20, "bit": 22, "mask": 2, "width": 1, "win": 0, "off": 0},
{"i": 5, "op": "rotl", "dst": 4, "src": 7, "src2": 7, "imm": "0x946f7818", "imm2": "0x45d3399e", "rot": 9, "bit": 2, "mask": 16, "width": 1, "win": 0, "off": 0},
{"i": 6, "op": "rotr", "dst": 0, "src": 2, "src2": 4, "imm": "0x5f6a0ed2", "imm2": "0x7043a636", "rot": 19, "bit": 31, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 7, "op": "load", "dst": 6, "src": 7, "src2": 1, "imm": "0x5c61dcf7", "imm2": "0x7466aa40", "rot": 19, "bit": 9, "mask": 2, "width": 1, "win": 2, "off": 1},
{"i": 8, "op": "load", "dst": 1, "src": 4, "src2": 5, "imm": "0x85668475", "imm2": "0xdb8cc483", "rot": 29, "bit": 7, "mask": 4, "width": 1, "win": 1, "off": 1},
{"i": 9, "op": "load", "dst": 1, "src": 2, "src2": 4, "imm": "0x4454980f", "imm2": "0xebd31581", "rot": 10, "bit": 28, "mask": 16, "width": 1, "win": 1, "off": 1},
{"i": 10, "op": "load", "dst": 7, "src": 0, "src2": 2, "imm": "0xe075297c", "imm2": "0x5779c44c", "rot": 10, "bit": 22, "mask": 2, "width": 1, "win": 1, "off": 1},
{"i": 11, "op": "load", "dst": 7, "src": 1, "src2": 6, "imm": "0x65aa4311", "imm2": "0x4fe48ea9", "rot": 15, "bit": 9, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 12, "op": "mul", "dst": 5, "src": 4, "src2": 2, "imm": "0x1383d3ad", "imm2": "0xf3094b29", "rot": 8, "bit": 9, "mask": 2, "width": 1, "win": 0, "off": 0},
{"i": 13, "op": "load", "dst": 7, "src": 6, "src2": 2, "imm": "0xed8a496f", "imm2": "0x3072c3c6", "rot": 28, "bit": 19, "mask": 8, "width": 1, "win": 1, "off": 1},
{"i": 14, "op": "shfl", "dst": 6, "src": 5, "src2": 0, "imm": "0x8b965b57", "imm2": "0xcfeca6c1", "rot": 12, "bit": 27, "mask": 16, "width": 1, "win": 0, "off": 0},
{"i": 15, "op": "shfl", "dst": 3, "src": 5, "src2": 3, "imm": "0x877c7586", "imm2": "0xa9cb2a03", "rot": 2, "bit": 29, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 16, "op": "or", "dst": 2, "src": 7, "src2": 3, "imm": "0xb740221a", "imm2": "0x89d38d6d", "rot": 6, "bit": 31, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 17, "op": "rotr", "dst": 0, "src": 6, "src2": 1, "imm": "0x26f3ad8a", "imm2": "0x27256f15", "rot": 18, "bit": 5, "mask": 2, "width": 1, "win": 0, "off": 0},
{"i": 18, "op": "add", "dst": 7, "src": 5, "src2": 7, "imm": "0xf66e7017", "imm2": "0xb9e3577e", "rot": 9, "bit": 12, "mask": 16, "width": 1, "win": 0, "off": 0},
{"i": 19, "op": "rotl", "dst": 7, "src": 0, "src2": 5, "imm": "0x849ae6ee", "imm2": "0x02b358f9", "rot": 24, "bit": 18, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 20, "op": "add", "dst": 6, "src": 7, "src2": 6, "imm": "0x52334d12", "imm2": "0x8c9f0ef8", "rot": 11, "bit": 23, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 21, "op": "or", "dst": 2, "src": 6, "src2": 5, "imm": "0xb1871e63", "imm2": "0xb2e40191", "rot": 5, "bit": 13, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 22, "op": "mad", "dst": 6, "src": 5, "src2": 4, "imm": "0x97df29e4", "imm2": "0xe60fea84", "rot": 11, "bit": 16, "mask": 2, "width": 1, "win": 0, "off": 0},
{"i": 23, "op": "or", "dst": 1, "src": 0, "src2": 3, "imm": "0x8f30d21d", "imm2": "0x2df685a0", "rot": 31, "bit": 0, "mask": 16, "width": 1, "win": 0, "off": 0},
{"i": 24, "op": "xor", "dst": 6, "src": 1, "src2": 2, "imm": "0xd2c4025f", "imm2": "0x5269eb4d", "rot": 31, "bit": 21, "mask": 16, "width": 1, "win": 0, "off": 0},
{"i": 25, "op": "add", "dst": 2, "src": 6, "src2": 5, "imm": "0x659fc3d3", "imm2": "0x9cec0e12", "rot": 6, "bit": 17, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 26, "op": "rotr", "dst": 7, "src": 0, "src2": 1, "imm": "0xd89ef484", "imm2": "0x20be3846", "rot": 12, "bit": 9, "mask": 16, "width": 1, "win": 0, "off": 0},
{"i": 27, "op": "mad", "dst": 4, "src": 5, "src2": 7, "imm": "0xb48420ae", "imm2": "0x3d1f2485", "rot": 14, "bit": 28, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 28, "op": "mad", "dst": 3, "src": 2, "src2": 4, "imm": "0xc5c46d76", "imm2": "0x700044b5", "rot": 22, "bit": 10, "mask": 2, "width": 1, "win": 0, "off": 0},
{"i": 29, "op": "load", "dst": 1, "src": 4, "src2": 3, "imm": "0x0fbaf177", "imm2": "0xfff4f2ed", "rot": 20, "bit": 31, "mask": 2, "width": 1, "win": 0, "off": 0},
{"i": 30, "op": "load", "dst": 2, "src": 3, "src2": 0, "imm": "0xe90eb2e5", "imm2": "0xb0f9eb79", "rot": 27, "bit": 14, "mask": 4, "width": 1, "win": 2, "off": 2},
{"i": 31, "op": "load", "dst": 1, "src": 5, "src2": 2, "imm": "0x97ba3fc3", "imm2": "0x7894e657", "rot": 3, "bit": 30, "mask": 16, "width": 1, "win": 1, "off": 1},
{"i": 32, "op": "add", "dst": 7, "src": 2, "src2": 7, "imm": "0x070888a8", "imm2": "0xe403240e", "rot": 2, "bit": 16, "mask": 2, "width": 1, "win": 0, "off": 0},
{"i": 33, "op": "add", "dst": 2, "src": 0, "src2": 5, "imm": "0xf2e46d55", "imm2": "0x29701828", "rot": 31, "bit": 28, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 34, "op": "add", "dst": 2, "src": 3, "src2": 0, "imm": "0x58f75b87", "imm2": "0x343b7aee", "rot": 12, "bit": 14, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 35, "op": "mulhi", "dst": 2, "src": 5, "src2": 6, "imm": "0x5892a9e6", "imm2": "0xc9824c94", "rot": 19, "bit": 26, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 36, "op": "xor", "dst": 4, "src": 2, "src2": 3, "imm": "0x7b5b5474", "imm2": "0x45cfc5dd", "rot": 17, "bit": 18, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 37, "op": "mul", "dst": 6, "src": 5, "src2": 7, "imm": "0xccf564a5", "imm2": "0x873ad101", "rot": 7, "bit": 11, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 38, "op": "xor", "dst": 7, "src": 0, "src2": 6, "imm": "0xcac8f06d", "imm2": "0x6b97c683", "rot": 18, "bit": 28, "mask": 2, "width": 1, "win": 0, "off": 0},
{"i": 39, "op": "add", "dst": 7, "src": 2, "src2": 4, "imm": "0xb8180e9d", "imm2": "0x32bbd117", "rot": 23, "bit": 19, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 40, "op": "rotr", "dst": 2, "src": 3, "src2": 0, "imm": "0x2d6070bc", "imm2": "0x68ff101e", "rot": 13, "bit": 18, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 41, "op": "sub", "dst": 7, "src": 0, "src2": 2, "imm": "0x0daf96ea", "imm2": "0x36f37be1", "rot": 5, "bit": 0, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 42, "op": "add", "dst": 4, "src": 3, "src2": 6, "imm": "0x6ced15b7", "imm2": "0x6df7aed4", "rot": 19, "bit": 4, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 43, "op": "shfl", "dst": 7, "src": 3, "src2": 5, "imm": "0x8ace05f3", "imm2": "0xd378ec12", "rot": 23, "bit": 10, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 44, "op": "load", "dst": 0, "src": 7, "src2": 4, "imm": "0xb0607786", "imm2": "0xc4acabbc", "rot": 13, "bit": 7, "mask": 16, "width": 1, "win": 1, "off": 1},
{"i": 45, "op": "add", "dst": 1, "src": 6, "src2": 5, "imm": "0xe09f54e9", "imm2": "0x83e825bf", "rot": 23, "bit": 14, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 46, "op": "load", "dst": 3, "src": 1, "src2": 0, "imm": "0x63cc1e4e", "imm2": "0xa1be8118", "rot": 12, "bit": 6, "mask": 2, "width": 1, "win": 2, "off": 0},
{"i": 47, "op": "load", "dst": 6, "src": 3, "src2": 7, "imm": "0x353f1d79", "imm2": "0x3b2e7456", "rot": 18, "bit": 18, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 48, "op": "mulhi", "dst": 4, "src": 2, "src2": 7, "imm": "0x00d8a3cd", "imm2": "0x231866d2", "rot": 21, "bit": 20, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 49, "op": "add", "dst": 5, "src": 0, "src2": 2, "imm": "0xa8bae6df", "imm2": "0xf572bdb9", "rot": 14, "bit": 7, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 50, "op": "shfl", "dst": 0, "src": 7, "src2": 7, "imm": "0x81ef22e1", "imm2": "0x74438fc5", "rot": 28, "bit": 18, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 51, "op": "add", "dst": 6, "src": 0, "src2": 6, "imm": "0x383b9260", "imm2": "0x11e17c61", "rot": 12, "bit": 19, "mask": 16, "width": 1, "win": 0, "off": 0},
{"i": 52, "op": "load", "dst": 5, "src": 2, "src2": 2, "imm": "0xfb84f451", "imm2": "0x11cd863e", "rot": 21, "bit": 20, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 53, "op": "rotl", "dst": 6, "src": 5, "src2": 4, "imm": "0xb1a7db6b", "imm2": "0x76686b9b", "rot": 12, "bit": 4, "mask": 16, "width": 1, "win": 0, "off": 0},
{"i": 54, "op": "rotl", "dst": 3, "src": 6, "src2": 3, "imm": "0x6f981f52", "imm2": "0xd99aeba2", "rot": 12, "bit": 27, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 55, "op": "add", "dst": 2, "src": 1, "src2": 2, "imm": "0xac6be8e3", "imm2": "0x18d67dbb", "rot": 26, "bit": 10, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 56, "op": "load", "dst": 1, "src": 4, "src2": 0, "imm": "0x7e7f6a00", "imm2": "0x6f0747da", "rot": 25, "bit": 28, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 57, "op": "add", "dst": 5, "src": 0, "src2": 4, "imm": "0xf03673fe", "imm2": "0xa75cd60d", "rot": 16, "bit": 12, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 58, "op": "load", "dst": 5, "src": 0, "src2": 2, "imm": "0x227f94a6", "imm2": "0x0e8344f9", "rot": 20, "bit": 10, "mask": 2, "width": 1, "win": 2, "off": 0},
{"i": 59, "op": "add", "dst": 1, "src": 4, "src2": 3, "imm": "0xdecd4794", "imm2": "0x8dfb96bb", "rot": 21, "bit": 7, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 60, "op": "mulhi", "dst": 3, "src": 2, "src2": 2, "imm": "0x0dd268e0", "imm2": "0x53034ca9", "rot": 1, "bit": 8, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 61, "op": "or", "dst": 6, "src": 4, "src2": 7, "imm": "0x3a45a321", "imm2": "0x9bc59a5f", "rot": 25, "bit": 11, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 62, "op": "shfl", "dst": 5, "src": 4, "src2": 2, "imm": "0x8f229cc1", "imm2": "0xcaac64a2", "rot": 17, "bit": 13, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 63, "op": "load", "dst": 3, "src": 6, "src2": 6, "imm": "0xb2574178", "imm2": "0xcbcc798d", "rot": 28, "bit": 0, "mask": 16, "width": 1, "win": 1, "off": 1}
]
}

View file

@ -0,0 +1,109 @@
#include <metal_stdlib>
using namespace metal;
#define MASK 0x0fffffffu
constant uint SEEDW[8] = { 0x667d0fbdu, 0x7b8e5963u, 0x31c67e5eu, 0x4529ddc6u, 0xef19d6d8u, 0xaccf6211u, 0xda0aed32u, 0xabc6df31u };
inline uint splitmix32(uint x) {
x ^= x >> 16; x *= 0x7feb352du;
x ^= x >> 15; x *= 0x846ca68bu;
x ^= x >> 16;
return x;
}
inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31
inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
inline uint ds_elem(uint i, uint d0, uint d1) {
uint x = i ^ d0;
x *= 0x9E3779B1u; x ^= x >> 15;
x += d1;
x *= 0x85EBCA77u; x ^= x >> 13;
x *= 0xC2B2AE3Du; x ^= x >> 16;
return x;
}
kernel void igneum_hash(device const uint* dataset [[buffer(0)]],
device ulong* out [[buffer(1)]],
constant uint& baseNonce [[buffer(2)]],
uint gid [[thread_position_in_grid]]) {
uint nonce = baseNonce + gid;
uint r0, r1, r2, r3, r4, r5, r6, r7;
{ uint x = nonce ^ SEEDW[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ SEEDW[1]; }
{ uint x = nonce ^ SEEDW[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ SEEDW[2]; }
{ uint x = nonce ^ SEEDW[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ SEEDW[3]; }
{ uint x = nonce ^ SEEDW[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ SEEDW[4]; }
{ uint x = nonce ^ SEEDW[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ SEEDW[5]; }
{ uint x = nonce ^ SEEDW[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ SEEDW[6]; }
{ uint x = nonce ^ SEEDW[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ SEEDW[7]; }
{ uint x = nonce ^ SEEDW[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ SEEDW[0]; }
for (uint it = 0u; it < 8u; ++it) {
uint sel = r0;
r4 = r4 + r5 + select(0xea86e152u, 0x5810667au, ((sel >> 13u) & 1u) != 0u); // 0
r7 = r7 ^ r0; // 1
r3 = r3 ^ simd_shuffle_xor(r6, (ushort)1); // 2
r4 = r4 - r1; // 3
r2 = r0 * r4 + r2; // 4
r4 = rotl_imm(r4, 9u); // 5
r0 = rotr_var(r0, r2); // 6
r6 = r6 ^ dataset[((rotl_imm(r7 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x04000000u) & MASK]; // 7
r1 = r1 ^ dataset[((rotl_imm(r4 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 8
r1 = r1 ^ dataset[((rotl_imm(r2 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 9
r7 = r7 ^ dataset[((rotl_imm(r0 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 10
r7 = r7 ^ dataset[((rotl_imm(r1 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 11
r5 = r5 * r4; // 12
r7 = r7 ^ dataset[((rotl_imm(r6 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 13
r6 = r6 ^ simd_shuffle_xor(r5, (ushort)16); // 14
r3 = r3 ^ simd_shuffle_xor(r5, (ushort)1); // 15
r2 = r2 | r7; // 16
r0 = rotr_var(r0, r6); // 17
r7 = r7 + r5 + select(0xf66e7017u, 0xb9e3577eu, ((sel >> 12u) & 1u) != 0u); // 18
r7 = rotl_imm(r7, 24u); // 19
r6 = r6 + r7 + select(0x52334d12u, 0x8c9f0ef8u, ((sel >> 23u) & 1u) != 0u); // 20
r2 = r2 | r6; // 21
r6 = r5 * r4 + r6; // 22
r1 = r1 | r0; // 23
r6 = r6 ^ r1; // 24
r2 = r2 + r6 + select(0x659fc3d3u, 0x9cec0e12u, ((sel >> 17u) & 1u) != 0u); // 25
r7 = rotr_var(r7, r0); // 26
r4 = r5 * r7 + r4; // 27
r3 = r2 * r4 + r3; // 28
r1 = r1 ^ dataset[((rotl_imm(r4 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 29
r2 = r2 ^ dataset[((rotl_imm(r3 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x08000000u) & MASK]; // 30
r1 = r1 ^ dataset[((rotl_imm(r5 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 31
r7 = r7 + r2 + select(0x070888a8u, 0xe403240eu, ((sel >> 16u) & 1u) != 0u); // 32
r2 = r2 + r0 + select(0xf2e46d55u, 0x29701828u, ((sel >> 28u) & 1u) != 0u); // 33
r2 = r2 + r3 + select(0x58f75b87u, 0x343b7aeeu, ((sel >> 14u) & 1u) != 0u); // 34
r2 = mulhi(r2, r5); // 35
r4 = r4 ^ r2; // 36
r6 = r6 * r5; // 37
r7 = r7 ^ r0; // 38
r7 = r7 + r2 + select(0xb8180e9du, 0x32bbd117u, ((sel >> 19u) & 1u) != 0u); // 39
r2 = rotr_var(r2, r3); // 40
r7 = r7 - r0; // 41
r4 = r4 + r3 + select(0x6ced15b7u, 0x6df7aed4u, ((sel >> 4u) & 1u) != 0u); // 42
r7 = r7 ^ simd_shuffle_xor(r3, (ushort)4); // 43
r0 = r0 ^ dataset[((rotl_imm(r7 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 44
r1 = r1 + r6 + select(0xe09f54e9u, 0x83e825bfu, ((sel >> 14u) & 1u) != 0u); // 45
r3 = r3 ^ dataset[((rotl_imm(r1 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 46
r6 = r6 ^ dataset[((rotl_imm(r3 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 47
r4 = mulhi(r4, r2); // 48
r5 = r5 + r0 + select(0xa8bae6dfu, 0xf572bdb9u, ((sel >> 7u) & 1u) != 0u); // 49
r0 = r0 ^ simd_shuffle_xor(r7, (ushort)4); // 50
r6 = r6 + r0 + select(0x383b9260u, 0x11e17c61u, ((sel >> 19u) & 1u) != 0u); // 51
r5 = r5 ^ dataset[((rotl_imm(r2 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 52
r6 = rotl_imm(r6, 12u); // 53
r3 = rotl_imm(r3, 12u); // 54
r2 = r2 + r1 + select(0xac6be8e3u, 0x18d67dbbu, ((sel >> 10u) & 1u) != 0u); // 55
r1 = r1 ^ dataset[((rotl_imm(r4 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 56
r5 = r5 + r0 + select(0xf03673feu, 0xa75cd60du, ((sel >> 12u) & 1u) != 0u); // 57
r5 = r5 ^ dataset[((rotl_imm(r0 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 58
r1 = r1 + r4 + select(0xdecd4794u, 0x8dfb96bbu, ((sel >> 7u) & 1u) != 0u); // 59
r3 = mulhi(r3, r2); // 60
r6 = r6 | r4; // 61
r5 = r5 ^ simd_shuffle_xor(r4, (ushort)8); // 62
r3 = r3 ^ dataset[((rotl_imm(r6 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 63
}
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((ulong)hi << 32) | (ulong)lo;
}

View file

@ -0,0 +1,111 @@
#include <metal_stdlib>
using namespace metal;
#define MASK 0x0fffffffu
constant uint SEEDW[8] = { 0x667d0fbdu, 0x7b8e5963u, 0x31c67e5eu, 0x4529ddc6u, 0xef19d6d8u, 0xaccf6211u, 0xda0aed32u, 0xabc6df31u };
inline uint splitmix32(uint x) {
x ^= x >> 16; x *= 0x7feb352du;
x ^= x >> 15; x *= 0x846ca68bu;
x ^= x >> 16;
return x;
}
inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31
inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
inline uint ds_elem(uint i, uint d0, uint d1) {
uint x = i ^ d0;
x *= 0x9E3779B1u; x ^= x >> 15;
x += d1;
x *= 0x85EBCA77u; x ^= x >> 13;
x *= 0xC2B2AE3Du; x ^= x >> 16;
return x;
}
// Header-bound variant: the init words come from buffer 3 (bind.rs), not from SEEDW.
kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]],
device ulong* out [[buffer(1)]],
constant uint& baseNonce [[buffer(2)]],
constant uint* initw [[buffer(3)]],
uint gid [[thread_position_in_grid]]) {
uint nonce = baseNonce + gid;
uint r0, r1, r2, r3, r4, r5, r6, r7;
{ uint x = nonce ^ initw[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ initw[1]; }
{ uint x = nonce ^ initw[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ initw[2]; }
{ uint x = nonce ^ initw[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ initw[3]; }
{ uint x = nonce ^ initw[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ initw[4]; }
{ uint x = nonce ^ initw[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ initw[5]; }
{ uint x = nonce ^ initw[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ initw[6]; }
{ uint x = nonce ^ initw[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ initw[7]; }
{ uint x = nonce ^ initw[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ initw[0]; }
for (uint it = 0u; it < 8u; ++it) {
uint sel = r0;
r4 = r4 + r5 + select(0xea86e152u, 0x5810667au, ((sel >> 13u) & 1u) != 0u); // 0
r7 = r7 ^ r0; // 1
r3 = r3 ^ simd_shuffle_xor(r6, (ushort)1); // 2
r4 = r4 - r1; // 3
r2 = r0 * r4 + r2; // 4
r4 = rotl_imm(r4, 9u); // 5
r0 = rotr_var(r0, r2); // 6
r6 = r6 ^ dataset[((rotl_imm(r7 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x04000000u) & MASK]; // 7
r1 = r1 ^ dataset[((rotl_imm(r4 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 8
r1 = r1 ^ dataset[((rotl_imm(r2 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 9
r7 = r7 ^ dataset[((rotl_imm(r0 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 10
r7 = r7 ^ dataset[((rotl_imm(r1 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 11
r5 = r5 * r4; // 12
r7 = r7 ^ dataset[((rotl_imm(r6 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 13
r6 = r6 ^ simd_shuffle_xor(r5, (ushort)16); // 14
r3 = r3 ^ simd_shuffle_xor(r5, (ushort)1); // 15
r2 = r2 | r7; // 16
r0 = rotr_var(r0, r6); // 17
r7 = r7 + r5 + select(0xf66e7017u, 0xb9e3577eu, ((sel >> 12u) & 1u) != 0u); // 18
r7 = rotl_imm(r7, 24u); // 19
r6 = r6 + r7 + select(0x52334d12u, 0x8c9f0ef8u, ((sel >> 23u) & 1u) != 0u); // 20
r2 = r2 | r6; // 21
r6 = r5 * r4 + r6; // 22
r1 = r1 | r0; // 23
r6 = r6 ^ r1; // 24
r2 = r2 + r6 + select(0x659fc3d3u, 0x9cec0e12u, ((sel >> 17u) & 1u) != 0u); // 25
r7 = rotr_var(r7, r0); // 26
r4 = r5 * r7 + r4; // 27
r3 = r2 * r4 + r3; // 28
r1 = r1 ^ dataset[((rotl_imm(r4 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 29
r2 = r2 ^ dataset[((rotl_imm(r3 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x08000000u) & MASK]; // 30
r1 = r1 ^ dataset[((rotl_imm(r5 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 31
r7 = r7 + r2 + select(0x070888a8u, 0xe403240eu, ((sel >> 16u) & 1u) != 0u); // 32
r2 = r2 + r0 + select(0xf2e46d55u, 0x29701828u, ((sel >> 28u) & 1u) != 0u); // 33
r2 = r2 + r3 + select(0x58f75b87u, 0x343b7aeeu, ((sel >> 14u) & 1u) != 0u); // 34
r2 = mulhi(r2, r5); // 35
r4 = r4 ^ r2; // 36
r6 = r6 * r5; // 37
r7 = r7 ^ r0; // 38
r7 = r7 + r2 + select(0xb8180e9du, 0x32bbd117u, ((sel >> 19u) & 1u) != 0u); // 39
r2 = rotr_var(r2, r3); // 40
r7 = r7 - r0; // 41
r4 = r4 + r3 + select(0x6ced15b7u, 0x6df7aed4u, ((sel >> 4u) & 1u) != 0u); // 42
r7 = r7 ^ simd_shuffle_xor(r3, (ushort)4); // 43
r0 = r0 ^ dataset[((rotl_imm(r7 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 44
r1 = r1 + r6 + select(0xe09f54e9u, 0x83e825bfu, ((sel >> 14u) & 1u) != 0u); // 45
r3 = r3 ^ dataset[((rotl_imm(r1 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 46
r6 = r6 ^ dataset[((rotl_imm(r3 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 47
r4 = mulhi(r4, r2); // 48
r5 = r5 + r0 + select(0xa8bae6dfu, 0xf572bdb9u, ((sel >> 7u) & 1u) != 0u); // 49
r0 = r0 ^ simd_shuffle_xor(r7, (ushort)4); // 50
r6 = r6 + r0 + select(0x383b9260u, 0x11e17c61u, ((sel >> 19u) & 1u) != 0u); // 51
r5 = r5 ^ dataset[((rotl_imm(r2 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 52
r6 = rotl_imm(r6, 12u); // 53
r3 = rotl_imm(r3, 12u); // 54
r2 = r2 + r1 + select(0xac6be8e3u, 0x18d67dbbu, ((sel >> 10u) & 1u) != 0u); // 55
r1 = r1 ^ dataset[((rotl_imm(r4 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 56
r5 = r5 + r0 + select(0xf03673feu, 0xa75cd60du, ((sel >> 12u) & 1u) != 0u); // 57
r5 = r5 ^ dataset[((rotl_imm(r0 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 58
r1 = r1 + r4 + select(0xdecd4794u, 0x8dfb96bbu, ((sel >> 7u) & 1u) != 0u); // 59
r3 = mulhi(r3, r2); // 60
r6 = r6 | r4; // 61
r5 = r5 ^ simd_shuffle_xor(r4, (ushort)8); // 62
r3 = r3 ^ dataset[((rotl_imm(r6 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 63
}
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((ulong)hi << 32) | (ulong)lo;
}

View file

@ -0,0 +1,5 @@
epoch_seed_hex edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07
day_seed_hex 69676e65756d2d6461792ffa50000000000000
epoch_index 0
day_index 20730
daa_score 0

View file

@ -0,0 +1,57 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v3, memory-hard dataset
#pragma once
#ifdef __cplusplus
#include <cstdint>
#else
#include <stdint.h>
#endif
#define IGNEUM_VEC_WARPS 3
static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u };
static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = {
{ // base nonce 0
0xdabe038c15c539f6ull, 0xda2a12ac4857cae1ull, 0x2768199f161cd4efull, 0x28e44246d90e9cffull, 0xa6e49bc2e7b5da30ull, 0x9797d70a66935020ull, 0x4820804005effffeull, 0x88662fd15b9fe6e0ull,
0xd7ffe7abc6bed954ull, 0x29baff8fcb44fa53ull, 0x5a090c72603a9301ull, 0x3c81aea5dac7be1eull, 0xfb1a9cc02b40b817ull, 0x3d0e8e32741a6091ull, 0xf7fd20269e5402beull, 0x89b31407de8924b6ull,
0xaae6fe83cc322e25ull, 0x076d6b1d7de1f125ull, 0xb4aa525302008bdfull, 0x860c033cd41fcd1dull, 0x20d77bcc09e08aa7ull, 0xe75793c696a71f74ull, 0xca05dc63ec4f5b6eull, 0x0c07b473ff504716ull,
0x8318ca28bfe8d189ull, 0xae11050d0adc5a87ull, 0x3b899c2da4ce3239ull, 0x815c563707ff1e0cull, 0x783799ce7b282d84ull, 0xc96488bfb2156fc9ull, 0x65d30aace55b4d9full, 0x0051015324bc1753ull
},
{ // base nonce 4096
0x09635fc8b43e490aull, 0x2cde2ee9fb2a0428ull, 0x7efadc8593c3b2ccull, 0x902068a0e0355bc3ull, 0xc3af55591a6b9a4cull, 0xb36f5ef5bef69733ull, 0x32d797b26f3ff603ull, 0x80e0374e701bb677ull,
0x08db70d60ac78f51ull, 0x23d6ee5b06233744ull, 0x2a4de99fd6f150e5ull, 0x5727b4c1f0db5162ull, 0x191bafa761798bacull, 0x94e47312bbf20681ull, 0xee9da67774b46e6aull, 0x1465ae37760b4e08ull,
0x2d96a0f805f370dcull, 0xc1a761ea03e0d500ull, 0x06317cfdbfcdc10eull, 0xef71c2112b4897dfull, 0x6e364b602701201dull, 0xdd0cfb3a907838ceull, 0x569e82cb5acd3801ull, 0x56723edf79db59fbull,
0x5826c8add612ef45ull, 0xc3fdce113372a83dull, 0x632b40402f52ff52ull, 0x443994dfc952a6abull, 0x5ffda84776dcf8eeull, 0x79fe9a93648e18dfull, 0x266faea56ca3ba5dull, 0x0277609022e6adc0ull
},
{ // base nonce 1000000
0x689af011faee80c1ull, 0x4e62e84a6a805665ull, 0x8828e37e049d9739ull, 0xaf733761893000dbull, 0xdf9aa4524c6c9075ull, 0x34af22f30b9827c9ull, 0xe7ae4db768b3bef8ull, 0x83bb23a187bf66c9ull,
0xad82b133ca30a3e5ull, 0x5a4c4ae80d988fd2ull, 0xbfc7229dca4c05bdull, 0xce783e7c7636ca10ull, 0xb18c78869943c79cull, 0x1023948ca66770cbull, 0x84dcd256dbf5e9dcull, 0x2e9ad8e9c0bcd6e4ull,
0xf2d3c29d300046bdull, 0x10713b9bd25e3583ull, 0x9a1fa2385a63b827ull, 0xb31ae3aeeaa46884ull, 0xb16e8a6c82c7bd1aull, 0xfed6df91f0f3cfb4ull, 0x864d1f1b8727aed8ull, 0x89e9ac61e6362617ull,
0x70a1df41038533bbull, 0x3d10c42046a6413bull, 0x31023c3566190ecfull, 0x2294dc505d480256ull, 0x24ff4fb403755cc7ull, 0x50ed84b9b4720072ull, 0x783690abb3addc85ull, 0x12eabf30075edb7bull
}
};
// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455).
static const uint32_t IGNEUM_DS_HEAD[16] = {
0x3dd50b1fu, 0x93123701u, 0x48edec90u, 0x3a7d2407u, 0xcdd88b45u, 0x2c37c6b0u, 0xb03a0ac0u, 0x10e2c6a5u,
0x5119540eu, 0x3358ef2du, 0x657ac748u, 0x09e2105cu, 0x25e2929cu, 0x5fa0cdccu, 0x4317cabdu, 0x6178ee52u
};
static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u;
static const uint32_t IGNEUM_DS_LAST = 0xf7b7180eu;
// 64 sampled dataset words (index, value) computed on the Mac.
#define IGNEUM_DS_SAMPLES 64
static const uint32_t IGNEUM_DS_SAMPLE_INDEX[IGNEUM_DS_SAMPLES] = {
59471966u, 217795994u, 208353206u, 42483309u, 172547758u, 148076330u, 183853158u, 214389424u, 267488061u, 169781097u, 184093494u, 153880993u, 84977930u, 46426879u, 3093825u, 225364072u, 44593546u, 260713159u, 168250303u, 52384140u, 223401610u, 45554030u, 95410555u, 175039924u, 79171087u, 267580473u, 24168642u, 37981670u, 171551130u, 195559979u, 204611762u, 140997658u, 138925853u, 86637313u, 20736778u, 219665210u, 160430336u, 264654675u, 8013395u, 228945585u, 213884386u, 104419827u, 44185464u, 142737231u, 99284897u, 132475900u, 61861762u, 132056166u, 262388043u, 91878046u, 117353561u, 124768597u, 71352993u, 190698941u, 46055428u, 55281366u, 165145231u, 106810753u, 171985651u, 232085256u, 159510492u, 40072060u, 209107596u, 39023794u
};
static const uint32_t IGNEUM_DS_SAMPLE_VALUE[IGNEUM_DS_SAMPLES] = {
0xf6e00c89u, 0x370d4c24u, 0x09260b8cu, 0x16208e9du, 0x4028c489u, 0x1386ca6bu, 0x88bf905cu, 0x88f76874u, 0xac034e2du, 0x88e89b31u, 0x22deecc6u, 0xd6022e44u, 0x4b18c60eu, 0x8c761ec7u, 0x7079d93du, 0x79341b4cu, 0x5b93bcc4u, 0xc659d408u, 0xd5883003u, 0xe6b03742u, 0x016da861u, 0xbbec95bfu, 0xc1647d24u, 0x88b20f49u, 0xcbcb1898u, 0x05d9a8aeu, 0x2bfaeba0u, 0x12094029u, 0xd2a03583u, 0xf8bc77efu, 0xeb7dd210u, 0x251ecc1du, 0xa8d8a48bu, 0x26c9e599u, 0x12927937u, 0x536d595eu, 0x1f437e50u, 0x9b52077au, 0x876dce0cu, 0x65092ed9u, 0xf89bbfddu, 0xf20ee5b8u, 0xb4205b97u, 0xd79db283u, 0x65869dcfu, 0xd6c0f997u, 0xa9effec4u, 0xd192a75au, 0x73d9e395u, 0x2f043a85u, 0xdc0c3fddu, 0x60b1d49au, 0xf5da472eu, 0x83f8e3d3u, 0xdb881ba6u, 0x2f00a5dcu, 0xbe7f358eu, 0x68a25926u, 0xf5852ad5u, 0xa731ea98u, 0xbb567051u, 0xa1b64724u, 0x1d8896b6u, 0x81ef09f4u
};
// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^26 words.
static const uint32_t IGNEUM_CACHE_HEAD[16] = {
0xebd9055cu, 0x7eaf6b21u, 0x9b610855u, 0x134a8562u, 0xb10059bau, 0x7f459e58u, 0x8f3c7873u, 0x3523e609u,
0x75b8f4cau, 0x750d31b4u, 0x0dae0782u, 0x26f10015u, 0x7ba87a41u, 0x0efd8543u, 0x691d6368u, 0x8929d967u
};
static const uint32_t IGNEUM_CACHE_LAST[16] = {
0x51474edfu, 0xc3cfba93u, 0xf21454cfu, 0x9b79baacu, 0x4d4fce23u, 0xd701dfd5u, 0x37357ab3u, 0x1be693fau,
0xa701cb3bu, 0x7467c620u, 0x428184e8u, 0xf4010df0u, 0xe33aa88fu, 0xc78d6d62u, 0xae9ba9c0u, 0xf4bba97eu
};
static const uint64_t IGNEUM_CACHE_FNV64 = 0x448274a57f508cbcull;

View file

@ -0,0 +1,36 @@
{
"seed": "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000",
"day": "bytes:69676e65756d2d6461792ffa50000000000000",
"dataset_mode": "memory-hard",
"dataset_log2_words": 28,
"mask": "0x0fffffff",
"lanes": 32,
"source": "igneum-pow (Rust) CPU interpreter, generator v3, memory-hard dataset",
"warps": [
{"base_nonce": 0, "expected": [
"0xdabe038c15c539f6", "0xda2a12ac4857cae1", "0x2768199f161cd4ef", "0x28e44246d90e9cff", "0xa6e49bc2e7b5da30", "0x9797d70a66935020", "0x4820804005effffe", "0x88662fd15b9fe6e0",
"0xd7ffe7abc6bed954", "0x29baff8fcb44fa53", "0x5a090c72603a9301", "0x3c81aea5dac7be1e", "0xfb1a9cc02b40b817", "0x3d0e8e32741a6091", "0xf7fd20269e5402be", "0x89b31407de8924b6",
"0xaae6fe83cc322e25", "0x076d6b1d7de1f125", "0xb4aa525302008bdf", "0x860c033cd41fcd1d", "0x20d77bcc09e08aa7", "0xe75793c696a71f74", "0xca05dc63ec4f5b6e", "0x0c07b473ff504716",
"0x8318ca28bfe8d189", "0xae11050d0adc5a87", "0x3b899c2da4ce3239", "0x815c563707ff1e0c", "0x783799ce7b282d84", "0xc96488bfb2156fc9", "0x65d30aace55b4d9f", "0x0051015324bc1753"
]},
{"base_nonce": 4096, "expected": [
"0x09635fc8b43e490a", "0x2cde2ee9fb2a0428", "0x7efadc8593c3b2cc", "0x902068a0e0355bc3", "0xc3af55591a6b9a4c", "0xb36f5ef5bef69733", "0x32d797b26f3ff603", "0x80e0374e701bb677",
"0x08db70d60ac78f51", "0x23d6ee5b06233744", "0x2a4de99fd6f150e5", "0x5727b4c1f0db5162", "0x191bafa761798bac", "0x94e47312bbf20681", "0xee9da67774b46e6a", "0x1465ae37760b4e08",
"0x2d96a0f805f370dc", "0xc1a761ea03e0d500", "0x06317cfdbfcdc10e", "0xef71c2112b4897df", "0x6e364b602701201d", "0xdd0cfb3a907838ce", "0x569e82cb5acd3801", "0x56723edf79db59fb",
"0x5826c8add612ef45", "0xc3fdce113372a83d", "0x632b40402f52ff52", "0x443994dfc952a6ab", "0x5ffda84776dcf8ee", "0x79fe9a93648e18df", "0x266faea56ca3ba5d", "0x0277609022e6adc0"
]},
{"base_nonce": 1000000, "expected": [
"0x689af011faee80c1", "0x4e62e84a6a805665", "0x8828e37e049d9739", "0xaf733761893000db", "0xdf9aa4524c6c9075", "0x34af22f30b9827c9", "0xe7ae4db768b3bef8", "0x83bb23a187bf66c9",
"0xad82b133ca30a3e5", "0x5a4c4ae80d988fd2", "0xbfc7229dca4c05bd", "0xce783e7c7636ca10", "0xb18c78869943c79c", "0x1023948ca66770cb", "0x84dcd256dbf5e9dc", "0x2e9ad8e9c0bcd6e4",
"0xf2d3c29d300046bd", "0x10713b9bd25e3583", "0x9a1fa2385a63b827", "0xb31ae3aeeaa46884", "0xb16e8a6c82c7bd1a", "0xfed6df91f0f3cfb4", "0x864d1f1b8727aed8", "0x89e9ac61e6362617",
"0x70a1df41038533bb", "0x3d10c42046a6413b", "0x31023c3566190ecf", "0x2294dc505d480256", "0x24ff4fb403755cc7", "0x50ed84b9b4720072", "0x783690abb3addc85", "0x12eabf30075edb7b"
]}
],
"dataset_head": ["0x3dd50b1f", "0x93123701", "0x48edec90", "0x3a7d2407", "0xcdd88b45", "0x2c37c6b0", "0xb03a0ac0", "0x10e2c6a5", "0x5119540e", "0x3358ef2d", "0x657ac748", "0x09e2105c", "0x25e2929c", "0x5fa0cdcc", "0x4317cabd", "0x6178ee52"],
"dataset_last_index": 268435455,
"dataset_last": "0xf7b7180e",
"dataset_samples": [{"index": 59471966, "value": "0xf6e00c89"}, {"index": 217795994, "value": "0x370d4c24"}, {"index": 208353206, "value": "0x09260b8c"}, {"index": 42483309, "value": "0x16208e9d"}, {"index": 172547758, "value": "0x4028c489"}, {"index": 148076330, "value": "0x1386ca6b"}, {"index": 183853158, "value": "0x88bf905c"}, {"index": 214389424, "value": "0x88f76874"}, {"index": 267488061, "value": "0xac034e2d"}, {"index": 169781097, "value": "0x88e89b31"}, {"index": 184093494, "value": "0x22deecc6"}, {"index": 153880993, "value": "0xd6022e44"}, {"index": 84977930, "value": "0x4b18c60e"}, {"index": 46426879, "value": "0x8c761ec7"}, {"index": 3093825, "value": "0x7079d93d"}, {"index": 225364072, "value": "0x79341b4c"}, {"index": 44593546, "value": "0x5b93bcc4"}, {"index": 260713159, "value": "0xc659d408"}, {"index": 168250303, "value": "0xd5883003"}, {"index": 52384140, "value": "0xe6b03742"}, {"index": 223401610, "value": "0x016da861"}, {"index": 45554030, "value": "0xbbec95bf"}, {"index": 95410555, "value": "0xc1647d24"}, {"index": 175039924, "value": "0x88b20f49"}, {"index": 79171087, "value": "0xcbcb1898"}, {"index": 267580473, "value": "0x05d9a8ae"}, {"index": 24168642, "value": "0x2bfaeba0"}, {"index": 37981670, "value": "0x12094029"}, {"index": 171551130, "value": "0xd2a03583"}, {"index": 195559979, "value": "0xf8bc77ef"}, {"index": 204611762, "value": "0xeb7dd210"}, {"index": 140997658, "value": "0x251ecc1d"}, {"index": 138925853, "value": "0xa8d8a48b"}, {"index": 86637313, "value": "0x26c9e599"}, {"index": 20736778, "value": "0x12927937"}, {"index": 219665210, "value": "0x536d595e"}, {"index": 160430336, "value": "0x1f437e50"}, {"index": 264654675, "value": "0x9b52077a"}, {"index": 8013395, "value": "0x876dce0c"}, {"index": 228945585, "value": "0x65092ed9"}, {"index": 213884386, "value": "0xf89bbfdd"}, {"index": 104419827, "value": "0xf20ee5b8"}, {"index": 44185464, "value": "0xb4205b97"}, {"index": 142737231, "value": "0xd79db283"}, {"index": 99284897, "value": "0x65869dcf"}, {"index": 132475900, "value": "0xd6c0f997"}, {"index": 61861762, "value": "0xa9effec4"}, {"index": 132056166, "value": "0xd192a75a"}, {"index": 262388043, "value": "0x73d9e395"}, {"index": 91878046, "value": "0x2f043a85"}, {"index": 117353561, "value": "0xdc0c3fdd"}, {"index": 124768597, "value": "0x60b1d49a"}, {"index": 71352993, "value": "0xf5da472e"}, {"index": 190698941, "value": "0x83f8e3d3"}, {"index": 46055428, "value": "0xdb881ba6"}, {"index": 55281366, "value": "0x2f00a5dc"}, {"index": 165145231, "value": "0xbe7f358e"}, {"index": 106810753, "value": "0x68a25926"}, {"index": 171985651, "value": "0xf5852ad5"}, {"index": 232085256, "value": "0xa731ea98"}, {"index": 159510492, "value": "0xbb567051"}, {"index": 40072060, "value": "0xa1b64724"}, {"index": 209107596, "value": "0x1d8896b6"}, {"index": 39023794, "value": "0x81ef09f4"}],
"cache_head": ["0xebd9055c", "0x7eaf6b21", "0x9b610855", "0x134a8562", "0xb10059ba", "0x7f459e58", "0x8f3c7873", "0x3523e609", "0x75b8f4ca", "0x750d31b4", "0x0dae0782", "0x26f10015", "0x7ba87a41", "0x0efd8543", "0x691d6368", "0x8929d967"],
"cache_last_line": ["0x51474edf", "0xc3cfba93", "0xf21454cf", "0x9b79baac", "0x4d4fce23", "0xd701dfd5", "0x37357ab3", "0x1be693fa", "0xa701cb3b", "0x7467c620", "0x428184e8", "0xf4010df0", "0xe33aa88f", "0xc78d6d62", "0xae9ba9c0", "0xf4bba97e"],
"cache_fnv1a64": "0x448274a57f508cbc"
}

View file

@ -0,0 +1,281 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal).
// Built from source at runtime by proto-opencl/host.c, which passes these defines:
// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit)
// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default)
// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32
// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition
// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units;
// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32.
#ifndef IGNEUM_GROUP
#define IGNEUM_GROUP 32
#endif
#ifndef IGNEUM_EXCHANGE
#define IGNEUM_EXCHANGE 0
#endif
#ifdef __OPENCL_VERSION__
#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1)))
#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n]
#if IGNEUM_EXCHANGE == 1
#ifdef cl_khr_subgroups
#pragma OPENCL EXTENSION cl_khr_subgroups : enable
#endif
#ifdef cl_khr_subgroup_shuffle
#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable
#endif
#elif IGNEUM_EXCHANGE == 2
#pragma OPENCL EXTENSION cl_intel_subgroups : enable
#endif
#else
// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros.
#include "emu_opencl.h"
#endif
#if IGNEUM_EXCHANGE == 1
#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m))
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
#elif IGNEUM_EXCHANGE == 2
#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m))
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
#else
// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per
// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane
// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's
// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier.
#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; }
#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; }
#endif
static inline uint splitmix32(uint x) {
x ^= x >> 16; x *= 0x7feb352du;
x ^= x >> 15; x *= 0x846ca68bu;
x ^= x >> 16;
return x;
}
// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32.
static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); }
// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x.
static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); }
static inline uint ds_elem(uint i, uint d0, uint d1) {
uint x = i ^ d0;
x *= 0x9E3779B1u; x ^= x >> 15;
x += d1;
x *= 0x85EBCA77u; x ^= x >> 13;
x *= 0xC2B2AE3Du; x ^= x >> 16;
return x;
}
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
#define MH_CACHE_LINE_MASK 0x003fffffu
#define MH_SEGMENT_LINES 64u
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
// y = ChaCha12 core(x) + x
static inline void mh_chacha_block(const uint* x, uint* y) {
for (uint i = 0u; i < 16u; ++i) y[i] = x[i];
for (uint r = 0u; r < 6u; ++r) {
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
}
for (uint i = 0u; i < 16u; ++i) y[i] += x[i];
}
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
static inline void mh_cache_segment(__global uint* cache, uint seg) {
uint prev[16]; uint x[16]; uint y[16];
for (uint i = 0u; i < 16u; ++i) prev[i] = 0u;
for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) {
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
x[4] = 0xceed56d7u ^ prev[4];
x[5] = 0x9ba270d2u ^ prev[5];
x[6] = 0x82caab2du ^ prev[6];
x[7] = 0x81ebce0eu ^ prev[7];
x[8] = 0x12b6ecf1u ^ prev[8];
x[9] = 0xd0f3fd7cu ^ prev[9];
x[10] = 0xd872eefeu ^ prev[10];
x[11] = 0xc158c7bdu ^ prev[11];
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
mh_chacha_block(x, y);
__global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
}
}
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
static inline void mh_mixer(uint* s, uint rk) {
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
}
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
static inline void mh_item(__global const uint* cache, uint t, uint* s) {
s[0] = 0xceed56d7u;
s[1] = 0x9ba270d2u;
s[2] = 0x82caab2du;
s[3] = 0x81ebce0eu;
s[4] = 0x12b6ecf1u;
s[5] = 0xd0f3fd7cu;
s[6] = 0xd872eefeu;
s[7] = 0xc158c7bdu;
s[8] = t * 0xf351d601u + 0xc6892460u;
s[9] = t * 0xa3bb398fu + 0x25b7228au;
s[10] = t * 0xb5a09e35u + 0xcd515004u;
s[11] = t * 0x7509c9c1u + 0x2846527au;
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
s[14] = t * 0xded91851u + 0x82961bacu;
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
for (uint r = 0u; r < 8u; ++r) {
mh_mixer(s, 0x9E3779B9u * (r + 1u));
__global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i];
}
mh_mixer(s, 0x9E3779B9u * 9u);
}
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 1 3 4 8 of w.
static inline uint mh_j(uint w) { return ((w >> 1u) & 1u) | (((w >> 3u) & 1u) << 1) | (((w >> 4u) & 1u) << 2) | (((w >> 8u) & 1u) << 3); }
static inline uint mh_t(uint w) { w = (w & 0x000000ffu) | ((w >> 9u) << 8u); w = (w & 0x0000000fu) | ((w >> 5u) << 4u); w = (w & 0x00000007u) | ((w >> 4u) << 3u); w = (w & 0x00000001u) | ((w >> 2u) << 1u); return w; }
static inline uint mh_addr(uint t, uint j) { uint w = t; w = ((w >> 1u) << 2u) | (w & 0x00000001u) | (((j >> 0u) & 1u) << 1u); w = ((w >> 3u) << 4u) | (w & 0x00000007u) | (((j >> 1u) & 1u) << 3u); w = ((w >> 4u) << 5u) | (w & 0x0000000fu) | (((j >> 2u) & 1u) << 4u); w = ((w >> 8u) << 9u) | (w & 0x000000ffu) | (((j >> 3u) & 1u) << 8u); return w; }
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }
// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item.
// The same constants as memhard.h in this pack (one emitter, three dialects).
__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) {
uint seg = (uint)get_global_id(0);
if (seg < nSegments) mh_cache_segment(cache, seg);
}
__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) {
uint t = (uint)get_global_id(0);
if (t < nItems) {
uint s[16];
mh_item(cache, t, s);
for (uint i = 0u; i < 16u; ++i) ds[(ulong)mh_addr(t, i)] = s[i];
}
}
// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the
// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and
// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all).
IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) {
uint gid = (uint)get_global_id(0);
uint lid = (uint)get_local_id(0);
uint nonce = baseNonce + gid;
uint r0, r1, r2, r3, r4, r5, r6, r7;
#if IGNEUM_EXCHANGE == 0
IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
uint xk = 0u;
#else
(void)lid;
#endif
{ uint x = nonce ^ 0x667d0fbdu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x7b8e5963u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
{ uint x = nonce ^ 0x7b8e5963u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x31c67e5eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
{ uint x = nonce ^ 0x31c67e5eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4529ddc6u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
{ uint x = nonce ^ 0x4529ddc6u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef19d6d8u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
{ uint x = nonce ^ 0xef19d6d8u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xaccf6211u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
{ uint x = nonce ^ 0xaccf6211u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xda0aed32u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
{ uint x = nonce ^ 0xda0aed32u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xabc6df31u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
{ uint x = nonce ^ 0xabc6df31u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x667d0fbdu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
for (uint it = 0u; it < 8u; ++it) {
uint sel = r0;
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
r7 = r7 ^ r0; // 1 xor
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 1u); r3 = r3 ^ t_; } // 2 shfl
r4 = r4 - r1; // 3 sub
r2 = r0 * r4 + r2; // 4 mad
r4 = rotl_imm(r4, 9u); // 5 rotl
r0 = rotr_var(r0, r2); // 6 rotr
r6 = r6 ^ ds[((rotl_imm(r7 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
r1 = r1 ^ ds[((rotl_imm(r4 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
r1 = r1 ^ ds[((rotl_imm(r2 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
r7 = r7 ^ ds[((rotl_imm(r0 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
r7 = r7 ^ ds[((rotl_imm(r1 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
r5 = r5 * r4; // 12 mul
r7 = r7 ^ ds[((rotl_imm(r6 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 16u); r6 = r6 ^ t_; } // 14 shfl
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r3 = r3 ^ t_; } // 15 shfl
r2 = r2 | r7; // 16 or
r0 = rotr_var(r0, r6); // 17 rotr
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
r7 = rotl_imm(r7, 24u); // 19 rotl
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
r2 = r2 | r6; // 21 or
r6 = r5 * r4 + r6; // 22 mad
r1 = r1 | r0; // 23 or
r6 = r6 ^ r1; // 24 xor
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
r7 = rotr_var(r7, r0); // 26 rotr
r4 = r5 * r7 + r4; // 27 mad
r3 = r2 * r4 + r3; // 28 mad
r1 = r1 ^ ds[((rotl_imm(r4 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
r2 = r2 ^ ds[((rotl_imm(r3 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
r1 = r1 ^ ds[((rotl_imm(r5 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
r2 = mul_hi(r2, r5); // 35 mulhi
r4 = r4 ^ r2; // 36 xor
r6 = r6 * r5; // 37 mul
r7 = r7 ^ r0; // 38 xor
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
r2 = rotr_var(r2, r3); // 40 rotr
r7 = r7 - r0; // 41 sub
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 43 shfl
r0 = r0 ^ ds[((rotl_imm(r7 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
r3 = r3 ^ ds[((rotl_imm(r1 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
r6 = r6 ^ ds[((rotl_imm(r3 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
r4 = mul_hi(r4, r2); // 48 mulhi
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r0 = r0 ^ t_; } // 50 shfl
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
r5 = r5 ^ ds[((rotl_imm(r2 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
r6 = rotl_imm(r6, 12u); // 53 rotl
r3 = rotl_imm(r3, 12u); // 54 rotl
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
r1 = r1 ^ ds[((rotl_imm(r4 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
r5 = r5 ^ ds[((rotl_imm(r0 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
r3 = mul_hi(r3, r2); // 60 mulhi
r6 = r6 | r4; // 61 or
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r5 = r5 ^ t_; } // 62 shfl
r3 = r3 ^ ds[((rotl_imm(r6 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
}
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((ulong)hi << 32) | (ulong)lo;
}
#if IGNEUM_EXCHANGE != 0
// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the
// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a
// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md.
IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) {
if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); }
}
#endif

View file

@ -0,0 +1,163 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal).
// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC.
#include <cuda_runtime.h>
#include <cstdint>
#include "program.h"
#include "memhard.h"
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
x ^= x >> 16; x *= 0x7feb352du;
x ^= x >> 15; x *= 0x846ca68bu;
x ^= x >> 16;
return x;
}
// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31.
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x.
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) {
uint32_t x = i ^ d0;
x *= 0x9E3779B1u; x ^= x >> 15;
x += d1;
x *= 0x85EBCA77u; x ^= x >> 13;
x *= 0xC2B2AE3Du; x ^= x >> 16;
return x;
}
// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item.
// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host.
__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) {
uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x;
if (seg < nSegments) mh_cache_segment(cache, seg);
}
__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
uint32_t t = blockIdx.x * blockDim.x + threadIdx.x;
if (t < nItems) {
uint32_t s[16];
mh_item(cache, t, s);
for (uint32_t i = 0u; i < 16u; ++i) ds[(size_t)mh_addr(t, i)] = s[i];
}
}
// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every
// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a
// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid.
__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) {
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
uint32_t nonce = baseNonce + gid;
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
{ uint32_t x = nonce ^ 0x667d0fbdu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x7b8e5963u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
{ uint32_t x = nonce ^ 0x7b8e5963u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x31c67e5eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
{ uint32_t x = nonce ^ 0x31c67e5eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4529ddc6u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
{ uint32_t x = nonce ^ 0x4529ddc6u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef19d6d8u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
{ uint32_t x = nonce ^ 0xef19d6d8u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xaccf6211u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
{ uint32_t x = nonce ^ 0xaccf6211u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xda0aed32u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
{ uint32_t x = nonce ^ 0xda0aed32u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xabc6df31u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
{ uint32_t x = nonce ^ 0xabc6df31u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x667d0fbdu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
for (uint32_t it = 0u; it < 8u; ++it) {
uint32_t sel = r0;
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
r7 = r7 ^ r0; // 1 xor
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r6, 1); // 2 shfl
r4 = r4 - r1; // 3 sub
r2 = r0 * r4 + r2; // 4 mad
r4 = rotl_imm(r4, 9u); // 5 rotl
r0 = rotr_var(r0, r2); // 6 rotr
r6 = r6 ^ ds[((rotl_imm(r7 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
r1 = r1 ^ ds[((rotl_imm(r4 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
r1 = r1 ^ ds[((rotl_imm(r2 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
r7 = r7 ^ ds[((rotl_imm(r0 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
r7 = r7 ^ ds[((rotl_imm(r1 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
r5 = r5 * r4; // 12 mul
r7 = r7 ^ ds[((rotl_imm(r6 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r5, 16); // 14 shfl
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r5, 1); // 15 shfl
r2 = r2 | r7; // 16 or
r0 = rotr_var(r0, r6); // 17 rotr
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
r7 = rotl_imm(r7, 24u); // 19 rotl
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
r2 = r2 | r6; // 21 or
r6 = r5 * r4 + r6; // 22 mad
r1 = r1 | r0; // 23 or
r6 = r6 ^ r1; // 24 xor
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
r7 = rotr_var(r7, r0); // 26 rotr
r4 = r5 * r7 + r4; // 27 mad
r3 = r2 * r4 + r3; // 28 mad
r1 = r1 ^ ds[((rotl_imm(r4 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
r2 = r2 ^ ds[((rotl_imm(r3 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
r1 = r1 ^ ds[((rotl_imm(r5 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
r2 = __umulhi(r2, r5); // 35 mulhi
r4 = r4 ^ r2; // 36 xor
r6 = r6 * r5; // 37 mul
r7 = r7 ^ r0; // 38 xor
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
r2 = rotr_var(r2, r3); // 40 rotr
r7 = r7 - r0; // 41 sub
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 43 shfl
r0 = r0 ^ ds[((rotl_imm(r7 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
r3 = r3 ^ ds[((rotl_imm(r1 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
r6 = r6 ^ ds[((rotl_imm(r3 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
r4 = __umulhi(r4, r2); // 48 mulhi
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r7, 4); // 50 shfl
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
r5 = r5 ^ ds[((rotl_imm(r2 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
r6 = rotl_imm(r6, 12u); // 53 rotl
r3 = rotl_imm(r3, 12u); // 54 rotl
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
r1 = r1 ^ ds[((rotl_imm(r4 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
r5 = r5 ^ ds[((rotl_imm(r0 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
r3 = __umulhi(r3, r2); // 60 mulhi
r6 = r6 | r4; // 61 or
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 62 shfl
r3 = r3 ^ ds[((rotl_imm(r6 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
}
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
}
// Host-side launch wrappers. Declared in program.h, called from host.cu.
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) {
if (nSegments == 0u) return cudaErrorInvalidValue;
uint32_t block = 256u;
uint32_t grid = (nSegments + block - 1u) / block;
igneum_cache_fill<<<grid, block>>>(cache, nSegments);
return cudaGetLastError();
}
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
if (nItems == 0u) return cudaErrorInvalidValue;
uint32_t block = 256u;
uint32_t grid = (nItems + block - 1u) / block;
igneum_build<<<grid, block>>>(ds, cache, nItems);
return cudaGetLastError();
}
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
uint32_t nonces, uint32_t blockWarps) {
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
uint32_t block = 32u * blockWarps;
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
igneum_hash<<<nonces / block, block>>>(ds, out, baseNonce, mask);
return cudaGetLastError();
}
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
cudaFuncAttributes attr;
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash);
if (e != cudaSuccess) return e;
*numRegs = attr.numRegs;
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0);
}

View file

@ -0,0 +1,375 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal).
// Built from source at runtime by proto-opencl/host.c, which passes these defines:
// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit)
// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default)
// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32
// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition
// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units;
// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32.
#ifndef IGNEUM_GROUP
#define IGNEUM_GROUP 32
#endif
#ifndef IGNEUM_EXCHANGE
#define IGNEUM_EXCHANGE 0
#endif
#ifdef __OPENCL_VERSION__
#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1)))
#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n]
#if IGNEUM_EXCHANGE == 1
#ifdef cl_khr_subgroups
#pragma OPENCL EXTENSION cl_khr_subgroups : enable
#endif
#ifdef cl_khr_subgroup_shuffle
#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable
#endif
#elif IGNEUM_EXCHANGE == 2
#pragma OPENCL EXTENSION cl_intel_subgroups : enable
#endif
#else
// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros.
#include "emu_opencl.h"
#endif
#if IGNEUM_EXCHANGE == 1
#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m))
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
#elif IGNEUM_EXCHANGE == 2
#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m))
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
#else
// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per
// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane
// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's
// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier.
#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; }
#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; }
#endif
static inline uint splitmix32(uint x) {
x ^= x >> 16; x *= 0x7feb352du;
x ^= x >> 15; x *= 0x846ca68bu;
x ^= x >> 16;
return x;
}
// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32.
static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); }
// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x.
static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); }
static inline uint ds_elem(uint i, uint d0, uint d1) {
uint x = i ^ d0;
x *= 0x9E3779B1u; x ^= x >> 15;
x += d1;
x *= 0x85EBCA77u; x ^= x >> 13;
x *= 0xC2B2AE3Du; x ^= x >> 16;
return x;
}
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
#define MH_CACHE_LINE_MASK 0x003fffffu
#define MH_SEGMENT_LINES 64u
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
// y = ChaCha12 core(x) + x
static inline void mh_chacha_block(const uint* x, uint* y) {
for (uint i = 0u; i < 16u; ++i) y[i] = x[i];
for (uint r = 0u; r < 6u; ++r) {
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
}
for (uint i = 0u; i < 16u; ++i) y[i] += x[i];
}
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
static inline void mh_cache_segment(__global uint* cache, uint seg) {
uint prev[16]; uint x[16]; uint y[16];
for (uint i = 0u; i < 16u; ++i) prev[i] = 0u;
for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) {
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
x[4] = 0xceed56d7u ^ prev[4];
x[5] = 0x9ba270d2u ^ prev[5];
x[6] = 0x82caab2du ^ prev[6];
x[7] = 0x81ebce0eu ^ prev[7];
x[8] = 0x12b6ecf1u ^ prev[8];
x[9] = 0xd0f3fd7cu ^ prev[9];
x[10] = 0xd872eefeu ^ prev[10];
x[11] = 0xc158c7bdu ^ prev[11];
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
mh_chacha_block(x, y);
__global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
}
}
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
static inline void mh_mixer(uint* s, uint rk) {
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
}
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
static inline void mh_item(__global const uint* cache, uint t, uint* s) {
s[0] = 0xceed56d7u;
s[1] = 0x9ba270d2u;
s[2] = 0x82caab2du;
s[3] = 0x81ebce0eu;
s[4] = 0x12b6ecf1u;
s[5] = 0xd0f3fd7cu;
s[6] = 0xd872eefeu;
s[7] = 0xc158c7bdu;
s[8] = t * 0xf351d601u + 0xc6892460u;
s[9] = t * 0xa3bb398fu + 0x25b7228au;
s[10] = t * 0xb5a09e35u + 0xcd515004u;
s[11] = t * 0x7509c9c1u + 0x2846527au;
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
s[14] = t * 0xded91851u + 0x82961bacu;
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
for (uint r = 0u; r < 8u; ++r) {
mh_mixer(s, 0x9E3779B9u * (r + 1u));
__global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i];
}
mh_mixer(s, 0x9E3779B9u * 9u);
}
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 1 3 4 8 of w.
static inline uint mh_j(uint w) { return ((w >> 1u) & 1u) | (((w >> 3u) & 1u) << 1) | (((w >> 4u) & 1u) << 2) | (((w >> 8u) & 1u) << 3); }
static inline uint mh_t(uint w) { w = (w & 0x000000ffu) | ((w >> 9u) << 8u); w = (w & 0x0000000fu) | ((w >> 5u) << 4u); w = (w & 0x00000007u) | ((w >> 4u) << 3u); w = (w & 0x00000001u) | ((w >> 2u) << 1u); return w; }
static inline uint mh_addr(uint t, uint j) { uint w = t; w = ((w >> 1u) << 2u) | (w & 0x00000001u) | (((j >> 0u) & 1u) << 1u); w = ((w >> 3u) << 4u) | (w & 0x00000007u) | (((j >> 1u) & 1u) << 3u); w = ((w >> 4u) << 5u) | (w & 0x0000000fu) | (((j >> 2u) & 1u) << 4u); w = ((w >> 8u) << 9u) | (w & 0x000000ffu) | (((j >> 3u) & 1u) << 8u); return w; }
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }
// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item.
// The same constants as memhard.h in this pack (one emitter, three dialects).
__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) {
uint seg = (uint)get_global_id(0);
if (seg < nSegments) mh_cache_segment(cache, seg);
}
__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) {
uint t = (uint)get_global_id(0);
if (t < nItems) {
uint s[16];
mh_item(cache, t, s);
for (uint i = 0u; i < 16u; ++i) ds[(ulong)mh_addr(t, i)] = s[i];
}
}
// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the
// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and
// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all).
IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) {
uint gid = (uint)get_global_id(0);
uint lid = (uint)get_local_id(0);
uint nonce = baseNonce + gid;
uint r0, r1, r2, r3, r4, r5, r6, r7;
#if IGNEUM_EXCHANGE == 0
IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
uint xk = 0u;
#else
(void)lid;
#endif
{ uint x = nonce ^ 0x667d0fbdu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x7b8e5963u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
{ uint x = nonce ^ 0x7b8e5963u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x31c67e5eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
{ uint x = nonce ^ 0x31c67e5eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4529ddc6u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
{ uint x = nonce ^ 0x4529ddc6u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef19d6d8u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
{ uint x = nonce ^ 0xef19d6d8u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xaccf6211u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
{ uint x = nonce ^ 0xaccf6211u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xda0aed32u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
{ uint x = nonce ^ 0xda0aed32u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xabc6df31u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
{ uint x = nonce ^ 0xabc6df31u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x667d0fbdu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
for (uint it = 0u; it < 8u; ++it) {
uint sel = r0;
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
r7 = r7 ^ r0; // 1 xor
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 1u); r3 = r3 ^ t_; } // 2 shfl
r4 = r4 - r1; // 3 sub
r2 = r0 * r4 + r2; // 4 mad
r4 = rotl_imm(r4, 9u); // 5 rotl
r0 = rotr_var(r0, r2); // 6 rotr
r6 = r6 ^ ds[((rotl_imm(r7 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
r1 = r1 ^ ds[((rotl_imm(r4 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
r1 = r1 ^ ds[((rotl_imm(r2 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
r7 = r7 ^ ds[((rotl_imm(r0 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
r7 = r7 ^ ds[((rotl_imm(r1 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
r5 = r5 * r4; // 12 mul
r7 = r7 ^ ds[((rotl_imm(r6 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 16u); r6 = r6 ^ t_; } // 14 shfl
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r3 = r3 ^ t_; } // 15 shfl
r2 = r2 | r7; // 16 or
r0 = rotr_var(r0, r6); // 17 rotr
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
r7 = rotl_imm(r7, 24u); // 19 rotl
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
r2 = r2 | r6; // 21 or
r6 = r5 * r4 + r6; // 22 mad
r1 = r1 | r0; // 23 or
r6 = r6 ^ r1; // 24 xor
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
r7 = rotr_var(r7, r0); // 26 rotr
r4 = r5 * r7 + r4; // 27 mad
r3 = r2 * r4 + r3; // 28 mad
r1 = r1 ^ ds[((rotl_imm(r4 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
r2 = r2 ^ ds[((rotl_imm(r3 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
r1 = r1 ^ ds[((rotl_imm(r5 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
r2 = mul_hi(r2, r5); // 35 mulhi
r4 = r4 ^ r2; // 36 xor
r6 = r6 * r5; // 37 mul
r7 = r7 ^ r0; // 38 xor
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
r2 = rotr_var(r2, r3); // 40 rotr
r7 = r7 - r0; // 41 sub
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 43 shfl
r0 = r0 ^ ds[((rotl_imm(r7 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
r3 = r3 ^ ds[((rotl_imm(r1 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
r6 = r6 ^ ds[((rotl_imm(r3 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
r4 = mul_hi(r4, r2); // 48 mulhi
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r0 = r0 ^ t_; } // 50 shfl
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
r5 = r5 ^ ds[((rotl_imm(r2 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
r6 = rotl_imm(r6, 12u); // 53 rotl
r3 = rotl_imm(r3, 12u); // 54 rotl
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
r1 = r1 ^ ds[((rotl_imm(r4 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
r5 = r5 ^ ds[((rotl_imm(r0 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
r3 = mul_hi(r3, r2); // 60 mulhi
r6 = r6 | r4; // 61 or
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r5 = r5 ^ t_; } // 62 shfl
r3 = r3 ^ ds[((rotl_imm(r6 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
}
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((ulong)hi << 32) | (ulong)lo;
}
#if IGNEUM_EXCHANGE != 0
// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the
// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a
// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md.
IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) {
if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); }
}
#endif
// Header-bound variant (bind.rs): the init words come from initw, not SEEDW. Same body as igneum_hash.
IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw) {
uint gid = (uint)get_global_id(0);
uint lid = (uint)get_local_id(0);
uint nonce = baseNonce + gid;
uint r0, r1, r2, r3, r4, r5, r6, r7;
uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7];
#if IGNEUM_EXCHANGE == 0
IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
uint xk = 0u;
#else
(void)lid;
#endif
{ uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; }
{ uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; }
{ uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; }
{ uint x = nonce ^ iw3; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw4; }
{ uint x = nonce ^ iw4; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw5; }
{ uint x = nonce ^ iw5; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw6; }
{ uint x = nonce ^ iw6; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw7; }
{ uint x = nonce ^ iw7; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw0; }
for (uint it = 0u; it < 8u; ++it) {
uint sel = r0;
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
r7 = r7 ^ r0; // 1 xor
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 1u); r3 = r3 ^ t_; } // 2 shfl
r4 = r4 - r1; // 3 sub
r2 = r0 * r4 + r2; // 4 mad
r4 = rotl_imm(r4, 9u); // 5 rotl
r0 = rotr_var(r0, r2); // 6 rotr
r6 = r6 ^ ds[((rotl_imm(r7 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
r1 = r1 ^ ds[((rotl_imm(r4 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
r1 = r1 ^ ds[((rotl_imm(r2 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
r7 = r7 ^ ds[((rotl_imm(r0 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
r7 = r7 ^ ds[((rotl_imm(r1 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
r5 = r5 * r4; // 12 mul
r7 = r7 ^ ds[((rotl_imm(r6 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 16u); r6 = r6 ^ t_; } // 14 shfl
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r3 = r3 ^ t_; } // 15 shfl
r2 = r2 | r7; // 16 or
r0 = rotr_var(r0, r6); // 17 rotr
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
r7 = rotl_imm(r7, 24u); // 19 rotl
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
r2 = r2 | r6; // 21 or
r6 = r5 * r4 + r6; // 22 mad
r1 = r1 | r0; // 23 or
r6 = r6 ^ r1; // 24 xor
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
r7 = rotr_var(r7, r0); // 26 rotr
r4 = r5 * r7 + r4; // 27 mad
r3 = r2 * r4 + r3; // 28 mad
r1 = r1 ^ ds[((rotl_imm(r4 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
r2 = r2 ^ ds[((rotl_imm(r3 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
r1 = r1 ^ ds[((rotl_imm(r5 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
r2 = mul_hi(r2, r5); // 35 mulhi
r4 = r4 ^ r2; // 36 xor
r6 = r6 * r5; // 37 mul
r7 = r7 ^ r0; // 38 xor
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
r2 = rotr_var(r2, r3); // 40 rotr
r7 = r7 - r0; // 41 sub
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 43 shfl
r0 = r0 ^ ds[((rotl_imm(r7 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
r3 = r3 ^ ds[((rotl_imm(r1 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
r6 = r6 ^ ds[((rotl_imm(r3 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
r4 = mul_hi(r4, r2); // 48 mulhi
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r0 = r0 ^ t_; } // 50 shfl
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
r5 = r5 ^ ds[((rotl_imm(r2 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
r6 = rotl_imm(r6, 12u); // 53 rotl
r3 = rotl_imm(r3, 12u); // 54 rotl
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
r1 = r1 ^ ds[((rotl_imm(r4 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
r5 = r5 ^ ds[((rotl_imm(r0 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
r3 = mul_hi(r3, r2); // 60 mulhi
r6 = r6 | r4; // 61 or
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r5 = r5 ^ t_; } // 62 shfl
r3 = r3 ^ ds[((rotl_imm(r6 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
}
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((ulong)hi << 32) | (ulong)lo;
}

View file

@ -0,0 +1,123 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
// Header-bound twin of igneum_hash in kernel.cu: the init words come from a kernel argument, not SEEDW.
// Host declarations (also in program_bound.h if present):
// struct IgneumInitWords { uint32_t w[8]; };
// cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
// IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps);
// cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
#include <cuda_runtime.h>
#include <cstdint>
#include "program.h"
struct IgneumInitWords { uint32_t w[8]; };
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
x ^= x >> 16; x *= 0x7feb352du;
x ^= x >> 15; x *= 0x846ca68bu;
x ^= x >> 16;
return x;
}
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw) {
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
uint32_t nonce = baseNonce + gid;
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
{ uint32_t x = nonce ^ iw.w[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw.w[1]; }
{ uint32_t x = nonce ^ iw.w[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw.w[2]; }
{ uint32_t x = nonce ^ iw.w[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw.w[3]; }
{ uint32_t x = nonce ^ iw.w[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw.w[4]; }
{ uint32_t x = nonce ^ iw.w[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw.w[5]; }
{ uint32_t x = nonce ^ iw.w[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw.w[6]; }
{ uint32_t x = nonce ^ iw.w[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw.w[7]; }
{ uint32_t x = nonce ^ iw.w[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw.w[0]; }
for (uint32_t it = 0u; it < 8u; ++it) {
uint32_t sel = r0;
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
r7 = r7 ^ r0; // 1 xor
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r6, 1); // 2 shfl
r4 = r4 - r1; // 3 sub
r2 = r0 * r4 + r2; // 4 mad
r4 = rotl_imm(r4, 9u); // 5 rotl
r0 = rotr_var(r0, r2); // 6 rotr
r6 = r6 ^ ds[((rotl_imm(r7 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
r1 = r1 ^ ds[((rotl_imm(r4 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
r1 = r1 ^ ds[((rotl_imm(r2 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
r7 = r7 ^ ds[((rotl_imm(r0 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
r7 = r7 ^ ds[((rotl_imm(r1 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
r5 = r5 * r4; // 12 mul
r7 = r7 ^ ds[((rotl_imm(r6 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r5, 16); // 14 shfl
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r5, 1); // 15 shfl
r2 = r2 | r7; // 16 or
r0 = rotr_var(r0, r6); // 17 rotr
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
r7 = rotl_imm(r7, 24u); // 19 rotl
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
r2 = r2 | r6; // 21 or
r6 = r5 * r4 + r6; // 22 mad
r1 = r1 | r0; // 23 or
r6 = r6 ^ r1; // 24 xor
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
r7 = rotr_var(r7, r0); // 26 rotr
r4 = r5 * r7 + r4; // 27 mad
r3 = r2 * r4 + r3; // 28 mad
r1 = r1 ^ ds[((rotl_imm(r4 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
r2 = r2 ^ ds[((rotl_imm(r3 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
r1 = r1 ^ ds[((rotl_imm(r5 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
r2 = __umulhi(r2, r5); // 35 mulhi
r4 = r4 ^ r2; // 36 xor
r6 = r6 * r5; // 37 mul
r7 = r7 ^ r0; // 38 xor
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
r2 = rotr_var(r2, r3); // 40 rotr
r7 = r7 - r0; // 41 sub
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 43 shfl
r0 = r0 ^ ds[((rotl_imm(r7 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
r3 = r3 ^ ds[((rotl_imm(r1 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
r6 = r6 ^ ds[((rotl_imm(r3 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
r4 = __umulhi(r4, r2); // 48 mulhi
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r7, 4); // 50 shfl
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
r5 = r5 ^ ds[((rotl_imm(r2 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
r6 = rotl_imm(r6, 12u); // 53 rotl
r3 = rotl_imm(r3, 12u); // 54 rotl
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
r1 = r1 ^ ds[((rotl_imm(r4 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
r5 = r5 ^ ds[((rotl_imm(r0 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
r3 = __umulhi(r3, r2); // 60 mulhi
r6 = r6 | r4; // 61 or
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 62 shfl
r3 = r3 ^ ds[((rotl_imm(r6 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
}
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
}
cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps) {
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
uint32_t block = 32u * blockWarps;
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
igneum_hash_bound<<<nonces / block, block>>>(ds, out, baseNonce, mask, iw);
return cudaGetLastError();
}
cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
cudaFuncAttributes attr;
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash_bound);
if (e != cudaSuccess) return e;
*numRegs = attr.numRegs;
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash_bound, (int)(32u * blockWarps), 0);
}

View file

@ -0,0 +1,112 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
// Memory-hard dataset core, the same text that the Mac's Metal kernels and CPU verifier were checked against.
// Included by kernel.cu (device), host.cu (host reference) and proto-opencl/host.c (C99 host reference).
// See proto-metal/MEMHARD.md for the construction. kernel.cl carries the same text in OpenCL C.
#pragma once
#ifdef __cplusplus
#include <cstdint>
#else
#include <stdint.h>
#endif
#if defined(__CUDACC__)
#define IGNEUM_HD __host__ __device__ __forceinline__
#elif defined(_MSC_VER) && !defined(__cplusplus)
#define IGNEUM_HD static __inline
#else
#define IGNEUM_HD static inline
#endif
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
#define MH_CACHE_LINE_MASK 0x003fffffu
#define MH_SEGMENT_LINES 64u
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
IGNEUM_HD uint32_t mh_rotl(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
// y = ChaCha12 core(x) + x
IGNEUM_HD void mh_chacha_block(const uint32_t* x, uint32_t* y) {
for (uint32_t i = 0u; i < 16u; ++i) y[i] = x[i];
for (uint32_t r = 0u; r < 6u; ++r) {
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
}
for (uint32_t i = 0u; i < 16u; ++i) y[i] += x[i];
}
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
IGNEUM_HD void mh_cache_segment(uint32_t* cache, uint32_t seg) {
uint32_t prev[16]; uint32_t x[16]; uint32_t y[16];
for (uint32_t i = 0u; i < 16u; ++i) prev[i] = 0u;
for (uint32_t j = 0u; j < MH_SEGMENT_LINES; ++j) {
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
x[4] = 0xceed56d7u ^ prev[4];
x[5] = 0x9ba270d2u ^ prev[5];
x[6] = 0x82caab2du ^ prev[6];
x[7] = 0x81ebce0eu ^ prev[7];
x[8] = 0x12b6ecf1u ^ prev[8];
x[9] = 0xd0f3fd7cu ^ prev[9];
x[10] = 0xd872eefeu ^ prev[10];
x[11] = 0xc158c7bdu ^ prev[11];
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
mh_chacha_block(x, y);
uint32_t* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
for (uint32_t i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
}
}
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
IGNEUM_HD void mh_mixer(uint32_t* s, uint32_t rk) {
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
}
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
IGNEUM_HD void mh_item(const uint32_t* cache, uint32_t t, uint32_t* s) {
s[0] = 0xceed56d7u;
s[1] = 0x9ba270d2u;
s[2] = 0x82caab2du;
s[3] = 0x81ebce0eu;
s[4] = 0x12b6ecf1u;
s[5] = 0xd0f3fd7cu;
s[6] = 0xd872eefeu;
s[7] = 0xc158c7bdu;
s[8] = t * 0xf351d601u + 0xc6892460u;
s[9] = t * 0xa3bb398fu + 0x25b7228au;
s[10] = t * 0xb5a09e35u + 0xcd515004u;
s[11] = t * 0x7509c9c1u + 0x2846527au;
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
s[14] = t * 0xded91851u + 0x82961bacu;
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
for (uint32_t r = 0u; r < 8u; ++r) {
mh_mixer(s, 0x9E3779B9u * (r + 1u));
const uint32_t* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= line[i];
}
mh_mixer(s, 0x9E3779B9u * 9u);
}
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 1 3 4 8 of w.
IGNEUM_HD uint32_t mh_j(uint32_t w) { return ((w >> 1u) & 1u) | (((w >> 3u) & 1u) << 1) | (((w >> 4u) & 1u) << 2) | (((w >> 8u) & 1u) << 3); }
IGNEUM_HD uint32_t mh_t(uint32_t w) { w = (w & 0x000000ffu) | ((w >> 9u) << 8u); w = (w & 0x0000000fu) | ((w >> 5u) << 4u); w = (w & 0x00000007u) | ((w >> 4u) << 3u); w = (w & 0x00000001u) | ((w >> 2u) << 1u); return w; }
IGNEUM_HD uint32_t mh_addr(uint32_t t, uint32_t j) { uint32_t w = t; w = ((w >> 1u) << 2u) | (w & 0x00000001u) | (((j >> 0u) & 1u) << 1u); w = ((w >> 3u) << 4u) | (w & 0x00000007u) | (((j >> 1u) & 1u) << 3u); w = ((w >> 4u) << 5u) | (w & 0x0000000fu) | (((j >> 2u) & 1u) << 4u); w = ((w >> 8u) << 9u) | (w & 0x000000ffu) | (((j >> 3u) & 1u) << 8u); return w; }
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
IGNEUM_HD uint32_t mh_word(const uint32_t* cache, uint32_t w) { uint32_t s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }

View file

@ -0,0 +1,109 @@
#include <metal_stdlib>
using namespace metal;
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
#define MH_CACHE_LINE_MASK 0x003fffffu
#define MH_SEGMENT_LINES 64u
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
// y = ChaCha12 core(x) + x
inline void mh_chacha_block(const thread uint* x, thread uint* y) {
for (uint i = 0u; i < 16u; ++i) y[i] = x[i];
for (uint r = 0u; r < 6u; ++r) {
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
}
for (uint i = 0u; i < 16u; ++i) y[i] += x[i];
}
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
inline void mh_cache_segment(device uint* cache, uint seg) {
uint prev[16]; uint x[16]; uint y[16];
for (uint i = 0u; i < 16u; ++i) prev[i] = 0u;
for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) {
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
x[4] = 0xceed56d7u ^ prev[4];
x[5] = 0x9ba270d2u ^ prev[5];
x[6] = 0x82caab2du ^ prev[6];
x[7] = 0x81ebce0eu ^ prev[7];
x[8] = 0x12b6ecf1u ^ prev[8];
x[9] = 0xd0f3fd7cu ^ prev[9];
x[10] = 0xd872eefeu ^ prev[10];
x[11] = 0xc158c7bdu ^ prev[11];
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
mh_chacha_block(x, y);
device uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
}
}
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
inline void mh_mixer(thread uint* s, uint rk) {
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
}
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
inline void mh_item(device const uint* cache, uint t, thread uint* s) {
s[0] = 0xceed56d7u;
s[1] = 0x9ba270d2u;
s[2] = 0x82caab2du;
s[3] = 0x81ebce0eu;
s[4] = 0x12b6ecf1u;
s[5] = 0xd0f3fd7cu;
s[6] = 0xd872eefeu;
s[7] = 0xc158c7bdu;
s[8] = t * 0xf351d601u + 0xc6892460u;
s[9] = t * 0xa3bb398fu + 0x25b7228au;
s[10] = t * 0xb5a09e35u + 0xcd515004u;
s[11] = t * 0x7509c9c1u + 0x2846527au;
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
s[14] = t * 0xded91851u + 0x82961bacu;
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
for (uint r = 0u; r < 8u; ++r) {
mh_mixer(s, 0x9E3779B9u * (r + 1u));
device const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i];
}
mh_mixer(s, 0x9E3779B9u * 9u);
}
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 1 3 4 8 of w.
inline uint mh_j(uint w) { return ((w >> 1u) & 1u) | (((w >> 3u) & 1u) << 1) | (((w >> 4u) & 1u) << 2) | (((w >> 8u) & 1u) << 3); }
inline uint mh_t(uint w) { w = (w & 0x000000ffu) | ((w >> 9u) << 8u); w = (w & 0x0000000fu) | ((w >> 5u) << 4u); w = (w & 0x00000007u) | ((w >> 4u) << 3u); w = (w & 0x00000001u) | ((w >> 2u) << 1u); return w; }
inline uint mh_addr(uint t, uint j) { uint w = t; w = ((w >> 1u) << 2u) | (w & 0x00000001u) | (((j >> 0u) & 1u) << 1u); w = ((w >> 3u) << 4u) | (w & 0x00000007u) | (((j >> 1u) & 1u) << 3u); w = ((w >> 4u) << 5u) | (w & 0x0000000fu) | (((j >> 2u) & 1u) << 4u); w = ((w >> 8u) << 9u) | (w & 0x000000ffu) | (((j >> 3u) & 1u) << 8u); return w; }
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
inline uint mh_word(device const uint* cache, uint w) { uint s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }
// One thread per segment (2^16 threads).
kernel void igneum_cache_fill(device uint* cache [[buffer(0)]], uint gid [[thread_position_in_grid]]) {
mh_cache_segment(cache, gid);
}
// One thread per 64-byte item (dataset words / 16 threads).
kernel void igneum_build(device const uint* cache [[buffer(0)]], device uint* dataset [[buffer(1)]],
uint gid [[thread_position_in_grid]]) {
uint s[16];
mh_item(cache, gid, s);
for (uint i = 0u; i < 16u; ++i) dataset[mh_addr(gid, i)] = s[i];
}

View file

@ -0,0 +1,76 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
// Program metadata for host.cu plus the launch wrappers defined in kernel.cu.
// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros.
#pragma once
#ifdef __cplusplus
#include <cstdint>
#else
#include <stdint.h>
#endif
#ifndef IGNEUM_NO_CUDA
#include <cuda_runtime.h>
#endif
#define IGNEUM_SEED_STRING "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000"
#define IGNEUM_SEED_BYTES_HEX "edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07"
#define IGNEUM_GENERATOR 3
#define IGNEUM_PROGRAM_ATTEMPT 0
#define IGNEUM_PROGRAM_ID 0x73bcbfe8ccf988f1ull
#define IGNEUM_DAY_STRING "bytes:69676e65756d2d6461792ffa50000000000000"
#define IGNEUM_DAY_BYTES_HEX "69676e65756d2d6461792ffa50000000000000"
#define IGNEUM_DAY0 0xceed56d7u
#define IGNEUM_DAY1 0x9ba270d2u
#define IGNEUM_DATASET_LOG2 28
#define IGNEUM_MASK 0x0fffffffu
#define IGNEUM_LANES 32
#define IGNEUM_ITERATIONS 8
#define IGNEUM_INSTR_COUNT 64
#define IGNEUM_LOADS_PER_HASH 128
#define IGNEUM_WIDE_LOADS_PER_HASH 0
#define IGNEUM_OP_MIX "load=16 add=15 shfl=6 mad=4 or=4 rotl=4 rotr=4 xor=4 mulhi=3 mul=2 sub=2"
// Program class v3 (Counter ASIC 2.0, docs/plans/counter-asic-2-rollout.md): generator version 3; a worker that
// runs another class refuses this pack, and a job line names the class it wants (class=v3 era=<hex>).
#define IGNEUM_PROGRAM_CLASS "v3"
#define IGNEUM_ERA_SEED_HEX "df57136f2ad5f410e6145023090eea148c1342ccef5a5f541afa590be7e44745"
// Class v3 construction (Counter ASIC 2.0, 5 October 2026, docs/plans/mixer-x4.md): version 2 loads; the dataset item
// derivation applies the mixer IGNEUM_MIXER_MULT times per round (memhard.h), and the cache follows the growth rule.
#define IGNEUM_LOAD_CLASS "w4-era843155d7"
#define IGNEUM_LOAD_SLOTS 16
#define IGNEUM_LOAD_MIX { 100, 0, 0 }
#define IGNEUM_LOAD_WIDTH_COUNTS { 16, 0, 0 } // loads of 4, 16, 64 bytes per program
#define IGNEUM_BYTES_PER_HASH 512
#define IGNEUM_FOLD_ROT 11
#define IGNEUM_FOLD_MUL 0x9e3779b1u
// Era layout (5 October 2026, docs/plans/era-layout.md): NOT the lottery hash. Every dataset load reads
// idx = ((rotl(src * STRIDE_MUL, STRIDE_ROT) & window mask) | window offset) & MASK; the window of a load site is the
// dataset, a half or a quarter of it (IGNEUM_ERA_WINDOWS: site:shrink:offset); dataset word w holds word j(w) of item
// t(w) with j's bits at the INTERLEAVE positions (memhard.h: mh_t, mh_j, mh_addr).
#define IGNEUM_ERA_LABEL "843155d7"
#define IGNEUM_ERA_SEED_WORDS { 0x843155d7u, 0x8fb2bbb8u, 0x3af89788u, 0xc80fe6f2u, 0xb265c95eu, 0x001f1648u, 0x236e8759u, 0x6cada520u }
#define IGNEUM_ERA_ALLOWED_WIDTHS { 1, 0, 0 } // words, ascending, 0 = unused; one entry pins the width
#define IGNEUM_ERA_WIDTH_WORDS 1
#define IGNEUM_ERA_STRIDE_MUL 0x2b4a5b97u
#define IGNEUM_ERA_STRIDE_ROT 28
#define IGNEUM_ERA_INTERLEAVE { 1, 3, 4, 8 }
#define IGNEUM_ERA_WINDOWS "7:2:1 8:1:1 9:1:1 10:1:1 11:0:0 13:1:1 29:0:0 30:2:2 31:1:1 44:1:1 46:2:0 47:0:0 52:0:0 56:0:0 58:2:0 63:1:1"
// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h)
#define IGNEUM_DATASET_MODE 1
#define IGNEUM_SEEDW_INIT { 0x667d0fbdu, 0x7b8e5963u, 0x31c67e5eu, 0x4529ddc6u, 0xef19d6d8u, 0xaccf6211u, 0xda0aed32u, 0xabc6df31u }
#define IGNEUM_KEY_INIT { 0xceed56d7u, 0x9ba270d2u, 0x82caab2du, 0x81ebce0eu, 0x12b6ecf1u, 0xd0f3fd7cu, 0xd872eefeu, 0xc158c7bdu }
#define IGNEUM_CACHE_LOG2_WORDS 26
#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6
#define IGNEUM_CACHE_SEGMENTS 65536u
#define IGNEUM_ITEM_ROUNDS 8
#define IGNEUM_MIX_ROT_INIT { 17u, 12u, 20u, 23u, 7u, 3u, 27u, 16u }
#define IGNEUM_MIX_MUL_INIT { 0xf351d601u, 0xa3bb398fu, 0xb5a09e35u, 0x7509c9c1u, 0x6bbf31e9u, 0xfc849a79u, 0xded91851u, 0x8d9113d1u, 0x0ff15225u, 0x3a5bdd41u, 0xab533435u, 0xe1c55ad5u, 0xe6d3bd0du, 0x9d9ffbbdu, 0xbb2a3cf3u, 0x50a7c08du }
#define IGNEUM_MIX_RC_INIT { 0xc6892460u, 0x25b7228au, 0xcd515004u, 0x2846527au, 0xa6324241u, 0x36e3ec53u, 0x82961bacu, 0x0f97ba7du, 0xb6f921a9u, 0x3ada24e5u, 0xde20ab91u, 0x5378eeb2u, 0x7d161662u, 0x89353cc1u, 0xb1aa03a2u, 0x788acae6u }
#ifndef IGNEUM_NO_CUDA
// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError().
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments);
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems);
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
uint32_t nonces, uint32_t blockWarps);
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
#endif

View file

@ -0,0 +1,143 @@
{
"format": "igneum-program-pack-3",
"generator": 3,
"attempt": 0,
"program_id": "0x73bcbfe8ccf988f1",
"program_id_derivation": "FNV-1a 64 over 'igneum-program/' || generator_le32 || seed_words as little-endian bytes || attempt_le32",
"dataset_mode": "memory-hard",
"seed": "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000",
"seed_bytes": "edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07",
"seed_words": ["0x667d0fbd", "0x7b8e5963", "0x31c67e5e", "0x4529ddc6", "0xef19d6d8", "0xaccf6211", "0xda0aed32", "0xabc6df31"],
"seed_derivation": "seed_words = FNV-1a 64 over seed_bytes (attempt 0) or seed_bytes || attempt_le32 (attempt k >= 1), basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32",
"generator_rule": "version 2: exactly 16 load slots drawn first from instructions 1..63 (partial Fisher-Yates), the other 48 ops from the ten non-load weights (sum 75); a load's source is drawn from the registers other than dst written by an earlier instruction and not read by a load since; the candidate must pass the acceptance rule of spec 01 section 1.4.6 (static: no cyclically stale load source, every register has an injecting write; dynamic: 64 units on the seed-keyed closed-form dataset with no constant register bit, no lane-constant load site, under 164 saturated final values, every output bit within 136 of 1024, distinct addresses above 245760), else the next attempt of the seed is tried",
"lanes": 32,
"registers": 8,
"iterations": 8,
"instruction_count": 64,
"loads_per_hash": 128,
"program_class": "v3",
"era_seed_bytes": "df57136f2ad5f410e6145023090eea148c1342ccef5a5f541afa590be7e44745",
"load_class": "w4-era843155d7",
"load_slots": 16,
"load_mix_percent_4_16_64": [100, 0, 0],
"load_width_counts_4_16_64": [16, 0, 0],
"bytes_per_hash": 512,
"wide_load": "read-width experiment (5 October 2026, docs/plans/read-width.md), NOT the lottery hash: a load of W words (width field, 4 or 16) reads dataset[b .. b + W) with b = (src & mask) & ~(W - 1) and folds every word into dst: x = dst ^ w[0]; for j in 1..W: x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x; width 1 is the plain load; the width is drawn per instruction from the class mix with one extra below(100) draw after the nine of version 2, and the program id is FNV-1a 64 over 'igneum-program-rw/' || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots",
"era": {
"label": "843155d7",
"seed_words": ["0x843155d7", "0x8fb2bbb8", "0x3af89788", "0xc80fe6f2", "0xb265c95e", "0x001f1648", "0x236e8759", "0x6cada520"],
"draw": "docs/plans/era-layout.md 1.1: SplitMix64 seeded with seed_words[0] | seed_words[1] << 32 of seed_words_from_bytes('igneum-era/' || n_le64 || E_n); width = allowed[below(|allowed|)], stride_mul = low32(next()) | 1, stride_rot = 1 + below(31), then four next() draws for a partial Fisher-Yates over positions log2(W)..15 of which 4 - log2(W) are used",
"allowed_widths": [1],
"width_words": 1,
"stride_mul": "0x2b4a5b97",
"stride_rot": 28,
"interleave": [1, 3, 4, 8],
"address": "y = rotl(src * stride_mul, stride_rot); k = min(win, D - 26); idx = ((y & (mask >> k)) | ((off & (2^k - 1)) << (D - k))) & mask; a wide load aligns idx down to W words",
"windows": "per instruction, after the width roll: win = below(3), off = low32(next()) & (2^win - 1); used on a load slot (the instruction's win and off fields)",
"dataset_word": "dataset[w] = item(t(w))[j(w)]: j(w) gathers the bits of w at the interleave positions, t(w) is w with those bits removed",
"program_id_suffix": "'era/' || allowed[3] || width_words || stride_mul_le32 || stride_rot_le32 || interleave[4]"
},
"op_mix": {"load": 16, "add": 15, "shfl": 6, "mad": 4, "or": 4, "rotl": 4, "rotr": 4, "xor": 4, "mulhi": 3, "mul": 2, "sub": 2},
"register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]",
"splitmix32": "x ^= x>>16; x *= 0x7feb352d; x ^= x>>15; x *= 0x846ca68b; x ^= x>>16",
"iteration": "sel = r0 sampled once at the top of each iteration, then all instructions in order",
"output": "lo = r0 ^ rotl(r1,7) ^ rotl(r2,14) ^ rotl(r3,21); hi = r4 ^ rotl(r5,9) ^ rotl(r6,18) ^ rotl(r7,27); out = (hi << 32) | lo",
"op_semantics": {
"add": "dst = dst + src + (bit `bit` of sel ? imm2 : imm)",
"sub": "dst = dst - src",
"mul": "dst = dst * src (low 32)",
"mulhi": "dst = high 32 bits of dst * src",
"xor": "dst = dst ^ src",
"or": "dst = dst | src",
"rotl": "dst = rotl(dst, rot), rot in 1..31",
"rotr": "dst = rotr(dst, src & 31)",
"mad": "dst = src * src2 + dst",
"shfl": "dst = dst ^ (src of lane (lane ^ mask)), mask in {1,2,4,8,16}, within the 32-lane warp",
"load": "dst = dst ^ dataset[src & dataset.mask]",
"wload": "base = (src of lane 0 & dataset.mask) & ~31; dst = dst ^ dataset[base + lane] (warp-coalesced 128-byte load, lever b, only when --wide-frac > 0)"
},
"dataset": {
"log2_words": 28,
"bytes": 1073741824,
"mask": "0x0fffffff",
"day": "bytes:69676e65756d2d6461792ffa50000000000000",
"day_bytes": "69676e65756d2d6461792ffa50000000000000",
"day_words_from": "seed_words_from_bytes(day_bytes)",
"d0": "0xceed56d7",
"d1": "0x9ba270d2",
"mode": "memory-hard",
"spec": "proto-metal/MEMHARD.md",
"key": ["0xceed56d7", "0x9ba270d2", "0x82caab2d", "0x81ebce0e", "0x12b6ecf1", "0xd0f3fd7c", "0xd872eefe", "0xc158c7bd"],
"key_derivation": "the 8 words of seed_words_from_bytes(day_bytes); d0, d1 are key[0], key[1]",
"cache": {"log2_words": 26, "bytes": 268435456, "line_words": 16, "segment_lines": 64, "segments": 65536, "block": "ChaCha12 core + feed-forward, rotations 16 12 8 7", "sigma": ["0x61707865", "0x3320646e", "0x79622d32", "0x6b206574"], "tag": ["0x49676e65", "0x756d4d48"], "chain": "in_j = prev_line ^ (sigma[0..3] || key[0..7] || seg || j || tag[0..1]); line_j = block(in_j); prev_0 = 0"},
"mixer": {"draw": "SplitMix64 seeded with key[0] | key[1] << 32: rot[0..7] = 1 + next() % 31, mul[0..15] = low32(next()) | 1, rc[0..15] = low32(next())", "rot": [17, 12, 20, 23, 7, 3, 27, 16], "mul": ["0xf351d601", "0xa3bb398f", "0xb5a09e35", "0x7509c9c1", "0x6bbf31e9", "0xfc849a79", "0xded91851", "0x8d9113d1", "0x0ff15225", "0x3a5bdd41", "0xab533435", "0xe1c55ad5", "0xe6d3bd0d", "0x9d9ffbbd", "0xbb2a3cf3", "0x50a7c08d"], "rc": ["0xc6892460", "0x25b7228a", "0xcd515004", "0x2846527a", "0xa6324241", "0x36e3ec53", "0x82961bac", "0x0f97ba7d", "0xb6f921a9", "0x3ada24e5", "0xde20ab91", "0x5378eeb2", "0x7d161662", "0x89353cc1", "0xb1aa03a2", "0x788acae6"], "round": "for i in 0..15: s[i] = (s[i] ^ (rc[i] + (r+1) * 0x9E3779B9)) * mul[i]; then quarter rounds on columns (0,4,8,12) (1,5,9,13) (2,6,10,14) (3,7,11,15) with rot[0..3] and diagonals (0,5,10,15) (1,6,11,12) (2,7,8,13) (3,4,9,14) with rot[4..7]", "quarter_round": "a += b; d ^= a; d = rotl(d, r1); c += d; b ^= c; b = rotl(b, r2); a += b; d ^= a; d = rotl(d, r3); c += d; b ^= c; b = rotl(b, r4)"},
"item": "s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..7: s = M_r(s); line = s[0] & 0x003fffff; s[i] ^= cache[line * 16 + i]; then s = M_8(s); item(t) = s",
"word": "dataset[w] = item(w >> 4)[w & 15]"
},
"instructions": [
{"i": 0, "op": "add", "dst": 4, "src": 5, "src2": 7, "imm": "0xea86e152", "imm2": "0x5810667a", "rot": 27, "bit": 13, "mask": 2, "width": 1, "win": 0, "off": 0},
{"i": 1, "op": "xor", "dst": 7, "src": 0, "src2": 6, "imm": "0xbaab6229", "imm2": "0xed861989", "rot": 26, "bit": 22, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 2, "op": "shfl", "dst": 3, "src": 6, "src2": 2, "imm": "0x5b623116", "imm2": "0xff12e5b2", "rot": 12, "bit": 24, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 3, "op": "sub", "dst": 4, "src": 1, "src2": 1, "imm": "0xe99741c7", "imm2": "0xf5fa5009", "rot": 1, "bit": 21, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 4, "op": "mad", "dst": 2, "src": 0, "src2": 4, "imm": "0x673c2157", "imm2": "0xee02465f", "rot": 20, "bit": 22, "mask": 2, "width": 1, "win": 0, "off": 0},
{"i": 5, "op": "rotl", "dst": 4, "src": 7, "src2": 7, "imm": "0x946f7818", "imm2": "0x45d3399e", "rot": 9, "bit": 2, "mask": 16, "width": 1, "win": 0, "off": 0},
{"i": 6, "op": "rotr", "dst": 0, "src": 2, "src2": 4, "imm": "0x5f6a0ed2", "imm2": "0x7043a636", "rot": 19, "bit": 31, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 7, "op": "load", "dst": 6, "src": 7, "src2": 1, "imm": "0x5c61dcf7", "imm2": "0x7466aa40", "rot": 19, "bit": 9, "mask": 2, "width": 1, "win": 2, "off": 1},
{"i": 8, "op": "load", "dst": 1, "src": 4, "src2": 5, "imm": "0x85668475", "imm2": "0xdb8cc483", "rot": 29, "bit": 7, "mask": 4, "width": 1, "win": 1, "off": 1},
{"i": 9, "op": "load", "dst": 1, "src": 2, "src2": 4, "imm": "0x4454980f", "imm2": "0xebd31581", "rot": 10, "bit": 28, "mask": 16, "width": 1, "win": 1, "off": 1},
{"i": 10, "op": "load", "dst": 7, "src": 0, "src2": 2, "imm": "0xe075297c", "imm2": "0x5779c44c", "rot": 10, "bit": 22, "mask": 2, "width": 1, "win": 1, "off": 1},
{"i": 11, "op": "load", "dst": 7, "src": 1, "src2": 6, "imm": "0x65aa4311", "imm2": "0x4fe48ea9", "rot": 15, "bit": 9, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 12, "op": "mul", "dst": 5, "src": 4, "src2": 2, "imm": "0x1383d3ad", "imm2": "0xf3094b29", "rot": 8, "bit": 9, "mask": 2, "width": 1, "win": 0, "off": 0},
{"i": 13, "op": "load", "dst": 7, "src": 6, "src2": 2, "imm": "0xed8a496f", "imm2": "0x3072c3c6", "rot": 28, "bit": 19, "mask": 8, "width": 1, "win": 1, "off": 1},
{"i": 14, "op": "shfl", "dst": 6, "src": 5, "src2": 0, "imm": "0x8b965b57", "imm2": "0xcfeca6c1", "rot": 12, "bit": 27, "mask": 16, "width": 1, "win": 0, "off": 0},
{"i": 15, "op": "shfl", "dst": 3, "src": 5, "src2": 3, "imm": "0x877c7586", "imm2": "0xa9cb2a03", "rot": 2, "bit": 29, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 16, "op": "or", "dst": 2, "src": 7, "src2": 3, "imm": "0xb740221a", "imm2": "0x89d38d6d", "rot": 6, "bit": 31, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 17, "op": "rotr", "dst": 0, "src": 6, "src2": 1, "imm": "0x26f3ad8a", "imm2": "0x27256f15", "rot": 18, "bit": 5, "mask": 2, "width": 1, "win": 0, "off": 0},
{"i": 18, "op": "add", "dst": 7, "src": 5, "src2": 7, "imm": "0xf66e7017", "imm2": "0xb9e3577e", "rot": 9, "bit": 12, "mask": 16, "width": 1, "win": 0, "off": 0},
{"i": 19, "op": "rotl", "dst": 7, "src": 0, "src2": 5, "imm": "0x849ae6ee", "imm2": "0x02b358f9", "rot": 24, "bit": 18, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 20, "op": "add", "dst": 6, "src": 7, "src2": 6, "imm": "0x52334d12", "imm2": "0x8c9f0ef8", "rot": 11, "bit": 23, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 21, "op": "or", "dst": 2, "src": 6, "src2": 5, "imm": "0xb1871e63", "imm2": "0xb2e40191", "rot": 5, "bit": 13, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 22, "op": "mad", "dst": 6, "src": 5, "src2": 4, "imm": "0x97df29e4", "imm2": "0xe60fea84", "rot": 11, "bit": 16, "mask": 2, "width": 1, "win": 0, "off": 0},
{"i": 23, "op": "or", "dst": 1, "src": 0, "src2": 3, "imm": "0x8f30d21d", "imm2": "0x2df685a0", "rot": 31, "bit": 0, "mask": 16, "width": 1, "win": 0, "off": 0},
{"i": 24, "op": "xor", "dst": 6, "src": 1, "src2": 2, "imm": "0xd2c4025f", "imm2": "0x5269eb4d", "rot": 31, "bit": 21, "mask": 16, "width": 1, "win": 0, "off": 0},
{"i": 25, "op": "add", "dst": 2, "src": 6, "src2": 5, "imm": "0x659fc3d3", "imm2": "0x9cec0e12", "rot": 6, "bit": 17, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 26, "op": "rotr", "dst": 7, "src": 0, "src2": 1, "imm": "0xd89ef484", "imm2": "0x20be3846", "rot": 12, "bit": 9, "mask": 16, "width": 1, "win": 0, "off": 0},
{"i": 27, "op": "mad", "dst": 4, "src": 5, "src2": 7, "imm": "0xb48420ae", "imm2": "0x3d1f2485", "rot": 14, "bit": 28, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 28, "op": "mad", "dst": 3, "src": 2, "src2": 4, "imm": "0xc5c46d76", "imm2": "0x700044b5", "rot": 22, "bit": 10, "mask": 2, "width": 1, "win": 0, "off": 0},
{"i": 29, "op": "load", "dst": 1, "src": 4, "src2": 3, "imm": "0x0fbaf177", "imm2": "0xfff4f2ed", "rot": 20, "bit": 31, "mask": 2, "width": 1, "win": 0, "off": 0},
{"i": 30, "op": "load", "dst": 2, "src": 3, "src2": 0, "imm": "0xe90eb2e5", "imm2": "0xb0f9eb79", "rot": 27, "bit": 14, "mask": 4, "width": 1, "win": 2, "off": 2},
{"i": 31, "op": "load", "dst": 1, "src": 5, "src2": 2, "imm": "0x97ba3fc3", "imm2": "0x7894e657", "rot": 3, "bit": 30, "mask": 16, "width": 1, "win": 1, "off": 1},
{"i": 32, "op": "add", "dst": 7, "src": 2, "src2": 7, "imm": "0x070888a8", "imm2": "0xe403240e", "rot": 2, "bit": 16, "mask": 2, "width": 1, "win": 0, "off": 0},
{"i": 33, "op": "add", "dst": 2, "src": 0, "src2": 5, "imm": "0xf2e46d55", "imm2": "0x29701828", "rot": 31, "bit": 28, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 34, "op": "add", "dst": 2, "src": 3, "src2": 0, "imm": "0x58f75b87", "imm2": "0x343b7aee", "rot": 12, "bit": 14, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 35, "op": "mulhi", "dst": 2, "src": 5, "src2": 6, "imm": "0x5892a9e6", "imm2": "0xc9824c94", "rot": 19, "bit": 26, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 36, "op": "xor", "dst": 4, "src": 2, "src2": 3, "imm": "0x7b5b5474", "imm2": "0x45cfc5dd", "rot": 17, "bit": 18, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 37, "op": "mul", "dst": 6, "src": 5, "src2": 7, "imm": "0xccf564a5", "imm2": "0x873ad101", "rot": 7, "bit": 11, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 38, "op": "xor", "dst": 7, "src": 0, "src2": 6, "imm": "0xcac8f06d", "imm2": "0x6b97c683", "rot": 18, "bit": 28, "mask": 2, "width": 1, "win": 0, "off": 0},
{"i": 39, "op": "add", "dst": 7, "src": 2, "src2": 4, "imm": "0xb8180e9d", "imm2": "0x32bbd117", "rot": 23, "bit": 19, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 40, "op": "rotr", "dst": 2, "src": 3, "src2": 0, "imm": "0x2d6070bc", "imm2": "0x68ff101e", "rot": 13, "bit": 18, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 41, "op": "sub", "dst": 7, "src": 0, "src2": 2, "imm": "0x0daf96ea", "imm2": "0x36f37be1", "rot": 5, "bit": 0, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 42, "op": "add", "dst": 4, "src": 3, "src2": 6, "imm": "0x6ced15b7", "imm2": "0x6df7aed4", "rot": 19, "bit": 4, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 43, "op": "shfl", "dst": 7, "src": 3, "src2": 5, "imm": "0x8ace05f3", "imm2": "0xd378ec12", "rot": 23, "bit": 10, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 44, "op": "load", "dst": 0, "src": 7, "src2": 4, "imm": "0xb0607786", "imm2": "0xc4acabbc", "rot": 13, "bit": 7, "mask": 16, "width": 1, "win": 1, "off": 1},
{"i": 45, "op": "add", "dst": 1, "src": 6, "src2": 5, "imm": "0xe09f54e9", "imm2": "0x83e825bf", "rot": 23, "bit": 14, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 46, "op": "load", "dst": 3, "src": 1, "src2": 0, "imm": "0x63cc1e4e", "imm2": "0xa1be8118", "rot": 12, "bit": 6, "mask": 2, "width": 1, "win": 2, "off": 0},
{"i": 47, "op": "load", "dst": 6, "src": 3, "src2": 7, "imm": "0x353f1d79", "imm2": "0x3b2e7456", "rot": 18, "bit": 18, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 48, "op": "mulhi", "dst": 4, "src": 2, "src2": 7, "imm": "0x00d8a3cd", "imm2": "0x231866d2", "rot": 21, "bit": 20, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 49, "op": "add", "dst": 5, "src": 0, "src2": 2, "imm": "0xa8bae6df", "imm2": "0xf572bdb9", "rot": 14, "bit": 7, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 50, "op": "shfl", "dst": 0, "src": 7, "src2": 7, "imm": "0x81ef22e1", "imm2": "0x74438fc5", "rot": 28, "bit": 18, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 51, "op": "add", "dst": 6, "src": 0, "src2": 6, "imm": "0x383b9260", "imm2": "0x11e17c61", "rot": 12, "bit": 19, "mask": 16, "width": 1, "win": 0, "off": 0},
{"i": 52, "op": "load", "dst": 5, "src": 2, "src2": 2, "imm": "0xfb84f451", "imm2": "0x11cd863e", "rot": 21, "bit": 20, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 53, "op": "rotl", "dst": 6, "src": 5, "src2": 4, "imm": "0xb1a7db6b", "imm2": "0x76686b9b", "rot": 12, "bit": 4, "mask": 16, "width": 1, "win": 0, "off": 0},
{"i": 54, "op": "rotl", "dst": 3, "src": 6, "src2": 3, "imm": "0x6f981f52", "imm2": "0xd99aeba2", "rot": 12, "bit": 27, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 55, "op": "add", "dst": 2, "src": 1, "src2": 2, "imm": "0xac6be8e3", "imm2": "0x18d67dbb", "rot": 26, "bit": 10, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 56, "op": "load", "dst": 1, "src": 4, "src2": 0, "imm": "0x7e7f6a00", "imm2": "0x6f0747da", "rot": 25, "bit": 28, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 57, "op": "add", "dst": 5, "src": 0, "src2": 4, "imm": "0xf03673fe", "imm2": "0xa75cd60d", "rot": 16, "bit": 12, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 58, "op": "load", "dst": 5, "src": 0, "src2": 2, "imm": "0x227f94a6", "imm2": "0x0e8344f9", "rot": 20, "bit": 10, "mask": 2, "width": 1, "win": 2, "off": 0},
{"i": 59, "op": "add", "dst": 1, "src": 4, "src2": 3, "imm": "0xdecd4794", "imm2": "0x8dfb96bb", "rot": 21, "bit": 7, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 60, "op": "mulhi", "dst": 3, "src": 2, "src2": 2, "imm": "0x0dd268e0", "imm2": "0x53034ca9", "rot": 1, "bit": 8, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 61, "op": "or", "dst": 6, "src": 4, "src2": 7, "imm": "0x3a45a321", "imm2": "0x9bc59a5f", "rot": 25, "bit": 11, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 62, "op": "shfl", "dst": 5, "src": 4, "src2": 2, "imm": "0x8f229cc1", "imm2": "0xcaac64a2", "rot": 17, "bit": 13, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 63, "op": "load", "dst": 3, "src": 6, "src2": 6, "imm": "0xb2574178", "imm2": "0xcbcc798d", "rot": 28, "bit": 0, "mask": 16, "width": 1, "win": 1, "off": 1}
]
}

View file

@ -0,0 +1,109 @@
#include <metal_stdlib>
using namespace metal;
#define MASK 0x0fffffffu
constant uint SEEDW[8] = { 0x667d0fbdu, 0x7b8e5963u, 0x31c67e5eu, 0x4529ddc6u, 0xef19d6d8u, 0xaccf6211u, 0xda0aed32u, 0xabc6df31u };
inline uint splitmix32(uint x) {
x ^= x >> 16; x *= 0x7feb352du;
x ^= x >> 15; x *= 0x846ca68bu;
x ^= x >> 16;
return x;
}
inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31
inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
inline uint ds_elem(uint i, uint d0, uint d1) {
uint x = i ^ d0;
x *= 0x9E3779B1u; x ^= x >> 15;
x += d1;
x *= 0x85EBCA77u; x ^= x >> 13;
x *= 0xC2B2AE3Du; x ^= x >> 16;
return x;
}
kernel void igneum_hash(device const uint* dataset [[buffer(0)]],
device ulong* out [[buffer(1)]],
constant uint& baseNonce [[buffer(2)]],
uint gid [[thread_position_in_grid]]) {
uint nonce = baseNonce + gid;
uint r0, r1, r2, r3, r4, r5, r6, r7;
{ uint x = nonce ^ SEEDW[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ SEEDW[1]; }
{ uint x = nonce ^ SEEDW[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ SEEDW[2]; }
{ uint x = nonce ^ SEEDW[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ SEEDW[3]; }
{ uint x = nonce ^ SEEDW[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ SEEDW[4]; }
{ uint x = nonce ^ SEEDW[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ SEEDW[5]; }
{ uint x = nonce ^ SEEDW[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ SEEDW[6]; }
{ uint x = nonce ^ SEEDW[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ SEEDW[7]; }
{ uint x = nonce ^ SEEDW[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ SEEDW[0]; }
for (uint it = 0u; it < 8u; ++it) {
uint sel = r0;
r4 = r4 + r5 + select(0xea86e152u, 0x5810667au, ((sel >> 13u) & 1u) != 0u); // 0
r7 = r7 ^ r0; // 1
r3 = r3 ^ simd_shuffle_xor(r6, (ushort)1); // 2
r4 = r4 - r1; // 3
r2 = r0 * r4 + r2; // 4
r4 = rotl_imm(r4, 9u); // 5
r0 = rotr_var(r0, r2); // 6
r6 = r6 ^ dataset[((rotl_imm(r7 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x04000000u) & MASK]; // 7
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 8
r1 = r1 ^ dataset[((rotl_imm(r2 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 9
r7 = r7 ^ dataset[((rotl_imm(r0 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 10
r7 = r7 ^ dataset[((rotl_imm(r1 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 11
r5 = r5 * r4; // 12
r7 = r7 ^ dataset[((rotl_imm(r6 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 13
r6 = r6 ^ simd_shuffle_xor(r5, (ushort)16); // 14
r3 = r3 ^ simd_shuffle_xor(r5, (ushort)1); // 15
r2 = r2 | r7; // 16
r0 = rotr_var(r0, r6); // 17
r7 = r7 + r5 + select(0xf66e7017u, 0xb9e3577eu, ((sel >> 12u) & 1u) != 0u); // 18
r7 = rotl_imm(r7, 24u); // 19
r6 = r6 + r7 + select(0x52334d12u, 0x8c9f0ef8u, ((sel >> 23u) & 1u) != 0u); // 20
r2 = r2 | r6; // 21
r6 = r5 * r4 + r6; // 22
r1 = r1 | r0; // 23
r6 = r6 ^ r1; // 24
r2 = r2 + r6 + select(0x659fc3d3u, 0x9cec0e12u, ((sel >> 17u) & 1u) != 0u); // 25
r7 = rotr_var(r7, r0); // 26
r4 = r5 * r7 + r4; // 27
r3 = r2 * r4 + r3; // 28
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 29
r2 = r2 ^ dataset[((rotl_imm(r3 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x08000000u) & MASK]; // 30
r1 = r1 ^ dataset[((rotl_imm(r5 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 31
r7 = r7 + r2 + select(0x070888a8u, 0xe403240eu, ((sel >> 16u) & 1u) != 0u); // 32
r2 = r2 + r0 + select(0xf2e46d55u, 0x29701828u, ((sel >> 28u) & 1u) != 0u); // 33
r2 = r2 + r3 + select(0x58f75b87u, 0x343b7aeeu, ((sel >> 14u) & 1u) != 0u); // 34
r2 = mulhi(r2, r5); // 35
r4 = r4 ^ r2; // 36
r6 = r6 * r5; // 37
r7 = r7 ^ r0; // 38
r7 = r7 + r2 + select(0xb8180e9du, 0x32bbd117u, ((sel >> 19u) & 1u) != 0u); // 39
r2 = rotr_var(r2, r3); // 40
r7 = r7 - r0; // 41
r4 = r4 + r3 + select(0x6ced15b7u, 0x6df7aed4u, ((sel >> 4u) & 1u) != 0u); // 42
r7 = r7 ^ simd_shuffle_xor(r3, (ushort)4); // 43
r0 = r0 ^ dataset[((rotl_imm(r7 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 44
r1 = r1 + r6 + select(0xe09f54e9u, 0x83e825bfu, ((sel >> 14u) & 1u) != 0u); // 45
r3 = r3 ^ dataset[((rotl_imm(r1 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 46
r6 = r6 ^ dataset[((rotl_imm(r3 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 47
r4 = mulhi(r4, r2); // 48
r5 = r5 + r0 + select(0xa8bae6dfu, 0xf572bdb9u, ((sel >> 7u) & 1u) != 0u); // 49
r0 = r0 ^ simd_shuffle_xor(r7, (ushort)4); // 50
r6 = r6 + r0 + select(0x383b9260u, 0x11e17c61u, ((sel >> 19u) & 1u) != 0u); // 51
r5 = r5 ^ dataset[((rotl_imm(r2 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 52
r6 = rotl_imm(r6, 12u); // 53
r3 = rotl_imm(r3, 12u); // 54
r2 = r2 + r1 + select(0xac6be8e3u, 0x18d67dbbu, ((sel >> 10u) & 1u) != 0u); // 55
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 56
r5 = r5 + r0 + select(0xf03673feu, 0xa75cd60du, ((sel >> 12u) & 1u) != 0u); // 57
r5 = r5 ^ dataset[((rotl_imm(r0 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 58
r1 = r1 + r4 + select(0xdecd4794u, 0x8dfb96bbu, ((sel >> 7u) & 1u) != 0u); // 59
r3 = mulhi(r3, r2); // 60
r6 = r6 | r4; // 61
r5 = r5 ^ simd_shuffle_xor(r4, (ushort)8); // 62
r3 = r3 ^ dataset[((rotl_imm(r6 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 63
}
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((ulong)hi << 32) | (ulong)lo;
}

View file

@ -0,0 +1,111 @@
#include <metal_stdlib>
using namespace metal;
#define MASK 0x0fffffffu
constant uint SEEDW[8] = { 0x667d0fbdu, 0x7b8e5963u, 0x31c67e5eu, 0x4529ddc6u, 0xef19d6d8u, 0xaccf6211u, 0xda0aed32u, 0xabc6df31u };
inline uint splitmix32(uint x) {
x ^= x >> 16; x *= 0x7feb352du;
x ^= x >> 15; x *= 0x846ca68bu;
x ^= x >> 16;
return x;
}
inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31
inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
inline uint ds_elem(uint i, uint d0, uint d1) {
uint x = i ^ d0;
x *= 0x9E3779B1u; x ^= x >> 15;
x += d1;
x *= 0x85EBCA77u; x ^= x >> 13;
x *= 0xC2B2AE3Du; x ^= x >> 16;
return x;
}
// Header-bound variant: the init words come from buffer 3 (bind.rs), not from SEEDW.
kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]],
device ulong* out [[buffer(1)]],
constant uint& baseNonce [[buffer(2)]],
constant uint* initw [[buffer(3)]],
uint gid [[thread_position_in_grid]]) {
uint nonce = baseNonce + gid;
uint r0, r1, r2, r3, r4, r5, r6, r7;
{ uint x = nonce ^ initw[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ initw[1]; }
{ uint x = nonce ^ initw[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ initw[2]; }
{ uint x = nonce ^ initw[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ initw[3]; }
{ uint x = nonce ^ initw[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ initw[4]; }
{ uint x = nonce ^ initw[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ initw[5]; }
{ uint x = nonce ^ initw[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ initw[6]; }
{ uint x = nonce ^ initw[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ initw[7]; }
{ uint x = nonce ^ initw[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ initw[0]; }
for (uint it = 0u; it < 8u; ++it) {
uint sel = r0;
r4 = r4 + r5 + select(0xea86e152u, 0x5810667au, ((sel >> 13u) & 1u) != 0u); // 0
r7 = r7 ^ r0; // 1
r3 = r3 ^ simd_shuffle_xor(r6, (ushort)1); // 2
r4 = r4 - r1; // 3
r2 = r0 * r4 + r2; // 4
r4 = rotl_imm(r4, 9u); // 5
r0 = rotr_var(r0, r2); // 6
r6 = r6 ^ dataset[((rotl_imm(r7 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x04000000u) & MASK]; // 7
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 8
r1 = r1 ^ dataset[((rotl_imm(r2 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 9
r7 = r7 ^ dataset[((rotl_imm(r0 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 10
r7 = r7 ^ dataset[((rotl_imm(r1 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 11
r5 = r5 * r4; // 12
r7 = r7 ^ dataset[((rotl_imm(r6 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 13
r6 = r6 ^ simd_shuffle_xor(r5, (ushort)16); // 14
r3 = r3 ^ simd_shuffle_xor(r5, (ushort)1); // 15
r2 = r2 | r7; // 16
r0 = rotr_var(r0, r6); // 17
r7 = r7 + r5 + select(0xf66e7017u, 0xb9e3577eu, ((sel >> 12u) & 1u) != 0u); // 18
r7 = rotl_imm(r7, 24u); // 19
r6 = r6 + r7 + select(0x52334d12u, 0x8c9f0ef8u, ((sel >> 23u) & 1u) != 0u); // 20
r2 = r2 | r6; // 21
r6 = r5 * r4 + r6; // 22
r1 = r1 | r0; // 23
r6 = r6 ^ r1; // 24
r2 = r2 + r6 + select(0x659fc3d3u, 0x9cec0e12u, ((sel >> 17u) & 1u) != 0u); // 25
r7 = rotr_var(r7, r0); // 26
r4 = r5 * r7 + r4; // 27
r3 = r2 * r4 + r3; // 28
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 29
r2 = r2 ^ dataset[((rotl_imm(r3 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x08000000u) & MASK]; // 30
r1 = r1 ^ dataset[((rotl_imm(r5 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 31
r7 = r7 + r2 + select(0x070888a8u, 0xe403240eu, ((sel >> 16u) & 1u) != 0u); // 32
r2 = r2 + r0 + select(0xf2e46d55u, 0x29701828u, ((sel >> 28u) & 1u) != 0u); // 33
r2 = r2 + r3 + select(0x58f75b87u, 0x343b7aeeu, ((sel >> 14u) & 1u) != 0u); // 34
r2 = mulhi(r2, r5); // 35
r4 = r4 ^ r2; // 36
r6 = r6 * r5; // 37
r7 = r7 ^ r0; // 38
r7 = r7 + r2 + select(0xb8180e9du, 0x32bbd117u, ((sel >> 19u) & 1u) != 0u); // 39
r2 = rotr_var(r2, r3); // 40
r7 = r7 - r0; // 41
r4 = r4 + r3 + select(0x6ced15b7u, 0x6df7aed4u, ((sel >> 4u) & 1u) != 0u); // 42
r7 = r7 ^ simd_shuffle_xor(r3, (ushort)4); // 43
r0 = r0 ^ dataset[((rotl_imm(r7 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 44
r1 = r1 + r6 + select(0xe09f54e9u, 0x83e825bfu, ((sel >> 14u) & 1u) != 0u); // 45
r3 = r3 ^ dataset[((rotl_imm(r1 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 46
r6 = r6 ^ dataset[((rotl_imm(r3 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 47
r4 = mulhi(r4, r2); // 48
r5 = r5 + r0 + select(0xa8bae6dfu, 0xf572bdb9u, ((sel >> 7u) & 1u) != 0u); // 49
r0 = r0 ^ simd_shuffle_xor(r7, (ushort)4); // 50
r6 = r6 + r0 + select(0x383b9260u, 0x11e17c61u, ((sel >> 19u) & 1u) != 0u); // 51
r5 = r5 ^ dataset[((rotl_imm(r2 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 52
r6 = rotl_imm(r6, 12u); // 53
r3 = rotl_imm(r3, 12u); // 54
r2 = r2 + r1 + select(0xac6be8e3u, 0x18d67dbbu, ((sel >> 10u) & 1u) != 0u); // 55
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 56
r5 = r5 + r0 + select(0xf03673feu, 0xa75cd60du, ((sel >> 12u) & 1u) != 0u); // 57
r5 = r5 ^ dataset[((rotl_imm(r0 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 58
r1 = r1 + r4 + select(0xdecd4794u, 0x8dfb96bbu, ((sel >> 7u) & 1u) != 0u); // 59
r3 = mulhi(r3, r2); // 60
r6 = r6 | r4; // 61
r5 = r5 ^ simd_shuffle_xor(r4, (ushort)8); // 62
r3 = r3 ^ dataset[((rotl_imm(r6 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 63
}
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((ulong)hi << 32) | (ulong)lo;
}

View file

@ -0,0 +1,5 @@
epoch_seed_hex edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07
day_seed_hex 69676e65756d2d6461792ffa50000000000000
epoch_index 0
day_index 20730
daa_score 0

View file

@ -0,0 +1,57 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v3, memory-hard dataset
#pragma once
#ifdef __cplusplus
#include <cstdint>
#else
#include <stdint.h>
#endif
#define IGNEUM_VEC_WARPS 3
static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u };
static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = {
{ // base nonce 0
0x2d96e3f8e9ad42e2ull, 0x0d36523b17ba1058ull, 0x899e2a6df6991c3dull, 0xdb382aa991ff18d6ull, 0xc896fe7b624c7722ull, 0x932cfc8892d24216ull, 0xd9ef4047f797129bull, 0x63fb054ee8cfe1d2ull,
0xbcbd287cfd7e12a9ull, 0x68e41482f26203c2ull, 0x5c368b15e9ec2bf5ull, 0x83484f0946bfda39ull, 0xc82c1aeadba75e3cull, 0x023cbca13fae681dull, 0x000fcab814a56dd9ull, 0x432082a86895e2a7ull,
0x635fb930f3cf454bull, 0x1104c1b855a17655ull, 0x2fd53820b88956b2ull, 0xc415d67f6420191dull, 0xeaf6853e1c859f5dull, 0x95dfdf66b21e0518ull, 0x7d67b2b734243bfeull, 0xb142ae8e4a4b5ca2ull,
0x71ec0aff12d4c515ull, 0x0b2207c2c51fc2faull, 0xbc3957dbbf51d260ull, 0x74ea737eebe8a3e4ull, 0xd49fd6ea590200e9ull, 0xd5096fd4e83b8fa1ull, 0x999b099025a53907ull, 0xa5692c50b79eac9dull
},
{ // base nonce 4096
0x9d4b6aaa8c292525ull, 0xb2ad4ab3941b3a3full, 0x17988da38edbf912ull, 0x566b31d7cceae84aull, 0x132f77e29f4e6d40ull, 0xbdc098ecac64b3dcull, 0x1670a7bfbc675d88ull, 0xda3a9960471e44bbull,
0x887e9fda5ec82271ull, 0x6137f598a67fc0d4ull, 0xd77b98d7d2025d26ull, 0x96b472c7560a777dull, 0x0c98d0ff42f589f0ull, 0x4f0f15c9e4eefef2ull, 0x6bae44572933ab20ull, 0x2772c0266bdd5308ull,
0x2fdb9e92d13d07baull, 0x6fe9239589ed122aull, 0x7e3955e5a91916fdull, 0x966d21685073c9b8ull, 0xa050194eb84c4104ull, 0xf7e095dae1d70633ull, 0x4f49616da323736bull, 0x1fcc532045f01cd4ull,
0x69549d64913199fcull, 0x6684a1b257c011b0ull, 0xff318d60e8970b9aull, 0x1429c63759eb426dull, 0x3f470be3f4d4815aull, 0x6e1d6c6b91b984e2ull, 0xda2dd7f31fea85efull, 0xb3944abe38dae250ull
},
{ // base nonce 1000000
0xd549a905f9121657ull, 0xd863810b66483e86ull, 0x65d735cf2e5447afull, 0x9cda91c71e0f5790ull, 0x682c4a1be66d2444ull, 0xcb64852ca144ffe1ull, 0x8e72112d5db27544ull, 0x1a9e6bb74ea3e563ull,
0x63c279f228faf1bcull, 0xf7b77f8435da26a2ull, 0x3c28bd5b463df46full, 0xd2cdd0e9e968d2bdull, 0x46eae240fc0c3f01ull, 0xbdd4c4e525220d3cull, 0x345178054b1eac9dull, 0x312f541e08ab8fe9ull,
0x0cbc6f22e65887adull, 0xf9a0d018f9ce95cfull, 0x994d29e263550764ull, 0x27a37ad3487c4018ull, 0xcdb7645aa05dac6aull, 0x8578f1e945156d7full, 0x8e135ab725d23599ull, 0xfa902aebc88a620cull,
0x5bdb23834f693ba8ull, 0xae53782ddf331358ull, 0x2304fcfaad2616fcull, 0x21e9216f6c0b14a6ull, 0xa88e90dc82b85f66ull, 0xba5aa73b369dd894ull, 0x97e37652c487cc5eull, 0x36cb34ed0c8ce81full
}
};
// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455).
static const uint32_t IGNEUM_DS_HEAD[16] = {
0x3dd50b1fu, 0x93123701u, 0x48edec90u, 0x3a7d2407u, 0xcdd88b45u, 0x2c37c6b0u, 0xb03a0ac0u, 0x10e2c6a5u,
0x5119540eu, 0x3358ef2du, 0x657ac748u, 0x09e2105cu, 0x25e2929cu, 0x5fa0cdccu, 0x4317cabdu, 0x6178ee52u
};
static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u;
static const uint32_t IGNEUM_DS_LAST = 0xf7b7180eu;
// 64 sampled dataset words (index, value) computed on the Mac.
#define IGNEUM_DS_SAMPLES 64
static const uint32_t IGNEUM_DS_SAMPLE_INDEX[IGNEUM_DS_SAMPLES] = {
59471966u, 217795994u, 208353206u, 42483309u, 172547758u, 148076330u, 183853158u, 214389424u, 267488061u, 169781097u, 184093494u, 153880993u, 84977930u, 46426879u, 3093825u, 225364072u, 44593546u, 260713159u, 168250303u, 52384140u, 223401610u, 45554030u, 95410555u, 175039924u, 79171087u, 267580473u, 24168642u, 37981670u, 171551130u, 195559979u, 204611762u, 140997658u, 138925853u, 86637313u, 20736778u, 219665210u, 160430336u, 264654675u, 8013395u, 228945585u, 213884386u, 104419827u, 44185464u, 142737231u, 99284897u, 132475900u, 61861762u, 132056166u, 262388043u, 91878046u, 117353561u, 124768597u, 71352993u, 190698941u, 46055428u, 55281366u, 165145231u, 106810753u, 171985651u, 232085256u, 159510492u, 40072060u, 209107596u, 39023794u
};
static const uint32_t IGNEUM_DS_SAMPLE_VALUE[IGNEUM_DS_SAMPLES] = {
0x88717593u, 0x79f4c2ebu, 0xa55c10fau, 0x1e48deedu, 0x825505dfu, 0x056e4b2au, 0x51c5cdc5u, 0x9133ff90u, 0xf3e1f630u, 0x8c0f4bcau, 0x13c805d6u, 0x7cd9cc39u, 0x54d3ad95u, 0x3142a64bu, 0xf690ebf9u, 0x5f05ecf5u, 0x708fd94du, 0x57e3a213u, 0xb60de9a6u, 0xc3cacc1fu, 0x72a0a173u, 0x2a13f6c7u, 0xfb5b4620u, 0xeae5f6c3u, 0x38e0ceb1u, 0x103d3fa4u, 0xcb03f2aeu, 0x562969ccu, 0xf1349e23u, 0x3fa57b5du, 0x4fb463b9u, 0xb2b70250u, 0x7f689f8cu, 0x27e5f93eu, 0xbaeafce2u, 0xaf3d3776u, 0x084c3724u, 0xffe6a16cu, 0xd4d0d18du, 0x5d78da4eu, 0xf0bd37c6u, 0x261f1c0du, 0x7bfc0ff6u, 0x08e5373bu, 0xd23a038cu, 0x124d74c0u, 0xdd74caceu, 0x5ab7c9bau, 0xc05d2942u, 0x2d1e2f73u, 0x52a27429u, 0x0c828d10u, 0x7514e589u, 0xdb8ca7d2u, 0xdb881ba6u, 0x0ce95e1du, 0x3c93a15eu, 0x73ec90f8u, 0xe47554edu, 0x496959f3u, 0xebf27832u, 0x11f5885eu, 0x69a1ff20u, 0x804542bdu
};
// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^26 words.
static const uint32_t IGNEUM_CACHE_HEAD[16] = {
0xebd9055cu, 0x7eaf6b21u, 0x9b610855u, 0x134a8562u, 0xb10059bau, 0x7f459e58u, 0x8f3c7873u, 0x3523e609u,
0x75b8f4cau, 0x750d31b4u, 0x0dae0782u, 0x26f10015u, 0x7ba87a41u, 0x0efd8543u, 0x691d6368u, 0x8929d967u
};
static const uint32_t IGNEUM_CACHE_LAST[16] = {
0x51474edfu, 0xc3cfba93u, 0xf21454cfu, 0x9b79baacu, 0x4d4fce23u, 0xd701dfd5u, 0x37357ab3u, 0x1be693fau,
0xa701cb3bu, 0x7467c620u, 0x428184e8u, 0xf4010df0u, 0xe33aa88fu, 0xc78d6d62u, 0xae9ba9c0u, 0xf4bba97eu
};
static const uint64_t IGNEUM_CACHE_FNV64 = 0x448274a57f508cbcull;

View file

@ -0,0 +1,36 @@
{
"seed": "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000",
"day": "bytes:69676e65756d2d6461792ffa50000000000000",
"dataset_mode": "memory-hard",
"dataset_log2_words": 28,
"mask": "0x0fffffff",
"lanes": 32,
"source": "igneum-pow (Rust) CPU interpreter, generator v3, memory-hard dataset",
"warps": [
{"base_nonce": 0, "expected": [
"0x2d96e3f8e9ad42e2", "0x0d36523b17ba1058", "0x899e2a6df6991c3d", "0xdb382aa991ff18d6", "0xc896fe7b624c7722", "0x932cfc8892d24216", "0xd9ef4047f797129b", "0x63fb054ee8cfe1d2",
"0xbcbd287cfd7e12a9", "0x68e41482f26203c2", "0x5c368b15e9ec2bf5", "0x83484f0946bfda39", "0xc82c1aeadba75e3c", "0x023cbca13fae681d", "0x000fcab814a56dd9", "0x432082a86895e2a7",
"0x635fb930f3cf454b", "0x1104c1b855a17655", "0x2fd53820b88956b2", "0xc415d67f6420191d", "0xeaf6853e1c859f5d", "0x95dfdf66b21e0518", "0x7d67b2b734243bfe", "0xb142ae8e4a4b5ca2",
"0x71ec0aff12d4c515", "0x0b2207c2c51fc2fa", "0xbc3957dbbf51d260", "0x74ea737eebe8a3e4", "0xd49fd6ea590200e9", "0xd5096fd4e83b8fa1", "0x999b099025a53907", "0xa5692c50b79eac9d"
]},
{"base_nonce": 4096, "expected": [
"0x9d4b6aaa8c292525", "0xb2ad4ab3941b3a3f", "0x17988da38edbf912", "0x566b31d7cceae84a", "0x132f77e29f4e6d40", "0xbdc098ecac64b3dc", "0x1670a7bfbc675d88", "0xda3a9960471e44bb",
"0x887e9fda5ec82271", "0x6137f598a67fc0d4", "0xd77b98d7d2025d26", "0x96b472c7560a777d", "0x0c98d0ff42f589f0", "0x4f0f15c9e4eefef2", "0x6bae44572933ab20", "0x2772c0266bdd5308",
"0x2fdb9e92d13d07ba", "0x6fe9239589ed122a", "0x7e3955e5a91916fd", "0x966d21685073c9b8", "0xa050194eb84c4104", "0xf7e095dae1d70633", "0x4f49616da323736b", "0x1fcc532045f01cd4",
"0x69549d64913199fc", "0x6684a1b257c011b0", "0xff318d60e8970b9a", "0x1429c63759eb426d", "0x3f470be3f4d4815a", "0x6e1d6c6b91b984e2", "0xda2dd7f31fea85ef", "0xb3944abe38dae250"
]},
{"base_nonce": 1000000, "expected": [
"0xd549a905f9121657", "0xd863810b66483e86", "0x65d735cf2e5447af", "0x9cda91c71e0f5790", "0x682c4a1be66d2444", "0xcb64852ca144ffe1", "0x8e72112d5db27544", "0x1a9e6bb74ea3e563",
"0x63c279f228faf1bc", "0xf7b77f8435da26a2", "0x3c28bd5b463df46f", "0xd2cdd0e9e968d2bd", "0x46eae240fc0c3f01", "0xbdd4c4e525220d3c", "0x345178054b1eac9d", "0x312f541e08ab8fe9",
"0x0cbc6f22e65887ad", "0xf9a0d018f9ce95cf", "0x994d29e263550764", "0x27a37ad3487c4018", "0xcdb7645aa05dac6a", "0x8578f1e945156d7f", "0x8e135ab725d23599", "0xfa902aebc88a620c",
"0x5bdb23834f693ba8", "0xae53782ddf331358", "0x2304fcfaad2616fc", "0x21e9216f6c0b14a6", "0xa88e90dc82b85f66", "0xba5aa73b369dd894", "0x97e37652c487cc5e", "0x36cb34ed0c8ce81f"
]}
],
"dataset_head": ["0x3dd50b1f", "0x93123701", "0x48edec90", "0x3a7d2407", "0xcdd88b45", "0x2c37c6b0", "0xb03a0ac0", "0x10e2c6a5", "0x5119540e", "0x3358ef2d", "0x657ac748", "0x09e2105c", "0x25e2929c", "0x5fa0cdcc", "0x4317cabd", "0x6178ee52"],
"dataset_last_index": 268435455,
"dataset_last": "0xf7b7180e",
"dataset_samples": [{"index": 59471966, "value": "0x88717593"}, {"index": 217795994, "value": "0x79f4c2eb"}, {"index": 208353206, "value": "0xa55c10fa"}, {"index": 42483309, "value": "0x1e48deed"}, {"index": 172547758, "value": "0x825505df"}, {"index": 148076330, "value": "0x056e4b2a"}, {"index": 183853158, "value": "0x51c5cdc5"}, {"index": 214389424, "value": "0x9133ff90"}, {"index": 267488061, "value": "0xf3e1f630"}, {"index": 169781097, "value": "0x8c0f4bca"}, {"index": 184093494, "value": "0x13c805d6"}, {"index": 153880993, "value": "0x7cd9cc39"}, {"index": 84977930, "value": "0x54d3ad95"}, {"index": 46426879, "value": "0x3142a64b"}, {"index": 3093825, "value": "0xf690ebf9"}, {"index": 225364072, "value": "0x5f05ecf5"}, {"index": 44593546, "value": "0x708fd94d"}, {"index": 260713159, "value": "0x57e3a213"}, {"index": 168250303, "value": "0xb60de9a6"}, {"index": 52384140, "value": "0xc3cacc1f"}, {"index": 223401610, "value": "0x72a0a173"}, {"index": 45554030, "value": "0x2a13f6c7"}, {"index": 95410555, "value": "0xfb5b4620"}, {"index": 175039924, "value": "0xeae5f6c3"}, {"index": 79171087, "value": "0x38e0ceb1"}, {"index": 267580473, "value": "0x103d3fa4"}, {"index": 24168642, "value": "0xcb03f2ae"}, {"index": 37981670, "value": "0x562969cc"}, {"index": 171551130, "value": "0xf1349e23"}, {"index": 195559979, "value": "0x3fa57b5d"}, {"index": 204611762, "value": "0x4fb463b9"}, {"index": 140997658, "value": "0xb2b70250"}, {"index": 138925853, "value": "0x7f689f8c"}, {"index": 86637313, "value": "0x27e5f93e"}, {"index": 20736778, "value": "0xbaeafce2"}, {"index": 219665210, "value": "0xaf3d3776"}, {"index": 160430336, "value": "0x084c3724"}, {"index": 264654675, "value": "0xffe6a16c"}, {"index": 8013395, "value": "0xd4d0d18d"}, {"index": 228945585, "value": "0x5d78da4e"}, {"index": 213884386, "value": "0xf0bd37c6"}, {"index": 104419827, "value": "0x261f1c0d"}, {"index": 44185464, "value": "0x7bfc0ff6"}, {"index": 142737231, "value": "0x08e5373b"}, {"index": 99284897, "value": "0xd23a038c"}, {"index": 132475900, "value": "0x124d74c0"}, {"index": 61861762, "value": "0xdd74cace"}, {"index": 132056166, "value": "0x5ab7c9ba"}, {"index": 262388043, "value": "0xc05d2942"}, {"index": 91878046, "value": "0x2d1e2f73"}, {"index": 117353561, "value": "0x52a27429"}, {"index": 124768597, "value": "0x0c828d10"}, {"index": 71352993, "value": "0x7514e589"}, {"index": 190698941, "value": "0xdb8ca7d2"}, {"index": 46055428, "value": "0xdb881ba6"}, {"index": 55281366, "value": "0x0ce95e1d"}, {"index": 165145231, "value": "0x3c93a15e"}, {"index": 106810753, "value": "0x73ec90f8"}, {"index": 171985651, "value": "0xe47554ed"}, {"index": 232085256, "value": "0x496959f3"}, {"index": 159510492, "value": "0xebf27832"}, {"index": 40072060, "value": "0x11f5885e"}, {"index": 209107596, "value": "0x69a1ff20"}, {"index": 39023794, "value": "0x804542bd"}],
"cache_head": ["0xebd9055c", "0x7eaf6b21", "0x9b610855", "0x134a8562", "0xb10059ba", "0x7f459e58", "0x8f3c7873", "0x3523e609", "0x75b8f4ca", "0x750d31b4", "0x0dae0782", "0x26f10015", "0x7ba87a41", "0x0efd8543", "0x691d6368", "0x8929d967"],
"cache_last_line": ["0x51474edf", "0xc3cfba93", "0xf21454cf", "0x9b79baac", "0x4d4fce23", "0xd701dfd5", "0x37357ab3", "0x1be693fa", "0xa701cb3b", "0x7467c620", "0x428184e8", "0xf4010df0", "0xe33aa88f", "0xc78d6d62", "0xae9ba9c0", "0xf4bba97e"],
"cache_fnv1a64": "0x448274a57f508cbc"
}

View file

@ -0,0 +1,281 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal).
// Built from source at runtime by proto-opencl/host.c, which passes these defines:
// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit)
// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default)
// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32
// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition
// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units;
// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32.
#ifndef IGNEUM_GROUP
#define IGNEUM_GROUP 32
#endif
#ifndef IGNEUM_EXCHANGE
#define IGNEUM_EXCHANGE 0
#endif
#ifdef __OPENCL_VERSION__
#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1)))
#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n]
#if IGNEUM_EXCHANGE == 1
#ifdef cl_khr_subgroups
#pragma OPENCL EXTENSION cl_khr_subgroups : enable
#endif
#ifdef cl_khr_subgroup_shuffle
#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable
#endif
#elif IGNEUM_EXCHANGE == 2
#pragma OPENCL EXTENSION cl_intel_subgroups : enable
#endif
#else
// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros.
#include "emu_opencl.h"
#endif
#if IGNEUM_EXCHANGE == 1
#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m))
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
#elif IGNEUM_EXCHANGE == 2
#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m))
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
#else
// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per
// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane
// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's
// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier.
#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; }
#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; }
#endif
static inline uint splitmix32(uint x) {
x ^= x >> 16; x *= 0x7feb352du;
x ^= x >> 15; x *= 0x846ca68bu;
x ^= x >> 16;
return x;
}
// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32.
static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); }
// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x.
static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); }
static inline uint ds_elem(uint i, uint d0, uint d1) {
uint x = i ^ d0;
x *= 0x9E3779B1u; x ^= x >> 15;
x += d1;
x *= 0x85EBCA77u; x ^= x >> 13;
x *= 0xC2B2AE3Du; x ^= x >> 16;
return x;
}
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
#define MH_CACHE_LINE_MASK 0x003fffffu
#define MH_SEGMENT_LINES 64u
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
// y = ChaCha12 core(x) + x
static inline void mh_chacha_block(const uint* x, uint* y) {
for (uint i = 0u; i < 16u; ++i) y[i] = x[i];
for (uint r = 0u; r < 6u; ++r) {
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
}
for (uint i = 0u; i < 16u; ++i) y[i] += x[i];
}
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
static inline void mh_cache_segment(__global uint* cache, uint seg) {
uint prev[16]; uint x[16]; uint y[16];
for (uint i = 0u; i < 16u; ++i) prev[i] = 0u;
for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) {
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
x[4] = 0xceed56d7u ^ prev[4];
x[5] = 0x9ba270d2u ^ prev[5];
x[6] = 0x82caab2du ^ prev[6];
x[7] = 0x81ebce0eu ^ prev[7];
x[8] = 0x12b6ecf1u ^ prev[8];
x[9] = 0xd0f3fd7cu ^ prev[9];
x[10] = 0xd872eefeu ^ prev[10];
x[11] = 0xc158c7bdu ^ prev[11];
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
mh_chacha_block(x, y);
__global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
}
}
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
static inline void mh_mixer(uint* s, uint rk) {
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
}
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
static inline void mh_item(__global const uint* cache, uint t, uint* s) {
s[0] = 0xceed56d7u;
s[1] = 0x9ba270d2u;
s[2] = 0x82caab2du;
s[3] = 0x81ebce0eu;
s[4] = 0x12b6ecf1u;
s[5] = 0xd0f3fd7cu;
s[6] = 0xd872eefeu;
s[7] = 0xc158c7bdu;
s[8] = t * 0xf351d601u + 0xc6892460u;
s[9] = t * 0xa3bb398fu + 0x25b7228au;
s[10] = t * 0xb5a09e35u + 0xcd515004u;
s[11] = t * 0x7509c9c1u + 0x2846527au;
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
s[14] = t * 0xded91851u + 0x82961bacu;
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
for (uint r = 0u; r < 8u; ++r) {
mh_mixer(s, 0x9E3779B9u * (r + 1u));
__global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i];
}
mh_mixer(s, 0x9E3779B9u * 9u);
}
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 2 3 8 13 of w.
static inline uint mh_j(uint w) { return ((w >> 2u) & 1u) | (((w >> 3u) & 1u) << 1) | (((w >> 8u) & 1u) << 2) | (((w >> 13u) & 1u) << 3); }
static inline uint mh_t(uint w) { w = (w & 0x00001fffu) | ((w >> 14u) << 13u); w = (w & 0x000000ffu) | ((w >> 9u) << 8u); w = (w & 0x00000007u) | ((w >> 4u) << 3u); w = (w & 0x00000003u) | ((w >> 3u) << 2u); return w; }
static inline uint mh_addr(uint t, uint j) { uint w = t; w = ((w >> 2u) << 3u) | (w & 0x00000003u) | (((j >> 0u) & 1u) << 2u); w = ((w >> 3u) << 4u) | (w & 0x00000007u) | (((j >> 1u) & 1u) << 3u); w = ((w >> 8u) << 9u) | (w & 0x000000ffu) | (((j >> 2u) & 1u) << 8u); w = ((w >> 13u) << 14u) | (w & 0x00001fffu) | (((j >> 3u) & 1u) << 13u); return w; }
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }
// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item.
// The same constants as memhard.h in this pack (one emitter, three dialects).
__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) {
uint seg = (uint)get_global_id(0);
if (seg < nSegments) mh_cache_segment(cache, seg);
}
__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) {
uint t = (uint)get_global_id(0);
if (t < nItems) {
uint s[16];
mh_item(cache, t, s);
for (uint i = 0u; i < 16u; ++i) ds[(ulong)mh_addr(t, i)] = s[i];
}
}
// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the
// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and
// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all).
IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) {
uint gid = (uint)get_global_id(0);
uint lid = (uint)get_local_id(0);
uint nonce = baseNonce + gid;
uint r0, r1, r2, r3, r4, r5, r6, r7;
#if IGNEUM_EXCHANGE == 0
IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
uint xk = 0u;
#else
(void)lid;
#endif
{ uint x = nonce ^ 0x667d0fbdu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x7b8e5963u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
{ uint x = nonce ^ 0x7b8e5963u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x31c67e5eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
{ uint x = nonce ^ 0x31c67e5eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4529ddc6u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
{ uint x = nonce ^ 0x4529ddc6u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef19d6d8u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
{ uint x = nonce ^ 0xef19d6d8u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xaccf6211u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
{ uint x = nonce ^ 0xaccf6211u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xda0aed32u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
{ uint x = nonce ^ 0xda0aed32u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xabc6df31u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
{ uint x = nonce ^ 0xabc6df31u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x667d0fbdu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
for (uint it = 0u; it < 8u; ++it) {
uint sel = r0;
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
r7 = r7 ^ r0; // 1 xor
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 1u); r3 = r3 ^ t_; } // 2 shfl
r4 = r4 - r1; // 3 sub
r2 = r0 * r4 + r2; // 4 mad
r4 = rotl_imm(r4, 9u); // 5 rotl
r0 = rotr_var(r0, r2); // 6 rotr
r6 = r6 ^ ds[((rotl_imm(r7 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
r1 = r1 ^ ds[((rotl_imm(r4 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
r1 = r1 ^ ds[((rotl_imm(r2 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
r7 = r7 ^ ds[((rotl_imm(r0 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
r7 = r7 ^ ds[((rotl_imm(r1 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
r5 = r5 * r4; // 12 mul
r7 = r7 ^ ds[((rotl_imm(r6 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 16u); r6 = r6 ^ t_; } // 14 shfl
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r3 = r3 ^ t_; } // 15 shfl
r2 = r2 | r7; // 16 or
r0 = rotr_var(r0, r6); // 17 rotr
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
r7 = rotl_imm(r7, 24u); // 19 rotl
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
r2 = r2 | r6; // 21 or
r6 = r5 * r4 + r6; // 22 mad
r1 = r1 | r0; // 23 or
r6 = r6 ^ r1; // 24 xor
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
r7 = rotr_var(r7, r0); // 26 rotr
r4 = r5 * r7 + r4; // 27 mad
r3 = r2 * r4 + r3; // 28 mad
r1 = r1 ^ ds[((rotl_imm(r4 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
r2 = r2 ^ ds[((rotl_imm(r3 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
r1 = r1 ^ ds[((rotl_imm(r5 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
r2 = mul_hi(r2, r5); // 35 mulhi
r4 = r4 ^ r2; // 36 xor
r6 = r6 * r5; // 37 mul
r7 = r7 ^ r0; // 38 xor
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
r2 = rotr_var(r2, r3); // 40 rotr
r7 = r7 - r0; // 41 sub
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 43 shfl
r0 = r0 ^ ds[((rotl_imm(r7 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
r3 = r3 ^ ds[((rotl_imm(r1 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
r6 = r6 ^ ds[((rotl_imm(r3 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
r4 = mul_hi(r4, r2); // 48 mulhi
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r0 = r0 ^ t_; } // 50 shfl
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
r5 = r5 ^ ds[((rotl_imm(r2 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
r6 = rotl_imm(r6, 12u); // 53 rotl
r3 = rotl_imm(r3, 12u); // 54 rotl
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
r1 = r1 ^ ds[((rotl_imm(r4 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
r5 = r5 ^ ds[((rotl_imm(r0 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
r3 = mul_hi(r3, r2); // 60 mulhi
r6 = r6 | r4; // 61 or
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r5 = r5 ^ t_; } // 62 shfl
r3 = r3 ^ ds[((rotl_imm(r6 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
}
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((ulong)hi << 32) | (ulong)lo;
}
#if IGNEUM_EXCHANGE != 0
// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the
// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a
// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md.
IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) {
if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); }
}
#endif

View file

@ -0,0 +1,163 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal).
// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC.
#include <cuda_runtime.h>
#include <cstdint>
#include "program.h"
#include "memhard.h"
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
x ^= x >> 16; x *= 0x7feb352du;
x ^= x >> 15; x *= 0x846ca68bu;
x ^= x >> 16;
return x;
}
// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31.
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x.
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) {
uint32_t x = i ^ d0;
x *= 0x9E3779B1u; x ^= x >> 15;
x += d1;
x *= 0x85EBCA77u; x ^= x >> 13;
x *= 0xC2B2AE3Du; x ^= x >> 16;
return x;
}
// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item.
// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host.
__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) {
uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x;
if (seg < nSegments) mh_cache_segment(cache, seg);
}
__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
uint32_t t = blockIdx.x * blockDim.x + threadIdx.x;
if (t < nItems) {
uint32_t s[16];
mh_item(cache, t, s);
for (uint32_t i = 0u; i < 16u; ++i) ds[(size_t)mh_addr(t, i)] = s[i];
}
}
// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every
// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a
// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid.
__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) {
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
uint32_t nonce = baseNonce + gid;
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
{ uint32_t x = nonce ^ 0x667d0fbdu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x7b8e5963u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
{ uint32_t x = nonce ^ 0x7b8e5963u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x31c67e5eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
{ uint32_t x = nonce ^ 0x31c67e5eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4529ddc6u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
{ uint32_t x = nonce ^ 0x4529ddc6u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef19d6d8u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
{ uint32_t x = nonce ^ 0xef19d6d8u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xaccf6211u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
{ uint32_t x = nonce ^ 0xaccf6211u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xda0aed32u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
{ uint32_t x = nonce ^ 0xda0aed32u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xabc6df31u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
{ uint32_t x = nonce ^ 0xabc6df31u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x667d0fbdu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
for (uint32_t it = 0u; it < 8u; ++it) {
uint32_t sel = r0;
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
r7 = r7 ^ r0; // 1 xor
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r6, 1); // 2 shfl
r4 = r4 - r1; // 3 sub
r2 = r0 * r4 + r2; // 4 mad
r4 = rotl_imm(r4, 9u); // 5 rotl
r0 = rotr_var(r0, r2); // 6 rotr
r6 = r6 ^ ds[((rotl_imm(r7 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
r1 = r1 ^ ds[((rotl_imm(r4 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
r1 = r1 ^ ds[((rotl_imm(r2 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
r7 = r7 ^ ds[((rotl_imm(r0 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
r7 = r7 ^ ds[((rotl_imm(r1 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
r5 = r5 * r4; // 12 mul
r7 = r7 ^ ds[((rotl_imm(r6 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r5, 16); // 14 shfl
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r5, 1); // 15 shfl
r2 = r2 | r7; // 16 or
r0 = rotr_var(r0, r6); // 17 rotr
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
r7 = rotl_imm(r7, 24u); // 19 rotl
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
r2 = r2 | r6; // 21 or
r6 = r5 * r4 + r6; // 22 mad
r1 = r1 | r0; // 23 or
r6 = r6 ^ r1; // 24 xor
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
r7 = rotr_var(r7, r0); // 26 rotr
r4 = r5 * r7 + r4; // 27 mad
r3 = r2 * r4 + r3; // 28 mad
r1 = r1 ^ ds[((rotl_imm(r4 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
r2 = r2 ^ ds[((rotl_imm(r3 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
r1 = r1 ^ ds[((rotl_imm(r5 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
r2 = __umulhi(r2, r5); // 35 mulhi
r4 = r4 ^ r2; // 36 xor
r6 = r6 * r5; // 37 mul
r7 = r7 ^ r0; // 38 xor
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
r2 = rotr_var(r2, r3); // 40 rotr
r7 = r7 - r0; // 41 sub
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 43 shfl
r0 = r0 ^ ds[((rotl_imm(r7 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
r3 = r3 ^ ds[((rotl_imm(r1 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
r6 = r6 ^ ds[((rotl_imm(r3 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
r4 = __umulhi(r4, r2); // 48 mulhi
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r7, 4); // 50 shfl
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
r5 = r5 ^ ds[((rotl_imm(r2 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
r6 = rotl_imm(r6, 12u); // 53 rotl
r3 = rotl_imm(r3, 12u); // 54 rotl
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
r1 = r1 ^ ds[((rotl_imm(r4 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
r5 = r5 ^ ds[((rotl_imm(r0 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
r3 = __umulhi(r3, r2); // 60 mulhi
r6 = r6 | r4; // 61 or
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 62 shfl
r3 = r3 ^ ds[((rotl_imm(r6 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
}
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
}
// Host-side launch wrappers. Declared in program.h, called from host.cu.
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) {
if (nSegments == 0u) return cudaErrorInvalidValue;
uint32_t block = 256u;
uint32_t grid = (nSegments + block - 1u) / block;
igneum_cache_fill<<<grid, block>>>(cache, nSegments);
return cudaGetLastError();
}
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
if (nItems == 0u) return cudaErrorInvalidValue;
uint32_t block = 256u;
uint32_t grid = (nItems + block - 1u) / block;
igneum_build<<<grid, block>>>(ds, cache, nItems);
return cudaGetLastError();
}
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
uint32_t nonces, uint32_t blockWarps) {
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
uint32_t block = 32u * blockWarps;
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
igneum_hash<<<nonces / block, block>>>(ds, out, baseNonce, mask);
return cudaGetLastError();
}
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
cudaFuncAttributes attr;
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash);
if (e != cudaSuccess) return e;
*numRegs = attr.numRegs;
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0);
}

View file

@ -0,0 +1,375 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal).
// Built from source at runtime by proto-opencl/host.c, which passes these defines:
// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit)
// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default)
// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32
// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition
// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units;
// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32.
#ifndef IGNEUM_GROUP
#define IGNEUM_GROUP 32
#endif
#ifndef IGNEUM_EXCHANGE
#define IGNEUM_EXCHANGE 0
#endif
#ifdef __OPENCL_VERSION__
#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1)))
#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n]
#if IGNEUM_EXCHANGE == 1
#ifdef cl_khr_subgroups
#pragma OPENCL EXTENSION cl_khr_subgroups : enable
#endif
#ifdef cl_khr_subgroup_shuffle
#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable
#endif
#elif IGNEUM_EXCHANGE == 2
#pragma OPENCL EXTENSION cl_intel_subgroups : enable
#endif
#else
// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros.
#include "emu_opencl.h"
#endif
#if IGNEUM_EXCHANGE == 1
#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m))
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
#elif IGNEUM_EXCHANGE == 2
#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m))
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
#else
// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per
// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane
// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's
// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier.
#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; }
#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; }
#endif
static inline uint splitmix32(uint x) {
x ^= x >> 16; x *= 0x7feb352du;
x ^= x >> 15; x *= 0x846ca68bu;
x ^= x >> 16;
return x;
}
// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32.
static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); }
// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x.
static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); }
static inline uint ds_elem(uint i, uint d0, uint d1) {
uint x = i ^ d0;
x *= 0x9E3779B1u; x ^= x >> 15;
x += d1;
x *= 0x85EBCA77u; x ^= x >> 13;
x *= 0xC2B2AE3Du; x ^= x >> 16;
return x;
}
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
#define MH_CACHE_LINE_MASK 0x003fffffu
#define MH_SEGMENT_LINES 64u
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
// y = ChaCha12 core(x) + x
static inline void mh_chacha_block(const uint* x, uint* y) {
for (uint i = 0u; i < 16u; ++i) y[i] = x[i];
for (uint r = 0u; r < 6u; ++r) {
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
}
for (uint i = 0u; i < 16u; ++i) y[i] += x[i];
}
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
static inline void mh_cache_segment(__global uint* cache, uint seg) {
uint prev[16]; uint x[16]; uint y[16];
for (uint i = 0u; i < 16u; ++i) prev[i] = 0u;
for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) {
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
x[4] = 0xceed56d7u ^ prev[4];
x[5] = 0x9ba270d2u ^ prev[5];
x[6] = 0x82caab2du ^ prev[6];
x[7] = 0x81ebce0eu ^ prev[7];
x[8] = 0x12b6ecf1u ^ prev[8];
x[9] = 0xd0f3fd7cu ^ prev[9];
x[10] = 0xd872eefeu ^ prev[10];
x[11] = 0xc158c7bdu ^ prev[11];
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
mh_chacha_block(x, y);
__global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
}
}
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
static inline void mh_mixer(uint* s, uint rk) {
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
}
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
static inline void mh_item(__global const uint* cache, uint t, uint* s) {
s[0] = 0xceed56d7u;
s[1] = 0x9ba270d2u;
s[2] = 0x82caab2du;
s[3] = 0x81ebce0eu;
s[4] = 0x12b6ecf1u;
s[5] = 0xd0f3fd7cu;
s[6] = 0xd872eefeu;
s[7] = 0xc158c7bdu;
s[8] = t * 0xf351d601u + 0xc6892460u;
s[9] = t * 0xa3bb398fu + 0x25b7228au;
s[10] = t * 0xb5a09e35u + 0xcd515004u;
s[11] = t * 0x7509c9c1u + 0x2846527au;
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
s[14] = t * 0xded91851u + 0x82961bacu;
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
for (uint r = 0u; r < 8u; ++r) {
mh_mixer(s, 0x9E3779B9u * (r + 1u));
__global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i];
}
mh_mixer(s, 0x9E3779B9u * 9u);
}
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 2 3 8 13 of w.
static inline uint mh_j(uint w) { return ((w >> 2u) & 1u) | (((w >> 3u) & 1u) << 1) | (((w >> 8u) & 1u) << 2) | (((w >> 13u) & 1u) << 3); }
static inline uint mh_t(uint w) { w = (w & 0x00001fffu) | ((w >> 14u) << 13u); w = (w & 0x000000ffu) | ((w >> 9u) << 8u); w = (w & 0x00000007u) | ((w >> 4u) << 3u); w = (w & 0x00000003u) | ((w >> 3u) << 2u); return w; }
static inline uint mh_addr(uint t, uint j) { uint w = t; w = ((w >> 2u) << 3u) | (w & 0x00000003u) | (((j >> 0u) & 1u) << 2u); w = ((w >> 3u) << 4u) | (w & 0x00000007u) | (((j >> 1u) & 1u) << 3u); w = ((w >> 8u) << 9u) | (w & 0x000000ffu) | (((j >> 2u) & 1u) << 8u); w = ((w >> 13u) << 14u) | (w & 0x00001fffu) | (((j >> 3u) & 1u) << 13u); return w; }
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }
// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item.
// The same constants as memhard.h in this pack (one emitter, three dialects).
__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) {
uint seg = (uint)get_global_id(0);
if (seg < nSegments) mh_cache_segment(cache, seg);
}
__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) {
uint t = (uint)get_global_id(0);
if (t < nItems) {
uint s[16];
mh_item(cache, t, s);
for (uint i = 0u; i < 16u; ++i) ds[(ulong)mh_addr(t, i)] = s[i];
}
}
// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the
// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and
// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all).
IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) {
uint gid = (uint)get_global_id(0);
uint lid = (uint)get_local_id(0);
uint nonce = baseNonce + gid;
uint r0, r1, r2, r3, r4, r5, r6, r7;
#if IGNEUM_EXCHANGE == 0
IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
uint xk = 0u;
#else
(void)lid;
#endif
{ uint x = nonce ^ 0x667d0fbdu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x7b8e5963u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
{ uint x = nonce ^ 0x7b8e5963u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x31c67e5eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
{ uint x = nonce ^ 0x31c67e5eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4529ddc6u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
{ uint x = nonce ^ 0x4529ddc6u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef19d6d8u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
{ uint x = nonce ^ 0xef19d6d8u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xaccf6211u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
{ uint x = nonce ^ 0xaccf6211u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xda0aed32u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
{ uint x = nonce ^ 0xda0aed32u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xabc6df31u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
{ uint x = nonce ^ 0xabc6df31u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x667d0fbdu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
for (uint it = 0u; it < 8u; ++it) {
uint sel = r0;
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
r7 = r7 ^ r0; // 1 xor
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 1u); r3 = r3 ^ t_; } // 2 shfl
r4 = r4 - r1; // 3 sub
r2 = r0 * r4 + r2; // 4 mad
r4 = rotl_imm(r4, 9u); // 5 rotl
r0 = rotr_var(r0, r2); // 6 rotr
r6 = r6 ^ ds[((rotl_imm(r7 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
r1 = r1 ^ ds[((rotl_imm(r4 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
r1 = r1 ^ ds[((rotl_imm(r2 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
r7 = r7 ^ ds[((rotl_imm(r0 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
r7 = r7 ^ ds[((rotl_imm(r1 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
r5 = r5 * r4; // 12 mul
r7 = r7 ^ ds[((rotl_imm(r6 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 16u); r6 = r6 ^ t_; } // 14 shfl
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r3 = r3 ^ t_; } // 15 shfl
r2 = r2 | r7; // 16 or
r0 = rotr_var(r0, r6); // 17 rotr
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
r7 = rotl_imm(r7, 24u); // 19 rotl
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
r2 = r2 | r6; // 21 or
r6 = r5 * r4 + r6; // 22 mad
r1 = r1 | r0; // 23 or
r6 = r6 ^ r1; // 24 xor
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
r7 = rotr_var(r7, r0); // 26 rotr
r4 = r5 * r7 + r4; // 27 mad
r3 = r2 * r4 + r3; // 28 mad
r1 = r1 ^ ds[((rotl_imm(r4 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
r2 = r2 ^ ds[((rotl_imm(r3 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
r1 = r1 ^ ds[((rotl_imm(r5 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
r2 = mul_hi(r2, r5); // 35 mulhi
r4 = r4 ^ r2; // 36 xor
r6 = r6 * r5; // 37 mul
r7 = r7 ^ r0; // 38 xor
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
r2 = rotr_var(r2, r3); // 40 rotr
r7 = r7 - r0; // 41 sub
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 43 shfl
r0 = r0 ^ ds[((rotl_imm(r7 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
r3 = r3 ^ ds[((rotl_imm(r1 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
r6 = r6 ^ ds[((rotl_imm(r3 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
r4 = mul_hi(r4, r2); // 48 mulhi
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r0 = r0 ^ t_; } // 50 shfl
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
r5 = r5 ^ ds[((rotl_imm(r2 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
r6 = rotl_imm(r6, 12u); // 53 rotl
r3 = rotl_imm(r3, 12u); // 54 rotl
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
r1 = r1 ^ ds[((rotl_imm(r4 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
r5 = r5 ^ ds[((rotl_imm(r0 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
r3 = mul_hi(r3, r2); // 60 mulhi
r6 = r6 | r4; // 61 or
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r5 = r5 ^ t_; } // 62 shfl
r3 = r3 ^ ds[((rotl_imm(r6 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
}
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((ulong)hi << 32) | (ulong)lo;
}
#if IGNEUM_EXCHANGE != 0
// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the
// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a
// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md.
IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) {
if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); }
}
#endif
// Header-bound variant (bind.rs): the init words come from initw, not SEEDW. Same body as igneum_hash.
IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw) {
uint gid = (uint)get_global_id(0);
uint lid = (uint)get_local_id(0);
uint nonce = baseNonce + gid;
uint r0, r1, r2, r3, r4, r5, r6, r7;
uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7];
#if IGNEUM_EXCHANGE == 0
IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
uint xk = 0u;
#else
(void)lid;
#endif
{ uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; }
{ uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; }
{ uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; }
{ uint x = nonce ^ iw3; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw4; }
{ uint x = nonce ^ iw4; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw5; }
{ uint x = nonce ^ iw5; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw6; }
{ uint x = nonce ^ iw6; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw7; }
{ uint x = nonce ^ iw7; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw0; }
for (uint it = 0u; it < 8u; ++it) {
uint sel = r0;
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
r7 = r7 ^ r0; // 1 xor
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 1u); r3 = r3 ^ t_; } // 2 shfl
r4 = r4 - r1; // 3 sub
r2 = r0 * r4 + r2; // 4 mad
r4 = rotl_imm(r4, 9u); // 5 rotl
r0 = rotr_var(r0, r2); // 6 rotr
r6 = r6 ^ ds[((rotl_imm(r7 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
r1 = r1 ^ ds[((rotl_imm(r4 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
r1 = r1 ^ ds[((rotl_imm(r2 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
r7 = r7 ^ ds[((rotl_imm(r0 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
r7 = r7 ^ ds[((rotl_imm(r1 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
r5 = r5 * r4; // 12 mul
r7 = r7 ^ ds[((rotl_imm(r6 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 16u); r6 = r6 ^ t_; } // 14 shfl
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r3 = r3 ^ t_; } // 15 shfl
r2 = r2 | r7; // 16 or
r0 = rotr_var(r0, r6); // 17 rotr
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
r7 = rotl_imm(r7, 24u); // 19 rotl
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
r2 = r2 | r6; // 21 or
r6 = r5 * r4 + r6; // 22 mad
r1 = r1 | r0; // 23 or
r6 = r6 ^ r1; // 24 xor
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
r7 = rotr_var(r7, r0); // 26 rotr
r4 = r5 * r7 + r4; // 27 mad
r3 = r2 * r4 + r3; // 28 mad
r1 = r1 ^ ds[((rotl_imm(r4 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
r2 = r2 ^ ds[((rotl_imm(r3 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
r1 = r1 ^ ds[((rotl_imm(r5 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
r2 = mul_hi(r2, r5); // 35 mulhi
r4 = r4 ^ r2; // 36 xor
r6 = r6 * r5; // 37 mul
r7 = r7 ^ r0; // 38 xor
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
r2 = rotr_var(r2, r3); // 40 rotr
r7 = r7 - r0; // 41 sub
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 43 shfl
r0 = r0 ^ ds[((rotl_imm(r7 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
r3 = r3 ^ ds[((rotl_imm(r1 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
r6 = r6 ^ ds[((rotl_imm(r3 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
r4 = mul_hi(r4, r2); // 48 mulhi
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r0 = r0 ^ t_; } // 50 shfl
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
r5 = r5 ^ ds[((rotl_imm(r2 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
r6 = rotl_imm(r6, 12u); // 53 rotl
r3 = rotl_imm(r3, 12u); // 54 rotl
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
r1 = r1 ^ ds[((rotl_imm(r4 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
r5 = r5 ^ ds[((rotl_imm(r0 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
r3 = mul_hi(r3, r2); // 60 mulhi
r6 = r6 | r4; // 61 or
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r5 = r5 ^ t_; } // 62 shfl
r3 = r3 ^ ds[((rotl_imm(r6 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
}
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((ulong)hi << 32) | (ulong)lo;
}

View file

@ -0,0 +1,123 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
// Header-bound twin of igneum_hash in kernel.cu: the init words come from a kernel argument, not SEEDW.
// Host declarations (also in program_bound.h if present):
// struct IgneumInitWords { uint32_t w[8]; };
// cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
// IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps);
// cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
#include <cuda_runtime.h>
#include <cstdint>
#include "program.h"
struct IgneumInitWords { uint32_t w[8]; };
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
x ^= x >> 16; x *= 0x7feb352du;
x ^= x >> 15; x *= 0x846ca68bu;
x ^= x >> 16;
return x;
}
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw) {
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
uint32_t nonce = baseNonce + gid;
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
{ uint32_t x = nonce ^ iw.w[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw.w[1]; }
{ uint32_t x = nonce ^ iw.w[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw.w[2]; }
{ uint32_t x = nonce ^ iw.w[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw.w[3]; }
{ uint32_t x = nonce ^ iw.w[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw.w[4]; }
{ uint32_t x = nonce ^ iw.w[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw.w[5]; }
{ uint32_t x = nonce ^ iw.w[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw.w[6]; }
{ uint32_t x = nonce ^ iw.w[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw.w[7]; }
{ uint32_t x = nonce ^ iw.w[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw.w[0]; }
for (uint32_t it = 0u; it < 8u; ++it) {
uint32_t sel = r0;
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
r7 = r7 ^ r0; // 1 xor
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r6, 1); // 2 shfl
r4 = r4 - r1; // 3 sub
r2 = r0 * r4 + r2; // 4 mad
r4 = rotl_imm(r4, 9u); // 5 rotl
r0 = rotr_var(r0, r2); // 6 rotr
r6 = r6 ^ ds[((rotl_imm(r7 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
r1 = r1 ^ ds[((rotl_imm(r4 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
r1 = r1 ^ ds[((rotl_imm(r2 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
r7 = r7 ^ ds[((rotl_imm(r0 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
r7 = r7 ^ ds[((rotl_imm(r1 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
r5 = r5 * r4; // 12 mul
r7 = r7 ^ ds[((rotl_imm(r6 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r5, 16); // 14 shfl
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r5, 1); // 15 shfl
r2 = r2 | r7; // 16 or
r0 = rotr_var(r0, r6); // 17 rotr
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
r7 = rotl_imm(r7, 24u); // 19 rotl
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
r2 = r2 | r6; // 21 or
r6 = r5 * r4 + r6; // 22 mad
r1 = r1 | r0; // 23 or
r6 = r6 ^ r1; // 24 xor
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
r7 = rotr_var(r7, r0); // 26 rotr
r4 = r5 * r7 + r4; // 27 mad
r3 = r2 * r4 + r3; // 28 mad
r1 = r1 ^ ds[((rotl_imm(r4 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
r2 = r2 ^ ds[((rotl_imm(r3 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
r1 = r1 ^ ds[((rotl_imm(r5 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
r2 = __umulhi(r2, r5); // 35 mulhi
r4 = r4 ^ r2; // 36 xor
r6 = r6 * r5; // 37 mul
r7 = r7 ^ r0; // 38 xor
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
r2 = rotr_var(r2, r3); // 40 rotr
r7 = r7 - r0; // 41 sub
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 43 shfl
r0 = r0 ^ ds[((rotl_imm(r7 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
r3 = r3 ^ ds[((rotl_imm(r1 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
r6 = r6 ^ ds[((rotl_imm(r3 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
r4 = __umulhi(r4, r2); // 48 mulhi
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r7, 4); // 50 shfl
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
r5 = r5 ^ ds[((rotl_imm(r2 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
r6 = rotl_imm(r6, 12u); // 53 rotl
r3 = rotl_imm(r3, 12u); // 54 rotl
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
r1 = r1 ^ ds[((rotl_imm(r4 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
r5 = r5 ^ ds[((rotl_imm(r0 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
r3 = __umulhi(r3, r2); // 60 mulhi
r6 = r6 | r4; // 61 or
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 62 shfl
r3 = r3 ^ ds[((rotl_imm(r6 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
}
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
}
cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps) {
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
uint32_t block = 32u * blockWarps;
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
igneum_hash_bound<<<nonces / block, block>>>(ds, out, baseNonce, mask, iw);
return cudaGetLastError();
}
cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
cudaFuncAttributes attr;
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash_bound);
if (e != cudaSuccess) return e;
*numRegs = attr.numRegs;
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash_bound, (int)(32u * blockWarps), 0);
}

View file

@ -0,0 +1,112 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
// Memory-hard dataset core, the same text that the Mac's Metal kernels and CPU verifier were checked against.
// Included by kernel.cu (device), host.cu (host reference) and proto-opencl/host.c (C99 host reference).
// See proto-metal/MEMHARD.md for the construction. kernel.cl carries the same text in OpenCL C.
#pragma once
#ifdef __cplusplus
#include <cstdint>
#else
#include <stdint.h>
#endif
#if defined(__CUDACC__)
#define IGNEUM_HD __host__ __device__ __forceinline__
#elif defined(_MSC_VER) && !defined(__cplusplus)
#define IGNEUM_HD static __inline
#else
#define IGNEUM_HD static inline
#endif
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
#define MH_CACHE_LINE_MASK 0x003fffffu
#define MH_SEGMENT_LINES 64u
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
IGNEUM_HD uint32_t mh_rotl(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
// y = ChaCha12 core(x) + x
IGNEUM_HD void mh_chacha_block(const uint32_t* x, uint32_t* y) {
for (uint32_t i = 0u; i < 16u; ++i) y[i] = x[i];
for (uint32_t r = 0u; r < 6u; ++r) {
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
}
for (uint32_t i = 0u; i < 16u; ++i) y[i] += x[i];
}
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
IGNEUM_HD void mh_cache_segment(uint32_t* cache, uint32_t seg) {
uint32_t prev[16]; uint32_t x[16]; uint32_t y[16];
for (uint32_t i = 0u; i < 16u; ++i) prev[i] = 0u;
for (uint32_t j = 0u; j < MH_SEGMENT_LINES; ++j) {
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
x[4] = 0xceed56d7u ^ prev[4];
x[5] = 0x9ba270d2u ^ prev[5];
x[6] = 0x82caab2du ^ prev[6];
x[7] = 0x81ebce0eu ^ prev[7];
x[8] = 0x12b6ecf1u ^ prev[8];
x[9] = 0xd0f3fd7cu ^ prev[9];
x[10] = 0xd872eefeu ^ prev[10];
x[11] = 0xc158c7bdu ^ prev[11];
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
mh_chacha_block(x, y);
uint32_t* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
for (uint32_t i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
}
}
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
IGNEUM_HD void mh_mixer(uint32_t* s, uint32_t rk) {
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
}
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
IGNEUM_HD void mh_item(const uint32_t* cache, uint32_t t, uint32_t* s) {
s[0] = 0xceed56d7u;
s[1] = 0x9ba270d2u;
s[2] = 0x82caab2du;
s[3] = 0x81ebce0eu;
s[4] = 0x12b6ecf1u;
s[5] = 0xd0f3fd7cu;
s[6] = 0xd872eefeu;
s[7] = 0xc158c7bdu;
s[8] = t * 0xf351d601u + 0xc6892460u;
s[9] = t * 0xa3bb398fu + 0x25b7228au;
s[10] = t * 0xb5a09e35u + 0xcd515004u;
s[11] = t * 0x7509c9c1u + 0x2846527au;
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
s[14] = t * 0xded91851u + 0x82961bacu;
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
for (uint32_t r = 0u; r < 8u; ++r) {
mh_mixer(s, 0x9E3779B9u * (r + 1u));
const uint32_t* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= line[i];
}
mh_mixer(s, 0x9E3779B9u * 9u);
}
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 2 3 8 13 of w.
IGNEUM_HD uint32_t mh_j(uint32_t w) { return ((w >> 2u) & 1u) | (((w >> 3u) & 1u) << 1) | (((w >> 8u) & 1u) << 2) | (((w >> 13u) & 1u) << 3); }
IGNEUM_HD uint32_t mh_t(uint32_t w) { w = (w & 0x00001fffu) | ((w >> 14u) << 13u); w = (w & 0x000000ffu) | ((w >> 9u) << 8u); w = (w & 0x00000007u) | ((w >> 4u) << 3u); w = (w & 0x00000003u) | ((w >> 3u) << 2u); return w; }
IGNEUM_HD uint32_t mh_addr(uint32_t t, uint32_t j) { uint32_t w = t; w = ((w >> 2u) << 3u) | (w & 0x00000003u) | (((j >> 0u) & 1u) << 2u); w = ((w >> 3u) << 4u) | (w & 0x00000007u) | (((j >> 1u) & 1u) << 3u); w = ((w >> 8u) << 9u) | (w & 0x000000ffu) | (((j >> 2u) & 1u) << 8u); w = ((w >> 13u) << 14u) | (w & 0x00001fffu) | (((j >> 3u) & 1u) << 13u); return w; }
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
IGNEUM_HD uint32_t mh_word(const uint32_t* cache, uint32_t w) { uint32_t s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }

View file

@ -0,0 +1,109 @@
#include <metal_stdlib>
using namespace metal;
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
#define MH_CACHE_LINE_MASK 0x003fffffu
#define MH_SEGMENT_LINES 64u
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
// y = ChaCha12 core(x) + x
inline void mh_chacha_block(const thread uint* x, thread uint* y) {
for (uint i = 0u; i < 16u; ++i) y[i] = x[i];
for (uint r = 0u; r < 6u; ++r) {
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
}
for (uint i = 0u; i < 16u; ++i) y[i] += x[i];
}
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
inline void mh_cache_segment(device uint* cache, uint seg) {
uint prev[16]; uint x[16]; uint y[16];
for (uint i = 0u; i < 16u; ++i) prev[i] = 0u;
for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) {
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
x[4] = 0xceed56d7u ^ prev[4];
x[5] = 0x9ba270d2u ^ prev[5];
x[6] = 0x82caab2du ^ prev[6];
x[7] = 0x81ebce0eu ^ prev[7];
x[8] = 0x12b6ecf1u ^ prev[8];
x[9] = 0xd0f3fd7cu ^ prev[9];
x[10] = 0xd872eefeu ^ prev[10];
x[11] = 0xc158c7bdu ^ prev[11];
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
mh_chacha_block(x, y);
device uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
}
}
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
inline void mh_mixer(thread uint* s, uint rk) {
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
}
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
inline void mh_item(device const uint* cache, uint t, thread uint* s) {
s[0] = 0xceed56d7u;
s[1] = 0x9ba270d2u;
s[2] = 0x82caab2du;
s[3] = 0x81ebce0eu;
s[4] = 0x12b6ecf1u;
s[5] = 0xd0f3fd7cu;
s[6] = 0xd872eefeu;
s[7] = 0xc158c7bdu;
s[8] = t * 0xf351d601u + 0xc6892460u;
s[9] = t * 0xa3bb398fu + 0x25b7228au;
s[10] = t * 0xb5a09e35u + 0xcd515004u;
s[11] = t * 0x7509c9c1u + 0x2846527au;
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
s[14] = t * 0xded91851u + 0x82961bacu;
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
for (uint r = 0u; r < 8u; ++r) {
mh_mixer(s, 0x9E3779B9u * (r + 1u));
device const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i];
}
mh_mixer(s, 0x9E3779B9u * 9u);
}
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 2 3 8 13 of w.
inline uint mh_j(uint w) { return ((w >> 2u) & 1u) | (((w >> 3u) & 1u) << 1) | (((w >> 8u) & 1u) << 2) | (((w >> 13u) & 1u) << 3); }
inline uint mh_t(uint w) { w = (w & 0x00001fffu) | ((w >> 14u) << 13u); w = (w & 0x000000ffu) | ((w >> 9u) << 8u); w = (w & 0x00000007u) | ((w >> 4u) << 3u); w = (w & 0x00000003u) | ((w >> 3u) << 2u); return w; }
inline uint mh_addr(uint t, uint j) { uint w = t; w = ((w >> 2u) << 3u) | (w & 0x00000003u) | (((j >> 0u) & 1u) << 2u); w = ((w >> 3u) << 4u) | (w & 0x00000007u) | (((j >> 1u) & 1u) << 3u); w = ((w >> 8u) << 9u) | (w & 0x000000ffu) | (((j >> 2u) & 1u) << 8u); w = ((w >> 13u) << 14u) | (w & 0x00001fffu) | (((j >> 3u) & 1u) << 13u); return w; }
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
inline uint mh_word(device const uint* cache, uint w) { uint s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }
// One thread per segment (2^16 threads).
kernel void igneum_cache_fill(device uint* cache [[buffer(0)]], uint gid [[thread_position_in_grid]]) {
mh_cache_segment(cache, gid);
}
// One thread per 64-byte item (dataset words / 16 threads).
kernel void igneum_build(device const uint* cache [[buffer(0)]], device uint* dataset [[buffer(1)]],
uint gid [[thread_position_in_grid]]) {
uint s[16];
mh_item(cache, gid, s);
for (uint i = 0u; i < 16u; ++i) dataset[mh_addr(gid, i)] = s[i];
}

View file

@ -0,0 +1,76 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
// Program metadata for host.cu plus the launch wrappers defined in kernel.cu.
// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros.
#pragma once
#ifdef __cplusplus
#include <cstdint>
#else
#include <stdint.h>
#endif
#ifndef IGNEUM_NO_CUDA
#include <cuda_runtime.h>
#endif
#define IGNEUM_SEED_STRING "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000"
#define IGNEUM_SEED_BYTES_HEX "edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07"
#define IGNEUM_GENERATOR 3
#define IGNEUM_PROGRAM_ATTEMPT 0
#define IGNEUM_PROGRAM_ID 0x73bcbfe8ccf988f1ull
#define IGNEUM_DAY_STRING "bytes:69676e65756d2d6461792ffa50000000000000"
#define IGNEUM_DAY_BYTES_HEX "69676e65756d2d6461792ffa50000000000000"
#define IGNEUM_DAY0 0xceed56d7u
#define IGNEUM_DAY1 0x9ba270d2u
#define IGNEUM_DATASET_LOG2 28
#define IGNEUM_MASK 0x0fffffffu
#define IGNEUM_LANES 32
#define IGNEUM_ITERATIONS 8
#define IGNEUM_INSTR_COUNT 64
#define IGNEUM_LOADS_PER_HASH 128
#define IGNEUM_WIDE_LOADS_PER_HASH 0
#define IGNEUM_OP_MIX "load=16 add=15 shfl=6 mad=4 or=4 rotl=4 rotr=4 xor=4 mulhi=3 mul=2 sub=2"
// Program class v3 (Counter ASIC 2.0, docs/plans/counter-asic-2-rollout.md): generator version 3; a worker that
// runs another class refuses this pack, and a job line names the class it wants (class=v3 era=<hex>).
#define IGNEUM_PROGRAM_CLASS "v3"
#define IGNEUM_ERA_SEED_HEX "e593fc1d48475c88e456632f6aa0a752a5fdbca3c33022afe37680e44b7dfd11"
// Class v3 construction (Counter ASIC 2.0, 5 October 2026, docs/plans/mixer-x4.md): version 2 loads; the dataset item
// derivation applies the mixer IGNEUM_MIXER_MULT times per round (memhard.h), and the cache follows the growth rule.
#define IGNEUM_LOAD_CLASS "w4-erad6367bfe"
#define IGNEUM_LOAD_SLOTS 16
#define IGNEUM_LOAD_MIX { 100, 0, 0 }
#define IGNEUM_LOAD_WIDTH_COUNTS { 16, 0, 0 } // loads of 4, 16, 64 bytes per program
#define IGNEUM_BYTES_PER_HASH 512
#define IGNEUM_FOLD_ROT 11
#define IGNEUM_FOLD_MUL 0x9e3779b1u
// Era layout (5 October 2026, docs/plans/era-layout.md): NOT the lottery hash. Every dataset load reads
// idx = ((rotl(src * STRIDE_MUL, STRIDE_ROT) & window mask) | window offset) & MASK; the window of a load site is the
// dataset, a half or a quarter of it (IGNEUM_ERA_WINDOWS: site:shrink:offset); dataset word w holds word j(w) of item
// t(w) with j's bits at the INTERLEAVE positions (memhard.h: mh_t, mh_j, mh_addr).
#define IGNEUM_ERA_LABEL "d6367bfe"
#define IGNEUM_ERA_SEED_WORDS { 0xd6367bfeu, 0x8bee0097u, 0x6c3c1b37u, 0x5dc9cefcu, 0x50b52d88u, 0x03833326u, 0x28ca58f9u, 0xa6e3b5c0u }
#define IGNEUM_ERA_ALLOWED_WIDTHS { 1, 0, 0 } // words, ascending, 0 = unused; one entry pins the width
#define IGNEUM_ERA_WIDTH_WORDS 1
#define IGNEUM_ERA_STRIDE_MUL 0x27ea7effu
#define IGNEUM_ERA_STRIDE_ROT 30
#define IGNEUM_ERA_INTERLEAVE { 2, 3, 8, 13 }
#define IGNEUM_ERA_WINDOWS "7:2:1 8:1:1 9:1:1 10:1:1 11:0:0 13:1:1 29:0:0 30:2:2 31:1:1 44:1:1 46:2:0 47:0:0 52:0:0 56:0:0 58:2:0 63:1:1"
// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h)
#define IGNEUM_DATASET_MODE 1
#define IGNEUM_SEEDW_INIT { 0x667d0fbdu, 0x7b8e5963u, 0x31c67e5eu, 0x4529ddc6u, 0xef19d6d8u, 0xaccf6211u, 0xda0aed32u, 0xabc6df31u }
#define IGNEUM_KEY_INIT { 0xceed56d7u, 0x9ba270d2u, 0x82caab2du, 0x81ebce0eu, 0x12b6ecf1u, 0xd0f3fd7cu, 0xd872eefeu, 0xc158c7bdu }
#define IGNEUM_CACHE_LOG2_WORDS 26
#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6
#define IGNEUM_CACHE_SEGMENTS 65536u
#define IGNEUM_ITEM_ROUNDS 8
#define IGNEUM_MIX_ROT_INIT { 17u, 12u, 20u, 23u, 7u, 3u, 27u, 16u }
#define IGNEUM_MIX_MUL_INIT { 0xf351d601u, 0xa3bb398fu, 0xb5a09e35u, 0x7509c9c1u, 0x6bbf31e9u, 0xfc849a79u, 0xded91851u, 0x8d9113d1u, 0x0ff15225u, 0x3a5bdd41u, 0xab533435u, 0xe1c55ad5u, 0xe6d3bd0du, 0x9d9ffbbdu, 0xbb2a3cf3u, 0x50a7c08du }
#define IGNEUM_MIX_RC_INIT { 0xc6892460u, 0x25b7228au, 0xcd515004u, 0x2846527au, 0xa6324241u, 0x36e3ec53u, 0x82961bacu, 0x0f97ba7du, 0xb6f921a9u, 0x3ada24e5u, 0xde20ab91u, 0x5378eeb2u, 0x7d161662u, 0x89353cc1u, 0xb1aa03a2u, 0x788acae6u }
#ifndef IGNEUM_NO_CUDA
// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError().
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments);
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems);
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
uint32_t nonces, uint32_t blockWarps);
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
#endif

View file

@ -0,0 +1,143 @@
{
"format": "igneum-program-pack-3",
"generator": 3,
"attempt": 0,
"program_id": "0x73bcbfe8ccf988f1",
"program_id_derivation": "FNV-1a 64 over 'igneum-program/' || generator_le32 || seed_words as little-endian bytes || attempt_le32",
"dataset_mode": "memory-hard",
"seed": "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000",
"seed_bytes": "edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07",
"seed_words": ["0x667d0fbd", "0x7b8e5963", "0x31c67e5e", "0x4529ddc6", "0xef19d6d8", "0xaccf6211", "0xda0aed32", "0xabc6df31"],
"seed_derivation": "seed_words = FNV-1a 64 over seed_bytes (attempt 0) or seed_bytes || attempt_le32 (attempt k >= 1), basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32",
"generator_rule": "version 2: exactly 16 load slots drawn first from instructions 1..63 (partial Fisher-Yates), the other 48 ops from the ten non-load weights (sum 75); a load's source is drawn from the registers other than dst written by an earlier instruction and not read by a load since; the candidate must pass the acceptance rule of spec 01 section 1.4.6 (static: no cyclically stale load source, every register has an injecting write; dynamic: 64 units on the seed-keyed closed-form dataset with no constant register bit, no lane-constant load site, under 164 saturated final values, every output bit within 136 of 1024, distinct addresses above 245760), else the next attempt of the seed is tried",
"lanes": 32,
"registers": 8,
"iterations": 8,
"instruction_count": 64,
"loads_per_hash": 128,
"program_class": "v3",
"era_seed_bytes": "e593fc1d48475c88e456632f6aa0a752a5fdbca3c33022afe37680e44b7dfd11",
"load_class": "w4-erad6367bfe",
"load_slots": 16,
"load_mix_percent_4_16_64": [100, 0, 0],
"load_width_counts_4_16_64": [16, 0, 0],
"bytes_per_hash": 512,
"wide_load": "read-width experiment (5 October 2026, docs/plans/read-width.md), NOT the lottery hash: a load of W words (width field, 4 or 16) reads dataset[b .. b + W) with b = (src & mask) & ~(W - 1) and folds every word into dst: x = dst ^ w[0]; for j in 1..W: x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x; width 1 is the plain load; the width is drawn per instruction from the class mix with one extra below(100) draw after the nine of version 2, and the program id is FNV-1a 64 over 'igneum-program-rw/' || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots",
"era": {
"label": "d6367bfe",
"seed_words": ["0xd6367bfe", "0x8bee0097", "0x6c3c1b37", "0x5dc9cefc", "0x50b52d88", "0x03833326", "0x28ca58f9", "0xa6e3b5c0"],
"draw": "docs/plans/era-layout.md 1.1: SplitMix64 seeded with seed_words[0] | seed_words[1] << 32 of seed_words_from_bytes('igneum-era/' || n_le64 || E_n); width = allowed[below(|allowed|)], stride_mul = low32(next()) | 1, stride_rot = 1 + below(31), then four next() draws for a partial Fisher-Yates over positions log2(W)..15 of which 4 - log2(W) are used",
"allowed_widths": [1],
"width_words": 1,
"stride_mul": "0x27ea7eff",
"stride_rot": 30,
"interleave": [2, 3, 8, 13],
"address": "y = rotl(src * stride_mul, stride_rot); k = min(win, D - 26); idx = ((y & (mask >> k)) | ((off & (2^k - 1)) << (D - k))) & mask; a wide load aligns idx down to W words",
"windows": "per instruction, after the width roll: win = below(3), off = low32(next()) & (2^win - 1); used on a load slot (the instruction's win and off fields)",
"dataset_word": "dataset[w] = item(t(w))[j(w)]: j(w) gathers the bits of w at the interleave positions, t(w) is w with those bits removed",
"program_id_suffix": "'era/' || allowed[3] || width_words || stride_mul_le32 || stride_rot_le32 || interleave[4]"
},
"op_mix": {"load": 16, "add": 15, "shfl": 6, "mad": 4, "or": 4, "rotl": 4, "rotr": 4, "xor": 4, "mulhi": 3, "mul": 2, "sub": 2},
"register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]",
"splitmix32": "x ^= x>>16; x *= 0x7feb352d; x ^= x>>15; x *= 0x846ca68b; x ^= x>>16",
"iteration": "sel = r0 sampled once at the top of each iteration, then all instructions in order",
"output": "lo = r0 ^ rotl(r1,7) ^ rotl(r2,14) ^ rotl(r3,21); hi = r4 ^ rotl(r5,9) ^ rotl(r6,18) ^ rotl(r7,27); out = (hi << 32) | lo",
"op_semantics": {
"add": "dst = dst + src + (bit `bit` of sel ? imm2 : imm)",
"sub": "dst = dst - src",
"mul": "dst = dst * src (low 32)",
"mulhi": "dst = high 32 bits of dst * src",
"xor": "dst = dst ^ src",
"or": "dst = dst | src",
"rotl": "dst = rotl(dst, rot), rot in 1..31",
"rotr": "dst = rotr(dst, src & 31)",
"mad": "dst = src * src2 + dst",
"shfl": "dst = dst ^ (src of lane (lane ^ mask)), mask in {1,2,4,8,16}, within the 32-lane warp",
"load": "dst = dst ^ dataset[src & dataset.mask]",
"wload": "base = (src of lane 0 & dataset.mask) & ~31; dst = dst ^ dataset[base + lane] (warp-coalesced 128-byte load, lever b, only when --wide-frac > 0)"
},
"dataset": {
"log2_words": 28,
"bytes": 1073741824,
"mask": "0x0fffffff",
"day": "bytes:69676e65756d2d6461792ffa50000000000000",
"day_bytes": "69676e65756d2d6461792ffa50000000000000",
"day_words_from": "seed_words_from_bytes(day_bytes)",
"d0": "0xceed56d7",
"d1": "0x9ba270d2",
"mode": "memory-hard",
"spec": "proto-metal/MEMHARD.md",
"key": ["0xceed56d7", "0x9ba270d2", "0x82caab2d", "0x81ebce0e", "0x12b6ecf1", "0xd0f3fd7c", "0xd872eefe", "0xc158c7bd"],
"key_derivation": "the 8 words of seed_words_from_bytes(day_bytes); d0, d1 are key[0], key[1]",
"cache": {"log2_words": 26, "bytes": 268435456, "line_words": 16, "segment_lines": 64, "segments": 65536, "block": "ChaCha12 core + feed-forward, rotations 16 12 8 7", "sigma": ["0x61707865", "0x3320646e", "0x79622d32", "0x6b206574"], "tag": ["0x49676e65", "0x756d4d48"], "chain": "in_j = prev_line ^ (sigma[0..3] || key[0..7] || seg || j || tag[0..1]); line_j = block(in_j); prev_0 = 0"},
"mixer": {"draw": "SplitMix64 seeded with key[0] | key[1] << 32: rot[0..7] = 1 + next() % 31, mul[0..15] = low32(next()) | 1, rc[0..15] = low32(next())", "rot": [17, 12, 20, 23, 7, 3, 27, 16], "mul": ["0xf351d601", "0xa3bb398f", "0xb5a09e35", "0x7509c9c1", "0x6bbf31e9", "0xfc849a79", "0xded91851", "0x8d9113d1", "0x0ff15225", "0x3a5bdd41", "0xab533435", "0xe1c55ad5", "0xe6d3bd0d", "0x9d9ffbbd", "0xbb2a3cf3", "0x50a7c08d"], "rc": ["0xc6892460", "0x25b7228a", "0xcd515004", "0x2846527a", "0xa6324241", "0x36e3ec53", "0x82961bac", "0x0f97ba7d", "0xb6f921a9", "0x3ada24e5", "0xde20ab91", "0x5378eeb2", "0x7d161662", "0x89353cc1", "0xb1aa03a2", "0x788acae6"], "round": "for i in 0..15: s[i] = (s[i] ^ (rc[i] + (r+1) * 0x9E3779B9)) * mul[i]; then quarter rounds on columns (0,4,8,12) (1,5,9,13) (2,6,10,14) (3,7,11,15) with rot[0..3] and diagonals (0,5,10,15) (1,6,11,12) (2,7,8,13) (3,4,9,14) with rot[4..7]", "quarter_round": "a += b; d ^= a; d = rotl(d, r1); c += d; b ^= c; b = rotl(b, r2); a += b; d ^= a; d = rotl(d, r3); c += d; b ^= c; b = rotl(b, r4)"},
"item": "s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..7: s = M_r(s); line = s[0] & 0x003fffff; s[i] ^= cache[line * 16 + i]; then s = M_8(s); item(t) = s",
"word": "dataset[w] = item(w >> 4)[w & 15]"
},
"instructions": [
{"i": 0, "op": "add", "dst": 4, "src": 5, "src2": 7, "imm": "0xea86e152", "imm2": "0x5810667a", "rot": 27, "bit": 13, "mask": 2, "width": 1, "win": 0, "off": 0},
{"i": 1, "op": "xor", "dst": 7, "src": 0, "src2": 6, "imm": "0xbaab6229", "imm2": "0xed861989", "rot": 26, "bit": 22, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 2, "op": "shfl", "dst": 3, "src": 6, "src2": 2, "imm": "0x5b623116", "imm2": "0xff12e5b2", "rot": 12, "bit": 24, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 3, "op": "sub", "dst": 4, "src": 1, "src2": 1, "imm": "0xe99741c7", "imm2": "0xf5fa5009", "rot": 1, "bit": 21, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 4, "op": "mad", "dst": 2, "src": 0, "src2": 4, "imm": "0x673c2157", "imm2": "0xee02465f", "rot": 20, "bit": 22, "mask": 2, "width": 1, "win": 0, "off": 0},
{"i": 5, "op": "rotl", "dst": 4, "src": 7, "src2": 7, "imm": "0x946f7818", "imm2": "0x45d3399e", "rot": 9, "bit": 2, "mask": 16, "width": 1, "win": 0, "off": 0},
{"i": 6, "op": "rotr", "dst": 0, "src": 2, "src2": 4, "imm": "0x5f6a0ed2", "imm2": "0x7043a636", "rot": 19, "bit": 31, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 7, "op": "load", "dst": 6, "src": 7, "src2": 1, "imm": "0x5c61dcf7", "imm2": "0x7466aa40", "rot": 19, "bit": 9, "mask": 2, "width": 1, "win": 2, "off": 1},
{"i": 8, "op": "load", "dst": 1, "src": 4, "src2": 5, "imm": "0x85668475", "imm2": "0xdb8cc483", "rot": 29, "bit": 7, "mask": 4, "width": 1, "win": 1, "off": 1},
{"i": 9, "op": "load", "dst": 1, "src": 2, "src2": 4, "imm": "0x4454980f", "imm2": "0xebd31581", "rot": 10, "bit": 28, "mask": 16, "width": 1, "win": 1, "off": 1},
{"i": 10, "op": "load", "dst": 7, "src": 0, "src2": 2, "imm": "0xe075297c", "imm2": "0x5779c44c", "rot": 10, "bit": 22, "mask": 2, "width": 1, "win": 1, "off": 1},
{"i": 11, "op": "load", "dst": 7, "src": 1, "src2": 6, "imm": "0x65aa4311", "imm2": "0x4fe48ea9", "rot": 15, "bit": 9, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 12, "op": "mul", "dst": 5, "src": 4, "src2": 2, "imm": "0x1383d3ad", "imm2": "0xf3094b29", "rot": 8, "bit": 9, "mask": 2, "width": 1, "win": 0, "off": 0},
{"i": 13, "op": "load", "dst": 7, "src": 6, "src2": 2, "imm": "0xed8a496f", "imm2": "0x3072c3c6", "rot": 28, "bit": 19, "mask": 8, "width": 1, "win": 1, "off": 1},
{"i": 14, "op": "shfl", "dst": 6, "src": 5, "src2": 0, "imm": "0x8b965b57", "imm2": "0xcfeca6c1", "rot": 12, "bit": 27, "mask": 16, "width": 1, "win": 0, "off": 0},
{"i": 15, "op": "shfl", "dst": 3, "src": 5, "src2": 3, "imm": "0x877c7586", "imm2": "0xa9cb2a03", "rot": 2, "bit": 29, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 16, "op": "or", "dst": 2, "src": 7, "src2": 3, "imm": "0xb740221a", "imm2": "0x89d38d6d", "rot": 6, "bit": 31, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 17, "op": "rotr", "dst": 0, "src": 6, "src2": 1, "imm": "0x26f3ad8a", "imm2": "0x27256f15", "rot": 18, "bit": 5, "mask": 2, "width": 1, "win": 0, "off": 0},
{"i": 18, "op": "add", "dst": 7, "src": 5, "src2": 7, "imm": "0xf66e7017", "imm2": "0xb9e3577e", "rot": 9, "bit": 12, "mask": 16, "width": 1, "win": 0, "off": 0},
{"i": 19, "op": "rotl", "dst": 7, "src": 0, "src2": 5, "imm": "0x849ae6ee", "imm2": "0x02b358f9", "rot": 24, "bit": 18, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 20, "op": "add", "dst": 6, "src": 7, "src2": 6, "imm": "0x52334d12", "imm2": "0x8c9f0ef8", "rot": 11, "bit": 23, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 21, "op": "or", "dst": 2, "src": 6, "src2": 5, "imm": "0xb1871e63", "imm2": "0xb2e40191", "rot": 5, "bit": 13, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 22, "op": "mad", "dst": 6, "src": 5, "src2": 4, "imm": "0x97df29e4", "imm2": "0xe60fea84", "rot": 11, "bit": 16, "mask": 2, "width": 1, "win": 0, "off": 0},
{"i": 23, "op": "or", "dst": 1, "src": 0, "src2": 3, "imm": "0x8f30d21d", "imm2": "0x2df685a0", "rot": 31, "bit": 0, "mask": 16, "width": 1, "win": 0, "off": 0},
{"i": 24, "op": "xor", "dst": 6, "src": 1, "src2": 2, "imm": "0xd2c4025f", "imm2": "0x5269eb4d", "rot": 31, "bit": 21, "mask": 16, "width": 1, "win": 0, "off": 0},
{"i": 25, "op": "add", "dst": 2, "src": 6, "src2": 5, "imm": "0x659fc3d3", "imm2": "0x9cec0e12", "rot": 6, "bit": 17, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 26, "op": "rotr", "dst": 7, "src": 0, "src2": 1, "imm": "0xd89ef484", "imm2": "0x20be3846", "rot": 12, "bit": 9, "mask": 16, "width": 1, "win": 0, "off": 0},
{"i": 27, "op": "mad", "dst": 4, "src": 5, "src2": 7, "imm": "0xb48420ae", "imm2": "0x3d1f2485", "rot": 14, "bit": 28, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 28, "op": "mad", "dst": 3, "src": 2, "src2": 4, "imm": "0xc5c46d76", "imm2": "0x700044b5", "rot": 22, "bit": 10, "mask": 2, "width": 1, "win": 0, "off": 0},
{"i": 29, "op": "load", "dst": 1, "src": 4, "src2": 3, "imm": "0x0fbaf177", "imm2": "0xfff4f2ed", "rot": 20, "bit": 31, "mask": 2, "width": 1, "win": 0, "off": 0},
{"i": 30, "op": "load", "dst": 2, "src": 3, "src2": 0, "imm": "0xe90eb2e5", "imm2": "0xb0f9eb79", "rot": 27, "bit": 14, "mask": 4, "width": 1, "win": 2, "off": 2},
{"i": 31, "op": "load", "dst": 1, "src": 5, "src2": 2, "imm": "0x97ba3fc3", "imm2": "0x7894e657", "rot": 3, "bit": 30, "mask": 16, "width": 1, "win": 1, "off": 1},
{"i": 32, "op": "add", "dst": 7, "src": 2, "src2": 7, "imm": "0x070888a8", "imm2": "0xe403240e", "rot": 2, "bit": 16, "mask": 2, "width": 1, "win": 0, "off": 0},
{"i": 33, "op": "add", "dst": 2, "src": 0, "src2": 5, "imm": "0xf2e46d55", "imm2": "0x29701828", "rot": 31, "bit": 28, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 34, "op": "add", "dst": 2, "src": 3, "src2": 0, "imm": "0x58f75b87", "imm2": "0x343b7aee", "rot": 12, "bit": 14, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 35, "op": "mulhi", "dst": 2, "src": 5, "src2": 6, "imm": "0x5892a9e6", "imm2": "0xc9824c94", "rot": 19, "bit": 26, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 36, "op": "xor", "dst": 4, "src": 2, "src2": 3, "imm": "0x7b5b5474", "imm2": "0x45cfc5dd", "rot": 17, "bit": 18, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 37, "op": "mul", "dst": 6, "src": 5, "src2": 7, "imm": "0xccf564a5", "imm2": "0x873ad101", "rot": 7, "bit": 11, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 38, "op": "xor", "dst": 7, "src": 0, "src2": 6, "imm": "0xcac8f06d", "imm2": "0x6b97c683", "rot": 18, "bit": 28, "mask": 2, "width": 1, "win": 0, "off": 0},
{"i": 39, "op": "add", "dst": 7, "src": 2, "src2": 4, "imm": "0xb8180e9d", "imm2": "0x32bbd117", "rot": 23, "bit": 19, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 40, "op": "rotr", "dst": 2, "src": 3, "src2": 0, "imm": "0x2d6070bc", "imm2": "0x68ff101e", "rot": 13, "bit": 18, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 41, "op": "sub", "dst": 7, "src": 0, "src2": 2, "imm": "0x0daf96ea", "imm2": "0x36f37be1", "rot": 5, "bit": 0, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 42, "op": "add", "dst": 4, "src": 3, "src2": 6, "imm": "0x6ced15b7", "imm2": "0x6df7aed4", "rot": 19, "bit": 4, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 43, "op": "shfl", "dst": 7, "src": 3, "src2": 5, "imm": "0x8ace05f3", "imm2": "0xd378ec12", "rot": 23, "bit": 10, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 44, "op": "load", "dst": 0, "src": 7, "src2": 4, "imm": "0xb0607786", "imm2": "0xc4acabbc", "rot": 13, "bit": 7, "mask": 16, "width": 1, "win": 1, "off": 1},
{"i": 45, "op": "add", "dst": 1, "src": 6, "src2": 5, "imm": "0xe09f54e9", "imm2": "0x83e825bf", "rot": 23, "bit": 14, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 46, "op": "load", "dst": 3, "src": 1, "src2": 0, "imm": "0x63cc1e4e", "imm2": "0xa1be8118", "rot": 12, "bit": 6, "mask": 2, "width": 1, "win": 2, "off": 0},
{"i": 47, "op": "load", "dst": 6, "src": 3, "src2": 7, "imm": "0x353f1d79", "imm2": "0x3b2e7456", "rot": 18, "bit": 18, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 48, "op": "mulhi", "dst": 4, "src": 2, "src2": 7, "imm": "0x00d8a3cd", "imm2": "0x231866d2", "rot": 21, "bit": 20, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 49, "op": "add", "dst": 5, "src": 0, "src2": 2, "imm": "0xa8bae6df", "imm2": "0xf572bdb9", "rot": 14, "bit": 7, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 50, "op": "shfl", "dst": 0, "src": 7, "src2": 7, "imm": "0x81ef22e1", "imm2": "0x74438fc5", "rot": 28, "bit": 18, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 51, "op": "add", "dst": 6, "src": 0, "src2": 6, "imm": "0x383b9260", "imm2": "0x11e17c61", "rot": 12, "bit": 19, "mask": 16, "width": 1, "win": 0, "off": 0},
{"i": 52, "op": "load", "dst": 5, "src": 2, "src2": 2, "imm": "0xfb84f451", "imm2": "0x11cd863e", "rot": 21, "bit": 20, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 53, "op": "rotl", "dst": 6, "src": 5, "src2": 4, "imm": "0xb1a7db6b", "imm2": "0x76686b9b", "rot": 12, "bit": 4, "mask": 16, "width": 1, "win": 0, "off": 0},
{"i": 54, "op": "rotl", "dst": 3, "src": 6, "src2": 3, "imm": "0x6f981f52", "imm2": "0xd99aeba2", "rot": 12, "bit": 27, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 55, "op": "add", "dst": 2, "src": 1, "src2": 2, "imm": "0xac6be8e3", "imm2": "0x18d67dbb", "rot": 26, "bit": 10, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 56, "op": "load", "dst": 1, "src": 4, "src2": 0, "imm": "0x7e7f6a00", "imm2": "0x6f0747da", "rot": 25, "bit": 28, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 57, "op": "add", "dst": 5, "src": 0, "src2": 4, "imm": "0xf03673fe", "imm2": "0xa75cd60d", "rot": 16, "bit": 12, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 58, "op": "load", "dst": 5, "src": 0, "src2": 2, "imm": "0x227f94a6", "imm2": "0x0e8344f9", "rot": 20, "bit": 10, "mask": 2, "width": 1, "win": 2, "off": 0},
{"i": 59, "op": "add", "dst": 1, "src": 4, "src2": 3, "imm": "0xdecd4794", "imm2": "0x8dfb96bb", "rot": 21, "bit": 7, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 60, "op": "mulhi", "dst": 3, "src": 2, "src2": 2, "imm": "0x0dd268e0", "imm2": "0x53034ca9", "rot": 1, "bit": 8, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 61, "op": "or", "dst": 6, "src": 4, "src2": 7, "imm": "0x3a45a321", "imm2": "0x9bc59a5f", "rot": 25, "bit": 11, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 62, "op": "shfl", "dst": 5, "src": 4, "src2": 2, "imm": "0x8f229cc1", "imm2": "0xcaac64a2", "rot": 17, "bit": 13, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 63, "op": "load", "dst": 3, "src": 6, "src2": 6, "imm": "0xb2574178", "imm2": "0xcbcc798d", "rot": 28, "bit": 0, "mask": 16, "width": 1, "win": 1, "off": 1}
]
}

View file

@ -0,0 +1,109 @@
#include <metal_stdlib>
using namespace metal;
#define MASK 0x0fffffffu
constant uint SEEDW[8] = { 0x667d0fbdu, 0x7b8e5963u, 0x31c67e5eu, 0x4529ddc6u, 0xef19d6d8u, 0xaccf6211u, 0xda0aed32u, 0xabc6df31u };
inline uint splitmix32(uint x) {
x ^= x >> 16; x *= 0x7feb352du;
x ^= x >> 15; x *= 0x846ca68bu;
x ^= x >> 16;
return x;
}
inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31
inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
inline uint ds_elem(uint i, uint d0, uint d1) {
uint x = i ^ d0;
x *= 0x9E3779B1u; x ^= x >> 15;
x += d1;
x *= 0x85EBCA77u; x ^= x >> 13;
x *= 0xC2B2AE3Du; x ^= x >> 16;
return x;
}
kernel void igneum_hash(device const uint* dataset [[buffer(0)]],
device ulong* out [[buffer(1)]],
constant uint& baseNonce [[buffer(2)]],
uint gid [[thread_position_in_grid]]) {
uint nonce = baseNonce + gid;
uint r0, r1, r2, r3, r4, r5, r6, r7;
{ uint x = nonce ^ SEEDW[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ SEEDW[1]; }
{ uint x = nonce ^ SEEDW[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ SEEDW[2]; }
{ uint x = nonce ^ SEEDW[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ SEEDW[3]; }
{ uint x = nonce ^ SEEDW[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ SEEDW[4]; }
{ uint x = nonce ^ SEEDW[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ SEEDW[5]; }
{ uint x = nonce ^ SEEDW[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ SEEDW[6]; }
{ uint x = nonce ^ SEEDW[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ SEEDW[7]; }
{ uint x = nonce ^ SEEDW[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ SEEDW[0]; }
for (uint it = 0u; it < 8u; ++it) {
uint sel = r0;
r4 = r4 + r5 + select(0xea86e152u, 0x5810667au, ((sel >> 13u) & 1u) != 0u); // 0
r7 = r7 ^ r0; // 1
r3 = r3 ^ simd_shuffle_xor(r6, (ushort)1); // 2
r4 = r4 - r1; // 3
r2 = r0 * r4 + r2; // 4
r4 = rotl_imm(r4, 9u); // 5
r0 = rotr_var(r0, r2); // 6
r6 = r6 ^ dataset[((rotl_imm(r7 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x04000000u) & MASK]; // 7
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 8
r1 = r1 ^ dataset[((rotl_imm(r2 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 9
r7 = r7 ^ dataset[((rotl_imm(r0 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 10
r7 = r7 ^ dataset[((rotl_imm(r1 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 11
r5 = r5 * r4; // 12
r7 = r7 ^ dataset[((rotl_imm(r6 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 13
r6 = r6 ^ simd_shuffle_xor(r5, (ushort)16); // 14
r3 = r3 ^ simd_shuffle_xor(r5, (ushort)1); // 15
r2 = r2 | r7; // 16
r0 = rotr_var(r0, r6); // 17
r7 = r7 + r5 + select(0xf66e7017u, 0xb9e3577eu, ((sel >> 12u) & 1u) != 0u); // 18
r7 = rotl_imm(r7, 24u); // 19
r6 = r6 + r7 + select(0x52334d12u, 0x8c9f0ef8u, ((sel >> 23u) & 1u) != 0u); // 20
r2 = r2 | r6; // 21
r6 = r5 * r4 + r6; // 22
r1 = r1 | r0; // 23
r6 = r6 ^ r1; // 24
r2 = r2 + r6 + select(0x659fc3d3u, 0x9cec0e12u, ((sel >> 17u) & 1u) != 0u); // 25
r7 = rotr_var(r7, r0); // 26
r4 = r5 * r7 + r4; // 27
r3 = r2 * r4 + r3; // 28
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 29
r2 = r2 ^ dataset[((rotl_imm(r3 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x08000000u) & MASK]; // 30
r1 = r1 ^ dataset[((rotl_imm(r5 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 31
r7 = r7 + r2 + select(0x070888a8u, 0xe403240eu, ((sel >> 16u) & 1u) != 0u); // 32
r2 = r2 + r0 + select(0xf2e46d55u, 0x29701828u, ((sel >> 28u) & 1u) != 0u); // 33
r2 = r2 + r3 + select(0x58f75b87u, 0x343b7aeeu, ((sel >> 14u) & 1u) != 0u); // 34
r2 = mulhi(r2, r5); // 35
r4 = r4 ^ r2; // 36
r6 = r6 * r5; // 37
r7 = r7 ^ r0; // 38
r7 = r7 + r2 + select(0xb8180e9du, 0x32bbd117u, ((sel >> 19u) & 1u) != 0u); // 39
r2 = rotr_var(r2, r3); // 40
r7 = r7 - r0; // 41
r4 = r4 + r3 + select(0x6ced15b7u, 0x6df7aed4u, ((sel >> 4u) & 1u) != 0u); // 42
r7 = r7 ^ simd_shuffle_xor(r3, (ushort)4); // 43
r0 = r0 ^ dataset[((rotl_imm(r7 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 44
r1 = r1 + r6 + select(0xe09f54e9u, 0x83e825bfu, ((sel >> 14u) & 1u) != 0u); // 45
r3 = r3 ^ dataset[((rotl_imm(r1 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 46
r6 = r6 ^ dataset[((rotl_imm(r3 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 47
r4 = mulhi(r4, r2); // 48
r5 = r5 + r0 + select(0xa8bae6dfu, 0xf572bdb9u, ((sel >> 7u) & 1u) != 0u); // 49
r0 = r0 ^ simd_shuffle_xor(r7, (ushort)4); // 50
r6 = r6 + r0 + select(0x383b9260u, 0x11e17c61u, ((sel >> 19u) & 1u) != 0u); // 51
r5 = r5 ^ dataset[((rotl_imm(r2 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 52
r6 = rotl_imm(r6, 12u); // 53
r3 = rotl_imm(r3, 12u); // 54
r2 = r2 + r1 + select(0xac6be8e3u, 0x18d67dbbu, ((sel >> 10u) & 1u) != 0u); // 55
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 56
r5 = r5 + r0 + select(0xf03673feu, 0xa75cd60du, ((sel >> 12u) & 1u) != 0u); // 57
r5 = r5 ^ dataset[((rotl_imm(r0 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 58
r1 = r1 + r4 + select(0xdecd4794u, 0x8dfb96bbu, ((sel >> 7u) & 1u) != 0u); // 59
r3 = mulhi(r3, r2); // 60
r6 = r6 | r4; // 61
r5 = r5 ^ simd_shuffle_xor(r4, (ushort)8); // 62
r3 = r3 ^ dataset[((rotl_imm(r6 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 63
}
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((ulong)hi << 32) | (ulong)lo;
}

View file

@ -0,0 +1,111 @@
#include <metal_stdlib>
using namespace metal;
#define MASK 0x0fffffffu
constant uint SEEDW[8] = { 0x667d0fbdu, 0x7b8e5963u, 0x31c67e5eu, 0x4529ddc6u, 0xef19d6d8u, 0xaccf6211u, 0xda0aed32u, 0xabc6df31u };
inline uint splitmix32(uint x) {
x ^= x >> 16; x *= 0x7feb352du;
x ^= x >> 15; x *= 0x846ca68bu;
x ^= x >> 16;
return x;
}
inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31
inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
inline uint ds_elem(uint i, uint d0, uint d1) {
uint x = i ^ d0;
x *= 0x9E3779B1u; x ^= x >> 15;
x += d1;
x *= 0x85EBCA77u; x ^= x >> 13;
x *= 0xC2B2AE3Du; x ^= x >> 16;
return x;
}
// Header-bound variant: the init words come from buffer 3 (bind.rs), not from SEEDW.
kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]],
device ulong* out [[buffer(1)]],
constant uint& baseNonce [[buffer(2)]],
constant uint* initw [[buffer(3)]],
uint gid [[thread_position_in_grid]]) {
uint nonce = baseNonce + gid;
uint r0, r1, r2, r3, r4, r5, r6, r7;
{ uint x = nonce ^ initw[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ initw[1]; }
{ uint x = nonce ^ initw[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ initw[2]; }
{ uint x = nonce ^ initw[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ initw[3]; }
{ uint x = nonce ^ initw[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ initw[4]; }
{ uint x = nonce ^ initw[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ initw[5]; }
{ uint x = nonce ^ initw[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ initw[6]; }
{ uint x = nonce ^ initw[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ initw[7]; }
{ uint x = nonce ^ initw[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ initw[0]; }
for (uint it = 0u; it < 8u; ++it) {
uint sel = r0;
r4 = r4 + r5 + select(0xea86e152u, 0x5810667au, ((sel >> 13u) & 1u) != 0u); // 0
r7 = r7 ^ r0; // 1
r3 = r3 ^ simd_shuffle_xor(r6, (ushort)1); // 2
r4 = r4 - r1; // 3
r2 = r0 * r4 + r2; // 4
r4 = rotl_imm(r4, 9u); // 5
r0 = rotr_var(r0, r2); // 6
r6 = r6 ^ dataset[((rotl_imm(r7 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x04000000u) & MASK]; // 7
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 8
r1 = r1 ^ dataset[((rotl_imm(r2 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 9
r7 = r7 ^ dataset[((rotl_imm(r0 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 10
r7 = r7 ^ dataset[((rotl_imm(r1 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 11
r5 = r5 * r4; // 12
r7 = r7 ^ dataset[((rotl_imm(r6 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 13
r6 = r6 ^ simd_shuffle_xor(r5, (ushort)16); // 14
r3 = r3 ^ simd_shuffle_xor(r5, (ushort)1); // 15
r2 = r2 | r7; // 16
r0 = rotr_var(r0, r6); // 17
r7 = r7 + r5 + select(0xf66e7017u, 0xb9e3577eu, ((sel >> 12u) & 1u) != 0u); // 18
r7 = rotl_imm(r7, 24u); // 19
r6 = r6 + r7 + select(0x52334d12u, 0x8c9f0ef8u, ((sel >> 23u) & 1u) != 0u); // 20
r2 = r2 | r6; // 21
r6 = r5 * r4 + r6; // 22
r1 = r1 | r0; // 23
r6 = r6 ^ r1; // 24
r2 = r2 + r6 + select(0x659fc3d3u, 0x9cec0e12u, ((sel >> 17u) & 1u) != 0u); // 25
r7 = rotr_var(r7, r0); // 26
r4 = r5 * r7 + r4; // 27
r3 = r2 * r4 + r3; // 28
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 29
r2 = r2 ^ dataset[((rotl_imm(r3 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x08000000u) & MASK]; // 30
r1 = r1 ^ dataset[((rotl_imm(r5 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 31
r7 = r7 + r2 + select(0x070888a8u, 0xe403240eu, ((sel >> 16u) & 1u) != 0u); // 32
r2 = r2 + r0 + select(0xf2e46d55u, 0x29701828u, ((sel >> 28u) & 1u) != 0u); // 33
r2 = r2 + r3 + select(0x58f75b87u, 0x343b7aeeu, ((sel >> 14u) & 1u) != 0u); // 34
r2 = mulhi(r2, r5); // 35
r4 = r4 ^ r2; // 36
r6 = r6 * r5; // 37
r7 = r7 ^ r0; // 38
r7 = r7 + r2 + select(0xb8180e9du, 0x32bbd117u, ((sel >> 19u) & 1u) != 0u); // 39
r2 = rotr_var(r2, r3); // 40
r7 = r7 - r0; // 41
r4 = r4 + r3 + select(0x6ced15b7u, 0x6df7aed4u, ((sel >> 4u) & 1u) != 0u); // 42
r7 = r7 ^ simd_shuffle_xor(r3, (ushort)4); // 43
r0 = r0 ^ dataset[((rotl_imm(r7 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 44
r1 = r1 + r6 + select(0xe09f54e9u, 0x83e825bfu, ((sel >> 14u) & 1u) != 0u); // 45
r3 = r3 ^ dataset[((rotl_imm(r1 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 46
r6 = r6 ^ dataset[((rotl_imm(r3 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 47
r4 = mulhi(r4, r2); // 48
r5 = r5 + r0 + select(0xa8bae6dfu, 0xf572bdb9u, ((sel >> 7u) & 1u) != 0u); // 49
r0 = r0 ^ simd_shuffle_xor(r7, (ushort)4); // 50
r6 = r6 + r0 + select(0x383b9260u, 0x11e17c61u, ((sel >> 19u) & 1u) != 0u); // 51
r5 = r5 ^ dataset[((rotl_imm(r2 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 52
r6 = rotl_imm(r6, 12u); // 53
r3 = rotl_imm(r3, 12u); // 54
r2 = r2 + r1 + select(0xac6be8e3u, 0x18d67dbbu, ((sel >> 10u) & 1u) != 0u); // 55
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 56
r5 = r5 + r0 + select(0xf03673feu, 0xa75cd60du, ((sel >> 12u) & 1u) != 0u); // 57
r5 = r5 ^ dataset[((rotl_imm(r0 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 58
r1 = r1 + r4 + select(0xdecd4794u, 0x8dfb96bbu, ((sel >> 7u) & 1u) != 0u); // 59
r3 = mulhi(r3, r2); // 60
r6 = r6 | r4; // 61
r5 = r5 ^ simd_shuffle_xor(r4, (ushort)8); // 62
r3 = r3 ^ dataset[((rotl_imm(r6 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 63
}
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((ulong)hi << 32) | (ulong)lo;
}

View file

@ -0,0 +1,5 @@
epoch_seed_hex edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07
day_seed_hex 69676e65756d2d6461792ffa50000000000000
epoch_index 0
day_index 20730
daa_score 0

View file

@ -0,0 +1,57 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v3, memory-hard dataset
#pragma once
#ifdef __cplusplus
#include <cstdint>
#else
#include <stdint.h>
#endif
#define IGNEUM_VEC_WARPS 3
static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u };
static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = {
{ // base nonce 0
0x482d9ca50850d913ull, 0x4ad212a8109dbf80ull, 0xfc22f1d2859b72dbull, 0x1b122111c880daddull, 0x590c0ed07d26f3beull, 0x14f9ccb1fe4a0aa0ull, 0xb9aa46ddbf3358f4ull, 0x7a2e7a0758a139dcull,
0x4ca0c6f57fc6a6f8ull, 0xa9303e49a53c55cfull, 0xd0eef090a205c2daull, 0xf916463a1fe71f3dull, 0xece100319c1ba260ull, 0x0d765a753bab6621ull, 0x7a8b38b0cd86e711ull, 0x6330fa995f69eac6ull,
0x31630dc8222e4cf3ull, 0x1ab386a5c4bfd048ull, 0xd1849ab5d7ccc035ull, 0x7e42d0b071e213b7ull, 0xc61d482f91f73a00ull, 0xb759a48c18732abcull, 0x7cb720597cea45f4ull, 0x9aa782a7d7153267ull,
0x3928c8ca30daa819ull, 0xf6624ec425bf44b6ull, 0xb92b06f50f5fca62ull, 0xb7ac77af13b44e09ull, 0x9bd4a23ed9a8d186ull, 0xed157fa7db51abdfull, 0x169dae4f9708b9b0ull, 0x3928f1c5bb85b8ffull
},
{ // base nonce 4096
0x68d28d84a9cacd37ull, 0x3de7667c504d33b1ull, 0x9cd9cbd429ecccbfull, 0xce28296aba64ec41ull, 0x7f25bfb162361f07ull, 0x305c8fcecda70a88ull, 0x4d26391bf3d1c1f0ull, 0x230a57b46fde607full,
0x35482e4e9ae6e874ull, 0x87f3bdd2250cb21eull, 0x71a6a482132c6938ull, 0x3c853fb23317a129ull, 0x5932d928549c095aull, 0xe90460bc189cdb34ull, 0x6dbe852db160461aull, 0x2abd73567e773ab6ull,
0x97be9fa85c3805ffull, 0xf7d138416c3ac1ccull, 0x888b84f204d9305eull, 0x92800629cdb2d056ull, 0x2cb98fc7e6f8b22bull, 0x085ba0e520fa1df6ull, 0x9b0ee619aa074133ull, 0xd92649b972786f58ull,
0x94e56fdfade518caull, 0xcae28f9eec9d5be5ull, 0x2e715314eaa160b6ull, 0xd3faf570104ca532ull, 0xeeecb2834a0e878bull, 0x9e978602b88d0a38ull, 0x3536c9dd7062e398ull, 0x5974adef999d72a2ull
},
{ // base nonce 1000000
0x86c2830118bf5d18ull, 0x0ca010e1f65efea9ull, 0x0e065db2ff5d7177ull, 0xb6c5de9453893f86ull, 0x4b2bca8ca7ff4dacull, 0x875e9e4b09262995ull, 0xfe1c3a11b0dc75abull, 0x9f83ac0ba031e6dfull,
0x5c4e3154cdf0d3beull, 0x812fd5f5be4159e6ull, 0xcbf87d1772e7a740ull, 0x251c1bb3b31e0de1ull, 0xba117c2b253c2d04ull, 0xf42a27fffa765e8aull, 0xb654d4c021ced921ull, 0x63fe5bc570bd9a6dull,
0xe3729cacfa2738b9ull, 0x0686bf7f029dd650ull, 0xa4da20856fbb382eull, 0xcacd5c3aa30da86cull, 0x7b47735366038a7aull, 0xe0b5603c8652702eull, 0x62b834fee790c2c5ull, 0x4c979e00c086c371ull,
0x9a94911a7eaa32d6ull, 0x377a5edc37bfb69full, 0xecebe1aab4f07982ull, 0xdd2bba556d4d8996ull, 0xb6e0be464e82691cull, 0xe4cc7f9840d754d9ull, 0x8e31fb423c108e46ull, 0x8164e7242a952c96ull
}
};
// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455).
static const uint32_t IGNEUM_DS_HEAD[16] = {
0x3dd50b1fu, 0x93123701u, 0xcdd88b45u, 0x2c37c6b0u, 0x48edec90u, 0x3a7d2407u, 0xb03a0ac0u, 0x10e2c6a5u,
0x5119540eu, 0x3358ef2du, 0x25e2929cu, 0x5fa0cdccu, 0x657ac748u, 0x09e2105cu, 0x4317cabdu, 0x6178ee52u
};
static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u;
static const uint32_t IGNEUM_DS_LAST = 0xf7b7180eu;
// 64 sampled dataset words (index, value) computed on the Mac.
#define IGNEUM_DS_SAMPLES 64
static const uint32_t IGNEUM_DS_SAMPLE_INDEX[IGNEUM_DS_SAMPLES] = {
59471966u, 217795994u, 208353206u, 42483309u, 172547758u, 148076330u, 183853158u, 214389424u, 267488061u, 169781097u, 184093494u, 153880993u, 84977930u, 46426879u, 3093825u, 225364072u, 44593546u, 260713159u, 168250303u, 52384140u, 223401610u, 45554030u, 95410555u, 175039924u, 79171087u, 267580473u, 24168642u, 37981670u, 171551130u, 195559979u, 204611762u, 140997658u, 138925853u, 86637313u, 20736778u, 219665210u, 160430336u, 264654675u, 8013395u, 228945585u, 213884386u, 104419827u, 44185464u, 142737231u, 99284897u, 132475900u, 61861762u, 132056166u, 262388043u, 91878046u, 117353561u, 124768597u, 71352993u, 190698941u, 46055428u, 55281366u, 165145231u, 106810753u, 171985651u, 232085256u, 159510492u, 40072060u, 209107596u, 39023794u
};
static const uint32_t IGNEUM_DS_SAMPLE_VALUE[IGNEUM_DS_SAMPLES] = {
0xf6e00c89u, 0xe206f72bu, 0x09260b8cu, 0x662c6558u, 0x4028c489u, 0xbccb5ab0u, 0x88bf905cu, 0x88f76874u, 0xea529b56u, 0x88e89b31u, 0x22deecc6u, 0xd6022e44u, 0x7a7923c9u, 0x8c761ec7u, 0x7079d93du, 0x79341b4cu, 0x560640c9u, 0xc659d408u, 0xd5883003u, 0x95e70092u, 0xbb0552dbu, 0xbbec95bfu, 0x1efacd50u, 0x8d4c6036u, 0xcbcb1898u, 0x05d9a8aeu, 0x0e5af59bu, 0x12094029u, 0xd6132e60u, 0xffc32f3eu, 0x3bae4991u, 0xbf6d57a0u, 0x7f03491eu, 0x26c9e599u, 0xa0c9ecc9u, 0xae96f173u, 0x1f437e50u, 0x1f641c65u, 0x30d7028au, 0x65092ed9u, 0xe390c42bu, 0x8f89bb6bu, 0xb4205b97u, 0xd79db283u, 0x65869dcfu, 0x4ac02d55u, 0x5f1b8dc3u, 0xd192a75au, 0xcca97073u, 0x2f043a85u, 0xdc0c3fddu, 0x19b12a30u, 0xf5da472eu, 0xbf7dd652u, 0xec1dc918u, 0x2f00a5dcu, 0xbe7f358eu, 0x68a25926u, 0x817ae6c9u, 0xa731ea98u, 0x6af869dau, 0xb84e7e41u, 0xd3519f00u, 0x3e6c426au
};
// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^26 words.
static const uint32_t IGNEUM_CACHE_HEAD[16] = {
0xebd9055cu, 0x7eaf6b21u, 0x9b610855u, 0x134a8562u, 0xb10059bau, 0x7f459e58u, 0x8f3c7873u, 0x3523e609u,
0x75b8f4cau, 0x750d31b4u, 0x0dae0782u, 0x26f10015u, 0x7ba87a41u, 0x0efd8543u, 0x691d6368u, 0x8929d967u
};
static const uint32_t IGNEUM_CACHE_LAST[16] = {
0x51474edfu, 0xc3cfba93u, 0xf21454cfu, 0x9b79baacu, 0x4d4fce23u, 0xd701dfd5u, 0x37357ab3u, 0x1be693fau,
0xa701cb3bu, 0x7467c620u, 0x428184e8u, 0xf4010df0u, 0xe33aa88fu, 0xc78d6d62u, 0xae9ba9c0u, 0xf4bba97eu
};
static const uint64_t IGNEUM_CACHE_FNV64 = 0x448274a57f508cbcull;

View file

@ -0,0 +1,36 @@
{
"seed": "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000",
"day": "bytes:69676e65756d2d6461792ffa50000000000000",
"dataset_mode": "memory-hard",
"dataset_log2_words": 28,
"mask": "0x0fffffff",
"lanes": 32,
"source": "igneum-pow (Rust) CPU interpreter, generator v3, memory-hard dataset",
"warps": [
{"base_nonce": 0, "expected": [
"0x482d9ca50850d913", "0x4ad212a8109dbf80", "0xfc22f1d2859b72db", "0x1b122111c880dadd", "0x590c0ed07d26f3be", "0x14f9ccb1fe4a0aa0", "0xb9aa46ddbf3358f4", "0x7a2e7a0758a139dc",
"0x4ca0c6f57fc6a6f8", "0xa9303e49a53c55cf", "0xd0eef090a205c2da", "0xf916463a1fe71f3d", "0xece100319c1ba260", "0x0d765a753bab6621", "0x7a8b38b0cd86e711", "0x6330fa995f69eac6",
"0x31630dc8222e4cf3", "0x1ab386a5c4bfd048", "0xd1849ab5d7ccc035", "0x7e42d0b071e213b7", "0xc61d482f91f73a00", "0xb759a48c18732abc", "0x7cb720597cea45f4", "0x9aa782a7d7153267",
"0x3928c8ca30daa819", "0xf6624ec425bf44b6", "0xb92b06f50f5fca62", "0xb7ac77af13b44e09", "0x9bd4a23ed9a8d186", "0xed157fa7db51abdf", "0x169dae4f9708b9b0", "0x3928f1c5bb85b8ff"
]},
{"base_nonce": 4096, "expected": [
"0x68d28d84a9cacd37", "0x3de7667c504d33b1", "0x9cd9cbd429ecccbf", "0xce28296aba64ec41", "0x7f25bfb162361f07", "0x305c8fcecda70a88", "0x4d26391bf3d1c1f0", "0x230a57b46fde607f",
"0x35482e4e9ae6e874", "0x87f3bdd2250cb21e", "0x71a6a482132c6938", "0x3c853fb23317a129", "0x5932d928549c095a", "0xe90460bc189cdb34", "0x6dbe852db160461a", "0x2abd73567e773ab6",
"0x97be9fa85c3805ff", "0xf7d138416c3ac1cc", "0x888b84f204d9305e", "0x92800629cdb2d056", "0x2cb98fc7e6f8b22b", "0x085ba0e520fa1df6", "0x9b0ee619aa074133", "0xd92649b972786f58",
"0x94e56fdfade518ca", "0xcae28f9eec9d5be5", "0x2e715314eaa160b6", "0xd3faf570104ca532", "0xeeecb2834a0e878b", "0x9e978602b88d0a38", "0x3536c9dd7062e398", "0x5974adef999d72a2"
]},
{"base_nonce": 1000000, "expected": [
"0x86c2830118bf5d18", "0x0ca010e1f65efea9", "0x0e065db2ff5d7177", "0xb6c5de9453893f86", "0x4b2bca8ca7ff4dac", "0x875e9e4b09262995", "0xfe1c3a11b0dc75ab", "0x9f83ac0ba031e6df",
"0x5c4e3154cdf0d3be", "0x812fd5f5be4159e6", "0xcbf87d1772e7a740", "0x251c1bb3b31e0de1", "0xba117c2b253c2d04", "0xf42a27fffa765e8a", "0xb654d4c021ced921", "0x63fe5bc570bd9a6d",
"0xe3729cacfa2738b9", "0x0686bf7f029dd650", "0xa4da20856fbb382e", "0xcacd5c3aa30da86c", "0x7b47735366038a7a", "0xe0b5603c8652702e", "0x62b834fee790c2c5", "0x4c979e00c086c371",
"0x9a94911a7eaa32d6", "0x377a5edc37bfb69f", "0xecebe1aab4f07982", "0xdd2bba556d4d8996", "0xb6e0be464e82691c", "0xe4cc7f9840d754d9", "0x8e31fb423c108e46", "0x8164e7242a952c96"
]}
],
"dataset_head": ["0x3dd50b1f", "0x93123701", "0xcdd88b45", "0x2c37c6b0", "0x48edec90", "0x3a7d2407", "0xb03a0ac0", "0x10e2c6a5", "0x5119540e", "0x3358ef2d", "0x25e2929c", "0x5fa0cdcc", "0x657ac748", "0x09e2105c", "0x4317cabd", "0x6178ee52"],
"dataset_last_index": 268435455,
"dataset_last": "0xf7b7180e",
"dataset_samples": [{"index": 59471966, "value": "0xf6e00c89"}, {"index": 217795994, "value": "0xe206f72b"}, {"index": 208353206, "value": "0x09260b8c"}, {"index": 42483309, "value": "0x662c6558"}, {"index": 172547758, "value": "0x4028c489"}, {"index": 148076330, "value": "0xbccb5ab0"}, {"index": 183853158, "value": "0x88bf905c"}, {"index": 214389424, "value": "0x88f76874"}, {"index": 267488061, "value": "0xea529b56"}, {"index": 169781097, "value": "0x88e89b31"}, {"index": 184093494, "value": "0x22deecc6"}, {"index": 153880993, "value": "0xd6022e44"}, {"index": 84977930, "value": "0x7a7923c9"}, {"index": 46426879, "value": "0x8c761ec7"}, {"index": 3093825, "value": "0x7079d93d"}, {"index": 225364072, "value": "0x79341b4c"}, {"index": 44593546, "value": "0x560640c9"}, {"index": 260713159, "value": "0xc659d408"}, {"index": 168250303, "value": "0xd5883003"}, {"index": 52384140, "value": "0x95e70092"}, {"index": 223401610, "value": "0xbb0552db"}, {"index": 45554030, "value": "0xbbec95bf"}, {"index": 95410555, "value": "0x1efacd50"}, {"index": 175039924, "value": "0x8d4c6036"}, {"index": 79171087, "value": "0xcbcb1898"}, {"index": 267580473, "value": "0x05d9a8ae"}, {"index": 24168642, "value": "0x0e5af59b"}, {"index": 37981670, "value": "0x12094029"}, {"index": 171551130, "value": "0xd6132e60"}, {"index": 195559979, "value": "0xffc32f3e"}, {"index": 204611762, "value": "0x3bae4991"}, {"index": 140997658, "value": "0xbf6d57a0"}, {"index": 138925853, "value": "0x7f03491e"}, {"index": 86637313, "value": "0x26c9e599"}, {"index": 20736778, "value": "0xa0c9ecc9"}, {"index": 219665210, "value": "0xae96f173"}, {"index": 160430336, "value": "0x1f437e50"}, {"index": 264654675, "value": "0x1f641c65"}, {"index": 8013395, "value": "0x30d7028a"}, {"index": 228945585, "value": "0x65092ed9"}, {"index": 213884386, "value": "0xe390c42b"}, {"index": 104419827, "value": "0x8f89bb6b"}, {"index": 44185464, "value": "0xb4205b97"}, {"index": 142737231, "value": "0xd79db283"}, {"index": 99284897, "value": "0x65869dcf"}, {"index": 132475900, "value": "0x4ac02d55"}, {"index": 61861762, "value": "0x5f1b8dc3"}, {"index": 132056166, "value": "0xd192a75a"}, {"index": 262388043, "value": "0xcca97073"}, {"index": 91878046, "value": "0x2f043a85"}, {"index": 117353561, "value": "0xdc0c3fdd"}, {"index": 124768597, "value": "0x19b12a30"}, {"index": 71352993, "value": "0xf5da472e"}, {"index": 190698941, "value": "0xbf7dd652"}, {"index": 46055428, "value": "0xec1dc918"}, {"index": 55281366, "value": "0x2f00a5dc"}, {"index": 165145231, "value": "0xbe7f358e"}, {"index": 106810753, "value": "0x68a25926"}, {"index": 171985651, "value": "0x817ae6c9"}, {"index": 232085256, "value": "0xa731ea98"}, {"index": 159510492, "value": "0x6af869da"}, {"index": 40072060, "value": "0xb84e7e41"}, {"index": 209107596, "value": "0xd3519f00"}, {"index": 39023794, "value": "0x3e6c426a"}],
"cache_head": ["0xebd9055c", "0x7eaf6b21", "0x9b610855", "0x134a8562", "0xb10059ba", "0x7f459e58", "0x8f3c7873", "0x3523e609", "0x75b8f4ca", "0x750d31b4", "0x0dae0782", "0x26f10015", "0x7ba87a41", "0x0efd8543", "0x691d6368", "0x8929d967"],
"cache_last_line": ["0x51474edf", "0xc3cfba93", "0xf21454cf", "0x9b79baac", "0x4d4fce23", "0xd701dfd5", "0x37357ab3", "0x1be693fa", "0xa701cb3b", "0x7467c620", "0x428184e8", "0xf4010df0", "0xe33aa88f", "0xc78d6d62", "0xae9ba9c0", "0xf4bba97e"],
"cache_fnv1a64": "0x448274a57f508cbc"
}

View file

@ -0,0 +1,281 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal).
// Built from source at runtime by proto-opencl/host.c, which passes these defines:
// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit)
// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default)
// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32
// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition
// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units;
// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32.
#ifndef IGNEUM_GROUP
#define IGNEUM_GROUP 32
#endif
#ifndef IGNEUM_EXCHANGE
#define IGNEUM_EXCHANGE 0
#endif
#ifdef __OPENCL_VERSION__
#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1)))
#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n]
#if IGNEUM_EXCHANGE == 1
#ifdef cl_khr_subgroups
#pragma OPENCL EXTENSION cl_khr_subgroups : enable
#endif
#ifdef cl_khr_subgroup_shuffle
#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable
#endif
#elif IGNEUM_EXCHANGE == 2
#pragma OPENCL EXTENSION cl_intel_subgroups : enable
#endif
#else
// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros.
#include "emu_opencl.h"
#endif
#if IGNEUM_EXCHANGE == 1
#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m))
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
#elif IGNEUM_EXCHANGE == 2
#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m))
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
#else
// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per
// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane
// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's
// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier.
#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; }
#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; }
#endif
static inline uint splitmix32(uint x) {
x ^= x >> 16; x *= 0x7feb352du;
x ^= x >> 15; x *= 0x846ca68bu;
x ^= x >> 16;
return x;
}
// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32.
static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); }
// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x.
static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); }
static inline uint ds_elem(uint i, uint d0, uint d1) {
uint x = i ^ d0;
x *= 0x9E3779B1u; x ^= x >> 15;
x += d1;
x *= 0x85EBCA77u; x ^= x >> 13;
x *= 0xC2B2AE3Du; x ^= x >> 16;
return x;
}
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
#define MH_CACHE_LINE_MASK 0x003fffffu
#define MH_SEGMENT_LINES 64u
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
// y = ChaCha12 core(x) + x
static inline void mh_chacha_block(const uint* x, uint* y) {
for (uint i = 0u; i < 16u; ++i) y[i] = x[i];
for (uint r = 0u; r < 6u; ++r) {
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
}
for (uint i = 0u; i < 16u; ++i) y[i] += x[i];
}
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
static inline void mh_cache_segment(__global uint* cache, uint seg) {
uint prev[16]; uint x[16]; uint y[16];
for (uint i = 0u; i < 16u; ++i) prev[i] = 0u;
for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) {
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
x[4] = 0xceed56d7u ^ prev[4];
x[5] = 0x9ba270d2u ^ prev[5];
x[6] = 0x82caab2du ^ prev[6];
x[7] = 0x81ebce0eu ^ prev[7];
x[8] = 0x12b6ecf1u ^ prev[8];
x[9] = 0xd0f3fd7cu ^ prev[9];
x[10] = 0xd872eefeu ^ prev[10];
x[11] = 0xc158c7bdu ^ prev[11];
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
mh_chacha_block(x, y);
__global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
}
}
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
static inline void mh_mixer(uint* s, uint rk) {
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
}
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
static inline void mh_item(__global const uint* cache, uint t, uint* s) {
s[0] = 0xceed56d7u;
s[1] = 0x9ba270d2u;
s[2] = 0x82caab2du;
s[3] = 0x81ebce0eu;
s[4] = 0x12b6ecf1u;
s[5] = 0xd0f3fd7cu;
s[6] = 0xd872eefeu;
s[7] = 0xc158c7bdu;
s[8] = t * 0xf351d601u + 0xc6892460u;
s[9] = t * 0xa3bb398fu + 0x25b7228au;
s[10] = t * 0xb5a09e35u + 0xcd515004u;
s[11] = t * 0x7509c9c1u + 0x2846527au;
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
s[14] = t * 0xded91851u + 0x82961bacu;
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
for (uint r = 0u; r < 8u; ++r) {
mh_mixer(s, 0x9E3779B9u * (r + 1u));
__global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i];
}
mh_mixer(s, 0x9E3779B9u * 9u);
}
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 2 9 13 15 of w.
static inline uint mh_j(uint w) { return ((w >> 2u) & 1u) | (((w >> 9u) & 1u) << 1) | (((w >> 13u) & 1u) << 2) | (((w >> 15u) & 1u) << 3); }
static inline uint mh_t(uint w) { w = (w & 0x00007fffu) | ((w >> 16u) << 15u); w = (w & 0x00001fffu) | ((w >> 14u) << 13u); w = (w & 0x000001ffu) | ((w >> 10u) << 9u); w = (w & 0x00000003u) | ((w >> 3u) << 2u); return w; }
static inline uint mh_addr(uint t, uint j) { uint w = t; w = ((w >> 2u) << 3u) | (w & 0x00000003u) | (((j >> 0u) & 1u) << 2u); w = ((w >> 9u) << 10u) | (w & 0x000001ffu) | (((j >> 1u) & 1u) << 9u); w = ((w >> 13u) << 14u) | (w & 0x00001fffu) | (((j >> 2u) & 1u) << 13u); w = ((w >> 15u) << 16u) | (w & 0x00007fffu) | (((j >> 3u) & 1u) << 15u); return w; }
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }
// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item.
// The same constants as memhard.h in this pack (one emitter, three dialects).
__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) {
uint seg = (uint)get_global_id(0);
if (seg < nSegments) mh_cache_segment(cache, seg);
}
__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) {
uint t = (uint)get_global_id(0);
if (t < nItems) {
uint s[16];
mh_item(cache, t, s);
for (uint i = 0u; i < 16u; ++i) ds[(ulong)mh_addr(t, i)] = s[i];
}
}
// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the
// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and
// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all).
IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) {
uint gid = (uint)get_global_id(0);
uint lid = (uint)get_local_id(0);
uint nonce = baseNonce + gid;
uint r0, r1, r2, r3, r4, r5, r6, r7;
#if IGNEUM_EXCHANGE == 0
IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
uint xk = 0u;
#else
(void)lid;
#endif
{ uint x = nonce ^ 0x667d0fbdu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x7b8e5963u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
{ uint x = nonce ^ 0x7b8e5963u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x31c67e5eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
{ uint x = nonce ^ 0x31c67e5eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4529ddc6u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
{ uint x = nonce ^ 0x4529ddc6u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef19d6d8u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
{ uint x = nonce ^ 0xef19d6d8u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xaccf6211u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
{ uint x = nonce ^ 0xaccf6211u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xda0aed32u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
{ uint x = nonce ^ 0xda0aed32u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xabc6df31u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
{ uint x = nonce ^ 0xabc6df31u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x667d0fbdu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
for (uint it = 0u; it < 8u; ++it) {
uint sel = r0;
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
r7 = r7 ^ r0; // 1 xor
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 1u); r3 = r3 ^ t_; } // 2 shfl
r4 = r4 - r1; // 3 sub
r2 = r0 * r4 + r2; // 4 mad
r4 = rotl_imm(r4, 9u); // 5 rotl
r0 = rotr_var(r0, r2); // 6 rotr
r6 = r6 ^ ds[((rotl_imm(r7 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
r1 = r1 ^ ds[((rotl_imm(r4 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
r1 = r1 ^ ds[((rotl_imm(r2 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
r7 = r7 ^ ds[((rotl_imm(r0 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
r7 = r7 ^ ds[((rotl_imm(r1 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
r5 = r5 * r4; // 12 mul
r7 = r7 ^ ds[((rotl_imm(r6 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 16u); r6 = r6 ^ t_; } // 14 shfl
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r3 = r3 ^ t_; } // 15 shfl
r2 = r2 | r7; // 16 or
r0 = rotr_var(r0, r6); // 17 rotr
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
r7 = rotl_imm(r7, 24u); // 19 rotl
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
r2 = r2 | r6; // 21 or
r6 = r5 * r4 + r6; // 22 mad
r1 = r1 | r0; // 23 or
r6 = r6 ^ r1; // 24 xor
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
r7 = rotr_var(r7, r0); // 26 rotr
r4 = r5 * r7 + r4; // 27 mad
r3 = r2 * r4 + r3; // 28 mad
r1 = r1 ^ ds[((rotl_imm(r4 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
r2 = r2 ^ ds[((rotl_imm(r3 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
r1 = r1 ^ ds[((rotl_imm(r5 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
r2 = mul_hi(r2, r5); // 35 mulhi
r4 = r4 ^ r2; // 36 xor
r6 = r6 * r5; // 37 mul
r7 = r7 ^ r0; // 38 xor
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
r2 = rotr_var(r2, r3); // 40 rotr
r7 = r7 - r0; // 41 sub
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 43 shfl
r0 = r0 ^ ds[((rotl_imm(r7 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
r3 = r3 ^ ds[((rotl_imm(r1 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
r6 = r6 ^ ds[((rotl_imm(r3 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
r4 = mul_hi(r4, r2); // 48 mulhi
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r0 = r0 ^ t_; } // 50 shfl
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
r5 = r5 ^ ds[((rotl_imm(r2 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
r6 = rotl_imm(r6, 12u); // 53 rotl
r3 = rotl_imm(r3, 12u); // 54 rotl
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
r1 = r1 ^ ds[((rotl_imm(r4 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
r5 = r5 ^ ds[((rotl_imm(r0 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
r3 = mul_hi(r3, r2); // 60 mulhi
r6 = r6 | r4; // 61 or
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r5 = r5 ^ t_; } // 62 shfl
r3 = r3 ^ ds[((rotl_imm(r6 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
}
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((ulong)hi << 32) | (ulong)lo;
}
#if IGNEUM_EXCHANGE != 0
// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the
// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a
// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md.
IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) {
if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); }
}
#endif

View file

@ -0,0 +1,163 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal).
// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC.
#include <cuda_runtime.h>
#include <cstdint>
#include "program.h"
#include "memhard.h"
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
x ^= x >> 16; x *= 0x7feb352du;
x ^= x >> 15; x *= 0x846ca68bu;
x ^= x >> 16;
return x;
}
// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31.
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x.
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) {
uint32_t x = i ^ d0;
x *= 0x9E3779B1u; x ^= x >> 15;
x += d1;
x *= 0x85EBCA77u; x ^= x >> 13;
x *= 0xC2B2AE3Du; x ^= x >> 16;
return x;
}
// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item.
// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host.
__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) {
uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x;
if (seg < nSegments) mh_cache_segment(cache, seg);
}
__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
uint32_t t = blockIdx.x * blockDim.x + threadIdx.x;
if (t < nItems) {
uint32_t s[16];
mh_item(cache, t, s);
for (uint32_t i = 0u; i < 16u; ++i) ds[(size_t)mh_addr(t, i)] = s[i];
}
}
// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every
// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a
// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid.
__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) {
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
uint32_t nonce = baseNonce + gid;
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
{ uint32_t x = nonce ^ 0x667d0fbdu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x7b8e5963u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
{ uint32_t x = nonce ^ 0x7b8e5963u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x31c67e5eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
{ uint32_t x = nonce ^ 0x31c67e5eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4529ddc6u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
{ uint32_t x = nonce ^ 0x4529ddc6u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef19d6d8u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
{ uint32_t x = nonce ^ 0xef19d6d8u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xaccf6211u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
{ uint32_t x = nonce ^ 0xaccf6211u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xda0aed32u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
{ uint32_t x = nonce ^ 0xda0aed32u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xabc6df31u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
{ uint32_t x = nonce ^ 0xabc6df31u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x667d0fbdu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
for (uint32_t it = 0u; it < 8u; ++it) {
uint32_t sel = r0;
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
r7 = r7 ^ r0; // 1 xor
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r6, 1); // 2 shfl
r4 = r4 - r1; // 3 sub
r2 = r0 * r4 + r2; // 4 mad
r4 = rotl_imm(r4, 9u); // 5 rotl
r0 = rotr_var(r0, r2); // 6 rotr
r6 = r6 ^ ds[((rotl_imm(r7 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
r1 = r1 ^ ds[((rotl_imm(r4 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
r1 = r1 ^ ds[((rotl_imm(r2 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
r7 = r7 ^ ds[((rotl_imm(r0 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
r7 = r7 ^ ds[((rotl_imm(r1 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
r5 = r5 * r4; // 12 mul
r7 = r7 ^ ds[((rotl_imm(r6 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r5, 16); // 14 shfl
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r5, 1); // 15 shfl
r2 = r2 | r7; // 16 or
r0 = rotr_var(r0, r6); // 17 rotr
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
r7 = rotl_imm(r7, 24u); // 19 rotl
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
r2 = r2 | r6; // 21 or
r6 = r5 * r4 + r6; // 22 mad
r1 = r1 | r0; // 23 or
r6 = r6 ^ r1; // 24 xor
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
r7 = rotr_var(r7, r0); // 26 rotr
r4 = r5 * r7 + r4; // 27 mad
r3 = r2 * r4 + r3; // 28 mad
r1 = r1 ^ ds[((rotl_imm(r4 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
r2 = r2 ^ ds[((rotl_imm(r3 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
r1 = r1 ^ ds[((rotl_imm(r5 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
r2 = __umulhi(r2, r5); // 35 mulhi
r4 = r4 ^ r2; // 36 xor
r6 = r6 * r5; // 37 mul
r7 = r7 ^ r0; // 38 xor
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
r2 = rotr_var(r2, r3); // 40 rotr
r7 = r7 - r0; // 41 sub
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 43 shfl
r0 = r0 ^ ds[((rotl_imm(r7 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
r3 = r3 ^ ds[((rotl_imm(r1 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
r6 = r6 ^ ds[((rotl_imm(r3 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
r4 = __umulhi(r4, r2); // 48 mulhi
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r7, 4); // 50 shfl
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
r5 = r5 ^ ds[((rotl_imm(r2 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
r6 = rotl_imm(r6, 12u); // 53 rotl
r3 = rotl_imm(r3, 12u); // 54 rotl
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
r1 = r1 ^ ds[((rotl_imm(r4 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
r5 = r5 ^ ds[((rotl_imm(r0 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
r3 = __umulhi(r3, r2); // 60 mulhi
r6 = r6 | r4; // 61 or
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 62 shfl
r3 = r3 ^ ds[((rotl_imm(r6 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
}
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
}
// Host-side launch wrappers. Declared in program.h, called from host.cu.
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) {
if (nSegments == 0u) return cudaErrorInvalidValue;
uint32_t block = 256u;
uint32_t grid = (nSegments + block - 1u) / block;
igneum_cache_fill<<<grid, block>>>(cache, nSegments);
return cudaGetLastError();
}
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
if (nItems == 0u) return cudaErrorInvalidValue;
uint32_t block = 256u;
uint32_t grid = (nItems + block - 1u) / block;
igneum_build<<<grid, block>>>(ds, cache, nItems);
return cudaGetLastError();
}
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
uint32_t nonces, uint32_t blockWarps) {
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
uint32_t block = 32u * blockWarps;
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
igneum_hash<<<nonces / block, block>>>(ds, out, baseNonce, mask);
return cudaGetLastError();
}
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
cudaFuncAttributes attr;
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash);
if (e != cudaSuccess) return e;
*numRegs = attr.numRegs;
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0);
}

View file

@ -0,0 +1,375 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal).
// Built from source at runtime by proto-opencl/host.c, which passes these defines:
// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit)
// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default)
// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32
// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition
// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units;
// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32.
#ifndef IGNEUM_GROUP
#define IGNEUM_GROUP 32
#endif
#ifndef IGNEUM_EXCHANGE
#define IGNEUM_EXCHANGE 0
#endif
#ifdef __OPENCL_VERSION__
#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1)))
#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n]
#if IGNEUM_EXCHANGE == 1
#ifdef cl_khr_subgroups
#pragma OPENCL EXTENSION cl_khr_subgroups : enable
#endif
#ifdef cl_khr_subgroup_shuffle
#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable
#endif
#elif IGNEUM_EXCHANGE == 2
#pragma OPENCL EXTENSION cl_intel_subgroups : enable
#endif
#else
// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros.
#include "emu_opencl.h"
#endif
#if IGNEUM_EXCHANGE == 1
#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m))
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
#elif IGNEUM_EXCHANGE == 2
#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m))
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
#else
// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per
// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane
// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's
// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier.
#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; }
#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; }
#endif
static inline uint splitmix32(uint x) {
x ^= x >> 16; x *= 0x7feb352du;
x ^= x >> 15; x *= 0x846ca68bu;
x ^= x >> 16;
return x;
}
// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32.
static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); }
// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x.
static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); }
static inline uint ds_elem(uint i, uint d0, uint d1) {
uint x = i ^ d0;
x *= 0x9E3779B1u; x ^= x >> 15;
x += d1;
x *= 0x85EBCA77u; x ^= x >> 13;
x *= 0xC2B2AE3Du; x ^= x >> 16;
return x;
}
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
#define MH_CACHE_LINE_MASK 0x003fffffu
#define MH_SEGMENT_LINES 64u
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
// y = ChaCha12 core(x) + x
static inline void mh_chacha_block(const uint* x, uint* y) {
for (uint i = 0u; i < 16u; ++i) y[i] = x[i];
for (uint r = 0u; r < 6u; ++r) {
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
}
for (uint i = 0u; i < 16u; ++i) y[i] += x[i];
}
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
static inline void mh_cache_segment(__global uint* cache, uint seg) {
uint prev[16]; uint x[16]; uint y[16];
for (uint i = 0u; i < 16u; ++i) prev[i] = 0u;
for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) {
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
x[4] = 0xceed56d7u ^ prev[4];
x[5] = 0x9ba270d2u ^ prev[5];
x[6] = 0x82caab2du ^ prev[6];
x[7] = 0x81ebce0eu ^ prev[7];
x[8] = 0x12b6ecf1u ^ prev[8];
x[9] = 0xd0f3fd7cu ^ prev[9];
x[10] = 0xd872eefeu ^ prev[10];
x[11] = 0xc158c7bdu ^ prev[11];
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
mh_chacha_block(x, y);
__global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
}
}
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
static inline void mh_mixer(uint* s, uint rk) {
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
}
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
static inline void mh_item(__global const uint* cache, uint t, uint* s) {
s[0] = 0xceed56d7u;
s[1] = 0x9ba270d2u;
s[2] = 0x82caab2du;
s[3] = 0x81ebce0eu;
s[4] = 0x12b6ecf1u;
s[5] = 0xd0f3fd7cu;
s[6] = 0xd872eefeu;
s[7] = 0xc158c7bdu;
s[8] = t * 0xf351d601u + 0xc6892460u;
s[9] = t * 0xa3bb398fu + 0x25b7228au;
s[10] = t * 0xb5a09e35u + 0xcd515004u;
s[11] = t * 0x7509c9c1u + 0x2846527au;
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
s[14] = t * 0xded91851u + 0x82961bacu;
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
for (uint r = 0u; r < 8u; ++r) {
mh_mixer(s, 0x9E3779B9u * (r + 1u));
__global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i];
}
mh_mixer(s, 0x9E3779B9u * 9u);
}
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 2 9 13 15 of w.
static inline uint mh_j(uint w) { return ((w >> 2u) & 1u) | (((w >> 9u) & 1u) << 1) | (((w >> 13u) & 1u) << 2) | (((w >> 15u) & 1u) << 3); }
static inline uint mh_t(uint w) { w = (w & 0x00007fffu) | ((w >> 16u) << 15u); w = (w & 0x00001fffu) | ((w >> 14u) << 13u); w = (w & 0x000001ffu) | ((w >> 10u) << 9u); w = (w & 0x00000003u) | ((w >> 3u) << 2u); return w; }
static inline uint mh_addr(uint t, uint j) { uint w = t; w = ((w >> 2u) << 3u) | (w & 0x00000003u) | (((j >> 0u) & 1u) << 2u); w = ((w >> 9u) << 10u) | (w & 0x000001ffu) | (((j >> 1u) & 1u) << 9u); w = ((w >> 13u) << 14u) | (w & 0x00001fffu) | (((j >> 2u) & 1u) << 13u); w = ((w >> 15u) << 16u) | (w & 0x00007fffu) | (((j >> 3u) & 1u) << 15u); return w; }
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }
// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item.
// The same constants as memhard.h in this pack (one emitter, three dialects).
__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) {
uint seg = (uint)get_global_id(0);
if (seg < nSegments) mh_cache_segment(cache, seg);
}
__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) {
uint t = (uint)get_global_id(0);
if (t < nItems) {
uint s[16];
mh_item(cache, t, s);
for (uint i = 0u; i < 16u; ++i) ds[(ulong)mh_addr(t, i)] = s[i];
}
}
// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the
// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and
// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all).
IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) {
uint gid = (uint)get_global_id(0);
uint lid = (uint)get_local_id(0);
uint nonce = baseNonce + gid;
uint r0, r1, r2, r3, r4, r5, r6, r7;
#if IGNEUM_EXCHANGE == 0
IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
uint xk = 0u;
#else
(void)lid;
#endif
{ uint x = nonce ^ 0x667d0fbdu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x7b8e5963u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
{ uint x = nonce ^ 0x7b8e5963u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x31c67e5eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
{ uint x = nonce ^ 0x31c67e5eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4529ddc6u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
{ uint x = nonce ^ 0x4529ddc6u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef19d6d8u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
{ uint x = nonce ^ 0xef19d6d8u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xaccf6211u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
{ uint x = nonce ^ 0xaccf6211u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xda0aed32u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
{ uint x = nonce ^ 0xda0aed32u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xabc6df31u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
{ uint x = nonce ^ 0xabc6df31u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x667d0fbdu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
for (uint it = 0u; it < 8u; ++it) {
uint sel = r0;
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
r7 = r7 ^ r0; // 1 xor
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 1u); r3 = r3 ^ t_; } // 2 shfl
r4 = r4 - r1; // 3 sub
r2 = r0 * r4 + r2; // 4 mad
r4 = rotl_imm(r4, 9u); // 5 rotl
r0 = rotr_var(r0, r2); // 6 rotr
r6 = r6 ^ ds[((rotl_imm(r7 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
r1 = r1 ^ ds[((rotl_imm(r4 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
r1 = r1 ^ ds[((rotl_imm(r2 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
r7 = r7 ^ ds[((rotl_imm(r0 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
r7 = r7 ^ ds[((rotl_imm(r1 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
r5 = r5 * r4; // 12 mul
r7 = r7 ^ ds[((rotl_imm(r6 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 16u); r6 = r6 ^ t_; } // 14 shfl
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r3 = r3 ^ t_; } // 15 shfl
r2 = r2 | r7; // 16 or
r0 = rotr_var(r0, r6); // 17 rotr
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
r7 = rotl_imm(r7, 24u); // 19 rotl
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
r2 = r2 | r6; // 21 or
r6 = r5 * r4 + r6; // 22 mad
r1 = r1 | r0; // 23 or
r6 = r6 ^ r1; // 24 xor
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
r7 = rotr_var(r7, r0); // 26 rotr
r4 = r5 * r7 + r4; // 27 mad
r3 = r2 * r4 + r3; // 28 mad
r1 = r1 ^ ds[((rotl_imm(r4 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
r2 = r2 ^ ds[((rotl_imm(r3 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
r1 = r1 ^ ds[((rotl_imm(r5 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
r2 = mul_hi(r2, r5); // 35 mulhi
r4 = r4 ^ r2; // 36 xor
r6 = r6 * r5; // 37 mul
r7 = r7 ^ r0; // 38 xor
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
r2 = rotr_var(r2, r3); // 40 rotr
r7 = r7 - r0; // 41 sub
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 43 shfl
r0 = r0 ^ ds[((rotl_imm(r7 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
r3 = r3 ^ ds[((rotl_imm(r1 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
r6 = r6 ^ ds[((rotl_imm(r3 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
r4 = mul_hi(r4, r2); // 48 mulhi
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r0 = r0 ^ t_; } // 50 shfl
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
r5 = r5 ^ ds[((rotl_imm(r2 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
r6 = rotl_imm(r6, 12u); // 53 rotl
r3 = rotl_imm(r3, 12u); // 54 rotl
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
r1 = r1 ^ ds[((rotl_imm(r4 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
r5 = r5 ^ ds[((rotl_imm(r0 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
r3 = mul_hi(r3, r2); // 60 mulhi
r6 = r6 | r4; // 61 or
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r5 = r5 ^ t_; } // 62 shfl
r3 = r3 ^ ds[((rotl_imm(r6 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
}
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((ulong)hi << 32) | (ulong)lo;
}
#if IGNEUM_EXCHANGE != 0
// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the
// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a
// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md.
IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) {
if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); }
}
#endif
// Header-bound variant (bind.rs): the init words come from initw, not SEEDW. Same body as igneum_hash.
IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw) {
uint gid = (uint)get_global_id(0);
uint lid = (uint)get_local_id(0);
uint nonce = baseNonce + gid;
uint r0, r1, r2, r3, r4, r5, r6, r7;
uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7];
#if IGNEUM_EXCHANGE == 0
IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
uint xk = 0u;
#else
(void)lid;
#endif
{ uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; }
{ uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; }
{ uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; }
{ uint x = nonce ^ iw3; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw4; }
{ uint x = nonce ^ iw4; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw5; }
{ uint x = nonce ^ iw5; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw6; }
{ uint x = nonce ^ iw6; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw7; }
{ uint x = nonce ^ iw7; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw0; }
for (uint it = 0u; it < 8u; ++it) {
uint sel = r0;
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
r7 = r7 ^ r0; // 1 xor
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 1u); r3 = r3 ^ t_; } // 2 shfl
r4 = r4 - r1; // 3 sub
r2 = r0 * r4 + r2; // 4 mad
r4 = rotl_imm(r4, 9u); // 5 rotl
r0 = rotr_var(r0, r2); // 6 rotr
r6 = r6 ^ ds[((rotl_imm(r7 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
r1 = r1 ^ ds[((rotl_imm(r4 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
r1 = r1 ^ ds[((rotl_imm(r2 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
r7 = r7 ^ ds[((rotl_imm(r0 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
r7 = r7 ^ ds[((rotl_imm(r1 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
r5 = r5 * r4; // 12 mul
r7 = r7 ^ ds[((rotl_imm(r6 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 16u); r6 = r6 ^ t_; } // 14 shfl
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r3 = r3 ^ t_; } // 15 shfl
r2 = r2 | r7; // 16 or
r0 = rotr_var(r0, r6); // 17 rotr
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
r7 = rotl_imm(r7, 24u); // 19 rotl
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
r2 = r2 | r6; // 21 or
r6 = r5 * r4 + r6; // 22 mad
r1 = r1 | r0; // 23 or
r6 = r6 ^ r1; // 24 xor
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
r7 = rotr_var(r7, r0); // 26 rotr
r4 = r5 * r7 + r4; // 27 mad
r3 = r2 * r4 + r3; // 28 mad
r1 = r1 ^ ds[((rotl_imm(r4 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
r2 = r2 ^ ds[((rotl_imm(r3 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
r1 = r1 ^ ds[((rotl_imm(r5 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
r2 = mul_hi(r2, r5); // 35 mulhi
r4 = r4 ^ r2; // 36 xor
r6 = r6 * r5; // 37 mul
r7 = r7 ^ r0; // 38 xor
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
r2 = rotr_var(r2, r3); // 40 rotr
r7 = r7 - r0; // 41 sub
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 43 shfl
r0 = r0 ^ ds[((rotl_imm(r7 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
r3 = r3 ^ ds[((rotl_imm(r1 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
r6 = r6 ^ ds[((rotl_imm(r3 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
r4 = mul_hi(r4, r2); // 48 mulhi
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r0 = r0 ^ t_; } // 50 shfl
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
r5 = r5 ^ ds[((rotl_imm(r2 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
r6 = rotl_imm(r6, 12u); // 53 rotl
r3 = rotl_imm(r3, 12u); // 54 rotl
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
r1 = r1 ^ ds[((rotl_imm(r4 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
r5 = r5 ^ ds[((rotl_imm(r0 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
r3 = mul_hi(r3, r2); // 60 mulhi
r6 = r6 | r4; // 61 or
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r5 = r5 ^ t_; } // 62 shfl
r3 = r3 ^ ds[((rotl_imm(r6 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
}
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((ulong)hi << 32) | (ulong)lo;
}

View file

@ -0,0 +1,123 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
// Header-bound twin of igneum_hash in kernel.cu: the init words come from a kernel argument, not SEEDW.
// Host declarations (also in program_bound.h if present):
// struct IgneumInitWords { uint32_t w[8]; };
// cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
// IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps);
// cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
#include <cuda_runtime.h>
#include <cstdint>
#include "program.h"
struct IgneumInitWords { uint32_t w[8]; };
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
x ^= x >> 16; x *= 0x7feb352du;
x ^= x >> 15; x *= 0x846ca68bu;
x ^= x >> 16;
return x;
}
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw) {
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
uint32_t nonce = baseNonce + gid;
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
{ uint32_t x = nonce ^ iw.w[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw.w[1]; }
{ uint32_t x = nonce ^ iw.w[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw.w[2]; }
{ uint32_t x = nonce ^ iw.w[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw.w[3]; }
{ uint32_t x = nonce ^ iw.w[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw.w[4]; }
{ uint32_t x = nonce ^ iw.w[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw.w[5]; }
{ uint32_t x = nonce ^ iw.w[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw.w[6]; }
{ uint32_t x = nonce ^ iw.w[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw.w[7]; }
{ uint32_t x = nonce ^ iw.w[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw.w[0]; }
for (uint32_t it = 0u; it < 8u; ++it) {
uint32_t sel = r0;
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
r7 = r7 ^ r0; // 1 xor
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r6, 1); // 2 shfl
r4 = r4 - r1; // 3 sub
r2 = r0 * r4 + r2; // 4 mad
r4 = rotl_imm(r4, 9u); // 5 rotl
r0 = rotr_var(r0, r2); // 6 rotr
r6 = r6 ^ ds[((rotl_imm(r7 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
r1 = r1 ^ ds[((rotl_imm(r4 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
r1 = r1 ^ ds[((rotl_imm(r2 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
r7 = r7 ^ ds[((rotl_imm(r0 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
r7 = r7 ^ ds[((rotl_imm(r1 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
r5 = r5 * r4; // 12 mul
r7 = r7 ^ ds[((rotl_imm(r6 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r5, 16); // 14 shfl
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r5, 1); // 15 shfl
r2 = r2 | r7; // 16 or
r0 = rotr_var(r0, r6); // 17 rotr
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
r7 = rotl_imm(r7, 24u); // 19 rotl
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
r2 = r2 | r6; // 21 or
r6 = r5 * r4 + r6; // 22 mad
r1 = r1 | r0; // 23 or
r6 = r6 ^ r1; // 24 xor
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
r7 = rotr_var(r7, r0); // 26 rotr
r4 = r5 * r7 + r4; // 27 mad
r3 = r2 * r4 + r3; // 28 mad
r1 = r1 ^ ds[((rotl_imm(r4 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
r2 = r2 ^ ds[((rotl_imm(r3 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
r1 = r1 ^ ds[((rotl_imm(r5 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
r2 = __umulhi(r2, r5); // 35 mulhi
r4 = r4 ^ r2; // 36 xor
r6 = r6 * r5; // 37 mul
r7 = r7 ^ r0; // 38 xor
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
r2 = rotr_var(r2, r3); // 40 rotr
r7 = r7 - r0; // 41 sub
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 43 shfl
r0 = r0 ^ ds[((rotl_imm(r7 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
r3 = r3 ^ ds[((rotl_imm(r1 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
r6 = r6 ^ ds[((rotl_imm(r3 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
r4 = __umulhi(r4, r2); // 48 mulhi
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r7, 4); // 50 shfl
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
r5 = r5 ^ ds[((rotl_imm(r2 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
r6 = rotl_imm(r6, 12u); // 53 rotl
r3 = rotl_imm(r3, 12u); // 54 rotl
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
r1 = r1 ^ ds[((rotl_imm(r4 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
r5 = r5 ^ ds[((rotl_imm(r0 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
r3 = __umulhi(r3, r2); // 60 mulhi
r6 = r6 | r4; // 61 or
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 62 shfl
r3 = r3 ^ ds[((rotl_imm(r6 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
}
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
}
cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps) {
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
uint32_t block = 32u * blockWarps;
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
igneum_hash_bound<<<nonces / block, block>>>(ds, out, baseNonce, mask, iw);
return cudaGetLastError();
}
cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
cudaFuncAttributes attr;
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash_bound);
if (e != cudaSuccess) return e;
*numRegs = attr.numRegs;
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash_bound, (int)(32u * blockWarps), 0);
}

View file

@ -0,0 +1,112 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
// Memory-hard dataset core, the same text that the Mac's Metal kernels and CPU verifier were checked against.
// Included by kernel.cu (device), host.cu (host reference) and proto-opencl/host.c (C99 host reference).
// See proto-metal/MEMHARD.md for the construction. kernel.cl carries the same text in OpenCL C.
#pragma once
#ifdef __cplusplus
#include <cstdint>
#else
#include <stdint.h>
#endif
#if defined(__CUDACC__)
#define IGNEUM_HD __host__ __device__ __forceinline__
#elif defined(_MSC_VER) && !defined(__cplusplus)
#define IGNEUM_HD static __inline
#else
#define IGNEUM_HD static inline
#endif
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
#define MH_CACHE_LINE_MASK 0x003fffffu
#define MH_SEGMENT_LINES 64u
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
IGNEUM_HD uint32_t mh_rotl(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
// y = ChaCha12 core(x) + x
IGNEUM_HD void mh_chacha_block(const uint32_t* x, uint32_t* y) {
for (uint32_t i = 0u; i < 16u; ++i) y[i] = x[i];
for (uint32_t r = 0u; r < 6u; ++r) {
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
}
for (uint32_t i = 0u; i < 16u; ++i) y[i] += x[i];
}
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
IGNEUM_HD void mh_cache_segment(uint32_t* cache, uint32_t seg) {
uint32_t prev[16]; uint32_t x[16]; uint32_t y[16];
for (uint32_t i = 0u; i < 16u; ++i) prev[i] = 0u;
for (uint32_t j = 0u; j < MH_SEGMENT_LINES; ++j) {
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
x[4] = 0xceed56d7u ^ prev[4];
x[5] = 0x9ba270d2u ^ prev[5];
x[6] = 0x82caab2du ^ prev[6];
x[7] = 0x81ebce0eu ^ prev[7];
x[8] = 0x12b6ecf1u ^ prev[8];
x[9] = 0xd0f3fd7cu ^ prev[9];
x[10] = 0xd872eefeu ^ prev[10];
x[11] = 0xc158c7bdu ^ prev[11];
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
mh_chacha_block(x, y);
uint32_t* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
for (uint32_t i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
}
}
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
IGNEUM_HD void mh_mixer(uint32_t* s, uint32_t rk) {
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
}
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
IGNEUM_HD void mh_item(const uint32_t* cache, uint32_t t, uint32_t* s) {
s[0] = 0xceed56d7u;
s[1] = 0x9ba270d2u;
s[2] = 0x82caab2du;
s[3] = 0x81ebce0eu;
s[4] = 0x12b6ecf1u;
s[5] = 0xd0f3fd7cu;
s[6] = 0xd872eefeu;
s[7] = 0xc158c7bdu;
s[8] = t * 0xf351d601u + 0xc6892460u;
s[9] = t * 0xa3bb398fu + 0x25b7228au;
s[10] = t * 0xb5a09e35u + 0xcd515004u;
s[11] = t * 0x7509c9c1u + 0x2846527au;
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
s[14] = t * 0xded91851u + 0x82961bacu;
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
for (uint32_t r = 0u; r < 8u; ++r) {
mh_mixer(s, 0x9E3779B9u * (r + 1u));
const uint32_t* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= line[i];
}
mh_mixer(s, 0x9E3779B9u * 9u);
}
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 2 9 13 15 of w.
IGNEUM_HD uint32_t mh_j(uint32_t w) { return ((w >> 2u) & 1u) | (((w >> 9u) & 1u) << 1) | (((w >> 13u) & 1u) << 2) | (((w >> 15u) & 1u) << 3); }
IGNEUM_HD uint32_t mh_t(uint32_t w) { w = (w & 0x00007fffu) | ((w >> 16u) << 15u); w = (w & 0x00001fffu) | ((w >> 14u) << 13u); w = (w & 0x000001ffu) | ((w >> 10u) << 9u); w = (w & 0x00000003u) | ((w >> 3u) << 2u); return w; }
IGNEUM_HD uint32_t mh_addr(uint32_t t, uint32_t j) { uint32_t w = t; w = ((w >> 2u) << 3u) | (w & 0x00000003u) | (((j >> 0u) & 1u) << 2u); w = ((w >> 9u) << 10u) | (w & 0x000001ffu) | (((j >> 1u) & 1u) << 9u); w = ((w >> 13u) << 14u) | (w & 0x00001fffu) | (((j >> 2u) & 1u) << 13u); w = ((w >> 15u) << 16u) | (w & 0x00007fffu) | (((j >> 3u) & 1u) << 15u); return w; }
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
IGNEUM_HD uint32_t mh_word(const uint32_t* cache, uint32_t w) { uint32_t s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }

View file

@ -0,0 +1,109 @@
#include <metal_stdlib>
using namespace metal;
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
#define MH_CACHE_LINE_MASK 0x003fffffu
#define MH_SEGMENT_LINES 64u
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
// y = ChaCha12 core(x) + x
inline void mh_chacha_block(const thread uint* x, thread uint* y) {
for (uint i = 0u; i < 16u; ++i) y[i] = x[i];
for (uint r = 0u; r < 6u; ++r) {
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
}
for (uint i = 0u; i < 16u; ++i) y[i] += x[i];
}
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
inline void mh_cache_segment(device uint* cache, uint seg) {
uint prev[16]; uint x[16]; uint y[16];
for (uint i = 0u; i < 16u; ++i) prev[i] = 0u;
for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) {
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
x[4] = 0xceed56d7u ^ prev[4];
x[5] = 0x9ba270d2u ^ prev[5];
x[6] = 0x82caab2du ^ prev[6];
x[7] = 0x81ebce0eu ^ prev[7];
x[8] = 0x12b6ecf1u ^ prev[8];
x[9] = 0xd0f3fd7cu ^ prev[9];
x[10] = 0xd872eefeu ^ prev[10];
x[11] = 0xc158c7bdu ^ prev[11];
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
mh_chacha_block(x, y);
device uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
}
}
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
inline void mh_mixer(thread uint* s, uint rk) {
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
}
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
inline void mh_item(device const uint* cache, uint t, thread uint* s) {
s[0] = 0xceed56d7u;
s[1] = 0x9ba270d2u;
s[2] = 0x82caab2du;
s[3] = 0x81ebce0eu;
s[4] = 0x12b6ecf1u;
s[5] = 0xd0f3fd7cu;
s[6] = 0xd872eefeu;
s[7] = 0xc158c7bdu;
s[8] = t * 0xf351d601u + 0xc6892460u;
s[9] = t * 0xa3bb398fu + 0x25b7228au;
s[10] = t * 0xb5a09e35u + 0xcd515004u;
s[11] = t * 0x7509c9c1u + 0x2846527au;
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
s[14] = t * 0xded91851u + 0x82961bacu;
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
for (uint r = 0u; r < 8u; ++r) {
mh_mixer(s, 0x9E3779B9u * (r + 1u));
device const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i];
}
mh_mixer(s, 0x9E3779B9u * 9u);
}
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 2 9 13 15 of w.
inline uint mh_j(uint w) { return ((w >> 2u) & 1u) | (((w >> 9u) & 1u) << 1) | (((w >> 13u) & 1u) << 2) | (((w >> 15u) & 1u) << 3); }
inline uint mh_t(uint w) { w = (w & 0x00007fffu) | ((w >> 16u) << 15u); w = (w & 0x00001fffu) | ((w >> 14u) << 13u); w = (w & 0x000001ffu) | ((w >> 10u) << 9u); w = (w & 0x00000003u) | ((w >> 3u) << 2u); return w; }
inline uint mh_addr(uint t, uint j) { uint w = t; w = ((w >> 2u) << 3u) | (w & 0x00000003u) | (((j >> 0u) & 1u) << 2u); w = ((w >> 9u) << 10u) | (w & 0x000001ffu) | (((j >> 1u) & 1u) << 9u); w = ((w >> 13u) << 14u) | (w & 0x00001fffu) | (((j >> 2u) & 1u) << 13u); w = ((w >> 15u) << 16u) | (w & 0x00007fffu) | (((j >> 3u) & 1u) << 15u); return w; }
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
inline uint mh_word(device const uint* cache, uint w) { uint s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }
// One thread per segment (2^16 threads).
kernel void igneum_cache_fill(device uint* cache [[buffer(0)]], uint gid [[thread_position_in_grid]]) {
mh_cache_segment(cache, gid);
}
// One thread per 64-byte item (dataset words / 16 threads).
kernel void igneum_build(device const uint* cache [[buffer(0)]], device uint* dataset [[buffer(1)]],
uint gid [[thread_position_in_grid]]) {
uint s[16];
mh_item(cache, gid, s);
for (uint i = 0u; i < 16u; ++i) dataset[mh_addr(gid, i)] = s[i];
}

View file

@ -0,0 +1,76 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
// Program metadata for host.cu plus the launch wrappers defined in kernel.cu.
// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros.
#pragma once
#ifdef __cplusplus
#include <cstdint>
#else
#include <stdint.h>
#endif
#ifndef IGNEUM_NO_CUDA
#include <cuda_runtime.h>
#endif
#define IGNEUM_SEED_STRING "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000"
#define IGNEUM_SEED_BYTES_HEX "edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07"
#define IGNEUM_GENERATOR 3
#define IGNEUM_PROGRAM_ATTEMPT 0
#define IGNEUM_PROGRAM_ID 0x73bcbfe8ccf988f1ull
#define IGNEUM_DAY_STRING "bytes:69676e65756d2d6461792ffa50000000000000"
#define IGNEUM_DAY_BYTES_HEX "69676e65756d2d6461792ffa50000000000000"
#define IGNEUM_DAY0 0xceed56d7u
#define IGNEUM_DAY1 0x9ba270d2u
#define IGNEUM_DATASET_LOG2 28
#define IGNEUM_MASK 0x0fffffffu
#define IGNEUM_LANES 32
#define IGNEUM_ITERATIONS 8
#define IGNEUM_INSTR_COUNT 64
#define IGNEUM_LOADS_PER_HASH 128
#define IGNEUM_WIDE_LOADS_PER_HASH 0
#define IGNEUM_OP_MIX "load=16 add=15 shfl=6 mad=4 or=4 rotl=4 rotr=4 xor=4 mulhi=3 mul=2 sub=2"
// Program class v3 (Counter ASIC 2.0, docs/plans/counter-asic-2-rollout.md): generator version 3; a worker that
// runs another class refuses this pack, and a job line names the class it wants (class=v3 era=<hex>).
#define IGNEUM_PROGRAM_CLASS "v3"
#define IGNEUM_ERA_SEED_HEX "5e0587f455a86e91e4990f5c481a34cca0044d3f3ae519adcae58ddc82885d31"
// Class v3 construction (Counter ASIC 2.0, 5 October 2026, docs/plans/mixer-x4.md): version 2 loads; the dataset item
// derivation applies the mixer IGNEUM_MIXER_MULT times per round (memhard.h), and the cache follows the growth rule.
#define IGNEUM_LOAD_CLASS "w4-era4488f3ed"
#define IGNEUM_LOAD_SLOTS 16
#define IGNEUM_LOAD_MIX { 100, 0, 0 }
#define IGNEUM_LOAD_WIDTH_COUNTS { 16, 0, 0 } // loads of 4, 16, 64 bytes per program
#define IGNEUM_BYTES_PER_HASH 512
#define IGNEUM_FOLD_ROT 11
#define IGNEUM_FOLD_MUL 0x9e3779b1u
// Era layout (5 October 2026, docs/plans/era-layout.md): NOT the lottery hash. Every dataset load reads
// idx = ((rotl(src * STRIDE_MUL, STRIDE_ROT) & window mask) | window offset) & MASK; the window of a load site is the
// dataset, a half or a quarter of it (IGNEUM_ERA_WINDOWS: site:shrink:offset); dataset word w holds word j(w) of item
// t(w) with j's bits at the INTERLEAVE positions (memhard.h: mh_t, mh_j, mh_addr).
#define IGNEUM_ERA_LABEL "4488f3ed"
#define IGNEUM_ERA_SEED_WORDS { 0x4488f3edu, 0x3cf22d2au, 0xb3e8271eu, 0x55754277u, 0xc2b1c4c7u, 0x8e627302u, 0x584d5acdu, 0xc31dd01du }
#define IGNEUM_ERA_ALLOWED_WIDTHS { 1, 0, 0 } // words, ascending, 0 = unused; one entry pins the width
#define IGNEUM_ERA_WIDTH_WORDS 1
#define IGNEUM_ERA_STRIDE_MUL 0x4d38603du
#define IGNEUM_ERA_STRIDE_ROT 10
#define IGNEUM_ERA_INTERLEAVE { 2, 9, 13, 15 }
#define IGNEUM_ERA_WINDOWS "7:2:1 8:1:1 9:1:1 10:1:1 11:0:0 13:1:1 29:0:0 30:2:2 31:1:1 44:1:1 46:2:0 47:0:0 52:0:0 56:0:0 58:2:0 63:1:1"
// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h)
#define IGNEUM_DATASET_MODE 1
#define IGNEUM_SEEDW_INIT { 0x667d0fbdu, 0x7b8e5963u, 0x31c67e5eu, 0x4529ddc6u, 0xef19d6d8u, 0xaccf6211u, 0xda0aed32u, 0xabc6df31u }
#define IGNEUM_KEY_INIT { 0xceed56d7u, 0x9ba270d2u, 0x82caab2du, 0x81ebce0eu, 0x12b6ecf1u, 0xd0f3fd7cu, 0xd872eefeu, 0xc158c7bdu }
#define IGNEUM_CACHE_LOG2_WORDS 26
#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6
#define IGNEUM_CACHE_SEGMENTS 65536u
#define IGNEUM_ITEM_ROUNDS 8
#define IGNEUM_MIX_ROT_INIT { 17u, 12u, 20u, 23u, 7u, 3u, 27u, 16u }
#define IGNEUM_MIX_MUL_INIT { 0xf351d601u, 0xa3bb398fu, 0xb5a09e35u, 0x7509c9c1u, 0x6bbf31e9u, 0xfc849a79u, 0xded91851u, 0x8d9113d1u, 0x0ff15225u, 0x3a5bdd41u, 0xab533435u, 0xe1c55ad5u, 0xe6d3bd0du, 0x9d9ffbbdu, 0xbb2a3cf3u, 0x50a7c08du }
#define IGNEUM_MIX_RC_INIT { 0xc6892460u, 0x25b7228au, 0xcd515004u, 0x2846527au, 0xa6324241u, 0x36e3ec53u, 0x82961bacu, 0x0f97ba7du, 0xb6f921a9u, 0x3ada24e5u, 0xde20ab91u, 0x5378eeb2u, 0x7d161662u, 0x89353cc1u, 0xb1aa03a2u, 0x788acae6u }
#ifndef IGNEUM_NO_CUDA
// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError().
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments);
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems);
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
uint32_t nonces, uint32_t blockWarps);
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
#endif

View file

@ -0,0 +1,143 @@
{
"format": "igneum-program-pack-3",
"generator": 3,
"attempt": 0,
"program_id": "0x73bcbfe8ccf988f1",
"program_id_derivation": "FNV-1a 64 over 'igneum-program/' || generator_le32 || seed_words as little-endian bytes || attempt_le32",
"dataset_mode": "memory-hard",
"seed": "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000",
"seed_bytes": "edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07",
"seed_words": ["0x667d0fbd", "0x7b8e5963", "0x31c67e5e", "0x4529ddc6", "0xef19d6d8", "0xaccf6211", "0xda0aed32", "0xabc6df31"],
"seed_derivation": "seed_words = FNV-1a 64 over seed_bytes (attempt 0) or seed_bytes || attempt_le32 (attempt k >= 1), basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32",
"generator_rule": "version 2: exactly 16 load slots drawn first from instructions 1..63 (partial Fisher-Yates), the other 48 ops from the ten non-load weights (sum 75); a load's source is drawn from the registers other than dst written by an earlier instruction and not read by a load since; the candidate must pass the acceptance rule of spec 01 section 1.4.6 (static: no cyclically stale load source, every register has an injecting write; dynamic: 64 units on the seed-keyed closed-form dataset with no constant register bit, no lane-constant load site, under 164 saturated final values, every output bit within 136 of 1024, distinct addresses above 245760), else the next attempt of the seed is tried",
"lanes": 32,
"registers": 8,
"iterations": 8,
"instruction_count": 64,
"loads_per_hash": 128,
"program_class": "v3",
"era_seed_bytes": "5e0587f455a86e91e4990f5c481a34cca0044d3f3ae519adcae58ddc82885d31",
"load_class": "w4-era4488f3ed",
"load_slots": 16,
"load_mix_percent_4_16_64": [100, 0, 0],
"load_width_counts_4_16_64": [16, 0, 0],
"bytes_per_hash": 512,
"wide_load": "read-width experiment (5 October 2026, docs/plans/read-width.md), NOT the lottery hash: a load of W words (width field, 4 or 16) reads dataset[b .. b + W) with b = (src & mask) & ~(W - 1) and folds every word into dst: x = dst ^ w[0]; for j in 1..W: x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x; width 1 is the plain load; the width is drawn per instruction from the class mix with one extra below(100) draw after the nine of version 2, and the program id is FNV-1a 64 over 'igneum-program-rw/' || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots",
"era": {
"label": "4488f3ed",
"seed_words": ["0x4488f3ed", "0x3cf22d2a", "0xb3e8271e", "0x55754277", "0xc2b1c4c7", "0x8e627302", "0x584d5acd", "0xc31dd01d"],
"draw": "docs/plans/era-layout.md 1.1: SplitMix64 seeded with seed_words[0] | seed_words[1] << 32 of seed_words_from_bytes('igneum-era/' || n_le64 || E_n); width = allowed[below(|allowed|)], stride_mul = low32(next()) | 1, stride_rot = 1 + below(31), then four next() draws for a partial Fisher-Yates over positions log2(W)..15 of which 4 - log2(W) are used",
"allowed_widths": [1],
"width_words": 1,
"stride_mul": "0x4d38603d",
"stride_rot": 10,
"interleave": [2, 9, 13, 15],
"address": "y = rotl(src * stride_mul, stride_rot); k = min(win, D - 26); idx = ((y & (mask >> k)) | ((off & (2^k - 1)) << (D - k))) & mask; a wide load aligns idx down to W words",
"windows": "per instruction, after the width roll: win = below(3), off = low32(next()) & (2^win - 1); used on a load slot (the instruction's win and off fields)",
"dataset_word": "dataset[w] = item(t(w))[j(w)]: j(w) gathers the bits of w at the interleave positions, t(w) is w with those bits removed",
"program_id_suffix": "'era/' || allowed[3] || width_words || stride_mul_le32 || stride_rot_le32 || interleave[4]"
},
"op_mix": {"load": 16, "add": 15, "shfl": 6, "mad": 4, "or": 4, "rotl": 4, "rotr": 4, "xor": 4, "mulhi": 3, "mul": 2, "sub": 2},
"register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]",
"splitmix32": "x ^= x>>16; x *= 0x7feb352d; x ^= x>>15; x *= 0x846ca68b; x ^= x>>16",
"iteration": "sel = r0 sampled once at the top of each iteration, then all instructions in order",
"output": "lo = r0 ^ rotl(r1,7) ^ rotl(r2,14) ^ rotl(r3,21); hi = r4 ^ rotl(r5,9) ^ rotl(r6,18) ^ rotl(r7,27); out = (hi << 32) | lo",
"op_semantics": {
"add": "dst = dst + src + (bit `bit` of sel ? imm2 : imm)",
"sub": "dst = dst - src",
"mul": "dst = dst * src (low 32)",
"mulhi": "dst = high 32 bits of dst * src",
"xor": "dst = dst ^ src",
"or": "dst = dst | src",
"rotl": "dst = rotl(dst, rot), rot in 1..31",
"rotr": "dst = rotr(dst, src & 31)",
"mad": "dst = src * src2 + dst",
"shfl": "dst = dst ^ (src of lane (lane ^ mask)), mask in {1,2,4,8,16}, within the 32-lane warp",
"load": "dst = dst ^ dataset[src & dataset.mask]",
"wload": "base = (src of lane 0 & dataset.mask) & ~31; dst = dst ^ dataset[base + lane] (warp-coalesced 128-byte load, lever b, only when --wide-frac > 0)"
},
"dataset": {
"log2_words": 28,
"bytes": 1073741824,
"mask": "0x0fffffff",
"day": "bytes:69676e65756d2d6461792ffa50000000000000",
"day_bytes": "69676e65756d2d6461792ffa50000000000000",
"day_words_from": "seed_words_from_bytes(day_bytes)",
"d0": "0xceed56d7",
"d1": "0x9ba270d2",
"mode": "memory-hard",
"spec": "proto-metal/MEMHARD.md",
"key": ["0xceed56d7", "0x9ba270d2", "0x82caab2d", "0x81ebce0e", "0x12b6ecf1", "0xd0f3fd7c", "0xd872eefe", "0xc158c7bd"],
"key_derivation": "the 8 words of seed_words_from_bytes(day_bytes); d0, d1 are key[0], key[1]",
"cache": {"log2_words": 26, "bytes": 268435456, "line_words": 16, "segment_lines": 64, "segments": 65536, "block": "ChaCha12 core + feed-forward, rotations 16 12 8 7", "sigma": ["0x61707865", "0x3320646e", "0x79622d32", "0x6b206574"], "tag": ["0x49676e65", "0x756d4d48"], "chain": "in_j = prev_line ^ (sigma[0..3] || key[0..7] || seg || j || tag[0..1]); line_j = block(in_j); prev_0 = 0"},
"mixer": {"draw": "SplitMix64 seeded with key[0] | key[1] << 32: rot[0..7] = 1 + next() % 31, mul[0..15] = low32(next()) | 1, rc[0..15] = low32(next())", "rot": [17, 12, 20, 23, 7, 3, 27, 16], "mul": ["0xf351d601", "0xa3bb398f", "0xb5a09e35", "0x7509c9c1", "0x6bbf31e9", "0xfc849a79", "0xded91851", "0x8d9113d1", "0x0ff15225", "0x3a5bdd41", "0xab533435", "0xe1c55ad5", "0xe6d3bd0d", "0x9d9ffbbd", "0xbb2a3cf3", "0x50a7c08d"], "rc": ["0xc6892460", "0x25b7228a", "0xcd515004", "0x2846527a", "0xa6324241", "0x36e3ec53", "0x82961bac", "0x0f97ba7d", "0xb6f921a9", "0x3ada24e5", "0xde20ab91", "0x5378eeb2", "0x7d161662", "0x89353cc1", "0xb1aa03a2", "0x788acae6"], "round": "for i in 0..15: s[i] = (s[i] ^ (rc[i] + (r+1) * 0x9E3779B9)) * mul[i]; then quarter rounds on columns (0,4,8,12) (1,5,9,13) (2,6,10,14) (3,7,11,15) with rot[0..3] and diagonals (0,5,10,15) (1,6,11,12) (2,7,8,13) (3,4,9,14) with rot[4..7]", "quarter_round": "a += b; d ^= a; d = rotl(d, r1); c += d; b ^= c; b = rotl(b, r2); a += b; d ^= a; d = rotl(d, r3); c += d; b ^= c; b = rotl(b, r4)"},
"item": "s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..7: s = M_r(s); line = s[0] & 0x003fffff; s[i] ^= cache[line * 16 + i]; then s = M_8(s); item(t) = s",
"word": "dataset[w] = item(w >> 4)[w & 15]"
},
"instructions": [
{"i": 0, "op": "add", "dst": 4, "src": 5, "src2": 7, "imm": "0xea86e152", "imm2": "0x5810667a", "rot": 27, "bit": 13, "mask": 2, "width": 1, "win": 0, "off": 0},
{"i": 1, "op": "xor", "dst": 7, "src": 0, "src2": 6, "imm": "0xbaab6229", "imm2": "0xed861989", "rot": 26, "bit": 22, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 2, "op": "shfl", "dst": 3, "src": 6, "src2": 2, "imm": "0x5b623116", "imm2": "0xff12e5b2", "rot": 12, "bit": 24, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 3, "op": "sub", "dst": 4, "src": 1, "src2": 1, "imm": "0xe99741c7", "imm2": "0xf5fa5009", "rot": 1, "bit": 21, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 4, "op": "mad", "dst": 2, "src": 0, "src2": 4, "imm": "0x673c2157", "imm2": "0xee02465f", "rot": 20, "bit": 22, "mask": 2, "width": 1, "win": 0, "off": 0},
{"i": 5, "op": "rotl", "dst": 4, "src": 7, "src2": 7, "imm": "0x946f7818", "imm2": "0x45d3399e", "rot": 9, "bit": 2, "mask": 16, "width": 1, "win": 0, "off": 0},
{"i": 6, "op": "rotr", "dst": 0, "src": 2, "src2": 4, "imm": "0x5f6a0ed2", "imm2": "0x7043a636", "rot": 19, "bit": 31, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 7, "op": "load", "dst": 6, "src": 7, "src2": 1, "imm": "0x5c61dcf7", "imm2": "0x7466aa40", "rot": 19, "bit": 9, "mask": 2, "width": 1, "win": 2, "off": 1},
{"i": 8, "op": "load", "dst": 1, "src": 4, "src2": 5, "imm": "0x85668475", "imm2": "0xdb8cc483", "rot": 29, "bit": 7, "mask": 4, "width": 1, "win": 1, "off": 1},
{"i": 9, "op": "load", "dst": 1, "src": 2, "src2": 4, "imm": "0x4454980f", "imm2": "0xebd31581", "rot": 10, "bit": 28, "mask": 16, "width": 1, "win": 1, "off": 1},
{"i": 10, "op": "load", "dst": 7, "src": 0, "src2": 2, "imm": "0xe075297c", "imm2": "0x5779c44c", "rot": 10, "bit": 22, "mask": 2, "width": 1, "win": 1, "off": 1},
{"i": 11, "op": "load", "dst": 7, "src": 1, "src2": 6, "imm": "0x65aa4311", "imm2": "0x4fe48ea9", "rot": 15, "bit": 9, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 12, "op": "mul", "dst": 5, "src": 4, "src2": 2, "imm": "0x1383d3ad", "imm2": "0xf3094b29", "rot": 8, "bit": 9, "mask": 2, "width": 1, "win": 0, "off": 0},
{"i": 13, "op": "load", "dst": 7, "src": 6, "src2": 2, "imm": "0xed8a496f", "imm2": "0x3072c3c6", "rot": 28, "bit": 19, "mask": 8, "width": 1, "win": 1, "off": 1},
{"i": 14, "op": "shfl", "dst": 6, "src": 5, "src2": 0, "imm": "0x8b965b57", "imm2": "0xcfeca6c1", "rot": 12, "bit": 27, "mask": 16, "width": 1, "win": 0, "off": 0},
{"i": 15, "op": "shfl", "dst": 3, "src": 5, "src2": 3, "imm": "0x877c7586", "imm2": "0xa9cb2a03", "rot": 2, "bit": 29, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 16, "op": "or", "dst": 2, "src": 7, "src2": 3, "imm": "0xb740221a", "imm2": "0x89d38d6d", "rot": 6, "bit": 31, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 17, "op": "rotr", "dst": 0, "src": 6, "src2": 1, "imm": "0x26f3ad8a", "imm2": "0x27256f15", "rot": 18, "bit": 5, "mask": 2, "width": 1, "win": 0, "off": 0},
{"i": 18, "op": "add", "dst": 7, "src": 5, "src2": 7, "imm": "0xf66e7017", "imm2": "0xb9e3577e", "rot": 9, "bit": 12, "mask": 16, "width": 1, "win": 0, "off": 0},
{"i": 19, "op": "rotl", "dst": 7, "src": 0, "src2": 5, "imm": "0x849ae6ee", "imm2": "0x02b358f9", "rot": 24, "bit": 18, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 20, "op": "add", "dst": 6, "src": 7, "src2": 6, "imm": "0x52334d12", "imm2": "0x8c9f0ef8", "rot": 11, "bit": 23, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 21, "op": "or", "dst": 2, "src": 6, "src2": 5, "imm": "0xb1871e63", "imm2": "0xb2e40191", "rot": 5, "bit": 13, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 22, "op": "mad", "dst": 6, "src": 5, "src2": 4, "imm": "0x97df29e4", "imm2": "0xe60fea84", "rot": 11, "bit": 16, "mask": 2, "width": 1, "win": 0, "off": 0},
{"i": 23, "op": "or", "dst": 1, "src": 0, "src2": 3, "imm": "0x8f30d21d", "imm2": "0x2df685a0", "rot": 31, "bit": 0, "mask": 16, "width": 1, "win": 0, "off": 0},
{"i": 24, "op": "xor", "dst": 6, "src": 1, "src2": 2, "imm": "0xd2c4025f", "imm2": "0x5269eb4d", "rot": 31, "bit": 21, "mask": 16, "width": 1, "win": 0, "off": 0},
{"i": 25, "op": "add", "dst": 2, "src": 6, "src2": 5, "imm": "0x659fc3d3", "imm2": "0x9cec0e12", "rot": 6, "bit": 17, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 26, "op": "rotr", "dst": 7, "src": 0, "src2": 1, "imm": "0xd89ef484", "imm2": "0x20be3846", "rot": 12, "bit": 9, "mask": 16, "width": 1, "win": 0, "off": 0},
{"i": 27, "op": "mad", "dst": 4, "src": 5, "src2": 7, "imm": "0xb48420ae", "imm2": "0x3d1f2485", "rot": 14, "bit": 28, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 28, "op": "mad", "dst": 3, "src": 2, "src2": 4, "imm": "0xc5c46d76", "imm2": "0x700044b5", "rot": 22, "bit": 10, "mask": 2, "width": 1, "win": 0, "off": 0},
{"i": 29, "op": "load", "dst": 1, "src": 4, "src2": 3, "imm": "0x0fbaf177", "imm2": "0xfff4f2ed", "rot": 20, "bit": 31, "mask": 2, "width": 1, "win": 0, "off": 0},
{"i": 30, "op": "load", "dst": 2, "src": 3, "src2": 0, "imm": "0xe90eb2e5", "imm2": "0xb0f9eb79", "rot": 27, "bit": 14, "mask": 4, "width": 1, "win": 2, "off": 2},
{"i": 31, "op": "load", "dst": 1, "src": 5, "src2": 2, "imm": "0x97ba3fc3", "imm2": "0x7894e657", "rot": 3, "bit": 30, "mask": 16, "width": 1, "win": 1, "off": 1},
{"i": 32, "op": "add", "dst": 7, "src": 2, "src2": 7, "imm": "0x070888a8", "imm2": "0xe403240e", "rot": 2, "bit": 16, "mask": 2, "width": 1, "win": 0, "off": 0},
{"i": 33, "op": "add", "dst": 2, "src": 0, "src2": 5, "imm": "0xf2e46d55", "imm2": "0x29701828", "rot": 31, "bit": 28, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 34, "op": "add", "dst": 2, "src": 3, "src2": 0, "imm": "0x58f75b87", "imm2": "0x343b7aee", "rot": 12, "bit": 14, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 35, "op": "mulhi", "dst": 2, "src": 5, "src2": 6, "imm": "0x5892a9e6", "imm2": "0xc9824c94", "rot": 19, "bit": 26, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 36, "op": "xor", "dst": 4, "src": 2, "src2": 3, "imm": "0x7b5b5474", "imm2": "0x45cfc5dd", "rot": 17, "bit": 18, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 37, "op": "mul", "dst": 6, "src": 5, "src2": 7, "imm": "0xccf564a5", "imm2": "0x873ad101", "rot": 7, "bit": 11, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 38, "op": "xor", "dst": 7, "src": 0, "src2": 6, "imm": "0xcac8f06d", "imm2": "0x6b97c683", "rot": 18, "bit": 28, "mask": 2, "width": 1, "win": 0, "off": 0},
{"i": 39, "op": "add", "dst": 7, "src": 2, "src2": 4, "imm": "0xb8180e9d", "imm2": "0x32bbd117", "rot": 23, "bit": 19, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 40, "op": "rotr", "dst": 2, "src": 3, "src2": 0, "imm": "0x2d6070bc", "imm2": "0x68ff101e", "rot": 13, "bit": 18, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 41, "op": "sub", "dst": 7, "src": 0, "src2": 2, "imm": "0x0daf96ea", "imm2": "0x36f37be1", "rot": 5, "bit": 0, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 42, "op": "add", "dst": 4, "src": 3, "src2": 6, "imm": "0x6ced15b7", "imm2": "0x6df7aed4", "rot": 19, "bit": 4, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 43, "op": "shfl", "dst": 7, "src": 3, "src2": 5, "imm": "0x8ace05f3", "imm2": "0xd378ec12", "rot": 23, "bit": 10, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 44, "op": "load", "dst": 0, "src": 7, "src2": 4, "imm": "0xb0607786", "imm2": "0xc4acabbc", "rot": 13, "bit": 7, "mask": 16, "width": 1, "win": 1, "off": 1},
{"i": 45, "op": "add", "dst": 1, "src": 6, "src2": 5, "imm": "0xe09f54e9", "imm2": "0x83e825bf", "rot": 23, "bit": 14, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 46, "op": "load", "dst": 3, "src": 1, "src2": 0, "imm": "0x63cc1e4e", "imm2": "0xa1be8118", "rot": 12, "bit": 6, "mask": 2, "width": 1, "win": 2, "off": 0},
{"i": 47, "op": "load", "dst": 6, "src": 3, "src2": 7, "imm": "0x353f1d79", "imm2": "0x3b2e7456", "rot": 18, "bit": 18, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 48, "op": "mulhi", "dst": 4, "src": 2, "src2": 7, "imm": "0x00d8a3cd", "imm2": "0x231866d2", "rot": 21, "bit": 20, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 49, "op": "add", "dst": 5, "src": 0, "src2": 2, "imm": "0xa8bae6df", "imm2": "0xf572bdb9", "rot": 14, "bit": 7, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 50, "op": "shfl", "dst": 0, "src": 7, "src2": 7, "imm": "0x81ef22e1", "imm2": "0x74438fc5", "rot": 28, "bit": 18, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 51, "op": "add", "dst": 6, "src": 0, "src2": 6, "imm": "0x383b9260", "imm2": "0x11e17c61", "rot": 12, "bit": 19, "mask": 16, "width": 1, "win": 0, "off": 0},
{"i": 52, "op": "load", "dst": 5, "src": 2, "src2": 2, "imm": "0xfb84f451", "imm2": "0x11cd863e", "rot": 21, "bit": 20, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 53, "op": "rotl", "dst": 6, "src": 5, "src2": 4, "imm": "0xb1a7db6b", "imm2": "0x76686b9b", "rot": 12, "bit": 4, "mask": 16, "width": 1, "win": 0, "off": 0},
{"i": 54, "op": "rotl", "dst": 3, "src": 6, "src2": 3, "imm": "0x6f981f52", "imm2": "0xd99aeba2", "rot": 12, "bit": 27, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 55, "op": "add", "dst": 2, "src": 1, "src2": 2, "imm": "0xac6be8e3", "imm2": "0x18d67dbb", "rot": 26, "bit": 10, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 56, "op": "load", "dst": 1, "src": 4, "src2": 0, "imm": "0x7e7f6a00", "imm2": "0x6f0747da", "rot": 25, "bit": 28, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 57, "op": "add", "dst": 5, "src": 0, "src2": 4, "imm": "0xf03673fe", "imm2": "0xa75cd60d", "rot": 16, "bit": 12, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 58, "op": "load", "dst": 5, "src": 0, "src2": 2, "imm": "0x227f94a6", "imm2": "0x0e8344f9", "rot": 20, "bit": 10, "mask": 2, "width": 1, "win": 2, "off": 0},
{"i": 59, "op": "add", "dst": 1, "src": 4, "src2": 3, "imm": "0xdecd4794", "imm2": "0x8dfb96bb", "rot": 21, "bit": 7, "mask": 4, "width": 1, "win": 0, "off": 0},
{"i": 60, "op": "mulhi", "dst": 3, "src": 2, "src2": 2, "imm": "0x0dd268e0", "imm2": "0x53034ca9", "rot": 1, "bit": 8, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 61, "op": "or", "dst": 6, "src": 4, "src2": 7, "imm": "0x3a45a321", "imm2": "0x9bc59a5f", "rot": 25, "bit": 11, "mask": 1, "width": 1, "win": 0, "off": 0},
{"i": 62, "op": "shfl", "dst": 5, "src": 4, "src2": 2, "imm": "0x8f229cc1", "imm2": "0xcaac64a2", "rot": 17, "bit": 13, "mask": 8, "width": 1, "win": 0, "off": 0},
{"i": 63, "op": "load", "dst": 3, "src": 6, "src2": 6, "imm": "0xb2574178", "imm2": "0xcbcc798d", "rot": 28, "bit": 0, "mask": 16, "width": 1, "win": 1, "off": 1}
]
}

View file

@ -0,0 +1,109 @@
#include <metal_stdlib>
using namespace metal;
#define MASK 0x0fffffffu
constant uint SEEDW[8] = { 0x667d0fbdu, 0x7b8e5963u, 0x31c67e5eu, 0x4529ddc6u, 0xef19d6d8u, 0xaccf6211u, 0xda0aed32u, 0xabc6df31u };
inline uint splitmix32(uint x) {
x ^= x >> 16; x *= 0x7feb352du;
x ^= x >> 15; x *= 0x846ca68bu;
x ^= x >> 16;
return x;
}
inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31
inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
inline uint ds_elem(uint i, uint d0, uint d1) {
uint x = i ^ d0;
x *= 0x9E3779B1u; x ^= x >> 15;
x += d1;
x *= 0x85EBCA77u; x ^= x >> 13;
x *= 0xC2B2AE3Du; x ^= x >> 16;
return x;
}
kernel void igneum_hash(device const uint* dataset [[buffer(0)]],
device ulong* out [[buffer(1)]],
constant uint& baseNonce [[buffer(2)]],
uint gid [[thread_position_in_grid]]) {
uint nonce = baseNonce + gid;
uint r0, r1, r2, r3, r4, r5, r6, r7;
{ uint x = nonce ^ SEEDW[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ SEEDW[1]; }
{ uint x = nonce ^ SEEDW[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ SEEDW[2]; }
{ uint x = nonce ^ SEEDW[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ SEEDW[3]; }
{ uint x = nonce ^ SEEDW[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ SEEDW[4]; }
{ uint x = nonce ^ SEEDW[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ SEEDW[5]; }
{ uint x = nonce ^ SEEDW[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ SEEDW[6]; }
{ uint x = nonce ^ SEEDW[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ SEEDW[7]; }
{ uint x = nonce ^ SEEDW[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ SEEDW[0]; }
for (uint it = 0u; it < 8u; ++it) {
uint sel = r0;
r4 = r4 + r5 + select(0xea86e152u, 0x5810667au, ((sel >> 13u) & 1u) != 0u); // 0
r7 = r7 ^ r0; // 1
r3 = r3 ^ simd_shuffle_xor(r6, (ushort)1); // 2
r4 = r4 - r1; // 3
r2 = r0 * r4 + r2; // 4
r4 = rotl_imm(r4, 9u); // 5
r0 = rotr_var(r0, r2); // 6
r6 = r6 ^ dataset[((rotl_imm(r7 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x04000000u) & MASK]; // 7
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 8
r1 = r1 ^ dataset[((rotl_imm(r2 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 9
r7 = r7 ^ dataset[((rotl_imm(r0 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 10
r7 = r7 ^ dataset[((rotl_imm(r1 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 11
r5 = r5 * r4; // 12
r7 = r7 ^ dataset[((rotl_imm(r6 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 13
r6 = r6 ^ simd_shuffle_xor(r5, (ushort)16); // 14
r3 = r3 ^ simd_shuffle_xor(r5, (ushort)1); // 15
r2 = r2 | r7; // 16
r0 = rotr_var(r0, r6); // 17
r7 = r7 + r5 + select(0xf66e7017u, 0xb9e3577eu, ((sel >> 12u) & 1u) != 0u); // 18
r7 = rotl_imm(r7, 24u); // 19
r6 = r6 + r7 + select(0x52334d12u, 0x8c9f0ef8u, ((sel >> 23u) & 1u) != 0u); // 20
r2 = r2 | r6; // 21
r6 = r5 * r4 + r6; // 22
r1 = r1 | r0; // 23
r6 = r6 ^ r1; // 24
r2 = r2 + r6 + select(0x659fc3d3u, 0x9cec0e12u, ((sel >> 17u) & 1u) != 0u); // 25
r7 = rotr_var(r7, r0); // 26
r4 = r5 * r7 + r4; // 27
r3 = r2 * r4 + r3; // 28
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 29
r2 = r2 ^ dataset[((rotl_imm(r3 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x08000000u) & MASK]; // 30
r1 = r1 ^ dataset[((rotl_imm(r5 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 31
r7 = r7 + r2 + select(0x070888a8u, 0xe403240eu, ((sel >> 16u) & 1u) != 0u); // 32
r2 = r2 + r0 + select(0xf2e46d55u, 0x29701828u, ((sel >> 28u) & 1u) != 0u); // 33
r2 = r2 + r3 + select(0x58f75b87u, 0x343b7aeeu, ((sel >> 14u) & 1u) != 0u); // 34
r2 = mulhi(r2, r5); // 35
r4 = r4 ^ r2; // 36
r6 = r6 * r5; // 37
r7 = r7 ^ r0; // 38
r7 = r7 + r2 + select(0xb8180e9du, 0x32bbd117u, ((sel >> 19u) & 1u) != 0u); // 39
r2 = rotr_var(r2, r3); // 40
r7 = r7 - r0; // 41
r4 = r4 + r3 + select(0x6ced15b7u, 0x6df7aed4u, ((sel >> 4u) & 1u) != 0u); // 42
r7 = r7 ^ simd_shuffle_xor(r3, (ushort)4); // 43
r0 = r0 ^ dataset[((rotl_imm(r7 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 44
r1 = r1 + r6 + select(0xe09f54e9u, 0x83e825bfu, ((sel >> 14u) & 1u) != 0u); // 45
r3 = r3 ^ dataset[((rotl_imm(r1 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 46
r6 = r6 ^ dataset[((rotl_imm(r3 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 47
r4 = mulhi(r4, r2); // 48
r5 = r5 + r0 + select(0xa8bae6dfu, 0xf572bdb9u, ((sel >> 7u) & 1u) != 0u); // 49
r0 = r0 ^ simd_shuffle_xor(r7, (ushort)4); // 50
r6 = r6 + r0 + select(0x383b9260u, 0x11e17c61u, ((sel >> 19u) & 1u) != 0u); // 51
r5 = r5 ^ dataset[((rotl_imm(r2 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 52
r6 = rotl_imm(r6, 12u); // 53
r3 = rotl_imm(r3, 12u); // 54
r2 = r2 + r1 + select(0xac6be8e3u, 0x18d67dbbu, ((sel >> 10u) & 1u) != 0u); // 55
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 56
r5 = r5 + r0 + select(0xf03673feu, 0xa75cd60du, ((sel >> 12u) & 1u) != 0u); // 57
r5 = r5 ^ dataset[((rotl_imm(r0 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 58
r1 = r1 + r4 + select(0xdecd4794u, 0x8dfb96bbu, ((sel >> 7u) & 1u) != 0u); // 59
r3 = mulhi(r3, r2); // 60
r6 = r6 | r4; // 61
r5 = r5 ^ simd_shuffle_xor(r4, (ushort)8); // 62
r3 = r3 ^ dataset[((rotl_imm(r6 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 63
}
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((ulong)hi << 32) | (ulong)lo;
}

View file

@ -0,0 +1,111 @@
#include <metal_stdlib>
using namespace metal;
#define MASK 0x0fffffffu
constant uint SEEDW[8] = { 0x667d0fbdu, 0x7b8e5963u, 0x31c67e5eu, 0x4529ddc6u, 0xef19d6d8u, 0xaccf6211u, 0xda0aed32u, 0xabc6df31u };
inline uint splitmix32(uint x) {
x ^= x >> 16; x *= 0x7feb352du;
x ^= x >> 15; x *= 0x846ca68bu;
x ^= x >> 16;
return x;
}
inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31
inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
inline uint ds_elem(uint i, uint d0, uint d1) {
uint x = i ^ d0;
x *= 0x9E3779B1u; x ^= x >> 15;
x += d1;
x *= 0x85EBCA77u; x ^= x >> 13;
x *= 0xC2B2AE3Du; x ^= x >> 16;
return x;
}
// Header-bound variant: the init words come from buffer 3 (bind.rs), not from SEEDW.
kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]],
device ulong* out [[buffer(1)]],
constant uint& baseNonce [[buffer(2)]],
constant uint* initw [[buffer(3)]],
uint gid [[thread_position_in_grid]]) {
uint nonce = baseNonce + gid;
uint r0, r1, r2, r3, r4, r5, r6, r7;
{ uint x = nonce ^ initw[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ initw[1]; }
{ uint x = nonce ^ initw[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ initw[2]; }
{ uint x = nonce ^ initw[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ initw[3]; }
{ uint x = nonce ^ initw[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ initw[4]; }
{ uint x = nonce ^ initw[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ initw[5]; }
{ uint x = nonce ^ initw[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ initw[6]; }
{ uint x = nonce ^ initw[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ initw[7]; }
{ uint x = nonce ^ initw[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ initw[0]; }
for (uint it = 0u; it < 8u; ++it) {
uint sel = r0;
r4 = r4 + r5 + select(0xea86e152u, 0x5810667au, ((sel >> 13u) & 1u) != 0u); // 0
r7 = r7 ^ r0; // 1
r3 = r3 ^ simd_shuffle_xor(r6, (ushort)1); // 2
r4 = r4 - r1; // 3
r2 = r0 * r4 + r2; // 4
r4 = rotl_imm(r4, 9u); // 5
r0 = rotr_var(r0, r2); // 6
r6 = r6 ^ dataset[((rotl_imm(r7 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x04000000u) & MASK]; // 7
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 8
r1 = r1 ^ dataset[((rotl_imm(r2 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 9
r7 = r7 ^ dataset[((rotl_imm(r0 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 10
r7 = r7 ^ dataset[((rotl_imm(r1 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 11
r5 = r5 * r4; // 12
r7 = r7 ^ dataset[((rotl_imm(r6 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 13
r6 = r6 ^ simd_shuffle_xor(r5, (ushort)16); // 14
r3 = r3 ^ simd_shuffle_xor(r5, (ushort)1); // 15
r2 = r2 | r7; // 16
r0 = rotr_var(r0, r6); // 17
r7 = r7 + r5 + select(0xf66e7017u, 0xb9e3577eu, ((sel >> 12u) & 1u) != 0u); // 18
r7 = rotl_imm(r7, 24u); // 19
r6 = r6 + r7 + select(0x52334d12u, 0x8c9f0ef8u, ((sel >> 23u) & 1u) != 0u); // 20
r2 = r2 | r6; // 21
r6 = r5 * r4 + r6; // 22
r1 = r1 | r0; // 23
r6 = r6 ^ r1; // 24
r2 = r2 + r6 + select(0x659fc3d3u, 0x9cec0e12u, ((sel >> 17u) & 1u) != 0u); // 25
r7 = rotr_var(r7, r0); // 26
r4 = r5 * r7 + r4; // 27
r3 = r2 * r4 + r3; // 28
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 29
r2 = r2 ^ dataset[((rotl_imm(r3 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x08000000u) & MASK]; // 30
r1 = r1 ^ dataset[((rotl_imm(r5 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 31
r7 = r7 + r2 + select(0x070888a8u, 0xe403240eu, ((sel >> 16u) & 1u) != 0u); // 32
r2 = r2 + r0 + select(0xf2e46d55u, 0x29701828u, ((sel >> 28u) & 1u) != 0u); // 33
r2 = r2 + r3 + select(0x58f75b87u, 0x343b7aeeu, ((sel >> 14u) & 1u) != 0u); // 34
r2 = mulhi(r2, r5); // 35
r4 = r4 ^ r2; // 36
r6 = r6 * r5; // 37
r7 = r7 ^ r0; // 38
r7 = r7 + r2 + select(0xb8180e9du, 0x32bbd117u, ((sel >> 19u) & 1u) != 0u); // 39
r2 = rotr_var(r2, r3); // 40
r7 = r7 - r0; // 41
r4 = r4 + r3 + select(0x6ced15b7u, 0x6df7aed4u, ((sel >> 4u) & 1u) != 0u); // 42
r7 = r7 ^ simd_shuffle_xor(r3, (ushort)4); // 43
r0 = r0 ^ dataset[((rotl_imm(r7 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 44
r1 = r1 + r6 + select(0xe09f54e9u, 0x83e825bfu, ((sel >> 14u) & 1u) != 0u); // 45
r3 = r3 ^ dataset[((rotl_imm(r1 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 46
r6 = r6 ^ dataset[((rotl_imm(r3 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 47
r4 = mulhi(r4, r2); // 48
r5 = r5 + r0 + select(0xa8bae6dfu, 0xf572bdb9u, ((sel >> 7u) & 1u) != 0u); // 49
r0 = r0 ^ simd_shuffle_xor(r7, (ushort)4); // 50
r6 = r6 + r0 + select(0x383b9260u, 0x11e17c61u, ((sel >> 19u) & 1u) != 0u); // 51
r5 = r5 ^ dataset[((rotl_imm(r2 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 52
r6 = rotl_imm(r6, 12u); // 53
r3 = rotl_imm(r3, 12u); // 54
r2 = r2 + r1 + select(0xac6be8e3u, 0x18d67dbbu, ((sel >> 10u) & 1u) != 0u); // 55
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 56
r5 = r5 + r0 + select(0xf03673feu, 0xa75cd60du, ((sel >> 12u) & 1u) != 0u); // 57
r5 = r5 ^ dataset[((rotl_imm(r0 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 58
r1 = r1 + r4 + select(0xdecd4794u, 0x8dfb96bbu, ((sel >> 7u) & 1u) != 0u); // 59
r3 = mulhi(r3, r2); // 60
r6 = r6 | r4; // 61
r5 = r5 ^ simd_shuffle_xor(r4, (ushort)8); // 62
r3 = r3 ^ dataset[((rotl_imm(r6 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 63
}
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((ulong)hi << 32) | (ulong)lo;
}

View file

@ -0,0 +1,5 @@
epoch_seed_hex edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07
day_seed_hex 69676e65756d2d6461792ffa50000000000000
epoch_index 0
day_index 20730
daa_score 0

View file

@ -0,0 +1,57 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v3, memory-hard dataset
#pragma once
#ifdef __cplusplus
#include <cstdint>
#else
#include <stdint.h>
#endif
#define IGNEUM_VEC_WARPS 3
static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u };
static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = {
{ // base nonce 0
0x8e32500ee72aa1f8ull, 0x710f1522fd52e28eull, 0x62fd41712fd5bd3eull, 0x3837ee0ed93c4015ull, 0x0338e814e0e8a4faull, 0xaab705ee7514af71ull, 0x4a3e143121ed41a6ull, 0xe843b01bee434ef6ull,
0x4f9974ee1f331bc6ull, 0x1e089106d4d04581ull, 0x8a986409754a08b8ull, 0x54bf0b893560e130ull, 0x1346ddc8074de7ddull, 0x5fdc54666289e864ull, 0xe9ed965ed22a9ea1ull, 0x7915337fc840db87ull,
0xaab900aea6b7c2cfull, 0x81d768fa492b1cf2ull, 0xfd62bcf1021629dfull, 0x6280dc76e5125eddull, 0x17055f35fc9a4ed3ull, 0xead760de60ae3e84ull, 0x14ea3ac017d95d57ull, 0x5f09e499a0096d83ull,
0xdf7c24624de53e02ull, 0x10e39b014ea6d6c6ull, 0xfcfac36ecf54fb8cull, 0x8b510ae1be613404ull, 0xa9a07c9f9fa58801ull, 0xba88c22112b80fd0ull, 0x673c5bf0fbb48d35ull, 0xeef7f536b1e0a810ull
},
{ // base nonce 4096
0x88df895d8d8539ffull, 0x052fb10eab66d76dull, 0x534663c60e7de2e5ull, 0xc76cb6763d97b1beull, 0x24c28c52fdbc5c86ull, 0x46a9b9ee877a5764ull, 0xcc000e068d59a8e0ull, 0x7e634381ee2664e7ull,
0x0e8d46d76bb5f04aull, 0x2a7caab33aaf260full, 0x11805382c57768fcull, 0x1050d656bfd66847ull, 0x6dac200ef67b302full, 0xfcebdae4acc594c2ull, 0x59ddcc381267cfb0ull, 0x15e9e31cd7347f92ull,
0xad4fb91f9c31a9d1ull, 0x01bdb6cd682f2d3aull, 0xd94caecd87d9bc10ull, 0x3bc7dd917ded4591ull, 0x1dc97ba9f06c33cbull, 0xb503609bab3d4397ull, 0xa6001b284f75c5e4ull, 0x56981fe3d9beddccull,
0xc1e3cf75a47fd944ull, 0x0d44d93176a1ab11ull, 0x31d24d65e0706657ull, 0xff7363ff80342341ull, 0x8060c156f3d360beull, 0xa75837c717e7eadaull, 0x86f2a525a3cf8d13ull, 0xac0d7a25d491c048ull
},
{ // base nonce 1000000
0x77e307c1d3c4a45dull, 0xc46db2c1181e7d29ull, 0x60d540ae23190d93ull, 0x03a6005b827a3b1eull, 0x549bac8bd1c1e201ull, 0x780dfc346d705778ull, 0xcfcbebfdbb68e6f4ull, 0x3c5ed8998530360dull,
0x2b88fe019717d013ull, 0xb70795cd85b32d67ull, 0x6decdbef0be191d7ull, 0x8f22242400f2cb78ull, 0xbbd28d1419989b8aull, 0xe16cffacddfb8292ull, 0xd8248122cb40a2a3ull, 0x2522a2642fae7251ull,
0xfeeea1da8b23ad1bull, 0x3be311612545a761ull, 0x4ddb4b56a8275591ull, 0x7cb9edbe2572ab8bull, 0xe9a8ab8f94c06690ull, 0xd853e9b7bc76760dull, 0x03b2b843379ca0e4ull, 0x1954de4a7d95cb4full,
0xad25e818a9aa212bull, 0x14be913efb5627feull, 0x2006c5d0f1ee99feull, 0xb1ce75c39bd2f557ull, 0xcd2f842898802fecull, 0xf44486cc9aa01340ull, 0x09a89237e25250eeull, 0xb342a4880eb580b3ull
}
};
// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455).
static const uint32_t IGNEUM_DS_HEAD[16] = {
0x3dd50b1fu, 0x93123701u, 0xcdd88b45u, 0x2c37c6b0u, 0x48edec90u, 0x3a7d2407u, 0xb03a0ac0u, 0x10e2c6a5u,
0xf46429c0u, 0x4b1f3b3au, 0xd69dcaffu, 0xa48808fdu, 0xc8b2db13u, 0xc555587fu, 0x53042758u, 0x1a71bda6u
};
static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u;
static const uint32_t IGNEUM_DS_LAST = 0xf7b7180eu;
// 64 sampled dataset words (index, value) computed on the Mac.
#define IGNEUM_DS_SAMPLES 64
static const uint32_t IGNEUM_DS_SAMPLE_INDEX[IGNEUM_DS_SAMPLES] = {
59471966u, 217795994u, 208353206u, 42483309u, 172547758u, 148076330u, 183853158u, 214389424u, 267488061u, 169781097u, 184093494u, 153880993u, 84977930u, 46426879u, 3093825u, 225364072u, 44593546u, 260713159u, 168250303u, 52384140u, 223401610u, 45554030u, 95410555u, 175039924u, 79171087u, 267580473u, 24168642u, 37981670u, 171551130u, 195559979u, 204611762u, 140997658u, 138925853u, 86637313u, 20736778u, 219665210u, 160430336u, 264654675u, 8013395u, 228945585u, 213884386u, 104419827u, 44185464u, 142737231u, 99284897u, 132475900u, 61861762u, 132056166u, 262388043u, 91878046u, 117353561u, 124768597u, 71352993u, 190698941u, 46055428u, 55281366u, 165145231u, 106810753u, 171985651u, 232085256u, 159510492u, 40072060u, 209107596u, 39023794u
};
static const uint32_t IGNEUM_DS_SAMPLE_VALUE[IGNEUM_DS_SAMPLES] = {
0x7f292d0cu, 0x73dd96cfu, 0x202b4ad2u, 0xf6d8038fu, 0x7835659bu, 0x03a3a00du, 0x022d1e04u, 0x89957f84u, 0xe3451874u, 0xc82651dau, 0x29759e8cu, 0xa7a8cca3u, 0x45098d58u, 0x60d618eau, 0xf42d0e2cu, 0xc3546ab2u, 0xa7de6336u, 0xd592fdf4u, 0x083d78d1u, 0x64bd8824u, 0x8a66c547u, 0x8bc97d55u, 0x1411385du, 0x485f6fc0u, 0xc8142e6cu, 0xd8c1b7ebu, 0x6671b76bu, 0x63d4b24eu, 0xd622091bu, 0xa6821c04u, 0x8c2949d6u, 0x6b064395u, 0x5fa8e992u, 0xf49195b8u, 0x990254beu, 0xc5e6169du, 0xfff9c259u, 0x0e3c45dau, 0x42d69754u, 0xe23a1172u, 0x6d7d42e0u, 0xc30766efu, 0xe8abd145u, 0x11348b7cu, 0x837ccca1u, 0xe65b14c7u, 0x550f2d75u, 0xfedbcd25u, 0xb7430a6au, 0x5bbbdc45u, 0x2f81bc7du, 0xebbd5147u, 0xe991ee42u, 0x74d6d9abu, 0xc0eb32eau, 0x589a8daau, 0xd50d6a45u, 0x3f0ec2ccu, 0x31377e5eu, 0xe22a256fu, 0x52e144d0u, 0xa3fd1811u, 0xd3be259du, 0xc9bbc553u
};
// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^26 words.
static const uint32_t IGNEUM_CACHE_HEAD[16] = {
0xebd9055cu, 0x7eaf6b21u, 0x9b610855u, 0x134a8562u, 0xb10059bau, 0x7f459e58u, 0x8f3c7873u, 0x3523e609u,
0x75b8f4cau, 0x750d31b4u, 0x0dae0782u, 0x26f10015u, 0x7ba87a41u, 0x0efd8543u, 0x691d6368u, 0x8929d967u
};
static const uint32_t IGNEUM_CACHE_LAST[16] = {
0x51474edfu, 0xc3cfba93u, 0xf21454cfu, 0x9b79baacu, 0x4d4fce23u, 0xd701dfd5u, 0x37357ab3u, 0x1be693fau,
0xa701cb3bu, 0x7467c620u, 0x428184e8u, 0xf4010df0u, 0xe33aa88fu, 0xc78d6d62u, 0xae9ba9c0u, 0xf4bba97eu
};
static const uint64_t IGNEUM_CACHE_FNV64 = 0x448274a57f508cbcull;

View file

@ -0,0 +1,36 @@
{
"seed": "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000",
"day": "bytes:69676e65756d2d6461792ffa50000000000000",
"dataset_mode": "memory-hard",
"dataset_log2_words": 28,
"mask": "0x0fffffff",
"lanes": 32,
"source": "igneum-pow (Rust) CPU interpreter, generator v3, memory-hard dataset",
"warps": [
{"base_nonce": 0, "expected": [
"0x8e32500ee72aa1f8", "0x710f1522fd52e28e", "0x62fd41712fd5bd3e", "0x3837ee0ed93c4015", "0x0338e814e0e8a4fa", "0xaab705ee7514af71", "0x4a3e143121ed41a6", "0xe843b01bee434ef6",
"0x4f9974ee1f331bc6", "0x1e089106d4d04581", "0x8a986409754a08b8", "0x54bf0b893560e130", "0x1346ddc8074de7dd", "0x5fdc54666289e864", "0xe9ed965ed22a9ea1", "0x7915337fc840db87",
"0xaab900aea6b7c2cf", "0x81d768fa492b1cf2", "0xfd62bcf1021629df", "0x6280dc76e5125edd", "0x17055f35fc9a4ed3", "0xead760de60ae3e84", "0x14ea3ac017d95d57", "0x5f09e499a0096d83",
"0xdf7c24624de53e02", "0x10e39b014ea6d6c6", "0xfcfac36ecf54fb8c", "0x8b510ae1be613404", "0xa9a07c9f9fa58801", "0xba88c22112b80fd0", "0x673c5bf0fbb48d35", "0xeef7f536b1e0a810"
]},
{"base_nonce": 4096, "expected": [
"0x88df895d8d8539ff", "0x052fb10eab66d76d", "0x534663c60e7de2e5", "0xc76cb6763d97b1be", "0x24c28c52fdbc5c86", "0x46a9b9ee877a5764", "0xcc000e068d59a8e0", "0x7e634381ee2664e7",
"0x0e8d46d76bb5f04a", "0x2a7caab33aaf260f", "0x11805382c57768fc", "0x1050d656bfd66847", "0x6dac200ef67b302f", "0xfcebdae4acc594c2", "0x59ddcc381267cfb0", "0x15e9e31cd7347f92",
"0xad4fb91f9c31a9d1", "0x01bdb6cd682f2d3a", "0xd94caecd87d9bc10", "0x3bc7dd917ded4591", "0x1dc97ba9f06c33cb", "0xb503609bab3d4397", "0xa6001b284f75c5e4", "0x56981fe3d9beddcc",
"0xc1e3cf75a47fd944", "0x0d44d93176a1ab11", "0x31d24d65e0706657", "0xff7363ff80342341", "0x8060c156f3d360be", "0xa75837c717e7eada", "0x86f2a525a3cf8d13", "0xac0d7a25d491c048"
]},
{"base_nonce": 1000000, "expected": [
"0x77e307c1d3c4a45d", "0xc46db2c1181e7d29", "0x60d540ae23190d93", "0x03a6005b827a3b1e", "0x549bac8bd1c1e201", "0x780dfc346d705778", "0xcfcbebfdbb68e6f4", "0x3c5ed8998530360d",
"0x2b88fe019717d013", "0xb70795cd85b32d67", "0x6decdbef0be191d7", "0x8f22242400f2cb78", "0xbbd28d1419989b8a", "0xe16cffacddfb8292", "0xd8248122cb40a2a3", "0x2522a2642fae7251",
"0xfeeea1da8b23ad1b", "0x3be311612545a761", "0x4ddb4b56a8275591", "0x7cb9edbe2572ab8b", "0xe9a8ab8f94c06690", "0xd853e9b7bc76760d", "0x03b2b843379ca0e4", "0x1954de4a7d95cb4f",
"0xad25e818a9aa212b", "0x14be913efb5627fe", "0x2006c5d0f1ee99fe", "0xb1ce75c39bd2f557", "0xcd2f842898802fec", "0xf44486cc9aa01340", "0x09a89237e25250ee", "0xb342a4880eb580b3"
]}
],
"dataset_head": ["0x3dd50b1f", "0x93123701", "0xcdd88b45", "0x2c37c6b0", "0x48edec90", "0x3a7d2407", "0xb03a0ac0", "0x10e2c6a5", "0xf46429c0", "0x4b1f3b3a", "0xd69dcaff", "0xa48808fd", "0xc8b2db13", "0xc555587f", "0x53042758", "0x1a71bda6"],
"dataset_last_index": 268435455,
"dataset_last": "0xf7b7180e",
"dataset_samples": [{"index": 59471966, "value": "0x7f292d0c"}, {"index": 217795994, "value": "0x73dd96cf"}, {"index": 208353206, "value": "0x202b4ad2"}, {"index": 42483309, "value": "0xf6d8038f"}, {"index": 172547758, "value": "0x7835659b"}, {"index": 148076330, "value": "0x03a3a00d"}, {"index": 183853158, "value": "0x022d1e04"}, {"index": 214389424, "value": "0x89957f84"}, {"index": 267488061, "value": "0xe3451874"}, {"index": 169781097, "value": "0xc82651da"}, {"index": 184093494, "value": "0x29759e8c"}, {"index": 153880993, "value": "0xa7a8cca3"}, {"index": 84977930, "value": "0x45098d58"}, {"index": 46426879, "value": "0x60d618ea"}, {"index": 3093825, "value": "0xf42d0e2c"}, {"index": 225364072, "value": "0xc3546ab2"}, {"index": 44593546, "value": "0xa7de6336"}, {"index": 260713159, "value": "0xd592fdf4"}, {"index": 168250303, "value": "0x083d78d1"}, {"index": 52384140, "value": "0x64bd8824"}, {"index": 223401610, "value": "0x8a66c547"}, {"index": 45554030, "value": "0x8bc97d55"}, {"index": 95410555, "value": "0x1411385d"}, {"index": 175039924, "value": "0x485f6fc0"}, {"index": 79171087, "value": "0xc8142e6c"}, {"index": 267580473, "value": "0xd8c1b7eb"}, {"index": 24168642, "value": "0x6671b76b"}, {"index": 37981670, "value": "0x63d4b24e"}, {"index": 171551130, "value": "0xd622091b"}, {"index": 195559979, "value": "0xa6821c04"}, {"index": 204611762, "value": "0x8c2949d6"}, {"index": 140997658, "value": "0x6b064395"}, {"index": 138925853, "value": "0x5fa8e992"}, {"index": 86637313, "value": "0xf49195b8"}, {"index": 20736778, "value": "0x990254be"}, {"index": 219665210, "value": "0xc5e6169d"}, {"index": 160430336, "value": "0xfff9c259"}, {"index": 264654675, "value": "0x0e3c45da"}, {"index": 8013395, "value": "0x42d69754"}, {"index": 228945585, "value": "0xe23a1172"}, {"index": 213884386, "value": "0x6d7d42e0"}, {"index": 104419827, "value": "0xc30766ef"}, {"index": 44185464, "value": "0xe8abd145"}, {"index": 142737231, "value": "0x11348b7c"}, {"index": 99284897, "value": "0x837ccca1"}, {"index": 132475900, "value": "0xe65b14c7"}, {"index": 61861762, "value": "0x550f2d75"}, {"index": 132056166, "value": "0xfedbcd25"}, {"index": 262388043, "value": "0xb7430a6a"}, {"index": 91878046, "value": "0x5bbbdc45"}, {"index": 117353561, "value": "0x2f81bc7d"}, {"index": 124768597, "value": "0xebbd5147"}, {"index": 71352993, "value": "0xe991ee42"}, {"index": 190698941, "value": "0x74d6d9ab"}, {"index": 46055428, "value": "0xc0eb32ea"}, {"index": 55281366, "value": "0x589a8daa"}, {"index": 165145231, "value": "0xd50d6a45"}, {"index": 106810753, "value": "0x3f0ec2cc"}, {"index": 171985651, "value": "0x31377e5e"}, {"index": 232085256, "value": "0xe22a256f"}, {"index": 159510492, "value": "0x52e144d0"}, {"index": 40072060, "value": "0xa3fd1811"}, {"index": 209107596, "value": "0xd3be259d"}, {"index": 39023794, "value": "0xc9bbc553"}],
"cache_head": ["0xebd9055c", "0x7eaf6b21", "0x9b610855", "0x134a8562", "0xb10059ba", "0x7f459e58", "0x8f3c7873", "0x3523e609", "0x75b8f4ca", "0x750d31b4", "0x0dae0782", "0x26f10015", "0x7ba87a41", "0x0efd8543", "0x691d6368", "0x8929d967"],
"cache_last_line": ["0x51474edf", "0xc3cfba93", "0xf21454cf", "0x9b79baac", "0x4d4fce23", "0xd701dfd5", "0x37357ab3", "0x1be693fa", "0xa701cb3b", "0x7467c620", "0x428184e8", "0xf4010df0", "0xe33aa88f", "0xc78d6d62", "0xae9ba9c0", "0xf4bba97e"],
"cache_fnv1a64": "0x448274a57f508cbc"
}

View file

@ -0,0 +1,281 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal).
// Built from source at runtime by proto-opencl/host.c, which passes these defines:
// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit)
// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default)
// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32
// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition
// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units;
// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32.
#ifndef IGNEUM_GROUP
#define IGNEUM_GROUP 32
#endif
#ifndef IGNEUM_EXCHANGE
#define IGNEUM_EXCHANGE 0
#endif
#ifdef __OPENCL_VERSION__
#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1)))
#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n]
#if IGNEUM_EXCHANGE == 1
#ifdef cl_khr_subgroups
#pragma OPENCL EXTENSION cl_khr_subgroups : enable
#endif
#ifdef cl_khr_subgroup_shuffle
#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable
#endif
#elif IGNEUM_EXCHANGE == 2
#pragma OPENCL EXTENSION cl_intel_subgroups : enable
#endif
#else
// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros.
#include "emu_opencl.h"
#endif
#if IGNEUM_EXCHANGE == 1
#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m))
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
#elif IGNEUM_EXCHANGE == 2
#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m))
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
#else
// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per
// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane
// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's
// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier.
#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; }
#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; }
#endif
static inline uint splitmix32(uint x) {
x ^= x >> 16; x *= 0x7feb352du;
x ^= x >> 15; x *= 0x846ca68bu;
x ^= x >> 16;
return x;
}
// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32.
static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); }
// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x.
static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); }
static inline uint ds_elem(uint i, uint d0, uint d1) {
uint x = i ^ d0;
x *= 0x9E3779B1u; x ^= x >> 15;
x += d1;
x *= 0x85EBCA77u; x ^= x >> 13;
x *= 0xC2B2AE3Du; x ^= x >> 16;
return x;
}
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
#define MH_CACHE_LINE_MASK 0x003fffffu
#define MH_SEGMENT_LINES 64u
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
// y = ChaCha12 core(x) + x
static inline void mh_chacha_block(const uint* x, uint* y) {
for (uint i = 0u; i < 16u; ++i) y[i] = x[i];
for (uint r = 0u; r < 6u; ++r) {
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
}
for (uint i = 0u; i < 16u; ++i) y[i] += x[i];
}
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
static inline void mh_cache_segment(__global uint* cache, uint seg) {
uint prev[16]; uint x[16]; uint y[16];
for (uint i = 0u; i < 16u; ++i) prev[i] = 0u;
for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) {
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
x[4] = 0xceed56d7u ^ prev[4];
x[5] = 0x9ba270d2u ^ prev[5];
x[6] = 0x82caab2du ^ prev[6];
x[7] = 0x81ebce0eu ^ prev[7];
x[8] = 0x12b6ecf1u ^ prev[8];
x[9] = 0xd0f3fd7cu ^ prev[9];
x[10] = 0xd872eefeu ^ prev[10];
x[11] = 0xc158c7bdu ^ prev[11];
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
mh_chacha_block(x, y);
__global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
}
}
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
static inline void mh_mixer(uint* s, uint rk) {
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
}
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
static inline void mh_item(__global const uint* cache, uint t, uint* s) {
s[0] = 0xceed56d7u;
s[1] = 0x9ba270d2u;
s[2] = 0x82caab2du;
s[3] = 0x81ebce0eu;
s[4] = 0x12b6ecf1u;
s[5] = 0xd0f3fd7cu;
s[6] = 0xd872eefeu;
s[7] = 0xc158c7bdu;
s[8] = t * 0xf351d601u + 0xc6892460u;
s[9] = t * 0xa3bb398fu + 0x25b7228au;
s[10] = t * 0xb5a09e35u + 0xcd515004u;
s[11] = t * 0x7509c9c1u + 0x2846527au;
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
s[14] = t * 0xded91851u + 0x82961bacu;
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
for (uint r = 0u; r < 8u; ++r) {
mh_mixer(s, 0x9E3779B9u * (r + 1u));
__global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i];
}
mh_mixer(s, 0x9E3779B9u * 9u);
}
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 0 2 10 13 of w.
static inline uint mh_j(uint w) { return ((w >> 0u) & 1u) | (((w >> 2u) & 1u) << 1) | (((w >> 10u) & 1u) << 2) | (((w >> 13u) & 1u) << 3); }
static inline uint mh_t(uint w) { w = (w & 0x00001fffu) | ((w >> 14u) << 13u); w = (w & 0x000003ffu) | ((w >> 11u) << 10u); w = (w & 0x00000003u) | ((w >> 3u) << 2u); w = (w & 0x00000000u) | ((w >> 1u) << 0u); return w; }
static inline uint mh_addr(uint t, uint j) { uint w = t; w = ((w >> 0u) << 1u) | (w & 0x00000000u) | (((j >> 0u) & 1u) << 0u); w = ((w >> 2u) << 3u) | (w & 0x00000003u) | (((j >> 1u) & 1u) << 2u); w = ((w >> 10u) << 11u) | (w & 0x000003ffu) | (((j >> 2u) & 1u) << 10u); w = ((w >> 13u) << 14u) | (w & 0x00001fffu) | (((j >> 3u) & 1u) << 13u); return w; }
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }
// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item.
// The same constants as memhard.h in this pack (one emitter, three dialects).
__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) {
uint seg = (uint)get_global_id(0);
if (seg < nSegments) mh_cache_segment(cache, seg);
}
__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) {
uint t = (uint)get_global_id(0);
if (t < nItems) {
uint s[16];
mh_item(cache, t, s);
for (uint i = 0u; i < 16u; ++i) ds[(ulong)mh_addr(t, i)] = s[i];
}
}
// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the
// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and
// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all).
IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) {
uint gid = (uint)get_global_id(0);
uint lid = (uint)get_local_id(0);
uint nonce = baseNonce + gid;
uint r0, r1, r2, r3, r4, r5, r6, r7;
#if IGNEUM_EXCHANGE == 0
IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
uint xk = 0u;
#else
(void)lid;
#endif
{ uint x = nonce ^ 0x667d0fbdu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x7b8e5963u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
{ uint x = nonce ^ 0x7b8e5963u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x31c67e5eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
{ uint x = nonce ^ 0x31c67e5eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4529ddc6u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
{ uint x = nonce ^ 0x4529ddc6u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef19d6d8u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
{ uint x = nonce ^ 0xef19d6d8u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xaccf6211u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
{ uint x = nonce ^ 0xaccf6211u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xda0aed32u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
{ uint x = nonce ^ 0xda0aed32u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xabc6df31u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
{ uint x = nonce ^ 0xabc6df31u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x667d0fbdu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
for (uint it = 0u; it < 8u; ++it) {
uint sel = r0;
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
r7 = r7 ^ r0; // 1 xor
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 1u); r3 = r3 ^ t_; } // 2 shfl
r4 = r4 - r1; // 3 sub
r2 = r0 * r4 + r2; // 4 mad
r4 = rotl_imm(r4, 9u); // 5 rotl
r0 = rotr_var(r0, r2); // 6 rotr
r6 = r6 ^ ds[((rotl_imm(r7 * 0x03ac37adu, 22u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
r1 = r1 ^ ds[((rotl_imm(r4 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
r1 = r1 ^ ds[((rotl_imm(r2 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
r7 = r7 ^ ds[((rotl_imm(r0 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
r7 = r7 ^ ds[((rotl_imm(r1 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
r5 = r5 * r4; // 12 mul
r7 = r7 ^ ds[((rotl_imm(r6 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 16u); r6 = r6 ^ t_; } // 14 shfl
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r3 = r3 ^ t_; } // 15 shfl
r2 = r2 | r7; // 16 or
r0 = rotr_var(r0, r6); // 17 rotr
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
r7 = rotl_imm(r7, 24u); // 19 rotl
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
r2 = r2 | r6; // 21 or
r6 = r5 * r4 + r6; // 22 mad
r1 = r1 | r0; // 23 or
r6 = r6 ^ r1; // 24 xor
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
r7 = rotr_var(r7, r0); // 26 rotr
r4 = r5 * r7 + r4; // 27 mad
r3 = r2 * r4 + r3; // 28 mad
r1 = r1 ^ ds[((rotl_imm(r4 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
r2 = r2 ^ ds[((rotl_imm(r3 * 0x03ac37adu, 22u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
r1 = r1 ^ ds[((rotl_imm(r5 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
r2 = mul_hi(r2, r5); // 35 mulhi
r4 = r4 ^ r2; // 36 xor
r6 = r6 * r5; // 37 mul
r7 = r7 ^ r0; // 38 xor
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
r2 = rotr_var(r2, r3); // 40 rotr
r7 = r7 - r0; // 41 sub
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 43 shfl
r0 = r0 ^ ds[((rotl_imm(r7 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
r3 = r3 ^ ds[((rotl_imm(r1 * 0x03ac37adu, 22u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
r6 = r6 ^ ds[((rotl_imm(r3 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
r4 = mul_hi(r4, r2); // 48 mulhi
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r0 = r0 ^ t_; } // 50 shfl
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
r5 = r5 ^ ds[((rotl_imm(r2 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
r6 = rotl_imm(r6, 12u); // 53 rotl
r3 = rotl_imm(r3, 12u); // 54 rotl
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
r1 = r1 ^ ds[((rotl_imm(r4 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
r5 = r5 ^ ds[((rotl_imm(r0 * 0x03ac37adu, 22u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
r3 = mul_hi(r3, r2); // 60 mulhi
r6 = r6 | r4; // 61 or
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r5 = r5 ^ t_; } // 62 shfl
r3 = r3 ^ ds[((rotl_imm(r6 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
}
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((ulong)hi << 32) | (ulong)lo;
}
#if IGNEUM_EXCHANGE != 0
// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the
// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a
// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md.
IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) {
if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); }
}
#endif

View file

@ -0,0 +1,163 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal).
// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC.
#include <cuda_runtime.h>
#include <cstdint>
#include "program.h"
#include "memhard.h"
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
x ^= x >> 16; x *= 0x7feb352du;
x ^= x >> 15; x *= 0x846ca68bu;
x ^= x >> 16;
return x;
}
// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31.
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x.
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) {
uint32_t x = i ^ d0;
x *= 0x9E3779B1u; x ^= x >> 15;
x += d1;
x *= 0x85EBCA77u; x ^= x >> 13;
x *= 0xC2B2AE3Du; x ^= x >> 16;
return x;
}
// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item.
// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host.
__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) {
uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x;
if (seg < nSegments) mh_cache_segment(cache, seg);
}
__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
uint32_t t = blockIdx.x * blockDim.x + threadIdx.x;
if (t < nItems) {
uint32_t s[16];
mh_item(cache, t, s);
for (uint32_t i = 0u; i < 16u; ++i) ds[(size_t)mh_addr(t, i)] = s[i];
}
}
// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every
// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a
// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid.
__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) {
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
uint32_t nonce = baseNonce + gid;
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
{ uint32_t x = nonce ^ 0x667d0fbdu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x7b8e5963u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
{ uint32_t x = nonce ^ 0x7b8e5963u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x31c67e5eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
{ uint32_t x = nonce ^ 0x31c67e5eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4529ddc6u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
{ uint32_t x = nonce ^ 0x4529ddc6u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef19d6d8u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
{ uint32_t x = nonce ^ 0xef19d6d8u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xaccf6211u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
{ uint32_t x = nonce ^ 0xaccf6211u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xda0aed32u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
{ uint32_t x = nonce ^ 0xda0aed32u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xabc6df31u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
{ uint32_t x = nonce ^ 0xabc6df31u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x667d0fbdu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
for (uint32_t it = 0u; it < 8u; ++it) {
uint32_t sel = r0;
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
r7 = r7 ^ r0; // 1 xor
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r6, 1); // 2 shfl
r4 = r4 - r1; // 3 sub
r2 = r0 * r4 + r2; // 4 mad
r4 = rotl_imm(r4, 9u); // 5 rotl
r0 = rotr_var(r0, r2); // 6 rotr
r6 = r6 ^ ds[((rotl_imm(r7 * 0x03ac37adu, 22u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
r1 = r1 ^ ds[((rotl_imm(r4 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
r1 = r1 ^ ds[((rotl_imm(r2 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
r7 = r7 ^ ds[((rotl_imm(r0 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
r7 = r7 ^ ds[((rotl_imm(r1 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
r5 = r5 * r4; // 12 mul
r7 = r7 ^ ds[((rotl_imm(r6 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r5, 16); // 14 shfl
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r5, 1); // 15 shfl
r2 = r2 | r7; // 16 or
r0 = rotr_var(r0, r6); // 17 rotr
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
r7 = rotl_imm(r7, 24u); // 19 rotl
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
r2 = r2 | r6; // 21 or
r6 = r5 * r4 + r6; // 22 mad
r1 = r1 | r0; // 23 or
r6 = r6 ^ r1; // 24 xor
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
r7 = rotr_var(r7, r0); // 26 rotr
r4 = r5 * r7 + r4; // 27 mad
r3 = r2 * r4 + r3; // 28 mad
r1 = r1 ^ ds[((rotl_imm(r4 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
r2 = r2 ^ ds[((rotl_imm(r3 * 0x03ac37adu, 22u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
r1 = r1 ^ ds[((rotl_imm(r5 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
r2 = __umulhi(r2, r5); // 35 mulhi
r4 = r4 ^ r2; // 36 xor
r6 = r6 * r5; // 37 mul
r7 = r7 ^ r0; // 38 xor
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
r2 = rotr_var(r2, r3); // 40 rotr
r7 = r7 - r0; // 41 sub
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 43 shfl
r0 = r0 ^ ds[((rotl_imm(r7 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
r3 = r3 ^ ds[((rotl_imm(r1 * 0x03ac37adu, 22u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
r6 = r6 ^ ds[((rotl_imm(r3 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
r4 = __umulhi(r4, r2); // 48 mulhi
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r7, 4); // 50 shfl
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
r5 = r5 ^ ds[((rotl_imm(r2 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
r6 = rotl_imm(r6, 12u); // 53 rotl
r3 = rotl_imm(r3, 12u); // 54 rotl
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
r1 = r1 ^ ds[((rotl_imm(r4 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
r5 = r5 ^ ds[((rotl_imm(r0 * 0x03ac37adu, 22u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
r3 = __umulhi(r3, r2); // 60 mulhi
r6 = r6 | r4; // 61 or
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 62 shfl
r3 = r3 ^ ds[((rotl_imm(r6 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
}
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
}
// Host-side launch wrappers. Declared in program.h, called from host.cu.
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) {
if (nSegments == 0u) return cudaErrorInvalidValue;
uint32_t block = 256u;
uint32_t grid = (nSegments + block - 1u) / block;
igneum_cache_fill<<<grid, block>>>(cache, nSegments);
return cudaGetLastError();
}
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
if (nItems == 0u) return cudaErrorInvalidValue;
uint32_t block = 256u;
uint32_t grid = (nItems + block - 1u) / block;
igneum_build<<<grid, block>>>(ds, cache, nItems);
return cudaGetLastError();
}
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
uint32_t nonces, uint32_t blockWarps) {
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
uint32_t block = 32u * blockWarps;
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
igneum_hash<<<nonces / block, block>>>(ds, out, baseNonce, mask);
return cudaGetLastError();
}
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
cudaFuncAttributes attr;
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash);
if (e != cudaSuccess) return e;
*numRegs = attr.numRegs;
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0);
}

View file

@ -0,0 +1,375 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal).
// Built from source at runtime by proto-opencl/host.c, which passes these defines:
// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit)
// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default)
// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32
// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition
// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units;
// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32.
#ifndef IGNEUM_GROUP
#define IGNEUM_GROUP 32
#endif
#ifndef IGNEUM_EXCHANGE
#define IGNEUM_EXCHANGE 0
#endif
#ifdef __OPENCL_VERSION__
#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1)))
#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n]
#if IGNEUM_EXCHANGE == 1
#ifdef cl_khr_subgroups
#pragma OPENCL EXTENSION cl_khr_subgroups : enable
#endif
#ifdef cl_khr_subgroup_shuffle
#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable
#endif
#elif IGNEUM_EXCHANGE == 2
#pragma OPENCL EXTENSION cl_intel_subgroups : enable
#endif
#else
// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros.
#include "emu_opencl.h"
#endif
#if IGNEUM_EXCHANGE == 1
#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m))
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
#elif IGNEUM_EXCHANGE == 2
#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m))
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
#else
// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per
// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane
// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's
// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier.
#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; }
#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; }
#endif
static inline uint splitmix32(uint x) {
x ^= x >> 16; x *= 0x7feb352du;
x ^= x >> 15; x *= 0x846ca68bu;
x ^= x >> 16;
return x;
}
// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32.
static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); }
// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x.
static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); }
static inline uint ds_elem(uint i, uint d0, uint d1) {
uint x = i ^ d0;
x *= 0x9E3779B1u; x ^= x >> 15;
x += d1;
x *= 0x85EBCA77u; x ^= x >> 13;
x *= 0xC2B2AE3Du; x ^= x >> 16;
return x;
}
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
#define MH_CACHE_LINE_MASK 0x003fffffu
#define MH_SEGMENT_LINES 64u
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
// y = ChaCha12 core(x) + x
static inline void mh_chacha_block(const uint* x, uint* y) {
for (uint i = 0u; i < 16u; ++i) y[i] = x[i];
for (uint r = 0u; r < 6u; ++r) {
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
}
for (uint i = 0u; i < 16u; ++i) y[i] += x[i];
}
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
static inline void mh_cache_segment(__global uint* cache, uint seg) {
uint prev[16]; uint x[16]; uint y[16];
for (uint i = 0u; i < 16u; ++i) prev[i] = 0u;
for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) {
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
x[4] = 0xceed56d7u ^ prev[4];
x[5] = 0x9ba270d2u ^ prev[5];
x[6] = 0x82caab2du ^ prev[6];
x[7] = 0x81ebce0eu ^ prev[7];
x[8] = 0x12b6ecf1u ^ prev[8];
x[9] = 0xd0f3fd7cu ^ prev[9];
x[10] = 0xd872eefeu ^ prev[10];
x[11] = 0xc158c7bdu ^ prev[11];
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
mh_chacha_block(x, y);
__global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
}
}
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
static inline void mh_mixer(uint* s, uint rk) {
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
}
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
static inline void mh_item(__global const uint* cache, uint t, uint* s) {
s[0] = 0xceed56d7u;
s[1] = 0x9ba270d2u;
s[2] = 0x82caab2du;
s[3] = 0x81ebce0eu;
s[4] = 0x12b6ecf1u;
s[5] = 0xd0f3fd7cu;
s[6] = 0xd872eefeu;
s[7] = 0xc158c7bdu;
s[8] = t * 0xf351d601u + 0xc6892460u;
s[9] = t * 0xa3bb398fu + 0x25b7228au;
s[10] = t * 0xb5a09e35u + 0xcd515004u;
s[11] = t * 0x7509c9c1u + 0x2846527au;
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
s[14] = t * 0xded91851u + 0x82961bacu;
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
for (uint r = 0u; r < 8u; ++r) {
mh_mixer(s, 0x9E3779B9u * (r + 1u));
__global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i];
}
mh_mixer(s, 0x9E3779B9u * 9u);
}
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 0 2 10 13 of w.
static inline uint mh_j(uint w) { return ((w >> 0u) & 1u) | (((w >> 2u) & 1u) << 1) | (((w >> 10u) & 1u) << 2) | (((w >> 13u) & 1u) << 3); }
static inline uint mh_t(uint w) { w = (w & 0x00001fffu) | ((w >> 14u) << 13u); w = (w & 0x000003ffu) | ((w >> 11u) << 10u); w = (w & 0x00000003u) | ((w >> 3u) << 2u); w = (w & 0x00000000u) | ((w >> 1u) << 0u); return w; }
static inline uint mh_addr(uint t, uint j) { uint w = t; w = ((w >> 0u) << 1u) | (w & 0x00000000u) | (((j >> 0u) & 1u) << 0u); w = ((w >> 2u) << 3u) | (w & 0x00000003u) | (((j >> 1u) & 1u) << 2u); w = ((w >> 10u) << 11u) | (w & 0x000003ffu) | (((j >> 2u) & 1u) << 10u); w = ((w >> 13u) << 14u) | (w & 0x00001fffu) | (((j >> 3u) & 1u) << 13u); return w; }
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }
// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item.
// The same constants as memhard.h in this pack (one emitter, three dialects).
__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) {
uint seg = (uint)get_global_id(0);
if (seg < nSegments) mh_cache_segment(cache, seg);
}
__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) {
uint t = (uint)get_global_id(0);
if (t < nItems) {
uint s[16];
mh_item(cache, t, s);
for (uint i = 0u; i < 16u; ++i) ds[(ulong)mh_addr(t, i)] = s[i];
}
}
// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the
// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and
// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all).
IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) {
uint gid = (uint)get_global_id(0);
uint lid = (uint)get_local_id(0);
uint nonce = baseNonce + gid;
uint r0, r1, r2, r3, r4, r5, r6, r7;
#if IGNEUM_EXCHANGE == 0
IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
uint xk = 0u;
#else
(void)lid;
#endif
{ uint x = nonce ^ 0x667d0fbdu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x7b8e5963u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
{ uint x = nonce ^ 0x7b8e5963u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x31c67e5eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
{ uint x = nonce ^ 0x31c67e5eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4529ddc6u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
{ uint x = nonce ^ 0x4529ddc6u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef19d6d8u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
{ uint x = nonce ^ 0xef19d6d8u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xaccf6211u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
{ uint x = nonce ^ 0xaccf6211u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xda0aed32u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
{ uint x = nonce ^ 0xda0aed32u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xabc6df31u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
{ uint x = nonce ^ 0xabc6df31u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x667d0fbdu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
for (uint it = 0u; it < 8u; ++it) {
uint sel = r0;
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
r7 = r7 ^ r0; // 1 xor
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 1u); r3 = r3 ^ t_; } // 2 shfl
r4 = r4 - r1; // 3 sub
r2 = r0 * r4 + r2; // 4 mad
r4 = rotl_imm(r4, 9u); // 5 rotl
r0 = rotr_var(r0, r2); // 6 rotr
r6 = r6 ^ ds[((rotl_imm(r7 * 0x03ac37adu, 22u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
r1 = r1 ^ ds[((rotl_imm(r4 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
r1 = r1 ^ ds[((rotl_imm(r2 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
r7 = r7 ^ ds[((rotl_imm(r0 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
r7 = r7 ^ ds[((rotl_imm(r1 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
r5 = r5 * r4; // 12 mul
r7 = r7 ^ ds[((rotl_imm(r6 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 16u); r6 = r6 ^ t_; } // 14 shfl
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r3 = r3 ^ t_; } // 15 shfl
r2 = r2 | r7; // 16 or
r0 = rotr_var(r0, r6); // 17 rotr
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
r7 = rotl_imm(r7, 24u); // 19 rotl
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
r2 = r2 | r6; // 21 or
r6 = r5 * r4 + r6; // 22 mad
r1 = r1 | r0; // 23 or
r6 = r6 ^ r1; // 24 xor
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
r7 = rotr_var(r7, r0); // 26 rotr
r4 = r5 * r7 + r4; // 27 mad
r3 = r2 * r4 + r3; // 28 mad
r1 = r1 ^ ds[((rotl_imm(r4 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
r2 = r2 ^ ds[((rotl_imm(r3 * 0x03ac37adu, 22u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
r1 = r1 ^ ds[((rotl_imm(r5 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
r2 = mul_hi(r2, r5); // 35 mulhi
r4 = r4 ^ r2; // 36 xor
r6 = r6 * r5; // 37 mul
r7 = r7 ^ r0; // 38 xor
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
r2 = rotr_var(r2, r3); // 40 rotr
r7 = r7 - r0; // 41 sub
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 43 shfl
r0 = r0 ^ ds[((rotl_imm(r7 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
r3 = r3 ^ ds[((rotl_imm(r1 * 0x03ac37adu, 22u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
r6 = r6 ^ ds[((rotl_imm(r3 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
r4 = mul_hi(r4, r2); // 48 mulhi
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r0 = r0 ^ t_; } // 50 shfl
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
r5 = r5 ^ ds[((rotl_imm(r2 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
r6 = rotl_imm(r6, 12u); // 53 rotl
r3 = rotl_imm(r3, 12u); // 54 rotl
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
r1 = r1 ^ ds[((rotl_imm(r4 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
r5 = r5 ^ ds[((rotl_imm(r0 * 0x03ac37adu, 22u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
r3 = mul_hi(r3, r2); // 60 mulhi
r6 = r6 | r4; // 61 or
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r5 = r5 ^ t_; } // 62 shfl
r3 = r3 ^ ds[((rotl_imm(r6 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
}
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((ulong)hi << 32) | (ulong)lo;
}
#if IGNEUM_EXCHANGE != 0
// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the
// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a
// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md.
IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) {
if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); }
}
#endif
// Header-bound variant (bind.rs): the init words come from initw, not SEEDW. Same body as igneum_hash.
IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw) {
uint gid = (uint)get_global_id(0);
uint lid = (uint)get_local_id(0);
uint nonce = baseNonce + gid;
uint r0, r1, r2, r3, r4, r5, r6, r7;
uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7];
#if IGNEUM_EXCHANGE == 0
IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
uint xk = 0u;
#else
(void)lid;
#endif
{ uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; }
{ uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; }
{ uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; }
{ uint x = nonce ^ iw3; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw4; }
{ uint x = nonce ^ iw4; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw5; }
{ uint x = nonce ^ iw5; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw6; }
{ uint x = nonce ^ iw6; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw7; }
{ uint x = nonce ^ iw7; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw0; }
for (uint it = 0u; it < 8u; ++it) {
uint sel = r0;
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
r7 = r7 ^ r0; // 1 xor
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 1u); r3 = r3 ^ t_; } // 2 shfl
r4 = r4 - r1; // 3 sub
r2 = r0 * r4 + r2; // 4 mad
r4 = rotl_imm(r4, 9u); // 5 rotl
r0 = rotr_var(r0, r2); // 6 rotr
r6 = r6 ^ ds[((rotl_imm(r7 * 0x03ac37adu, 22u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
r1 = r1 ^ ds[((rotl_imm(r4 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
r1 = r1 ^ ds[((rotl_imm(r2 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
r7 = r7 ^ ds[((rotl_imm(r0 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
r7 = r7 ^ ds[((rotl_imm(r1 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
r5 = r5 * r4; // 12 mul
r7 = r7 ^ ds[((rotl_imm(r6 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 16u); r6 = r6 ^ t_; } // 14 shfl
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r3 = r3 ^ t_; } // 15 shfl
r2 = r2 | r7; // 16 or
r0 = rotr_var(r0, r6); // 17 rotr
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
r7 = rotl_imm(r7, 24u); // 19 rotl
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
r2 = r2 | r6; // 21 or
r6 = r5 * r4 + r6; // 22 mad
r1 = r1 | r0; // 23 or
r6 = r6 ^ r1; // 24 xor
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
r7 = rotr_var(r7, r0); // 26 rotr
r4 = r5 * r7 + r4; // 27 mad
r3 = r2 * r4 + r3; // 28 mad
r1 = r1 ^ ds[((rotl_imm(r4 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
r2 = r2 ^ ds[((rotl_imm(r3 * 0x03ac37adu, 22u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
r1 = r1 ^ ds[((rotl_imm(r5 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
r2 = mul_hi(r2, r5); // 35 mulhi
r4 = r4 ^ r2; // 36 xor
r6 = r6 * r5; // 37 mul
r7 = r7 ^ r0; // 38 xor
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
r2 = rotr_var(r2, r3); // 40 rotr
r7 = r7 - r0; // 41 sub
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 43 shfl
r0 = r0 ^ ds[((rotl_imm(r7 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
r3 = r3 ^ ds[((rotl_imm(r1 * 0x03ac37adu, 22u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
r6 = r6 ^ ds[((rotl_imm(r3 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
r4 = mul_hi(r4, r2); // 48 mulhi
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r0 = r0 ^ t_; } // 50 shfl
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
r5 = r5 ^ ds[((rotl_imm(r2 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
r6 = rotl_imm(r6, 12u); // 53 rotl
r3 = rotl_imm(r3, 12u); // 54 rotl
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
r1 = r1 ^ ds[((rotl_imm(r4 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
r5 = r5 ^ ds[((rotl_imm(r0 * 0x03ac37adu, 22u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
r3 = mul_hi(r3, r2); // 60 mulhi
r6 = r6 | r4; // 61 or
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r5 = r5 ^ t_; } // 62 shfl
r3 = r3 ^ ds[((rotl_imm(r6 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
}
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((ulong)hi << 32) | (ulong)lo;
}

Some files were not shown because too many files have changed in this diff Show more