era layout (Counter ASIC 2.0 layers 4 and 8) behind the class flag, on the ca2-v3 seam: the 7-draw era stream from E_n (stride, interleave, width pinned at 4 B), per-site window draws (dataset, half, quarter at a 256 MiB floor), the strided windowed load address in the interpreter, the acceptance mirror and the three emitters, the interleaved dataset layout riding with the program (memhard::Layout, mh_t/mh_j/mh_addr, Epoch::dataset_word), V3_CLASS with the era drawn inside by generate_from_seed_bytes_program_class, --era / --era-widths on the CLI, six class v3 era packs (packs-ca2-era), tests, the design doc, host.cu/host.c deriving host words through the pack's mh_word, emu/test-layout.sh, the PC 1 playbook
Squashed from six commits (tag ca2-era-pre-squash) for one merge. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
88dafbc92e
commit
b105a5533c
118 changed files with 12573 additions and 822 deletions
|
|
@ -1524,3 +1524,50 @@ What is measured: one BLS12-381 aggregate signature over 16 summed G1 keys plus
|
|||
| on, split 90 s | v3 | 0 / 2 | none / 3 | 278 / 265 | apart | none | 3 on n0 | 2 (n0 reconnected 6 s after the heal, A's chain at about 58 DAA, inside the table) |
|
||||
|
||||
Reading (the NEW finding, ledger C4). With the module off GHOSTDAG alone converges on the heavier chain and the losing side's records re-determine (F24 works when the chain moves). With the module on the overlay holds during the split (A, with 30% of the frozen table, locks nothing; B locks 7 and 8) and then fails at the heal in the shipped node: B's certificates for blocks off n0's chain are "kept pending until the chain decides (no lock at this index)", n0's chain never decides because GHOSTDAG keeps its heavier tip and nothing turns the certificate into a fork-choice constraint, and once n0's last lock (index 7, DAA 209) is one window old (DAA 329) the frozen table stops applying on A's chain ("no frozen table (no lock on this chain inside the window)"), A's two keys are 100% of A's own window (B's post-cut blocks are red there) and n0 locks 10, 11, 12 alone; B's certificates for 10 and 11 then log CONFLICTING on n0 (n0 log, 17:27:04 to 17:29:54 BST). A finality fork from a 96-s honest partition, no attacker, table intact at the heal; the 150-s run and the v2 control end the same way. The spec's fork choice ("GHOSTDAG among tips through all certified checkpoints", 3.5) is therefore implemented only for certificates over blocks already on the node's chain. Fix named in the ledger entry: verify an off-chain certificate against the table at its own block and let it constrain fork choice (a certificate-driven reorg), then re-determine. Raw: `scratchpad fud-a/c4-results-*.md`, node logs `c4-on90-tmp/`, `c4-v2-control-tmp/`.
|
||||
|
||||
## 5 October 2026 (night), read width of the lottery hash: 4, 16 and 64-byte loads, a per-load mix, a written scratch; three cards (gate 1 experiment, cryptographer)
|
||||
|
||||
Branch `readwidth` (commits 019b014, b970dda, 4badcee, a9e002c, d0018cf and the entry commit); plan and recommendation in `docs/plans/read-width.md`. Nothing here changes consensus: every class sits behind `igneum-pow --class` and the default class is generator version 2 byte for byte (`igneum-pow/tests/packs.rs` passes on the four pinned packs after every commit). Question (Josh, after "the 9070 XT on the eGPU" above): would wider reads keep the latency-bound random-access property while closing the vendor gap. Additions from the coordinator: a per-load width drawn from an era-fixed mix, and a written per-warp scratch (measurement only, no soundness claim).
|
||||
|
||||
**What a class does** (`igneum-pow/src/generator.rs` `LoadClass`, `verify::fold_words`, the three emitters): a load of W words reads the W-word-aligned address `(src AND MASK) AND NOT (W - 1)` and folds every word into `dst` (`x = dst ^ w0; x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]`); W = 1 is the lottery hash exactly (`w4` = pack `bcc1248b10cc90f2`). A mix class draws W per load with one extra `below(100)` roll per instruction. A scratch class `scr<k>k<kb>` turns `k` of the 16 memory slots into read-modify-writes of a 16-byte slot of the lane's share of a `kb` KiB per-warp scratch (kernels run persistent warps, one per block or work-group; a slot reads as a seed-and-base fill until the unit writes it, behind a per-unit tag). Program ids carry the class. Dependent chain and 32-lane unit unchanged.
|
||||
|
||||
**Correctness**: 23 packs (`proto-cuda/packs-readwidth/`, Rust CPU reference vectors). Every pack passed its three vector units and the cache and dataset checks on Metal (M5 Max, `proto-metal/packbench`), Apple OpenCL (`--bench-pack`), the RTX 5090 (NVRTC, `igneum-worker-cuda --bench`) and, the 16 width and mix packs, the RX 9070 XT (`igneum-worker-opencl --bench-pack`); the 2^24 batch fingerprints agree across all four runtimes on every pack (for example w16 `e7c890445b47af60`, w64 `836e56e7d496e980`, mixB-2 `a18ac73098c76007`). The clang CUDA emulation (w16, w64, w64x4, mixA-0, mixB-0: 3 of 3 units standalone and 2 of 2 in batch at 2 warps per block) and the clang OpenCL emulation (the same five plus scr2k32 and scr8k128, sub-group 32 and, width packs, wave64 with sub-group shuffles) pass with equal fingerprints per configuration. Acceptance rule on the classes: 60 candidates per class, rejection 0 to 14 of 60 (w16 and w64 as v2; the mixes the same; the scratch classes' distinct-address bound now covers dataset loads only, since a 64-slot lane scratch repeats slots by design). CPU verifier (M5 Max, one core, avg of 50 units, `igneum-pow bench --class`): v2 0.604 ms, w16 0.610, w64 0.630, w64x4 0.160, mix50-35-15 0.620, mix25-50-25 0.614, scr0k32 0.600 (1.004 on a loaded re-run), scr2k32 0.657, scr4k32 0.458, scr8k32 0.317, scr2k128 0.535, scr4k128 0.458, scr8k128 0.311; per hash divide by 32. The wide reads cost the verifier nothing (a lane's words lie in one item); scratch ops replace item derivations and make it cheaper.
|
||||
|
||||
**Probes** (`--memprobe`, dependent random reads at 1024 MiB, G reads/s, best over lanes in flight; 4 B = the hash's pattern; the 5090 and 9070 XT with the card off in the app, the Mac through Apple OpenCL under a load average of 5 to 10):
|
||||
|
||||
| Card | 4 B chase | 16 B | 64 B | 64 B as GB/s | coalesced stream GB/s (rated) | integer chain |
|
||||
|---|---|---|---|---|---|---|
|
||||
| RTX 5090 (PC 2, CUDA) | 17.5 to 18.2 | 18.0 to 19.9 | 9.1 to 15.7 (9.1 at 4 M lanes) | 584 | 1,579 (1,792) | 39.0 T op/s |
|
||||
| RX 9070 XT (PC 1, eGPU, OpenCL) | 2.42 to 2.66 | 2.43 to 2.73 | 2.47 to 2.87 | 158 | 636 (640) | 6.2 T op/s |
|
||||
| Apple M5 Max (Apple OpenCL, approximate) | 3.50 | 3.51 | 3.51 | 225 | 522 | |
|
||||
|
||||
Reading: on the 9070 XT and the M5 Max a 64-byte dependent read costs exactly what a 4-byte one costs (the line is fetched either way); on the 5090 a 64-byte read costs about two 4-byte reads (two 32-byte sectors) and the 64 B chase at full occupancy sits at 584 GB/s, a third of the stream.
|
||||
|
||||
**Hash rates** (5 timed dispatches of 2^24 nonces after a warm-up; Metal and the 9070 XT by device time, the 5090 by wall time around the stream sync; the PC cards switched off in the app for the run and restored, PC 1's 5090 and the integrated chip kept mining; the Mac under other agents' builds, load 4 to 9, so its absolute numbers carry that; the share = measured / (the card's probe ceiling at the class's widths / loads per hash)):
|
||||
|
||||
| Class | dataset B/hash | RTX 5090 MH/s (share) | RX 9070 XT MH/s (share) | M5 Max Metal MH/s (share) | 5090 / 9070 |
|
||||
|---|---|---|---|---|---|
|
||||
| v2 (w4, the lottery hash) | 512 | 136.1 (0.96) | 18.15 (0.87) | 27.74 (1.01) | 7.5x |
|
||||
| w16 | 2,048 | 139.8 (0.90) | 17.90 (0.84) | 28.26 (1.03) | 7.8x |
|
||||
| w64 | 8,192 | 71.9 (0.58) | 17.59 (0.78) | 28.27 (1.03) | 4.1x |
|
||||
| w64x4 (32 loads) | 2,048 | 275.3 (0.56) | 75.19 (0.84) | 109.7 (1.00) | 3.7x |
|
||||
| mix50-35-15, 6 programs: min / median / max (spread of median) | 1,664 to 3,680 | 99.5 / 114.2 / 121.0 (18.8%) | 17.45 / 18.76 / 18.83 (7.4%) | 25.36 / 27.26 / 28.43 (11.3%) | 6.1x |
|
||||
| mix25-50-25, 6 programs | 2,240 to 5,024 | 95.9 / 107.3 / 119.8 (22.3%) | 17.84 / 18.45 / 18.85 (5.5%) | 23.21 / 24.68 / 25.21 (8.1%) | 5.8x |
|
||||
|
||||
Scratch (variant 5; N persistent warps; 5090: 2,048 warps launched against a resident capacity of 4,080 = 24 blocks/SM x 1 warp/block x 170 SMs at `--block-warps 1`, the occupancy query unchanged by the allocation (24 before and after); Metal: 2,048 to 16,384 warps swept, best shown; arena = N x per-warp size; the whole working set = 1 GiB dataset + 256 MiB cache + 128 MiB output + arena, under 2 GB on every row):
|
||||
|
||||
| Class (k of 16 slots, KiB per warp) | scratch ops/hash | dataset B/hash | RTX 5090 MH/s (vs scr0, share) | M5 Max Metal MH/s (vs scr0) | RX 9070 XT MH/s | 5090 arena / working set |
|
||||
|---|---|---|---|---|---|---|
|
||||
| scr0k32 (control, persistent loop, no RMW) | 0 | 512 | 139.1 (0, 0.98) | 28.25 (0) | 17.88 (control, 0.86) | 64 MiB / 1.4 GiB |
|
||||
| scr2k32 (12.5%) | 16 | 448 | 114.4 (-18%, 0.80) | 26.14 (-7%) | 14.65 (-18%) | 64 MiB / 1.4 GiB |
|
||||
| scr4k32 (25%) | 32 | 384 | 109.8 (-21%, 0.76) | 31.74 (+12%) | 14.00 (-22%) | 64 MiB / 1.4 GiB |
|
||||
| scr8k32 (50%) | 64 | 256 | 122.1 (-12%, 0.82) | 49.08 (+74%) | 14.17 (-21%) | 64 MiB / 1.4 GiB |
|
||||
| scr2k128 (12.5%) | 16 | 448 | 110.1 (-21%, 0.77) | 26.24 (-7%) | 14.07 (-21%) | 256 MiB / 1.6 GiB |
|
||||
| scr4k128 (25%) | 32 | 384 | 98.0 (-30%, 0.68) | 28.08 (-1%) | 13.14 (-27%) | 256 MiB / 1.6 GiB |
|
||||
| scr8k128 (50%) | 64 | 256 | 72.8 (-48%, 0.49) | 35.44 (+25%) | 12.03 (-33%) | 256 MiB / 1.6 GiB |
|
||||
|
||||
The 9070 XT rows are 2,048 persistent warps (4,096 within 1 percent), arena 64 MiB at 32 KiB and 256 MiB at 128 KiB, working set 1.4 and 1.6 GiB; its control (17.88, the persistent loop) equals its v2 rate (18.15) within 2 percent, and every RMW share costs it 18 to 33 percent: on AMD a scratch op is a dependent 16-byte read plus a write into a region the 64 MB Infinity Cache does not hold for 2,048 warps, so it is memory work there as on the 5090, not the cached op it is on Apple. Apple OpenCL on the same scratch packs (wall time, `--bench-pack --warps 2048`): scr0k32 27.85, scr2k32 28.58, scr4k32 32.43, scr8k32 47.93, scr2k128 25.67, scr4k128 27.24, scr8k128 32.93 MH/s, the Metal shape within 4 percent, fingerprints equal. Bytes moved per scratch op: 16 read + 16 written (the tag word included); per hash at 50 percent, 1,024 read + 1,024 written beside 256 of dataset reads. The 5090 at 4,096 launched warps (above its 4,080 resident) lost 2 to 26 percent (scr8k32 90.0 MH/s), so the rows above are the in-capacity launch.
|
||||
|
||||
**Readings.** (1) Same count, wider: the vendor gap does not move at 16 B (7.8x) because on the 9070 XT a 4-byte read already costs a 64-byte line and on the 5090 a 16-byte read costs one 32-byte sector, the same as 4 bytes: the memory systems do identical work, only the fold's input grows. At 64 B the gap closes to 4.1x, entirely by the 5090 losing half its rate (its share falls to 0.58 and its DRAM traffic reaches 589 GB/s, 37 percent of the stream: bandwidth, not latency, bounds it), while the 9070 XT and the M5 Max do not move. (2) Fewer, wider (w64x4): 3.7x, but every card runs 4x faster because the dependent chain is 32 loads long instead of 128; the 5090 sits at a 0.56 share (bandwidth), so a chip with more bandwidth per dollar than a GPU gains, which is the Ethash shape the design avoids. (3) The mix: the hour-to-hour spread is 7 to 22 percent of the median per card (the 5090 the widest, because its 64-byte loads are the expensive ones and their count per program runs 2 to 8 of 16); the programs with many 64-byte loads (mixA-3, mixA-5, mixB-2) are the slow hours on the 5090 and the fast ones nowhere. (4) The scratch: on the 5090 every RMW share costs 12 to 48 percent against the persistent control, the 32 KiB arena less than the 128 KiB one (the smaller arena, 64 MiB over 2,048 warps, sits inside the 96 MB L2); on the M5 Max the 32 KiB rows are FASTER than the control (+12 and +74 percent at 25 and 50 percent), because the arena (128 MiB over 4,096 warps) lives in the chip's caches and a scratch op is cheaper than a dataset read, so replacing dataset loads raises the rate: the scratch at these sizes is not memory work on Apple and is partly cached on NVIDIA. The chip row for these variants comes from the ca2-soundness branch; what this entry gives is the GPU cost and the share. (5) Latency-bound shares: v2 0.87 to 1.01 on the three cards, w16 0.84 to 1.03, w64 0.58 (5090) and 0.78 (9070 XT); the Mac's shares above 1 are an Apple OpenCL probe under load against a Metal rate.
|
||||
|
||||
Jobs: `run-readwidth-5090-20261005` and `run-readwidth-9070-20261005` (probes; the packs refused for their string seeds, fixed in a9e002c), `run-readwidth-5090-20261005c`, `run-readwidth-9070-20261005c` (benches), `run-readwidth-9070-scratch-20261005d` (the scratch packs after the `__local` fix d0018cf, AMD's compiler requires the exchange buffer at the kernel's outermost scope); read back with `node tools/jobs.mjs <id> --all`. Mac commands and logs: `docs/plans/read-width.md` section 3. The worker exes for the jobs: `proto-cuda/nvrtc/build-windows.sh` on this branch (mingw), sha256 of the CUDA one `6f46336f...defe1`.
|
||||
|
|
|
|||
177
docs/plans/era-layout.md
Normal file
177
docs/plans/era-layout.md
Normal file
|
|
@ -0,0 +1,177 @@
|
|||
# Era layout: table layout and working set drawn per era and per program (Counter ASIC 2.0, layers 4 and 8)
|
||||
|
||||
5 October 2026, branch `ca2-era`, worker "ca2-era". Status: Designed and Implemented behind the class flag (`LoadClass::era`, not the lottery hash); nothing here changes the default generator, the pinned packs or any live program. Measured sections are marked as such; everything else is design.
|
||||
|
||||
Plan: `docs/plans/counter-asic-2.md`, layers 4 ("table layout drawn per era: item size, stride, interleave") and 8 ("working-set size drawn per program"). The era seed is `E_n` of spec 04 section 4.4. Confirmed on 5 October 2026 by grep over `igneum-pow/src` on branches master, readwidth and opencl-rdna4-telemetry: no era draw existed in code before this branch (`igneum-era`, `EraParams`, `era_seed`: no match).
|
||||
|
||||
## 1. What is drawn, and from what
|
||||
|
||||
Two streams, nothing else:
|
||||
|
||||
| Stream | Seeded from | Draws | Sets |
|
||||
|---|---|---|---|
|
||||
| Era stream | `seed_words_from_bytes("igneum-era/" \|\| E_n)` words 0 and 1 (spec 01 section 1.13.1 wrote `n_le64 \|\| E_n`; the index is dropped here because `E_n` already commits to `n` through the VDF input of section 4.4 step 2, and the node's seam hands the generator the era bytes alone: `Epoch::from_chain_seeds(epoch, day, era, class, label)`, branch ca2-v3) | 7 per era, fixed | the load width `W`, the stride `(M, R)`, the interleave `pos[0..3]` |
|
||||
| Program stream | the epoch seed words as today (spec 01 section 1.3.3) | 11 per instruction instead of 9: the 9 of version 2, then 2 window draws (a class whose loads are not version 2's takes the read-width width roll between them, ca2-v3's rule) | per load site: the window shrink `k_off` and its offset `o` |
|
||||
|
||||
The era parameters change the class of every program of the era. The program stream does not see the era parameters (two eras with the same epoch seed draw the same instruction list and the same windows, and differ in width, stride and layout); this is what makes the six era packs below a controlled comparison.
|
||||
|
||||
### 1.1 Era draw (proposed spec text for section 1.13.1, replacing its parameter table)
|
||||
|
||||
One SplitMix64 stream `S` seeded with `lo = words[0] | (words[1] << 32)` of `seed_words_from_bytes("igneum-era/" || E_n)`. Seven draws, in this order, whether or not a value is used:
|
||||
|
||||
1. `W = allowed[below(|allowed|)]`: the width in words of every dataset load of the era, drawn from the genesis-fixed ascending set `allowed`, a subset of {1, 4, 16} (4, 16 or 64 bytes). A set of one element pins the width; the draw is still consumed. The set is `{1}` (4 bytes, v2's load): the read-width decision of 5 October 2026 (`docs/plans/read-width.md`, "keep v2; w16 the only width that passes the rules and closes nothing") and the adoption rule of the same evening (a draw that changes the bytes per hash changes the rate; the six-era hash-rate spread must stay under 5 percent per card). The set is recorded in every pack (`IGNEUM_ERA_ALLOWED_WIDTHS`, `program.json` `era.allowed_widths`) and enters the program id. The code keeps the draw general so the set can be widened at genesis without a new derivation.
|
||||
2. `M = low32(next()) OR 1`: the stride multiplier, odd, so `x -> x * M` is a bijection on 32-bit words.
|
||||
3. `R = 1 + below(31)`: the stride rotation, in 1..31 (never 0: spec 01 section 1.14 item 3).
|
||||
4. to 7. `r_i = next()` for `i` in 0..3: the interleave draws. Let `b = log2(W)` (0, 2 or 4) and `free = 4 - b`. Let `c = [b, b + 1, ..., 15]` (16 - b candidates). For `i` in `0..free`: `j = i + (r_i mod (16 - b - i))`, swap `c[i]` and `c[j]`. The interleave is `pos = [0, ..., b - 1] ++ sort(c[0..free])`, four ascending bit positions in 0..15. Draws `r_free..r_3` are consumed and ignored.
|
||||
|
||||
The era parameters are `(W, M, R, pos)`. Era 0 of the devnet packs is listed in section 5.
|
||||
|
||||
### 1.2 Dataset mapping with the interleave (replaces the last sentence of section 1.8.5)
|
||||
|
||||
Word `w` of the dataset holds word `j(w)` of item `t(w)`, where `j(w)` is the 4-bit number whose bit `i` is bit `pos[i]` of `w`, and `t(w)` is `w` with bits `pos[0..3]` removed (the remaining bits in order). With `pos = [0, 1, 2, 3]` this is today's `dataset[w] = item(w >> 4)[w AND 15]`, byte for byte.
|
||||
|
||||
Properties kept:
|
||||
|
||||
- An item has the same value at every dataset size (item derivation is untouched), and because every `pos[i] < 16`, `dataset[w]` is the same at every dataset size of at least 2^16 words. The 1 GiB vectors of an era remain valid for words below 2^28 at any larger size, as today.
|
||||
- The low `b = log2(W)` positions are 0..b-1, so the `W` words of one aligned load lie in one item (`t` is the same for all of them and `j` runs `j0 .. j0 + W - 1`). The verifier derives one item per lane per load, as today: the 4,096-item bound of section 1.11 holds (16 loads x 8 iterations x 32 lanes, whatever the width).
|
||||
- The dataset build writes 16 words of one item to 16 addresses `w(t, j)` (a scatter of 4-byte writes instead of one 64-byte line when `pos != [0, 1, 2, 3]`). This is the only GPU cost of the interleave and is paid once per day; section 6 measures it.
|
||||
|
||||
What the interleave does and does not buy. A chip that hard-wires today's layout (64-byte items, a 64-byte line per item) reads the wrong 15 words with every word once the era draws another layout; the layout changes every 180 days inside rules fixed at genesis. A chip whose address decoder can permute 28 address lines under firmware control pays nothing for it. The honest claim is the first sentence only. The stride below is the same kind of lever: two integer operations per load on a chip, nothing on a GPU.
|
||||
|
||||
### 1.3 Load address (replaces "a load reads one 4-byte word at `src AND MASK`" in section 1.5 for era programs)
|
||||
|
||||
For a load site with window draws `(k_off, o)` (section 1.4), a dataset of `2^D` words (`MASK = 2^D - 1`), and register value `x`:
|
||||
|
||||
```
|
||||
k = min(k_off, D - 26) (0 when D <= 26)
|
||||
y = rotl(x * M, R)
|
||||
idx = ((y AND (MASK >> k)) OR ((o AND (2^k - 1)) << (D - k))) AND MASK
|
||||
base = idx AND NOT (W - 1) (W words from base are folded as verify::fold_words)
|
||||
```
|
||||
|
||||
Uniformity: `x * M` with `M` odd and `rotl` are bijections of the 32-bit word, so `y` is uniform when `x` is; `y AND (MASK >> k)` is uniform on the window; the offset picks which of the `2^k` aligned windows. Branch-free, integer only, three operations before the mask (multiply, rotate, and-or) against one today. The emitted text has one form per dialect, checkable by text search (section 1.14 item 2): CUDA and OpenCL `ds[((rotl_imm(rN * 0x........u, Ru) & 0x........u) | 0x........u) & mask]`, Metal the same with `dataset[` and `& MASK]`.
|
||||
|
||||
### 1.4 Window draw per load site (layer 8; proposed text for section 1.4.3)
|
||||
|
||||
After the nine draws of version 2 (and the width roll, for a class whose loads are not version 2's), every instruction takes two more draws, used only on a load slot:
|
||||
|
||||
```
|
||||
k_off = below(3) window = the dataset, a half or a quarter of it (2^(D - k_off) words)
|
||||
o = low32(next()) AND (2^k_off - 1) which aligned window
|
||||
```
|
||||
|
||||
Bounds: the window never goes below `2^26` words (256 MiB; `k = min(k_off, D - 26)` in 1.3), which exceeds the largest on-chip cache of any card in the benchmark (the RTX 5090's 96 MiB L2, the RX 9070 XT's 64 MB Infinity Cache, vendor figures), and never above the dataset. At the prototype dataset (2^28) the windows are 1 GiB, 512 MiB and 256 MiB; at the genesis dataset (2^29) 2 GiB, 1 GiB and 512 MiB. The dataset grows by the step schedule recommended to Josh (spec 01 section 1.13.3 option (b), `docs/analysis/card-lifetime-2026-10-05.md`: power-of-two steps, 4 GiB at year 4, 8 GiB at year 12, 16 GiB at year 28, 32 GiB at year 60, every index `AND MASK`), so the window ceiling follows the steps and the floor stays the genesis constant 2^26 words; nothing in the address of 1.3 needs a range reduction. A program has 16 load sites and so up to 16 windows; the set a program reads is their union (section 7 computes its distribution). The verifier bound is unchanged (1.2).
|
||||
|
||||
Why per load site and not per program: a per-program window of a quarter of the dataset would hand a 256 MiB SRAM mirror a third of the hours at the prototype size. Sixteen sites with drawn offsets cover the dataset with high probability (section 7), so the mirror a chip would need is the whole dataset in every hour, and the hour-to-hour variation lands on the memory design (which quarter, which half, how many distinct windows), not on its size.
|
||||
|
||||
### 1.5 Program stream (replaces "592 draws per program" in section 1.4.3 for era programs)
|
||||
|
||||
16 slot draws, then 64 x 11 = 704: 720 draws per program (768 and 784 for a class with the width roll). On the chain an era program is a class v3 program (branch ca2-v3's seam: `ProgramClass::V3`, generator version 3): its id is `program_id(3, seed words, attempt)` as that branch defines it, and the era it was drawn under is identified beside the id by `IGNEUM_ERA_SEED_HEX` (`packcheck::verify_pack_dir_chain` refuses a pack whose era is not the job's), so the pair (id, era seed) names the program. The experiment classes that are not class v3 (`--class <other> --era ...`) carry the era inside the class id instead: `FNV-1a-64("igneum-program-rw/" || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots || "era/" || allowed[3] || W || M_le32 || R || pos[4])`. The class name is the base name with `-era<first stream word as hex>` (`w4-era401998a5`).
|
||||
|
||||
### 1.5.1 The seam (branch ca2-v3, kept as its signatures stand)
|
||||
|
||||
`V3_CLASS` is the base class of class v3 (16 loads of 4 bytes, the lottery hash's load; the era rides inside it), `V3_ALLOWED = [1]` its width set. `generate_from_seed_bytes_program_class(label, seed, V3, Some(era))` draws `LoadClass::era(V3_CLASS, era, &V3_ALLOWED)` and stamps generator 3; without era bytes (a template before the era is known) the bare `V3_CLASS` stands. `Epoch::chain_dataset(day, class)` is unchanged: the layout of 1.2 is a property of the program (`program.class.layout()`), applied by the interpreter and by `Epoch::dataset_word`, so the day's cache is shared by every era of a day and the engine keys its caches on `(day, class)` as before.
|
||||
|
||||
### 1.6 Acceptance
|
||||
|
||||
The rule of section 1.4.6 is unchanged in its tests. Its interpreter mirrors 1.3 at the rule's constant `D = 28` (`idx` as above with `MASK = 0x0fffffff`), as `igneum-pow/src/accept.rs` does. The distinct-address bound counts dataset loads as the read-width branch defines it.
|
||||
|
||||
## 2. The era seed on the devnet (stand-in for `E_n`)
|
||||
|
||||
Until the 1-hour VDF of section 4.4 is in the node, in the shape of the epoch seed's stand-in (`docs/fork-divergence.md` "Epoch seed"):
|
||||
|
||||
| Era | `E_n` |
|
||||
|---|---|
|
||||
| 0 | the genesis block hash (32 bytes) |
|
||||
| n >= 1 | the hash of the last selected-chain block whose DAA score is below `15,552,000 n - 7,200` (the 2-hour lead of section 4.4 step 1) |
|
||||
|
||||
The era of a block is `floor(DAA score / 15,552,000)`, a function of the header alone. `E_n` for `n >= 1` is known 7,200 DAA seconds before the era starts, which covers the 1-hour VDF when it arrives and the kernel compile and dataset rebuild now. Test seeds for packs and tests: `E_n` = the 32 bytes (little-endian words) of `seed_words_from_bytes("igneum-era-test/<n>")` (`igneum-pow ... --era igneum-era-test/<n>`, the number only names the pack); raw bytes with `--era <n>:<64 hex>`.
|
||||
|
||||
## 3. Memory budget (Josh, 5 October 2026: under 6 GB on an 8 GB card)
|
||||
|
||||
The era layout adds no resident memory: a window is a mask and an offset in the kernel text, the interleave is address arithmetic, the stride is two operations. The whole working set on a card, every item from this branch and the others:
|
||||
|
||||
| Item | Bytes | Who |
|
||||
|---|---|---|
|
||||
| Dataset (prototype) | 1 GiB | existing |
|
||||
| Cache (256 MiB, resident only while the day's dataset is built, then free; the decided reading, confirmed by the coordinator on 5 October 2026, `docs/analysis/card-lifetime-2026-10-05.md` carries the per-tier working set with the cache freed as the best case) | 256 MiB peak | existing |
|
||||
| Layer 5 hot table | the ca2-cache worker's figure | not this branch |
|
||||
| Scratch per resident warp (read-width variant 5, 32 or 128 KiB per warp) | 2,048 warps x 128 KiB = 256 MiB at most | readwidth branch |
|
||||
| Output and read-back buffers | 2^24 nonces x 8 B = 128 MiB per dispatch | existing harness |
|
||||
| Era windows, stride, interleave | 0 | this branch |
|
||||
|
||||
The era window never exceeds the dataset, so it never grows the footprint.
|
||||
|
||||
## 4. Implementation (behind the flag)
|
||||
|
||||
| Piece | Where | What |
|
||||
|---|---|---|
|
||||
| `EraParams`, `era_draw`, `LoadClass::era` | `igneum-pow/src/generator.rs` | the 7-draw era stream of 1.1; the class carries `era: Option<EraParams>` beside `mix`, `load_slots`, `scratch`, `scratch_kb`; the two window draws per instruction (`Instr::win`, `Instr::off`); the program id of 1.5 |
|
||||
| `Layout` | `igneum-pow/src/memhard.rs` | `split(w) -> (t, j)`, `join(t, j) -> w`, `LINEAR = [0, 1, 2, 3]`; `MemhardCpu` and `DatasetSource` carry it |
|
||||
| `load_index` | `igneum-pow/src/verify.rs` | the address of 1.3, shared by the interpreter and the acceptance mirror |
|
||||
| Emitters | `igneum-pow/src/emit.rs` | the one load form of 1.3 in Metal, CUDA and OpenCL C; `mh_word`, the three `igneum_build` kernels and the Metal build kernel with `mh_t`, `mh_j`, `mh_addr` when the layout is not linear; `IGNEUM_ERA_*` in `program.h`, an `"era"` object in `program.json` |
|
||||
| CLI | `igneum-pow/src/main.rs` | `--era <igneum-era-test/n \| n:hex>` and `--era-widths 4,16,64` (one width pins) on every command |
|
||||
| Tests | `igneum-pow/src/*.rs`, `igneum-pow/tests/packs.rs` | the draw is deterministic and within bounds; `split`/`join` are inverse and the dataset is a prefix at every size; six era programs pass the generator contract and the acceptance rule; vectors round-trip; the pinned packs are byte-identical; the six era packs match the emitters and every load has the form of 1.3 |
|
||||
|
||||
The default class is untouched: `LoadClass::V2` has `era: None`, every emitter branch on `era` keeps today's text, and `tests/packs.rs` diffs the pinned packs (`igneum-genesis-mh`, `igneum-devnet-v4-epoch0`) against the emitters as before. Section 6 records the diff of a fresh export against the checked-in files.
|
||||
|
||||
## 5. The six era packs
|
||||
|
||||
`proto-cuda/packs-ca2-era/era-<n>`, `n` in 0..5: the devnet's 32-byte epoch seed `edc4fa84...fb07` and day bytes `igneum-day/20730` (the seeds of the pinned pack `igneum-devnet-v4-epoch0`, which is the v2 baseline with the same program seed), dataset 2^28 words, era seed `igneum-era-test/<n>`, width pinned at 4 bytes (`--era-widths 4`, the default). The program seed is held fixed so that the six packs differ in the era parameters only (section 1); each carries `seeds.txt` for the one-click workers. Every pack is attempt 1 (attempt 0 of this seed is rejected under the era class: 12 draws per instruction give a different stream from v2's). The drawn parameters (`igneum-pow show --epoch-hex edc4... --era igneum-era-test/<n>`, 5 October 2026):
|
||||
|
||||
| Pack | Class | Era seed `E_n` (first 16 hex) | W (bytes) | M | R | pos |
|
||||
|---|---|---|---|---|---|---|
|
||||
| era-0 | w4-erab2ed8a89 | 5e0587f455a86e91 | 4 | 0x625e5ab3 | 19 | 0, 2, 10, 15 |
|
||||
| era-1 | w4-era676a17fc | df57136f2ad5f410 | 4 | 0xb2a9d70d | 6 | 1, 3, 8, 13 |
|
||||
| era-2 | w4-era843155d7 | 7f450623297a954f | 4 | 0x2b4a5b97 | 28 | 1, 3, 4, 8 |
|
||||
| era-3 | w4-erad6367bfe | 8bffdd3366b9c3ff | 4 | 0x27ea7eff | 30 | 2, 3, 8, 13 |
|
||||
| era-4 | w4-era4488f3ed | e593fc1d48475c88 | 4 | 0x4d38603d | 10 | 2, 9, 13, 15 |
|
||||
| era-5 | w4-eraf897c84e | ff87ad96a1b53f36 | 4 | 0x03ac37ad | 22 | 0, 2, 10, 13 |
|
||||
|
||||
All six are class v3 packs (generator 3, `IGNEUM_PROGRAM_CLASS "v3"`, `IGNEUM_ERA_SEED_HEX`), attempt 0, program id `73bcbfe8ccf988f1` in every pack (the seam's `program_id(3, seed, attempt)`; the era seed beside it names the program), the era layout over version 2's item construction (mixer x1, the genesis cache) so that the v2 baseline pack is the same dataset; the chain's class v3 composes the same draw over `LoadClass::MX4` (mixer x4, growth), and the integration re-exports these packs on it after the PC rows. The era-seed-to-pack assignment above is from `program.h` of each pack; the test-seed numbering is only the pack name.
|
||||
|
||||
The windows are a property of the program, so they are the same in all six packs (site:shrink:offset): `1:2:3 4:2:2 6:0:0 10:1:0 12:1:1 20:2:0 27:1:0 30:0:0 35:0:0 40:0:0 41:0:0 43:2:3 45:1:0 52:0:0 53:2:2 54:0:0`: 7 sites read the whole dataset, 5 a half, 4 a quarter; the union is the whole dataset.
|
||||
|
||||
## 6. Measurements
|
||||
|
||||
Pending at the time of this commit; each table below says the machine, the date, the harness and the command when filled.
|
||||
|
||||
### 6.1 Byte-identical default path
|
||||
|
||||
### 6.2 Bit-exactness on the Mac (Metal, Apple OpenCL, CUDA emulation)
|
||||
|
||||
### 6.3 Hash rate per era on the M5 Max (Metal) and the CPU verifier
|
||||
|
||||
### 6.4 PCs (RTX 5090 CUDA, RX 9070 XT OpenCL): prepared, waiting for the go
|
||||
|
||||
## 7. The chip-model line per draw
|
||||
|
||||
From `docs/analysis/m16-recompute-attacker-2026-10-05.md` and the random-read ceilings of `docs/bench-log.md` ("the 9070 XT on the eGPU": 9070 XT 2.42 to 2.68 G loads/s at 1 GiB, 5090 16.4 to 18.0, M5 Max 3.41 to 3.49; every random 4-byte read costs AMD a 64-byte line):
|
||||
|
||||
| Quantity | Formula |
|
||||
|---|---|
|
||||
| Bytes read per hash | 128 loads x W bytes |
|
||||
| Distinct 64-byte lines per hash | 128 (one line per load at every W up to 64 bytes; the census's 120 to 128 distinct addresses per hash) |
|
||||
| SRAM a chip needs to mirror what the hash reads | the union of the program's 16 windows (a distribution over programs; section 7.1) |
|
||||
| Latency-bound share | measured rate / (the card's 4-byte random-read ceiling / 128) |
|
||||
|
||||
The latency-bound share uses the 4-byte ceiling for every width because the 9070 XT line probe showed the same count per second for 4-byte and 64-byte random reads; the 5090's 16-byte and 64-byte ceilings are the read-width branch's measurement, cited when they land.
|
||||
|
||||
### 7.1 Union of windows per program
|
||||
|
||||
Filled from a CPU census over programs (section 6).
|
||||
|
||||
## 7.2 Found on the way (harness defects, both fixed on this branch)
|
||||
|
||||
| Where | Defect | Fix | Checked |
|
||||
|---|---|---|---|
|
||||
| `proto-cuda/nvrtc/packfile.h` (the one-click workers) | the seed words were re-derived from the bare epoch seed, so every pack of attempt 1 or higher was refused ("the epoch seed bytes do not give the pack's IGNEUM_SEEDW_INIT"); 5.14 percent of epochs under v2, all six era packs, and the epoch 34 fleet outage of 18:23Z on 5 October 2026 (branch pack-loop af983a7, which this branch takes: `pf_program_words`) | the pack-loop derivation merged over the readwidth packfile (class fields and string seeds kept) | the devnet pack (attempt 0) loads, the six era packs (attempt 1) load, a copy of era-0 with the attempt tampered to 0 is refused on the re-derivation (section 6) |
|
||||
| `proto-cuda/host.cu`, `proto-opencl/host.c` | the host-side dataset word was `mh_item(w >> 4)[w AND 15]`, the harness's own copy of the linear layout; under an interleaved layout the "64 random points vs host derivation" check failed while the Mac samples and the vectors passed | `host_ds_word` calls the pack's `mh_word` (memhard.h), which carries the layout | `proto-cuda/emu/test-layout.sh`: the CUDA emulation on era-1 (interleaved) and the devnet pack (linear) must pass the random-point check; it failed on era-1, era-3 and era-5 before the fix (section 6) |
|
||||
|
||||
## 8. What is unverified
|
||||
|
||||
- Everything in section 6 marked pending.
|
||||
- The 1-hour VDF does not exist; the devnet stand-in of section 2 is a proposal.
|
||||
- The interleave's value against a chip with a programmable address decoder is nil (1.2); the claim is limited to hard-wired layouts.
|
||||
- The window floor of 2^26 words is set by the 5090's L2 (96 MiB) and the 9070 XT's Infinity Cache (64 MB, vendor figures); a future card with a larger cache moves the floor, which is a genesis constant.
|
||||
- No cryptanalysis of the stride (a multiply and a rotate before the mask); it is a bijection, so the address distribution is that of the register value, as today.
|
||||
109
docs/plans/read-width.md
Normal file
109
docs/plans/read-width.md
Normal file
|
|
@ -0,0 +1,109 @@
|
|||
# Read width of the lottery hash: 4, 16 and 64-byte loads, a per-load mix, and a written scratch (gate 1 experiment)
|
||||
|
||||
5 October 2026. Branch `readwidth` (worktree `../igneum-wt-readwidth`), commits 019b014 and b970dda plus the measurement commit. Nothing here changes consensus, the live generator, the pinned vectors or a shipped binary: every class sits behind `--class` in `igneum-pow` and the default class is generator version 2 byte for byte (`igneum-pow/tests/packs.rs` still compares the four pinned packs against the emitters). Numbers and a recommendation; the decision is Josh's.
|
||||
|
||||
## 1. The question
|
||||
|
||||
The bench-log entry "the 9070 XT on the eGPU" (5 October 2026) found the hash bound by dependent random 4-byte reads over the 1 GiB dataset, 128 per hash: the RX 9070 XT finishes 2.4 to 2.7 G such reads a second (18 MH/s), the RTX 5090 16 to 18 G (127 MH/s), the M5 Max 3.45 G (23 to 28 MH/s). AMD fetches a 64-byte line per 4-byte read, so 94 percent of its memory traffic is unused; NVIDIA fetches a 32-byte sector and its 96 MB L2 catches a share. Josh's question: would wider reads keep the chip-resistance property (latency-bound, random access) while closing the vendor gap? Two additions from the coordinator: a per-load width drawn from an era-fixed mix so no chip is built for one width, and a written per-warp scratch so part of the memory work cannot be mirrored into read-only SRAM.
|
||||
|
||||
## 2. What was built (all behind the flag)
|
||||
|
||||
| Class (`--class`) | Loads per hash | What a load does | Dataset bytes per hash |
|
||||
|---|---|---|---|
|
||||
| `v2` (= `w4`, the lottery hash) | 128 | `dst ^= dataset[src & MASK]`, one 4-byte word | 512 |
|
||||
| `w16` | 128 | the 16-byte-aligned group of 4 words at `src & MASK`, every word folded into `dst` | 2,048 |
|
||||
| `w64` | 128 | the 64-byte-aligned item (16 words), every word folded | 8,192 |
|
||||
| `w64x4` | 32 (4 load slots) | as `w64`; the same bytes per hash as 512 loads of 4 bytes | 2,048 |
|
||||
| `mix50-35-15` | 128 | per load, width 4, 16 or 64 bytes drawn from the program stream with probabilities 50/35/15 | 1,664 to 3,680 over the six programs measured (expected 2,202) |
|
||||
| `mix25-50-25` | 128 | the same with 25/50/25 | 2,240 to 5,024 (expected 3,200) |
|
||||
| `scr<k>k<kb>` | 128 memory operations | `k` of the 16 slots are scratch read-modify-writes into a `kb` KiB per-warp scratch (16-byte slots, lane-major); the other `16 - k` are 4-byte loads | 4 x (16 - k) x 8 reads plus 16 B read and 16 B written per scratch op |
|
||||
|
||||
The fold. A load of W words reads the W-word-aligned address `b = (src AND MASK) AND NOT (W - 1)` and sets `x = dst XOR w[0]; for j in 1..W: x = (rotl(x, 11) * 0x9e3779b1) XOR w[j]; dst = x` (`verify::fold_words`, mirrored in the three kernel dialects). The rotate-multiply between the words makes the fold state-dependent: two different lines give two different maps of `dst` (the multiply by an odd constant is not xor-linear), so no function of the line alone can stand in for it, and a dataset of pre-folded lines cannot replace the dataset. The dependent chain is unchanged: the next load's address comes from a register that the fold wrote. Width 1 is the lottery hash's xor of one word, so `w4` is the pinned program `bcc1248b10cc90f2` exactly.
|
||||
|
||||
The per-load draw. Every class other than `v2` takes one extra draw per instruction (`below(100)`, the width roll, consumed on every slot so the stream stays uniform) after the nine draws of spec 01 section 1.4.3; on a load slot the width is the first entry of the mix whose cumulative weight exceeds the roll. The program id of a class is `FNV-1a-64("igneum-program-rw/" || 2 || seed words || attempt || mix[3] || slots [|| "scratch/" k kb])`, so no class program can pass for a version 2 program. The acceptance rule of 1.4.6 runs unchanged on the aligned addresses (lane-constant sites and the distinct-address bound, scaled to the dataset loads per hash).
|
||||
|
||||
The scratch (variant 5, measurement only). The kernel runs N persistent warps (one per block or work-group of 32); warp `w` owns scratch `w` and runs units `w, w + N, ...` of the launch. A scratch op reads the lane's 16-byte slot `src AND (slots - 1)`: three data words behind a per-unit tag; a slot whose tag is not this unit's reads as its fill `splitmix32(((base + lane) XOR seed[j]) + slot x 0x9e3779b1 + (j + 1) x 0x85ebca77)`, the three words are folded into `dst` as above, and the slot is rewritten `(tag, x XOR w1, rotl(x, 7) XOR w2, x + w0)`. The CPU verifier holds the touched slots of one unit (at most 32 x k x 8) and nothing else. The scratch is per lane (a 32 KiB warp scratch is 64 slots per lane, 128 KiB is 256), so two lanes never race on a slot and the result is a function of (program, day, unit) alone; the GPU's tags make the lazy fill exact as long as a tag is not reused within the arena's history (the salt advances per unit; it wraps after 2^32 units, a measurement caveat, not a design).
|
||||
|
||||
## 3. Method
|
||||
|
||||
| Step | Command (every figure in the bench log carries its command) |
|
||||
|---|---|
|
||||
| Packs | `igneum-pow export --seed <s> --class <c> --out proto-cuda/packs-readwidth/<name>` (23 packs; the six mix seeds per mix are `igneum-readwidth/A/<k>` and `/B/<k>` with attempt 0 accepted, `A/4` skipped: rejected at attempt 0) |
|
||||
| CPU verifier | `igneum-pow bench --seed igneum-genesis --class <c> --warps 50` (M5 Max, one core; load average 4 to 9 from other agents' builds during the run) |
|
||||
| Metal | `proto-metal/packbench --pack <dir> --batches 5 --batch-log2 24 --group 256 [--warps N]` (new harness: runs the pack's own text; vectors, cache FNV, dataset words, 2^24 fingerprint, MH/s by GPU time) under the measure lock |
|
||||
| Apple OpenCL | `proto-opencl/igneum-bench-cl-rw --bench-pack --pack <dir> --batches 5 --batch-log2 24` and `--memprobe` (Apple's OpenCL, a correctness check and an approximate rate) |
|
||||
| Emulators | `proto-cuda/emu/emu.sh ../packs-readwidth/<p> --batch-log2 13 --batches 1 --block-warps 2` (the CUDA text as C++); `proto-opencl/emu/emu.sh ../packs-readwidth/<p> 0 32 --sg 32` and `1 64 --sg 64` (the OpenCL text, wave32 and wave64) |
|
||||
| RTX 5090 | job `run-readwidth-5090-20261005` on PC 2 (`relay/playbooks/readwidth-5090.ps1`): the NVIDIA card switched off in the app through `POST app.url/api/cards` and restored after; `igneum-worker-cuda --memprobe`, then `--bench --pack <dir> --batches 5 --batch-log2 24` per pack (NVRTC, the pack's own text, vectors through the bound kernel, 2^24 fingerprint) |
|
||||
| RX 9070 XT | job `run-readwidth-9070-20261005` on PC 1 (`relay/playbooks/readwidth-9070.ps1`): only the gfx1201 card switched off; `igneum-worker-opencl --device D --memprobe`, then `--bench-pack --pack <dir> --batches 5 --batch-log2 24` per pack |
|
||||
|
||||
Latency-bound share = measured MH/s x loads per hash / the card's dependent-read ceiling for that width from its own probe at 1024 MiB (for a mix, the harmonic combination of the widths' ceilings weighted by the program's width counts). A share near 1 means the hash runs at the card's random-access limit, the property the design wants; a share well under 1 means something else bounds it (bandwidth, ALU, occupancy).
|
||||
|
||||
## 4. Results (full tables with commands in `docs/bench-log.md`, "read width of the lottery hash")
|
||||
|
||||
Probe ceilings at 1024 MiB (G dependent reads/s): RTX 5090 4 B 17.5, 16 B 18.0, 64 B 9.1 (584 GB/s), stream 1,579 GB/s; RX 9070 XT 4 B 2.42, 16 B 2.43, 64 B 2.47 (158 GB/s), stream 636; M5 Max (Apple OpenCL, approximate) 3.50 / 3.51 / 3.51, stream 522.
|
||||
|
||||
| Class | dataset B/hash | RTX 5090 MH/s (latency-bound share) | RX 9070 XT (share) | M5 Max Metal (share) | 5090 / 9070 | DRAM bytes moved per hash, NVIDIA 32 B sector / AMD 64 B line | CPU verify ms per unit |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| v2 = w4 (today) | 512 | 136.1 (0.96) | 18.15 (0.87) | 27.74 (1.01) | 7.5x | 4,096 / 8,192 | 0.604 |
|
||||
| w16 | 2,048 | 139.8 (0.90) | 17.90 (0.84) | 28.26 (1.03) | 7.8x | 4,096 / 8,192 | 0.610 |
|
||||
| w64 | 8,192 | 71.9 (0.58) | 17.59 (0.78) | 28.27 (1.03) | 4.1x | 8,192 / 8,192 | 0.630 |
|
||||
| w64x4 (32 loads) | 2,048 | 275.3 (0.56) | 75.19 (0.84) | 109.7 (1.00) | 3.7x | 2,048 / 2,048 | 0.160 |
|
||||
| mix50-35-15 (6 programs, min / median / max) | 1,664 to 3,680 | 99.5 / 114.2 / 121.0, spread 18.8% | 17.45 / 18.76 / 18.83, 7.4% | 25.36 / 27.26 / 28.43, 11.3% | 6.1x | 5,939 / 8,192 expected | 0.620 |
|
||||
| mix25-50-25 (6 programs) | 2,240 to 5,024 | 95.9 / 107.3 / 119.8, 22.3% | 17.84 / 18.45 / 18.85, 5.5% | 23.21 / 24.68 / 25.21, 8.1% | 5.8x | 7,168 / 8,192 expected | 0.614 |
|
||||
|
||||
Scratch, variant 5 (N persistent warps; GPU cost against the persistent control scr0k32; working set = 1 GiB + 256 MiB + 128 MiB output + N x size):
|
||||
|
||||
| Class | RMW share | dataset B/hash | scratch B/hash read + written | RTX 5090 MH/s, 2,048 warps of 4,080 resident (vs control, share) | M5 Max Metal (vs control) | RX 9070 XT, 4,096 warps | working set 5090 / 9070 / Mac |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| scr0k32 | 0 | 512 | 0 | 139.1 (control, 0.98) | 28.25 (control) | 17.88 (control, 0.86) | 1.4 GiB / 1.5 GiB / 1.5 GiB |
|
||||
| scr2k32 | 12.5% | 448 | 256 + 256 | 114.4 (-18%, 0.80) | 26.14 (-7%) | 14.65 (-18%) | same |
|
||||
| scr4k32 | 25% | 384 | 512 + 512 | 109.8 (-21%, 0.76) | 31.74 (+12%) | 14.00 (-22%) | same |
|
||||
| scr8k32 | 50% | 256 | 1,024 + 1,024 | 122.1 (-12%, 0.82) | 49.08 (+74%) | 14.17 (-21%) | same |
|
||||
| scr2k128 | 12.5% | 448 | 256 + 256 | 110.1 (-21%, 0.77) | 26.24 (-7%) | 14.07 (-21%) | 1.6 GiB / 1.9 GiB / 1.9 GiB |
|
||||
| scr4k128 | 25% | 384 | 512 + 512 | 98.0 (-30%, 0.68) | 28.08 (-1%) | 13.14 (-27%) | same |
|
||||
| scr8k128 | 50% | 256 | 1,024 + 1,024 | 72.8 (-48%, 0.49) | 35.44 (+25%) | 12.03 (-33%) | same |
|
||||
|
||||
Resident warps and the cap: the 5090 holds 4,080 warps at one warp per block (24 blocks per SM x 170 SMs; 8,160 at 8 warps per block), so 128 KiB each is 510 MiB and the whole working set 1.9 GiB; a 1 MB scratch would have been 4.0 GiB at this geometry and 10.6 GiB at the 64-warp-per-SM figure, which is why the cap moved the size to the tens of kilobytes. The occupancy query returned 24 blocks per SM before and after the arena allocation: the allocation did not change it. The 9070 XT's OpenCL runtime has no occupancy query; 4,096 persistent warps were launched (64 per compute unit over 64 CUs, approximate) and the arena is 128 MiB at 32 KiB, 512 MiB at 128 KiB. The Mac's residency is not reported; 2,048 to 16,384 warps were swept and the best row kept.
|
||||
|
||||
Chip model, re-run with the measured widths (the M16 arithmetic of `docs/analysis/m16-recompute-attacker-2026-10-05.md`; the on-die-cache recompute chip's row per scratch variant is the ca2-soundness branch's, as agreed with the Counter ASIC 2.0 coordinator):
|
||||
|
||||
| Class | what a chip with its own DRAM controller gains over the GPU's memory system | what a chip with on-die SRAM gains |
|
||||
|---|---|---|
|
||||
| v2 | the GPU fetches 8 to 16x the bytes it uses (AMD 64 B, NVIDIA 32 B per 4 B); a chip fetching 32 B bursts moves 4,096 B per hash, the 5090's figure, so nothing over NVIDIA and 2x over AMD in traffic, none in latency (the chain is 128 dependent DRAM latencies on either) | the recompute attacker of M16: 150,000 integer ops per hash against the 256 MiB cache; 2.4x at equal silicon before a fixed-function factor (unchanged by the width) |
|
||||
| w16 | the same: 4,096 / 8,192 bytes moved, 2,048 used; traffic efficiency 50 percent on NVIDIA, 25 on AMD | unchanged: the fold uses every byte, so the chip recomputes 128 items per hash exactly as before; the SRAM mirror of the read-only dataset (1 GiB) stays out of reach |
|
||||
| w64 | every byte moved is used on both vendors (8,192 moved, 8,192 used); the 5090 is bandwidth-bound at 589 GB/s, so a chip with HBM3 class bandwidth (several TB/s, approximate) is bandwidth-advantaged: the Ethash shape | unchanged in op count; but the chain of 128 loads now moves 8 KB, so a chip's advantage shifts from latency to bandwidth per dollar, which is the wrong direction for the design's 2x target |
|
||||
| w64x4 | 2,048 moved and used; 32 latencies per hash; every card 4x faster; a bandwidth-rich chip gains as above | 32 items per hash: the recompute attacker's op count falls 4x (37,500 per hash), so the M16 gain rises 4x: fails the 2x target by arithmetic |
|
||||
| mixes | between v2 and w64 per program; the chip cannot be built for one width, but the GPU pays the 64-byte hours (the 5090 loses up to 27 percent in a heavy hour) | as v2 per item; the recompute attacker is indifferent to the width |
|
||||
| scratch | a chip must provide writable memory for N units in flight: 32 KiB x N at the GPU's geometry (128 MiB at 4,080), against the 256 MiB read-only cache it could mirror into SRAM (54 to 83 mm^2 at a leading node, the coordinator's figure, approximate); but a unit touches at most k x 8 x 32 slots (4 KiB at 50 percent), the fill is a function and the tags are per unit, so a chip need only hold the touched set per unit in flight (the soundness caveat below) | the dataset reads replaced by scratch ops are reads the chip no longer has to serve from the 1 GiB; at 50 percent the recompute attacker computes 64 items instead of 128 |
|
||||
|
||||
Soundness (variant 5, measurement only, as instructed; the chip row is the ca2-soundness branch's, a465881: the on-die-cache recompute chip's gain is 2.4x at 0, 12.5, 25 and 50 percent, replaced or added, 32 or 128 KB, so the scratch does not move it): the per-unit scratch starts from a fill that any implementation can compute, and a unit writes at most `k x 8` slots per lane; an implementation that keeps only the touched slots of each unit in flight (the CPU verifier does exactly this) needs 16 B x touched slots, not the nominal arena, so the "real memory a chip must provide" is bounded by units in flight x touched slots, not by N x 32 KiB. The variant forces memory that is written, which SRAM can hold as well as DRAM; it does not force memory that is large. A written region that outlives the unit (state carried across units) would, and the CPU verifier could not replay it. This is the finding, not a recommendation.
|
||||
|
||||
## 5. Recommendation (the decision is Josh's)
|
||||
|
||||
Josh's rules, as passed by the coordinator: width = the widest read that keeps every card latency-bound (achieved within 90 percent of the probe ceiling at that width) with margin on the 5090 (bytes per hash x rate under a third of the 1,579 GB/s stream); the mix is in only if the six-program spread is under 5 percent per card; the scratch share is the smallest at which the chip model's gain falls under 1.5x at the lowest GPU cost within the 6 GB cap.
|
||||
|
||||
| Variant | Verdict under the rules | Numbers |
|
||||
|---|---|---|
|
||||
| w16 (16-byte loads, 128 per hash) | PASSES the rules: shares 0.90 / 0.84 / 1.03 (the 9070 XT's 0.84 equals its v2 share of 0.87 within noise: the card is at its ceiling in both), 286 GB/s on the 5090 = 18 percent of the stream. It does NOT close the vendor gap (7.8x against 7.5x), because the memory systems already move a sector or a line per load; it changes what the fold consumes, nothing the DRAM does | the only width row that passes; a no-cost change in rate (+2.7 percent 5090, -1.4 percent 9070 XT, +1.9 percent M5 Max) |
|
||||
| w64 | FAILS: 5090 share 0.58, 37 percent of the stream; closes the gap to 4.1x only by making the 5090 bandwidth-bound | |
|
||||
| w64x4 | FAILS: shares 0.56 / 0.84 / 1.00, the recompute gain rises 4x | |
|
||||
| mix 50/35/15 and 25/50/25 | OUT: spreads 18.8 and 22.3 percent on the 5090, 7.4 and 5.5 on the 9070 XT, 11.3 and 8.1 on the M5 Max, all over 5 percent; a chip is not built for a width anyway (see the model: the width does not change the recompute attacker) | |
|
||||
| scratch | OUT: every share costs the 5090 12 to 48 percent and the 9070 XT 18 to 33 percent, and raises the M5 Max's rate (the arena is cached there); the soundness branch's chip row (ca2-soundness a465881, the on-die-cache recompute chip) stays at 2.4x at every share, 32 or 128 KB, because the verifier resets the scratch per unit and the live state is the hash's own read-modify-writes, which a chip keeps in 80 to 320 B per lane; under the rule the share is 0. The rows stay as the measurement that decided it | |
|
||||
|
||||
Recommendation: keep 128 loads per hash and 4 bytes per load (v2) for the devnet; if a width change is wanted for the fold's sake (every byte of the sector consumed, which removes the "94 percent waste" statement from the AMD entry without changing what the card does), w16 is the one that passes every rule and costs nothing measurable, and it is the only width worth a vector re-cut. The AMD gap is a random-access gap (2.4 G against 17.5 G dependent reads per second at 1 GiB on the cards we own), and no read width closes it without turning the 5090 bandwidth-bound; the levers that act on the gap are the ones outside this experiment (the AMD card's memory path, and the dataset size against the 5090's 96 MB L2 share, which the probe's 64 MiB rows show at 9 G reads/s against 2.4 at 1 GiB). The per-load mix is out on stability; the scratch is out on GPU cost and on the soundness caveat.
|
||||
|
||||
## 6. What w16 would change if adopted (not done; the decision is Josh's)
|
||||
|
||||
| Where | Change |
|
||||
|---|---|
|
||||
| `docs/spec/01-lottery-hash.md` 1.4.1 | `load`: `dst = fold(dst, dataset[b .. b + 4))`, `b = (src AND MASK) AND NOT 3`, with the fold written out; 1.4.3 unchanged (no width draw for a fixed width); 1.4.6 unchanged (the aligned address is the address the rule sees) |
|
||||
| 1.5 | "A load reads one 4-byte word" becomes 16 bytes aligned; the single-form text-search rule of 1.14 item 2 becomes the wide form; the item size (64 B) and `dataset[w] = item(w >> 4)[w AND 15]` unchanged |
|
||||
| 1.11 | unchanged in count (4,096 items per unit; the verifier derives the same items) |
|
||||
| 1.15, 1.17 | every vector re-cut (new program ids: the class enters the id or the generator version steps to 3); the four pinned packs replaced; the conformance fuzz re-run on Metal, CUDA and OpenCL (this branch's 23 packs and the three emulators are the template) |
|
||||
| Litepaper, Mining ("random reads over a multi-gigabyte dataset") and the vs-RandomX "128 dataset addresses" rows | "128 reads of 16 bytes"; `site/bench.html` sector arithmetic (32 B per 4 B) becomes 32 B per 16 B |
|
||||
| Workers | no host change: the kernel text carries the loads; `proto-cuda/host.cu`'s static mask check (`TESTS.md` section 5) learns the wide form |
|
||||
| Cost on the 5090 | none measured (+2.7 percent); on the 9070 XT -1.4 percent; verifier +1 percent |
|
||||
|
||||
## 7. Files
|
||||
|
||||
`igneum-pow/src/{generator,verify,accept,emit,memhard,main}.rs` (the classes, behind `--class`), `proto-cuda/packs-readwidth/` (23 packs), `proto-metal/packbench.swift` (Metal from a pack's files), `proto-opencl/host.c` (`--bench-pack`, `--warps`, the 16-byte probe row, the scratch arguments), `proto-cuda/nvrtc/worker.cpp` (`--bench`, `--memprobe`, the scratch arena), `proto-cuda/nvrtc/packfile.h` (class fields; string-seed packs), `proto-cuda/emu/cuda_runtime.h` and `proto-opencl/emu/{emu_opencl.h,emu_main.cpp}` (vector types, the persistent launch), `relay/playbooks/readwidth-*.ps1` (the PC jobs: the card under test off in the app and restored, never the other card).
|
||||
|
|
@ -16,7 +16,7 @@
|
|||
|
||||
use crate::generator::{Instr, Op, Program, INSTR_COUNT, ITERATIONS, LANES};
|
||||
use crate::seed::{fnv1a64, SplitMix64};
|
||||
use crate::verify::{dataset_elem, fold_words, splitmix32, ScratchModel};
|
||||
use crate::verify::{dataset_elem, fold_words, load_index, splitmix32, ScratchModel};
|
||||
|
||||
/// Units (32-lane warps) the dynamic test interprets.
|
||||
pub const ACCEPT_UNITS: usize = 64;
|
||||
|
|
@ -193,6 +193,7 @@ fn run_unit(p: &Program, unit: usize, base: u32, acc: &mut Acc, lane_addrs: &mut
|
|||
let mut nload = 0usize;
|
||||
let mut scratch = if p.has_scratch() { Some(ScratchModel::new(p.class.scratch_slots_per_lane())) } else { None };
|
||||
let slot_mask = p.class.scratch_slot_mask();
|
||||
let era = p.class.era;
|
||||
for it in 0..ITERATIONS {
|
||||
let sel = r[0];
|
||||
for (k, ins) in p.instrs.iter().enumerate() {
|
||||
|
|
@ -285,7 +286,7 @@ fn run_unit(p: &Program, unit: usize, base: u32, acc: &mut Acc, lane_addrs: &mut
|
|||
let width = ins.width as usize;
|
||||
let align = !(ins.width as u32 - 1);
|
||||
for lane in 0..LANES {
|
||||
idx[lane] = (r[a][lane] & mask) & align;
|
||||
idx[lane] = load_index(era.as_ref(), ins, r[a][lane], mask, ACCEPT_DATASET_LOG2) & align;
|
||||
}
|
||||
if idx.iter().all(|&x| x == idx[0]) {
|
||||
return Err(Reject::LaneConstantSite { iteration: it as u8, instr: k as u8, unit: unit as u8 });
|
||||
|
|
|
|||
|
|
@ -9,13 +9,90 @@
|
|||
//! One deliberate difference from the Swift: `program_json` writes the cache line mask inside the `"item"` string
|
||||
//! as a bare `0x003fffff`. The Swift writes it quoted (`jhex`), which is not valid JSON.
|
||||
|
||||
use crate::generator::{Instr, Op, Program, ProgramClass, GENERATOR_VERSION, INSTR_COUNT, ITERATIONS, LOAD_SLOTS};
|
||||
use crate::generator::{EraParams, Instr, Op, Program, ProgramClass, GENERATOR_VERSION, INSTR_COUNT, ITERATIONS, LOAD_SLOTS};
|
||||
use crate::memhard::{
|
||||
MixParams, Shape, CACHE_LINES_PER_SEGMENT, CACHE_SEGMENT_LOG2_LINES, CACHE_TAG, CHACHA_ROUNDS, CHACHA_SIGMA,
|
||||
Layout, MixParams, Shape, CACHE_LINES_PER_SEGMENT, CACHE_SEGMENT_LOG2_LINES, CACHE_TAG, CHACHA_ROUNDS, CHACHA_SIGMA,
|
||||
ITEM_ROUNDS,
|
||||
};
|
||||
use crate::seed::SplitMix64;
|
||||
use crate::verify::{DatasetMode, DatasetSource, Epoch, FOLD_MUL, FOLD_ROT};
|
||||
use crate::verify::{window, DatasetMode, DatasetSource, Epoch, DEFAULT_DATASET_LOG2, FOLD_MUL, FOLD_ROT};
|
||||
|
||||
/// The index expression of a dataset load (era layout, `docs/plans/era-layout.md` section 1.3). For every class
|
||||
/// without an era it is the lottery hash's `rN & MASK`; for an era program it is the one form
|
||||
/// `((rotl_imm(rN * M, R) & WM) | OFF) & MASK` with the site's window constants at the pack's dataset size.
|
||||
fn load_index_expr(dialect: CoreDialect, era: Option<&EraParams>, ins: &Instr, a: &str, dataset_log2: u32) -> String {
|
||||
let mask_name = match dialect {
|
||||
CoreDialect::Metal => "MASK",
|
||||
_ => "mask",
|
||||
};
|
||||
match era {
|
||||
None => format!("{a} & {mask_name}"),
|
||||
Some(e) => {
|
||||
let (wm, off) = window(ins, mask_for(dataset_log2), dataset_log2);
|
||||
format!("((rotl_imm({a} * {}, {}u) & {}) | {}) & {mask_name}", hex(e.stride_mul), e.stride_rot, hex(wm), hex(off))
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// The era lines of program.h (empty without an era).
|
||||
fn era_header_lines(p: &Program) -> String {
|
||||
let Some(e) = p.class.era else { return String::new() };
|
||||
let mut s = String::new();
|
||||
s.push_str("// Era layout (5 October 2026, docs/plans/era-layout.md): NOT the lottery hash. Every dataset load reads\n");
|
||||
s.push_str("// idx = ((rotl(src * STRIDE_MUL, STRIDE_ROT) & window mask) | window offset) & MASK; the window of a load site is the\n");
|
||||
s.push_str("// dataset, a half or a quarter of it (IGNEUM_ERA_WINDOWS: site:shrink:offset); dataset word w holds word j(w) of item\n");
|
||||
s.push_str("// t(w) with j's bits at the INTERLEAVE positions (memhard.h: mh_t, mh_j, mh_addr).\n");
|
||||
s.push_str(&format!("#define IGNEUM_ERA_LABEL {}\n", jstr(&e.label())));
|
||||
s.push_str(&format!("#define IGNEUM_ERA_SEED_WORDS {{ {} }}\n", join_hex(&e.words)));
|
||||
s.push_str(&format!("#define IGNEUM_ERA_ALLOWED_WIDTHS {{ {}, {}, {} }} // words, ascending, 0 = unused; one entry pins the width\n", e.allowed[0], e.allowed[1], e.allowed[2]));
|
||||
s.push_str(&format!("#define IGNEUM_ERA_WIDTH_WORDS {}\n", e.width_words));
|
||||
s.push_str(&format!("#define IGNEUM_ERA_STRIDE_MUL {}\n", hex(e.stride_mul)));
|
||||
s.push_str(&format!("#define IGNEUM_ERA_STRIDE_ROT {}\n", e.stride_rot));
|
||||
s.push_str(&format!("#define IGNEUM_ERA_INTERLEAVE {{ {}, {}, {}, {} }}\n", e.pos[0], e.pos[1], e.pos[2], e.pos[3]));
|
||||
s.push_str(&format!("#define IGNEUM_ERA_WINDOWS {}\n", jstr(&era_windows(p))));
|
||||
s
|
||||
}
|
||||
|
||||
/// "site:shrink:offset" for every load site of an era program, space separated.
|
||||
fn era_windows(p: &Program) -> String {
|
||||
p.instrs
|
||||
.iter()
|
||||
.enumerate()
|
||||
.filter(|(_, i)| i.op == Op::Load)
|
||||
.map(|(k, i)| format!("{k}:{}:{}", i.win, i.off))
|
||||
.collect::<Vec<_>>()
|
||||
.join(" ")
|
||||
}
|
||||
|
||||
/// The layout helpers of the memory-hard core for a non-linear layout: `mh_j(w)`, `mh_t(w)` and `mh_addr(t, j)`
|
||||
/// (`Layout::split` and `Layout::join` as text). Empty for the linear layout, so the pinned packs do not change.
|
||||
fn layout_helpers(layout: Layout, u: &str, fn_: &str) -> String {
|
||||
if layout.is_linear() {
|
||||
return String::new();
|
||||
}
|
||||
let p = layout.pos;
|
||||
let low = |q: u8| hex(((1u64 << q) - 1) as u32);
|
||||
let mut s = String::new();
|
||||
s.push_str(&format!(
|
||||
"// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions {} {} {} {} of w.\n",
|
||||
p[0], p[1], p[2], p[3]
|
||||
));
|
||||
s.push_str(&format!(
|
||||
"{fn_} {u} mh_j({u} w) {{ return ((w >> {}u) & 1u) | (((w >> {}u) & 1u) << 1) | (((w >> {}u) & 1u) << 2) | (((w >> {}u) & 1u) << 3); }}\n",
|
||||
p[0], p[1], p[2], p[3]
|
||||
));
|
||||
s.push_str(&format!("{fn_} {u} mh_t({u} w) {{"));
|
||||
for &q in p.iter().rev() {
|
||||
s.push_str(&format!(" w = (w & {}) | ((w >> {}u) << {}u);", low(q), q + 1, q));
|
||||
}
|
||||
s.push_str(" return w; }\n");
|
||||
s.push_str(&format!("{fn_} {u} mh_addr({u} t, {u} j) {{ {u} w = t;"));
|
||||
for (i, &q) in p.iter().enumerate() {
|
||||
s.push_str(&format!(" w = ((w >> {}u) << {}u) | (w & {}) | (((j >> {}u) & 1u) << {}u);", q, q + 1, low(q), i, q));
|
||||
}
|
||||
s.push_str(" return w; }\n");
|
||||
s
|
||||
}
|
||||
|
||||
/// Where the words of a wide load come from (read-width experiment).
|
||||
#[derive(Clone, Copy, PartialEq, Eq)]
|
||||
|
|
@ -32,16 +109,16 @@ enum WideSource {
|
|||
/// address aligned down to `width` words, folded into `dst` as `verify::fold_words`. The emitted text is the
|
||||
/// same shape in the three dialects: the vector loads differ (`uint4` pointer on Metal and CUDA, `vload4` on
|
||||
/// OpenCL C 1.2). For `width == 1` the caller emits the lottery hash's one-word form instead.
|
||||
fn wide_load_stmt(dialect: CoreDialect, d: &str, a: &str, width: u8, src: WideSource, closed: Option<(u32, u32)>) -> String {
|
||||
fn wide_load_stmt(dialect: CoreDialect, d: &str, idx: &str, width: u8, src: WideSource, closed: Option<(u32, u32)>) -> String {
|
||||
debug_assert!(width == 4 || width == 16);
|
||||
let (u, mask, base_ptr) = match dialect {
|
||||
CoreDialect::Metal => ("uint", "MASK", "dataset"),
|
||||
CoreDialect::Cuda => ("uint32_t", "mask", "ds"),
|
||||
CoreDialect::OpenCl => ("uint", "mask", "ds"),
|
||||
let (u, base_ptr) = match dialect {
|
||||
CoreDialect::Metal => ("uint", "dataset"),
|
||||
CoreDialect::Cuda => ("uint32_t", "ds"),
|
||||
CoreDialect::OpenCl => ("uint", "ds"),
|
||||
};
|
||||
let vectors = width as usize / 4;
|
||||
let mut s = String::with_capacity(400);
|
||||
s.push_str(&format!("{{ {u} b_ = ({a} & {mask}) & ~{}u; ", width as u32 - 1));
|
||||
s.push_str(&format!("{{ {u} b_ = ({idx}) & ~{}u; ", width as u32 - 1));
|
||||
match src {
|
||||
WideSource::Stored => match dialect {
|
||||
CoreDialect::Metal => s.push_str(&format!("device const uint4* l_ = (device const uint4*)({base_ptr} + b_); ")),
|
||||
|
|
@ -49,7 +126,7 @@ fn wide_load_stmt(dialect: CoreDialect, d: &str, a: &str, width: u8, src: WideSo
|
|||
CoreDialect::OpenCl => {}
|
||||
},
|
||||
WideSource::InlineClosed => {}
|
||||
WideSource::InlineMemhard => s.push_str("uint s_[16]; mh_item(cache, b_ >> 4u, s_); "),
|
||||
WideSource::InlineMemhard => s.push_str("uint s_[16]; mh_item(cache, mh_t(b_), s_); "),
|
||||
}
|
||||
let word = |j: usize| -> String {
|
||||
match src {
|
||||
|
|
@ -58,7 +135,7 @@ fn wide_load_stmt(dialect: CoreDialect, d: &str, a: &str, width: u8, src: WideSo
|
|||
let (d0, d1) = closed.expect("closed-form words need d0, d1");
|
||||
format!("ds_elem(b_ + {j}u, {}, {})", hex(d0), hex(d1))
|
||||
}
|
||||
WideSource::InlineMemhard => format!("s_[(b_ & 15u) + {j}u]"),
|
||||
WideSource::InlineMemhard => format!("s_[mh_j(b_) + {j}u]"),
|
||||
}
|
||||
};
|
||||
if src == WideSource::Stored {
|
||||
|
|
@ -187,6 +264,14 @@ fn scratch_stmt(dialect: CoreDialect, d: &str, a: &str, slot_mask: u32) -> Strin
|
|||
/// lottery hash's text is unchanged: `gid` is the unit's first output index plus the lane. The host MUST launch
|
||||
/// `groups` as a multiple of N (a uniform trip count: the OpenCL local-memory exchange carries a barrier).
|
||||
fn persistent_prologue(dialect: CoreDialect, words_per_lane: usize) -> String {
|
||||
let (a, b) = persistent_prologue_parts(dialect, words_per_lane);
|
||||
a + &b
|
||||
}
|
||||
|
||||
/// The prologue in two parts: the warp's identity and arena, then the unit loop. OpenCL C requires a `__local`
|
||||
/// variable at the outermost scope of the kernel (AMD's compiler enforces it, 5 October 2026, round 3 on the
|
||||
/// 9070 XT), so the OpenCL kernels declare the exchange buffer between the two parts.
|
||||
fn persistent_prologue_parts(dialect: CoreDialect, words_per_lane: usize) -> (String, String) {
|
||||
let (u, tid, nthreads, ptr) = match dialect {
|
||||
CoreDialect::Metal => ("uint", "tid", "nthreads", "device uint*"),
|
||||
CoreDialect::Cuda => ("uint32_t", "(blockIdx.x * blockDim.x + threadIdx.x)", "(gridDim.x * blockDim.x)", "uint32_t*"),
|
||||
|
|
@ -197,11 +282,12 @@ fn persistent_prologue(dialect: CoreDialect, words_per_lane: usize) -> String {
|
|||
s.push_str(&format!(" {u} warp_ = {tid} >> 5;\n"));
|
||||
s.push_str(&format!(" {u} nwarps_ = {nthreads} >> 5;\n"));
|
||||
s.push_str(&format!(" {ptr} arena = scratch + ((size_t)warp_ * 32u + lane) * {words_per_lane}u;\n"));
|
||||
s.push_str(&format!(" for ({u} g_ = warp_; g_ < groups; g_ += nwarps_) {{\n"));
|
||||
s.push_str(&format!(" {u} gid = g_ * 32u + lane;\n"));
|
||||
s.push_str(&format!(" {u} gbase = baseNonce + g_ * 32u;\n"));
|
||||
s.push_str(&format!(" {u} tag = salt + g_;\n"));
|
||||
s
|
||||
let mut l = String::new();
|
||||
l.push_str(&format!(" for ({u} g_ = warp_; g_ < groups; g_ += nwarps_) {{\n"));
|
||||
l.push_str(&format!(" {u} gid = g_ * 32u + lane;\n"));
|
||||
l.push_str(&format!(" {u} gbase = baseNonce + g_ * 32u;\n"));
|
||||
l.push_str(&format!(" {u} tag = salt + g_;\n"));
|
||||
(s, l)
|
||||
}
|
||||
|
||||
/// The scratch lines of program.h (variant 5).
|
||||
|
|
@ -274,11 +360,14 @@ fn log2_segments(shape: &Shape) -> usize {
|
|||
shape.log2_segments() as usize
|
||||
}
|
||||
|
||||
/// The memory-hard core as source text (`emitMemhardCore`). Every parameter is a literal: the mixer constants of
|
||||
/// the day and, from `mp.shape`, the cache size and the mixer multiplier `m`. Under `m = 1` the text is version
|
||||
/// 2's byte for byte; under `m > 1` the item loop applies `mh_mixer` `m` times per round with the keys
|
||||
/// `round_key(r m + j)` (spec 01 section 1.8.5 under class v3).
|
||||
/// The memory-hard core as source text (`emitMemhardCore`). Every parameter is a literal. Linear layout.
|
||||
pub fn emit_memhard_core(mp: &MixParams, dialect: CoreDialect) -> String {
|
||||
emit_memhard_core_layout(mp, dialect, Layout::LINEAR)
|
||||
}
|
||||
|
||||
/// [`emit_memhard_core`] with the dataset layout: with a non-linear layout `mh_word` and the build kernels go
|
||||
/// through `mh_t`, `mh_j` and `mh_addr` (era layout); with the linear layout the text is unchanged.
|
||||
pub fn emit_memhard_core_layout(mp: &MixParams, dialect: CoreDialect, layout: Layout) -> String {
|
||||
let shape = &mp.shape;
|
||||
let m = shape.mixer_mult;
|
||||
let cache_log2_words = shape.cache_log2_words;
|
||||
|
|
@ -414,19 +503,52 @@ pub fn emit_memhard_core(mp: &MixParams, dialect: CoreDialect) -> String {
|
|||
));
|
||||
}
|
||||
s.push_str("}\n");
|
||||
s.push_str("// dataset[w] without the dataset: derive item w >> 4 and take word w & 15.\n");
|
||||
s.push_str(&format!(
|
||||
"{fn_} {u} mh_word({cptr} cache, {u} w) {{ {u} s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; }}\n"
|
||||
));
|
||||
if layout.is_linear() {
|
||||
s.push_str("// dataset[w] without the dataset: derive item w >> 4 and take word w & 15.\n");
|
||||
s.push_str(&format!(
|
||||
"{fn_} {u} mh_word({cptr} cache, {u} w) {{ {u} s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; }}\n"
|
||||
));
|
||||
} else {
|
||||
s.push_str(&layout_helpers(layout, u, fn_));
|
||||
s.push_str("// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).\n");
|
||||
s.push_str(&format!(
|
||||
"{fn_} {u} mh_word({cptr} cache, {u} w) {{ {u} s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }}\n"
|
||||
));
|
||||
}
|
||||
s
|
||||
}
|
||||
|
||||
/// The dataset store of one item in a build kernel: `d[i] = s[i]` at `ds + t * 16` for the linear layout, else the
|
||||
/// scatter `ds[mh_addr(t, i)] = s[i]` (era layout).
|
||||
fn build_store(layout: Layout, dialect: CoreDialect, ds: &str, t: &str) -> String {
|
||||
let (u, wptr, cast) = match dialect {
|
||||
CoreDialect::Metal => ("uint", "device uint*", ""),
|
||||
CoreDialect::Cuda => ("uint32_t", "uint32_t*", "(size_t)"),
|
||||
CoreDialect::OpenCl => ("uint", "__global uint*", "(ulong)"),
|
||||
};
|
||||
if layout.is_linear() {
|
||||
match dialect {
|
||||
CoreDialect::Metal => format!(" {wptr} d = {ds} + {t} * 16u;\n for ({u} i = 0u; i < 16u; ++i) d[i] = s[i];\n"),
|
||||
CoreDialect::Cuda => format!(" {wptr} d = {ds} + (size_t){t} * 16u;\n for ({u} i = 0u; i < 16u; ++i) d[i] = s[i];\n"),
|
||||
CoreDialect::OpenCl => format!(" {wptr} d = {ds} + ((ulong){t} * 16u);\n for ({u} i = 0u; i < 16u; ++i) d[i] = s[i];\n"),
|
||||
}
|
||||
} else {
|
||||
let indent = if dialect == CoreDialect::Metal { " " } else { " " };
|
||||
format!("{indent}for ({u} i = 0u; i < 16u; ++i) {ds}[{cast}mh_addr({t}, i)] = s[i];\n")
|
||||
}
|
||||
}
|
||||
|
||||
/// Metal library with the cache fill and dataset build kernels for one day key (`memhardMSL`, memhard.metal).
|
||||
pub fn metal_memhard(mp: &MixParams) -> String {
|
||||
metal_memhard_layout(mp, Layout::LINEAR)
|
||||
}
|
||||
|
||||
/// [`metal_memhard`] with the dataset layout (era layout).
|
||||
pub fn metal_memhard_layout(mp: &MixParams, layout: Layout) -> String {
|
||||
let mut s = String::new();
|
||||
s.push_str("#include <metal_stdlib>\n");
|
||||
s.push_str("using namespace metal;\n");
|
||||
s.push_str(&emit_memhard_core(mp, CoreDialect::Metal));
|
||||
s.push_str(&emit_memhard_core_layout(mp, CoreDialect::Metal, layout));
|
||||
s.push('\n');
|
||||
s.push_str(&format!("// One thread per segment (2^{} threads).\n", log2_segments(&mp.shape)));
|
||||
s.push_str(
|
||||
|
|
@ -441,8 +563,7 @@ pub fn metal_memhard(mp: &MixParams) -> String {
|
|||
s.push_str(" uint gid [[thread_position_in_grid]]) {\n");
|
||||
s.push_str(" uint s[16];\n");
|
||||
s.push_str(" mh_item(cache, gid, s);\n");
|
||||
s.push_str(" device uint* d = dataset + gid * 16u;\n");
|
||||
s.push_str(" for (uint i = 0u; i < 16u; ++i) d[i] = s[i];\n");
|
||||
s.push_str(&build_store(layout, CoreDialect::Metal, "dataset", "gid"));
|
||||
s.push_str("}\n");
|
||||
s
|
||||
}
|
||||
|
|
@ -487,7 +608,7 @@ fn metal_program_impl(p: &Program, dataset_log2: u32, source: LoadSource, bound:
|
|||
}
|
||||
let mut buffer0 = "device const uint* dataset [[buffer(0)]]";
|
||||
if let LoadSource::InlineMemhard(mp) = &source {
|
||||
s.push_str(&emit_memhard_core(mp, CoreDialect::Metal));
|
||||
s.push_str(&emit_memhard_core_layout(mp, CoreDialect::Metal, p.class.layout()));
|
||||
s.push('\n');
|
||||
buffer0 = "device const uint* cache [[buffer(0)]]";
|
||||
}
|
||||
|
|
@ -528,11 +649,12 @@ fn metal_program_impl(p: &Program, dataset_log2: u32, source: LoadSource, bound:
|
|||
));
|
||||
}
|
||||
s.push_str(&format!("\n for (uint it = 0u; it < {ITERATIONS}u; ++it) {{\n uint sel = r0;\n"));
|
||||
let word_index = |a: &str, wide: bool| -> String {
|
||||
let era = p.class.era;
|
||||
let word_index = |a: &str, wide: bool, ins: &Instr| -> String {
|
||||
if wide {
|
||||
format!("(simd_broadcast({a}, 0) & WMASK) + lane")
|
||||
} else {
|
||||
format!("{a} & MASK")
|
||||
load_index_expr(CoreDialect::Metal, era.as_ref(), ins, a, dataset_log2)
|
||||
}
|
||||
};
|
||||
let fetch = |idx: String| -> String {
|
||||
|
|
@ -568,10 +690,10 @@ fn metal_program_impl(p: &Program, dataset_log2: u32, source: LoadSource, bound:
|
|||
LoadSource::InlineClosed(d0, d1) => (WideSource::InlineClosed, Some((*d0, *d1))),
|
||||
LoadSource::InlineMemhard(_) => (WideSource::InlineMemhard, None),
|
||||
};
|
||||
wide_load_stmt(CoreDialect::Metal, &d, &a, ins.width, src, closed)
|
||||
wide_load_stmt(CoreDialect::Metal, &d, &word_index(&a, false, ins), ins.width, src, closed)
|
||||
}
|
||||
Op::Load => format!("{d} = {d} ^ {};", fetch(word_index(&a, false))),
|
||||
Op::WLoad => format!("{d} = {d} ^ {};", fetch(word_index(&a, true))),
|
||||
Op::Load => format!("{d} = {d} ^ {};", fetch(word_index(&a, false, ins))),
|
||||
Op::WLoad => format!("{d} = {d} ^ {};", fetch(word_index(&a, true, ins))),
|
||||
Op::Scratch => scratch_stmt(CoreDialect::Metal, &d, &a, p.class.scratch_slot_mask()),
|
||||
};
|
||||
s.push_str(&format!(" {line} // {k}\n"));
|
||||
|
|
@ -611,8 +733,9 @@ fn init_line(p: &Program, u: &str, i: usize) -> String {
|
|||
}
|
||||
|
||||
/// The instruction lines of the CUDA hash kernel body (shared by `igneum_hash` and `igneum_hash_bound`).
|
||||
fn cuda_instr_lines(p: &Program) -> String {
|
||||
fn cuda_instr_lines(p: &Program, dataset_log2: u32) -> String {
|
||||
let mut s = String::with_capacity(6000);
|
||||
let era = p.class.era;
|
||||
for (k, ins) in p.instrs.iter().enumerate() {
|
||||
let d = format!("r{}", ins.dst);
|
||||
let a = format!("r{}", ins.src);
|
||||
|
|
@ -634,8 +757,10 @@ fn cuda_instr_lines(p: &Program) -> String {
|
|||
Op::Rotr => format!("{d} = rotr_var({d}, {a});"),
|
||||
Op::Mad => format!("{d} = {a} * {b} + {d};"),
|
||||
Op::Shfl => format!("{d} = {d} ^ __shfl_xor_sync(0xffffffffu, {a}, {});", ins.mask),
|
||||
Op::Load if load_width(ins) > 1 => wide_load_stmt(CoreDialect::Cuda, &d, &a, ins.width, WideSource::Stored, None),
|
||||
Op::Load => format!("{d} = {d} ^ ds[{a} & mask];"),
|
||||
Op::Load if load_width(ins) > 1 => {
|
||||
wide_load_stmt(CoreDialect::Cuda, &d, &load_index_expr(CoreDialect::Cuda, era.as_ref(), ins, &a, dataset_log2), ins.width, WideSource::Stored, None)
|
||||
}
|
||||
Op::Load => format!("{d} = {d} ^ ds[{}];", load_index_expr(CoreDialect::Cuda, era.as_ref(), ins, &a, dataset_log2)),
|
||||
Op::WLoad => format!("{d} = {d} ^ ds[(__shfl_sync(0xffffffffu, {a}, 0) & wmask) + lane];"),
|
||||
Op::Scratch => scratch_stmt(CoreDialect::Cuda, &d, &a, p.class.scratch_slot_mask()),
|
||||
};
|
||||
|
|
@ -646,6 +771,13 @@ fn cuda_instr_lines(p: &Program) -> String {
|
|||
|
||||
/// The CUDA kernel (`generateCUDA`, kernel.cu). `memhard` is `None` for a closed-form pack.
|
||||
pub fn cuda_kernel(p: &Program, memhard: Option<&MixParams>) -> String {
|
||||
cuda_kernel_at(p, memhard, DEFAULT_DATASET_LOG2)
|
||||
}
|
||||
|
||||
/// [`cuda_kernel`] at a dataset size (an era program's window constants are literals of the pack's size; every
|
||||
/// other class ignores it).
|
||||
pub fn cuda_kernel_at(p: &Program, memhard: Option<&MixParams>, dataset_log2: u32) -> String {
|
||||
let layout = p.class.layout();
|
||||
let mut s = String::with_capacity(9000);
|
||||
s.push_str(&generated_by(&p.seed_string));
|
||||
s.push_str(
|
||||
|
|
@ -696,8 +828,7 @@ pub fn cuda_kernel(p: &Program, memhard: Option<&MixParams>) -> String {
|
|||
s.push_str(" if (t < nItems) {\n");
|
||||
s.push_str(" uint32_t s[16];\n");
|
||||
s.push_str(" mh_item(cache, t, s);\n");
|
||||
s.push_str(" uint32_t* d = ds + (size_t)t * 16u;\n");
|
||||
s.push_str(" for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i];\n");
|
||||
s.push_str(&build_store(layout, CoreDialect::Cuda, "ds", "t"));
|
||||
s.push_str(" }\n");
|
||||
s.push_str("}\n");
|
||||
s.push('\n');
|
||||
|
|
@ -722,7 +853,7 @@ pub fn cuda_kernel(p: &Program, memhard: Option<&MixParams>) -> String {
|
|||
s.push_str(&init_line(p, "uint32_t", i));
|
||||
}
|
||||
s.push_str(&format!("\n for (uint32_t it = 0u; it < {ITERATIONS}u; ++it) {{\n uint32_t sel = r0;\n"));
|
||||
s.push_str(&cuda_instr_lines(p));
|
||||
s.push_str(&cuda_instr_lines(p, dataset_log2));
|
||||
s.push_str(" }\n");
|
||||
s.push_str(" uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);\n");
|
||||
s.push_str(" uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);\n");
|
||||
|
|
@ -798,6 +929,11 @@ pub fn cuda_kernel(p: &Program, memhard: Option<&MixParams>) -> String {
|
|||
/// `I` arrive by value in `IgneumInitWords` (`bind::block_init_words`), the instruction text is that of
|
||||
/// `igneum_hash`. Declarations for the host are at the top of the file.
|
||||
pub fn cuda_kernel_bound(p: &Program, memhard: Option<&MixParams>) -> String {
|
||||
cuda_kernel_bound_at(p, memhard, DEFAULT_DATASET_LOG2)
|
||||
}
|
||||
|
||||
/// [`cuda_kernel_bound`] at a dataset size (see [`cuda_kernel_at`]).
|
||||
pub fn cuda_kernel_bound_at(p: &Program, memhard: Option<&MixParams>, dataset_log2: u32) -> String {
|
||||
let mut s = String::with_capacity(9000);
|
||||
s.push_str(&generated_by(&p.seed_string));
|
||||
s.push_str(
|
||||
|
|
@ -847,7 +983,7 @@ pub fn cuda_kernel_bound(p: &Program, memhard: Option<&MixParams>) -> String {
|
|||
));
|
||||
}
|
||||
s.push_str(&format!("\n for (uint32_t it = 0u; it < {ITERATIONS}u; ++it) {{\n uint32_t sel = r0;\n"));
|
||||
s.push_str(&cuda_instr_lines(p));
|
||||
s.push_str(&cuda_instr_lines(p, dataset_log2));
|
||||
s.push_str(" }\n");
|
||||
s.push_str(" uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);\n");
|
||||
s.push_str(" uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);\n");
|
||||
|
|
@ -891,8 +1027,9 @@ pub fn cuda_kernel_bound(p: &Program, memhard: Option<&MixParams>) -> String {
|
|||
}
|
||||
|
||||
/// The instruction lines of the OpenCL hash kernel body (shared by `igneum_hash` and `igneum_hash_bound`).
|
||||
fn opencl_instr_lines(p: &Program) -> String {
|
||||
fn opencl_instr_lines(p: &Program, dataset_log2: u32) -> String {
|
||||
let mut s = String::with_capacity(6000);
|
||||
let era = p.class.era;
|
||||
for (k, ins) in p.instrs.iter().enumerate() {
|
||||
let d = format!("r{}", ins.dst);
|
||||
let a = format!("r{}", ins.src);
|
||||
|
|
@ -913,8 +1050,10 @@ fn opencl_instr_lines(p: &Program) -> String {
|
|||
Op::Rotr => format!("{d} = rotr_var({d}, {a});"),
|
||||
Op::Mad => format!("{d} = {a} * {b} + {d};"),
|
||||
Op::Shfl => format!("{{ uint t_; IGNEUM_SHFL_XOR(t_, {a}, {}u); {d} = {d} ^ t_; }}", ins.mask),
|
||||
Op::Load if load_width(ins) > 1 => wide_load_stmt(CoreDialect::OpenCl, &d, &a, ins.width, WideSource::Stored, None),
|
||||
Op::Load => format!("{d} = {d} ^ ds[{a} & mask];"),
|
||||
Op::Load if load_width(ins) > 1 => {
|
||||
wide_load_stmt(CoreDialect::OpenCl, &d, &load_index_expr(CoreDialect::OpenCl, era.as_ref(), ins, &a, dataset_log2), ins.width, WideSource::Stored, None)
|
||||
}
|
||||
Op::Load => format!("{d} = {d} ^ ds[{}];", load_index_expr(CoreDialect::OpenCl, era.as_ref(), ins, &a, dataset_log2)),
|
||||
Op::WLoad => format!("{{ uint t_; IGNEUM_BCAST0(t_, {a}); {d} = {d} ^ ds[(t_ & wmask) + lane]; }}"),
|
||||
Op::Scratch => scratch_stmt(CoreDialect::OpenCl, &d, &a, p.class.scratch_slot_mask()),
|
||||
};
|
||||
|
|
@ -927,28 +1066,48 @@ fn opencl_instr_lines(p: &Program) -> String {
|
|||
/// fifth argument (`__global const uint* initw`, 8 words, `bind::block_init_words`). One source file so the serve
|
||||
/// mode of proto-opencl/host.c builds cache fill, dataset build and the bound hash from it at runtime.
|
||||
pub fn opencl_kernel_bound(p: &Program, memhard: Option<&MixParams>) -> String {
|
||||
let mut s = opencl_kernel(p, memhard);
|
||||
opencl_kernel_bound_at(p, memhard, DEFAULT_DATASET_LOG2)
|
||||
}
|
||||
|
||||
/// [`opencl_kernel_bound`] at a dataset size (see [`cuda_kernel_at`]).
|
||||
pub fn opencl_kernel_bound_at(p: &Program, memhard: Option<&MixParams>, dataset_log2: u32) -> String {
|
||||
let mut s = opencl_kernel_at(p, memhard, dataset_log2);
|
||||
s.push('\n');
|
||||
s.push_str(
|
||||
"// Header-bound variant (bind.rs): the init words come from initw, not SEEDW. Same body as igneum_hash.\n",
|
||||
);
|
||||
let scratch_args = if p.has_scratch() { ", __global uint* scratch, uint groups, uint salt" } else { "" };
|
||||
s.push_str(&format!("IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw{scratch_args}) {{\n"));
|
||||
let (setup, unit_loop) = persistent_prologue_parts(CoreDialect::OpenCl, p.class.scratch_words_per_lane());
|
||||
if p.has_scratch() {
|
||||
s.push_str(&persistent_prologue(CoreDialect::OpenCl, p.class.scratch_words_per_lane()));
|
||||
s.push_str(&setup);
|
||||
} else {
|
||||
s.push_str(" uint gid = (uint)get_global_id(0);\n");
|
||||
}
|
||||
s.push_str(" uint lid = (uint)get_local_id(0);\n");
|
||||
s.push_str(" uint nonce = baseNonce + gid;\n");
|
||||
s.push_str(" uint r0, r1, r2, r3, r4, r5, r6, r7;\n");
|
||||
s.push_str(" uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7];\n");
|
||||
s.push_str("#if IGNEUM_EXCHANGE == 0\n");
|
||||
s.push_str(" IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);\n");
|
||||
s.push_str(" uint xk = 0u;\n");
|
||||
s.push_str("#else\n");
|
||||
s.push_str(" (void)lid;\n");
|
||||
s.push_str("#endif\n");
|
||||
if p.has_scratch() {
|
||||
// the __local exchange buffer must sit at the kernel's outermost scope: declare it, then open the unit loop
|
||||
s.push_str("#if IGNEUM_EXCHANGE == 0\n");
|
||||
s.push_str(" IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);\n");
|
||||
s.push_str(" uint xk = 0u;\n");
|
||||
s.push_str("#else\n");
|
||||
s.push_str(" (void)lid;\n");
|
||||
s.push_str("#endif\n");
|
||||
s.push_str(&unit_loop);
|
||||
s.push_str(" uint nonce = baseNonce + gid;\n");
|
||||
s.push_str(" uint r0, r1, r2, r3, r4, r5, r6, r7;\n");
|
||||
s.push_str(" uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7];\n");
|
||||
} else {
|
||||
s.push_str(" uint nonce = baseNonce + gid;\n");
|
||||
s.push_str(" uint r0, r1, r2, r3, r4, r5, r6, r7;\n");
|
||||
s.push_str(" uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7];\n");
|
||||
s.push_str("#if IGNEUM_EXCHANGE == 0\n");
|
||||
s.push_str(" IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);\n");
|
||||
s.push_str(" uint xk = 0u;\n");
|
||||
s.push_str("#else\n");
|
||||
s.push_str(" (void)lid;\n");
|
||||
s.push_str("#endif\n");
|
||||
}
|
||||
if p.has_wide() {
|
||||
s.push_str(" uint lane = lid & 31u;\n uint wmask = mask & ~31u;\n");
|
||||
}
|
||||
|
|
@ -960,7 +1119,7 @@ pub fn opencl_kernel_bound(p: &Program, memhard: Option<&MixParams>) -> String {
|
|||
));
|
||||
}
|
||||
s.push_str(&format!("\n for (uint it = 0u; it < {ITERATIONS}u; ++it) {{\n uint sel = r0;\n"));
|
||||
s.push_str(&opencl_instr_lines(p));
|
||||
s.push_str(&opencl_instr_lines(p, dataset_log2));
|
||||
s.push_str(" }\n");
|
||||
s.push_str(" uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);\n");
|
||||
s.push_str(" uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);\n");
|
||||
|
|
@ -974,6 +1133,12 @@ pub fn opencl_kernel_bound(p: &Program, memhard: Option<&MixParams>) -> String {
|
|||
|
||||
/// The OpenCL C 1.2 kernel (`generateOpenCL`, kernel.cl).
|
||||
pub fn opencl_kernel(p: &Program, memhard: Option<&MixParams>) -> String {
|
||||
opencl_kernel_at(p, memhard, DEFAULT_DATASET_LOG2)
|
||||
}
|
||||
|
||||
/// [`opencl_kernel`] at a dataset size (see [`cuda_kernel_at`]).
|
||||
pub fn opencl_kernel_at(p: &Program, memhard: Option<&MixParams>, dataset_log2: u32) -> String {
|
||||
let layout = p.class.layout();
|
||||
let mut s = String::with_capacity(14000);
|
||||
s.push_str(&generated_by(&p.seed_string));
|
||||
s.push_str("// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal).\n");
|
||||
|
|
@ -1034,7 +1199,7 @@ pub fn opencl_kernel(p: &Program, memhard: Option<&MixParams>) -> String {
|
|||
s.push_str(DS_ELEM_BODY);
|
||||
s.push('\n');
|
||||
if let Some(mp) = memhard {
|
||||
s.push_str(&emit_memhard_core(mp, CoreDialect::OpenCl));
|
||||
s.push_str(&emit_memhard_core_layout(mp, CoreDialect::OpenCl, layout));
|
||||
s.push('\n');
|
||||
s.push_str("// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item.\n");
|
||||
s.push_str("// The same constants as memhard.h in this pack (one emitter, three dialects).\n");
|
||||
|
|
@ -1047,8 +1212,7 @@ pub fn opencl_kernel(p: &Program, memhard: Option<&MixParams>) -> String {
|
|||
s.push_str(" if (t < nItems) {\n");
|
||||
s.push_str(" uint s[16];\n");
|
||||
s.push_str(" mh_item(cache, t, s);\n");
|
||||
s.push_str(" __global uint* d = ds + ((ulong)t * 16u);\n");
|
||||
s.push_str(" for (uint i = 0u; i < 16u; ++i) d[i] = s[i];\n");
|
||||
s.push_str(&build_store(layout, CoreDialect::OpenCl, "ds", "t"));
|
||||
s.push_str(" }\n");
|
||||
s.push_str("}\n");
|
||||
s.push('\n');
|
||||
|
|
@ -1066,20 +1230,33 @@ pub fn opencl_kernel(p: &Program, memhard: Option<&MixParams>) -> String {
|
|||
s.push_str(&scratch_prelude(p, CoreDialect::OpenCl));
|
||||
let scratch_args = if p.has_scratch() { ", __global uint* scratch, uint groups, uint salt" } else { "" };
|
||||
s.push_str(&format!("IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask{scratch_args}) {{\n"));
|
||||
let (setup, unit_loop) = persistent_prologue_parts(CoreDialect::OpenCl, p.class.scratch_words_per_lane());
|
||||
if p.has_scratch() {
|
||||
s.push_str(&persistent_prologue(CoreDialect::OpenCl, p.class.scratch_words_per_lane()));
|
||||
s.push_str(&setup);
|
||||
} else {
|
||||
s.push_str(" uint gid = (uint)get_global_id(0);\n");
|
||||
}
|
||||
s.push_str(" uint lid = (uint)get_local_id(0);\n");
|
||||
s.push_str(" uint nonce = baseNonce + gid;\n");
|
||||
s.push_str(" uint r0, r1, r2, r3, r4, r5, r6, r7;\n");
|
||||
s.push_str("#if IGNEUM_EXCHANGE == 0\n");
|
||||
s.push_str(" IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);\n");
|
||||
s.push_str(" uint xk = 0u;\n");
|
||||
s.push_str("#else\n");
|
||||
s.push_str(" (void)lid;\n");
|
||||
s.push_str("#endif\n");
|
||||
if p.has_scratch() {
|
||||
s.push_str("#if IGNEUM_EXCHANGE == 0\n");
|
||||
s.push_str(" IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);\n");
|
||||
s.push_str(" uint xk = 0u;\n");
|
||||
s.push_str("#else\n");
|
||||
s.push_str(" (void)lid;\n");
|
||||
s.push_str("#endif\n");
|
||||
s.push_str(&unit_loop);
|
||||
s.push_str(" uint nonce = baseNonce + gid;\n");
|
||||
s.push_str(" uint r0, r1, r2, r3, r4, r5, r6, r7;\n");
|
||||
} else {
|
||||
s.push_str(" uint nonce = baseNonce + gid;\n");
|
||||
s.push_str(" uint r0, r1, r2, r3, r4, r5, r6, r7;\n");
|
||||
s.push_str("#if IGNEUM_EXCHANGE == 0\n");
|
||||
s.push_str(" IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);\n");
|
||||
s.push_str(" uint xk = 0u;\n");
|
||||
s.push_str("#else\n");
|
||||
s.push_str(" (void)lid;\n");
|
||||
s.push_str("#endif\n");
|
||||
}
|
||||
if p.has_wide() {
|
||||
s.push_str(" uint lane = lid & 31u;\n uint wmask = mask & ~31u;\n");
|
||||
}
|
||||
|
|
@ -1087,7 +1264,7 @@ pub fn opencl_kernel(p: &Program, memhard: Option<&MixParams>) -> String {
|
|||
s.push_str(&init_line(p, "uint", i));
|
||||
}
|
||||
s.push_str(&format!("\n for (uint it = 0u; it < {ITERATIONS}u; ++it) {{\n uint sel = r0;\n"));
|
||||
s.push_str(&opencl_instr_lines(p));
|
||||
s.push_str(&opencl_instr_lines(p, dataset_log2));
|
||||
s.push_str(" }\n");
|
||||
s.push_str(" uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);\n");
|
||||
s.push_str(" uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);\n");
|
||||
|
|
@ -1152,6 +1329,7 @@ pub fn program_header(p: &Program, day: &str, ds: &DatasetSource) -> String {
|
|||
s.push_str(&program_class_header_lines(p));
|
||||
s.push_str(&class_header_lines(p));
|
||||
s.push_str(&scratch_header_lines(p));
|
||||
s.push_str(&era_header_lines(p));
|
||||
s.push_str("// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h)\n");
|
||||
s.push_str(&format!("#define IGNEUM_DATASET_MODE {}\n", if memhard.is_some() { 1 } else { 0 }));
|
||||
s.push('\n');
|
||||
|
|
@ -1213,7 +1391,7 @@ pub fn cuda_memhard_header(p: &Program, mp: &MixParams) -> String {
|
|||
s.push_str("#else\n");
|
||||
s.push_str("#define IGNEUM_HD static inline\n");
|
||||
s.push_str("#endif\n");
|
||||
s.push_str(&emit_memhard_core(mp, CoreDialect::Cuda));
|
||||
s.push_str(&emit_memhard_core_layout(mp, CoreDialect::Cuda, p.class.layout()));
|
||||
s
|
||||
}
|
||||
|
||||
|
|
@ -1363,6 +1541,22 @@ pub fn program_json(p: &Program, day: &str, ds: &DatasetSource) -> String {
|
|||
s.push_str(&format!(" \"scratch\": \"variant 5 (measurement only): persistent warps; a {kb} KiB scratch per warp of {slots} 16-byte slots per lane (lane-major); slot = src & 0x{smask:x}; a slot reads as its fill (scratch_fill(seed words, unit base nonce, lane, slot, j) = splitmix32(((base + lane) ^ seed[j]) + slot * 0x9e3779b1 + (j + 1) * 0x85ebca77), j in 0..2) until the unit writes it; read w0 w1 w2 (behind a per-unit tag on the GPU), x = fold(dst, w0, w1, w2), dst = x, rewrite (x ^ w1, rotl(x, 7) ^ w2, x + w0)\",\n", kb = p.class.scratch_kb, slots = p.class.scratch_slots_per_lane(), smask = p.class.scratch_slot_mask()));
|
||||
}
|
||||
s.push_str(&format!(" \"wide_load\": \"read-width experiment (5 October 2026, docs/plans/read-width.md), NOT the lottery hash: a load of W words (width field, 4 or 16) reads dataset[b .. b + W) with b = (src & mask) & ~(W - 1) and folds every word into dst: x = dst ^ w[0]; for j in 1..W: x = (rotl(x, {FOLD_ROT}) * 0x{FOLD_MUL:08x}) ^ w[j]; dst = x; width 1 is the plain load; the width is drawn per instruction from the class mix with one extra below(100) draw after the nine of version 2, and the program id is FNV-1a 64 over 'igneum-program-rw/' || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots\",\n"));
|
||||
if let Some(e) = p.class.era {
|
||||
s.push_str(" \"era\": {\n");
|
||||
s.push_str(&format!(" \"label\": {},\n", jstr(&e.label())));
|
||||
s.push_str(&format!(" \"seed_words\": [{}],\n", join_jhex(&e.words)));
|
||||
s.push_str(" \"draw\": \"docs/plans/era-layout.md 1.1: SplitMix64 seeded with seed_words[0] | seed_words[1] << 32 of seed_words_from_bytes('igneum-era/' || n_le64 || E_n); width = allowed[below(|allowed|)], stride_mul = low32(next()) | 1, stride_rot = 1 + below(31), then four next() draws for a partial Fisher-Yates over positions log2(W)..15 of which 4 - log2(W) are used\",\n");
|
||||
s.push_str(&format!(" \"allowed_widths\": [{}],\n", e.allowed_set().iter().map(|w| w.to_string()).collect::<Vec<_>>().join(", ")));
|
||||
s.push_str(&format!(" \"width_words\": {},\n", e.width_words));
|
||||
s.push_str(&format!(" \"stride_mul\": {},\n", jhex(e.stride_mul)));
|
||||
s.push_str(&format!(" \"stride_rot\": {},\n", e.stride_rot));
|
||||
s.push_str(&format!(" \"interleave\": [{}, {}, {}, {}],\n", e.pos[0], e.pos[1], e.pos[2], e.pos[3]));
|
||||
s.push_str(" \"address\": \"y = rotl(src * stride_mul, stride_rot); k = min(win, D - 26); idx = ((y & (mask >> k)) | ((off & (2^k - 1)) << (D - k))) & mask; a wide load aligns idx down to W words\",\n");
|
||||
s.push_str(" \"windows\": \"per instruction, after the width roll: win = below(3), off = low32(next()) & (2^win - 1); used on a load slot (the instruction's win and off fields)\",\n");
|
||||
s.push_str(" \"dataset_word\": \"dataset[w] = item(t(w))[j(w)]: j(w) gathers the bits of w at the interleave positions, t(w) is w with those bits removed\",\n");
|
||||
s.push_str(" \"program_id_suffix\": \"'era/' || allowed[3] || width_words || stride_mul_le32 || stride_rot_le32 || interleave[4]\"\n");
|
||||
s.push_str(" },\n");
|
||||
}
|
||||
}
|
||||
s.push_str(&format!(
|
||||
" \"op_mix\": {{{}}},\n",
|
||||
|
|
@ -1447,8 +1641,9 @@ pub fn program_json(p: &Program, day: &str, ds: &DatasetSource) -> String {
|
|||
let n = p.instrs.len();
|
||||
for (k, ins) in p.instrs.iter().enumerate() {
|
||||
if !p.class.is_v2() {
|
||||
let era_fields = if p.class.era.is_some() { format!(", \"win\": {}, \"off\": {}", ins.win, ins.off) } else { String::new() };
|
||||
s.push_str(&format!(
|
||||
" {{\"i\": {k}, \"op\": {}, \"dst\": {}, \"src\": {}, \"src2\": {}, \"imm\": {}, \"imm2\": {}, \"rot\": {}, \"bit\": {}, \"mask\": {}, \"width\": {}}}",
|
||||
" {{\"i\": {k}, \"op\": {}, \"dst\": {}, \"src\": {}, \"src2\": {}, \"imm\": {}, \"imm2\": {}, \"rot\": {}, \"bit\": {}, \"mask\": {}, \"width\": {}{era_fields}}}",
|
||||
jstr(ins.op.name()),
|
||||
ins.dst,
|
||||
ins.src,
|
||||
|
|
@ -1563,13 +1758,14 @@ pub fn export_pack(epoch: &Epoch, day: &str, source: &str) -> Pack {
|
|||
let memhard = ds.memhard().map(|m| &m.params);
|
||||
let bases = PACK_VECTOR_BASES.to_vec();
|
||||
let outs: Vec<[u64; 32]> = bases.iter().map(|&b| epoch.hash_warp(b)).collect();
|
||||
// the self-test words under the program's layout (era layout; linear for every other class)
|
||||
let mut v = PackVectors {
|
||||
head: (0..16).map(|i| ds.word(i)).collect(),
|
||||
last: ds.word(mask),
|
||||
head: (0..16).map(|i| epoch.dataset_word(i)).collect(),
|
||||
last: epoch.dataset_word(mask),
|
||||
sample_idx: sample_indices(mask),
|
||||
..Default::default()
|
||||
};
|
||||
v.sample_val = v.sample_idx.iter().map(|&i| ds.word(i)).collect();
|
||||
v.sample_val = v.sample_idx.iter().map(|&i| epoch.dataset_word(i)).collect();
|
||||
if let Some(m) = ds.memhard() {
|
||||
let w = m.cache.words();
|
||||
v.cache_head = w[..16].to_vec();
|
||||
|
|
@ -1581,19 +1777,19 @@ pub fn export_pack(epoch: &Epoch, day: &str, source: &str) -> Pack {
|
|||
let mut files = vec![
|
||||
("program.json".to_string(), program_json(p, day, ds)),
|
||||
("vectors.json".to_string(), vectors_json(p, day, ds.log2_words, &bases, &outs, &v, mask, source, is_mh)),
|
||||
("kernel.cu".to_string(), cuda_kernel(p, memhard)),
|
||||
("kernel.cl".to_string(), opencl_kernel(p, memhard)),
|
||||
("kernel.cu".to_string(), cuda_kernel_at(p, memhard, ds.log2_words)),
|
||||
("kernel.cl".to_string(), opencl_kernel_at(p, memhard, ds.log2_words)),
|
||||
("program.h".to_string(), program_header(p, day, ds)),
|
||||
("vectors.h".to_string(), vectors_header(p, &bases, &outs, &v, mask, source, is_mh)),
|
||||
("program.metal".to_string(), metal_program(p, ds.log2_words, LoadSource::Stored)),
|
||||
// Header-bound kernels (3 October 2026, bind.rs): new files, the seven above are unchanged.
|
||||
("program_bound.metal".to_string(), metal_program_bound(p, ds.log2_words)),
|
||||
("kernel_bound.cu".to_string(), cuda_kernel_bound(p, memhard)),
|
||||
("kernel_bound.cl".to_string(), opencl_kernel_bound(p, memhard)),
|
||||
("kernel_bound.cu".to_string(), cuda_kernel_bound_at(p, memhard, ds.log2_words)),
|
||||
("kernel_bound.cl".to_string(), opencl_kernel_bound_at(p, memhard, ds.log2_words)),
|
||||
];
|
||||
if let Some(mp) = memhard {
|
||||
files.push(("memhard.h".to_string(), cuda_memhard_header(p, mp)));
|
||||
files.push(("memhard.metal".to_string(), metal_memhard(mp)));
|
||||
files.push(("memhard.metal".to_string(), metal_memhard_layout(mp, p.class.layout())));
|
||||
}
|
||||
Pack { files, bases, outs, vectors: v }
|
||||
}
|
||||
|
|
|
|||
|
|
@ -19,6 +19,12 @@
|
|||
//! kept as [`generate_v1`] for the census tool and the lever measurements of `proto-metal/MEMHARD.md`. Its
|
||||
//! programs are not the lottery hash and no pack or vector of version 1 is current.
|
||||
//!
|
||||
//! Era layout (5 October 2026, Counter ASIC 2.0 layers 4 and 8, `docs/plans/era-layout.md`; NOT the lottery hash,
|
||||
//! behind [`LoadClass::era`]): [`EraParams`] drawn from the era seed `E_n` by [`era_draw`] (the load width, a stride
|
||||
//! multiplier and rotation, the interleave of item words over the dataset), and per load site two more draws (a
|
||||
//! window of the dataset: a half, a quarter or all of it, at a drawn offset). The load address is
|
||||
//! [`crate::verify::load_index`]. An era class takes 12 draws per instruction, so its stream differs from version 2.
|
||||
//!
|
||||
//! Read-width experiment (5 October 2026, gate 1, `docs/plans/read-width.md`; NOT the lottery hash, behind
|
||||
//! [`LoadClass`]): a program class whose `load` reads `W` bytes (4, 16 or 64: 1, 4 or 16 words, aligned to `W`)
|
||||
//! and folds every word into `dst` (`verify::fold_words`), with the width fixed per class or drawn per load from
|
||||
|
|
@ -27,7 +33,7 @@
|
|||
//! version 2 and its program id carries the class.
|
||||
|
||||
use crate::accept::{check, Reject};
|
||||
use crate::seed::{fnv1a64, program_rng, seed_words_from_bytes};
|
||||
use crate::seed::{fnv1a64, program_rng, seed_words_from_bytes, SplitMix64};
|
||||
|
||||
/// Iterations of the instruction list per hash.
|
||||
pub const ITERATIONS: usize = 8;
|
||||
|
|
@ -141,6 +147,11 @@ pub struct Instr {
|
|||
/// Words read by a `load`: 1 (the lottery hash, 4 bytes), 4 or 16 (the read-width experiment). 1 on every
|
||||
/// other op.
|
||||
pub width: u8,
|
||||
/// Era layout, layer 8: the window shrink of this load site, 0..2 (the dataset, a half, a quarter). 0 on every
|
||||
/// op of every other class.
|
||||
pub win: u8,
|
||||
/// Era layout, layer 8: which aligned window, below `2^win`. 0 on every op of every other class.
|
||||
pub off: u8,
|
||||
}
|
||||
|
||||
#[derive(Clone, Debug, PartialEq, Eq)]
|
||||
|
|
@ -188,6 +199,121 @@ pub struct LoadClass {
|
|||
/// Cache growth rule, option C (`memhard::growth_doublings`): the cache doubles when the dataset doubles. `false`
|
||||
/// for version 2 (the cache is 2^26 words on every day), `true` for v3.
|
||||
pub growth: bool,
|
||||
/// Era layout (`docs/plans/era-layout.md`): `Some` turns on the strided, windowed load address and the
|
||||
/// interleaved dataset mapping with the parameters drawn from the era seed. `None` for every other class.
|
||||
pub era: Option<EraParams>,
|
||||
}
|
||||
|
||||
/// The parameters one era draws from its seed `E_n` (`docs/plans/era-layout.md` section 1.1, the proposed text of
|
||||
/// spec 01 section 1.13.1). `Copy` so the class stays `Copy`; the eight words of the era stream's seed and the era
|
||||
/// index are carried so a pack can say where the draw came from.
|
||||
#[derive(Clone, Copy, Debug, PartialEq, Eq, Hash)]
|
||||
pub struct EraParams {
|
||||
/// `seed_words_from_bytes("igneum-era/" || E_n)` (`E_n` commits to the era index through the VDF input of spec
|
||||
/// 04 section 4.4 step 2, so the index does not enter the draw).
|
||||
pub words: [u32; 8],
|
||||
/// The genesis-fixed set the width is drawn from, ascending, zero-padded (`[1, 0, 0]` pins 4 bytes: the
|
||||
/// read-width decision of 5 October 2026, v2's 128 x 4 B stays).
|
||||
pub allowed: [u8; 3],
|
||||
/// Layer 4, item size: the words one load folds (1, 4 or 16), drawn from the allowed set.
|
||||
pub width_words: u8,
|
||||
/// Layer 4, stride: `y = rotl(x * stride_mul, stride_rot)`; the multiplier is odd, the rotation in 1..31.
|
||||
pub stride_mul: u32,
|
||||
pub stride_rot: u32,
|
||||
/// Layer 4, interleave: the four ascending bit positions (0..15) of the word-within-item bits in the word
|
||||
/// index; the first `log2(width_words)` are `0..`, so one aligned load stays inside one item.
|
||||
pub pos: [u8; 4],
|
||||
}
|
||||
|
||||
/// Domain tag of the era stream seed.
|
||||
pub const ERA_TAG: &[u8] = b"igneum-era/";
|
||||
|
||||
impl EraParams {
|
||||
/// `seed_words_from_bytes("igneum-era/" || era_bytes)`.
|
||||
pub fn stream_words(era_bytes: &[u8]) -> [u32; 8] {
|
||||
let mut b = Vec::with_capacity(ERA_TAG.len() + era_bytes.len());
|
||||
b.extend_from_slice(ERA_TAG);
|
||||
b.extend_from_slice(era_bytes);
|
||||
seed_words_from_bytes(&b)
|
||||
}
|
||||
|
||||
/// The short label of the era in class names and pack lines: the first stream word as hex.
|
||||
pub fn label(&self) -> String {
|
||||
format!("{:08x}", self.words[0])
|
||||
}
|
||||
|
||||
/// The era bytes of a test seed string: the 32 bytes (little-endian words) of `seed_words_from_bytes(s)`.
|
||||
pub fn test_era_bytes(s: &str) -> [u8; 32] {
|
||||
let w = seed_words_from_bytes(s.as_bytes());
|
||||
let mut out = [0u8; 32];
|
||||
for (i, x) in w.iter().enumerate() {
|
||||
out[i * 4..i * 4 + 4].copy_from_slice(&x.to_le_bytes());
|
||||
}
|
||||
out
|
||||
}
|
||||
|
||||
/// The dataset layout this era's loads and build use.
|
||||
pub fn layout(&self) -> crate::memhard::Layout {
|
||||
crate::memhard::Layout { pos: self.pos }
|
||||
}
|
||||
|
||||
/// The allowed widths as a slice (the non-zero entries).
|
||||
pub fn allowed_set(&self) -> Vec<u8> {
|
||||
self.allowed.iter().copied().filter(|&w| w != 0).collect()
|
||||
}
|
||||
|
||||
/// The bytes that enter the program id after `"era/"`: the allowed set, width, multiplier, rotation, positions.
|
||||
pub fn id_bytes(&self) -> Vec<u8> {
|
||||
let mut b = Vec::with_capacity(3 + 1 + 4 + 4 + 4);
|
||||
b.extend_from_slice(&self.allowed);
|
||||
b.push(self.width_words);
|
||||
b.extend_from_slice(&self.stride_mul.to_le_bytes());
|
||||
b.extend_from_slice(&self.stride_rot.to_le_bytes());
|
||||
b.extend_from_slice(&self.pos);
|
||||
b
|
||||
}
|
||||
}
|
||||
|
||||
/// The era draw (`docs/plans/era-layout.md` section 1.1): seven draws from one SplitMix64 stream seeded with words 0
|
||||
/// and 1 of [`EraParams::stream_words`], in this order: the width from `allowed` (ascending, a non-empty subset of
|
||||
/// [`WIDTH_WORDS`]; one element pins it, the draw is still consumed), the odd stride multiplier, the stride rotation
|
||||
/// in 1..31, then four draws for the interleave (a partial Fisher-Yates over the candidate positions `log2(W)..15`,
|
||||
/// `4 - log2(W)` of them used, the rest consumed).
|
||||
pub fn era_draw(era_bytes: &[u8], allowed: &[u8]) -> EraParams {
|
||||
assert!(!allowed.is_empty() && allowed.len() <= 3, "the allowed width set has 1 to 3 entries");
|
||||
for (i, &w) in allowed.iter().enumerate() {
|
||||
assert!(WIDTH_WORDS.contains(&w), "allowed width {w} is not 1, 4 or 16 words");
|
||||
assert!(i == 0 || allowed[i - 1] < w, "the allowed width set is ascending");
|
||||
}
|
||||
let words = EraParams::stream_words(era_bytes);
|
||||
let mut s = SplitMix64::new(words[0] as u64 | ((words[1] as u64) << 32));
|
||||
let width_words = allowed[s.below(allowed.len() as u64) as usize];
|
||||
let stride_mul = (s.next() as u32) | 1;
|
||||
let stride_rot = 1 + s.below(31) as u32;
|
||||
let b = width_words.trailing_zeros() as usize; // 0, 2 or 4
|
||||
let free = 4 - b;
|
||||
let mut c: Vec<u8> = (b as u8..16).collect();
|
||||
let mut r = [0u64; 4];
|
||||
for x in r.iter_mut() {
|
||||
*x = s.next();
|
||||
}
|
||||
for i in 0..free {
|
||||
let n = c.len() - i;
|
||||
let j = i + (r[i] % n as u64) as usize;
|
||||
c.swap(i, j);
|
||||
}
|
||||
let mut chosen: Vec<u8> = c[..free].to_vec();
|
||||
chosen.sort_unstable();
|
||||
let mut pos = [0u8; 4];
|
||||
for i in 0..b {
|
||||
pos[i] = i as u8;
|
||||
}
|
||||
for (i, p) in chosen.iter().enumerate() {
|
||||
pos[b + i] = *p;
|
||||
}
|
||||
let mut al = [0u8; 3];
|
||||
al[..allowed.len()].copy_from_slice(allowed);
|
||||
EraParams { words, allowed: al, width_words, stride_mul, stride_rot, pos }
|
||||
}
|
||||
|
||||
/// Scratch geometry (variant 5): 16-byte slots, lane-major, 32 lanes per warp; `scratch_kb` KiB per warp gives
|
||||
|
|
@ -213,13 +339,40 @@ impl LoadClass {
|
|||
impl LoadClass {
|
||||
/// Generator version 2 as adopted on 4 October 2026: 16 loads of one word. The lottery hash.
|
||||
pub const V2: LoadClass =
|
||||
LoadClass { mix: [100, 0, 0], load_slots: LOAD_SLOTS as u8, scratch: None, scratch_kb: 0, mixer_mult: 1, growth: false };
|
||||
LoadClass { mix: [100, 0, 0], load_slots: LOAD_SLOTS as u8, scratch: None, scratch_kb: 0, mixer_mult: 1, growth: false, era: None };
|
||||
|
||||
/// The construction decided for program class v3 on 5 October 2026 (Counter ASIC 2.0, `docs/plans/mixer-x4.md`):
|
||||
/// version 2 loads (16 slots of one word, no scratch, no width roll, so the program stream is version 2's), the
|
||||
/// mixer applied 4 times per round, and the cache growth rule. Name "mx4".
|
||||
pub const MX4: LoadClass =
|
||||
LoadClass { mix: [100, 0, 0], load_slots: LOAD_SLOTS as u8, scratch: None, scratch_kb: 0, mixer_mult: 4, growth: true };
|
||||
LoadClass { mix: [100, 0, 0], load_slots: LOAD_SLOTS as u8, scratch: None, scratch_kb: 0, mixer_mult: 4, growth: true, era: None };
|
||||
|
||||
/// The era class over `base` (`docs/plans/era-layout.md`): the parameters drawn by [`era_draw`]; when `allowed`
|
||||
/// has more than one width the drawn width becomes the class mix (every load that width), otherwise the base
|
||||
/// class's mix stands (the width rule of the read-width branch, pinned) and the era's `width_words` is the widest
|
||||
/// width that mix can draw.
|
||||
pub fn era(base: LoadClass, era_bytes: &[u8], allowed: &[u8]) -> LoadClass {
|
||||
let mut e = era_draw(era_bytes, allowed);
|
||||
let mut c = base;
|
||||
if allowed.len() > 1 {
|
||||
let i = WIDTH_WORDS.iter().position(|&w| w == e.width_words).unwrap();
|
||||
c.mix = [0, 0, 0];
|
||||
c.mix[i] = 100;
|
||||
} else {
|
||||
let widest = (0..3).rev().find(|&i| c.mix[i] > 0).map(|i| WIDTH_WORDS[i]).unwrap_or(1);
|
||||
if widest != e.width_words {
|
||||
// the interleave must keep the widest load inside one item: redraw the positions for that width
|
||||
e = era_draw(era_bytes, &[widest]);
|
||||
}
|
||||
}
|
||||
c.era = Some(e);
|
||||
c
|
||||
}
|
||||
|
||||
/// The dataset layout of this class ([`crate::memhard::Layout::LINEAR`] without an era).
|
||||
pub fn layout(&self) -> crate::memhard::Layout {
|
||||
self.era.map(|e| e.layout()).unwrap_or(crate::memhard::Layout::LINEAR)
|
||||
}
|
||||
|
||||
/// A fixed width (1, 4 or 16 words) with `load_slots` loads per program.
|
||||
pub fn fixed(width_words: u8, load_slots: u8) -> LoadClass {
|
||||
|
|
@ -338,7 +491,13 @@ impl LoadClass {
|
|||
|
||||
/// "v2", "w4", "w16", "w64", "w64x4", "mix50-35-15", "mix25-50-25x8", "scr4k32"; "mx4" for the v3 construction;
|
||||
/// any other mixer setting appends "m<mult>" and, with the growth rule, "g" ("v2m2", "w16m4g").
|
||||
/// An era class is the base name with "-era<first stream word as hex>" appended ("w4-era401998a5", "mx4-era...").
|
||||
pub fn name(&self) -> String {
|
||||
if let Some(e) = self.era {
|
||||
let base = LoadClass { era: None, ..*self };
|
||||
let base_name = if base.is_v2() { "w4".to_string() } else { base.name() };
|
||||
return format!("{base_name}-era{}", e.label());
|
||||
}
|
||||
if self.is_v2() {
|
||||
return "v2".to_string();
|
||||
}
|
||||
|
|
@ -417,7 +576,23 @@ pub enum ProgramClass {
|
|||
/// "22:00 decided", `docs/plans/mixer-x4.md`): [`LoadClass::MX4`], version 2 loads (the width stays 4 bytes, the
|
||||
/// per-load mix and the scratch share are out), the mixer applied 4 times per round and the cache growth rule. The
|
||||
/// placeholder of the seam (w16) is replaced here; nothing else in the seam names the class.
|
||||
pub const V3_CLASS: LoadClass = LoadClass::MX4;
|
||||
/// Composed on 5 October 2026 (branch ca2-era): the era layout of `docs/plans/era-layout.md` is drawn inside this class by
|
||||
/// [`generate_from_seed_bytes_program_class`] (`LoadClass::era(V3_CLASS, era, &V3_ALLOWED)`); here `era` is `None`.
|
||||
pub const V3_CLASS: LoadClass = LoadClass { era: None, ..LoadClass::MX4 };
|
||||
|
||||
/// The width set class v3's era draw chooses from: 4 bytes only (the read-width decision of 5 October 2026; the
|
||||
/// draw is consumed, so widening the set at genesis keeps the derivation).
|
||||
pub const V3_ALLOWED: [u8; 1] = [1];
|
||||
|
||||
/// The program of an era class (generator version 3): `base` with the era parameters drawn from `era_bytes`
|
||||
/// over the width set `allowed`, generator 3 stamped and the era bytes recorded (`docs/plans/era-layout.md`).
|
||||
/// The chain's path is this with `base = V3_CLASS` and `allowed = V3_ALLOWED`.
|
||||
pub fn generate_era(seed_string: &str, seed_bytes: &[u8], base: LoadClass, era_bytes: &[u8], allowed: &[u8]) -> Program {
|
||||
let mut p = generate_from_seed_bytes_class(seed_string, seed_bytes, LoadClass::era(base, era_bytes, allowed));
|
||||
p.generator = GENERATOR_VERSION_V3;
|
||||
p.era_bytes = Some(era_bytes.to_vec());
|
||||
p
|
||||
}
|
||||
|
||||
impl ProgramClass {
|
||||
/// The load class this program class draws from.
|
||||
|
|
@ -581,6 +756,10 @@ pub fn program_id_class(generator: u32, seed: &[u32; 8], attempt: u32, class: &L
|
|||
b.push(class.mixer_mult);
|
||||
b.push(class.growth as u8);
|
||||
}
|
||||
if let Some(e) = class.era {
|
||||
b.extend_from_slice(b"era/");
|
||||
b.extend_from_slice(&e.id_bytes());
|
||||
}
|
||||
fnv1a64(&b)
|
||||
}
|
||||
|
||||
|
|
@ -715,11 +894,23 @@ pub fn candidate_from_words_class(
|
|||
// Version 2 loads take no width roll, so a mixer class with version 2 loads draws the version 2 program
|
||||
let width = if class.takes_width_roll() { class.width_for_roll(rng.below(100)) } else { 1 };
|
||||
let width = if op == Op::Load { width } else { 1 };
|
||||
// Era layout, layer 8: two window draws per instruction (drawn on every slot, used on a load slot).
|
||||
let (win, off) = if class.era.is_some() {
|
||||
let k = rng.below(3) as u8;
|
||||
let o = (rng.next() as u32 & ((1u32 << k) - 1)) as u8;
|
||||
if op == Op::Load {
|
||||
(k, o)
|
||||
} else {
|
||||
(0, 0)
|
||||
}
|
||||
} else {
|
||||
(0, 0)
|
||||
};
|
||||
if op.is_load() {
|
||||
fresh[src as usize] = false;
|
||||
}
|
||||
fresh[dst as usize] = true;
|
||||
instrs.push(Instr { op, dst: dst as u8, src: src as u8, src2: b as u8, imm, imm2, rot, bit: bit as u8, mask, width });
|
||||
instrs.push(Instr { op, dst: dst as u8, src: src as u8, src2: b as u8, imm, imm2, rot, bit: bit as u8, mask, width, win, off });
|
||||
}
|
||||
Program {
|
||||
seed_string: seed_string.to_string(),
|
||||
|
|
@ -793,6 +984,11 @@ pub fn generate_from_seed_bytes_class(seed_string: &str, seed_bytes: &[u8], clas
|
|||
/// [`generate_from_seed_bytes`] exactly; class v3 draws from [`V3_CLASS`] and stamps generator version 3 on the
|
||||
/// program, so its packs and its id say generator 3 (spec 01 sections 1.4.5 and 1.4.6).
|
||||
pub fn generate_from_seed_bytes_program_class(seed_string: &str, seed_bytes: &[u8], class: ProgramClass, era_bytes: Option<&[u8]>) -> Program {
|
||||
// Class v3 with an era seed: the era layout (docs/plans/era-layout.md) drawn from the era bytes inside V3_CLASS.
|
||||
// Without era bytes (a template before the era is known) the bare V3_CLASS stands.
|
||||
if let (ProgramClass::V3, Some(era)) = (class, era_bytes) {
|
||||
return generate_era(seed_string, seed_bytes, V3_CLASS, era, &V3_ALLOWED);
|
||||
}
|
||||
let mut p = generate_from_seed_bytes_class(seed_string, seed_bytes, class.load_class());
|
||||
p.generator = class.generator_version();
|
||||
// The era is a property of class v3 programs; a version 2 program never records one, so the pinned v2 packs
|
||||
|
|
@ -908,7 +1104,7 @@ pub fn generate_v1_from_words(seed_string: &str, seed: [u32; 8], cfg: &Generator
|
|||
if op == Op::Load && bit * 100 < cfg.wide_frac * 32 {
|
||||
op = Op::WLoad;
|
||||
}
|
||||
instrs.push(Instr { op, dst: dst as u8, src: a as u8, src2: b as u8, imm, imm2, rot, bit: bit as u8, mask, width: 1 });
|
||||
instrs.push(Instr { op, dst: dst as u8, src: a as u8, src2: b as u8, imm, imm2, rot, bit: bit as u8, mask, width: 1, win: 0, off: 0 });
|
||||
}
|
||||
Program {
|
||||
seed_string: seed_string.to_string(),
|
||||
|
|
@ -1114,10 +1310,14 @@ mod tests {
|
|||
assert_eq!(v3.generator, GENERATOR_VERSION_V3);
|
||||
assert_eq!(v3.era_bytes.as_deref(), Some(&[7u8; 32][..]));
|
||||
let v3_no_era = generate_from_seed_bytes_program_class("igneum-genesis", b"igneum-genesis", ProgramClass::V3, None);
|
||||
assert_eq!(v3_no_era.instrs, v3.instrs, "the placeholder class does not read the era");
|
||||
assert_eq!(v3_no_era.program_id(), v3.program_id());
|
||||
assert_ne!(v3_no_era.instrs, v3.instrs, "the era class takes two more draws per instruction");
|
||||
assert_eq!(v3_no_era.class, V3_CLASS);
|
||||
assert_eq!(v3.program_class(), ProgramClass::V3);
|
||||
assert_eq!(v3.class, V3_CLASS);
|
||||
assert_eq!(LoadClass { era: None, ..v3.class }, V3_CLASS, "the era rides inside V3_CLASS");
|
||||
assert_eq!(v3.class, LoadClass::era(V3_CLASS, &[7u8; 32], &V3_ALLOWED));
|
||||
assert_eq!(v3.class.era.unwrap().width_words, 1);
|
||||
assert_eq!(v3.class.layout(), v3.class.era.unwrap().layout());
|
||||
assert_ne!(generate_from_seed_bytes_program_class("igneum-genesis", b"igneum-genesis", ProgramClass::V3, Some(&[8u8; 32])).class, v3.class);
|
||||
assert!(check(&v3).is_ok());
|
||||
assert_eq!(v3.program_id(), program_id(GENERATOR_VERSION_V3, &v3.seed, v3.attempt));
|
||||
assert_ne!(v3.program_id(), program_id(GENERATOR_VERSION, &v3.seed, v3.attempt));
|
||||
|
|
@ -1134,6 +1334,106 @@ mod tests {
|
|||
assert_eq!(ProgramClass::V2.load_class(), LoadClass::V2);
|
||||
}
|
||||
|
||||
/// Era layout: the draw is deterministic, within bounds, and the six test eras are pinned; an era program has
|
||||
/// 16 loads with window draws in bounds and nothing drawn on ALU slots; the class names and program ids separate
|
||||
/// the eras from each other and from every other class.
|
||||
#[test]
|
||||
fn era_draw_deterministic_and_bounded() {
|
||||
let all = [1u8, 4, 16];
|
||||
for n in 0..200u64 {
|
||||
let eb = EraParams::test_era_bytes(&format!("igneum-era-test/{n}"));
|
||||
let e = era_draw(&eb, &all);
|
||||
assert_eq!(e, era_draw(&eb, &all));
|
||||
assert_eq!(e.words, EraParams::stream_words(&eb));
|
||||
assert!(all.contains(&e.width_words));
|
||||
assert_eq!(e.stride_mul & 1, 1);
|
||||
assert!((1..=31).contains(&e.stride_rot));
|
||||
assert!(e.layout().is_valid(), "{:?}", e.pos);
|
||||
let b = e.width_words.trailing_zeros() as usize;
|
||||
for i in 0..b {
|
||||
assert_eq!(e.pos[i], i as u8, "the low positions are the identity for width {}", e.width_words);
|
||||
}
|
||||
// a pinned set consumes the draw and keeps the stride draws in step
|
||||
let pinned = era_draw(&eb, &[1]);
|
||||
assert_eq!(pinned.width_words, 1);
|
||||
assert_eq!((pinned.stride_mul, pinned.stride_rot), (e.stride_mul, e.stride_rot));
|
||||
assert_ne!(era_draw(&EraParams::test_era_bytes(&format!("igneum-era-test/{}", n + 1)), &all).words, e.words);
|
||||
}
|
||||
// the six test eras pinned at 4 bytes (docs/plans/era-layout.md section 5): ERA_VECTORS
|
||||
for (n, mul, rot, pos) in ERA_VECTORS {
|
||||
let e = era_draw(&EraParams::test_era_bytes(&format!("igneum-era-test/{n}")), &V3_ALLOWED);
|
||||
assert_eq!((e.width_words, e.stride_mul, e.stride_rot, e.pos), (1, mul, rot, pos), "era test seed {n}");
|
||||
}
|
||||
// era programs
|
||||
let mut ids = std::collections::HashSet::new();
|
||||
ids.insert(candidate("igneum-genesis", b"igneum-genesis", 0).program_id());
|
||||
ids.insert(candidate_class("igneum-genesis", b"igneum-genesis", 0, LoadClass::fixed(16, 16)).program_id());
|
||||
for n in 0..6u64 {
|
||||
let eb = EraParams::test_era_bytes(&format!("igneum-era-test/{n}"));
|
||||
let c = LoadClass::era(LoadClass::V2, &eb, &all);
|
||||
let e = c.era.unwrap();
|
||||
assert_eq!(c.mix.iter().position(|&m| m == 100).map(|i| WIDTH_WORDS[i]), Some(e.width_words));
|
||||
let base = match e.width_words {
|
||||
1 => "w4",
|
||||
4 => "w16",
|
||||
_ => "w64",
|
||||
};
|
||||
assert_eq!(c.name(), format!("{base}-era{}", e.label()));
|
||||
let p = candidate_class("igneum-genesis", b"igneum-genesis", 0, c);
|
||||
assert_eq!(p.class, c);
|
||||
assert_eq!(p.loads_per_hash(), 128);
|
||||
assert_eq!(p.bytes_per_hash(), 128 * 4 * e.width_words as usize);
|
||||
assert_ne!(p.instrs[0].op, Op::Load);
|
||||
for ins in &p.instrs {
|
||||
assert_ne!(ins.dst, ins.src);
|
||||
assert!((1..=31).contains(&ins.rot));
|
||||
if ins.op == Op::Load {
|
||||
assert_eq!(ins.width, e.width_words);
|
||||
assert!(ins.win <= 2 && (ins.off as u32) < (1u32 << ins.win), "win {} off {}", ins.win, ins.off);
|
||||
} else {
|
||||
assert_eq!((ins.width, ins.win, ins.off), (1, 0, 0));
|
||||
}
|
||||
}
|
||||
assert!(ids.insert(p.program_id()), "era {n}: program id collides");
|
||||
// the pinned form keeps the base mix and redraws the interleave for the base's widest width
|
||||
let pinned = LoadClass::era(LoadClass::mixed([50, 35, 15]), &eb, &[1]);
|
||||
assert_eq!(pinned.mix, [50, 35, 15]);
|
||||
assert_eq!(pinned.era.unwrap().width_words, 16);
|
||||
assert_eq!(pinned.era.unwrap().pos, [0, 1, 2, 3]);
|
||||
assert!(ids.insert(candidate_class("igneum-genesis", b"igneum-genesis", 0, pinned).program_id()));
|
||||
}
|
||||
// the window draws take two more draws per instruction: the stream differs from the read-width class
|
||||
let w64 = candidate_class("igneum-genesis", b"igneum-genesis", 0, LoadClass::fixed(16, 16));
|
||||
let era0 = candidate_class("igneum-genesis", b"igneum-genesis", 0, LoadClass::era(LoadClass::V2, &EraParams::test_era_bytes("igneum-era-test/0"), &all));
|
||||
assert_ne!(w64.instrs, era0.instrs);
|
||||
}
|
||||
|
||||
/// The era draws of the six test seeds at the 4-byte width: (test seed, stride multiplier, rotation, positions),
|
||||
/// as `igneum-pow show --era igneum-era-test/<n>` printed them on 5 October 2026 (docs/plans/era-layout.md 5).
|
||||
const ERA_VECTORS: [(u64, u32, u32, [u8; 4]); 6] = [
|
||||
(0, 0x625e5ab3, 19, [0, 2, 10, 15]),
|
||||
(1, 0xb2a9d70d, 6, [1, 3, 8, 13]),
|
||||
(2, 0x2b4a5b97, 28, [1, 3, 4, 8]),
|
||||
(3, 0x27ea7eff, 30, [2, 3, 8, 13]),
|
||||
(4, 0x4d38603d, 10, [2, 9, 13, 15]),
|
||||
(5, 0x03ac37ad, 22, [0, 2, 10, 13]),
|
||||
];
|
||||
|
||||
/// The six test eras generate accepted programs through the chain's class v3 path (the acceptance rule with the
|
||||
/// era address mirror); generator 3, the era bytes recorded, the class V3_CLASS with the era inside.
|
||||
#[test]
|
||||
fn era_programs_are_accepted() {
|
||||
for n in 0..6u64 {
|
||||
let eb = EraParams::test_era_bytes(&format!("igneum-era-test/{n}"));
|
||||
let p = generate_from_seed_bytes_program_class("igneum-genesis", b"igneum-genesis", ProgramClass::V3, Some(&eb));
|
||||
assert!(check(&p).is_ok(), "era {n}");
|
||||
assert!(p.attempt < MAX_ATTEMPTS);
|
||||
assert_eq!(p.generator, GENERATOR_VERSION_V3);
|
||||
assert_eq!(p.era_bytes.as_deref(), Some(&eb[..]));
|
||||
assert_eq!(p, generate_era("igneum-genesis", b"igneum-genesis", V3_CLASS, &eb, &V3_ALLOWED));
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn generate_returns_an_accepted_program() {
|
||||
let p = generate("igneum-genesis");
|
||||
|
|
|
|||
|
|
@ -10,6 +10,13 @@
|
|||
//!
|
||||
//! Read-width experiment (5 October 2026, docs/plans/read-width.md): `--class v2|w4|w16|w64|w64x4|p4,p16,p64[xN]`
|
||||
//! on every command selects the load class (default v2, the lottery hash). Nothing in a v2 run changes.
|
||||
//!
|
||||
//! Era layout (5 October 2026, docs/plans/era-layout.md): `--era igneum-era-test/<n>` (a test era seed: the 32 bytes
|
||||
//! of seed_words_from_bytes of the string, era index n) or `--era <n>:<64 hex>` (the chain's 32-byte era seed E_n)
|
||||
//! turns the chosen class into its era class; `--era-widths 4` (default: the read-width decision of 5 October 2026
|
||||
//! keeps v2's 4-byte load) is the allowed width set the era draws from; `4,16,64` lets the era draw the width.
|
||||
|
||||
use igneum_pow::generator::{EraParams, GENERATOR_VERSION_V3};
|
||||
|
||||
use igneum_pow::emit::export_pack;
|
||||
use igneum_pow::generator::{LoadClass, ProgramClass};
|
||||
|
|
@ -37,6 +44,44 @@ struct Args {
|
|||
program_class: Option<ProgramClass>,
|
||||
/// The era seed bytes a class v3 chain program records (`--era-hex`).
|
||||
era_hex: Option<String>,
|
||||
/// Era layout: `--era igneum-era-test/<n>` or `--era <n>:<64 hex>` composes the era class over `--class` with
|
||||
/// generator 3 and the era bytes recorded (the measurement packs: v2's mixer under the era layout).
|
||||
era: Option<(u64, Vec<u8>, String)>,
|
||||
era_widths: Vec<u8>,
|
||||
}
|
||||
|
||||
/// `igneum-era-test/<n>` or `<n>:<64 hex>` -> (index, 32 era bytes, label).
|
||||
fn parse_era(s: &str) -> Option<(u64, Vec<u8>, String)> {
|
||||
if let Some(n) = s.strip_prefix("igneum-era-test/") {
|
||||
let index: u64 = n.parse().ok()?;
|
||||
return Some((index, EraParams::test_era_bytes(s).to_vec(), s.to_string()));
|
||||
}
|
||||
let (n, hex) = s.split_once(':')?;
|
||||
let index: u64 = n.parse().ok()?;
|
||||
let bytes = igneum_pow::bind::unhex(hex)?;
|
||||
if bytes.len() != 32 {
|
||||
return None;
|
||||
}
|
||||
Some((index, bytes, format!("igneum-era/{index}/{hex}")))
|
||||
}
|
||||
|
||||
/// "4,16,64" (bytes) -> ascending words.
|
||||
fn parse_widths(s: &str) -> Option<Vec<u8>> {
|
||||
let mut v: Vec<u8> = s
|
||||
.split(',')
|
||||
.map(|x| match x.trim() {
|
||||
"4" => Some(1u8),
|
||||
"16" => Some(4),
|
||||
"64" => Some(16),
|
||||
_ => None,
|
||||
})
|
||||
.collect::<Option<Vec<_>>>()?;
|
||||
v.sort_unstable();
|
||||
v.dedup();
|
||||
if v.is_empty() {
|
||||
return None;
|
||||
}
|
||||
Some(v)
|
||||
}
|
||||
|
||||
fn usage() -> ! {
|
||||
|
|
@ -50,7 +95,9 @@ fn usage() -> ! {
|
|||
\x20 show the accepted program, one instruction per line\n\
|
||||
\x20 --class C load class: v2 (default), mx4 (class v3: mixer x4, cache growth), w4, w16, w64, w64x4, p4,p16,p64[xN], <class>m<mult>[g]\n\
|
||||
\x20 --days N days since genesis for the cache growth rule of a class with it (default 0: the 2^26-word cache)\n\
|
||||
\x20 --program-class v2|v3 the program class of the seam (v3 = generator 3 on V3_CLASS, the chain's own derivation; --era-hex records the era seed)"
|
||||
\x20 --program-class v2|v3 the program class of the seam (v3 = generator 3 on V3_CLASS, the chain's own derivation; --era-hex records the era seed)\n\
|
||||
\x20 --era E era layout over --class: igneum-era-test/<n> or <n>:<64 hex> (the 32-byte era seed E_n)\n\
|
||||
\x20 --era-widths 4[,16,64] the width set the era draws from, in bytes (default 4: pinned; more lets the era draw it)"
|
||||
);
|
||||
std::process::exit(2)
|
||||
}
|
||||
|
|
@ -72,6 +119,8 @@ fn parse() -> Args {
|
|||
days: 0,
|
||||
program_class: None,
|
||||
era_hex: None,
|
||||
era: None,
|
||||
era_widths: vec![1],
|
||||
};
|
||||
let mut it = std::env::args().skip(1);
|
||||
a.cmd = it.next().unwrap_or_else(|| usage());
|
||||
|
|
@ -92,12 +141,26 @@ fn parse() -> Args {
|
|||
"--days" => a.days = val().parse().unwrap_or_else(|_| usage()),
|
||||
"--program-class" => a.program_class = Some(ProgramClass::parse(&val()).unwrap_or_else(|| usage())),
|
||||
"--era-hex" => a.era_hex = Some(val()),
|
||||
"--era" => a.era = Some(parse_era(&val()).unwrap_or_else(|| usage())),
|
||||
"--era-widths" => a.era_widths = parse_widths(&val()).unwrap_or_else(|| usage()),
|
||||
_ => usage(),
|
||||
}
|
||||
}
|
||||
if let Some((_, bytes, _)) = &a.era {
|
||||
a.class = LoadClass::era(a.class, bytes, &a.era_widths);
|
||||
}
|
||||
a
|
||||
}
|
||||
|
||||
/// An era program is a class v3 program: generator 3 and the era bytes recorded (what the chain's
|
||||
/// `Epoch::from_chain_seeds` does); the pack then carries IGNEUM_PROGRAM_CLASS "v3" and IGNEUM_ERA_SEED_HEX.
|
||||
fn stamp_era(e: &mut Epoch, a: &Args) {
|
||||
if let Some((_, bytes, _)) = &a.era {
|
||||
e.program.generator = GENERATOR_VERSION_V3;
|
||||
e.program.era_bytes = Some(bytes.clone());
|
||||
}
|
||||
}
|
||||
|
||||
fn main() {
|
||||
let a = parse();
|
||||
let mode = if a.closed_form { DatasetMode::ClosedForm } else { DatasetMode::MemoryHard };
|
||||
|
|
@ -129,6 +192,12 @@ fn main() {
|
|||
/// dataset for `--days` through `Epoch::chain_dataset_day`; `--class` is ignored under a program class (the class
|
||||
/// names the load class). Closed-form mode is only for string seeds under the default class.
|
||||
fn epoch_of(a: &Args, mode: DatasetMode) -> (Epoch, String) {
|
||||
let (mut e, label) = epoch_of_class(a, mode);
|
||||
stamp_era(&mut e, a);
|
||||
(e, label)
|
||||
}
|
||||
|
||||
fn epoch_of_class(a: &Args, mode: DatasetMode) -> (Epoch, String) {
|
||||
let era = a.era_hex.as_ref().map(|h| igneum_pow::bind::unhex(h).unwrap_or_else(|| usage()));
|
||||
match (&a.epoch_hex, &a.day_hex) {
|
||||
(Some(eh), Some(dh)) => {
|
||||
|
|
@ -291,7 +360,11 @@ fn accept(a: &Args) {
|
|||
|
||||
fn show(a: &Args) {
|
||||
let (label, bytes) = seed_bytes_of(a);
|
||||
let p = igneum_pow::generator::generate_from_seed_bytes_class(&label, &bytes, a.class);
|
||||
let mut p = igneum_pow::generator::generate_from_seed_bytes_class(&label, &bytes, a.class);
|
||||
if let Some((_, eb, _)) = &a.era {
|
||||
p.generator = GENERATOR_VERSION_V3;
|
||||
p.era_bytes = Some(eb.clone());
|
||||
}
|
||||
println!(
|
||||
"seed \"{}\" generator v{} class {} attempt {} program id {:016x} seed words {}",
|
||||
p.seed_string,
|
||||
|
|
@ -302,6 +375,24 @@ fn show(a: &Args) {
|
|||
p.seed.iter().map(|w| format!("{w:08x}")).collect::<Vec<_>>().join(" ")
|
||||
);
|
||||
println!("op mix {} loads/hash {} bytes/hash {}", p.op_mix(), p.loads_per_hash(), p.bytes_per_hash());
|
||||
if let Some(e) = p.class.era {
|
||||
println!(
|
||||
"era {} ({}): width {} B, stride mul {:#010x} rot {}, interleave {:?}, windows (site:shrink:offset) {}",
|
||||
e.label(),
|
||||
a.era.as_ref().map(|x| x.2.as_str()).unwrap_or("?"),
|
||||
e.width_words as u32 * 4,
|
||||
e.stride_mul,
|
||||
e.stride_rot,
|
||||
e.pos,
|
||||
p.instrs
|
||||
.iter()
|
||||
.enumerate()
|
||||
.filter(|(_, i)| i.op == igneum_pow::generator::Op::Load)
|
||||
.map(|(k, i)| format!("{k}:{}:{}", i.win, i.off))
|
||||
.collect::<Vec<_>>()
|
||||
.join(" ")
|
||||
);
|
||||
}
|
||||
for (k, i) in p.instrs.iter().enumerate() {
|
||||
println!(
|
||||
"{k:2}: {:5} dst={} src={} src2={} imm={:#010x} imm2={:#010x} rot={} bit={} mask={}{}",
|
||||
|
|
|
|||
|
|
@ -198,6 +198,71 @@ impl MixParams {
|
|||
}
|
||||
}
|
||||
|
||||
/// The dataset layout (era layout, `docs/plans/era-layout.md` section 1.2): word `w` of the dataset holds word
|
||||
/// `j(w)` of item `t(w)`, where `j(w)` gathers the four bits of `w` at the ascending positions `pos` and `t(w)` is
|
||||
/// `w` with those bits removed. [`Layout::LINEAR`] (`pos = [0, 1, 2, 3]`) is `dataset[w] = item(w >> 4)[w & 15]`,
|
||||
/// the lottery hash's mapping. Every position is below 16, so the mapping is the same at every dataset size of
|
||||
/// at least 2^16 words and an item keeps its value at every size.
|
||||
#[derive(Clone, Copy, Debug, PartialEq, Eq, Hash)]
|
||||
pub struct Layout {
|
||||
pub pos: [u8; 4],
|
||||
}
|
||||
|
||||
impl Layout {
|
||||
pub const LINEAR: Layout = Layout { pos: [0, 1, 2, 3] };
|
||||
|
||||
pub fn is_linear(&self) -> bool {
|
||||
self.pos == [0, 1, 2, 3]
|
||||
}
|
||||
|
||||
/// Positions ascending, distinct, below 16.
|
||||
pub fn is_valid(&self) -> bool {
|
||||
self.pos.iter().all(|&p| p < 16) && (1..4).all(|i| self.pos[i] > self.pos[i - 1])
|
||||
}
|
||||
|
||||
/// `(t, j)` of word index `w`.
|
||||
#[inline(always)]
|
||||
pub fn split(&self, w: u32) -> (u32, u32) {
|
||||
if self.is_linear() {
|
||||
return (w >> 4, w & 15);
|
||||
}
|
||||
let mut j = 0u32;
|
||||
for (i, &p) in self.pos.iter().enumerate() {
|
||||
j |= ((w >> p) & 1) << i;
|
||||
}
|
||||
// remove the highest position first so the lower ones stay where they are
|
||||
let mut t = w;
|
||||
for &p in self.pos.iter().rev() {
|
||||
let p = p as u32;
|
||||
let low = (1u32 << p) - 1;
|
||||
t = (t & low) | ((t >> (p + 1)) << p);
|
||||
}
|
||||
(t, j)
|
||||
}
|
||||
|
||||
/// The word index of word `j` of item `t`: the inverse of [`Layout::split`].
|
||||
#[inline(always)]
|
||||
pub fn join(&self, t: u32, j: u32) -> u32 {
|
||||
if self.is_linear() {
|
||||
return (t << 4) | (j & 15);
|
||||
}
|
||||
// insert the lowest position first: every later position counts the bit just inserted
|
||||
let mut w = t;
|
||||
for (i, &p) in self.pos.iter().enumerate() {
|
||||
let p = p as u32;
|
||||
let low = (1u32 << p) - 1;
|
||||
w = ((w >> p) << (p + 1)) | (w & low) | (((j >> i) & 1) << p);
|
||||
}
|
||||
w
|
||||
}
|
||||
}
|
||||
|
||||
impl Default for Layout {
|
||||
fn default() -> Self {
|
||||
Layout::LINEAR
|
||||
}
|
||||
}
|
||||
|
||||
/// Round key `(r + 1) * 0x9E3779B9` mod 2^32.
|
||||
#[inline(always)]
|
||||
pub fn round_key(r: usize) -> u32 {
|
||||
|
|
@ -379,20 +444,28 @@ impl MemhardCpu {
|
|||
pub fn shape(&self) -> Shape {
|
||||
self.params.shape
|
||||
}
|
||||
/// `dataset[w] = item(w >> 4)[w & 15]`.
|
||||
/// `dataset[w] = item(w >> 4)[w & 15]` (the linear layout).
|
||||
pub fn word(&self, w: u32) -> u32 {
|
||||
derive_item(w >> 4, &self.params, &self.cache)[(w & 15) as usize]
|
||||
self.word_at(Layout::LINEAR, w)
|
||||
}
|
||||
/// `dataset[w] = item(t(w))[j(w)]` under `layout` (era layout; the layout is the program's, the cache the
|
||||
/// day's, so one cache serves every era of a day).
|
||||
pub fn word_at(&self, layout: Layout, w: u32) -> u32 {
|
||||
let (t, j) = layout.split(w);
|
||||
derive_item(t, &self.params, &self.cache)[j as usize]
|
||||
}
|
||||
/// `out[k] = dataset[idx[k]]` for every k, `idx.len() <= FETCH_MAX`. Equal items are derived once.
|
||||
/// Returns the number of distinct items derived.
|
||||
pub fn fetch(&self, idx: &[u32], out: &mut [u32]) -> usize {
|
||||
pub fn fetch(&self, idx: &[u32], out: &mut [u32], layout: Layout) -> usize {
|
||||
let n = idx.len();
|
||||
assert!(n <= FETCH_MAX && out.len() >= n);
|
||||
let mut uniq = [0u32; FETCH_MAX];
|
||||
let mut slot = [0u8; FETCH_MAX];
|
||||
let mut word = [0u8; FETCH_MAX];
|
||||
let mut u = 0usize;
|
||||
for k in 0..n {
|
||||
let t = idx[k] >> 4;
|
||||
let (t, j) = layout.split(idx[k]);
|
||||
word[k] = j as u8;
|
||||
let found = uniq[..u].iter().position(|&x| x == t);
|
||||
let j = match found {
|
||||
Some(j) => j,
|
||||
|
|
@ -407,20 +480,24 @@ impl MemhardCpu {
|
|||
let mut items = [[0u32; 16]; FETCH_MAX];
|
||||
derive_items(&uniq[..u], &self.params, &self.cache, &mut items);
|
||||
for k in 0..n {
|
||||
out[k] = items[slot[k] as usize][(idx[k] & 15) as usize];
|
||||
out[k] = items[slot[k] as usize][word[k] as usize];
|
||||
}
|
||||
u
|
||||
}
|
||||
/// `out[k][j] = dataset[base[k] + j]` for `j < width` (read-width experiment): `base[k]` is aligned to `width`
|
||||
/// words, so every lane's words lie in one item, derived once per distinct item. Returns the distinct items.
|
||||
pub fn fetch_wide(&self, base: &[u32], width: usize, out: &mut [[u32; 16]]) -> usize {
|
||||
/// words and the layout's low `log2(width)` positions are the identity, so every lane's words lie in one item
|
||||
/// at consecutive word offsets, derived once per distinct item. Returns the distinct items.
|
||||
pub fn fetch_wide(&self, base: &[u32], width: usize, out: &mut [[u32; 16]], layout: Layout) -> usize {
|
||||
let n = base.len();
|
||||
assert!(n <= FETCH_MAX && out.len() >= n && width <= 16);
|
||||
debug_assert!((0..width.trailing_zeros() as usize).all(|i| layout.pos[i] == i as u8), "a wide load needs the identity on its low positions");
|
||||
let mut uniq = [0u32; FETCH_MAX];
|
||||
let mut slot = [0u8; FETCH_MAX];
|
||||
let mut word = [0u8; FETCH_MAX];
|
||||
let mut u = 0usize;
|
||||
for k in 0..n {
|
||||
let t = base[k] >> 4;
|
||||
let (t, j0) = layout.split(base[k]);
|
||||
word[k] = j0 as u8;
|
||||
let j = match uniq[..u].iter().position(|&x| x == t) {
|
||||
Some(j) => j,
|
||||
None => {
|
||||
|
|
@ -434,7 +511,7 @@ impl MemhardCpu {
|
|||
let mut items = [[0u32; 16]; FETCH_MAX];
|
||||
derive_items(&uniq[..u], &self.params, &self.cache, &mut items);
|
||||
for k in 0..n {
|
||||
let o = (base[k] & 15) as usize;
|
||||
let o = word[k] as usize;
|
||||
out[k][..width].copy_from_slice(&items[slot[k] as usize][o..o + width]);
|
||||
}
|
||||
u
|
||||
|
|
@ -458,6 +535,39 @@ mod tests {
|
|||
assert_eq!(mp.shape, Shape::V2);
|
||||
}
|
||||
|
||||
/// Era layout: split and join are inverse, the linear layout is today's mapping, and an interleaved layout
|
||||
/// keeps every position below 16 so the mapping is the same at every size of at least 2^16 words.
|
||||
#[test]
|
||||
fn layout_split_join() {
|
||||
let lin = Layout::LINEAR;
|
||||
assert!(lin.is_linear() && lin.is_valid());
|
||||
for w in [0u32, 1, 15, 16, 17, 0x0fff_ffff, 0xffff_ffff] {
|
||||
assert_eq!(lin.split(w), (w >> 4, w & 15));
|
||||
assert_eq!(lin.join(w >> 4, w & 15), w);
|
||||
}
|
||||
let l = Layout { pos: [0, 1, 7, 12] };
|
||||
assert!(!l.is_linear() && l.is_valid());
|
||||
for w in [0u32, 1, 2, 3, 4, 127, 128, 129, 4095, 4096, 0x0fff_ffff, 0x1234_5678, 0xffff_ffff] {
|
||||
let (t, j) = l.split(w);
|
||||
assert!(j < 16);
|
||||
assert_eq!(l.join(t, j), w, "w {w:#x}");
|
||||
}
|
||||
// bits: j0 = bit 0, j1 = bit 1, j2 = bit 7, j3 = bit 12; t = the other 28 bits in order
|
||||
assert_eq!(l.split(0b1_0000_0000_0000), (0, 8));
|
||||
assert_eq!(l.split(1 << 7), (0, 4));
|
||||
assert_eq!(l.split(0b100), (1, 0));
|
||||
// every t in 0..2^(D-4) appears exactly once among w < 2^D (D = 16), with every j
|
||||
let mut seen = vec![0u32; 1 << 12];
|
||||
for w in 0..(1u32 << 16) {
|
||||
let (t, j) = l.split(w);
|
||||
seen[t as usize] |= 1 << j;
|
||||
}
|
||||
assert!(seen.iter().all(|&s| s == 0xffff));
|
||||
assert!(!Layout { pos: [0, 1, 1, 5] }.is_valid());
|
||||
assert!(!Layout { pos: [0, 1, 2, 16] }.is_valid());
|
||||
assert!(!Layout { pos: [1, 0, 2, 3] }.is_valid());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn chacha_block_is_a_permutation_plus_feedforward() {
|
||||
let x = [1u32; 16];
|
||||
|
|
|
|||
|
|
@ -1,10 +1,36 @@
|
|||
//! The CPU reference interpreter for one 32-lane warp (`cpuWarpTraced` in the Swift) and the API the node
|
||||
//! calls. Dataset words come from the memory-hard cache (default) or from the closed form (old packs).
|
||||
|
||||
use crate::generator::{generate, generate_class, Instr, LoadClass, Op, Program, ProgramClass, ITERATIONS, LANES};
|
||||
use crate::memhard::{MemhardCpu, Shape};
|
||||
use crate::generator::{generate, generate_class, EraParams, Instr, LoadClass, Op, Program, ProgramClass, ITERATIONS, LANES};
|
||||
use crate::memhard::{Layout, MemhardCpu, Shape};
|
||||
use crate::seed::day_key;
|
||||
|
||||
/// The load address of an era program (`docs/plans/era-layout.md` section 1.3): `y = rotl(x * M, R)`, then the
|
||||
/// window of the load site, `k = min(win, D - 26)` (0 when `D <= 26`), `idx = ((y & (MASK >> k)) | ((off &
|
||||
/// (2^k - 1)) << (D - k))) & MASK`. For every other class `idx = x & MASK`, the lottery hash's address. `mask` is
|
||||
/// `2^D - 1`. The acceptance mirror calls this at the rule's constant `D = 28`.
|
||||
#[inline(always)]
|
||||
pub fn load_index(era: Option<&EraParams>, ins: &Instr, x: u32, mask: u32, log2: u32) -> u32 {
|
||||
match era {
|
||||
None => x & mask,
|
||||
Some(e) => {
|
||||
let (wm, off) = window(ins, mask, log2);
|
||||
let y = x.wrapping_mul(e.stride_mul).rotate_left(e.stride_rot);
|
||||
((y & wm) | off) & mask
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// The window of a load site at a dataset of `2^log2` words: `(window mask, offset)` such that
|
||||
/// `idx = (y & window mask) | offset` lies in the site's aligned window of `2^(log2 - k)` words.
|
||||
#[inline(always)]
|
||||
pub fn window(ins: &Instr, mask: u32, log2: u32) -> (u32, u32) {
|
||||
let k = (ins.win as u32).min(log2.saturating_sub(26));
|
||||
let wm = mask >> k;
|
||||
let off = ((ins.off as u32) & ((1u32 << k) - 1)) << (log2 - k);
|
||||
(wm, off)
|
||||
}
|
||||
|
||||
/// Read-width experiment (5 October 2026): a `load` of `W` words folds every word into `dst`:
|
||||
/// `x = dst XOR w[0]; for j in 1..W: x = (rotl(x, FOLD_ROT) * FOLD_MUL) XOR w[j]; dst = x`. For `W = 1` this is the
|
||||
/// lottery hash's `dst XOR dataset[...]`. The fold is state-dependent (the rotate-multiply sits between the words),
|
||||
|
|
@ -195,18 +221,23 @@ impl DatasetSource {
|
|||
}
|
||||
}
|
||||
|
||||
/// `dataset[w & mask]`.
|
||||
/// `dataset[w & mask]` under the linear layout (the lottery hash).
|
||||
pub fn word(&self, w: u32) -> u32 {
|
||||
self.word_at(Layout::LINEAR, w)
|
||||
}
|
||||
|
||||
/// `dataset[w & mask]` under a program's layout (era layout). The closed form has no items and ignores it.
|
||||
pub fn word_at(&self, layout: Layout, w: u32) -> u32 {
|
||||
let w = w & self.mask;
|
||||
match &self.dataset {
|
||||
Dataset::ClosedForm { d0, d1 } => dataset_elem(w, *d0, *d1),
|
||||
Dataset::MemoryHard(m) => m.word(w),
|
||||
Dataset::MemoryHard(m) => m.word_at(layout, w),
|
||||
}
|
||||
}
|
||||
|
||||
/// `out[k] = dataset[idx[k]]`; indices are already masked. Returns items derived (0 for the closed form).
|
||||
#[inline]
|
||||
fn fetch(&self, idx: &[u32; LANES], out: &mut [u32; LANES]) -> usize {
|
||||
fn fetch(&self, idx: &[u32; LANES], out: &mut [u32; LANES], layout: Layout) -> usize {
|
||||
match &self.dataset {
|
||||
Dataset::ClosedForm { d0, d1 } => {
|
||||
for k in 0..LANES {
|
||||
|
|
@ -214,14 +245,14 @@ impl DatasetSource {
|
|||
}
|
||||
0
|
||||
}
|
||||
Dataset::MemoryHard(m) => m.fetch(idx, out),
|
||||
Dataset::MemoryHard(m) => m.fetch(idx, out, layout),
|
||||
}
|
||||
}
|
||||
|
||||
/// `out[k][j] = dataset[base[k] + j]` for `j < width`; bases are masked and aligned to `width` words
|
||||
/// (`width` 4 or 16, so a lane's words lie in one item). Returns items derived (0 for the closed form).
|
||||
#[inline]
|
||||
fn fetch_wide(&self, base: &[u32; LANES], width: usize, out: &mut [[u32; 16]; LANES]) -> usize {
|
||||
fn fetch_wide(&self, base: &[u32; LANES], width: usize, out: &mut [[u32; 16]; LANES], layout: Layout) -> usize {
|
||||
match &self.dataset {
|
||||
Dataset::ClosedForm { d0, d1 } => {
|
||||
for k in 0..LANES {
|
||||
|
|
@ -231,7 +262,7 @@ impl DatasetSource {
|
|||
}
|
||||
0
|
||||
}
|
||||
Dataset::MemoryHard(m) => m.fetch_wide(base, width, out),
|
||||
Dataset::MemoryHard(m) => m.fetch_wide(base, width, out, layout),
|
||||
}
|
||||
}
|
||||
}
|
||||
|
|
@ -260,6 +291,9 @@ pub fn interpret_warp(program: &Program, base_nonce: u32, ds: &DatasetSource) ->
|
|||
/// a block uses `I = bind::block_init_words(H, nonce)`.
|
||||
pub fn interpret_warp_init(program: &Program, seed: &[u32; 8], base_nonce: u32, ds: &DatasetSource) -> WarpResult {
|
||||
let mask = ds.mask;
|
||||
let log2 = ds.log2_words;
|
||||
let era = program.class.era;
|
||||
let layout = program.class.layout();
|
||||
let mut r = [[0u32; LANES]; 8];
|
||||
for lane in 0..LANES {
|
||||
let nonce = base_nonce.wrapping_add(lane as u32);
|
||||
|
|
@ -278,7 +312,7 @@ pub fn interpret_warp_init(program: &Program, seed: &[u32; 8], base_nonce: u32,
|
|||
for _ in 0..ITERATIONS {
|
||||
let sel = r[0];
|
||||
for ins in &program.instrs {
|
||||
step(ins, &mut r, &sel, mask, ds, &mut idx, &mut val, &mut items_derived);
|
||||
step(ins, &mut r, &sel, mask, log2, era.as_ref(), layout, ds, &mut idx, &mut val, &mut items_derived);
|
||||
if ins.op == Op::Scratch {
|
||||
let m = scratch.as_mut().expect("a scratch op needs a scratch class");
|
||||
let (d, a) = (ins.dst as usize, ins.src as usize);
|
||||
|
|
@ -305,6 +339,9 @@ fn step(
|
|||
r: &mut [[u32; LANES]; 8],
|
||||
sel: &[u32; LANES],
|
||||
mask: u32,
|
||||
log2: u32,
|
||||
era: Option<&EraParams>,
|
||||
layout: Layout,
|
||||
ds: &DatasetSource,
|
||||
idx: &mut [u32; LANES],
|
||||
val: &mut [u32; LANES],
|
||||
|
|
@ -380,9 +417,9 @@ fn step(
|
|||
}
|
||||
Op::Load if ins.width == 1 => {
|
||||
for lane in 0..LANES {
|
||||
idx[lane] = r[a][lane] & mask;
|
||||
idx[lane] = load_index(era, ins, r[a][lane], mask, log2);
|
||||
}
|
||||
*items_derived += ds.fetch(idx, val);
|
||||
*items_derived += ds.fetch(idx, val, layout);
|
||||
for lane in 0..LANES {
|
||||
r[d][lane] ^= val[lane];
|
||||
}
|
||||
|
|
@ -392,10 +429,10 @@ fn step(
|
|||
let width = ins.width as usize;
|
||||
let align = !(ins.width as u32 - 1);
|
||||
for lane in 0..LANES {
|
||||
idx[lane] = (r[a][lane] & mask) & align;
|
||||
idx[lane] = load_index(era, ins, r[a][lane], mask, log2) & align;
|
||||
}
|
||||
let mut vals = [[0u32; 16]; LANES];
|
||||
*items_derived += ds.fetch_wide(idx, width, &mut vals);
|
||||
*items_derived += ds.fetch_wide(idx, width, &mut vals, layout);
|
||||
for lane in 0..LANES {
|
||||
r[d][lane] = fold_words(r[d][lane], &vals[lane][..width]);
|
||||
}
|
||||
|
|
@ -409,7 +446,7 @@ fn step(
|
|||
for lane in 0..LANES {
|
||||
idx[lane] = base + lane as u32;
|
||||
}
|
||||
*items_derived += ds.fetch(idx, val);
|
||||
*items_derived += ds.fetch(idx, val, Layout::LINEAR);
|
||||
for lane in 0..LANES {
|
||||
r[d][lane] ^= val[lane];
|
||||
}
|
||||
|
|
@ -457,6 +494,11 @@ impl Epoch {
|
|||
Self { program: generate_class(seed, class), dataset: DatasetSource::new_shape(day, mode, dataset_log2, shape) }
|
||||
}
|
||||
|
||||
/// `dataset[w]` as this epoch's program reads it: under the program's layout (era layout; linear for v2).
|
||||
pub fn dataset_word(&self, w: u32) -> u32 {
|
||||
self.dataset.word_at(self.program.class.layout(), w)
|
||||
}
|
||||
|
||||
/// The production shape: memory-hard, 1 GiB dataset.
|
||||
pub fn memory_hard(seed: &str, day: &str) -> Self {
|
||||
Self::new(seed, day, DatasetMode::MemoryHard, DEFAULT_DATASET_LOG2)
|
||||
|
|
@ -587,10 +629,10 @@ mod tests {
|
|||
let ds = DatasetSource::new("2026-10-03", DatasetMode::MemoryHard, 20);
|
||||
let mut base = [0u32; LANES];
|
||||
for (k, b) in base.iter_mut().enumerate() {
|
||||
*b = ((k as u32 * 0x9E37_79B1) & ds.mask) & !15;
|
||||
*b = ((k as u32).wrapping_mul(0x9E37_79B1) & ds.mask) & !15;
|
||||
}
|
||||
let mut out = [[0u32; 16]; LANES];
|
||||
let items = ds.fetch_wide(&base, 16, &mut out);
|
||||
let items = ds.fetch_wide(&base, 16, &mut out, Layout::LINEAR);
|
||||
assert!(items >= 1 && items <= LANES);
|
||||
for k in 0..LANES {
|
||||
for j in 0..16 {
|
||||
|
|
@ -601,7 +643,7 @@ mod tests {
|
|||
for b in base4.iter_mut() {
|
||||
*b += 8;
|
||||
}
|
||||
let items4 = ds.fetch_wide(&base4, 4, &mut out);
|
||||
let items4 = ds.fetch_wide(&base4, 4, &mut out, Layout::LINEAR);
|
||||
assert_eq!(items4, items);
|
||||
for k in 0..LANES {
|
||||
for j in 0..4 {
|
||||
|
|
@ -632,6 +674,83 @@ mod tests {
|
|||
assert_eq!(e.hash_warp(0), e.hash_warp(0));
|
||||
}
|
||||
|
||||
/// Era layout: the load address stays inside the site's window and below the mask at every dataset size (the
|
||||
/// window floor of 2^26 words clamps the shrink), the interleaved memory-hard dataset reads item(t(w))[j(w)] and
|
||||
/// is the same prefix at 2^20 and 2^22 words, the wide fetch agrees word for word, and an era epoch hashes
|
||||
/// deterministically through the interpreter and the single-nonce API.
|
||||
#[test]
|
||||
fn era_windows_layout_and_epochs() {
|
||||
let eb = EraParams::test_era_bytes("igneum-era-test/1");
|
||||
let c = LoadClass::era(LoadClass::V2, &eb, &[1]);
|
||||
let e = c.era.unwrap();
|
||||
let mut ins = Instr { op: Op::Load, dst: 0, src: 1, src2: 0, imm: 0, imm2: 0, rot: 1, bit: 0, mask: 1, width: 1, win: 2, off: 3 };
|
||||
let mut s = crate::seed::SplitMix64::new(7);
|
||||
for log2 in [20u32, 26, 27, 28, 29] {
|
||||
let mask = (1u64 << log2) as u32 - 1;
|
||||
let k = ins.win.min(log2.saturating_sub(26) as u8) as u32;
|
||||
for _ in 0..1000 {
|
||||
let x = s.next() as u32;
|
||||
let idx = load_index(Some(&e), &ins, x, mask, log2);
|
||||
assert!(idx <= mask);
|
||||
let (wm, off) = window(&ins, mask, log2);
|
||||
assert_eq!(idx & !wm, off, "log2 {log2}");
|
||||
assert_eq!(wm, mask >> k);
|
||||
assert_eq!(idx, ((x.wrapping_mul(e.stride_mul).rotate_left(e.stride_rot) & wm) | off) & mask);
|
||||
}
|
||||
}
|
||||
ins.win = 0;
|
||||
assert_eq!(load_index(None, &ins, 0xdead_beef, 0x0fff_ffff, 28), 0xdead_beef & 0x0fff_ffff);
|
||||
// the interleaved dataset: one day cache, the layout per program
|
||||
let l = e.layout();
|
||||
assert_eq!(l.pos, [1, 3, 8, 13]);
|
||||
let small = DatasetSource::new("2026-10-03", DatasetMode::MemoryHard, 20);
|
||||
let big = DatasetSource::new("2026-10-03", DatasetMode::MemoryHard, 22);
|
||||
let m = small.memhard().unwrap();
|
||||
for w in [0u32, 1, 4, 5, 255, 256, 4095, 8192, 0x0f_ffff] {
|
||||
let (t, j) = l.split(w);
|
||||
assert_eq!(small.word_at(l, w), crate::memhard::derive_item(t, &m.params, &m.cache)[j as usize], "w {w}");
|
||||
assert_eq!(small.word_at(l, w), big.word_at(l, w), "prefix at w {w}");
|
||||
assert_eq!(small.word(w), small.word_at(Layout::LINEAR, w));
|
||||
}
|
||||
let mut idx = [0u32; LANES];
|
||||
for (k, i) in idx.iter_mut().enumerate() {
|
||||
*i = (k as u32).wrapping_mul(0x9E37_79B1) & small.mask;
|
||||
}
|
||||
let mut out = [0u32; LANES];
|
||||
small.fetch(&idx, &mut out, l);
|
||||
for k in 0..LANES {
|
||||
assert_eq!(out[k], small.word_at(l, idx[k]));
|
||||
}
|
||||
// a 16-byte era: the wide fetch keeps a lane's four words in one item
|
||||
let eb3 = EraParams::test_era_bytes("igneum-era-test/3");
|
||||
let c3 = LoadClass::era(LoadClass::fixed(4, 16), &eb3, &[4]);
|
||||
let l3 = c3.layout();
|
||||
assert_eq!(l3.pos[..2], [0, 1]);
|
||||
let mut base = [0u32; LANES];
|
||||
for (k, b) in base.iter_mut().enumerate() {
|
||||
*b = ((k as u32).wrapping_mul(0x9E37_79B1) & small.mask) & !3;
|
||||
}
|
||||
let mut wide = [[0u32; 16]; LANES];
|
||||
small.fetch_wide(&base, 4, &mut wide, l3);
|
||||
for k in 0..LANES {
|
||||
for j in 0..4 {
|
||||
assert_eq!(wide[k][j], small.word_at(l3, base[k] + j as u32), "lane {k} word {j}");
|
||||
}
|
||||
}
|
||||
// era epochs hash deterministically, differ per era, and the single-nonce API agrees with the warp
|
||||
let mut seen = std::collections::HashSet::new();
|
||||
for (n, c) in [(1u64, c), (3, c3)] {
|
||||
let ep = Epoch::new_class("igneum-genesis", "2026-10-03", DatasetMode::MemoryHard, 20, c);
|
||||
assert_eq!(ep.dataset_word(5), ep.dataset.word_at(c.layout(), 5));
|
||||
let a = ep.hash_warp(64);
|
||||
assert_eq!(a, ep.hash_warp(64));
|
||||
assert_eq!(ep.hash(64 + 5), a[5]);
|
||||
assert!(seen.insert(a[0]), "era {n}");
|
||||
let closed = Epoch::new_class("igneum-genesis", "2026-10-03", DatasetMode::ClosedForm, 20, c);
|
||||
assert_ne!(closed.hash_warp(64), a);
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn wide_class_epochs_hash() {
|
||||
for name in ["w16", "w64x4", "50,35,15"] {
|
||||
|
|
|
|||
|
|
@ -15,7 +15,7 @@ use igneum_pow::emit::{
|
|||
cuda_kernel, cuda_kernel_bound, cuda_memhard_header, export_pack, metal_memhard, metal_program,
|
||||
metal_program_bound, opencl_kernel, opencl_kernel_bound, program_header, program_json, LoadSource,
|
||||
};
|
||||
use igneum_pow::generator::{generate_from_seed_bytes, generate_from_seed_bytes_program_class, Op, ProgramClass, GENERATOR_VERSION, GENERATOR_VERSION_V3, LOAD_SLOTS, V3_CLASS};
|
||||
use igneum_pow::generator::{generate_from_seed_bytes, generate_from_seed_bytes_program_class, LoadClass, Op, ProgramClass, GENERATOR_VERSION, GENERATOR_VERSION_V3, LOAD_SLOTS, V3_CLASS};
|
||||
use igneum_pow::memhard::{Shape, CACHE_WORDS};
|
||||
use igneum_pow::verify::{DatasetMode, DatasetSource, Epoch};
|
||||
use serde_json::Value;
|
||||
|
|
@ -191,19 +191,21 @@ fn cache_matches_vectors() {
|
|||
|
||||
fn check_dataset_words(pack: &str) {
|
||||
let v = json(pack, "vectors.json");
|
||||
let ds = &epoch(pack).dataset;
|
||||
let e = epoch(pack);
|
||||
let ds = &e.dataset;
|
||||
// a pack's self-test words are read under its program's layout (the era layout; linear for every v2 pack)
|
||||
let head: Vec<u32> = v["dataset_head"].as_array().unwrap().iter().map(hex32).collect();
|
||||
for (i, h) in head.iter().enumerate() {
|
||||
assert_eq!(ds.word(i as u32), *h, "{pack}: dataset[{i}]");
|
||||
assert_eq!(e.dataset_word(i as u32), *h, "{pack}: dataset[{i}]");
|
||||
}
|
||||
let last_index = v["dataset_last_index"].as_u64().unwrap() as u32;
|
||||
assert_eq!(last_index, ds.mask);
|
||||
assert_eq!(ds.word(last_index), hex32(&v["dataset_last"]), "{pack}: dataset[MASK]");
|
||||
assert_eq!(e.dataset_word(last_index), hex32(&v["dataset_last"]), "{pack}: dataset[MASK]");
|
||||
let samples = v["dataset_samples"].as_array().unwrap();
|
||||
assert_eq!(samples.len(), 64);
|
||||
for s in samples {
|
||||
let idx = s["index"].as_u64().unwrap() as u32;
|
||||
assert_eq!(ds.word(idx), hex32(&s["value"]), "{pack}: dataset[{idx}]");
|
||||
assert_eq!(e.dataset_word(idx), hex32(&s["value"]), "{pack}: dataset[{idx}]");
|
||||
}
|
||||
}
|
||||
|
||||
|
|
@ -286,12 +288,15 @@ fn check_sources(pack: &str) {
|
|||
assert_same_text(pack, "program.h", &program_header(p, &day, &e.dataset));
|
||||
if let Some(mp) = mp {
|
||||
assert_same_text(pack, "memhard.h", &cuda_memhard_header(p, mp));
|
||||
assert_same_text(pack, "memhard.metal", &metal_memhard(mp));
|
||||
assert_same_text(pack, "memhard.metal", &igneum_pow::emit::metal_memhard_layout(mp, p.class.layout()));
|
||||
}
|
||||
let got = program_json(p, &day, &e.dataset);
|
||||
assert_same_text(pack, "program.json", &got);
|
||||
let _: Value = serde_json::from_str(&got).expect("program.json is valid JSON");
|
||||
// Every load in every emitted hash kernel has the masked form, and there are exactly 16 of them.
|
||||
// Every load in every emitted hash kernel has the masked form, and there are exactly 16 of them (a class with
|
||||
// the era layout inside has the era form instead: `((rotl_imm(rN * M, R) & WM) | OFF) & mask`, checked by
|
||||
// era_emitted_sources_match_and_loads_have_the_era_form over the era packs, and here by the same count).
|
||||
let era_load = if p.class.era.is_some() { "((rotl_imm(r" } else { "" };
|
||||
for (file, load, masked) in [
|
||||
("kernel.cu", "ds[r", " & mask]"),
|
||||
("kernel_bound.cu", "ds[r", " & mask]"),
|
||||
|
|
@ -299,6 +304,7 @@ fn check_sources(pack: &str) {
|
|||
("program_bound.metal", "dataset[r", " & MASK]"),
|
||||
] {
|
||||
let text = read(pack, file);
|
||||
let load = if era_load.is_empty() { load } else { era_load };
|
||||
assert_eq!(text.matches(load).count(), LOAD_SLOTS, "{pack}/{file}: 16 loads");
|
||||
assert_eq!(text.matches(masked).count(), LOAD_SLOTS, "{pack}/{file}: 16 masked loads");
|
||||
}
|
||||
|
|
@ -326,14 +332,23 @@ fn v3_packs_are_the_v2_seeds_under_mixer_x4() {
|
|||
let e2 = epoch(v2);
|
||||
let j = json(v3, "program.json");
|
||||
assert_eq!(j["program_class"].as_str().unwrap(), "v3");
|
||||
assert_eq!(j["load_class"].as_str().unwrap(), "mx4");
|
||||
assert!(j["load_class"].as_str().unwrap().starts_with("mx4"), "{v3}: mx4, or mx4 with the era inside");
|
||||
assert_eq!(j["mixer_mult"].as_u64().unwrap(), 4);
|
||||
assert_eq!(j["cache_growth"].as_bool().unwrap(), true);
|
||||
assert_eq!(j["dataset"]["mixer_mult"].as_u64().unwrap(), 4);
|
||||
assert_eq!(j["dataset"]["cache"]["log2_words"].as_u64().unwrap(), 26);
|
||||
assert_eq!(e3.program.generator, GENERATOR_VERSION_V3);
|
||||
assert_eq!(e3.program.class, V3_CLASS);
|
||||
assert_eq!(e3.program.instrs, e2.program.instrs, "{v3}: the v2 program under the v3 construction");
|
||||
if e3.program.era_bytes.is_some() {
|
||||
// the era layout composed into class v3 (docs/plans/era-layout.md, 5 October 2026): a chain pack carries an
|
||||
// era, so its class is V3_CLASS with the era drawn inside and its stream takes two window draws per
|
||||
// instruction; the v2 program carries over only in the seed, the attempt and the day
|
||||
assert_eq!(LoadClass { era: None, ..e3.program.class }, V3_CLASS, "{v3}: the composed class");
|
||||
assert!(e3.program.class.era.is_some());
|
||||
assert_ne!(e3.program.instrs, e2.program.instrs, "{v3}: the era windows change the stream");
|
||||
} else {
|
||||
assert_eq!(e3.program.class, V3_CLASS);
|
||||
assert_eq!(e3.program.instrs, e2.program.instrs, "{v3}: the v2 program under the v3 construction");
|
||||
}
|
||||
assert_eq!(e3.program.seed, e2.program.seed);
|
||||
assert_eq!(e3.program.attempt, e2.program.attempt);
|
||||
assert_ne!(e3.program.program_id(), e2.program.program_id());
|
||||
|
|
@ -352,7 +367,7 @@ fn v3_packs_are_the_v2_seeds_under_mixer_x4() {
|
|||
assert!(h.contains("#define IGNEUM_MIXER_MULT 4"));
|
||||
assert!(h.contains("#define IGNEUM_CACHE_GROWTH 1"));
|
||||
assert!(h.contains("#define IGNEUM_CACHE_LOG2_WORDS 26\n"));
|
||||
assert!(h.contains("#define IGNEUM_LOAD_CLASS \"mx4\"\n"));
|
||||
assert!(h.contains("#define IGNEUM_LOAD_CLASS \"mx4"), "mx4, or mx4 with the era inside");
|
||||
for file in ["memhard.h", "memhard.metal", "kernel.cl"] {
|
||||
let text = read(v3, file);
|
||||
assert_eq!(text.matches("j < 4u; ++j) mh_mixer(s, 0x9E3779B9u * (r * 4u + j + 1u))").count(), 1, "{v3}/{file}");
|
||||
|
|
@ -433,3 +448,230 @@ fn devnet_pack_is_the_chain_derivation() {
|
|||
assert_eq!(e.program.instrs, epoch("igneum-devnet-v4-epoch0").program.instrs);
|
||||
assert_eq!(e.hash_warp(0), epoch("igneum-devnet-v4-epoch0").hash_warp(0));
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------------------------------------
|
||||
// Era layout packs (5 October 2026, docs/plans/era-layout.md): proto-cuda/packs-ca2-era/era-<n>, n in 0..5, the
|
||||
// devnet epoch seed and day bytes under the era class of test era seed igneum-era-test/<n>, the width pinned at
|
||||
// 4 bytes (allowed_widths in program.json). Checked like the pinned packs, plus the one load form of 1.3 by text search.
|
||||
// ---------------------------------------------------------------------------------------------------------
|
||||
|
||||
use igneum_pow::generator::{generate_era, EraParams, V3_ALLOWED};
|
||||
|
||||
const ERA_PACKS: [&str; 6] = ["era-0", "era-1", "era-2", "era-3", "era-4", "era-5"];
|
||||
|
||||
fn era_packs_dir() -> PathBuf {
|
||||
PathBuf::from(env!("CARGO_MANIFEST_DIR")).join("../proto-cuda/packs-ca2-era")
|
||||
}
|
||||
|
||||
fn era_read(pack: &str, file: &str) -> String {
|
||||
let p = era_packs_dir().join(pack).join(file);
|
||||
std::fs::read_to_string(&p).unwrap_or_else(|e| panic!("read {}: {e}", p.display()))
|
||||
}
|
||||
|
||||
fn era_json(pack: &str, file: &str) -> Value {
|
||||
serde_json::from_str(&era_read(pack, file)).unwrap_or_else(|e| panic!("{pack}/{file}: {e}"))
|
||||
}
|
||||
|
||||
/// The era class of a pack: the era bytes from program.json (`era_seed_bytes`, the chain's `E_n`; the test seed
|
||||
/// `igneum-era-test/<n>` of the pack's number gives the same bytes), the allowed set from `era.allowed_widths`
|
||||
/// (class v3's `V3_ALLOWED`); the pack's recorded stream words and draw must be the class's.
|
||||
fn era_class(pack: &str) -> LoadClass {
|
||||
let j = era_json(pack, "program.json");
|
||||
let n: u64 = pack.trim_start_matches("era-").parse().unwrap();
|
||||
let eb = unhex(&j["era_seed_bytes"]);
|
||||
assert_eq!(eb, EraParams::test_era_bytes(&format!("igneum-era-test/{n}")).to_vec(), "{pack}: the era bytes of test seed {n}");
|
||||
let allowed: Vec<u8> = j["era"]["allowed_widths"].as_array().unwrap().iter().map(|v| v.as_u64().unwrap() as u8).collect();
|
||||
assert_eq!(allowed, V3_ALLOWED.to_vec(), "{pack}: class v3's width set");
|
||||
// the measurement packs of 5 October 2026: the era layout over version 2's construction (mixer x1, the genesis
|
||||
// cache), generator 3 and the era bytes recorded; the chain's class v3 composes the same draw over LoadClass::MX4
|
||||
// (generator tests era_programs_are_accepted and program_classes), and the integration re-exports these packs
|
||||
let c = LoadClass::era(LoadClass::V2, &eb, &allowed);
|
||||
let e = c.era.unwrap();
|
||||
let words: Vec<u32> = j["era"]["seed_words"].as_array().unwrap().iter().map(hex32).collect();
|
||||
assert_eq!(e.words.to_vec(), words, "{pack}: era seed words");
|
||||
assert_eq!(e.width_words as u64, j["era"]["width_words"].as_u64().unwrap(), "{pack}: width");
|
||||
assert_eq!(e.stride_mul, hex32(&j["era"]["stride_mul"]), "{pack}: stride mul");
|
||||
assert_eq!(e.stride_rot as u64, j["era"]["stride_rot"].as_u64().unwrap(), "{pack}: stride rot");
|
||||
let pos: Vec<u8> = j["era"]["interleave"].as_array().unwrap().iter().map(|v| v.as_u64().unwrap() as u8).collect();
|
||||
assert_eq!(e.pos.to_vec(), pos, "{pack}: interleave");
|
||||
c
|
||||
}
|
||||
|
||||
fn era_epoch(pack: &str) -> &'static Epoch {
|
||||
static E: OnceLock<Vec<(String, Epoch)>> = OnceLock::new();
|
||||
let all = E.get_or_init(|| {
|
||||
ERA_PACKS
|
||||
.iter()
|
||||
.map(|p| {
|
||||
let j = era_json(p, "program.json");
|
||||
let seed = j["seed"].as_str().unwrap();
|
||||
let seed_bytes = unhex(&j["seed_bytes"]);
|
||||
let day_bytes = unhex(&j["dataset"]["day_bytes"]);
|
||||
assert_eq!(j["dataset_mode"].as_str().unwrap(), "memory-hard");
|
||||
let log2 = j["dataset"]["log2_words"].as_u64().unwrap() as u32;
|
||||
let class = era_class(p);
|
||||
assert_eq!(log2, igneum_pow::verify::DEFAULT_DATASET_LOG2);
|
||||
let eb = unhex(&j["era_seed_bytes"]);
|
||||
let program = generate_era(seed, &seed_bytes, LoadClass::V2, &eb, &V3_ALLOWED);
|
||||
assert_eq!(program.class, class, "{p}: the pack's era class");
|
||||
assert_eq!(program.generator, GENERATOR_VERSION_V3);
|
||||
assert_eq!(program.era_bytes.as_deref(), Some(&eb[..]));
|
||||
let mut dataset = DatasetSource::from_key(igneum_pow::seed::seed_words_from_bytes(&day_bytes), DatasetMode::MemoryHard, log2);
|
||||
dataset.key_bytes = day_bytes;
|
||||
(p.to_string(), Epoch { program, dataset })
|
||||
})
|
||||
.collect()
|
||||
});
|
||||
&all.iter().find(|(n, _)| n == pack).unwrap().1
|
||||
}
|
||||
|
||||
/// program.json of an era pack: generator, attempt, id, class, every instruction with width, win and off, and the
|
||||
/// program passes the acceptance rule.
|
||||
#[test]
|
||||
fn era_program_json_matches() {
|
||||
for pack in ERA_PACKS {
|
||||
let j = era_json(pack, "program.json");
|
||||
let p = &era_epoch(pack).program;
|
||||
assert_eq!(j["generator"].as_u64().unwrap() as u32, GENERATOR_VERSION_V3, "{pack}: a class v3 pack");
|
||||
assert_eq!(j["program_class"].as_str().unwrap(), "v3");
|
||||
assert_eq!(j["attempt"].as_u64().unwrap() as u32, p.attempt, "{pack}: attempt");
|
||||
assert_eq!(hex64(&j["program_id"]), p.program_id(), "{pack}: program id");
|
||||
assert_eq!(j["load_class"].as_str().unwrap(), p.class.name(), "{pack}: class");
|
||||
assert_eq!(j["bytes_per_hash"].as_u64().unwrap() as usize, p.bytes_per_hash());
|
||||
assert_eq!(p.loads_per_hash(), 8 * LOAD_SLOTS);
|
||||
assert!(accept::check(p).is_ok(), "{pack}: acceptance");
|
||||
let instrs = j["instructions"].as_array().unwrap();
|
||||
assert_eq!(instrs.len(), p.instrs.len());
|
||||
for (k, (ins, ji)) in p.instrs.iter().zip(instrs).enumerate() {
|
||||
assert_eq!(Op::from_name(ji["op"].as_str().unwrap()).unwrap(), ins.op, "{pack} #{k} op");
|
||||
assert_eq!(ji["dst"].as_u64().unwrap(), ins.dst as u64);
|
||||
assert_eq!(ji["src"].as_u64().unwrap(), ins.src as u64);
|
||||
assert_eq!(hex32(&ji["imm"]), ins.imm);
|
||||
assert_eq!(ji["width"].as_u64().unwrap(), ins.width as u64, "{pack} #{k} width");
|
||||
assert_eq!(ji["win"].as_u64().unwrap(), ins.win as u64, "{pack} #{k} win");
|
||||
assert_eq!(ji["off"].as_u64().unwrap(), ins.off as u64, "{pack} #{k} off");
|
||||
if ins.op == Op::Load {
|
||||
assert_eq!(ins.width, p.class.era.unwrap().width_words);
|
||||
assert!(ins.win <= 2 && (ins.off as u32) < (1u32 << ins.win));
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// The six era packs are the same program seed under six draws: the instruction lists agree, the widths and layouts
|
||||
/// follow the draw, and the dataset words differ from the linear layout exactly when the interleave is not linear.
|
||||
#[test]
|
||||
fn era_packs_share_the_program_and_differ_in_layout() {
|
||||
let linear = &epoch("igneum-devnet-v4-epoch0").dataset;
|
||||
for pack in ERA_PACKS {
|
||||
let e = era_epoch(pack);
|
||||
assert_eq!(e.program.seed_bytes, epoch("igneum-devnet-v4-epoch0").program.seed_bytes, "{pack}: the devnet seed");
|
||||
let strip = |p: &igneum_pow::generator::Program| {
|
||||
p.instrs.iter().map(|i| (i.op, i.dst, i.src, i.src2, i.imm, i.imm2, i.rot, i.bit, i.mask, i.win, i.off)).collect::<Vec<_>>()
|
||||
};
|
||||
assert_eq!(strip(&e.program), strip(&era_epoch("era-0").program), "{pack}: same stream as era-0");
|
||||
let l = e.program.class.layout();
|
||||
let same_at_1 = (0..64u32).all(|w| e.dataset_word(w * 977 + 1) == linear.word(w * 977 + 1));
|
||||
assert_eq!(same_at_1, l.is_linear(), "{pack}: layout {:?}", l.pos);
|
||||
assert_eq!(e.dataset_word(0), linear.word(0), "{pack}: word 0 is item 0 word 0 in every layout");
|
||||
// the chain's shared day cache serves every era: the day's dataset source is the pinned pack's, bit for bit
|
||||
assert_eq!(e.dataset.key, linear.key);
|
||||
assert_eq!(e.dataset.memhard().unwrap().cache.fnv1a64(), linear.memhard().unwrap().cache.fnv1a64());
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn era_dataset_words_and_vectors_match() {
|
||||
for pack in ERA_PACKS {
|
||||
let v = era_json(pack, "vectors.json");
|
||||
let e = era_epoch(pack);
|
||||
let ds = &e.dataset;
|
||||
let head: Vec<u32> = v["dataset_head"].as_array().unwrap().iter().map(hex32).collect();
|
||||
for (i, h) in head.iter().enumerate() {
|
||||
assert_eq!(e.dataset_word(i as u32), *h, "{pack}: dataset[{i}]");
|
||||
}
|
||||
assert_eq!(e.dataset_word(ds.mask), hex32(&v["dataset_last"]), "{pack}: dataset[MASK]");
|
||||
for s in v["dataset_samples"].as_array().unwrap() {
|
||||
let idx = s["index"].as_u64().unwrap() as u32;
|
||||
assert_eq!(e.dataset_word(idx), hex32(&s["value"]), "{pack}: dataset[{idx}]");
|
||||
}
|
||||
assert_eq!(ds.memhard().unwrap().cache.fnv1a64(), hex64(&v["cache_fnv1a64"]));
|
||||
let mut n = 0;
|
||||
for w in v["warps"].as_array().unwrap() {
|
||||
let base = w["base_nonce"].as_u64().unwrap() as u32;
|
||||
let expected: Vec<u64> = w["expected"].as_array().unwrap().iter().map(hex64).collect();
|
||||
let got = e.hash_warp(base);
|
||||
for lane in 0..32 {
|
||||
assert_eq!(got[lane], expected[lane], "{pack}: base {base} lane {lane}");
|
||||
n += 1;
|
||||
}
|
||||
assert_eq!(e.hash(base + 7), expected[7]);
|
||||
}
|
||||
assert_eq!(n, 96, "{pack}");
|
||||
}
|
||||
}
|
||||
|
||||
/// Every emitted file of every era pack matches the emitters byte for byte, the export reproduces vectors.json and
|
||||
/// vectors.h, and every dataset load in every hash kernel has the one era form (no plain `ds[rN & mask]` remains).
|
||||
#[test]
|
||||
fn era_emitted_sources_match_and_loads_have_the_era_form() {
|
||||
for pack in ERA_PACKS {
|
||||
let e = era_epoch(pack);
|
||||
let day = era_json(pack, "program.json")["dataset"]["day"].as_str().unwrap().to_string();
|
||||
let source = era_json(pack, "vectors.json")["source"].as_str().unwrap().to_string();
|
||||
let out = export_pack(e, &day, &source);
|
||||
for (name, text) in &out.files {
|
||||
let want = era_read(pack, name);
|
||||
assert!(text == &want, "{pack}/{name} differs from the emitter");
|
||||
}
|
||||
let mut on_disk: Vec<String> = std::fs::read_dir(era_packs_dir().join(pack))
|
||||
.unwrap()
|
||||
.map(|d| d.unwrap().file_name().to_string_lossy().to_string())
|
||||
.filter(|n| !n.starts_with('.') && n != "seeds.txt")
|
||||
.collect();
|
||||
on_disk.sort();
|
||||
let mut want: Vec<String> = out.files.iter().map(|(n, _)| n.clone()).collect();
|
||||
want.sort();
|
||||
assert_eq!(on_disk, want, "{pack}: the pack holds the export's files and seeds.txt only");
|
||||
let era = e.program.class.era.unwrap();
|
||||
let mul = format!("0x{:08x}u", era.stride_mul);
|
||||
for (file, mask) in [
|
||||
("kernel.cu", "mask"),
|
||||
("kernel_bound.cu", "mask"),
|
||||
("kernel.cl", "mask"),
|
||||
("kernel_bound.cl", "mask"),
|
||||
("program.metal", "MASK"),
|
||||
("program_bound.metal", "MASK"),
|
||||
] {
|
||||
let text = era_read(pack, file);
|
||||
let kernels = if file.starts_with("kernel_bound") || file == "kernel.cu" || file == "kernel.cl" || file.starts_with("program") { 1 } else { 1 };
|
||||
// kernel_bound.cl carries igneum_hash and igneum_hash_bound: two kernels
|
||||
let kernels = if file == "kernel_bound.cl" { 2 } else { kernels };
|
||||
let era_form: usize = text
|
||||
.lines()
|
||||
.filter(|l| l.contains("rotl_imm(r") && l.contains(&format!(" * {mul}, {}u) & ", era.stride_rot)) && l.contains(&format!(") & {mask}")))
|
||||
.filter(|l| l.contains("ds[") || l.contains("dataset[") || l.contains("b_ = "))
|
||||
.count();
|
||||
assert_eq!(era_form, LOAD_SLOTS * kernels, "{pack}/{file}: {} loads of the era form", LOAD_SLOTS * kernels);
|
||||
let plain = text.lines().filter(|l| l.contains("ds[r") || l.contains("dataset[r")).count();
|
||||
assert_eq!(plain, 0, "{pack}/{file}: a load without the era form");
|
||||
}
|
||||
// the layout helpers appear exactly when the layout is not linear
|
||||
let mh = era_read(pack, "memhard.h");
|
||||
assert_eq!(mh.contains("mh_addr("), !era.layout().is_linear(), "{pack}: memhard.h layout helpers");
|
||||
}
|
||||
}
|
||||
|
||||
/// An era pack's dataset is a prefix at every size of at least 2^16 words: the 2^20-word source gives the pack's
|
||||
/// words below 2^20.
|
||||
#[test]
|
||||
fn era_dataset_is_a_prefix_at_smaller_sizes() {
|
||||
for pack in ["era-1", "era-3"] {
|
||||
let e = era_epoch(pack);
|
||||
let small = DatasetSource::from_key(e.dataset.key, DatasetMode::MemoryHard, 20);
|
||||
let l = e.program.class.layout();
|
||||
for w in [0u32, 1, 2, 3, 16, 255, 4096, 65_535, 65_536, 0x000f_ffff] {
|
||||
assert_eq!(small.word_at(l, w), e.dataset_word(w), "{pack}: w {w}");
|
||||
}
|
||||
}
|
||||
}
|
||||
|
|
|
|||
28
proto-cuda/emu/test-layout.sh
Executable file
28
proto-cuda/emu/test-layout.sh
Executable file
|
|
@ -0,0 +1,28 @@
|
|||
#!/usr/bin/env bash
|
||||
# The host-side dataset derivation of the harnesses under a non-linear item-to-word layout (era layout,
|
||||
# docs/plans/era-layout.md 1.2): proto-cuda/host.cu (and proto-opencl/host.c, the same function) derive dataset words
|
||||
# on the host through the pack's own mh_word, which carries the layout. Until 5 October 2026 both derived
|
||||
# mh_item(w >> 4)[w & 15] and the "64 random points vs host derivation" check failed on every interleaved pack while
|
||||
# the Mac's samples and the vectors passed (the known-failed case, docs/bench-log.md, era layout entry). This script
|
||||
# runs the CUDA CPU emulation on an interleaved era pack and a linear one and requires OVERALL: PASS on both; it is
|
||||
# the test that fails on the old derivation. Usage: emu/test-layout.sh [pack dir ...]
|
||||
set -euo pipefail
|
||||
HERE="$(cd "$(dirname "$0")" && pwd)"
|
||||
CUDA_DIR="$(cd "$HERE/.." && pwd)"
|
||||
PACKS=("$@")
|
||||
[ ${#PACKS[@]} -gt 0 ] || PACKS=("$CUDA_DIR/packs-ca2-era/era-1" "$CUDA_DIR/packs/igneum-devnet-v4-epoch0")
|
||||
CXX="${CXX:-c++}"
|
||||
for P in "${PACKS[@]}"; do
|
||||
grep -q "mh_addr(" "$P/memhard.h" && kind=interleaved || kind=linear
|
||||
BUILD="$HERE/build-layout-$(basename "$P")"
|
||||
mkdir -p "$BUILD"
|
||||
cp "$CUDA_DIR/host.cu" "$BUILD/host_emu.cpp"
|
||||
sed -E 's/([A-Za-z_0-9]+)<<<([^,]+), ([^>]+)>>>\(/emu_launch(\1, \2, \3, /' "$P/kernel.cu" > "$BUILD/kernel_emu.cpp"
|
||||
sed -E 's/([A-Za-z_0-9]+)<<<([^,]+), ([^>]+)>>>\(/emu_launch(\1, \2, \3, /' "$P/kernel_bound.cu" > "$BUILD/kernel_bound_emu.cpp"
|
||||
"$CXX" -std=c++17 -O2 -w -I "$HERE" -I "$P" -o "$BUILD/igneum-emu" "$BUILD/host_emu.cpp" "$BUILD/kernel_emu.cpp" -DIGNEUM_BOUND "$BUILD/kernel_bound_emu.cpp" "$HERE/shim.cpp" -pthread
|
||||
out="$("$BUILD/igneum-emu" --batch-log2 13 --batches 1 2>&1)"
|
||||
echo "$out" | grep -E "dataset self-test|OVERALL" | sed "s|^|$(basename "$P") ($kind): |"
|
||||
echo "$out" | grep -q "^OVERALL: PASS" || { echo "FAIL: $P ($kind layout) did not pass the emulated harness"; exit 1; }
|
||||
echo "$out" | grep -q "64 random points vs host derivation PASS" || { echo "FAIL: $P ($kind layout): the host derivation disagrees with the pack's layout"; exit 1; }
|
||||
done
|
||||
echo "PASS: the host derivation follows the pack's layout on ${#PACKS[@]} pack(s)"
|
||||
|
|
@ -115,11 +115,11 @@ static bool setupCache() {
|
|||
return gCachePass;
|
||||
}
|
||||
|
||||
// dataset[w] derived on the host from the host cache, exactly as proto-metal's verifier does it.
|
||||
// dataset[w] derived on the host from the host cache through the pack's own mh_word (memhard.h), which carries the
|
||||
// pack's item-to-word layout (era layout, 5 October 2026: the harness's former w >> 4 / w & 15 failed the random
|
||||
// points of every interleaved pack while the Mac samples and vectors passed).
|
||||
static uint32_t host_ds_word(uint32_t w) {
|
||||
uint32_t s[16];
|
||||
mh_item(hCache.data(), w >> 4u, s);
|
||||
return s[w & 15u];
|
||||
return mh_word(hCache.data(), w);
|
||||
}
|
||||
#endif
|
||||
|
||||
|
|
|
|||
281
proto-cuda/packs-ca2-era/era-0/kernel.cl
Normal file
281
proto-cuda/packs-ca2-era/era-0/kernel.cl
Normal file
|
|
@ -0,0 +1,281 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal).
|
||||
// Built from source at runtime by proto-opencl/host.c, which passes these defines:
|
||||
// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit)
|
||||
// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default)
|
||||
// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32
|
||||
// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition
|
||||
// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units;
|
||||
// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32.
|
||||
#ifndef IGNEUM_GROUP
|
||||
#define IGNEUM_GROUP 32
|
||||
#endif
|
||||
#ifndef IGNEUM_EXCHANGE
|
||||
#define IGNEUM_EXCHANGE 0
|
||||
#endif
|
||||
#ifdef __OPENCL_VERSION__
|
||||
#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1)))
|
||||
#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n]
|
||||
#if IGNEUM_EXCHANGE == 1
|
||||
#ifdef cl_khr_subgroups
|
||||
#pragma OPENCL EXTENSION cl_khr_subgroups : enable
|
||||
#endif
|
||||
#ifdef cl_khr_subgroup_shuffle
|
||||
#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable
|
||||
#endif
|
||||
#elif IGNEUM_EXCHANGE == 2
|
||||
#pragma OPENCL EXTENSION cl_intel_subgroups : enable
|
||||
#endif
|
||||
#else
|
||||
// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros.
|
||||
#include "emu_opencl.h"
|
||||
#endif
|
||||
|
||||
#if IGNEUM_EXCHANGE == 1
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m))
|
||||
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
|
||||
#elif IGNEUM_EXCHANGE == 2
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m))
|
||||
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
|
||||
#else
|
||||
// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per
|
||||
// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane
|
||||
// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's
|
||||
// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier.
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; }
|
||||
#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; }
|
||||
#endif
|
||||
|
||||
static inline uint splitmix32(uint x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32.
|
||||
static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); }
|
||||
// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x.
|
||||
static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); }
|
||||
static inline uint ds_elem(uint i, uint d0, uint d1) {
|
||||
uint x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
|
||||
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
|
||||
#define MH_CACHE_LINE_MASK 0x003fffffu
|
||||
#define MH_SEGMENT_LINES 64u
|
||||
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
|
||||
static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
|
||||
|
||||
// y = ChaCha12 core(x) + x
|
||||
static inline void mh_chacha_block(const uint* x, uint* y) {
|
||||
for (uint i = 0u; i < 16u; ++i) y[i] = x[i];
|
||||
for (uint r = 0u; r < 6u; ++r) {
|
||||
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
|
||||
}
|
||||
for (uint i = 0u; i < 16u; ++i) y[i] += x[i];
|
||||
}
|
||||
|
||||
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
|
||||
static inline void mh_cache_segment(__global uint* cache, uint seg) {
|
||||
uint prev[16]; uint x[16]; uint y[16];
|
||||
for (uint i = 0u; i < 16u; ++i) prev[i] = 0u;
|
||||
for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) {
|
||||
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
|
||||
x[4] = 0xceed56d7u ^ prev[4];
|
||||
x[5] = 0x9ba270d2u ^ prev[5];
|
||||
x[6] = 0x82caab2du ^ prev[6];
|
||||
x[7] = 0x81ebce0eu ^ prev[7];
|
||||
x[8] = 0x12b6ecf1u ^ prev[8];
|
||||
x[9] = 0xd0f3fd7cu ^ prev[9];
|
||||
x[10] = 0xd872eefeu ^ prev[10];
|
||||
x[11] = 0xc158c7bdu ^ prev[11];
|
||||
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
|
||||
mh_chacha_block(x, y);
|
||||
__global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
|
||||
}
|
||||
}
|
||||
|
||||
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
|
||||
static inline void mh_mixer(uint* s, uint rk) {
|
||||
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
|
||||
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
|
||||
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
|
||||
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
|
||||
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
|
||||
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
|
||||
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
|
||||
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
|
||||
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
|
||||
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
|
||||
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
|
||||
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
|
||||
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
|
||||
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
|
||||
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
|
||||
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
|
||||
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
|
||||
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
|
||||
}
|
||||
|
||||
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
|
||||
static inline void mh_item(__global const uint* cache, uint t, uint* s) {
|
||||
s[0] = 0xceed56d7u;
|
||||
s[1] = 0x9ba270d2u;
|
||||
s[2] = 0x82caab2du;
|
||||
s[3] = 0x81ebce0eu;
|
||||
s[4] = 0x12b6ecf1u;
|
||||
s[5] = 0xd0f3fd7cu;
|
||||
s[6] = 0xd872eefeu;
|
||||
s[7] = 0xc158c7bdu;
|
||||
s[8] = t * 0xf351d601u + 0xc6892460u;
|
||||
s[9] = t * 0xa3bb398fu + 0x25b7228au;
|
||||
s[10] = t * 0xb5a09e35u + 0xcd515004u;
|
||||
s[11] = t * 0x7509c9c1u + 0x2846527au;
|
||||
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
|
||||
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
|
||||
s[14] = t * 0xded91851u + 0x82961bacu;
|
||||
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
|
||||
for (uint r = 0u; r < 8u; ++r) {
|
||||
mh_mixer(s, 0x9E3779B9u * (r + 1u));
|
||||
__global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i];
|
||||
}
|
||||
mh_mixer(s, 0x9E3779B9u * 9u);
|
||||
}
|
||||
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 0 2 10 15 of w.
|
||||
static inline uint mh_j(uint w) { return ((w >> 0u) & 1u) | (((w >> 2u) & 1u) << 1) | (((w >> 10u) & 1u) << 2) | (((w >> 15u) & 1u) << 3); }
|
||||
static inline uint mh_t(uint w) { w = (w & 0x00007fffu) | ((w >> 16u) << 15u); w = (w & 0x000003ffu) | ((w >> 11u) << 10u); w = (w & 0x00000003u) | ((w >> 3u) << 2u); w = (w & 0x00000000u) | ((w >> 1u) << 0u); return w; }
|
||||
static inline uint mh_addr(uint t, uint j) { uint w = t; w = ((w >> 0u) << 1u) | (w & 0x00000000u) | (((j >> 0u) & 1u) << 0u); w = ((w >> 2u) << 3u) | (w & 0x00000003u) | (((j >> 1u) & 1u) << 2u); w = ((w >> 10u) << 11u) | (w & 0x000003ffu) | (((j >> 2u) & 1u) << 10u); w = ((w >> 15u) << 16u) | (w & 0x00007fffu) | (((j >> 3u) & 1u) << 15u); return w; }
|
||||
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
|
||||
static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }
|
||||
|
||||
// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item.
|
||||
// The same constants as memhard.h in this pack (one emitter, three dialects).
|
||||
__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) {
|
||||
uint seg = (uint)get_global_id(0);
|
||||
if (seg < nSegments) mh_cache_segment(cache, seg);
|
||||
}
|
||||
__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) {
|
||||
uint t = (uint)get_global_id(0);
|
||||
if (t < nItems) {
|
||||
uint s[16];
|
||||
mh_item(cache, t, s);
|
||||
for (uint i = 0u; i < 16u; ++i) ds[(ulong)mh_addr(t, i)] = s[i];
|
||||
}
|
||||
}
|
||||
|
||||
// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the
|
||||
// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and
|
||||
// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all).
|
||||
IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) {
|
||||
uint gid = (uint)get_global_id(0);
|
||||
uint lid = (uint)get_local_id(0);
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
#if IGNEUM_EXCHANGE == 0
|
||||
IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
|
||||
uint xk = 0u;
|
||||
#else
|
||||
(void)lid;
|
||||
#endif
|
||||
{ uint x = nonce ^ 0x667d0fbdu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x7b8e5963u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
|
||||
{ uint x = nonce ^ 0x7b8e5963u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x31c67e5eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
|
||||
{ uint x = nonce ^ 0x31c67e5eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4529ddc6u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
|
||||
{ uint x = nonce ^ 0x4529ddc6u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef19d6d8u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
|
||||
{ uint x = nonce ^ 0xef19d6d8u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xaccf6211u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
|
||||
{ uint x = nonce ^ 0xaccf6211u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xda0aed32u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
|
||||
{ uint x = nonce ^ 0xda0aed32u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xabc6df31u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
|
||||
{ uint x = nonce ^ 0xabc6df31u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x667d0fbdu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
|
||||
r7 = r7 ^ r0; // 1 xor
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 1u); r3 = r3 ^ t_; } // 2 shfl
|
||||
r4 = r4 - r1; // 3 sub
|
||||
r2 = r0 * r4 + r2; // 4 mad
|
||||
r4 = rotl_imm(r4, 9u); // 5 rotl
|
||||
r0 = rotr_var(r0, r2); // 6 rotr
|
||||
r6 = r6 ^ ds[((rotl_imm(r7 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r2 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r0 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r1 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
|
||||
r5 = r5 * r4; // 12 mul
|
||||
r7 = r7 ^ ds[((rotl_imm(r6 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 16u); r6 = r6 ^ t_; } // 14 shfl
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r3 = r3 ^ t_; } // 15 shfl
|
||||
r2 = r2 | r7; // 16 or
|
||||
r0 = rotr_var(r0, r6); // 17 rotr
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
|
||||
r7 = rotl_imm(r7, 24u); // 19 rotl
|
||||
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
|
||||
r2 = r2 | r6; // 21 or
|
||||
r6 = r5 * r4 + r6; // 22 mad
|
||||
r1 = r1 | r0; // 23 or
|
||||
r6 = r6 ^ r1; // 24 xor
|
||||
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
|
||||
r7 = rotr_var(r7, r0); // 26 rotr
|
||||
r4 = r5 * r7 + r4; // 27 mad
|
||||
r3 = r2 * r4 + r3; // 28 mad
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
|
||||
r2 = r2 ^ ds[((rotl_imm(r3 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r5 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
|
||||
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
|
||||
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
|
||||
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
|
||||
r2 = mul_hi(r2, r5); // 35 mulhi
|
||||
r4 = r4 ^ r2; // 36 xor
|
||||
r6 = r6 * r5; // 37 mul
|
||||
r7 = r7 ^ r0; // 38 xor
|
||||
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
|
||||
r2 = rotr_var(r2, r3); // 40 rotr
|
||||
r7 = r7 - r0; // 41 sub
|
||||
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 43 shfl
|
||||
r0 = r0 ^ ds[((rotl_imm(r7 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
|
||||
r3 = r3 ^ ds[((rotl_imm(r1 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
|
||||
r6 = r6 ^ ds[((rotl_imm(r3 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
|
||||
r4 = mul_hi(r4, r2); // 48 mulhi
|
||||
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r0 = r0 ^ t_; } // 50 shfl
|
||||
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r2 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
|
||||
r6 = rotl_imm(r6, 12u); // 53 rotl
|
||||
r3 = rotl_imm(r3, 12u); // 54 rotl
|
||||
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
|
||||
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r0 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
|
||||
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
|
||||
r3 = mul_hi(r3, r2); // 60 mulhi
|
||||
r6 = r6 | r4; // 61 or
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r5 = r5 ^ t_; } // 62 shfl
|
||||
r3 = r3 ^ ds[((rotl_imm(r6 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
|
||||
#if IGNEUM_EXCHANGE != 0
|
||||
// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the
|
||||
// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a
|
||||
// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md.
|
||||
IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) {
|
||||
if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); }
|
||||
}
|
||||
#endif
|
||||
163
proto-cuda/packs-ca2-era/era-0/kernel.cu
Normal file
163
proto-cuda/packs-ca2-era/era-0/kernel.cu
Normal file
|
|
@ -0,0 +1,163 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal).
|
||||
// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC.
|
||||
#include <cuda_runtime.h>
|
||||
#include <cstdint>
|
||||
#include "program.h"
|
||||
#include "memhard.h"
|
||||
|
||||
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31.
|
||||
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
|
||||
// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x.
|
||||
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) {
|
||||
uint32_t x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item.
|
||||
// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host.
|
||||
__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) {
|
||||
uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
if (seg < nSegments) mh_cache_segment(cache, seg);
|
||||
}
|
||||
__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
|
||||
uint32_t t = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
if (t < nItems) {
|
||||
uint32_t s[16];
|
||||
mh_item(cache, t, s);
|
||||
for (uint32_t i = 0u; i < 16u; ++i) ds[(size_t)mh_addr(t, i)] = s[i];
|
||||
}
|
||||
}
|
||||
|
||||
// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every
|
||||
// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a
|
||||
// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid.
|
||||
__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) {
|
||||
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
uint32_t nonce = baseNonce + gid;
|
||||
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint32_t x = nonce ^ 0x667d0fbdu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x7b8e5963u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
|
||||
{ uint32_t x = nonce ^ 0x7b8e5963u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x31c67e5eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
|
||||
{ uint32_t x = nonce ^ 0x31c67e5eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4529ddc6u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
|
||||
{ uint32_t x = nonce ^ 0x4529ddc6u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef19d6d8u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
|
||||
{ uint32_t x = nonce ^ 0xef19d6d8u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xaccf6211u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
|
||||
{ uint32_t x = nonce ^ 0xaccf6211u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xda0aed32u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
|
||||
{ uint32_t x = nonce ^ 0xda0aed32u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xabc6df31u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
|
||||
{ uint32_t x = nonce ^ 0xabc6df31u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x667d0fbdu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
|
||||
|
||||
for (uint32_t it = 0u; it < 8u; ++it) {
|
||||
uint32_t sel = r0;
|
||||
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
|
||||
r7 = r7 ^ r0; // 1 xor
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r6, 1); // 2 shfl
|
||||
r4 = r4 - r1; // 3 sub
|
||||
r2 = r0 * r4 + r2; // 4 mad
|
||||
r4 = rotl_imm(r4, 9u); // 5 rotl
|
||||
r0 = rotr_var(r0, r2); // 6 rotr
|
||||
r6 = r6 ^ ds[((rotl_imm(r7 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r2 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r0 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r1 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
|
||||
r5 = r5 * r4; // 12 mul
|
||||
r7 = r7 ^ ds[((rotl_imm(r6 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
|
||||
r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r5, 16); // 14 shfl
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r5, 1); // 15 shfl
|
||||
r2 = r2 | r7; // 16 or
|
||||
r0 = rotr_var(r0, r6); // 17 rotr
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
|
||||
r7 = rotl_imm(r7, 24u); // 19 rotl
|
||||
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
|
||||
r2 = r2 | r6; // 21 or
|
||||
r6 = r5 * r4 + r6; // 22 mad
|
||||
r1 = r1 | r0; // 23 or
|
||||
r6 = r6 ^ r1; // 24 xor
|
||||
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
|
||||
r7 = rotr_var(r7, r0); // 26 rotr
|
||||
r4 = r5 * r7 + r4; // 27 mad
|
||||
r3 = r2 * r4 + r3; // 28 mad
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
|
||||
r2 = r2 ^ ds[((rotl_imm(r3 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r5 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
|
||||
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
|
||||
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
|
||||
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
|
||||
r2 = __umulhi(r2, r5); // 35 mulhi
|
||||
r4 = r4 ^ r2; // 36 xor
|
||||
r6 = r6 * r5; // 37 mul
|
||||
r7 = r7 ^ r0; // 38 xor
|
||||
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
|
||||
r2 = rotr_var(r2, r3); // 40 rotr
|
||||
r7 = r7 - r0; // 41 sub
|
||||
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 43 shfl
|
||||
r0 = r0 ^ ds[((rotl_imm(r7 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
|
||||
r3 = r3 ^ ds[((rotl_imm(r1 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
|
||||
r6 = r6 ^ ds[((rotl_imm(r3 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
|
||||
r4 = __umulhi(r4, r2); // 48 mulhi
|
||||
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
|
||||
r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r7, 4); // 50 shfl
|
||||
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r2 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
|
||||
r6 = rotl_imm(r6, 12u); // 53 rotl
|
||||
r3 = rotl_imm(r3, 12u); // 54 rotl
|
||||
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
|
||||
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r0 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
|
||||
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
|
||||
r3 = __umulhi(r3, r2); // 60 mulhi
|
||||
r6 = r6 | r4; // 61 or
|
||||
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 62 shfl
|
||||
r3 = r3 ^ ds[((rotl_imm(r6 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
|
||||
}
|
||||
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
|
||||
}
|
||||
|
||||
// Host-side launch wrappers. Declared in program.h, called from host.cu.
|
||||
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) {
|
||||
if (nSegments == 0u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 256u;
|
||||
uint32_t grid = (nSegments + block - 1u) / block;
|
||||
igneum_cache_fill<<<grid, block>>>(cache, nSegments);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
|
||||
if (nItems == 0u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 256u;
|
||||
uint32_t grid = (nItems + block - 1u) / block;
|
||||
igneum_build<<<grid, block>>>(ds, cache, nItems);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
uint32_t nonces, uint32_t blockWarps) {
|
||||
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 32u * blockWarps;
|
||||
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
|
||||
igneum_hash<<<nonces / block, block>>>(ds, out, baseNonce, mask);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
|
||||
cudaFuncAttributes attr;
|
||||
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash);
|
||||
if (e != cudaSuccess) return e;
|
||||
*numRegs = attr.numRegs;
|
||||
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0);
|
||||
}
|
||||
375
proto-cuda/packs-ca2-era/era-0/kernel_bound.cl
Normal file
375
proto-cuda/packs-ca2-era/era-0/kernel_bound.cl
Normal file
|
|
@ -0,0 +1,375 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal).
|
||||
// Built from source at runtime by proto-opencl/host.c, which passes these defines:
|
||||
// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit)
|
||||
// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default)
|
||||
// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32
|
||||
// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition
|
||||
// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units;
|
||||
// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32.
|
||||
#ifndef IGNEUM_GROUP
|
||||
#define IGNEUM_GROUP 32
|
||||
#endif
|
||||
#ifndef IGNEUM_EXCHANGE
|
||||
#define IGNEUM_EXCHANGE 0
|
||||
#endif
|
||||
#ifdef __OPENCL_VERSION__
|
||||
#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1)))
|
||||
#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n]
|
||||
#if IGNEUM_EXCHANGE == 1
|
||||
#ifdef cl_khr_subgroups
|
||||
#pragma OPENCL EXTENSION cl_khr_subgroups : enable
|
||||
#endif
|
||||
#ifdef cl_khr_subgroup_shuffle
|
||||
#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable
|
||||
#endif
|
||||
#elif IGNEUM_EXCHANGE == 2
|
||||
#pragma OPENCL EXTENSION cl_intel_subgroups : enable
|
||||
#endif
|
||||
#else
|
||||
// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros.
|
||||
#include "emu_opencl.h"
|
||||
#endif
|
||||
|
||||
#if IGNEUM_EXCHANGE == 1
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m))
|
||||
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
|
||||
#elif IGNEUM_EXCHANGE == 2
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m))
|
||||
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
|
||||
#else
|
||||
// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per
|
||||
// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane
|
||||
// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's
|
||||
// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier.
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; }
|
||||
#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; }
|
||||
#endif
|
||||
|
||||
static inline uint splitmix32(uint x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32.
|
||||
static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); }
|
||||
// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x.
|
||||
static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); }
|
||||
static inline uint ds_elem(uint i, uint d0, uint d1) {
|
||||
uint x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
|
||||
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
|
||||
#define MH_CACHE_LINE_MASK 0x003fffffu
|
||||
#define MH_SEGMENT_LINES 64u
|
||||
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
|
||||
static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
|
||||
|
||||
// y = ChaCha12 core(x) + x
|
||||
static inline void mh_chacha_block(const uint* x, uint* y) {
|
||||
for (uint i = 0u; i < 16u; ++i) y[i] = x[i];
|
||||
for (uint r = 0u; r < 6u; ++r) {
|
||||
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
|
||||
}
|
||||
for (uint i = 0u; i < 16u; ++i) y[i] += x[i];
|
||||
}
|
||||
|
||||
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
|
||||
static inline void mh_cache_segment(__global uint* cache, uint seg) {
|
||||
uint prev[16]; uint x[16]; uint y[16];
|
||||
for (uint i = 0u; i < 16u; ++i) prev[i] = 0u;
|
||||
for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) {
|
||||
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
|
||||
x[4] = 0xceed56d7u ^ prev[4];
|
||||
x[5] = 0x9ba270d2u ^ prev[5];
|
||||
x[6] = 0x82caab2du ^ prev[6];
|
||||
x[7] = 0x81ebce0eu ^ prev[7];
|
||||
x[8] = 0x12b6ecf1u ^ prev[8];
|
||||
x[9] = 0xd0f3fd7cu ^ prev[9];
|
||||
x[10] = 0xd872eefeu ^ prev[10];
|
||||
x[11] = 0xc158c7bdu ^ prev[11];
|
||||
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
|
||||
mh_chacha_block(x, y);
|
||||
__global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
|
||||
}
|
||||
}
|
||||
|
||||
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
|
||||
static inline void mh_mixer(uint* s, uint rk) {
|
||||
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
|
||||
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
|
||||
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
|
||||
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
|
||||
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
|
||||
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
|
||||
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
|
||||
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
|
||||
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
|
||||
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
|
||||
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
|
||||
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
|
||||
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
|
||||
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
|
||||
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
|
||||
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
|
||||
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
|
||||
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
|
||||
}
|
||||
|
||||
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
|
||||
static inline void mh_item(__global const uint* cache, uint t, uint* s) {
|
||||
s[0] = 0xceed56d7u;
|
||||
s[1] = 0x9ba270d2u;
|
||||
s[2] = 0x82caab2du;
|
||||
s[3] = 0x81ebce0eu;
|
||||
s[4] = 0x12b6ecf1u;
|
||||
s[5] = 0xd0f3fd7cu;
|
||||
s[6] = 0xd872eefeu;
|
||||
s[7] = 0xc158c7bdu;
|
||||
s[8] = t * 0xf351d601u + 0xc6892460u;
|
||||
s[9] = t * 0xa3bb398fu + 0x25b7228au;
|
||||
s[10] = t * 0xb5a09e35u + 0xcd515004u;
|
||||
s[11] = t * 0x7509c9c1u + 0x2846527au;
|
||||
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
|
||||
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
|
||||
s[14] = t * 0xded91851u + 0x82961bacu;
|
||||
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
|
||||
for (uint r = 0u; r < 8u; ++r) {
|
||||
mh_mixer(s, 0x9E3779B9u * (r + 1u));
|
||||
__global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i];
|
||||
}
|
||||
mh_mixer(s, 0x9E3779B9u * 9u);
|
||||
}
|
||||
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 0 2 10 15 of w.
|
||||
static inline uint mh_j(uint w) { return ((w >> 0u) & 1u) | (((w >> 2u) & 1u) << 1) | (((w >> 10u) & 1u) << 2) | (((w >> 15u) & 1u) << 3); }
|
||||
static inline uint mh_t(uint w) { w = (w & 0x00007fffu) | ((w >> 16u) << 15u); w = (w & 0x000003ffu) | ((w >> 11u) << 10u); w = (w & 0x00000003u) | ((w >> 3u) << 2u); w = (w & 0x00000000u) | ((w >> 1u) << 0u); return w; }
|
||||
static inline uint mh_addr(uint t, uint j) { uint w = t; w = ((w >> 0u) << 1u) | (w & 0x00000000u) | (((j >> 0u) & 1u) << 0u); w = ((w >> 2u) << 3u) | (w & 0x00000003u) | (((j >> 1u) & 1u) << 2u); w = ((w >> 10u) << 11u) | (w & 0x000003ffu) | (((j >> 2u) & 1u) << 10u); w = ((w >> 15u) << 16u) | (w & 0x00007fffu) | (((j >> 3u) & 1u) << 15u); return w; }
|
||||
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
|
||||
static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }
|
||||
|
||||
// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item.
|
||||
// The same constants as memhard.h in this pack (one emitter, three dialects).
|
||||
__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) {
|
||||
uint seg = (uint)get_global_id(0);
|
||||
if (seg < nSegments) mh_cache_segment(cache, seg);
|
||||
}
|
||||
__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) {
|
||||
uint t = (uint)get_global_id(0);
|
||||
if (t < nItems) {
|
||||
uint s[16];
|
||||
mh_item(cache, t, s);
|
||||
for (uint i = 0u; i < 16u; ++i) ds[(ulong)mh_addr(t, i)] = s[i];
|
||||
}
|
||||
}
|
||||
|
||||
// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the
|
||||
// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and
|
||||
// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all).
|
||||
IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) {
|
||||
uint gid = (uint)get_global_id(0);
|
||||
uint lid = (uint)get_local_id(0);
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
#if IGNEUM_EXCHANGE == 0
|
||||
IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
|
||||
uint xk = 0u;
|
||||
#else
|
||||
(void)lid;
|
||||
#endif
|
||||
{ uint x = nonce ^ 0x667d0fbdu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x7b8e5963u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
|
||||
{ uint x = nonce ^ 0x7b8e5963u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x31c67e5eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
|
||||
{ uint x = nonce ^ 0x31c67e5eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4529ddc6u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
|
||||
{ uint x = nonce ^ 0x4529ddc6u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef19d6d8u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
|
||||
{ uint x = nonce ^ 0xef19d6d8u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xaccf6211u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
|
||||
{ uint x = nonce ^ 0xaccf6211u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xda0aed32u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
|
||||
{ uint x = nonce ^ 0xda0aed32u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xabc6df31u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
|
||||
{ uint x = nonce ^ 0xabc6df31u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x667d0fbdu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
|
||||
r7 = r7 ^ r0; // 1 xor
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 1u); r3 = r3 ^ t_; } // 2 shfl
|
||||
r4 = r4 - r1; // 3 sub
|
||||
r2 = r0 * r4 + r2; // 4 mad
|
||||
r4 = rotl_imm(r4, 9u); // 5 rotl
|
||||
r0 = rotr_var(r0, r2); // 6 rotr
|
||||
r6 = r6 ^ ds[((rotl_imm(r7 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r2 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r0 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r1 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
|
||||
r5 = r5 * r4; // 12 mul
|
||||
r7 = r7 ^ ds[((rotl_imm(r6 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 16u); r6 = r6 ^ t_; } // 14 shfl
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r3 = r3 ^ t_; } // 15 shfl
|
||||
r2 = r2 | r7; // 16 or
|
||||
r0 = rotr_var(r0, r6); // 17 rotr
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
|
||||
r7 = rotl_imm(r7, 24u); // 19 rotl
|
||||
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
|
||||
r2 = r2 | r6; // 21 or
|
||||
r6 = r5 * r4 + r6; // 22 mad
|
||||
r1 = r1 | r0; // 23 or
|
||||
r6 = r6 ^ r1; // 24 xor
|
||||
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
|
||||
r7 = rotr_var(r7, r0); // 26 rotr
|
||||
r4 = r5 * r7 + r4; // 27 mad
|
||||
r3 = r2 * r4 + r3; // 28 mad
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
|
||||
r2 = r2 ^ ds[((rotl_imm(r3 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r5 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
|
||||
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
|
||||
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
|
||||
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
|
||||
r2 = mul_hi(r2, r5); // 35 mulhi
|
||||
r4 = r4 ^ r2; // 36 xor
|
||||
r6 = r6 * r5; // 37 mul
|
||||
r7 = r7 ^ r0; // 38 xor
|
||||
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
|
||||
r2 = rotr_var(r2, r3); // 40 rotr
|
||||
r7 = r7 - r0; // 41 sub
|
||||
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 43 shfl
|
||||
r0 = r0 ^ ds[((rotl_imm(r7 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
|
||||
r3 = r3 ^ ds[((rotl_imm(r1 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
|
||||
r6 = r6 ^ ds[((rotl_imm(r3 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
|
||||
r4 = mul_hi(r4, r2); // 48 mulhi
|
||||
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r0 = r0 ^ t_; } // 50 shfl
|
||||
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r2 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
|
||||
r6 = rotl_imm(r6, 12u); // 53 rotl
|
||||
r3 = rotl_imm(r3, 12u); // 54 rotl
|
||||
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
|
||||
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r0 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
|
||||
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
|
||||
r3 = mul_hi(r3, r2); // 60 mulhi
|
||||
r6 = r6 | r4; // 61 or
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r5 = r5 ^ t_; } // 62 shfl
|
||||
r3 = r3 ^ ds[((rotl_imm(r6 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
|
||||
#if IGNEUM_EXCHANGE != 0
|
||||
// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the
|
||||
// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a
|
||||
// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md.
|
||||
IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) {
|
||||
if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); }
|
||||
}
|
||||
#endif
|
||||
|
||||
// Header-bound variant (bind.rs): the init words come from initw, not SEEDW. Same body as igneum_hash.
|
||||
IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw) {
|
||||
uint gid = (uint)get_global_id(0);
|
||||
uint lid = (uint)get_local_id(0);
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7];
|
||||
#if IGNEUM_EXCHANGE == 0
|
||||
IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
|
||||
uint xk = 0u;
|
||||
#else
|
||||
(void)lid;
|
||||
#endif
|
||||
{ uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; }
|
||||
{ uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; }
|
||||
{ uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; }
|
||||
{ uint x = nonce ^ iw3; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw4; }
|
||||
{ uint x = nonce ^ iw4; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw5; }
|
||||
{ uint x = nonce ^ iw5; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw6; }
|
||||
{ uint x = nonce ^ iw6; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw7; }
|
||||
{ uint x = nonce ^ iw7; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw0; }
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
|
||||
r7 = r7 ^ r0; // 1 xor
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 1u); r3 = r3 ^ t_; } // 2 shfl
|
||||
r4 = r4 - r1; // 3 sub
|
||||
r2 = r0 * r4 + r2; // 4 mad
|
||||
r4 = rotl_imm(r4, 9u); // 5 rotl
|
||||
r0 = rotr_var(r0, r2); // 6 rotr
|
||||
r6 = r6 ^ ds[((rotl_imm(r7 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r2 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r0 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r1 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
|
||||
r5 = r5 * r4; // 12 mul
|
||||
r7 = r7 ^ ds[((rotl_imm(r6 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 16u); r6 = r6 ^ t_; } // 14 shfl
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r3 = r3 ^ t_; } // 15 shfl
|
||||
r2 = r2 | r7; // 16 or
|
||||
r0 = rotr_var(r0, r6); // 17 rotr
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
|
||||
r7 = rotl_imm(r7, 24u); // 19 rotl
|
||||
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
|
||||
r2 = r2 | r6; // 21 or
|
||||
r6 = r5 * r4 + r6; // 22 mad
|
||||
r1 = r1 | r0; // 23 or
|
||||
r6 = r6 ^ r1; // 24 xor
|
||||
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
|
||||
r7 = rotr_var(r7, r0); // 26 rotr
|
||||
r4 = r5 * r7 + r4; // 27 mad
|
||||
r3 = r2 * r4 + r3; // 28 mad
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
|
||||
r2 = r2 ^ ds[((rotl_imm(r3 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r5 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
|
||||
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
|
||||
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
|
||||
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
|
||||
r2 = mul_hi(r2, r5); // 35 mulhi
|
||||
r4 = r4 ^ r2; // 36 xor
|
||||
r6 = r6 * r5; // 37 mul
|
||||
r7 = r7 ^ r0; // 38 xor
|
||||
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
|
||||
r2 = rotr_var(r2, r3); // 40 rotr
|
||||
r7 = r7 - r0; // 41 sub
|
||||
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 43 shfl
|
||||
r0 = r0 ^ ds[((rotl_imm(r7 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
|
||||
r3 = r3 ^ ds[((rotl_imm(r1 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
|
||||
r6 = r6 ^ ds[((rotl_imm(r3 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
|
||||
r4 = mul_hi(r4, r2); // 48 mulhi
|
||||
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r0 = r0 ^ t_; } // 50 shfl
|
||||
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r2 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
|
||||
r6 = rotl_imm(r6, 12u); // 53 rotl
|
||||
r3 = rotl_imm(r3, 12u); // 54 rotl
|
||||
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
|
||||
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r0 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
|
||||
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
|
||||
r3 = mul_hi(r3, r2); // 60 mulhi
|
||||
r6 = r6 | r4; // 61 or
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r5 = r5 ^ t_; } // 62 shfl
|
||||
r3 = r3 ^ ds[((rotl_imm(r6 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
123
proto-cuda/packs-ca2-era/era-0/kernel_bound.cu
Normal file
123
proto-cuda/packs-ca2-era/era-0/kernel_bound.cu
Normal file
|
|
@ -0,0 +1,123 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// Header-bound twin of igneum_hash in kernel.cu: the init words come from a kernel argument, not SEEDW.
|
||||
// Host declarations (also in program_bound.h if present):
|
||||
// struct IgneumInitWords { uint32_t w[8]; };
|
||||
// cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
// IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps);
|
||||
// cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
|
||||
#include <cuda_runtime.h>
|
||||
#include <cstdint>
|
||||
#include "program.h"
|
||||
|
||||
struct IgneumInitWords { uint32_t w[8]; };
|
||||
|
||||
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
|
||||
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
|
||||
__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw) {
|
||||
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
uint32_t nonce = baseNonce + gid;
|
||||
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint32_t x = nonce ^ iw.w[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw.w[1]; }
|
||||
{ uint32_t x = nonce ^ iw.w[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw.w[2]; }
|
||||
{ uint32_t x = nonce ^ iw.w[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw.w[3]; }
|
||||
{ uint32_t x = nonce ^ iw.w[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw.w[4]; }
|
||||
{ uint32_t x = nonce ^ iw.w[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw.w[5]; }
|
||||
{ uint32_t x = nonce ^ iw.w[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw.w[6]; }
|
||||
{ uint32_t x = nonce ^ iw.w[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw.w[7]; }
|
||||
{ uint32_t x = nonce ^ iw.w[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw.w[0]; }
|
||||
|
||||
for (uint32_t it = 0u; it < 8u; ++it) {
|
||||
uint32_t sel = r0;
|
||||
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
|
||||
r7 = r7 ^ r0; // 1 xor
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r6, 1); // 2 shfl
|
||||
r4 = r4 - r1; // 3 sub
|
||||
r2 = r0 * r4 + r2; // 4 mad
|
||||
r4 = rotl_imm(r4, 9u); // 5 rotl
|
||||
r0 = rotr_var(r0, r2); // 6 rotr
|
||||
r6 = r6 ^ ds[((rotl_imm(r7 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r2 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r0 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r1 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
|
||||
r5 = r5 * r4; // 12 mul
|
||||
r7 = r7 ^ ds[((rotl_imm(r6 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
|
||||
r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r5, 16); // 14 shfl
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r5, 1); // 15 shfl
|
||||
r2 = r2 | r7; // 16 or
|
||||
r0 = rotr_var(r0, r6); // 17 rotr
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
|
||||
r7 = rotl_imm(r7, 24u); // 19 rotl
|
||||
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
|
||||
r2 = r2 | r6; // 21 or
|
||||
r6 = r5 * r4 + r6; // 22 mad
|
||||
r1 = r1 | r0; // 23 or
|
||||
r6 = r6 ^ r1; // 24 xor
|
||||
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
|
||||
r7 = rotr_var(r7, r0); // 26 rotr
|
||||
r4 = r5 * r7 + r4; // 27 mad
|
||||
r3 = r2 * r4 + r3; // 28 mad
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
|
||||
r2 = r2 ^ ds[((rotl_imm(r3 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r5 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
|
||||
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
|
||||
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
|
||||
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
|
||||
r2 = __umulhi(r2, r5); // 35 mulhi
|
||||
r4 = r4 ^ r2; // 36 xor
|
||||
r6 = r6 * r5; // 37 mul
|
||||
r7 = r7 ^ r0; // 38 xor
|
||||
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
|
||||
r2 = rotr_var(r2, r3); // 40 rotr
|
||||
r7 = r7 - r0; // 41 sub
|
||||
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 43 shfl
|
||||
r0 = r0 ^ ds[((rotl_imm(r7 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
|
||||
r3 = r3 ^ ds[((rotl_imm(r1 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
|
||||
r6 = r6 ^ ds[((rotl_imm(r3 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
|
||||
r4 = __umulhi(r4, r2); // 48 mulhi
|
||||
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
|
||||
r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r7, 4); // 50 shfl
|
||||
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r2 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
|
||||
r6 = rotl_imm(r6, 12u); // 53 rotl
|
||||
r3 = rotl_imm(r3, 12u); // 54 rotl
|
||||
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
|
||||
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r0 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
|
||||
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
|
||||
r3 = __umulhi(r3, r2); // 60 mulhi
|
||||
r6 = r6 | r4; // 61 or
|
||||
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 62 shfl
|
||||
r3 = r3 ^ ds[((rotl_imm(r6 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
|
||||
}
|
||||
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps) {
|
||||
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 32u * blockWarps;
|
||||
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
|
||||
igneum_hash_bound<<<nonces / block, block>>>(ds, out, baseNonce, mask, iw);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
|
||||
cudaFuncAttributes attr;
|
||||
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash_bound);
|
||||
if (e != cudaSuccess) return e;
|
||||
*numRegs = attr.numRegs;
|
||||
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash_bound, (int)(32u * blockWarps), 0);
|
||||
}
|
||||
112
proto-cuda/packs-ca2-era/era-0/memhard.h
Normal file
112
proto-cuda/packs-ca2-era/era-0/memhard.h
Normal file
|
|
@ -0,0 +1,112 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// Memory-hard dataset core, the same text that the Mac's Metal kernels and CPU verifier were checked against.
|
||||
// Included by kernel.cu (device), host.cu (host reference) and proto-opencl/host.c (C99 host reference).
|
||||
// See proto-metal/MEMHARD.md for the construction. kernel.cl carries the same text in OpenCL C.
|
||||
#pragma once
|
||||
#ifdef __cplusplus
|
||||
#include <cstdint>
|
||||
#else
|
||||
#include <stdint.h>
|
||||
#endif
|
||||
#if defined(__CUDACC__)
|
||||
#define IGNEUM_HD __host__ __device__ __forceinline__
|
||||
#elif defined(_MSC_VER) && !defined(__cplusplus)
|
||||
#define IGNEUM_HD static __inline
|
||||
#else
|
||||
#define IGNEUM_HD static inline
|
||||
#endif
|
||||
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
|
||||
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
|
||||
#define MH_CACHE_LINE_MASK 0x003fffffu
|
||||
#define MH_SEGMENT_LINES 64u
|
||||
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
|
||||
IGNEUM_HD uint32_t mh_rotl(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
|
||||
|
||||
// y = ChaCha12 core(x) + x
|
||||
IGNEUM_HD void mh_chacha_block(const uint32_t* x, uint32_t* y) {
|
||||
for (uint32_t i = 0u; i < 16u; ++i) y[i] = x[i];
|
||||
for (uint32_t r = 0u; r < 6u; ++r) {
|
||||
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
|
||||
}
|
||||
for (uint32_t i = 0u; i < 16u; ++i) y[i] += x[i];
|
||||
}
|
||||
|
||||
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
|
||||
IGNEUM_HD void mh_cache_segment(uint32_t* cache, uint32_t seg) {
|
||||
uint32_t prev[16]; uint32_t x[16]; uint32_t y[16];
|
||||
for (uint32_t i = 0u; i < 16u; ++i) prev[i] = 0u;
|
||||
for (uint32_t j = 0u; j < MH_SEGMENT_LINES; ++j) {
|
||||
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
|
||||
x[4] = 0xceed56d7u ^ prev[4];
|
||||
x[5] = 0x9ba270d2u ^ prev[5];
|
||||
x[6] = 0x82caab2du ^ prev[6];
|
||||
x[7] = 0x81ebce0eu ^ prev[7];
|
||||
x[8] = 0x12b6ecf1u ^ prev[8];
|
||||
x[9] = 0xd0f3fd7cu ^ prev[9];
|
||||
x[10] = 0xd872eefeu ^ prev[10];
|
||||
x[11] = 0xc158c7bdu ^ prev[11];
|
||||
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
|
||||
mh_chacha_block(x, y);
|
||||
uint32_t* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
|
||||
for (uint32_t i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
|
||||
}
|
||||
}
|
||||
|
||||
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
|
||||
IGNEUM_HD void mh_mixer(uint32_t* s, uint32_t rk) {
|
||||
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
|
||||
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
|
||||
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
|
||||
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
|
||||
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
|
||||
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
|
||||
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
|
||||
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
|
||||
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
|
||||
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
|
||||
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
|
||||
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
|
||||
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
|
||||
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
|
||||
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
|
||||
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
|
||||
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
|
||||
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
|
||||
}
|
||||
|
||||
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
|
||||
IGNEUM_HD void mh_item(const uint32_t* cache, uint32_t t, uint32_t* s) {
|
||||
s[0] = 0xceed56d7u;
|
||||
s[1] = 0x9ba270d2u;
|
||||
s[2] = 0x82caab2du;
|
||||
s[3] = 0x81ebce0eu;
|
||||
s[4] = 0x12b6ecf1u;
|
||||
s[5] = 0xd0f3fd7cu;
|
||||
s[6] = 0xd872eefeu;
|
||||
s[7] = 0xc158c7bdu;
|
||||
s[8] = t * 0xf351d601u + 0xc6892460u;
|
||||
s[9] = t * 0xa3bb398fu + 0x25b7228au;
|
||||
s[10] = t * 0xb5a09e35u + 0xcd515004u;
|
||||
s[11] = t * 0x7509c9c1u + 0x2846527au;
|
||||
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
|
||||
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
|
||||
s[14] = t * 0xded91851u + 0x82961bacu;
|
||||
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
|
||||
for (uint32_t r = 0u; r < 8u; ++r) {
|
||||
mh_mixer(s, 0x9E3779B9u * (r + 1u));
|
||||
const uint32_t* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
|
||||
for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= line[i];
|
||||
}
|
||||
mh_mixer(s, 0x9E3779B9u * 9u);
|
||||
}
|
||||
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 0 2 10 15 of w.
|
||||
IGNEUM_HD uint32_t mh_j(uint32_t w) { return ((w >> 0u) & 1u) | (((w >> 2u) & 1u) << 1) | (((w >> 10u) & 1u) << 2) | (((w >> 15u) & 1u) << 3); }
|
||||
IGNEUM_HD uint32_t mh_t(uint32_t w) { w = (w & 0x00007fffu) | ((w >> 16u) << 15u); w = (w & 0x000003ffu) | ((w >> 11u) << 10u); w = (w & 0x00000003u) | ((w >> 3u) << 2u); w = (w & 0x00000000u) | ((w >> 1u) << 0u); return w; }
|
||||
IGNEUM_HD uint32_t mh_addr(uint32_t t, uint32_t j) { uint32_t w = t; w = ((w >> 0u) << 1u) | (w & 0x00000000u) | (((j >> 0u) & 1u) << 0u); w = ((w >> 2u) << 3u) | (w & 0x00000003u) | (((j >> 1u) & 1u) << 2u); w = ((w >> 10u) << 11u) | (w & 0x000003ffu) | (((j >> 2u) & 1u) << 10u); w = ((w >> 15u) << 16u) | (w & 0x00007fffu) | (((j >> 3u) & 1u) << 15u); return w; }
|
||||
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
|
||||
IGNEUM_HD uint32_t mh_word(const uint32_t* cache, uint32_t w) { uint32_t s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }
|
||||
109
proto-cuda/packs-ca2-era/era-0/memhard.metal
Normal file
109
proto-cuda/packs-ca2-era/era-0/memhard.metal
Normal file
|
|
@ -0,0 +1,109 @@
|
|||
#include <metal_stdlib>
|
||||
using namespace metal;
|
||||
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
|
||||
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
|
||||
#define MH_CACHE_LINE_MASK 0x003fffffu
|
||||
#define MH_SEGMENT_LINES 64u
|
||||
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
|
||||
inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
|
||||
|
||||
// y = ChaCha12 core(x) + x
|
||||
inline void mh_chacha_block(const thread uint* x, thread uint* y) {
|
||||
for (uint i = 0u; i < 16u; ++i) y[i] = x[i];
|
||||
for (uint r = 0u; r < 6u; ++r) {
|
||||
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
|
||||
}
|
||||
for (uint i = 0u; i < 16u; ++i) y[i] += x[i];
|
||||
}
|
||||
|
||||
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
|
||||
inline void mh_cache_segment(device uint* cache, uint seg) {
|
||||
uint prev[16]; uint x[16]; uint y[16];
|
||||
for (uint i = 0u; i < 16u; ++i) prev[i] = 0u;
|
||||
for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) {
|
||||
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
|
||||
x[4] = 0xceed56d7u ^ prev[4];
|
||||
x[5] = 0x9ba270d2u ^ prev[5];
|
||||
x[6] = 0x82caab2du ^ prev[6];
|
||||
x[7] = 0x81ebce0eu ^ prev[7];
|
||||
x[8] = 0x12b6ecf1u ^ prev[8];
|
||||
x[9] = 0xd0f3fd7cu ^ prev[9];
|
||||
x[10] = 0xd872eefeu ^ prev[10];
|
||||
x[11] = 0xc158c7bdu ^ prev[11];
|
||||
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
|
||||
mh_chacha_block(x, y);
|
||||
device uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
|
||||
}
|
||||
}
|
||||
|
||||
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
|
||||
inline void mh_mixer(thread uint* s, uint rk) {
|
||||
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
|
||||
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
|
||||
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
|
||||
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
|
||||
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
|
||||
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
|
||||
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
|
||||
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
|
||||
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
|
||||
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
|
||||
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
|
||||
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
|
||||
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
|
||||
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
|
||||
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
|
||||
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
|
||||
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
|
||||
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
|
||||
}
|
||||
|
||||
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
|
||||
inline void mh_item(device const uint* cache, uint t, thread uint* s) {
|
||||
s[0] = 0xceed56d7u;
|
||||
s[1] = 0x9ba270d2u;
|
||||
s[2] = 0x82caab2du;
|
||||
s[3] = 0x81ebce0eu;
|
||||
s[4] = 0x12b6ecf1u;
|
||||
s[5] = 0xd0f3fd7cu;
|
||||
s[6] = 0xd872eefeu;
|
||||
s[7] = 0xc158c7bdu;
|
||||
s[8] = t * 0xf351d601u + 0xc6892460u;
|
||||
s[9] = t * 0xa3bb398fu + 0x25b7228au;
|
||||
s[10] = t * 0xb5a09e35u + 0xcd515004u;
|
||||
s[11] = t * 0x7509c9c1u + 0x2846527au;
|
||||
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
|
||||
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
|
||||
s[14] = t * 0xded91851u + 0x82961bacu;
|
||||
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
|
||||
for (uint r = 0u; r < 8u; ++r) {
|
||||
mh_mixer(s, 0x9E3779B9u * (r + 1u));
|
||||
device const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i];
|
||||
}
|
||||
mh_mixer(s, 0x9E3779B9u * 9u);
|
||||
}
|
||||
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 0 2 10 15 of w.
|
||||
inline uint mh_j(uint w) { return ((w >> 0u) & 1u) | (((w >> 2u) & 1u) << 1) | (((w >> 10u) & 1u) << 2) | (((w >> 15u) & 1u) << 3); }
|
||||
inline uint mh_t(uint w) { w = (w & 0x00007fffu) | ((w >> 16u) << 15u); w = (w & 0x000003ffu) | ((w >> 11u) << 10u); w = (w & 0x00000003u) | ((w >> 3u) << 2u); w = (w & 0x00000000u) | ((w >> 1u) << 0u); return w; }
|
||||
inline uint mh_addr(uint t, uint j) { uint w = t; w = ((w >> 0u) << 1u) | (w & 0x00000000u) | (((j >> 0u) & 1u) << 0u); w = ((w >> 2u) << 3u) | (w & 0x00000003u) | (((j >> 1u) & 1u) << 2u); w = ((w >> 10u) << 11u) | (w & 0x000003ffu) | (((j >> 2u) & 1u) << 10u); w = ((w >> 15u) << 16u) | (w & 0x00007fffu) | (((j >> 3u) & 1u) << 15u); return w; }
|
||||
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
|
||||
inline uint mh_word(device const uint* cache, uint w) { uint s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }
|
||||
|
||||
// One thread per segment (2^16 threads).
|
||||
kernel void igneum_cache_fill(device uint* cache [[buffer(0)]], uint gid [[thread_position_in_grid]]) {
|
||||
mh_cache_segment(cache, gid);
|
||||
}
|
||||
// One thread per 64-byte item (dataset words / 16 threads).
|
||||
kernel void igneum_build(device const uint* cache [[buffer(0)]], device uint* dataset [[buffer(1)]],
|
||||
uint gid [[thread_position_in_grid]]) {
|
||||
uint s[16];
|
||||
mh_item(cache, gid, s);
|
||||
for (uint i = 0u; i < 16u; ++i) dataset[mh_addr(gid, i)] = s[i];
|
||||
}
|
||||
76
proto-cuda/packs-ca2-era/era-0/program.h
Normal file
76
proto-cuda/packs-ca2-era/era-0/program.h
Normal file
|
|
@ -0,0 +1,76 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// Program metadata for host.cu plus the launch wrappers defined in kernel.cu.
|
||||
// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros.
|
||||
#pragma once
|
||||
#ifdef __cplusplus
|
||||
#include <cstdint>
|
||||
#else
|
||||
#include <stdint.h>
|
||||
#endif
|
||||
#ifndef IGNEUM_NO_CUDA
|
||||
#include <cuda_runtime.h>
|
||||
#endif
|
||||
|
||||
#define IGNEUM_SEED_STRING "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000"
|
||||
#define IGNEUM_SEED_BYTES_HEX "edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07"
|
||||
#define IGNEUM_GENERATOR 3
|
||||
#define IGNEUM_PROGRAM_ATTEMPT 0
|
||||
#define IGNEUM_PROGRAM_ID 0x73bcbfe8ccf988f1ull
|
||||
#define IGNEUM_DAY_STRING "bytes:69676e65756d2d6461792ffa50000000000000"
|
||||
#define IGNEUM_DAY_BYTES_HEX "69676e65756d2d6461792ffa50000000000000"
|
||||
#define IGNEUM_DAY0 0xceed56d7u
|
||||
#define IGNEUM_DAY1 0x9ba270d2u
|
||||
#define IGNEUM_DATASET_LOG2 28
|
||||
#define IGNEUM_MASK 0x0fffffffu
|
||||
#define IGNEUM_LANES 32
|
||||
#define IGNEUM_ITERATIONS 8
|
||||
#define IGNEUM_INSTR_COUNT 64
|
||||
#define IGNEUM_LOADS_PER_HASH 128
|
||||
#define IGNEUM_WIDE_LOADS_PER_HASH 0
|
||||
#define IGNEUM_OP_MIX "load=16 add=15 shfl=6 mad=4 or=4 rotl=4 rotr=4 xor=4 mulhi=3 mul=2 sub=2"
|
||||
// Program class v3 (Counter ASIC 2.0, docs/plans/counter-asic-2-rollout.md): generator version 3; a worker that
|
||||
// runs another class refuses this pack, and a job line names the class it wants (class=v3 era=<hex>).
|
||||
#define IGNEUM_PROGRAM_CLASS "v3"
|
||||
#define IGNEUM_ERA_SEED_HEX "8bffdd3366b9c3ffe89c1231e91538d531cfa717307df5c48ccba43096e17212"
|
||||
// Class v3 construction (Counter ASIC 2.0, 5 October 2026, docs/plans/mixer-x4.md): version 2 loads; the dataset item
|
||||
// derivation applies the mixer IGNEUM_MIXER_MULT times per round (memhard.h), and the cache follows the growth rule.
|
||||
#define IGNEUM_LOAD_CLASS "w4-erab2ed8a89"
|
||||
#define IGNEUM_LOAD_SLOTS 16
|
||||
#define IGNEUM_LOAD_MIX { 100, 0, 0 }
|
||||
#define IGNEUM_LOAD_WIDTH_COUNTS { 16, 0, 0 } // loads of 4, 16, 64 bytes per program
|
||||
#define IGNEUM_BYTES_PER_HASH 512
|
||||
#define IGNEUM_FOLD_ROT 11
|
||||
#define IGNEUM_FOLD_MUL 0x9e3779b1u
|
||||
// Era layout (5 October 2026, docs/plans/era-layout.md): NOT the lottery hash. Every dataset load reads
|
||||
// idx = ((rotl(src * STRIDE_MUL, STRIDE_ROT) & window mask) | window offset) & MASK; the window of a load site is the
|
||||
// dataset, a half or a quarter of it (IGNEUM_ERA_WINDOWS: site:shrink:offset); dataset word w holds word j(w) of item
|
||||
// t(w) with j's bits at the INTERLEAVE positions (memhard.h: mh_t, mh_j, mh_addr).
|
||||
#define IGNEUM_ERA_LABEL "b2ed8a89"
|
||||
#define IGNEUM_ERA_SEED_WORDS { 0xb2ed8a89u, 0xb023f2bau, 0x6bbf405eu, 0x98153cddu, 0x49428e54u, 0xfbfa65eeu, 0xbcb76d0du, 0x2f509891u }
|
||||
#define IGNEUM_ERA_ALLOWED_WIDTHS { 1, 0, 0 } // words, ascending, 0 = unused; one entry pins the width
|
||||
#define IGNEUM_ERA_WIDTH_WORDS 1
|
||||
#define IGNEUM_ERA_STRIDE_MUL 0x625e5ab3u
|
||||
#define IGNEUM_ERA_STRIDE_ROT 19
|
||||
#define IGNEUM_ERA_INTERLEAVE { 0, 2, 10, 15 }
|
||||
#define IGNEUM_ERA_WINDOWS "7:2:1 8:1:1 9:1:1 10:1:1 11:0:0 13:1:1 29:0:0 30:2:2 31:1:1 44:1:1 46:2:0 47:0:0 52:0:0 56:0:0 58:2:0 63:1:1"
|
||||
// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h)
|
||||
#define IGNEUM_DATASET_MODE 1
|
||||
|
||||
#define IGNEUM_SEEDW_INIT { 0x667d0fbdu, 0x7b8e5963u, 0x31c67e5eu, 0x4529ddc6u, 0xef19d6d8u, 0xaccf6211u, 0xda0aed32u, 0xabc6df31u }
|
||||
#define IGNEUM_KEY_INIT { 0xceed56d7u, 0x9ba270d2u, 0x82caab2du, 0x81ebce0eu, 0x12b6ecf1u, 0xd0f3fd7cu, 0xd872eefeu, 0xc158c7bdu }
|
||||
#define IGNEUM_CACHE_LOG2_WORDS 26
|
||||
#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6
|
||||
#define IGNEUM_CACHE_SEGMENTS 65536u
|
||||
#define IGNEUM_ITEM_ROUNDS 8
|
||||
#define IGNEUM_MIX_ROT_INIT { 17u, 12u, 20u, 23u, 7u, 3u, 27u, 16u }
|
||||
#define IGNEUM_MIX_MUL_INIT { 0xf351d601u, 0xa3bb398fu, 0xb5a09e35u, 0x7509c9c1u, 0x6bbf31e9u, 0xfc849a79u, 0xded91851u, 0x8d9113d1u, 0x0ff15225u, 0x3a5bdd41u, 0xab533435u, 0xe1c55ad5u, 0xe6d3bd0du, 0x9d9ffbbdu, 0xbb2a3cf3u, 0x50a7c08du }
|
||||
#define IGNEUM_MIX_RC_INIT { 0xc6892460u, 0x25b7228au, 0xcd515004u, 0x2846527au, 0xa6324241u, 0x36e3ec53u, 0x82961bacu, 0x0f97ba7du, 0xb6f921a9u, 0x3ada24e5u, 0xde20ab91u, 0x5378eeb2u, 0x7d161662u, 0x89353cc1u, 0xb1aa03a2u, 0x788acae6u }
|
||||
|
||||
#ifndef IGNEUM_NO_CUDA
|
||||
// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError().
|
||||
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments);
|
||||
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems);
|
||||
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
uint32_t nonces, uint32_t blockWarps);
|
||||
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
|
||||
#endif
|
||||
143
proto-cuda/packs-ca2-era/era-0/program.json
Normal file
143
proto-cuda/packs-ca2-era/era-0/program.json
Normal file
|
|
@ -0,0 +1,143 @@
|
|||
{
|
||||
"format": "igneum-program-pack-3",
|
||||
"generator": 3,
|
||||
"attempt": 0,
|
||||
"program_id": "0x73bcbfe8ccf988f1",
|
||||
"program_id_derivation": "FNV-1a 64 over 'igneum-program/' || generator_le32 || seed_words as little-endian bytes || attempt_le32",
|
||||
"dataset_mode": "memory-hard",
|
||||
"seed": "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000",
|
||||
"seed_bytes": "edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07",
|
||||
"seed_words": ["0x667d0fbd", "0x7b8e5963", "0x31c67e5e", "0x4529ddc6", "0xef19d6d8", "0xaccf6211", "0xda0aed32", "0xabc6df31"],
|
||||
"seed_derivation": "seed_words = FNV-1a 64 over seed_bytes (attempt 0) or seed_bytes || attempt_le32 (attempt k >= 1), basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32",
|
||||
"generator_rule": "version 2: exactly 16 load slots drawn first from instructions 1..63 (partial Fisher-Yates), the other 48 ops from the ten non-load weights (sum 75); a load's source is drawn from the registers other than dst written by an earlier instruction and not read by a load since; the candidate must pass the acceptance rule of spec 01 section 1.4.6 (static: no cyclically stale load source, every register has an injecting write; dynamic: 64 units on the seed-keyed closed-form dataset with no constant register bit, no lane-constant load site, under 164 saturated final values, every output bit within 136 of 1024, distinct addresses above 245760), else the next attempt of the seed is tried",
|
||||
"lanes": 32,
|
||||
"registers": 8,
|
||||
"iterations": 8,
|
||||
"instruction_count": 64,
|
||||
"loads_per_hash": 128,
|
||||
"program_class": "v3",
|
||||
"era_seed_bytes": "8bffdd3366b9c3ffe89c1231e91538d531cfa717307df5c48ccba43096e17212",
|
||||
"load_class": "w4-erab2ed8a89",
|
||||
"load_slots": 16,
|
||||
"load_mix_percent_4_16_64": [100, 0, 0],
|
||||
"load_width_counts_4_16_64": [16, 0, 0],
|
||||
"bytes_per_hash": 512,
|
||||
"wide_load": "read-width experiment (5 October 2026, docs/plans/read-width.md), NOT the lottery hash: a load of W words (width field, 4 or 16) reads dataset[b .. b + W) with b = (src & mask) & ~(W - 1) and folds every word into dst: x = dst ^ w[0]; for j in 1..W: x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x; width 1 is the plain load; the width is drawn per instruction from the class mix with one extra below(100) draw after the nine of version 2, and the program id is FNV-1a 64 over 'igneum-program-rw/' || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots",
|
||||
"era": {
|
||||
"label": "b2ed8a89",
|
||||
"seed_words": ["0xb2ed8a89", "0xb023f2ba", "0x6bbf405e", "0x98153cdd", "0x49428e54", "0xfbfa65ee", "0xbcb76d0d", "0x2f509891"],
|
||||
"draw": "docs/plans/era-layout.md 1.1: SplitMix64 seeded with seed_words[0] | seed_words[1] << 32 of seed_words_from_bytes('igneum-era/' || n_le64 || E_n); width = allowed[below(|allowed|)], stride_mul = low32(next()) | 1, stride_rot = 1 + below(31), then four next() draws for a partial Fisher-Yates over positions log2(W)..15 of which 4 - log2(W) are used",
|
||||
"allowed_widths": [1],
|
||||
"width_words": 1,
|
||||
"stride_mul": "0x625e5ab3",
|
||||
"stride_rot": 19,
|
||||
"interleave": [0, 2, 10, 15],
|
||||
"address": "y = rotl(src * stride_mul, stride_rot); k = min(win, D - 26); idx = ((y & (mask >> k)) | ((off & (2^k - 1)) << (D - k))) & mask; a wide load aligns idx down to W words",
|
||||
"windows": "per instruction, after the width roll: win = below(3), off = low32(next()) & (2^win - 1); used on a load slot (the instruction's win and off fields)",
|
||||
"dataset_word": "dataset[w] = item(t(w))[j(w)]: j(w) gathers the bits of w at the interleave positions, t(w) is w with those bits removed",
|
||||
"program_id_suffix": "'era/' || allowed[3] || width_words || stride_mul_le32 || stride_rot_le32 || interleave[4]"
|
||||
},
|
||||
"op_mix": {"load": 16, "add": 15, "shfl": 6, "mad": 4, "or": 4, "rotl": 4, "rotr": 4, "xor": 4, "mulhi": 3, "mul": 2, "sub": 2},
|
||||
"register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]",
|
||||
"splitmix32": "x ^= x>>16; x *= 0x7feb352d; x ^= x>>15; x *= 0x846ca68b; x ^= x>>16",
|
||||
"iteration": "sel = r0 sampled once at the top of each iteration, then all instructions in order",
|
||||
"output": "lo = r0 ^ rotl(r1,7) ^ rotl(r2,14) ^ rotl(r3,21); hi = r4 ^ rotl(r5,9) ^ rotl(r6,18) ^ rotl(r7,27); out = (hi << 32) | lo",
|
||||
"op_semantics": {
|
||||
"add": "dst = dst + src + (bit `bit` of sel ? imm2 : imm)",
|
||||
"sub": "dst = dst - src",
|
||||
"mul": "dst = dst * src (low 32)",
|
||||
"mulhi": "dst = high 32 bits of dst * src",
|
||||
"xor": "dst = dst ^ src",
|
||||
"or": "dst = dst | src",
|
||||
"rotl": "dst = rotl(dst, rot), rot in 1..31",
|
||||
"rotr": "dst = rotr(dst, src & 31)",
|
||||
"mad": "dst = src * src2 + dst",
|
||||
"shfl": "dst = dst ^ (src of lane (lane ^ mask)), mask in {1,2,4,8,16}, within the 32-lane warp",
|
||||
"load": "dst = dst ^ dataset[src & dataset.mask]",
|
||||
"wload": "base = (src of lane 0 & dataset.mask) & ~31; dst = dst ^ dataset[base + lane] (warp-coalesced 128-byte load, lever b, only when --wide-frac > 0)"
|
||||
},
|
||||
"dataset": {
|
||||
"log2_words": 28,
|
||||
"bytes": 1073741824,
|
||||
"mask": "0x0fffffff",
|
||||
"day": "bytes:69676e65756d2d6461792ffa50000000000000",
|
||||
"day_bytes": "69676e65756d2d6461792ffa50000000000000",
|
||||
"day_words_from": "seed_words_from_bytes(day_bytes)",
|
||||
"d0": "0xceed56d7",
|
||||
"d1": "0x9ba270d2",
|
||||
"mode": "memory-hard",
|
||||
"spec": "proto-metal/MEMHARD.md",
|
||||
"key": ["0xceed56d7", "0x9ba270d2", "0x82caab2d", "0x81ebce0e", "0x12b6ecf1", "0xd0f3fd7c", "0xd872eefe", "0xc158c7bd"],
|
||||
"key_derivation": "the 8 words of seed_words_from_bytes(day_bytes); d0, d1 are key[0], key[1]",
|
||||
"cache": {"log2_words": 26, "bytes": 268435456, "line_words": 16, "segment_lines": 64, "segments": 65536, "block": "ChaCha12 core + feed-forward, rotations 16 12 8 7", "sigma": ["0x61707865", "0x3320646e", "0x79622d32", "0x6b206574"], "tag": ["0x49676e65", "0x756d4d48"], "chain": "in_j = prev_line ^ (sigma[0..3] || key[0..7] || seg || j || tag[0..1]); line_j = block(in_j); prev_0 = 0"},
|
||||
"mixer": {"draw": "SplitMix64 seeded with key[0] | key[1] << 32: rot[0..7] = 1 + next() % 31, mul[0..15] = low32(next()) | 1, rc[0..15] = low32(next())", "rot": [17, 12, 20, 23, 7, 3, 27, 16], "mul": ["0xf351d601", "0xa3bb398f", "0xb5a09e35", "0x7509c9c1", "0x6bbf31e9", "0xfc849a79", "0xded91851", "0x8d9113d1", "0x0ff15225", "0x3a5bdd41", "0xab533435", "0xe1c55ad5", "0xe6d3bd0d", "0x9d9ffbbd", "0xbb2a3cf3", "0x50a7c08d"], "rc": ["0xc6892460", "0x25b7228a", "0xcd515004", "0x2846527a", "0xa6324241", "0x36e3ec53", "0x82961bac", "0x0f97ba7d", "0xb6f921a9", "0x3ada24e5", "0xde20ab91", "0x5378eeb2", "0x7d161662", "0x89353cc1", "0xb1aa03a2", "0x788acae6"], "round": "for i in 0..15: s[i] = (s[i] ^ (rc[i] + (r+1) * 0x9E3779B9)) * mul[i]; then quarter rounds on columns (0,4,8,12) (1,5,9,13) (2,6,10,14) (3,7,11,15) with rot[0..3] and diagonals (0,5,10,15) (1,6,11,12) (2,7,8,13) (3,4,9,14) with rot[4..7]", "quarter_round": "a += b; d ^= a; d = rotl(d, r1); c += d; b ^= c; b = rotl(b, r2); a += b; d ^= a; d = rotl(d, r3); c += d; b ^= c; b = rotl(b, r4)"},
|
||||
"item": "s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..7: s = M_r(s); line = s[0] & 0x003fffff; s[i] ^= cache[line * 16 + i]; then s = M_8(s); item(t) = s",
|
||||
"word": "dataset[w] = item(w >> 4)[w & 15]"
|
||||
},
|
||||
"instructions": [
|
||||
{"i": 0, "op": "add", "dst": 4, "src": 5, "src2": 7, "imm": "0xea86e152", "imm2": "0x5810667a", "rot": 27, "bit": 13, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 1, "op": "xor", "dst": 7, "src": 0, "src2": 6, "imm": "0xbaab6229", "imm2": "0xed861989", "rot": 26, "bit": 22, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 2, "op": "shfl", "dst": 3, "src": 6, "src2": 2, "imm": "0x5b623116", "imm2": "0xff12e5b2", "rot": 12, "bit": 24, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 3, "op": "sub", "dst": 4, "src": 1, "src2": 1, "imm": "0xe99741c7", "imm2": "0xf5fa5009", "rot": 1, "bit": 21, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 4, "op": "mad", "dst": 2, "src": 0, "src2": 4, "imm": "0x673c2157", "imm2": "0xee02465f", "rot": 20, "bit": 22, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 5, "op": "rotl", "dst": 4, "src": 7, "src2": 7, "imm": "0x946f7818", "imm2": "0x45d3399e", "rot": 9, "bit": 2, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 6, "op": "rotr", "dst": 0, "src": 2, "src2": 4, "imm": "0x5f6a0ed2", "imm2": "0x7043a636", "rot": 19, "bit": 31, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 7, "op": "load", "dst": 6, "src": 7, "src2": 1, "imm": "0x5c61dcf7", "imm2": "0x7466aa40", "rot": 19, "bit": 9, "mask": 2, "width": 1, "win": 2, "off": 1},
|
||||
{"i": 8, "op": "load", "dst": 1, "src": 4, "src2": 5, "imm": "0x85668475", "imm2": "0xdb8cc483", "rot": 29, "bit": 7, "mask": 4, "width": 1, "win": 1, "off": 1},
|
||||
{"i": 9, "op": "load", "dst": 1, "src": 2, "src2": 4, "imm": "0x4454980f", "imm2": "0xebd31581", "rot": 10, "bit": 28, "mask": 16, "width": 1, "win": 1, "off": 1},
|
||||
{"i": 10, "op": "load", "dst": 7, "src": 0, "src2": 2, "imm": "0xe075297c", "imm2": "0x5779c44c", "rot": 10, "bit": 22, "mask": 2, "width": 1, "win": 1, "off": 1},
|
||||
{"i": 11, "op": "load", "dst": 7, "src": 1, "src2": 6, "imm": "0x65aa4311", "imm2": "0x4fe48ea9", "rot": 15, "bit": 9, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 12, "op": "mul", "dst": 5, "src": 4, "src2": 2, "imm": "0x1383d3ad", "imm2": "0xf3094b29", "rot": 8, "bit": 9, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 13, "op": "load", "dst": 7, "src": 6, "src2": 2, "imm": "0xed8a496f", "imm2": "0x3072c3c6", "rot": 28, "bit": 19, "mask": 8, "width": 1, "win": 1, "off": 1},
|
||||
{"i": 14, "op": "shfl", "dst": 6, "src": 5, "src2": 0, "imm": "0x8b965b57", "imm2": "0xcfeca6c1", "rot": 12, "bit": 27, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 15, "op": "shfl", "dst": 3, "src": 5, "src2": 3, "imm": "0x877c7586", "imm2": "0xa9cb2a03", "rot": 2, "bit": 29, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 16, "op": "or", "dst": 2, "src": 7, "src2": 3, "imm": "0xb740221a", "imm2": "0x89d38d6d", "rot": 6, "bit": 31, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 17, "op": "rotr", "dst": 0, "src": 6, "src2": 1, "imm": "0x26f3ad8a", "imm2": "0x27256f15", "rot": 18, "bit": 5, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 18, "op": "add", "dst": 7, "src": 5, "src2": 7, "imm": "0xf66e7017", "imm2": "0xb9e3577e", "rot": 9, "bit": 12, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 19, "op": "rotl", "dst": 7, "src": 0, "src2": 5, "imm": "0x849ae6ee", "imm2": "0x02b358f9", "rot": 24, "bit": 18, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 20, "op": "add", "dst": 6, "src": 7, "src2": 6, "imm": "0x52334d12", "imm2": "0x8c9f0ef8", "rot": 11, "bit": 23, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 21, "op": "or", "dst": 2, "src": 6, "src2": 5, "imm": "0xb1871e63", "imm2": "0xb2e40191", "rot": 5, "bit": 13, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 22, "op": "mad", "dst": 6, "src": 5, "src2": 4, "imm": "0x97df29e4", "imm2": "0xe60fea84", "rot": 11, "bit": 16, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 23, "op": "or", "dst": 1, "src": 0, "src2": 3, "imm": "0x8f30d21d", "imm2": "0x2df685a0", "rot": 31, "bit": 0, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 24, "op": "xor", "dst": 6, "src": 1, "src2": 2, "imm": "0xd2c4025f", "imm2": "0x5269eb4d", "rot": 31, "bit": 21, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 25, "op": "add", "dst": 2, "src": 6, "src2": 5, "imm": "0x659fc3d3", "imm2": "0x9cec0e12", "rot": 6, "bit": 17, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 26, "op": "rotr", "dst": 7, "src": 0, "src2": 1, "imm": "0xd89ef484", "imm2": "0x20be3846", "rot": 12, "bit": 9, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 27, "op": "mad", "dst": 4, "src": 5, "src2": 7, "imm": "0xb48420ae", "imm2": "0x3d1f2485", "rot": 14, "bit": 28, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 28, "op": "mad", "dst": 3, "src": 2, "src2": 4, "imm": "0xc5c46d76", "imm2": "0x700044b5", "rot": 22, "bit": 10, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 29, "op": "load", "dst": 1, "src": 4, "src2": 3, "imm": "0x0fbaf177", "imm2": "0xfff4f2ed", "rot": 20, "bit": 31, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 30, "op": "load", "dst": 2, "src": 3, "src2": 0, "imm": "0xe90eb2e5", "imm2": "0xb0f9eb79", "rot": 27, "bit": 14, "mask": 4, "width": 1, "win": 2, "off": 2},
|
||||
{"i": 31, "op": "load", "dst": 1, "src": 5, "src2": 2, "imm": "0x97ba3fc3", "imm2": "0x7894e657", "rot": 3, "bit": 30, "mask": 16, "width": 1, "win": 1, "off": 1},
|
||||
{"i": 32, "op": "add", "dst": 7, "src": 2, "src2": 7, "imm": "0x070888a8", "imm2": "0xe403240e", "rot": 2, "bit": 16, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 33, "op": "add", "dst": 2, "src": 0, "src2": 5, "imm": "0xf2e46d55", "imm2": "0x29701828", "rot": 31, "bit": 28, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 34, "op": "add", "dst": 2, "src": 3, "src2": 0, "imm": "0x58f75b87", "imm2": "0x343b7aee", "rot": 12, "bit": 14, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 35, "op": "mulhi", "dst": 2, "src": 5, "src2": 6, "imm": "0x5892a9e6", "imm2": "0xc9824c94", "rot": 19, "bit": 26, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 36, "op": "xor", "dst": 4, "src": 2, "src2": 3, "imm": "0x7b5b5474", "imm2": "0x45cfc5dd", "rot": 17, "bit": 18, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 37, "op": "mul", "dst": 6, "src": 5, "src2": 7, "imm": "0xccf564a5", "imm2": "0x873ad101", "rot": 7, "bit": 11, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 38, "op": "xor", "dst": 7, "src": 0, "src2": 6, "imm": "0xcac8f06d", "imm2": "0x6b97c683", "rot": 18, "bit": 28, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 39, "op": "add", "dst": 7, "src": 2, "src2": 4, "imm": "0xb8180e9d", "imm2": "0x32bbd117", "rot": 23, "bit": 19, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 40, "op": "rotr", "dst": 2, "src": 3, "src2": 0, "imm": "0x2d6070bc", "imm2": "0x68ff101e", "rot": 13, "bit": 18, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 41, "op": "sub", "dst": 7, "src": 0, "src2": 2, "imm": "0x0daf96ea", "imm2": "0x36f37be1", "rot": 5, "bit": 0, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 42, "op": "add", "dst": 4, "src": 3, "src2": 6, "imm": "0x6ced15b7", "imm2": "0x6df7aed4", "rot": 19, "bit": 4, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 43, "op": "shfl", "dst": 7, "src": 3, "src2": 5, "imm": "0x8ace05f3", "imm2": "0xd378ec12", "rot": 23, "bit": 10, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 44, "op": "load", "dst": 0, "src": 7, "src2": 4, "imm": "0xb0607786", "imm2": "0xc4acabbc", "rot": 13, "bit": 7, "mask": 16, "width": 1, "win": 1, "off": 1},
|
||||
{"i": 45, "op": "add", "dst": 1, "src": 6, "src2": 5, "imm": "0xe09f54e9", "imm2": "0x83e825bf", "rot": 23, "bit": 14, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 46, "op": "load", "dst": 3, "src": 1, "src2": 0, "imm": "0x63cc1e4e", "imm2": "0xa1be8118", "rot": 12, "bit": 6, "mask": 2, "width": 1, "win": 2, "off": 0},
|
||||
{"i": 47, "op": "load", "dst": 6, "src": 3, "src2": 7, "imm": "0x353f1d79", "imm2": "0x3b2e7456", "rot": 18, "bit": 18, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 48, "op": "mulhi", "dst": 4, "src": 2, "src2": 7, "imm": "0x00d8a3cd", "imm2": "0x231866d2", "rot": 21, "bit": 20, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 49, "op": "add", "dst": 5, "src": 0, "src2": 2, "imm": "0xa8bae6df", "imm2": "0xf572bdb9", "rot": 14, "bit": 7, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 50, "op": "shfl", "dst": 0, "src": 7, "src2": 7, "imm": "0x81ef22e1", "imm2": "0x74438fc5", "rot": 28, "bit": 18, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 51, "op": "add", "dst": 6, "src": 0, "src2": 6, "imm": "0x383b9260", "imm2": "0x11e17c61", "rot": 12, "bit": 19, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 52, "op": "load", "dst": 5, "src": 2, "src2": 2, "imm": "0xfb84f451", "imm2": "0x11cd863e", "rot": 21, "bit": 20, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 53, "op": "rotl", "dst": 6, "src": 5, "src2": 4, "imm": "0xb1a7db6b", "imm2": "0x76686b9b", "rot": 12, "bit": 4, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 54, "op": "rotl", "dst": 3, "src": 6, "src2": 3, "imm": "0x6f981f52", "imm2": "0xd99aeba2", "rot": 12, "bit": 27, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 55, "op": "add", "dst": 2, "src": 1, "src2": 2, "imm": "0xac6be8e3", "imm2": "0x18d67dbb", "rot": 26, "bit": 10, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 56, "op": "load", "dst": 1, "src": 4, "src2": 0, "imm": "0x7e7f6a00", "imm2": "0x6f0747da", "rot": 25, "bit": 28, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 57, "op": "add", "dst": 5, "src": 0, "src2": 4, "imm": "0xf03673fe", "imm2": "0xa75cd60d", "rot": 16, "bit": 12, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 58, "op": "load", "dst": 5, "src": 0, "src2": 2, "imm": "0x227f94a6", "imm2": "0x0e8344f9", "rot": 20, "bit": 10, "mask": 2, "width": 1, "win": 2, "off": 0},
|
||||
{"i": 59, "op": "add", "dst": 1, "src": 4, "src2": 3, "imm": "0xdecd4794", "imm2": "0x8dfb96bb", "rot": 21, "bit": 7, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 60, "op": "mulhi", "dst": 3, "src": 2, "src2": 2, "imm": "0x0dd268e0", "imm2": "0x53034ca9", "rot": 1, "bit": 8, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 61, "op": "or", "dst": 6, "src": 4, "src2": 7, "imm": "0x3a45a321", "imm2": "0x9bc59a5f", "rot": 25, "bit": 11, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 62, "op": "shfl", "dst": 5, "src": 4, "src2": 2, "imm": "0x8f229cc1", "imm2": "0xcaac64a2", "rot": 17, "bit": 13, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 63, "op": "load", "dst": 3, "src": 6, "src2": 6, "imm": "0xb2574178", "imm2": "0xcbcc798d", "rot": 28, "bit": 0, "mask": 16, "width": 1, "win": 1, "off": 1}
|
||||
]
|
||||
}
|
||||
109
proto-cuda/packs-ca2-era/era-0/program.metal
Normal file
109
proto-cuda/packs-ca2-era/era-0/program.metal
Normal file
|
|
@ -0,0 +1,109 @@
|
|||
#include <metal_stdlib>
|
||||
using namespace metal;
|
||||
|
||||
#define MASK 0x0fffffffu
|
||||
constant uint SEEDW[8] = { 0x667d0fbdu, 0x7b8e5963u, 0x31c67e5eu, 0x4529ddc6u, 0xef19d6d8u, 0xaccf6211u, 0xda0aed32u, 0xabc6df31u };
|
||||
|
||||
inline uint splitmix32(uint x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31
|
||||
inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
inline uint ds_elem(uint i, uint d0, uint d1) {
|
||||
uint x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
kernel void igneum_hash(device const uint* dataset [[buffer(0)]],
|
||||
device ulong* out [[buffer(1)]],
|
||||
constant uint& baseNonce [[buffer(2)]],
|
||||
uint gid [[thread_position_in_grid]]) {
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint x = nonce ^ SEEDW[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ SEEDW[1]; }
|
||||
{ uint x = nonce ^ SEEDW[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ SEEDW[2]; }
|
||||
{ uint x = nonce ^ SEEDW[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ SEEDW[3]; }
|
||||
{ uint x = nonce ^ SEEDW[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ SEEDW[4]; }
|
||||
{ uint x = nonce ^ SEEDW[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ SEEDW[5]; }
|
||||
{ uint x = nonce ^ SEEDW[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ SEEDW[6]; }
|
||||
{ uint x = nonce ^ SEEDW[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ SEEDW[7]; }
|
||||
{ uint x = nonce ^ SEEDW[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ SEEDW[0]; }
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = r4 + r5 + select(0xea86e152u, 0x5810667au, ((sel >> 13u) & 1u) != 0u); // 0
|
||||
r7 = r7 ^ r0; // 1
|
||||
r3 = r3 ^ simd_shuffle_xor(r6, (ushort)1); // 2
|
||||
r4 = r4 - r1; // 3
|
||||
r2 = r0 * r4 + r2; // 4
|
||||
r4 = rotl_imm(r4, 9u); // 5
|
||||
r0 = rotr_var(r0, r2); // 6
|
||||
r6 = r6 ^ dataset[((rotl_imm(r7 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x04000000u) & MASK]; // 7
|
||||
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 8
|
||||
r1 = r1 ^ dataset[((rotl_imm(r2 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 9
|
||||
r7 = r7 ^ dataset[((rotl_imm(r0 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 10
|
||||
r7 = r7 ^ dataset[((rotl_imm(r1 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 11
|
||||
r5 = r5 * r4; // 12
|
||||
r7 = r7 ^ dataset[((rotl_imm(r6 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 13
|
||||
r6 = r6 ^ simd_shuffle_xor(r5, (ushort)16); // 14
|
||||
r3 = r3 ^ simd_shuffle_xor(r5, (ushort)1); // 15
|
||||
r2 = r2 | r7; // 16
|
||||
r0 = rotr_var(r0, r6); // 17
|
||||
r7 = r7 + r5 + select(0xf66e7017u, 0xb9e3577eu, ((sel >> 12u) & 1u) != 0u); // 18
|
||||
r7 = rotl_imm(r7, 24u); // 19
|
||||
r6 = r6 + r7 + select(0x52334d12u, 0x8c9f0ef8u, ((sel >> 23u) & 1u) != 0u); // 20
|
||||
r2 = r2 | r6; // 21
|
||||
r6 = r5 * r4 + r6; // 22
|
||||
r1 = r1 | r0; // 23
|
||||
r6 = r6 ^ r1; // 24
|
||||
r2 = r2 + r6 + select(0x659fc3d3u, 0x9cec0e12u, ((sel >> 17u) & 1u) != 0u); // 25
|
||||
r7 = rotr_var(r7, r0); // 26
|
||||
r4 = r5 * r7 + r4; // 27
|
||||
r3 = r2 * r4 + r3; // 28
|
||||
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 29
|
||||
r2 = r2 ^ dataset[((rotl_imm(r3 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x08000000u) & MASK]; // 30
|
||||
r1 = r1 ^ dataset[((rotl_imm(r5 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 31
|
||||
r7 = r7 + r2 + select(0x070888a8u, 0xe403240eu, ((sel >> 16u) & 1u) != 0u); // 32
|
||||
r2 = r2 + r0 + select(0xf2e46d55u, 0x29701828u, ((sel >> 28u) & 1u) != 0u); // 33
|
||||
r2 = r2 + r3 + select(0x58f75b87u, 0x343b7aeeu, ((sel >> 14u) & 1u) != 0u); // 34
|
||||
r2 = mulhi(r2, r5); // 35
|
||||
r4 = r4 ^ r2; // 36
|
||||
r6 = r6 * r5; // 37
|
||||
r7 = r7 ^ r0; // 38
|
||||
r7 = r7 + r2 + select(0xb8180e9du, 0x32bbd117u, ((sel >> 19u) & 1u) != 0u); // 39
|
||||
r2 = rotr_var(r2, r3); // 40
|
||||
r7 = r7 - r0; // 41
|
||||
r4 = r4 + r3 + select(0x6ced15b7u, 0x6df7aed4u, ((sel >> 4u) & 1u) != 0u); // 42
|
||||
r7 = r7 ^ simd_shuffle_xor(r3, (ushort)4); // 43
|
||||
r0 = r0 ^ dataset[((rotl_imm(r7 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 44
|
||||
r1 = r1 + r6 + select(0xe09f54e9u, 0x83e825bfu, ((sel >> 14u) & 1u) != 0u); // 45
|
||||
r3 = r3 ^ dataset[((rotl_imm(r1 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 46
|
||||
r6 = r6 ^ dataset[((rotl_imm(r3 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 47
|
||||
r4 = mulhi(r4, r2); // 48
|
||||
r5 = r5 + r0 + select(0xa8bae6dfu, 0xf572bdb9u, ((sel >> 7u) & 1u) != 0u); // 49
|
||||
r0 = r0 ^ simd_shuffle_xor(r7, (ushort)4); // 50
|
||||
r6 = r6 + r0 + select(0x383b9260u, 0x11e17c61u, ((sel >> 19u) & 1u) != 0u); // 51
|
||||
r5 = r5 ^ dataset[((rotl_imm(r2 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 52
|
||||
r6 = rotl_imm(r6, 12u); // 53
|
||||
r3 = rotl_imm(r3, 12u); // 54
|
||||
r2 = r2 + r1 + select(0xac6be8e3u, 0x18d67dbbu, ((sel >> 10u) & 1u) != 0u); // 55
|
||||
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 56
|
||||
r5 = r5 + r0 + select(0xf03673feu, 0xa75cd60du, ((sel >> 12u) & 1u) != 0u); // 57
|
||||
r5 = r5 ^ dataset[((rotl_imm(r0 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 58
|
||||
r1 = r1 + r4 + select(0xdecd4794u, 0x8dfb96bbu, ((sel >> 7u) & 1u) != 0u); // 59
|
||||
r3 = mulhi(r3, r2); // 60
|
||||
r6 = r6 | r4; // 61
|
||||
r5 = r5 ^ simd_shuffle_xor(r4, (ushort)8); // 62
|
||||
r3 = r3 ^ dataset[((rotl_imm(r6 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 63
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
111
proto-cuda/packs-ca2-era/era-0/program_bound.metal
Normal file
111
proto-cuda/packs-ca2-era/era-0/program_bound.metal
Normal file
|
|
@ -0,0 +1,111 @@
|
|||
#include <metal_stdlib>
|
||||
using namespace metal;
|
||||
|
||||
#define MASK 0x0fffffffu
|
||||
constant uint SEEDW[8] = { 0x667d0fbdu, 0x7b8e5963u, 0x31c67e5eu, 0x4529ddc6u, 0xef19d6d8u, 0xaccf6211u, 0xda0aed32u, 0xabc6df31u };
|
||||
|
||||
inline uint splitmix32(uint x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31
|
||||
inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
inline uint ds_elem(uint i, uint d0, uint d1) {
|
||||
uint x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// Header-bound variant: the init words come from buffer 3 (bind.rs), not from SEEDW.
|
||||
kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]],
|
||||
device ulong* out [[buffer(1)]],
|
||||
constant uint& baseNonce [[buffer(2)]],
|
||||
constant uint* initw [[buffer(3)]],
|
||||
uint gid [[thread_position_in_grid]]) {
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint x = nonce ^ initw[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ initw[1]; }
|
||||
{ uint x = nonce ^ initw[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ initw[2]; }
|
||||
{ uint x = nonce ^ initw[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ initw[3]; }
|
||||
{ uint x = nonce ^ initw[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ initw[4]; }
|
||||
{ uint x = nonce ^ initw[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ initw[5]; }
|
||||
{ uint x = nonce ^ initw[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ initw[6]; }
|
||||
{ uint x = nonce ^ initw[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ initw[7]; }
|
||||
{ uint x = nonce ^ initw[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ initw[0]; }
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = r4 + r5 + select(0xea86e152u, 0x5810667au, ((sel >> 13u) & 1u) != 0u); // 0
|
||||
r7 = r7 ^ r0; // 1
|
||||
r3 = r3 ^ simd_shuffle_xor(r6, (ushort)1); // 2
|
||||
r4 = r4 - r1; // 3
|
||||
r2 = r0 * r4 + r2; // 4
|
||||
r4 = rotl_imm(r4, 9u); // 5
|
||||
r0 = rotr_var(r0, r2); // 6
|
||||
r6 = r6 ^ dataset[((rotl_imm(r7 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x04000000u) & MASK]; // 7
|
||||
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 8
|
||||
r1 = r1 ^ dataset[((rotl_imm(r2 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 9
|
||||
r7 = r7 ^ dataset[((rotl_imm(r0 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 10
|
||||
r7 = r7 ^ dataset[((rotl_imm(r1 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 11
|
||||
r5 = r5 * r4; // 12
|
||||
r7 = r7 ^ dataset[((rotl_imm(r6 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 13
|
||||
r6 = r6 ^ simd_shuffle_xor(r5, (ushort)16); // 14
|
||||
r3 = r3 ^ simd_shuffle_xor(r5, (ushort)1); // 15
|
||||
r2 = r2 | r7; // 16
|
||||
r0 = rotr_var(r0, r6); // 17
|
||||
r7 = r7 + r5 + select(0xf66e7017u, 0xb9e3577eu, ((sel >> 12u) & 1u) != 0u); // 18
|
||||
r7 = rotl_imm(r7, 24u); // 19
|
||||
r6 = r6 + r7 + select(0x52334d12u, 0x8c9f0ef8u, ((sel >> 23u) & 1u) != 0u); // 20
|
||||
r2 = r2 | r6; // 21
|
||||
r6 = r5 * r4 + r6; // 22
|
||||
r1 = r1 | r0; // 23
|
||||
r6 = r6 ^ r1; // 24
|
||||
r2 = r2 + r6 + select(0x659fc3d3u, 0x9cec0e12u, ((sel >> 17u) & 1u) != 0u); // 25
|
||||
r7 = rotr_var(r7, r0); // 26
|
||||
r4 = r5 * r7 + r4; // 27
|
||||
r3 = r2 * r4 + r3; // 28
|
||||
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 29
|
||||
r2 = r2 ^ dataset[((rotl_imm(r3 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x08000000u) & MASK]; // 30
|
||||
r1 = r1 ^ dataset[((rotl_imm(r5 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 31
|
||||
r7 = r7 + r2 + select(0x070888a8u, 0xe403240eu, ((sel >> 16u) & 1u) != 0u); // 32
|
||||
r2 = r2 + r0 + select(0xf2e46d55u, 0x29701828u, ((sel >> 28u) & 1u) != 0u); // 33
|
||||
r2 = r2 + r3 + select(0x58f75b87u, 0x343b7aeeu, ((sel >> 14u) & 1u) != 0u); // 34
|
||||
r2 = mulhi(r2, r5); // 35
|
||||
r4 = r4 ^ r2; // 36
|
||||
r6 = r6 * r5; // 37
|
||||
r7 = r7 ^ r0; // 38
|
||||
r7 = r7 + r2 + select(0xb8180e9du, 0x32bbd117u, ((sel >> 19u) & 1u) != 0u); // 39
|
||||
r2 = rotr_var(r2, r3); // 40
|
||||
r7 = r7 - r0; // 41
|
||||
r4 = r4 + r3 + select(0x6ced15b7u, 0x6df7aed4u, ((sel >> 4u) & 1u) != 0u); // 42
|
||||
r7 = r7 ^ simd_shuffle_xor(r3, (ushort)4); // 43
|
||||
r0 = r0 ^ dataset[((rotl_imm(r7 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 44
|
||||
r1 = r1 + r6 + select(0xe09f54e9u, 0x83e825bfu, ((sel >> 14u) & 1u) != 0u); // 45
|
||||
r3 = r3 ^ dataset[((rotl_imm(r1 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 46
|
||||
r6 = r6 ^ dataset[((rotl_imm(r3 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 47
|
||||
r4 = mulhi(r4, r2); // 48
|
||||
r5 = r5 + r0 + select(0xa8bae6dfu, 0xf572bdb9u, ((sel >> 7u) & 1u) != 0u); // 49
|
||||
r0 = r0 ^ simd_shuffle_xor(r7, (ushort)4); // 50
|
||||
r6 = r6 + r0 + select(0x383b9260u, 0x11e17c61u, ((sel >> 19u) & 1u) != 0u); // 51
|
||||
r5 = r5 ^ dataset[((rotl_imm(r2 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 52
|
||||
r6 = rotl_imm(r6, 12u); // 53
|
||||
r3 = rotl_imm(r3, 12u); // 54
|
||||
r2 = r2 + r1 + select(0xac6be8e3u, 0x18d67dbbu, ((sel >> 10u) & 1u) != 0u); // 55
|
||||
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x625e5ab3u, 19u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 56
|
||||
r5 = r5 + r0 + select(0xf03673feu, 0xa75cd60du, ((sel >> 12u) & 1u) != 0u); // 57
|
||||
r5 = r5 ^ dataset[((rotl_imm(r0 * 0x625e5ab3u, 19u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 58
|
||||
r1 = r1 + r4 + select(0xdecd4794u, 0x8dfb96bbu, ((sel >> 7u) & 1u) != 0u); // 59
|
||||
r3 = mulhi(r3, r2); // 60
|
||||
r6 = r6 | r4; // 61
|
||||
r5 = r5 ^ simd_shuffle_xor(r4, (ushort)8); // 62
|
||||
r3 = r3 ^ dataset[((rotl_imm(r6 * 0x625e5ab3u, 19u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 63
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
5
proto-cuda/packs-ca2-era/era-0/seeds.txt
Normal file
5
proto-cuda/packs-ca2-era/era-0/seeds.txt
Normal file
|
|
@ -0,0 +1,5 @@
|
|||
epoch_seed_hex edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07
|
||||
day_seed_hex 69676e65756d2d6461792ffa50000000000000
|
||||
epoch_index 0
|
||||
day_index 20730
|
||||
daa_score 0
|
||||
57
proto-cuda/packs-ca2-era/era-0/vectors.h
Normal file
57
proto-cuda/packs-ca2-era/era-0/vectors.h
Normal file
|
|
@ -0,0 +1,57 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v3, memory-hard dataset
|
||||
#pragma once
|
||||
#ifdef __cplusplus
|
||||
#include <cstdint>
|
||||
#else
|
||||
#include <stdint.h>
|
||||
#endif
|
||||
|
||||
#define IGNEUM_VEC_WARPS 3
|
||||
static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u };
|
||||
static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = {
|
||||
{ // base nonce 0
|
||||
0xa32ad2dd06264089ull, 0x854fea26361763aaull, 0xd958908d36ee63b5ull, 0x33641fe988c44ad9ull, 0xb5290b7892dbc7baull, 0xd815f329e6f806c0ull, 0x9fb28ff3fc29e39cull, 0xdbc28a5d8bf42ab4ull,
|
||||
0x06f955808ff71bb9ull, 0x1ec10fff7bcdeaddull, 0x1acbc289cddd83cdull, 0x395e1564cda10137ull, 0x8a2aac2aa5b190d2ull, 0x499b8401e1f3e412ull, 0x9e2320319c85fc43ull, 0x66ba9661bc421ad6ull,
|
||||
0x3ba69971f1c9d8bdull, 0x00d7a64d460bfefeull, 0xc0f6487a1fbc5978ull, 0x1a261ac8673e1146ull, 0xf5a95a7caaaafa0aull, 0x7bbdabd107842744ull, 0x49dd3bf9b77a16cfull, 0xf7026cc441495a05ull,
|
||||
0x49fc48c1cba2be91ull, 0x84c2fc11de378cf9ull, 0x96de9e40c9da551cull, 0x3d20f01e363aaf61ull, 0xef36922a5fcf96caull, 0x31f1b0c4b0b42aafull, 0x33c6406d516d5992ull, 0x847af3a4248a972full
|
||||
},
|
||||
{ // base nonce 4096
|
||||
0x2039f40a61003341ull, 0x8238174df3761142ull, 0x32a72a04c4ca4a3cull, 0x16aa0fd3f65da6b3ull, 0x502bf9db9bd98782ull, 0x109f60bc6167ff30ull, 0x3873c59792d6b552ull, 0xea98fba68fc32149ull,
|
||||
0x56a8440169cb0f83ull, 0x7362edc3a7c98252ull, 0x8f393a7ad9f4ff40ull, 0x6b3cb3a1e0b02453ull, 0xdcbdc40aedfa06c0ull, 0x2bc381227147a2f3ull, 0x1570ef95b0c8e31full, 0xf3ca64d0f9ca8c91ull,
|
||||
0x1c767a63cd6a0bf8ull, 0x8e1a13a1237a948bull, 0xff2241b0b53b3efdull, 0x95949c528978b1fbull, 0xce3519499b78db7eull, 0x56a0fcecaaa62096ull, 0x0a62c2bddc7041b2ull, 0x5ec463510d31d7c0ull,
|
||||
0x2fc74273ed47b17full, 0xedc5740b0c5c191eull, 0xe62f737c106216cbull, 0x185df496808afd70ull, 0x29070790e64a0bc7ull, 0xc0ddaeac8b22b14bull, 0x302982837b6d96b4ull, 0x523972727aa906b5ull
|
||||
},
|
||||
{ // base nonce 1000000
|
||||
0x3527292a5f4afb4eull, 0x8c4128a03d946030ull, 0xb5b598dd18eb207dull, 0x9f1b21b5bf9dce8full, 0xa791a53fa1a4733eull, 0x12d2f9629aaca3b8ull, 0x400e844a00ba254cull, 0x1e9f98bee0272859ull,
|
||||
0x9983064542aabb54ull, 0x017873e17fab719cull, 0xf3e747e327968d2bull, 0x562380f8016b9ca7ull, 0x3f5007c1ec21b4a8ull, 0x5612489676e34c2bull, 0x97e13b52bc125dabull, 0xe15cf6d8d12f9b00ull,
|
||||
0xd8ba508db7b12919ull, 0x2e16d66db0bd1837ull, 0x0026e228553d7b18ull, 0x0c6d22f5018958ddull, 0xe76da566fc36580dull, 0x7824813a77faff44ull, 0x663811c5504ff04dull, 0x75efd5493353d65aull,
|
||||
0x926b72ebe963caa2ull, 0x392723945438b76dull, 0x4df61663eb2bfe05ull, 0x9a305821ba9ceb59ull, 0x317f4a6ac3f65a88ull, 0xdbae9b884f6829a9ull, 0x97cc8c8866b1a0cbull, 0x5db142b6d7777df1ull
|
||||
}
|
||||
};
|
||||
|
||||
// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455).
|
||||
static const uint32_t IGNEUM_DS_HEAD[16] = {
|
||||
0x3dd50b1fu, 0x48edec90u, 0x93123701u, 0x3a7d2407u, 0x5119540eu, 0x657ac748u, 0x3358ef2du, 0x09e2105cu,
|
||||
0xcdd88b45u, 0xb03a0ac0u, 0x2c37c6b0u, 0x10e2c6a5u, 0x25e2929cu, 0x4317cabdu, 0x5fa0cdccu, 0x6178ee52u
|
||||
};
|
||||
static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u;
|
||||
static const uint32_t IGNEUM_DS_LAST = 0xf7b7180eu;
|
||||
// 64 sampled dataset words (index, value) computed on the Mac.
|
||||
#define IGNEUM_DS_SAMPLES 64
|
||||
static const uint32_t IGNEUM_DS_SAMPLE_INDEX[IGNEUM_DS_SAMPLES] = {
|
||||
59471966u, 217795994u, 208353206u, 42483309u, 172547758u, 148076330u, 183853158u, 214389424u, 267488061u, 169781097u, 184093494u, 153880993u, 84977930u, 46426879u, 3093825u, 225364072u, 44593546u, 260713159u, 168250303u, 52384140u, 223401610u, 45554030u, 95410555u, 175039924u, 79171087u, 267580473u, 24168642u, 37981670u, 171551130u, 195559979u, 204611762u, 140997658u, 138925853u, 86637313u, 20736778u, 219665210u, 160430336u, 264654675u, 8013395u, 228945585u, 213884386u, 104419827u, 44185464u, 142737231u, 99284897u, 132475900u, 61861762u, 132056166u, 262388043u, 91878046u, 117353561u, 124768597u, 71352993u, 190698941u, 46055428u, 55281366u, 165145231u, 106810753u, 171985651u, 232085256u, 159510492u, 40072060u, 209107596u, 39023794u
|
||||
};
|
||||
static const uint32_t IGNEUM_DS_SAMPLE_VALUE[IGNEUM_DS_SAMPLES] = {
|
||||
0x1faef64au, 0x068e4a54u, 0x24551c62u, 0x1234f539u, 0xc15100f5u, 0x2f53b6b0u, 0xbab5fb76u, 0x06edf897u, 0xb0880fc6u, 0xccfddc4du, 0xf640c0f4u, 0xf145a054u, 0xa2425e8fu, 0x0a218118u, 0xd310512au, 0xb612d7ffu, 0x9a290b4du, 0xaf0227edu, 0x81b9c33cu, 0x5f352daeu, 0x46889c3du, 0x8ddc5d17u, 0x368b4a6cu, 0x16acb57au, 0x103f86c1u, 0x7144efafu, 0x87cd0dc1u, 0x1d9132b1u, 0x28b0d87cu, 0x01c44610u, 0x908b6bacu, 0x9e127820u, 0x9e695a6fu, 0xfbb80109u, 0x39ad9ba2u, 0x0c867f51u, 0xac1e69e5u, 0xf16f763du, 0x2de60ff0u, 0xfabaf858u, 0x53a35511u, 0x249e753fu, 0x88f84300u, 0xd79db283u, 0x7ac395cau, 0x5df23d8bu, 0xe3f98318u, 0x33ad1217u, 0x2dd19e3fu, 0xf1a97b54u, 0xb1fabd1eu, 0x1bc27b43u, 0xb208c2e5u, 0x1f109813u, 0xccb6a56bu, 0x64b0398eu, 0x9ac3815eu, 0x14970614u, 0x6139a686u, 0x6a5eeda3u, 0xbb567051u, 0xd2cd3e4du, 0x08c46214u, 0xc15f2c97u
|
||||
};
|
||||
// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^26 words.
|
||||
static const uint32_t IGNEUM_CACHE_HEAD[16] = {
|
||||
0xebd9055cu, 0x7eaf6b21u, 0x9b610855u, 0x134a8562u, 0xb10059bau, 0x7f459e58u, 0x8f3c7873u, 0x3523e609u,
|
||||
0x75b8f4cau, 0x750d31b4u, 0x0dae0782u, 0x26f10015u, 0x7ba87a41u, 0x0efd8543u, 0x691d6368u, 0x8929d967u
|
||||
};
|
||||
static const uint32_t IGNEUM_CACHE_LAST[16] = {
|
||||
0x51474edfu, 0xc3cfba93u, 0xf21454cfu, 0x9b79baacu, 0x4d4fce23u, 0xd701dfd5u, 0x37357ab3u, 0x1be693fau,
|
||||
0xa701cb3bu, 0x7467c620u, 0x428184e8u, 0xf4010df0u, 0xe33aa88fu, 0xc78d6d62u, 0xae9ba9c0u, 0xf4bba97eu
|
||||
};
|
||||
static const uint64_t IGNEUM_CACHE_FNV64 = 0x448274a57f508cbcull;
|
||||
36
proto-cuda/packs-ca2-era/era-0/vectors.json
Normal file
36
proto-cuda/packs-ca2-era/era-0/vectors.json
Normal file
|
|
@ -0,0 +1,36 @@
|
|||
{
|
||||
"seed": "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000",
|
||||
"day": "bytes:69676e65756d2d6461792ffa50000000000000",
|
||||
"dataset_mode": "memory-hard",
|
||||
"dataset_log2_words": 28,
|
||||
"mask": "0x0fffffff",
|
||||
"lanes": 32,
|
||||
"source": "igneum-pow (Rust) CPU interpreter, generator v3, memory-hard dataset",
|
||||
"warps": [
|
||||
{"base_nonce": 0, "expected": [
|
||||
"0xa32ad2dd06264089", "0x854fea26361763aa", "0xd958908d36ee63b5", "0x33641fe988c44ad9", "0xb5290b7892dbc7ba", "0xd815f329e6f806c0", "0x9fb28ff3fc29e39c", "0xdbc28a5d8bf42ab4",
|
||||
"0x06f955808ff71bb9", "0x1ec10fff7bcdeadd", "0x1acbc289cddd83cd", "0x395e1564cda10137", "0x8a2aac2aa5b190d2", "0x499b8401e1f3e412", "0x9e2320319c85fc43", "0x66ba9661bc421ad6",
|
||||
"0x3ba69971f1c9d8bd", "0x00d7a64d460bfefe", "0xc0f6487a1fbc5978", "0x1a261ac8673e1146", "0xf5a95a7caaaafa0a", "0x7bbdabd107842744", "0x49dd3bf9b77a16cf", "0xf7026cc441495a05",
|
||||
"0x49fc48c1cba2be91", "0x84c2fc11de378cf9", "0x96de9e40c9da551c", "0x3d20f01e363aaf61", "0xef36922a5fcf96ca", "0x31f1b0c4b0b42aaf", "0x33c6406d516d5992", "0x847af3a4248a972f"
|
||||
]},
|
||||
{"base_nonce": 4096, "expected": [
|
||||
"0x2039f40a61003341", "0x8238174df3761142", "0x32a72a04c4ca4a3c", "0x16aa0fd3f65da6b3", "0x502bf9db9bd98782", "0x109f60bc6167ff30", "0x3873c59792d6b552", "0xea98fba68fc32149",
|
||||
"0x56a8440169cb0f83", "0x7362edc3a7c98252", "0x8f393a7ad9f4ff40", "0x6b3cb3a1e0b02453", "0xdcbdc40aedfa06c0", "0x2bc381227147a2f3", "0x1570ef95b0c8e31f", "0xf3ca64d0f9ca8c91",
|
||||
"0x1c767a63cd6a0bf8", "0x8e1a13a1237a948b", "0xff2241b0b53b3efd", "0x95949c528978b1fb", "0xce3519499b78db7e", "0x56a0fcecaaa62096", "0x0a62c2bddc7041b2", "0x5ec463510d31d7c0",
|
||||
"0x2fc74273ed47b17f", "0xedc5740b0c5c191e", "0xe62f737c106216cb", "0x185df496808afd70", "0x29070790e64a0bc7", "0xc0ddaeac8b22b14b", "0x302982837b6d96b4", "0x523972727aa906b5"
|
||||
]},
|
||||
{"base_nonce": 1000000, "expected": [
|
||||
"0x3527292a5f4afb4e", "0x8c4128a03d946030", "0xb5b598dd18eb207d", "0x9f1b21b5bf9dce8f", "0xa791a53fa1a4733e", "0x12d2f9629aaca3b8", "0x400e844a00ba254c", "0x1e9f98bee0272859",
|
||||
"0x9983064542aabb54", "0x017873e17fab719c", "0xf3e747e327968d2b", "0x562380f8016b9ca7", "0x3f5007c1ec21b4a8", "0x5612489676e34c2b", "0x97e13b52bc125dab", "0xe15cf6d8d12f9b00",
|
||||
"0xd8ba508db7b12919", "0x2e16d66db0bd1837", "0x0026e228553d7b18", "0x0c6d22f5018958dd", "0xe76da566fc36580d", "0x7824813a77faff44", "0x663811c5504ff04d", "0x75efd5493353d65a",
|
||||
"0x926b72ebe963caa2", "0x392723945438b76d", "0x4df61663eb2bfe05", "0x9a305821ba9ceb59", "0x317f4a6ac3f65a88", "0xdbae9b884f6829a9", "0x97cc8c8866b1a0cb", "0x5db142b6d7777df1"
|
||||
]}
|
||||
],
|
||||
"dataset_head": ["0x3dd50b1f", "0x48edec90", "0x93123701", "0x3a7d2407", "0x5119540e", "0x657ac748", "0x3358ef2d", "0x09e2105c", "0xcdd88b45", "0xb03a0ac0", "0x2c37c6b0", "0x10e2c6a5", "0x25e2929c", "0x4317cabd", "0x5fa0cdcc", "0x6178ee52"],
|
||||
"dataset_last_index": 268435455,
|
||||
"dataset_last": "0xf7b7180e",
|
||||
"dataset_samples": [{"index": 59471966, "value": "0x1faef64a"}, {"index": 217795994, "value": "0x068e4a54"}, {"index": 208353206, "value": "0x24551c62"}, {"index": 42483309, "value": "0x1234f539"}, {"index": 172547758, "value": "0xc15100f5"}, {"index": 148076330, "value": "0x2f53b6b0"}, {"index": 183853158, "value": "0xbab5fb76"}, {"index": 214389424, "value": "0x06edf897"}, {"index": 267488061, "value": "0xb0880fc6"}, {"index": 169781097, "value": "0xccfddc4d"}, {"index": 184093494, "value": "0xf640c0f4"}, {"index": 153880993, "value": "0xf145a054"}, {"index": 84977930, "value": "0xa2425e8f"}, {"index": 46426879, "value": "0x0a218118"}, {"index": 3093825, "value": "0xd310512a"}, {"index": 225364072, "value": "0xb612d7ff"}, {"index": 44593546, "value": "0x9a290b4d"}, {"index": 260713159, "value": "0xaf0227ed"}, {"index": 168250303, "value": "0x81b9c33c"}, {"index": 52384140, "value": "0x5f352dae"}, {"index": 223401610, "value": "0x46889c3d"}, {"index": 45554030, "value": "0x8ddc5d17"}, {"index": 95410555, "value": "0x368b4a6c"}, {"index": 175039924, "value": "0x16acb57a"}, {"index": 79171087, "value": "0x103f86c1"}, {"index": 267580473, "value": "0x7144efaf"}, {"index": 24168642, "value": "0x87cd0dc1"}, {"index": 37981670, "value": "0x1d9132b1"}, {"index": 171551130, "value": "0x28b0d87c"}, {"index": 195559979, "value": "0x01c44610"}, {"index": 204611762, "value": "0x908b6bac"}, {"index": 140997658, "value": "0x9e127820"}, {"index": 138925853, "value": "0x9e695a6f"}, {"index": 86637313, "value": "0xfbb80109"}, {"index": 20736778, "value": "0x39ad9ba2"}, {"index": 219665210, "value": "0x0c867f51"}, {"index": 160430336, "value": "0xac1e69e5"}, {"index": 264654675, "value": "0xf16f763d"}, {"index": 8013395, "value": "0x2de60ff0"}, {"index": 228945585, "value": "0xfabaf858"}, {"index": 213884386, "value": "0x53a35511"}, {"index": 104419827, "value": "0x249e753f"}, {"index": 44185464, "value": "0x88f84300"}, {"index": 142737231, "value": "0xd79db283"}, {"index": 99284897, "value": "0x7ac395ca"}, {"index": 132475900, "value": "0x5df23d8b"}, {"index": 61861762, "value": "0xe3f98318"}, {"index": 132056166, "value": "0x33ad1217"}, {"index": 262388043, "value": "0x2dd19e3f"}, {"index": 91878046, "value": "0xf1a97b54"}, {"index": 117353561, "value": "0xb1fabd1e"}, {"index": 124768597, "value": "0x1bc27b43"}, {"index": 71352993, "value": "0xb208c2e5"}, {"index": 190698941, "value": "0x1f109813"}, {"index": 46055428, "value": "0xccb6a56b"}, {"index": 55281366, "value": "0x64b0398e"}, {"index": 165145231, "value": "0x9ac3815e"}, {"index": 106810753, "value": "0x14970614"}, {"index": 171985651, "value": "0x6139a686"}, {"index": 232085256, "value": "0x6a5eeda3"}, {"index": 159510492, "value": "0xbb567051"}, {"index": 40072060, "value": "0xd2cd3e4d"}, {"index": 209107596, "value": "0x08c46214"}, {"index": 39023794, "value": "0xc15f2c97"}],
|
||||
"cache_head": ["0xebd9055c", "0x7eaf6b21", "0x9b610855", "0x134a8562", "0xb10059ba", "0x7f459e58", "0x8f3c7873", "0x3523e609", "0x75b8f4ca", "0x750d31b4", "0x0dae0782", "0x26f10015", "0x7ba87a41", "0x0efd8543", "0x691d6368", "0x8929d967"],
|
||||
"cache_last_line": ["0x51474edf", "0xc3cfba93", "0xf21454cf", "0x9b79baac", "0x4d4fce23", "0xd701dfd5", "0x37357ab3", "0x1be693fa", "0xa701cb3b", "0x7467c620", "0x428184e8", "0xf4010df0", "0xe33aa88f", "0xc78d6d62", "0xae9ba9c0", "0xf4bba97e"],
|
||||
"cache_fnv1a64": "0x448274a57f508cbc"
|
||||
}
|
||||
281
proto-cuda/packs-ca2-era/era-1/kernel.cl
Normal file
281
proto-cuda/packs-ca2-era/era-1/kernel.cl
Normal file
|
|
@ -0,0 +1,281 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal).
|
||||
// Built from source at runtime by proto-opencl/host.c, which passes these defines:
|
||||
// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit)
|
||||
// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default)
|
||||
// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32
|
||||
// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition
|
||||
// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units;
|
||||
// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32.
|
||||
#ifndef IGNEUM_GROUP
|
||||
#define IGNEUM_GROUP 32
|
||||
#endif
|
||||
#ifndef IGNEUM_EXCHANGE
|
||||
#define IGNEUM_EXCHANGE 0
|
||||
#endif
|
||||
#ifdef __OPENCL_VERSION__
|
||||
#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1)))
|
||||
#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n]
|
||||
#if IGNEUM_EXCHANGE == 1
|
||||
#ifdef cl_khr_subgroups
|
||||
#pragma OPENCL EXTENSION cl_khr_subgroups : enable
|
||||
#endif
|
||||
#ifdef cl_khr_subgroup_shuffle
|
||||
#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable
|
||||
#endif
|
||||
#elif IGNEUM_EXCHANGE == 2
|
||||
#pragma OPENCL EXTENSION cl_intel_subgroups : enable
|
||||
#endif
|
||||
#else
|
||||
// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros.
|
||||
#include "emu_opencl.h"
|
||||
#endif
|
||||
|
||||
#if IGNEUM_EXCHANGE == 1
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m))
|
||||
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
|
||||
#elif IGNEUM_EXCHANGE == 2
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m))
|
||||
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
|
||||
#else
|
||||
// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per
|
||||
// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane
|
||||
// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's
|
||||
// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier.
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; }
|
||||
#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; }
|
||||
#endif
|
||||
|
||||
static inline uint splitmix32(uint x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32.
|
||||
static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); }
|
||||
// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x.
|
||||
static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); }
|
||||
static inline uint ds_elem(uint i, uint d0, uint d1) {
|
||||
uint x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
|
||||
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
|
||||
#define MH_CACHE_LINE_MASK 0x003fffffu
|
||||
#define MH_SEGMENT_LINES 64u
|
||||
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
|
||||
static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
|
||||
|
||||
// y = ChaCha12 core(x) + x
|
||||
static inline void mh_chacha_block(const uint* x, uint* y) {
|
||||
for (uint i = 0u; i < 16u; ++i) y[i] = x[i];
|
||||
for (uint r = 0u; r < 6u; ++r) {
|
||||
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
|
||||
}
|
||||
for (uint i = 0u; i < 16u; ++i) y[i] += x[i];
|
||||
}
|
||||
|
||||
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
|
||||
static inline void mh_cache_segment(__global uint* cache, uint seg) {
|
||||
uint prev[16]; uint x[16]; uint y[16];
|
||||
for (uint i = 0u; i < 16u; ++i) prev[i] = 0u;
|
||||
for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) {
|
||||
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
|
||||
x[4] = 0xceed56d7u ^ prev[4];
|
||||
x[5] = 0x9ba270d2u ^ prev[5];
|
||||
x[6] = 0x82caab2du ^ prev[6];
|
||||
x[7] = 0x81ebce0eu ^ prev[7];
|
||||
x[8] = 0x12b6ecf1u ^ prev[8];
|
||||
x[9] = 0xd0f3fd7cu ^ prev[9];
|
||||
x[10] = 0xd872eefeu ^ prev[10];
|
||||
x[11] = 0xc158c7bdu ^ prev[11];
|
||||
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
|
||||
mh_chacha_block(x, y);
|
||||
__global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
|
||||
}
|
||||
}
|
||||
|
||||
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
|
||||
static inline void mh_mixer(uint* s, uint rk) {
|
||||
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
|
||||
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
|
||||
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
|
||||
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
|
||||
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
|
||||
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
|
||||
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
|
||||
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
|
||||
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
|
||||
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
|
||||
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
|
||||
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
|
||||
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
|
||||
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
|
||||
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
|
||||
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
|
||||
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
|
||||
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
|
||||
}
|
||||
|
||||
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
|
||||
static inline void mh_item(__global const uint* cache, uint t, uint* s) {
|
||||
s[0] = 0xceed56d7u;
|
||||
s[1] = 0x9ba270d2u;
|
||||
s[2] = 0x82caab2du;
|
||||
s[3] = 0x81ebce0eu;
|
||||
s[4] = 0x12b6ecf1u;
|
||||
s[5] = 0xd0f3fd7cu;
|
||||
s[6] = 0xd872eefeu;
|
||||
s[7] = 0xc158c7bdu;
|
||||
s[8] = t * 0xf351d601u + 0xc6892460u;
|
||||
s[9] = t * 0xa3bb398fu + 0x25b7228au;
|
||||
s[10] = t * 0xb5a09e35u + 0xcd515004u;
|
||||
s[11] = t * 0x7509c9c1u + 0x2846527au;
|
||||
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
|
||||
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
|
||||
s[14] = t * 0xded91851u + 0x82961bacu;
|
||||
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
|
||||
for (uint r = 0u; r < 8u; ++r) {
|
||||
mh_mixer(s, 0x9E3779B9u * (r + 1u));
|
||||
__global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i];
|
||||
}
|
||||
mh_mixer(s, 0x9E3779B9u * 9u);
|
||||
}
|
||||
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 1 3 8 13 of w.
|
||||
static inline uint mh_j(uint w) { return ((w >> 1u) & 1u) | (((w >> 3u) & 1u) << 1) | (((w >> 8u) & 1u) << 2) | (((w >> 13u) & 1u) << 3); }
|
||||
static inline uint mh_t(uint w) { w = (w & 0x00001fffu) | ((w >> 14u) << 13u); w = (w & 0x000000ffu) | ((w >> 9u) << 8u); w = (w & 0x00000007u) | ((w >> 4u) << 3u); w = (w & 0x00000001u) | ((w >> 2u) << 1u); return w; }
|
||||
static inline uint mh_addr(uint t, uint j) { uint w = t; w = ((w >> 1u) << 2u) | (w & 0x00000001u) | (((j >> 0u) & 1u) << 1u); w = ((w >> 3u) << 4u) | (w & 0x00000007u) | (((j >> 1u) & 1u) << 3u); w = ((w >> 8u) << 9u) | (w & 0x000000ffu) | (((j >> 2u) & 1u) << 8u); w = ((w >> 13u) << 14u) | (w & 0x00001fffu) | (((j >> 3u) & 1u) << 13u); return w; }
|
||||
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
|
||||
static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }
|
||||
|
||||
// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item.
|
||||
// The same constants as memhard.h in this pack (one emitter, three dialects).
|
||||
__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) {
|
||||
uint seg = (uint)get_global_id(0);
|
||||
if (seg < nSegments) mh_cache_segment(cache, seg);
|
||||
}
|
||||
__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) {
|
||||
uint t = (uint)get_global_id(0);
|
||||
if (t < nItems) {
|
||||
uint s[16];
|
||||
mh_item(cache, t, s);
|
||||
for (uint i = 0u; i < 16u; ++i) ds[(ulong)mh_addr(t, i)] = s[i];
|
||||
}
|
||||
}
|
||||
|
||||
// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the
|
||||
// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and
|
||||
// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all).
|
||||
IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) {
|
||||
uint gid = (uint)get_global_id(0);
|
||||
uint lid = (uint)get_local_id(0);
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
#if IGNEUM_EXCHANGE == 0
|
||||
IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
|
||||
uint xk = 0u;
|
||||
#else
|
||||
(void)lid;
|
||||
#endif
|
||||
{ uint x = nonce ^ 0x667d0fbdu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x7b8e5963u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
|
||||
{ uint x = nonce ^ 0x7b8e5963u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x31c67e5eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
|
||||
{ uint x = nonce ^ 0x31c67e5eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4529ddc6u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
|
||||
{ uint x = nonce ^ 0x4529ddc6u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef19d6d8u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
|
||||
{ uint x = nonce ^ 0xef19d6d8u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xaccf6211u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
|
||||
{ uint x = nonce ^ 0xaccf6211u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xda0aed32u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
|
||||
{ uint x = nonce ^ 0xda0aed32u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xabc6df31u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
|
||||
{ uint x = nonce ^ 0xabc6df31u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x667d0fbdu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
|
||||
r7 = r7 ^ r0; // 1 xor
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 1u); r3 = r3 ^ t_; } // 2 shfl
|
||||
r4 = r4 - r1; // 3 sub
|
||||
r2 = r0 * r4 + r2; // 4 mad
|
||||
r4 = rotl_imm(r4, 9u); // 5 rotl
|
||||
r0 = rotr_var(r0, r2); // 6 rotr
|
||||
r6 = r6 ^ ds[((rotl_imm(r7 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r2 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r0 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r1 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
|
||||
r5 = r5 * r4; // 12 mul
|
||||
r7 = r7 ^ ds[((rotl_imm(r6 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 16u); r6 = r6 ^ t_; } // 14 shfl
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r3 = r3 ^ t_; } // 15 shfl
|
||||
r2 = r2 | r7; // 16 or
|
||||
r0 = rotr_var(r0, r6); // 17 rotr
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
|
||||
r7 = rotl_imm(r7, 24u); // 19 rotl
|
||||
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
|
||||
r2 = r2 | r6; // 21 or
|
||||
r6 = r5 * r4 + r6; // 22 mad
|
||||
r1 = r1 | r0; // 23 or
|
||||
r6 = r6 ^ r1; // 24 xor
|
||||
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
|
||||
r7 = rotr_var(r7, r0); // 26 rotr
|
||||
r4 = r5 * r7 + r4; // 27 mad
|
||||
r3 = r2 * r4 + r3; // 28 mad
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
|
||||
r2 = r2 ^ ds[((rotl_imm(r3 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r5 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
|
||||
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
|
||||
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
|
||||
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
|
||||
r2 = mul_hi(r2, r5); // 35 mulhi
|
||||
r4 = r4 ^ r2; // 36 xor
|
||||
r6 = r6 * r5; // 37 mul
|
||||
r7 = r7 ^ r0; // 38 xor
|
||||
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
|
||||
r2 = rotr_var(r2, r3); // 40 rotr
|
||||
r7 = r7 - r0; // 41 sub
|
||||
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 43 shfl
|
||||
r0 = r0 ^ ds[((rotl_imm(r7 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
|
||||
r3 = r3 ^ ds[((rotl_imm(r1 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
|
||||
r6 = r6 ^ ds[((rotl_imm(r3 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
|
||||
r4 = mul_hi(r4, r2); // 48 mulhi
|
||||
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r0 = r0 ^ t_; } // 50 shfl
|
||||
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r2 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
|
||||
r6 = rotl_imm(r6, 12u); // 53 rotl
|
||||
r3 = rotl_imm(r3, 12u); // 54 rotl
|
||||
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
|
||||
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r0 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
|
||||
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
|
||||
r3 = mul_hi(r3, r2); // 60 mulhi
|
||||
r6 = r6 | r4; // 61 or
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r5 = r5 ^ t_; } // 62 shfl
|
||||
r3 = r3 ^ ds[((rotl_imm(r6 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
|
||||
#if IGNEUM_EXCHANGE != 0
|
||||
// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the
|
||||
// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a
|
||||
// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md.
|
||||
IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) {
|
||||
if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); }
|
||||
}
|
||||
#endif
|
||||
163
proto-cuda/packs-ca2-era/era-1/kernel.cu
Normal file
163
proto-cuda/packs-ca2-era/era-1/kernel.cu
Normal file
|
|
@ -0,0 +1,163 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal).
|
||||
// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC.
|
||||
#include <cuda_runtime.h>
|
||||
#include <cstdint>
|
||||
#include "program.h"
|
||||
#include "memhard.h"
|
||||
|
||||
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31.
|
||||
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
|
||||
// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x.
|
||||
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) {
|
||||
uint32_t x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item.
|
||||
// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host.
|
||||
__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) {
|
||||
uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
if (seg < nSegments) mh_cache_segment(cache, seg);
|
||||
}
|
||||
__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
|
||||
uint32_t t = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
if (t < nItems) {
|
||||
uint32_t s[16];
|
||||
mh_item(cache, t, s);
|
||||
for (uint32_t i = 0u; i < 16u; ++i) ds[(size_t)mh_addr(t, i)] = s[i];
|
||||
}
|
||||
}
|
||||
|
||||
// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every
|
||||
// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a
|
||||
// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid.
|
||||
__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) {
|
||||
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
uint32_t nonce = baseNonce + gid;
|
||||
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint32_t x = nonce ^ 0x667d0fbdu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x7b8e5963u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
|
||||
{ uint32_t x = nonce ^ 0x7b8e5963u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x31c67e5eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
|
||||
{ uint32_t x = nonce ^ 0x31c67e5eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4529ddc6u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
|
||||
{ uint32_t x = nonce ^ 0x4529ddc6u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef19d6d8u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
|
||||
{ uint32_t x = nonce ^ 0xef19d6d8u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xaccf6211u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
|
||||
{ uint32_t x = nonce ^ 0xaccf6211u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xda0aed32u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
|
||||
{ uint32_t x = nonce ^ 0xda0aed32u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xabc6df31u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
|
||||
{ uint32_t x = nonce ^ 0xabc6df31u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x667d0fbdu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
|
||||
|
||||
for (uint32_t it = 0u; it < 8u; ++it) {
|
||||
uint32_t sel = r0;
|
||||
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
|
||||
r7 = r7 ^ r0; // 1 xor
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r6, 1); // 2 shfl
|
||||
r4 = r4 - r1; // 3 sub
|
||||
r2 = r0 * r4 + r2; // 4 mad
|
||||
r4 = rotl_imm(r4, 9u); // 5 rotl
|
||||
r0 = rotr_var(r0, r2); // 6 rotr
|
||||
r6 = r6 ^ ds[((rotl_imm(r7 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r2 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r0 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r1 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
|
||||
r5 = r5 * r4; // 12 mul
|
||||
r7 = r7 ^ ds[((rotl_imm(r6 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
|
||||
r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r5, 16); // 14 shfl
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r5, 1); // 15 shfl
|
||||
r2 = r2 | r7; // 16 or
|
||||
r0 = rotr_var(r0, r6); // 17 rotr
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
|
||||
r7 = rotl_imm(r7, 24u); // 19 rotl
|
||||
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
|
||||
r2 = r2 | r6; // 21 or
|
||||
r6 = r5 * r4 + r6; // 22 mad
|
||||
r1 = r1 | r0; // 23 or
|
||||
r6 = r6 ^ r1; // 24 xor
|
||||
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
|
||||
r7 = rotr_var(r7, r0); // 26 rotr
|
||||
r4 = r5 * r7 + r4; // 27 mad
|
||||
r3 = r2 * r4 + r3; // 28 mad
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
|
||||
r2 = r2 ^ ds[((rotl_imm(r3 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r5 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
|
||||
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
|
||||
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
|
||||
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
|
||||
r2 = __umulhi(r2, r5); // 35 mulhi
|
||||
r4 = r4 ^ r2; // 36 xor
|
||||
r6 = r6 * r5; // 37 mul
|
||||
r7 = r7 ^ r0; // 38 xor
|
||||
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
|
||||
r2 = rotr_var(r2, r3); // 40 rotr
|
||||
r7 = r7 - r0; // 41 sub
|
||||
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 43 shfl
|
||||
r0 = r0 ^ ds[((rotl_imm(r7 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
|
||||
r3 = r3 ^ ds[((rotl_imm(r1 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
|
||||
r6 = r6 ^ ds[((rotl_imm(r3 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
|
||||
r4 = __umulhi(r4, r2); // 48 mulhi
|
||||
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
|
||||
r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r7, 4); // 50 shfl
|
||||
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r2 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
|
||||
r6 = rotl_imm(r6, 12u); // 53 rotl
|
||||
r3 = rotl_imm(r3, 12u); // 54 rotl
|
||||
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
|
||||
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r0 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
|
||||
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
|
||||
r3 = __umulhi(r3, r2); // 60 mulhi
|
||||
r6 = r6 | r4; // 61 or
|
||||
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 62 shfl
|
||||
r3 = r3 ^ ds[((rotl_imm(r6 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
|
||||
}
|
||||
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
|
||||
}
|
||||
|
||||
// Host-side launch wrappers. Declared in program.h, called from host.cu.
|
||||
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) {
|
||||
if (nSegments == 0u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 256u;
|
||||
uint32_t grid = (nSegments + block - 1u) / block;
|
||||
igneum_cache_fill<<<grid, block>>>(cache, nSegments);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
|
||||
if (nItems == 0u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 256u;
|
||||
uint32_t grid = (nItems + block - 1u) / block;
|
||||
igneum_build<<<grid, block>>>(ds, cache, nItems);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
uint32_t nonces, uint32_t blockWarps) {
|
||||
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 32u * blockWarps;
|
||||
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
|
||||
igneum_hash<<<nonces / block, block>>>(ds, out, baseNonce, mask);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
|
||||
cudaFuncAttributes attr;
|
||||
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash);
|
||||
if (e != cudaSuccess) return e;
|
||||
*numRegs = attr.numRegs;
|
||||
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0);
|
||||
}
|
||||
375
proto-cuda/packs-ca2-era/era-1/kernel_bound.cl
Normal file
375
proto-cuda/packs-ca2-era/era-1/kernel_bound.cl
Normal file
|
|
@ -0,0 +1,375 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal).
|
||||
// Built from source at runtime by proto-opencl/host.c, which passes these defines:
|
||||
// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit)
|
||||
// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default)
|
||||
// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32
|
||||
// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition
|
||||
// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units;
|
||||
// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32.
|
||||
#ifndef IGNEUM_GROUP
|
||||
#define IGNEUM_GROUP 32
|
||||
#endif
|
||||
#ifndef IGNEUM_EXCHANGE
|
||||
#define IGNEUM_EXCHANGE 0
|
||||
#endif
|
||||
#ifdef __OPENCL_VERSION__
|
||||
#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1)))
|
||||
#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n]
|
||||
#if IGNEUM_EXCHANGE == 1
|
||||
#ifdef cl_khr_subgroups
|
||||
#pragma OPENCL EXTENSION cl_khr_subgroups : enable
|
||||
#endif
|
||||
#ifdef cl_khr_subgroup_shuffle
|
||||
#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable
|
||||
#endif
|
||||
#elif IGNEUM_EXCHANGE == 2
|
||||
#pragma OPENCL EXTENSION cl_intel_subgroups : enable
|
||||
#endif
|
||||
#else
|
||||
// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros.
|
||||
#include "emu_opencl.h"
|
||||
#endif
|
||||
|
||||
#if IGNEUM_EXCHANGE == 1
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m))
|
||||
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
|
||||
#elif IGNEUM_EXCHANGE == 2
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m))
|
||||
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
|
||||
#else
|
||||
// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per
|
||||
// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane
|
||||
// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's
|
||||
// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier.
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; }
|
||||
#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; }
|
||||
#endif
|
||||
|
||||
static inline uint splitmix32(uint x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32.
|
||||
static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); }
|
||||
// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x.
|
||||
static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); }
|
||||
static inline uint ds_elem(uint i, uint d0, uint d1) {
|
||||
uint x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
|
||||
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
|
||||
#define MH_CACHE_LINE_MASK 0x003fffffu
|
||||
#define MH_SEGMENT_LINES 64u
|
||||
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
|
||||
static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
|
||||
|
||||
// y = ChaCha12 core(x) + x
|
||||
static inline void mh_chacha_block(const uint* x, uint* y) {
|
||||
for (uint i = 0u; i < 16u; ++i) y[i] = x[i];
|
||||
for (uint r = 0u; r < 6u; ++r) {
|
||||
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
|
||||
}
|
||||
for (uint i = 0u; i < 16u; ++i) y[i] += x[i];
|
||||
}
|
||||
|
||||
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
|
||||
static inline void mh_cache_segment(__global uint* cache, uint seg) {
|
||||
uint prev[16]; uint x[16]; uint y[16];
|
||||
for (uint i = 0u; i < 16u; ++i) prev[i] = 0u;
|
||||
for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) {
|
||||
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
|
||||
x[4] = 0xceed56d7u ^ prev[4];
|
||||
x[5] = 0x9ba270d2u ^ prev[5];
|
||||
x[6] = 0x82caab2du ^ prev[6];
|
||||
x[7] = 0x81ebce0eu ^ prev[7];
|
||||
x[8] = 0x12b6ecf1u ^ prev[8];
|
||||
x[9] = 0xd0f3fd7cu ^ prev[9];
|
||||
x[10] = 0xd872eefeu ^ prev[10];
|
||||
x[11] = 0xc158c7bdu ^ prev[11];
|
||||
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
|
||||
mh_chacha_block(x, y);
|
||||
__global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
|
||||
}
|
||||
}
|
||||
|
||||
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
|
||||
static inline void mh_mixer(uint* s, uint rk) {
|
||||
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
|
||||
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
|
||||
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
|
||||
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
|
||||
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
|
||||
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
|
||||
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
|
||||
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
|
||||
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
|
||||
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
|
||||
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
|
||||
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
|
||||
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
|
||||
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
|
||||
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
|
||||
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
|
||||
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
|
||||
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
|
||||
}
|
||||
|
||||
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
|
||||
static inline void mh_item(__global const uint* cache, uint t, uint* s) {
|
||||
s[0] = 0xceed56d7u;
|
||||
s[1] = 0x9ba270d2u;
|
||||
s[2] = 0x82caab2du;
|
||||
s[3] = 0x81ebce0eu;
|
||||
s[4] = 0x12b6ecf1u;
|
||||
s[5] = 0xd0f3fd7cu;
|
||||
s[6] = 0xd872eefeu;
|
||||
s[7] = 0xc158c7bdu;
|
||||
s[8] = t * 0xf351d601u + 0xc6892460u;
|
||||
s[9] = t * 0xa3bb398fu + 0x25b7228au;
|
||||
s[10] = t * 0xb5a09e35u + 0xcd515004u;
|
||||
s[11] = t * 0x7509c9c1u + 0x2846527au;
|
||||
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
|
||||
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
|
||||
s[14] = t * 0xded91851u + 0x82961bacu;
|
||||
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
|
||||
for (uint r = 0u; r < 8u; ++r) {
|
||||
mh_mixer(s, 0x9E3779B9u * (r + 1u));
|
||||
__global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i];
|
||||
}
|
||||
mh_mixer(s, 0x9E3779B9u * 9u);
|
||||
}
|
||||
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 1 3 8 13 of w.
|
||||
static inline uint mh_j(uint w) { return ((w >> 1u) & 1u) | (((w >> 3u) & 1u) << 1) | (((w >> 8u) & 1u) << 2) | (((w >> 13u) & 1u) << 3); }
|
||||
static inline uint mh_t(uint w) { w = (w & 0x00001fffu) | ((w >> 14u) << 13u); w = (w & 0x000000ffu) | ((w >> 9u) << 8u); w = (w & 0x00000007u) | ((w >> 4u) << 3u); w = (w & 0x00000001u) | ((w >> 2u) << 1u); return w; }
|
||||
static inline uint mh_addr(uint t, uint j) { uint w = t; w = ((w >> 1u) << 2u) | (w & 0x00000001u) | (((j >> 0u) & 1u) << 1u); w = ((w >> 3u) << 4u) | (w & 0x00000007u) | (((j >> 1u) & 1u) << 3u); w = ((w >> 8u) << 9u) | (w & 0x000000ffu) | (((j >> 2u) & 1u) << 8u); w = ((w >> 13u) << 14u) | (w & 0x00001fffu) | (((j >> 3u) & 1u) << 13u); return w; }
|
||||
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
|
||||
static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }
|
||||
|
||||
// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item.
|
||||
// The same constants as memhard.h in this pack (one emitter, three dialects).
|
||||
__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) {
|
||||
uint seg = (uint)get_global_id(0);
|
||||
if (seg < nSegments) mh_cache_segment(cache, seg);
|
||||
}
|
||||
__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) {
|
||||
uint t = (uint)get_global_id(0);
|
||||
if (t < nItems) {
|
||||
uint s[16];
|
||||
mh_item(cache, t, s);
|
||||
for (uint i = 0u; i < 16u; ++i) ds[(ulong)mh_addr(t, i)] = s[i];
|
||||
}
|
||||
}
|
||||
|
||||
// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the
|
||||
// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and
|
||||
// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all).
|
||||
IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) {
|
||||
uint gid = (uint)get_global_id(0);
|
||||
uint lid = (uint)get_local_id(0);
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
#if IGNEUM_EXCHANGE == 0
|
||||
IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
|
||||
uint xk = 0u;
|
||||
#else
|
||||
(void)lid;
|
||||
#endif
|
||||
{ uint x = nonce ^ 0x667d0fbdu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x7b8e5963u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
|
||||
{ uint x = nonce ^ 0x7b8e5963u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x31c67e5eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
|
||||
{ uint x = nonce ^ 0x31c67e5eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4529ddc6u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
|
||||
{ uint x = nonce ^ 0x4529ddc6u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef19d6d8u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
|
||||
{ uint x = nonce ^ 0xef19d6d8u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xaccf6211u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
|
||||
{ uint x = nonce ^ 0xaccf6211u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xda0aed32u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
|
||||
{ uint x = nonce ^ 0xda0aed32u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xabc6df31u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
|
||||
{ uint x = nonce ^ 0xabc6df31u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x667d0fbdu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
|
||||
r7 = r7 ^ r0; // 1 xor
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 1u); r3 = r3 ^ t_; } // 2 shfl
|
||||
r4 = r4 - r1; // 3 sub
|
||||
r2 = r0 * r4 + r2; // 4 mad
|
||||
r4 = rotl_imm(r4, 9u); // 5 rotl
|
||||
r0 = rotr_var(r0, r2); // 6 rotr
|
||||
r6 = r6 ^ ds[((rotl_imm(r7 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r2 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r0 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r1 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
|
||||
r5 = r5 * r4; // 12 mul
|
||||
r7 = r7 ^ ds[((rotl_imm(r6 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 16u); r6 = r6 ^ t_; } // 14 shfl
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r3 = r3 ^ t_; } // 15 shfl
|
||||
r2 = r2 | r7; // 16 or
|
||||
r0 = rotr_var(r0, r6); // 17 rotr
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
|
||||
r7 = rotl_imm(r7, 24u); // 19 rotl
|
||||
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
|
||||
r2 = r2 | r6; // 21 or
|
||||
r6 = r5 * r4 + r6; // 22 mad
|
||||
r1 = r1 | r0; // 23 or
|
||||
r6 = r6 ^ r1; // 24 xor
|
||||
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
|
||||
r7 = rotr_var(r7, r0); // 26 rotr
|
||||
r4 = r5 * r7 + r4; // 27 mad
|
||||
r3 = r2 * r4 + r3; // 28 mad
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
|
||||
r2 = r2 ^ ds[((rotl_imm(r3 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r5 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
|
||||
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
|
||||
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
|
||||
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
|
||||
r2 = mul_hi(r2, r5); // 35 mulhi
|
||||
r4 = r4 ^ r2; // 36 xor
|
||||
r6 = r6 * r5; // 37 mul
|
||||
r7 = r7 ^ r0; // 38 xor
|
||||
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
|
||||
r2 = rotr_var(r2, r3); // 40 rotr
|
||||
r7 = r7 - r0; // 41 sub
|
||||
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 43 shfl
|
||||
r0 = r0 ^ ds[((rotl_imm(r7 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
|
||||
r3 = r3 ^ ds[((rotl_imm(r1 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
|
||||
r6 = r6 ^ ds[((rotl_imm(r3 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
|
||||
r4 = mul_hi(r4, r2); // 48 mulhi
|
||||
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r0 = r0 ^ t_; } // 50 shfl
|
||||
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r2 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
|
||||
r6 = rotl_imm(r6, 12u); // 53 rotl
|
||||
r3 = rotl_imm(r3, 12u); // 54 rotl
|
||||
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
|
||||
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r0 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
|
||||
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
|
||||
r3 = mul_hi(r3, r2); // 60 mulhi
|
||||
r6 = r6 | r4; // 61 or
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r5 = r5 ^ t_; } // 62 shfl
|
||||
r3 = r3 ^ ds[((rotl_imm(r6 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
|
||||
#if IGNEUM_EXCHANGE != 0
|
||||
// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the
|
||||
// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a
|
||||
// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md.
|
||||
IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) {
|
||||
if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); }
|
||||
}
|
||||
#endif
|
||||
|
||||
// Header-bound variant (bind.rs): the init words come from initw, not SEEDW. Same body as igneum_hash.
|
||||
IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw) {
|
||||
uint gid = (uint)get_global_id(0);
|
||||
uint lid = (uint)get_local_id(0);
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7];
|
||||
#if IGNEUM_EXCHANGE == 0
|
||||
IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
|
||||
uint xk = 0u;
|
||||
#else
|
||||
(void)lid;
|
||||
#endif
|
||||
{ uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; }
|
||||
{ uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; }
|
||||
{ uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; }
|
||||
{ uint x = nonce ^ iw3; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw4; }
|
||||
{ uint x = nonce ^ iw4; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw5; }
|
||||
{ uint x = nonce ^ iw5; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw6; }
|
||||
{ uint x = nonce ^ iw6; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw7; }
|
||||
{ uint x = nonce ^ iw7; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw0; }
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
|
||||
r7 = r7 ^ r0; // 1 xor
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 1u); r3 = r3 ^ t_; } // 2 shfl
|
||||
r4 = r4 - r1; // 3 sub
|
||||
r2 = r0 * r4 + r2; // 4 mad
|
||||
r4 = rotl_imm(r4, 9u); // 5 rotl
|
||||
r0 = rotr_var(r0, r2); // 6 rotr
|
||||
r6 = r6 ^ ds[((rotl_imm(r7 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r2 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r0 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r1 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
|
||||
r5 = r5 * r4; // 12 mul
|
||||
r7 = r7 ^ ds[((rotl_imm(r6 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 16u); r6 = r6 ^ t_; } // 14 shfl
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r3 = r3 ^ t_; } // 15 shfl
|
||||
r2 = r2 | r7; // 16 or
|
||||
r0 = rotr_var(r0, r6); // 17 rotr
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
|
||||
r7 = rotl_imm(r7, 24u); // 19 rotl
|
||||
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
|
||||
r2 = r2 | r6; // 21 or
|
||||
r6 = r5 * r4 + r6; // 22 mad
|
||||
r1 = r1 | r0; // 23 or
|
||||
r6 = r6 ^ r1; // 24 xor
|
||||
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
|
||||
r7 = rotr_var(r7, r0); // 26 rotr
|
||||
r4 = r5 * r7 + r4; // 27 mad
|
||||
r3 = r2 * r4 + r3; // 28 mad
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
|
||||
r2 = r2 ^ ds[((rotl_imm(r3 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r5 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
|
||||
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
|
||||
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
|
||||
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
|
||||
r2 = mul_hi(r2, r5); // 35 mulhi
|
||||
r4 = r4 ^ r2; // 36 xor
|
||||
r6 = r6 * r5; // 37 mul
|
||||
r7 = r7 ^ r0; // 38 xor
|
||||
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
|
||||
r2 = rotr_var(r2, r3); // 40 rotr
|
||||
r7 = r7 - r0; // 41 sub
|
||||
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 43 shfl
|
||||
r0 = r0 ^ ds[((rotl_imm(r7 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
|
||||
r3 = r3 ^ ds[((rotl_imm(r1 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
|
||||
r6 = r6 ^ ds[((rotl_imm(r3 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
|
||||
r4 = mul_hi(r4, r2); // 48 mulhi
|
||||
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r0 = r0 ^ t_; } // 50 shfl
|
||||
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r2 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
|
||||
r6 = rotl_imm(r6, 12u); // 53 rotl
|
||||
r3 = rotl_imm(r3, 12u); // 54 rotl
|
||||
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
|
||||
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r0 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
|
||||
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
|
||||
r3 = mul_hi(r3, r2); // 60 mulhi
|
||||
r6 = r6 | r4; // 61 or
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r5 = r5 ^ t_; } // 62 shfl
|
||||
r3 = r3 ^ ds[((rotl_imm(r6 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
123
proto-cuda/packs-ca2-era/era-1/kernel_bound.cu
Normal file
123
proto-cuda/packs-ca2-era/era-1/kernel_bound.cu
Normal file
|
|
@ -0,0 +1,123 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// Header-bound twin of igneum_hash in kernel.cu: the init words come from a kernel argument, not SEEDW.
|
||||
// Host declarations (also in program_bound.h if present):
|
||||
// struct IgneumInitWords { uint32_t w[8]; };
|
||||
// cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
// IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps);
|
||||
// cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
|
||||
#include <cuda_runtime.h>
|
||||
#include <cstdint>
|
||||
#include "program.h"
|
||||
|
||||
struct IgneumInitWords { uint32_t w[8]; };
|
||||
|
||||
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
|
||||
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
|
||||
__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw) {
|
||||
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
uint32_t nonce = baseNonce + gid;
|
||||
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint32_t x = nonce ^ iw.w[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw.w[1]; }
|
||||
{ uint32_t x = nonce ^ iw.w[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw.w[2]; }
|
||||
{ uint32_t x = nonce ^ iw.w[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw.w[3]; }
|
||||
{ uint32_t x = nonce ^ iw.w[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw.w[4]; }
|
||||
{ uint32_t x = nonce ^ iw.w[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw.w[5]; }
|
||||
{ uint32_t x = nonce ^ iw.w[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw.w[6]; }
|
||||
{ uint32_t x = nonce ^ iw.w[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw.w[7]; }
|
||||
{ uint32_t x = nonce ^ iw.w[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw.w[0]; }
|
||||
|
||||
for (uint32_t it = 0u; it < 8u; ++it) {
|
||||
uint32_t sel = r0;
|
||||
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
|
||||
r7 = r7 ^ r0; // 1 xor
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r6, 1); // 2 shfl
|
||||
r4 = r4 - r1; // 3 sub
|
||||
r2 = r0 * r4 + r2; // 4 mad
|
||||
r4 = rotl_imm(r4, 9u); // 5 rotl
|
||||
r0 = rotr_var(r0, r2); // 6 rotr
|
||||
r6 = r6 ^ ds[((rotl_imm(r7 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r2 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r0 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r1 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
|
||||
r5 = r5 * r4; // 12 mul
|
||||
r7 = r7 ^ ds[((rotl_imm(r6 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
|
||||
r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r5, 16); // 14 shfl
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r5, 1); // 15 shfl
|
||||
r2 = r2 | r7; // 16 or
|
||||
r0 = rotr_var(r0, r6); // 17 rotr
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
|
||||
r7 = rotl_imm(r7, 24u); // 19 rotl
|
||||
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
|
||||
r2 = r2 | r6; // 21 or
|
||||
r6 = r5 * r4 + r6; // 22 mad
|
||||
r1 = r1 | r0; // 23 or
|
||||
r6 = r6 ^ r1; // 24 xor
|
||||
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
|
||||
r7 = rotr_var(r7, r0); // 26 rotr
|
||||
r4 = r5 * r7 + r4; // 27 mad
|
||||
r3 = r2 * r4 + r3; // 28 mad
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
|
||||
r2 = r2 ^ ds[((rotl_imm(r3 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r5 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
|
||||
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
|
||||
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
|
||||
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
|
||||
r2 = __umulhi(r2, r5); // 35 mulhi
|
||||
r4 = r4 ^ r2; // 36 xor
|
||||
r6 = r6 * r5; // 37 mul
|
||||
r7 = r7 ^ r0; // 38 xor
|
||||
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
|
||||
r2 = rotr_var(r2, r3); // 40 rotr
|
||||
r7 = r7 - r0; // 41 sub
|
||||
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 43 shfl
|
||||
r0 = r0 ^ ds[((rotl_imm(r7 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
|
||||
r3 = r3 ^ ds[((rotl_imm(r1 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
|
||||
r6 = r6 ^ ds[((rotl_imm(r3 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
|
||||
r4 = __umulhi(r4, r2); // 48 mulhi
|
||||
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
|
||||
r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r7, 4); // 50 shfl
|
||||
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r2 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
|
||||
r6 = rotl_imm(r6, 12u); // 53 rotl
|
||||
r3 = rotl_imm(r3, 12u); // 54 rotl
|
||||
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
|
||||
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r0 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
|
||||
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
|
||||
r3 = __umulhi(r3, r2); // 60 mulhi
|
||||
r6 = r6 | r4; // 61 or
|
||||
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 62 shfl
|
||||
r3 = r3 ^ ds[((rotl_imm(r6 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
|
||||
}
|
||||
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps) {
|
||||
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 32u * blockWarps;
|
||||
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
|
||||
igneum_hash_bound<<<nonces / block, block>>>(ds, out, baseNonce, mask, iw);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
|
||||
cudaFuncAttributes attr;
|
||||
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash_bound);
|
||||
if (e != cudaSuccess) return e;
|
||||
*numRegs = attr.numRegs;
|
||||
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash_bound, (int)(32u * blockWarps), 0);
|
||||
}
|
||||
112
proto-cuda/packs-ca2-era/era-1/memhard.h
Normal file
112
proto-cuda/packs-ca2-era/era-1/memhard.h
Normal file
|
|
@ -0,0 +1,112 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// Memory-hard dataset core, the same text that the Mac's Metal kernels and CPU verifier were checked against.
|
||||
// Included by kernel.cu (device), host.cu (host reference) and proto-opencl/host.c (C99 host reference).
|
||||
// See proto-metal/MEMHARD.md for the construction. kernel.cl carries the same text in OpenCL C.
|
||||
#pragma once
|
||||
#ifdef __cplusplus
|
||||
#include <cstdint>
|
||||
#else
|
||||
#include <stdint.h>
|
||||
#endif
|
||||
#if defined(__CUDACC__)
|
||||
#define IGNEUM_HD __host__ __device__ __forceinline__
|
||||
#elif defined(_MSC_VER) && !defined(__cplusplus)
|
||||
#define IGNEUM_HD static __inline
|
||||
#else
|
||||
#define IGNEUM_HD static inline
|
||||
#endif
|
||||
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
|
||||
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
|
||||
#define MH_CACHE_LINE_MASK 0x003fffffu
|
||||
#define MH_SEGMENT_LINES 64u
|
||||
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
|
||||
IGNEUM_HD uint32_t mh_rotl(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
|
||||
|
||||
// y = ChaCha12 core(x) + x
|
||||
IGNEUM_HD void mh_chacha_block(const uint32_t* x, uint32_t* y) {
|
||||
for (uint32_t i = 0u; i < 16u; ++i) y[i] = x[i];
|
||||
for (uint32_t r = 0u; r < 6u; ++r) {
|
||||
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
|
||||
}
|
||||
for (uint32_t i = 0u; i < 16u; ++i) y[i] += x[i];
|
||||
}
|
||||
|
||||
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
|
||||
IGNEUM_HD void mh_cache_segment(uint32_t* cache, uint32_t seg) {
|
||||
uint32_t prev[16]; uint32_t x[16]; uint32_t y[16];
|
||||
for (uint32_t i = 0u; i < 16u; ++i) prev[i] = 0u;
|
||||
for (uint32_t j = 0u; j < MH_SEGMENT_LINES; ++j) {
|
||||
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
|
||||
x[4] = 0xceed56d7u ^ prev[4];
|
||||
x[5] = 0x9ba270d2u ^ prev[5];
|
||||
x[6] = 0x82caab2du ^ prev[6];
|
||||
x[7] = 0x81ebce0eu ^ prev[7];
|
||||
x[8] = 0x12b6ecf1u ^ prev[8];
|
||||
x[9] = 0xd0f3fd7cu ^ prev[9];
|
||||
x[10] = 0xd872eefeu ^ prev[10];
|
||||
x[11] = 0xc158c7bdu ^ prev[11];
|
||||
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
|
||||
mh_chacha_block(x, y);
|
||||
uint32_t* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
|
||||
for (uint32_t i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
|
||||
}
|
||||
}
|
||||
|
||||
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
|
||||
IGNEUM_HD void mh_mixer(uint32_t* s, uint32_t rk) {
|
||||
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
|
||||
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
|
||||
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
|
||||
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
|
||||
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
|
||||
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
|
||||
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
|
||||
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
|
||||
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
|
||||
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
|
||||
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
|
||||
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
|
||||
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
|
||||
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
|
||||
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
|
||||
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
|
||||
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
|
||||
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
|
||||
}
|
||||
|
||||
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
|
||||
IGNEUM_HD void mh_item(const uint32_t* cache, uint32_t t, uint32_t* s) {
|
||||
s[0] = 0xceed56d7u;
|
||||
s[1] = 0x9ba270d2u;
|
||||
s[2] = 0x82caab2du;
|
||||
s[3] = 0x81ebce0eu;
|
||||
s[4] = 0x12b6ecf1u;
|
||||
s[5] = 0xd0f3fd7cu;
|
||||
s[6] = 0xd872eefeu;
|
||||
s[7] = 0xc158c7bdu;
|
||||
s[8] = t * 0xf351d601u + 0xc6892460u;
|
||||
s[9] = t * 0xa3bb398fu + 0x25b7228au;
|
||||
s[10] = t * 0xb5a09e35u + 0xcd515004u;
|
||||
s[11] = t * 0x7509c9c1u + 0x2846527au;
|
||||
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
|
||||
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
|
||||
s[14] = t * 0xded91851u + 0x82961bacu;
|
||||
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
|
||||
for (uint32_t r = 0u; r < 8u; ++r) {
|
||||
mh_mixer(s, 0x9E3779B9u * (r + 1u));
|
||||
const uint32_t* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
|
||||
for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= line[i];
|
||||
}
|
||||
mh_mixer(s, 0x9E3779B9u * 9u);
|
||||
}
|
||||
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 1 3 8 13 of w.
|
||||
IGNEUM_HD uint32_t mh_j(uint32_t w) { return ((w >> 1u) & 1u) | (((w >> 3u) & 1u) << 1) | (((w >> 8u) & 1u) << 2) | (((w >> 13u) & 1u) << 3); }
|
||||
IGNEUM_HD uint32_t mh_t(uint32_t w) { w = (w & 0x00001fffu) | ((w >> 14u) << 13u); w = (w & 0x000000ffu) | ((w >> 9u) << 8u); w = (w & 0x00000007u) | ((w >> 4u) << 3u); w = (w & 0x00000001u) | ((w >> 2u) << 1u); return w; }
|
||||
IGNEUM_HD uint32_t mh_addr(uint32_t t, uint32_t j) { uint32_t w = t; w = ((w >> 1u) << 2u) | (w & 0x00000001u) | (((j >> 0u) & 1u) << 1u); w = ((w >> 3u) << 4u) | (w & 0x00000007u) | (((j >> 1u) & 1u) << 3u); w = ((w >> 8u) << 9u) | (w & 0x000000ffu) | (((j >> 2u) & 1u) << 8u); w = ((w >> 13u) << 14u) | (w & 0x00001fffu) | (((j >> 3u) & 1u) << 13u); return w; }
|
||||
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
|
||||
IGNEUM_HD uint32_t mh_word(const uint32_t* cache, uint32_t w) { uint32_t s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }
|
||||
109
proto-cuda/packs-ca2-era/era-1/memhard.metal
Normal file
109
proto-cuda/packs-ca2-era/era-1/memhard.metal
Normal file
|
|
@ -0,0 +1,109 @@
|
|||
#include <metal_stdlib>
|
||||
using namespace metal;
|
||||
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
|
||||
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
|
||||
#define MH_CACHE_LINE_MASK 0x003fffffu
|
||||
#define MH_SEGMENT_LINES 64u
|
||||
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
|
||||
inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
|
||||
|
||||
// y = ChaCha12 core(x) + x
|
||||
inline void mh_chacha_block(const thread uint* x, thread uint* y) {
|
||||
for (uint i = 0u; i < 16u; ++i) y[i] = x[i];
|
||||
for (uint r = 0u; r < 6u; ++r) {
|
||||
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
|
||||
}
|
||||
for (uint i = 0u; i < 16u; ++i) y[i] += x[i];
|
||||
}
|
||||
|
||||
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
|
||||
inline void mh_cache_segment(device uint* cache, uint seg) {
|
||||
uint prev[16]; uint x[16]; uint y[16];
|
||||
for (uint i = 0u; i < 16u; ++i) prev[i] = 0u;
|
||||
for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) {
|
||||
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
|
||||
x[4] = 0xceed56d7u ^ prev[4];
|
||||
x[5] = 0x9ba270d2u ^ prev[5];
|
||||
x[6] = 0x82caab2du ^ prev[6];
|
||||
x[7] = 0x81ebce0eu ^ prev[7];
|
||||
x[8] = 0x12b6ecf1u ^ prev[8];
|
||||
x[9] = 0xd0f3fd7cu ^ prev[9];
|
||||
x[10] = 0xd872eefeu ^ prev[10];
|
||||
x[11] = 0xc158c7bdu ^ prev[11];
|
||||
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
|
||||
mh_chacha_block(x, y);
|
||||
device uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
|
||||
}
|
||||
}
|
||||
|
||||
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
|
||||
inline void mh_mixer(thread uint* s, uint rk) {
|
||||
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
|
||||
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
|
||||
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
|
||||
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
|
||||
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
|
||||
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
|
||||
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
|
||||
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
|
||||
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
|
||||
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
|
||||
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
|
||||
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
|
||||
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
|
||||
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
|
||||
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
|
||||
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
|
||||
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
|
||||
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
|
||||
}
|
||||
|
||||
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
|
||||
inline void mh_item(device const uint* cache, uint t, thread uint* s) {
|
||||
s[0] = 0xceed56d7u;
|
||||
s[1] = 0x9ba270d2u;
|
||||
s[2] = 0x82caab2du;
|
||||
s[3] = 0x81ebce0eu;
|
||||
s[4] = 0x12b6ecf1u;
|
||||
s[5] = 0xd0f3fd7cu;
|
||||
s[6] = 0xd872eefeu;
|
||||
s[7] = 0xc158c7bdu;
|
||||
s[8] = t * 0xf351d601u + 0xc6892460u;
|
||||
s[9] = t * 0xa3bb398fu + 0x25b7228au;
|
||||
s[10] = t * 0xb5a09e35u + 0xcd515004u;
|
||||
s[11] = t * 0x7509c9c1u + 0x2846527au;
|
||||
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
|
||||
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
|
||||
s[14] = t * 0xded91851u + 0x82961bacu;
|
||||
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
|
||||
for (uint r = 0u; r < 8u; ++r) {
|
||||
mh_mixer(s, 0x9E3779B9u * (r + 1u));
|
||||
device const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i];
|
||||
}
|
||||
mh_mixer(s, 0x9E3779B9u * 9u);
|
||||
}
|
||||
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 1 3 8 13 of w.
|
||||
inline uint mh_j(uint w) { return ((w >> 1u) & 1u) | (((w >> 3u) & 1u) << 1) | (((w >> 8u) & 1u) << 2) | (((w >> 13u) & 1u) << 3); }
|
||||
inline uint mh_t(uint w) { w = (w & 0x00001fffu) | ((w >> 14u) << 13u); w = (w & 0x000000ffu) | ((w >> 9u) << 8u); w = (w & 0x00000007u) | ((w >> 4u) << 3u); w = (w & 0x00000001u) | ((w >> 2u) << 1u); return w; }
|
||||
inline uint mh_addr(uint t, uint j) { uint w = t; w = ((w >> 1u) << 2u) | (w & 0x00000001u) | (((j >> 0u) & 1u) << 1u); w = ((w >> 3u) << 4u) | (w & 0x00000007u) | (((j >> 1u) & 1u) << 3u); w = ((w >> 8u) << 9u) | (w & 0x000000ffu) | (((j >> 2u) & 1u) << 8u); w = ((w >> 13u) << 14u) | (w & 0x00001fffu) | (((j >> 3u) & 1u) << 13u); return w; }
|
||||
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
|
||||
inline uint mh_word(device const uint* cache, uint w) { uint s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }
|
||||
|
||||
// One thread per segment (2^16 threads).
|
||||
kernel void igneum_cache_fill(device uint* cache [[buffer(0)]], uint gid [[thread_position_in_grid]]) {
|
||||
mh_cache_segment(cache, gid);
|
||||
}
|
||||
// One thread per 64-byte item (dataset words / 16 threads).
|
||||
kernel void igneum_build(device const uint* cache [[buffer(0)]], device uint* dataset [[buffer(1)]],
|
||||
uint gid [[thread_position_in_grid]]) {
|
||||
uint s[16];
|
||||
mh_item(cache, gid, s);
|
||||
for (uint i = 0u; i < 16u; ++i) dataset[mh_addr(gid, i)] = s[i];
|
||||
}
|
||||
76
proto-cuda/packs-ca2-era/era-1/program.h
Normal file
76
proto-cuda/packs-ca2-era/era-1/program.h
Normal file
|
|
@ -0,0 +1,76 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// Program metadata for host.cu plus the launch wrappers defined in kernel.cu.
|
||||
// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros.
|
||||
#pragma once
|
||||
#ifdef __cplusplus
|
||||
#include <cstdint>
|
||||
#else
|
||||
#include <stdint.h>
|
||||
#endif
|
||||
#ifndef IGNEUM_NO_CUDA
|
||||
#include <cuda_runtime.h>
|
||||
#endif
|
||||
|
||||
#define IGNEUM_SEED_STRING "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000"
|
||||
#define IGNEUM_SEED_BYTES_HEX "edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07"
|
||||
#define IGNEUM_GENERATOR 3
|
||||
#define IGNEUM_PROGRAM_ATTEMPT 0
|
||||
#define IGNEUM_PROGRAM_ID 0x73bcbfe8ccf988f1ull
|
||||
#define IGNEUM_DAY_STRING "bytes:69676e65756d2d6461792ffa50000000000000"
|
||||
#define IGNEUM_DAY_BYTES_HEX "69676e65756d2d6461792ffa50000000000000"
|
||||
#define IGNEUM_DAY0 0xceed56d7u
|
||||
#define IGNEUM_DAY1 0x9ba270d2u
|
||||
#define IGNEUM_DATASET_LOG2 28
|
||||
#define IGNEUM_MASK 0x0fffffffu
|
||||
#define IGNEUM_LANES 32
|
||||
#define IGNEUM_ITERATIONS 8
|
||||
#define IGNEUM_INSTR_COUNT 64
|
||||
#define IGNEUM_LOADS_PER_HASH 128
|
||||
#define IGNEUM_WIDE_LOADS_PER_HASH 0
|
||||
#define IGNEUM_OP_MIX "load=16 add=15 shfl=6 mad=4 or=4 rotl=4 rotr=4 xor=4 mulhi=3 mul=2 sub=2"
|
||||
// Program class v3 (Counter ASIC 2.0, docs/plans/counter-asic-2-rollout.md): generator version 3; a worker that
|
||||
// runs another class refuses this pack, and a job line names the class it wants (class=v3 era=<hex>).
|
||||
#define IGNEUM_PROGRAM_CLASS "v3"
|
||||
#define IGNEUM_ERA_SEED_HEX "ff87ad96a1b53f367a95d5ca123bab64211bd9aa57fc7ad3c5e2d496c20138d4"
|
||||
// Class v3 construction (Counter ASIC 2.0, 5 October 2026, docs/plans/mixer-x4.md): version 2 loads; the dataset item
|
||||
// derivation applies the mixer IGNEUM_MIXER_MULT times per round (memhard.h), and the cache follows the growth rule.
|
||||
#define IGNEUM_LOAD_CLASS "w4-era676a17fc"
|
||||
#define IGNEUM_LOAD_SLOTS 16
|
||||
#define IGNEUM_LOAD_MIX { 100, 0, 0 }
|
||||
#define IGNEUM_LOAD_WIDTH_COUNTS { 16, 0, 0 } // loads of 4, 16, 64 bytes per program
|
||||
#define IGNEUM_BYTES_PER_HASH 512
|
||||
#define IGNEUM_FOLD_ROT 11
|
||||
#define IGNEUM_FOLD_MUL 0x9e3779b1u
|
||||
// Era layout (5 October 2026, docs/plans/era-layout.md): NOT the lottery hash. Every dataset load reads
|
||||
// idx = ((rotl(src * STRIDE_MUL, STRIDE_ROT) & window mask) | window offset) & MASK; the window of a load site is the
|
||||
// dataset, a half or a quarter of it (IGNEUM_ERA_WINDOWS: site:shrink:offset); dataset word w holds word j(w) of item
|
||||
// t(w) with j's bits at the INTERLEAVE positions (memhard.h: mh_t, mh_j, mh_addr).
|
||||
#define IGNEUM_ERA_LABEL "676a17fc"
|
||||
#define IGNEUM_ERA_SEED_WORDS { 0x676a17fcu, 0x60bc956eu, 0x3e9865f4u, 0x68ae9e61u, 0xc0b3f442u, 0xaf406bebu, 0xb6126b9au, 0xaa30959bu }
|
||||
#define IGNEUM_ERA_ALLOWED_WIDTHS { 1, 0, 0 } // words, ascending, 0 = unused; one entry pins the width
|
||||
#define IGNEUM_ERA_WIDTH_WORDS 1
|
||||
#define IGNEUM_ERA_STRIDE_MUL 0xb2a9d70du
|
||||
#define IGNEUM_ERA_STRIDE_ROT 6
|
||||
#define IGNEUM_ERA_INTERLEAVE { 1, 3, 8, 13 }
|
||||
#define IGNEUM_ERA_WINDOWS "7:2:1 8:1:1 9:1:1 10:1:1 11:0:0 13:1:1 29:0:0 30:2:2 31:1:1 44:1:1 46:2:0 47:0:0 52:0:0 56:0:0 58:2:0 63:1:1"
|
||||
// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h)
|
||||
#define IGNEUM_DATASET_MODE 1
|
||||
|
||||
#define IGNEUM_SEEDW_INIT { 0x667d0fbdu, 0x7b8e5963u, 0x31c67e5eu, 0x4529ddc6u, 0xef19d6d8u, 0xaccf6211u, 0xda0aed32u, 0xabc6df31u }
|
||||
#define IGNEUM_KEY_INIT { 0xceed56d7u, 0x9ba270d2u, 0x82caab2du, 0x81ebce0eu, 0x12b6ecf1u, 0xd0f3fd7cu, 0xd872eefeu, 0xc158c7bdu }
|
||||
#define IGNEUM_CACHE_LOG2_WORDS 26
|
||||
#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6
|
||||
#define IGNEUM_CACHE_SEGMENTS 65536u
|
||||
#define IGNEUM_ITEM_ROUNDS 8
|
||||
#define IGNEUM_MIX_ROT_INIT { 17u, 12u, 20u, 23u, 7u, 3u, 27u, 16u }
|
||||
#define IGNEUM_MIX_MUL_INIT { 0xf351d601u, 0xa3bb398fu, 0xb5a09e35u, 0x7509c9c1u, 0x6bbf31e9u, 0xfc849a79u, 0xded91851u, 0x8d9113d1u, 0x0ff15225u, 0x3a5bdd41u, 0xab533435u, 0xe1c55ad5u, 0xe6d3bd0du, 0x9d9ffbbdu, 0xbb2a3cf3u, 0x50a7c08du }
|
||||
#define IGNEUM_MIX_RC_INIT { 0xc6892460u, 0x25b7228au, 0xcd515004u, 0x2846527au, 0xa6324241u, 0x36e3ec53u, 0x82961bacu, 0x0f97ba7du, 0xb6f921a9u, 0x3ada24e5u, 0xde20ab91u, 0x5378eeb2u, 0x7d161662u, 0x89353cc1u, 0xb1aa03a2u, 0x788acae6u }
|
||||
|
||||
#ifndef IGNEUM_NO_CUDA
|
||||
// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError().
|
||||
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments);
|
||||
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems);
|
||||
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
uint32_t nonces, uint32_t blockWarps);
|
||||
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
|
||||
#endif
|
||||
143
proto-cuda/packs-ca2-era/era-1/program.json
Normal file
143
proto-cuda/packs-ca2-era/era-1/program.json
Normal file
|
|
@ -0,0 +1,143 @@
|
|||
{
|
||||
"format": "igneum-program-pack-3",
|
||||
"generator": 3,
|
||||
"attempt": 0,
|
||||
"program_id": "0x73bcbfe8ccf988f1",
|
||||
"program_id_derivation": "FNV-1a 64 over 'igneum-program/' || generator_le32 || seed_words as little-endian bytes || attempt_le32",
|
||||
"dataset_mode": "memory-hard",
|
||||
"seed": "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000",
|
||||
"seed_bytes": "edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07",
|
||||
"seed_words": ["0x667d0fbd", "0x7b8e5963", "0x31c67e5e", "0x4529ddc6", "0xef19d6d8", "0xaccf6211", "0xda0aed32", "0xabc6df31"],
|
||||
"seed_derivation": "seed_words = FNV-1a 64 over seed_bytes (attempt 0) or seed_bytes || attempt_le32 (attempt k >= 1), basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32",
|
||||
"generator_rule": "version 2: exactly 16 load slots drawn first from instructions 1..63 (partial Fisher-Yates), the other 48 ops from the ten non-load weights (sum 75); a load's source is drawn from the registers other than dst written by an earlier instruction and not read by a load since; the candidate must pass the acceptance rule of spec 01 section 1.4.6 (static: no cyclically stale load source, every register has an injecting write; dynamic: 64 units on the seed-keyed closed-form dataset with no constant register bit, no lane-constant load site, under 164 saturated final values, every output bit within 136 of 1024, distinct addresses above 245760), else the next attempt of the seed is tried",
|
||||
"lanes": 32,
|
||||
"registers": 8,
|
||||
"iterations": 8,
|
||||
"instruction_count": 64,
|
||||
"loads_per_hash": 128,
|
||||
"program_class": "v3",
|
||||
"era_seed_bytes": "ff87ad96a1b53f367a95d5ca123bab64211bd9aa57fc7ad3c5e2d496c20138d4",
|
||||
"load_class": "w4-era676a17fc",
|
||||
"load_slots": 16,
|
||||
"load_mix_percent_4_16_64": [100, 0, 0],
|
||||
"load_width_counts_4_16_64": [16, 0, 0],
|
||||
"bytes_per_hash": 512,
|
||||
"wide_load": "read-width experiment (5 October 2026, docs/plans/read-width.md), NOT the lottery hash: a load of W words (width field, 4 or 16) reads dataset[b .. b + W) with b = (src & mask) & ~(W - 1) and folds every word into dst: x = dst ^ w[0]; for j in 1..W: x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x; width 1 is the plain load; the width is drawn per instruction from the class mix with one extra below(100) draw after the nine of version 2, and the program id is FNV-1a 64 over 'igneum-program-rw/' || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots",
|
||||
"era": {
|
||||
"label": "676a17fc",
|
||||
"seed_words": ["0x676a17fc", "0x60bc956e", "0x3e9865f4", "0x68ae9e61", "0xc0b3f442", "0xaf406beb", "0xb6126b9a", "0xaa30959b"],
|
||||
"draw": "docs/plans/era-layout.md 1.1: SplitMix64 seeded with seed_words[0] | seed_words[1] << 32 of seed_words_from_bytes('igneum-era/' || n_le64 || E_n); width = allowed[below(|allowed|)], stride_mul = low32(next()) | 1, stride_rot = 1 + below(31), then four next() draws for a partial Fisher-Yates over positions log2(W)..15 of which 4 - log2(W) are used",
|
||||
"allowed_widths": [1],
|
||||
"width_words": 1,
|
||||
"stride_mul": "0xb2a9d70d",
|
||||
"stride_rot": 6,
|
||||
"interleave": [1, 3, 8, 13],
|
||||
"address": "y = rotl(src * stride_mul, stride_rot); k = min(win, D - 26); idx = ((y & (mask >> k)) | ((off & (2^k - 1)) << (D - k))) & mask; a wide load aligns idx down to W words",
|
||||
"windows": "per instruction, after the width roll: win = below(3), off = low32(next()) & (2^win - 1); used on a load slot (the instruction's win and off fields)",
|
||||
"dataset_word": "dataset[w] = item(t(w))[j(w)]: j(w) gathers the bits of w at the interleave positions, t(w) is w with those bits removed",
|
||||
"program_id_suffix": "'era/' || allowed[3] || width_words || stride_mul_le32 || stride_rot_le32 || interleave[4]"
|
||||
},
|
||||
"op_mix": {"load": 16, "add": 15, "shfl": 6, "mad": 4, "or": 4, "rotl": 4, "rotr": 4, "xor": 4, "mulhi": 3, "mul": 2, "sub": 2},
|
||||
"register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]",
|
||||
"splitmix32": "x ^= x>>16; x *= 0x7feb352d; x ^= x>>15; x *= 0x846ca68b; x ^= x>>16",
|
||||
"iteration": "sel = r0 sampled once at the top of each iteration, then all instructions in order",
|
||||
"output": "lo = r0 ^ rotl(r1,7) ^ rotl(r2,14) ^ rotl(r3,21); hi = r4 ^ rotl(r5,9) ^ rotl(r6,18) ^ rotl(r7,27); out = (hi << 32) | lo",
|
||||
"op_semantics": {
|
||||
"add": "dst = dst + src + (bit `bit` of sel ? imm2 : imm)",
|
||||
"sub": "dst = dst - src",
|
||||
"mul": "dst = dst * src (low 32)",
|
||||
"mulhi": "dst = high 32 bits of dst * src",
|
||||
"xor": "dst = dst ^ src",
|
||||
"or": "dst = dst | src",
|
||||
"rotl": "dst = rotl(dst, rot), rot in 1..31",
|
||||
"rotr": "dst = rotr(dst, src & 31)",
|
||||
"mad": "dst = src * src2 + dst",
|
||||
"shfl": "dst = dst ^ (src of lane (lane ^ mask)), mask in {1,2,4,8,16}, within the 32-lane warp",
|
||||
"load": "dst = dst ^ dataset[src & dataset.mask]",
|
||||
"wload": "base = (src of lane 0 & dataset.mask) & ~31; dst = dst ^ dataset[base + lane] (warp-coalesced 128-byte load, lever b, only when --wide-frac > 0)"
|
||||
},
|
||||
"dataset": {
|
||||
"log2_words": 28,
|
||||
"bytes": 1073741824,
|
||||
"mask": "0x0fffffff",
|
||||
"day": "bytes:69676e65756d2d6461792ffa50000000000000",
|
||||
"day_bytes": "69676e65756d2d6461792ffa50000000000000",
|
||||
"day_words_from": "seed_words_from_bytes(day_bytes)",
|
||||
"d0": "0xceed56d7",
|
||||
"d1": "0x9ba270d2",
|
||||
"mode": "memory-hard",
|
||||
"spec": "proto-metal/MEMHARD.md",
|
||||
"key": ["0xceed56d7", "0x9ba270d2", "0x82caab2d", "0x81ebce0e", "0x12b6ecf1", "0xd0f3fd7c", "0xd872eefe", "0xc158c7bd"],
|
||||
"key_derivation": "the 8 words of seed_words_from_bytes(day_bytes); d0, d1 are key[0], key[1]",
|
||||
"cache": {"log2_words": 26, "bytes": 268435456, "line_words": 16, "segment_lines": 64, "segments": 65536, "block": "ChaCha12 core + feed-forward, rotations 16 12 8 7", "sigma": ["0x61707865", "0x3320646e", "0x79622d32", "0x6b206574"], "tag": ["0x49676e65", "0x756d4d48"], "chain": "in_j = prev_line ^ (sigma[0..3] || key[0..7] || seg || j || tag[0..1]); line_j = block(in_j); prev_0 = 0"},
|
||||
"mixer": {"draw": "SplitMix64 seeded with key[0] | key[1] << 32: rot[0..7] = 1 + next() % 31, mul[0..15] = low32(next()) | 1, rc[0..15] = low32(next())", "rot": [17, 12, 20, 23, 7, 3, 27, 16], "mul": ["0xf351d601", "0xa3bb398f", "0xb5a09e35", "0x7509c9c1", "0x6bbf31e9", "0xfc849a79", "0xded91851", "0x8d9113d1", "0x0ff15225", "0x3a5bdd41", "0xab533435", "0xe1c55ad5", "0xe6d3bd0d", "0x9d9ffbbd", "0xbb2a3cf3", "0x50a7c08d"], "rc": ["0xc6892460", "0x25b7228a", "0xcd515004", "0x2846527a", "0xa6324241", "0x36e3ec53", "0x82961bac", "0x0f97ba7d", "0xb6f921a9", "0x3ada24e5", "0xde20ab91", "0x5378eeb2", "0x7d161662", "0x89353cc1", "0xb1aa03a2", "0x788acae6"], "round": "for i in 0..15: s[i] = (s[i] ^ (rc[i] + (r+1) * 0x9E3779B9)) * mul[i]; then quarter rounds on columns (0,4,8,12) (1,5,9,13) (2,6,10,14) (3,7,11,15) with rot[0..3] and diagonals (0,5,10,15) (1,6,11,12) (2,7,8,13) (3,4,9,14) with rot[4..7]", "quarter_round": "a += b; d ^= a; d = rotl(d, r1); c += d; b ^= c; b = rotl(b, r2); a += b; d ^= a; d = rotl(d, r3); c += d; b ^= c; b = rotl(b, r4)"},
|
||||
"item": "s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..7: s = M_r(s); line = s[0] & 0x003fffff; s[i] ^= cache[line * 16 + i]; then s = M_8(s); item(t) = s",
|
||||
"word": "dataset[w] = item(w >> 4)[w & 15]"
|
||||
},
|
||||
"instructions": [
|
||||
{"i": 0, "op": "add", "dst": 4, "src": 5, "src2": 7, "imm": "0xea86e152", "imm2": "0x5810667a", "rot": 27, "bit": 13, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 1, "op": "xor", "dst": 7, "src": 0, "src2": 6, "imm": "0xbaab6229", "imm2": "0xed861989", "rot": 26, "bit": 22, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 2, "op": "shfl", "dst": 3, "src": 6, "src2": 2, "imm": "0x5b623116", "imm2": "0xff12e5b2", "rot": 12, "bit": 24, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 3, "op": "sub", "dst": 4, "src": 1, "src2": 1, "imm": "0xe99741c7", "imm2": "0xf5fa5009", "rot": 1, "bit": 21, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 4, "op": "mad", "dst": 2, "src": 0, "src2": 4, "imm": "0x673c2157", "imm2": "0xee02465f", "rot": 20, "bit": 22, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 5, "op": "rotl", "dst": 4, "src": 7, "src2": 7, "imm": "0x946f7818", "imm2": "0x45d3399e", "rot": 9, "bit": 2, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 6, "op": "rotr", "dst": 0, "src": 2, "src2": 4, "imm": "0x5f6a0ed2", "imm2": "0x7043a636", "rot": 19, "bit": 31, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 7, "op": "load", "dst": 6, "src": 7, "src2": 1, "imm": "0x5c61dcf7", "imm2": "0x7466aa40", "rot": 19, "bit": 9, "mask": 2, "width": 1, "win": 2, "off": 1},
|
||||
{"i": 8, "op": "load", "dst": 1, "src": 4, "src2": 5, "imm": "0x85668475", "imm2": "0xdb8cc483", "rot": 29, "bit": 7, "mask": 4, "width": 1, "win": 1, "off": 1},
|
||||
{"i": 9, "op": "load", "dst": 1, "src": 2, "src2": 4, "imm": "0x4454980f", "imm2": "0xebd31581", "rot": 10, "bit": 28, "mask": 16, "width": 1, "win": 1, "off": 1},
|
||||
{"i": 10, "op": "load", "dst": 7, "src": 0, "src2": 2, "imm": "0xe075297c", "imm2": "0x5779c44c", "rot": 10, "bit": 22, "mask": 2, "width": 1, "win": 1, "off": 1},
|
||||
{"i": 11, "op": "load", "dst": 7, "src": 1, "src2": 6, "imm": "0x65aa4311", "imm2": "0x4fe48ea9", "rot": 15, "bit": 9, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 12, "op": "mul", "dst": 5, "src": 4, "src2": 2, "imm": "0x1383d3ad", "imm2": "0xf3094b29", "rot": 8, "bit": 9, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 13, "op": "load", "dst": 7, "src": 6, "src2": 2, "imm": "0xed8a496f", "imm2": "0x3072c3c6", "rot": 28, "bit": 19, "mask": 8, "width": 1, "win": 1, "off": 1},
|
||||
{"i": 14, "op": "shfl", "dst": 6, "src": 5, "src2": 0, "imm": "0x8b965b57", "imm2": "0xcfeca6c1", "rot": 12, "bit": 27, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 15, "op": "shfl", "dst": 3, "src": 5, "src2": 3, "imm": "0x877c7586", "imm2": "0xa9cb2a03", "rot": 2, "bit": 29, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 16, "op": "or", "dst": 2, "src": 7, "src2": 3, "imm": "0xb740221a", "imm2": "0x89d38d6d", "rot": 6, "bit": 31, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 17, "op": "rotr", "dst": 0, "src": 6, "src2": 1, "imm": "0x26f3ad8a", "imm2": "0x27256f15", "rot": 18, "bit": 5, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 18, "op": "add", "dst": 7, "src": 5, "src2": 7, "imm": "0xf66e7017", "imm2": "0xb9e3577e", "rot": 9, "bit": 12, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 19, "op": "rotl", "dst": 7, "src": 0, "src2": 5, "imm": "0x849ae6ee", "imm2": "0x02b358f9", "rot": 24, "bit": 18, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 20, "op": "add", "dst": 6, "src": 7, "src2": 6, "imm": "0x52334d12", "imm2": "0x8c9f0ef8", "rot": 11, "bit": 23, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 21, "op": "or", "dst": 2, "src": 6, "src2": 5, "imm": "0xb1871e63", "imm2": "0xb2e40191", "rot": 5, "bit": 13, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 22, "op": "mad", "dst": 6, "src": 5, "src2": 4, "imm": "0x97df29e4", "imm2": "0xe60fea84", "rot": 11, "bit": 16, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 23, "op": "or", "dst": 1, "src": 0, "src2": 3, "imm": "0x8f30d21d", "imm2": "0x2df685a0", "rot": 31, "bit": 0, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 24, "op": "xor", "dst": 6, "src": 1, "src2": 2, "imm": "0xd2c4025f", "imm2": "0x5269eb4d", "rot": 31, "bit": 21, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 25, "op": "add", "dst": 2, "src": 6, "src2": 5, "imm": "0x659fc3d3", "imm2": "0x9cec0e12", "rot": 6, "bit": 17, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 26, "op": "rotr", "dst": 7, "src": 0, "src2": 1, "imm": "0xd89ef484", "imm2": "0x20be3846", "rot": 12, "bit": 9, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 27, "op": "mad", "dst": 4, "src": 5, "src2": 7, "imm": "0xb48420ae", "imm2": "0x3d1f2485", "rot": 14, "bit": 28, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 28, "op": "mad", "dst": 3, "src": 2, "src2": 4, "imm": "0xc5c46d76", "imm2": "0x700044b5", "rot": 22, "bit": 10, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 29, "op": "load", "dst": 1, "src": 4, "src2": 3, "imm": "0x0fbaf177", "imm2": "0xfff4f2ed", "rot": 20, "bit": 31, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 30, "op": "load", "dst": 2, "src": 3, "src2": 0, "imm": "0xe90eb2e5", "imm2": "0xb0f9eb79", "rot": 27, "bit": 14, "mask": 4, "width": 1, "win": 2, "off": 2},
|
||||
{"i": 31, "op": "load", "dst": 1, "src": 5, "src2": 2, "imm": "0x97ba3fc3", "imm2": "0x7894e657", "rot": 3, "bit": 30, "mask": 16, "width": 1, "win": 1, "off": 1},
|
||||
{"i": 32, "op": "add", "dst": 7, "src": 2, "src2": 7, "imm": "0x070888a8", "imm2": "0xe403240e", "rot": 2, "bit": 16, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 33, "op": "add", "dst": 2, "src": 0, "src2": 5, "imm": "0xf2e46d55", "imm2": "0x29701828", "rot": 31, "bit": 28, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 34, "op": "add", "dst": 2, "src": 3, "src2": 0, "imm": "0x58f75b87", "imm2": "0x343b7aee", "rot": 12, "bit": 14, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 35, "op": "mulhi", "dst": 2, "src": 5, "src2": 6, "imm": "0x5892a9e6", "imm2": "0xc9824c94", "rot": 19, "bit": 26, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 36, "op": "xor", "dst": 4, "src": 2, "src2": 3, "imm": "0x7b5b5474", "imm2": "0x45cfc5dd", "rot": 17, "bit": 18, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 37, "op": "mul", "dst": 6, "src": 5, "src2": 7, "imm": "0xccf564a5", "imm2": "0x873ad101", "rot": 7, "bit": 11, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 38, "op": "xor", "dst": 7, "src": 0, "src2": 6, "imm": "0xcac8f06d", "imm2": "0x6b97c683", "rot": 18, "bit": 28, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 39, "op": "add", "dst": 7, "src": 2, "src2": 4, "imm": "0xb8180e9d", "imm2": "0x32bbd117", "rot": 23, "bit": 19, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 40, "op": "rotr", "dst": 2, "src": 3, "src2": 0, "imm": "0x2d6070bc", "imm2": "0x68ff101e", "rot": 13, "bit": 18, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 41, "op": "sub", "dst": 7, "src": 0, "src2": 2, "imm": "0x0daf96ea", "imm2": "0x36f37be1", "rot": 5, "bit": 0, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 42, "op": "add", "dst": 4, "src": 3, "src2": 6, "imm": "0x6ced15b7", "imm2": "0x6df7aed4", "rot": 19, "bit": 4, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 43, "op": "shfl", "dst": 7, "src": 3, "src2": 5, "imm": "0x8ace05f3", "imm2": "0xd378ec12", "rot": 23, "bit": 10, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 44, "op": "load", "dst": 0, "src": 7, "src2": 4, "imm": "0xb0607786", "imm2": "0xc4acabbc", "rot": 13, "bit": 7, "mask": 16, "width": 1, "win": 1, "off": 1},
|
||||
{"i": 45, "op": "add", "dst": 1, "src": 6, "src2": 5, "imm": "0xe09f54e9", "imm2": "0x83e825bf", "rot": 23, "bit": 14, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 46, "op": "load", "dst": 3, "src": 1, "src2": 0, "imm": "0x63cc1e4e", "imm2": "0xa1be8118", "rot": 12, "bit": 6, "mask": 2, "width": 1, "win": 2, "off": 0},
|
||||
{"i": 47, "op": "load", "dst": 6, "src": 3, "src2": 7, "imm": "0x353f1d79", "imm2": "0x3b2e7456", "rot": 18, "bit": 18, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 48, "op": "mulhi", "dst": 4, "src": 2, "src2": 7, "imm": "0x00d8a3cd", "imm2": "0x231866d2", "rot": 21, "bit": 20, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 49, "op": "add", "dst": 5, "src": 0, "src2": 2, "imm": "0xa8bae6df", "imm2": "0xf572bdb9", "rot": 14, "bit": 7, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 50, "op": "shfl", "dst": 0, "src": 7, "src2": 7, "imm": "0x81ef22e1", "imm2": "0x74438fc5", "rot": 28, "bit": 18, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 51, "op": "add", "dst": 6, "src": 0, "src2": 6, "imm": "0x383b9260", "imm2": "0x11e17c61", "rot": 12, "bit": 19, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 52, "op": "load", "dst": 5, "src": 2, "src2": 2, "imm": "0xfb84f451", "imm2": "0x11cd863e", "rot": 21, "bit": 20, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 53, "op": "rotl", "dst": 6, "src": 5, "src2": 4, "imm": "0xb1a7db6b", "imm2": "0x76686b9b", "rot": 12, "bit": 4, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 54, "op": "rotl", "dst": 3, "src": 6, "src2": 3, "imm": "0x6f981f52", "imm2": "0xd99aeba2", "rot": 12, "bit": 27, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 55, "op": "add", "dst": 2, "src": 1, "src2": 2, "imm": "0xac6be8e3", "imm2": "0x18d67dbb", "rot": 26, "bit": 10, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 56, "op": "load", "dst": 1, "src": 4, "src2": 0, "imm": "0x7e7f6a00", "imm2": "0x6f0747da", "rot": 25, "bit": 28, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 57, "op": "add", "dst": 5, "src": 0, "src2": 4, "imm": "0xf03673fe", "imm2": "0xa75cd60d", "rot": 16, "bit": 12, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 58, "op": "load", "dst": 5, "src": 0, "src2": 2, "imm": "0x227f94a6", "imm2": "0x0e8344f9", "rot": 20, "bit": 10, "mask": 2, "width": 1, "win": 2, "off": 0},
|
||||
{"i": 59, "op": "add", "dst": 1, "src": 4, "src2": 3, "imm": "0xdecd4794", "imm2": "0x8dfb96bb", "rot": 21, "bit": 7, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 60, "op": "mulhi", "dst": 3, "src": 2, "src2": 2, "imm": "0x0dd268e0", "imm2": "0x53034ca9", "rot": 1, "bit": 8, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 61, "op": "or", "dst": 6, "src": 4, "src2": 7, "imm": "0x3a45a321", "imm2": "0x9bc59a5f", "rot": 25, "bit": 11, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 62, "op": "shfl", "dst": 5, "src": 4, "src2": 2, "imm": "0x8f229cc1", "imm2": "0xcaac64a2", "rot": 17, "bit": 13, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 63, "op": "load", "dst": 3, "src": 6, "src2": 6, "imm": "0xb2574178", "imm2": "0xcbcc798d", "rot": 28, "bit": 0, "mask": 16, "width": 1, "win": 1, "off": 1}
|
||||
]
|
||||
}
|
||||
109
proto-cuda/packs-ca2-era/era-1/program.metal
Normal file
109
proto-cuda/packs-ca2-era/era-1/program.metal
Normal file
|
|
@ -0,0 +1,109 @@
|
|||
#include <metal_stdlib>
|
||||
using namespace metal;
|
||||
|
||||
#define MASK 0x0fffffffu
|
||||
constant uint SEEDW[8] = { 0x667d0fbdu, 0x7b8e5963u, 0x31c67e5eu, 0x4529ddc6u, 0xef19d6d8u, 0xaccf6211u, 0xda0aed32u, 0xabc6df31u };
|
||||
|
||||
inline uint splitmix32(uint x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31
|
||||
inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
inline uint ds_elem(uint i, uint d0, uint d1) {
|
||||
uint x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
kernel void igneum_hash(device const uint* dataset [[buffer(0)]],
|
||||
device ulong* out [[buffer(1)]],
|
||||
constant uint& baseNonce [[buffer(2)]],
|
||||
uint gid [[thread_position_in_grid]]) {
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint x = nonce ^ SEEDW[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ SEEDW[1]; }
|
||||
{ uint x = nonce ^ SEEDW[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ SEEDW[2]; }
|
||||
{ uint x = nonce ^ SEEDW[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ SEEDW[3]; }
|
||||
{ uint x = nonce ^ SEEDW[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ SEEDW[4]; }
|
||||
{ uint x = nonce ^ SEEDW[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ SEEDW[5]; }
|
||||
{ uint x = nonce ^ SEEDW[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ SEEDW[6]; }
|
||||
{ uint x = nonce ^ SEEDW[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ SEEDW[7]; }
|
||||
{ uint x = nonce ^ SEEDW[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ SEEDW[0]; }
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = r4 + r5 + select(0xea86e152u, 0x5810667au, ((sel >> 13u) & 1u) != 0u); // 0
|
||||
r7 = r7 ^ r0; // 1
|
||||
r3 = r3 ^ simd_shuffle_xor(r6, (ushort)1); // 2
|
||||
r4 = r4 - r1; // 3
|
||||
r2 = r0 * r4 + r2; // 4
|
||||
r4 = rotl_imm(r4, 9u); // 5
|
||||
r0 = rotr_var(r0, r2); // 6
|
||||
r6 = r6 ^ dataset[((rotl_imm(r7 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x04000000u) & MASK]; // 7
|
||||
r1 = r1 ^ dataset[((rotl_imm(r4 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 8
|
||||
r1 = r1 ^ dataset[((rotl_imm(r2 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 9
|
||||
r7 = r7 ^ dataset[((rotl_imm(r0 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 10
|
||||
r7 = r7 ^ dataset[((rotl_imm(r1 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 11
|
||||
r5 = r5 * r4; // 12
|
||||
r7 = r7 ^ dataset[((rotl_imm(r6 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 13
|
||||
r6 = r6 ^ simd_shuffle_xor(r5, (ushort)16); // 14
|
||||
r3 = r3 ^ simd_shuffle_xor(r5, (ushort)1); // 15
|
||||
r2 = r2 | r7; // 16
|
||||
r0 = rotr_var(r0, r6); // 17
|
||||
r7 = r7 + r5 + select(0xf66e7017u, 0xb9e3577eu, ((sel >> 12u) & 1u) != 0u); // 18
|
||||
r7 = rotl_imm(r7, 24u); // 19
|
||||
r6 = r6 + r7 + select(0x52334d12u, 0x8c9f0ef8u, ((sel >> 23u) & 1u) != 0u); // 20
|
||||
r2 = r2 | r6; // 21
|
||||
r6 = r5 * r4 + r6; // 22
|
||||
r1 = r1 | r0; // 23
|
||||
r6 = r6 ^ r1; // 24
|
||||
r2 = r2 + r6 + select(0x659fc3d3u, 0x9cec0e12u, ((sel >> 17u) & 1u) != 0u); // 25
|
||||
r7 = rotr_var(r7, r0); // 26
|
||||
r4 = r5 * r7 + r4; // 27
|
||||
r3 = r2 * r4 + r3; // 28
|
||||
r1 = r1 ^ dataset[((rotl_imm(r4 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 29
|
||||
r2 = r2 ^ dataset[((rotl_imm(r3 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x08000000u) & MASK]; // 30
|
||||
r1 = r1 ^ dataset[((rotl_imm(r5 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 31
|
||||
r7 = r7 + r2 + select(0x070888a8u, 0xe403240eu, ((sel >> 16u) & 1u) != 0u); // 32
|
||||
r2 = r2 + r0 + select(0xf2e46d55u, 0x29701828u, ((sel >> 28u) & 1u) != 0u); // 33
|
||||
r2 = r2 + r3 + select(0x58f75b87u, 0x343b7aeeu, ((sel >> 14u) & 1u) != 0u); // 34
|
||||
r2 = mulhi(r2, r5); // 35
|
||||
r4 = r4 ^ r2; // 36
|
||||
r6 = r6 * r5; // 37
|
||||
r7 = r7 ^ r0; // 38
|
||||
r7 = r7 + r2 + select(0xb8180e9du, 0x32bbd117u, ((sel >> 19u) & 1u) != 0u); // 39
|
||||
r2 = rotr_var(r2, r3); // 40
|
||||
r7 = r7 - r0; // 41
|
||||
r4 = r4 + r3 + select(0x6ced15b7u, 0x6df7aed4u, ((sel >> 4u) & 1u) != 0u); // 42
|
||||
r7 = r7 ^ simd_shuffle_xor(r3, (ushort)4); // 43
|
||||
r0 = r0 ^ dataset[((rotl_imm(r7 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 44
|
||||
r1 = r1 + r6 + select(0xe09f54e9u, 0x83e825bfu, ((sel >> 14u) & 1u) != 0u); // 45
|
||||
r3 = r3 ^ dataset[((rotl_imm(r1 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 46
|
||||
r6 = r6 ^ dataset[((rotl_imm(r3 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 47
|
||||
r4 = mulhi(r4, r2); // 48
|
||||
r5 = r5 + r0 + select(0xa8bae6dfu, 0xf572bdb9u, ((sel >> 7u) & 1u) != 0u); // 49
|
||||
r0 = r0 ^ simd_shuffle_xor(r7, (ushort)4); // 50
|
||||
r6 = r6 + r0 + select(0x383b9260u, 0x11e17c61u, ((sel >> 19u) & 1u) != 0u); // 51
|
||||
r5 = r5 ^ dataset[((rotl_imm(r2 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 52
|
||||
r6 = rotl_imm(r6, 12u); // 53
|
||||
r3 = rotl_imm(r3, 12u); // 54
|
||||
r2 = r2 + r1 + select(0xac6be8e3u, 0x18d67dbbu, ((sel >> 10u) & 1u) != 0u); // 55
|
||||
r1 = r1 ^ dataset[((rotl_imm(r4 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 56
|
||||
r5 = r5 + r0 + select(0xf03673feu, 0xa75cd60du, ((sel >> 12u) & 1u) != 0u); // 57
|
||||
r5 = r5 ^ dataset[((rotl_imm(r0 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 58
|
||||
r1 = r1 + r4 + select(0xdecd4794u, 0x8dfb96bbu, ((sel >> 7u) & 1u) != 0u); // 59
|
||||
r3 = mulhi(r3, r2); // 60
|
||||
r6 = r6 | r4; // 61
|
||||
r5 = r5 ^ simd_shuffle_xor(r4, (ushort)8); // 62
|
||||
r3 = r3 ^ dataset[((rotl_imm(r6 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 63
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
111
proto-cuda/packs-ca2-era/era-1/program_bound.metal
Normal file
111
proto-cuda/packs-ca2-era/era-1/program_bound.metal
Normal file
|
|
@ -0,0 +1,111 @@
|
|||
#include <metal_stdlib>
|
||||
using namespace metal;
|
||||
|
||||
#define MASK 0x0fffffffu
|
||||
constant uint SEEDW[8] = { 0x667d0fbdu, 0x7b8e5963u, 0x31c67e5eu, 0x4529ddc6u, 0xef19d6d8u, 0xaccf6211u, 0xda0aed32u, 0xabc6df31u };
|
||||
|
||||
inline uint splitmix32(uint x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31
|
||||
inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
inline uint ds_elem(uint i, uint d0, uint d1) {
|
||||
uint x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// Header-bound variant: the init words come from buffer 3 (bind.rs), not from SEEDW.
|
||||
kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]],
|
||||
device ulong* out [[buffer(1)]],
|
||||
constant uint& baseNonce [[buffer(2)]],
|
||||
constant uint* initw [[buffer(3)]],
|
||||
uint gid [[thread_position_in_grid]]) {
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint x = nonce ^ initw[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ initw[1]; }
|
||||
{ uint x = nonce ^ initw[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ initw[2]; }
|
||||
{ uint x = nonce ^ initw[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ initw[3]; }
|
||||
{ uint x = nonce ^ initw[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ initw[4]; }
|
||||
{ uint x = nonce ^ initw[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ initw[5]; }
|
||||
{ uint x = nonce ^ initw[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ initw[6]; }
|
||||
{ uint x = nonce ^ initw[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ initw[7]; }
|
||||
{ uint x = nonce ^ initw[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ initw[0]; }
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = r4 + r5 + select(0xea86e152u, 0x5810667au, ((sel >> 13u) & 1u) != 0u); // 0
|
||||
r7 = r7 ^ r0; // 1
|
||||
r3 = r3 ^ simd_shuffle_xor(r6, (ushort)1); // 2
|
||||
r4 = r4 - r1; // 3
|
||||
r2 = r0 * r4 + r2; // 4
|
||||
r4 = rotl_imm(r4, 9u); // 5
|
||||
r0 = rotr_var(r0, r2); // 6
|
||||
r6 = r6 ^ dataset[((rotl_imm(r7 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x04000000u) & MASK]; // 7
|
||||
r1 = r1 ^ dataset[((rotl_imm(r4 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 8
|
||||
r1 = r1 ^ dataset[((rotl_imm(r2 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 9
|
||||
r7 = r7 ^ dataset[((rotl_imm(r0 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 10
|
||||
r7 = r7 ^ dataset[((rotl_imm(r1 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 11
|
||||
r5 = r5 * r4; // 12
|
||||
r7 = r7 ^ dataset[((rotl_imm(r6 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 13
|
||||
r6 = r6 ^ simd_shuffle_xor(r5, (ushort)16); // 14
|
||||
r3 = r3 ^ simd_shuffle_xor(r5, (ushort)1); // 15
|
||||
r2 = r2 | r7; // 16
|
||||
r0 = rotr_var(r0, r6); // 17
|
||||
r7 = r7 + r5 + select(0xf66e7017u, 0xb9e3577eu, ((sel >> 12u) & 1u) != 0u); // 18
|
||||
r7 = rotl_imm(r7, 24u); // 19
|
||||
r6 = r6 + r7 + select(0x52334d12u, 0x8c9f0ef8u, ((sel >> 23u) & 1u) != 0u); // 20
|
||||
r2 = r2 | r6; // 21
|
||||
r6 = r5 * r4 + r6; // 22
|
||||
r1 = r1 | r0; // 23
|
||||
r6 = r6 ^ r1; // 24
|
||||
r2 = r2 + r6 + select(0x659fc3d3u, 0x9cec0e12u, ((sel >> 17u) & 1u) != 0u); // 25
|
||||
r7 = rotr_var(r7, r0); // 26
|
||||
r4 = r5 * r7 + r4; // 27
|
||||
r3 = r2 * r4 + r3; // 28
|
||||
r1 = r1 ^ dataset[((rotl_imm(r4 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 29
|
||||
r2 = r2 ^ dataset[((rotl_imm(r3 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x08000000u) & MASK]; // 30
|
||||
r1 = r1 ^ dataset[((rotl_imm(r5 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 31
|
||||
r7 = r7 + r2 + select(0x070888a8u, 0xe403240eu, ((sel >> 16u) & 1u) != 0u); // 32
|
||||
r2 = r2 + r0 + select(0xf2e46d55u, 0x29701828u, ((sel >> 28u) & 1u) != 0u); // 33
|
||||
r2 = r2 + r3 + select(0x58f75b87u, 0x343b7aeeu, ((sel >> 14u) & 1u) != 0u); // 34
|
||||
r2 = mulhi(r2, r5); // 35
|
||||
r4 = r4 ^ r2; // 36
|
||||
r6 = r6 * r5; // 37
|
||||
r7 = r7 ^ r0; // 38
|
||||
r7 = r7 + r2 + select(0xb8180e9du, 0x32bbd117u, ((sel >> 19u) & 1u) != 0u); // 39
|
||||
r2 = rotr_var(r2, r3); // 40
|
||||
r7 = r7 - r0; // 41
|
||||
r4 = r4 + r3 + select(0x6ced15b7u, 0x6df7aed4u, ((sel >> 4u) & 1u) != 0u); // 42
|
||||
r7 = r7 ^ simd_shuffle_xor(r3, (ushort)4); // 43
|
||||
r0 = r0 ^ dataset[((rotl_imm(r7 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 44
|
||||
r1 = r1 + r6 + select(0xe09f54e9u, 0x83e825bfu, ((sel >> 14u) & 1u) != 0u); // 45
|
||||
r3 = r3 ^ dataset[((rotl_imm(r1 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 46
|
||||
r6 = r6 ^ dataset[((rotl_imm(r3 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 47
|
||||
r4 = mulhi(r4, r2); // 48
|
||||
r5 = r5 + r0 + select(0xa8bae6dfu, 0xf572bdb9u, ((sel >> 7u) & 1u) != 0u); // 49
|
||||
r0 = r0 ^ simd_shuffle_xor(r7, (ushort)4); // 50
|
||||
r6 = r6 + r0 + select(0x383b9260u, 0x11e17c61u, ((sel >> 19u) & 1u) != 0u); // 51
|
||||
r5 = r5 ^ dataset[((rotl_imm(r2 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 52
|
||||
r6 = rotl_imm(r6, 12u); // 53
|
||||
r3 = rotl_imm(r3, 12u); // 54
|
||||
r2 = r2 + r1 + select(0xac6be8e3u, 0x18d67dbbu, ((sel >> 10u) & 1u) != 0u); // 55
|
||||
r1 = r1 ^ dataset[((rotl_imm(r4 * 0xb2a9d70du, 6u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 56
|
||||
r5 = r5 + r0 + select(0xf03673feu, 0xa75cd60du, ((sel >> 12u) & 1u) != 0u); // 57
|
||||
r5 = r5 ^ dataset[((rotl_imm(r0 * 0xb2a9d70du, 6u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 58
|
||||
r1 = r1 + r4 + select(0xdecd4794u, 0x8dfb96bbu, ((sel >> 7u) & 1u) != 0u); // 59
|
||||
r3 = mulhi(r3, r2); // 60
|
||||
r6 = r6 | r4; // 61
|
||||
r5 = r5 ^ simd_shuffle_xor(r4, (ushort)8); // 62
|
||||
r3 = r3 ^ dataset[((rotl_imm(r6 * 0xb2a9d70du, 6u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 63
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
5
proto-cuda/packs-ca2-era/era-1/seeds.txt
Normal file
5
proto-cuda/packs-ca2-era/era-1/seeds.txt
Normal file
|
|
@ -0,0 +1,5 @@
|
|||
epoch_seed_hex edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07
|
||||
day_seed_hex 69676e65756d2d6461792ffa50000000000000
|
||||
epoch_index 0
|
||||
day_index 20730
|
||||
daa_score 0
|
||||
57
proto-cuda/packs-ca2-era/era-1/vectors.h
Normal file
57
proto-cuda/packs-ca2-era/era-1/vectors.h
Normal file
|
|
@ -0,0 +1,57 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v3, memory-hard dataset
|
||||
#pragma once
|
||||
#ifdef __cplusplus
|
||||
#include <cstdint>
|
||||
#else
|
||||
#include <stdint.h>
|
||||
#endif
|
||||
|
||||
#define IGNEUM_VEC_WARPS 3
|
||||
static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u };
|
||||
static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = {
|
||||
{ // base nonce 0
|
||||
0xdabe038c15c539f6ull, 0xda2a12ac4857cae1ull, 0x2768199f161cd4efull, 0x28e44246d90e9cffull, 0xa6e49bc2e7b5da30ull, 0x9797d70a66935020ull, 0x4820804005effffeull, 0x88662fd15b9fe6e0ull,
|
||||
0xd7ffe7abc6bed954ull, 0x29baff8fcb44fa53ull, 0x5a090c72603a9301ull, 0x3c81aea5dac7be1eull, 0xfb1a9cc02b40b817ull, 0x3d0e8e32741a6091ull, 0xf7fd20269e5402beull, 0x89b31407de8924b6ull,
|
||||
0xaae6fe83cc322e25ull, 0x076d6b1d7de1f125ull, 0xb4aa525302008bdfull, 0x860c033cd41fcd1dull, 0x20d77bcc09e08aa7ull, 0xe75793c696a71f74ull, 0xca05dc63ec4f5b6eull, 0x0c07b473ff504716ull,
|
||||
0x8318ca28bfe8d189ull, 0xae11050d0adc5a87ull, 0x3b899c2da4ce3239ull, 0x815c563707ff1e0cull, 0x783799ce7b282d84ull, 0xc96488bfb2156fc9ull, 0x65d30aace55b4d9full, 0x0051015324bc1753ull
|
||||
},
|
||||
{ // base nonce 4096
|
||||
0x09635fc8b43e490aull, 0x2cde2ee9fb2a0428ull, 0x7efadc8593c3b2ccull, 0x902068a0e0355bc3ull, 0xc3af55591a6b9a4cull, 0xb36f5ef5bef69733ull, 0x32d797b26f3ff603ull, 0x80e0374e701bb677ull,
|
||||
0x08db70d60ac78f51ull, 0x23d6ee5b06233744ull, 0x2a4de99fd6f150e5ull, 0x5727b4c1f0db5162ull, 0x191bafa761798bacull, 0x94e47312bbf20681ull, 0xee9da67774b46e6aull, 0x1465ae37760b4e08ull,
|
||||
0x2d96a0f805f370dcull, 0xc1a761ea03e0d500ull, 0x06317cfdbfcdc10eull, 0xef71c2112b4897dfull, 0x6e364b602701201dull, 0xdd0cfb3a907838ceull, 0x569e82cb5acd3801ull, 0x56723edf79db59fbull,
|
||||
0x5826c8add612ef45ull, 0xc3fdce113372a83dull, 0x632b40402f52ff52ull, 0x443994dfc952a6abull, 0x5ffda84776dcf8eeull, 0x79fe9a93648e18dfull, 0x266faea56ca3ba5dull, 0x0277609022e6adc0ull
|
||||
},
|
||||
{ // base nonce 1000000
|
||||
0x689af011faee80c1ull, 0x4e62e84a6a805665ull, 0x8828e37e049d9739ull, 0xaf733761893000dbull, 0xdf9aa4524c6c9075ull, 0x34af22f30b9827c9ull, 0xe7ae4db768b3bef8ull, 0x83bb23a187bf66c9ull,
|
||||
0xad82b133ca30a3e5ull, 0x5a4c4ae80d988fd2ull, 0xbfc7229dca4c05bdull, 0xce783e7c7636ca10ull, 0xb18c78869943c79cull, 0x1023948ca66770cbull, 0x84dcd256dbf5e9dcull, 0x2e9ad8e9c0bcd6e4ull,
|
||||
0xf2d3c29d300046bdull, 0x10713b9bd25e3583ull, 0x9a1fa2385a63b827ull, 0xb31ae3aeeaa46884ull, 0xb16e8a6c82c7bd1aull, 0xfed6df91f0f3cfb4ull, 0x864d1f1b8727aed8ull, 0x89e9ac61e6362617ull,
|
||||
0x70a1df41038533bbull, 0x3d10c42046a6413bull, 0x31023c3566190ecfull, 0x2294dc505d480256ull, 0x24ff4fb403755cc7ull, 0x50ed84b9b4720072ull, 0x783690abb3addc85ull, 0x12eabf30075edb7bull
|
||||
}
|
||||
};
|
||||
|
||||
// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455).
|
||||
static const uint32_t IGNEUM_DS_HEAD[16] = {
|
||||
0x3dd50b1fu, 0x93123701u, 0x48edec90u, 0x3a7d2407u, 0xcdd88b45u, 0x2c37c6b0u, 0xb03a0ac0u, 0x10e2c6a5u,
|
||||
0x5119540eu, 0x3358ef2du, 0x657ac748u, 0x09e2105cu, 0x25e2929cu, 0x5fa0cdccu, 0x4317cabdu, 0x6178ee52u
|
||||
};
|
||||
static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u;
|
||||
static const uint32_t IGNEUM_DS_LAST = 0xf7b7180eu;
|
||||
// 64 sampled dataset words (index, value) computed on the Mac.
|
||||
#define IGNEUM_DS_SAMPLES 64
|
||||
static const uint32_t IGNEUM_DS_SAMPLE_INDEX[IGNEUM_DS_SAMPLES] = {
|
||||
59471966u, 217795994u, 208353206u, 42483309u, 172547758u, 148076330u, 183853158u, 214389424u, 267488061u, 169781097u, 184093494u, 153880993u, 84977930u, 46426879u, 3093825u, 225364072u, 44593546u, 260713159u, 168250303u, 52384140u, 223401610u, 45554030u, 95410555u, 175039924u, 79171087u, 267580473u, 24168642u, 37981670u, 171551130u, 195559979u, 204611762u, 140997658u, 138925853u, 86637313u, 20736778u, 219665210u, 160430336u, 264654675u, 8013395u, 228945585u, 213884386u, 104419827u, 44185464u, 142737231u, 99284897u, 132475900u, 61861762u, 132056166u, 262388043u, 91878046u, 117353561u, 124768597u, 71352993u, 190698941u, 46055428u, 55281366u, 165145231u, 106810753u, 171985651u, 232085256u, 159510492u, 40072060u, 209107596u, 39023794u
|
||||
};
|
||||
static const uint32_t IGNEUM_DS_SAMPLE_VALUE[IGNEUM_DS_SAMPLES] = {
|
||||
0xf6e00c89u, 0x370d4c24u, 0x09260b8cu, 0x16208e9du, 0x4028c489u, 0x1386ca6bu, 0x88bf905cu, 0x88f76874u, 0xac034e2du, 0x88e89b31u, 0x22deecc6u, 0xd6022e44u, 0x4b18c60eu, 0x8c761ec7u, 0x7079d93du, 0x79341b4cu, 0x5b93bcc4u, 0xc659d408u, 0xd5883003u, 0xe6b03742u, 0x016da861u, 0xbbec95bfu, 0xc1647d24u, 0x88b20f49u, 0xcbcb1898u, 0x05d9a8aeu, 0x2bfaeba0u, 0x12094029u, 0xd2a03583u, 0xf8bc77efu, 0xeb7dd210u, 0x251ecc1du, 0xa8d8a48bu, 0x26c9e599u, 0x12927937u, 0x536d595eu, 0x1f437e50u, 0x9b52077au, 0x876dce0cu, 0x65092ed9u, 0xf89bbfddu, 0xf20ee5b8u, 0xb4205b97u, 0xd79db283u, 0x65869dcfu, 0xd6c0f997u, 0xa9effec4u, 0xd192a75au, 0x73d9e395u, 0x2f043a85u, 0xdc0c3fddu, 0x60b1d49au, 0xf5da472eu, 0x83f8e3d3u, 0xdb881ba6u, 0x2f00a5dcu, 0xbe7f358eu, 0x68a25926u, 0xf5852ad5u, 0xa731ea98u, 0xbb567051u, 0xa1b64724u, 0x1d8896b6u, 0x81ef09f4u
|
||||
};
|
||||
// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^26 words.
|
||||
static const uint32_t IGNEUM_CACHE_HEAD[16] = {
|
||||
0xebd9055cu, 0x7eaf6b21u, 0x9b610855u, 0x134a8562u, 0xb10059bau, 0x7f459e58u, 0x8f3c7873u, 0x3523e609u,
|
||||
0x75b8f4cau, 0x750d31b4u, 0x0dae0782u, 0x26f10015u, 0x7ba87a41u, 0x0efd8543u, 0x691d6368u, 0x8929d967u
|
||||
};
|
||||
static const uint32_t IGNEUM_CACHE_LAST[16] = {
|
||||
0x51474edfu, 0xc3cfba93u, 0xf21454cfu, 0x9b79baacu, 0x4d4fce23u, 0xd701dfd5u, 0x37357ab3u, 0x1be693fau,
|
||||
0xa701cb3bu, 0x7467c620u, 0x428184e8u, 0xf4010df0u, 0xe33aa88fu, 0xc78d6d62u, 0xae9ba9c0u, 0xf4bba97eu
|
||||
};
|
||||
static const uint64_t IGNEUM_CACHE_FNV64 = 0x448274a57f508cbcull;
|
||||
36
proto-cuda/packs-ca2-era/era-1/vectors.json
Normal file
36
proto-cuda/packs-ca2-era/era-1/vectors.json
Normal file
|
|
@ -0,0 +1,36 @@
|
|||
{
|
||||
"seed": "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000",
|
||||
"day": "bytes:69676e65756d2d6461792ffa50000000000000",
|
||||
"dataset_mode": "memory-hard",
|
||||
"dataset_log2_words": 28,
|
||||
"mask": "0x0fffffff",
|
||||
"lanes": 32,
|
||||
"source": "igneum-pow (Rust) CPU interpreter, generator v3, memory-hard dataset",
|
||||
"warps": [
|
||||
{"base_nonce": 0, "expected": [
|
||||
"0xdabe038c15c539f6", "0xda2a12ac4857cae1", "0x2768199f161cd4ef", "0x28e44246d90e9cff", "0xa6e49bc2e7b5da30", "0x9797d70a66935020", "0x4820804005effffe", "0x88662fd15b9fe6e0",
|
||||
"0xd7ffe7abc6bed954", "0x29baff8fcb44fa53", "0x5a090c72603a9301", "0x3c81aea5dac7be1e", "0xfb1a9cc02b40b817", "0x3d0e8e32741a6091", "0xf7fd20269e5402be", "0x89b31407de8924b6",
|
||||
"0xaae6fe83cc322e25", "0x076d6b1d7de1f125", "0xb4aa525302008bdf", "0x860c033cd41fcd1d", "0x20d77bcc09e08aa7", "0xe75793c696a71f74", "0xca05dc63ec4f5b6e", "0x0c07b473ff504716",
|
||||
"0x8318ca28bfe8d189", "0xae11050d0adc5a87", "0x3b899c2da4ce3239", "0x815c563707ff1e0c", "0x783799ce7b282d84", "0xc96488bfb2156fc9", "0x65d30aace55b4d9f", "0x0051015324bc1753"
|
||||
]},
|
||||
{"base_nonce": 4096, "expected": [
|
||||
"0x09635fc8b43e490a", "0x2cde2ee9fb2a0428", "0x7efadc8593c3b2cc", "0x902068a0e0355bc3", "0xc3af55591a6b9a4c", "0xb36f5ef5bef69733", "0x32d797b26f3ff603", "0x80e0374e701bb677",
|
||||
"0x08db70d60ac78f51", "0x23d6ee5b06233744", "0x2a4de99fd6f150e5", "0x5727b4c1f0db5162", "0x191bafa761798bac", "0x94e47312bbf20681", "0xee9da67774b46e6a", "0x1465ae37760b4e08",
|
||||
"0x2d96a0f805f370dc", "0xc1a761ea03e0d500", "0x06317cfdbfcdc10e", "0xef71c2112b4897df", "0x6e364b602701201d", "0xdd0cfb3a907838ce", "0x569e82cb5acd3801", "0x56723edf79db59fb",
|
||||
"0x5826c8add612ef45", "0xc3fdce113372a83d", "0x632b40402f52ff52", "0x443994dfc952a6ab", "0x5ffda84776dcf8ee", "0x79fe9a93648e18df", "0x266faea56ca3ba5d", "0x0277609022e6adc0"
|
||||
]},
|
||||
{"base_nonce": 1000000, "expected": [
|
||||
"0x689af011faee80c1", "0x4e62e84a6a805665", "0x8828e37e049d9739", "0xaf733761893000db", "0xdf9aa4524c6c9075", "0x34af22f30b9827c9", "0xe7ae4db768b3bef8", "0x83bb23a187bf66c9",
|
||||
"0xad82b133ca30a3e5", "0x5a4c4ae80d988fd2", "0xbfc7229dca4c05bd", "0xce783e7c7636ca10", "0xb18c78869943c79c", "0x1023948ca66770cb", "0x84dcd256dbf5e9dc", "0x2e9ad8e9c0bcd6e4",
|
||||
"0xf2d3c29d300046bd", "0x10713b9bd25e3583", "0x9a1fa2385a63b827", "0xb31ae3aeeaa46884", "0xb16e8a6c82c7bd1a", "0xfed6df91f0f3cfb4", "0x864d1f1b8727aed8", "0x89e9ac61e6362617",
|
||||
"0x70a1df41038533bb", "0x3d10c42046a6413b", "0x31023c3566190ecf", "0x2294dc505d480256", "0x24ff4fb403755cc7", "0x50ed84b9b4720072", "0x783690abb3addc85", "0x12eabf30075edb7b"
|
||||
]}
|
||||
],
|
||||
"dataset_head": ["0x3dd50b1f", "0x93123701", "0x48edec90", "0x3a7d2407", "0xcdd88b45", "0x2c37c6b0", "0xb03a0ac0", "0x10e2c6a5", "0x5119540e", "0x3358ef2d", "0x657ac748", "0x09e2105c", "0x25e2929c", "0x5fa0cdcc", "0x4317cabd", "0x6178ee52"],
|
||||
"dataset_last_index": 268435455,
|
||||
"dataset_last": "0xf7b7180e",
|
||||
"dataset_samples": [{"index": 59471966, "value": "0xf6e00c89"}, {"index": 217795994, "value": "0x370d4c24"}, {"index": 208353206, "value": "0x09260b8c"}, {"index": 42483309, "value": "0x16208e9d"}, {"index": 172547758, "value": "0x4028c489"}, {"index": 148076330, "value": "0x1386ca6b"}, {"index": 183853158, "value": "0x88bf905c"}, {"index": 214389424, "value": "0x88f76874"}, {"index": 267488061, "value": "0xac034e2d"}, {"index": 169781097, "value": "0x88e89b31"}, {"index": 184093494, "value": "0x22deecc6"}, {"index": 153880993, "value": "0xd6022e44"}, {"index": 84977930, "value": "0x4b18c60e"}, {"index": 46426879, "value": "0x8c761ec7"}, {"index": 3093825, "value": "0x7079d93d"}, {"index": 225364072, "value": "0x79341b4c"}, {"index": 44593546, "value": "0x5b93bcc4"}, {"index": 260713159, "value": "0xc659d408"}, {"index": 168250303, "value": "0xd5883003"}, {"index": 52384140, "value": "0xe6b03742"}, {"index": 223401610, "value": "0x016da861"}, {"index": 45554030, "value": "0xbbec95bf"}, {"index": 95410555, "value": "0xc1647d24"}, {"index": 175039924, "value": "0x88b20f49"}, {"index": 79171087, "value": "0xcbcb1898"}, {"index": 267580473, "value": "0x05d9a8ae"}, {"index": 24168642, "value": "0x2bfaeba0"}, {"index": 37981670, "value": "0x12094029"}, {"index": 171551130, "value": "0xd2a03583"}, {"index": 195559979, "value": "0xf8bc77ef"}, {"index": 204611762, "value": "0xeb7dd210"}, {"index": 140997658, "value": "0x251ecc1d"}, {"index": 138925853, "value": "0xa8d8a48b"}, {"index": 86637313, "value": "0x26c9e599"}, {"index": 20736778, "value": "0x12927937"}, {"index": 219665210, "value": "0x536d595e"}, {"index": 160430336, "value": "0x1f437e50"}, {"index": 264654675, "value": "0x9b52077a"}, {"index": 8013395, "value": "0x876dce0c"}, {"index": 228945585, "value": "0x65092ed9"}, {"index": 213884386, "value": "0xf89bbfdd"}, {"index": 104419827, "value": "0xf20ee5b8"}, {"index": 44185464, "value": "0xb4205b97"}, {"index": 142737231, "value": "0xd79db283"}, {"index": 99284897, "value": "0x65869dcf"}, {"index": 132475900, "value": "0xd6c0f997"}, {"index": 61861762, "value": "0xa9effec4"}, {"index": 132056166, "value": "0xd192a75a"}, {"index": 262388043, "value": "0x73d9e395"}, {"index": 91878046, "value": "0x2f043a85"}, {"index": 117353561, "value": "0xdc0c3fdd"}, {"index": 124768597, "value": "0x60b1d49a"}, {"index": 71352993, "value": "0xf5da472e"}, {"index": 190698941, "value": "0x83f8e3d3"}, {"index": 46055428, "value": "0xdb881ba6"}, {"index": 55281366, "value": "0x2f00a5dc"}, {"index": 165145231, "value": "0xbe7f358e"}, {"index": 106810753, "value": "0x68a25926"}, {"index": 171985651, "value": "0xf5852ad5"}, {"index": 232085256, "value": "0xa731ea98"}, {"index": 159510492, "value": "0xbb567051"}, {"index": 40072060, "value": "0xa1b64724"}, {"index": 209107596, "value": "0x1d8896b6"}, {"index": 39023794, "value": "0x81ef09f4"}],
|
||||
"cache_head": ["0xebd9055c", "0x7eaf6b21", "0x9b610855", "0x134a8562", "0xb10059ba", "0x7f459e58", "0x8f3c7873", "0x3523e609", "0x75b8f4ca", "0x750d31b4", "0x0dae0782", "0x26f10015", "0x7ba87a41", "0x0efd8543", "0x691d6368", "0x8929d967"],
|
||||
"cache_last_line": ["0x51474edf", "0xc3cfba93", "0xf21454cf", "0x9b79baac", "0x4d4fce23", "0xd701dfd5", "0x37357ab3", "0x1be693fa", "0xa701cb3b", "0x7467c620", "0x428184e8", "0xf4010df0", "0xe33aa88f", "0xc78d6d62", "0xae9ba9c0", "0xf4bba97e"],
|
||||
"cache_fnv1a64": "0x448274a57f508cbc"
|
||||
}
|
||||
281
proto-cuda/packs-ca2-era/era-2/kernel.cl
Normal file
281
proto-cuda/packs-ca2-era/era-2/kernel.cl
Normal file
|
|
@ -0,0 +1,281 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal).
|
||||
// Built from source at runtime by proto-opencl/host.c, which passes these defines:
|
||||
// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit)
|
||||
// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default)
|
||||
// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32
|
||||
// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition
|
||||
// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units;
|
||||
// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32.
|
||||
#ifndef IGNEUM_GROUP
|
||||
#define IGNEUM_GROUP 32
|
||||
#endif
|
||||
#ifndef IGNEUM_EXCHANGE
|
||||
#define IGNEUM_EXCHANGE 0
|
||||
#endif
|
||||
#ifdef __OPENCL_VERSION__
|
||||
#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1)))
|
||||
#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n]
|
||||
#if IGNEUM_EXCHANGE == 1
|
||||
#ifdef cl_khr_subgroups
|
||||
#pragma OPENCL EXTENSION cl_khr_subgroups : enable
|
||||
#endif
|
||||
#ifdef cl_khr_subgroup_shuffle
|
||||
#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable
|
||||
#endif
|
||||
#elif IGNEUM_EXCHANGE == 2
|
||||
#pragma OPENCL EXTENSION cl_intel_subgroups : enable
|
||||
#endif
|
||||
#else
|
||||
// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros.
|
||||
#include "emu_opencl.h"
|
||||
#endif
|
||||
|
||||
#if IGNEUM_EXCHANGE == 1
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m))
|
||||
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
|
||||
#elif IGNEUM_EXCHANGE == 2
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m))
|
||||
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
|
||||
#else
|
||||
// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per
|
||||
// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane
|
||||
// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's
|
||||
// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier.
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; }
|
||||
#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; }
|
||||
#endif
|
||||
|
||||
static inline uint splitmix32(uint x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32.
|
||||
static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); }
|
||||
// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x.
|
||||
static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); }
|
||||
static inline uint ds_elem(uint i, uint d0, uint d1) {
|
||||
uint x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
|
||||
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
|
||||
#define MH_CACHE_LINE_MASK 0x003fffffu
|
||||
#define MH_SEGMENT_LINES 64u
|
||||
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
|
||||
static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
|
||||
|
||||
// y = ChaCha12 core(x) + x
|
||||
static inline void mh_chacha_block(const uint* x, uint* y) {
|
||||
for (uint i = 0u; i < 16u; ++i) y[i] = x[i];
|
||||
for (uint r = 0u; r < 6u; ++r) {
|
||||
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
|
||||
}
|
||||
for (uint i = 0u; i < 16u; ++i) y[i] += x[i];
|
||||
}
|
||||
|
||||
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
|
||||
static inline void mh_cache_segment(__global uint* cache, uint seg) {
|
||||
uint prev[16]; uint x[16]; uint y[16];
|
||||
for (uint i = 0u; i < 16u; ++i) prev[i] = 0u;
|
||||
for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) {
|
||||
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
|
||||
x[4] = 0xceed56d7u ^ prev[4];
|
||||
x[5] = 0x9ba270d2u ^ prev[5];
|
||||
x[6] = 0x82caab2du ^ prev[6];
|
||||
x[7] = 0x81ebce0eu ^ prev[7];
|
||||
x[8] = 0x12b6ecf1u ^ prev[8];
|
||||
x[9] = 0xd0f3fd7cu ^ prev[9];
|
||||
x[10] = 0xd872eefeu ^ prev[10];
|
||||
x[11] = 0xc158c7bdu ^ prev[11];
|
||||
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
|
||||
mh_chacha_block(x, y);
|
||||
__global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
|
||||
}
|
||||
}
|
||||
|
||||
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
|
||||
static inline void mh_mixer(uint* s, uint rk) {
|
||||
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
|
||||
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
|
||||
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
|
||||
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
|
||||
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
|
||||
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
|
||||
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
|
||||
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
|
||||
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
|
||||
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
|
||||
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
|
||||
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
|
||||
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
|
||||
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
|
||||
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
|
||||
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
|
||||
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
|
||||
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
|
||||
}
|
||||
|
||||
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
|
||||
static inline void mh_item(__global const uint* cache, uint t, uint* s) {
|
||||
s[0] = 0xceed56d7u;
|
||||
s[1] = 0x9ba270d2u;
|
||||
s[2] = 0x82caab2du;
|
||||
s[3] = 0x81ebce0eu;
|
||||
s[4] = 0x12b6ecf1u;
|
||||
s[5] = 0xd0f3fd7cu;
|
||||
s[6] = 0xd872eefeu;
|
||||
s[7] = 0xc158c7bdu;
|
||||
s[8] = t * 0xf351d601u + 0xc6892460u;
|
||||
s[9] = t * 0xa3bb398fu + 0x25b7228au;
|
||||
s[10] = t * 0xb5a09e35u + 0xcd515004u;
|
||||
s[11] = t * 0x7509c9c1u + 0x2846527au;
|
||||
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
|
||||
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
|
||||
s[14] = t * 0xded91851u + 0x82961bacu;
|
||||
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
|
||||
for (uint r = 0u; r < 8u; ++r) {
|
||||
mh_mixer(s, 0x9E3779B9u * (r + 1u));
|
||||
__global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i];
|
||||
}
|
||||
mh_mixer(s, 0x9E3779B9u * 9u);
|
||||
}
|
||||
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 1 3 4 8 of w.
|
||||
static inline uint mh_j(uint w) { return ((w >> 1u) & 1u) | (((w >> 3u) & 1u) << 1) | (((w >> 4u) & 1u) << 2) | (((w >> 8u) & 1u) << 3); }
|
||||
static inline uint mh_t(uint w) { w = (w & 0x000000ffu) | ((w >> 9u) << 8u); w = (w & 0x0000000fu) | ((w >> 5u) << 4u); w = (w & 0x00000007u) | ((w >> 4u) << 3u); w = (w & 0x00000001u) | ((w >> 2u) << 1u); return w; }
|
||||
static inline uint mh_addr(uint t, uint j) { uint w = t; w = ((w >> 1u) << 2u) | (w & 0x00000001u) | (((j >> 0u) & 1u) << 1u); w = ((w >> 3u) << 4u) | (w & 0x00000007u) | (((j >> 1u) & 1u) << 3u); w = ((w >> 4u) << 5u) | (w & 0x0000000fu) | (((j >> 2u) & 1u) << 4u); w = ((w >> 8u) << 9u) | (w & 0x000000ffu) | (((j >> 3u) & 1u) << 8u); return w; }
|
||||
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
|
||||
static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }
|
||||
|
||||
// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item.
|
||||
// The same constants as memhard.h in this pack (one emitter, three dialects).
|
||||
__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) {
|
||||
uint seg = (uint)get_global_id(0);
|
||||
if (seg < nSegments) mh_cache_segment(cache, seg);
|
||||
}
|
||||
__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) {
|
||||
uint t = (uint)get_global_id(0);
|
||||
if (t < nItems) {
|
||||
uint s[16];
|
||||
mh_item(cache, t, s);
|
||||
for (uint i = 0u; i < 16u; ++i) ds[(ulong)mh_addr(t, i)] = s[i];
|
||||
}
|
||||
}
|
||||
|
||||
// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the
|
||||
// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and
|
||||
// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all).
|
||||
IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) {
|
||||
uint gid = (uint)get_global_id(0);
|
||||
uint lid = (uint)get_local_id(0);
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
#if IGNEUM_EXCHANGE == 0
|
||||
IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
|
||||
uint xk = 0u;
|
||||
#else
|
||||
(void)lid;
|
||||
#endif
|
||||
{ uint x = nonce ^ 0x667d0fbdu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x7b8e5963u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
|
||||
{ uint x = nonce ^ 0x7b8e5963u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x31c67e5eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
|
||||
{ uint x = nonce ^ 0x31c67e5eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4529ddc6u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
|
||||
{ uint x = nonce ^ 0x4529ddc6u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef19d6d8u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
|
||||
{ uint x = nonce ^ 0xef19d6d8u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xaccf6211u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
|
||||
{ uint x = nonce ^ 0xaccf6211u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xda0aed32u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
|
||||
{ uint x = nonce ^ 0xda0aed32u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xabc6df31u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
|
||||
{ uint x = nonce ^ 0xabc6df31u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x667d0fbdu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
|
||||
r7 = r7 ^ r0; // 1 xor
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 1u); r3 = r3 ^ t_; } // 2 shfl
|
||||
r4 = r4 - r1; // 3 sub
|
||||
r2 = r0 * r4 + r2; // 4 mad
|
||||
r4 = rotl_imm(r4, 9u); // 5 rotl
|
||||
r0 = rotr_var(r0, r2); // 6 rotr
|
||||
r6 = r6 ^ ds[((rotl_imm(r7 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r2 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r0 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r1 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
|
||||
r5 = r5 * r4; // 12 mul
|
||||
r7 = r7 ^ ds[((rotl_imm(r6 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 16u); r6 = r6 ^ t_; } // 14 shfl
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r3 = r3 ^ t_; } // 15 shfl
|
||||
r2 = r2 | r7; // 16 or
|
||||
r0 = rotr_var(r0, r6); // 17 rotr
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
|
||||
r7 = rotl_imm(r7, 24u); // 19 rotl
|
||||
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
|
||||
r2 = r2 | r6; // 21 or
|
||||
r6 = r5 * r4 + r6; // 22 mad
|
||||
r1 = r1 | r0; // 23 or
|
||||
r6 = r6 ^ r1; // 24 xor
|
||||
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
|
||||
r7 = rotr_var(r7, r0); // 26 rotr
|
||||
r4 = r5 * r7 + r4; // 27 mad
|
||||
r3 = r2 * r4 + r3; // 28 mad
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
|
||||
r2 = r2 ^ ds[((rotl_imm(r3 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r5 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
|
||||
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
|
||||
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
|
||||
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
|
||||
r2 = mul_hi(r2, r5); // 35 mulhi
|
||||
r4 = r4 ^ r2; // 36 xor
|
||||
r6 = r6 * r5; // 37 mul
|
||||
r7 = r7 ^ r0; // 38 xor
|
||||
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
|
||||
r2 = rotr_var(r2, r3); // 40 rotr
|
||||
r7 = r7 - r0; // 41 sub
|
||||
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 43 shfl
|
||||
r0 = r0 ^ ds[((rotl_imm(r7 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
|
||||
r3 = r3 ^ ds[((rotl_imm(r1 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
|
||||
r6 = r6 ^ ds[((rotl_imm(r3 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
|
||||
r4 = mul_hi(r4, r2); // 48 mulhi
|
||||
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r0 = r0 ^ t_; } // 50 shfl
|
||||
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r2 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
|
||||
r6 = rotl_imm(r6, 12u); // 53 rotl
|
||||
r3 = rotl_imm(r3, 12u); // 54 rotl
|
||||
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
|
||||
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r0 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
|
||||
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
|
||||
r3 = mul_hi(r3, r2); // 60 mulhi
|
||||
r6 = r6 | r4; // 61 or
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r5 = r5 ^ t_; } // 62 shfl
|
||||
r3 = r3 ^ ds[((rotl_imm(r6 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
|
||||
#if IGNEUM_EXCHANGE != 0
|
||||
// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the
|
||||
// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a
|
||||
// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md.
|
||||
IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) {
|
||||
if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); }
|
||||
}
|
||||
#endif
|
||||
163
proto-cuda/packs-ca2-era/era-2/kernel.cu
Normal file
163
proto-cuda/packs-ca2-era/era-2/kernel.cu
Normal file
|
|
@ -0,0 +1,163 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal).
|
||||
// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC.
|
||||
#include <cuda_runtime.h>
|
||||
#include <cstdint>
|
||||
#include "program.h"
|
||||
#include "memhard.h"
|
||||
|
||||
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31.
|
||||
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
|
||||
// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x.
|
||||
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) {
|
||||
uint32_t x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item.
|
||||
// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host.
|
||||
__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) {
|
||||
uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
if (seg < nSegments) mh_cache_segment(cache, seg);
|
||||
}
|
||||
__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
|
||||
uint32_t t = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
if (t < nItems) {
|
||||
uint32_t s[16];
|
||||
mh_item(cache, t, s);
|
||||
for (uint32_t i = 0u; i < 16u; ++i) ds[(size_t)mh_addr(t, i)] = s[i];
|
||||
}
|
||||
}
|
||||
|
||||
// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every
|
||||
// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a
|
||||
// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid.
|
||||
__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) {
|
||||
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
uint32_t nonce = baseNonce + gid;
|
||||
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint32_t x = nonce ^ 0x667d0fbdu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x7b8e5963u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
|
||||
{ uint32_t x = nonce ^ 0x7b8e5963u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x31c67e5eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
|
||||
{ uint32_t x = nonce ^ 0x31c67e5eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4529ddc6u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
|
||||
{ uint32_t x = nonce ^ 0x4529ddc6u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef19d6d8u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
|
||||
{ uint32_t x = nonce ^ 0xef19d6d8u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xaccf6211u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
|
||||
{ uint32_t x = nonce ^ 0xaccf6211u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xda0aed32u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
|
||||
{ uint32_t x = nonce ^ 0xda0aed32u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xabc6df31u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
|
||||
{ uint32_t x = nonce ^ 0xabc6df31u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x667d0fbdu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
|
||||
|
||||
for (uint32_t it = 0u; it < 8u; ++it) {
|
||||
uint32_t sel = r0;
|
||||
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
|
||||
r7 = r7 ^ r0; // 1 xor
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r6, 1); // 2 shfl
|
||||
r4 = r4 - r1; // 3 sub
|
||||
r2 = r0 * r4 + r2; // 4 mad
|
||||
r4 = rotl_imm(r4, 9u); // 5 rotl
|
||||
r0 = rotr_var(r0, r2); // 6 rotr
|
||||
r6 = r6 ^ ds[((rotl_imm(r7 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r2 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r0 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r1 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
|
||||
r5 = r5 * r4; // 12 mul
|
||||
r7 = r7 ^ ds[((rotl_imm(r6 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
|
||||
r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r5, 16); // 14 shfl
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r5, 1); // 15 shfl
|
||||
r2 = r2 | r7; // 16 or
|
||||
r0 = rotr_var(r0, r6); // 17 rotr
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
|
||||
r7 = rotl_imm(r7, 24u); // 19 rotl
|
||||
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
|
||||
r2 = r2 | r6; // 21 or
|
||||
r6 = r5 * r4 + r6; // 22 mad
|
||||
r1 = r1 | r0; // 23 or
|
||||
r6 = r6 ^ r1; // 24 xor
|
||||
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
|
||||
r7 = rotr_var(r7, r0); // 26 rotr
|
||||
r4 = r5 * r7 + r4; // 27 mad
|
||||
r3 = r2 * r4 + r3; // 28 mad
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
|
||||
r2 = r2 ^ ds[((rotl_imm(r3 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r5 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
|
||||
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
|
||||
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
|
||||
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
|
||||
r2 = __umulhi(r2, r5); // 35 mulhi
|
||||
r4 = r4 ^ r2; // 36 xor
|
||||
r6 = r6 * r5; // 37 mul
|
||||
r7 = r7 ^ r0; // 38 xor
|
||||
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
|
||||
r2 = rotr_var(r2, r3); // 40 rotr
|
||||
r7 = r7 - r0; // 41 sub
|
||||
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 43 shfl
|
||||
r0 = r0 ^ ds[((rotl_imm(r7 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
|
||||
r3 = r3 ^ ds[((rotl_imm(r1 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
|
||||
r6 = r6 ^ ds[((rotl_imm(r3 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
|
||||
r4 = __umulhi(r4, r2); // 48 mulhi
|
||||
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
|
||||
r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r7, 4); // 50 shfl
|
||||
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r2 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
|
||||
r6 = rotl_imm(r6, 12u); // 53 rotl
|
||||
r3 = rotl_imm(r3, 12u); // 54 rotl
|
||||
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
|
||||
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r0 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
|
||||
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
|
||||
r3 = __umulhi(r3, r2); // 60 mulhi
|
||||
r6 = r6 | r4; // 61 or
|
||||
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 62 shfl
|
||||
r3 = r3 ^ ds[((rotl_imm(r6 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
|
||||
}
|
||||
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
|
||||
}
|
||||
|
||||
// Host-side launch wrappers. Declared in program.h, called from host.cu.
|
||||
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) {
|
||||
if (nSegments == 0u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 256u;
|
||||
uint32_t grid = (nSegments + block - 1u) / block;
|
||||
igneum_cache_fill<<<grid, block>>>(cache, nSegments);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
|
||||
if (nItems == 0u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 256u;
|
||||
uint32_t grid = (nItems + block - 1u) / block;
|
||||
igneum_build<<<grid, block>>>(ds, cache, nItems);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
uint32_t nonces, uint32_t blockWarps) {
|
||||
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 32u * blockWarps;
|
||||
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
|
||||
igneum_hash<<<nonces / block, block>>>(ds, out, baseNonce, mask);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
|
||||
cudaFuncAttributes attr;
|
||||
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash);
|
||||
if (e != cudaSuccess) return e;
|
||||
*numRegs = attr.numRegs;
|
||||
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0);
|
||||
}
|
||||
375
proto-cuda/packs-ca2-era/era-2/kernel_bound.cl
Normal file
375
proto-cuda/packs-ca2-era/era-2/kernel_bound.cl
Normal file
|
|
@ -0,0 +1,375 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal).
|
||||
// Built from source at runtime by proto-opencl/host.c, which passes these defines:
|
||||
// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit)
|
||||
// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default)
|
||||
// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32
|
||||
// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition
|
||||
// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units;
|
||||
// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32.
|
||||
#ifndef IGNEUM_GROUP
|
||||
#define IGNEUM_GROUP 32
|
||||
#endif
|
||||
#ifndef IGNEUM_EXCHANGE
|
||||
#define IGNEUM_EXCHANGE 0
|
||||
#endif
|
||||
#ifdef __OPENCL_VERSION__
|
||||
#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1)))
|
||||
#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n]
|
||||
#if IGNEUM_EXCHANGE == 1
|
||||
#ifdef cl_khr_subgroups
|
||||
#pragma OPENCL EXTENSION cl_khr_subgroups : enable
|
||||
#endif
|
||||
#ifdef cl_khr_subgroup_shuffle
|
||||
#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable
|
||||
#endif
|
||||
#elif IGNEUM_EXCHANGE == 2
|
||||
#pragma OPENCL EXTENSION cl_intel_subgroups : enable
|
||||
#endif
|
||||
#else
|
||||
// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros.
|
||||
#include "emu_opencl.h"
|
||||
#endif
|
||||
|
||||
#if IGNEUM_EXCHANGE == 1
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m))
|
||||
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
|
||||
#elif IGNEUM_EXCHANGE == 2
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m))
|
||||
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
|
||||
#else
|
||||
// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per
|
||||
// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane
|
||||
// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's
|
||||
// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier.
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; }
|
||||
#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; }
|
||||
#endif
|
||||
|
||||
static inline uint splitmix32(uint x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32.
|
||||
static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); }
|
||||
// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x.
|
||||
static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); }
|
||||
static inline uint ds_elem(uint i, uint d0, uint d1) {
|
||||
uint x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
|
||||
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
|
||||
#define MH_CACHE_LINE_MASK 0x003fffffu
|
||||
#define MH_SEGMENT_LINES 64u
|
||||
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
|
||||
static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
|
||||
|
||||
// y = ChaCha12 core(x) + x
|
||||
static inline void mh_chacha_block(const uint* x, uint* y) {
|
||||
for (uint i = 0u; i < 16u; ++i) y[i] = x[i];
|
||||
for (uint r = 0u; r < 6u; ++r) {
|
||||
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
|
||||
}
|
||||
for (uint i = 0u; i < 16u; ++i) y[i] += x[i];
|
||||
}
|
||||
|
||||
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
|
||||
static inline void mh_cache_segment(__global uint* cache, uint seg) {
|
||||
uint prev[16]; uint x[16]; uint y[16];
|
||||
for (uint i = 0u; i < 16u; ++i) prev[i] = 0u;
|
||||
for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) {
|
||||
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
|
||||
x[4] = 0xceed56d7u ^ prev[4];
|
||||
x[5] = 0x9ba270d2u ^ prev[5];
|
||||
x[6] = 0x82caab2du ^ prev[6];
|
||||
x[7] = 0x81ebce0eu ^ prev[7];
|
||||
x[8] = 0x12b6ecf1u ^ prev[8];
|
||||
x[9] = 0xd0f3fd7cu ^ prev[9];
|
||||
x[10] = 0xd872eefeu ^ prev[10];
|
||||
x[11] = 0xc158c7bdu ^ prev[11];
|
||||
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
|
||||
mh_chacha_block(x, y);
|
||||
__global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
|
||||
}
|
||||
}
|
||||
|
||||
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
|
||||
static inline void mh_mixer(uint* s, uint rk) {
|
||||
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
|
||||
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
|
||||
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
|
||||
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
|
||||
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
|
||||
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
|
||||
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
|
||||
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
|
||||
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
|
||||
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
|
||||
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
|
||||
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
|
||||
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
|
||||
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
|
||||
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
|
||||
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
|
||||
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
|
||||
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
|
||||
}
|
||||
|
||||
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
|
||||
static inline void mh_item(__global const uint* cache, uint t, uint* s) {
|
||||
s[0] = 0xceed56d7u;
|
||||
s[1] = 0x9ba270d2u;
|
||||
s[2] = 0x82caab2du;
|
||||
s[3] = 0x81ebce0eu;
|
||||
s[4] = 0x12b6ecf1u;
|
||||
s[5] = 0xd0f3fd7cu;
|
||||
s[6] = 0xd872eefeu;
|
||||
s[7] = 0xc158c7bdu;
|
||||
s[8] = t * 0xf351d601u + 0xc6892460u;
|
||||
s[9] = t * 0xa3bb398fu + 0x25b7228au;
|
||||
s[10] = t * 0xb5a09e35u + 0xcd515004u;
|
||||
s[11] = t * 0x7509c9c1u + 0x2846527au;
|
||||
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
|
||||
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
|
||||
s[14] = t * 0xded91851u + 0x82961bacu;
|
||||
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
|
||||
for (uint r = 0u; r < 8u; ++r) {
|
||||
mh_mixer(s, 0x9E3779B9u * (r + 1u));
|
||||
__global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i];
|
||||
}
|
||||
mh_mixer(s, 0x9E3779B9u * 9u);
|
||||
}
|
||||
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 1 3 4 8 of w.
|
||||
static inline uint mh_j(uint w) { return ((w >> 1u) & 1u) | (((w >> 3u) & 1u) << 1) | (((w >> 4u) & 1u) << 2) | (((w >> 8u) & 1u) << 3); }
|
||||
static inline uint mh_t(uint w) { w = (w & 0x000000ffu) | ((w >> 9u) << 8u); w = (w & 0x0000000fu) | ((w >> 5u) << 4u); w = (w & 0x00000007u) | ((w >> 4u) << 3u); w = (w & 0x00000001u) | ((w >> 2u) << 1u); return w; }
|
||||
static inline uint mh_addr(uint t, uint j) { uint w = t; w = ((w >> 1u) << 2u) | (w & 0x00000001u) | (((j >> 0u) & 1u) << 1u); w = ((w >> 3u) << 4u) | (w & 0x00000007u) | (((j >> 1u) & 1u) << 3u); w = ((w >> 4u) << 5u) | (w & 0x0000000fu) | (((j >> 2u) & 1u) << 4u); w = ((w >> 8u) << 9u) | (w & 0x000000ffu) | (((j >> 3u) & 1u) << 8u); return w; }
|
||||
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
|
||||
static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }
|
||||
|
||||
// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item.
|
||||
// The same constants as memhard.h in this pack (one emitter, three dialects).
|
||||
__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) {
|
||||
uint seg = (uint)get_global_id(0);
|
||||
if (seg < nSegments) mh_cache_segment(cache, seg);
|
||||
}
|
||||
__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) {
|
||||
uint t = (uint)get_global_id(0);
|
||||
if (t < nItems) {
|
||||
uint s[16];
|
||||
mh_item(cache, t, s);
|
||||
for (uint i = 0u; i < 16u; ++i) ds[(ulong)mh_addr(t, i)] = s[i];
|
||||
}
|
||||
}
|
||||
|
||||
// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the
|
||||
// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and
|
||||
// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all).
|
||||
IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) {
|
||||
uint gid = (uint)get_global_id(0);
|
||||
uint lid = (uint)get_local_id(0);
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
#if IGNEUM_EXCHANGE == 0
|
||||
IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
|
||||
uint xk = 0u;
|
||||
#else
|
||||
(void)lid;
|
||||
#endif
|
||||
{ uint x = nonce ^ 0x667d0fbdu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x7b8e5963u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
|
||||
{ uint x = nonce ^ 0x7b8e5963u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x31c67e5eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
|
||||
{ uint x = nonce ^ 0x31c67e5eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4529ddc6u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
|
||||
{ uint x = nonce ^ 0x4529ddc6u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef19d6d8u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
|
||||
{ uint x = nonce ^ 0xef19d6d8u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xaccf6211u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
|
||||
{ uint x = nonce ^ 0xaccf6211u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xda0aed32u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
|
||||
{ uint x = nonce ^ 0xda0aed32u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xabc6df31u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
|
||||
{ uint x = nonce ^ 0xabc6df31u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x667d0fbdu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
|
||||
r7 = r7 ^ r0; // 1 xor
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 1u); r3 = r3 ^ t_; } // 2 shfl
|
||||
r4 = r4 - r1; // 3 sub
|
||||
r2 = r0 * r4 + r2; // 4 mad
|
||||
r4 = rotl_imm(r4, 9u); // 5 rotl
|
||||
r0 = rotr_var(r0, r2); // 6 rotr
|
||||
r6 = r6 ^ ds[((rotl_imm(r7 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r2 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r0 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r1 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
|
||||
r5 = r5 * r4; // 12 mul
|
||||
r7 = r7 ^ ds[((rotl_imm(r6 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 16u); r6 = r6 ^ t_; } // 14 shfl
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r3 = r3 ^ t_; } // 15 shfl
|
||||
r2 = r2 | r7; // 16 or
|
||||
r0 = rotr_var(r0, r6); // 17 rotr
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
|
||||
r7 = rotl_imm(r7, 24u); // 19 rotl
|
||||
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
|
||||
r2 = r2 | r6; // 21 or
|
||||
r6 = r5 * r4 + r6; // 22 mad
|
||||
r1 = r1 | r0; // 23 or
|
||||
r6 = r6 ^ r1; // 24 xor
|
||||
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
|
||||
r7 = rotr_var(r7, r0); // 26 rotr
|
||||
r4 = r5 * r7 + r4; // 27 mad
|
||||
r3 = r2 * r4 + r3; // 28 mad
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
|
||||
r2 = r2 ^ ds[((rotl_imm(r3 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r5 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
|
||||
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
|
||||
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
|
||||
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
|
||||
r2 = mul_hi(r2, r5); // 35 mulhi
|
||||
r4 = r4 ^ r2; // 36 xor
|
||||
r6 = r6 * r5; // 37 mul
|
||||
r7 = r7 ^ r0; // 38 xor
|
||||
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
|
||||
r2 = rotr_var(r2, r3); // 40 rotr
|
||||
r7 = r7 - r0; // 41 sub
|
||||
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 43 shfl
|
||||
r0 = r0 ^ ds[((rotl_imm(r7 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
|
||||
r3 = r3 ^ ds[((rotl_imm(r1 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
|
||||
r6 = r6 ^ ds[((rotl_imm(r3 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
|
||||
r4 = mul_hi(r4, r2); // 48 mulhi
|
||||
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r0 = r0 ^ t_; } // 50 shfl
|
||||
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r2 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
|
||||
r6 = rotl_imm(r6, 12u); // 53 rotl
|
||||
r3 = rotl_imm(r3, 12u); // 54 rotl
|
||||
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
|
||||
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r0 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
|
||||
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
|
||||
r3 = mul_hi(r3, r2); // 60 mulhi
|
||||
r6 = r6 | r4; // 61 or
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r5 = r5 ^ t_; } // 62 shfl
|
||||
r3 = r3 ^ ds[((rotl_imm(r6 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
|
||||
#if IGNEUM_EXCHANGE != 0
|
||||
// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the
|
||||
// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a
|
||||
// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md.
|
||||
IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) {
|
||||
if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); }
|
||||
}
|
||||
#endif
|
||||
|
||||
// Header-bound variant (bind.rs): the init words come from initw, not SEEDW. Same body as igneum_hash.
|
||||
IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw) {
|
||||
uint gid = (uint)get_global_id(0);
|
||||
uint lid = (uint)get_local_id(0);
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7];
|
||||
#if IGNEUM_EXCHANGE == 0
|
||||
IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
|
||||
uint xk = 0u;
|
||||
#else
|
||||
(void)lid;
|
||||
#endif
|
||||
{ uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; }
|
||||
{ uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; }
|
||||
{ uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; }
|
||||
{ uint x = nonce ^ iw3; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw4; }
|
||||
{ uint x = nonce ^ iw4; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw5; }
|
||||
{ uint x = nonce ^ iw5; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw6; }
|
||||
{ uint x = nonce ^ iw6; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw7; }
|
||||
{ uint x = nonce ^ iw7; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw0; }
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
|
||||
r7 = r7 ^ r0; // 1 xor
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 1u); r3 = r3 ^ t_; } // 2 shfl
|
||||
r4 = r4 - r1; // 3 sub
|
||||
r2 = r0 * r4 + r2; // 4 mad
|
||||
r4 = rotl_imm(r4, 9u); // 5 rotl
|
||||
r0 = rotr_var(r0, r2); // 6 rotr
|
||||
r6 = r6 ^ ds[((rotl_imm(r7 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r2 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r0 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r1 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
|
||||
r5 = r5 * r4; // 12 mul
|
||||
r7 = r7 ^ ds[((rotl_imm(r6 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 16u); r6 = r6 ^ t_; } // 14 shfl
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r3 = r3 ^ t_; } // 15 shfl
|
||||
r2 = r2 | r7; // 16 or
|
||||
r0 = rotr_var(r0, r6); // 17 rotr
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
|
||||
r7 = rotl_imm(r7, 24u); // 19 rotl
|
||||
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
|
||||
r2 = r2 | r6; // 21 or
|
||||
r6 = r5 * r4 + r6; // 22 mad
|
||||
r1 = r1 | r0; // 23 or
|
||||
r6 = r6 ^ r1; // 24 xor
|
||||
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
|
||||
r7 = rotr_var(r7, r0); // 26 rotr
|
||||
r4 = r5 * r7 + r4; // 27 mad
|
||||
r3 = r2 * r4 + r3; // 28 mad
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
|
||||
r2 = r2 ^ ds[((rotl_imm(r3 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r5 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
|
||||
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
|
||||
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
|
||||
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
|
||||
r2 = mul_hi(r2, r5); // 35 mulhi
|
||||
r4 = r4 ^ r2; // 36 xor
|
||||
r6 = r6 * r5; // 37 mul
|
||||
r7 = r7 ^ r0; // 38 xor
|
||||
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
|
||||
r2 = rotr_var(r2, r3); // 40 rotr
|
||||
r7 = r7 - r0; // 41 sub
|
||||
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 43 shfl
|
||||
r0 = r0 ^ ds[((rotl_imm(r7 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
|
||||
r3 = r3 ^ ds[((rotl_imm(r1 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
|
||||
r6 = r6 ^ ds[((rotl_imm(r3 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
|
||||
r4 = mul_hi(r4, r2); // 48 mulhi
|
||||
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r0 = r0 ^ t_; } // 50 shfl
|
||||
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r2 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
|
||||
r6 = rotl_imm(r6, 12u); // 53 rotl
|
||||
r3 = rotl_imm(r3, 12u); // 54 rotl
|
||||
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
|
||||
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r0 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
|
||||
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
|
||||
r3 = mul_hi(r3, r2); // 60 mulhi
|
||||
r6 = r6 | r4; // 61 or
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r5 = r5 ^ t_; } // 62 shfl
|
||||
r3 = r3 ^ ds[((rotl_imm(r6 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
123
proto-cuda/packs-ca2-era/era-2/kernel_bound.cu
Normal file
123
proto-cuda/packs-ca2-era/era-2/kernel_bound.cu
Normal file
|
|
@ -0,0 +1,123 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// Header-bound twin of igneum_hash in kernel.cu: the init words come from a kernel argument, not SEEDW.
|
||||
// Host declarations (also in program_bound.h if present):
|
||||
// struct IgneumInitWords { uint32_t w[8]; };
|
||||
// cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
// IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps);
|
||||
// cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
|
||||
#include <cuda_runtime.h>
|
||||
#include <cstdint>
|
||||
#include "program.h"
|
||||
|
||||
struct IgneumInitWords { uint32_t w[8]; };
|
||||
|
||||
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
|
||||
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
|
||||
__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw) {
|
||||
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
uint32_t nonce = baseNonce + gid;
|
||||
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint32_t x = nonce ^ iw.w[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw.w[1]; }
|
||||
{ uint32_t x = nonce ^ iw.w[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw.w[2]; }
|
||||
{ uint32_t x = nonce ^ iw.w[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw.w[3]; }
|
||||
{ uint32_t x = nonce ^ iw.w[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw.w[4]; }
|
||||
{ uint32_t x = nonce ^ iw.w[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw.w[5]; }
|
||||
{ uint32_t x = nonce ^ iw.w[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw.w[6]; }
|
||||
{ uint32_t x = nonce ^ iw.w[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw.w[7]; }
|
||||
{ uint32_t x = nonce ^ iw.w[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw.w[0]; }
|
||||
|
||||
for (uint32_t it = 0u; it < 8u; ++it) {
|
||||
uint32_t sel = r0;
|
||||
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
|
||||
r7 = r7 ^ r0; // 1 xor
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r6, 1); // 2 shfl
|
||||
r4 = r4 - r1; // 3 sub
|
||||
r2 = r0 * r4 + r2; // 4 mad
|
||||
r4 = rotl_imm(r4, 9u); // 5 rotl
|
||||
r0 = rotr_var(r0, r2); // 6 rotr
|
||||
r6 = r6 ^ ds[((rotl_imm(r7 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r2 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r0 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r1 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
|
||||
r5 = r5 * r4; // 12 mul
|
||||
r7 = r7 ^ ds[((rotl_imm(r6 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
|
||||
r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r5, 16); // 14 shfl
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r5, 1); // 15 shfl
|
||||
r2 = r2 | r7; // 16 or
|
||||
r0 = rotr_var(r0, r6); // 17 rotr
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
|
||||
r7 = rotl_imm(r7, 24u); // 19 rotl
|
||||
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
|
||||
r2 = r2 | r6; // 21 or
|
||||
r6 = r5 * r4 + r6; // 22 mad
|
||||
r1 = r1 | r0; // 23 or
|
||||
r6 = r6 ^ r1; // 24 xor
|
||||
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
|
||||
r7 = rotr_var(r7, r0); // 26 rotr
|
||||
r4 = r5 * r7 + r4; // 27 mad
|
||||
r3 = r2 * r4 + r3; // 28 mad
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
|
||||
r2 = r2 ^ ds[((rotl_imm(r3 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r5 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
|
||||
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
|
||||
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
|
||||
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
|
||||
r2 = __umulhi(r2, r5); // 35 mulhi
|
||||
r4 = r4 ^ r2; // 36 xor
|
||||
r6 = r6 * r5; // 37 mul
|
||||
r7 = r7 ^ r0; // 38 xor
|
||||
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
|
||||
r2 = rotr_var(r2, r3); // 40 rotr
|
||||
r7 = r7 - r0; // 41 sub
|
||||
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 43 shfl
|
||||
r0 = r0 ^ ds[((rotl_imm(r7 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
|
||||
r3 = r3 ^ ds[((rotl_imm(r1 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
|
||||
r6 = r6 ^ ds[((rotl_imm(r3 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
|
||||
r4 = __umulhi(r4, r2); // 48 mulhi
|
||||
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
|
||||
r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r7, 4); // 50 shfl
|
||||
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r2 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
|
||||
r6 = rotl_imm(r6, 12u); // 53 rotl
|
||||
r3 = rotl_imm(r3, 12u); // 54 rotl
|
||||
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
|
||||
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r0 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
|
||||
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
|
||||
r3 = __umulhi(r3, r2); // 60 mulhi
|
||||
r6 = r6 | r4; // 61 or
|
||||
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 62 shfl
|
||||
r3 = r3 ^ ds[((rotl_imm(r6 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
|
||||
}
|
||||
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps) {
|
||||
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 32u * blockWarps;
|
||||
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
|
||||
igneum_hash_bound<<<nonces / block, block>>>(ds, out, baseNonce, mask, iw);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
|
||||
cudaFuncAttributes attr;
|
||||
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash_bound);
|
||||
if (e != cudaSuccess) return e;
|
||||
*numRegs = attr.numRegs;
|
||||
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash_bound, (int)(32u * blockWarps), 0);
|
||||
}
|
||||
112
proto-cuda/packs-ca2-era/era-2/memhard.h
Normal file
112
proto-cuda/packs-ca2-era/era-2/memhard.h
Normal file
|
|
@ -0,0 +1,112 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// Memory-hard dataset core, the same text that the Mac's Metal kernels and CPU verifier were checked against.
|
||||
// Included by kernel.cu (device), host.cu (host reference) and proto-opencl/host.c (C99 host reference).
|
||||
// See proto-metal/MEMHARD.md for the construction. kernel.cl carries the same text in OpenCL C.
|
||||
#pragma once
|
||||
#ifdef __cplusplus
|
||||
#include <cstdint>
|
||||
#else
|
||||
#include <stdint.h>
|
||||
#endif
|
||||
#if defined(__CUDACC__)
|
||||
#define IGNEUM_HD __host__ __device__ __forceinline__
|
||||
#elif defined(_MSC_VER) && !defined(__cplusplus)
|
||||
#define IGNEUM_HD static __inline
|
||||
#else
|
||||
#define IGNEUM_HD static inline
|
||||
#endif
|
||||
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
|
||||
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
|
||||
#define MH_CACHE_LINE_MASK 0x003fffffu
|
||||
#define MH_SEGMENT_LINES 64u
|
||||
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
|
||||
IGNEUM_HD uint32_t mh_rotl(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
|
||||
|
||||
// y = ChaCha12 core(x) + x
|
||||
IGNEUM_HD void mh_chacha_block(const uint32_t* x, uint32_t* y) {
|
||||
for (uint32_t i = 0u; i < 16u; ++i) y[i] = x[i];
|
||||
for (uint32_t r = 0u; r < 6u; ++r) {
|
||||
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
|
||||
}
|
||||
for (uint32_t i = 0u; i < 16u; ++i) y[i] += x[i];
|
||||
}
|
||||
|
||||
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
|
||||
IGNEUM_HD void mh_cache_segment(uint32_t* cache, uint32_t seg) {
|
||||
uint32_t prev[16]; uint32_t x[16]; uint32_t y[16];
|
||||
for (uint32_t i = 0u; i < 16u; ++i) prev[i] = 0u;
|
||||
for (uint32_t j = 0u; j < MH_SEGMENT_LINES; ++j) {
|
||||
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
|
||||
x[4] = 0xceed56d7u ^ prev[4];
|
||||
x[5] = 0x9ba270d2u ^ prev[5];
|
||||
x[6] = 0x82caab2du ^ prev[6];
|
||||
x[7] = 0x81ebce0eu ^ prev[7];
|
||||
x[8] = 0x12b6ecf1u ^ prev[8];
|
||||
x[9] = 0xd0f3fd7cu ^ prev[9];
|
||||
x[10] = 0xd872eefeu ^ prev[10];
|
||||
x[11] = 0xc158c7bdu ^ prev[11];
|
||||
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
|
||||
mh_chacha_block(x, y);
|
||||
uint32_t* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
|
||||
for (uint32_t i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
|
||||
}
|
||||
}
|
||||
|
||||
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
|
||||
IGNEUM_HD void mh_mixer(uint32_t* s, uint32_t rk) {
|
||||
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
|
||||
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
|
||||
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
|
||||
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
|
||||
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
|
||||
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
|
||||
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
|
||||
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
|
||||
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
|
||||
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
|
||||
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
|
||||
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
|
||||
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
|
||||
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
|
||||
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
|
||||
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
|
||||
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
|
||||
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
|
||||
}
|
||||
|
||||
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
|
||||
IGNEUM_HD void mh_item(const uint32_t* cache, uint32_t t, uint32_t* s) {
|
||||
s[0] = 0xceed56d7u;
|
||||
s[1] = 0x9ba270d2u;
|
||||
s[2] = 0x82caab2du;
|
||||
s[3] = 0x81ebce0eu;
|
||||
s[4] = 0x12b6ecf1u;
|
||||
s[5] = 0xd0f3fd7cu;
|
||||
s[6] = 0xd872eefeu;
|
||||
s[7] = 0xc158c7bdu;
|
||||
s[8] = t * 0xf351d601u + 0xc6892460u;
|
||||
s[9] = t * 0xa3bb398fu + 0x25b7228au;
|
||||
s[10] = t * 0xb5a09e35u + 0xcd515004u;
|
||||
s[11] = t * 0x7509c9c1u + 0x2846527au;
|
||||
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
|
||||
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
|
||||
s[14] = t * 0xded91851u + 0x82961bacu;
|
||||
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
|
||||
for (uint32_t r = 0u; r < 8u; ++r) {
|
||||
mh_mixer(s, 0x9E3779B9u * (r + 1u));
|
||||
const uint32_t* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
|
||||
for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= line[i];
|
||||
}
|
||||
mh_mixer(s, 0x9E3779B9u * 9u);
|
||||
}
|
||||
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 1 3 4 8 of w.
|
||||
IGNEUM_HD uint32_t mh_j(uint32_t w) { return ((w >> 1u) & 1u) | (((w >> 3u) & 1u) << 1) | (((w >> 4u) & 1u) << 2) | (((w >> 8u) & 1u) << 3); }
|
||||
IGNEUM_HD uint32_t mh_t(uint32_t w) { w = (w & 0x000000ffu) | ((w >> 9u) << 8u); w = (w & 0x0000000fu) | ((w >> 5u) << 4u); w = (w & 0x00000007u) | ((w >> 4u) << 3u); w = (w & 0x00000001u) | ((w >> 2u) << 1u); return w; }
|
||||
IGNEUM_HD uint32_t mh_addr(uint32_t t, uint32_t j) { uint32_t w = t; w = ((w >> 1u) << 2u) | (w & 0x00000001u) | (((j >> 0u) & 1u) << 1u); w = ((w >> 3u) << 4u) | (w & 0x00000007u) | (((j >> 1u) & 1u) << 3u); w = ((w >> 4u) << 5u) | (w & 0x0000000fu) | (((j >> 2u) & 1u) << 4u); w = ((w >> 8u) << 9u) | (w & 0x000000ffu) | (((j >> 3u) & 1u) << 8u); return w; }
|
||||
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
|
||||
IGNEUM_HD uint32_t mh_word(const uint32_t* cache, uint32_t w) { uint32_t s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }
|
||||
109
proto-cuda/packs-ca2-era/era-2/memhard.metal
Normal file
109
proto-cuda/packs-ca2-era/era-2/memhard.metal
Normal file
|
|
@ -0,0 +1,109 @@
|
|||
#include <metal_stdlib>
|
||||
using namespace metal;
|
||||
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
|
||||
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
|
||||
#define MH_CACHE_LINE_MASK 0x003fffffu
|
||||
#define MH_SEGMENT_LINES 64u
|
||||
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
|
||||
inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
|
||||
|
||||
// y = ChaCha12 core(x) + x
|
||||
inline void mh_chacha_block(const thread uint* x, thread uint* y) {
|
||||
for (uint i = 0u; i < 16u; ++i) y[i] = x[i];
|
||||
for (uint r = 0u; r < 6u; ++r) {
|
||||
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
|
||||
}
|
||||
for (uint i = 0u; i < 16u; ++i) y[i] += x[i];
|
||||
}
|
||||
|
||||
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
|
||||
inline void mh_cache_segment(device uint* cache, uint seg) {
|
||||
uint prev[16]; uint x[16]; uint y[16];
|
||||
for (uint i = 0u; i < 16u; ++i) prev[i] = 0u;
|
||||
for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) {
|
||||
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
|
||||
x[4] = 0xceed56d7u ^ prev[4];
|
||||
x[5] = 0x9ba270d2u ^ prev[5];
|
||||
x[6] = 0x82caab2du ^ prev[6];
|
||||
x[7] = 0x81ebce0eu ^ prev[7];
|
||||
x[8] = 0x12b6ecf1u ^ prev[8];
|
||||
x[9] = 0xd0f3fd7cu ^ prev[9];
|
||||
x[10] = 0xd872eefeu ^ prev[10];
|
||||
x[11] = 0xc158c7bdu ^ prev[11];
|
||||
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
|
||||
mh_chacha_block(x, y);
|
||||
device uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
|
||||
}
|
||||
}
|
||||
|
||||
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
|
||||
inline void mh_mixer(thread uint* s, uint rk) {
|
||||
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
|
||||
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
|
||||
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
|
||||
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
|
||||
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
|
||||
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
|
||||
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
|
||||
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
|
||||
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
|
||||
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
|
||||
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
|
||||
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
|
||||
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
|
||||
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
|
||||
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
|
||||
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
|
||||
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
|
||||
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
|
||||
}
|
||||
|
||||
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
|
||||
inline void mh_item(device const uint* cache, uint t, thread uint* s) {
|
||||
s[0] = 0xceed56d7u;
|
||||
s[1] = 0x9ba270d2u;
|
||||
s[2] = 0x82caab2du;
|
||||
s[3] = 0x81ebce0eu;
|
||||
s[4] = 0x12b6ecf1u;
|
||||
s[5] = 0xd0f3fd7cu;
|
||||
s[6] = 0xd872eefeu;
|
||||
s[7] = 0xc158c7bdu;
|
||||
s[8] = t * 0xf351d601u + 0xc6892460u;
|
||||
s[9] = t * 0xa3bb398fu + 0x25b7228au;
|
||||
s[10] = t * 0xb5a09e35u + 0xcd515004u;
|
||||
s[11] = t * 0x7509c9c1u + 0x2846527au;
|
||||
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
|
||||
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
|
||||
s[14] = t * 0xded91851u + 0x82961bacu;
|
||||
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
|
||||
for (uint r = 0u; r < 8u; ++r) {
|
||||
mh_mixer(s, 0x9E3779B9u * (r + 1u));
|
||||
device const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i];
|
||||
}
|
||||
mh_mixer(s, 0x9E3779B9u * 9u);
|
||||
}
|
||||
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 1 3 4 8 of w.
|
||||
inline uint mh_j(uint w) { return ((w >> 1u) & 1u) | (((w >> 3u) & 1u) << 1) | (((w >> 4u) & 1u) << 2) | (((w >> 8u) & 1u) << 3); }
|
||||
inline uint mh_t(uint w) { w = (w & 0x000000ffu) | ((w >> 9u) << 8u); w = (w & 0x0000000fu) | ((w >> 5u) << 4u); w = (w & 0x00000007u) | ((w >> 4u) << 3u); w = (w & 0x00000001u) | ((w >> 2u) << 1u); return w; }
|
||||
inline uint mh_addr(uint t, uint j) { uint w = t; w = ((w >> 1u) << 2u) | (w & 0x00000001u) | (((j >> 0u) & 1u) << 1u); w = ((w >> 3u) << 4u) | (w & 0x00000007u) | (((j >> 1u) & 1u) << 3u); w = ((w >> 4u) << 5u) | (w & 0x0000000fu) | (((j >> 2u) & 1u) << 4u); w = ((w >> 8u) << 9u) | (w & 0x000000ffu) | (((j >> 3u) & 1u) << 8u); return w; }
|
||||
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
|
||||
inline uint mh_word(device const uint* cache, uint w) { uint s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }
|
||||
|
||||
// One thread per segment (2^16 threads).
|
||||
kernel void igneum_cache_fill(device uint* cache [[buffer(0)]], uint gid [[thread_position_in_grid]]) {
|
||||
mh_cache_segment(cache, gid);
|
||||
}
|
||||
// One thread per 64-byte item (dataset words / 16 threads).
|
||||
kernel void igneum_build(device const uint* cache [[buffer(0)]], device uint* dataset [[buffer(1)]],
|
||||
uint gid [[thread_position_in_grid]]) {
|
||||
uint s[16];
|
||||
mh_item(cache, gid, s);
|
||||
for (uint i = 0u; i < 16u; ++i) dataset[mh_addr(gid, i)] = s[i];
|
||||
}
|
||||
76
proto-cuda/packs-ca2-era/era-2/program.h
Normal file
76
proto-cuda/packs-ca2-era/era-2/program.h
Normal file
|
|
@ -0,0 +1,76 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// Program metadata for host.cu plus the launch wrappers defined in kernel.cu.
|
||||
// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros.
|
||||
#pragma once
|
||||
#ifdef __cplusplus
|
||||
#include <cstdint>
|
||||
#else
|
||||
#include <stdint.h>
|
||||
#endif
|
||||
#ifndef IGNEUM_NO_CUDA
|
||||
#include <cuda_runtime.h>
|
||||
#endif
|
||||
|
||||
#define IGNEUM_SEED_STRING "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000"
|
||||
#define IGNEUM_SEED_BYTES_HEX "edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07"
|
||||
#define IGNEUM_GENERATOR 3
|
||||
#define IGNEUM_PROGRAM_ATTEMPT 0
|
||||
#define IGNEUM_PROGRAM_ID 0x73bcbfe8ccf988f1ull
|
||||
#define IGNEUM_DAY_STRING "bytes:69676e65756d2d6461792ffa50000000000000"
|
||||
#define IGNEUM_DAY_BYTES_HEX "69676e65756d2d6461792ffa50000000000000"
|
||||
#define IGNEUM_DAY0 0xceed56d7u
|
||||
#define IGNEUM_DAY1 0x9ba270d2u
|
||||
#define IGNEUM_DATASET_LOG2 28
|
||||
#define IGNEUM_MASK 0x0fffffffu
|
||||
#define IGNEUM_LANES 32
|
||||
#define IGNEUM_ITERATIONS 8
|
||||
#define IGNEUM_INSTR_COUNT 64
|
||||
#define IGNEUM_LOADS_PER_HASH 128
|
||||
#define IGNEUM_WIDE_LOADS_PER_HASH 0
|
||||
#define IGNEUM_OP_MIX "load=16 add=15 shfl=6 mad=4 or=4 rotl=4 rotr=4 xor=4 mulhi=3 mul=2 sub=2"
|
||||
// Program class v3 (Counter ASIC 2.0, docs/plans/counter-asic-2-rollout.md): generator version 3; a worker that
|
||||
// runs another class refuses this pack, and a job line names the class it wants (class=v3 era=<hex>).
|
||||
#define IGNEUM_PROGRAM_CLASS "v3"
|
||||
#define IGNEUM_ERA_SEED_HEX "df57136f2ad5f410e6145023090eea148c1342ccef5a5f541afa590be7e44745"
|
||||
// Class v3 construction (Counter ASIC 2.0, 5 October 2026, docs/plans/mixer-x4.md): version 2 loads; the dataset item
|
||||
// derivation applies the mixer IGNEUM_MIXER_MULT times per round (memhard.h), and the cache follows the growth rule.
|
||||
#define IGNEUM_LOAD_CLASS "w4-era843155d7"
|
||||
#define IGNEUM_LOAD_SLOTS 16
|
||||
#define IGNEUM_LOAD_MIX { 100, 0, 0 }
|
||||
#define IGNEUM_LOAD_WIDTH_COUNTS { 16, 0, 0 } // loads of 4, 16, 64 bytes per program
|
||||
#define IGNEUM_BYTES_PER_HASH 512
|
||||
#define IGNEUM_FOLD_ROT 11
|
||||
#define IGNEUM_FOLD_MUL 0x9e3779b1u
|
||||
// Era layout (5 October 2026, docs/plans/era-layout.md): NOT the lottery hash. Every dataset load reads
|
||||
// idx = ((rotl(src * STRIDE_MUL, STRIDE_ROT) & window mask) | window offset) & MASK; the window of a load site is the
|
||||
// dataset, a half or a quarter of it (IGNEUM_ERA_WINDOWS: site:shrink:offset); dataset word w holds word j(w) of item
|
||||
// t(w) with j's bits at the INTERLEAVE positions (memhard.h: mh_t, mh_j, mh_addr).
|
||||
#define IGNEUM_ERA_LABEL "843155d7"
|
||||
#define IGNEUM_ERA_SEED_WORDS { 0x843155d7u, 0x8fb2bbb8u, 0x3af89788u, 0xc80fe6f2u, 0xb265c95eu, 0x001f1648u, 0x236e8759u, 0x6cada520u }
|
||||
#define IGNEUM_ERA_ALLOWED_WIDTHS { 1, 0, 0 } // words, ascending, 0 = unused; one entry pins the width
|
||||
#define IGNEUM_ERA_WIDTH_WORDS 1
|
||||
#define IGNEUM_ERA_STRIDE_MUL 0x2b4a5b97u
|
||||
#define IGNEUM_ERA_STRIDE_ROT 28
|
||||
#define IGNEUM_ERA_INTERLEAVE { 1, 3, 4, 8 }
|
||||
#define IGNEUM_ERA_WINDOWS "7:2:1 8:1:1 9:1:1 10:1:1 11:0:0 13:1:1 29:0:0 30:2:2 31:1:1 44:1:1 46:2:0 47:0:0 52:0:0 56:0:0 58:2:0 63:1:1"
|
||||
// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h)
|
||||
#define IGNEUM_DATASET_MODE 1
|
||||
|
||||
#define IGNEUM_SEEDW_INIT { 0x667d0fbdu, 0x7b8e5963u, 0x31c67e5eu, 0x4529ddc6u, 0xef19d6d8u, 0xaccf6211u, 0xda0aed32u, 0xabc6df31u }
|
||||
#define IGNEUM_KEY_INIT { 0xceed56d7u, 0x9ba270d2u, 0x82caab2du, 0x81ebce0eu, 0x12b6ecf1u, 0xd0f3fd7cu, 0xd872eefeu, 0xc158c7bdu }
|
||||
#define IGNEUM_CACHE_LOG2_WORDS 26
|
||||
#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6
|
||||
#define IGNEUM_CACHE_SEGMENTS 65536u
|
||||
#define IGNEUM_ITEM_ROUNDS 8
|
||||
#define IGNEUM_MIX_ROT_INIT { 17u, 12u, 20u, 23u, 7u, 3u, 27u, 16u }
|
||||
#define IGNEUM_MIX_MUL_INIT { 0xf351d601u, 0xa3bb398fu, 0xb5a09e35u, 0x7509c9c1u, 0x6bbf31e9u, 0xfc849a79u, 0xded91851u, 0x8d9113d1u, 0x0ff15225u, 0x3a5bdd41u, 0xab533435u, 0xe1c55ad5u, 0xe6d3bd0du, 0x9d9ffbbdu, 0xbb2a3cf3u, 0x50a7c08du }
|
||||
#define IGNEUM_MIX_RC_INIT { 0xc6892460u, 0x25b7228au, 0xcd515004u, 0x2846527au, 0xa6324241u, 0x36e3ec53u, 0x82961bacu, 0x0f97ba7du, 0xb6f921a9u, 0x3ada24e5u, 0xde20ab91u, 0x5378eeb2u, 0x7d161662u, 0x89353cc1u, 0xb1aa03a2u, 0x788acae6u }
|
||||
|
||||
#ifndef IGNEUM_NO_CUDA
|
||||
// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError().
|
||||
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments);
|
||||
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems);
|
||||
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
uint32_t nonces, uint32_t blockWarps);
|
||||
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
|
||||
#endif
|
||||
143
proto-cuda/packs-ca2-era/era-2/program.json
Normal file
143
proto-cuda/packs-ca2-era/era-2/program.json
Normal file
|
|
@ -0,0 +1,143 @@
|
|||
{
|
||||
"format": "igneum-program-pack-3",
|
||||
"generator": 3,
|
||||
"attempt": 0,
|
||||
"program_id": "0x73bcbfe8ccf988f1",
|
||||
"program_id_derivation": "FNV-1a 64 over 'igneum-program/' || generator_le32 || seed_words as little-endian bytes || attempt_le32",
|
||||
"dataset_mode": "memory-hard",
|
||||
"seed": "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000",
|
||||
"seed_bytes": "edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07",
|
||||
"seed_words": ["0x667d0fbd", "0x7b8e5963", "0x31c67e5e", "0x4529ddc6", "0xef19d6d8", "0xaccf6211", "0xda0aed32", "0xabc6df31"],
|
||||
"seed_derivation": "seed_words = FNV-1a 64 over seed_bytes (attempt 0) or seed_bytes || attempt_le32 (attempt k >= 1), basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32",
|
||||
"generator_rule": "version 2: exactly 16 load slots drawn first from instructions 1..63 (partial Fisher-Yates), the other 48 ops from the ten non-load weights (sum 75); a load's source is drawn from the registers other than dst written by an earlier instruction and not read by a load since; the candidate must pass the acceptance rule of spec 01 section 1.4.6 (static: no cyclically stale load source, every register has an injecting write; dynamic: 64 units on the seed-keyed closed-form dataset with no constant register bit, no lane-constant load site, under 164 saturated final values, every output bit within 136 of 1024, distinct addresses above 245760), else the next attempt of the seed is tried",
|
||||
"lanes": 32,
|
||||
"registers": 8,
|
||||
"iterations": 8,
|
||||
"instruction_count": 64,
|
||||
"loads_per_hash": 128,
|
||||
"program_class": "v3",
|
||||
"era_seed_bytes": "df57136f2ad5f410e6145023090eea148c1342ccef5a5f541afa590be7e44745",
|
||||
"load_class": "w4-era843155d7",
|
||||
"load_slots": 16,
|
||||
"load_mix_percent_4_16_64": [100, 0, 0],
|
||||
"load_width_counts_4_16_64": [16, 0, 0],
|
||||
"bytes_per_hash": 512,
|
||||
"wide_load": "read-width experiment (5 October 2026, docs/plans/read-width.md), NOT the lottery hash: a load of W words (width field, 4 or 16) reads dataset[b .. b + W) with b = (src & mask) & ~(W - 1) and folds every word into dst: x = dst ^ w[0]; for j in 1..W: x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x; width 1 is the plain load; the width is drawn per instruction from the class mix with one extra below(100) draw after the nine of version 2, and the program id is FNV-1a 64 over 'igneum-program-rw/' || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots",
|
||||
"era": {
|
||||
"label": "843155d7",
|
||||
"seed_words": ["0x843155d7", "0x8fb2bbb8", "0x3af89788", "0xc80fe6f2", "0xb265c95e", "0x001f1648", "0x236e8759", "0x6cada520"],
|
||||
"draw": "docs/plans/era-layout.md 1.1: SplitMix64 seeded with seed_words[0] | seed_words[1] << 32 of seed_words_from_bytes('igneum-era/' || n_le64 || E_n); width = allowed[below(|allowed|)], stride_mul = low32(next()) | 1, stride_rot = 1 + below(31), then four next() draws for a partial Fisher-Yates over positions log2(W)..15 of which 4 - log2(W) are used",
|
||||
"allowed_widths": [1],
|
||||
"width_words": 1,
|
||||
"stride_mul": "0x2b4a5b97",
|
||||
"stride_rot": 28,
|
||||
"interleave": [1, 3, 4, 8],
|
||||
"address": "y = rotl(src * stride_mul, stride_rot); k = min(win, D - 26); idx = ((y & (mask >> k)) | ((off & (2^k - 1)) << (D - k))) & mask; a wide load aligns idx down to W words",
|
||||
"windows": "per instruction, after the width roll: win = below(3), off = low32(next()) & (2^win - 1); used on a load slot (the instruction's win and off fields)",
|
||||
"dataset_word": "dataset[w] = item(t(w))[j(w)]: j(w) gathers the bits of w at the interleave positions, t(w) is w with those bits removed",
|
||||
"program_id_suffix": "'era/' || allowed[3] || width_words || stride_mul_le32 || stride_rot_le32 || interleave[4]"
|
||||
},
|
||||
"op_mix": {"load": 16, "add": 15, "shfl": 6, "mad": 4, "or": 4, "rotl": 4, "rotr": 4, "xor": 4, "mulhi": 3, "mul": 2, "sub": 2},
|
||||
"register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]",
|
||||
"splitmix32": "x ^= x>>16; x *= 0x7feb352d; x ^= x>>15; x *= 0x846ca68b; x ^= x>>16",
|
||||
"iteration": "sel = r0 sampled once at the top of each iteration, then all instructions in order",
|
||||
"output": "lo = r0 ^ rotl(r1,7) ^ rotl(r2,14) ^ rotl(r3,21); hi = r4 ^ rotl(r5,9) ^ rotl(r6,18) ^ rotl(r7,27); out = (hi << 32) | lo",
|
||||
"op_semantics": {
|
||||
"add": "dst = dst + src + (bit `bit` of sel ? imm2 : imm)",
|
||||
"sub": "dst = dst - src",
|
||||
"mul": "dst = dst * src (low 32)",
|
||||
"mulhi": "dst = high 32 bits of dst * src",
|
||||
"xor": "dst = dst ^ src",
|
||||
"or": "dst = dst | src",
|
||||
"rotl": "dst = rotl(dst, rot), rot in 1..31",
|
||||
"rotr": "dst = rotr(dst, src & 31)",
|
||||
"mad": "dst = src * src2 + dst",
|
||||
"shfl": "dst = dst ^ (src of lane (lane ^ mask)), mask in {1,2,4,8,16}, within the 32-lane warp",
|
||||
"load": "dst = dst ^ dataset[src & dataset.mask]",
|
||||
"wload": "base = (src of lane 0 & dataset.mask) & ~31; dst = dst ^ dataset[base + lane] (warp-coalesced 128-byte load, lever b, only when --wide-frac > 0)"
|
||||
},
|
||||
"dataset": {
|
||||
"log2_words": 28,
|
||||
"bytes": 1073741824,
|
||||
"mask": "0x0fffffff",
|
||||
"day": "bytes:69676e65756d2d6461792ffa50000000000000",
|
||||
"day_bytes": "69676e65756d2d6461792ffa50000000000000",
|
||||
"day_words_from": "seed_words_from_bytes(day_bytes)",
|
||||
"d0": "0xceed56d7",
|
||||
"d1": "0x9ba270d2",
|
||||
"mode": "memory-hard",
|
||||
"spec": "proto-metal/MEMHARD.md",
|
||||
"key": ["0xceed56d7", "0x9ba270d2", "0x82caab2d", "0x81ebce0e", "0x12b6ecf1", "0xd0f3fd7c", "0xd872eefe", "0xc158c7bd"],
|
||||
"key_derivation": "the 8 words of seed_words_from_bytes(day_bytes); d0, d1 are key[0], key[1]",
|
||||
"cache": {"log2_words": 26, "bytes": 268435456, "line_words": 16, "segment_lines": 64, "segments": 65536, "block": "ChaCha12 core + feed-forward, rotations 16 12 8 7", "sigma": ["0x61707865", "0x3320646e", "0x79622d32", "0x6b206574"], "tag": ["0x49676e65", "0x756d4d48"], "chain": "in_j = prev_line ^ (sigma[0..3] || key[0..7] || seg || j || tag[0..1]); line_j = block(in_j); prev_0 = 0"},
|
||||
"mixer": {"draw": "SplitMix64 seeded with key[0] | key[1] << 32: rot[0..7] = 1 + next() % 31, mul[0..15] = low32(next()) | 1, rc[0..15] = low32(next())", "rot": [17, 12, 20, 23, 7, 3, 27, 16], "mul": ["0xf351d601", "0xa3bb398f", "0xb5a09e35", "0x7509c9c1", "0x6bbf31e9", "0xfc849a79", "0xded91851", "0x8d9113d1", "0x0ff15225", "0x3a5bdd41", "0xab533435", "0xe1c55ad5", "0xe6d3bd0d", "0x9d9ffbbd", "0xbb2a3cf3", "0x50a7c08d"], "rc": ["0xc6892460", "0x25b7228a", "0xcd515004", "0x2846527a", "0xa6324241", "0x36e3ec53", "0x82961bac", "0x0f97ba7d", "0xb6f921a9", "0x3ada24e5", "0xde20ab91", "0x5378eeb2", "0x7d161662", "0x89353cc1", "0xb1aa03a2", "0x788acae6"], "round": "for i in 0..15: s[i] = (s[i] ^ (rc[i] + (r+1) * 0x9E3779B9)) * mul[i]; then quarter rounds on columns (0,4,8,12) (1,5,9,13) (2,6,10,14) (3,7,11,15) with rot[0..3] and diagonals (0,5,10,15) (1,6,11,12) (2,7,8,13) (3,4,9,14) with rot[4..7]", "quarter_round": "a += b; d ^= a; d = rotl(d, r1); c += d; b ^= c; b = rotl(b, r2); a += b; d ^= a; d = rotl(d, r3); c += d; b ^= c; b = rotl(b, r4)"},
|
||||
"item": "s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..7: s = M_r(s); line = s[0] & 0x003fffff; s[i] ^= cache[line * 16 + i]; then s = M_8(s); item(t) = s",
|
||||
"word": "dataset[w] = item(w >> 4)[w & 15]"
|
||||
},
|
||||
"instructions": [
|
||||
{"i": 0, "op": "add", "dst": 4, "src": 5, "src2": 7, "imm": "0xea86e152", "imm2": "0x5810667a", "rot": 27, "bit": 13, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 1, "op": "xor", "dst": 7, "src": 0, "src2": 6, "imm": "0xbaab6229", "imm2": "0xed861989", "rot": 26, "bit": 22, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 2, "op": "shfl", "dst": 3, "src": 6, "src2": 2, "imm": "0x5b623116", "imm2": "0xff12e5b2", "rot": 12, "bit": 24, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 3, "op": "sub", "dst": 4, "src": 1, "src2": 1, "imm": "0xe99741c7", "imm2": "0xf5fa5009", "rot": 1, "bit": 21, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 4, "op": "mad", "dst": 2, "src": 0, "src2": 4, "imm": "0x673c2157", "imm2": "0xee02465f", "rot": 20, "bit": 22, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 5, "op": "rotl", "dst": 4, "src": 7, "src2": 7, "imm": "0x946f7818", "imm2": "0x45d3399e", "rot": 9, "bit": 2, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 6, "op": "rotr", "dst": 0, "src": 2, "src2": 4, "imm": "0x5f6a0ed2", "imm2": "0x7043a636", "rot": 19, "bit": 31, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 7, "op": "load", "dst": 6, "src": 7, "src2": 1, "imm": "0x5c61dcf7", "imm2": "0x7466aa40", "rot": 19, "bit": 9, "mask": 2, "width": 1, "win": 2, "off": 1},
|
||||
{"i": 8, "op": "load", "dst": 1, "src": 4, "src2": 5, "imm": "0x85668475", "imm2": "0xdb8cc483", "rot": 29, "bit": 7, "mask": 4, "width": 1, "win": 1, "off": 1},
|
||||
{"i": 9, "op": "load", "dst": 1, "src": 2, "src2": 4, "imm": "0x4454980f", "imm2": "0xebd31581", "rot": 10, "bit": 28, "mask": 16, "width": 1, "win": 1, "off": 1},
|
||||
{"i": 10, "op": "load", "dst": 7, "src": 0, "src2": 2, "imm": "0xe075297c", "imm2": "0x5779c44c", "rot": 10, "bit": 22, "mask": 2, "width": 1, "win": 1, "off": 1},
|
||||
{"i": 11, "op": "load", "dst": 7, "src": 1, "src2": 6, "imm": "0x65aa4311", "imm2": "0x4fe48ea9", "rot": 15, "bit": 9, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 12, "op": "mul", "dst": 5, "src": 4, "src2": 2, "imm": "0x1383d3ad", "imm2": "0xf3094b29", "rot": 8, "bit": 9, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 13, "op": "load", "dst": 7, "src": 6, "src2": 2, "imm": "0xed8a496f", "imm2": "0x3072c3c6", "rot": 28, "bit": 19, "mask": 8, "width": 1, "win": 1, "off": 1},
|
||||
{"i": 14, "op": "shfl", "dst": 6, "src": 5, "src2": 0, "imm": "0x8b965b57", "imm2": "0xcfeca6c1", "rot": 12, "bit": 27, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 15, "op": "shfl", "dst": 3, "src": 5, "src2": 3, "imm": "0x877c7586", "imm2": "0xa9cb2a03", "rot": 2, "bit": 29, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 16, "op": "or", "dst": 2, "src": 7, "src2": 3, "imm": "0xb740221a", "imm2": "0x89d38d6d", "rot": 6, "bit": 31, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 17, "op": "rotr", "dst": 0, "src": 6, "src2": 1, "imm": "0x26f3ad8a", "imm2": "0x27256f15", "rot": 18, "bit": 5, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 18, "op": "add", "dst": 7, "src": 5, "src2": 7, "imm": "0xf66e7017", "imm2": "0xb9e3577e", "rot": 9, "bit": 12, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 19, "op": "rotl", "dst": 7, "src": 0, "src2": 5, "imm": "0x849ae6ee", "imm2": "0x02b358f9", "rot": 24, "bit": 18, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 20, "op": "add", "dst": 6, "src": 7, "src2": 6, "imm": "0x52334d12", "imm2": "0x8c9f0ef8", "rot": 11, "bit": 23, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 21, "op": "or", "dst": 2, "src": 6, "src2": 5, "imm": "0xb1871e63", "imm2": "0xb2e40191", "rot": 5, "bit": 13, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 22, "op": "mad", "dst": 6, "src": 5, "src2": 4, "imm": "0x97df29e4", "imm2": "0xe60fea84", "rot": 11, "bit": 16, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 23, "op": "or", "dst": 1, "src": 0, "src2": 3, "imm": "0x8f30d21d", "imm2": "0x2df685a0", "rot": 31, "bit": 0, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 24, "op": "xor", "dst": 6, "src": 1, "src2": 2, "imm": "0xd2c4025f", "imm2": "0x5269eb4d", "rot": 31, "bit": 21, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 25, "op": "add", "dst": 2, "src": 6, "src2": 5, "imm": "0x659fc3d3", "imm2": "0x9cec0e12", "rot": 6, "bit": 17, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 26, "op": "rotr", "dst": 7, "src": 0, "src2": 1, "imm": "0xd89ef484", "imm2": "0x20be3846", "rot": 12, "bit": 9, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 27, "op": "mad", "dst": 4, "src": 5, "src2": 7, "imm": "0xb48420ae", "imm2": "0x3d1f2485", "rot": 14, "bit": 28, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 28, "op": "mad", "dst": 3, "src": 2, "src2": 4, "imm": "0xc5c46d76", "imm2": "0x700044b5", "rot": 22, "bit": 10, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 29, "op": "load", "dst": 1, "src": 4, "src2": 3, "imm": "0x0fbaf177", "imm2": "0xfff4f2ed", "rot": 20, "bit": 31, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 30, "op": "load", "dst": 2, "src": 3, "src2": 0, "imm": "0xe90eb2e5", "imm2": "0xb0f9eb79", "rot": 27, "bit": 14, "mask": 4, "width": 1, "win": 2, "off": 2},
|
||||
{"i": 31, "op": "load", "dst": 1, "src": 5, "src2": 2, "imm": "0x97ba3fc3", "imm2": "0x7894e657", "rot": 3, "bit": 30, "mask": 16, "width": 1, "win": 1, "off": 1},
|
||||
{"i": 32, "op": "add", "dst": 7, "src": 2, "src2": 7, "imm": "0x070888a8", "imm2": "0xe403240e", "rot": 2, "bit": 16, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 33, "op": "add", "dst": 2, "src": 0, "src2": 5, "imm": "0xf2e46d55", "imm2": "0x29701828", "rot": 31, "bit": 28, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 34, "op": "add", "dst": 2, "src": 3, "src2": 0, "imm": "0x58f75b87", "imm2": "0x343b7aee", "rot": 12, "bit": 14, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 35, "op": "mulhi", "dst": 2, "src": 5, "src2": 6, "imm": "0x5892a9e6", "imm2": "0xc9824c94", "rot": 19, "bit": 26, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 36, "op": "xor", "dst": 4, "src": 2, "src2": 3, "imm": "0x7b5b5474", "imm2": "0x45cfc5dd", "rot": 17, "bit": 18, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 37, "op": "mul", "dst": 6, "src": 5, "src2": 7, "imm": "0xccf564a5", "imm2": "0x873ad101", "rot": 7, "bit": 11, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 38, "op": "xor", "dst": 7, "src": 0, "src2": 6, "imm": "0xcac8f06d", "imm2": "0x6b97c683", "rot": 18, "bit": 28, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 39, "op": "add", "dst": 7, "src": 2, "src2": 4, "imm": "0xb8180e9d", "imm2": "0x32bbd117", "rot": 23, "bit": 19, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 40, "op": "rotr", "dst": 2, "src": 3, "src2": 0, "imm": "0x2d6070bc", "imm2": "0x68ff101e", "rot": 13, "bit": 18, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 41, "op": "sub", "dst": 7, "src": 0, "src2": 2, "imm": "0x0daf96ea", "imm2": "0x36f37be1", "rot": 5, "bit": 0, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 42, "op": "add", "dst": 4, "src": 3, "src2": 6, "imm": "0x6ced15b7", "imm2": "0x6df7aed4", "rot": 19, "bit": 4, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 43, "op": "shfl", "dst": 7, "src": 3, "src2": 5, "imm": "0x8ace05f3", "imm2": "0xd378ec12", "rot": 23, "bit": 10, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 44, "op": "load", "dst": 0, "src": 7, "src2": 4, "imm": "0xb0607786", "imm2": "0xc4acabbc", "rot": 13, "bit": 7, "mask": 16, "width": 1, "win": 1, "off": 1},
|
||||
{"i": 45, "op": "add", "dst": 1, "src": 6, "src2": 5, "imm": "0xe09f54e9", "imm2": "0x83e825bf", "rot": 23, "bit": 14, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 46, "op": "load", "dst": 3, "src": 1, "src2": 0, "imm": "0x63cc1e4e", "imm2": "0xa1be8118", "rot": 12, "bit": 6, "mask": 2, "width": 1, "win": 2, "off": 0},
|
||||
{"i": 47, "op": "load", "dst": 6, "src": 3, "src2": 7, "imm": "0x353f1d79", "imm2": "0x3b2e7456", "rot": 18, "bit": 18, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 48, "op": "mulhi", "dst": 4, "src": 2, "src2": 7, "imm": "0x00d8a3cd", "imm2": "0x231866d2", "rot": 21, "bit": 20, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 49, "op": "add", "dst": 5, "src": 0, "src2": 2, "imm": "0xa8bae6df", "imm2": "0xf572bdb9", "rot": 14, "bit": 7, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 50, "op": "shfl", "dst": 0, "src": 7, "src2": 7, "imm": "0x81ef22e1", "imm2": "0x74438fc5", "rot": 28, "bit": 18, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 51, "op": "add", "dst": 6, "src": 0, "src2": 6, "imm": "0x383b9260", "imm2": "0x11e17c61", "rot": 12, "bit": 19, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 52, "op": "load", "dst": 5, "src": 2, "src2": 2, "imm": "0xfb84f451", "imm2": "0x11cd863e", "rot": 21, "bit": 20, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 53, "op": "rotl", "dst": 6, "src": 5, "src2": 4, "imm": "0xb1a7db6b", "imm2": "0x76686b9b", "rot": 12, "bit": 4, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 54, "op": "rotl", "dst": 3, "src": 6, "src2": 3, "imm": "0x6f981f52", "imm2": "0xd99aeba2", "rot": 12, "bit": 27, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 55, "op": "add", "dst": 2, "src": 1, "src2": 2, "imm": "0xac6be8e3", "imm2": "0x18d67dbb", "rot": 26, "bit": 10, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 56, "op": "load", "dst": 1, "src": 4, "src2": 0, "imm": "0x7e7f6a00", "imm2": "0x6f0747da", "rot": 25, "bit": 28, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 57, "op": "add", "dst": 5, "src": 0, "src2": 4, "imm": "0xf03673fe", "imm2": "0xa75cd60d", "rot": 16, "bit": 12, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 58, "op": "load", "dst": 5, "src": 0, "src2": 2, "imm": "0x227f94a6", "imm2": "0x0e8344f9", "rot": 20, "bit": 10, "mask": 2, "width": 1, "win": 2, "off": 0},
|
||||
{"i": 59, "op": "add", "dst": 1, "src": 4, "src2": 3, "imm": "0xdecd4794", "imm2": "0x8dfb96bb", "rot": 21, "bit": 7, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 60, "op": "mulhi", "dst": 3, "src": 2, "src2": 2, "imm": "0x0dd268e0", "imm2": "0x53034ca9", "rot": 1, "bit": 8, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 61, "op": "or", "dst": 6, "src": 4, "src2": 7, "imm": "0x3a45a321", "imm2": "0x9bc59a5f", "rot": 25, "bit": 11, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 62, "op": "shfl", "dst": 5, "src": 4, "src2": 2, "imm": "0x8f229cc1", "imm2": "0xcaac64a2", "rot": 17, "bit": 13, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 63, "op": "load", "dst": 3, "src": 6, "src2": 6, "imm": "0xb2574178", "imm2": "0xcbcc798d", "rot": 28, "bit": 0, "mask": 16, "width": 1, "win": 1, "off": 1}
|
||||
]
|
||||
}
|
||||
109
proto-cuda/packs-ca2-era/era-2/program.metal
Normal file
109
proto-cuda/packs-ca2-era/era-2/program.metal
Normal file
|
|
@ -0,0 +1,109 @@
|
|||
#include <metal_stdlib>
|
||||
using namespace metal;
|
||||
|
||||
#define MASK 0x0fffffffu
|
||||
constant uint SEEDW[8] = { 0x667d0fbdu, 0x7b8e5963u, 0x31c67e5eu, 0x4529ddc6u, 0xef19d6d8u, 0xaccf6211u, 0xda0aed32u, 0xabc6df31u };
|
||||
|
||||
inline uint splitmix32(uint x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31
|
||||
inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
inline uint ds_elem(uint i, uint d0, uint d1) {
|
||||
uint x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
kernel void igneum_hash(device const uint* dataset [[buffer(0)]],
|
||||
device ulong* out [[buffer(1)]],
|
||||
constant uint& baseNonce [[buffer(2)]],
|
||||
uint gid [[thread_position_in_grid]]) {
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint x = nonce ^ SEEDW[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ SEEDW[1]; }
|
||||
{ uint x = nonce ^ SEEDW[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ SEEDW[2]; }
|
||||
{ uint x = nonce ^ SEEDW[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ SEEDW[3]; }
|
||||
{ uint x = nonce ^ SEEDW[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ SEEDW[4]; }
|
||||
{ uint x = nonce ^ SEEDW[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ SEEDW[5]; }
|
||||
{ uint x = nonce ^ SEEDW[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ SEEDW[6]; }
|
||||
{ uint x = nonce ^ SEEDW[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ SEEDW[7]; }
|
||||
{ uint x = nonce ^ SEEDW[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ SEEDW[0]; }
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = r4 + r5 + select(0xea86e152u, 0x5810667au, ((sel >> 13u) & 1u) != 0u); // 0
|
||||
r7 = r7 ^ r0; // 1
|
||||
r3 = r3 ^ simd_shuffle_xor(r6, (ushort)1); // 2
|
||||
r4 = r4 - r1; // 3
|
||||
r2 = r0 * r4 + r2; // 4
|
||||
r4 = rotl_imm(r4, 9u); // 5
|
||||
r0 = rotr_var(r0, r2); // 6
|
||||
r6 = r6 ^ dataset[((rotl_imm(r7 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x04000000u) & MASK]; // 7
|
||||
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 8
|
||||
r1 = r1 ^ dataset[((rotl_imm(r2 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 9
|
||||
r7 = r7 ^ dataset[((rotl_imm(r0 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 10
|
||||
r7 = r7 ^ dataset[((rotl_imm(r1 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 11
|
||||
r5 = r5 * r4; // 12
|
||||
r7 = r7 ^ dataset[((rotl_imm(r6 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 13
|
||||
r6 = r6 ^ simd_shuffle_xor(r5, (ushort)16); // 14
|
||||
r3 = r3 ^ simd_shuffle_xor(r5, (ushort)1); // 15
|
||||
r2 = r2 | r7; // 16
|
||||
r0 = rotr_var(r0, r6); // 17
|
||||
r7 = r7 + r5 + select(0xf66e7017u, 0xb9e3577eu, ((sel >> 12u) & 1u) != 0u); // 18
|
||||
r7 = rotl_imm(r7, 24u); // 19
|
||||
r6 = r6 + r7 + select(0x52334d12u, 0x8c9f0ef8u, ((sel >> 23u) & 1u) != 0u); // 20
|
||||
r2 = r2 | r6; // 21
|
||||
r6 = r5 * r4 + r6; // 22
|
||||
r1 = r1 | r0; // 23
|
||||
r6 = r6 ^ r1; // 24
|
||||
r2 = r2 + r6 + select(0x659fc3d3u, 0x9cec0e12u, ((sel >> 17u) & 1u) != 0u); // 25
|
||||
r7 = rotr_var(r7, r0); // 26
|
||||
r4 = r5 * r7 + r4; // 27
|
||||
r3 = r2 * r4 + r3; // 28
|
||||
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 29
|
||||
r2 = r2 ^ dataset[((rotl_imm(r3 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x08000000u) & MASK]; // 30
|
||||
r1 = r1 ^ dataset[((rotl_imm(r5 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 31
|
||||
r7 = r7 + r2 + select(0x070888a8u, 0xe403240eu, ((sel >> 16u) & 1u) != 0u); // 32
|
||||
r2 = r2 + r0 + select(0xf2e46d55u, 0x29701828u, ((sel >> 28u) & 1u) != 0u); // 33
|
||||
r2 = r2 + r3 + select(0x58f75b87u, 0x343b7aeeu, ((sel >> 14u) & 1u) != 0u); // 34
|
||||
r2 = mulhi(r2, r5); // 35
|
||||
r4 = r4 ^ r2; // 36
|
||||
r6 = r6 * r5; // 37
|
||||
r7 = r7 ^ r0; // 38
|
||||
r7 = r7 + r2 + select(0xb8180e9du, 0x32bbd117u, ((sel >> 19u) & 1u) != 0u); // 39
|
||||
r2 = rotr_var(r2, r3); // 40
|
||||
r7 = r7 - r0; // 41
|
||||
r4 = r4 + r3 + select(0x6ced15b7u, 0x6df7aed4u, ((sel >> 4u) & 1u) != 0u); // 42
|
||||
r7 = r7 ^ simd_shuffle_xor(r3, (ushort)4); // 43
|
||||
r0 = r0 ^ dataset[((rotl_imm(r7 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 44
|
||||
r1 = r1 + r6 + select(0xe09f54e9u, 0x83e825bfu, ((sel >> 14u) & 1u) != 0u); // 45
|
||||
r3 = r3 ^ dataset[((rotl_imm(r1 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 46
|
||||
r6 = r6 ^ dataset[((rotl_imm(r3 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 47
|
||||
r4 = mulhi(r4, r2); // 48
|
||||
r5 = r5 + r0 + select(0xa8bae6dfu, 0xf572bdb9u, ((sel >> 7u) & 1u) != 0u); // 49
|
||||
r0 = r0 ^ simd_shuffle_xor(r7, (ushort)4); // 50
|
||||
r6 = r6 + r0 + select(0x383b9260u, 0x11e17c61u, ((sel >> 19u) & 1u) != 0u); // 51
|
||||
r5 = r5 ^ dataset[((rotl_imm(r2 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 52
|
||||
r6 = rotl_imm(r6, 12u); // 53
|
||||
r3 = rotl_imm(r3, 12u); // 54
|
||||
r2 = r2 + r1 + select(0xac6be8e3u, 0x18d67dbbu, ((sel >> 10u) & 1u) != 0u); // 55
|
||||
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 56
|
||||
r5 = r5 + r0 + select(0xf03673feu, 0xa75cd60du, ((sel >> 12u) & 1u) != 0u); // 57
|
||||
r5 = r5 ^ dataset[((rotl_imm(r0 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 58
|
||||
r1 = r1 + r4 + select(0xdecd4794u, 0x8dfb96bbu, ((sel >> 7u) & 1u) != 0u); // 59
|
||||
r3 = mulhi(r3, r2); // 60
|
||||
r6 = r6 | r4; // 61
|
||||
r5 = r5 ^ simd_shuffle_xor(r4, (ushort)8); // 62
|
||||
r3 = r3 ^ dataset[((rotl_imm(r6 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 63
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
111
proto-cuda/packs-ca2-era/era-2/program_bound.metal
Normal file
111
proto-cuda/packs-ca2-era/era-2/program_bound.metal
Normal file
|
|
@ -0,0 +1,111 @@
|
|||
#include <metal_stdlib>
|
||||
using namespace metal;
|
||||
|
||||
#define MASK 0x0fffffffu
|
||||
constant uint SEEDW[8] = { 0x667d0fbdu, 0x7b8e5963u, 0x31c67e5eu, 0x4529ddc6u, 0xef19d6d8u, 0xaccf6211u, 0xda0aed32u, 0xabc6df31u };
|
||||
|
||||
inline uint splitmix32(uint x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31
|
||||
inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
inline uint ds_elem(uint i, uint d0, uint d1) {
|
||||
uint x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// Header-bound variant: the init words come from buffer 3 (bind.rs), not from SEEDW.
|
||||
kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]],
|
||||
device ulong* out [[buffer(1)]],
|
||||
constant uint& baseNonce [[buffer(2)]],
|
||||
constant uint* initw [[buffer(3)]],
|
||||
uint gid [[thread_position_in_grid]]) {
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint x = nonce ^ initw[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ initw[1]; }
|
||||
{ uint x = nonce ^ initw[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ initw[2]; }
|
||||
{ uint x = nonce ^ initw[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ initw[3]; }
|
||||
{ uint x = nonce ^ initw[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ initw[4]; }
|
||||
{ uint x = nonce ^ initw[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ initw[5]; }
|
||||
{ uint x = nonce ^ initw[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ initw[6]; }
|
||||
{ uint x = nonce ^ initw[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ initw[7]; }
|
||||
{ uint x = nonce ^ initw[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ initw[0]; }
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = r4 + r5 + select(0xea86e152u, 0x5810667au, ((sel >> 13u) & 1u) != 0u); // 0
|
||||
r7 = r7 ^ r0; // 1
|
||||
r3 = r3 ^ simd_shuffle_xor(r6, (ushort)1); // 2
|
||||
r4 = r4 - r1; // 3
|
||||
r2 = r0 * r4 + r2; // 4
|
||||
r4 = rotl_imm(r4, 9u); // 5
|
||||
r0 = rotr_var(r0, r2); // 6
|
||||
r6 = r6 ^ dataset[((rotl_imm(r7 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x04000000u) & MASK]; // 7
|
||||
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 8
|
||||
r1 = r1 ^ dataset[((rotl_imm(r2 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 9
|
||||
r7 = r7 ^ dataset[((rotl_imm(r0 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 10
|
||||
r7 = r7 ^ dataset[((rotl_imm(r1 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 11
|
||||
r5 = r5 * r4; // 12
|
||||
r7 = r7 ^ dataset[((rotl_imm(r6 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 13
|
||||
r6 = r6 ^ simd_shuffle_xor(r5, (ushort)16); // 14
|
||||
r3 = r3 ^ simd_shuffle_xor(r5, (ushort)1); // 15
|
||||
r2 = r2 | r7; // 16
|
||||
r0 = rotr_var(r0, r6); // 17
|
||||
r7 = r7 + r5 + select(0xf66e7017u, 0xb9e3577eu, ((sel >> 12u) & 1u) != 0u); // 18
|
||||
r7 = rotl_imm(r7, 24u); // 19
|
||||
r6 = r6 + r7 + select(0x52334d12u, 0x8c9f0ef8u, ((sel >> 23u) & 1u) != 0u); // 20
|
||||
r2 = r2 | r6; // 21
|
||||
r6 = r5 * r4 + r6; // 22
|
||||
r1 = r1 | r0; // 23
|
||||
r6 = r6 ^ r1; // 24
|
||||
r2 = r2 + r6 + select(0x659fc3d3u, 0x9cec0e12u, ((sel >> 17u) & 1u) != 0u); // 25
|
||||
r7 = rotr_var(r7, r0); // 26
|
||||
r4 = r5 * r7 + r4; // 27
|
||||
r3 = r2 * r4 + r3; // 28
|
||||
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 29
|
||||
r2 = r2 ^ dataset[((rotl_imm(r3 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x08000000u) & MASK]; // 30
|
||||
r1 = r1 ^ dataset[((rotl_imm(r5 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 31
|
||||
r7 = r7 + r2 + select(0x070888a8u, 0xe403240eu, ((sel >> 16u) & 1u) != 0u); // 32
|
||||
r2 = r2 + r0 + select(0xf2e46d55u, 0x29701828u, ((sel >> 28u) & 1u) != 0u); // 33
|
||||
r2 = r2 + r3 + select(0x58f75b87u, 0x343b7aeeu, ((sel >> 14u) & 1u) != 0u); // 34
|
||||
r2 = mulhi(r2, r5); // 35
|
||||
r4 = r4 ^ r2; // 36
|
||||
r6 = r6 * r5; // 37
|
||||
r7 = r7 ^ r0; // 38
|
||||
r7 = r7 + r2 + select(0xb8180e9du, 0x32bbd117u, ((sel >> 19u) & 1u) != 0u); // 39
|
||||
r2 = rotr_var(r2, r3); // 40
|
||||
r7 = r7 - r0; // 41
|
||||
r4 = r4 + r3 + select(0x6ced15b7u, 0x6df7aed4u, ((sel >> 4u) & 1u) != 0u); // 42
|
||||
r7 = r7 ^ simd_shuffle_xor(r3, (ushort)4); // 43
|
||||
r0 = r0 ^ dataset[((rotl_imm(r7 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 44
|
||||
r1 = r1 + r6 + select(0xe09f54e9u, 0x83e825bfu, ((sel >> 14u) & 1u) != 0u); // 45
|
||||
r3 = r3 ^ dataset[((rotl_imm(r1 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 46
|
||||
r6 = r6 ^ dataset[((rotl_imm(r3 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 47
|
||||
r4 = mulhi(r4, r2); // 48
|
||||
r5 = r5 + r0 + select(0xa8bae6dfu, 0xf572bdb9u, ((sel >> 7u) & 1u) != 0u); // 49
|
||||
r0 = r0 ^ simd_shuffle_xor(r7, (ushort)4); // 50
|
||||
r6 = r6 + r0 + select(0x383b9260u, 0x11e17c61u, ((sel >> 19u) & 1u) != 0u); // 51
|
||||
r5 = r5 ^ dataset[((rotl_imm(r2 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 52
|
||||
r6 = rotl_imm(r6, 12u); // 53
|
||||
r3 = rotl_imm(r3, 12u); // 54
|
||||
r2 = r2 + r1 + select(0xac6be8e3u, 0x18d67dbbu, ((sel >> 10u) & 1u) != 0u); // 55
|
||||
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x2b4a5b97u, 28u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 56
|
||||
r5 = r5 + r0 + select(0xf03673feu, 0xa75cd60du, ((sel >> 12u) & 1u) != 0u); // 57
|
||||
r5 = r5 ^ dataset[((rotl_imm(r0 * 0x2b4a5b97u, 28u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 58
|
||||
r1 = r1 + r4 + select(0xdecd4794u, 0x8dfb96bbu, ((sel >> 7u) & 1u) != 0u); // 59
|
||||
r3 = mulhi(r3, r2); // 60
|
||||
r6 = r6 | r4; // 61
|
||||
r5 = r5 ^ simd_shuffle_xor(r4, (ushort)8); // 62
|
||||
r3 = r3 ^ dataset[((rotl_imm(r6 * 0x2b4a5b97u, 28u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 63
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
5
proto-cuda/packs-ca2-era/era-2/seeds.txt
Normal file
5
proto-cuda/packs-ca2-era/era-2/seeds.txt
Normal file
|
|
@ -0,0 +1,5 @@
|
|||
epoch_seed_hex edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07
|
||||
day_seed_hex 69676e65756d2d6461792ffa50000000000000
|
||||
epoch_index 0
|
||||
day_index 20730
|
||||
daa_score 0
|
||||
57
proto-cuda/packs-ca2-era/era-2/vectors.h
Normal file
57
proto-cuda/packs-ca2-era/era-2/vectors.h
Normal file
|
|
@ -0,0 +1,57 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v3, memory-hard dataset
|
||||
#pragma once
|
||||
#ifdef __cplusplus
|
||||
#include <cstdint>
|
||||
#else
|
||||
#include <stdint.h>
|
||||
#endif
|
||||
|
||||
#define IGNEUM_VEC_WARPS 3
|
||||
static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u };
|
||||
static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = {
|
||||
{ // base nonce 0
|
||||
0x2d96e3f8e9ad42e2ull, 0x0d36523b17ba1058ull, 0x899e2a6df6991c3dull, 0xdb382aa991ff18d6ull, 0xc896fe7b624c7722ull, 0x932cfc8892d24216ull, 0xd9ef4047f797129bull, 0x63fb054ee8cfe1d2ull,
|
||||
0xbcbd287cfd7e12a9ull, 0x68e41482f26203c2ull, 0x5c368b15e9ec2bf5ull, 0x83484f0946bfda39ull, 0xc82c1aeadba75e3cull, 0x023cbca13fae681dull, 0x000fcab814a56dd9ull, 0x432082a86895e2a7ull,
|
||||
0x635fb930f3cf454bull, 0x1104c1b855a17655ull, 0x2fd53820b88956b2ull, 0xc415d67f6420191dull, 0xeaf6853e1c859f5dull, 0x95dfdf66b21e0518ull, 0x7d67b2b734243bfeull, 0xb142ae8e4a4b5ca2ull,
|
||||
0x71ec0aff12d4c515ull, 0x0b2207c2c51fc2faull, 0xbc3957dbbf51d260ull, 0x74ea737eebe8a3e4ull, 0xd49fd6ea590200e9ull, 0xd5096fd4e83b8fa1ull, 0x999b099025a53907ull, 0xa5692c50b79eac9dull
|
||||
},
|
||||
{ // base nonce 4096
|
||||
0x9d4b6aaa8c292525ull, 0xb2ad4ab3941b3a3full, 0x17988da38edbf912ull, 0x566b31d7cceae84aull, 0x132f77e29f4e6d40ull, 0xbdc098ecac64b3dcull, 0x1670a7bfbc675d88ull, 0xda3a9960471e44bbull,
|
||||
0x887e9fda5ec82271ull, 0x6137f598a67fc0d4ull, 0xd77b98d7d2025d26ull, 0x96b472c7560a777dull, 0x0c98d0ff42f589f0ull, 0x4f0f15c9e4eefef2ull, 0x6bae44572933ab20ull, 0x2772c0266bdd5308ull,
|
||||
0x2fdb9e92d13d07baull, 0x6fe9239589ed122aull, 0x7e3955e5a91916fdull, 0x966d21685073c9b8ull, 0xa050194eb84c4104ull, 0xf7e095dae1d70633ull, 0x4f49616da323736bull, 0x1fcc532045f01cd4ull,
|
||||
0x69549d64913199fcull, 0x6684a1b257c011b0ull, 0xff318d60e8970b9aull, 0x1429c63759eb426dull, 0x3f470be3f4d4815aull, 0x6e1d6c6b91b984e2ull, 0xda2dd7f31fea85efull, 0xb3944abe38dae250ull
|
||||
},
|
||||
{ // base nonce 1000000
|
||||
0xd549a905f9121657ull, 0xd863810b66483e86ull, 0x65d735cf2e5447afull, 0x9cda91c71e0f5790ull, 0x682c4a1be66d2444ull, 0xcb64852ca144ffe1ull, 0x8e72112d5db27544ull, 0x1a9e6bb74ea3e563ull,
|
||||
0x63c279f228faf1bcull, 0xf7b77f8435da26a2ull, 0x3c28bd5b463df46full, 0xd2cdd0e9e968d2bdull, 0x46eae240fc0c3f01ull, 0xbdd4c4e525220d3cull, 0x345178054b1eac9dull, 0x312f541e08ab8fe9ull,
|
||||
0x0cbc6f22e65887adull, 0xf9a0d018f9ce95cfull, 0x994d29e263550764ull, 0x27a37ad3487c4018ull, 0xcdb7645aa05dac6aull, 0x8578f1e945156d7full, 0x8e135ab725d23599ull, 0xfa902aebc88a620cull,
|
||||
0x5bdb23834f693ba8ull, 0xae53782ddf331358ull, 0x2304fcfaad2616fcull, 0x21e9216f6c0b14a6ull, 0xa88e90dc82b85f66ull, 0xba5aa73b369dd894ull, 0x97e37652c487cc5eull, 0x36cb34ed0c8ce81full
|
||||
}
|
||||
};
|
||||
|
||||
// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455).
|
||||
static const uint32_t IGNEUM_DS_HEAD[16] = {
|
||||
0x3dd50b1fu, 0x93123701u, 0x48edec90u, 0x3a7d2407u, 0xcdd88b45u, 0x2c37c6b0u, 0xb03a0ac0u, 0x10e2c6a5u,
|
||||
0x5119540eu, 0x3358ef2du, 0x657ac748u, 0x09e2105cu, 0x25e2929cu, 0x5fa0cdccu, 0x4317cabdu, 0x6178ee52u
|
||||
};
|
||||
static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u;
|
||||
static const uint32_t IGNEUM_DS_LAST = 0xf7b7180eu;
|
||||
// 64 sampled dataset words (index, value) computed on the Mac.
|
||||
#define IGNEUM_DS_SAMPLES 64
|
||||
static const uint32_t IGNEUM_DS_SAMPLE_INDEX[IGNEUM_DS_SAMPLES] = {
|
||||
59471966u, 217795994u, 208353206u, 42483309u, 172547758u, 148076330u, 183853158u, 214389424u, 267488061u, 169781097u, 184093494u, 153880993u, 84977930u, 46426879u, 3093825u, 225364072u, 44593546u, 260713159u, 168250303u, 52384140u, 223401610u, 45554030u, 95410555u, 175039924u, 79171087u, 267580473u, 24168642u, 37981670u, 171551130u, 195559979u, 204611762u, 140997658u, 138925853u, 86637313u, 20736778u, 219665210u, 160430336u, 264654675u, 8013395u, 228945585u, 213884386u, 104419827u, 44185464u, 142737231u, 99284897u, 132475900u, 61861762u, 132056166u, 262388043u, 91878046u, 117353561u, 124768597u, 71352993u, 190698941u, 46055428u, 55281366u, 165145231u, 106810753u, 171985651u, 232085256u, 159510492u, 40072060u, 209107596u, 39023794u
|
||||
};
|
||||
static const uint32_t IGNEUM_DS_SAMPLE_VALUE[IGNEUM_DS_SAMPLES] = {
|
||||
0x88717593u, 0x79f4c2ebu, 0xa55c10fau, 0x1e48deedu, 0x825505dfu, 0x056e4b2au, 0x51c5cdc5u, 0x9133ff90u, 0xf3e1f630u, 0x8c0f4bcau, 0x13c805d6u, 0x7cd9cc39u, 0x54d3ad95u, 0x3142a64bu, 0xf690ebf9u, 0x5f05ecf5u, 0x708fd94du, 0x57e3a213u, 0xb60de9a6u, 0xc3cacc1fu, 0x72a0a173u, 0x2a13f6c7u, 0xfb5b4620u, 0xeae5f6c3u, 0x38e0ceb1u, 0x103d3fa4u, 0xcb03f2aeu, 0x562969ccu, 0xf1349e23u, 0x3fa57b5du, 0x4fb463b9u, 0xb2b70250u, 0x7f689f8cu, 0x27e5f93eu, 0xbaeafce2u, 0xaf3d3776u, 0x084c3724u, 0xffe6a16cu, 0xd4d0d18du, 0x5d78da4eu, 0xf0bd37c6u, 0x261f1c0du, 0x7bfc0ff6u, 0x08e5373bu, 0xd23a038cu, 0x124d74c0u, 0xdd74caceu, 0x5ab7c9bau, 0xc05d2942u, 0x2d1e2f73u, 0x52a27429u, 0x0c828d10u, 0x7514e589u, 0xdb8ca7d2u, 0xdb881ba6u, 0x0ce95e1du, 0x3c93a15eu, 0x73ec90f8u, 0xe47554edu, 0x496959f3u, 0xebf27832u, 0x11f5885eu, 0x69a1ff20u, 0x804542bdu
|
||||
};
|
||||
// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^26 words.
|
||||
static const uint32_t IGNEUM_CACHE_HEAD[16] = {
|
||||
0xebd9055cu, 0x7eaf6b21u, 0x9b610855u, 0x134a8562u, 0xb10059bau, 0x7f459e58u, 0x8f3c7873u, 0x3523e609u,
|
||||
0x75b8f4cau, 0x750d31b4u, 0x0dae0782u, 0x26f10015u, 0x7ba87a41u, 0x0efd8543u, 0x691d6368u, 0x8929d967u
|
||||
};
|
||||
static const uint32_t IGNEUM_CACHE_LAST[16] = {
|
||||
0x51474edfu, 0xc3cfba93u, 0xf21454cfu, 0x9b79baacu, 0x4d4fce23u, 0xd701dfd5u, 0x37357ab3u, 0x1be693fau,
|
||||
0xa701cb3bu, 0x7467c620u, 0x428184e8u, 0xf4010df0u, 0xe33aa88fu, 0xc78d6d62u, 0xae9ba9c0u, 0xf4bba97eu
|
||||
};
|
||||
static const uint64_t IGNEUM_CACHE_FNV64 = 0x448274a57f508cbcull;
|
||||
36
proto-cuda/packs-ca2-era/era-2/vectors.json
Normal file
36
proto-cuda/packs-ca2-era/era-2/vectors.json
Normal file
|
|
@ -0,0 +1,36 @@
|
|||
{
|
||||
"seed": "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000",
|
||||
"day": "bytes:69676e65756d2d6461792ffa50000000000000",
|
||||
"dataset_mode": "memory-hard",
|
||||
"dataset_log2_words": 28,
|
||||
"mask": "0x0fffffff",
|
||||
"lanes": 32,
|
||||
"source": "igneum-pow (Rust) CPU interpreter, generator v3, memory-hard dataset",
|
||||
"warps": [
|
||||
{"base_nonce": 0, "expected": [
|
||||
"0x2d96e3f8e9ad42e2", "0x0d36523b17ba1058", "0x899e2a6df6991c3d", "0xdb382aa991ff18d6", "0xc896fe7b624c7722", "0x932cfc8892d24216", "0xd9ef4047f797129b", "0x63fb054ee8cfe1d2",
|
||||
"0xbcbd287cfd7e12a9", "0x68e41482f26203c2", "0x5c368b15e9ec2bf5", "0x83484f0946bfda39", "0xc82c1aeadba75e3c", "0x023cbca13fae681d", "0x000fcab814a56dd9", "0x432082a86895e2a7",
|
||||
"0x635fb930f3cf454b", "0x1104c1b855a17655", "0x2fd53820b88956b2", "0xc415d67f6420191d", "0xeaf6853e1c859f5d", "0x95dfdf66b21e0518", "0x7d67b2b734243bfe", "0xb142ae8e4a4b5ca2",
|
||||
"0x71ec0aff12d4c515", "0x0b2207c2c51fc2fa", "0xbc3957dbbf51d260", "0x74ea737eebe8a3e4", "0xd49fd6ea590200e9", "0xd5096fd4e83b8fa1", "0x999b099025a53907", "0xa5692c50b79eac9d"
|
||||
]},
|
||||
{"base_nonce": 4096, "expected": [
|
||||
"0x9d4b6aaa8c292525", "0xb2ad4ab3941b3a3f", "0x17988da38edbf912", "0x566b31d7cceae84a", "0x132f77e29f4e6d40", "0xbdc098ecac64b3dc", "0x1670a7bfbc675d88", "0xda3a9960471e44bb",
|
||||
"0x887e9fda5ec82271", "0x6137f598a67fc0d4", "0xd77b98d7d2025d26", "0x96b472c7560a777d", "0x0c98d0ff42f589f0", "0x4f0f15c9e4eefef2", "0x6bae44572933ab20", "0x2772c0266bdd5308",
|
||||
"0x2fdb9e92d13d07ba", "0x6fe9239589ed122a", "0x7e3955e5a91916fd", "0x966d21685073c9b8", "0xa050194eb84c4104", "0xf7e095dae1d70633", "0x4f49616da323736b", "0x1fcc532045f01cd4",
|
||||
"0x69549d64913199fc", "0x6684a1b257c011b0", "0xff318d60e8970b9a", "0x1429c63759eb426d", "0x3f470be3f4d4815a", "0x6e1d6c6b91b984e2", "0xda2dd7f31fea85ef", "0xb3944abe38dae250"
|
||||
]},
|
||||
{"base_nonce": 1000000, "expected": [
|
||||
"0xd549a905f9121657", "0xd863810b66483e86", "0x65d735cf2e5447af", "0x9cda91c71e0f5790", "0x682c4a1be66d2444", "0xcb64852ca144ffe1", "0x8e72112d5db27544", "0x1a9e6bb74ea3e563",
|
||||
"0x63c279f228faf1bc", "0xf7b77f8435da26a2", "0x3c28bd5b463df46f", "0xd2cdd0e9e968d2bd", "0x46eae240fc0c3f01", "0xbdd4c4e525220d3c", "0x345178054b1eac9d", "0x312f541e08ab8fe9",
|
||||
"0x0cbc6f22e65887ad", "0xf9a0d018f9ce95cf", "0x994d29e263550764", "0x27a37ad3487c4018", "0xcdb7645aa05dac6a", "0x8578f1e945156d7f", "0x8e135ab725d23599", "0xfa902aebc88a620c",
|
||||
"0x5bdb23834f693ba8", "0xae53782ddf331358", "0x2304fcfaad2616fc", "0x21e9216f6c0b14a6", "0xa88e90dc82b85f66", "0xba5aa73b369dd894", "0x97e37652c487cc5e", "0x36cb34ed0c8ce81f"
|
||||
]}
|
||||
],
|
||||
"dataset_head": ["0x3dd50b1f", "0x93123701", "0x48edec90", "0x3a7d2407", "0xcdd88b45", "0x2c37c6b0", "0xb03a0ac0", "0x10e2c6a5", "0x5119540e", "0x3358ef2d", "0x657ac748", "0x09e2105c", "0x25e2929c", "0x5fa0cdcc", "0x4317cabd", "0x6178ee52"],
|
||||
"dataset_last_index": 268435455,
|
||||
"dataset_last": "0xf7b7180e",
|
||||
"dataset_samples": [{"index": 59471966, "value": "0x88717593"}, {"index": 217795994, "value": "0x79f4c2eb"}, {"index": 208353206, "value": "0xa55c10fa"}, {"index": 42483309, "value": "0x1e48deed"}, {"index": 172547758, "value": "0x825505df"}, {"index": 148076330, "value": "0x056e4b2a"}, {"index": 183853158, "value": "0x51c5cdc5"}, {"index": 214389424, "value": "0x9133ff90"}, {"index": 267488061, "value": "0xf3e1f630"}, {"index": 169781097, "value": "0x8c0f4bca"}, {"index": 184093494, "value": "0x13c805d6"}, {"index": 153880993, "value": "0x7cd9cc39"}, {"index": 84977930, "value": "0x54d3ad95"}, {"index": 46426879, "value": "0x3142a64b"}, {"index": 3093825, "value": "0xf690ebf9"}, {"index": 225364072, "value": "0x5f05ecf5"}, {"index": 44593546, "value": "0x708fd94d"}, {"index": 260713159, "value": "0x57e3a213"}, {"index": 168250303, "value": "0xb60de9a6"}, {"index": 52384140, "value": "0xc3cacc1f"}, {"index": 223401610, "value": "0x72a0a173"}, {"index": 45554030, "value": "0x2a13f6c7"}, {"index": 95410555, "value": "0xfb5b4620"}, {"index": 175039924, "value": "0xeae5f6c3"}, {"index": 79171087, "value": "0x38e0ceb1"}, {"index": 267580473, "value": "0x103d3fa4"}, {"index": 24168642, "value": "0xcb03f2ae"}, {"index": 37981670, "value": "0x562969cc"}, {"index": 171551130, "value": "0xf1349e23"}, {"index": 195559979, "value": "0x3fa57b5d"}, {"index": 204611762, "value": "0x4fb463b9"}, {"index": 140997658, "value": "0xb2b70250"}, {"index": 138925853, "value": "0x7f689f8c"}, {"index": 86637313, "value": "0x27e5f93e"}, {"index": 20736778, "value": "0xbaeafce2"}, {"index": 219665210, "value": "0xaf3d3776"}, {"index": 160430336, "value": "0x084c3724"}, {"index": 264654675, "value": "0xffe6a16c"}, {"index": 8013395, "value": "0xd4d0d18d"}, {"index": 228945585, "value": "0x5d78da4e"}, {"index": 213884386, "value": "0xf0bd37c6"}, {"index": 104419827, "value": "0x261f1c0d"}, {"index": 44185464, "value": "0x7bfc0ff6"}, {"index": 142737231, "value": "0x08e5373b"}, {"index": 99284897, "value": "0xd23a038c"}, {"index": 132475900, "value": "0x124d74c0"}, {"index": 61861762, "value": "0xdd74cace"}, {"index": 132056166, "value": "0x5ab7c9ba"}, {"index": 262388043, "value": "0xc05d2942"}, {"index": 91878046, "value": "0x2d1e2f73"}, {"index": 117353561, "value": "0x52a27429"}, {"index": 124768597, "value": "0x0c828d10"}, {"index": 71352993, "value": "0x7514e589"}, {"index": 190698941, "value": "0xdb8ca7d2"}, {"index": 46055428, "value": "0xdb881ba6"}, {"index": 55281366, "value": "0x0ce95e1d"}, {"index": 165145231, "value": "0x3c93a15e"}, {"index": 106810753, "value": "0x73ec90f8"}, {"index": 171985651, "value": "0xe47554ed"}, {"index": 232085256, "value": "0x496959f3"}, {"index": 159510492, "value": "0xebf27832"}, {"index": 40072060, "value": "0x11f5885e"}, {"index": 209107596, "value": "0x69a1ff20"}, {"index": 39023794, "value": "0x804542bd"}],
|
||||
"cache_head": ["0xebd9055c", "0x7eaf6b21", "0x9b610855", "0x134a8562", "0xb10059ba", "0x7f459e58", "0x8f3c7873", "0x3523e609", "0x75b8f4ca", "0x750d31b4", "0x0dae0782", "0x26f10015", "0x7ba87a41", "0x0efd8543", "0x691d6368", "0x8929d967"],
|
||||
"cache_last_line": ["0x51474edf", "0xc3cfba93", "0xf21454cf", "0x9b79baac", "0x4d4fce23", "0xd701dfd5", "0x37357ab3", "0x1be693fa", "0xa701cb3b", "0x7467c620", "0x428184e8", "0xf4010df0", "0xe33aa88f", "0xc78d6d62", "0xae9ba9c0", "0xf4bba97e"],
|
||||
"cache_fnv1a64": "0x448274a57f508cbc"
|
||||
}
|
||||
281
proto-cuda/packs-ca2-era/era-3/kernel.cl
Normal file
281
proto-cuda/packs-ca2-era/era-3/kernel.cl
Normal file
|
|
@ -0,0 +1,281 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal).
|
||||
// Built from source at runtime by proto-opencl/host.c, which passes these defines:
|
||||
// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit)
|
||||
// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default)
|
||||
// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32
|
||||
// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition
|
||||
// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units;
|
||||
// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32.
|
||||
#ifndef IGNEUM_GROUP
|
||||
#define IGNEUM_GROUP 32
|
||||
#endif
|
||||
#ifndef IGNEUM_EXCHANGE
|
||||
#define IGNEUM_EXCHANGE 0
|
||||
#endif
|
||||
#ifdef __OPENCL_VERSION__
|
||||
#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1)))
|
||||
#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n]
|
||||
#if IGNEUM_EXCHANGE == 1
|
||||
#ifdef cl_khr_subgroups
|
||||
#pragma OPENCL EXTENSION cl_khr_subgroups : enable
|
||||
#endif
|
||||
#ifdef cl_khr_subgroup_shuffle
|
||||
#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable
|
||||
#endif
|
||||
#elif IGNEUM_EXCHANGE == 2
|
||||
#pragma OPENCL EXTENSION cl_intel_subgroups : enable
|
||||
#endif
|
||||
#else
|
||||
// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros.
|
||||
#include "emu_opencl.h"
|
||||
#endif
|
||||
|
||||
#if IGNEUM_EXCHANGE == 1
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m))
|
||||
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
|
||||
#elif IGNEUM_EXCHANGE == 2
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m))
|
||||
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
|
||||
#else
|
||||
// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per
|
||||
// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane
|
||||
// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's
|
||||
// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier.
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; }
|
||||
#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; }
|
||||
#endif
|
||||
|
||||
static inline uint splitmix32(uint x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32.
|
||||
static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); }
|
||||
// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x.
|
||||
static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); }
|
||||
static inline uint ds_elem(uint i, uint d0, uint d1) {
|
||||
uint x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
|
||||
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
|
||||
#define MH_CACHE_LINE_MASK 0x003fffffu
|
||||
#define MH_SEGMENT_LINES 64u
|
||||
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
|
||||
static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
|
||||
|
||||
// y = ChaCha12 core(x) + x
|
||||
static inline void mh_chacha_block(const uint* x, uint* y) {
|
||||
for (uint i = 0u; i < 16u; ++i) y[i] = x[i];
|
||||
for (uint r = 0u; r < 6u; ++r) {
|
||||
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
|
||||
}
|
||||
for (uint i = 0u; i < 16u; ++i) y[i] += x[i];
|
||||
}
|
||||
|
||||
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
|
||||
static inline void mh_cache_segment(__global uint* cache, uint seg) {
|
||||
uint prev[16]; uint x[16]; uint y[16];
|
||||
for (uint i = 0u; i < 16u; ++i) prev[i] = 0u;
|
||||
for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) {
|
||||
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
|
||||
x[4] = 0xceed56d7u ^ prev[4];
|
||||
x[5] = 0x9ba270d2u ^ prev[5];
|
||||
x[6] = 0x82caab2du ^ prev[6];
|
||||
x[7] = 0x81ebce0eu ^ prev[7];
|
||||
x[8] = 0x12b6ecf1u ^ prev[8];
|
||||
x[9] = 0xd0f3fd7cu ^ prev[9];
|
||||
x[10] = 0xd872eefeu ^ prev[10];
|
||||
x[11] = 0xc158c7bdu ^ prev[11];
|
||||
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
|
||||
mh_chacha_block(x, y);
|
||||
__global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
|
||||
}
|
||||
}
|
||||
|
||||
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
|
||||
static inline void mh_mixer(uint* s, uint rk) {
|
||||
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
|
||||
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
|
||||
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
|
||||
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
|
||||
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
|
||||
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
|
||||
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
|
||||
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
|
||||
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
|
||||
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
|
||||
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
|
||||
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
|
||||
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
|
||||
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
|
||||
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
|
||||
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
|
||||
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
|
||||
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
|
||||
}
|
||||
|
||||
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
|
||||
static inline void mh_item(__global const uint* cache, uint t, uint* s) {
|
||||
s[0] = 0xceed56d7u;
|
||||
s[1] = 0x9ba270d2u;
|
||||
s[2] = 0x82caab2du;
|
||||
s[3] = 0x81ebce0eu;
|
||||
s[4] = 0x12b6ecf1u;
|
||||
s[5] = 0xd0f3fd7cu;
|
||||
s[6] = 0xd872eefeu;
|
||||
s[7] = 0xc158c7bdu;
|
||||
s[8] = t * 0xf351d601u + 0xc6892460u;
|
||||
s[9] = t * 0xa3bb398fu + 0x25b7228au;
|
||||
s[10] = t * 0xb5a09e35u + 0xcd515004u;
|
||||
s[11] = t * 0x7509c9c1u + 0x2846527au;
|
||||
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
|
||||
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
|
||||
s[14] = t * 0xded91851u + 0x82961bacu;
|
||||
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
|
||||
for (uint r = 0u; r < 8u; ++r) {
|
||||
mh_mixer(s, 0x9E3779B9u * (r + 1u));
|
||||
__global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i];
|
||||
}
|
||||
mh_mixer(s, 0x9E3779B9u * 9u);
|
||||
}
|
||||
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 2 3 8 13 of w.
|
||||
static inline uint mh_j(uint w) { return ((w >> 2u) & 1u) | (((w >> 3u) & 1u) << 1) | (((w >> 8u) & 1u) << 2) | (((w >> 13u) & 1u) << 3); }
|
||||
static inline uint mh_t(uint w) { w = (w & 0x00001fffu) | ((w >> 14u) << 13u); w = (w & 0x000000ffu) | ((w >> 9u) << 8u); w = (w & 0x00000007u) | ((w >> 4u) << 3u); w = (w & 0x00000003u) | ((w >> 3u) << 2u); return w; }
|
||||
static inline uint mh_addr(uint t, uint j) { uint w = t; w = ((w >> 2u) << 3u) | (w & 0x00000003u) | (((j >> 0u) & 1u) << 2u); w = ((w >> 3u) << 4u) | (w & 0x00000007u) | (((j >> 1u) & 1u) << 3u); w = ((w >> 8u) << 9u) | (w & 0x000000ffu) | (((j >> 2u) & 1u) << 8u); w = ((w >> 13u) << 14u) | (w & 0x00001fffu) | (((j >> 3u) & 1u) << 13u); return w; }
|
||||
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
|
||||
static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }
|
||||
|
||||
// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item.
|
||||
// The same constants as memhard.h in this pack (one emitter, three dialects).
|
||||
__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) {
|
||||
uint seg = (uint)get_global_id(0);
|
||||
if (seg < nSegments) mh_cache_segment(cache, seg);
|
||||
}
|
||||
__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) {
|
||||
uint t = (uint)get_global_id(0);
|
||||
if (t < nItems) {
|
||||
uint s[16];
|
||||
mh_item(cache, t, s);
|
||||
for (uint i = 0u; i < 16u; ++i) ds[(ulong)mh_addr(t, i)] = s[i];
|
||||
}
|
||||
}
|
||||
|
||||
// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the
|
||||
// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and
|
||||
// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all).
|
||||
IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) {
|
||||
uint gid = (uint)get_global_id(0);
|
||||
uint lid = (uint)get_local_id(0);
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
#if IGNEUM_EXCHANGE == 0
|
||||
IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
|
||||
uint xk = 0u;
|
||||
#else
|
||||
(void)lid;
|
||||
#endif
|
||||
{ uint x = nonce ^ 0x667d0fbdu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x7b8e5963u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
|
||||
{ uint x = nonce ^ 0x7b8e5963u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x31c67e5eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
|
||||
{ uint x = nonce ^ 0x31c67e5eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4529ddc6u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
|
||||
{ uint x = nonce ^ 0x4529ddc6u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef19d6d8u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
|
||||
{ uint x = nonce ^ 0xef19d6d8u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xaccf6211u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
|
||||
{ uint x = nonce ^ 0xaccf6211u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xda0aed32u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
|
||||
{ uint x = nonce ^ 0xda0aed32u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xabc6df31u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
|
||||
{ uint x = nonce ^ 0xabc6df31u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x667d0fbdu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
|
||||
r7 = r7 ^ r0; // 1 xor
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 1u); r3 = r3 ^ t_; } // 2 shfl
|
||||
r4 = r4 - r1; // 3 sub
|
||||
r2 = r0 * r4 + r2; // 4 mad
|
||||
r4 = rotl_imm(r4, 9u); // 5 rotl
|
||||
r0 = rotr_var(r0, r2); // 6 rotr
|
||||
r6 = r6 ^ ds[((rotl_imm(r7 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r2 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r0 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r1 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
|
||||
r5 = r5 * r4; // 12 mul
|
||||
r7 = r7 ^ ds[((rotl_imm(r6 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 16u); r6 = r6 ^ t_; } // 14 shfl
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r3 = r3 ^ t_; } // 15 shfl
|
||||
r2 = r2 | r7; // 16 or
|
||||
r0 = rotr_var(r0, r6); // 17 rotr
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
|
||||
r7 = rotl_imm(r7, 24u); // 19 rotl
|
||||
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
|
||||
r2 = r2 | r6; // 21 or
|
||||
r6 = r5 * r4 + r6; // 22 mad
|
||||
r1 = r1 | r0; // 23 or
|
||||
r6 = r6 ^ r1; // 24 xor
|
||||
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
|
||||
r7 = rotr_var(r7, r0); // 26 rotr
|
||||
r4 = r5 * r7 + r4; // 27 mad
|
||||
r3 = r2 * r4 + r3; // 28 mad
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
|
||||
r2 = r2 ^ ds[((rotl_imm(r3 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r5 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
|
||||
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
|
||||
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
|
||||
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
|
||||
r2 = mul_hi(r2, r5); // 35 mulhi
|
||||
r4 = r4 ^ r2; // 36 xor
|
||||
r6 = r6 * r5; // 37 mul
|
||||
r7 = r7 ^ r0; // 38 xor
|
||||
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
|
||||
r2 = rotr_var(r2, r3); // 40 rotr
|
||||
r7 = r7 - r0; // 41 sub
|
||||
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 43 shfl
|
||||
r0 = r0 ^ ds[((rotl_imm(r7 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
|
||||
r3 = r3 ^ ds[((rotl_imm(r1 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
|
||||
r6 = r6 ^ ds[((rotl_imm(r3 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
|
||||
r4 = mul_hi(r4, r2); // 48 mulhi
|
||||
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r0 = r0 ^ t_; } // 50 shfl
|
||||
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r2 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
|
||||
r6 = rotl_imm(r6, 12u); // 53 rotl
|
||||
r3 = rotl_imm(r3, 12u); // 54 rotl
|
||||
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
|
||||
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r0 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
|
||||
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
|
||||
r3 = mul_hi(r3, r2); // 60 mulhi
|
||||
r6 = r6 | r4; // 61 or
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r5 = r5 ^ t_; } // 62 shfl
|
||||
r3 = r3 ^ ds[((rotl_imm(r6 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
|
||||
#if IGNEUM_EXCHANGE != 0
|
||||
// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the
|
||||
// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a
|
||||
// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md.
|
||||
IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) {
|
||||
if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); }
|
||||
}
|
||||
#endif
|
||||
163
proto-cuda/packs-ca2-era/era-3/kernel.cu
Normal file
163
proto-cuda/packs-ca2-era/era-3/kernel.cu
Normal file
|
|
@ -0,0 +1,163 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal).
|
||||
// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC.
|
||||
#include <cuda_runtime.h>
|
||||
#include <cstdint>
|
||||
#include "program.h"
|
||||
#include "memhard.h"
|
||||
|
||||
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31.
|
||||
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
|
||||
// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x.
|
||||
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) {
|
||||
uint32_t x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item.
|
||||
// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host.
|
||||
__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) {
|
||||
uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
if (seg < nSegments) mh_cache_segment(cache, seg);
|
||||
}
|
||||
__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
|
||||
uint32_t t = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
if (t < nItems) {
|
||||
uint32_t s[16];
|
||||
mh_item(cache, t, s);
|
||||
for (uint32_t i = 0u; i < 16u; ++i) ds[(size_t)mh_addr(t, i)] = s[i];
|
||||
}
|
||||
}
|
||||
|
||||
// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every
|
||||
// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a
|
||||
// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid.
|
||||
__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) {
|
||||
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
uint32_t nonce = baseNonce + gid;
|
||||
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint32_t x = nonce ^ 0x667d0fbdu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x7b8e5963u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
|
||||
{ uint32_t x = nonce ^ 0x7b8e5963u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x31c67e5eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
|
||||
{ uint32_t x = nonce ^ 0x31c67e5eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4529ddc6u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
|
||||
{ uint32_t x = nonce ^ 0x4529ddc6u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef19d6d8u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
|
||||
{ uint32_t x = nonce ^ 0xef19d6d8u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xaccf6211u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
|
||||
{ uint32_t x = nonce ^ 0xaccf6211u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xda0aed32u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
|
||||
{ uint32_t x = nonce ^ 0xda0aed32u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xabc6df31u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
|
||||
{ uint32_t x = nonce ^ 0xabc6df31u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x667d0fbdu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
|
||||
|
||||
for (uint32_t it = 0u; it < 8u; ++it) {
|
||||
uint32_t sel = r0;
|
||||
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
|
||||
r7 = r7 ^ r0; // 1 xor
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r6, 1); // 2 shfl
|
||||
r4 = r4 - r1; // 3 sub
|
||||
r2 = r0 * r4 + r2; // 4 mad
|
||||
r4 = rotl_imm(r4, 9u); // 5 rotl
|
||||
r0 = rotr_var(r0, r2); // 6 rotr
|
||||
r6 = r6 ^ ds[((rotl_imm(r7 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r2 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r0 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r1 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
|
||||
r5 = r5 * r4; // 12 mul
|
||||
r7 = r7 ^ ds[((rotl_imm(r6 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
|
||||
r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r5, 16); // 14 shfl
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r5, 1); // 15 shfl
|
||||
r2 = r2 | r7; // 16 or
|
||||
r0 = rotr_var(r0, r6); // 17 rotr
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
|
||||
r7 = rotl_imm(r7, 24u); // 19 rotl
|
||||
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
|
||||
r2 = r2 | r6; // 21 or
|
||||
r6 = r5 * r4 + r6; // 22 mad
|
||||
r1 = r1 | r0; // 23 or
|
||||
r6 = r6 ^ r1; // 24 xor
|
||||
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
|
||||
r7 = rotr_var(r7, r0); // 26 rotr
|
||||
r4 = r5 * r7 + r4; // 27 mad
|
||||
r3 = r2 * r4 + r3; // 28 mad
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
|
||||
r2 = r2 ^ ds[((rotl_imm(r3 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r5 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
|
||||
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
|
||||
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
|
||||
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
|
||||
r2 = __umulhi(r2, r5); // 35 mulhi
|
||||
r4 = r4 ^ r2; // 36 xor
|
||||
r6 = r6 * r5; // 37 mul
|
||||
r7 = r7 ^ r0; // 38 xor
|
||||
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
|
||||
r2 = rotr_var(r2, r3); // 40 rotr
|
||||
r7 = r7 - r0; // 41 sub
|
||||
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 43 shfl
|
||||
r0 = r0 ^ ds[((rotl_imm(r7 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
|
||||
r3 = r3 ^ ds[((rotl_imm(r1 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
|
||||
r6 = r6 ^ ds[((rotl_imm(r3 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
|
||||
r4 = __umulhi(r4, r2); // 48 mulhi
|
||||
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
|
||||
r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r7, 4); // 50 shfl
|
||||
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r2 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
|
||||
r6 = rotl_imm(r6, 12u); // 53 rotl
|
||||
r3 = rotl_imm(r3, 12u); // 54 rotl
|
||||
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
|
||||
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r0 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
|
||||
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
|
||||
r3 = __umulhi(r3, r2); // 60 mulhi
|
||||
r6 = r6 | r4; // 61 or
|
||||
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 62 shfl
|
||||
r3 = r3 ^ ds[((rotl_imm(r6 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
|
||||
}
|
||||
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
|
||||
}
|
||||
|
||||
// Host-side launch wrappers. Declared in program.h, called from host.cu.
|
||||
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) {
|
||||
if (nSegments == 0u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 256u;
|
||||
uint32_t grid = (nSegments + block - 1u) / block;
|
||||
igneum_cache_fill<<<grid, block>>>(cache, nSegments);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
|
||||
if (nItems == 0u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 256u;
|
||||
uint32_t grid = (nItems + block - 1u) / block;
|
||||
igneum_build<<<grid, block>>>(ds, cache, nItems);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
uint32_t nonces, uint32_t blockWarps) {
|
||||
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 32u * blockWarps;
|
||||
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
|
||||
igneum_hash<<<nonces / block, block>>>(ds, out, baseNonce, mask);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
|
||||
cudaFuncAttributes attr;
|
||||
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash);
|
||||
if (e != cudaSuccess) return e;
|
||||
*numRegs = attr.numRegs;
|
||||
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0);
|
||||
}
|
||||
375
proto-cuda/packs-ca2-era/era-3/kernel_bound.cl
Normal file
375
proto-cuda/packs-ca2-era/era-3/kernel_bound.cl
Normal file
|
|
@ -0,0 +1,375 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal).
|
||||
// Built from source at runtime by proto-opencl/host.c, which passes these defines:
|
||||
// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit)
|
||||
// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default)
|
||||
// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32
|
||||
// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition
|
||||
// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units;
|
||||
// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32.
|
||||
#ifndef IGNEUM_GROUP
|
||||
#define IGNEUM_GROUP 32
|
||||
#endif
|
||||
#ifndef IGNEUM_EXCHANGE
|
||||
#define IGNEUM_EXCHANGE 0
|
||||
#endif
|
||||
#ifdef __OPENCL_VERSION__
|
||||
#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1)))
|
||||
#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n]
|
||||
#if IGNEUM_EXCHANGE == 1
|
||||
#ifdef cl_khr_subgroups
|
||||
#pragma OPENCL EXTENSION cl_khr_subgroups : enable
|
||||
#endif
|
||||
#ifdef cl_khr_subgroup_shuffle
|
||||
#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable
|
||||
#endif
|
||||
#elif IGNEUM_EXCHANGE == 2
|
||||
#pragma OPENCL EXTENSION cl_intel_subgroups : enable
|
||||
#endif
|
||||
#else
|
||||
// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros.
|
||||
#include "emu_opencl.h"
|
||||
#endif
|
||||
|
||||
#if IGNEUM_EXCHANGE == 1
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m))
|
||||
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
|
||||
#elif IGNEUM_EXCHANGE == 2
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m))
|
||||
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
|
||||
#else
|
||||
// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per
|
||||
// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane
|
||||
// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's
|
||||
// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier.
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; }
|
||||
#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; }
|
||||
#endif
|
||||
|
||||
static inline uint splitmix32(uint x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32.
|
||||
static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); }
|
||||
// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x.
|
||||
static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); }
|
||||
static inline uint ds_elem(uint i, uint d0, uint d1) {
|
||||
uint x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
|
||||
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
|
||||
#define MH_CACHE_LINE_MASK 0x003fffffu
|
||||
#define MH_SEGMENT_LINES 64u
|
||||
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
|
||||
static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
|
||||
|
||||
// y = ChaCha12 core(x) + x
|
||||
static inline void mh_chacha_block(const uint* x, uint* y) {
|
||||
for (uint i = 0u; i < 16u; ++i) y[i] = x[i];
|
||||
for (uint r = 0u; r < 6u; ++r) {
|
||||
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
|
||||
}
|
||||
for (uint i = 0u; i < 16u; ++i) y[i] += x[i];
|
||||
}
|
||||
|
||||
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
|
||||
static inline void mh_cache_segment(__global uint* cache, uint seg) {
|
||||
uint prev[16]; uint x[16]; uint y[16];
|
||||
for (uint i = 0u; i < 16u; ++i) prev[i] = 0u;
|
||||
for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) {
|
||||
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
|
||||
x[4] = 0xceed56d7u ^ prev[4];
|
||||
x[5] = 0x9ba270d2u ^ prev[5];
|
||||
x[6] = 0x82caab2du ^ prev[6];
|
||||
x[7] = 0x81ebce0eu ^ prev[7];
|
||||
x[8] = 0x12b6ecf1u ^ prev[8];
|
||||
x[9] = 0xd0f3fd7cu ^ prev[9];
|
||||
x[10] = 0xd872eefeu ^ prev[10];
|
||||
x[11] = 0xc158c7bdu ^ prev[11];
|
||||
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
|
||||
mh_chacha_block(x, y);
|
||||
__global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
|
||||
}
|
||||
}
|
||||
|
||||
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
|
||||
static inline void mh_mixer(uint* s, uint rk) {
|
||||
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
|
||||
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
|
||||
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
|
||||
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
|
||||
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
|
||||
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
|
||||
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
|
||||
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
|
||||
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
|
||||
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
|
||||
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
|
||||
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
|
||||
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
|
||||
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
|
||||
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
|
||||
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
|
||||
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
|
||||
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
|
||||
}
|
||||
|
||||
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
|
||||
static inline void mh_item(__global const uint* cache, uint t, uint* s) {
|
||||
s[0] = 0xceed56d7u;
|
||||
s[1] = 0x9ba270d2u;
|
||||
s[2] = 0x82caab2du;
|
||||
s[3] = 0x81ebce0eu;
|
||||
s[4] = 0x12b6ecf1u;
|
||||
s[5] = 0xd0f3fd7cu;
|
||||
s[6] = 0xd872eefeu;
|
||||
s[7] = 0xc158c7bdu;
|
||||
s[8] = t * 0xf351d601u + 0xc6892460u;
|
||||
s[9] = t * 0xa3bb398fu + 0x25b7228au;
|
||||
s[10] = t * 0xb5a09e35u + 0xcd515004u;
|
||||
s[11] = t * 0x7509c9c1u + 0x2846527au;
|
||||
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
|
||||
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
|
||||
s[14] = t * 0xded91851u + 0x82961bacu;
|
||||
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
|
||||
for (uint r = 0u; r < 8u; ++r) {
|
||||
mh_mixer(s, 0x9E3779B9u * (r + 1u));
|
||||
__global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i];
|
||||
}
|
||||
mh_mixer(s, 0x9E3779B9u * 9u);
|
||||
}
|
||||
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 2 3 8 13 of w.
|
||||
static inline uint mh_j(uint w) { return ((w >> 2u) & 1u) | (((w >> 3u) & 1u) << 1) | (((w >> 8u) & 1u) << 2) | (((w >> 13u) & 1u) << 3); }
|
||||
static inline uint mh_t(uint w) { w = (w & 0x00001fffu) | ((w >> 14u) << 13u); w = (w & 0x000000ffu) | ((w >> 9u) << 8u); w = (w & 0x00000007u) | ((w >> 4u) << 3u); w = (w & 0x00000003u) | ((w >> 3u) << 2u); return w; }
|
||||
static inline uint mh_addr(uint t, uint j) { uint w = t; w = ((w >> 2u) << 3u) | (w & 0x00000003u) | (((j >> 0u) & 1u) << 2u); w = ((w >> 3u) << 4u) | (w & 0x00000007u) | (((j >> 1u) & 1u) << 3u); w = ((w >> 8u) << 9u) | (w & 0x000000ffu) | (((j >> 2u) & 1u) << 8u); w = ((w >> 13u) << 14u) | (w & 0x00001fffu) | (((j >> 3u) & 1u) << 13u); return w; }
|
||||
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
|
||||
static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }
|
||||
|
||||
// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item.
|
||||
// The same constants as memhard.h in this pack (one emitter, three dialects).
|
||||
__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) {
|
||||
uint seg = (uint)get_global_id(0);
|
||||
if (seg < nSegments) mh_cache_segment(cache, seg);
|
||||
}
|
||||
__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) {
|
||||
uint t = (uint)get_global_id(0);
|
||||
if (t < nItems) {
|
||||
uint s[16];
|
||||
mh_item(cache, t, s);
|
||||
for (uint i = 0u; i < 16u; ++i) ds[(ulong)mh_addr(t, i)] = s[i];
|
||||
}
|
||||
}
|
||||
|
||||
// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the
|
||||
// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and
|
||||
// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all).
|
||||
IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) {
|
||||
uint gid = (uint)get_global_id(0);
|
||||
uint lid = (uint)get_local_id(0);
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
#if IGNEUM_EXCHANGE == 0
|
||||
IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
|
||||
uint xk = 0u;
|
||||
#else
|
||||
(void)lid;
|
||||
#endif
|
||||
{ uint x = nonce ^ 0x667d0fbdu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x7b8e5963u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
|
||||
{ uint x = nonce ^ 0x7b8e5963u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x31c67e5eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
|
||||
{ uint x = nonce ^ 0x31c67e5eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4529ddc6u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
|
||||
{ uint x = nonce ^ 0x4529ddc6u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef19d6d8u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
|
||||
{ uint x = nonce ^ 0xef19d6d8u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xaccf6211u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
|
||||
{ uint x = nonce ^ 0xaccf6211u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xda0aed32u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
|
||||
{ uint x = nonce ^ 0xda0aed32u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xabc6df31u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
|
||||
{ uint x = nonce ^ 0xabc6df31u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x667d0fbdu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
|
||||
r7 = r7 ^ r0; // 1 xor
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 1u); r3 = r3 ^ t_; } // 2 shfl
|
||||
r4 = r4 - r1; // 3 sub
|
||||
r2 = r0 * r4 + r2; // 4 mad
|
||||
r4 = rotl_imm(r4, 9u); // 5 rotl
|
||||
r0 = rotr_var(r0, r2); // 6 rotr
|
||||
r6 = r6 ^ ds[((rotl_imm(r7 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r2 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r0 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r1 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
|
||||
r5 = r5 * r4; // 12 mul
|
||||
r7 = r7 ^ ds[((rotl_imm(r6 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 16u); r6 = r6 ^ t_; } // 14 shfl
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r3 = r3 ^ t_; } // 15 shfl
|
||||
r2 = r2 | r7; // 16 or
|
||||
r0 = rotr_var(r0, r6); // 17 rotr
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
|
||||
r7 = rotl_imm(r7, 24u); // 19 rotl
|
||||
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
|
||||
r2 = r2 | r6; // 21 or
|
||||
r6 = r5 * r4 + r6; // 22 mad
|
||||
r1 = r1 | r0; // 23 or
|
||||
r6 = r6 ^ r1; // 24 xor
|
||||
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
|
||||
r7 = rotr_var(r7, r0); // 26 rotr
|
||||
r4 = r5 * r7 + r4; // 27 mad
|
||||
r3 = r2 * r4 + r3; // 28 mad
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
|
||||
r2 = r2 ^ ds[((rotl_imm(r3 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r5 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
|
||||
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
|
||||
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
|
||||
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
|
||||
r2 = mul_hi(r2, r5); // 35 mulhi
|
||||
r4 = r4 ^ r2; // 36 xor
|
||||
r6 = r6 * r5; // 37 mul
|
||||
r7 = r7 ^ r0; // 38 xor
|
||||
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
|
||||
r2 = rotr_var(r2, r3); // 40 rotr
|
||||
r7 = r7 - r0; // 41 sub
|
||||
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 43 shfl
|
||||
r0 = r0 ^ ds[((rotl_imm(r7 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
|
||||
r3 = r3 ^ ds[((rotl_imm(r1 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
|
||||
r6 = r6 ^ ds[((rotl_imm(r3 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
|
||||
r4 = mul_hi(r4, r2); // 48 mulhi
|
||||
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r0 = r0 ^ t_; } // 50 shfl
|
||||
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r2 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
|
||||
r6 = rotl_imm(r6, 12u); // 53 rotl
|
||||
r3 = rotl_imm(r3, 12u); // 54 rotl
|
||||
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
|
||||
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r0 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
|
||||
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
|
||||
r3 = mul_hi(r3, r2); // 60 mulhi
|
||||
r6 = r6 | r4; // 61 or
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r5 = r5 ^ t_; } // 62 shfl
|
||||
r3 = r3 ^ ds[((rotl_imm(r6 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
|
||||
#if IGNEUM_EXCHANGE != 0
|
||||
// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the
|
||||
// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a
|
||||
// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md.
|
||||
IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) {
|
||||
if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); }
|
||||
}
|
||||
#endif
|
||||
|
||||
// Header-bound variant (bind.rs): the init words come from initw, not SEEDW. Same body as igneum_hash.
|
||||
IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw) {
|
||||
uint gid = (uint)get_global_id(0);
|
||||
uint lid = (uint)get_local_id(0);
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7];
|
||||
#if IGNEUM_EXCHANGE == 0
|
||||
IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
|
||||
uint xk = 0u;
|
||||
#else
|
||||
(void)lid;
|
||||
#endif
|
||||
{ uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; }
|
||||
{ uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; }
|
||||
{ uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; }
|
||||
{ uint x = nonce ^ iw3; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw4; }
|
||||
{ uint x = nonce ^ iw4; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw5; }
|
||||
{ uint x = nonce ^ iw5; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw6; }
|
||||
{ uint x = nonce ^ iw6; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw7; }
|
||||
{ uint x = nonce ^ iw7; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw0; }
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
|
||||
r7 = r7 ^ r0; // 1 xor
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 1u); r3 = r3 ^ t_; } // 2 shfl
|
||||
r4 = r4 - r1; // 3 sub
|
||||
r2 = r0 * r4 + r2; // 4 mad
|
||||
r4 = rotl_imm(r4, 9u); // 5 rotl
|
||||
r0 = rotr_var(r0, r2); // 6 rotr
|
||||
r6 = r6 ^ ds[((rotl_imm(r7 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r2 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r0 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r1 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
|
||||
r5 = r5 * r4; // 12 mul
|
||||
r7 = r7 ^ ds[((rotl_imm(r6 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 16u); r6 = r6 ^ t_; } // 14 shfl
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r3 = r3 ^ t_; } // 15 shfl
|
||||
r2 = r2 | r7; // 16 or
|
||||
r0 = rotr_var(r0, r6); // 17 rotr
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
|
||||
r7 = rotl_imm(r7, 24u); // 19 rotl
|
||||
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
|
||||
r2 = r2 | r6; // 21 or
|
||||
r6 = r5 * r4 + r6; // 22 mad
|
||||
r1 = r1 | r0; // 23 or
|
||||
r6 = r6 ^ r1; // 24 xor
|
||||
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
|
||||
r7 = rotr_var(r7, r0); // 26 rotr
|
||||
r4 = r5 * r7 + r4; // 27 mad
|
||||
r3 = r2 * r4 + r3; // 28 mad
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
|
||||
r2 = r2 ^ ds[((rotl_imm(r3 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r5 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
|
||||
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
|
||||
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
|
||||
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
|
||||
r2 = mul_hi(r2, r5); // 35 mulhi
|
||||
r4 = r4 ^ r2; // 36 xor
|
||||
r6 = r6 * r5; // 37 mul
|
||||
r7 = r7 ^ r0; // 38 xor
|
||||
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
|
||||
r2 = rotr_var(r2, r3); // 40 rotr
|
||||
r7 = r7 - r0; // 41 sub
|
||||
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 43 shfl
|
||||
r0 = r0 ^ ds[((rotl_imm(r7 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
|
||||
r3 = r3 ^ ds[((rotl_imm(r1 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
|
||||
r6 = r6 ^ ds[((rotl_imm(r3 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
|
||||
r4 = mul_hi(r4, r2); // 48 mulhi
|
||||
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r0 = r0 ^ t_; } // 50 shfl
|
||||
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r2 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
|
||||
r6 = rotl_imm(r6, 12u); // 53 rotl
|
||||
r3 = rotl_imm(r3, 12u); // 54 rotl
|
||||
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
|
||||
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r0 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
|
||||
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
|
||||
r3 = mul_hi(r3, r2); // 60 mulhi
|
||||
r6 = r6 | r4; // 61 or
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r5 = r5 ^ t_; } // 62 shfl
|
||||
r3 = r3 ^ ds[((rotl_imm(r6 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
123
proto-cuda/packs-ca2-era/era-3/kernel_bound.cu
Normal file
123
proto-cuda/packs-ca2-era/era-3/kernel_bound.cu
Normal file
|
|
@ -0,0 +1,123 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// Header-bound twin of igneum_hash in kernel.cu: the init words come from a kernel argument, not SEEDW.
|
||||
// Host declarations (also in program_bound.h if present):
|
||||
// struct IgneumInitWords { uint32_t w[8]; };
|
||||
// cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
// IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps);
|
||||
// cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
|
||||
#include <cuda_runtime.h>
|
||||
#include <cstdint>
|
||||
#include "program.h"
|
||||
|
||||
struct IgneumInitWords { uint32_t w[8]; };
|
||||
|
||||
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
|
||||
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
|
||||
__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw) {
|
||||
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
uint32_t nonce = baseNonce + gid;
|
||||
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint32_t x = nonce ^ iw.w[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw.w[1]; }
|
||||
{ uint32_t x = nonce ^ iw.w[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw.w[2]; }
|
||||
{ uint32_t x = nonce ^ iw.w[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw.w[3]; }
|
||||
{ uint32_t x = nonce ^ iw.w[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw.w[4]; }
|
||||
{ uint32_t x = nonce ^ iw.w[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw.w[5]; }
|
||||
{ uint32_t x = nonce ^ iw.w[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw.w[6]; }
|
||||
{ uint32_t x = nonce ^ iw.w[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw.w[7]; }
|
||||
{ uint32_t x = nonce ^ iw.w[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw.w[0]; }
|
||||
|
||||
for (uint32_t it = 0u; it < 8u; ++it) {
|
||||
uint32_t sel = r0;
|
||||
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
|
||||
r7 = r7 ^ r0; // 1 xor
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r6, 1); // 2 shfl
|
||||
r4 = r4 - r1; // 3 sub
|
||||
r2 = r0 * r4 + r2; // 4 mad
|
||||
r4 = rotl_imm(r4, 9u); // 5 rotl
|
||||
r0 = rotr_var(r0, r2); // 6 rotr
|
||||
r6 = r6 ^ ds[((rotl_imm(r7 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r2 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r0 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r1 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
|
||||
r5 = r5 * r4; // 12 mul
|
||||
r7 = r7 ^ ds[((rotl_imm(r6 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
|
||||
r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r5, 16); // 14 shfl
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r5, 1); // 15 shfl
|
||||
r2 = r2 | r7; // 16 or
|
||||
r0 = rotr_var(r0, r6); // 17 rotr
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
|
||||
r7 = rotl_imm(r7, 24u); // 19 rotl
|
||||
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
|
||||
r2 = r2 | r6; // 21 or
|
||||
r6 = r5 * r4 + r6; // 22 mad
|
||||
r1 = r1 | r0; // 23 or
|
||||
r6 = r6 ^ r1; // 24 xor
|
||||
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
|
||||
r7 = rotr_var(r7, r0); // 26 rotr
|
||||
r4 = r5 * r7 + r4; // 27 mad
|
||||
r3 = r2 * r4 + r3; // 28 mad
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
|
||||
r2 = r2 ^ ds[((rotl_imm(r3 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r5 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
|
||||
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
|
||||
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
|
||||
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
|
||||
r2 = __umulhi(r2, r5); // 35 mulhi
|
||||
r4 = r4 ^ r2; // 36 xor
|
||||
r6 = r6 * r5; // 37 mul
|
||||
r7 = r7 ^ r0; // 38 xor
|
||||
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
|
||||
r2 = rotr_var(r2, r3); // 40 rotr
|
||||
r7 = r7 - r0; // 41 sub
|
||||
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 43 shfl
|
||||
r0 = r0 ^ ds[((rotl_imm(r7 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
|
||||
r3 = r3 ^ ds[((rotl_imm(r1 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
|
||||
r6 = r6 ^ ds[((rotl_imm(r3 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
|
||||
r4 = __umulhi(r4, r2); // 48 mulhi
|
||||
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
|
||||
r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r7, 4); // 50 shfl
|
||||
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r2 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
|
||||
r6 = rotl_imm(r6, 12u); // 53 rotl
|
||||
r3 = rotl_imm(r3, 12u); // 54 rotl
|
||||
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
|
||||
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r0 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
|
||||
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
|
||||
r3 = __umulhi(r3, r2); // 60 mulhi
|
||||
r6 = r6 | r4; // 61 or
|
||||
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 62 shfl
|
||||
r3 = r3 ^ ds[((rotl_imm(r6 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
|
||||
}
|
||||
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps) {
|
||||
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 32u * blockWarps;
|
||||
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
|
||||
igneum_hash_bound<<<nonces / block, block>>>(ds, out, baseNonce, mask, iw);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
|
||||
cudaFuncAttributes attr;
|
||||
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash_bound);
|
||||
if (e != cudaSuccess) return e;
|
||||
*numRegs = attr.numRegs;
|
||||
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash_bound, (int)(32u * blockWarps), 0);
|
||||
}
|
||||
112
proto-cuda/packs-ca2-era/era-3/memhard.h
Normal file
112
proto-cuda/packs-ca2-era/era-3/memhard.h
Normal file
|
|
@ -0,0 +1,112 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// Memory-hard dataset core, the same text that the Mac's Metal kernels and CPU verifier were checked against.
|
||||
// Included by kernel.cu (device), host.cu (host reference) and proto-opencl/host.c (C99 host reference).
|
||||
// See proto-metal/MEMHARD.md for the construction. kernel.cl carries the same text in OpenCL C.
|
||||
#pragma once
|
||||
#ifdef __cplusplus
|
||||
#include <cstdint>
|
||||
#else
|
||||
#include <stdint.h>
|
||||
#endif
|
||||
#if defined(__CUDACC__)
|
||||
#define IGNEUM_HD __host__ __device__ __forceinline__
|
||||
#elif defined(_MSC_VER) && !defined(__cplusplus)
|
||||
#define IGNEUM_HD static __inline
|
||||
#else
|
||||
#define IGNEUM_HD static inline
|
||||
#endif
|
||||
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
|
||||
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
|
||||
#define MH_CACHE_LINE_MASK 0x003fffffu
|
||||
#define MH_SEGMENT_LINES 64u
|
||||
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
|
||||
IGNEUM_HD uint32_t mh_rotl(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
|
||||
|
||||
// y = ChaCha12 core(x) + x
|
||||
IGNEUM_HD void mh_chacha_block(const uint32_t* x, uint32_t* y) {
|
||||
for (uint32_t i = 0u; i < 16u; ++i) y[i] = x[i];
|
||||
for (uint32_t r = 0u; r < 6u; ++r) {
|
||||
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
|
||||
}
|
||||
for (uint32_t i = 0u; i < 16u; ++i) y[i] += x[i];
|
||||
}
|
||||
|
||||
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
|
||||
IGNEUM_HD void mh_cache_segment(uint32_t* cache, uint32_t seg) {
|
||||
uint32_t prev[16]; uint32_t x[16]; uint32_t y[16];
|
||||
for (uint32_t i = 0u; i < 16u; ++i) prev[i] = 0u;
|
||||
for (uint32_t j = 0u; j < MH_SEGMENT_LINES; ++j) {
|
||||
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
|
||||
x[4] = 0xceed56d7u ^ prev[4];
|
||||
x[5] = 0x9ba270d2u ^ prev[5];
|
||||
x[6] = 0x82caab2du ^ prev[6];
|
||||
x[7] = 0x81ebce0eu ^ prev[7];
|
||||
x[8] = 0x12b6ecf1u ^ prev[8];
|
||||
x[9] = 0xd0f3fd7cu ^ prev[9];
|
||||
x[10] = 0xd872eefeu ^ prev[10];
|
||||
x[11] = 0xc158c7bdu ^ prev[11];
|
||||
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
|
||||
mh_chacha_block(x, y);
|
||||
uint32_t* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
|
||||
for (uint32_t i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
|
||||
}
|
||||
}
|
||||
|
||||
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
|
||||
IGNEUM_HD void mh_mixer(uint32_t* s, uint32_t rk) {
|
||||
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
|
||||
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
|
||||
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
|
||||
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
|
||||
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
|
||||
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
|
||||
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
|
||||
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
|
||||
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
|
||||
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
|
||||
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
|
||||
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
|
||||
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
|
||||
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
|
||||
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
|
||||
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
|
||||
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
|
||||
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
|
||||
}
|
||||
|
||||
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
|
||||
IGNEUM_HD void mh_item(const uint32_t* cache, uint32_t t, uint32_t* s) {
|
||||
s[0] = 0xceed56d7u;
|
||||
s[1] = 0x9ba270d2u;
|
||||
s[2] = 0x82caab2du;
|
||||
s[3] = 0x81ebce0eu;
|
||||
s[4] = 0x12b6ecf1u;
|
||||
s[5] = 0xd0f3fd7cu;
|
||||
s[6] = 0xd872eefeu;
|
||||
s[7] = 0xc158c7bdu;
|
||||
s[8] = t * 0xf351d601u + 0xc6892460u;
|
||||
s[9] = t * 0xa3bb398fu + 0x25b7228au;
|
||||
s[10] = t * 0xb5a09e35u + 0xcd515004u;
|
||||
s[11] = t * 0x7509c9c1u + 0x2846527au;
|
||||
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
|
||||
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
|
||||
s[14] = t * 0xded91851u + 0x82961bacu;
|
||||
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
|
||||
for (uint32_t r = 0u; r < 8u; ++r) {
|
||||
mh_mixer(s, 0x9E3779B9u * (r + 1u));
|
||||
const uint32_t* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
|
||||
for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= line[i];
|
||||
}
|
||||
mh_mixer(s, 0x9E3779B9u * 9u);
|
||||
}
|
||||
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 2 3 8 13 of w.
|
||||
IGNEUM_HD uint32_t mh_j(uint32_t w) { return ((w >> 2u) & 1u) | (((w >> 3u) & 1u) << 1) | (((w >> 8u) & 1u) << 2) | (((w >> 13u) & 1u) << 3); }
|
||||
IGNEUM_HD uint32_t mh_t(uint32_t w) { w = (w & 0x00001fffu) | ((w >> 14u) << 13u); w = (w & 0x000000ffu) | ((w >> 9u) << 8u); w = (w & 0x00000007u) | ((w >> 4u) << 3u); w = (w & 0x00000003u) | ((w >> 3u) << 2u); return w; }
|
||||
IGNEUM_HD uint32_t mh_addr(uint32_t t, uint32_t j) { uint32_t w = t; w = ((w >> 2u) << 3u) | (w & 0x00000003u) | (((j >> 0u) & 1u) << 2u); w = ((w >> 3u) << 4u) | (w & 0x00000007u) | (((j >> 1u) & 1u) << 3u); w = ((w >> 8u) << 9u) | (w & 0x000000ffu) | (((j >> 2u) & 1u) << 8u); w = ((w >> 13u) << 14u) | (w & 0x00001fffu) | (((j >> 3u) & 1u) << 13u); return w; }
|
||||
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
|
||||
IGNEUM_HD uint32_t mh_word(const uint32_t* cache, uint32_t w) { uint32_t s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }
|
||||
109
proto-cuda/packs-ca2-era/era-3/memhard.metal
Normal file
109
proto-cuda/packs-ca2-era/era-3/memhard.metal
Normal file
|
|
@ -0,0 +1,109 @@
|
|||
#include <metal_stdlib>
|
||||
using namespace metal;
|
||||
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
|
||||
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
|
||||
#define MH_CACHE_LINE_MASK 0x003fffffu
|
||||
#define MH_SEGMENT_LINES 64u
|
||||
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
|
||||
inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
|
||||
|
||||
// y = ChaCha12 core(x) + x
|
||||
inline void mh_chacha_block(const thread uint* x, thread uint* y) {
|
||||
for (uint i = 0u; i < 16u; ++i) y[i] = x[i];
|
||||
for (uint r = 0u; r < 6u; ++r) {
|
||||
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
|
||||
}
|
||||
for (uint i = 0u; i < 16u; ++i) y[i] += x[i];
|
||||
}
|
||||
|
||||
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
|
||||
inline void mh_cache_segment(device uint* cache, uint seg) {
|
||||
uint prev[16]; uint x[16]; uint y[16];
|
||||
for (uint i = 0u; i < 16u; ++i) prev[i] = 0u;
|
||||
for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) {
|
||||
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
|
||||
x[4] = 0xceed56d7u ^ prev[4];
|
||||
x[5] = 0x9ba270d2u ^ prev[5];
|
||||
x[6] = 0x82caab2du ^ prev[6];
|
||||
x[7] = 0x81ebce0eu ^ prev[7];
|
||||
x[8] = 0x12b6ecf1u ^ prev[8];
|
||||
x[9] = 0xd0f3fd7cu ^ prev[9];
|
||||
x[10] = 0xd872eefeu ^ prev[10];
|
||||
x[11] = 0xc158c7bdu ^ prev[11];
|
||||
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
|
||||
mh_chacha_block(x, y);
|
||||
device uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
|
||||
}
|
||||
}
|
||||
|
||||
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
|
||||
inline void mh_mixer(thread uint* s, uint rk) {
|
||||
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
|
||||
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
|
||||
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
|
||||
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
|
||||
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
|
||||
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
|
||||
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
|
||||
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
|
||||
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
|
||||
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
|
||||
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
|
||||
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
|
||||
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
|
||||
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
|
||||
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
|
||||
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
|
||||
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
|
||||
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
|
||||
}
|
||||
|
||||
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
|
||||
inline void mh_item(device const uint* cache, uint t, thread uint* s) {
|
||||
s[0] = 0xceed56d7u;
|
||||
s[1] = 0x9ba270d2u;
|
||||
s[2] = 0x82caab2du;
|
||||
s[3] = 0x81ebce0eu;
|
||||
s[4] = 0x12b6ecf1u;
|
||||
s[5] = 0xd0f3fd7cu;
|
||||
s[6] = 0xd872eefeu;
|
||||
s[7] = 0xc158c7bdu;
|
||||
s[8] = t * 0xf351d601u + 0xc6892460u;
|
||||
s[9] = t * 0xa3bb398fu + 0x25b7228au;
|
||||
s[10] = t * 0xb5a09e35u + 0xcd515004u;
|
||||
s[11] = t * 0x7509c9c1u + 0x2846527au;
|
||||
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
|
||||
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
|
||||
s[14] = t * 0xded91851u + 0x82961bacu;
|
||||
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
|
||||
for (uint r = 0u; r < 8u; ++r) {
|
||||
mh_mixer(s, 0x9E3779B9u * (r + 1u));
|
||||
device const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i];
|
||||
}
|
||||
mh_mixer(s, 0x9E3779B9u * 9u);
|
||||
}
|
||||
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 2 3 8 13 of w.
|
||||
inline uint mh_j(uint w) { return ((w >> 2u) & 1u) | (((w >> 3u) & 1u) << 1) | (((w >> 8u) & 1u) << 2) | (((w >> 13u) & 1u) << 3); }
|
||||
inline uint mh_t(uint w) { w = (w & 0x00001fffu) | ((w >> 14u) << 13u); w = (w & 0x000000ffu) | ((w >> 9u) << 8u); w = (w & 0x00000007u) | ((w >> 4u) << 3u); w = (w & 0x00000003u) | ((w >> 3u) << 2u); return w; }
|
||||
inline uint mh_addr(uint t, uint j) { uint w = t; w = ((w >> 2u) << 3u) | (w & 0x00000003u) | (((j >> 0u) & 1u) << 2u); w = ((w >> 3u) << 4u) | (w & 0x00000007u) | (((j >> 1u) & 1u) << 3u); w = ((w >> 8u) << 9u) | (w & 0x000000ffu) | (((j >> 2u) & 1u) << 8u); w = ((w >> 13u) << 14u) | (w & 0x00001fffu) | (((j >> 3u) & 1u) << 13u); return w; }
|
||||
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
|
||||
inline uint mh_word(device const uint* cache, uint w) { uint s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }
|
||||
|
||||
// One thread per segment (2^16 threads).
|
||||
kernel void igneum_cache_fill(device uint* cache [[buffer(0)]], uint gid [[thread_position_in_grid]]) {
|
||||
mh_cache_segment(cache, gid);
|
||||
}
|
||||
// One thread per 64-byte item (dataset words / 16 threads).
|
||||
kernel void igneum_build(device const uint* cache [[buffer(0)]], device uint* dataset [[buffer(1)]],
|
||||
uint gid [[thread_position_in_grid]]) {
|
||||
uint s[16];
|
||||
mh_item(cache, gid, s);
|
||||
for (uint i = 0u; i < 16u; ++i) dataset[mh_addr(gid, i)] = s[i];
|
||||
}
|
||||
76
proto-cuda/packs-ca2-era/era-3/program.h
Normal file
76
proto-cuda/packs-ca2-era/era-3/program.h
Normal file
|
|
@ -0,0 +1,76 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// Program metadata for host.cu plus the launch wrappers defined in kernel.cu.
|
||||
// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros.
|
||||
#pragma once
|
||||
#ifdef __cplusplus
|
||||
#include <cstdint>
|
||||
#else
|
||||
#include <stdint.h>
|
||||
#endif
|
||||
#ifndef IGNEUM_NO_CUDA
|
||||
#include <cuda_runtime.h>
|
||||
#endif
|
||||
|
||||
#define IGNEUM_SEED_STRING "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000"
|
||||
#define IGNEUM_SEED_BYTES_HEX "edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07"
|
||||
#define IGNEUM_GENERATOR 3
|
||||
#define IGNEUM_PROGRAM_ATTEMPT 0
|
||||
#define IGNEUM_PROGRAM_ID 0x73bcbfe8ccf988f1ull
|
||||
#define IGNEUM_DAY_STRING "bytes:69676e65756d2d6461792ffa50000000000000"
|
||||
#define IGNEUM_DAY_BYTES_HEX "69676e65756d2d6461792ffa50000000000000"
|
||||
#define IGNEUM_DAY0 0xceed56d7u
|
||||
#define IGNEUM_DAY1 0x9ba270d2u
|
||||
#define IGNEUM_DATASET_LOG2 28
|
||||
#define IGNEUM_MASK 0x0fffffffu
|
||||
#define IGNEUM_LANES 32
|
||||
#define IGNEUM_ITERATIONS 8
|
||||
#define IGNEUM_INSTR_COUNT 64
|
||||
#define IGNEUM_LOADS_PER_HASH 128
|
||||
#define IGNEUM_WIDE_LOADS_PER_HASH 0
|
||||
#define IGNEUM_OP_MIX "load=16 add=15 shfl=6 mad=4 or=4 rotl=4 rotr=4 xor=4 mulhi=3 mul=2 sub=2"
|
||||
// Program class v3 (Counter ASIC 2.0, docs/plans/counter-asic-2-rollout.md): generator version 3; a worker that
|
||||
// runs another class refuses this pack, and a job line names the class it wants (class=v3 era=<hex>).
|
||||
#define IGNEUM_PROGRAM_CLASS "v3"
|
||||
#define IGNEUM_ERA_SEED_HEX "e593fc1d48475c88e456632f6aa0a752a5fdbca3c33022afe37680e44b7dfd11"
|
||||
// Class v3 construction (Counter ASIC 2.0, 5 October 2026, docs/plans/mixer-x4.md): version 2 loads; the dataset item
|
||||
// derivation applies the mixer IGNEUM_MIXER_MULT times per round (memhard.h), and the cache follows the growth rule.
|
||||
#define IGNEUM_LOAD_CLASS "w4-erad6367bfe"
|
||||
#define IGNEUM_LOAD_SLOTS 16
|
||||
#define IGNEUM_LOAD_MIX { 100, 0, 0 }
|
||||
#define IGNEUM_LOAD_WIDTH_COUNTS { 16, 0, 0 } // loads of 4, 16, 64 bytes per program
|
||||
#define IGNEUM_BYTES_PER_HASH 512
|
||||
#define IGNEUM_FOLD_ROT 11
|
||||
#define IGNEUM_FOLD_MUL 0x9e3779b1u
|
||||
// Era layout (5 October 2026, docs/plans/era-layout.md): NOT the lottery hash. Every dataset load reads
|
||||
// idx = ((rotl(src * STRIDE_MUL, STRIDE_ROT) & window mask) | window offset) & MASK; the window of a load site is the
|
||||
// dataset, a half or a quarter of it (IGNEUM_ERA_WINDOWS: site:shrink:offset); dataset word w holds word j(w) of item
|
||||
// t(w) with j's bits at the INTERLEAVE positions (memhard.h: mh_t, mh_j, mh_addr).
|
||||
#define IGNEUM_ERA_LABEL "d6367bfe"
|
||||
#define IGNEUM_ERA_SEED_WORDS { 0xd6367bfeu, 0x8bee0097u, 0x6c3c1b37u, 0x5dc9cefcu, 0x50b52d88u, 0x03833326u, 0x28ca58f9u, 0xa6e3b5c0u }
|
||||
#define IGNEUM_ERA_ALLOWED_WIDTHS { 1, 0, 0 } // words, ascending, 0 = unused; one entry pins the width
|
||||
#define IGNEUM_ERA_WIDTH_WORDS 1
|
||||
#define IGNEUM_ERA_STRIDE_MUL 0x27ea7effu
|
||||
#define IGNEUM_ERA_STRIDE_ROT 30
|
||||
#define IGNEUM_ERA_INTERLEAVE { 2, 3, 8, 13 }
|
||||
#define IGNEUM_ERA_WINDOWS "7:2:1 8:1:1 9:1:1 10:1:1 11:0:0 13:1:1 29:0:0 30:2:2 31:1:1 44:1:1 46:2:0 47:0:0 52:0:0 56:0:0 58:2:0 63:1:1"
|
||||
// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h)
|
||||
#define IGNEUM_DATASET_MODE 1
|
||||
|
||||
#define IGNEUM_SEEDW_INIT { 0x667d0fbdu, 0x7b8e5963u, 0x31c67e5eu, 0x4529ddc6u, 0xef19d6d8u, 0xaccf6211u, 0xda0aed32u, 0xabc6df31u }
|
||||
#define IGNEUM_KEY_INIT { 0xceed56d7u, 0x9ba270d2u, 0x82caab2du, 0x81ebce0eu, 0x12b6ecf1u, 0xd0f3fd7cu, 0xd872eefeu, 0xc158c7bdu }
|
||||
#define IGNEUM_CACHE_LOG2_WORDS 26
|
||||
#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6
|
||||
#define IGNEUM_CACHE_SEGMENTS 65536u
|
||||
#define IGNEUM_ITEM_ROUNDS 8
|
||||
#define IGNEUM_MIX_ROT_INIT { 17u, 12u, 20u, 23u, 7u, 3u, 27u, 16u }
|
||||
#define IGNEUM_MIX_MUL_INIT { 0xf351d601u, 0xa3bb398fu, 0xb5a09e35u, 0x7509c9c1u, 0x6bbf31e9u, 0xfc849a79u, 0xded91851u, 0x8d9113d1u, 0x0ff15225u, 0x3a5bdd41u, 0xab533435u, 0xe1c55ad5u, 0xe6d3bd0du, 0x9d9ffbbdu, 0xbb2a3cf3u, 0x50a7c08du }
|
||||
#define IGNEUM_MIX_RC_INIT { 0xc6892460u, 0x25b7228au, 0xcd515004u, 0x2846527au, 0xa6324241u, 0x36e3ec53u, 0x82961bacu, 0x0f97ba7du, 0xb6f921a9u, 0x3ada24e5u, 0xde20ab91u, 0x5378eeb2u, 0x7d161662u, 0x89353cc1u, 0xb1aa03a2u, 0x788acae6u }
|
||||
|
||||
#ifndef IGNEUM_NO_CUDA
|
||||
// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError().
|
||||
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments);
|
||||
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems);
|
||||
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
uint32_t nonces, uint32_t blockWarps);
|
||||
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
|
||||
#endif
|
||||
143
proto-cuda/packs-ca2-era/era-3/program.json
Normal file
143
proto-cuda/packs-ca2-era/era-3/program.json
Normal file
|
|
@ -0,0 +1,143 @@
|
|||
{
|
||||
"format": "igneum-program-pack-3",
|
||||
"generator": 3,
|
||||
"attempt": 0,
|
||||
"program_id": "0x73bcbfe8ccf988f1",
|
||||
"program_id_derivation": "FNV-1a 64 over 'igneum-program/' || generator_le32 || seed_words as little-endian bytes || attempt_le32",
|
||||
"dataset_mode": "memory-hard",
|
||||
"seed": "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000",
|
||||
"seed_bytes": "edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07",
|
||||
"seed_words": ["0x667d0fbd", "0x7b8e5963", "0x31c67e5e", "0x4529ddc6", "0xef19d6d8", "0xaccf6211", "0xda0aed32", "0xabc6df31"],
|
||||
"seed_derivation": "seed_words = FNV-1a 64 over seed_bytes (attempt 0) or seed_bytes || attempt_le32 (attempt k >= 1), basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32",
|
||||
"generator_rule": "version 2: exactly 16 load slots drawn first from instructions 1..63 (partial Fisher-Yates), the other 48 ops from the ten non-load weights (sum 75); a load's source is drawn from the registers other than dst written by an earlier instruction and not read by a load since; the candidate must pass the acceptance rule of spec 01 section 1.4.6 (static: no cyclically stale load source, every register has an injecting write; dynamic: 64 units on the seed-keyed closed-form dataset with no constant register bit, no lane-constant load site, under 164 saturated final values, every output bit within 136 of 1024, distinct addresses above 245760), else the next attempt of the seed is tried",
|
||||
"lanes": 32,
|
||||
"registers": 8,
|
||||
"iterations": 8,
|
||||
"instruction_count": 64,
|
||||
"loads_per_hash": 128,
|
||||
"program_class": "v3",
|
||||
"era_seed_bytes": "e593fc1d48475c88e456632f6aa0a752a5fdbca3c33022afe37680e44b7dfd11",
|
||||
"load_class": "w4-erad6367bfe",
|
||||
"load_slots": 16,
|
||||
"load_mix_percent_4_16_64": [100, 0, 0],
|
||||
"load_width_counts_4_16_64": [16, 0, 0],
|
||||
"bytes_per_hash": 512,
|
||||
"wide_load": "read-width experiment (5 October 2026, docs/plans/read-width.md), NOT the lottery hash: a load of W words (width field, 4 or 16) reads dataset[b .. b + W) with b = (src & mask) & ~(W - 1) and folds every word into dst: x = dst ^ w[0]; for j in 1..W: x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x; width 1 is the plain load; the width is drawn per instruction from the class mix with one extra below(100) draw after the nine of version 2, and the program id is FNV-1a 64 over 'igneum-program-rw/' || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots",
|
||||
"era": {
|
||||
"label": "d6367bfe",
|
||||
"seed_words": ["0xd6367bfe", "0x8bee0097", "0x6c3c1b37", "0x5dc9cefc", "0x50b52d88", "0x03833326", "0x28ca58f9", "0xa6e3b5c0"],
|
||||
"draw": "docs/plans/era-layout.md 1.1: SplitMix64 seeded with seed_words[0] | seed_words[1] << 32 of seed_words_from_bytes('igneum-era/' || n_le64 || E_n); width = allowed[below(|allowed|)], stride_mul = low32(next()) | 1, stride_rot = 1 + below(31), then four next() draws for a partial Fisher-Yates over positions log2(W)..15 of which 4 - log2(W) are used",
|
||||
"allowed_widths": [1],
|
||||
"width_words": 1,
|
||||
"stride_mul": "0x27ea7eff",
|
||||
"stride_rot": 30,
|
||||
"interleave": [2, 3, 8, 13],
|
||||
"address": "y = rotl(src * stride_mul, stride_rot); k = min(win, D - 26); idx = ((y & (mask >> k)) | ((off & (2^k - 1)) << (D - k))) & mask; a wide load aligns idx down to W words",
|
||||
"windows": "per instruction, after the width roll: win = below(3), off = low32(next()) & (2^win - 1); used on a load slot (the instruction's win and off fields)",
|
||||
"dataset_word": "dataset[w] = item(t(w))[j(w)]: j(w) gathers the bits of w at the interleave positions, t(w) is w with those bits removed",
|
||||
"program_id_suffix": "'era/' || allowed[3] || width_words || stride_mul_le32 || stride_rot_le32 || interleave[4]"
|
||||
},
|
||||
"op_mix": {"load": 16, "add": 15, "shfl": 6, "mad": 4, "or": 4, "rotl": 4, "rotr": 4, "xor": 4, "mulhi": 3, "mul": 2, "sub": 2},
|
||||
"register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]",
|
||||
"splitmix32": "x ^= x>>16; x *= 0x7feb352d; x ^= x>>15; x *= 0x846ca68b; x ^= x>>16",
|
||||
"iteration": "sel = r0 sampled once at the top of each iteration, then all instructions in order",
|
||||
"output": "lo = r0 ^ rotl(r1,7) ^ rotl(r2,14) ^ rotl(r3,21); hi = r4 ^ rotl(r5,9) ^ rotl(r6,18) ^ rotl(r7,27); out = (hi << 32) | lo",
|
||||
"op_semantics": {
|
||||
"add": "dst = dst + src + (bit `bit` of sel ? imm2 : imm)",
|
||||
"sub": "dst = dst - src",
|
||||
"mul": "dst = dst * src (low 32)",
|
||||
"mulhi": "dst = high 32 bits of dst * src",
|
||||
"xor": "dst = dst ^ src",
|
||||
"or": "dst = dst | src",
|
||||
"rotl": "dst = rotl(dst, rot), rot in 1..31",
|
||||
"rotr": "dst = rotr(dst, src & 31)",
|
||||
"mad": "dst = src * src2 + dst",
|
||||
"shfl": "dst = dst ^ (src of lane (lane ^ mask)), mask in {1,2,4,8,16}, within the 32-lane warp",
|
||||
"load": "dst = dst ^ dataset[src & dataset.mask]",
|
||||
"wload": "base = (src of lane 0 & dataset.mask) & ~31; dst = dst ^ dataset[base + lane] (warp-coalesced 128-byte load, lever b, only when --wide-frac > 0)"
|
||||
},
|
||||
"dataset": {
|
||||
"log2_words": 28,
|
||||
"bytes": 1073741824,
|
||||
"mask": "0x0fffffff",
|
||||
"day": "bytes:69676e65756d2d6461792ffa50000000000000",
|
||||
"day_bytes": "69676e65756d2d6461792ffa50000000000000",
|
||||
"day_words_from": "seed_words_from_bytes(day_bytes)",
|
||||
"d0": "0xceed56d7",
|
||||
"d1": "0x9ba270d2",
|
||||
"mode": "memory-hard",
|
||||
"spec": "proto-metal/MEMHARD.md",
|
||||
"key": ["0xceed56d7", "0x9ba270d2", "0x82caab2d", "0x81ebce0e", "0x12b6ecf1", "0xd0f3fd7c", "0xd872eefe", "0xc158c7bd"],
|
||||
"key_derivation": "the 8 words of seed_words_from_bytes(day_bytes); d0, d1 are key[0], key[1]",
|
||||
"cache": {"log2_words": 26, "bytes": 268435456, "line_words": 16, "segment_lines": 64, "segments": 65536, "block": "ChaCha12 core + feed-forward, rotations 16 12 8 7", "sigma": ["0x61707865", "0x3320646e", "0x79622d32", "0x6b206574"], "tag": ["0x49676e65", "0x756d4d48"], "chain": "in_j = prev_line ^ (sigma[0..3] || key[0..7] || seg || j || tag[0..1]); line_j = block(in_j); prev_0 = 0"},
|
||||
"mixer": {"draw": "SplitMix64 seeded with key[0] | key[1] << 32: rot[0..7] = 1 + next() % 31, mul[0..15] = low32(next()) | 1, rc[0..15] = low32(next())", "rot": [17, 12, 20, 23, 7, 3, 27, 16], "mul": ["0xf351d601", "0xa3bb398f", "0xb5a09e35", "0x7509c9c1", "0x6bbf31e9", "0xfc849a79", "0xded91851", "0x8d9113d1", "0x0ff15225", "0x3a5bdd41", "0xab533435", "0xe1c55ad5", "0xe6d3bd0d", "0x9d9ffbbd", "0xbb2a3cf3", "0x50a7c08d"], "rc": ["0xc6892460", "0x25b7228a", "0xcd515004", "0x2846527a", "0xa6324241", "0x36e3ec53", "0x82961bac", "0x0f97ba7d", "0xb6f921a9", "0x3ada24e5", "0xde20ab91", "0x5378eeb2", "0x7d161662", "0x89353cc1", "0xb1aa03a2", "0x788acae6"], "round": "for i in 0..15: s[i] = (s[i] ^ (rc[i] + (r+1) * 0x9E3779B9)) * mul[i]; then quarter rounds on columns (0,4,8,12) (1,5,9,13) (2,6,10,14) (3,7,11,15) with rot[0..3] and diagonals (0,5,10,15) (1,6,11,12) (2,7,8,13) (3,4,9,14) with rot[4..7]", "quarter_round": "a += b; d ^= a; d = rotl(d, r1); c += d; b ^= c; b = rotl(b, r2); a += b; d ^= a; d = rotl(d, r3); c += d; b ^= c; b = rotl(b, r4)"},
|
||||
"item": "s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..7: s = M_r(s); line = s[0] & 0x003fffff; s[i] ^= cache[line * 16 + i]; then s = M_8(s); item(t) = s",
|
||||
"word": "dataset[w] = item(w >> 4)[w & 15]"
|
||||
},
|
||||
"instructions": [
|
||||
{"i": 0, "op": "add", "dst": 4, "src": 5, "src2": 7, "imm": "0xea86e152", "imm2": "0x5810667a", "rot": 27, "bit": 13, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 1, "op": "xor", "dst": 7, "src": 0, "src2": 6, "imm": "0xbaab6229", "imm2": "0xed861989", "rot": 26, "bit": 22, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 2, "op": "shfl", "dst": 3, "src": 6, "src2": 2, "imm": "0x5b623116", "imm2": "0xff12e5b2", "rot": 12, "bit": 24, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 3, "op": "sub", "dst": 4, "src": 1, "src2": 1, "imm": "0xe99741c7", "imm2": "0xf5fa5009", "rot": 1, "bit": 21, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 4, "op": "mad", "dst": 2, "src": 0, "src2": 4, "imm": "0x673c2157", "imm2": "0xee02465f", "rot": 20, "bit": 22, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 5, "op": "rotl", "dst": 4, "src": 7, "src2": 7, "imm": "0x946f7818", "imm2": "0x45d3399e", "rot": 9, "bit": 2, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 6, "op": "rotr", "dst": 0, "src": 2, "src2": 4, "imm": "0x5f6a0ed2", "imm2": "0x7043a636", "rot": 19, "bit": 31, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 7, "op": "load", "dst": 6, "src": 7, "src2": 1, "imm": "0x5c61dcf7", "imm2": "0x7466aa40", "rot": 19, "bit": 9, "mask": 2, "width": 1, "win": 2, "off": 1},
|
||||
{"i": 8, "op": "load", "dst": 1, "src": 4, "src2": 5, "imm": "0x85668475", "imm2": "0xdb8cc483", "rot": 29, "bit": 7, "mask": 4, "width": 1, "win": 1, "off": 1},
|
||||
{"i": 9, "op": "load", "dst": 1, "src": 2, "src2": 4, "imm": "0x4454980f", "imm2": "0xebd31581", "rot": 10, "bit": 28, "mask": 16, "width": 1, "win": 1, "off": 1},
|
||||
{"i": 10, "op": "load", "dst": 7, "src": 0, "src2": 2, "imm": "0xe075297c", "imm2": "0x5779c44c", "rot": 10, "bit": 22, "mask": 2, "width": 1, "win": 1, "off": 1},
|
||||
{"i": 11, "op": "load", "dst": 7, "src": 1, "src2": 6, "imm": "0x65aa4311", "imm2": "0x4fe48ea9", "rot": 15, "bit": 9, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 12, "op": "mul", "dst": 5, "src": 4, "src2": 2, "imm": "0x1383d3ad", "imm2": "0xf3094b29", "rot": 8, "bit": 9, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 13, "op": "load", "dst": 7, "src": 6, "src2": 2, "imm": "0xed8a496f", "imm2": "0x3072c3c6", "rot": 28, "bit": 19, "mask": 8, "width": 1, "win": 1, "off": 1},
|
||||
{"i": 14, "op": "shfl", "dst": 6, "src": 5, "src2": 0, "imm": "0x8b965b57", "imm2": "0xcfeca6c1", "rot": 12, "bit": 27, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 15, "op": "shfl", "dst": 3, "src": 5, "src2": 3, "imm": "0x877c7586", "imm2": "0xa9cb2a03", "rot": 2, "bit": 29, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 16, "op": "or", "dst": 2, "src": 7, "src2": 3, "imm": "0xb740221a", "imm2": "0x89d38d6d", "rot": 6, "bit": 31, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 17, "op": "rotr", "dst": 0, "src": 6, "src2": 1, "imm": "0x26f3ad8a", "imm2": "0x27256f15", "rot": 18, "bit": 5, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 18, "op": "add", "dst": 7, "src": 5, "src2": 7, "imm": "0xf66e7017", "imm2": "0xb9e3577e", "rot": 9, "bit": 12, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 19, "op": "rotl", "dst": 7, "src": 0, "src2": 5, "imm": "0x849ae6ee", "imm2": "0x02b358f9", "rot": 24, "bit": 18, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 20, "op": "add", "dst": 6, "src": 7, "src2": 6, "imm": "0x52334d12", "imm2": "0x8c9f0ef8", "rot": 11, "bit": 23, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 21, "op": "or", "dst": 2, "src": 6, "src2": 5, "imm": "0xb1871e63", "imm2": "0xb2e40191", "rot": 5, "bit": 13, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 22, "op": "mad", "dst": 6, "src": 5, "src2": 4, "imm": "0x97df29e4", "imm2": "0xe60fea84", "rot": 11, "bit": 16, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 23, "op": "or", "dst": 1, "src": 0, "src2": 3, "imm": "0x8f30d21d", "imm2": "0x2df685a0", "rot": 31, "bit": 0, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 24, "op": "xor", "dst": 6, "src": 1, "src2": 2, "imm": "0xd2c4025f", "imm2": "0x5269eb4d", "rot": 31, "bit": 21, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 25, "op": "add", "dst": 2, "src": 6, "src2": 5, "imm": "0x659fc3d3", "imm2": "0x9cec0e12", "rot": 6, "bit": 17, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 26, "op": "rotr", "dst": 7, "src": 0, "src2": 1, "imm": "0xd89ef484", "imm2": "0x20be3846", "rot": 12, "bit": 9, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 27, "op": "mad", "dst": 4, "src": 5, "src2": 7, "imm": "0xb48420ae", "imm2": "0x3d1f2485", "rot": 14, "bit": 28, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 28, "op": "mad", "dst": 3, "src": 2, "src2": 4, "imm": "0xc5c46d76", "imm2": "0x700044b5", "rot": 22, "bit": 10, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 29, "op": "load", "dst": 1, "src": 4, "src2": 3, "imm": "0x0fbaf177", "imm2": "0xfff4f2ed", "rot": 20, "bit": 31, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 30, "op": "load", "dst": 2, "src": 3, "src2": 0, "imm": "0xe90eb2e5", "imm2": "0xb0f9eb79", "rot": 27, "bit": 14, "mask": 4, "width": 1, "win": 2, "off": 2},
|
||||
{"i": 31, "op": "load", "dst": 1, "src": 5, "src2": 2, "imm": "0x97ba3fc3", "imm2": "0x7894e657", "rot": 3, "bit": 30, "mask": 16, "width": 1, "win": 1, "off": 1},
|
||||
{"i": 32, "op": "add", "dst": 7, "src": 2, "src2": 7, "imm": "0x070888a8", "imm2": "0xe403240e", "rot": 2, "bit": 16, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 33, "op": "add", "dst": 2, "src": 0, "src2": 5, "imm": "0xf2e46d55", "imm2": "0x29701828", "rot": 31, "bit": 28, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 34, "op": "add", "dst": 2, "src": 3, "src2": 0, "imm": "0x58f75b87", "imm2": "0x343b7aee", "rot": 12, "bit": 14, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 35, "op": "mulhi", "dst": 2, "src": 5, "src2": 6, "imm": "0x5892a9e6", "imm2": "0xc9824c94", "rot": 19, "bit": 26, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 36, "op": "xor", "dst": 4, "src": 2, "src2": 3, "imm": "0x7b5b5474", "imm2": "0x45cfc5dd", "rot": 17, "bit": 18, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 37, "op": "mul", "dst": 6, "src": 5, "src2": 7, "imm": "0xccf564a5", "imm2": "0x873ad101", "rot": 7, "bit": 11, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 38, "op": "xor", "dst": 7, "src": 0, "src2": 6, "imm": "0xcac8f06d", "imm2": "0x6b97c683", "rot": 18, "bit": 28, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 39, "op": "add", "dst": 7, "src": 2, "src2": 4, "imm": "0xb8180e9d", "imm2": "0x32bbd117", "rot": 23, "bit": 19, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 40, "op": "rotr", "dst": 2, "src": 3, "src2": 0, "imm": "0x2d6070bc", "imm2": "0x68ff101e", "rot": 13, "bit": 18, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 41, "op": "sub", "dst": 7, "src": 0, "src2": 2, "imm": "0x0daf96ea", "imm2": "0x36f37be1", "rot": 5, "bit": 0, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 42, "op": "add", "dst": 4, "src": 3, "src2": 6, "imm": "0x6ced15b7", "imm2": "0x6df7aed4", "rot": 19, "bit": 4, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 43, "op": "shfl", "dst": 7, "src": 3, "src2": 5, "imm": "0x8ace05f3", "imm2": "0xd378ec12", "rot": 23, "bit": 10, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 44, "op": "load", "dst": 0, "src": 7, "src2": 4, "imm": "0xb0607786", "imm2": "0xc4acabbc", "rot": 13, "bit": 7, "mask": 16, "width": 1, "win": 1, "off": 1},
|
||||
{"i": 45, "op": "add", "dst": 1, "src": 6, "src2": 5, "imm": "0xe09f54e9", "imm2": "0x83e825bf", "rot": 23, "bit": 14, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 46, "op": "load", "dst": 3, "src": 1, "src2": 0, "imm": "0x63cc1e4e", "imm2": "0xa1be8118", "rot": 12, "bit": 6, "mask": 2, "width": 1, "win": 2, "off": 0},
|
||||
{"i": 47, "op": "load", "dst": 6, "src": 3, "src2": 7, "imm": "0x353f1d79", "imm2": "0x3b2e7456", "rot": 18, "bit": 18, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 48, "op": "mulhi", "dst": 4, "src": 2, "src2": 7, "imm": "0x00d8a3cd", "imm2": "0x231866d2", "rot": 21, "bit": 20, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 49, "op": "add", "dst": 5, "src": 0, "src2": 2, "imm": "0xa8bae6df", "imm2": "0xf572bdb9", "rot": 14, "bit": 7, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 50, "op": "shfl", "dst": 0, "src": 7, "src2": 7, "imm": "0x81ef22e1", "imm2": "0x74438fc5", "rot": 28, "bit": 18, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 51, "op": "add", "dst": 6, "src": 0, "src2": 6, "imm": "0x383b9260", "imm2": "0x11e17c61", "rot": 12, "bit": 19, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 52, "op": "load", "dst": 5, "src": 2, "src2": 2, "imm": "0xfb84f451", "imm2": "0x11cd863e", "rot": 21, "bit": 20, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 53, "op": "rotl", "dst": 6, "src": 5, "src2": 4, "imm": "0xb1a7db6b", "imm2": "0x76686b9b", "rot": 12, "bit": 4, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 54, "op": "rotl", "dst": 3, "src": 6, "src2": 3, "imm": "0x6f981f52", "imm2": "0xd99aeba2", "rot": 12, "bit": 27, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 55, "op": "add", "dst": 2, "src": 1, "src2": 2, "imm": "0xac6be8e3", "imm2": "0x18d67dbb", "rot": 26, "bit": 10, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 56, "op": "load", "dst": 1, "src": 4, "src2": 0, "imm": "0x7e7f6a00", "imm2": "0x6f0747da", "rot": 25, "bit": 28, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 57, "op": "add", "dst": 5, "src": 0, "src2": 4, "imm": "0xf03673fe", "imm2": "0xa75cd60d", "rot": 16, "bit": 12, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 58, "op": "load", "dst": 5, "src": 0, "src2": 2, "imm": "0x227f94a6", "imm2": "0x0e8344f9", "rot": 20, "bit": 10, "mask": 2, "width": 1, "win": 2, "off": 0},
|
||||
{"i": 59, "op": "add", "dst": 1, "src": 4, "src2": 3, "imm": "0xdecd4794", "imm2": "0x8dfb96bb", "rot": 21, "bit": 7, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 60, "op": "mulhi", "dst": 3, "src": 2, "src2": 2, "imm": "0x0dd268e0", "imm2": "0x53034ca9", "rot": 1, "bit": 8, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 61, "op": "or", "dst": 6, "src": 4, "src2": 7, "imm": "0x3a45a321", "imm2": "0x9bc59a5f", "rot": 25, "bit": 11, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 62, "op": "shfl", "dst": 5, "src": 4, "src2": 2, "imm": "0x8f229cc1", "imm2": "0xcaac64a2", "rot": 17, "bit": 13, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 63, "op": "load", "dst": 3, "src": 6, "src2": 6, "imm": "0xb2574178", "imm2": "0xcbcc798d", "rot": 28, "bit": 0, "mask": 16, "width": 1, "win": 1, "off": 1}
|
||||
]
|
||||
}
|
||||
109
proto-cuda/packs-ca2-era/era-3/program.metal
Normal file
109
proto-cuda/packs-ca2-era/era-3/program.metal
Normal file
|
|
@ -0,0 +1,109 @@
|
|||
#include <metal_stdlib>
|
||||
using namespace metal;
|
||||
|
||||
#define MASK 0x0fffffffu
|
||||
constant uint SEEDW[8] = { 0x667d0fbdu, 0x7b8e5963u, 0x31c67e5eu, 0x4529ddc6u, 0xef19d6d8u, 0xaccf6211u, 0xda0aed32u, 0xabc6df31u };
|
||||
|
||||
inline uint splitmix32(uint x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31
|
||||
inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
inline uint ds_elem(uint i, uint d0, uint d1) {
|
||||
uint x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
kernel void igneum_hash(device const uint* dataset [[buffer(0)]],
|
||||
device ulong* out [[buffer(1)]],
|
||||
constant uint& baseNonce [[buffer(2)]],
|
||||
uint gid [[thread_position_in_grid]]) {
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint x = nonce ^ SEEDW[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ SEEDW[1]; }
|
||||
{ uint x = nonce ^ SEEDW[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ SEEDW[2]; }
|
||||
{ uint x = nonce ^ SEEDW[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ SEEDW[3]; }
|
||||
{ uint x = nonce ^ SEEDW[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ SEEDW[4]; }
|
||||
{ uint x = nonce ^ SEEDW[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ SEEDW[5]; }
|
||||
{ uint x = nonce ^ SEEDW[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ SEEDW[6]; }
|
||||
{ uint x = nonce ^ SEEDW[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ SEEDW[7]; }
|
||||
{ uint x = nonce ^ SEEDW[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ SEEDW[0]; }
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = r4 + r5 + select(0xea86e152u, 0x5810667au, ((sel >> 13u) & 1u) != 0u); // 0
|
||||
r7 = r7 ^ r0; // 1
|
||||
r3 = r3 ^ simd_shuffle_xor(r6, (ushort)1); // 2
|
||||
r4 = r4 - r1; // 3
|
||||
r2 = r0 * r4 + r2; // 4
|
||||
r4 = rotl_imm(r4, 9u); // 5
|
||||
r0 = rotr_var(r0, r2); // 6
|
||||
r6 = r6 ^ dataset[((rotl_imm(r7 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x04000000u) & MASK]; // 7
|
||||
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 8
|
||||
r1 = r1 ^ dataset[((rotl_imm(r2 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 9
|
||||
r7 = r7 ^ dataset[((rotl_imm(r0 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 10
|
||||
r7 = r7 ^ dataset[((rotl_imm(r1 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 11
|
||||
r5 = r5 * r4; // 12
|
||||
r7 = r7 ^ dataset[((rotl_imm(r6 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 13
|
||||
r6 = r6 ^ simd_shuffle_xor(r5, (ushort)16); // 14
|
||||
r3 = r3 ^ simd_shuffle_xor(r5, (ushort)1); // 15
|
||||
r2 = r2 | r7; // 16
|
||||
r0 = rotr_var(r0, r6); // 17
|
||||
r7 = r7 + r5 + select(0xf66e7017u, 0xb9e3577eu, ((sel >> 12u) & 1u) != 0u); // 18
|
||||
r7 = rotl_imm(r7, 24u); // 19
|
||||
r6 = r6 + r7 + select(0x52334d12u, 0x8c9f0ef8u, ((sel >> 23u) & 1u) != 0u); // 20
|
||||
r2 = r2 | r6; // 21
|
||||
r6 = r5 * r4 + r6; // 22
|
||||
r1 = r1 | r0; // 23
|
||||
r6 = r6 ^ r1; // 24
|
||||
r2 = r2 + r6 + select(0x659fc3d3u, 0x9cec0e12u, ((sel >> 17u) & 1u) != 0u); // 25
|
||||
r7 = rotr_var(r7, r0); // 26
|
||||
r4 = r5 * r7 + r4; // 27
|
||||
r3 = r2 * r4 + r3; // 28
|
||||
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 29
|
||||
r2 = r2 ^ dataset[((rotl_imm(r3 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x08000000u) & MASK]; // 30
|
||||
r1 = r1 ^ dataset[((rotl_imm(r5 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 31
|
||||
r7 = r7 + r2 + select(0x070888a8u, 0xe403240eu, ((sel >> 16u) & 1u) != 0u); // 32
|
||||
r2 = r2 + r0 + select(0xf2e46d55u, 0x29701828u, ((sel >> 28u) & 1u) != 0u); // 33
|
||||
r2 = r2 + r3 + select(0x58f75b87u, 0x343b7aeeu, ((sel >> 14u) & 1u) != 0u); // 34
|
||||
r2 = mulhi(r2, r5); // 35
|
||||
r4 = r4 ^ r2; // 36
|
||||
r6 = r6 * r5; // 37
|
||||
r7 = r7 ^ r0; // 38
|
||||
r7 = r7 + r2 + select(0xb8180e9du, 0x32bbd117u, ((sel >> 19u) & 1u) != 0u); // 39
|
||||
r2 = rotr_var(r2, r3); // 40
|
||||
r7 = r7 - r0; // 41
|
||||
r4 = r4 + r3 + select(0x6ced15b7u, 0x6df7aed4u, ((sel >> 4u) & 1u) != 0u); // 42
|
||||
r7 = r7 ^ simd_shuffle_xor(r3, (ushort)4); // 43
|
||||
r0 = r0 ^ dataset[((rotl_imm(r7 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 44
|
||||
r1 = r1 + r6 + select(0xe09f54e9u, 0x83e825bfu, ((sel >> 14u) & 1u) != 0u); // 45
|
||||
r3 = r3 ^ dataset[((rotl_imm(r1 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 46
|
||||
r6 = r6 ^ dataset[((rotl_imm(r3 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 47
|
||||
r4 = mulhi(r4, r2); // 48
|
||||
r5 = r5 + r0 + select(0xa8bae6dfu, 0xf572bdb9u, ((sel >> 7u) & 1u) != 0u); // 49
|
||||
r0 = r0 ^ simd_shuffle_xor(r7, (ushort)4); // 50
|
||||
r6 = r6 + r0 + select(0x383b9260u, 0x11e17c61u, ((sel >> 19u) & 1u) != 0u); // 51
|
||||
r5 = r5 ^ dataset[((rotl_imm(r2 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 52
|
||||
r6 = rotl_imm(r6, 12u); // 53
|
||||
r3 = rotl_imm(r3, 12u); // 54
|
||||
r2 = r2 + r1 + select(0xac6be8e3u, 0x18d67dbbu, ((sel >> 10u) & 1u) != 0u); // 55
|
||||
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 56
|
||||
r5 = r5 + r0 + select(0xf03673feu, 0xa75cd60du, ((sel >> 12u) & 1u) != 0u); // 57
|
||||
r5 = r5 ^ dataset[((rotl_imm(r0 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 58
|
||||
r1 = r1 + r4 + select(0xdecd4794u, 0x8dfb96bbu, ((sel >> 7u) & 1u) != 0u); // 59
|
||||
r3 = mulhi(r3, r2); // 60
|
||||
r6 = r6 | r4; // 61
|
||||
r5 = r5 ^ simd_shuffle_xor(r4, (ushort)8); // 62
|
||||
r3 = r3 ^ dataset[((rotl_imm(r6 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 63
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
111
proto-cuda/packs-ca2-era/era-3/program_bound.metal
Normal file
111
proto-cuda/packs-ca2-era/era-3/program_bound.metal
Normal file
|
|
@ -0,0 +1,111 @@
|
|||
#include <metal_stdlib>
|
||||
using namespace metal;
|
||||
|
||||
#define MASK 0x0fffffffu
|
||||
constant uint SEEDW[8] = { 0x667d0fbdu, 0x7b8e5963u, 0x31c67e5eu, 0x4529ddc6u, 0xef19d6d8u, 0xaccf6211u, 0xda0aed32u, 0xabc6df31u };
|
||||
|
||||
inline uint splitmix32(uint x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31
|
||||
inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
inline uint ds_elem(uint i, uint d0, uint d1) {
|
||||
uint x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// Header-bound variant: the init words come from buffer 3 (bind.rs), not from SEEDW.
|
||||
kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]],
|
||||
device ulong* out [[buffer(1)]],
|
||||
constant uint& baseNonce [[buffer(2)]],
|
||||
constant uint* initw [[buffer(3)]],
|
||||
uint gid [[thread_position_in_grid]]) {
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint x = nonce ^ initw[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ initw[1]; }
|
||||
{ uint x = nonce ^ initw[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ initw[2]; }
|
||||
{ uint x = nonce ^ initw[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ initw[3]; }
|
||||
{ uint x = nonce ^ initw[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ initw[4]; }
|
||||
{ uint x = nonce ^ initw[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ initw[5]; }
|
||||
{ uint x = nonce ^ initw[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ initw[6]; }
|
||||
{ uint x = nonce ^ initw[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ initw[7]; }
|
||||
{ uint x = nonce ^ initw[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ initw[0]; }
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = r4 + r5 + select(0xea86e152u, 0x5810667au, ((sel >> 13u) & 1u) != 0u); // 0
|
||||
r7 = r7 ^ r0; // 1
|
||||
r3 = r3 ^ simd_shuffle_xor(r6, (ushort)1); // 2
|
||||
r4 = r4 - r1; // 3
|
||||
r2 = r0 * r4 + r2; // 4
|
||||
r4 = rotl_imm(r4, 9u); // 5
|
||||
r0 = rotr_var(r0, r2); // 6
|
||||
r6 = r6 ^ dataset[((rotl_imm(r7 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x04000000u) & MASK]; // 7
|
||||
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 8
|
||||
r1 = r1 ^ dataset[((rotl_imm(r2 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 9
|
||||
r7 = r7 ^ dataset[((rotl_imm(r0 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 10
|
||||
r7 = r7 ^ dataset[((rotl_imm(r1 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 11
|
||||
r5 = r5 * r4; // 12
|
||||
r7 = r7 ^ dataset[((rotl_imm(r6 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 13
|
||||
r6 = r6 ^ simd_shuffle_xor(r5, (ushort)16); // 14
|
||||
r3 = r3 ^ simd_shuffle_xor(r5, (ushort)1); // 15
|
||||
r2 = r2 | r7; // 16
|
||||
r0 = rotr_var(r0, r6); // 17
|
||||
r7 = r7 + r5 + select(0xf66e7017u, 0xb9e3577eu, ((sel >> 12u) & 1u) != 0u); // 18
|
||||
r7 = rotl_imm(r7, 24u); // 19
|
||||
r6 = r6 + r7 + select(0x52334d12u, 0x8c9f0ef8u, ((sel >> 23u) & 1u) != 0u); // 20
|
||||
r2 = r2 | r6; // 21
|
||||
r6 = r5 * r4 + r6; // 22
|
||||
r1 = r1 | r0; // 23
|
||||
r6 = r6 ^ r1; // 24
|
||||
r2 = r2 + r6 + select(0x659fc3d3u, 0x9cec0e12u, ((sel >> 17u) & 1u) != 0u); // 25
|
||||
r7 = rotr_var(r7, r0); // 26
|
||||
r4 = r5 * r7 + r4; // 27
|
||||
r3 = r2 * r4 + r3; // 28
|
||||
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 29
|
||||
r2 = r2 ^ dataset[((rotl_imm(r3 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x08000000u) & MASK]; // 30
|
||||
r1 = r1 ^ dataset[((rotl_imm(r5 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 31
|
||||
r7 = r7 + r2 + select(0x070888a8u, 0xe403240eu, ((sel >> 16u) & 1u) != 0u); // 32
|
||||
r2 = r2 + r0 + select(0xf2e46d55u, 0x29701828u, ((sel >> 28u) & 1u) != 0u); // 33
|
||||
r2 = r2 + r3 + select(0x58f75b87u, 0x343b7aeeu, ((sel >> 14u) & 1u) != 0u); // 34
|
||||
r2 = mulhi(r2, r5); // 35
|
||||
r4 = r4 ^ r2; // 36
|
||||
r6 = r6 * r5; // 37
|
||||
r7 = r7 ^ r0; // 38
|
||||
r7 = r7 + r2 + select(0xb8180e9du, 0x32bbd117u, ((sel >> 19u) & 1u) != 0u); // 39
|
||||
r2 = rotr_var(r2, r3); // 40
|
||||
r7 = r7 - r0; // 41
|
||||
r4 = r4 + r3 + select(0x6ced15b7u, 0x6df7aed4u, ((sel >> 4u) & 1u) != 0u); // 42
|
||||
r7 = r7 ^ simd_shuffle_xor(r3, (ushort)4); // 43
|
||||
r0 = r0 ^ dataset[((rotl_imm(r7 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 44
|
||||
r1 = r1 + r6 + select(0xe09f54e9u, 0x83e825bfu, ((sel >> 14u) & 1u) != 0u); // 45
|
||||
r3 = r3 ^ dataset[((rotl_imm(r1 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 46
|
||||
r6 = r6 ^ dataset[((rotl_imm(r3 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 47
|
||||
r4 = mulhi(r4, r2); // 48
|
||||
r5 = r5 + r0 + select(0xa8bae6dfu, 0xf572bdb9u, ((sel >> 7u) & 1u) != 0u); // 49
|
||||
r0 = r0 ^ simd_shuffle_xor(r7, (ushort)4); // 50
|
||||
r6 = r6 + r0 + select(0x383b9260u, 0x11e17c61u, ((sel >> 19u) & 1u) != 0u); // 51
|
||||
r5 = r5 ^ dataset[((rotl_imm(r2 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 52
|
||||
r6 = rotl_imm(r6, 12u); // 53
|
||||
r3 = rotl_imm(r3, 12u); // 54
|
||||
r2 = r2 + r1 + select(0xac6be8e3u, 0x18d67dbbu, ((sel >> 10u) & 1u) != 0u); // 55
|
||||
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x27ea7effu, 30u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 56
|
||||
r5 = r5 + r0 + select(0xf03673feu, 0xa75cd60du, ((sel >> 12u) & 1u) != 0u); // 57
|
||||
r5 = r5 ^ dataset[((rotl_imm(r0 * 0x27ea7effu, 30u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 58
|
||||
r1 = r1 + r4 + select(0xdecd4794u, 0x8dfb96bbu, ((sel >> 7u) & 1u) != 0u); // 59
|
||||
r3 = mulhi(r3, r2); // 60
|
||||
r6 = r6 | r4; // 61
|
||||
r5 = r5 ^ simd_shuffle_xor(r4, (ushort)8); // 62
|
||||
r3 = r3 ^ dataset[((rotl_imm(r6 * 0x27ea7effu, 30u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 63
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
5
proto-cuda/packs-ca2-era/era-3/seeds.txt
Normal file
5
proto-cuda/packs-ca2-era/era-3/seeds.txt
Normal file
|
|
@ -0,0 +1,5 @@
|
|||
epoch_seed_hex edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07
|
||||
day_seed_hex 69676e65756d2d6461792ffa50000000000000
|
||||
epoch_index 0
|
||||
day_index 20730
|
||||
daa_score 0
|
||||
57
proto-cuda/packs-ca2-era/era-3/vectors.h
Normal file
57
proto-cuda/packs-ca2-era/era-3/vectors.h
Normal file
|
|
@ -0,0 +1,57 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v3, memory-hard dataset
|
||||
#pragma once
|
||||
#ifdef __cplusplus
|
||||
#include <cstdint>
|
||||
#else
|
||||
#include <stdint.h>
|
||||
#endif
|
||||
|
||||
#define IGNEUM_VEC_WARPS 3
|
||||
static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u };
|
||||
static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = {
|
||||
{ // base nonce 0
|
||||
0x482d9ca50850d913ull, 0x4ad212a8109dbf80ull, 0xfc22f1d2859b72dbull, 0x1b122111c880daddull, 0x590c0ed07d26f3beull, 0x14f9ccb1fe4a0aa0ull, 0xb9aa46ddbf3358f4ull, 0x7a2e7a0758a139dcull,
|
||||
0x4ca0c6f57fc6a6f8ull, 0xa9303e49a53c55cfull, 0xd0eef090a205c2daull, 0xf916463a1fe71f3dull, 0xece100319c1ba260ull, 0x0d765a753bab6621ull, 0x7a8b38b0cd86e711ull, 0x6330fa995f69eac6ull,
|
||||
0x31630dc8222e4cf3ull, 0x1ab386a5c4bfd048ull, 0xd1849ab5d7ccc035ull, 0x7e42d0b071e213b7ull, 0xc61d482f91f73a00ull, 0xb759a48c18732abcull, 0x7cb720597cea45f4ull, 0x9aa782a7d7153267ull,
|
||||
0x3928c8ca30daa819ull, 0xf6624ec425bf44b6ull, 0xb92b06f50f5fca62ull, 0xb7ac77af13b44e09ull, 0x9bd4a23ed9a8d186ull, 0xed157fa7db51abdfull, 0x169dae4f9708b9b0ull, 0x3928f1c5bb85b8ffull
|
||||
},
|
||||
{ // base nonce 4096
|
||||
0x68d28d84a9cacd37ull, 0x3de7667c504d33b1ull, 0x9cd9cbd429ecccbfull, 0xce28296aba64ec41ull, 0x7f25bfb162361f07ull, 0x305c8fcecda70a88ull, 0x4d26391bf3d1c1f0ull, 0x230a57b46fde607full,
|
||||
0x35482e4e9ae6e874ull, 0x87f3bdd2250cb21eull, 0x71a6a482132c6938ull, 0x3c853fb23317a129ull, 0x5932d928549c095aull, 0xe90460bc189cdb34ull, 0x6dbe852db160461aull, 0x2abd73567e773ab6ull,
|
||||
0x97be9fa85c3805ffull, 0xf7d138416c3ac1ccull, 0x888b84f204d9305eull, 0x92800629cdb2d056ull, 0x2cb98fc7e6f8b22bull, 0x085ba0e520fa1df6ull, 0x9b0ee619aa074133ull, 0xd92649b972786f58ull,
|
||||
0x94e56fdfade518caull, 0xcae28f9eec9d5be5ull, 0x2e715314eaa160b6ull, 0xd3faf570104ca532ull, 0xeeecb2834a0e878bull, 0x9e978602b88d0a38ull, 0x3536c9dd7062e398ull, 0x5974adef999d72a2ull
|
||||
},
|
||||
{ // base nonce 1000000
|
||||
0x86c2830118bf5d18ull, 0x0ca010e1f65efea9ull, 0x0e065db2ff5d7177ull, 0xb6c5de9453893f86ull, 0x4b2bca8ca7ff4dacull, 0x875e9e4b09262995ull, 0xfe1c3a11b0dc75abull, 0x9f83ac0ba031e6dfull,
|
||||
0x5c4e3154cdf0d3beull, 0x812fd5f5be4159e6ull, 0xcbf87d1772e7a740ull, 0x251c1bb3b31e0de1ull, 0xba117c2b253c2d04ull, 0xf42a27fffa765e8aull, 0xb654d4c021ced921ull, 0x63fe5bc570bd9a6dull,
|
||||
0xe3729cacfa2738b9ull, 0x0686bf7f029dd650ull, 0xa4da20856fbb382eull, 0xcacd5c3aa30da86cull, 0x7b47735366038a7aull, 0xe0b5603c8652702eull, 0x62b834fee790c2c5ull, 0x4c979e00c086c371ull,
|
||||
0x9a94911a7eaa32d6ull, 0x377a5edc37bfb69full, 0xecebe1aab4f07982ull, 0xdd2bba556d4d8996ull, 0xb6e0be464e82691cull, 0xe4cc7f9840d754d9ull, 0x8e31fb423c108e46ull, 0x8164e7242a952c96ull
|
||||
}
|
||||
};
|
||||
|
||||
// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455).
|
||||
static const uint32_t IGNEUM_DS_HEAD[16] = {
|
||||
0x3dd50b1fu, 0x93123701u, 0xcdd88b45u, 0x2c37c6b0u, 0x48edec90u, 0x3a7d2407u, 0xb03a0ac0u, 0x10e2c6a5u,
|
||||
0x5119540eu, 0x3358ef2du, 0x25e2929cu, 0x5fa0cdccu, 0x657ac748u, 0x09e2105cu, 0x4317cabdu, 0x6178ee52u
|
||||
};
|
||||
static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u;
|
||||
static const uint32_t IGNEUM_DS_LAST = 0xf7b7180eu;
|
||||
// 64 sampled dataset words (index, value) computed on the Mac.
|
||||
#define IGNEUM_DS_SAMPLES 64
|
||||
static const uint32_t IGNEUM_DS_SAMPLE_INDEX[IGNEUM_DS_SAMPLES] = {
|
||||
59471966u, 217795994u, 208353206u, 42483309u, 172547758u, 148076330u, 183853158u, 214389424u, 267488061u, 169781097u, 184093494u, 153880993u, 84977930u, 46426879u, 3093825u, 225364072u, 44593546u, 260713159u, 168250303u, 52384140u, 223401610u, 45554030u, 95410555u, 175039924u, 79171087u, 267580473u, 24168642u, 37981670u, 171551130u, 195559979u, 204611762u, 140997658u, 138925853u, 86637313u, 20736778u, 219665210u, 160430336u, 264654675u, 8013395u, 228945585u, 213884386u, 104419827u, 44185464u, 142737231u, 99284897u, 132475900u, 61861762u, 132056166u, 262388043u, 91878046u, 117353561u, 124768597u, 71352993u, 190698941u, 46055428u, 55281366u, 165145231u, 106810753u, 171985651u, 232085256u, 159510492u, 40072060u, 209107596u, 39023794u
|
||||
};
|
||||
static const uint32_t IGNEUM_DS_SAMPLE_VALUE[IGNEUM_DS_SAMPLES] = {
|
||||
0xf6e00c89u, 0xe206f72bu, 0x09260b8cu, 0x662c6558u, 0x4028c489u, 0xbccb5ab0u, 0x88bf905cu, 0x88f76874u, 0xea529b56u, 0x88e89b31u, 0x22deecc6u, 0xd6022e44u, 0x7a7923c9u, 0x8c761ec7u, 0x7079d93du, 0x79341b4cu, 0x560640c9u, 0xc659d408u, 0xd5883003u, 0x95e70092u, 0xbb0552dbu, 0xbbec95bfu, 0x1efacd50u, 0x8d4c6036u, 0xcbcb1898u, 0x05d9a8aeu, 0x0e5af59bu, 0x12094029u, 0xd6132e60u, 0xffc32f3eu, 0x3bae4991u, 0xbf6d57a0u, 0x7f03491eu, 0x26c9e599u, 0xa0c9ecc9u, 0xae96f173u, 0x1f437e50u, 0x1f641c65u, 0x30d7028au, 0x65092ed9u, 0xe390c42bu, 0x8f89bb6bu, 0xb4205b97u, 0xd79db283u, 0x65869dcfu, 0x4ac02d55u, 0x5f1b8dc3u, 0xd192a75au, 0xcca97073u, 0x2f043a85u, 0xdc0c3fddu, 0x19b12a30u, 0xf5da472eu, 0xbf7dd652u, 0xec1dc918u, 0x2f00a5dcu, 0xbe7f358eu, 0x68a25926u, 0x817ae6c9u, 0xa731ea98u, 0x6af869dau, 0xb84e7e41u, 0xd3519f00u, 0x3e6c426au
|
||||
};
|
||||
// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^26 words.
|
||||
static const uint32_t IGNEUM_CACHE_HEAD[16] = {
|
||||
0xebd9055cu, 0x7eaf6b21u, 0x9b610855u, 0x134a8562u, 0xb10059bau, 0x7f459e58u, 0x8f3c7873u, 0x3523e609u,
|
||||
0x75b8f4cau, 0x750d31b4u, 0x0dae0782u, 0x26f10015u, 0x7ba87a41u, 0x0efd8543u, 0x691d6368u, 0x8929d967u
|
||||
};
|
||||
static const uint32_t IGNEUM_CACHE_LAST[16] = {
|
||||
0x51474edfu, 0xc3cfba93u, 0xf21454cfu, 0x9b79baacu, 0x4d4fce23u, 0xd701dfd5u, 0x37357ab3u, 0x1be693fau,
|
||||
0xa701cb3bu, 0x7467c620u, 0x428184e8u, 0xf4010df0u, 0xe33aa88fu, 0xc78d6d62u, 0xae9ba9c0u, 0xf4bba97eu
|
||||
};
|
||||
static const uint64_t IGNEUM_CACHE_FNV64 = 0x448274a57f508cbcull;
|
||||
36
proto-cuda/packs-ca2-era/era-3/vectors.json
Normal file
36
proto-cuda/packs-ca2-era/era-3/vectors.json
Normal file
|
|
@ -0,0 +1,36 @@
|
|||
{
|
||||
"seed": "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000",
|
||||
"day": "bytes:69676e65756d2d6461792ffa50000000000000",
|
||||
"dataset_mode": "memory-hard",
|
||||
"dataset_log2_words": 28,
|
||||
"mask": "0x0fffffff",
|
||||
"lanes": 32,
|
||||
"source": "igneum-pow (Rust) CPU interpreter, generator v3, memory-hard dataset",
|
||||
"warps": [
|
||||
{"base_nonce": 0, "expected": [
|
||||
"0x482d9ca50850d913", "0x4ad212a8109dbf80", "0xfc22f1d2859b72db", "0x1b122111c880dadd", "0x590c0ed07d26f3be", "0x14f9ccb1fe4a0aa0", "0xb9aa46ddbf3358f4", "0x7a2e7a0758a139dc",
|
||||
"0x4ca0c6f57fc6a6f8", "0xa9303e49a53c55cf", "0xd0eef090a205c2da", "0xf916463a1fe71f3d", "0xece100319c1ba260", "0x0d765a753bab6621", "0x7a8b38b0cd86e711", "0x6330fa995f69eac6",
|
||||
"0x31630dc8222e4cf3", "0x1ab386a5c4bfd048", "0xd1849ab5d7ccc035", "0x7e42d0b071e213b7", "0xc61d482f91f73a00", "0xb759a48c18732abc", "0x7cb720597cea45f4", "0x9aa782a7d7153267",
|
||||
"0x3928c8ca30daa819", "0xf6624ec425bf44b6", "0xb92b06f50f5fca62", "0xb7ac77af13b44e09", "0x9bd4a23ed9a8d186", "0xed157fa7db51abdf", "0x169dae4f9708b9b0", "0x3928f1c5bb85b8ff"
|
||||
]},
|
||||
{"base_nonce": 4096, "expected": [
|
||||
"0x68d28d84a9cacd37", "0x3de7667c504d33b1", "0x9cd9cbd429ecccbf", "0xce28296aba64ec41", "0x7f25bfb162361f07", "0x305c8fcecda70a88", "0x4d26391bf3d1c1f0", "0x230a57b46fde607f",
|
||||
"0x35482e4e9ae6e874", "0x87f3bdd2250cb21e", "0x71a6a482132c6938", "0x3c853fb23317a129", "0x5932d928549c095a", "0xe90460bc189cdb34", "0x6dbe852db160461a", "0x2abd73567e773ab6",
|
||||
"0x97be9fa85c3805ff", "0xf7d138416c3ac1cc", "0x888b84f204d9305e", "0x92800629cdb2d056", "0x2cb98fc7e6f8b22b", "0x085ba0e520fa1df6", "0x9b0ee619aa074133", "0xd92649b972786f58",
|
||||
"0x94e56fdfade518ca", "0xcae28f9eec9d5be5", "0x2e715314eaa160b6", "0xd3faf570104ca532", "0xeeecb2834a0e878b", "0x9e978602b88d0a38", "0x3536c9dd7062e398", "0x5974adef999d72a2"
|
||||
]},
|
||||
{"base_nonce": 1000000, "expected": [
|
||||
"0x86c2830118bf5d18", "0x0ca010e1f65efea9", "0x0e065db2ff5d7177", "0xb6c5de9453893f86", "0x4b2bca8ca7ff4dac", "0x875e9e4b09262995", "0xfe1c3a11b0dc75ab", "0x9f83ac0ba031e6df",
|
||||
"0x5c4e3154cdf0d3be", "0x812fd5f5be4159e6", "0xcbf87d1772e7a740", "0x251c1bb3b31e0de1", "0xba117c2b253c2d04", "0xf42a27fffa765e8a", "0xb654d4c021ced921", "0x63fe5bc570bd9a6d",
|
||||
"0xe3729cacfa2738b9", "0x0686bf7f029dd650", "0xa4da20856fbb382e", "0xcacd5c3aa30da86c", "0x7b47735366038a7a", "0xe0b5603c8652702e", "0x62b834fee790c2c5", "0x4c979e00c086c371",
|
||||
"0x9a94911a7eaa32d6", "0x377a5edc37bfb69f", "0xecebe1aab4f07982", "0xdd2bba556d4d8996", "0xb6e0be464e82691c", "0xe4cc7f9840d754d9", "0x8e31fb423c108e46", "0x8164e7242a952c96"
|
||||
]}
|
||||
],
|
||||
"dataset_head": ["0x3dd50b1f", "0x93123701", "0xcdd88b45", "0x2c37c6b0", "0x48edec90", "0x3a7d2407", "0xb03a0ac0", "0x10e2c6a5", "0x5119540e", "0x3358ef2d", "0x25e2929c", "0x5fa0cdcc", "0x657ac748", "0x09e2105c", "0x4317cabd", "0x6178ee52"],
|
||||
"dataset_last_index": 268435455,
|
||||
"dataset_last": "0xf7b7180e",
|
||||
"dataset_samples": [{"index": 59471966, "value": "0xf6e00c89"}, {"index": 217795994, "value": "0xe206f72b"}, {"index": 208353206, "value": "0x09260b8c"}, {"index": 42483309, "value": "0x662c6558"}, {"index": 172547758, "value": "0x4028c489"}, {"index": 148076330, "value": "0xbccb5ab0"}, {"index": 183853158, "value": "0x88bf905c"}, {"index": 214389424, "value": "0x88f76874"}, {"index": 267488061, "value": "0xea529b56"}, {"index": 169781097, "value": "0x88e89b31"}, {"index": 184093494, "value": "0x22deecc6"}, {"index": 153880993, "value": "0xd6022e44"}, {"index": 84977930, "value": "0x7a7923c9"}, {"index": 46426879, "value": "0x8c761ec7"}, {"index": 3093825, "value": "0x7079d93d"}, {"index": 225364072, "value": "0x79341b4c"}, {"index": 44593546, "value": "0x560640c9"}, {"index": 260713159, "value": "0xc659d408"}, {"index": 168250303, "value": "0xd5883003"}, {"index": 52384140, "value": "0x95e70092"}, {"index": 223401610, "value": "0xbb0552db"}, {"index": 45554030, "value": "0xbbec95bf"}, {"index": 95410555, "value": "0x1efacd50"}, {"index": 175039924, "value": "0x8d4c6036"}, {"index": 79171087, "value": "0xcbcb1898"}, {"index": 267580473, "value": "0x05d9a8ae"}, {"index": 24168642, "value": "0x0e5af59b"}, {"index": 37981670, "value": "0x12094029"}, {"index": 171551130, "value": "0xd6132e60"}, {"index": 195559979, "value": "0xffc32f3e"}, {"index": 204611762, "value": "0x3bae4991"}, {"index": 140997658, "value": "0xbf6d57a0"}, {"index": 138925853, "value": "0x7f03491e"}, {"index": 86637313, "value": "0x26c9e599"}, {"index": 20736778, "value": "0xa0c9ecc9"}, {"index": 219665210, "value": "0xae96f173"}, {"index": 160430336, "value": "0x1f437e50"}, {"index": 264654675, "value": "0x1f641c65"}, {"index": 8013395, "value": "0x30d7028a"}, {"index": 228945585, "value": "0x65092ed9"}, {"index": 213884386, "value": "0xe390c42b"}, {"index": 104419827, "value": "0x8f89bb6b"}, {"index": 44185464, "value": "0xb4205b97"}, {"index": 142737231, "value": "0xd79db283"}, {"index": 99284897, "value": "0x65869dcf"}, {"index": 132475900, "value": "0x4ac02d55"}, {"index": 61861762, "value": "0x5f1b8dc3"}, {"index": 132056166, "value": "0xd192a75a"}, {"index": 262388043, "value": "0xcca97073"}, {"index": 91878046, "value": "0x2f043a85"}, {"index": 117353561, "value": "0xdc0c3fdd"}, {"index": 124768597, "value": "0x19b12a30"}, {"index": 71352993, "value": "0xf5da472e"}, {"index": 190698941, "value": "0xbf7dd652"}, {"index": 46055428, "value": "0xec1dc918"}, {"index": 55281366, "value": "0x2f00a5dc"}, {"index": 165145231, "value": "0xbe7f358e"}, {"index": 106810753, "value": "0x68a25926"}, {"index": 171985651, "value": "0x817ae6c9"}, {"index": 232085256, "value": "0xa731ea98"}, {"index": 159510492, "value": "0x6af869da"}, {"index": 40072060, "value": "0xb84e7e41"}, {"index": 209107596, "value": "0xd3519f00"}, {"index": 39023794, "value": "0x3e6c426a"}],
|
||||
"cache_head": ["0xebd9055c", "0x7eaf6b21", "0x9b610855", "0x134a8562", "0xb10059ba", "0x7f459e58", "0x8f3c7873", "0x3523e609", "0x75b8f4ca", "0x750d31b4", "0x0dae0782", "0x26f10015", "0x7ba87a41", "0x0efd8543", "0x691d6368", "0x8929d967"],
|
||||
"cache_last_line": ["0x51474edf", "0xc3cfba93", "0xf21454cf", "0x9b79baac", "0x4d4fce23", "0xd701dfd5", "0x37357ab3", "0x1be693fa", "0xa701cb3b", "0x7467c620", "0x428184e8", "0xf4010df0", "0xe33aa88f", "0xc78d6d62", "0xae9ba9c0", "0xf4bba97e"],
|
||||
"cache_fnv1a64": "0x448274a57f508cbc"
|
||||
}
|
||||
281
proto-cuda/packs-ca2-era/era-4/kernel.cl
Normal file
281
proto-cuda/packs-ca2-era/era-4/kernel.cl
Normal file
|
|
@ -0,0 +1,281 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal).
|
||||
// Built from source at runtime by proto-opencl/host.c, which passes these defines:
|
||||
// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit)
|
||||
// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default)
|
||||
// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32
|
||||
// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition
|
||||
// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units;
|
||||
// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32.
|
||||
#ifndef IGNEUM_GROUP
|
||||
#define IGNEUM_GROUP 32
|
||||
#endif
|
||||
#ifndef IGNEUM_EXCHANGE
|
||||
#define IGNEUM_EXCHANGE 0
|
||||
#endif
|
||||
#ifdef __OPENCL_VERSION__
|
||||
#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1)))
|
||||
#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n]
|
||||
#if IGNEUM_EXCHANGE == 1
|
||||
#ifdef cl_khr_subgroups
|
||||
#pragma OPENCL EXTENSION cl_khr_subgroups : enable
|
||||
#endif
|
||||
#ifdef cl_khr_subgroup_shuffle
|
||||
#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable
|
||||
#endif
|
||||
#elif IGNEUM_EXCHANGE == 2
|
||||
#pragma OPENCL EXTENSION cl_intel_subgroups : enable
|
||||
#endif
|
||||
#else
|
||||
// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros.
|
||||
#include "emu_opencl.h"
|
||||
#endif
|
||||
|
||||
#if IGNEUM_EXCHANGE == 1
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m))
|
||||
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
|
||||
#elif IGNEUM_EXCHANGE == 2
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m))
|
||||
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
|
||||
#else
|
||||
// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per
|
||||
// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane
|
||||
// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's
|
||||
// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier.
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; }
|
||||
#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; }
|
||||
#endif
|
||||
|
||||
static inline uint splitmix32(uint x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32.
|
||||
static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); }
|
||||
// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x.
|
||||
static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); }
|
||||
static inline uint ds_elem(uint i, uint d0, uint d1) {
|
||||
uint x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
|
||||
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
|
||||
#define MH_CACHE_LINE_MASK 0x003fffffu
|
||||
#define MH_SEGMENT_LINES 64u
|
||||
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
|
||||
static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
|
||||
|
||||
// y = ChaCha12 core(x) + x
|
||||
static inline void mh_chacha_block(const uint* x, uint* y) {
|
||||
for (uint i = 0u; i < 16u; ++i) y[i] = x[i];
|
||||
for (uint r = 0u; r < 6u; ++r) {
|
||||
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
|
||||
}
|
||||
for (uint i = 0u; i < 16u; ++i) y[i] += x[i];
|
||||
}
|
||||
|
||||
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
|
||||
static inline void mh_cache_segment(__global uint* cache, uint seg) {
|
||||
uint prev[16]; uint x[16]; uint y[16];
|
||||
for (uint i = 0u; i < 16u; ++i) prev[i] = 0u;
|
||||
for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) {
|
||||
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
|
||||
x[4] = 0xceed56d7u ^ prev[4];
|
||||
x[5] = 0x9ba270d2u ^ prev[5];
|
||||
x[6] = 0x82caab2du ^ prev[6];
|
||||
x[7] = 0x81ebce0eu ^ prev[7];
|
||||
x[8] = 0x12b6ecf1u ^ prev[8];
|
||||
x[9] = 0xd0f3fd7cu ^ prev[9];
|
||||
x[10] = 0xd872eefeu ^ prev[10];
|
||||
x[11] = 0xc158c7bdu ^ prev[11];
|
||||
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
|
||||
mh_chacha_block(x, y);
|
||||
__global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
|
||||
}
|
||||
}
|
||||
|
||||
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
|
||||
static inline void mh_mixer(uint* s, uint rk) {
|
||||
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
|
||||
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
|
||||
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
|
||||
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
|
||||
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
|
||||
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
|
||||
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
|
||||
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
|
||||
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
|
||||
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
|
||||
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
|
||||
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
|
||||
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
|
||||
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
|
||||
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
|
||||
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
|
||||
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
|
||||
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
|
||||
}
|
||||
|
||||
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
|
||||
static inline void mh_item(__global const uint* cache, uint t, uint* s) {
|
||||
s[0] = 0xceed56d7u;
|
||||
s[1] = 0x9ba270d2u;
|
||||
s[2] = 0x82caab2du;
|
||||
s[3] = 0x81ebce0eu;
|
||||
s[4] = 0x12b6ecf1u;
|
||||
s[5] = 0xd0f3fd7cu;
|
||||
s[6] = 0xd872eefeu;
|
||||
s[7] = 0xc158c7bdu;
|
||||
s[8] = t * 0xf351d601u + 0xc6892460u;
|
||||
s[9] = t * 0xa3bb398fu + 0x25b7228au;
|
||||
s[10] = t * 0xb5a09e35u + 0xcd515004u;
|
||||
s[11] = t * 0x7509c9c1u + 0x2846527au;
|
||||
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
|
||||
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
|
||||
s[14] = t * 0xded91851u + 0x82961bacu;
|
||||
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
|
||||
for (uint r = 0u; r < 8u; ++r) {
|
||||
mh_mixer(s, 0x9E3779B9u * (r + 1u));
|
||||
__global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i];
|
||||
}
|
||||
mh_mixer(s, 0x9E3779B9u * 9u);
|
||||
}
|
||||
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 2 9 13 15 of w.
|
||||
static inline uint mh_j(uint w) { return ((w >> 2u) & 1u) | (((w >> 9u) & 1u) << 1) | (((w >> 13u) & 1u) << 2) | (((w >> 15u) & 1u) << 3); }
|
||||
static inline uint mh_t(uint w) { w = (w & 0x00007fffu) | ((w >> 16u) << 15u); w = (w & 0x00001fffu) | ((w >> 14u) << 13u); w = (w & 0x000001ffu) | ((w >> 10u) << 9u); w = (w & 0x00000003u) | ((w >> 3u) << 2u); return w; }
|
||||
static inline uint mh_addr(uint t, uint j) { uint w = t; w = ((w >> 2u) << 3u) | (w & 0x00000003u) | (((j >> 0u) & 1u) << 2u); w = ((w >> 9u) << 10u) | (w & 0x000001ffu) | (((j >> 1u) & 1u) << 9u); w = ((w >> 13u) << 14u) | (w & 0x00001fffu) | (((j >> 2u) & 1u) << 13u); w = ((w >> 15u) << 16u) | (w & 0x00007fffu) | (((j >> 3u) & 1u) << 15u); return w; }
|
||||
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
|
||||
static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }
|
||||
|
||||
// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item.
|
||||
// The same constants as memhard.h in this pack (one emitter, three dialects).
|
||||
__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) {
|
||||
uint seg = (uint)get_global_id(0);
|
||||
if (seg < nSegments) mh_cache_segment(cache, seg);
|
||||
}
|
||||
__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) {
|
||||
uint t = (uint)get_global_id(0);
|
||||
if (t < nItems) {
|
||||
uint s[16];
|
||||
mh_item(cache, t, s);
|
||||
for (uint i = 0u; i < 16u; ++i) ds[(ulong)mh_addr(t, i)] = s[i];
|
||||
}
|
||||
}
|
||||
|
||||
// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the
|
||||
// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and
|
||||
// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all).
|
||||
IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) {
|
||||
uint gid = (uint)get_global_id(0);
|
||||
uint lid = (uint)get_local_id(0);
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
#if IGNEUM_EXCHANGE == 0
|
||||
IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
|
||||
uint xk = 0u;
|
||||
#else
|
||||
(void)lid;
|
||||
#endif
|
||||
{ uint x = nonce ^ 0x667d0fbdu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x7b8e5963u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
|
||||
{ uint x = nonce ^ 0x7b8e5963u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x31c67e5eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
|
||||
{ uint x = nonce ^ 0x31c67e5eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4529ddc6u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
|
||||
{ uint x = nonce ^ 0x4529ddc6u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef19d6d8u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
|
||||
{ uint x = nonce ^ 0xef19d6d8u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xaccf6211u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
|
||||
{ uint x = nonce ^ 0xaccf6211u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xda0aed32u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
|
||||
{ uint x = nonce ^ 0xda0aed32u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xabc6df31u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
|
||||
{ uint x = nonce ^ 0xabc6df31u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x667d0fbdu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
|
||||
r7 = r7 ^ r0; // 1 xor
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 1u); r3 = r3 ^ t_; } // 2 shfl
|
||||
r4 = r4 - r1; // 3 sub
|
||||
r2 = r0 * r4 + r2; // 4 mad
|
||||
r4 = rotl_imm(r4, 9u); // 5 rotl
|
||||
r0 = rotr_var(r0, r2); // 6 rotr
|
||||
r6 = r6 ^ ds[((rotl_imm(r7 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r2 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r0 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r1 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
|
||||
r5 = r5 * r4; // 12 mul
|
||||
r7 = r7 ^ ds[((rotl_imm(r6 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 16u); r6 = r6 ^ t_; } // 14 shfl
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r3 = r3 ^ t_; } // 15 shfl
|
||||
r2 = r2 | r7; // 16 or
|
||||
r0 = rotr_var(r0, r6); // 17 rotr
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
|
||||
r7 = rotl_imm(r7, 24u); // 19 rotl
|
||||
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
|
||||
r2 = r2 | r6; // 21 or
|
||||
r6 = r5 * r4 + r6; // 22 mad
|
||||
r1 = r1 | r0; // 23 or
|
||||
r6 = r6 ^ r1; // 24 xor
|
||||
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
|
||||
r7 = rotr_var(r7, r0); // 26 rotr
|
||||
r4 = r5 * r7 + r4; // 27 mad
|
||||
r3 = r2 * r4 + r3; // 28 mad
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
|
||||
r2 = r2 ^ ds[((rotl_imm(r3 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r5 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
|
||||
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
|
||||
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
|
||||
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
|
||||
r2 = mul_hi(r2, r5); // 35 mulhi
|
||||
r4 = r4 ^ r2; // 36 xor
|
||||
r6 = r6 * r5; // 37 mul
|
||||
r7 = r7 ^ r0; // 38 xor
|
||||
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
|
||||
r2 = rotr_var(r2, r3); // 40 rotr
|
||||
r7 = r7 - r0; // 41 sub
|
||||
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 43 shfl
|
||||
r0 = r0 ^ ds[((rotl_imm(r7 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
|
||||
r3 = r3 ^ ds[((rotl_imm(r1 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
|
||||
r6 = r6 ^ ds[((rotl_imm(r3 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
|
||||
r4 = mul_hi(r4, r2); // 48 mulhi
|
||||
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r0 = r0 ^ t_; } // 50 shfl
|
||||
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r2 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
|
||||
r6 = rotl_imm(r6, 12u); // 53 rotl
|
||||
r3 = rotl_imm(r3, 12u); // 54 rotl
|
||||
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
|
||||
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r0 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
|
||||
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
|
||||
r3 = mul_hi(r3, r2); // 60 mulhi
|
||||
r6 = r6 | r4; // 61 or
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r5 = r5 ^ t_; } // 62 shfl
|
||||
r3 = r3 ^ ds[((rotl_imm(r6 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
|
||||
#if IGNEUM_EXCHANGE != 0
|
||||
// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the
|
||||
// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a
|
||||
// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md.
|
||||
IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) {
|
||||
if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); }
|
||||
}
|
||||
#endif
|
||||
163
proto-cuda/packs-ca2-era/era-4/kernel.cu
Normal file
163
proto-cuda/packs-ca2-era/era-4/kernel.cu
Normal file
|
|
@ -0,0 +1,163 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal).
|
||||
// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC.
|
||||
#include <cuda_runtime.h>
|
||||
#include <cstdint>
|
||||
#include "program.h"
|
||||
#include "memhard.h"
|
||||
|
||||
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31.
|
||||
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
|
||||
// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x.
|
||||
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) {
|
||||
uint32_t x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item.
|
||||
// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host.
|
||||
__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) {
|
||||
uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
if (seg < nSegments) mh_cache_segment(cache, seg);
|
||||
}
|
||||
__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
|
||||
uint32_t t = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
if (t < nItems) {
|
||||
uint32_t s[16];
|
||||
mh_item(cache, t, s);
|
||||
for (uint32_t i = 0u; i < 16u; ++i) ds[(size_t)mh_addr(t, i)] = s[i];
|
||||
}
|
||||
}
|
||||
|
||||
// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every
|
||||
// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a
|
||||
// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid.
|
||||
__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) {
|
||||
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
uint32_t nonce = baseNonce + gid;
|
||||
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint32_t x = nonce ^ 0x667d0fbdu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x7b8e5963u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
|
||||
{ uint32_t x = nonce ^ 0x7b8e5963u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x31c67e5eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
|
||||
{ uint32_t x = nonce ^ 0x31c67e5eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4529ddc6u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
|
||||
{ uint32_t x = nonce ^ 0x4529ddc6u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef19d6d8u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
|
||||
{ uint32_t x = nonce ^ 0xef19d6d8u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xaccf6211u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
|
||||
{ uint32_t x = nonce ^ 0xaccf6211u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xda0aed32u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
|
||||
{ uint32_t x = nonce ^ 0xda0aed32u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xabc6df31u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
|
||||
{ uint32_t x = nonce ^ 0xabc6df31u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x667d0fbdu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
|
||||
|
||||
for (uint32_t it = 0u; it < 8u; ++it) {
|
||||
uint32_t sel = r0;
|
||||
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
|
||||
r7 = r7 ^ r0; // 1 xor
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r6, 1); // 2 shfl
|
||||
r4 = r4 - r1; // 3 sub
|
||||
r2 = r0 * r4 + r2; // 4 mad
|
||||
r4 = rotl_imm(r4, 9u); // 5 rotl
|
||||
r0 = rotr_var(r0, r2); // 6 rotr
|
||||
r6 = r6 ^ ds[((rotl_imm(r7 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r2 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r0 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r1 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
|
||||
r5 = r5 * r4; // 12 mul
|
||||
r7 = r7 ^ ds[((rotl_imm(r6 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
|
||||
r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r5, 16); // 14 shfl
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r5, 1); // 15 shfl
|
||||
r2 = r2 | r7; // 16 or
|
||||
r0 = rotr_var(r0, r6); // 17 rotr
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
|
||||
r7 = rotl_imm(r7, 24u); // 19 rotl
|
||||
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
|
||||
r2 = r2 | r6; // 21 or
|
||||
r6 = r5 * r4 + r6; // 22 mad
|
||||
r1 = r1 | r0; // 23 or
|
||||
r6 = r6 ^ r1; // 24 xor
|
||||
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
|
||||
r7 = rotr_var(r7, r0); // 26 rotr
|
||||
r4 = r5 * r7 + r4; // 27 mad
|
||||
r3 = r2 * r4 + r3; // 28 mad
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
|
||||
r2 = r2 ^ ds[((rotl_imm(r3 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r5 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
|
||||
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
|
||||
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
|
||||
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
|
||||
r2 = __umulhi(r2, r5); // 35 mulhi
|
||||
r4 = r4 ^ r2; // 36 xor
|
||||
r6 = r6 * r5; // 37 mul
|
||||
r7 = r7 ^ r0; // 38 xor
|
||||
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
|
||||
r2 = rotr_var(r2, r3); // 40 rotr
|
||||
r7 = r7 - r0; // 41 sub
|
||||
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 43 shfl
|
||||
r0 = r0 ^ ds[((rotl_imm(r7 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
|
||||
r3 = r3 ^ ds[((rotl_imm(r1 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
|
||||
r6 = r6 ^ ds[((rotl_imm(r3 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
|
||||
r4 = __umulhi(r4, r2); // 48 mulhi
|
||||
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
|
||||
r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r7, 4); // 50 shfl
|
||||
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r2 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
|
||||
r6 = rotl_imm(r6, 12u); // 53 rotl
|
||||
r3 = rotl_imm(r3, 12u); // 54 rotl
|
||||
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
|
||||
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r0 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
|
||||
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
|
||||
r3 = __umulhi(r3, r2); // 60 mulhi
|
||||
r6 = r6 | r4; // 61 or
|
||||
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 62 shfl
|
||||
r3 = r3 ^ ds[((rotl_imm(r6 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
|
||||
}
|
||||
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
|
||||
}
|
||||
|
||||
// Host-side launch wrappers. Declared in program.h, called from host.cu.
|
||||
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) {
|
||||
if (nSegments == 0u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 256u;
|
||||
uint32_t grid = (nSegments + block - 1u) / block;
|
||||
igneum_cache_fill<<<grid, block>>>(cache, nSegments);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
|
||||
if (nItems == 0u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 256u;
|
||||
uint32_t grid = (nItems + block - 1u) / block;
|
||||
igneum_build<<<grid, block>>>(ds, cache, nItems);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
uint32_t nonces, uint32_t blockWarps) {
|
||||
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 32u * blockWarps;
|
||||
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
|
||||
igneum_hash<<<nonces / block, block>>>(ds, out, baseNonce, mask);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
|
||||
cudaFuncAttributes attr;
|
||||
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash);
|
||||
if (e != cudaSuccess) return e;
|
||||
*numRegs = attr.numRegs;
|
||||
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0);
|
||||
}
|
||||
375
proto-cuda/packs-ca2-era/era-4/kernel_bound.cl
Normal file
375
proto-cuda/packs-ca2-era/era-4/kernel_bound.cl
Normal file
|
|
@ -0,0 +1,375 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal).
|
||||
// Built from source at runtime by proto-opencl/host.c, which passes these defines:
|
||||
// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit)
|
||||
// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default)
|
||||
// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32
|
||||
// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition
|
||||
// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units;
|
||||
// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32.
|
||||
#ifndef IGNEUM_GROUP
|
||||
#define IGNEUM_GROUP 32
|
||||
#endif
|
||||
#ifndef IGNEUM_EXCHANGE
|
||||
#define IGNEUM_EXCHANGE 0
|
||||
#endif
|
||||
#ifdef __OPENCL_VERSION__
|
||||
#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1)))
|
||||
#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n]
|
||||
#if IGNEUM_EXCHANGE == 1
|
||||
#ifdef cl_khr_subgroups
|
||||
#pragma OPENCL EXTENSION cl_khr_subgroups : enable
|
||||
#endif
|
||||
#ifdef cl_khr_subgroup_shuffle
|
||||
#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable
|
||||
#endif
|
||||
#elif IGNEUM_EXCHANGE == 2
|
||||
#pragma OPENCL EXTENSION cl_intel_subgroups : enable
|
||||
#endif
|
||||
#else
|
||||
// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros.
|
||||
#include "emu_opencl.h"
|
||||
#endif
|
||||
|
||||
#if IGNEUM_EXCHANGE == 1
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m))
|
||||
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
|
||||
#elif IGNEUM_EXCHANGE == 2
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m))
|
||||
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
|
||||
#else
|
||||
// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per
|
||||
// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane
|
||||
// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's
|
||||
// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier.
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; }
|
||||
#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; }
|
||||
#endif
|
||||
|
||||
static inline uint splitmix32(uint x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32.
|
||||
static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); }
|
||||
// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x.
|
||||
static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); }
|
||||
static inline uint ds_elem(uint i, uint d0, uint d1) {
|
||||
uint x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
|
||||
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
|
||||
#define MH_CACHE_LINE_MASK 0x003fffffu
|
||||
#define MH_SEGMENT_LINES 64u
|
||||
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
|
||||
static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
|
||||
|
||||
// y = ChaCha12 core(x) + x
|
||||
static inline void mh_chacha_block(const uint* x, uint* y) {
|
||||
for (uint i = 0u; i < 16u; ++i) y[i] = x[i];
|
||||
for (uint r = 0u; r < 6u; ++r) {
|
||||
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
|
||||
}
|
||||
for (uint i = 0u; i < 16u; ++i) y[i] += x[i];
|
||||
}
|
||||
|
||||
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
|
||||
static inline void mh_cache_segment(__global uint* cache, uint seg) {
|
||||
uint prev[16]; uint x[16]; uint y[16];
|
||||
for (uint i = 0u; i < 16u; ++i) prev[i] = 0u;
|
||||
for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) {
|
||||
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
|
||||
x[4] = 0xceed56d7u ^ prev[4];
|
||||
x[5] = 0x9ba270d2u ^ prev[5];
|
||||
x[6] = 0x82caab2du ^ prev[6];
|
||||
x[7] = 0x81ebce0eu ^ prev[7];
|
||||
x[8] = 0x12b6ecf1u ^ prev[8];
|
||||
x[9] = 0xd0f3fd7cu ^ prev[9];
|
||||
x[10] = 0xd872eefeu ^ prev[10];
|
||||
x[11] = 0xc158c7bdu ^ prev[11];
|
||||
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
|
||||
mh_chacha_block(x, y);
|
||||
__global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
|
||||
}
|
||||
}
|
||||
|
||||
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
|
||||
static inline void mh_mixer(uint* s, uint rk) {
|
||||
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
|
||||
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
|
||||
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
|
||||
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
|
||||
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
|
||||
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
|
||||
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
|
||||
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
|
||||
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
|
||||
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
|
||||
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
|
||||
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
|
||||
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
|
||||
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
|
||||
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
|
||||
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
|
||||
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
|
||||
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
|
||||
}
|
||||
|
||||
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
|
||||
static inline void mh_item(__global const uint* cache, uint t, uint* s) {
|
||||
s[0] = 0xceed56d7u;
|
||||
s[1] = 0x9ba270d2u;
|
||||
s[2] = 0x82caab2du;
|
||||
s[3] = 0x81ebce0eu;
|
||||
s[4] = 0x12b6ecf1u;
|
||||
s[5] = 0xd0f3fd7cu;
|
||||
s[6] = 0xd872eefeu;
|
||||
s[7] = 0xc158c7bdu;
|
||||
s[8] = t * 0xf351d601u + 0xc6892460u;
|
||||
s[9] = t * 0xa3bb398fu + 0x25b7228au;
|
||||
s[10] = t * 0xb5a09e35u + 0xcd515004u;
|
||||
s[11] = t * 0x7509c9c1u + 0x2846527au;
|
||||
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
|
||||
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
|
||||
s[14] = t * 0xded91851u + 0x82961bacu;
|
||||
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
|
||||
for (uint r = 0u; r < 8u; ++r) {
|
||||
mh_mixer(s, 0x9E3779B9u * (r + 1u));
|
||||
__global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i];
|
||||
}
|
||||
mh_mixer(s, 0x9E3779B9u * 9u);
|
||||
}
|
||||
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 2 9 13 15 of w.
|
||||
static inline uint mh_j(uint w) { return ((w >> 2u) & 1u) | (((w >> 9u) & 1u) << 1) | (((w >> 13u) & 1u) << 2) | (((w >> 15u) & 1u) << 3); }
|
||||
static inline uint mh_t(uint w) { w = (w & 0x00007fffu) | ((w >> 16u) << 15u); w = (w & 0x00001fffu) | ((w >> 14u) << 13u); w = (w & 0x000001ffu) | ((w >> 10u) << 9u); w = (w & 0x00000003u) | ((w >> 3u) << 2u); return w; }
|
||||
static inline uint mh_addr(uint t, uint j) { uint w = t; w = ((w >> 2u) << 3u) | (w & 0x00000003u) | (((j >> 0u) & 1u) << 2u); w = ((w >> 9u) << 10u) | (w & 0x000001ffu) | (((j >> 1u) & 1u) << 9u); w = ((w >> 13u) << 14u) | (w & 0x00001fffu) | (((j >> 2u) & 1u) << 13u); w = ((w >> 15u) << 16u) | (w & 0x00007fffu) | (((j >> 3u) & 1u) << 15u); return w; }
|
||||
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
|
||||
static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }
|
||||
|
||||
// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item.
|
||||
// The same constants as memhard.h in this pack (one emitter, three dialects).
|
||||
__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) {
|
||||
uint seg = (uint)get_global_id(0);
|
||||
if (seg < nSegments) mh_cache_segment(cache, seg);
|
||||
}
|
||||
__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) {
|
||||
uint t = (uint)get_global_id(0);
|
||||
if (t < nItems) {
|
||||
uint s[16];
|
||||
mh_item(cache, t, s);
|
||||
for (uint i = 0u; i < 16u; ++i) ds[(ulong)mh_addr(t, i)] = s[i];
|
||||
}
|
||||
}
|
||||
|
||||
// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the
|
||||
// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and
|
||||
// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all).
|
||||
IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) {
|
||||
uint gid = (uint)get_global_id(0);
|
||||
uint lid = (uint)get_local_id(0);
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
#if IGNEUM_EXCHANGE == 0
|
||||
IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
|
||||
uint xk = 0u;
|
||||
#else
|
||||
(void)lid;
|
||||
#endif
|
||||
{ uint x = nonce ^ 0x667d0fbdu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x7b8e5963u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
|
||||
{ uint x = nonce ^ 0x7b8e5963u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x31c67e5eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
|
||||
{ uint x = nonce ^ 0x31c67e5eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4529ddc6u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
|
||||
{ uint x = nonce ^ 0x4529ddc6u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef19d6d8u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
|
||||
{ uint x = nonce ^ 0xef19d6d8u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xaccf6211u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
|
||||
{ uint x = nonce ^ 0xaccf6211u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xda0aed32u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
|
||||
{ uint x = nonce ^ 0xda0aed32u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xabc6df31u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
|
||||
{ uint x = nonce ^ 0xabc6df31u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x667d0fbdu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
|
||||
r7 = r7 ^ r0; // 1 xor
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 1u); r3 = r3 ^ t_; } // 2 shfl
|
||||
r4 = r4 - r1; // 3 sub
|
||||
r2 = r0 * r4 + r2; // 4 mad
|
||||
r4 = rotl_imm(r4, 9u); // 5 rotl
|
||||
r0 = rotr_var(r0, r2); // 6 rotr
|
||||
r6 = r6 ^ ds[((rotl_imm(r7 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r2 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r0 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r1 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
|
||||
r5 = r5 * r4; // 12 mul
|
||||
r7 = r7 ^ ds[((rotl_imm(r6 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 16u); r6 = r6 ^ t_; } // 14 shfl
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r3 = r3 ^ t_; } // 15 shfl
|
||||
r2 = r2 | r7; // 16 or
|
||||
r0 = rotr_var(r0, r6); // 17 rotr
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
|
||||
r7 = rotl_imm(r7, 24u); // 19 rotl
|
||||
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
|
||||
r2 = r2 | r6; // 21 or
|
||||
r6 = r5 * r4 + r6; // 22 mad
|
||||
r1 = r1 | r0; // 23 or
|
||||
r6 = r6 ^ r1; // 24 xor
|
||||
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
|
||||
r7 = rotr_var(r7, r0); // 26 rotr
|
||||
r4 = r5 * r7 + r4; // 27 mad
|
||||
r3 = r2 * r4 + r3; // 28 mad
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
|
||||
r2 = r2 ^ ds[((rotl_imm(r3 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r5 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
|
||||
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
|
||||
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
|
||||
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
|
||||
r2 = mul_hi(r2, r5); // 35 mulhi
|
||||
r4 = r4 ^ r2; // 36 xor
|
||||
r6 = r6 * r5; // 37 mul
|
||||
r7 = r7 ^ r0; // 38 xor
|
||||
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
|
||||
r2 = rotr_var(r2, r3); // 40 rotr
|
||||
r7 = r7 - r0; // 41 sub
|
||||
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 43 shfl
|
||||
r0 = r0 ^ ds[((rotl_imm(r7 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
|
||||
r3 = r3 ^ ds[((rotl_imm(r1 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
|
||||
r6 = r6 ^ ds[((rotl_imm(r3 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
|
||||
r4 = mul_hi(r4, r2); // 48 mulhi
|
||||
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r0 = r0 ^ t_; } // 50 shfl
|
||||
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r2 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
|
||||
r6 = rotl_imm(r6, 12u); // 53 rotl
|
||||
r3 = rotl_imm(r3, 12u); // 54 rotl
|
||||
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
|
||||
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r0 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
|
||||
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
|
||||
r3 = mul_hi(r3, r2); // 60 mulhi
|
||||
r6 = r6 | r4; // 61 or
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r5 = r5 ^ t_; } // 62 shfl
|
||||
r3 = r3 ^ ds[((rotl_imm(r6 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
|
||||
#if IGNEUM_EXCHANGE != 0
|
||||
// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the
|
||||
// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a
|
||||
// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md.
|
||||
IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) {
|
||||
if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); }
|
||||
}
|
||||
#endif
|
||||
|
||||
// Header-bound variant (bind.rs): the init words come from initw, not SEEDW. Same body as igneum_hash.
|
||||
IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw) {
|
||||
uint gid = (uint)get_global_id(0);
|
||||
uint lid = (uint)get_local_id(0);
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7];
|
||||
#if IGNEUM_EXCHANGE == 0
|
||||
IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
|
||||
uint xk = 0u;
|
||||
#else
|
||||
(void)lid;
|
||||
#endif
|
||||
{ uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; }
|
||||
{ uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; }
|
||||
{ uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; }
|
||||
{ uint x = nonce ^ iw3; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw4; }
|
||||
{ uint x = nonce ^ iw4; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw5; }
|
||||
{ uint x = nonce ^ iw5; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw6; }
|
||||
{ uint x = nonce ^ iw6; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw7; }
|
||||
{ uint x = nonce ^ iw7; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw0; }
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
|
||||
r7 = r7 ^ r0; // 1 xor
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 1u); r3 = r3 ^ t_; } // 2 shfl
|
||||
r4 = r4 - r1; // 3 sub
|
||||
r2 = r0 * r4 + r2; // 4 mad
|
||||
r4 = rotl_imm(r4, 9u); // 5 rotl
|
||||
r0 = rotr_var(r0, r2); // 6 rotr
|
||||
r6 = r6 ^ ds[((rotl_imm(r7 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r2 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r0 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r1 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
|
||||
r5 = r5 * r4; // 12 mul
|
||||
r7 = r7 ^ ds[((rotl_imm(r6 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 16u); r6 = r6 ^ t_; } // 14 shfl
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r3 = r3 ^ t_; } // 15 shfl
|
||||
r2 = r2 | r7; // 16 or
|
||||
r0 = rotr_var(r0, r6); // 17 rotr
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
|
||||
r7 = rotl_imm(r7, 24u); // 19 rotl
|
||||
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
|
||||
r2 = r2 | r6; // 21 or
|
||||
r6 = r5 * r4 + r6; // 22 mad
|
||||
r1 = r1 | r0; // 23 or
|
||||
r6 = r6 ^ r1; // 24 xor
|
||||
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
|
||||
r7 = rotr_var(r7, r0); // 26 rotr
|
||||
r4 = r5 * r7 + r4; // 27 mad
|
||||
r3 = r2 * r4 + r3; // 28 mad
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
|
||||
r2 = r2 ^ ds[((rotl_imm(r3 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r5 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
|
||||
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
|
||||
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
|
||||
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
|
||||
r2 = mul_hi(r2, r5); // 35 mulhi
|
||||
r4 = r4 ^ r2; // 36 xor
|
||||
r6 = r6 * r5; // 37 mul
|
||||
r7 = r7 ^ r0; // 38 xor
|
||||
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
|
||||
r2 = rotr_var(r2, r3); // 40 rotr
|
||||
r7 = r7 - r0; // 41 sub
|
||||
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 43 shfl
|
||||
r0 = r0 ^ ds[((rotl_imm(r7 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
|
||||
r3 = r3 ^ ds[((rotl_imm(r1 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
|
||||
r6 = r6 ^ ds[((rotl_imm(r3 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
|
||||
r4 = mul_hi(r4, r2); // 48 mulhi
|
||||
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r0 = r0 ^ t_; } // 50 shfl
|
||||
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r2 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
|
||||
r6 = rotl_imm(r6, 12u); // 53 rotl
|
||||
r3 = rotl_imm(r3, 12u); // 54 rotl
|
||||
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
|
||||
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r0 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
|
||||
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
|
||||
r3 = mul_hi(r3, r2); // 60 mulhi
|
||||
r6 = r6 | r4; // 61 or
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r5 = r5 ^ t_; } // 62 shfl
|
||||
r3 = r3 ^ ds[((rotl_imm(r6 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
123
proto-cuda/packs-ca2-era/era-4/kernel_bound.cu
Normal file
123
proto-cuda/packs-ca2-era/era-4/kernel_bound.cu
Normal file
|
|
@ -0,0 +1,123 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// Header-bound twin of igneum_hash in kernel.cu: the init words come from a kernel argument, not SEEDW.
|
||||
// Host declarations (also in program_bound.h if present):
|
||||
// struct IgneumInitWords { uint32_t w[8]; };
|
||||
// cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
// IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps);
|
||||
// cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
|
||||
#include <cuda_runtime.h>
|
||||
#include <cstdint>
|
||||
#include "program.h"
|
||||
|
||||
struct IgneumInitWords { uint32_t w[8]; };
|
||||
|
||||
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
|
||||
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
|
||||
__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw) {
|
||||
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
uint32_t nonce = baseNonce + gid;
|
||||
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint32_t x = nonce ^ iw.w[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw.w[1]; }
|
||||
{ uint32_t x = nonce ^ iw.w[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw.w[2]; }
|
||||
{ uint32_t x = nonce ^ iw.w[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw.w[3]; }
|
||||
{ uint32_t x = nonce ^ iw.w[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw.w[4]; }
|
||||
{ uint32_t x = nonce ^ iw.w[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw.w[5]; }
|
||||
{ uint32_t x = nonce ^ iw.w[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw.w[6]; }
|
||||
{ uint32_t x = nonce ^ iw.w[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw.w[7]; }
|
||||
{ uint32_t x = nonce ^ iw.w[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw.w[0]; }
|
||||
|
||||
for (uint32_t it = 0u; it < 8u; ++it) {
|
||||
uint32_t sel = r0;
|
||||
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
|
||||
r7 = r7 ^ r0; // 1 xor
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r6, 1); // 2 shfl
|
||||
r4 = r4 - r1; // 3 sub
|
||||
r2 = r0 * r4 + r2; // 4 mad
|
||||
r4 = rotl_imm(r4, 9u); // 5 rotl
|
||||
r0 = rotr_var(r0, r2); // 6 rotr
|
||||
r6 = r6 ^ ds[((rotl_imm(r7 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r2 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r0 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r1 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
|
||||
r5 = r5 * r4; // 12 mul
|
||||
r7 = r7 ^ ds[((rotl_imm(r6 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
|
||||
r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r5, 16); // 14 shfl
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r5, 1); // 15 shfl
|
||||
r2 = r2 | r7; // 16 or
|
||||
r0 = rotr_var(r0, r6); // 17 rotr
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
|
||||
r7 = rotl_imm(r7, 24u); // 19 rotl
|
||||
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
|
||||
r2 = r2 | r6; // 21 or
|
||||
r6 = r5 * r4 + r6; // 22 mad
|
||||
r1 = r1 | r0; // 23 or
|
||||
r6 = r6 ^ r1; // 24 xor
|
||||
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
|
||||
r7 = rotr_var(r7, r0); // 26 rotr
|
||||
r4 = r5 * r7 + r4; // 27 mad
|
||||
r3 = r2 * r4 + r3; // 28 mad
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
|
||||
r2 = r2 ^ ds[((rotl_imm(r3 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r5 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
|
||||
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
|
||||
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
|
||||
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
|
||||
r2 = __umulhi(r2, r5); // 35 mulhi
|
||||
r4 = r4 ^ r2; // 36 xor
|
||||
r6 = r6 * r5; // 37 mul
|
||||
r7 = r7 ^ r0; // 38 xor
|
||||
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
|
||||
r2 = rotr_var(r2, r3); // 40 rotr
|
||||
r7 = r7 - r0; // 41 sub
|
||||
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 43 shfl
|
||||
r0 = r0 ^ ds[((rotl_imm(r7 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
|
||||
r3 = r3 ^ ds[((rotl_imm(r1 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
|
||||
r6 = r6 ^ ds[((rotl_imm(r3 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
|
||||
r4 = __umulhi(r4, r2); // 48 mulhi
|
||||
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
|
||||
r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r7, 4); // 50 shfl
|
||||
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r2 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
|
||||
r6 = rotl_imm(r6, 12u); // 53 rotl
|
||||
r3 = rotl_imm(r3, 12u); // 54 rotl
|
||||
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
|
||||
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r0 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
|
||||
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
|
||||
r3 = __umulhi(r3, r2); // 60 mulhi
|
||||
r6 = r6 | r4; // 61 or
|
||||
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 62 shfl
|
||||
r3 = r3 ^ ds[((rotl_imm(r6 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
|
||||
}
|
||||
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps) {
|
||||
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 32u * blockWarps;
|
||||
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
|
||||
igneum_hash_bound<<<nonces / block, block>>>(ds, out, baseNonce, mask, iw);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
|
||||
cudaFuncAttributes attr;
|
||||
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash_bound);
|
||||
if (e != cudaSuccess) return e;
|
||||
*numRegs = attr.numRegs;
|
||||
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash_bound, (int)(32u * blockWarps), 0);
|
||||
}
|
||||
112
proto-cuda/packs-ca2-era/era-4/memhard.h
Normal file
112
proto-cuda/packs-ca2-era/era-4/memhard.h
Normal file
|
|
@ -0,0 +1,112 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// Memory-hard dataset core, the same text that the Mac's Metal kernels and CPU verifier were checked against.
|
||||
// Included by kernel.cu (device), host.cu (host reference) and proto-opencl/host.c (C99 host reference).
|
||||
// See proto-metal/MEMHARD.md for the construction. kernel.cl carries the same text in OpenCL C.
|
||||
#pragma once
|
||||
#ifdef __cplusplus
|
||||
#include <cstdint>
|
||||
#else
|
||||
#include <stdint.h>
|
||||
#endif
|
||||
#if defined(__CUDACC__)
|
||||
#define IGNEUM_HD __host__ __device__ __forceinline__
|
||||
#elif defined(_MSC_VER) && !defined(__cplusplus)
|
||||
#define IGNEUM_HD static __inline
|
||||
#else
|
||||
#define IGNEUM_HD static inline
|
||||
#endif
|
||||
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
|
||||
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
|
||||
#define MH_CACHE_LINE_MASK 0x003fffffu
|
||||
#define MH_SEGMENT_LINES 64u
|
||||
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
|
||||
IGNEUM_HD uint32_t mh_rotl(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
|
||||
|
||||
// y = ChaCha12 core(x) + x
|
||||
IGNEUM_HD void mh_chacha_block(const uint32_t* x, uint32_t* y) {
|
||||
for (uint32_t i = 0u; i < 16u; ++i) y[i] = x[i];
|
||||
for (uint32_t r = 0u; r < 6u; ++r) {
|
||||
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
|
||||
}
|
||||
for (uint32_t i = 0u; i < 16u; ++i) y[i] += x[i];
|
||||
}
|
||||
|
||||
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
|
||||
IGNEUM_HD void mh_cache_segment(uint32_t* cache, uint32_t seg) {
|
||||
uint32_t prev[16]; uint32_t x[16]; uint32_t y[16];
|
||||
for (uint32_t i = 0u; i < 16u; ++i) prev[i] = 0u;
|
||||
for (uint32_t j = 0u; j < MH_SEGMENT_LINES; ++j) {
|
||||
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
|
||||
x[4] = 0xceed56d7u ^ prev[4];
|
||||
x[5] = 0x9ba270d2u ^ prev[5];
|
||||
x[6] = 0x82caab2du ^ prev[6];
|
||||
x[7] = 0x81ebce0eu ^ prev[7];
|
||||
x[8] = 0x12b6ecf1u ^ prev[8];
|
||||
x[9] = 0xd0f3fd7cu ^ prev[9];
|
||||
x[10] = 0xd872eefeu ^ prev[10];
|
||||
x[11] = 0xc158c7bdu ^ prev[11];
|
||||
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
|
||||
mh_chacha_block(x, y);
|
||||
uint32_t* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
|
||||
for (uint32_t i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
|
||||
}
|
||||
}
|
||||
|
||||
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
|
||||
IGNEUM_HD void mh_mixer(uint32_t* s, uint32_t rk) {
|
||||
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
|
||||
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
|
||||
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
|
||||
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
|
||||
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
|
||||
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
|
||||
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
|
||||
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
|
||||
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
|
||||
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
|
||||
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
|
||||
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
|
||||
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
|
||||
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
|
||||
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
|
||||
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
|
||||
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
|
||||
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
|
||||
}
|
||||
|
||||
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
|
||||
IGNEUM_HD void mh_item(const uint32_t* cache, uint32_t t, uint32_t* s) {
|
||||
s[0] = 0xceed56d7u;
|
||||
s[1] = 0x9ba270d2u;
|
||||
s[2] = 0x82caab2du;
|
||||
s[3] = 0x81ebce0eu;
|
||||
s[4] = 0x12b6ecf1u;
|
||||
s[5] = 0xd0f3fd7cu;
|
||||
s[6] = 0xd872eefeu;
|
||||
s[7] = 0xc158c7bdu;
|
||||
s[8] = t * 0xf351d601u + 0xc6892460u;
|
||||
s[9] = t * 0xa3bb398fu + 0x25b7228au;
|
||||
s[10] = t * 0xb5a09e35u + 0xcd515004u;
|
||||
s[11] = t * 0x7509c9c1u + 0x2846527au;
|
||||
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
|
||||
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
|
||||
s[14] = t * 0xded91851u + 0x82961bacu;
|
||||
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
|
||||
for (uint32_t r = 0u; r < 8u; ++r) {
|
||||
mh_mixer(s, 0x9E3779B9u * (r + 1u));
|
||||
const uint32_t* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
|
||||
for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= line[i];
|
||||
}
|
||||
mh_mixer(s, 0x9E3779B9u * 9u);
|
||||
}
|
||||
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 2 9 13 15 of w.
|
||||
IGNEUM_HD uint32_t mh_j(uint32_t w) { return ((w >> 2u) & 1u) | (((w >> 9u) & 1u) << 1) | (((w >> 13u) & 1u) << 2) | (((w >> 15u) & 1u) << 3); }
|
||||
IGNEUM_HD uint32_t mh_t(uint32_t w) { w = (w & 0x00007fffu) | ((w >> 16u) << 15u); w = (w & 0x00001fffu) | ((w >> 14u) << 13u); w = (w & 0x000001ffu) | ((w >> 10u) << 9u); w = (w & 0x00000003u) | ((w >> 3u) << 2u); return w; }
|
||||
IGNEUM_HD uint32_t mh_addr(uint32_t t, uint32_t j) { uint32_t w = t; w = ((w >> 2u) << 3u) | (w & 0x00000003u) | (((j >> 0u) & 1u) << 2u); w = ((w >> 9u) << 10u) | (w & 0x000001ffu) | (((j >> 1u) & 1u) << 9u); w = ((w >> 13u) << 14u) | (w & 0x00001fffu) | (((j >> 2u) & 1u) << 13u); w = ((w >> 15u) << 16u) | (w & 0x00007fffu) | (((j >> 3u) & 1u) << 15u); return w; }
|
||||
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
|
||||
IGNEUM_HD uint32_t mh_word(const uint32_t* cache, uint32_t w) { uint32_t s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }
|
||||
109
proto-cuda/packs-ca2-era/era-4/memhard.metal
Normal file
109
proto-cuda/packs-ca2-era/era-4/memhard.metal
Normal file
|
|
@ -0,0 +1,109 @@
|
|||
#include <metal_stdlib>
|
||||
using namespace metal;
|
||||
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
|
||||
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
|
||||
#define MH_CACHE_LINE_MASK 0x003fffffu
|
||||
#define MH_SEGMENT_LINES 64u
|
||||
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
|
||||
inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
|
||||
|
||||
// y = ChaCha12 core(x) + x
|
||||
inline void mh_chacha_block(const thread uint* x, thread uint* y) {
|
||||
for (uint i = 0u; i < 16u; ++i) y[i] = x[i];
|
||||
for (uint r = 0u; r < 6u; ++r) {
|
||||
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
|
||||
}
|
||||
for (uint i = 0u; i < 16u; ++i) y[i] += x[i];
|
||||
}
|
||||
|
||||
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
|
||||
inline void mh_cache_segment(device uint* cache, uint seg) {
|
||||
uint prev[16]; uint x[16]; uint y[16];
|
||||
for (uint i = 0u; i < 16u; ++i) prev[i] = 0u;
|
||||
for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) {
|
||||
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
|
||||
x[4] = 0xceed56d7u ^ prev[4];
|
||||
x[5] = 0x9ba270d2u ^ prev[5];
|
||||
x[6] = 0x82caab2du ^ prev[6];
|
||||
x[7] = 0x81ebce0eu ^ prev[7];
|
||||
x[8] = 0x12b6ecf1u ^ prev[8];
|
||||
x[9] = 0xd0f3fd7cu ^ prev[9];
|
||||
x[10] = 0xd872eefeu ^ prev[10];
|
||||
x[11] = 0xc158c7bdu ^ prev[11];
|
||||
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
|
||||
mh_chacha_block(x, y);
|
||||
device uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
|
||||
}
|
||||
}
|
||||
|
||||
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
|
||||
inline void mh_mixer(thread uint* s, uint rk) {
|
||||
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
|
||||
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
|
||||
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
|
||||
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
|
||||
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
|
||||
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
|
||||
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
|
||||
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
|
||||
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
|
||||
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
|
||||
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
|
||||
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
|
||||
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
|
||||
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
|
||||
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
|
||||
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
|
||||
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
|
||||
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
|
||||
}
|
||||
|
||||
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
|
||||
inline void mh_item(device const uint* cache, uint t, thread uint* s) {
|
||||
s[0] = 0xceed56d7u;
|
||||
s[1] = 0x9ba270d2u;
|
||||
s[2] = 0x82caab2du;
|
||||
s[3] = 0x81ebce0eu;
|
||||
s[4] = 0x12b6ecf1u;
|
||||
s[5] = 0xd0f3fd7cu;
|
||||
s[6] = 0xd872eefeu;
|
||||
s[7] = 0xc158c7bdu;
|
||||
s[8] = t * 0xf351d601u + 0xc6892460u;
|
||||
s[9] = t * 0xa3bb398fu + 0x25b7228au;
|
||||
s[10] = t * 0xb5a09e35u + 0xcd515004u;
|
||||
s[11] = t * 0x7509c9c1u + 0x2846527au;
|
||||
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
|
||||
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
|
||||
s[14] = t * 0xded91851u + 0x82961bacu;
|
||||
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
|
||||
for (uint r = 0u; r < 8u; ++r) {
|
||||
mh_mixer(s, 0x9E3779B9u * (r + 1u));
|
||||
device const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i];
|
||||
}
|
||||
mh_mixer(s, 0x9E3779B9u * 9u);
|
||||
}
|
||||
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 2 9 13 15 of w.
|
||||
inline uint mh_j(uint w) { return ((w >> 2u) & 1u) | (((w >> 9u) & 1u) << 1) | (((w >> 13u) & 1u) << 2) | (((w >> 15u) & 1u) << 3); }
|
||||
inline uint mh_t(uint w) { w = (w & 0x00007fffu) | ((w >> 16u) << 15u); w = (w & 0x00001fffu) | ((w >> 14u) << 13u); w = (w & 0x000001ffu) | ((w >> 10u) << 9u); w = (w & 0x00000003u) | ((w >> 3u) << 2u); return w; }
|
||||
inline uint mh_addr(uint t, uint j) { uint w = t; w = ((w >> 2u) << 3u) | (w & 0x00000003u) | (((j >> 0u) & 1u) << 2u); w = ((w >> 9u) << 10u) | (w & 0x000001ffu) | (((j >> 1u) & 1u) << 9u); w = ((w >> 13u) << 14u) | (w & 0x00001fffu) | (((j >> 2u) & 1u) << 13u); w = ((w >> 15u) << 16u) | (w & 0x00007fffu) | (((j >> 3u) & 1u) << 15u); return w; }
|
||||
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
|
||||
inline uint mh_word(device const uint* cache, uint w) { uint s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }
|
||||
|
||||
// One thread per segment (2^16 threads).
|
||||
kernel void igneum_cache_fill(device uint* cache [[buffer(0)]], uint gid [[thread_position_in_grid]]) {
|
||||
mh_cache_segment(cache, gid);
|
||||
}
|
||||
// One thread per 64-byte item (dataset words / 16 threads).
|
||||
kernel void igneum_build(device const uint* cache [[buffer(0)]], device uint* dataset [[buffer(1)]],
|
||||
uint gid [[thread_position_in_grid]]) {
|
||||
uint s[16];
|
||||
mh_item(cache, gid, s);
|
||||
for (uint i = 0u; i < 16u; ++i) dataset[mh_addr(gid, i)] = s[i];
|
||||
}
|
||||
76
proto-cuda/packs-ca2-era/era-4/program.h
Normal file
76
proto-cuda/packs-ca2-era/era-4/program.h
Normal file
|
|
@ -0,0 +1,76 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// Program metadata for host.cu plus the launch wrappers defined in kernel.cu.
|
||||
// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros.
|
||||
#pragma once
|
||||
#ifdef __cplusplus
|
||||
#include <cstdint>
|
||||
#else
|
||||
#include <stdint.h>
|
||||
#endif
|
||||
#ifndef IGNEUM_NO_CUDA
|
||||
#include <cuda_runtime.h>
|
||||
#endif
|
||||
|
||||
#define IGNEUM_SEED_STRING "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000"
|
||||
#define IGNEUM_SEED_BYTES_HEX "edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07"
|
||||
#define IGNEUM_GENERATOR 3
|
||||
#define IGNEUM_PROGRAM_ATTEMPT 0
|
||||
#define IGNEUM_PROGRAM_ID 0x73bcbfe8ccf988f1ull
|
||||
#define IGNEUM_DAY_STRING "bytes:69676e65756d2d6461792ffa50000000000000"
|
||||
#define IGNEUM_DAY_BYTES_HEX "69676e65756d2d6461792ffa50000000000000"
|
||||
#define IGNEUM_DAY0 0xceed56d7u
|
||||
#define IGNEUM_DAY1 0x9ba270d2u
|
||||
#define IGNEUM_DATASET_LOG2 28
|
||||
#define IGNEUM_MASK 0x0fffffffu
|
||||
#define IGNEUM_LANES 32
|
||||
#define IGNEUM_ITERATIONS 8
|
||||
#define IGNEUM_INSTR_COUNT 64
|
||||
#define IGNEUM_LOADS_PER_HASH 128
|
||||
#define IGNEUM_WIDE_LOADS_PER_HASH 0
|
||||
#define IGNEUM_OP_MIX "load=16 add=15 shfl=6 mad=4 or=4 rotl=4 rotr=4 xor=4 mulhi=3 mul=2 sub=2"
|
||||
// Program class v3 (Counter ASIC 2.0, docs/plans/counter-asic-2-rollout.md): generator version 3; a worker that
|
||||
// runs another class refuses this pack, and a job line names the class it wants (class=v3 era=<hex>).
|
||||
#define IGNEUM_PROGRAM_CLASS "v3"
|
||||
#define IGNEUM_ERA_SEED_HEX "5e0587f455a86e91e4990f5c481a34cca0044d3f3ae519adcae58ddc82885d31"
|
||||
// Class v3 construction (Counter ASIC 2.0, 5 October 2026, docs/plans/mixer-x4.md): version 2 loads; the dataset item
|
||||
// derivation applies the mixer IGNEUM_MIXER_MULT times per round (memhard.h), and the cache follows the growth rule.
|
||||
#define IGNEUM_LOAD_CLASS "w4-era4488f3ed"
|
||||
#define IGNEUM_LOAD_SLOTS 16
|
||||
#define IGNEUM_LOAD_MIX { 100, 0, 0 }
|
||||
#define IGNEUM_LOAD_WIDTH_COUNTS { 16, 0, 0 } // loads of 4, 16, 64 bytes per program
|
||||
#define IGNEUM_BYTES_PER_HASH 512
|
||||
#define IGNEUM_FOLD_ROT 11
|
||||
#define IGNEUM_FOLD_MUL 0x9e3779b1u
|
||||
// Era layout (5 October 2026, docs/plans/era-layout.md): NOT the lottery hash. Every dataset load reads
|
||||
// idx = ((rotl(src * STRIDE_MUL, STRIDE_ROT) & window mask) | window offset) & MASK; the window of a load site is the
|
||||
// dataset, a half or a quarter of it (IGNEUM_ERA_WINDOWS: site:shrink:offset); dataset word w holds word j(w) of item
|
||||
// t(w) with j's bits at the INTERLEAVE positions (memhard.h: mh_t, mh_j, mh_addr).
|
||||
#define IGNEUM_ERA_LABEL "4488f3ed"
|
||||
#define IGNEUM_ERA_SEED_WORDS { 0x4488f3edu, 0x3cf22d2au, 0xb3e8271eu, 0x55754277u, 0xc2b1c4c7u, 0x8e627302u, 0x584d5acdu, 0xc31dd01du }
|
||||
#define IGNEUM_ERA_ALLOWED_WIDTHS { 1, 0, 0 } // words, ascending, 0 = unused; one entry pins the width
|
||||
#define IGNEUM_ERA_WIDTH_WORDS 1
|
||||
#define IGNEUM_ERA_STRIDE_MUL 0x4d38603du
|
||||
#define IGNEUM_ERA_STRIDE_ROT 10
|
||||
#define IGNEUM_ERA_INTERLEAVE { 2, 9, 13, 15 }
|
||||
#define IGNEUM_ERA_WINDOWS "7:2:1 8:1:1 9:1:1 10:1:1 11:0:0 13:1:1 29:0:0 30:2:2 31:1:1 44:1:1 46:2:0 47:0:0 52:0:0 56:0:0 58:2:0 63:1:1"
|
||||
// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h)
|
||||
#define IGNEUM_DATASET_MODE 1
|
||||
|
||||
#define IGNEUM_SEEDW_INIT { 0x667d0fbdu, 0x7b8e5963u, 0x31c67e5eu, 0x4529ddc6u, 0xef19d6d8u, 0xaccf6211u, 0xda0aed32u, 0xabc6df31u }
|
||||
#define IGNEUM_KEY_INIT { 0xceed56d7u, 0x9ba270d2u, 0x82caab2du, 0x81ebce0eu, 0x12b6ecf1u, 0xd0f3fd7cu, 0xd872eefeu, 0xc158c7bdu }
|
||||
#define IGNEUM_CACHE_LOG2_WORDS 26
|
||||
#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6
|
||||
#define IGNEUM_CACHE_SEGMENTS 65536u
|
||||
#define IGNEUM_ITEM_ROUNDS 8
|
||||
#define IGNEUM_MIX_ROT_INIT { 17u, 12u, 20u, 23u, 7u, 3u, 27u, 16u }
|
||||
#define IGNEUM_MIX_MUL_INIT { 0xf351d601u, 0xa3bb398fu, 0xb5a09e35u, 0x7509c9c1u, 0x6bbf31e9u, 0xfc849a79u, 0xded91851u, 0x8d9113d1u, 0x0ff15225u, 0x3a5bdd41u, 0xab533435u, 0xe1c55ad5u, 0xe6d3bd0du, 0x9d9ffbbdu, 0xbb2a3cf3u, 0x50a7c08du }
|
||||
#define IGNEUM_MIX_RC_INIT { 0xc6892460u, 0x25b7228au, 0xcd515004u, 0x2846527au, 0xa6324241u, 0x36e3ec53u, 0x82961bacu, 0x0f97ba7du, 0xb6f921a9u, 0x3ada24e5u, 0xde20ab91u, 0x5378eeb2u, 0x7d161662u, 0x89353cc1u, 0xb1aa03a2u, 0x788acae6u }
|
||||
|
||||
#ifndef IGNEUM_NO_CUDA
|
||||
// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError().
|
||||
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments);
|
||||
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems);
|
||||
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
uint32_t nonces, uint32_t blockWarps);
|
||||
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
|
||||
#endif
|
||||
143
proto-cuda/packs-ca2-era/era-4/program.json
Normal file
143
proto-cuda/packs-ca2-era/era-4/program.json
Normal file
|
|
@ -0,0 +1,143 @@
|
|||
{
|
||||
"format": "igneum-program-pack-3",
|
||||
"generator": 3,
|
||||
"attempt": 0,
|
||||
"program_id": "0x73bcbfe8ccf988f1",
|
||||
"program_id_derivation": "FNV-1a 64 over 'igneum-program/' || generator_le32 || seed_words as little-endian bytes || attempt_le32",
|
||||
"dataset_mode": "memory-hard",
|
||||
"seed": "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000",
|
||||
"seed_bytes": "edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07",
|
||||
"seed_words": ["0x667d0fbd", "0x7b8e5963", "0x31c67e5e", "0x4529ddc6", "0xef19d6d8", "0xaccf6211", "0xda0aed32", "0xabc6df31"],
|
||||
"seed_derivation": "seed_words = FNV-1a 64 over seed_bytes (attempt 0) or seed_bytes || attempt_le32 (attempt k >= 1), basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32",
|
||||
"generator_rule": "version 2: exactly 16 load slots drawn first from instructions 1..63 (partial Fisher-Yates), the other 48 ops from the ten non-load weights (sum 75); a load's source is drawn from the registers other than dst written by an earlier instruction and not read by a load since; the candidate must pass the acceptance rule of spec 01 section 1.4.6 (static: no cyclically stale load source, every register has an injecting write; dynamic: 64 units on the seed-keyed closed-form dataset with no constant register bit, no lane-constant load site, under 164 saturated final values, every output bit within 136 of 1024, distinct addresses above 245760), else the next attempt of the seed is tried",
|
||||
"lanes": 32,
|
||||
"registers": 8,
|
||||
"iterations": 8,
|
||||
"instruction_count": 64,
|
||||
"loads_per_hash": 128,
|
||||
"program_class": "v3",
|
||||
"era_seed_bytes": "5e0587f455a86e91e4990f5c481a34cca0044d3f3ae519adcae58ddc82885d31",
|
||||
"load_class": "w4-era4488f3ed",
|
||||
"load_slots": 16,
|
||||
"load_mix_percent_4_16_64": [100, 0, 0],
|
||||
"load_width_counts_4_16_64": [16, 0, 0],
|
||||
"bytes_per_hash": 512,
|
||||
"wide_load": "read-width experiment (5 October 2026, docs/plans/read-width.md), NOT the lottery hash: a load of W words (width field, 4 or 16) reads dataset[b .. b + W) with b = (src & mask) & ~(W - 1) and folds every word into dst: x = dst ^ w[0]; for j in 1..W: x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x; width 1 is the plain load; the width is drawn per instruction from the class mix with one extra below(100) draw after the nine of version 2, and the program id is FNV-1a 64 over 'igneum-program-rw/' || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots",
|
||||
"era": {
|
||||
"label": "4488f3ed",
|
||||
"seed_words": ["0x4488f3ed", "0x3cf22d2a", "0xb3e8271e", "0x55754277", "0xc2b1c4c7", "0x8e627302", "0x584d5acd", "0xc31dd01d"],
|
||||
"draw": "docs/plans/era-layout.md 1.1: SplitMix64 seeded with seed_words[0] | seed_words[1] << 32 of seed_words_from_bytes('igneum-era/' || n_le64 || E_n); width = allowed[below(|allowed|)], stride_mul = low32(next()) | 1, stride_rot = 1 + below(31), then four next() draws for a partial Fisher-Yates over positions log2(W)..15 of which 4 - log2(W) are used",
|
||||
"allowed_widths": [1],
|
||||
"width_words": 1,
|
||||
"stride_mul": "0x4d38603d",
|
||||
"stride_rot": 10,
|
||||
"interleave": [2, 9, 13, 15],
|
||||
"address": "y = rotl(src * stride_mul, stride_rot); k = min(win, D - 26); idx = ((y & (mask >> k)) | ((off & (2^k - 1)) << (D - k))) & mask; a wide load aligns idx down to W words",
|
||||
"windows": "per instruction, after the width roll: win = below(3), off = low32(next()) & (2^win - 1); used on a load slot (the instruction's win and off fields)",
|
||||
"dataset_word": "dataset[w] = item(t(w))[j(w)]: j(w) gathers the bits of w at the interleave positions, t(w) is w with those bits removed",
|
||||
"program_id_suffix": "'era/' || allowed[3] || width_words || stride_mul_le32 || stride_rot_le32 || interleave[4]"
|
||||
},
|
||||
"op_mix": {"load": 16, "add": 15, "shfl": 6, "mad": 4, "or": 4, "rotl": 4, "rotr": 4, "xor": 4, "mulhi": 3, "mul": 2, "sub": 2},
|
||||
"register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]",
|
||||
"splitmix32": "x ^= x>>16; x *= 0x7feb352d; x ^= x>>15; x *= 0x846ca68b; x ^= x>>16",
|
||||
"iteration": "sel = r0 sampled once at the top of each iteration, then all instructions in order",
|
||||
"output": "lo = r0 ^ rotl(r1,7) ^ rotl(r2,14) ^ rotl(r3,21); hi = r4 ^ rotl(r5,9) ^ rotl(r6,18) ^ rotl(r7,27); out = (hi << 32) | lo",
|
||||
"op_semantics": {
|
||||
"add": "dst = dst + src + (bit `bit` of sel ? imm2 : imm)",
|
||||
"sub": "dst = dst - src",
|
||||
"mul": "dst = dst * src (low 32)",
|
||||
"mulhi": "dst = high 32 bits of dst * src",
|
||||
"xor": "dst = dst ^ src",
|
||||
"or": "dst = dst | src",
|
||||
"rotl": "dst = rotl(dst, rot), rot in 1..31",
|
||||
"rotr": "dst = rotr(dst, src & 31)",
|
||||
"mad": "dst = src * src2 + dst",
|
||||
"shfl": "dst = dst ^ (src of lane (lane ^ mask)), mask in {1,2,4,8,16}, within the 32-lane warp",
|
||||
"load": "dst = dst ^ dataset[src & dataset.mask]",
|
||||
"wload": "base = (src of lane 0 & dataset.mask) & ~31; dst = dst ^ dataset[base + lane] (warp-coalesced 128-byte load, lever b, only when --wide-frac > 0)"
|
||||
},
|
||||
"dataset": {
|
||||
"log2_words": 28,
|
||||
"bytes": 1073741824,
|
||||
"mask": "0x0fffffff",
|
||||
"day": "bytes:69676e65756d2d6461792ffa50000000000000",
|
||||
"day_bytes": "69676e65756d2d6461792ffa50000000000000",
|
||||
"day_words_from": "seed_words_from_bytes(day_bytes)",
|
||||
"d0": "0xceed56d7",
|
||||
"d1": "0x9ba270d2",
|
||||
"mode": "memory-hard",
|
||||
"spec": "proto-metal/MEMHARD.md",
|
||||
"key": ["0xceed56d7", "0x9ba270d2", "0x82caab2d", "0x81ebce0e", "0x12b6ecf1", "0xd0f3fd7c", "0xd872eefe", "0xc158c7bd"],
|
||||
"key_derivation": "the 8 words of seed_words_from_bytes(day_bytes); d0, d1 are key[0], key[1]",
|
||||
"cache": {"log2_words": 26, "bytes": 268435456, "line_words": 16, "segment_lines": 64, "segments": 65536, "block": "ChaCha12 core + feed-forward, rotations 16 12 8 7", "sigma": ["0x61707865", "0x3320646e", "0x79622d32", "0x6b206574"], "tag": ["0x49676e65", "0x756d4d48"], "chain": "in_j = prev_line ^ (sigma[0..3] || key[0..7] || seg || j || tag[0..1]); line_j = block(in_j); prev_0 = 0"},
|
||||
"mixer": {"draw": "SplitMix64 seeded with key[0] | key[1] << 32: rot[0..7] = 1 + next() % 31, mul[0..15] = low32(next()) | 1, rc[0..15] = low32(next())", "rot": [17, 12, 20, 23, 7, 3, 27, 16], "mul": ["0xf351d601", "0xa3bb398f", "0xb5a09e35", "0x7509c9c1", "0x6bbf31e9", "0xfc849a79", "0xded91851", "0x8d9113d1", "0x0ff15225", "0x3a5bdd41", "0xab533435", "0xe1c55ad5", "0xe6d3bd0d", "0x9d9ffbbd", "0xbb2a3cf3", "0x50a7c08d"], "rc": ["0xc6892460", "0x25b7228a", "0xcd515004", "0x2846527a", "0xa6324241", "0x36e3ec53", "0x82961bac", "0x0f97ba7d", "0xb6f921a9", "0x3ada24e5", "0xde20ab91", "0x5378eeb2", "0x7d161662", "0x89353cc1", "0xb1aa03a2", "0x788acae6"], "round": "for i in 0..15: s[i] = (s[i] ^ (rc[i] + (r+1) * 0x9E3779B9)) * mul[i]; then quarter rounds on columns (0,4,8,12) (1,5,9,13) (2,6,10,14) (3,7,11,15) with rot[0..3] and diagonals (0,5,10,15) (1,6,11,12) (2,7,8,13) (3,4,9,14) with rot[4..7]", "quarter_round": "a += b; d ^= a; d = rotl(d, r1); c += d; b ^= c; b = rotl(b, r2); a += b; d ^= a; d = rotl(d, r3); c += d; b ^= c; b = rotl(b, r4)"},
|
||||
"item": "s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..7: s = M_r(s); line = s[0] & 0x003fffff; s[i] ^= cache[line * 16 + i]; then s = M_8(s); item(t) = s",
|
||||
"word": "dataset[w] = item(w >> 4)[w & 15]"
|
||||
},
|
||||
"instructions": [
|
||||
{"i": 0, "op": "add", "dst": 4, "src": 5, "src2": 7, "imm": "0xea86e152", "imm2": "0x5810667a", "rot": 27, "bit": 13, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 1, "op": "xor", "dst": 7, "src": 0, "src2": 6, "imm": "0xbaab6229", "imm2": "0xed861989", "rot": 26, "bit": 22, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 2, "op": "shfl", "dst": 3, "src": 6, "src2": 2, "imm": "0x5b623116", "imm2": "0xff12e5b2", "rot": 12, "bit": 24, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 3, "op": "sub", "dst": 4, "src": 1, "src2": 1, "imm": "0xe99741c7", "imm2": "0xf5fa5009", "rot": 1, "bit": 21, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 4, "op": "mad", "dst": 2, "src": 0, "src2": 4, "imm": "0x673c2157", "imm2": "0xee02465f", "rot": 20, "bit": 22, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 5, "op": "rotl", "dst": 4, "src": 7, "src2": 7, "imm": "0x946f7818", "imm2": "0x45d3399e", "rot": 9, "bit": 2, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 6, "op": "rotr", "dst": 0, "src": 2, "src2": 4, "imm": "0x5f6a0ed2", "imm2": "0x7043a636", "rot": 19, "bit": 31, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 7, "op": "load", "dst": 6, "src": 7, "src2": 1, "imm": "0x5c61dcf7", "imm2": "0x7466aa40", "rot": 19, "bit": 9, "mask": 2, "width": 1, "win": 2, "off": 1},
|
||||
{"i": 8, "op": "load", "dst": 1, "src": 4, "src2": 5, "imm": "0x85668475", "imm2": "0xdb8cc483", "rot": 29, "bit": 7, "mask": 4, "width": 1, "win": 1, "off": 1},
|
||||
{"i": 9, "op": "load", "dst": 1, "src": 2, "src2": 4, "imm": "0x4454980f", "imm2": "0xebd31581", "rot": 10, "bit": 28, "mask": 16, "width": 1, "win": 1, "off": 1},
|
||||
{"i": 10, "op": "load", "dst": 7, "src": 0, "src2": 2, "imm": "0xe075297c", "imm2": "0x5779c44c", "rot": 10, "bit": 22, "mask": 2, "width": 1, "win": 1, "off": 1},
|
||||
{"i": 11, "op": "load", "dst": 7, "src": 1, "src2": 6, "imm": "0x65aa4311", "imm2": "0x4fe48ea9", "rot": 15, "bit": 9, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 12, "op": "mul", "dst": 5, "src": 4, "src2": 2, "imm": "0x1383d3ad", "imm2": "0xf3094b29", "rot": 8, "bit": 9, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 13, "op": "load", "dst": 7, "src": 6, "src2": 2, "imm": "0xed8a496f", "imm2": "0x3072c3c6", "rot": 28, "bit": 19, "mask": 8, "width": 1, "win": 1, "off": 1},
|
||||
{"i": 14, "op": "shfl", "dst": 6, "src": 5, "src2": 0, "imm": "0x8b965b57", "imm2": "0xcfeca6c1", "rot": 12, "bit": 27, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 15, "op": "shfl", "dst": 3, "src": 5, "src2": 3, "imm": "0x877c7586", "imm2": "0xa9cb2a03", "rot": 2, "bit": 29, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 16, "op": "or", "dst": 2, "src": 7, "src2": 3, "imm": "0xb740221a", "imm2": "0x89d38d6d", "rot": 6, "bit": 31, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 17, "op": "rotr", "dst": 0, "src": 6, "src2": 1, "imm": "0x26f3ad8a", "imm2": "0x27256f15", "rot": 18, "bit": 5, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 18, "op": "add", "dst": 7, "src": 5, "src2": 7, "imm": "0xf66e7017", "imm2": "0xb9e3577e", "rot": 9, "bit": 12, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 19, "op": "rotl", "dst": 7, "src": 0, "src2": 5, "imm": "0x849ae6ee", "imm2": "0x02b358f9", "rot": 24, "bit": 18, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 20, "op": "add", "dst": 6, "src": 7, "src2": 6, "imm": "0x52334d12", "imm2": "0x8c9f0ef8", "rot": 11, "bit": 23, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 21, "op": "or", "dst": 2, "src": 6, "src2": 5, "imm": "0xb1871e63", "imm2": "0xb2e40191", "rot": 5, "bit": 13, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 22, "op": "mad", "dst": 6, "src": 5, "src2": 4, "imm": "0x97df29e4", "imm2": "0xe60fea84", "rot": 11, "bit": 16, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 23, "op": "or", "dst": 1, "src": 0, "src2": 3, "imm": "0x8f30d21d", "imm2": "0x2df685a0", "rot": 31, "bit": 0, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 24, "op": "xor", "dst": 6, "src": 1, "src2": 2, "imm": "0xd2c4025f", "imm2": "0x5269eb4d", "rot": 31, "bit": 21, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 25, "op": "add", "dst": 2, "src": 6, "src2": 5, "imm": "0x659fc3d3", "imm2": "0x9cec0e12", "rot": 6, "bit": 17, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 26, "op": "rotr", "dst": 7, "src": 0, "src2": 1, "imm": "0xd89ef484", "imm2": "0x20be3846", "rot": 12, "bit": 9, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 27, "op": "mad", "dst": 4, "src": 5, "src2": 7, "imm": "0xb48420ae", "imm2": "0x3d1f2485", "rot": 14, "bit": 28, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 28, "op": "mad", "dst": 3, "src": 2, "src2": 4, "imm": "0xc5c46d76", "imm2": "0x700044b5", "rot": 22, "bit": 10, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 29, "op": "load", "dst": 1, "src": 4, "src2": 3, "imm": "0x0fbaf177", "imm2": "0xfff4f2ed", "rot": 20, "bit": 31, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 30, "op": "load", "dst": 2, "src": 3, "src2": 0, "imm": "0xe90eb2e5", "imm2": "0xb0f9eb79", "rot": 27, "bit": 14, "mask": 4, "width": 1, "win": 2, "off": 2},
|
||||
{"i": 31, "op": "load", "dst": 1, "src": 5, "src2": 2, "imm": "0x97ba3fc3", "imm2": "0x7894e657", "rot": 3, "bit": 30, "mask": 16, "width": 1, "win": 1, "off": 1},
|
||||
{"i": 32, "op": "add", "dst": 7, "src": 2, "src2": 7, "imm": "0x070888a8", "imm2": "0xe403240e", "rot": 2, "bit": 16, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 33, "op": "add", "dst": 2, "src": 0, "src2": 5, "imm": "0xf2e46d55", "imm2": "0x29701828", "rot": 31, "bit": 28, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 34, "op": "add", "dst": 2, "src": 3, "src2": 0, "imm": "0x58f75b87", "imm2": "0x343b7aee", "rot": 12, "bit": 14, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 35, "op": "mulhi", "dst": 2, "src": 5, "src2": 6, "imm": "0x5892a9e6", "imm2": "0xc9824c94", "rot": 19, "bit": 26, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 36, "op": "xor", "dst": 4, "src": 2, "src2": 3, "imm": "0x7b5b5474", "imm2": "0x45cfc5dd", "rot": 17, "bit": 18, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 37, "op": "mul", "dst": 6, "src": 5, "src2": 7, "imm": "0xccf564a5", "imm2": "0x873ad101", "rot": 7, "bit": 11, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 38, "op": "xor", "dst": 7, "src": 0, "src2": 6, "imm": "0xcac8f06d", "imm2": "0x6b97c683", "rot": 18, "bit": 28, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 39, "op": "add", "dst": 7, "src": 2, "src2": 4, "imm": "0xb8180e9d", "imm2": "0x32bbd117", "rot": 23, "bit": 19, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 40, "op": "rotr", "dst": 2, "src": 3, "src2": 0, "imm": "0x2d6070bc", "imm2": "0x68ff101e", "rot": 13, "bit": 18, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 41, "op": "sub", "dst": 7, "src": 0, "src2": 2, "imm": "0x0daf96ea", "imm2": "0x36f37be1", "rot": 5, "bit": 0, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 42, "op": "add", "dst": 4, "src": 3, "src2": 6, "imm": "0x6ced15b7", "imm2": "0x6df7aed4", "rot": 19, "bit": 4, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 43, "op": "shfl", "dst": 7, "src": 3, "src2": 5, "imm": "0x8ace05f3", "imm2": "0xd378ec12", "rot": 23, "bit": 10, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 44, "op": "load", "dst": 0, "src": 7, "src2": 4, "imm": "0xb0607786", "imm2": "0xc4acabbc", "rot": 13, "bit": 7, "mask": 16, "width": 1, "win": 1, "off": 1},
|
||||
{"i": 45, "op": "add", "dst": 1, "src": 6, "src2": 5, "imm": "0xe09f54e9", "imm2": "0x83e825bf", "rot": 23, "bit": 14, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 46, "op": "load", "dst": 3, "src": 1, "src2": 0, "imm": "0x63cc1e4e", "imm2": "0xa1be8118", "rot": 12, "bit": 6, "mask": 2, "width": 1, "win": 2, "off": 0},
|
||||
{"i": 47, "op": "load", "dst": 6, "src": 3, "src2": 7, "imm": "0x353f1d79", "imm2": "0x3b2e7456", "rot": 18, "bit": 18, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 48, "op": "mulhi", "dst": 4, "src": 2, "src2": 7, "imm": "0x00d8a3cd", "imm2": "0x231866d2", "rot": 21, "bit": 20, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 49, "op": "add", "dst": 5, "src": 0, "src2": 2, "imm": "0xa8bae6df", "imm2": "0xf572bdb9", "rot": 14, "bit": 7, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 50, "op": "shfl", "dst": 0, "src": 7, "src2": 7, "imm": "0x81ef22e1", "imm2": "0x74438fc5", "rot": 28, "bit": 18, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 51, "op": "add", "dst": 6, "src": 0, "src2": 6, "imm": "0x383b9260", "imm2": "0x11e17c61", "rot": 12, "bit": 19, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 52, "op": "load", "dst": 5, "src": 2, "src2": 2, "imm": "0xfb84f451", "imm2": "0x11cd863e", "rot": 21, "bit": 20, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 53, "op": "rotl", "dst": 6, "src": 5, "src2": 4, "imm": "0xb1a7db6b", "imm2": "0x76686b9b", "rot": 12, "bit": 4, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 54, "op": "rotl", "dst": 3, "src": 6, "src2": 3, "imm": "0x6f981f52", "imm2": "0xd99aeba2", "rot": 12, "bit": 27, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 55, "op": "add", "dst": 2, "src": 1, "src2": 2, "imm": "0xac6be8e3", "imm2": "0x18d67dbb", "rot": 26, "bit": 10, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 56, "op": "load", "dst": 1, "src": 4, "src2": 0, "imm": "0x7e7f6a00", "imm2": "0x6f0747da", "rot": 25, "bit": 28, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 57, "op": "add", "dst": 5, "src": 0, "src2": 4, "imm": "0xf03673fe", "imm2": "0xa75cd60d", "rot": 16, "bit": 12, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 58, "op": "load", "dst": 5, "src": 0, "src2": 2, "imm": "0x227f94a6", "imm2": "0x0e8344f9", "rot": 20, "bit": 10, "mask": 2, "width": 1, "win": 2, "off": 0},
|
||||
{"i": 59, "op": "add", "dst": 1, "src": 4, "src2": 3, "imm": "0xdecd4794", "imm2": "0x8dfb96bb", "rot": 21, "bit": 7, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 60, "op": "mulhi", "dst": 3, "src": 2, "src2": 2, "imm": "0x0dd268e0", "imm2": "0x53034ca9", "rot": 1, "bit": 8, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 61, "op": "or", "dst": 6, "src": 4, "src2": 7, "imm": "0x3a45a321", "imm2": "0x9bc59a5f", "rot": 25, "bit": 11, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 62, "op": "shfl", "dst": 5, "src": 4, "src2": 2, "imm": "0x8f229cc1", "imm2": "0xcaac64a2", "rot": 17, "bit": 13, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 63, "op": "load", "dst": 3, "src": 6, "src2": 6, "imm": "0xb2574178", "imm2": "0xcbcc798d", "rot": 28, "bit": 0, "mask": 16, "width": 1, "win": 1, "off": 1}
|
||||
]
|
||||
}
|
||||
109
proto-cuda/packs-ca2-era/era-4/program.metal
Normal file
109
proto-cuda/packs-ca2-era/era-4/program.metal
Normal file
|
|
@ -0,0 +1,109 @@
|
|||
#include <metal_stdlib>
|
||||
using namespace metal;
|
||||
|
||||
#define MASK 0x0fffffffu
|
||||
constant uint SEEDW[8] = { 0x667d0fbdu, 0x7b8e5963u, 0x31c67e5eu, 0x4529ddc6u, 0xef19d6d8u, 0xaccf6211u, 0xda0aed32u, 0xabc6df31u };
|
||||
|
||||
inline uint splitmix32(uint x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31
|
||||
inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
inline uint ds_elem(uint i, uint d0, uint d1) {
|
||||
uint x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
kernel void igneum_hash(device const uint* dataset [[buffer(0)]],
|
||||
device ulong* out [[buffer(1)]],
|
||||
constant uint& baseNonce [[buffer(2)]],
|
||||
uint gid [[thread_position_in_grid]]) {
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint x = nonce ^ SEEDW[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ SEEDW[1]; }
|
||||
{ uint x = nonce ^ SEEDW[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ SEEDW[2]; }
|
||||
{ uint x = nonce ^ SEEDW[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ SEEDW[3]; }
|
||||
{ uint x = nonce ^ SEEDW[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ SEEDW[4]; }
|
||||
{ uint x = nonce ^ SEEDW[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ SEEDW[5]; }
|
||||
{ uint x = nonce ^ SEEDW[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ SEEDW[6]; }
|
||||
{ uint x = nonce ^ SEEDW[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ SEEDW[7]; }
|
||||
{ uint x = nonce ^ SEEDW[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ SEEDW[0]; }
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = r4 + r5 + select(0xea86e152u, 0x5810667au, ((sel >> 13u) & 1u) != 0u); // 0
|
||||
r7 = r7 ^ r0; // 1
|
||||
r3 = r3 ^ simd_shuffle_xor(r6, (ushort)1); // 2
|
||||
r4 = r4 - r1; // 3
|
||||
r2 = r0 * r4 + r2; // 4
|
||||
r4 = rotl_imm(r4, 9u); // 5
|
||||
r0 = rotr_var(r0, r2); // 6
|
||||
r6 = r6 ^ dataset[((rotl_imm(r7 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x04000000u) & MASK]; // 7
|
||||
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 8
|
||||
r1 = r1 ^ dataset[((rotl_imm(r2 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 9
|
||||
r7 = r7 ^ dataset[((rotl_imm(r0 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 10
|
||||
r7 = r7 ^ dataset[((rotl_imm(r1 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 11
|
||||
r5 = r5 * r4; // 12
|
||||
r7 = r7 ^ dataset[((rotl_imm(r6 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 13
|
||||
r6 = r6 ^ simd_shuffle_xor(r5, (ushort)16); // 14
|
||||
r3 = r3 ^ simd_shuffle_xor(r5, (ushort)1); // 15
|
||||
r2 = r2 | r7; // 16
|
||||
r0 = rotr_var(r0, r6); // 17
|
||||
r7 = r7 + r5 + select(0xf66e7017u, 0xb9e3577eu, ((sel >> 12u) & 1u) != 0u); // 18
|
||||
r7 = rotl_imm(r7, 24u); // 19
|
||||
r6 = r6 + r7 + select(0x52334d12u, 0x8c9f0ef8u, ((sel >> 23u) & 1u) != 0u); // 20
|
||||
r2 = r2 | r6; // 21
|
||||
r6 = r5 * r4 + r6; // 22
|
||||
r1 = r1 | r0; // 23
|
||||
r6 = r6 ^ r1; // 24
|
||||
r2 = r2 + r6 + select(0x659fc3d3u, 0x9cec0e12u, ((sel >> 17u) & 1u) != 0u); // 25
|
||||
r7 = rotr_var(r7, r0); // 26
|
||||
r4 = r5 * r7 + r4; // 27
|
||||
r3 = r2 * r4 + r3; // 28
|
||||
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 29
|
||||
r2 = r2 ^ dataset[((rotl_imm(r3 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x08000000u) & MASK]; // 30
|
||||
r1 = r1 ^ dataset[((rotl_imm(r5 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 31
|
||||
r7 = r7 + r2 + select(0x070888a8u, 0xe403240eu, ((sel >> 16u) & 1u) != 0u); // 32
|
||||
r2 = r2 + r0 + select(0xf2e46d55u, 0x29701828u, ((sel >> 28u) & 1u) != 0u); // 33
|
||||
r2 = r2 + r3 + select(0x58f75b87u, 0x343b7aeeu, ((sel >> 14u) & 1u) != 0u); // 34
|
||||
r2 = mulhi(r2, r5); // 35
|
||||
r4 = r4 ^ r2; // 36
|
||||
r6 = r6 * r5; // 37
|
||||
r7 = r7 ^ r0; // 38
|
||||
r7 = r7 + r2 + select(0xb8180e9du, 0x32bbd117u, ((sel >> 19u) & 1u) != 0u); // 39
|
||||
r2 = rotr_var(r2, r3); // 40
|
||||
r7 = r7 - r0; // 41
|
||||
r4 = r4 + r3 + select(0x6ced15b7u, 0x6df7aed4u, ((sel >> 4u) & 1u) != 0u); // 42
|
||||
r7 = r7 ^ simd_shuffle_xor(r3, (ushort)4); // 43
|
||||
r0 = r0 ^ dataset[((rotl_imm(r7 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 44
|
||||
r1 = r1 + r6 + select(0xe09f54e9u, 0x83e825bfu, ((sel >> 14u) & 1u) != 0u); // 45
|
||||
r3 = r3 ^ dataset[((rotl_imm(r1 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 46
|
||||
r6 = r6 ^ dataset[((rotl_imm(r3 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 47
|
||||
r4 = mulhi(r4, r2); // 48
|
||||
r5 = r5 + r0 + select(0xa8bae6dfu, 0xf572bdb9u, ((sel >> 7u) & 1u) != 0u); // 49
|
||||
r0 = r0 ^ simd_shuffle_xor(r7, (ushort)4); // 50
|
||||
r6 = r6 + r0 + select(0x383b9260u, 0x11e17c61u, ((sel >> 19u) & 1u) != 0u); // 51
|
||||
r5 = r5 ^ dataset[((rotl_imm(r2 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 52
|
||||
r6 = rotl_imm(r6, 12u); // 53
|
||||
r3 = rotl_imm(r3, 12u); // 54
|
||||
r2 = r2 + r1 + select(0xac6be8e3u, 0x18d67dbbu, ((sel >> 10u) & 1u) != 0u); // 55
|
||||
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 56
|
||||
r5 = r5 + r0 + select(0xf03673feu, 0xa75cd60du, ((sel >> 12u) & 1u) != 0u); // 57
|
||||
r5 = r5 ^ dataset[((rotl_imm(r0 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 58
|
||||
r1 = r1 + r4 + select(0xdecd4794u, 0x8dfb96bbu, ((sel >> 7u) & 1u) != 0u); // 59
|
||||
r3 = mulhi(r3, r2); // 60
|
||||
r6 = r6 | r4; // 61
|
||||
r5 = r5 ^ simd_shuffle_xor(r4, (ushort)8); // 62
|
||||
r3 = r3 ^ dataset[((rotl_imm(r6 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 63
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
111
proto-cuda/packs-ca2-era/era-4/program_bound.metal
Normal file
111
proto-cuda/packs-ca2-era/era-4/program_bound.metal
Normal file
|
|
@ -0,0 +1,111 @@
|
|||
#include <metal_stdlib>
|
||||
using namespace metal;
|
||||
|
||||
#define MASK 0x0fffffffu
|
||||
constant uint SEEDW[8] = { 0x667d0fbdu, 0x7b8e5963u, 0x31c67e5eu, 0x4529ddc6u, 0xef19d6d8u, 0xaccf6211u, 0xda0aed32u, 0xabc6df31u };
|
||||
|
||||
inline uint splitmix32(uint x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31
|
||||
inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
inline uint ds_elem(uint i, uint d0, uint d1) {
|
||||
uint x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// Header-bound variant: the init words come from buffer 3 (bind.rs), not from SEEDW.
|
||||
kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]],
|
||||
device ulong* out [[buffer(1)]],
|
||||
constant uint& baseNonce [[buffer(2)]],
|
||||
constant uint* initw [[buffer(3)]],
|
||||
uint gid [[thread_position_in_grid]]) {
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint x = nonce ^ initw[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ initw[1]; }
|
||||
{ uint x = nonce ^ initw[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ initw[2]; }
|
||||
{ uint x = nonce ^ initw[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ initw[3]; }
|
||||
{ uint x = nonce ^ initw[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ initw[4]; }
|
||||
{ uint x = nonce ^ initw[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ initw[5]; }
|
||||
{ uint x = nonce ^ initw[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ initw[6]; }
|
||||
{ uint x = nonce ^ initw[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ initw[7]; }
|
||||
{ uint x = nonce ^ initw[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ initw[0]; }
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = r4 + r5 + select(0xea86e152u, 0x5810667au, ((sel >> 13u) & 1u) != 0u); // 0
|
||||
r7 = r7 ^ r0; // 1
|
||||
r3 = r3 ^ simd_shuffle_xor(r6, (ushort)1); // 2
|
||||
r4 = r4 - r1; // 3
|
||||
r2 = r0 * r4 + r2; // 4
|
||||
r4 = rotl_imm(r4, 9u); // 5
|
||||
r0 = rotr_var(r0, r2); // 6
|
||||
r6 = r6 ^ dataset[((rotl_imm(r7 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x04000000u) & MASK]; // 7
|
||||
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 8
|
||||
r1 = r1 ^ dataset[((rotl_imm(r2 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 9
|
||||
r7 = r7 ^ dataset[((rotl_imm(r0 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 10
|
||||
r7 = r7 ^ dataset[((rotl_imm(r1 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 11
|
||||
r5 = r5 * r4; // 12
|
||||
r7 = r7 ^ dataset[((rotl_imm(r6 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 13
|
||||
r6 = r6 ^ simd_shuffle_xor(r5, (ushort)16); // 14
|
||||
r3 = r3 ^ simd_shuffle_xor(r5, (ushort)1); // 15
|
||||
r2 = r2 | r7; // 16
|
||||
r0 = rotr_var(r0, r6); // 17
|
||||
r7 = r7 + r5 + select(0xf66e7017u, 0xb9e3577eu, ((sel >> 12u) & 1u) != 0u); // 18
|
||||
r7 = rotl_imm(r7, 24u); // 19
|
||||
r6 = r6 + r7 + select(0x52334d12u, 0x8c9f0ef8u, ((sel >> 23u) & 1u) != 0u); // 20
|
||||
r2 = r2 | r6; // 21
|
||||
r6 = r5 * r4 + r6; // 22
|
||||
r1 = r1 | r0; // 23
|
||||
r6 = r6 ^ r1; // 24
|
||||
r2 = r2 + r6 + select(0x659fc3d3u, 0x9cec0e12u, ((sel >> 17u) & 1u) != 0u); // 25
|
||||
r7 = rotr_var(r7, r0); // 26
|
||||
r4 = r5 * r7 + r4; // 27
|
||||
r3 = r2 * r4 + r3; // 28
|
||||
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 29
|
||||
r2 = r2 ^ dataset[((rotl_imm(r3 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x08000000u) & MASK]; // 30
|
||||
r1 = r1 ^ dataset[((rotl_imm(r5 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 31
|
||||
r7 = r7 + r2 + select(0x070888a8u, 0xe403240eu, ((sel >> 16u) & 1u) != 0u); // 32
|
||||
r2 = r2 + r0 + select(0xf2e46d55u, 0x29701828u, ((sel >> 28u) & 1u) != 0u); // 33
|
||||
r2 = r2 + r3 + select(0x58f75b87u, 0x343b7aeeu, ((sel >> 14u) & 1u) != 0u); // 34
|
||||
r2 = mulhi(r2, r5); // 35
|
||||
r4 = r4 ^ r2; // 36
|
||||
r6 = r6 * r5; // 37
|
||||
r7 = r7 ^ r0; // 38
|
||||
r7 = r7 + r2 + select(0xb8180e9du, 0x32bbd117u, ((sel >> 19u) & 1u) != 0u); // 39
|
||||
r2 = rotr_var(r2, r3); // 40
|
||||
r7 = r7 - r0; // 41
|
||||
r4 = r4 + r3 + select(0x6ced15b7u, 0x6df7aed4u, ((sel >> 4u) & 1u) != 0u); // 42
|
||||
r7 = r7 ^ simd_shuffle_xor(r3, (ushort)4); // 43
|
||||
r0 = r0 ^ dataset[((rotl_imm(r7 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 44
|
||||
r1 = r1 + r6 + select(0xe09f54e9u, 0x83e825bfu, ((sel >> 14u) & 1u) != 0u); // 45
|
||||
r3 = r3 ^ dataset[((rotl_imm(r1 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 46
|
||||
r6 = r6 ^ dataset[((rotl_imm(r3 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 47
|
||||
r4 = mulhi(r4, r2); // 48
|
||||
r5 = r5 + r0 + select(0xa8bae6dfu, 0xf572bdb9u, ((sel >> 7u) & 1u) != 0u); // 49
|
||||
r0 = r0 ^ simd_shuffle_xor(r7, (ushort)4); // 50
|
||||
r6 = r6 + r0 + select(0x383b9260u, 0x11e17c61u, ((sel >> 19u) & 1u) != 0u); // 51
|
||||
r5 = r5 ^ dataset[((rotl_imm(r2 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 52
|
||||
r6 = rotl_imm(r6, 12u); // 53
|
||||
r3 = rotl_imm(r3, 12u); // 54
|
||||
r2 = r2 + r1 + select(0xac6be8e3u, 0x18d67dbbu, ((sel >> 10u) & 1u) != 0u); // 55
|
||||
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x4d38603du, 10u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 56
|
||||
r5 = r5 + r0 + select(0xf03673feu, 0xa75cd60du, ((sel >> 12u) & 1u) != 0u); // 57
|
||||
r5 = r5 ^ dataset[((rotl_imm(r0 * 0x4d38603du, 10u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 58
|
||||
r1 = r1 + r4 + select(0xdecd4794u, 0x8dfb96bbu, ((sel >> 7u) & 1u) != 0u); // 59
|
||||
r3 = mulhi(r3, r2); // 60
|
||||
r6 = r6 | r4; // 61
|
||||
r5 = r5 ^ simd_shuffle_xor(r4, (ushort)8); // 62
|
||||
r3 = r3 ^ dataset[((rotl_imm(r6 * 0x4d38603du, 10u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 63
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
5
proto-cuda/packs-ca2-era/era-4/seeds.txt
Normal file
5
proto-cuda/packs-ca2-era/era-4/seeds.txt
Normal file
|
|
@ -0,0 +1,5 @@
|
|||
epoch_seed_hex edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07
|
||||
day_seed_hex 69676e65756d2d6461792ffa50000000000000
|
||||
epoch_index 0
|
||||
day_index 20730
|
||||
daa_score 0
|
||||
57
proto-cuda/packs-ca2-era/era-4/vectors.h
Normal file
57
proto-cuda/packs-ca2-era/era-4/vectors.h
Normal file
|
|
@ -0,0 +1,57 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v3, memory-hard dataset
|
||||
#pragma once
|
||||
#ifdef __cplusplus
|
||||
#include <cstdint>
|
||||
#else
|
||||
#include <stdint.h>
|
||||
#endif
|
||||
|
||||
#define IGNEUM_VEC_WARPS 3
|
||||
static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u };
|
||||
static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = {
|
||||
{ // base nonce 0
|
||||
0x8e32500ee72aa1f8ull, 0x710f1522fd52e28eull, 0x62fd41712fd5bd3eull, 0x3837ee0ed93c4015ull, 0x0338e814e0e8a4faull, 0xaab705ee7514af71ull, 0x4a3e143121ed41a6ull, 0xe843b01bee434ef6ull,
|
||||
0x4f9974ee1f331bc6ull, 0x1e089106d4d04581ull, 0x8a986409754a08b8ull, 0x54bf0b893560e130ull, 0x1346ddc8074de7ddull, 0x5fdc54666289e864ull, 0xe9ed965ed22a9ea1ull, 0x7915337fc840db87ull,
|
||||
0xaab900aea6b7c2cfull, 0x81d768fa492b1cf2ull, 0xfd62bcf1021629dfull, 0x6280dc76e5125eddull, 0x17055f35fc9a4ed3ull, 0xead760de60ae3e84ull, 0x14ea3ac017d95d57ull, 0x5f09e499a0096d83ull,
|
||||
0xdf7c24624de53e02ull, 0x10e39b014ea6d6c6ull, 0xfcfac36ecf54fb8cull, 0x8b510ae1be613404ull, 0xa9a07c9f9fa58801ull, 0xba88c22112b80fd0ull, 0x673c5bf0fbb48d35ull, 0xeef7f536b1e0a810ull
|
||||
},
|
||||
{ // base nonce 4096
|
||||
0x88df895d8d8539ffull, 0x052fb10eab66d76dull, 0x534663c60e7de2e5ull, 0xc76cb6763d97b1beull, 0x24c28c52fdbc5c86ull, 0x46a9b9ee877a5764ull, 0xcc000e068d59a8e0ull, 0x7e634381ee2664e7ull,
|
||||
0x0e8d46d76bb5f04aull, 0x2a7caab33aaf260full, 0x11805382c57768fcull, 0x1050d656bfd66847ull, 0x6dac200ef67b302full, 0xfcebdae4acc594c2ull, 0x59ddcc381267cfb0ull, 0x15e9e31cd7347f92ull,
|
||||
0xad4fb91f9c31a9d1ull, 0x01bdb6cd682f2d3aull, 0xd94caecd87d9bc10ull, 0x3bc7dd917ded4591ull, 0x1dc97ba9f06c33cbull, 0xb503609bab3d4397ull, 0xa6001b284f75c5e4ull, 0x56981fe3d9beddccull,
|
||||
0xc1e3cf75a47fd944ull, 0x0d44d93176a1ab11ull, 0x31d24d65e0706657ull, 0xff7363ff80342341ull, 0x8060c156f3d360beull, 0xa75837c717e7eadaull, 0x86f2a525a3cf8d13ull, 0xac0d7a25d491c048ull
|
||||
},
|
||||
{ // base nonce 1000000
|
||||
0x77e307c1d3c4a45dull, 0xc46db2c1181e7d29ull, 0x60d540ae23190d93ull, 0x03a6005b827a3b1eull, 0x549bac8bd1c1e201ull, 0x780dfc346d705778ull, 0xcfcbebfdbb68e6f4ull, 0x3c5ed8998530360dull,
|
||||
0x2b88fe019717d013ull, 0xb70795cd85b32d67ull, 0x6decdbef0be191d7ull, 0x8f22242400f2cb78ull, 0xbbd28d1419989b8aull, 0xe16cffacddfb8292ull, 0xd8248122cb40a2a3ull, 0x2522a2642fae7251ull,
|
||||
0xfeeea1da8b23ad1bull, 0x3be311612545a761ull, 0x4ddb4b56a8275591ull, 0x7cb9edbe2572ab8bull, 0xe9a8ab8f94c06690ull, 0xd853e9b7bc76760dull, 0x03b2b843379ca0e4ull, 0x1954de4a7d95cb4full,
|
||||
0xad25e818a9aa212bull, 0x14be913efb5627feull, 0x2006c5d0f1ee99feull, 0xb1ce75c39bd2f557ull, 0xcd2f842898802fecull, 0xf44486cc9aa01340ull, 0x09a89237e25250eeull, 0xb342a4880eb580b3ull
|
||||
}
|
||||
};
|
||||
|
||||
// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455).
|
||||
static const uint32_t IGNEUM_DS_HEAD[16] = {
|
||||
0x3dd50b1fu, 0x93123701u, 0xcdd88b45u, 0x2c37c6b0u, 0x48edec90u, 0x3a7d2407u, 0xb03a0ac0u, 0x10e2c6a5u,
|
||||
0xf46429c0u, 0x4b1f3b3au, 0xd69dcaffu, 0xa48808fdu, 0xc8b2db13u, 0xc555587fu, 0x53042758u, 0x1a71bda6u
|
||||
};
|
||||
static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u;
|
||||
static const uint32_t IGNEUM_DS_LAST = 0xf7b7180eu;
|
||||
// 64 sampled dataset words (index, value) computed on the Mac.
|
||||
#define IGNEUM_DS_SAMPLES 64
|
||||
static const uint32_t IGNEUM_DS_SAMPLE_INDEX[IGNEUM_DS_SAMPLES] = {
|
||||
59471966u, 217795994u, 208353206u, 42483309u, 172547758u, 148076330u, 183853158u, 214389424u, 267488061u, 169781097u, 184093494u, 153880993u, 84977930u, 46426879u, 3093825u, 225364072u, 44593546u, 260713159u, 168250303u, 52384140u, 223401610u, 45554030u, 95410555u, 175039924u, 79171087u, 267580473u, 24168642u, 37981670u, 171551130u, 195559979u, 204611762u, 140997658u, 138925853u, 86637313u, 20736778u, 219665210u, 160430336u, 264654675u, 8013395u, 228945585u, 213884386u, 104419827u, 44185464u, 142737231u, 99284897u, 132475900u, 61861762u, 132056166u, 262388043u, 91878046u, 117353561u, 124768597u, 71352993u, 190698941u, 46055428u, 55281366u, 165145231u, 106810753u, 171985651u, 232085256u, 159510492u, 40072060u, 209107596u, 39023794u
|
||||
};
|
||||
static const uint32_t IGNEUM_DS_SAMPLE_VALUE[IGNEUM_DS_SAMPLES] = {
|
||||
0x7f292d0cu, 0x73dd96cfu, 0x202b4ad2u, 0xf6d8038fu, 0x7835659bu, 0x03a3a00du, 0x022d1e04u, 0x89957f84u, 0xe3451874u, 0xc82651dau, 0x29759e8cu, 0xa7a8cca3u, 0x45098d58u, 0x60d618eau, 0xf42d0e2cu, 0xc3546ab2u, 0xa7de6336u, 0xd592fdf4u, 0x083d78d1u, 0x64bd8824u, 0x8a66c547u, 0x8bc97d55u, 0x1411385du, 0x485f6fc0u, 0xc8142e6cu, 0xd8c1b7ebu, 0x6671b76bu, 0x63d4b24eu, 0xd622091bu, 0xa6821c04u, 0x8c2949d6u, 0x6b064395u, 0x5fa8e992u, 0xf49195b8u, 0x990254beu, 0xc5e6169du, 0xfff9c259u, 0x0e3c45dau, 0x42d69754u, 0xe23a1172u, 0x6d7d42e0u, 0xc30766efu, 0xe8abd145u, 0x11348b7cu, 0x837ccca1u, 0xe65b14c7u, 0x550f2d75u, 0xfedbcd25u, 0xb7430a6au, 0x5bbbdc45u, 0x2f81bc7du, 0xebbd5147u, 0xe991ee42u, 0x74d6d9abu, 0xc0eb32eau, 0x589a8daau, 0xd50d6a45u, 0x3f0ec2ccu, 0x31377e5eu, 0xe22a256fu, 0x52e144d0u, 0xa3fd1811u, 0xd3be259du, 0xc9bbc553u
|
||||
};
|
||||
// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^26 words.
|
||||
static const uint32_t IGNEUM_CACHE_HEAD[16] = {
|
||||
0xebd9055cu, 0x7eaf6b21u, 0x9b610855u, 0x134a8562u, 0xb10059bau, 0x7f459e58u, 0x8f3c7873u, 0x3523e609u,
|
||||
0x75b8f4cau, 0x750d31b4u, 0x0dae0782u, 0x26f10015u, 0x7ba87a41u, 0x0efd8543u, 0x691d6368u, 0x8929d967u
|
||||
};
|
||||
static const uint32_t IGNEUM_CACHE_LAST[16] = {
|
||||
0x51474edfu, 0xc3cfba93u, 0xf21454cfu, 0x9b79baacu, 0x4d4fce23u, 0xd701dfd5u, 0x37357ab3u, 0x1be693fau,
|
||||
0xa701cb3bu, 0x7467c620u, 0x428184e8u, 0xf4010df0u, 0xe33aa88fu, 0xc78d6d62u, 0xae9ba9c0u, 0xf4bba97eu
|
||||
};
|
||||
static const uint64_t IGNEUM_CACHE_FNV64 = 0x448274a57f508cbcull;
|
||||
36
proto-cuda/packs-ca2-era/era-4/vectors.json
Normal file
36
proto-cuda/packs-ca2-era/era-4/vectors.json
Normal file
|
|
@ -0,0 +1,36 @@
|
|||
{
|
||||
"seed": "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000",
|
||||
"day": "bytes:69676e65756d2d6461792ffa50000000000000",
|
||||
"dataset_mode": "memory-hard",
|
||||
"dataset_log2_words": 28,
|
||||
"mask": "0x0fffffff",
|
||||
"lanes": 32,
|
||||
"source": "igneum-pow (Rust) CPU interpreter, generator v3, memory-hard dataset",
|
||||
"warps": [
|
||||
{"base_nonce": 0, "expected": [
|
||||
"0x8e32500ee72aa1f8", "0x710f1522fd52e28e", "0x62fd41712fd5bd3e", "0x3837ee0ed93c4015", "0x0338e814e0e8a4fa", "0xaab705ee7514af71", "0x4a3e143121ed41a6", "0xe843b01bee434ef6",
|
||||
"0x4f9974ee1f331bc6", "0x1e089106d4d04581", "0x8a986409754a08b8", "0x54bf0b893560e130", "0x1346ddc8074de7dd", "0x5fdc54666289e864", "0xe9ed965ed22a9ea1", "0x7915337fc840db87",
|
||||
"0xaab900aea6b7c2cf", "0x81d768fa492b1cf2", "0xfd62bcf1021629df", "0x6280dc76e5125edd", "0x17055f35fc9a4ed3", "0xead760de60ae3e84", "0x14ea3ac017d95d57", "0x5f09e499a0096d83",
|
||||
"0xdf7c24624de53e02", "0x10e39b014ea6d6c6", "0xfcfac36ecf54fb8c", "0x8b510ae1be613404", "0xa9a07c9f9fa58801", "0xba88c22112b80fd0", "0x673c5bf0fbb48d35", "0xeef7f536b1e0a810"
|
||||
]},
|
||||
{"base_nonce": 4096, "expected": [
|
||||
"0x88df895d8d8539ff", "0x052fb10eab66d76d", "0x534663c60e7de2e5", "0xc76cb6763d97b1be", "0x24c28c52fdbc5c86", "0x46a9b9ee877a5764", "0xcc000e068d59a8e0", "0x7e634381ee2664e7",
|
||||
"0x0e8d46d76bb5f04a", "0x2a7caab33aaf260f", "0x11805382c57768fc", "0x1050d656bfd66847", "0x6dac200ef67b302f", "0xfcebdae4acc594c2", "0x59ddcc381267cfb0", "0x15e9e31cd7347f92",
|
||||
"0xad4fb91f9c31a9d1", "0x01bdb6cd682f2d3a", "0xd94caecd87d9bc10", "0x3bc7dd917ded4591", "0x1dc97ba9f06c33cb", "0xb503609bab3d4397", "0xa6001b284f75c5e4", "0x56981fe3d9beddcc",
|
||||
"0xc1e3cf75a47fd944", "0x0d44d93176a1ab11", "0x31d24d65e0706657", "0xff7363ff80342341", "0x8060c156f3d360be", "0xa75837c717e7eada", "0x86f2a525a3cf8d13", "0xac0d7a25d491c048"
|
||||
]},
|
||||
{"base_nonce": 1000000, "expected": [
|
||||
"0x77e307c1d3c4a45d", "0xc46db2c1181e7d29", "0x60d540ae23190d93", "0x03a6005b827a3b1e", "0x549bac8bd1c1e201", "0x780dfc346d705778", "0xcfcbebfdbb68e6f4", "0x3c5ed8998530360d",
|
||||
"0x2b88fe019717d013", "0xb70795cd85b32d67", "0x6decdbef0be191d7", "0x8f22242400f2cb78", "0xbbd28d1419989b8a", "0xe16cffacddfb8292", "0xd8248122cb40a2a3", "0x2522a2642fae7251",
|
||||
"0xfeeea1da8b23ad1b", "0x3be311612545a761", "0x4ddb4b56a8275591", "0x7cb9edbe2572ab8b", "0xe9a8ab8f94c06690", "0xd853e9b7bc76760d", "0x03b2b843379ca0e4", "0x1954de4a7d95cb4f",
|
||||
"0xad25e818a9aa212b", "0x14be913efb5627fe", "0x2006c5d0f1ee99fe", "0xb1ce75c39bd2f557", "0xcd2f842898802fec", "0xf44486cc9aa01340", "0x09a89237e25250ee", "0xb342a4880eb580b3"
|
||||
]}
|
||||
],
|
||||
"dataset_head": ["0x3dd50b1f", "0x93123701", "0xcdd88b45", "0x2c37c6b0", "0x48edec90", "0x3a7d2407", "0xb03a0ac0", "0x10e2c6a5", "0xf46429c0", "0x4b1f3b3a", "0xd69dcaff", "0xa48808fd", "0xc8b2db13", "0xc555587f", "0x53042758", "0x1a71bda6"],
|
||||
"dataset_last_index": 268435455,
|
||||
"dataset_last": "0xf7b7180e",
|
||||
"dataset_samples": [{"index": 59471966, "value": "0x7f292d0c"}, {"index": 217795994, "value": "0x73dd96cf"}, {"index": 208353206, "value": "0x202b4ad2"}, {"index": 42483309, "value": "0xf6d8038f"}, {"index": 172547758, "value": "0x7835659b"}, {"index": 148076330, "value": "0x03a3a00d"}, {"index": 183853158, "value": "0x022d1e04"}, {"index": 214389424, "value": "0x89957f84"}, {"index": 267488061, "value": "0xe3451874"}, {"index": 169781097, "value": "0xc82651da"}, {"index": 184093494, "value": "0x29759e8c"}, {"index": 153880993, "value": "0xa7a8cca3"}, {"index": 84977930, "value": "0x45098d58"}, {"index": 46426879, "value": "0x60d618ea"}, {"index": 3093825, "value": "0xf42d0e2c"}, {"index": 225364072, "value": "0xc3546ab2"}, {"index": 44593546, "value": "0xa7de6336"}, {"index": 260713159, "value": "0xd592fdf4"}, {"index": 168250303, "value": "0x083d78d1"}, {"index": 52384140, "value": "0x64bd8824"}, {"index": 223401610, "value": "0x8a66c547"}, {"index": 45554030, "value": "0x8bc97d55"}, {"index": 95410555, "value": "0x1411385d"}, {"index": 175039924, "value": "0x485f6fc0"}, {"index": 79171087, "value": "0xc8142e6c"}, {"index": 267580473, "value": "0xd8c1b7eb"}, {"index": 24168642, "value": "0x6671b76b"}, {"index": 37981670, "value": "0x63d4b24e"}, {"index": 171551130, "value": "0xd622091b"}, {"index": 195559979, "value": "0xa6821c04"}, {"index": 204611762, "value": "0x8c2949d6"}, {"index": 140997658, "value": "0x6b064395"}, {"index": 138925853, "value": "0x5fa8e992"}, {"index": 86637313, "value": "0xf49195b8"}, {"index": 20736778, "value": "0x990254be"}, {"index": 219665210, "value": "0xc5e6169d"}, {"index": 160430336, "value": "0xfff9c259"}, {"index": 264654675, "value": "0x0e3c45da"}, {"index": 8013395, "value": "0x42d69754"}, {"index": 228945585, "value": "0xe23a1172"}, {"index": 213884386, "value": "0x6d7d42e0"}, {"index": 104419827, "value": "0xc30766ef"}, {"index": 44185464, "value": "0xe8abd145"}, {"index": 142737231, "value": "0x11348b7c"}, {"index": 99284897, "value": "0x837ccca1"}, {"index": 132475900, "value": "0xe65b14c7"}, {"index": 61861762, "value": "0x550f2d75"}, {"index": 132056166, "value": "0xfedbcd25"}, {"index": 262388043, "value": "0xb7430a6a"}, {"index": 91878046, "value": "0x5bbbdc45"}, {"index": 117353561, "value": "0x2f81bc7d"}, {"index": 124768597, "value": "0xebbd5147"}, {"index": 71352993, "value": "0xe991ee42"}, {"index": 190698941, "value": "0x74d6d9ab"}, {"index": 46055428, "value": "0xc0eb32ea"}, {"index": 55281366, "value": "0x589a8daa"}, {"index": 165145231, "value": "0xd50d6a45"}, {"index": 106810753, "value": "0x3f0ec2cc"}, {"index": 171985651, "value": "0x31377e5e"}, {"index": 232085256, "value": "0xe22a256f"}, {"index": 159510492, "value": "0x52e144d0"}, {"index": 40072060, "value": "0xa3fd1811"}, {"index": 209107596, "value": "0xd3be259d"}, {"index": 39023794, "value": "0xc9bbc553"}],
|
||||
"cache_head": ["0xebd9055c", "0x7eaf6b21", "0x9b610855", "0x134a8562", "0xb10059ba", "0x7f459e58", "0x8f3c7873", "0x3523e609", "0x75b8f4ca", "0x750d31b4", "0x0dae0782", "0x26f10015", "0x7ba87a41", "0x0efd8543", "0x691d6368", "0x8929d967"],
|
||||
"cache_last_line": ["0x51474edf", "0xc3cfba93", "0xf21454cf", "0x9b79baac", "0x4d4fce23", "0xd701dfd5", "0x37357ab3", "0x1be693fa", "0xa701cb3b", "0x7467c620", "0x428184e8", "0xf4010df0", "0xe33aa88f", "0xc78d6d62", "0xae9ba9c0", "0xf4bba97e"],
|
||||
"cache_fnv1a64": "0x448274a57f508cbc"
|
||||
}
|
||||
281
proto-cuda/packs-ca2-era/era-5/kernel.cl
Normal file
281
proto-cuda/packs-ca2-era/era-5/kernel.cl
Normal file
|
|
@ -0,0 +1,281 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal).
|
||||
// Built from source at runtime by proto-opencl/host.c, which passes these defines:
|
||||
// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit)
|
||||
// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default)
|
||||
// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32
|
||||
// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition
|
||||
// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units;
|
||||
// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32.
|
||||
#ifndef IGNEUM_GROUP
|
||||
#define IGNEUM_GROUP 32
|
||||
#endif
|
||||
#ifndef IGNEUM_EXCHANGE
|
||||
#define IGNEUM_EXCHANGE 0
|
||||
#endif
|
||||
#ifdef __OPENCL_VERSION__
|
||||
#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1)))
|
||||
#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n]
|
||||
#if IGNEUM_EXCHANGE == 1
|
||||
#ifdef cl_khr_subgroups
|
||||
#pragma OPENCL EXTENSION cl_khr_subgroups : enable
|
||||
#endif
|
||||
#ifdef cl_khr_subgroup_shuffle
|
||||
#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable
|
||||
#endif
|
||||
#elif IGNEUM_EXCHANGE == 2
|
||||
#pragma OPENCL EXTENSION cl_intel_subgroups : enable
|
||||
#endif
|
||||
#else
|
||||
// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros.
|
||||
#include "emu_opencl.h"
|
||||
#endif
|
||||
|
||||
#if IGNEUM_EXCHANGE == 1
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m))
|
||||
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
|
||||
#elif IGNEUM_EXCHANGE == 2
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m))
|
||||
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
|
||||
#else
|
||||
// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per
|
||||
// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane
|
||||
// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's
|
||||
// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier.
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; }
|
||||
#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; }
|
||||
#endif
|
||||
|
||||
static inline uint splitmix32(uint x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32.
|
||||
static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); }
|
||||
// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x.
|
||||
static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); }
|
||||
static inline uint ds_elem(uint i, uint d0, uint d1) {
|
||||
uint x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
|
||||
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
|
||||
#define MH_CACHE_LINE_MASK 0x003fffffu
|
||||
#define MH_SEGMENT_LINES 64u
|
||||
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
|
||||
static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
|
||||
|
||||
// y = ChaCha12 core(x) + x
|
||||
static inline void mh_chacha_block(const uint* x, uint* y) {
|
||||
for (uint i = 0u; i < 16u; ++i) y[i] = x[i];
|
||||
for (uint r = 0u; r < 6u; ++r) {
|
||||
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
|
||||
}
|
||||
for (uint i = 0u; i < 16u; ++i) y[i] += x[i];
|
||||
}
|
||||
|
||||
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
|
||||
static inline void mh_cache_segment(__global uint* cache, uint seg) {
|
||||
uint prev[16]; uint x[16]; uint y[16];
|
||||
for (uint i = 0u; i < 16u; ++i) prev[i] = 0u;
|
||||
for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) {
|
||||
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
|
||||
x[4] = 0xceed56d7u ^ prev[4];
|
||||
x[5] = 0x9ba270d2u ^ prev[5];
|
||||
x[6] = 0x82caab2du ^ prev[6];
|
||||
x[7] = 0x81ebce0eu ^ prev[7];
|
||||
x[8] = 0x12b6ecf1u ^ prev[8];
|
||||
x[9] = 0xd0f3fd7cu ^ prev[9];
|
||||
x[10] = 0xd872eefeu ^ prev[10];
|
||||
x[11] = 0xc158c7bdu ^ prev[11];
|
||||
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
|
||||
mh_chacha_block(x, y);
|
||||
__global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
|
||||
}
|
||||
}
|
||||
|
||||
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
|
||||
static inline void mh_mixer(uint* s, uint rk) {
|
||||
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
|
||||
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
|
||||
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
|
||||
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
|
||||
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
|
||||
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
|
||||
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
|
||||
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
|
||||
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
|
||||
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
|
||||
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
|
||||
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
|
||||
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
|
||||
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
|
||||
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
|
||||
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
|
||||
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
|
||||
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
|
||||
}
|
||||
|
||||
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
|
||||
static inline void mh_item(__global const uint* cache, uint t, uint* s) {
|
||||
s[0] = 0xceed56d7u;
|
||||
s[1] = 0x9ba270d2u;
|
||||
s[2] = 0x82caab2du;
|
||||
s[3] = 0x81ebce0eu;
|
||||
s[4] = 0x12b6ecf1u;
|
||||
s[5] = 0xd0f3fd7cu;
|
||||
s[6] = 0xd872eefeu;
|
||||
s[7] = 0xc158c7bdu;
|
||||
s[8] = t * 0xf351d601u + 0xc6892460u;
|
||||
s[9] = t * 0xa3bb398fu + 0x25b7228au;
|
||||
s[10] = t * 0xb5a09e35u + 0xcd515004u;
|
||||
s[11] = t * 0x7509c9c1u + 0x2846527au;
|
||||
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
|
||||
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
|
||||
s[14] = t * 0xded91851u + 0x82961bacu;
|
||||
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
|
||||
for (uint r = 0u; r < 8u; ++r) {
|
||||
mh_mixer(s, 0x9E3779B9u * (r + 1u));
|
||||
__global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i];
|
||||
}
|
||||
mh_mixer(s, 0x9E3779B9u * 9u);
|
||||
}
|
||||
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 0 2 10 13 of w.
|
||||
static inline uint mh_j(uint w) { return ((w >> 0u) & 1u) | (((w >> 2u) & 1u) << 1) | (((w >> 10u) & 1u) << 2) | (((w >> 13u) & 1u) << 3); }
|
||||
static inline uint mh_t(uint w) { w = (w & 0x00001fffu) | ((w >> 14u) << 13u); w = (w & 0x000003ffu) | ((w >> 11u) << 10u); w = (w & 0x00000003u) | ((w >> 3u) << 2u); w = (w & 0x00000000u) | ((w >> 1u) << 0u); return w; }
|
||||
static inline uint mh_addr(uint t, uint j) { uint w = t; w = ((w >> 0u) << 1u) | (w & 0x00000000u) | (((j >> 0u) & 1u) << 0u); w = ((w >> 2u) << 3u) | (w & 0x00000003u) | (((j >> 1u) & 1u) << 2u); w = ((w >> 10u) << 11u) | (w & 0x000003ffu) | (((j >> 2u) & 1u) << 10u); w = ((w >> 13u) << 14u) | (w & 0x00001fffu) | (((j >> 3u) & 1u) << 13u); return w; }
|
||||
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
|
||||
static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }
|
||||
|
||||
// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item.
|
||||
// The same constants as memhard.h in this pack (one emitter, three dialects).
|
||||
__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) {
|
||||
uint seg = (uint)get_global_id(0);
|
||||
if (seg < nSegments) mh_cache_segment(cache, seg);
|
||||
}
|
||||
__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) {
|
||||
uint t = (uint)get_global_id(0);
|
||||
if (t < nItems) {
|
||||
uint s[16];
|
||||
mh_item(cache, t, s);
|
||||
for (uint i = 0u; i < 16u; ++i) ds[(ulong)mh_addr(t, i)] = s[i];
|
||||
}
|
||||
}
|
||||
|
||||
// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the
|
||||
// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and
|
||||
// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all).
|
||||
IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) {
|
||||
uint gid = (uint)get_global_id(0);
|
||||
uint lid = (uint)get_local_id(0);
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
#if IGNEUM_EXCHANGE == 0
|
||||
IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
|
||||
uint xk = 0u;
|
||||
#else
|
||||
(void)lid;
|
||||
#endif
|
||||
{ uint x = nonce ^ 0x667d0fbdu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x7b8e5963u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
|
||||
{ uint x = nonce ^ 0x7b8e5963u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x31c67e5eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
|
||||
{ uint x = nonce ^ 0x31c67e5eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4529ddc6u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
|
||||
{ uint x = nonce ^ 0x4529ddc6u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef19d6d8u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
|
||||
{ uint x = nonce ^ 0xef19d6d8u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xaccf6211u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
|
||||
{ uint x = nonce ^ 0xaccf6211u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xda0aed32u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
|
||||
{ uint x = nonce ^ 0xda0aed32u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xabc6df31u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
|
||||
{ uint x = nonce ^ 0xabc6df31u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x667d0fbdu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
|
||||
r7 = r7 ^ r0; // 1 xor
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 1u); r3 = r3 ^ t_; } // 2 shfl
|
||||
r4 = r4 - r1; // 3 sub
|
||||
r2 = r0 * r4 + r2; // 4 mad
|
||||
r4 = rotl_imm(r4, 9u); // 5 rotl
|
||||
r0 = rotr_var(r0, r2); // 6 rotr
|
||||
r6 = r6 ^ ds[((rotl_imm(r7 * 0x03ac37adu, 22u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r2 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r0 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r1 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
|
||||
r5 = r5 * r4; // 12 mul
|
||||
r7 = r7 ^ ds[((rotl_imm(r6 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 16u); r6 = r6 ^ t_; } // 14 shfl
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r3 = r3 ^ t_; } // 15 shfl
|
||||
r2 = r2 | r7; // 16 or
|
||||
r0 = rotr_var(r0, r6); // 17 rotr
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
|
||||
r7 = rotl_imm(r7, 24u); // 19 rotl
|
||||
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
|
||||
r2 = r2 | r6; // 21 or
|
||||
r6 = r5 * r4 + r6; // 22 mad
|
||||
r1 = r1 | r0; // 23 or
|
||||
r6 = r6 ^ r1; // 24 xor
|
||||
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
|
||||
r7 = rotr_var(r7, r0); // 26 rotr
|
||||
r4 = r5 * r7 + r4; // 27 mad
|
||||
r3 = r2 * r4 + r3; // 28 mad
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
|
||||
r2 = r2 ^ ds[((rotl_imm(r3 * 0x03ac37adu, 22u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r5 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
|
||||
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
|
||||
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
|
||||
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
|
||||
r2 = mul_hi(r2, r5); // 35 mulhi
|
||||
r4 = r4 ^ r2; // 36 xor
|
||||
r6 = r6 * r5; // 37 mul
|
||||
r7 = r7 ^ r0; // 38 xor
|
||||
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
|
||||
r2 = rotr_var(r2, r3); // 40 rotr
|
||||
r7 = r7 - r0; // 41 sub
|
||||
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 43 shfl
|
||||
r0 = r0 ^ ds[((rotl_imm(r7 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
|
||||
r3 = r3 ^ ds[((rotl_imm(r1 * 0x03ac37adu, 22u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
|
||||
r6 = r6 ^ ds[((rotl_imm(r3 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
|
||||
r4 = mul_hi(r4, r2); // 48 mulhi
|
||||
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r0 = r0 ^ t_; } // 50 shfl
|
||||
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r2 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
|
||||
r6 = rotl_imm(r6, 12u); // 53 rotl
|
||||
r3 = rotl_imm(r3, 12u); // 54 rotl
|
||||
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
|
||||
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r0 * 0x03ac37adu, 22u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
|
||||
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
|
||||
r3 = mul_hi(r3, r2); // 60 mulhi
|
||||
r6 = r6 | r4; // 61 or
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r5 = r5 ^ t_; } // 62 shfl
|
||||
r3 = r3 ^ ds[((rotl_imm(r6 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
|
||||
#if IGNEUM_EXCHANGE != 0
|
||||
// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the
|
||||
// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a
|
||||
// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md.
|
||||
IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) {
|
||||
if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); }
|
||||
}
|
||||
#endif
|
||||
163
proto-cuda/packs-ca2-era/era-5/kernel.cu
Normal file
163
proto-cuda/packs-ca2-era/era-5/kernel.cu
Normal file
|
|
@ -0,0 +1,163 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal).
|
||||
// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC.
|
||||
#include <cuda_runtime.h>
|
||||
#include <cstdint>
|
||||
#include "program.h"
|
||||
#include "memhard.h"
|
||||
|
||||
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31.
|
||||
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
|
||||
// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x.
|
||||
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) {
|
||||
uint32_t x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item.
|
||||
// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host.
|
||||
__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) {
|
||||
uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
if (seg < nSegments) mh_cache_segment(cache, seg);
|
||||
}
|
||||
__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
|
||||
uint32_t t = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
if (t < nItems) {
|
||||
uint32_t s[16];
|
||||
mh_item(cache, t, s);
|
||||
for (uint32_t i = 0u; i < 16u; ++i) ds[(size_t)mh_addr(t, i)] = s[i];
|
||||
}
|
||||
}
|
||||
|
||||
// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every
|
||||
// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a
|
||||
// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid.
|
||||
__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) {
|
||||
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
uint32_t nonce = baseNonce + gid;
|
||||
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint32_t x = nonce ^ 0x667d0fbdu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x7b8e5963u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
|
||||
{ uint32_t x = nonce ^ 0x7b8e5963u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x31c67e5eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
|
||||
{ uint32_t x = nonce ^ 0x31c67e5eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4529ddc6u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
|
||||
{ uint32_t x = nonce ^ 0x4529ddc6u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef19d6d8u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
|
||||
{ uint32_t x = nonce ^ 0xef19d6d8u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xaccf6211u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
|
||||
{ uint32_t x = nonce ^ 0xaccf6211u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xda0aed32u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
|
||||
{ uint32_t x = nonce ^ 0xda0aed32u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xabc6df31u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
|
||||
{ uint32_t x = nonce ^ 0xabc6df31u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x667d0fbdu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
|
||||
|
||||
for (uint32_t it = 0u; it < 8u; ++it) {
|
||||
uint32_t sel = r0;
|
||||
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
|
||||
r7 = r7 ^ r0; // 1 xor
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r6, 1); // 2 shfl
|
||||
r4 = r4 - r1; // 3 sub
|
||||
r2 = r0 * r4 + r2; // 4 mad
|
||||
r4 = rotl_imm(r4, 9u); // 5 rotl
|
||||
r0 = rotr_var(r0, r2); // 6 rotr
|
||||
r6 = r6 ^ ds[((rotl_imm(r7 * 0x03ac37adu, 22u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r2 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r0 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r1 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
|
||||
r5 = r5 * r4; // 12 mul
|
||||
r7 = r7 ^ ds[((rotl_imm(r6 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
|
||||
r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r5, 16); // 14 shfl
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r5, 1); // 15 shfl
|
||||
r2 = r2 | r7; // 16 or
|
||||
r0 = rotr_var(r0, r6); // 17 rotr
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
|
||||
r7 = rotl_imm(r7, 24u); // 19 rotl
|
||||
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
|
||||
r2 = r2 | r6; // 21 or
|
||||
r6 = r5 * r4 + r6; // 22 mad
|
||||
r1 = r1 | r0; // 23 or
|
||||
r6 = r6 ^ r1; // 24 xor
|
||||
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
|
||||
r7 = rotr_var(r7, r0); // 26 rotr
|
||||
r4 = r5 * r7 + r4; // 27 mad
|
||||
r3 = r2 * r4 + r3; // 28 mad
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
|
||||
r2 = r2 ^ ds[((rotl_imm(r3 * 0x03ac37adu, 22u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r5 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
|
||||
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
|
||||
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
|
||||
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
|
||||
r2 = __umulhi(r2, r5); // 35 mulhi
|
||||
r4 = r4 ^ r2; // 36 xor
|
||||
r6 = r6 * r5; // 37 mul
|
||||
r7 = r7 ^ r0; // 38 xor
|
||||
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
|
||||
r2 = rotr_var(r2, r3); // 40 rotr
|
||||
r7 = r7 - r0; // 41 sub
|
||||
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 43 shfl
|
||||
r0 = r0 ^ ds[((rotl_imm(r7 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
|
||||
r3 = r3 ^ ds[((rotl_imm(r1 * 0x03ac37adu, 22u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
|
||||
r6 = r6 ^ ds[((rotl_imm(r3 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
|
||||
r4 = __umulhi(r4, r2); // 48 mulhi
|
||||
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
|
||||
r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r7, 4); // 50 shfl
|
||||
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r2 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
|
||||
r6 = rotl_imm(r6, 12u); // 53 rotl
|
||||
r3 = rotl_imm(r3, 12u); // 54 rotl
|
||||
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
|
||||
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r0 * 0x03ac37adu, 22u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
|
||||
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
|
||||
r3 = __umulhi(r3, r2); // 60 mulhi
|
||||
r6 = r6 | r4; // 61 or
|
||||
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 62 shfl
|
||||
r3 = r3 ^ ds[((rotl_imm(r6 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
|
||||
}
|
||||
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
|
||||
}
|
||||
|
||||
// Host-side launch wrappers. Declared in program.h, called from host.cu.
|
||||
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) {
|
||||
if (nSegments == 0u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 256u;
|
||||
uint32_t grid = (nSegments + block - 1u) / block;
|
||||
igneum_cache_fill<<<grid, block>>>(cache, nSegments);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
|
||||
if (nItems == 0u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 256u;
|
||||
uint32_t grid = (nItems + block - 1u) / block;
|
||||
igneum_build<<<grid, block>>>(ds, cache, nItems);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
uint32_t nonces, uint32_t blockWarps) {
|
||||
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 32u * blockWarps;
|
||||
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
|
||||
igneum_hash<<<nonces / block, block>>>(ds, out, baseNonce, mask);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
|
||||
cudaFuncAttributes attr;
|
||||
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash);
|
||||
if (e != cudaSuccess) return e;
|
||||
*numRegs = attr.numRegs;
|
||||
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0);
|
||||
}
|
||||
375
proto-cuda/packs-ca2-era/era-5/kernel_bound.cl
Normal file
375
proto-cuda/packs-ca2-era/era-5/kernel_bound.cl
Normal file
|
|
@ -0,0 +1,375 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal).
|
||||
// Built from source at runtime by proto-opencl/host.c, which passes these defines:
|
||||
// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit)
|
||||
// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default)
|
||||
// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32
|
||||
// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition
|
||||
// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units;
|
||||
// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32.
|
||||
#ifndef IGNEUM_GROUP
|
||||
#define IGNEUM_GROUP 32
|
||||
#endif
|
||||
#ifndef IGNEUM_EXCHANGE
|
||||
#define IGNEUM_EXCHANGE 0
|
||||
#endif
|
||||
#ifdef __OPENCL_VERSION__
|
||||
#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1)))
|
||||
#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n]
|
||||
#if IGNEUM_EXCHANGE == 1
|
||||
#ifdef cl_khr_subgroups
|
||||
#pragma OPENCL EXTENSION cl_khr_subgroups : enable
|
||||
#endif
|
||||
#ifdef cl_khr_subgroup_shuffle
|
||||
#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable
|
||||
#endif
|
||||
#elif IGNEUM_EXCHANGE == 2
|
||||
#pragma OPENCL EXTENSION cl_intel_subgroups : enable
|
||||
#endif
|
||||
#else
|
||||
// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros.
|
||||
#include "emu_opencl.h"
|
||||
#endif
|
||||
|
||||
#if IGNEUM_EXCHANGE == 1
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m))
|
||||
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
|
||||
#elif IGNEUM_EXCHANGE == 2
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m))
|
||||
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
|
||||
#else
|
||||
// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per
|
||||
// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane
|
||||
// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's
|
||||
// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier.
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; }
|
||||
#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; }
|
||||
#endif
|
||||
|
||||
static inline uint splitmix32(uint x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32.
|
||||
static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); }
|
||||
// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x.
|
||||
static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); }
|
||||
static inline uint ds_elem(uint i, uint d0, uint d1) {
|
||||
uint x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
|
||||
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
|
||||
#define MH_CACHE_LINE_MASK 0x003fffffu
|
||||
#define MH_SEGMENT_LINES 64u
|
||||
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
|
||||
static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
|
||||
|
||||
// y = ChaCha12 core(x) + x
|
||||
static inline void mh_chacha_block(const uint* x, uint* y) {
|
||||
for (uint i = 0u; i < 16u; ++i) y[i] = x[i];
|
||||
for (uint r = 0u; r < 6u; ++r) {
|
||||
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
|
||||
}
|
||||
for (uint i = 0u; i < 16u; ++i) y[i] += x[i];
|
||||
}
|
||||
|
||||
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
|
||||
static inline void mh_cache_segment(__global uint* cache, uint seg) {
|
||||
uint prev[16]; uint x[16]; uint y[16];
|
||||
for (uint i = 0u; i < 16u; ++i) prev[i] = 0u;
|
||||
for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) {
|
||||
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
|
||||
x[4] = 0xceed56d7u ^ prev[4];
|
||||
x[5] = 0x9ba270d2u ^ prev[5];
|
||||
x[6] = 0x82caab2du ^ prev[6];
|
||||
x[7] = 0x81ebce0eu ^ prev[7];
|
||||
x[8] = 0x12b6ecf1u ^ prev[8];
|
||||
x[9] = 0xd0f3fd7cu ^ prev[9];
|
||||
x[10] = 0xd872eefeu ^ prev[10];
|
||||
x[11] = 0xc158c7bdu ^ prev[11];
|
||||
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
|
||||
mh_chacha_block(x, y);
|
||||
__global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
|
||||
}
|
||||
}
|
||||
|
||||
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
|
||||
static inline void mh_mixer(uint* s, uint rk) {
|
||||
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
|
||||
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
|
||||
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
|
||||
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
|
||||
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
|
||||
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
|
||||
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
|
||||
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
|
||||
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
|
||||
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
|
||||
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
|
||||
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
|
||||
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
|
||||
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
|
||||
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
|
||||
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
|
||||
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
|
||||
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
|
||||
}
|
||||
|
||||
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
|
||||
static inline void mh_item(__global const uint* cache, uint t, uint* s) {
|
||||
s[0] = 0xceed56d7u;
|
||||
s[1] = 0x9ba270d2u;
|
||||
s[2] = 0x82caab2du;
|
||||
s[3] = 0x81ebce0eu;
|
||||
s[4] = 0x12b6ecf1u;
|
||||
s[5] = 0xd0f3fd7cu;
|
||||
s[6] = 0xd872eefeu;
|
||||
s[7] = 0xc158c7bdu;
|
||||
s[8] = t * 0xf351d601u + 0xc6892460u;
|
||||
s[9] = t * 0xa3bb398fu + 0x25b7228au;
|
||||
s[10] = t * 0xb5a09e35u + 0xcd515004u;
|
||||
s[11] = t * 0x7509c9c1u + 0x2846527au;
|
||||
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
|
||||
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
|
||||
s[14] = t * 0xded91851u + 0x82961bacu;
|
||||
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
|
||||
for (uint r = 0u; r < 8u; ++r) {
|
||||
mh_mixer(s, 0x9E3779B9u * (r + 1u));
|
||||
__global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i];
|
||||
}
|
||||
mh_mixer(s, 0x9E3779B9u * 9u);
|
||||
}
|
||||
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 0 2 10 13 of w.
|
||||
static inline uint mh_j(uint w) { return ((w >> 0u) & 1u) | (((w >> 2u) & 1u) << 1) | (((w >> 10u) & 1u) << 2) | (((w >> 13u) & 1u) << 3); }
|
||||
static inline uint mh_t(uint w) { w = (w & 0x00001fffu) | ((w >> 14u) << 13u); w = (w & 0x000003ffu) | ((w >> 11u) << 10u); w = (w & 0x00000003u) | ((w >> 3u) << 2u); w = (w & 0x00000000u) | ((w >> 1u) << 0u); return w; }
|
||||
static inline uint mh_addr(uint t, uint j) { uint w = t; w = ((w >> 0u) << 1u) | (w & 0x00000000u) | (((j >> 0u) & 1u) << 0u); w = ((w >> 2u) << 3u) | (w & 0x00000003u) | (((j >> 1u) & 1u) << 2u); w = ((w >> 10u) << 11u) | (w & 0x000003ffu) | (((j >> 2u) & 1u) << 10u); w = ((w >> 13u) << 14u) | (w & 0x00001fffu) | (((j >> 3u) & 1u) << 13u); return w; }
|
||||
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
|
||||
static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }
|
||||
|
||||
// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item.
|
||||
// The same constants as memhard.h in this pack (one emitter, three dialects).
|
||||
__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) {
|
||||
uint seg = (uint)get_global_id(0);
|
||||
if (seg < nSegments) mh_cache_segment(cache, seg);
|
||||
}
|
||||
__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) {
|
||||
uint t = (uint)get_global_id(0);
|
||||
if (t < nItems) {
|
||||
uint s[16];
|
||||
mh_item(cache, t, s);
|
||||
for (uint i = 0u; i < 16u; ++i) ds[(ulong)mh_addr(t, i)] = s[i];
|
||||
}
|
||||
}
|
||||
|
||||
// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the
|
||||
// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and
|
||||
// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all).
|
||||
IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) {
|
||||
uint gid = (uint)get_global_id(0);
|
||||
uint lid = (uint)get_local_id(0);
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
#if IGNEUM_EXCHANGE == 0
|
||||
IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
|
||||
uint xk = 0u;
|
||||
#else
|
||||
(void)lid;
|
||||
#endif
|
||||
{ uint x = nonce ^ 0x667d0fbdu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x7b8e5963u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
|
||||
{ uint x = nonce ^ 0x7b8e5963u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x31c67e5eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
|
||||
{ uint x = nonce ^ 0x31c67e5eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4529ddc6u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
|
||||
{ uint x = nonce ^ 0x4529ddc6u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef19d6d8u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
|
||||
{ uint x = nonce ^ 0xef19d6d8u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xaccf6211u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
|
||||
{ uint x = nonce ^ 0xaccf6211u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xda0aed32u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
|
||||
{ uint x = nonce ^ 0xda0aed32u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xabc6df31u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
|
||||
{ uint x = nonce ^ 0xabc6df31u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x667d0fbdu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
|
||||
r7 = r7 ^ r0; // 1 xor
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 1u); r3 = r3 ^ t_; } // 2 shfl
|
||||
r4 = r4 - r1; // 3 sub
|
||||
r2 = r0 * r4 + r2; // 4 mad
|
||||
r4 = rotl_imm(r4, 9u); // 5 rotl
|
||||
r0 = rotr_var(r0, r2); // 6 rotr
|
||||
r6 = r6 ^ ds[((rotl_imm(r7 * 0x03ac37adu, 22u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r2 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r0 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r1 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
|
||||
r5 = r5 * r4; // 12 mul
|
||||
r7 = r7 ^ ds[((rotl_imm(r6 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 16u); r6 = r6 ^ t_; } // 14 shfl
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r3 = r3 ^ t_; } // 15 shfl
|
||||
r2 = r2 | r7; // 16 or
|
||||
r0 = rotr_var(r0, r6); // 17 rotr
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
|
||||
r7 = rotl_imm(r7, 24u); // 19 rotl
|
||||
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
|
||||
r2 = r2 | r6; // 21 or
|
||||
r6 = r5 * r4 + r6; // 22 mad
|
||||
r1 = r1 | r0; // 23 or
|
||||
r6 = r6 ^ r1; // 24 xor
|
||||
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
|
||||
r7 = rotr_var(r7, r0); // 26 rotr
|
||||
r4 = r5 * r7 + r4; // 27 mad
|
||||
r3 = r2 * r4 + r3; // 28 mad
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
|
||||
r2 = r2 ^ ds[((rotl_imm(r3 * 0x03ac37adu, 22u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r5 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
|
||||
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
|
||||
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
|
||||
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
|
||||
r2 = mul_hi(r2, r5); // 35 mulhi
|
||||
r4 = r4 ^ r2; // 36 xor
|
||||
r6 = r6 * r5; // 37 mul
|
||||
r7 = r7 ^ r0; // 38 xor
|
||||
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
|
||||
r2 = rotr_var(r2, r3); // 40 rotr
|
||||
r7 = r7 - r0; // 41 sub
|
||||
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 43 shfl
|
||||
r0 = r0 ^ ds[((rotl_imm(r7 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
|
||||
r3 = r3 ^ ds[((rotl_imm(r1 * 0x03ac37adu, 22u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
|
||||
r6 = r6 ^ ds[((rotl_imm(r3 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
|
||||
r4 = mul_hi(r4, r2); // 48 mulhi
|
||||
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r0 = r0 ^ t_; } // 50 shfl
|
||||
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r2 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
|
||||
r6 = rotl_imm(r6, 12u); // 53 rotl
|
||||
r3 = rotl_imm(r3, 12u); // 54 rotl
|
||||
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
|
||||
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r0 * 0x03ac37adu, 22u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
|
||||
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
|
||||
r3 = mul_hi(r3, r2); // 60 mulhi
|
||||
r6 = r6 | r4; // 61 or
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r5 = r5 ^ t_; } // 62 shfl
|
||||
r3 = r3 ^ ds[((rotl_imm(r6 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
|
||||
#if IGNEUM_EXCHANGE != 0
|
||||
// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the
|
||||
// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a
|
||||
// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md.
|
||||
IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) {
|
||||
if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); }
|
||||
}
|
||||
#endif
|
||||
|
||||
// Header-bound variant (bind.rs): the init words come from initw, not SEEDW. Same body as igneum_hash.
|
||||
IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw) {
|
||||
uint gid = (uint)get_global_id(0);
|
||||
uint lid = (uint)get_local_id(0);
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7];
|
||||
#if IGNEUM_EXCHANGE == 0
|
||||
IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
|
||||
uint xk = 0u;
|
||||
#else
|
||||
(void)lid;
|
||||
#endif
|
||||
{ uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; }
|
||||
{ uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; }
|
||||
{ uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; }
|
||||
{ uint x = nonce ^ iw3; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw4; }
|
||||
{ uint x = nonce ^ iw4; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw5; }
|
||||
{ uint x = nonce ^ iw5; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw6; }
|
||||
{ uint x = nonce ^ iw6; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw7; }
|
||||
{ uint x = nonce ^ iw7; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw0; }
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
|
||||
r7 = r7 ^ r0; // 1 xor
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 1u); r3 = r3 ^ t_; } // 2 shfl
|
||||
r4 = r4 - r1; // 3 sub
|
||||
r2 = r0 * r4 + r2; // 4 mad
|
||||
r4 = rotl_imm(r4, 9u); // 5 rotl
|
||||
r0 = rotr_var(r0, r2); // 6 rotr
|
||||
r6 = r6 ^ ds[((rotl_imm(r7 * 0x03ac37adu, 22u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r2 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r0 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r1 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
|
||||
r5 = r5 * r4; // 12 mul
|
||||
r7 = r7 ^ ds[((rotl_imm(r6 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 16u); r6 = r6 ^ t_; } // 14 shfl
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r3 = r3 ^ t_; } // 15 shfl
|
||||
r2 = r2 | r7; // 16 or
|
||||
r0 = rotr_var(r0, r6); // 17 rotr
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
|
||||
r7 = rotl_imm(r7, 24u); // 19 rotl
|
||||
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
|
||||
r2 = r2 | r6; // 21 or
|
||||
r6 = r5 * r4 + r6; // 22 mad
|
||||
r1 = r1 | r0; // 23 or
|
||||
r6 = r6 ^ r1; // 24 xor
|
||||
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
|
||||
r7 = rotr_var(r7, r0); // 26 rotr
|
||||
r4 = r5 * r7 + r4; // 27 mad
|
||||
r3 = r2 * r4 + r3; // 28 mad
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
|
||||
r2 = r2 ^ ds[((rotl_imm(r3 * 0x03ac37adu, 22u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r5 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
|
||||
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
|
||||
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
|
||||
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
|
||||
r2 = mul_hi(r2, r5); // 35 mulhi
|
||||
r4 = r4 ^ r2; // 36 xor
|
||||
r6 = r6 * r5; // 37 mul
|
||||
r7 = r7 ^ r0; // 38 xor
|
||||
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
|
||||
r2 = rotr_var(r2, r3); // 40 rotr
|
||||
r7 = r7 - r0; // 41 sub
|
||||
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 43 shfl
|
||||
r0 = r0 ^ ds[((rotl_imm(r7 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
|
||||
r3 = r3 ^ ds[((rotl_imm(r1 * 0x03ac37adu, 22u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
|
||||
r6 = r6 ^ ds[((rotl_imm(r3 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
|
||||
r4 = mul_hi(r4, r2); // 48 mulhi
|
||||
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r0 = r0 ^ t_; } // 50 shfl
|
||||
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r2 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
|
||||
r6 = rotl_imm(r6, 12u); // 53 rotl
|
||||
r3 = rotl_imm(r3, 12u); // 54 rotl
|
||||
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
|
||||
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r0 * 0x03ac37adu, 22u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
|
||||
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
|
||||
r3 = mul_hi(r3, r2); // 60 mulhi
|
||||
r6 = r6 | r4; // 61 or
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r5 = r5 ^ t_; } // 62 shfl
|
||||
r3 = r3 ^ ds[((rotl_imm(r6 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
123
proto-cuda/packs-ca2-era/era-5/kernel_bound.cu
Normal file
123
proto-cuda/packs-ca2-era/era-5/kernel_bound.cu
Normal file
|
|
@ -0,0 +1,123 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// Header-bound twin of igneum_hash in kernel.cu: the init words come from a kernel argument, not SEEDW.
|
||||
// Host declarations (also in program_bound.h if present):
|
||||
// struct IgneumInitWords { uint32_t w[8]; };
|
||||
// cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
// IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps);
|
||||
// cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
|
||||
#include <cuda_runtime.h>
|
||||
#include <cstdint>
|
||||
#include "program.h"
|
||||
|
||||
struct IgneumInitWords { uint32_t w[8]; };
|
||||
|
||||
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
|
||||
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
|
||||
__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw) {
|
||||
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
uint32_t nonce = baseNonce + gid;
|
||||
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint32_t x = nonce ^ iw.w[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw.w[1]; }
|
||||
{ uint32_t x = nonce ^ iw.w[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw.w[2]; }
|
||||
{ uint32_t x = nonce ^ iw.w[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw.w[3]; }
|
||||
{ uint32_t x = nonce ^ iw.w[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw.w[4]; }
|
||||
{ uint32_t x = nonce ^ iw.w[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw.w[5]; }
|
||||
{ uint32_t x = nonce ^ iw.w[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw.w[6]; }
|
||||
{ uint32_t x = nonce ^ iw.w[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw.w[7]; }
|
||||
{ uint32_t x = nonce ^ iw.w[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw.w[0]; }
|
||||
|
||||
for (uint32_t it = 0u; it < 8u; ++it) {
|
||||
uint32_t sel = r0;
|
||||
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
|
||||
r7 = r7 ^ r0; // 1 xor
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r6, 1); // 2 shfl
|
||||
r4 = r4 - r1; // 3 sub
|
||||
r2 = r0 * r4 + r2; // 4 mad
|
||||
r4 = rotl_imm(r4, 9u); // 5 rotl
|
||||
r0 = rotr_var(r0, r2); // 6 rotr
|
||||
r6 = r6 ^ ds[((rotl_imm(r7 * 0x03ac37adu, 22u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r2 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r0 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r1 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
|
||||
r5 = r5 * r4; // 12 mul
|
||||
r7 = r7 ^ ds[((rotl_imm(r6 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
|
||||
r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r5, 16); // 14 shfl
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r5, 1); // 15 shfl
|
||||
r2 = r2 | r7; // 16 or
|
||||
r0 = rotr_var(r0, r6); // 17 rotr
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
|
||||
r7 = rotl_imm(r7, 24u); // 19 rotl
|
||||
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
|
||||
r2 = r2 | r6; // 21 or
|
||||
r6 = r5 * r4 + r6; // 22 mad
|
||||
r1 = r1 | r0; // 23 or
|
||||
r6 = r6 ^ r1; // 24 xor
|
||||
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
|
||||
r7 = rotr_var(r7, r0); // 26 rotr
|
||||
r4 = r5 * r7 + r4; // 27 mad
|
||||
r3 = r2 * r4 + r3; // 28 mad
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
|
||||
r2 = r2 ^ ds[((rotl_imm(r3 * 0x03ac37adu, 22u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r5 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
|
||||
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
|
||||
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
|
||||
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
|
||||
r2 = __umulhi(r2, r5); // 35 mulhi
|
||||
r4 = r4 ^ r2; // 36 xor
|
||||
r6 = r6 * r5; // 37 mul
|
||||
r7 = r7 ^ r0; // 38 xor
|
||||
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
|
||||
r2 = rotr_var(r2, r3); // 40 rotr
|
||||
r7 = r7 - r0; // 41 sub
|
||||
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 43 shfl
|
||||
r0 = r0 ^ ds[((rotl_imm(r7 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
|
||||
r3 = r3 ^ ds[((rotl_imm(r1 * 0x03ac37adu, 22u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
|
||||
r6 = r6 ^ ds[((rotl_imm(r3 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
|
||||
r4 = __umulhi(r4, r2); // 48 mulhi
|
||||
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
|
||||
r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r7, 4); // 50 shfl
|
||||
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r2 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
|
||||
r6 = rotl_imm(r6, 12u); // 53 rotl
|
||||
r3 = rotl_imm(r3, 12u); // 54 rotl
|
||||
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
|
||||
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r0 * 0x03ac37adu, 22u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
|
||||
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
|
||||
r3 = __umulhi(r3, r2); // 60 mulhi
|
||||
r6 = r6 | r4; // 61 or
|
||||
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 62 shfl
|
||||
r3 = r3 ^ ds[((rotl_imm(r6 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
|
||||
}
|
||||
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps) {
|
||||
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 32u * blockWarps;
|
||||
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
|
||||
igneum_hash_bound<<<nonces / block, block>>>(ds, out, baseNonce, mask, iw);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
|
||||
cudaFuncAttributes attr;
|
||||
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash_bound);
|
||||
if (e != cudaSuccess) return e;
|
||||
*numRegs = attr.numRegs;
|
||||
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash_bound, (int)(32u * blockWarps), 0);
|
||||
}
|
||||
112
proto-cuda/packs-ca2-era/era-5/memhard.h
Normal file
112
proto-cuda/packs-ca2-era/era-5/memhard.h
Normal file
|
|
@ -0,0 +1,112 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// Memory-hard dataset core, the same text that the Mac's Metal kernels and CPU verifier were checked against.
|
||||
// Included by kernel.cu (device), host.cu (host reference) and proto-opencl/host.c (C99 host reference).
|
||||
// See proto-metal/MEMHARD.md for the construction. kernel.cl carries the same text in OpenCL C.
|
||||
#pragma once
|
||||
#ifdef __cplusplus
|
||||
#include <cstdint>
|
||||
#else
|
||||
#include <stdint.h>
|
||||
#endif
|
||||
#if defined(__CUDACC__)
|
||||
#define IGNEUM_HD __host__ __device__ __forceinline__
|
||||
#elif defined(_MSC_VER) && !defined(__cplusplus)
|
||||
#define IGNEUM_HD static __inline
|
||||
#else
|
||||
#define IGNEUM_HD static inline
|
||||
#endif
|
||||
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
|
||||
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
|
||||
#define MH_CACHE_LINE_MASK 0x003fffffu
|
||||
#define MH_SEGMENT_LINES 64u
|
||||
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
|
||||
IGNEUM_HD uint32_t mh_rotl(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
|
||||
|
||||
// y = ChaCha12 core(x) + x
|
||||
IGNEUM_HD void mh_chacha_block(const uint32_t* x, uint32_t* y) {
|
||||
for (uint32_t i = 0u; i < 16u; ++i) y[i] = x[i];
|
||||
for (uint32_t r = 0u; r < 6u; ++r) {
|
||||
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
|
||||
}
|
||||
for (uint32_t i = 0u; i < 16u; ++i) y[i] += x[i];
|
||||
}
|
||||
|
||||
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
|
||||
IGNEUM_HD void mh_cache_segment(uint32_t* cache, uint32_t seg) {
|
||||
uint32_t prev[16]; uint32_t x[16]; uint32_t y[16];
|
||||
for (uint32_t i = 0u; i < 16u; ++i) prev[i] = 0u;
|
||||
for (uint32_t j = 0u; j < MH_SEGMENT_LINES; ++j) {
|
||||
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
|
||||
x[4] = 0xceed56d7u ^ prev[4];
|
||||
x[5] = 0x9ba270d2u ^ prev[5];
|
||||
x[6] = 0x82caab2du ^ prev[6];
|
||||
x[7] = 0x81ebce0eu ^ prev[7];
|
||||
x[8] = 0x12b6ecf1u ^ prev[8];
|
||||
x[9] = 0xd0f3fd7cu ^ prev[9];
|
||||
x[10] = 0xd872eefeu ^ prev[10];
|
||||
x[11] = 0xc158c7bdu ^ prev[11];
|
||||
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
|
||||
mh_chacha_block(x, y);
|
||||
uint32_t* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
|
||||
for (uint32_t i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
|
||||
}
|
||||
}
|
||||
|
||||
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
|
||||
IGNEUM_HD void mh_mixer(uint32_t* s, uint32_t rk) {
|
||||
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
|
||||
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
|
||||
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
|
||||
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
|
||||
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
|
||||
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
|
||||
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
|
||||
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
|
||||
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
|
||||
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
|
||||
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
|
||||
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
|
||||
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
|
||||
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
|
||||
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
|
||||
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
|
||||
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
|
||||
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
|
||||
}
|
||||
|
||||
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
|
||||
IGNEUM_HD void mh_item(const uint32_t* cache, uint32_t t, uint32_t* s) {
|
||||
s[0] = 0xceed56d7u;
|
||||
s[1] = 0x9ba270d2u;
|
||||
s[2] = 0x82caab2du;
|
||||
s[3] = 0x81ebce0eu;
|
||||
s[4] = 0x12b6ecf1u;
|
||||
s[5] = 0xd0f3fd7cu;
|
||||
s[6] = 0xd872eefeu;
|
||||
s[7] = 0xc158c7bdu;
|
||||
s[8] = t * 0xf351d601u + 0xc6892460u;
|
||||
s[9] = t * 0xa3bb398fu + 0x25b7228au;
|
||||
s[10] = t * 0xb5a09e35u + 0xcd515004u;
|
||||
s[11] = t * 0x7509c9c1u + 0x2846527au;
|
||||
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
|
||||
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
|
||||
s[14] = t * 0xded91851u + 0x82961bacu;
|
||||
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
|
||||
for (uint32_t r = 0u; r < 8u; ++r) {
|
||||
mh_mixer(s, 0x9E3779B9u * (r + 1u));
|
||||
const uint32_t* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
|
||||
for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= line[i];
|
||||
}
|
||||
mh_mixer(s, 0x9E3779B9u * 9u);
|
||||
}
|
||||
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 0 2 10 13 of w.
|
||||
IGNEUM_HD uint32_t mh_j(uint32_t w) { return ((w >> 0u) & 1u) | (((w >> 2u) & 1u) << 1) | (((w >> 10u) & 1u) << 2) | (((w >> 13u) & 1u) << 3); }
|
||||
IGNEUM_HD uint32_t mh_t(uint32_t w) { w = (w & 0x00001fffu) | ((w >> 14u) << 13u); w = (w & 0x000003ffu) | ((w >> 11u) << 10u); w = (w & 0x00000003u) | ((w >> 3u) << 2u); w = (w & 0x00000000u) | ((w >> 1u) << 0u); return w; }
|
||||
IGNEUM_HD uint32_t mh_addr(uint32_t t, uint32_t j) { uint32_t w = t; w = ((w >> 0u) << 1u) | (w & 0x00000000u) | (((j >> 0u) & 1u) << 0u); w = ((w >> 2u) << 3u) | (w & 0x00000003u) | (((j >> 1u) & 1u) << 2u); w = ((w >> 10u) << 11u) | (w & 0x000003ffu) | (((j >> 2u) & 1u) << 10u); w = ((w >> 13u) << 14u) | (w & 0x00001fffu) | (((j >> 3u) & 1u) << 13u); return w; }
|
||||
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
|
||||
IGNEUM_HD uint32_t mh_word(const uint32_t* cache, uint32_t w) { uint32_t s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }
|
||||
109
proto-cuda/packs-ca2-era/era-5/memhard.metal
Normal file
109
proto-cuda/packs-ca2-era/era-5/memhard.metal
Normal file
|
|
@ -0,0 +1,109 @@
|
|||
#include <metal_stdlib>
|
||||
using namespace metal;
|
||||
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
|
||||
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
|
||||
#define MH_CACHE_LINE_MASK 0x003fffffu
|
||||
#define MH_SEGMENT_LINES 64u
|
||||
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
|
||||
inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
|
||||
|
||||
// y = ChaCha12 core(x) + x
|
||||
inline void mh_chacha_block(const thread uint* x, thread uint* y) {
|
||||
for (uint i = 0u; i < 16u; ++i) y[i] = x[i];
|
||||
for (uint r = 0u; r < 6u; ++r) {
|
||||
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
|
||||
}
|
||||
for (uint i = 0u; i < 16u; ++i) y[i] += x[i];
|
||||
}
|
||||
|
||||
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
|
||||
inline void mh_cache_segment(device uint* cache, uint seg) {
|
||||
uint prev[16]; uint x[16]; uint y[16];
|
||||
for (uint i = 0u; i < 16u; ++i) prev[i] = 0u;
|
||||
for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) {
|
||||
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
|
||||
x[4] = 0xceed56d7u ^ prev[4];
|
||||
x[5] = 0x9ba270d2u ^ prev[5];
|
||||
x[6] = 0x82caab2du ^ prev[6];
|
||||
x[7] = 0x81ebce0eu ^ prev[7];
|
||||
x[8] = 0x12b6ecf1u ^ prev[8];
|
||||
x[9] = 0xd0f3fd7cu ^ prev[9];
|
||||
x[10] = 0xd872eefeu ^ prev[10];
|
||||
x[11] = 0xc158c7bdu ^ prev[11];
|
||||
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
|
||||
mh_chacha_block(x, y);
|
||||
device uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
|
||||
}
|
||||
}
|
||||
|
||||
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
|
||||
inline void mh_mixer(thread uint* s, uint rk) {
|
||||
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
|
||||
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
|
||||
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
|
||||
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
|
||||
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
|
||||
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
|
||||
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
|
||||
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
|
||||
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
|
||||
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
|
||||
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
|
||||
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
|
||||
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
|
||||
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
|
||||
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
|
||||
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
|
||||
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
|
||||
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
|
||||
}
|
||||
|
||||
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
|
||||
inline void mh_item(device const uint* cache, uint t, thread uint* s) {
|
||||
s[0] = 0xceed56d7u;
|
||||
s[1] = 0x9ba270d2u;
|
||||
s[2] = 0x82caab2du;
|
||||
s[3] = 0x81ebce0eu;
|
||||
s[4] = 0x12b6ecf1u;
|
||||
s[5] = 0xd0f3fd7cu;
|
||||
s[6] = 0xd872eefeu;
|
||||
s[7] = 0xc158c7bdu;
|
||||
s[8] = t * 0xf351d601u + 0xc6892460u;
|
||||
s[9] = t * 0xa3bb398fu + 0x25b7228au;
|
||||
s[10] = t * 0xb5a09e35u + 0xcd515004u;
|
||||
s[11] = t * 0x7509c9c1u + 0x2846527au;
|
||||
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
|
||||
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
|
||||
s[14] = t * 0xded91851u + 0x82961bacu;
|
||||
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
|
||||
for (uint r = 0u; r < 8u; ++r) {
|
||||
mh_mixer(s, 0x9E3779B9u * (r + 1u));
|
||||
device const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i];
|
||||
}
|
||||
mh_mixer(s, 0x9E3779B9u * 9u);
|
||||
}
|
||||
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 0 2 10 13 of w.
|
||||
inline uint mh_j(uint w) { return ((w >> 0u) & 1u) | (((w >> 2u) & 1u) << 1) | (((w >> 10u) & 1u) << 2) | (((w >> 13u) & 1u) << 3); }
|
||||
inline uint mh_t(uint w) { w = (w & 0x00001fffu) | ((w >> 14u) << 13u); w = (w & 0x000003ffu) | ((w >> 11u) << 10u); w = (w & 0x00000003u) | ((w >> 3u) << 2u); w = (w & 0x00000000u) | ((w >> 1u) << 0u); return w; }
|
||||
inline uint mh_addr(uint t, uint j) { uint w = t; w = ((w >> 0u) << 1u) | (w & 0x00000000u) | (((j >> 0u) & 1u) << 0u); w = ((w >> 2u) << 3u) | (w & 0x00000003u) | (((j >> 1u) & 1u) << 2u); w = ((w >> 10u) << 11u) | (w & 0x000003ffu) | (((j >> 2u) & 1u) << 10u); w = ((w >> 13u) << 14u) | (w & 0x00001fffu) | (((j >> 3u) & 1u) << 13u); return w; }
|
||||
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
|
||||
inline uint mh_word(device const uint* cache, uint w) { uint s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }
|
||||
|
||||
// One thread per segment (2^16 threads).
|
||||
kernel void igneum_cache_fill(device uint* cache [[buffer(0)]], uint gid [[thread_position_in_grid]]) {
|
||||
mh_cache_segment(cache, gid);
|
||||
}
|
||||
// One thread per 64-byte item (dataset words / 16 threads).
|
||||
kernel void igneum_build(device const uint* cache [[buffer(0)]], device uint* dataset [[buffer(1)]],
|
||||
uint gid [[thread_position_in_grid]]) {
|
||||
uint s[16];
|
||||
mh_item(cache, gid, s);
|
||||
for (uint i = 0u; i < 16u; ++i) dataset[mh_addr(gid, i)] = s[i];
|
||||
}
|
||||
76
proto-cuda/packs-ca2-era/era-5/program.h
Normal file
76
proto-cuda/packs-ca2-era/era-5/program.h
Normal file
|
|
@ -0,0 +1,76 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// Program metadata for host.cu plus the launch wrappers defined in kernel.cu.
|
||||
// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros.
|
||||
#pragma once
|
||||
#ifdef __cplusplus
|
||||
#include <cstdint>
|
||||
#else
|
||||
#include <stdint.h>
|
||||
#endif
|
||||
#ifndef IGNEUM_NO_CUDA
|
||||
#include <cuda_runtime.h>
|
||||
#endif
|
||||
|
||||
#define IGNEUM_SEED_STRING "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000"
|
||||
#define IGNEUM_SEED_BYTES_HEX "edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07"
|
||||
#define IGNEUM_GENERATOR 3
|
||||
#define IGNEUM_PROGRAM_ATTEMPT 0
|
||||
#define IGNEUM_PROGRAM_ID 0x73bcbfe8ccf988f1ull
|
||||
#define IGNEUM_DAY_STRING "bytes:69676e65756d2d6461792ffa50000000000000"
|
||||
#define IGNEUM_DAY_BYTES_HEX "69676e65756d2d6461792ffa50000000000000"
|
||||
#define IGNEUM_DAY0 0xceed56d7u
|
||||
#define IGNEUM_DAY1 0x9ba270d2u
|
||||
#define IGNEUM_DATASET_LOG2 28
|
||||
#define IGNEUM_MASK 0x0fffffffu
|
||||
#define IGNEUM_LANES 32
|
||||
#define IGNEUM_ITERATIONS 8
|
||||
#define IGNEUM_INSTR_COUNT 64
|
||||
#define IGNEUM_LOADS_PER_HASH 128
|
||||
#define IGNEUM_WIDE_LOADS_PER_HASH 0
|
||||
#define IGNEUM_OP_MIX "load=16 add=15 shfl=6 mad=4 or=4 rotl=4 rotr=4 xor=4 mulhi=3 mul=2 sub=2"
|
||||
// Program class v3 (Counter ASIC 2.0, docs/plans/counter-asic-2-rollout.md): generator version 3; a worker that
|
||||
// runs another class refuses this pack, and a job line names the class it wants (class=v3 era=<hex>).
|
||||
#define IGNEUM_PROGRAM_CLASS "v3"
|
||||
#define IGNEUM_ERA_SEED_HEX "7f450623297a954f493ca08abcdfe7142bb753900abcee38256266007b4b48b7"
|
||||
// Class v3 construction (Counter ASIC 2.0, 5 October 2026, docs/plans/mixer-x4.md): version 2 loads; the dataset item
|
||||
// derivation applies the mixer IGNEUM_MIXER_MULT times per round (memhard.h), and the cache follows the growth rule.
|
||||
#define IGNEUM_LOAD_CLASS "w4-eraf897c84e"
|
||||
#define IGNEUM_LOAD_SLOTS 16
|
||||
#define IGNEUM_LOAD_MIX { 100, 0, 0 }
|
||||
#define IGNEUM_LOAD_WIDTH_COUNTS { 16, 0, 0 } // loads of 4, 16, 64 bytes per program
|
||||
#define IGNEUM_BYTES_PER_HASH 512
|
||||
#define IGNEUM_FOLD_ROT 11
|
||||
#define IGNEUM_FOLD_MUL 0x9e3779b1u
|
||||
// Era layout (5 October 2026, docs/plans/era-layout.md): NOT the lottery hash. Every dataset load reads
|
||||
// idx = ((rotl(src * STRIDE_MUL, STRIDE_ROT) & window mask) | window offset) & MASK; the window of a load site is the
|
||||
// dataset, a half or a quarter of it (IGNEUM_ERA_WINDOWS: site:shrink:offset); dataset word w holds word j(w) of item
|
||||
// t(w) with j's bits at the INTERLEAVE positions (memhard.h: mh_t, mh_j, mh_addr).
|
||||
#define IGNEUM_ERA_LABEL "f897c84e"
|
||||
#define IGNEUM_ERA_SEED_WORDS { 0xf897c84eu, 0xa4296ce5u, 0x3fb7ee16u, 0xd2354d94u, 0x82937d5fu, 0xe8f3ff88u, 0xe41af6f4u, 0x4f10d757u }
|
||||
#define IGNEUM_ERA_ALLOWED_WIDTHS { 1, 0, 0 } // words, ascending, 0 = unused; one entry pins the width
|
||||
#define IGNEUM_ERA_WIDTH_WORDS 1
|
||||
#define IGNEUM_ERA_STRIDE_MUL 0x03ac37adu
|
||||
#define IGNEUM_ERA_STRIDE_ROT 22
|
||||
#define IGNEUM_ERA_INTERLEAVE { 0, 2, 10, 13 }
|
||||
#define IGNEUM_ERA_WINDOWS "7:2:1 8:1:1 9:1:1 10:1:1 11:0:0 13:1:1 29:0:0 30:2:2 31:1:1 44:1:1 46:2:0 47:0:0 52:0:0 56:0:0 58:2:0 63:1:1"
|
||||
// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h)
|
||||
#define IGNEUM_DATASET_MODE 1
|
||||
|
||||
#define IGNEUM_SEEDW_INIT { 0x667d0fbdu, 0x7b8e5963u, 0x31c67e5eu, 0x4529ddc6u, 0xef19d6d8u, 0xaccf6211u, 0xda0aed32u, 0xabc6df31u }
|
||||
#define IGNEUM_KEY_INIT { 0xceed56d7u, 0x9ba270d2u, 0x82caab2du, 0x81ebce0eu, 0x12b6ecf1u, 0xd0f3fd7cu, 0xd872eefeu, 0xc158c7bdu }
|
||||
#define IGNEUM_CACHE_LOG2_WORDS 26
|
||||
#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6
|
||||
#define IGNEUM_CACHE_SEGMENTS 65536u
|
||||
#define IGNEUM_ITEM_ROUNDS 8
|
||||
#define IGNEUM_MIX_ROT_INIT { 17u, 12u, 20u, 23u, 7u, 3u, 27u, 16u }
|
||||
#define IGNEUM_MIX_MUL_INIT { 0xf351d601u, 0xa3bb398fu, 0xb5a09e35u, 0x7509c9c1u, 0x6bbf31e9u, 0xfc849a79u, 0xded91851u, 0x8d9113d1u, 0x0ff15225u, 0x3a5bdd41u, 0xab533435u, 0xe1c55ad5u, 0xe6d3bd0du, 0x9d9ffbbdu, 0xbb2a3cf3u, 0x50a7c08du }
|
||||
#define IGNEUM_MIX_RC_INIT { 0xc6892460u, 0x25b7228au, 0xcd515004u, 0x2846527au, 0xa6324241u, 0x36e3ec53u, 0x82961bacu, 0x0f97ba7du, 0xb6f921a9u, 0x3ada24e5u, 0xde20ab91u, 0x5378eeb2u, 0x7d161662u, 0x89353cc1u, 0xb1aa03a2u, 0x788acae6u }
|
||||
|
||||
#ifndef IGNEUM_NO_CUDA
|
||||
// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError().
|
||||
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments);
|
||||
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems);
|
||||
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
uint32_t nonces, uint32_t blockWarps);
|
||||
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
|
||||
#endif
|
||||
143
proto-cuda/packs-ca2-era/era-5/program.json
Normal file
143
proto-cuda/packs-ca2-era/era-5/program.json
Normal file
|
|
@ -0,0 +1,143 @@
|
|||
{
|
||||
"format": "igneum-program-pack-3",
|
||||
"generator": 3,
|
||||
"attempt": 0,
|
||||
"program_id": "0x73bcbfe8ccf988f1",
|
||||
"program_id_derivation": "FNV-1a 64 over 'igneum-program/' || generator_le32 || seed_words as little-endian bytes || attempt_le32",
|
||||
"dataset_mode": "memory-hard",
|
||||
"seed": "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000",
|
||||
"seed_bytes": "edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07",
|
||||
"seed_words": ["0x667d0fbd", "0x7b8e5963", "0x31c67e5e", "0x4529ddc6", "0xef19d6d8", "0xaccf6211", "0xda0aed32", "0xabc6df31"],
|
||||
"seed_derivation": "seed_words = FNV-1a 64 over seed_bytes (attempt 0) or seed_bytes || attempt_le32 (attempt k >= 1), basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32",
|
||||
"generator_rule": "version 2: exactly 16 load slots drawn first from instructions 1..63 (partial Fisher-Yates), the other 48 ops from the ten non-load weights (sum 75); a load's source is drawn from the registers other than dst written by an earlier instruction and not read by a load since; the candidate must pass the acceptance rule of spec 01 section 1.4.6 (static: no cyclically stale load source, every register has an injecting write; dynamic: 64 units on the seed-keyed closed-form dataset with no constant register bit, no lane-constant load site, under 164 saturated final values, every output bit within 136 of 1024, distinct addresses above 245760), else the next attempt of the seed is tried",
|
||||
"lanes": 32,
|
||||
"registers": 8,
|
||||
"iterations": 8,
|
||||
"instruction_count": 64,
|
||||
"loads_per_hash": 128,
|
||||
"program_class": "v3",
|
||||
"era_seed_bytes": "7f450623297a954f493ca08abcdfe7142bb753900abcee38256266007b4b48b7",
|
||||
"load_class": "w4-eraf897c84e",
|
||||
"load_slots": 16,
|
||||
"load_mix_percent_4_16_64": [100, 0, 0],
|
||||
"load_width_counts_4_16_64": [16, 0, 0],
|
||||
"bytes_per_hash": 512,
|
||||
"wide_load": "read-width experiment (5 October 2026, docs/plans/read-width.md), NOT the lottery hash: a load of W words (width field, 4 or 16) reads dataset[b .. b + W) with b = (src & mask) & ~(W - 1) and folds every word into dst: x = dst ^ w[0]; for j in 1..W: x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x; width 1 is the plain load; the width is drawn per instruction from the class mix with one extra below(100) draw after the nine of version 2, and the program id is FNV-1a 64 over 'igneum-program-rw/' || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots",
|
||||
"era": {
|
||||
"label": "f897c84e",
|
||||
"seed_words": ["0xf897c84e", "0xa4296ce5", "0x3fb7ee16", "0xd2354d94", "0x82937d5f", "0xe8f3ff88", "0xe41af6f4", "0x4f10d757"],
|
||||
"draw": "docs/plans/era-layout.md 1.1: SplitMix64 seeded with seed_words[0] | seed_words[1] << 32 of seed_words_from_bytes('igneum-era/' || n_le64 || E_n); width = allowed[below(|allowed|)], stride_mul = low32(next()) | 1, stride_rot = 1 + below(31), then four next() draws for a partial Fisher-Yates over positions log2(W)..15 of which 4 - log2(W) are used",
|
||||
"allowed_widths": [1],
|
||||
"width_words": 1,
|
||||
"stride_mul": "0x03ac37ad",
|
||||
"stride_rot": 22,
|
||||
"interleave": [0, 2, 10, 13],
|
||||
"address": "y = rotl(src * stride_mul, stride_rot); k = min(win, D - 26); idx = ((y & (mask >> k)) | ((off & (2^k - 1)) << (D - k))) & mask; a wide load aligns idx down to W words",
|
||||
"windows": "per instruction, after the width roll: win = below(3), off = low32(next()) & (2^win - 1); used on a load slot (the instruction's win and off fields)",
|
||||
"dataset_word": "dataset[w] = item(t(w))[j(w)]: j(w) gathers the bits of w at the interleave positions, t(w) is w with those bits removed",
|
||||
"program_id_suffix": "'era/' || allowed[3] || width_words || stride_mul_le32 || stride_rot_le32 || interleave[4]"
|
||||
},
|
||||
"op_mix": {"load": 16, "add": 15, "shfl": 6, "mad": 4, "or": 4, "rotl": 4, "rotr": 4, "xor": 4, "mulhi": 3, "mul": 2, "sub": 2},
|
||||
"register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]",
|
||||
"splitmix32": "x ^= x>>16; x *= 0x7feb352d; x ^= x>>15; x *= 0x846ca68b; x ^= x>>16",
|
||||
"iteration": "sel = r0 sampled once at the top of each iteration, then all instructions in order",
|
||||
"output": "lo = r0 ^ rotl(r1,7) ^ rotl(r2,14) ^ rotl(r3,21); hi = r4 ^ rotl(r5,9) ^ rotl(r6,18) ^ rotl(r7,27); out = (hi << 32) | lo",
|
||||
"op_semantics": {
|
||||
"add": "dst = dst + src + (bit `bit` of sel ? imm2 : imm)",
|
||||
"sub": "dst = dst - src",
|
||||
"mul": "dst = dst * src (low 32)",
|
||||
"mulhi": "dst = high 32 bits of dst * src",
|
||||
"xor": "dst = dst ^ src",
|
||||
"or": "dst = dst | src",
|
||||
"rotl": "dst = rotl(dst, rot), rot in 1..31",
|
||||
"rotr": "dst = rotr(dst, src & 31)",
|
||||
"mad": "dst = src * src2 + dst",
|
||||
"shfl": "dst = dst ^ (src of lane (lane ^ mask)), mask in {1,2,4,8,16}, within the 32-lane warp",
|
||||
"load": "dst = dst ^ dataset[src & dataset.mask]",
|
||||
"wload": "base = (src of lane 0 & dataset.mask) & ~31; dst = dst ^ dataset[base + lane] (warp-coalesced 128-byte load, lever b, only when --wide-frac > 0)"
|
||||
},
|
||||
"dataset": {
|
||||
"log2_words": 28,
|
||||
"bytes": 1073741824,
|
||||
"mask": "0x0fffffff",
|
||||
"day": "bytes:69676e65756d2d6461792ffa50000000000000",
|
||||
"day_bytes": "69676e65756d2d6461792ffa50000000000000",
|
||||
"day_words_from": "seed_words_from_bytes(day_bytes)",
|
||||
"d0": "0xceed56d7",
|
||||
"d1": "0x9ba270d2",
|
||||
"mode": "memory-hard",
|
||||
"spec": "proto-metal/MEMHARD.md",
|
||||
"key": ["0xceed56d7", "0x9ba270d2", "0x82caab2d", "0x81ebce0e", "0x12b6ecf1", "0xd0f3fd7c", "0xd872eefe", "0xc158c7bd"],
|
||||
"key_derivation": "the 8 words of seed_words_from_bytes(day_bytes); d0, d1 are key[0], key[1]",
|
||||
"cache": {"log2_words": 26, "bytes": 268435456, "line_words": 16, "segment_lines": 64, "segments": 65536, "block": "ChaCha12 core + feed-forward, rotations 16 12 8 7", "sigma": ["0x61707865", "0x3320646e", "0x79622d32", "0x6b206574"], "tag": ["0x49676e65", "0x756d4d48"], "chain": "in_j = prev_line ^ (sigma[0..3] || key[0..7] || seg || j || tag[0..1]); line_j = block(in_j); prev_0 = 0"},
|
||||
"mixer": {"draw": "SplitMix64 seeded with key[0] | key[1] << 32: rot[0..7] = 1 + next() % 31, mul[0..15] = low32(next()) | 1, rc[0..15] = low32(next())", "rot": [17, 12, 20, 23, 7, 3, 27, 16], "mul": ["0xf351d601", "0xa3bb398f", "0xb5a09e35", "0x7509c9c1", "0x6bbf31e9", "0xfc849a79", "0xded91851", "0x8d9113d1", "0x0ff15225", "0x3a5bdd41", "0xab533435", "0xe1c55ad5", "0xe6d3bd0d", "0x9d9ffbbd", "0xbb2a3cf3", "0x50a7c08d"], "rc": ["0xc6892460", "0x25b7228a", "0xcd515004", "0x2846527a", "0xa6324241", "0x36e3ec53", "0x82961bac", "0x0f97ba7d", "0xb6f921a9", "0x3ada24e5", "0xde20ab91", "0x5378eeb2", "0x7d161662", "0x89353cc1", "0xb1aa03a2", "0x788acae6"], "round": "for i in 0..15: s[i] = (s[i] ^ (rc[i] + (r+1) * 0x9E3779B9)) * mul[i]; then quarter rounds on columns (0,4,8,12) (1,5,9,13) (2,6,10,14) (3,7,11,15) with rot[0..3] and diagonals (0,5,10,15) (1,6,11,12) (2,7,8,13) (3,4,9,14) with rot[4..7]", "quarter_round": "a += b; d ^= a; d = rotl(d, r1); c += d; b ^= c; b = rotl(b, r2); a += b; d ^= a; d = rotl(d, r3); c += d; b ^= c; b = rotl(b, r4)"},
|
||||
"item": "s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..7: s = M_r(s); line = s[0] & 0x003fffff; s[i] ^= cache[line * 16 + i]; then s = M_8(s); item(t) = s",
|
||||
"word": "dataset[w] = item(w >> 4)[w & 15]"
|
||||
},
|
||||
"instructions": [
|
||||
{"i": 0, "op": "add", "dst": 4, "src": 5, "src2": 7, "imm": "0xea86e152", "imm2": "0x5810667a", "rot": 27, "bit": 13, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 1, "op": "xor", "dst": 7, "src": 0, "src2": 6, "imm": "0xbaab6229", "imm2": "0xed861989", "rot": 26, "bit": 22, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 2, "op": "shfl", "dst": 3, "src": 6, "src2": 2, "imm": "0x5b623116", "imm2": "0xff12e5b2", "rot": 12, "bit": 24, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 3, "op": "sub", "dst": 4, "src": 1, "src2": 1, "imm": "0xe99741c7", "imm2": "0xf5fa5009", "rot": 1, "bit": 21, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 4, "op": "mad", "dst": 2, "src": 0, "src2": 4, "imm": "0x673c2157", "imm2": "0xee02465f", "rot": 20, "bit": 22, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 5, "op": "rotl", "dst": 4, "src": 7, "src2": 7, "imm": "0x946f7818", "imm2": "0x45d3399e", "rot": 9, "bit": 2, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 6, "op": "rotr", "dst": 0, "src": 2, "src2": 4, "imm": "0x5f6a0ed2", "imm2": "0x7043a636", "rot": 19, "bit": 31, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 7, "op": "load", "dst": 6, "src": 7, "src2": 1, "imm": "0x5c61dcf7", "imm2": "0x7466aa40", "rot": 19, "bit": 9, "mask": 2, "width": 1, "win": 2, "off": 1},
|
||||
{"i": 8, "op": "load", "dst": 1, "src": 4, "src2": 5, "imm": "0x85668475", "imm2": "0xdb8cc483", "rot": 29, "bit": 7, "mask": 4, "width": 1, "win": 1, "off": 1},
|
||||
{"i": 9, "op": "load", "dst": 1, "src": 2, "src2": 4, "imm": "0x4454980f", "imm2": "0xebd31581", "rot": 10, "bit": 28, "mask": 16, "width": 1, "win": 1, "off": 1},
|
||||
{"i": 10, "op": "load", "dst": 7, "src": 0, "src2": 2, "imm": "0xe075297c", "imm2": "0x5779c44c", "rot": 10, "bit": 22, "mask": 2, "width": 1, "win": 1, "off": 1},
|
||||
{"i": 11, "op": "load", "dst": 7, "src": 1, "src2": 6, "imm": "0x65aa4311", "imm2": "0x4fe48ea9", "rot": 15, "bit": 9, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 12, "op": "mul", "dst": 5, "src": 4, "src2": 2, "imm": "0x1383d3ad", "imm2": "0xf3094b29", "rot": 8, "bit": 9, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 13, "op": "load", "dst": 7, "src": 6, "src2": 2, "imm": "0xed8a496f", "imm2": "0x3072c3c6", "rot": 28, "bit": 19, "mask": 8, "width": 1, "win": 1, "off": 1},
|
||||
{"i": 14, "op": "shfl", "dst": 6, "src": 5, "src2": 0, "imm": "0x8b965b57", "imm2": "0xcfeca6c1", "rot": 12, "bit": 27, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 15, "op": "shfl", "dst": 3, "src": 5, "src2": 3, "imm": "0x877c7586", "imm2": "0xa9cb2a03", "rot": 2, "bit": 29, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 16, "op": "or", "dst": 2, "src": 7, "src2": 3, "imm": "0xb740221a", "imm2": "0x89d38d6d", "rot": 6, "bit": 31, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 17, "op": "rotr", "dst": 0, "src": 6, "src2": 1, "imm": "0x26f3ad8a", "imm2": "0x27256f15", "rot": 18, "bit": 5, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 18, "op": "add", "dst": 7, "src": 5, "src2": 7, "imm": "0xf66e7017", "imm2": "0xb9e3577e", "rot": 9, "bit": 12, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 19, "op": "rotl", "dst": 7, "src": 0, "src2": 5, "imm": "0x849ae6ee", "imm2": "0x02b358f9", "rot": 24, "bit": 18, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 20, "op": "add", "dst": 6, "src": 7, "src2": 6, "imm": "0x52334d12", "imm2": "0x8c9f0ef8", "rot": 11, "bit": 23, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 21, "op": "or", "dst": 2, "src": 6, "src2": 5, "imm": "0xb1871e63", "imm2": "0xb2e40191", "rot": 5, "bit": 13, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 22, "op": "mad", "dst": 6, "src": 5, "src2": 4, "imm": "0x97df29e4", "imm2": "0xe60fea84", "rot": 11, "bit": 16, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 23, "op": "or", "dst": 1, "src": 0, "src2": 3, "imm": "0x8f30d21d", "imm2": "0x2df685a0", "rot": 31, "bit": 0, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 24, "op": "xor", "dst": 6, "src": 1, "src2": 2, "imm": "0xd2c4025f", "imm2": "0x5269eb4d", "rot": 31, "bit": 21, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 25, "op": "add", "dst": 2, "src": 6, "src2": 5, "imm": "0x659fc3d3", "imm2": "0x9cec0e12", "rot": 6, "bit": 17, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 26, "op": "rotr", "dst": 7, "src": 0, "src2": 1, "imm": "0xd89ef484", "imm2": "0x20be3846", "rot": 12, "bit": 9, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 27, "op": "mad", "dst": 4, "src": 5, "src2": 7, "imm": "0xb48420ae", "imm2": "0x3d1f2485", "rot": 14, "bit": 28, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 28, "op": "mad", "dst": 3, "src": 2, "src2": 4, "imm": "0xc5c46d76", "imm2": "0x700044b5", "rot": 22, "bit": 10, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 29, "op": "load", "dst": 1, "src": 4, "src2": 3, "imm": "0x0fbaf177", "imm2": "0xfff4f2ed", "rot": 20, "bit": 31, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 30, "op": "load", "dst": 2, "src": 3, "src2": 0, "imm": "0xe90eb2e5", "imm2": "0xb0f9eb79", "rot": 27, "bit": 14, "mask": 4, "width": 1, "win": 2, "off": 2},
|
||||
{"i": 31, "op": "load", "dst": 1, "src": 5, "src2": 2, "imm": "0x97ba3fc3", "imm2": "0x7894e657", "rot": 3, "bit": 30, "mask": 16, "width": 1, "win": 1, "off": 1},
|
||||
{"i": 32, "op": "add", "dst": 7, "src": 2, "src2": 7, "imm": "0x070888a8", "imm2": "0xe403240e", "rot": 2, "bit": 16, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 33, "op": "add", "dst": 2, "src": 0, "src2": 5, "imm": "0xf2e46d55", "imm2": "0x29701828", "rot": 31, "bit": 28, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 34, "op": "add", "dst": 2, "src": 3, "src2": 0, "imm": "0x58f75b87", "imm2": "0x343b7aee", "rot": 12, "bit": 14, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 35, "op": "mulhi", "dst": 2, "src": 5, "src2": 6, "imm": "0x5892a9e6", "imm2": "0xc9824c94", "rot": 19, "bit": 26, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 36, "op": "xor", "dst": 4, "src": 2, "src2": 3, "imm": "0x7b5b5474", "imm2": "0x45cfc5dd", "rot": 17, "bit": 18, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 37, "op": "mul", "dst": 6, "src": 5, "src2": 7, "imm": "0xccf564a5", "imm2": "0x873ad101", "rot": 7, "bit": 11, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 38, "op": "xor", "dst": 7, "src": 0, "src2": 6, "imm": "0xcac8f06d", "imm2": "0x6b97c683", "rot": 18, "bit": 28, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 39, "op": "add", "dst": 7, "src": 2, "src2": 4, "imm": "0xb8180e9d", "imm2": "0x32bbd117", "rot": 23, "bit": 19, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 40, "op": "rotr", "dst": 2, "src": 3, "src2": 0, "imm": "0x2d6070bc", "imm2": "0x68ff101e", "rot": 13, "bit": 18, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 41, "op": "sub", "dst": 7, "src": 0, "src2": 2, "imm": "0x0daf96ea", "imm2": "0x36f37be1", "rot": 5, "bit": 0, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 42, "op": "add", "dst": 4, "src": 3, "src2": 6, "imm": "0x6ced15b7", "imm2": "0x6df7aed4", "rot": 19, "bit": 4, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 43, "op": "shfl", "dst": 7, "src": 3, "src2": 5, "imm": "0x8ace05f3", "imm2": "0xd378ec12", "rot": 23, "bit": 10, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 44, "op": "load", "dst": 0, "src": 7, "src2": 4, "imm": "0xb0607786", "imm2": "0xc4acabbc", "rot": 13, "bit": 7, "mask": 16, "width": 1, "win": 1, "off": 1},
|
||||
{"i": 45, "op": "add", "dst": 1, "src": 6, "src2": 5, "imm": "0xe09f54e9", "imm2": "0x83e825bf", "rot": 23, "bit": 14, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 46, "op": "load", "dst": 3, "src": 1, "src2": 0, "imm": "0x63cc1e4e", "imm2": "0xa1be8118", "rot": 12, "bit": 6, "mask": 2, "width": 1, "win": 2, "off": 0},
|
||||
{"i": 47, "op": "load", "dst": 6, "src": 3, "src2": 7, "imm": "0x353f1d79", "imm2": "0x3b2e7456", "rot": 18, "bit": 18, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 48, "op": "mulhi", "dst": 4, "src": 2, "src2": 7, "imm": "0x00d8a3cd", "imm2": "0x231866d2", "rot": 21, "bit": 20, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 49, "op": "add", "dst": 5, "src": 0, "src2": 2, "imm": "0xa8bae6df", "imm2": "0xf572bdb9", "rot": 14, "bit": 7, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 50, "op": "shfl", "dst": 0, "src": 7, "src2": 7, "imm": "0x81ef22e1", "imm2": "0x74438fc5", "rot": 28, "bit": 18, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 51, "op": "add", "dst": 6, "src": 0, "src2": 6, "imm": "0x383b9260", "imm2": "0x11e17c61", "rot": 12, "bit": 19, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 52, "op": "load", "dst": 5, "src": 2, "src2": 2, "imm": "0xfb84f451", "imm2": "0x11cd863e", "rot": 21, "bit": 20, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 53, "op": "rotl", "dst": 6, "src": 5, "src2": 4, "imm": "0xb1a7db6b", "imm2": "0x76686b9b", "rot": 12, "bit": 4, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 54, "op": "rotl", "dst": 3, "src": 6, "src2": 3, "imm": "0x6f981f52", "imm2": "0xd99aeba2", "rot": 12, "bit": 27, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 55, "op": "add", "dst": 2, "src": 1, "src2": 2, "imm": "0xac6be8e3", "imm2": "0x18d67dbb", "rot": 26, "bit": 10, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 56, "op": "load", "dst": 1, "src": 4, "src2": 0, "imm": "0x7e7f6a00", "imm2": "0x6f0747da", "rot": 25, "bit": 28, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 57, "op": "add", "dst": 5, "src": 0, "src2": 4, "imm": "0xf03673fe", "imm2": "0xa75cd60d", "rot": 16, "bit": 12, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 58, "op": "load", "dst": 5, "src": 0, "src2": 2, "imm": "0x227f94a6", "imm2": "0x0e8344f9", "rot": 20, "bit": 10, "mask": 2, "width": 1, "win": 2, "off": 0},
|
||||
{"i": 59, "op": "add", "dst": 1, "src": 4, "src2": 3, "imm": "0xdecd4794", "imm2": "0x8dfb96bb", "rot": 21, "bit": 7, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 60, "op": "mulhi", "dst": 3, "src": 2, "src2": 2, "imm": "0x0dd268e0", "imm2": "0x53034ca9", "rot": 1, "bit": 8, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 61, "op": "or", "dst": 6, "src": 4, "src2": 7, "imm": "0x3a45a321", "imm2": "0x9bc59a5f", "rot": 25, "bit": 11, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 62, "op": "shfl", "dst": 5, "src": 4, "src2": 2, "imm": "0x8f229cc1", "imm2": "0xcaac64a2", "rot": 17, "bit": 13, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 63, "op": "load", "dst": 3, "src": 6, "src2": 6, "imm": "0xb2574178", "imm2": "0xcbcc798d", "rot": 28, "bit": 0, "mask": 16, "width": 1, "win": 1, "off": 1}
|
||||
]
|
||||
}
|
||||
109
proto-cuda/packs-ca2-era/era-5/program.metal
Normal file
109
proto-cuda/packs-ca2-era/era-5/program.metal
Normal file
|
|
@ -0,0 +1,109 @@
|
|||
#include <metal_stdlib>
|
||||
using namespace metal;
|
||||
|
||||
#define MASK 0x0fffffffu
|
||||
constant uint SEEDW[8] = { 0x667d0fbdu, 0x7b8e5963u, 0x31c67e5eu, 0x4529ddc6u, 0xef19d6d8u, 0xaccf6211u, 0xda0aed32u, 0xabc6df31u };
|
||||
|
||||
inline uint splitmix32(uint x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31
|
||||
inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
inline uint ds_elem(uint i, uint d0, uint d1) {
|
||||
uint x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
kernel void igneum_hash(device const uint* dataset [[buffer(0)]],
|
||||
device ulong* out [[buffer(1)]],
|
||||
constant uint& baseNonce [[buffer(2)]],
|
||||
uint gid [[thread_position_in_grid]]) {
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint x = nonce ^ SEEDW[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ SEEDW[1]; }
|
||||
{ uint x = nonce ^ SEEDW[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ SEEDW[2]; }
|
||||
{ uint x = nonce ^ SEEDW[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ SEEDW[3]; }
|
||||
{ uint x = nonce ^ SEEDW[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ SEEDW[4]; }
|
||||
{ uint x = nonce ^ SEEDW[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ SEEDW[5]; }
|
||||
{ uint x = nonce ^ SEEDW[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ SEEDW[6]; }
|
||||
{ uint x = nonce ^ SEEDW[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ SEEDW[7]; }
|
||||
{ uint x = nonce ^ SEEDW[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ SEEDW[0]; }
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = r4 + r5 + select(0xea86e152u, 0x5810667au, ((sel >> 13u) & 1u) != 0u); // 0
|
||||
r7 = r7 ^ r0; // 1
|
||||
r3 = r3 ^ simd_shuffle_xor(r6, (ushort)1); // 2
|
||||
r4 = r4 - r1; // 3
|
||||
r2 = r0 * r4 + r2; // 4
|
||||
r4 = rotl_imm(r4, 9u); // 5
|
||||
r0 = rotr_var(r0, r2); // 6
|
||||
r6 = r6 ^ dataset[((rotl_imm(r7 * 0x03ac37adu, 22u) & 0x03ffffffu) | 0x04000000u) & MASK]; // 7
|
||||
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 8
|
||||
r1 = r1 ^ dataset[((rotl_imm(r2 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 9
|
||||
r7 = r7 ^ dataset[((rotl_imm(r0 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 10
|
||||
r7 = r7 ^ dataset[((rotl_imm(r1 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 11
|
||||
r5 = r5 * r4; // 12
|
||||
r7 = r7 ^ dataset[((rotl_imm(r6 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 13
|
||||
r6 = r6 ^ simd_shuffle_xor(r5, (ushort)16); // 14
|
||||
r3 = r3 ^ simd_shuffle_xor(r5, (ushort)1); // 15
|
||||
r2 = r2 | r7; // 16
|
||||
r0 = rotr_var(r0, r6); // 17
|
||||
r7 = r7 + r5 + select(0xf66e7017u, 0xb9e3577eu, ((sel >> 12u) & 1u) != 0u); // 18
|
||||
r7 = rotl_imm(r7, 24u); // 19
|
||||
r6 = r6 + r7 + select(0x52334d12u, 0x8c9f0ef8u, ((sel >> 23u) & 1u) != 0u); // 20
|
||||
r2 = r2 | r6; // 21
|
||||
r6 = r5 * r4 + r6; // 22
|
||||
r1 = r1 | r0; // 23
|
||||
r6 = r6 ^ r1; // 24
|
||||
r2 = r2 + r6 + select(0x659fc3d3u, 0x9cec0e12u, ((sel >> 17u) & 1u) != 0u); // 25
|
||||
r7 = rotr_var(r7, r0); // 26
|
||||
r4 = r5 * r7 + r4; // 27
|
||||
r3 = r2 * r4 + r3; // 28
|
||||
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 29
|
||||
r2 = r2 ^ dataset[((rotl_imm(r3 * 0x03ac37adu, 22u) & 0x03ffffffu) | 0x08000000u) & MASK]; // 30
|
||||
r1 = r1 ^ dataset[((rotl_imm(r5 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 31
|
||||
r7 = r7 + r2 + select(0x070888a8u, 0xe403240eu, ((sel >> 16u) & 1u) != 0u); // 32
|
||||
r2 = r2 + r0 + select(0xf2e46d55u, 0x29701828u, ((sel >> 28u) & 1u) != 0u); // 33
|
||||
r2 = r2 + r3 + select(0x58f75b87u, 0x343b7aeeu, ((sel >> 14u) & 1u) != 0u); // 34
|
||||
r2 = mulhi(r2, r5); // 35
|
||||
r4 = r4 ^ r2; // 36
|
||||
r6 = r6 * r5; // 37
|
||||
r7 = r7 ^ r0; // 38
|
||||
r7 = r7 + r2 + select(0xb8180e9du, 0x32bbd117u, ((sel >> 19u) & 1u) != 0u); // 39
|
||||
r2 = rotr_var(r2, r3); // 40
|
||||
r7 = r7 - r0; // 41
|
||||
r4 = r4 + r3 + select(0x6ced15b7u, 0x6df7aed4u, ((sel >> 4u) & 1u) != 0u); // 42
|
||||
r7 = r7 ^ simd_shuffle_xor(r3, (ushort)4); // 43
|
||||
r0 = r0 ^ dataset[((rotl_imm(r7 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 44
|
||||
r1 = r1 + r6 + select(0xe09f54e9u, 0x83e825bfu, ((sel >> 14u) & 1u) != 0u); // 45
|
||||
r3 = r3 ^ dataset[((rotl_imm(r1 * 0x03ac37adu, 22u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 46
|
||||
r6 = r6 ^ dataset[((rotl_imm(r3 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 47
|
||||
r4 = mulhi(r4, r2); // 48
|
||||
r5 = r5 + r0 + select(0xa8bae6dfu, 0xf572bdb9u, ((sel >> 7u) & 1u) != 0u); // 49
|
||||
r0 = r0 ^ simd_shuffle_xor(r7, (ushort)4); // 50
|
||||
r6 = r6 + r0 + select(0x383b9260u, 0x11e17c61u, ((sel >> 19u) & 1u) != 0u); // 51
|
||||
r5 = r5 ^ dataset[((rotl_imm(r2 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 52
|
||||
r6 = rotl_imm(r6, 12u); // 53
|
||||
r3 = rotl_imm(r3, 12u); // 54
|
||||
r2 = r2 + r1 + select(0xac6be8e3u, 0x18d67dbbu, ((sel >> 10u) & 1u) != 0u); // 55
|
||||
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 56
|
||||
r5 = r5 + r0 + select(0xf03673feu, 0xa75cd60du, ((sel >> 12u) & 1u) != 0u); // 57
|
||||
r5 = r5 ^ dataset[((rotl_imm(r0 * 0x03ac37adu, 22u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 58
|
||||
r1 = r1 + r4 + select(0xdecd4794u, 0x8dfb96bbu, ((sel >> 7u) & 1u) != 0u); // 59
|
||||
r3 = mulhi(r3, r2); // 60
|
||||
r6 = r6 | r4; // 61
|
||||
r5 = r5 ^ simd_shuffle_xor(r4, (ushort)8); // 62
|
||||
r3 = r3 ^ dataset[((rotl_imm(r6 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 63
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
111
proto-cuda/packs-ca2-era/era-5/program_bound.metal
Normal file
111
proto-cuda/packs-ca2-era/era-5/program_bound.metal
Normal file
|
|
@ -0,0 +1,111 @@
|
|||
#include <metal_stdlib>
|
||||
using namespace metal;
|
||||
|
||||
#define MASK 0x0fffffffu
|
||||
constant uint SEEDW[8] = { 0x667d0fbdu, 0x7b8e5963u, 0x31c67e5eu, 0x4529ddc6u, 0xef19d6d8u, 0xaccf6211u, 0xda0aed32u, 0xabc6df31u };
|
||||
|
||||
inline uint splitmix32(uint x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31
|
||||
inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
inline uint ds_elem(uint i, uint d0, uint d1) {
|
||||
uint x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// Header-bound variant: the init words come from buffer 3 (bind.rs), not from SEEDW.
|
||||
kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]],
|
||||
device ulong* out [[buffer(1)]],
|
||||
constant uint& baseNonce [[buffer(2)]],
|
||||
constant uint* initw [[buffer(3)]],
|
||||
uint gid [[thread_position_in_grid]]) {
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint x = nonce ^ initw[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ initw[1]; }
|
||||
{ uint x = nonce ^ initw[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ initw[2]; }
|
||||
{ uint x = nonce ^ initw[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ initw[3]; }
|
||||
{ uint x = nonce ^ initw[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ initw[4]; }
|
||||
{ uint x = nonce ^ initw[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ initw[5]; }
|
||||
{ uint x = nonce ^ initw[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ initw[6]; }
|
||||
{ uint x = nonce ^ initw[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ initw[7]; }
|
||||
{ uint x = nonce ^ initw[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ initw[0]; }
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = r4 + r5 + select(0xea86e152u, 0x5810667au, ((sel >> 13u) & 1u) != 0u); // 0
|
||||
r7 = r7 ^ r0; // 1
|
||||
r3 = r3 ^ simd_shuffle_xor(r6, (ushort)1); // 2
|
||||
r4 = r4 - r1; // 3
|
||||
r2 = r0 * r4 + r2; // 4
|
||||
r4 = rotl_imm(r4, 9u); // 5
|
||||
r0 = rotr_var(r0, r2); // 6
|
||||
r6 = r6 ^ dataset[((rotl_imm(r7 * 0x03ac37adu, 22u) & 0x03ffffffu) | 0x04000000u) & MASK]; // 7
|
||||
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 8
|
||||
r1 = r1 ^ dataset[((rotl_imm(r2 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 9
|
||||
r7 = r7 ^ dataset[((rotl_imm(r0 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 10
|
||||
r7 = r7 ^ dataset[((rotl_imm(r1 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 11
|
||||
r5 = r5 * r4; // 12
|
||||
r7 = r7 ^ dataset[((rotl_imm(r6 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 13
|
||||
r6 = r6 ^ simd_shuffle_xor(r5, (ushort)16); // 14
|
||||
r3 = r3 ^ simd_shuffle_xor(r5, (ushort)1); // 15
|
||||
r2 = r2 | r7; // 16
|
||||
r0 = rotr_var(r0, r6); // 17
|
||||
r7 = r7 + r5 + select(0xf66e7017u, 0xb9e3577eu, ((sel >> 12u) & 1u) != 0u); // 18
|
||||
r7 = rotl_imm(r7, 24u); // 19
|
||||
r6 = r6 + r7 + select(0x52334d12u, 0x8c9f0ef8u, ((sel >> 23u) & 1u) != 0u); // 20
|
||||
r2 = r2 | r6; // 21
|
||||
r6 = r5 * r4 + r6; // 22
|
||||
r1 = r1 | r0; // 23
|
||||
r6 = r6 ^ r1; // 24
|
||||
r2 = r2 + r6 + select(0x659fc3d3u, 0x9cec0e12u, ((sel >> 17u) & 1u) != 0u); // 25
|
||||
r7 = rotr_var(r7, r0); // 26
|
||||
r4 = r5 * r7 + r4; // 27
|
||||
r3 = r2 * r4 + r3; // 28
|
||||
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 29
|
||||
r2 = r2 ^ dataset[((rotl_imm(r3 * 0x03ac37adu, 22u) & 0x03ffffffu) | 0x08000000u) & MASK]; // 30
|
||||
r1 = r1 ^ dataset[((rotl_imm(r5 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 31
|
||||
r7 = r7 + r2 + select(0x070888a8u, 0xe403240eu, ((sel >> 16u) & 1u) != 0u); // 32
|
||||
r2 = r2 + r0 + select(0xf2e46d55u, 0x29701828u, ((sel >> 28u) & 1u) != 0u); // 33
|
||||
r2 = r2 + r3 + select(0x58f75b87u, 0x343b7aeeu, ((sel >> 14u) & 1u) != 0u); // 34
|
||||
r2 = mulhi(r2, r5); // 35
|
||||
r4 = r4 ^ r2; // 36
|
||||
r6 = r6 * r5; // 37
|
||||
r7 = r7 ^ r0; // 38
|
||||
r7 = r7 + r2 + select(0xb8180e9du, 0x32bbd117u, ((sel >> 19u) & 1u) != 0u); // 39
|
||||
r2 = rotr_var(r2, r3); // 40
|
||||
r7 = r7 - r0; // 41
|
||||
r4 = r4 + r3 + select(0x6ced15b7u, 0x6df7aed4u, ((sel >> 4u) & 1u) != 0u); // 42
|
||||
r7 = r7 ^ simd_shuffle_xor(r3, (ushort)4); // 43
|
||||
r0 = r0 ^ dataset[((rotl_imm(r7 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 44
|
||||
r1 = r1 + r6 + select(0xe09f54e9u, 0x83e825bfu, ((sel >> 14u) & 1u) != 0u); // 45
|
||||
r3 = r3 ^ dataset[((rotl_imm(r1 * 0x03ac37adu, 22u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 46
|
||||
r6 = r6 ^ dataset[((rotl_imm(r3 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 47
|
||||
r4 = mulhi(r4, r2); // 48
|
||||
r5 = r5 + r0 + select(0xa8bae6dfu, 0xf572bdb9u, ((sel >> 7u) & 1u) != 0u); // 49
|
||||
r0 = r0 ^ simd_shuffle_xor(r7, (ushort)4); // 50
|
||||
r6 = r6 + r0 + select(0x383b9260u, 0x11e17c61u, ((sel >> 19u) & 1u) != 0u); // 51
|
||||
r5 = r5 ^ dataset[((rotl_imm(r2 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 52
|
||||
r6 = rotl_imm(r6, 12u); // 53
|
||||
r3 = rotl_imm(r3, 12u); // 54
|
||||
r2 = r2 + r1 + select(0xac6be8e3u, 0x18d67dbbu, ((sel >> 10u) & 1u) != 0u); // 55
|
||||
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x03ac37adu, 22u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 56
|
||||
r5 = r5 + r0 + select(0xf03673feu, 0xa75cd60du, ((sel >> 12u) & 1u) != 0u); // 57
|
||||
r5 = r5 ^ dataset[((rotl_imm(r0 * 0x03ac37adu, 22u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 58
|
||||
r1 = r1 + r4 + select(0xdecd4794u, 0x8dfb96bbu, ((sel >> 7u) & 1u) != 0u); // 59
|
||||
r3 = mulhi(r3, r2); // 60
|
||||
r6 = r6 | r4; // 61
|
||||
r5 = r5 ^ simd_shuffle_xor(r4, (ushort)8); // 62
|
||||
r3 = r3 ^ dataset[((rotl_imm(r6 * 0x03ac37adu, 22u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 63
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
5
proto-cuda/packs-ca2-era/era-5/seeds.txt
Normal file
5
proto-cuda/packs-ca2-era/era-5/seeds.txt
Normal file
|
|
@ -0,0 +1,5 @@
|
|||
epoch_seed_hex edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07
|
||||
day_seed_hex 69676e65756d2d6461792ffa50000000000000
|
||||
epoch_index 0
|
||||
day_index 20730
|
||||
daa_score 0
|
||||
57
proto-cuda/packs-ca2-era/era-5/vectors.h
Normal file
57
proto-cuda/packs-ca2-era/era-5/vectors.h
Normal file
|
|
@ -0,0 +1,57 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v3, memory-hard dataset
|
||||
#pragma once
|
||||
#ifdef __cplusplus
|
||||
#include <cstdint>
|
||||
#else
|
||||
#include <stdint.h>
|
||||
#endif
|
||||
|
||||
#define IGNEUM_VEC_WARPS 3
|
||||
static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u };
|
||||
static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = {
|
||||
{ // base nonce 0
|
||||
0x2db41cf6c7af96f3ull, 0x7fb10b72307a9a54ull, 0x29745e1b3faac3bcull, 0xfac26f39a0c10352ull, 0x27a840a938abfab5ull, 0xbff83c8c7002dd7dull, 0x75eeaa3644794177ull, 0x27af0df4449a9af8ull,
|
||||
0x20f63f8623b4d27bull, 0x5f185cdf257139a2ull, 0xfe856bbdc14a9d00ull, 0x75afbd5bf6aff171ull, 0xbeb4747caff3825bull, 0x732e6a1d507f5553ull, 0xcebad71cc5667e98ull, 0x9a348da211240839ull,
|
||||
0xcd7aef8110941186ull, 0x07de91b90080abdaull, 0xf57c41187ed07aa4ull, 0x333465d0d32dbbf3ull, 0x1af724ae2d46212full, 0x32401faf581fbf3aull, 0x5e17e5f8a8a041f2ull, 0x341a99051ce73344ull,
|
||||
0xca1e16b1e0667d5dull, 0x29a00e23b166e107ull, 0x80340e9bf913c10aull, 0x504a277409796603ull, 0x30a37c2ee897d721ull, 0xad8651e4bf99a237ull, 0x58daf3a88e4b45f7ull, 0x9b8a4b4bff0bf113ull
|
||||
},
|
||||
{ // base nonce 4096
|
||||
0x0bf4da3e489363e3ull, 0x0dbea592d74bd704ull, 0xb71f74d5fc74b3acull, 0x4484fc773d34c715ull, 0xcdccd9508c1d4b18ull, 0x63bd8af4608eddb0ull, 0x821fb75fe7f477d8ull, 0x3b05158071d781bfull,
|
||||
0x314d0fdf84dc1253ull, 0x05d72102d3c691b1ull, 0x4e952bc2080517ffull, 0xa7fd18dbbc098402ull, 0x61c8e269f3d5a985ull, 0x10d6cd5a58b5a7c6ull, 0xfbbe92eec88b825full, 0x7883d15dc89e4ec8ull,
|
||||
0xe9c2a972de6ec682ull, 0x1e8cc7731a411d89ull, 0xe08a27e57fff9792ull, 0x7e5f53655eb81585ull, 0x871553b59413784cull, 0x8148130a58f0e5feull, 0xd8531aada9d6c55cull, 0xe42fb76fdf380eecull,
|
||||
0x4d7b7adc743db968ull, 0xba4705cf0928c39full, 0x9c0b6ad26b4f6ff5ull, 0x7ab1c01b8c88fb16ull, 0x586bd03f8f60d8d7ull, 0xdee397b42dfe2e6dull, 0x0db04bc1e66f099bull, 0x583f1670cd1dc91cull
|
||||
},
|
||||
{ // base nonce 1000000
|
||||
0xbb874c9300805389ull, 0xaf7416b01b0059e6ull, 0x6d90b71470bbf4b5ull, 0x397fe67caedc8c53ull, 0x29cf7d38ad70857full, 0xe9d0c6f8f957524eull, 0x0ac6384db3ac4e90ull, 0x57f02bb957e51a77ull,
|
||||
0xc2cfded2133a57daull, 0xbcf718fb072d2052ull, 0x31cc1574fa6472c0ull, 0xbdbfb3be16c0c797ull, 0x173f764f53e1e0dfull, 0xeac42ae3a6b3dd9bull, 0x03f478695ad6abe4ull, 0xfd3913d32e21ac01ull,
|
||||
0x4b95c2bd3f95ff21ull, 0x8c677d2917e25e7aull, 0x309fa642634958f4ull, 0xcbac756299a8cf64ull, 0x49ec74ef6fba7b27ull, 0xe4d291c0294abd05ull, 0x1a4ae0aae503257cull, 0xeb8b228788d3dac3ull,
|
||||
0x0439d39ca4b6e471ull, 0xd7da9d74d91ee421ull, 0x6377779157d20db9ull, 0xec8c6fc70df81aa9ull, 0xea251e44b4c4bf67ull, 0x209ba0e64cdd2443ull, 0x3aa5cb567a6ae4f3ull, 0x9bff7b0427c95ed9ull
|
||||
}
|
||||
};
|
||||
|
||||
// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455).
|
||||
static const uint32_t IGNEUM_DS_HEAD[16] = {
|
||||
0x3dd50b1fu, 0x48edec90u, 0x93123701u, 0x3a7d2407u, 0x5119540eu, 0x657ac748u, 0x3358ef2du, 0x09e2105cu,
|
||||
0xcdd88b45u, 0xb03a0ac0u, 0x2c37c6b0u, 0x10e2c6a5u, 0x25e2929cu, 0x4317cabdu, 0x5fa0cdccu, 0x6178ee52u
|
||||
};
|
||||
static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u;
|
||||
static const uint32_t IGNEUM_DS_LAST = 0xf7b7180eu;
|
||||
// 64 sampled dataset words (index, value) computed on the Mac.
|
||||
#define IGNEUM_DS_SAMPLES 64
|
||||
static const uint32_t IGNEUM_DS_SAMPLE_INDEX[IGNEUM_DS_SAMPLES] = {
|
||||
59471966u, 217795994u, 208353206u, 42483309u, 172547758u, 148076330u, 183853158u, 214389424u, 267488061u, 169781097u, 184093494u, 153880993u, 84977930u, 46426879u, 3093825u, 225364072u, 44593546u, 260713159u, 168250303u, 52384140u, 223401610u, 45554030u, 95410555u, 175039924u, 79171087u, 267580473u, 24168642u, 37981670u, 171551130u, 195559979u, 204611762u, 140997658u, 138925853u, 86637313u, 20736778u, 219665210u, 160430336u, 264654675u, 8013395u, 228945585u, 213884386u, 104419827u, 44185464u, 142737231u, 99284897u, 132475900u, 61861762u, 132056166u, 262388043u, 91878046u, 117353561u, 124768597u, 71352993u, 190698941u, 46055428u, 55281366u, 165145231u, 106810753u, 171985651u, 232085256u, 159510492u, 40072060u, 209107596u, 39023794u
|
||||
};
|
||||
static const uint32_t IGNEUM_DS_SAMPLE_VALUE[IGNEUM_DS_SAMPLES] = {
|
||||
0x1a8d5842u, 0x02457408u, 0x9ee5ee59u, 0x1fadd2a3u, 0x046827bcu, 0xcb31518fu, 0xc502a5cbu, 0xc030deeeu, 0xa690604cu, 0x89831ed7u, 0xf640c0f4u, 0xf145a054u, 0xe4d9ed8du, 0x6abc511bu, 0xfd62884cu, 0x641fb4bdu, 0x252f7254u, 0xa224439cu, 0x395e9a21u, 0xb5e60d24u, 0x87ff915au, 0x8ddc5d17u, 0xe7da99a3u, 0x16acb57au, 0x103f86c1u, 0x7144efafu, 0x98e820c1u, 0x893347ebu, 0x0c063619u, 0x01c44610u, 0xab88c304u, 0xcd17d46au, 0x4dc37dbbu, 0xfbb80109u, 0x3d5154cbu, 0x406cb5bcu, 0xac1e69e5u, 0x9b52077au, 0x7b09b6fcu, 0xa1a5f8a9u, 0x14a583a6u, 0xc866af23u, 0xfaddd4d2u, 0xd79db283u, 0x7ac395cau, 0xc0d1c920u, 0xe3f98318u, 0x33ad1217u, 0x976191a9u, 0xf1a97b54u, 0x50ce7f4cu, 0x003f843cu, 0xfcdd3e77u, 0xe153a1c9u, 0x73bde4b8u, 0xcc5d5e5cu, 0x9ac3815eu, 0xa54ddfd1u, 0x8a9ab3adu, 0x788a9ca6u, 0xbb567051u, 0x8805df35u, 0x896453cdu, 0x4c77ba49u
|
||||
};
|
||||
// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^26 words.
|
||||
static const uint32_t IGNEUM_CACHE_HEAD[16] = {
|
||||
0xebd9055cu, 0x7eaf6b21u, 0x9b610855u, 0x134a8562u, 0xb10059bau, 0x7f459e58u, 0x8f3c7873u, 0x3523e609u,
|
||||
0x75b8f4cau, 0x750d31b4u, 0x0dae0782u, 0x26f10015u, 0x7ba87a41u, 0x0efd8543u, 0x691d6368u, 0x8929d967u
|
||||
};
|
||||
static const uint32_t IGNEUM_CACHE_LAST[16] = {
|
||||
0x51474edfu, 0xc3cfba93u, 0xf21454cfu, 0x9b79baacu, 0x4d4fce23u, 0xd701dfd5u, 0x37357ab3u, 0x1be693fau,
|
||||
0xa701cb3bu, 0x7467c620u, 0x428184e8u, 0xf4010df0u, 0xe33aa88fu, 0xc78d6d62u, 0xae9ba9c0u, 0xf4bba97eu
|
||||
};
|
||||
static const uint64_t IGNEUM_CACHE_FNV64 = 0x448274a57f508cbcull;
|
||||
36
proto-cuda/packs-ca2-era/era-5/vectors.json
Normal file
36
proto-cuda/packs-ca2-era/era-5/vectors.json
Normal file
|
|
@ -0,0 +1,36 @@
|
|||
{
|
||||
"seed": "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000",
|
||||
"day": "bytes:69676e65756d2d6461792ffa50000000000000",
|
||||
"dataset_mode": "memory-hard",
|
||||
"dataset_log2_words": 28,
|
||||
"mask": "0x0fffffff",
|
||||
"lanes": 32,
|
||||
"source": "igneum-pow (Rust) CPU interpreter, generator v3, memory-hard dataset",
|
||||
"warps": [
|
||||
{"base_nonce": 0, "expected": [
|
||||
"0x2db41cf6c7af96f3", "0x7fb10b72307a9a54", "0x29745e1b3faac3bc", "0xfac26f39a0c10352", "0x27a840a938abfab5", "0xbff83c8c7002dd7d", "0x75eeaa3644794177", "0x27af0df4449a9af8",
|
||||
"0x20f63f8623b4d27b", "0x5f185cdf257139a2", "0xfe856bbdc14a9d00", "0x75afbd5bf6aff171", "0xbeb4747caff3825b", "0x732e6a1d507f5553", "0xcebad71cc5667e98", "0x9a348da211240839",
|
||||
"0xcd7aef8110941186", "0x07de91b90080abda", "0xf57c41187ed07aa4", "0x333465d0d32dbbf3", "0x1af724ae2d46212f", "0x32401faf581fbf3a", "0x5e17e5f8a8a041f2", "0x341a99051ce73344",
|
||||
"0xca1e16b1e0667d5d", "0x29a00e23b166e107", "0x80340e9bf913c10a", "0x504a277409796603", "0x30a37c2ee897d721", "0xad8651e4bf99a237", "0x58daf3a88e4b45f7", "0x9b8a4b4bff0bf113"
|
||||
]},
|
||||
{"base_nonce": 4096, "expected": [
|
||||
"0x0bf4da3e489363e3", "0x0dbea592d74bd704", "0xb71f74d5fc74b3ac", "0x4484fc773d34c715", "0xcdccd9508c1d4b18", "0x63bd8af4608eddb0", "0x821fb75fe7f477d8", "0x3b05158071d781bf",
|
||||
"0x314d0fdf84dc1253", "0x05d72102d3c691b1", "0x4e952bc2080517ff", "0xa7fd18dbbc098402", "0x61c8e269f3d5a985", "0x10d6cd5a58b5a7c6", "0xfbbe92eec88b825f", "0x7883d15dc89e4ec8",
|
||||
"0xe9c2a972de6ec682", "0x1e8cc7731a411d89", "0xe08a27e57fff9792", "0x7e5f53655eb81585", "0x871553b59413784c", "0x8148130a58f0e5fe", "0xd8531aada9d6c55c", "0xe42fb76fdf380eec",
|
||||
"0x4d7b7adc743db968", "0xba4705cf0928c39f", "0x9c0b6ad26b4f6ff5", "0x7ab1c01b8c88fb16", "0x586bd03f8f60d8d7", "0xdee397b42dfe2e6d", "0x0db04bc1e66f099b", "0x583f1670cd1dc91c"
|
||||
]},
|
||||
{"base_nonce": 1000000, "expected": [
|
||||
"0xbb874c9300805389", "0xaf7416b01b0059e6", "0x6d90b71470bbf4b5", "0x397fe67caedc8c53", "0x29cf7d38ad70857f", "0xe9d0c6f8f957524e", "0x0ac6384db3ac4e90", "0x57f02bb957e51a77",
|
||||
"0xc2cfded2133a57da", "0xbcf718fb072d2052", "0x31cc1574fa6472c0", "0xbdbfb3be16c0c797", "0x173f764f53e1e0df", "0xeac42ae3a6b3dd9b", "0x03f478695ad6abe4", "0xfd3913d32e21ac01",
|
||||
"0x4b95c2bd3f95ff21", "0x8c677d2917e25e7a", "0x309fa642634958f4", "0xcbac756299a8cf64", "0x49ec74ef6fba7b27", "0xe4d291c0294abd05", "0x1a4ae0aae503257c", "0xeb8b228788d3dac3",
|
||||
"0x0439d39ca4b6e471", "0xd7da9d74d91ee421", "0x6377779157d20db9", "0xec8c6fc70df81aa9", "0xea251e44b4c4bf67", "0x209ba0e64cdd2443", "0x3aa5cb567a6ae4f3", "0x9bff7b0427c95ed9"
|
||||
]}
|
||||
],
|
||||
"dataset_head": ["0x3dd50b1f", "0x48edec90", "0x93123701", "0x3a7d2407", "0x5119540e", "0x657ac748", "0x3358ef2d", "0x09e2105c", "0xcdd88b45", "0xb03a0ac0", "0x2c37c6b0", "0x10e2c6a5", "0x25e2929c", "0x4317cabd", "0x5fa0cdcc", "0x6178ee52"],
|
||||
"dataset_last_index": 268435455,
|
||||
"dataset_last": "0xf7b7180e",
|
||||
"dataset_samples": [{"index": 59471966, "value": "0x1a8d5842"}, {"index": 217795994, "value": "0x02457408"}, {"index": 208353206, "value": "0x9ee5ee59"}, {"index": 42483309, "value": "0x1fadd2a3"}, {"index": 172547758, "value": "0x046827bc"}, {"index": 148076330, "value": "0xcb31518f"}, {"index": 183853158, "value": "0xc502a5cb"}, {"index": 214389424, "value": "0xc030deee"}, {"index": 267488061, "value": "0xa690604c"}, {"index": 169781097, "value": "0x89831ed7"}, {"index": 184093494, "value": "0xf640c0f4"}, {"index": 153880993, "value": "0xf145a054"}, {"index": 84977930, "value": "0xe4d9ed8d"}, {"index": 46426879, "value": "0x6abc511b"}, {"index": 3093825, "value": "0xfd62884c"}, {"index": 225364072, "value": "0x641fb4bd"}, {"index": 44593546, "value": "0x252f7254"}, {"index": 260713159, "value": "0xa224439c"}, {"index": 168250303, "value": "0x395e9a21"}, {"index": 52384140, "value": "0xb5e60d24"}, {"index": 223401610, "value": "0x87ff915a"}, {"index": 45554030, "value": "0x8ddc5d17"}, {"index": 95410555, "value": "0xe7da99a3"}, {"index": 175039924, "value": "0x16acb57a"}, {"index": 79171087, "value": "0x103f86c1"}, {"index": 267580473, "value": "0x7144efaf"}, {"index": 24168642, "value": "0x98e820c1"}, {"index": 37981670, "value": "0x893347eb"}, {"index": 171551130, "value": "0x0c063619"}, {"index": 195559979, "value": "0x01c44610"}, {"index": 204611762, "value": "0xab88c304"}, {"index": 140997658, "value": "0xcd17d46a"}, {"index": 138925853, "value": "0x4dc37dbb"}, {"index": 86637313, "value": "0xfbb80109"}, {"index": 20736778, "value": "0x3d5154cb"}, {"index": 219665210, "value": "0x406cb5bc"}, {"index": 160430336, "value": "0xac1e69e5"}, {"index": 264654675, "value": "0x9b52077a"}, {"index": 8013395, "value": "0x7b09b6fc"}, {"index": 228945585, "value": "0xa1a5f8a9"}, {"index": 213884386, "value": "0x14a583a6"}, {"index": 104419827, "value": "0xc866af23"}, {"index": 44185464, "value": "0xfaddd4d2"}, {"index": 142737231, "value": "0xd79db283"}, {"index": 99284897, "value": "0x7ac395ca"}, {"index": 132475900, "value": "0xc0d1c920"}, {"index": 61861762, "value": "0xe3f98318"}, {"index": 132056166, "value": "0x33ad1217"}, {"index": 262388043, "value": "0x976191a9"}, {"index": 91878046, "value": "0xf1a97b54"}, {"index": 117353561, "value": "0x50ce7f4c"}, {"index": 124768597, "value": "0x003f843c"}, {"index": 71352993, "value": "0xfcdd3e77"}, {"index": 190698941, "value": "0xe153a1c9"}, {"index": 46055428, "value": "0x73bde4b8"}, {"index": 55281366, "value": "0xcc5d5e5c"}, {"index": 165145231, "value": "0x9ac3815e"}, {"index": 106810753, "value": "0xa54ddfd1"}, {"index": 171985651, "value": "0x8a9ab3ad"}, {"index": 232085256, "value": "0x788a9ca6"}, {"index": 159510492, "value": "0xbb567051"}, {"index": 40072060, "value": "0x8805df35"}, {"index": 209107596, "value": "0x896453cd"}, {"index": 39023794, "value": "0x4c77ba49"}],
|
||||
"cache_head": ["0xebd9055c", "0x7eaf6b21", "0x9b610855", "0x134a8562", "0xb10059ba", "0x7f459e58", "0x8f3c7873", "0x3523e609", "0x75b8f4ca", "0x750d31b4", "0x0dae0782", "0x26f10015", "0x7ba87a41", "0x0efd8543", "0x691d6368", "0x8929d967"],
|
||||
"cache_last_line": ["0x51474edf", "0xc3cfba93", "0xf21454cf", "0x9b79baac", "0x4d4fce23", "0xd701dfd5", "0x37357ab3", "0x1be693fa", "0xa701cb3b", "0x7467c620", "0x428184e8", "0xf4010df0", "0xe33aa88f", "0xc78d6d62", "0xae9ba9c0", "0xf4bba97e"],
|
||||
"cache_fnv1a64": "0x448274a57f508cbc"
|
||||
}
|
||||
|
|
@ -155,8 +155,12 @@ static inline void mh_item(__global const uint* cache, uint t, uint* s) {
|
|||
}
|
||||
for (uint j = 0u; j < 4u; ++j) mh_mixer(s, 0x9E3779B9u * (32u + j + 1u));
|
||||
}
|
||||
// dataset[w] without the dataset: derive item w >> 4 and take word w & 15.
|
||||
static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; }
|
||||
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 0 2 12 13 of w.
|
||||
static inline uint mh_j(uint w) { return ((w >> 0u) & 1u) | (((w >> 2u) & 1u) << 1) | (((w >> 12u) & 1u) << 2) | (((w >> 13u) & 1u) << 3); }
|
||||
static inline uint mh_t(uint w) { w = (w & 0x00001fffu) | ((w >> 14u) << 13u); w = (w & 0x00000fffu) | ((w >> 13u) << 12u); w = (w & 0x00000003u) | ((w >> 3u) << 2u); w = (w & 0x00000000u) | ((w >> 1u) << 0u); return w; }
|
||||
static inline uint mh_addr(uint t, uint j) { uint w = t; w = ((w >> 0u) << 1u) | (w & 0x00000000u) | (((j >> 0u) & 1u) << 0u); w = ((w >> 2u) << 3u) | (w & 0x00000003u) | (((j >> 1u) & 1u) << 2u); w = ((w >> 12u) << 13u) | (w & 0x00000fffu) | (((j >> 2u) & 1u) << 12u); w = ((w >> 13u) << 14u) | (w & 0x00001fffu) | (((j >> 3u) & 1u) << 13u); return w; }
|
||||
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
|
||||
static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }
|
||||
|
||||
// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item.
|
||||
// The same constants as memhard.h in this pack (one emitter, three dialects).
|
||||
|
|
@ -169,8 +173,7 @@ __kernel void igneum_build(__global uint* ds, __global const uint* cache, uint n
|
|||
if (t < nItems) {
|
||||
uint s[16];
|
||||
mh_item(cache, t, s);
|
||||
__global uint* d = ds + ((ulong)t * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) d[i] = s[i];
|
||||
for (uint i = 0u; i < 16u; ++i) ds[(ulong)mh_addr(t, i)] = s[i];
|
||||
}
|
||||
}
|
||||
|
||||
|
|
@ -200,69 +203,69 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out
|
|||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r0, 4u); r2 = r2 ^ t_; } // 1 shfl
|
||||
r3 = r3 + r2 + ((((sel >> 10u) & 1u) != 0u) ? 0x642e66dbu : 0x2cccb6cau); // 2 add
|
||||
r0 = rotl_imm(r0, 19u); // 3 rotl
|
||||
r7 = rotr_var(r7, r6); // 4 rotr
|
||||
r7 = r7 + r4 + ((((sel >> 21u) & 1u) != 0u) ? 0xc1535555u : 0xee02465fu); // 5 add
|
||||
r1 = mul_hi(r1, r7); // 6 mulhi
|
||||
r4 = r4 ^ ds[r2 & mask]; // 7 load
|
||||
r7 = r7 ^ ds[r4 & mask]; // 8 load
|
||||
r0 = r0 ^ ds[r3 & mask]; // 9 load
|
||||
r5 = r5 ^ ds[r1 & mask]; // 10 load
|
||||
r1 = r1 ^ ds[r5 & mask]; // 11 load
|
||||
r3 = mul_hi(r3, r5); // 12 mulhi
|
||||
r1 = r1 ^ ds[r3 & mask]; // 13 load
|
||||
r0 = r0 - r3; // 14 sub
|
||||
r5 = r1 * r3 + r5; // 15 mad
|
||||
r6 = mul_hi(r6, r1); // 16 mulhi
|
||||
r5 = r5 + r2 + ((((sel >> 28u) & 1u) != 0u) ? 0x8b965b57u : 0x697b3d00u); // 17 add
|
||||
r0 = mul_hi(r0, r6); // 18 mulhi
|
||||
r5 = rotr_var(r5, r3); // 19 rotr
|
||||
r5 = mul_hi(r5, r2); // 20 mulhi
|
||||
r1 = r1 + r0 + ((((sel >> 1u) & 1u) != 0u) ? 0x6d7e8d05u : 0xebcf247au); // 21 add
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 22 add
|
||||
r1 = mul_hi(r1, r5); // 23 mulhi
|
||||
r2 = r2 - r5; // 24 sub
|
||||
r7 = r7 + r4 + ((((sel >> 2u) & 1u) != 0u) ? 0x699ef1bbu : 0x08ffa6c7u); // 25 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 26 shfl
|
||||
r7 = r7 + r1 + ((((sel >> 14u) & 1u) != 0u) ? 0xb4ead2fbu : 0xe60fea84u); // 27 add
|
||||
r3 = r3 + r1 + ((((sel >> 6u) & 1u) != 0u) ? 0x8f30d21du : 0x65c76dabu); // 28 add
|
||||
r2 = r2 ^ ds[r1 & mask]; // 29 load
|
||||
r5 = r5 ^ ds[r7 & mask]; // 30 load
|
||||
r2 = r2 ^ ds[r5 & mask]; // 31 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r1 = r1 ^ t_; } // 32 shfl
|
||||
r4 = r5 * r7 + r4; // 33 mad
|
||||
r4 = r4 + r2 + ((((sel >> 21u) & 1u) != 0u) ? 0xc7ce690cu : 0x0480debeu); // 34 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 8u); r3 = r3 ^ t_; } // 35 shfl
|
||||
r7 = r7 + r1 + ((((sel >> 2u) & 1u) != 0u) ? 0xe10c2c95u : 0xc53b542eu); // 36 add
|
||||
r5 = r5 ^ r7; // 37 xor
|
||||
r2 = r2 | r1; // 38 or
|
||||
r1 = mul_hi(r1, r0); // 39 mulhi
|
||||
r6 = rotl_imm(r6, 19u); // 40 rotl
|
||||
r4 = mul_hi(r4, r6); // 41 mulhi
|
||||
r6 = r6 - r0; // 42 sub
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r6 = r6 ^ t_; } // 43 shfl
|
||||
r4 = r4 ^ ds[r2 & mask]; // 44 load
|
||||
r1 = r1 ^ r3; // 45 xor
|
||||
r7 = r7 ^ ds[r0 & mask]; // 46 load
|
||||
r3 = r3 ^ ds[r1 & mask]; // 47 load
|
||||
r5 = r5 * r3; // 48 mul
|
||||
r1 = r1 - r5; // 49 sub
|
||||
r2 = rotl_imm(r2, 8u); // 50 rotl
|
||||
r1 = r1 + r5 + ((((sel >> 23u) & 1u) != 0u) ? 0x77b9bd43u : 0xa900fec4u); // 51 add
|
||||
r4 = r4 ^ ds[r7 & mask]; // 52 load
|
||||
r2 = r2 - r7; // 53 sub
|
||||
r4 = r4 ^ r0; // 54 xor
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 55 add
|
||||
r2 = r2 ^ ds[r4 & mask]; // 56 load
|
||||
r0 = r1 * r4 + r0; // 57 mad
|
||||
r3 = r3 ^ ds[r5 & mask]; // 58 load
|
||||
r5 = r5 | r6; // 59 or
|
||||
r6 = r5 * r7 + r6; // 60 mad
|
||||
r4 = rotl_imm(r4, 28u); // 61 rotl
|
||||
r5 = mul_hi(r5, r0); // 62 mulhi
|
||||
r3 = r3 ^ ds[r6 & mask]; // 63 load
|
||||
r7 = r7 ^ r0; // 1 xor
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 1u); r3 = r3 ^ t_; } // 2 shfl
|
||||
r4 = r4 - r1; // 3 sub
|
||||
r2 = r0 * r4 + r2; // 4 mad
|
||||
r4 = rotl_imm(r4, 9u); // 5 rotl
|
||||
r0 = rotr_var(r0, r2); // 6 rotr
|
||||
r6 = r6 ^ ds[((rotl_imm(r7 * 0x9ad30d99u, 29u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r2 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r0 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r1 * 0x9ad30d99u, 29u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
|
||||
r5 = r5 * r4; // 12 mul
|
||||
r7 = r7 ^ ds[((rotl_imm(r6 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 16u); r6 = r6 ^ t_; } // 14 shfl
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r3 = r3 ^ t_; } // 15 shfl
|
||||
r2 = r2 | r7; // 16 or
|
||||
r0 = rotr_var(r0, r6); // 17 rotr
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
|
||||
r7 = rotl_imm(r7, 24u); // 19 rotl
|
||||
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
|
||||
r2 = r2 | r6; // 21 or
|
||||
r6 = r5 * r4 + r6; // 22 mad
|
||||
r1 = r1 | r0; // 23 or
|
||||
r6 = r6 ^ r1; // 24 xor
|
||||
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
|
||||
r7 = rotr_var(r7, r0); // 26 rotr
|
||||
r4 = r5 * r7 + r4; // 27 mad
|
||||
r3 = r2 * r4 + r3; // 28 mad
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x9ad30d99u, 29u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
|
||||
r2 = r2 ^ ds[((rotl_imm(r3 * 0x9ad30d99u, 29u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r5 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
|
||||
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
|
||||
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
|
||||
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
|
||||
r2 = mul_hi(r2, r5); // 35 mulhi
|
||||
r4 = r4 ^ r2; // 36 xor
|
||||
r6 = r6 * r5; // 37 mul
|
||||
r7 = r7 ^ r0; // 38 xor
|
||||
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
|
||||
r2 = rotr_var(r2, r3); // 40 rotr
|
||||
r7 = r7 - r0; // 41 sub
|
||||
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 43 shfl
|
||||
r0 = r0 ^ ds[((rotl_imm(r7 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
|
||||
r3 = r3 ^ ds[((rotl_imm(r1 * 0x9ad30d99u, 29u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
|
||||
r6 = r6 ^ ds[((rotl_imm(r3 * 0x9ad30d99u, 29u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
|
||||
r4 = mul_hi(r4, r2); // 48 mulhi
|
||||
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r0 = r0 ^ t_; } // 50 shfl
|
||||
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r2 * 0x9ad30d99u, 29u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
|
||||
r6 = rotl_imm(r6, 12u); // 53 rotl
|
||||
r3 = rotl_imm(r3, 12u); // 54 rotl
|
||||
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x9ad30d99u, 29u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
|
||||
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r0 * 0x9ad30d99u, 29u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
|
||||
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
|
||||
r3 = mul_hi(r3, r2); // 60 mulhi
|
||||
r6 = r6 | r4; // 61 or
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r5 = r5 ^ t_; } // 62 shfl
|
||||
r3 = r3 ^ ds[((rotl_imm(r6 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
|
|
|
|||
|
|
@ -36,8 +36,7 @@ __global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItem
|
|||
if (t < nItems) {
|
||||
uint32_t s[16];
|
||||
mh_item(cache, t, s);
|
||||
uint32_t* d = ds + (size_t)t * 16u;
|
||||
for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i];
|
||||
for (uint32_t i = 0u; i < 16u; ++i) ds[(size_t)mh_addr(t, i)] = s[i];
|
||||
}
|
||||
}
|
||||
|
||||
|
|
@ -60,69 +59,69 @@ __global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonc
|
|||
for (uint32_t it = 0u; it < 8u; ++it) {
|
||||
uint32_t sel = r0;
|
||||
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
|
||||
r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r0, 4); // 1 shfl
|
||||
r3 = r3 + r2 + ((((sel >> 10u) & 1u) != 0u) ? 0x642e66dbu : 0x2cccb6cau); // 2 add
|
||||
r0 = rotl_imm(r0, 19u); // 3 rotl
|
||||
r7 = rotr_var(r7, r6); // 4 rotr
|
||||
r7 = r7 + r4 + ((((sel >> 21u) & 1u) != 0u) ? 0xc1535555u : 0xee02465fu); // 5 add
|
||||
r1 = __umulhi(r1, r7); // 6 mulhi
|
||||
r4 = r4 ^ ds[r2 & mask]; // 7 load
|
||||
r7 = r7 ^ ds[r4 & mask]; // 8 load
|
||||
r0 = r0 ^ ds[r3 & mask]; // 9 load
|
||||
r5 = r5 ^ ds[r1 & mask]; // 10 load
|
||||
r1 = r1 ^ ds[r5 & mask]; // 11 load
|
||||
r3 = __umulhi(r3, r5); // 12 mulhi
|
||||
r1 = r1 ^ ds[r3 & mask]; // 13 load
|
||||
r0 = r0 - r3; // 14 sub
|
||||
r5 = r1 * r3 + r5; // 15 mad
|
||||
r6 = __umulhi(r6, r1); // 16 mulhi
|
||||
r5 = r5 + r2 + ((((sel >> 28u) & 1u) != 0u) ? 0x8b965b57u : 0x697b3d00u); // 17 add
|
||||
r0 = __umulhi(r0, r6); // 18 mulhi
|
||||
r5 = rotr_var(r5, r3); // 19 rotr
|
||||
r5 = __umulhi(r5, r2); // 20 mulhi
|
||||
r1 = r1 + r0 + ((((sel >> 1u) & 1u) != 0u) ? 0x6d7e8d05u : 0xebcf247au); // 21 add
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 22 add
|
||||
r1 = __umulhi(r1, r5); // 23 mulhi
|
||||
r2 = r2 - r5; // 24 sub
|
||||
r7 = r7 + r4 + ((((sel >> 2u) & 1u) != 0u) ? 0x699ef1bbu : 0x08ffa6c7u); // 25 add
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 26 shfl
|
||||
r7 = r7 + r1 + ((((sel >> 14u) & 1u) != 0u) ? 0xb4ead2fbu : 0xe60fea84u); // 27 add
|
||||
r3 = r3 + r1 + ((((sel >> 6u) & 1u) != 0u) ? 0x8f30d21du : 0x65c76dabu); // 28 add
|
||||
r2 = r2 ^ ds[r1 & mask]; // 29 load
|
||||
r5 = r5 ^ ds[r7 & mask]; // 30 load
|
||||
r2 = r2 ^ ds[r5 & mask]; // 31 load
|
||||
r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r7, 4); // 32 shfl
|
||||
r4 = r5 * r7 + r4; // 33 mad
|
||||
r4 = r4 + r2 + ((((sel >> 21u) & 1u) != 0u) ? 0xc7ce690cu : 0x0480debeu); // 34 add
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r7, 8); // 35 shfl
|
||||
r7 = r7 + r1 + ((((sel >> 2u) & 1u) != 0u) ? 0xe10c2c95u : 0xc53b542eu); // 36 add
|
||||
r5 = r5 ^ r7; // 37 xor
|
||||
r2 = r2 | r1; // 38 or
|
||||
r1 = __umulhi(r1, r0); // 39 mulhi
|
||||
r6 = rotl_imm(r6, 19u); // 40 rotl
|
||||
r4 = __umulhi(r4, r6); // 41 mulhi
|
||||
r6 = r6 - r0; // 42 sub
|
||||
r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 43 shfl
|
||||
r4 = r4 ^ ds[r2 & mask]; // 44 load
|
||||
r1 = r1 ^ r3; // 45 xor
|
||||
r7 = r7 ^ ds[r0 & mask]; // 46 load
|
||||
r3 = r3 ^ ds[r1 & mask]; // 47 load
|
||||
r5 = r5 * r3; // 48 mul
|
||||
r1 = r1 - r5; // 49 sub
|
||||
r2 = rotl_imm(r2, 8u); // 50 rotl
|
||||
r1 = r1 + r5 + ((((sel >> 23u) & 1u) != 0u) ? 0x77b9bd43u : 0xa900fec4u); // 51 add
|
||||
r4 = r4 ^ ds[r7 & mask]; // 52 load
|
||||
r2 = r2 - r7; // 53 sub
|
||||
r4 = r4 ^ r0; // 54 xor
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 55 add
|
||||
r2 = r2 ^ ds[r4 & mask]; // 56 load
|
||||
r0 = r1 * r4 + r0; // 57 mad
|
||||
r3 = r3 ^ ds[r5 & mask]; // 58 load
|
||||
r5 = r5 | r6; // 59 or
|
||||
r6 = r5 * r7 + r6; // 60 mad
|
||||
r4 = rotl_imm(r4, 28u); // 61 rotl
|
||||
r5 = __umulhi(r5, r0); // 62 mulhi
|
||||
r3 = r3 ^ ds[r6 & mask]; // 63 load
|
||||
r7 = r7 ^ r0; // 1 xor
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r6, 1); // 2 shfl
|
||||
r4 = r4 - r1; // 3 sub
|
||||
r2 = r0 * r4 + r2; // 4 mad
|
||||
r4 = rotl_imm(r4, 9u); // 5 rotl
|
||||
r0 = rotr_var(r0, r2); // 6 rotr
|
||||
r6 = r6 ^ ds[((rotl_imm(r7 * 0x9ad30d99u, 29u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r2 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r0 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r1 * 0x9ad30d99u, 29u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
|
||||
r5 = r5 * r4; // 12 mul
|
||||
r7 = r7 ^ ds[((rotl_imm(r6 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
|
||||
r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r5, 16); // 14 shfl
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r5, 1); // 15 shfl
|
||||
r2 = r2 | r7; // 16 or
|
||||
r0 = rotr_var(r0, r6); // 17 rotr
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
|
||||
r7 = rotl_imm(r7, 24u); // 19 rotl
|
||||
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
|
||||
r2 = r2 | r6; // 21 or
|
||||
r6 = r5 * r4 + r6; // 22 mad
|
||||
r1 = r1 | r0; // 23 or
|
||||
r6 = r6 ^ r1; // 24 xor
|
||||
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
|
||||
r7 = rotr_var(r7, r0); // 26 rotr
|
||||
r4 = r5 * r7 + r4; // 27 mad
|
||||
r3 = r2 * r4 + r3; // 28 mad
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x9ad30d99u, 29u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
|
||||
r2 = r2 ^ ds[((rotl_imm(r3 * 0x9ad30d99u, 29u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r5 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
|
||||
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
|
||||
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
|
||||
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
|
||||
r2 = __umulhi(r2, r5); // 35 mulhi
|
||||
r4 = r4 ^ r2; // 36 xor
|
||||
r6 = r6 * r5; // 37 mul
|
||||
r7 = r7 ^ r0; // 38 xor
|
||||
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
|
||||
r2 = rotr_var(r2, r3); // 40 rotr
|
||||
r7 = r7 - r0; // 41 sub
|
||||
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 43 shfl
|
||||
r0 = r0 ^ ds[((rotl_imm(r7 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
|
||||
r3 = r3 ^ ds[((rotl_imm(r1 * 0x9ad30d99u, 29u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
|
||||
r6 = r6 ^ ds[((rotl_imm(r3 * 0x9ad30d99u, 29u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
|
||||
r4 = __umulhi(r4, r2); // 48 mulhi
|
||||
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
|
||||
r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r7, 4); // 50 shfl
|
||||
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r2 * 0x9ad30d99u, 29u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
|
||||
r6 = rotl_imm(r6, 12u); // 53 rotl
|
||||
r3 = rotl_imm(r3, 12u); // 54 rotl
|
||||
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x9ad30d99u, 29u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
|
||||
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r0 * 0x9ad30d99u, 29u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
|
||||
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
|
||||
r3 = __umulhi(r3, r2); // 60 mulhi
|
||||
r6 = r6 | r4; // 61 or
|
||||
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 62 shfl
|
||||
r3 = r3 ^ ds[((rotl_imm(r6 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
|
||||
}
|
||||
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
|
|
|
|||
|
|
@ -155,8 +155,12 @@ static inline void mh_item(__global const uint* cache, uint t, uint* s) {
|
|||
}
|
||||
for (uint j = 0u; j < 4u; ++j) mh_mixer(s, 0x9E3779B9u * (32u + j + 1u));
|
||||
}
|
||||
// dataset[w] without the dataset: derive item w >> 4 and take word w & 15.
|
||||
static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; }
|
||||
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 0 2 12 13 of w.
|
||||
static inline uint mh_j(uint w) { return ((w >> 0u) & 1u) | (((w >> 2u) & 1u) << 1) | (((w >> 12u) & 1u) << 2) | (((w >> 13u) & 1u) << 3); }
|
||||
static inline uint mh_t(uint w) { w = (w & 0x00001fffu) | ((w >> 14u) << 13u); w = (w & 0x00000fffu) | ((w >> 13u) << 12u); w = (w & 0x00000003u) | ((w >> 3u) << 2u); w = (w & 0x00000000u) | ((w >> 1u) << 0u); return w; }
|
||||
static inline uint mh_addr(uint t, uint j) { uint w = t; w = ((w >> 0u) << 1u) | (w & 0x00000000u) | (((j >> 0u) & 1u) << 0u); w = ((w >> 2u) << 3u) | (w & 0x00000003u) | (((j >> 1u) & 1u) << 2u); w = ((w >> 12u) << 13u) | (w & 0x00000fffu) | (((j >> 2u) & 1u) << 12u); w = ((w >> 13u) << 14u) | (w & 0x00001fffu) | (((j >> 3u) & 1u) << 13u); return w; }
|
||||
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
|
||||
static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }
|
||||
|
||||
// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item.
|
||||
// The same constants as memhard.h in this pack (one emitter, three dialects).
|
||||
|
|
@ -169,8 +173,7 @@ __kernel void igneum_build(__global uint* ds, __global const uint* cache, uint n
|
|||
if (t < nItems) {
|
||||
uint s[16];
|
||||
mh_item(cache, t, s);
|
||||
__global uint* d = ds + ((ulong)t * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) d[i] = s[i];
|
||||
for (uint i = 0u; i < 16u; ++i) ds[(ulong)mh_addr(t, i)] = s[i];
|
||||
}
|
||||
}
|
||||
|
||||
|
|
@ -200,69 +203,69 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out
|
|||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r0, 4u); r2 = r2 ^ t_; } // 1 shfl
|
||||
r3 = r3 + r2 + ((((sel >> 10u) & 1u) != 0u) ? 0x642e66dbu : 0x2cccb6cau); // 2 add
|
||||
r0 = rotl_imm(r0, 19u); // 3 rotl
|
||||
r7 = rotr_var(r7, r6); // 4 rotr
|
||||
r7 = r7 + r4 + ((((sel >> 21u) & 1u) != 0u) ? 0xc1535555u : 0xee02465fu); // 5 add
|
||||
r1 = mul_hi(r1, r7); // 6 mulhi
|
||||
r4 = r4 ^ ds[r2 & mask]; // 7 load
|
||||
r7 = r7 ^ ds[r4 & mask]; // 8 load
|
||||
r0 = r0 ^ ds[r3 & mask]; // 9 load
|
||||
r5 = r5 ^ ds[r1 & mask]; // 10 load
|
||||
r1 = r1 ^ ds[r5 & mask]; // 11 load
|
||||
r3 = mul_hi(r3, r5); // 12 mulhi
|
||||
r1 = r1 ^ ds[r3 & mask]; // 13 load
|
||||
r0 = r0 - r3; // 14 sub
|
||||
r5 = r1 * r3 + r5; // 15 mad
|
||||
r6 = mul_hi(r6, r1); // 16 mulhi
|
||||
r5 = r5 + r2 + ((((sel >> 28u) & 1u) != 0u) ? 0x8b965b57u : 0x697b3d00u); // 17 add
|
||||
r0 = mul_hi(r0, r6); // 18 mulhi
|
||||
r5 = rotr_var(r5, r3); // 19 rotr
|
||||
r5 = mul_hi(r5, r2); // 20 mulhi
|
||||
r1 = r1 + r0 + ((((sel >> 1u) & 1u) != 0u) ? 0x6d7e8d05u : 0xebcf247au); // 21 add
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 22 add
|
||||
r1 = mul_hi(r1, r5); // 23 mulhi
|
||||
r2 = r2 - r5; // 24 sub
|
||||
r7 = r7 + r4 + ((((sel >> 2u) & 1u) != 0u) ? 0x699ef1bbu : 0x08ffa6c7u); // 25 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 26 shfl
|
||||
r7 = r7 + r1 + ((((sel >> 14u) & 1u) != 0u) ? 0xb4ead2fbu : 0xe60fea84u); // 27 add
|
||||
r3 = r3 + r1 + ((((sel >> 6u) & 1u) != 0u) ? 0x8f30d21du : 0x65c76dabu); // 28 add
|
||||
r2 = r2 ^ ds[r1 & mask]; // 29 load
|
||||
r5 = r5 ^ ds[r7 & mask]; // 30 load
|
||||
r2 = r2 ^ ds[r5 & mask]; // 31 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r1 = r1 ^ t_; } // 32 shfl
|
||||
r4 = r5 * r7 + r4; // 33 mad
|
||||
r4 = r4 + r2 + ((((sel >> 21u) & 1u) != 0u) ? 0xc7ce690cu : 0x0480debeu); // 34 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 8u); r3 = r3 ^ t_; } // 35 shfl
|
||||
r7 = r7 + r1 + ((((sel >> 2u) & 1u) != 0u) ? 0xe10c2c95u : 0xc53b542eu); // 36 add
|
||||
r5 = r5 ^ r7; // 37 xor
|
||||
r2 = r2 | r1; // 38 or
|
||||
r1 = mul_hi(r1, r0); // 39 mulhi
|
||||
r6 = rotl_imm(r6, 19u); // 40 rotl
|
||||
r4 = mul_hi(r4, r6); // 41 mulhi
|
||||
r6 = r6 - r0; // 42 sub
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r6 = r6 ^ t_; } // 43 shfl
|
||||
r4 = r4 ^ ds[r2 & mask]; // 44 load
|
||||
r1 = r1 ^ r3; // 45 xor
|
||||
r7 = r7 ^ ds[r0 & mask]; // 46 load
|
||||
r3 = r3 ^ ds[r1 & mask]; // 47 load
|
||||
r5 = r5 * r3; // 48 mul
|
||||
r1 = r1 - r5; // 49 sub
|
||||
r2 = rotl_imm(r2, 8u); // 50 rotl
|
||||
r1 = r1 + r5 + ((((sel >> 23u) & 1u) != 0u) ? 0x77b9bd43u : 0xa900fec4u); // 51 add
|
||||
r4 = r4 ^ ds[r7 & mask]; // 52 load
|
||||
r2 = r2 - r7; // 53 sub
|
||||
r4 = r4 ^ r0; // 54 xor
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 55 add
|
||||
r2 = r2 ^ ds[r4 & mask]; // 56 load
|
||||
r0 = r1 * r4 + r0; // 57 mad
|
||||
r3 = r3 ^ ds[r5 & mask]; // 58 load
|
||||
r5 = r5 | r6; // 59 or
|
||||
r6 = r5 * r7 + r6; // 60 mad
|
||||
r4 = rotl_imm(r4, 28u); // 61 rotl
|
||||
r5 = mul_hi(r5, r0); // 62 mulhi
|
||||
r3 = r3 ^ ds[r6 & mask]; // 63 load
|
||||
r7 = r7 ^ r0; // 1 xor
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 1u); r3 = r3 ^ t_; } // 2 shfl
|
||||
r4 = r4 - r1; // 3 sub
|
||||
r2 = r0 * r4 + r2; // 4 mad
|
||||
r4 = rotl_imm(r4, 9u); // 5 rotl
|
||||
r0 = rotr_var(r0, r2); // 6 rotr
|
||||
r6 = r6 ^ ds[((rotl_imm(r7 * 0x9ad30d99u, 29u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r2 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r0 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r1 * 0x9ad30d99u, 29u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
|
||||
r5 = r5 * r4; // 12 mul
|
||||
r7 = r7 ^ ds[((rotl_imm(r6 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 16u); r6 = r6 ^ t_; } // 14 shfl
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r3 = r3 ^ t_; } // 15 shfl
|
||||
r2 = r2 | r7; // 16 or
|
||||
r0 = rotr_var(r0, r6); // 17 rotr
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
|
||||
r7 = rotl_imm(r7, 24u); // 19 rotl
|
||||
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
|
||||
r2 = r2 | r6; // 21 or
|
||||
r6 = r5 * r4 + r6; // 22 mad
|
||||
r1 = r1 | r0; // 23 or
|
||||
r6 = r6 ^ r1; // 24 xor
|
||||
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
|
||||
r7 = rotr_var(r7, r0); // 26 rotr
|
||||
r4 = r5 * r7 + r4; // 27 mad
|
||||
r3 = r2 * r4 + r3; // 28 mad
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x9ad30d99u, 29u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
|
||||
r2 = r2 ^ ds[((rotl_imm(r3 * 0x9ad30d99u, 29u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r5 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
|
||||
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
|
||||
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
|
||||
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
|
||||
r2 = mul_hi(r2, r5); // 35 mulhi
|
||||
r4 = r4 ^ r2; // 36 xor
|
||||
r6 = r6 * r5; // 37 mul
|
||||
r7 = r7 ^ r0; // 38 xor
|
||||
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
|
||||
r2 = rotr_var(r2, r3); // 40 rotr
|
||||
r7 = r7 - r0; // 41 sub
|
||||
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 43 shfl
|
||||
r0 = r0 ^ ds[((rotl_imm(r7 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
|
||||
r3 = r3 ^ ds[((rotl_imm(r1 * 0x9ad30d99u, 29u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
|
||||
r6 = r6 ^ ds[((rotl_imm(r3 * 0x9ad30d99u, 29u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
|
||||
r4 = mul_hi(r4, r2); // 48 mulhi
|
||||
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r0 = r0 ^ t_; } // 50 shfl
|
||||
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r2 * 0x9ad30d99u, 29u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
|
||||
r6 = rotl_imm(r6, 12u); // 53 rotl
|
||||
r3 = rotl_imm(r3, 12u); // 54 rotl
|
||||
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x9ad30d99u, 29u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
|
||||
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r0 * 0x9ad30d99u, 29u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
|
||||
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
|
||||
r3 = mul_hi(r3, r2); // 60 mulhi
|
||||
r6 = r6 | r4; // 61 or
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r5 = r5 ^ t_; } // 62 shfl
|
||||
r3 = r3 ^ ds[((rotl_imm(r6 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
|
|
@ -303,69 +306,69 @@ IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulon
|
|||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r0, 4u); r2 = r2 ^ t_; } // 1 shfl
|
||||
r3 = r3 + r2 + ((((sel >> 10u) & 1u) != 0u) ? 0x642e66dbu : 0x2cccb6cau); // 2 add
|
||||
r0 = rotl_imm(r0, 19u); // 3 rotl
|
||||
r7 = rotr_var(r7, r6); // 4 rotr
|
||||
r7 = r7 + r4 + ((((sel >> 21u) & 1u) != 0u) ? 0xc1535555u : 0xee02465fu); // 5 add
|
||||
r1 = mul_hi(r1, r7); // 6 mulhi
|
||||
r4 = r4 ^ ds[r2 & mask]; // 7 load
|
||||
r7 = r7 ^ ds[r4 & mask]; // 8 load
|
||||
r0 = r0 ^ ds[r3 & mask]; // 9 load
|
||||
r5 = r5 ^ ds[r1 & mask]; // 10 load
|
||||
r1 = r1 ^ ds[r5 & mask]; // 11 load
|
||||
r3 = mul_hi(r3, r5); // 12 mulhi
|
||||
r1 = r1 ^ ds[r3 & mask]; // 13 load
|
||||
r0 = r0 - r3; // 14 sub
|
||||
r5 = r1 * r3 + r5; // 15 mad
|
||||
r6 = mul_hi(r6, r1); // 16 mulhi
|
||||
r5 = r5 + r2 + ((((sel >> 28u) & 1u) != 0u) ? 0x8b965b57u : 0x697b3d00u); // 17 add
|
||||
r0 = mul_hi(r0, r6); // 18 mulhi
|
||||
r5 = rotr_var(r5, r3); // 19 rotr
|
||||
r5 = mul_hi(r5, r2); // 20 mulhi
|
||||
r1 = r1 + r0 + ((((sel >> 1u) & 1u) != 0u) ? 0x6d7e8d05u : 0xebcf247au); // 21 add
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 22 add
|
||||
r1 = mul_hi(r1, r5); // 23 mulhi
|
||||
r2 = r2 - r5; // 24 sub
|
||||
r7 = r7 + r4 + ((((sel >> 2u) & 1u) != 0u) ? 0x699ef1bbu : 0x08ffa6c7u); // 25 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 26 shfl
|
||||
r7 = r7 + r1 + ((((sel >> 14u) & 1u) != 0u) ? 0xb4ead2fbu : 0xe60fea84u); // 27 add
|
||||
r3 = r3 + r1 + ((((sel >> 6u) & 1u) != 0u) ? 0x8f30d21du : 0x65c76dabu); // 28 add
|
||||
r2 = r2 ^ ds[r1 & mask]; // 29 load
|
||||
r5 = r5 ^ ds[r7 & mask]; // 30 load
|
||||
r2 = r2 ^ ds[r5 & mask]; // 31 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r1 = r1 ^ t_; } // 32 shfl
|
||||
r4 = r5 * r7 + r4; // 33 mad
|
||||
r4 = r4 + r2 + ((((sel >> 21u) & 1u) != 0u) ? 0xc7ce690cu : 0x0480debeu); // 34 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 8u); r3 = r3 ^ t_; } // 35 shfl
|
||||
r7 = r7 + r1 + ((((sel >> 2u) & 1u) != 0u) ? 0xe10c2c95u : 0xc53b542eu); // 36 add
|
||||
r5 = r5 ^ r7; // 37 xor
|
||||
r2 = r2 | r1; // 38 or
|
||||
r1 = mul_hi(r1, r0); // 39 mulhi
|
||||
r6 = rotl_imm(r6, 19u); // 40 rotl
|
||||
r4 = mul_hi(r4, r6); // 41 mulhi
|
||||
r6 = r6 - r0; // 42 sub
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r6 = r6 ^ t_; } // 43 shfl
|
||||
r4 = r4 ^ ds[r2 & mask]; // 44 load
|
||||
r1 = r1 ^ r3; // 45 xor
|
||||
r7 = r7 ^ ds[r0 & mask]; // 46 load
|
||||
r3 = r3 ^ ds[r1 & mask]; // 47 load
|
||||
r5 = r5 * r3; // 48 mul
|
||||
r1 = r1 - r5; // 49 sub
|
||||
r2 = rotl_imm(r2, 8u); // 50 rotl
|
||||
r1 = r1 + r5 + ((((sel >> 23u) & 1u) != 0u) ? 0x77b9bd43u : 0xa900fec4u); // 51 add
|
||||
r4 = r4 ^ ds[r7 & mask]; // 52 load
|
||||
r2 = r2 - r7; // 53 sub
|
||||
r4 = r4 ^ r0; // 54 xor
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 55 add
|
||||
r2 = r2 ^ ds[r4 & mask]; // 56 load
|
||||
r0 = r1 * r4 + r0; // 57 mad
|
||||
r3 = r3 ^ ds[r5 & mask]; // 58 load
|
||||
r5 = r5 | r6; // 59 or
|
||||
r6 = r5 * r7 + r6; // 60 mad
|
||||
r4 = rotl_imm(r4, 28u); // 61 rotl
|
||||
r5 = mul_hi(r5, r0); // 62 mulhi
|
||||
r3 = r3 ^ ds[r6 & mask]; // 63 load
|
||||
r7 = r7 ^ r0; // 1 xor
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 1u); r3 = r3 ^ t_; } // 2 shfl
|
||||
r4 = r4 - r1; // 3 sub
|
||||
r2 = r0 * r4 + r2; // 4 mad
|
||||
r4 = rotl_imm(r4, 9u); // 5 rotl
|
||||
r0 = rotr_var(r0, r2); // 6 rotr
|
||||
r6 = r6 ^ ds[((rotl_imm(r7 * 0x9ad30d99u, 29u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r2 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r0 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r1 * 0x9ad30d99u, 29u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
|
||||
r5 = r5 * r4; // 12 mul
|
||||
r7 = r7 ^ ds[((rotl_imm(r6 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 16u); r6 = r6 ^ t_; } // 14 shfl
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 1u); r3 = r3 ^ t_; } // 15 shfl
|
||||
r2 = r2 | r7; // 16 or
|
||||
r0 = rotr_var(r0, r6); // 17 rotr
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
|
||||
r7 = rotl_imm(r7, 24u); // 19 rotl
|
||||
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
|
||||
r2 = r2 | r6; // 21 or
|
||||
r6 = r5 * r4 + r6; // 22 mad
|
||||
r1 = r1 | r0; // 23 or
|
||||
r6 = r6 ^ r1; // 24 xor
|
||||
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
|
||||
r7 = rotr_var(r7, r0); // 26 rotr
|
||||
r4 = r5 * r7 + r4; // 27 mad
|
||||
r3 = r2 * r4 + r3; // 28 mad
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x9ad30d99u, 29u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
|
||||
r2 = r2 ^ ds[((rotl_imm(r3 * 0x9ad30d99u, 29u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r5 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
|
||||
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
|
||||
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
|
||||
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
|
||||
r2 = mul_hi(r2, r5); // 35 mulhi
|
||||
r4 = r4 ^ r2; // 36 xor
|
||||
r6 = r6 * r5; // 37 mul
|
||||
r7 = r7 ^ r0; // 38 xor
|
||||
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
|
||||
r2 = rotr_var(r2, r3); // 40 rotr
|
||||
r7 = r7 - r0; // 41 sub
|
||||
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 43 shfl
|
||||
r0 = r0 ^ ds[((rotl_imm(r7 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
|
||||
r3 = r3 ^ ds[((rotl_imm(r1 * 0x9ad30d99u, 29u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
|
||||
r6 = r6 ^ ds[((rotl_imm(r3 * 0x9ad30d99u, 29u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
|
||||
r4 = mul_hi(r4, r2); // 48 mulhi
|
||||
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r0 = r0 ^ t_; } // 50 shfl
|
||||
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r2 * 0x9ad30d99u, 29u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
|
||||
r6 = rotl_imm(r6, 12u); // 53 rotl
|
||||
r3 = rotl_imm(r3, 12u); // 54 rotl
|
||||
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x9ad30d99u, 29u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
|
||||
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r0 * 0x9ad30d99u, 29u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
|
||||
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
|
||||
r3 = mul_hi(r3, r2); // 60 mulhi
|
||||
r6 = r6 | r4; // 61 or
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r5 = r5 ^ t_; } // 62 shfl
|
||||
r3 = r3 ^ ds[((rotl_imm(r6 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
|
|
|
|||
|
|
@ -36,69 +36,69 @@ __global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t ba
|
|||
for (uint32_t it = 0u; it < 8u; ++it) {
|
||||
uint32_t sel = r0;
|
||||
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
|
||||
r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r0, 4); // 1 shfl
|
||||
r3 = r3 + r2 + ((((sel >> 10u) & 1u) != 0u) ? 0x642e66dbu : 0x2cccb6cau); // 2 add
|
||||
r0 = rotl_imm(r0, 19u); // 3 rotl
|
||||
r7 = rotr_var(r7, r6); // 4 rotr
|
||||
r7 = r7 + r4 + ((((sel >> 21u) & 1u) != 0u) ? 0xc1535555u : 0xee02465fu); // 5 add
|
||||
r1 = __umulhi(r1, r7); // 6 mulhi
|
||||
r4 = r4 ^ ds[r2 & mask]; // 7 load
|
||||
r7 = r7 ^ ds[r4 & mask]; // 8 load
|
||||
r0 = r0 ^ ds[r3 & mask]; // 9 load
|
||||
r5 = r5 ^ ds[r1 & mask]; // 10 load
|
||||
r1 = r1 ^ ds[r5 & mask]; // 11 load
|
||||
r3 = __umulhi(r3, r5); // 12 mulhi
|
||||
r1 = r1 ^ ds[r3 & mask]; // 13 load
|
||||
r0 = r0 - r3; // 14 sub
|
||||
r5 = r1 * r3 + r5; // 15 mad
|
||||
r6 = __umulhi(r6, r1); // 16 mulhi
|
||||
r5 = r5 + r2 + ((((sel >> 28u) & 1u) != 0u) ? 0x8b965b57u : 0x697b3d00u); // 17 add
|
||||
r0 = __umulhi(r0, r6); // 18 mulhi
|
||||
r5 = rotr_var(r5, r3); // 19 rotr
|
||||
r5 = __umulhi(r5, r2); // 20 mulhi
|
||||
r1 = r1 + r0 + ((((sel >> 1u) & 1u) != 0u) ? 0x6d7e8d05u : 0xebcf247au); // 21 add
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 22 add
|
||||
r1 = __umulhi(r1, r5); // 23 mulhi
|
||||
r2 = r2 - r5; // 24 sub
|
||||
r7 = r7 + r4 + ((((sel >> 2u) & 1u) != 0u) ? 0x699ef1bbu : 0x08ffa6c7u); // 25 add
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 26 shfl
|
||||
r7 = r7 + r1 + ((((sel >> 14u) & 1u) != 0u) ? 0xb4ead2fbu : 0xe60fea84u); // 27 add
|
||||
r3 = r3 + r1 + ((((sel >> 6u) & 1u) != 0u) ? 0x8f30d21du : 0x65c76dabu); // 28 add
|
||||
r2 = r2 ^ ds[r1 & mask]; // 29 load
|
||||
r5 = r5 ^ ds[r7 & mask]; // 30 load
|
||||
r2 = r2 ^ ds[r5 & mask]; // 31 load
|
||||
r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r7, 4); // 32 shfl
|
||||
r4 = r5 * r7 + r4; // 33 mad
|
||||
r4 = r4 + r2 + ((((sel >> 21u) & 1u) != 0u) ? 0xc7ce690cu : 0x0480debeu); // 34 add
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r7, 8); // 35 shfl
|
||||
r7 = r7 + r1 + ((((sel >> 2u) & 1u) != 0u) ? 0xe10c2c95u : 0xc53b542eu); // 36 add
|
||||
r5 = r5 ^ r7; // 37 xor
|
||||
r2 = r2 | r1; // 38 or
|
||||
r1 = __umulhi(r1, r0); // 39 mulhi
|
||||
r6 = rotl_imm(r6, 19u); // 40 rotl
|
||||
r4 = __umulhi(r4, r6); // 41 mulhi
|
||||
r6 = r6 - r0; // 42 sub
|
||||
r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 43 shfl
|
||||
r4 = r4 ^ ds[r2 & mask]; // 44 load
|
||||
r1 = r1 ^ r3; // 45 xor
|
||||
r7 = r7 ^ ds[r0 & mask]; // 46 load
|
||||
r3 = r3 ^ ds[r1 & mask]; // 47 load
|
||||
r5 = r5 * r3; // 48 mul
|
||||
r1 = r1 - r5; // 49 sub
|
||||
r2 = rotl_imm(r2, 8u); // 50 rotl
|
||||
r1 = r1 + r5 + ((((sel >> 23u) & 1u) != 0u) ? 0x77b9bd43u : 0xa900fec4u); // 51 add
|
||||
r4 = r4 ^ ds[r7 & mask]; // 52 load
|
||||
r2 = r2 - r7; // 53 sub
|
||||
r4 = r4 ^ r0; // 54 xor
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 55 add
|
||||
r2 = r2 ^ ds[r4 & mask]; // 56 load
|
||||
r0 = r1 * r4 + r0; // 57 mad
|
||||
r3 = r3 ^ ds[r5 & mask]; // 58 load
|
||||
r5 = r5 | r6; // 59 or
|
||||
r6 = r5 * r7 + r6; // 60 mad
|
||||
r4 = rotl_imm(r4, 28u); // 61 rotl
|
||||
r5 = __umulhi(r5, r0); // 62 mulhi
|
||||
r3 = r3 ^ ds[r6 & mask]; // 63 load
|
||||
r7 = r7 ^ r0; // 1 xor
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r6, 1); // 2 shfl
|
||||
r4 = r4 - r1; // 3 sub
|
||||
r2 = r0 * r4 + r2; // 4 mad
|
||||
r4 = rotl_imm(r4, 9u); // 5 rotl
|
||||
r0 = rotr_var(r0, r2); // 6 rotr
|
||||
r6 = r6 ^ ds[((rotl_imm(r7 * 0x9ad30d99u, 29u) & 0x03ffffffu) | 0x04000000u) & mask]; // 7 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & mask]; // 8 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r2 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & mask]; // 9 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r0 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & mask]; // 10 load
|
||||
r7 = r7 ^ ds[((rotl_imm(r1 * 0x9ad30d99u, 29u) & 0x0fffffffu) | 0x00000000u) & mask]; // 11 load
|
||||
r5 = r5 * r4; // 12 mul
|
||||
r7 = r7 ^ ds[((rotl_imm(r6 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & mask]; // 13 load
|
||||
r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r5, 16); // 14 shfl
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r5, 1); // 15 shfl
|
||||
r2 = r2 | r7; // 16 or
|
||||
r0 = rotr_var(r0, r6); // 17 rotr
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 18 add
|
||||
r7 = rotl_imm(r7, 24u); // 19 rotl
|
||||
r6 = r6 + r7 + ((((sel >> 23u) & 1u) != 0u) ? 0x8c9f0ef8u : 0x52334d12u); // 20 add
|
||||
r2 = r2 | r6; // 21 or
|
||||
r6 = r5 * r4 + r6; // 22 mad
|
||||
r1 = r1 | r0; // 23 or
|
||||
r6 = r6 ^ r1; // 24 xor
|
||||
r2 = r2 + r6 + ((((sel >> 17u) & 1u) != 0u) ? 0x9cec0e12u : 0x659fc3d3u); // 25 add
|
||||
r7 = rotr_var(r7, r0); // 26 rotr
|
||||
r4 = r5 * r7 + r4; // 27 mad
|
||||
r3 = r2 * r4 + r3; // 28 mad
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x9ad30d99u, 29u) & 0x0fffffffu) | 0x00000000u) & mask]; // 29 load
|
||||
r2 = r2 ^ ds[((rotl_imm(r3 * 0x9ad30d99u, 29u) & 0x03ffffffu) | 0x08000000u) & mask]; // 30 load
|
||||
r1 = r1 ^ ds[((rotl_imm(r5 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & mask]; // 31 load
|
||||
r7 = r7 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xe403240eu : 0x070888a8u); // 32 add
|
||||
r2 = r2 + r0 + ((((sel >> 28u) & 1u) != 0u) ? 0x29701828u : 0xf2e46d55u); // 33 add
|
||||
r2 = r2 + r3 + ((((sel >> 14u) & 1u) != 0u) ? 0x343b7aeeu : 0x58f75b87u); // 34 add
|
||||
r2 = __umulhi(r2, r5); // 35 mulhi
|
||||
r4 = r4 ^ r2; // 36 xor
|
||||
r6 = r6 * r5; // 37 mul
|
||||
r7 = r7 ^ r0; // 38 xor
|
||||
r7 = r7 + r2 + ((((sel >> 19u) & 1u) != 0u) ? 0x32bbd117u : 0xb8180e9du); // 39 add
|
||||
r2 = rotr_var(r2, r3); // 40 rotr
|
||||
r7 = r7 - r0; // 41 sub
|
||||
r4 = r4 + r3 + ((((sel >> 4u) & 1u) != 0u) ? 0x6df7aed4u : 0x6ced15b7u); // 42 add
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 43 shfl
|
||||
r0 = r0 ^ ds[((rotl_imm(r7 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & mask]; // 44 load
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 45 add
|
||||
r3 = r3 ^ ds[((rotl_imm(r1 * 0x9ad30d99u, 29u) & 0x03ffffffu) | 0x00000000u) & mask]; // 46 load
|
||||
r6 = r6 ^ ds[((rotl_imm(r3 * 0x9ad30d99u, 29u) & 0x0fffffffu) | 0x00000000u) & mask]; // 47 load
|
||||
r4 = __umulhi(r4, r2); // 48 mulhi
|
||||
r5 = r5 + r0 + ((((sel >> 7u) & 1u) != 0u) ? 0xf572bdb9u : 0xa8bae6dfu); // 49 add
|
||||
r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r7, 4); // 50 shfl
|
||||
r6 = r6 + r0 + ((((sel >> 19u) & 1u) != 0u) ? 0x11e17c61u : 0x383b9260u); // 51 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r2 * 0x9ad30d99u, 29u) & 0x0fffffffu) | 0x00000000u) & mask]; // 52 load
|
||||
r6 = rotl_imm(r6, 12u); // 53 rotl
|
||||
r3 = rotl_imm(r3, 12u); // 54 rotl
|
||||
r2 = r2 + r1 + ((((sel >> 10u) & 1u) != 0u) ? 0x18d67dbbu : 0xac6be8e3u); // 55 add
|
||||
r1 = r1 ^ ds[((rotl_imm(r4 * 0x9ad30d99u, 29u) & 0x0fffffffu) | 0x00000000u) & mask]; // 56 load
|
||||
r5 = r5 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0xa75cd60du : 0xf03673feu); // 57 add
|
||||
r5 = r5 ^ ds[((rotl_imm(r0 * 0x9ad30d99u, 29u) & 0x03ffffffu) | 0x00000000u) & mask]; // 58 load
|
||||
r1 = r1 + r4 + ((((sel >> 7u) & 1u) != 0u) ? 0x8dfb96bbu : 0xdecd4794u); // 59 add
|
||||
r3 = __umulhi(r3, r2); // 60 mulhi
|
||||
r6 = r6 | r4; // 61 or
|
||||
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 62 shfl
|
||||
r3 = r3 ^ ds[((rotl_imm(r6 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & mask]; // 63 load
|
||||
}
|
||||
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
|
|
|
|||
|
|
@ -105,5 +105,9 @@ IGNEUM_HD void mh_item(const uint32_t* cache, uint32_t t, uint32_t* s) {
|
|||
}
|
||||
for (uint32_t j = 0u; j < 4u; ++j) mh_mixer(s, 0x9E3779B9u * (32u + j + 1u));
|
||||
}
|
||||
// dataset[w] without the dataset: derive item w >> 4 and take word w & 15.
|
||||
IGNEUM_HD uint32_t mh_word(const uint32_t* cache, uint32_t w) { uint32_t s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; }
|
||||
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 0 2 12 13 of w.
|
||||
IGNEUM_HD uint32_t mh_j(uint32_t w) { return ((w >> 0u) & 1u) | (((w >> 2u) & 1u) << 1) | (((w >> 12u) & 1u) << 2) | (((w >> 13u) & 1u) << 3); }
|
||||
IGNEUM_HD uint32_t mh_t(uint32_t w) { w = (w & 0x00001fffu) | ((w >> 14u) << 13u); w = (w & 0x00000fffu) | ((w >> 13u) << 12u); w = (w & 0x00000003u) | ((w >> 3u) << 2u); w = (w & 0x00000000u) | ((w >> 1u) << 0u); return w; }
|
||||
IGNEUM_HD uint32_t mh_addr(uint32_t t, uint32_t j) { uint32_t w = t; w = ((w >> 0u) << 1u) | (w & 0x00000000u) | (((j >> 0u) & 1u) << 0u); w = ((w >> 2u) << 3u) | (w & 0x00000003u) | (((j >> 1u) & 1u) << 2u); w = ((w >> 12u) << 13u) | (w & 0x00000fffu) | (((j >> 2u) & 1u) << 12u); w = ((w >> 13u) << 14u) | (w & 0x00001fffu) | (((j >> 3u) & 1u) << 13u); return w; }
|
||||
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
|
||||
IGNEUM_HD uint32_t mh_word(const uint32_t* cache, uint32_t w) { uint32_t s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }
|
||||
|
|
|
|||
|
|
@ -90,8 +90,12 @@ inline void mh_item(device const uint* cache, uint t, thread uint* s) {
|
|||
}
|
||||
for (uint j = 0u; j < 4u; ++j) mh_mixer(s, 0x9E3779B9u * (32u + j + 1u));
|
||||
}
|
||||
// dataset[w] without the dataset: derive item w >> 4 and take word w & 15.
|
||||
inline uint mh_word(device const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; }
|
||||
// Era layout (docs/plans/era-layout.md 1.2): dataset word w holds word mh_j(w) of item mh_t(w); j's bits sit at positions 0 2 12 13 of w.
|
||||
inline uint mh_j(uint w) { return ((w >> 0u) & 1u) | (((w >> 2u) & 1u) << 1) | (((w >> 12u) & 1u) << 2) | (((w >> 13u) & 1u) << 3); }
|
||||
inline uint mh_t(uint w) { w = (w & 0x00001fffu) | ((w >> 14u) << 13u); w = (w & 0x00000fffu) | ((w >> 13u) << 12u); w = (w & 0x00000003u) | ((w >> 3u) << 2u); w = (w & 0x00000000u) | ((w >> 1u) << 0u); return w; }
|
||||
inline uint mh_addr(uint t, uint j) { uint w = t; w = ((w >> 0u) << 1u) | (w & 0x00000000u) | (((j >> 0u) & 1u) << 0u); w = ((w >> 2u) << 3u) | (w & 0x00000003u) | (((j >> 1u) & 1u) << 2u); w = ((w >> 12u) << 13u) | (w & 0x00000fffu) | (((j >> 2u) & 1u) << 12u); w = ((w >> 13u) << 14u) | (w & 0x00001fffu) | (((j >> 3u) & 1u) << 13u); return w; }
|
||||
// dataset[w] without the dataset: derive item mh_t(w) and take word mh_j(w).
|
||||
inline uint mh_word(device const uint* cache, uint w) { uint s[16]; mh_item(cache, mh_t(w), s); return s[mh_j(w)]; }
|
||||
|
||||
// One thread per segment (2^16 threads).
|
||||
kernel void igneum_cache_fill(device uint* cache [[buffer(0)]], uint gid [[thread_position_in_grid]]) {
|
||||
|
|
@ -102,6 +106,5 @@ kernel void igneum_build(device const uint* cache [[buffer(0)]], device uint* da
|
|||
uint gid [[thread_position_in_grid]]) {
|
||||
uint s[16];
|
||||
mh_item(cache, gid, s);
|
||||
device uint* d = dataset + gid * 16u;
|
||||
for (uint i = 0u; i < 16u; ++i) d[i] = s[i];
|
||||
for (uint i = 0u; i < 16u; ++i) dataset[mh_addr(gid, i)] = s[i];
|
||||
}
|
||||
|
|
|
|||
|
|
@ -27,14 +27,14 @@
|
|||
#define IGNEUM_INSTR_COUNT 64
|
||||
#define IGNEUM_LOADS_PER_HASH 128
|
||||
#define IGNEUM_WIDE_LOADS_PER_HASH 0
|
||||
#define IGNEUM_OP_MIX "load=16 add=13 mulhi=9 shfl=5 sub=5 mad=4 rotl=4 xor=3 or=2 rotr=2 mul=1"
|
||||
#define IGNEUM_OP_MIX "load=16 add=15 shfl=6 mad=4 or=4 rotl=4 rotr=4 xor=4 mulhi=3 mul=2 sub=2"
|
||||
// Program class v3 (Counter ASIC 2.0, docs/plans/counter-asic-2-rollout.md): generator version 3; a worker that
|
||||
// runs another class refuses this pack, and a job line names the class it wants (class=v3 era=<hex>).
|
||||
#define IGNEUM_PROGRAM_CLASS "v3"
|
||||
#define IGNEUM_ERA_SEED_HEX "edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07"
|
||||
// Class v3 construction (Counter ASIC 2.0, 5 October 2026, docs/plans/mixer-x4.md): version 2 loads; the dataset item
|
||||
// derivation applies the mixer IGNEUM_MIXER_MULT times per round (memhard.h), and the cache follows the growth rule.
|
||||
#define IGNEUM_LOAD_CLASS "mx4"
|
||||
#define IGNEUM_LOAD_CLASS "mx4-erad810f22d"
|
||||
#define IGNEUM_CLASS_MIXER_MULT 4
|
||||
#define IGNEUM_CACHE_GROWTH 1 // 1: cache words = 2^(26 + doublings(day)), doublings = floor(log2(1 + day / 1460))
|
||||
#define IGNEUM_LOAD_SLOTS 16
|
||||
|
|
@ -43,6 +43,18 @@
|
|||
#define IGNEUM_BYTES_PER_HASH 512
|
||||
#define IGNEUM_FOLD_ROT 11
|
||||
#define IGNEUM_FOLD_MUL 0x9e3779b1u
|
||||
// Era layout (5 October 2026, docs/plans/era-layout.md): NOT the lottery hash. Every dataset load reads
|
||||
// idx = ((rotl(src * STRIDE_MUL, STRIDE_ROT) & window mask) | window offset) & MASK; the window of a load site is the
|
||||
// dataset, a half or a quarter of it (IGNEUM_ERA_WINDOWS: site:shrink:offset); dataset word w holds word j(w) of item
|
||||
// t(w) with j's bits at the INTERLEAVE positions (memhard.h: mh_t, mh_j, mh_addr).
|
||||
#define IGNEUM_ERA_LABEL "d810f22d"
|
||||
#define IGNEUM_ERA_SEED_WORDS { 0xd810f22du, 0xcf9dc883u, 0x5f3e571fu, 0x2b7d97fcu, 0x810012c4u, 0x002d9798u, 0xa4180c41u, 0xa4562c21u }
|
||||
#define IGNEUM_ERA_ALLOWED_WIDTHS { 1, 0, 0 } // words, ascending, 0 = unused; one entry pins the width
|
||||
#define IGNEUM_ERA_WIDTH_WORDS 1
|
||||
#define IGNEUM_ERA_STRIDE_MUL 0x9ad30d99u
|
||||
#define IGNEUM_ERA_STRIDE_ROT 29
|
||||
#define IGNEUM_ERA_INTERLEAVE { 0, 2, 12, 13 }
|
||||
#define IGNEUM_ERA_WINDOWS "7:2:1 8:1:1 9:1:1 10:1:1 11:0:0 13:1:1 29:0:0 30:2:2 31:1:1 44:1:1 46:2:0 47:0:0 52:0:0 56:0:0 58:2:0 63:1:1"
|
||||
// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h)
|
||||
#define IGNEUM_DATASET_MODE 1
|
||||
|
||||
|
|
|
|||
|
|
@ -17,7 +17,7 @@
|
|||
"loads_per_hash": 128,
|
||||
"program_class": "v3",
|
||||
"era_seed_bytes": "edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07",
|
||||
"load_class": "mx4",
|
||||
"load_class": "mx4-erad810f22d",
|
||||
"mixer_mult": 4,
|
||||
"cache_growth": true,
|
||||
"mixer": "class v3 (Counter ASIC 2.0, 5 October 2026, docs/plans/mixer-x4.md): every mixer application of the item derivation is 4 applications with round keys (r * 4 + j + 1) * 0x9E3779B9, the 8 dependent cache reads per item unchanged; cache growth rule option C: cache words = 2^(26 + doublings(day)), dataset words = 2^(genesis_log2 + doublings(day)), doublings(day) = floor(log2(1 + day / 1460)) for day = days since genesis",
|
||||
|
|
@ -26,7 +26,21 @@
|
|||
"load_width_counts_4_16_64": [16, 0, 0],
|
||||
"bytes_per_hash": 512,
|
||||
"wide_load": "read-width experiment (5 October 2026, docs/plans/read-width.md), NOT the lottery hash: a load of W words (width field, 4 or 16) reads dataset[b .. b + W) with b = (src & mask) & ~(W - 1) and folds every word into dst: x = dst ^ w[0]; for j in 1..W: x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]; dst = x; width 1 is the plain load; the width is drawn per instruction from the class mix with one extra below(100) draw after the nine of version 2, and the program id is FNV-1a 64 over 'igneum-program-rw/' || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots",
|
||||
"op_mix": {"load": 16, "add": 13, "mulhi": 9, "shfl": 5, "sub": 5, "mad": 4, "rotl": 4, "xor": 3, "or": 2, "rotr": 2, "mul": 1},
|
||||
"era": {
|
||||
"label": "d810f22d",
|
||||
"seed_words": ["0xd810f22d", "0xcf9dc883", "0x5f3e571f", "0x2b7d97fc", "0x810012c4", "0x002d9798", "0xa4180c41", "0xa4562c21"],
|
||||
"draw": "docs/plans/era-layout.md 1.1: SplitMix64 seeded with seed_words[0] | seed_words[1] << 32 of seed_words_from_bytes('igneum-era/' || n_le64 || E_n); width = allowed[below(|allowed|)], stride_mul = low32(next()) | 1, stride_rot = 1 + below(31), then four next() draws for a partial Fisher-Yates over positions log2(W)..15 of which 4 - log2(W) are used",
|
||||
"allowed_widths": [1],
|
||||
"width_words": 1,
|
||||
"stride_mul": "0x9ad30d99",
|
||||
"stride_rot": 29,
|
||||
"interleave": [0, 2, 12, 13],
|
||||
"address": "y = rotl(src * stride_mul, stride_rot); k = min(win, D - 26); idx = ((y & (mask >> k)) | ((off & (2^k - 1)) << (D - k))) & mask; a wide load aligns idx down to W words",
|
||||
"windows": "per instruction, after the width roll: win = below(3), off = low32(next()) & (2^win - 1); used on a load slot (the instruction's win and off fields)",
|
||||
"dataset_word": "dataset[w] = item(t(w))[j(w)]: j(w) gathers the bits of w at the interleave positions, t(w) is w with those bits removed",
|
||||
"program_id_suffix": "'era/' || allowed[3] || width_words || stride_mul_le32 || stride_rot_le32 || interleave[4]"
|
||||
},
|
||||
"op_mix": {"load": 16, "add": 15, "shfl": 6, "mad": 4, "or": 4, "rotl": 4, "rotr": 4, "xor": 4, "mulhi": 3, "mul": 2, "sub": 2},
|
||||
"register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]",
|
||||
"splitmix32": "x ^= x>>16; x *= 0x7feb352d; x ^= x>>15; x *= 0x846ca68b; x ^= x>>16",
|
||||
"iteration": "sel = r0 sampled once at the top of each iteration, then all instructions in order",
|
||||
|
|
@ -65,69 +79,69 @@
|
|||
"word": "dataset[w] = item(w >> 4)[w & 15]"
|
||||
},
|
||||
"instructions": [
|
||||
{"i": 0, "op": "add", "dst": 4, "src": 5, "src2": 7, "imm": "0xea86e152", "imm2": "0x5810667a", "rot": 27, "bit": 13, "mask": 2, "width": 1},
|
||||
{"i": 1, "op": "shfl", "dst": 2, "src": 0, "src2": 7, "imm": "0xe3c2f9cb", "imm2": "0xde0bea4e", "rot": 6, "bit": 9, "mask": 4, "width": 1},
|
||||
{"i": 2, "op": "add", "dst": 3, "src": 2, "src2": 0, "imm": "0x2cccb6ca", "imm2": "0x642e66db", "rot": 27, "bit": 10, "mask": 2, "width": 1},
|
||||
{"i": 3, "op": "rotl", "dst": 0, "src": 2, "src2": 1, "imm": "0xb59e83b2", "imm2": "0x19c26fb9", "rot": 19, "bit": 20, "mask": 4, "width": 1},
|
||||
{"i": 4, "op": "rotr", "dst": 7, "src": 6, "src2": 5, "imm": "0x6f055f55", "imm2": "0x550e4ea1", "rot": 31, "bit": 20, "mask": 4, "width": 1},
|
||||
{"i": 5, "op": "add", "dst": 7, "src": 4, "src2": 7, "imm": "0xee02465f", "imm2": "0xc1535555", "rot": 31, "bit": 21, "mask": 4, "width": 1},
|
||||
{"i": 6, "op": "mulhi", "dst": 1, "src": 7, "src2": 5, "imm": "0x7826a6a7", "imm2": "0x946f7818", "rot": 18, "bit": 21, "mask": 8, "width": 1},
|
||||
{"i": 7, "op": "load", "dst": 4, "src": 2, "src2": 0, "imm": "0x5d080878", "imm2": "0xdf885578", "rot": 5, "bit": 18, "mask": 4, "width": 1},
|
||||
{"i": 8, "op": "load", "dst": 7, "src": 4, "src2": 1, "imm": "0x875bbb36", "imm2": "0x594a838f", "rot": 24, "bit": 4, "mask": 2, "width": 1},
|
||||
{"i": 9, "op": "load", "dst": 0, "src": 3, "src2": 1, "imm": "0xe5e607c2", "imm2": "0xa5cd9f75", "rot": 26, "bit": 28, "mask": 4, "width": 1},
|
||||
{"i": 10, "op": "load", "dst": 5, "src": 1, "src2": 3, "imm": "0xea3f7b43", "imm2": "0x10c8d4e7", "rot": 5, "bit": 29, "mask": 8, "width": 1},
|
||||
{"i": 11, "op": "load", "dst": 1, "src": 5, "src2": 4, "imm": "0x4454980f", "imm2": "0xebd31581", "rot": 10, "bit": 28, "mask": 16, "width": 1},
|
||||
{"i": 12, "op": "mulhi", "dst": 3, "src": 5, "src2": 7, "imm": "0xbfd5615c", "imm2": "0xd8224ae2", "rot": 21, "bit": 12, "mask": 2, "width": 1},
|
||||
{"i": 13, "op": "load", "dst": 1, "src": 3, "src2": 5, "imm": "0x35c07cc5", "imm2": "0xe82db54f", "rot": 7, "bit": 6, "mask": 8, "width": 1},
|
||||
{"i": 14, "op": "sub", "dst": 0, "src": 3, "src2": 0, "imm": "0x587ee0f1", "imm2": "0xfd23eefd", "rot": 16, "bit": 21, "mask": 16, "width": 1},
|
||||
{"i": 15, "op": "mad", "dst": 5, "src": 1, "src2": 3, "imm": "0x6974dd29", "imm2": "0xc8148960", "rot": 3, "bit": 11, "mask": 2, "width": 1},
|
||||
{"i": 16, "op": "mulhi", "dst": 6, "src": 1, "src2": 7, "imm": "0x3072c3c6", "imm2": "0x55ee21f8", "rot": 26, "bit": 1, "mask": 1, "width": 1},
|
||||
{"i": 17, "op": "add", "dst": 5, "src": 2, "src2": 2, "imm": "0x697b3d00", "imm2": "0x8b965b57", "rot": 9, "bit": 28, "mask": 1, "width": 1},
|
||||
{"i": 18, "op": "mulhi", "dst": 0, "src": 6, "src2": 3, "imm": "0x2910cacb", "imm2": "0x6ac79431", "rot": 7, "bit": 6, "mask": 4, "width": 1},
|
||||
{"i": 19, "op": "rotr", "dst": 5, "src": 3, "src2": 0, "imm": "0xed96a94a", "imm2": "0x4c988c10", "rot": 24, "bit": 21, "mask": 8, "width": 1},
|
||||
{"i": 20, "op": "mulhi", "dst": 5, "src": 2, "src2": 7, "imm": "0x40d2fc76", "imm2": "0x2f7c7eca", "rot": 13, "bit": 8, "mask": 4, "width": 1},
|
||||
{"i": 21, "op": "add", "dst": 1, "src": 0, "src2": 5, "imm": "0xebcf247a", "imm2": "0x6d7e8d05", "rot": 31, "bit": 1, "mask": 16, "width": 1},
|
||||
{"i": 22, "op": "add", "dst": 7, "src": 5, "src2": 7, "imm": "0xf66e7017", "imm2": "0xb9e3577e", "rot": 9, "bit": 12, "mask": 16, "width": 1},
|
||||
{"i": 23, "op": "mulhi", "dst": 1, "src": 5, "src2": 7, "imm": "0xfdfe72fd", "imm2": "0x735eeb8d", "rot": 30, "bit": 25, "mask": 8, "width": 1},
|
||||
{"i": 24, "op": "sub", "dst": 2, "src": 5, "src2": 0, "imm": "0x32cf1258", "imm2": "0xd813deb6", "rot": 30, "bit": 6, "mask": 16, "width": 1},
|
||||
{"i": 25, "op": "add", "dst": 7, "src": 4, "src2": 2, "imm": "0x08ffa6c7", "imm2": "0x699ef1bb", "rot": 7, "bit": 2, "mask": 16, "width": 1},
|
||||
{"i": 26, "op": "shfl", "dst": 3, "src": 4, "src2": 5, "imm": "0x6f53c70d", "imm2": "0x3357513f", "rot": 3, "bit": 26, "mask": 2, "width": 1},
|
||||
{"i": 27, "op": "add", "dst": 7, "src": 1, "src2": 4, "imm": "0xe60fea84", "imm2": "0xb4ead2fb", "rot": 14, "bit": 14, "mask": 1, "width": 1},
|
||||
{"i": 28, "op": "add", "dst": 3, "src": 1, "src2": 0, "imm": "0x65c76dab", "imm2": "0x8f30d21d", "rot": 24, "bit": 6, "mask": 1, "width": 1},
|
||||
{"i": 29, "op": "load", "dst": 2, "src": 1, "src2": 2, "imm": "0x82fad9a6", "imm2": "0x8c6358db", "rot": 7, "bit": 31, "mask": 2, "width": 1},
|
||||
{"i": 30, "op": "load", "dst": 5, "src": 7, "src2": 3, "imm": "0x6e947ee0", "imm2": "0xaf9a2dda", "rot": 2, "bit": 30, "mask": 4, "width": 1},
|
||||
{"i": 31, "op": "load", "dst": 2, "src": 5, "src2": 1, "imm": "0x608bb7ce", "imm2": "0x4be663db", "rot": 19, "bit": 1, "mask": 16, "width": 1},
|
||||
{"i": 32, "op": "shfl", "dst": 1, "src": 7, "src2": 6, "imm": "0x88e52e20", "imm2": "0x77647269", "rot": 20, "bit": 25, "mask": 4, "width": 1},
|
||||
{"i": 33, "op": "mad", "dst": 4, "src": 5, "src2": 7, "imm": "0xb48420ae", "imm2": "0x3d1f2485", "rot": 14, "bit": 28, "mask": 1, "width": 1},
|
||||
{"i": 34, "op": "add", "dst": 4, "src": 2, "src2": 3, "imm": "0x0480debe", "imm2": "0xc7ce690c", "rot": 12, "bit": 21, "mask": 16, "width": 1},
|
||||
{"i": 35, "op": "shfl", "dst": 3, "src": 7, "src2": 5, "imm": "0xa73f59de", "imm2": "0x84b9e329", "rot": 21, "bit": 27, "mask": 8, "width": 1},
|
||||
{"i": 36, "op": "add", "dst": 7, "src": 1, "src2": 7, "imm": "0xc53b542e", "imm2": "0xe10c2c95", "rot": 21, "bit": 2, "mask": 4, "width": 1},
|
||||
{"i": 37, "op": "xor", "dst": 5, "src": 7, "src2": 4, "imm": "0x81cd7b0e", "imm2": "0x21a51823", "rot": 12, "bit": 10, "mask": 2, "width": 1},
|
||||
{"i": 38, "op": "or", "dst": 2, "src": 1, "src2": 3, "imm": "0x7894e657", "imm2": "0xf8e4b972", "rot": 18, "bit": 9, "mask": 4, "width": 1},
|
||||
{"i": 39, "op": "mulhi", "dst": 1, "src": 0, "src2": 4, "imm": "0xbee8421f", "imm2": "0x070888a8", "rot": 20, "bit": 28, "mask": 4, "width": 1},
|
||||
{"i": 40, "op": "rotl", "dst": 6, "src": 1, "src2": 4, "imm": "0x4609857a", "imm2": "0xaeecb156", "rot": 19, "bit": 21, "mask": 1, "width": 1},
|
||||
{"i": 41, "op": "mulhi", "dst": 4, "src": 6, "src2": 0, "imm": "0x1a84e1e9", "imm2": "0x9b26bb72", "rot": 28, "bit": 19, "mask": 1, "width": 1},
|
||||
{"i": 42, "op": "sub", "dst": 6, "src": 0, "src2": 6, "imm": "0xbf62908e", "imm2": "0xdf03ea88", "rot": 11, "bit": 27, "mask": 16, "width": 1},
|
||||
{"i": 43, "op": "shfl", "dst": 6, "src": 3, "src2": 4, "imm": "0xa966241c", "imm2": "0x9c639efa", "rot": 12, "bit": 4, "mask": 4, "width": 1},
|
||||
{"i": 44, "op": "load", "dst": 4, "src": 2, "src2": 3, "imm": "0x7b5b5474", "imm2": "0x45cfc5dd", "rot": 17, "bit": 18, "mask": 8, "width": 1},
|
||||
{"i": 45, "op": "xor", "dst": 1, "src": 3, "src2": 6, "imm": "0x006193d0", "imm2": "0xfc2acc3f", "rot": 25, "bit": 1, "mask": 1, "width": 1},
|
||||
{"i": 46, "op": "load", "dst": 7, "src": 0, "src2": 6, "imm": "0x4777c4f8", "imm2": "0x3cf0a02f", "rot": 11, "bit": 14, "mask": 4, "width": 1},
|
||||
{"i": 47, "op": "load", "dst": 3, "src": 1, "src2": 3, "imm": "0x10692532", "imm2": "0x1a292ea5", "rot": 20, "bit": 23, "mask": 1, "width": 1},
|
||||
{"i": 48, "op": "mul", "dst": 5, "src": 3, "src2": 7, "imm": "0xadf5bd13", "imm2": "0xb999de2e", "rot": 23, "bit": 10, "mask": 4, "width": 1},
|
||||
{"i": 49, "op": "sub", "dst": 1, "src": 5, "src2": 4, "imm": "0x68ff101e", "imm2": "0xbdaaf46a", "rot": 25, "bit": 23, "mask": 16, "width": 1},
|
||||
{"i": 50, "op": "rotl", "dst": 2, "src": 6, "src2": 4, "imm": "0x92d9a412", "imm2": "0x0daf96ea", "rot": 8, "bit": 4, "mask": 4, "width": 1},
|
||||
{"i": 51, "op": "add", "dst": 1, "src": 5, "src2": 0, "imm": "0xa900fec4", "imm2": "0x77b9bd43", "rot": 6, "bit": 23, "mask": 2, "width": 1},
|
||||
{"i": 52, "op": "load", "dst": 4, "src": 7, "src2": 1, "imm": "0x51392a72", "imm2": "0x99e8bb36", "rot": 11, "bit": 9, "mask": 1, "width": 1},
|
||||
{"i": 53, "op": "sub", "dst": 2, "src": 7, "src2": 2, "imm": "0x0ffe2ac7", "imm2": "0x030743df", "rot": 9, "bit": 30, "mask": 16, "width": 1},
|
||||
{"i": 54, "op": "xor", "dst": 4, "src": 0, "src2": 4, "imm": "0x9123ff15", "imm2": "0x10c329a7", "rot": 28, "bit": 15, "mask": 2, "width": 1},
|
||||
{"i": 55, "op": "add", "dst": 1, "src": 6, "src2": 5, "imm": "0xe09f54e9", "imm2": "0x83e825bf", "rot": 23, "bit": 14, "mask": 8, "width": 1},
|
||||
{"i": 56, "op": "load", "dst": 2, "src": 4, "src2": 3, "imm": "0x239de52c", "imm2": "0xf80bae18", "rot": 15, "bit": 24, "mask": 1, "width": 1},
|
||||
{"i": 57, "op": "mad", "dst": 0, "src": 1, "src2": 4, "imm": "0x31dede8e", "imm2": "0xd3f619e6", "rot": 29, "bit": 7, "mask": 2, "width": 1},
|
||||
{"i": 58, "op": "load", "dst": 3, "src": 5, "src2": 3, "imm": "0x206437d6", "imm2": "0x28d1c290", "rot": 17, "bit": 28, "mask": 4, "width": 1},
|
||||
{"i": 59, "op": "or", "dst": 5, "src": 6, "src2": 3, "imm": "0x8fffd674", "imm2": "0x0507903a", "rot": 26, "bit": 27, "mask": 2, "width": 1},
|
||||
{"i": 60, "op": "mad", "dst": 6, "src": 5, "src2": 7, "imm": "0xf572bdb9", "imm2": "0xeda2af31", "rot": 21, "bit": 8, "mask": 2, "width": 1},
|
||||
{"i": 61, "op": "rotl", "dst": 4, "src": 2, "src2": 7, "imm": "0x84f12ddf", "imm2": "0x81ef22e1", "rot": 28, "bit": 30, "mask": 1, "width": 1},
|
||||
{"i": 62, "op": "mulhi", "dst": 5, "src": 0, "src2": 6, "imm": "0xf6bb45ee", "imm2": "0x6bfb632d", "rot": 22, "bit": 0, "mask": 4, "width": 1},
|
||||
{"i": 63, "op": "load", "dst": 3, "src": 6, "src2": 6, "imm": "0x50ec702a", "imm2": "0xae6ee96e", "rot": 20, "bit": 25, "mask": 2, "width": 1}
|
||||
{"i": 0, "op": "add", "dst": 4, "src": 5, "src2": 7, "imm": "0xea86e152", "imm2": "0x5810667a", "rot": 27, "bit": 13, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 1, "op": "xor", "dst": 7, "src": 0, "src2": 6, "imm": "0xbaab6229", "imm2": "0xed861989", "rot": 26, "bit": 22, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 2, "op": "shfl", "dst": 3, "src": 6, "src2": 2, "imm": "0x5b623116", "imm2": "0xff12e5b2", "rot": 12, "bit": 24, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 3, "op": "sub", "dst": 4, "src": 1, "src2": 1, "imm": "0xe99741c7", "imm2": "0xf5fa5009", "rot": 1, "bit": 21, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 4, "op": "mad", "dst": 2, "src": 0, "src2": 4, "imm": "0x673c2157", "imm2": "0xee02465f", "rot": 20, "bit": 22, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 5, "op": "rotl", "dst": 4, "src": 7, "src2": 7, "imm": "0x946f7818", "imm2": "0x45d3399e", "rot": 9, "bit": 2, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 6, "op": "rotr", "dst": 0, "src": 2, "src2": 4, "imm": "0x5f6a0ed2", "imm2": "0x7043a636", "rot": 19, "bit": 31, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 7, "op": "load", "dst": 6, "src": 7, "src2": 1, "imm": "0x5c61dcf7", "imm2": "0x7466aa40", "rot": 19, "bit": 9, "mask": 2, "width": 1, "win": 2, "off": 1},
|
||||
{"i": 8, "op": "load", "dst": 1, "src": 4, "src2": 5, "imm": "0x85668475", "imm2": "0xdb8cc483", "rot": 29, "bit": 7, "mask": 4, "width": 1, "win": 1, "off": 1},
|
||||
{"i": 9, "op": "load", "dst": 1, "src": 2, "src2": 4, "imm": "0x4454980f", "imm2": "0xebd31581", "rot": 10, "bit": 28, "mask": 16, "width": 1, "win": 1, "off": 1},
|
||||
{"i": 10, "op": "load", "dst": 7, "src": 0, "src2": 2, "imm": "0xe075297c", "imm2": "0x5779c44c", "rot": 10, "bit": 22, "mask": 2, "width": 1, "win": 1, "off": 1},
|
||||
{"i": 11, "op": "load", "dst": 7, "src": 1, "src2": 6, "imm": "0x65aa4311", "imm2": "0x4fe48ea9", "rot": 15, "bit": 9, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 12, "op": "mul", "dst": 5, "src": 4, "src2": 2, "imm": "0x1383d3ad", "imm2": "0xf3094b29", "rot": 8, "bit": 9, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 13, "op": "load", "dst": 7, "src": 6, "src2": 2, "imm": "0xed8a496f", "imm2": "0x3072c3c6", "rot": 28, "bit": 19, "mask": 8, "width": 1, "win": 1, "off": 1},
|
||||
{"i": 14, "op": "shfl", "dst": 6, "src": 5, "src2": 0, "imm": "0x8b965b57", "imm2": "0xcfeca6c1", "rot": 12, "bit": 27, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 15, "op": "shfl", "dst": 3, "src": 5, "src2": 3, "imm": "0x877c7586", "imm2": "0xa9cb2a03", "rot": 2, "bit": 29, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 16, "op": "or", "dst": 2, "src": 7, "src2": 3, "imm": "0xb740221a", "imm2": "0x89d38d6d", "rot": 6, "bit": 31, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 17, "op": "rotr", "dst": 0, "src": 6, "src2": 1, "imm": "0x26f3ad8a", "imm2": "0x27256f15", "rot": 18, "bit": 5, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 18, "op": "add", "dst": 7, "src": 5, "src2": 7, "imm": "0xf66e7017", "imm2": "0xb9e3577e", "rot": 9, "bit": 12, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 19, "op": "rotl", "dst": 7, "src": 0, "src2": 5, "imm": "0x849ae6ee", "imm2": "0x02b358f9", "rot": 24, "bit": 18, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 20, "op": "add", "dst": 6, "src": 7, "src2": 6, "imm": "0x52334d12", "imm2": "0x8c9f0ef8", "rot": 11, "bit": 23, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 21, "op": "or", "dst": 2, "src": 6, "src2": 5, "imm": "0xb1871e63", "imm2": "0xb2e40191", "rot": 5, "bit": 13, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 22, "op": "mad", "dst": 6, "src": 5, "src2": 4, "imm": "0x97df29e4", "imm2": "0xe60fea84", "rot": 11, "bit": 16, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 23, "op": "or", "dst": 1, "src": 0, "src2": 3, "imm": "0x8f30d21d", "imm2": "0x2df685a0", "rot": 31, "bit": 0, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 24, "op": "xor", "dst": 6, "src": 1, "src2": 2, "imm": "0xd2c4025f", "imm2": "0x5269eb4d", "rot": 31, "bit": 21, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 25, "op": "add", "dst": 2, "src": 6, "src2": 5, "imm": "0x659fc3d3", "imm2": "0x9cec0e12", "rot": 6, "bit": 17, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 26, "op": "rotr", "dst": 7, "src": 0, "src2": 1, "imm": "0xd89ef484", "imm2": "0x20be3846", "rot": 12, "bit": 9, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 27, "op": "mad", "dst": 4, "src": 5, "src2": 7, "imm": "0xb48420ae", "imm2": "0x3d1f2485", "rot": 14, "bit": 28, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 28, "op": "mad", "dst": 3, "src": 2, "src2": 4, "imm": "0xc5c46d76", "imm2": "0x700044b5", "rot": 22, "bit": 10, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 29, "op": "load", "dst": 1, "src": 4, "src2": 3, "imm": "0x0fbaf177", "imm2": "0xfff4f2ed", "rot": 20, "bit": 31, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 30, "op": "load", "dst": 2, "src": 3, "src2": 0, "imm": "0xe90eb2e5", "imm2": "0xb0f9eb79", "rot": 27, "bit": 14, "mask": 4, "width": 1, "win": 2, "off": 2},
|
||||
{"i": 31, "op": "load", "dst": 1, "src": 5, "src2": 2, "imm": "0x97ba3fc3", "imm2": "0x7894e657", "rot": 3, "bit": 30, "mask": 16, "width": 1, "win": 1, "off": 1},
|
||||
{"i": 32, "op": "add", "dst": 7, "src": 2, "src2": 7, "imm": "0x070888a8", "imm2": "0xe403240e", "rot": 2, "bit": 16, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 33, "op": "add", "dst": 2, "src": 0, "src2": 5, "imm": "0xf2e46d55", "imm2": "0x29701828", "rot": 31, "bit": 28, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 34, "op": "add", "dst": 2, "src": 3, "src2": 0, "imm": "0x58f75b87", "imm2": "0x343b7aee", "rot": 12, "bit": 14, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 35, "op": "mulhi", "dst": 2, "src": 5, "src2": 6, "imm": "0x5892a9e6", "imm2": "0xc9824c94", "rot": 19, "bit": 26, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 36, "op": "xor", "dst": 4, "src": 2, "src2": 3, "imm": "0x7b5b5474", "imm2": "0x45cfc5dd", "rot": 17, "bit": 18, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 37, "op": "mul", "dst": 6, "src": 5, "src2": 7, "imm": "0xccf564a5", "imm2": "0x873ad101", "rot": 7, "bit": 11, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 38, "op": "xor", "dst": 7, "src": 0, "src2": 6, "imm": "0xcac8f06d", "imm2": "0x6b97c683", "rot": 18, "bit": 28, "mask": 2, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 39, "op": "add", "dst": 7, "src": 2, "src2": 4, "imm": "0xb8180e9d", "imm2": "0x32bbd117", "rot": 23, "bit": 19, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 40, "op": "rotr", "dst": 2, "src": 3, "src2": 0, "imm": "0x2d6070bc", "imm2": "0x68ff101e", "rot": 13, "bit": 18, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 41, "op": "sub", "dst": 7, "src": 0, "src2": 2, "imm": "0x0daf96ea", "imm2": "0x36f37be1", "rot": 5, "bit": 0, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 42, "op": "add", "dst": 4, "src": 3, "src2": 6, "imm": "0x6ced15b7", "imm2": "0x6df7aed4", "rot": 19, "bit": 4, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 43, "op": "shfl", "dst": 7, "src": 3, "src2": 5, "imm": "0x8ace05f3", "imm2": "0xd378ec12", "rot": 23, "bit": 10, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 44, "op": "load", "dst": 0, "src": 7, "src2": 4, "imm": "0xb0607786", "imm2": "0xc4acabbc", "rot": 13, "bit": 7, "mask": 16, "width": 1, "win": 1, "off": 1},
|
||||
{"i": 45, "op": "add", "dst": 1, "src": 6, "src2": 5, "imm": "0xe09f54e9", "imm2": "0x83e825bf", "rot": 23, "bit": 14, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 46, "op": "load", "dst": 3, "src": 1, "src2": 0, "imm": "0x63cc1e4e", "imm2": "0xa1be8118", "rot": 12, "bit": 6, "mask": 2, "width": 1, "win": 2, "off": 0},
|
||||
{"i": 47, "op": "load", "dst": 6, "src": 3, "src2": 7, "imm": "0x353f1d79", "imm2": "0x3b2e7456", "rot": 18, "bit": 18, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 48, "op": "mulhi", "dst": 4, "src": 2, "src2": 7, "imm": "0x00d8a3cd", "imm2": "0x231866d2", "rot": 21, "bit": 20, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 49, "op": "add", "dst": 5, "src": 0, "src2": 2, "imm": "0xa8bae6df", "imm2": "0xf572bdb9", "rot": 14, "bit": 7, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 50, "op": "shfl", "dst": 0, "src": 7, "src2": 7, "imm": "0x81ef22e1", "imm2": "0x74438fc5", "rot": 28, "bit": 18, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 51, "op": "add", "dst": 6, "src": 0, "src2": 6, "imm": "0x383b9260", "imm2": "0x11e17c61", "rot": 12, "bit": 19, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 52, "op": "load", "dst": 5, "src": 2, "src2": 2, "imm": "0xfb84f451", "imm2": "0x11cd863e", "rot": 21, "bit": 20, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 53, "op": "rotl", "dst": 6, "src": 5, "src2": 4, "imm": "0xb1a7db6b", "imm2": "0x76686b9b", "rot": 12, "bit": 4, "mask": 16, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 54, "op": "rotl", "dst": 3, "src": 6, "src2": 3, "imm": "0x6f981f52", "imm2": "0xd99aeba2", "rot": 12, "bit": 27, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 55, "op": "add", "dst": 2, "src": 1, "src2": 2, "imm": "0xac6be8e3", "imm2": "0x18d67dbb", "rot": 26, "bit": 10, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 56, "op": "load", "dst": 1, "src": 4, "src2": 0, "imm": "0x7e7f6a00", "imm2": "0x6f0747da", "rot": 25, "bit": 28, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 57, "op": "add", "dst": 5, "src": 0, "src2": 4, "imm": "0xf03673fe", "imm2": "0xa75cd60d", "rot": 16, "bit": 12, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 58, "op": "load", "dst": 5, "src": 0, "src2": 2, "imm": "0x227f94a6", "imm2": "0x0e8344f9", "rot": 20, "bit": 10, "mask": 2, "width": 1, "win": 2, "off": 0},
|
||||
{"i": 59, "op": "add", "dst": 1, "src": 4, "src2": 3, "imm": "0xdecd4794", "imm2": "0x8dfb96bb", "rot": 21, "bit": 7, "mask": 4, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 60, "op": "mulhi", "dst": 3, "src": 2, "src2": 2, "imm": "0x0dd268e0", "imm2": "0x53034ca9", "rot": 1, "bit": 8, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 61, "op": "or", "dst": 6, "src": 4, "src2": 7, "imm": "0x3a45a321", "imm2": "0x9bc59a5f", "rot": 25, "bit": 11, "mask": 1, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 62, "op": "shfl", "dst": 5, "src": 4, "src2": 2, "imm": "0x8f229cc1", "imm2": "0xcaac64a2", "rot": 17, "bit": 13, "mask": 8, "width": 1, "win": 0, "off": 0},
|
||||
{"i": 63, "op": "load", "dst": 3, "src": 6, "src2": 6, "imm": "0xb2574178", "imm2": "0xcbcc798d", "rot": 28, "bit": 0, "mask": 16, "width": 1, "win": 1, "off": 1}
|
||||
]
|
||||
}
|
||||
|
|
|
|||
|
|
@ -39,69 +39,69 @@ kernel void igneum_hash(device const uint* dataset [[buffer(0)]],
|
|||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = r4 + r5 + select(0xea86e152u, 0x5810667au, ((sel >> 13u) & 1u) != 0u); // 0
|
||||
r2 = r2 ^ simd_shuffle_xor(r0, (ushort)4); // 1
|
||||
r3 = r3 + r2 + select(0x2cccb6cau, 0x642e66dbu, ((sel >> 10u) & 1u) != 0u); // 2
|
||||
r0 = rotl_imm(r0, 19u); // 3
|
||||
r7 = rotr_var(r7, r6); // 4
|
||||
r7 = r7 + r4 + select(0xee02465fu, 0xc1535555u, ((sel >> 21u) & 1u) != 0u); // 5
|
||||
r1 = mulhi(r1, r7); // 6
|
||||
r4 = r4 ^ dataset[r2 & MASK]; // 7
|
||||
r7 = r7 ^ dataset[r4 & MASK]; // 8
|
||||
r0 = r0 ^ dataset[r3 & MASK]; // 9
|
||||
r5 = r5 ^ dataset[r1 & MASK]; // 10
|
||||
r1 = r1 ^ dataset[r5 & MASK]; // 11
|
||||
r3 = mulhi(r3, r5); // 12
|
||||
r1 = r1 ^ dataset[r3 & MASK]; // 13
|
||||
r0 = r0 - r3; // 14
|
||||
r5 = r1 * r3 + r5; // 15
|
||||
r6 = mulhi(r6, r1); // 16
|
||||
r5 = r5 + r2 + select(0x697b3d00u, 0x8b965b57u, ((sel >> 28u) & 1u) != 0u); // 17
|
||||
r0 = mulhi(r0, r6); // 18
|
||||
r5 = rotr_var(r5, r3); // 19
|
||||
r5 = mulhi(r5, r2); // 20
|
||||
r1 = r1 + r0 + select(0xebcf247au, 0x6d7e8d05u, ((sel >> 1u) & 1u) != 0u); // 21
|
||||
r7 = r7 + r5 + select(0xf66e7017u, 0xb9e3577eu, ((sel >> 12u) & 1u) != 0u); // 22
|
||||
r1 = mulhi(r1, r5); // 23
|
||||
r2 = r2 - r5; // 24
|
||||
r7 = r7 + r4 + select(0x08ffa6c7u, 0x699ef1bbu, ((sel >> 2u) & 1u) != 0u); // 25
|
||||
r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 26
|
||||
r7 = r7 + r1 + select(0xe60fea84u, 0xb4ead2fbu, ((sel >> 14u) & 1u) != 0u); // 27
|
||||
r3 = r3 + r1 + select(0x65c76dabu, 0x8f30d21du, ((sel >> 6u) & 1u) != 0u); // 28
|
||||
r2 = r2 ^ dataset[r1 & MASK]; // 29
|
||||
r5 = r5 ^ dataset[r7 & MASK]; // 30
|
||||
r2 = r2 ^ dataset[r5 & MASK]; // 31
|
||||
r1 = r1 ^ simd_shuffle_xor(r7, (ushort)4); // 32
|
||||
r4 = r5 * r7 + r4; // 33
|
||||
r4 = r4 + r2 + select(0x0480debeu, 0xc7ce690cu, ((sel >> 21u) & 1u) != 0u); // 34
|
||||
r3 = r3 ^ simd_shuffle_xor(r7, (ushort)8); // 35
|
||||
r7 = r7 + r1 + select(0xc53b542eu, 0xe10c2c95u, ((sel >> 2u) & 1u) != 0u); // 36
|
||||
r5 = r5 ^ r7; // 37
|
||||
r2 = r2 | r1; // 38
|
||||
r1 = mulhi(r1, r0); // 39
|
||||
r6 = rotl_imm(r6, 19u); // 40
|
||||
r4 = mulhi(r4, r6); // 41
|
||||
r6 = r6 - r0; // 42
|
||||
r6 = r6 ^ simd_shuffle_xor(r3, (ushort)4); // 43
|
||||
r4 = r4 ^ dataset[r2 & MASK]; // 44
|
||||
r1 = r1 ^ r3; // 45
|
||||
r7 = r7 ^ dataset[r0 & MASK]; // 46
|
||||
r3 = r3 ^ dataset[r1 & MASK]; // 47
|
||||
r5 = r5 * r3; // 48
|
||||
r1 = r1 - r5; // 49
|
||||
r2 = rotl_imm(r2, 8u); // 50
|
||||
r1 = r1 + r5 + select(0xa900fec4u, 0x77b9bd43u, ((sel >> 23u) & 1u) != 0u); // 51
|
||||
r4 = r4 ^ dataset[r7 & MASK]; // 52
|
||||
r2 = r2 - r7; // 53
|
||||
r4 = r4 ^ r0; // 54
|
||||
r1 = r1 + r6 + select(0xe09f54e9u, 0x83e825bfu, ((sel >> 14u) & 1u) != 0u); // 55
|
||||
r2 = r2 ^ dataset[r4 & MASK]; // 56
|
||||
r0 = r1 * r4 + r0; // 57
|
||||
r3 = r3 ^ dataset[r5 & MASK]; // 58
|
||||
r5 = r5 | r6; // 59
|
||||
r6 = r5 * r7 + r6; // 60
|
||||
r4 = rotl_imm(r4, 28u); // 61
|
||||
r5 = mulhi(r5, r0); // 62
|
||||
r3 = r3 ^ dataset[r6 & MASK]; // 63
|
||||
r7 = r7 ^ r0; // 1
|
||||
r3 = r3 ^ simd_shuffle_xor(r6, (ushort)1); // 2
|
||||
r4 = r4 - r1; // 3
|
||||
r2 = r0 * r4 + r2; // 4
|
||||
r4 = rotl_imm(r4, 9u); // 5
|
||||
r0 = rotr_var(r0, r2); // 6
|
||||
r6 = r6 ^ dataset[((rotl_imm(r7 * 0x9ad30d99u, 29u) & 0x03ffffffu) | 0x04000000u) & MASK]; // 7
|
||||
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 8
|
||||
r1 = r1 ^ dataset[((rotl_imm(r2 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 9
|
||||
r7 = r7 ^ dataset[((rotl_imm(r0 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 10
|
||||
r7 = r7 ^ dataset[((rotl_imm(r1 * 0x9ad30d99u, 29u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 11
|
||||
r5 = r5 * r4; // 12
|
||||
r7 = r7 ^ dataset[((rotl_imm(r6 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 13
|
||||
r6 = r6 ^ simd_shuffle_xor(r5, (ushort)16); // 14
|
||||
r3 = r3 ^ simd_shuffle_xor(r5, (ushort)1); // 15
|
||||
r2 = r2 | r7; // 16
|
||||
r0 = rotr_var(r0, r6); // 17
|
||||
r7 = r7 + r5 + select(0xf66e7017u, 0xb9e3577eu, ((sel >> 12u) & 1u) != 0u); // 18
|
||||
r7 = rotl_imm(r7, 24u); // 19
|
||||
r6 = r6 + r7 + select(0x52334d12u, 0x8c9f0ef8u, ((sel >> 23u) & 1u) != 0u); // 20
|
||||
r2 = r2 | r6; // 21
|
||||
r6 = r5 * r4 + r6; // 22
|
||||
r1 = r1 | r0; // 23
|
||||
r6 = r6 ^ r1; // 24
|
||||
r2 = r2 + r6 + select(0x659fc3d3u, 0x9cec0e12u, ((sel >> 17u) & 1u) != 0u); // 25
|
||||
r7 = rotr_var(r7, r0); // 26
|
||||
r4 = r5 * r7 + r4; // 27
|
||||
r3 = r2 * r4 + r3; // 28
|
||||
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x9ad30d99u, 29u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 29
|
||||
r2 = r2 ^ dataset[((rotl_imm(r3 * 0x9ad30d99u, 29u) & 0x03ffffffu) | 0x08000000u) & MASK]; // 30
|
||||
r1 = r1 ^ dataset[((rotl_imm(r5 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 31
|
||||
r7 = r7 + r2 + select(0x070888a8u, 0xe403240eu, ((sel >> 16u) & 1u) != 0u); // 32
|
||||
r2 = r2 + r0 + select(0xf2e46d55u, 0x29701828u, ((sel >> 28u) & 1u) != 0u); // 33
|
||||
r2 = r2 + r3 + select(0x58f75b87u, 0x343b7aeeu, ((sel >> 14u) & 1u) != 0u); // 34
|
||||
r2 = mulhi(r2, r5); // 35
|
||||
r4 = r4 ^ r2; // 36
|
||||
r6 = r6 * r5; // 37
|
||||
r7 = r7 ^ r0; // 38
|
||||
r7 = r7 + r2 + select(0xb8180e9du, 0x32bbd117u, ((sel >> 19u) & 1u) != 0u); // 39
|
||||
r2 = rotr_var(r2, r3); // 40
|
||||
r7 = r7 - r0; // 41
|
||||
r4 = r4 + r3 + select(0x6ced15b7u, 0x6df7aed4u, ((sel >> 4u) & 1u) != 0u); // 42
|
||||
r7 = r7 ^ simd_shuffle_xor(r3, (ushort)4); // 43
|
||||
r0 = r0 ^ dataset[((rotl_imm(r7 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 44
|
||||
r1 = r1 + r6 + select(0xe09f54e9u, 0x83e825bfu, ((sel >> 14u) & 1u) != 0u); // 45
|
||||
r3 = r3 ^ dataset[((rotl_imm(r1 * 0x9ad30d99u, 29u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 46
|
||||
r6 = r6 ^ dataset[((rotl_imm(r3 * 0x9ad30d99u, 29u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 47
|
||||
r4 = mulhi(r4, r2); // 48
|
||||
r5 = r5 + r0 + select(0xa8bae6dfu, 0xf572bdb9u, ((sel >> 7u) & 1u) != 0u); // 49
|
||||
r0 = r0 ^ simd_shuffle_xor(r7, (ushort)4); // 50
|
||||
r6 = r6 + r0 + select(0x383b9260u, 0x11e17c61u, ((sel >> 19u) & 1u) != 0u); // 51
|
||||
r5 = r5 ^ dataset[((rotl_imm(r2 * 0x9ad30d99u, 29u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 52
|
||||
r6 = rotl_imm(r6, 12u); // 53
|
||||
r3 = rotl_imm(r3, 12u); // 54
|
||||
r2 = r2 + r1 + select(0xac6be8e3u, 0x18d67dbbu, ((sel >> 10u) & 1u) != 0u); // 55
|
||||
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x9ad30d99u, 29u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 56
|
||||
r5 = r5 + r0 + select(0xf03673feu, 0xa75cd60du, ((sel >> 12u) & 1u) != 0u); // 57
|
||||
r5 = r5 ^ dataset[((rotl_imm(r0 * 0x9ad30d99u, 29u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 58
|
||||
r1 = r1 + r4 + select(0xdecd4794u, 0x8dfb96bbu, ((sel >> 7u) & 1u) != 0u); // 59
|
||||
r3 = mulhi(r3, r2); // 60
|
||||
r6 = r6 | r4; // 61
|
||||
r5 = r5 ^ simd_shuffle_xor(r4, (ushort)8); // 62
|
||||
r3 = r3 ^ dataset[((rotl_imm(r6 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 63
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
|
|
|
|||
|
|
@ -41,69 +41,69 @@ kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]],
|
|||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = r4 + r5 + select(0xea86e152u, 0x5810667au, ((sel >> 13u) & 1u) != 0u); // 0
|
||||
r2 = r2 ^ simd_shuffle_xor(r0, (ushort)4); // 1
|
||||
r3 = r3 + r2 + select(0x2cccb6cau, 0x642e66dbu, ((sel >> 10u) & 1u) != 0u); // 2
|
||||
r0 = rotl_imm(r0, 19u); // 3
|
||||
r7 = rotr_var(r7, r6); // 4
|
||||
r7 = r7 + r4 + select(0xee02465fu, 0xc1535555u, ((sel >> 21u) & 1u) != 0u); // 5
|
||||
r1 = mulhi(r1, r7); // 6
|
||||
r4 = r4 ^ dataset[r2 & MASK]; // 7
|
||||
r7 = r7 ^ dataset[r4 & MASK]; // 8
|
||||
r0 = r0 ^ dataset[r3 & MASK]; // 9
|
||||
r5 = r5 ^ dataset[r1 & MASK]; // 10
|
||||
r1 = r1 ^ dataset[r5 & MASK]; // 11
|
||||
r3 = mulhi(r3, r5); // 12
|
||||
r1 = r1 ^ dataset[r3 & MASK]; // 13
|
||||
r0 = r0 - r3; // 14
|
||||
r5 = r1 * r3 + r5; // 15
|
||||
r6 = mulhi(r6, r1); // 16
|
||||
r5 = r5 + r2 + select(0x697b3d00u, 0x8b965b57u, ((sel >> 28u) & 1u) != 0u); // 17
|
||||
r0 = mulhi(r0, r6); // 18
|
||||
r5 = rotr_var(r5, r3); // 19
|
||||
r5 = mulhi(r5, r2); // 20
|
||||
r1 = r1 + r0 + select(0xebcf247au, 0x6d7e8d05u, ((sel >> 1u) & 1u) != 0u); // 21
|
||||
r7 = r7 + r5 + select(0xf66e7017u, 0xb9e3577eu, ((sel >> 12u) & 1u) != 0u); // 22
|
||||
r1 = mulhi(r1, r5); // 23
|
||||
r2 = r2 - r5; // 24
|
||||
r7 = r7 + r4 + select(0x08ffa6c7u, 0x699ef1bbu, ((sel >> 2u) & 1u) != 0u); // 25
|
||||
r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 26
|
||||
r7 = r7 + r1 + select(0xe60fea84u, 0xb4ead2fbu, ((sel >> 14u) & 1u) != 0u); // 27
|
||||
r3 = r3 + r1 + select(0x65c76dabu, 0x8f30d21du, ((sel >> 6u) & 1u) != 0u); // 28
|
||||
r2 = r2 ^ dataset[r1 & MASK]; // 29
|
||||
r5 = r5 ^ dataset[r7 & MASK]; // 30
|
||||
r2 = r2 ^ dataset[r5 & MASK]; // 31
|
||||
r1 = r1 ^ simd_shuffle_xor(r7, (ushort)4); // 32
|
||||
r4 = r5 * r7 + r4; // 33
|
||||
r4 = r4 + r2 + select(0x0480debeu, 0xc7ce690cu, ((sel >> 21u) & 1u) != 0u); // 34
|
||||
r3 = r3 ^ simd_shuffle_xor(r7, (ushort)8); // 35
|
||||
r7 = r7 + r1 + select(0xc53b542eu, 0xe10c2c95u, ((sel >> 2u) & 1u) != 0u); // 36
|
||||
r5 = r5 ^ r7; // 37
|
||||
r2 = r2 | r1; // 38
|
||||
r1 = mulhi(r1, r0); // 39
|
||||
r6 = rotl_imm(r6, 19u); // 40
|
||||
r4 = mulhi(r4, r6); // 41
|
||||
r6 = r6 - r0; // 42
|
||||
r6 = r6 ^ simd_shuffle_xor(r3, (ushort)4); // 43
|
||||
r4 = r4 ^ dataset[r2 & MASK]; // 44
|
||||
r1 = r1 ^ r3; // 45
|
||||
r7 = r7 ^ dataset[r0 & MASK]; // 46
|
||||
r3 = r3 ^ dataset[r1 & MASK]; // 47
|
||||
r5 = r5 * r3; // 48
|
||||
r1 = r1 - r5; // 49
|
||||
r2 = rotl_imm(r2, 8u); // 50
|
||||
r1 = r1 + r5 + select(0xa900fec4u, 0x77b9bd43u, ((sel >> 23u) & 1u) != 0u); // 51
|
||||
r4 = r4 ^ dataset[r7 & MASK]; // 52
|
||||
r2 = r2 - r7; // 53
|
||||
r4 = r4 ^ r0; // 54
|
||||
r1 = r1 + r6 + select(0xe09f54e9u, 0x83e825bfu, ((sel >> 14u) & 1u) != 0u); // 55
|
||||
r2 = r2 ^ dataset[r4 & MASK]; // 56
|
||||
r0 = r1 * r4 + r0; // 57
|
||||
r3 = r3 ^ dataset[r5 & MASK]; // 58
|
||||
r5 = r5 | r6; // 59
|
||||
r6 = r5 * r7 + r6; // 60
|
||||
r4 = rotl_imm(r4, 28u); // 61
|
||||
r5 = mulhi(r5, r0); // 62
|
||||
r3 = r3 ^ dataset[r6 & MASK]; // 63
|
||||
r7 = r7 ^ r0; // 1
|
||||
r3 = r3 ^ simd_shuffle_xor(r6, (ushort)1); // 2
|
||||
r4 = r4 - r1; // 3
|
||||
r2 = r0 * r4 + r2; // 4
|
||||
r4 = rotl_imm(r4, 9u); // 5
|
||||
r0 = rotr_var(r0, r2); // 6
|
||||
r6 = r6 ^ dataset[((rotl_imm(r7 * 0x9ad30d99u, 29u) & 0x03ffffffu) | 0x04000000u) & MASK]; // 7
|
||||
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 8
|
||||
r1 = r1 ^ dataset[((rotl_imm(r2 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 9
|
||||
r7 = r7 ^ dataset[((rotl_imm(r0 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 10
|
||||
r7 = r7 ^ dataset[((rotl_imm(r1 * 0x9ad30d99u, 29u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 11
|
||||
r5 = r5 * r4; // 12
|
||||
r7 = r7 ^ dataset[((rotl_imm(r6 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 13
|
||||
r6 = r6 ^ simd_shuffle_xor(r5, (ushort)16); // 14
|
||||
r3 = r3 ^ simd_shuffle_xor(r5, (ushort)1); // 15
|
||||
r2 = r2 | r7; // 16
|
||||
r0 = rotr_var(r0, r6); // 17
|
||||
r7 = r7 + r5 + select(0xf66e7017u, 0xb9e3577eu, ((sel >> 12u) & 1u) != 0u); // 18
|
||||
r7 = rotl_imm(r7, 24u); // 19
|
||||
r6 = r6 + r7 + select(0x52334d12u, 0x8c9f0ef8u, ((sel >> 23u) & 1u) != 0u); // 20
|
||||
r2 = r2 | r6; // 21
|
||||
r6 = r5 * r4 + r6; // 22
|
||||
r1 = r1 | r0; // 23
|
||||
r6 = r6 ^ r1; // 24
|
||||
r2 = r2 + r6 + select(0x659fc3d3u, 0x9cec0e12u, ((sel >> 17u) & 1u) != 0u); // 25
|
||||
r7 = rotr_var(r7, r0); // 26
|
||||
r4 = r5 * r7 + r4; // 27
|
||||
r3 = r2 * r4 + r3; // 28
|
||||
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x9ad30d99u, 29u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 29
|
||||
r2 = r2 ^ dataset[((rotl_imm(r3 * 0x9ad30d99u, 29u) & 0x03ffffffu) | 0x08000000u) & MASK]; // 30
|
||||
r1 = r1 ^ dataset[((rotl_imm(r5 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 31
|
||||
r7 = r7 + r2 + select(0x070888a8u, 0xe403240eu, ((sel >> 16u) & 1u) != 0u); // 32
|
||||
r2 = r2 + r0 + select(0xf2e46d55u, 0x29701828u, ((sel >> 28u) & 1u) != 0u); // 33
|
||||
r2 = r2 + r3 + select(0x58f75b87u, 0x343b7aeeu, ((sel >> 14u) & 1u) != 0u); // 34
|
||||
r2 = mulhi(r2, r5); // 35
|
||||
r4 = r4 ^ r2; // 36
|
||||
r6 = r6 * r5; // 37
|
||||
r7 = r7 ^ r0; // 38
|
||||
r7 = r7 + r2 + select(0xb8180e9du, 0x32bbd117u, ((sel >> 19u) & 1u) != 0u); // 39
|
||||
r2 = rotr_var(r2, r3); // 40
|
||||
r7 = r7 - r0; // 41
|
||||
r4 = r4 + r3 + select(0x6ced15b7u, 0x6df7aed4u, ((sel >> 4u) & 1u) != 0u); // 42
|
||||
r7 = r7 ^ simd_shuffle_xor(r3, (ushort)4); // 43
|
||||
r0 = r0 ^ dataset[((rotl_imm(r7 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 44
|
||||
r1 = r1 + r6 + select(0xe09f54e9u, 0x83e825bfu, ((sel >> 14u) & 1u) != 0u); // 45
|
||||
r3 = r3 ^ dataset[((rotl_imm(r1 * 0x9ad30d99u, 29u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 46
|
||||
r6 = r6 ^ dataset[((rotl_imm(r3 * 0x9ad30d99u, 29u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 47
|
||||
r4 = mulhi(r4, r2); // 48
|
||||
r5 = r5 + r0 + select(0xa8bae6dfu, 0xf572bdb9u, ((sel >> 7u) & 1u) != 0u); // 49
|
||||
r0 = r0 ^ simd_shuffle_xor(r7, (ushort)4); // 50
|
||||
r6 = r6 + r0 + select(0x383b9260u, 0x11e17c61u, ((sel >> 19u) & 1u) != 0u); // 51
|
||||
r5 = r5 ^ dataset[((rotl_imm(r2 * 0x9ad30d99u, 29u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 52
|
||||
r6 = rotl_imm(r6, 12u); // 53
|
||||
r3 = rotl_imm(r3, 12u); // 54
|
||||
r2 = r2 + r1 + select(0xac6be8e3u, 0x18d67dbbu, ((sel >> 10u) & 1u) != 0u); // 55
|
||||
r1 = r1 ^ dataset[((rotl_imm(r4 * 0x9ad30d99u, 29u) & 0x0fffffffu) | 0x00000000u) & MASK]; // 56
|
||||
r5 = r5 + r0 + select(0xf03673feu, 0xa75cd60du, ((sel >> 12u) & 1u) != 0u); // 57
|
||||
r5 = r5 ^ dataset[((rotl_imm(r0 * 0x9ad30d99u, 29u) & 0x03ffffffu) | 0x00000000u) & MASK]; // 58
|
||||
r1 = r1 + r4 + select(0xdecd4794u, 0x8dfb96bbu, ((sel >> 7u) & 1u) != 0u); // 59
|
||||
r3 = mulhi(r3, r2); // 60
|
||||
r6 = r6 | r4; // 61
|
||||
r5 = r5 ^ simd_shuffle_xor(r4, (ushort)8); // 62
|
||||
r3 = r3 ^ dataset[((rotl_imm(r6 * 0x9ad30d99u, 29u) & 0x07ffffffu) | 0x08000000u) & MASK]; // 63
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
|
|
|
|||
Some files were not shown because too many files have changed in this diff Show more
Loading…
Reference in a new issue