Lottery hash: generator version 2 (16 load slots, fresh sources, acceptance rule), every vector re-cut, packs regenerated, three workers re-checked, 20,000-program census
igneum-pow 0.2.0: generator v2 draws exactly 16 load slots from instructions 1..63, a load's source from the registers written earlier and not read by a load since, the other 48 ops from the ten non-load weights; accept.rs is spec 01 section 1.4.6 (static: no stale load source, every register injected; dynamic: 64 units on the seed-keyed closed-form dataset, no constant bit, no lane-constant site, under 164 saturated, bias within 136 of 1024, distinct addresses above 245,760); a rejected candidate is replaced by the next attempt of the seed (seed || k_le32), 32 a consensus fault. Packs carry the generator version, attempt and program id. Version 1 kept as generate_v1 for the census. Packs: igneum-genesis, igneum-hourly, igneum-genesis-mh regenerated by igneum-pow export; new igneum-devnet-v4-epoch0 (devnet genesis hash, day bytes 20730). Checks: Rust 39 of 39 tests; Metal natively via the Swift port (export cross-check 3 of 3 warps, identical programs and vectors on five seeds incl. three with attempt 1, fuzz 2,000 of 2,000); CUDA emu 4 of 4 packs; OpenCL emu 2 packs x 2 configurations; Apple OpenCL 4 of 4 packs at 27.9 Mhash/s. Census 20,000: 5.225 percent rejected, accepted distinct mean 127.887. Spec 01 0.2 (1.4.2, 1.4.3, 1.4.6, 1.11, 1.15, 1.16, 1.17), igneum-pow README, the CUDA, OpenCL and Metal test notes, bench-log entry, ledger M5 and M6 Fixed. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
cdc01c85ab
commit
b27da39f9b
62 changed files with 5716 additions and 1283 deletions
|
|
@ -523,3 +523,22 @@ Test network, 02:02:27 to 02:18:53 BST: 3 `igneumd` on `igneum-devnet-880` (gRPC
|
|||
| Harness scenario 2 (copy adapted to the merged rules; the repo copy probes 132 s and `pmt+1`) | live: floor-2 and floor-1 rejected, floor and floor+1 accepted where floor = max(pmt + 1, parent - 10 s); future flip between +10.00 and +10.02 s; sim: honest 0.908 b/s, ahead +0.2%, oscillate -0.7% over 1,200 virtual s: PASS |
|
||||
|
||||
Binaries: `target-integration/release/{igneumd 40,463,680 B, igneum-miner 7,916,096 B, igneum-exec-diff, igneum-inject, igneum-p2p-probe, igneum-harness-sim}`; `target-integration/x86_64-pc-windows-gnu/release/{igneumd.exe 50,169,344 B, igneum-miner.exe 10,045,440 B}`; package `/tmp/igneum-integration/igneum-node-windows-v4.zip` (32,946,104 B). Not done: the GPU `prepare` hot-swap path on the merged miner (CPU miners only here), the cut-over itself, v4 builds for the Mac seed relay and igneum-seed-1, the repo harness's scenario 2 rules and its scenario 5 summary text (stale "built a cache on HEAD" wording while the per-case data says 0 builds).
|
||||
|
||||
## 4 October 2026, generator version 2 adopted: exact load count, fresh-source loads, program acceptance; every vector re-cut, three workers re-checked, 20,000-program census, devnet-v4 binaries rebuilt (cryptographer)
|
||||
|
||||
Machine: Apple M5 Max, idle at the start (load 3), rustc 1.99.0 (rustup), Swift 5.8.1, builds at nice 10. Decision (Josh, this morning): adopt the census's generator rule before any public vector ships. Rule as implemented (`igneum-pow/src/generator.rs`, `src/accept.rs`, spec 01 sections 1.4.2, 1.4.3 and 1.4.6; mirrored in `proto-metal/main.swift` as `generateProgramV2` and `acceptProgram` because the Metal worker derives its program from the seed itself): G1 exactly 16 load slots, a uniform subset of instructions 1..63 drawn first by partial Fisher-Yates, the other 48 ops from the ten non-load weights (sum 75); G2 a load's source is drawn from the registers other than `dst` written by an earlier instruction and not read by a load since; R (a) no cyclically stale load source, (b) every register has an injecting write, (c) 64 units at base nonces from `SplitMix64(FNV-1a-64("igneum-accept/" || seed words LE))` on the seed-keyed closed-form dataset at 2^28 words with init words = seed words: no constant register bit, no lane-constant load site in any unit, fewer than 164 saturated final values, every output bit within 136 of 1,024, distinct addresses above 245,760 over the 2,048 hashes; a rejected candidate is replaced by `seed_words_from_bytes(seed || k_le32)`, `k = 1, 2, ...`, 32 consecutive rejections a consensus fault. Every pack carries `generator 2`, the attempt and the program id `FNV-1a-64("igneum-program/" || 2_le32 || seed words LE || attempt_le32)`; `igneum-pow` is 0.2.0 and the node's engine reports `igneum-lottery-v2-bound`. The crate is now the pack source (`igneum-pow export`); the Swift exporter is the Metal cross-check. Version 1 stays as `generate_v1` for the census and MEMHARD.md levers; its vectors are retired.
|
||||
|
||||
Packs regenerated (`proto-cuda/packs/`): `igneum-genesis` and `igneum-hourly` (closed form), `igneum-genesis-mh` (memory-hard, day 2026-10-03, cache FNV unchanged `48c4f5bf24166b2e`), and new `igneum-devnet-v4-epoch0` (epoch seed = devnet genesis hash `edc4fa84...fb07`, day bytes `igneum-day/20730`, cache FNV `448274a57f508cbc`). `igneum-genesis` attempt 0, program id `bcc1248b10cc90f2`, op mix `load=16 add=8 shfl=8 xor=6 mad=5 mul=5 mulhi=5 sub=4 rotl=3 rotr=3 or=1`, lane 0 at base 0 `42246ba99fc58e4f`, lane 31 `b08446b1f2de7793`; devnet pack id `4be132dd1f2ff270`, lane 0 `285a83011e7ac3fc`. Bound vectors re-cut (`igneum-pow/README.md`: H zero, nonce 0 gives `746c567b090acf6a`). `cargo test --release`: 39 of 39 (28 unit, 11 pack).
|
||||
|
||||
| Check | Result |
|
||||
|---|---|
|
||||
| Rust CPU reference (`igneum-pow`) | the packs by construction; verify 0.631 ms per unit (avg of 20), cold 0.67 to 0.81 ms, 4,096 items per unit, cache fill 179 ms; acceptance 1.3 to 3.4 ms per candidate |
|
||||
| Apple Metal, natively (`proto-metal/igneum-bench`, Swift v2 generator) | `--export-pack igneum-genesis` memory-hard: GPU cache == CPU cache, Metal cross-check PASS 3 of 3 warps, instruction list and 96 vectors identical to the Rust pack; closed-form exports of `igneum-genesis`, `igneum-hourly`, `igneum-census-2026-10-03/22`, `/37`, `/51` (the last three have attempt 0 rejected: (b) r7, (c) 119.74 distinct, (b) r4; attempt 1 accepted, ids `22ed0609d079f4cf`, `947705cc4eb1df0a`, `9869afcc028bf9f1`): instructions, seed words and 96 vectors identical to Rust on all five; fuzz `--fuzz 2000 --fuzz-seed igneum-fuzz-gen2-2026-10-04`: 2,000 of 2,000 PASS, 8,000 warps, loads per hash 128 to 128, compile avg 21.8 ms, wall 91.4 s (`proto-metal/TESTS.md` section 9) |
|
||||
| CUDA through the clang emulation shim (`proto-cuda/emu/emu.sh`, `--batch-log2 13 --block-warps 2`) | all four packs OVERALL PASS: dataset self-test, 3 warps standalone, 2 warps per block in batch; memory-hard packs cache check 67,108,864 of 67,108,864 words, host fill 174 to 178 ms |
|
||||
| OpenCL through the clang emulation (`proto-opencl/emu/emu.sh`), sub-group 32 local exchange and wave64 sub-group shuffle | `igneum-genesis-mh` and `igneum-devnet-v4-epoch0`: 96 of 96 in both configurations, fingerprints `f2a95d5bb84d961e` and `8e22ad069cb2a8c3` at 2^13, identical across configurations |
|
||||
| Apple OpenCL on the M5 Max (`proto-opencl/host.c`) | all four packs: cache check PASS, 96 of 96 standalone and in batch (also `--group-warps 2`), fingerprint `f2a95d5bb84d961e` at 2^13 = the emulator's; 2^24 fingerprints `25f96e7dce90bd4e` (genesis-mh), `3cc4fbf90fa6366c` (devnet); rate 27.5 to 27.9 Mhash/s, 14.1 to 14.3 GB/s useful on every pack (version 1 genesis: 45.0 at 80 distinct loads; the census projected 28 at 128) |
|
||||
| igneum-census, 20,000 programs, `--gen v2 --warps 64`, memory-hard day 2026-10-03, 8 threads, 254.8 s | rejected 5.225 percent (static 4.130, dynamic 1.095), 1.0551 candidates per epoch; accepted programs: distinct addresses per hash mean 127.887, min 120.127, p1 126.897, p50 127.999, max 128.000; static loads 128 on every program (`census-v2-20k.tsv` and its summary in the session scratchpad, not checked in). Against the 100,000-seed figure of 3 October: 5.14 percent |
|
||||
| devnet-v4 node and miner (`vendor/igneum-node-v4`, path dependency bumped to igneum-pow 0.2.0, engine name v2) | `cargo build --release -p kaspad -p igneum-miner --features igneum-pow` into `vendor/igneum-node/target-integration` (1 min 17 s warm): `release/igneumd` 40,480,112 B, `release/igneum-miner` 7,932,880 B (08:41 BST); `cargo test --release -p kaspa-pow --features igneum-pow` 11 of 11; Windows cross-build (`proto-cuda/windows-node/cross-build.sh vendor/igneum-node-v4 6`, 4 min 53 s): `target-integration/x86_64-pc-windows-gnu/release/igneumd.exe` 50,179,072 B, `igneum-miner.exe` 10,065,920 B (libstdc++-6.dll import as before) |
|
||||
| 2-node test network on the real engine (ports 29000 to 29012, `/tmp/igneum-gen2`, `igneum-devnet-900`, `IGNEUM_DEVNET_GENESIS_BITS=0x1f010000`, `IGNEUM_POW_EPOCH_BLOCKS=100`, `IGNEUM_POW_EPOCH_LEAD=20`, one 3-thread CPU miner per node for 300 s) | 338 blocks accepted on both nodes, 0 rejected, 0 invalid, sink identical at 10 of 10 samples; four epochs crossed (DAA 0, 100, 200, 300; epoch seeds `234e08...`, `d3f427...`, `de316c...`, `971384...`, all attempt 0, ids `8f8806638d59850f`, `c015349db63beb2c`, `d7d52120407a0b69`, `512527bb7a528476`), program and cache ready in 192 to 284 ms on the miners, 4 cache builds per node; m1 158 and m2 180 blocks at 0.046 MH/s each; one WARN per node (eth JSON-RPC port 26790 held by the live devnet node, harmless) |
|
||||
|
||||
Not done: no NVIDIA or AMD hardware has run a version 2 pack (the RTX 5090's 192 of 192 and the gfx1036 run of 3 October were version 1; the kernel text is unchanged); the edge, stats, determinism and memcheck sections of TESTS.md were not re-run (they do not depend on the generator); the live devnet (v3, version 1 programs) was not touched, so the cut-over is where version 2 goes live; `proto-metal/main.swift` carries the version 2 port uncommitted next to the hot-swap working-tree changes (not in this agent's file list), and the Mac app's Metal worker must be rebuilt from it before the cut-over or Mac GPU shares will fail the CPU re-check; the v4 binaries above were built from the worktree as found, which also holds another agent's uncommitted finality floor change (2/3 of total, O-3.15); the GPU `prepare` hot-swap path was not exercised here (CPU miners only). The ten non-load weights and the 6-sigma bias threshold remain prototype values (spec 1.16).
|
||||
|
|
|
|||
|
|
@ -67,7 +67,7 @@ Evidence: none in the repository yet; the ProgPoW specification and the Ravencoi
|
|||
### M5. Your load count varies 6x between programs
|
||||
"TESTS.md: loads per hash ranged 40 to 232 across 10,000 programs. A 40-load program is ALU-bound and favours a chip for that hour. Your litepaper says the memory footprint and instruction count are fixed."
|
||||
|
||||
Status: Open, experiment scheduled.
|
||||
Status: Fixed (4 October 2026): generator version 2 draws exactly 16 load slots per program (spec 01 section 1.4.2, `igneum-pow/src/generator.rs`), and the fresh-source rule of 1.4.3 with the acceptance rule of 1.4.6 fixes the distinct count too, which the census showed is what the GPU pays for: every accepted program does 128 loads per hash of which at least 120 and typically 128 are distinct (20,000-program confirmation: mean 127.887, min 120.127). Apple OpenCL on the M5 Max runs every version 2 pack within 1 percent of the same rate (27.5 to 27.9 Mhash/s). Every vector was re-cut and all three workers re-checked (`docs/bench-log.md`, 4 October 2026 "generator version 2"). Still owed: the first RTX 5090 run on a version 2 pack. Was: Open, experiment scheduled.
|
||||
|
||||
Answer: Correct and a real gap. The instruction count is fixed (64 x 8); the load count is not, and the hash rate scales with it (104 loads gave 228 Mhash/s, 128 loads gave 185 on the 5090). The generator must fix the load count per program, or bound it tightly, so every hour is equally memory-bound and difficulty does not whiplash on the hour. This goes into the specification in phase 1 and is re-fuzzed. Until then the litepaper's sentence about fixed footprint is ahead of the prototype.
|
||||
|
||||
|
|
@ -76,7 +76,7 @@ Evidence: `proto-metal/TESTS.md` section 1 (loads per hash 40 to 232), `docs/ben
|
|||
### M6. Weak programs
|
||||
"Some hours the generator will emit a program whose OR chain saturates a register or whose load addresses collapse. That hour is both biased and shortcut-able. You have measured 3 seeds for bias out of an infinite population."
|
||||
|
||||
Status: Open, experiment scheduled.
|
||||
Status: Fixed (4 October 2026): the acceptance rule of spec 01 section 1.4.6 (`igneum-pow/src/accept.rs`, mirrored in `proto-metal/main.swift`) rejects a candidate with a stale load source, a register without an injecting write, a nonce-independent register bit, a lane-constant load site, more than 1 percent saturated final values, an output bit past 6 sigma, or fewer than 120 distinct addresses per hash on average, over 64 fixed units on the seed-keyed closed-form dataset; a rejected candidate is replaced by the next attempt of the seed, so every node agrees. Measured: 5.225 percent of 20,000 candidates rejected (4.130 static, 1.095 dynamic), 1.055 candidates per epoch; the rule costs 1.3 to 3.4 ms. The remaining question, whether 6 sigma at 2,048 nonces is the right bias threshold, is a prototype value of spec 1.16. Was: Open, experiment scheduled.
|
||||
|
||||
Answer: Correct. Three seeds were measured for bias (max deviation 2.90 sigma over 192 bit positions, avalanche mean 32.0, std 4.0, zero duplicates) and the population was not. Nothing yet rejects a weak program. The scheduled experiment is a weak-program census of at least 10^5 programs on the CPU interpreter measuring bias, distinct load addresses, OR saturation and nonce-independent registers, then a rejection rule written into the generator. Phase 1, before the spec is final.
|
||||
|
||||
|
|
@ -794,7 +794,8 @@ Evidence: `site/journey.json`.
|
|||
|---|---|---|
|
||||
| Answered with evidence | 2 | M8, M10 |
|
||||
| Answered by design | 13 | M3, M12, F4, F9, F11, F12, F13, P9, E2, E3, C1, C12, L6 |
|
||||
| Open, experiment or decision scheduled | 11 | M1, M5, M6, M11, F7, P3, C4, L1, L2, L4, L5 |
|
||||
| Open, experiment or decision scheduled | 9 | M1, M11, F7, P3, C4, L1, L2, L4, L5 |
|
||||
| Fixed (4 October 2026, generator version 2) | 2 | M5, M6 |
|
||||
| Closed by rule, Decided or Closed by removal (3 October 2026) | 9 | E4, F2, F3, P5, P8, C9, E7, G7, X6 |
|
||||
| Conceded, stated in the litepaper or design doc | 13 | F6, F8, P2, P6, E1, E6, E8, G3, G5, G8, C7, X3, X4 |
|
||||
| Conceded, not yet stated (fix in the overclaims list) | 32 | M2, M4, M7, M9, M13, F1, F5, F10, P1, P4, P7, P10, E5, G1, G2, G4, G6, C2, C3, C5, C6, C8, C10, C11, L3, X1, X2, X5, X7, X8, X9, X10 |
|
||||
|
|
|
|||
|
|
@ -1,8 +1,8 @@
|
|||
# Igneum protocol specification, section 1: the lottery hash
|
||||
|
||||
Spec version 0.1, 3 October 2026. Status of this section: Measured for the construction as implemented (Implemented values with test vectors on three GPU vendors and a CPU reference); Designed for the header binding, the day key, the dataset growth and the era schedule; Open where marked.
|
||||
Spec version 0.2, 4 October 2026 (0.1 on 3 October 2026). Status of this section: Measured for the construction as implemented (Implemented values with test vectors on three GPU vendors and a CPU reference); Designed for the header binding, the day key, the dataset growth and the era schedule; Open where marked. Version 0.2 adopts generator version 2 (sections 1.4.2, 1.4.3 and 1.4.6: exact load count, fresh-source loads, program acceptance) from the weak-program census of `docs/analysis/weak-program-census-2026-10-03.md`; every vector of version 0.1 is retired and re-cut (section 1.17).
|
||||
|
||||
Normative implementation: `igneum-pow/src/{seed,generator,memhard,verify}.rs`. Where this text and that code disagree, the code and its test vectors win until this text is corrected (section 0.4). The Swift prototype `proto-metal/main.swift` is bit-exact with the crate on every pack (`docs/bench-log.md`, entry "igneum-pow: Rust crate bit-exact with proto-metal").
|
||||
Normative implementation: `igneum-pow/src/{seed,generator,accept,memhard,verify}.rs`. Where this text and that code disagree, the code and its test vectors win until this text is corrected (section 0.4). The Swift prototype `proto-metal/main.swift` carries the same generator and acceptance rule and is bit-exact with the crate on every pack (`docs/bench-log.md`, entries "igneum-pow: Rust crate bit-exact with proto-metal" and "generator version 2").
|
||||
|
||||
Every parameter marked "prototype value, to be fixed at gate 1" is carried by the implementation today, is part of the test vectors, and is confirmed or replaced by the named measurement before the hash is frozen (section 1.16).
|
||||
|
||||
|
|
@ -11,7 +11,7 @@ Every parameter marked "prototype value, to be fixed at gate 1" is carried by th
|
|||
The lottery hash decides who produces the next block. It MUST be:
|
||||
|
||||
1. Deterministic and bit-exact on every conforming implementation: GPU kernels on any vendor, the CPU verifier, and any future implementation (section 1.14).
|
||||
2. Cheap to verify on one CPU core without the dataset: one 32-lane unit of work in under 10 ms (Target, Measured at 0.41 to 0.58 ms, section 1.11).
|
||||
2. Cheap to verify on one CPU core without the dataset: one 32-lane unit of work in under 10 ms (Target, Measured at 0.63 ms steady and 0.67 to 0.81 ms cold under generator version 2, section 1.11).
|
||||
3. Bound by random access to a dataset larger than any on-chip cache, so that computing dataset words is slower than loading them (Measured 4.8x slower on Apple, section 1.8.4; not measured on NVIDIA or AMD).
|
||||
4. Unknowable until shortly before it is needed, so that a miner cannot grind the seed (section 4).
|
||||
|
||||
|
|
@ -67,7 +67,7 @@ Implemented (`igneum-pow/src/generator.rs`). A program is a list of `INSTR_COUNT
|
|||
| Parameter | Value | Label |
|
||||
|---|---|---|
|
||||
| Registers per lane | 8 x u32 | prototype value, to be fixed at gate 1 (fixed by the ASIC-gain target and the register-pressure measurement of 1.16) |
|
||||
| Instructions per program | 64 | prototype value, to be fixed at gate 1 (fixed by the CPU-verify measurement on a 2019-class core and the load-count rule) |
|
||||
| Instructions per program | 64, of which exactly 16 are `load` (section 1.4.2) | prototype value, to be fixed at gate 1 (fixed by the CPU-verify measurement on a 2019-class core); the load count is Definition since 4 October 2026 |
|
||||
| Iterations per hash | 8 | prototype value, to be fixed at gate 1 (same measurement) |
|
||||
| Lanes per unit of work | 32 | Definition. Not a tuning parameter (section 1.9) |
|
||||
| Shuffle masks | {1, 2, 4, 8, 16} | Definition, follows from 32 lanes |
|
||||
|
|
@ -95,13 +95,12 @@ Eleven families. `dst`, `src`, `src2` name registers; `src != dst` always; `src2
|
|||
|
||||
A twelfth family `wload` (warp-coalesced 128-byte load) exists in the code as lever (b) and is never emitted at the default configuration (`wide_frac = 0`). It is NOT part of the lottery hash. It is retained only so the measurement in `proto-metal/MEMHARD.md` section 2.4 stays reproducible, and the recommendation there is not to adopt it.
|
||||
|
||||
### 1.4.2 Op weights
|
||||
### 1.4.2 Op weights and the load count
|
||||
|
||||
Implemented, prototype value, to be fixed at gate 1. Weights sum to 100 and are applied in this order:
|
||||
Implemented (generator version 2, 4 October 2026, `igneum-pow/src/generator.rs`; ledger M5 Fixed). Every program contains exactly 16 `load` instructions (`LOAD_SLOTS`, Definition), so every hash performs 128 loads and a 32-lane unit derives at most 4,096 dataset items, the bound of section 1.11. The other 48 instructions are drawn from the ten non-load families with these weights (sum 75, prototype value, to be fixed at gate 1), applied in this order:
|
||||
|
||||
| Op | Weight |
|
||||
|---|---|
|
||||
| load | 25 |
|
||||
| add | 12 |
|
||||
| xor | 10 |
|
||||
| mul | 8 |
|
||||
|
|
@ -113,16 +112,27 @@ Implemented, prototype value, to be fixed at gate 1. Weights sum to 100 and are
|
|||
| rotr | 6 |
|
||||
| or | 4 |
|
||||
|
||||
What fixes them: (1) the load count per program is not fixed today, only its expectation (16 loads per 64 instructions, 128 per hash); over 10,000 programs it ranged 40 to 232 loads per hash (`proto-metal/TESTS.md` section 1) and the GPU rate scales with it (104 loads: 228 Mhash/s, 128 loads: 185 Mhash/s on the RTX 5090, `docs/bench-log.md` RTX 5090 sweep entry). Ledger M5. The gate 1 rule MUST fix the load count per program exactly (candidate: draw exactly 16 load positions, then draw the other 48 ops from the remaining weights), so every epoch is equally memory-bound and difficulty does not step on the hour. (2) The lever (a) measurement (`load_weight = 17`, `MEMHARD.md` section 2.4) is the fallback if a slower verifier core ever threatens the 10 ms gate: it cut CPU time to 0.46 to 0.51 ms per warp on the Swift verifier and left the kernel memory bound.
|
||||
The load weight 25 of version 1 is retired; it survives only as the ratio 16 of 64. Why the count is fixed and not merely expected: under version 1 the static count ran 24 to 256 loads per hash over 100,000 programs and the GPU rate tracked the number of distinct addresses, 24 to 200, because 19.9 percent of all loads re-read an address the same hash had already read (census sections 3 and 5). Fixing the static count alone would not fix the memory work; the fresh-source rule of 1.4.3 fixes the distinct count by construction, and 1.4.6 rejects the few programs where it cannot.
|
||||
|
||||
The version 1 lever measurement (`load_weight = 17`, `proto-metal/MEMHARD.md` section 2.4) is kept in the code as `generate_v1` for reproduction only; its programs are not the lottery hash.
|
||||
|
||||
### 1.4.3 Draw order
|
||||
|
||||
For each of the 64 instructions, in order, draw from the program stream of 1.3.3 exactly these values in exactly this order, whether or not the op uses them:
|
||||
From the program stream of 1.3.3, in this order, whether or not an op uses a value.
|
||||
|
||||
(1) Load slots. Let `p[0..62] = 1..63` (instruction 0 is never a load: nothing is fresh before it). For `i` in 0..15 draw `j = i + below(63 - i)` and swap `p[i]` and `p[j]`. The load slots are `p[0..15]`, a uniform 16-subset of 1..63.
|
||||
|
||||
(2) For each instruction `k` in 0..63, nine draws:
|
||||
|
||||
```
|
||||
roll = below(100); op = first entry of the weight table whose cumulative weight exceeds roll
|
||||
roll = below(75); op = first entry of the table of 1.4.2 whose cumulative weight exceeds roll
|
||||
(on a load slot the roll is drawn and ignored and op = load)
|
||||
dst = below(8)
|
||||
a = below(7); if a >= dst then a = a + 1 (src, never equal to dst)
|
||||
src: on an ALU slot a = below(7); src = a + (a >= dst)
|
||||
on a load slot E = the registers other than dst, in register order, that an earlier instruction of this
|
||||
program has written and that no later load has used as its source (E is empty before
|
||||
instruction 0); if E is not empty, a = below(|E|) and src = E[a];
|
||||
if E is empty, a = below(7) and src = a + (a >= dst), and 1.4.6 (a) rejects the program
|
||||
b = below(8) (src2)
|
||||
imm = low32(next())
|
||||
imm2 = low32(next())
|
||||
|
|
@ -131,30 +141,58 @@ bit = below(32)
|
|||
mask = 1 << below(5)
|
||||
```
|
||||
|
||||
Nine draws per instruction, 576 per program. A program is fully determined by its eight seed words.
|
||||
16 slot draws plus 64 x 9: 592 draws per program. A program is fully determined by its eight seed words. A load's source holds a value written in the same iteration that no earlier load has read, so no load repeats the address of an earlier load of the same hash, across the iteration boundary included (the first form of the rule, with every register eligible at instruction 0, left the wrap open and failed 1.4.6 (a) on 36 percent of programs; census section 7.1).
|
||||
|
||||
Test vector: for seed `igneum-genesis` the first eight instructions are (`proto-cuda/packs/igneum-genesis-mh/program.json`):
|
||||
Test vector: for seed `igneum-genesis` (attempt 0, program id `bcc1248b10cc90f2`, section 1.4.6) the first eight instructions are (`proto-cuda/packs/igneum-genesis-mh/program.json`):
|
||||
|
||||
```
|
||||
0: rotl dst=4 src=2 src2=7 imm=0x20699878 imm2=0x6f1a6170 rot=25 bit=7 mask=16
|
||||
1: sub dst=0 src=5 src2=4 imm=0xf81a0b9d imm2=0xf0505e88 rot=1 bit=4 mask=1
|
||||
2: load dst=4 src=3 src2=6 imm=0xb1978a0b imm2=0x2ca4e162 rot=10 bit=21 mask=1
|
||||
3: rotl dst=1 src=6 src2=7 imm=0xc3bd2355 imm2=0xa8c5f27e rot=1 bit=22 mask=1
|
||||
4: add dst=2 src=3 src2=0 imm=0x61f0b51c imm2=0x2735a174 rot=4 bit=26 mask=2
|
||||
5: load dst=5 src=3 src2=3 imm=0x4d183796 imm2=0x679648a8 rot=4 bit=30 mask=4
|
||||
6: sub dst=5 src=7 src2=7 imm=0x265677dc imm2=0x9043323e rot=30 bit=7 mask=4
|
||||
7: add dst=3 src=4 src2=6 imm=0x52f2dbf4 imm2=0x5a069596 rot=5 bit=26 mask=2
|
||||
0: mad dst=2 src=3 src2=4 imm=0xbf7b174d imm2=0x337b762e rot=17 bit=2 mask=2
|
||||
1: mad dst=2 src=1 src2=1 imm=0xdd04a5da imm2=0x42da7657 rot=15 bit=30 mask=16
|
||||
2: mad dst=2 src=3 src2=2 imm=0x734003fa imm2=0x5bb67700 rot=3 bit=20 mask=1
|
||||
3: xor dst=3 src=5 src2=5 imm=0xc55a1b1c imm2=0xa19720f3 rot=7 bit=8 mask=1
|
||||
4: load dst=7 src=2 src2=5 imm=0xad572dd7 imm2=0x9ceb3ea7 rot=18 bit=30 mask=2
|
||||
5: load dst=5 src=7 src2=3 imm=0x769a53be imm2=0x80f9067e rot=12 bit=22 mask=1
|
||||
6: shfl dst=1 src=4 src2=0 imm=0xd3613d88 imm2=0x262fb219 rot=10 bit=30 mask=8
|
||||
7: shfl dst=7 src=3 src2=4 imm=0xce38e42f imm2=0xb868b818 rot=11 bit=8 mask=8
|
||||
```
|
||||
|
||||
The whole program has op mix `load=13 xor=13 sub=7 shfl=6 add=5 mulhi=5 mad=4 rotr=4 mul=3 rotl=3 or=1`, 104 loads per hash.
|
||||
The whole program has op mix `load=16 add=8 shfl=8 xor=6 mad=5 mul=5 mulhi=5 sub=4 rotl=3 rotr=3 or=1`, 128 loads per hash. Instruction 4 reads r2, written by instructions 0 to 2; instruction 5 reads r7, written by instruction 4.
|
||||
|
||||
### 1.4.4 Generator contract
|
||||
|
||||
Every emitted instruction satisfies: `rot` in 1..31, `mask` in {1, 2, 4, 8, 16}, `src != dst`. This held on every instruction of 10,200 fuzzed programs (`TESTS.md` section 1). A kernel emitter MAY rely on it; an interpreter MUST NOT accept a program that violates it.
|
||||
Every emitted instruction satisfies: `rot` in 1..31, `mask` in {1, 2, 4, 8, 16}, `src != dst`; every program has exactly 16 `load` instructions and none at instruction 0. This held on every instruction of 10,200 fuzzed version 1 programs (`TESTS.md` section 1) and of 2,000 fuzzed version 2 programs (`TESTS.md` section 9, 4 October 2026). A kernel emitter MAY rely on it; an interpreter MUST NOT accept a program that violates it.
|
||||
|
||||
### 1.4.5 Encoding
|
||||
|
||||
A program is transmitted as the seed words, never as instructions. A node hands a miner the pack it emits itself (`igneum-pow/src/emit.rs`: `kernel.cu`, `kernel.cl`, `program.metal`, `program.h`, `program.json`, `memhard.h`, `vectors.*`), and a miner MAY regenerate everything from the seed. `program.json` is the interchange form; its field names are those of `Instr` in `generator.rs`.
|
||||
A program is transmitted as the seed bytes, never as instructions. A node hands a miner the pack it emits itself (`igneum-pow/src/emit.rs`: `kernel.cu`, `kernel.cl`, `program.metal`, `program.h`, `program.json`, `memhard.h`, `vectors.*`, the three `*_bound` kernels), and a miner MAY regenerate everything from the seed bytes by the procedure of 1.4.6. `program.json` (format `igneum-program-pack-3`) is the interchange form; its field names are those of `Instr` in `generator.rs`, and it carries `generator` (2), `attempt`, `program_id` and `seed_bytes`. `program.h` carries the same as `IGNEUM_GENERATOR`, `IGNEUM_PROGRAM_ATTEMPT`, `IGNEUM_PROGRAM_ID` and `IGNEUM_SEED_BYTES_HEX`. An implementation MUST refuse a pack whose generator version is not its own.
|
||||
|
||||
### 1.4.6 Program acceptance
|
||||
|
||||
Implemented (`igneum-pow/src/accept.rs`, `proto-metal/main.swift`; ledger M6 Fixed). A candidate program is accepted only if all of the following hold, and every conforming implementation MUST evaluate them identically.
|
||||
|
||||
(a) For every `load`, some instruction between the previous `load` from the same source register and this one, in cyclic order over the 64 instructions, writes that register.
|
||||
|
||||
(b) Every register `r0..r7` is the destination of at least one `add`, `sub`, `xor`, `mad`, `shfl` or `load`.
|
||||
|
||||
(c) The program is interpreted (section 1.7) for 64 units at base nonces `low32(next()) AND NOT 31` from a SplitMix64 stream seeded with `FNV-1a-64("igneum-accept/" || seed words as little-endian bytes)`, with init words `I` equal to the seed words and dataset words `dataset_elem(idx, S[0], S[1])` of `verify.rs` (the six-operation closed form of the version 0.1 packs) at 2^28 words (`idx = src AND 0x0fffffff`, a constant of this rule whatever the live dataset size) in place of the memory-hard dataset. Over those 2,048 evaluations: no register has a bit equal in every final value; no `load` site (iteration, instruction) reads one address in all 32 lanes of any unit; the number of final register values equal to 0 or 2^32 - 1 is below 164 (1 percent of 16,384); every output bit's ones count is within 136 of 1,024 (6 sigma); and the number of distinct masked dataset addresses read by one lane in one evaluation, summed over the 2,048 evaluations, exceeds 245,760 (a mean above 120 of the 128 loads).
|
||||
|
||||
Attempts. Attempt 0 of a program seed `b` (the 32-byte epoch seed, or the UTF-8 of a seed string) is the candidate drawn from `seed_words_from_bytes(b)`. If it fails, attempt `k = 1, 2, ...` is drawn from `seed_words_from_bytes(b || k_le32)`; the first accepted candidate is the program of the epoch. Measured rejection rate under this generator: 5.14 percent over 100,000 seeds (census section 7) and the 20,000-seed confirmation of `docs/bench-log.md` (4 October 2026), so the probability that 32 consecutive candidates fail is below 2^-136, and an implementation MAY treat 32 consecutive failures as a consensus fault (`MAX_ATTEMPTS`).
|
||||
|
||||
Program id. `FNV-1a-64("igneum-program/" || generator_le32 || seed words as little-endian bytes || attempt_le32)` with `generator = 2`, written into every pack. Two implementations that agree on the id agree on the generator version, the seed words and the attempt.
|
||||
|
||||
Why the closed form: the test is then a pure function of the program (no cache, no day), costs 1.3 to 3.4 ms on one core, and the census checked on 100,000 programs that its verdict agrees with the memory-hard dataset's on all but 39 threshold-edge cases (section 7.3). What the three parts catch: (a) the empty-list fallback of 1.4.3; (b) registers that saturate to all ones (2.4 percent of candidates); (c) zero-absorbing register sets, lane-constant load sites, output bias and value-level address repeats (2.1 percent). Not in the rule, and why: a contraction as the last write (80 percent of programs) and the `or` count are too common and (c) already catches the cases that matter; the load critical path is a hash-rate question, not a weakness.
|
||||
|
||||
Test vectors for the rule (`igneum-pow accept --seed ...`):
|
||||
|
||||
| Seed | Attempt 0 | Attempt 1 |
|
||||
|---|---|---|
|
||||
| `igneum-genesis` | accepted, program id `bcc1248b10cc90f2`, 128.000 distinct addresses per hash, 0 saturated, bias max 54 | |
|
||||
| `igneum-hourly` | accepted, `a4c4d00961c855df` | |
|
||||
| `igneum-census-2026-10-03/22` | rejected, (b) r7 has no injecting write | accepted, `22ed0609d079f4cf` |
|
||||
| `igneum-census-2026-10-03/37` | rejected, (c) 245,230 distinct addresses (mean 119.74) | accepted, `947705cc4eb1df0a` |
|
||||
| `igneum-census-2026-10-03/51` | rejected, (b) r4 has no injecting write | accepted, `9869afcc028bf9f1` |
|
||||
|
||||
The Rust crate and the Swift prototype derive identical instruction lists and identical 96-vector sets on all five seeds (`docs/bench-log.md`, 4 October 2026).
|
||||
|
||||
## 1.5 Fixed memory footprint
|
||||
|
||||
|
|
@ -339,17 +377,18 @@ Implemented (`verify.rs`, `memhard.rs`). A verifier holds the program for the ep
|
|||
2. At each `load`, gather the 32 masked indices, deduplicate by item (`idx >> 4`), derive the distinct items with all chains interleaved round by round (all mixers for round `r`, then all cache-line XORs for round `r`), and hand each lane its word. Interleaving lets the eight dependent misses of each item overlap across up to 32 items; without it the verifier pays about 8 x 100 ns of DRAM latency per item in series.
|
||||
3. Fold and compare lane `n AND 31` against `target64`.
|
||||
|
||||
The verifier does at most 104 x 32 = 3,328 item derivations for a 104-load program (fewer with duplicates); the design bound is 4,096 items per unit (design document, Lottery seeds item 3), which a 128-load program meets exactly and a 144-load program (4,608 items) exceeds. The load-count rule of 1.4.2 is what will enforce the bound.
|
||||
The verifier does at most 128 x 32 = 4,096 item derivations per unit, the design bound (design document, Lottery seeds item 3), because every program has exactly 16 loads (1.4.2) and an accepted program reads at least 120 and typically 128 distinct addresses per hash (1.4.6). Under version 1 the count ran from 3,328 to 4,608 items.
|
||||
|
||||
| Verifier | ms per 32-lane unit, steady (avg of 20) | Worst cold single unit | Source |
|
||||
|---|---|---|---|
|
||||
| Rust, one M5 Max performance core, 104 loads | 0.441 | 0.41 to 0.87 across five seeds | Measured, `docs/bench-log.md`, igneum-pow entry |
|
||||
| Rust, 144 loads, 4,608 items | 0.579 | | same |
|
||||
| Rust, one M5 Max performance core, generator v2, 128 loads, 4,096 items | 0.631 | 0.67 to 0.81 across the three vector units | Measured 4 October 2026, `igneum-pow/README.md` |
|
||||
| Rust, version 1, 104 loads, 3,328 items (retired) | 0.441 | 0.41 to 0.87 across five seeds | Measured, `docs/bench-log.md`, igneum-pow entry |
|
||||
| Rust, version 1, 144 loads, 4,608 items (retired) | 0.579 | | same |
|
||||
| Swift, 104 loads | 0.649 | 1.16 to 2.11 | Measured, memory-hard dataset entry |
|
||||
| Swift, 144 loads | 1.205 | | same |
|
||||
| Closed-form dataset (not memory-hard, for scale) | 0.002 (Rust), 0.017 (Swift) | | same entries |
|
||||
|
||||
The 10 ms gate (Target) is met with a margin of about 17x steady and 11x worst-cold on this core. Not measured: a 2019-class laptop core (design document, "Three experiments before gate 3"), which is what fixes the gate.
|
||||
The 10 ms gate (Target) is met with a margin of about 16x steady and 12x worst-cold on this core. Not measured: a 2019-class laptop core (design document, "Three experiments before gate 3"), which is what fixes the gate.
|
||||
|
||||
## 1.12 Schedules: epoch, day, era
|
||||
|
||||
|
|
@ -361,7 +400,9 @@ All times are DAA seconds since genesis (section 0.6). At 1 block per second one
|
|||
| Day | 86,400 DAA s | The day key, hence the cache and the dataset | Designed |
|
||||
| Era | 15,552,000 DAA s (180 days) | Era parameters and one instruction-family unlock, section 1.13 | Designed; the length is a prototype value (the design says "every 6 months") |
|
||||
|
||||
Epoch `e` covers DAA scores `[3,600 e, 3,600 (e + 1))`. The epoch of a block is the epoch of its own DAA score, so "which program was this block mined under" is a function of the header alone once the seed is known. The program for epoch `e` is `generate_from_words(S_e)` with `S_e = seed_words_from_bytes(program_seed_e)` and `program_seed_e` the 32-byte VDF output of section 4.3.
|
||||
Epoch `e` covers DAA scores `[3,600 e, 3,600 (e + 1))`. The epoch of a block is the epoch of its own DAA score, so "which program was this block mined under" is a function of the header alone once the seed is known. The program for epoch `e` is `generate_from_seed_bytes(program_seed_e)`: attempt 0 is drawn from `S_e = seed_words_from_bytes(program_seed_e)`, and a rejected attempt is replaced as 1.4.6 says; `program_seed_e` is the 32-byte VDF output of section 4.3.
|
||||
|
||||
Implementation note (devnet, 3 October 2026, `docs/fork-divergence.md` "Epoch seed"): until the VDF of section 4 is in the node, `program_seed_e` is the hash of the last selected-chain block whose DAA score is below `3,600 e - 600`. The 600-DAA-score lead stands in for section 4.3's 20-minute lead: the program of epoch `e` is knowable about 10 minutes before it starts, every block template reports it (`pow_epoch.next_epoch_seed`), and a GPU worker compiles it in the background and swaps at the boundary with no pause (serve protocol `prepare`, `proto-metal/main.swift`, `proto-cuda/host.cu`, `proto-opencl/host.c`). Measured across boundaries on a short-epoch test network in `docs/bench-log.md` (hot-swap entry). The program schedule is a protocol constant; a miner that cannot compile ahead sees the same seed at the same time as everyone else, only later.
|
||||
|
||||
Day `d` covers DAA scores `[86,400 d, 86,400 (d + 1))`. The design document names a day seed and does not say how it is derived. Proposed (Designed, Open, O-1.10): `day_bytes = "igneum-day/" || d_le64 || program_seed of the first epoch of day d`, so the day key is as unpredictable as the epoch seed and known 20 minutes before the day starts (section 4.5), which is enough for a 0.2 s CPU cache fill or a 2 ms GPU one plus a 13 to 30 ms GPU dataset build (Measured, section 1.8.3 and `docs/bench-log.md` RTX 5090 memory-hard entry: 13.4 ms for 1 GiB).
|
||||
|
||||
|
|
@ -378,7 +419,7 @@ The era seed `E_n` is the 32-byte output of the 1-hour VDF of section 4.4. One S
|
|||
| Parameter | Base (era 0) | Draw | Bound |
|
||||
|---|---|---|---|
|
||||
| Op weights for the ten non-load ops | table 1.4.2 | each perturbed by `below(2 * B + 1) - B` points, then renormalised by largest remainder to 100 minus the load weight | B = 2 points, proposed |
|
||||
| Load weight | 25 | not drawn | fixed, so every era is equally memory-bound |
|
||||
| Load count | 16 of 64 | not drawn | fixed, so every era is equally memory-bound |
|
||||
| Output fold rotations | (7, 14, 21), (9, 18, 27) | each `1 + below(31)` | 1..31 |
|
||||
| Mixer round count | 8 | not drawn | fixed, so the verify budget holds |
|
||||
|
||||
|
|
@ -406,7 +447,7 @@ evaluated in integers (bytes), with one year = 31,536,000 DAA seconds. The datas
|
|||
A conforming implementation MUST:
|
||||
|
||||
1. Use only integer arithmetic. No floating point anywhere, including in index computation and in the mixer (floating point rounds differently per vendor and would split the chain; design document, hostile review table row 2).
|
||||
2. Mask or range-reduce every dataset index exactly as 1.5 and 1.13.3 state, and never read outside the dataset. Every load in emitted source MUST have the single form `dataset[rN AND MASK]` (or the adopted range reduction), checkable by text search (`TESTS.md` section 5: 13 of 13 loads at three sizes, 416 of 416 indices exceeded MASK before masking).
|
||||
2. Mask or range-reduce every dataset index exactly as 1.5 and 1.13.3 state, and never read outside the dataset. Every load in emitted source MUST have the single form `dataset[rN AND MASK]` (or the adopted range reduction), checkable by text search (`igneum-pow/tests/packs.rs`: 16 of 16 loads masked in every emitted kernel of every pack; `TESTS.md` section 5 for the version 1 run).
|
||||
3. Implement `rotr` by `src AND 31` and `rotl` by an immediate in 1..31; a rotate by 0 or 32 through the immediate path is undefined and MUST NOT occur.
|
||||
4. Compute `mulhi` as the exact high 32 bits of the 64-bit product (`__umulhi`, `mulhi`, `mul_hi`).
|
||||
5. Wrap on overflow everywhere (add, sub, mul, mad, the SplitMix and FNV state).
|
||||
|
|
@ -415,27 +456,28 @@ A conforming implementation MUST:
|
|||
|
||||
## 1.15 Conformance procedure for a miner implementation
|
||||
|
||||
A miner, kernel emitter or verifier conforms when all of the following pass. Each is a command that exists today; the AMD discrete-card rows are pending.
|
||||
A miner, kernel emitter or verifier conforms when all of the following pass. Each is a command that exists today; the AMD discrete-card rows are pending. Every vector below is of generator version 2 (4 October 2026); the version 0.1 vectors are retired and MUST NOT be used.
|
||||
|
||||
1. Cache: fill the 256 MiB cache for day `2026-10-03` and reproduce FNV-1a 64 `48c4f5bf24166b2e`, cache line 0 and cache line 4,194,303 of section 1.8.5, word for word.
|
||||
1. Cache: fill the 256 MiB cache for day `2026-10-03` and reproduce FNV-1a 64 `48c4f5bf24166b2e`, cache line 0 and cache line 4,194,303 of section 1.8.5, word for word (unchanged by version 2). For the devnet day bytes `"igneum-day/" || 20730_le64`: `448274a57f508cbc`.
|
||||
2. Dataset self-test at 1 GiB: words 0..15, word `0x0fffffff`, and the 64 sampled words of `vectors.json` (four are quoted in 1.8.5).
|
||||
3. Vectors: all 96 outputs of the three units at base nonces 0, 4,096 and 1,000,000 for pack `igneum-genesis-mh`, standalone (one unit per launch) and in batch (many units per launch, at least two units per work-group or block). Section 1.17 lists them.
|
||||
4. Batch fingerprint: FNV-1a 64 over the 2^13 outputs at base nonce 0 = `f99fb375b3abeaf5`, and over the 2^24 outputs = `98af644e993239e2`.
|
||||
5. Fuzz: at least 200 random programs (`--fuzz 200` or the Rust equivalent when it exists), 4 units each at base nonces drawn from the full 32-bit range including wraps past 2^32, at 64 MiB, 256 MiB and 1 GiB, zero mismatches against the CPU interpreter, zero compile failures, generator contract asserted on every instruction. Reference: 200 of 200 and 10,000 of 10,000 (`TESTS.md` section 1; re-run on the memory-hard dataset, `MEMHARD.md` section 2.5).
|
||||
6. Edge: the 14 hand-built programs of `TESTS.md` section 2 (rotates by 0 and 31 through the register path, `mulhi` at the extremes, every shuffle mask, loads at index 0 and at MASK through in-range and out-of-range registers, wraparound on add, sub, mul, mad, zero loads, 64 loads), 128 of 128 lanes each.
|
||||
7. Static mask check on every emitted kernel (1.14 item 2).
|
||||
8. Exchange rule: on any device whose sub-group size is not exactly 32, or cannot be queried per kernel, the local-memory path is used and the run says so.
|
||||
3. Generator and acceptance: derive the five seeds of the 1.4.6 table to the same attempt and program id, and the `igneum-genesis` instruction list of 1.4.3.
|
||||
4. Vectors: all 96 outputs of the three units at base nonces 0, 4,096 and 1,000,000 for pack `igneum-genesis-mh`, standalone (one unit per launch) and in batch (many units per launch, at least two units per work-group or block). Section 1.17 lists them. The same for `igneum-devnet-v4-epoch0`, the chain's own derivation (epoch seed = the devnet genesis hash, day bytes of 2026-10-04).
|
||||
5. Batch fingerprint: FNV-1a 64 over the 2^13 outputs at base nonce 0 = `f2a95d5bb84d961e` (igneum-genesis-mh; `8e22ad069cb2a8c3` for igneum-devnet-v4-epoch0), and over the 2^24 outputs = `25f96e7dce90bd4e` (igneum-genesis-mh; `3cc4fbf90fa6366c` devnet).
|
||||
6. Fuzz: at least 200 random programs of the current generator (`--fuzz 200` or the Rust equivalent when it exists), 4 units each at base nonces drawn from the full 32-bit range including wraps past 2^32, at 64 MiB, 256 MiB and 1 GiB, zero mismatches against the CPU interpreter, zero compile failures, generator contract asserted on every instruction. Reference: 2,000 of 2,000 under version 2 (`TESTS.md` section 9), 10,200 of 10,200 under version 1.
|
||||
7. Edge: the 14 hand-built programs of `TESTS.md` section 2 (rotates by 0 and 31 through the register path, `mulhi` at the extremes, every shuffle mask, loads at index 0 and at MASK through in-range and out-of-range registers, wraparound on add, sub, mul, mad, zero loads, 64 loads), 128 of 128 lanes each. These are hand-built and bypass the generator; they test the interpreter and the kernels, not the rule.
|
||||
8. Static mask check on every emitted kernel (1.14 item 2).
|
||||
9. Exchange rule: on any device whose sub-group size is not exactly 32, or cannot be queried per kernel, the local-memory path is used and the run says so.
|
||||
|
||||
Measured conformance to date (`docs/bench-log.md`): Apple Metal, Apple OpenCL, pocl (CPU), the clang emulators, NVIDIA RTX 5090 via CUDA and via NVIDIA OpenCL, AMD gfx1036 via AMD OpenCL, and the Rust and Swift CPU references all pass items 1 to 4; items 5 and 6 have run on Metal and the CPU references only. A discrete AMD card has not run anything (ledger M8).
|
||||
Measured conformance to date (`docs/bench-log.md`), version 2 vectors: Apple Metal natively (3 of 3 units on the exported `igneum-genesis-mh`, fuzz 2,000 of 2,000, and the Swift generator identical to the Rust one on the five seeds of 1.4.6), Apple OpenCL (96 of 96 on all four packs, fingerprints as in item 5), the clang CUDA emulation (96 of 96 on all four packs, 2 warps per block), the clang OpenCL emulation (96 of 96 on the two memory-hard packs in sub-group 32 and wave64 configurations), and the Rust CPU reference. Not yet run on version 2 vectors: the RTX 5090 (CUDA and NVIDIA OpenCL), AMD gfx1036, pocl; the version 1 runs on those devices (3 October 2026) stand as evidence that the kernel text, which version 2 did not change, agrees across vendors. A discrete AMD card has not run anything (ledger M8).
|
||||
|
||||
## 1.16 Parameters marked "prototype value, to be fixed at gate 1" and what fixes them
|
||||
|
||||
| Parameter | Prototype value | Measurement or decision that fixes it |
|
||||
|---|---|---|
|
||||
| Seed derivation (FNV-1a plus SplitMix64) | section 1.3 | Replace with a standard hash (ledger M7); re-run the weak-program census on the result |
|
||||
| Seed derivation (FNV-1a plus SplitMix64) | section 1.3 | Replace with a standard hash (ledger M7); re-run the weak-program census and re-cut every vector on the result (the attempt derivation and the program id of 1.4.6 go through the same function) |
|
||||
| Instructions per program, iterations | 64 x 8 | CPU verify on a 2019-class laptop core under 10 ms with the memory-hard dataset; register-pressure and occupancy on the three vendors |
|
||||
| Registers per lane | 8 | same |
|
||||
| Op weights, including the 25% load weight | table 1.4.2 | Weak-program census of at least 10^5 programs (bias, avalanche, distinct load addresses, `or` saturation, nonce-independent registers), then the rejection rule; and the exact-load-count rule (ledger M5) |
|
||||
| Op weights, including the 25% load weight | table 1.4.2 | Fixed 4 October 2026 by the weak-program census (`docs/analysis/weak-program-census-2026-10-03.md`): exact load count 16 (1.4.2), fresh-source rule (1.4.3), acceptance rule (1.4.6). The ten non-load weights remain a prototype value for the era draw bounds of 1.13.1 |
|
||||
| Output fold rotations | (7, 14, 21), (9, 18, 27) | Stats run of `TESTS.md` section 3 on the chosen fold |
|
||||
| Cache size, lines per segment, ChaCha rounds | 256 MiB, 64, 12 | Shortcut ratio and time-memory curve on an RTX 5090 and an AMD discrete card; external review of the chained-block cache |
|
||||
| Item rounds, mixer shape | 8, section 1.8.4 | Same measurement; external review of the mixer; weak-key check on `ROT` |
|
||||
|
|
@ -446,21 +488,25 @@ Measured conformance to date (`docs/bench-log.md`): Apple Metal, Apple OpenCL, p
|
|||
|
||||
## 1.17 Test vectors
|
||||
|
||||
Pack `proto-cuda/packs/igneum-genesis-mh/` (seed `igneum-genesis`, day `2026-10-03`, memory-hard, 2^28 words, MASK `0x0fffffff`, 32 lanes). Produced by the Swift CPU interpreter, cross-checked by Metal before the pack was written, then reproduced by every implementation listed in 1.15. The two closed-form packs (`igneum-genesis`, `igneum-hourly`) are regression vectors for the interpreter only and are not the lottery hash.
|
||||
Pack `proto-cuda/packs/igneum-genesis-mh/` (seed `igneum-genesis`, generator version 2, attempt 0, program id `bcc1248b10cc90f2`, day `2026-10-03`, memory-hard, 2^28 words, MASK `0x0fffffff`, 32 lanes). Produced by `igneum-pow export` (the Rust CPU interpreter) on 4 October 2026, reproduced by the Swift CPU interpreter and cross-checked by Metal, Apple OpenCL and the two clang emulations the same day (section 1.15). The version 0.1 vectors of 3 October 2026 (lane 0 `1fb0b3bbc1ac8279`) are retired: they belong to a generator that no longer exists in the protocol. The two closed-form packs (`igneum-genesis`, `igneum-hourly`) are regression vectors for the interpreter only and are not the lottery hash.
|
||||
|
||||
Unit at base nonce 0, lanes 0..31:
|
||||
|
||||
```
|
||||
1fb0b3bbc1ac8279 61533759ca995ac8 97cd15ec31242001 3c19e71e2dbff828
|
||||
c7aba60b6c2ab016 8f13725d2a59bf84 7b826b82fdec5a2f 3131d7817418c04e
|
||||
d7c7e57c6aaac948 b3b639431ebdd3d6 400af448658a1d56 17f5760b8cadb8fd
|
||||
01ded5c4894c411c d6b3fdbc57bdb128 888efb103cecb983 ac73538353a340c4
|
||||
3f9914ad9445f052 8a860b28afd3dd42 c2a122ef9c1fa330 96c4e58663f82dce
|
||||
baae4b9d6a3c4320 cd6c2dd653e07743 7d3d00f0fb46325b 7636224050132b09
|
||||
dc5d76c741cc60f4 aef3031b0d4b701c 735987104a975f5f a58fce4010bf99dd
|
||||
b16c863d2ffdc0dd 0337b56e39529c90 802ac0bc1d0c0696 fa052263a854f3de
|
||||
42246ba99fc58e4f 19561eec0db4f9f4 11a4fb70ea7b688f 8872bfadec960949
|
||||
085e6cf2ff6b8303 780e0d76e504fe6c 7bb05a713f8283fa fef7beeffd5ab1d1
|
||||
1b83afb12e78d65b 2576c5384a2e1cae 549d620a0735d6d8 2eb4e32f32e8d5f1
|
||||
81bb414492f69586 092d17324b465e01 8bde350ee2354b5a d64b47e9c9e0ec07
|
||||
dbbf1b78b9979a3f bee382abde89c111 6b598d99c16a8c70 69d72da59fb3b6e9
|
||||
1e309d2a549632fa 98fa8255ab65f005 2f48ab1bb516110c 7d2af17cadd18bea
|
||||
547c5978bd005e02 25ea2e21e88ea9d3 1c575ec6e43efc58 077b80cb079958b4
|
||||
a96eebd0c8634981 44b18bdf19eb7838 c3ccb9fe9ef5953c b08446b1f2de7793
|
||||
```
|
||||
|
||||
Unit at base nonce 4,096: lane 0 `f219cf7ecf6ec450`, lane 1 `b575ab4b388c01f7`, lane 31 `d5b198435d9c40da`. Unit at base nonce 1,000,000: lane 0 `8a6f7a32b06edb52`, lane 1 `362d2b789a0c0ebf`, lane 31 `7040299873672a35`. The remaining 58 values are in `vectors.json` and `vectors.h` of the pack.
|
||||
Unit at base nonce 4,096: lane 0 `3d3903e310ca038f`, lane 1 `9e72b9a86ebb29e4`, lane 31 `61c242509efdccdd`. Unit at base nonce 1,000,000: lane 0 `f218c1bd58e6dfe0`, lane 1 `8c5a362ee98971c2`, lane 31 `6c3b2c11adfbfcac`. The remaining 58 values are in `vectors.json` and `vectors.h` of the pack.
|
||||
|
||||
Cache, dataset and mixer vectors: section 1.8.4 and 1.8.5. Seed words: section 1.3.1. Generator: section 1.4.3. Batch fingerprints: section 1.15 item 4.
|
||||
Pack `proto-cuda/packs/igneum-devnet-v4-epoch0/`: epoch seed bytes `edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07` (the devnet genesis hash), day bytes `69676e65756d2d6461792ffa50000000000000` (`"igneum-day/" || 20730_le64`, 2026-10-04), attempt 0, program id `4be132dd1f2ff270`, cache FNV-1a 64 `448274a57f508cbc`, dataset words 0 and 1 `3dd50b1f 48edec90`, word `[MASK]` `f7b7180e`; unit at base nonce 0: lane 0 `285a83011e7ac3fc`, lane 31 `6f1136558c20e3d8`.
|
||||
|
||||
Header-bound vectors (section 1.6 rule, seed `igneum-genesis`, day `2026-10-03`): `igneum-pow/README.md`, eight values, for example H = 32 zero bytes and nonce 0 give `746c567b090acf6a`.
|
||||
|
||||
Cache, dataset and mixer vectors: section 1.8.4 and 1.8.5. Seed words: section 1.3.1. Generator: section 1.4.3. Acceptance: section 1.4.6. Batch fingerprints: section 1.15 item 5.
|
||||
|
|
|
|||
2
igneum-census/Cargo.lock
generated
2
igneum-census/Cargo.lock
generated
|
|
@ -11,4 +11,4 @@ dependencies = [
|
|||
|
||||
[[package]]
|
||||
name = "igneum-pow"
|
||||
version = "0.1.0"
|
||||
version = "0.2.0"
|
||||
|
|
|
|||
|
|
@ -6,8 +6,10 @@
|
|||
//! warp), and one line of metrics per program is written to a TSV. `summarise` turns the TSV into the tables
|
||||
//! of `docs/analysis/weak-program-census-2026-10-03.md`.
|
||||
//!
|
||||
//! The crate reads igneum-pow as a path dependency and changes nothing in it. The candidate generator rules
|
||||
//! (`--gen fixed16`, `--gen fixed16-fresh`) live here until the spec adopts one.
|
||||
//! The crate reads igneum-pow as a path dependency and changes nothing in it. Since 4 October 2026 the adopted
|
||||
//! generator (version 2) and its acceptance rule live in igneum-pow (`--gen v2` draws its candidates and the
|
||||
//! `accept` columns record the rule's verdict); `--gen default` is the retired version 1 generator and the two
|
||||
//! intermediate forms of the census (`fixed16`, `fixed16-fresh`) are kept here so the analysis document reproduces.
|
||||
|
||||
use std::collections::HashMap;
|
||||
use std::env;
|
||||
|
|
@ -17,7 +19,8 @@ use std::sync::atomic::{AtomicUsize, Ordering};
|
|||
use std::sync::Mutex;
|
||||
use std::time::Instant;
|
||||
|
||||
use igneum_pow::generator::{generate_from_words, GeneratorConfig, Instr, Op, Program, INSTR_COUNT, ITERATIONS, LANES, OP_WEIGHTS};
|
||||
use igneum_pow::accept::{check as accept_check, Reject};
|
||||
use igneum_pow::generator::{candidate_from_words, generate_v1_from_words, GeneratorConfig, Instr, Op, Program, GENERATOR_VERSION, INSTR_COUNT, ITERATIONS, LANES, OP_WEIGHTS};
|
||||
use igneum_pow::seed::{fnv1a64, program_rng, seed_words, SplitMix64};
|
||||
use igneum_pow::verify::{dataset_elem, hash_warp, splitmix32, DatasetMode, DatasetSource};
|
||||
|
||||
|
|
@ -27,7 +30,7 @@ use igneum_pow::verify::{dataset_elem, hash_warp, splitmix32, DatasetMode, Datas
|
|||
|
||||
#[derive(Clone, Copy, PartialEq, Eq, Debug)]
|
||||
enum Gen {
|
||||
/// igneum-pow `generate`: op rolled per instruction, load weight 25, load count free.
|
||||
/// The retired version 1 generator (igneum-pow `generate_v1`): op rolled per instruction, load weight 25, load count free.
|
||||
Default,
|
||||
/// Exactly 16 load slots drawn first (partial Fisher-Yates), the other 48 ops from the ten-op table.
|
||||
Fixed16,
|
||||
|
|
@ -37,17 +40,34 @@ enum Gen {
|
|||
Fixed16Fresh,
|
||||
/// Fixed16 plus the fresh-source rule, second form: eligible means written by an earlier instruction of this
|
||||
/// program and not read by a load since that write (nothing is eligible at instruction 0, so the 16 load
|
||||
/// slots are drawn from 1..63). Cyclically fresh by construction.
|
||||
/// slots are drawn from 1..63). Cyclically fresh by construction. Adopted as generator version 2 on
|
||||
/// 4 October 2026: this draws attempt 0 of each seed through igneum-pow's own `candidate_from_words`.
|
||||
Fixed16Fresh2,
|
||||
}
|
||||
|
||||
/// Which acceptance verdict a candidate got (`accept` columns).
|
||||
#[derive(Clone, Copy, PartialEq, Eq, Debug)]
|
||||
enum Verdict {
|
||||
Accepted,
|
||||
Static,
|
||||
Dynamic,
|
||||
}
|
||||
|
||||
fn verdict(p: &Program) -> (Verdict, Option<Reject>) {
|
||||
match accept_check(p) {
|
||||
Ok(_) => (Verdict::Accepted, None),
|
||||
Err(r @ (Reject::StaleLoadSource { .. } | Reject::NoInjectingWrite { .. })) => (Verdict::Static, Some(r)),
|
||||
Err(r) => (Verdict::Dynamic, Some(r)),
|
||||
}
|
||||
}
|
||||
|
||||
impl Gen {
|
||||
fn parse(s: &str) -> Option<Gen> {
|
||||
Some(match s {
|
||||
"default" => Gen::Default,
|
||||
"fixed16" => Gen::Fixed16,
|
||||
"fixed16-fresh" => Gen::Fixed16Fresh,
|
||||
"fixed16-fresh2" => Gen::Fixed16Fresh2,
|
||||
"fixed16-fresh2" | "v2" => Gen::Fixed16Fresh2,
|
||||
_ => return None,
|
||||
})
|
||||
}
|
||||
|
|
@ -56,7 +76,7 @@ impl Gen {
|
|||
Gen::Default => "default",
|
||||
Gen::Fixed16 => "fixed16",
|
||||
Gen::Fixed16Fresh => "fixed16-fresh",
|
||||
Gen::Fixed16Fresh2 => "fixed16-fresh2",
|
||||
Gen::Fixed16Fresh2 => "v2",
|
||||
}
|
||||
}
|
||||
}
|
||||
|
|
@ -67,10 +87,14 @@ const LOAD_SLOTS: usize = 16;
|
|||
fn gen_program(seed_string: &str, gen: Gen) -> Program {
|
||||
let seed = seed_words(seed_string);
|
||||
match gen {
|
||||
Gen::Default => generate_from_words(seed_string, seed, &GeneratorConfig::default()),
|
||||
Gen::Default => generate_v1_from_words(seed_string, seed, &GeneratorConfig::default()),
|
||||
Gen::Fixed16 => generate_fixed16(seed_string, seed, Fresh::None),
|
||||
Gen::Fixed16Fresh => generate_fixed16(seed_string, seed, Fresh::V1),
|
||||
Gen::Fixed16Fresh2 => generate_fixed16(seed_string, seed, Fresh::V2),
|
||||
Gen::Fixed16Fresh2 => {
|
||||
let p = candidate_from_words(seed_string, seed_string.as_bytes(), seed, 0);
|
||||
debug_assert_eq!(p.instrs, generate_fixed16(seed_string, seed, Fresh::V2).instrs, "igneum-pow v2 is the census's fixed16-fresh2");
|
||||
p
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
|
|
@ -158,7 +182,7 @@ fn generate_fixed16(seed_string: &str, seed: [u32; 8], fresh: Fresh) -> Program
|
|||
fresh_reg[dst as usize] = true;
|
||||
instrs.push(Instr { op, dst: dst as u8, src: src as u8, src2: b as u8, imm, imm2, rot, bit: bit as u8, mask });
|
||||
}
|
||||
Program { seed_string: seed_string.to_string(), seed, instrs }
|
||||
Program { seed_string: seed_string.to_string(), seed_bytes: seed_string.as_bytes().to_vec(), seed, generator: GENERATOR_VERSION, attempt: 0, instrs }
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------------------------------------
|
||||
|
|
@ -617,6 +641,9 @@ const COLUMNS: &[&str] = &[
|
|||
"flip_max",
|
||||
"aval_hi_mean",
|
||||
"dups",
|
||||
"accept",
|
||||
"reject_static",
|
||||
"reject_dynamic",
|
||||
];
|
||||
|
||||
struct Worker {
|
||||
|
|
@ -627,6 +654,7 @@ struct Worker {
|
|||
fn census_one(idx: usize, seed_string: &str, gen: Gen, warps: usize, ds: &DatasetSource, w: &mut Worker, check: bool) -> String {
|
||||
let p = gen_program(seed_string, gen);
|
||||
let st = analyse(&p);
|
||||
let (v, _) = verdict(&p);
|
||||
let lph = p.loads_per_hash();
|
||||
let acc = &mut w.acc;
|
||||
acc.reset();
|
||||
|
|
@ -700,7 +728,7 @@ fn census_one(idx: usize, seed_string: &str, gen: Gen, warps: usize, ds: &Datase
|
|||
let or_sat_frac = if acc.or_exec == 0 { 0.0 } else { acc.or_sat as f64 / acc.or_exec as f64 };
|
||||
let mlp = if st.depth == 0 { 0.0 } else { lph as f64 / st.depth as f64 };
|
||||
format!(
|
||||
"{idx}\t{seed_string}\t{}\t{lph}\t{}\t{}\t{}\t{}\t{}\t{}\t{}\t{mlp:.3}\t{}\t{}\t{}\t{dist_mean:.3}\t{}\t{:.5}\t{:.6}\t{sites_lane_const}\t{sites_nonce_const}\t{end_sat_frac:.6}\t{}\t{}\t{or_sat_frac:.6}\t{const_bits_max}\t{const_bits_sum}\t{nonce_indep}\t{bias_max:.5}\t{bias_z_max:.3}\t{aval_mean:.4}\t{aval_std:.4}\t{flip_min:.4}\t{flip_max:.4}\t{aval_hi_mean:.4}\t{dups}",
|
||||
"{idx}\t{seed_string}\t{}\t{lph}\t{}\t{}\t{}\t{}\t{}\t{}\t{}\t{mlp:.3}\t{}\t{}\t{}\t{dist_mean:.3}\t{}\t{:.5}\t{:.6}\t{sites_lane_const}\t{sites_nonce_const}\t{end_sat_frac:.6}\t{}\t{}\t{or_sat_frac:.6}\t{const_bits_max}\t{const_bits_sum}\t{nonce_indep}\t{bias_max:.5}\t{bias_z_max:.3}\t{aval_mean:.4}\t{aval_std:.4}\t{flip_min:.4}\t{flip_max:.4}\t{aval_hi_mean:.4}\t{dups}\t{}\t{}\t{}",
|
||||
st.loads,
|
||||
st.ors,
|
||||
st.never_written,
|
||||
|
|
@ -717,6 +745,9 @@ fn census_one(idx: usize, seed_string: &str, gen: Gen, warps: usize, ds: &Datase
|
|||
if total_loads == 0 { 1.0 } else { dist_total as f64 / total_loads as f64 },
|
||||
acc.end_zero,
|
||||
acc.end_ones,
|
||||
(v == Verdict::Accepted) as u8,
|
||||
(v == Verdict::Static) as u8,
|
||||
(v == Verdict::Dynamic) as u8,
|
||||
)
|
||||
}
|
||||
|
||||
|
|
@ -744,7 +775,7 @@ fn usage() -> ! {
|
|||
eprintln!(
|
||||
"igneum-census <command> [options]\n\
|
||||
\x20 run census: --root <s> --count <n> [--start <i>] [--warps <w>] [--threads <t>] [--day <d>]\n\
|
||||
\x20 [--closed-form] [--dataset-log2 <n>] [--gen default|fixed16|fixed16-fresh] --out <tsv>\n\
|
||||
\x20 [--closed-form] [--dataset-log2 <n>] [--gen v2|default|fixed16|fixed16-fresh] --out <tsv>\n\
|
||||
\x20 probe one line per named seed: --seed <s> [--seed <s> ...] [--warps <w>] [--gen ...]\n\
|
||||
\x20 summarise tables from a census TSV: --in <tsv>\n\
|
||||
\x20 show print the program for a seed: --seed <s> [--gen ...]"
|
||||
|
|
@ -810,15 +841,16 @@ fn build_dataset(a: &Args) -> DatasetSource {
|
|||
ds
|
||||
}
|
||||
|
||||
/// The spec vector for igneum-genesis lane 0, memory-hard, 2^28 words, day 2026-10-03.
|
||||
/// The spec vector for igneum-genesis lane 0, memory-hard, 2^28 words, day 2026-10-03 (generator v2, attempt 0).
|
||||
fn self_check(ds: &DatasetSource, a: &Args) {
|
||||
if !a.closed_form && a.dataset_log2 == 28 && a.day == "2026-10-03" {
|
||||
let p = gen_program("igneum-genesis", Gen::Default);
|
||||
let p = gen_program("igneum-genesis", Gen::Fixed16Fresh2);
|
||||
assert_eq!(verdict(&p).0, Verdict::Accepted, "igneum-genesis attempt 0 is the accepted program");
|
||||
let mut w = Worker { acc: Acc::new(), lane_addrs: Vec::new() };
|
||||
let h = run_warp(&p, 0, ds, &mut w.acc, &mut w.lane_addrs, p.loads_per_hash());
|
||||
assert_eq!(h[0], 0x1fb0b3bbc1ac8279, "igneum-genesis vector lane 0");
|
||||
assert_eq!(h[31], 0xfa052263a854f3de, "igneum-genesis vector lane 31");
|
||||
eprintln!("self-check: igneum-genesis warp 0 matches the spec vectors (lanes 0 and 31)");
|
||||
assert_eq!(h[0], 0x42246ba99fc58e4f, "igneum-genesis vector lane 0");
|
||||
assert_eq!(h[31], 0xb08446b1f2de7793, "igneum-genesis vector lane 31");
|
||||
eprintln!("self-check: igneum-genesis warp 0 matches the spec 01 section 1.17 vectors (lanes 0 and 31)");
|
||||
}
|
||||
}
|
||||
|
||||
|
|
@ -884,6 +916,11 @@ fn show(a: &Args) {
|
|||
let p = gen_program(s, a.gen);
|
||||
let st = analyse(&p);
|
||||
println!("seed {s} gen {} op mix {} loads/hash {}", a.gen.name(), p.op_mix(), p.loads_per_hash());
|
||||
match verdict(&p) {
|
||||
(Verdict::Accepted, _) => println!("acceptance rule: ACCEPTED (program id {:016x})", p.program_id()),
|
||||
(_, Some(r)) => println!("acceptance rule: REJECTED, {r}"),
|
||||
_ => {}
|
||||
}
|
||||
println!("{st:?}");
|
||||
for (k, i) in p.instrs.iter().enumerate() {
|
||||
println!("{k:2}: {:5} dst={} src={} src2={} rot={} bit={} mask={}", i.op.name(), i.dst, i.src, i.src2, i.rot, i.bit, i.mask);
|
||||
|
|
@ -1119,6 +1156,30 @@ fn summarise(a: &Args) {
|
|||
println!("| {name} | {c} | {:.4}% |", 100.0 * c as f64 / n as f64);
|
||||
flag_hits.push(hits);
|
||||
}
|
||||
// The acceptance rule's own verdict (igneum-pow accept::check, recorded per program when the TSV has the columns).
|
||||
if t.cols.iter().any(|c| c == "accept") {
|
||||
let acc = col(&t, "accept");
|
||||
let rs = col(&t, "reject_static");
|
||||
let rd = col(&t, "reject_dynamic");
|
||||
let na = acc.iter().filter(|&&x| x == 1.0).count();
|
||||
let ns = rs.iter().filter(|&&x| x == 1.0).count();
|
||||
let nd = rd.iter().filter(|&&x| x == 1.0).count();
|
||||
println!("\n## Acceptance rule (igneum-pow accept::check, spec 01 section 1.4.6)\n");
|
||||
println!("| Verdict | programs | share |");
|
||||
println!("|---|---|---|");
|
||||
println!("| accepted | {na} | {:.3}% |", 100.0 * na as f64 / n as f64);
|
||||
println!("| rejected, static (a or b) | {ns} | {:.3}% |", 100.0 * ns as f64 / n as f64);
|
||||
println!("| rejected, dynamic (c) | {nd} | {:.3}% |", 100.0 * nd as f64 / n as f64);
|
||||
println!("| rejected, all | {} | {:.3}% |", ns + nd, 100.0 * (ns + nd) as f64 / n as f64);
|
||||
println!("| expected candidates per epoch | | {:.4} |", n as f64 / na.max(1) as f64);
|
||||
let dm = col(&t, "dist_mean");
|
||||
let mut w: Vec<f64> = (0..n).filter(|&i| acc[i] == 1.0).map(|i| dm[i]).collect();
|
||||
if !w.is_empty() {
|
||||
let mean = w.iter().sum::<f64>() / w.len() as f64;
|
||||
w.sort_by(|x, y| x.partial_cmp(y).unwrap());
|
||||
println!("\nDistinct addresses per hash over the {} accepted programs (this census's own measurement on its dataset): mean {:.3}, min {:.3}, p1 {:.3}, p50 {:.3}, max {:.3}", w.len(), mean, w[0], pct(&w, 0.01), pct(&w, 0.5), w[w.len() - 1]);
|
||||
}
|
||||
}
|
||||
// Static features recomputed from the seed (so older TSVs get the newer columns), then the candidate
|
||||
// rules, each crossed with every dynamic flag.
|
||||
let gen = t.header_note.split_whitespace().find_map(|w| w.strip_prefix("gen=")).and_then(Gen::parse).unwrap_or(Gen::Default);
|
||||
|
|
|
|||
2
igneum-pow/Cargo.lock
generated
2
igneum-pow/Cargo.lock
generated
|
|
@ -4,7 +4,7 @@ version = 4
|
|||
|
||||
[[package]]
|
||||
name = "igneum-pow"
|
||||
version = "0.1.0"
|
||||
version = "0.2.0"
|
||||
dependencies = [
|
||||
"serde_json",
|
||||
]
|
||||
|
|
|
|||
|
|
@ -1,8 +1,8 @@
|
|||
[package]
|
||||
name = "igneum-pow"
|
||||
version = "0.1.0"
|
||||
version = "0.2.0"
|
||||
edition = "2021"
|
||||
description = "Igneum random-program GPU proof-of-work: seed, program generator, memory-hard dataset, CPU warp verifier and kernel emitters, bit-exact with proto-metal"
|
||||
description = "Igneum random-program GPU proof-of-work: seed, program generator (version 2: fixed load count, fresh sources, acceptance rule), memory-hard dataset, CPU warp verifier and kernel emitters; the source of every program pack"
|
||||
license = "MIT"
|
||||
publish = false
|
||||
|
||||
|
|
|
|||
|
|
@ -1,37 +1,60 @@
|
|||
# igneum-pow
|
||||
|
||||
The Igneum lottery hash in Rust, bit-exact with the Swift prototype in `proto-metal/main.swift`. This is the
|
||||
crate the rusty-kaspa fork will call (`docs/fork-map.md`, rows a1 to a3) so a node written in Rust can verify any
|
||||
block and hand miners the kernel source for the epoch. No dependency outside the standard library; `serde_json`
|
||||
is a dev-dependency for reading the packs in the tests.
|
||||
The Igneum lottery hash in Rust: the generator, the acceptance rule, the memory-hard dataset, the CPU verifier and the
|
||||
kernel emitters. This is the crate the rusty-kaspa fork calls (`docs/fork-map.md`, rows a1 to a3) and, since
|
||||
4 October 2026, the source of every program pack in `proto-cuda/packs/`. No dependency outside the standard
|
||||
library; `serde_json` is a dev-dependency for reading the packs in the tests.
|
||||
|
||||
Date: 3 October 2026. Toolchain: rustc 1.99.0 via rustup (the Homebrew 1.69 on PATH is too old; use
|
||||
`~/.cargo/bin/cargo`).
|
||||
Dates: 3 October 2026 (crate, bit-exact with the Swift prototype), 4 October 2026 (generator version 2 and the
|
||||
acceptance rule; every vector re-cut). Toolchain: rustc 1.99.0 via rustup (the Homebrew 1.69 on PATH is too old;
|
||||
use `~/.cargo/bin/cargo`). Crate version 0.2.0.
|
||||
|
||||
## Modules
|
||||
|
||||
| Module | What it is | Swift namesake |
|
||||
|---|---|---|
|
||||
| `seed` | 32-byte seed words from a string (FNV-1a 64, four salts, finalised); `seed_words_from_bytes` is the boundary where the chain will feed the VDF output; SplitMix64 | `seedWords`, `SplitMix64` |
|
||||
| `generator` | the 64-instruction program for a seed (op, dst, src, src2, imm, imm2, rot, bit, mask); levers `load_weight` and `wide_frac` | `generateProgram`, `GeneratorConfig` |
|
||||
| `seed` | 32-byte seed words from bytes (FNV-1a 64, four salts, finalised); `seed_words_from_bytes` is the boundary where the chain feeds the epoch seed; SplitMix64 | `seedWordsBytes`, `SplitMix64` |
|
||||
| `generator` | version 2: 16 load slots drawn first from instructions 1..63, fresh-source loads, the other 48 ops from the ten non-load weights; attempts `k = 0, 1, ...` of a seed; the program id. The retired version 1 generator stays as `generate_v1` for the census and the lever measurements | `generateProgramV2`, `candidateProgram`, `generateProgramV1` |
|
||||
| `accept` | the acceptance rule of spec 01 section 1.4.6: two static tests and the 64-unit dynamic test on the seed-keyed closed-form dataset | `acceptProgram` |
|
||||
| `memhard` | 256 MiB cache (2^16 chains of 64 ChaCha12 blocks), mixer parameters, 8-round item derivation with the 32 lanes interleaved, `MemhardCpu::fetch` | `cpuFillCache`, `MixParams`, `deriveItems`, `MemhardCPU` |
|
||||
| `verify` | the 32-lane warp interpreter, `DatasetMode::{ClosedForm, MemoryHard}`, `Epoch`, `hash_warp`, `verify_block` | `cpuWarp`, `DatasetSource` |
|
||||
| `emit` | Metal, CUDA and OpenCL source, program.h, memhard.h, vectors.h, program.json, vectors.json, `export_pack`; since 3 October 2026 also the header-bound kernels `program_bound.metal` and `kernel_bound.cu` | `generateMSL`, `memhardMSL`, `emitMemhardCore`, `generateCUDA`, `generateOpenCL`, `exportPack` |
|
||||
| `bind` | header binding (spec 01 section 1.6, O-1.9): init words from `"igneum-block/" \|\| H \|\| nonce_hi_le32`, bound hash API on `Epoch`, the 256-bit pow mapping, the interim day seed bytes | none (new) |
|
||||
| `emit` | Metal, CUDA and OpenCL source, program.h, memhard.h, vectors.h, program.json, vectors.json, the header-bound kernels, `export_pack` | `generateMSL`, `generateCUDA`, `generateOpenCL`, `exportPack` |
|
||||
| `bind` | header binding (spec 01 section 1.6): init words from `"igneum-block/" \|\| H \|\| nonce_hi_le32`, bound hash API on `Epoch`, the 256-bit pow mapping, the interim day seed bytes | `blockInitWords` |
|
||||
|
||||
## Generator version 2 (4 October 2026)
|
||||
|
||||
Adopted from `docs/analysis/weak-program-census-2026-10-03.md` (ledger M5 and M6). Three parts, all in this crate
|
||||
and mirrored in `proto-metal/main.swift` so the Metal worker derives the same program from the same seed:
|
||||
|
||||
| Part | Rule | Where |
|
||||
|---|---|---|
|
||||
| G1, exact load count | 16 `load` instructions per program, a uniform 16-subset of slots 1..63 drawn first by partial Fisher-Yates over the program stream; the other 48 ops from `add 12, xor 10, mul 8, mad 8, shfl 8, rotl 7, sub 6, mulhi 6, rotr 6, or 4` (sum 75). 128 loads per hash, 4,096 items per unit | `generator::candidate_from_words` |
|
||||
| G2, fresh source | a load's source is drawn from the registers other than `dst` written by an earlier instruction and not read by a load since, so no load repeats an earlier load's address in the hash | same |
|
||||
| R, acceptance | (a) no load whose source is unwritten since the previous load from it, cyclically; (b) every register has an `add`, `sub`, `xor`, `mad`, `shfl` or `load` write; (c) 64 units at base nonces from `SplitMix64(FNV-1a-64("igneum-accept/" \|\| seed words LE))`, init words = seed words, closed-form dataset `dataset_elem(idx, S[0], S[1])` at 2^28 words: no constant register bit, no lane-constant load site in any unit, fewer than 164 saturated final values, every output bit within 136 of 1,024, distinct addresses above 245,760 over the 2,048 hashes | `accept::check` |
|
||||
| Attempts | a rejected candidate is replaced by `seed_words_from_bytes(seed \|\| k_le32)` for `k = 1, 2, ...`; 32 consecutive rejections are a consensus fault (probability below 2^-136 at the measured 5 percent rate) | `generator::generate_from_seed_bytes` |
|
||||
| Program id | `FNV-1a-64("igneum-program/" \|\| 2_le32 \|\| seed words LE \|\| attempt_le32)`, written into program.json and program.h with the generator version and the attempt, so a version 1 pack or another attempt can never pass for the current program | `Program::program_id` |
|
||||
|
||||
Measured on this crate (`igneum-pow accept`): `igneum-genesis` attempt 0 accepted, program id `bcc1248b10cc90f2`,
|
||||
op mix `load=16 add=8 shfl=8 xor=6 mad=5 mul=5 mulhi=5 sub=4 rotl=3 rotr=3 or=1`, 128.000 distinct addresses per
|
||||
hash; `igneum-hourly` attempt 0, id `a4c4d00961c855df`; seeds `igneum-census-2026-10-03/22`, `/37` and `/51` have
|
||||
attempt 0 rejected ((b) r7 without an injecting write; (c) 119.74 distinct addresses; (b) r4) and attempt 1
|
||||
accepted (ids `22ed0609d079f4cf`, `947705cc4eb1df0a`, `9869afcc028bf9f1`); those three are the conformance vectors
|
||||
for the attempt rule. The rule costs 1.3 to 3.4 ms per seed on one core. The 20,000-program census under this
|
||||
generator is in `docs/bench-log.md` (4 October 2026 entry).
|
||||
|
||||
## The API the fork calls
|
||||
|
||||
```rust
|
||||
use igneum_pow::{Epoch, DatasetMode};
|
||||
|
||||
// Once per epoch and day: generates the program and fills the 256 MiB cache (about 0.18 s on one core).
|
||||
// Once per epoch and day: derives the accepted program and fills the 256 MiB cache (about 0.2 s on one core).
|
||||
let epoch = Epoch::memory_hard("igneum-genesis", "2026-10-03");
|
||||
|
||||
let h: u64 = epoch.hash(nonce); // one nonce (computes its aligned 32-nonce warp)
|
||||
let w: [u64; 32] = epoch.hash_warp(base_nonce); // one warp
|
||||
let ok: bool = epoch.verify_block(nonce, target_u64);
|
||||
|
||||
// Miner programs for the epoch, byte-identical to the Swift exporter.
|
||||
// Miner programs for the epoch: the pack every worker compiles.
|
||||
let pack = igneum_pow::emit::export_pack(&epoch, "2026-10-03", "igneum node");
|
||||
pack.write_to(std::path::Path::new("out"))?; // kernel.cu, kernel.cl, program.metal, memhard.h, ...
|
||||
```
|
||||
|
|
@ -40,7 +63,7 @@ pack.write_to(std::path::Path::new("out"))?; // kernel.cu, kernel.cl, program.
|
|||
|
||||
The pack form above initialises the lane registers from the program's own seed words, so one nonce has one
|
||||
hash per epoch whatever block is mined. On the chain the init words commit to the block (spec 01 section 1.6,
|
||||
open item O-1.9, implemented 3 October 2026 in `src/bind.rs`):
|
||||
`src/bind.rs`):
|
||||
|
||||
```
|
||||
H = header hash with the nonce field zeroed, every other field as mined
|
||||
|
|
@ -52,11 +75,10 @@ pow256 = hash in the top 64 bits, low 192 bits zero (little-endian bytes 24..3
|
|||
valid = pow256 <= target256, which is exactly hash <= target256 >> 192
|
||||
```
|
||||
|
||||
Choices the spec left open and how they were fixed: `H` keeps the timestamp (Kaspa zeroes it in the pre-PoW hash
|
||||
and absorbs it in cSHAKE afterwards; the lane hash has no afterwards, so a nonce would otherwise be reusable across
|
||||
timestamps). The pow value puts the lane in the top 64 bits with zero low bits so a GPU worker and the node compare
|
||||
the same 64-bit numbers. The interim day seed is `"igneum-day/" || day_le64` with `day = timestamp_ms / 86,400,000`
|
||||
(O-1.10's proposal needs the VDF schedule). The epoch seed bytes are the 32 bytes of the epoch block hash (devnet v0).
|
||||
`H` keeps the timestamp (Kaspa zeroes it in the pre-PoW hash and absorbs it in cSHAKE afterwards; the lane hash has
|
||||
no afterwards, so a nonce would otherwise be reusable across timestamps). The interim day seed is
|
||||
`"igneum-day/" || day_le64` with `day = timestamp_ms / 86,400,000`. The epoch seed bytes are the 32 bytes of the
|
||||
epoch block hash (devnet); the program is `generate_from_seed_bytes(epoch_seed)`, attempts included.
|
||||
|
||||
```rust
|
||||
use igneum_pow::{bind, Epoch};
|
||||
|
|
@ -71,29 +93,28 @@ let ok = epoch.verify_block_bound(&prehash, nonce, bind::target64_from_le256(&ta
|
|||
```
|
||||
|
||||
On the GPU the init words are a kernel argument: `igneum_hash_bound` in `program_bound.metal` takes
|
||||
`constant uint* initw [[buffer(3)]]`, and `kernel_bound.cu` takes `IgneumInitWords iw` by value. Both are emitted
|
||||
into every pack next to the unchanged `igneum_hash` and differ from it only in the kernel name, the argument and the
|
||||
eight init lines (`initw[i]` / `iw.w[i]` instead of `SEEDW[i]`). The lane nonce stays `baseNonce + gid`.
|
||||
`constant uint* initw [[buffer(3)]]`, `kernel_bound.cu` takes `IgneumInitWords iw` by value, `kernel_bound.cl` an
|
||||
`initw` buffer. All three differ from `igneum_hash` only in the kernel name, the argument and the eight init lines.
|
||||
|
||||
Bound vectors (seed `igneum-genesis`, day `2026-10-03`, memory-hard, 2^28 words; `igneum-pow hash-bound`):
|
||||
Bound vectors (seed `igneum-genesis`, day `2026-10-03`, memory-hard, 2^28 words, generator v2; `igneum-pow hash-bound`):
|
||||
|
||||
| H | nonce | hash_bound |
|
||||
|---|---|---|
|
||||
| 32 zero bytes | 0 | `2c619692d823263b` |
|
||||
| 32 zero bytes | 1 | `55d21ed545735ce8` |
|
||||
| 32 zero bytes | 31 | `862eebe7fbda564e` |
|
||||
| 32 zero bytes | 4096 | `37aadc51f95725df` |
|
||||
| 32 zero bytes | 4294967296 (1 << 32) | `b62e28b8a90e554f` |
|
||||
| bytes 00 01 02 .. 1f | 0 | `9b2437118e087833` |
|
||||
| bytes 00 01 02 .. 1f | 4294967301 ((1 << 32) + 5) | `714ae31e369e0446` |
|
||||
| bytes 00 01 02 .. 1f | 18446744073709551615 (u64::MAX) | `4ca4f84079025a13` |
|
||||
| 32 zero bytes | 0 | `746c567b090acf6a` |
|
||||
| 32 zero bytes | 1 | `45a619f860880c73` |
|
||||
| 32 zero bytes | 31 | `9aa495e43dedbfe6` |
|
||||
| 32 zero bytes | 4096 | `2e6ffd7624d3cba2` |
|
||||
| 32 zero bytes | 4294967296 (1 << 32) | `38a5cea1fb01431a` |
|
||||
| bytes 00 01 02 .. 1f | 0 | `2a79c5e4797bf6aa` |
|
||||
| bytes 00 01 02 .. 1f | 4294967301 ((1 << 32) + 5) | `a243e0c61aa1b82e` |
|
||||
| bytes 00 01 02 .. 1f | 18446744073709551615 (u64::MAX) | `9c7bbfbd064fe1a4` |
|
||||
|
||||
Init words for H = 32 zero bytes, nonce 0: `595a8f8a 37647e95 faadade1 cbbcf2a4 54f7cc13 f6851b5e 8c68ca04 7991ea9c`.
|
||||
The eight are pinned in `bind::tests::bound_vectors`. The pack vectors (96 per pack) are unchanged.
|
||||
Init words for H = 32 zero bytes, nonce 0: `595a8f8a 37647e95 faadade1 cbbcf2a4 54f7cc13 f6851b5e 8c68ca04 7991ea9c`
|
||||
(unchanged by v2: the binding does not depend on the program). The eight are pinned in `bind::tests::bound_vectors`.
|
||||
The version 1 bound vectors of 3 October 2026 are retired.
|
||||
|
||||
`Epoch` is `Send + Sync`; build one and share it. `DatasetMode::ClosedForm` reproduces the two old packs
|
||||
(`igneum-genesis`, `igneum-hourly`) and is not memory-hard. The hash is 64 bits; the fork maps it into its
|
||||
256-bit target space in `consensus/pow/src/lib.rs`.
|
||||
`Epoch` is `Send + Sync`; build one and share it. `DatasetMode::ClosedForm` is the interpreter regression dataset
|
||||
of the packs `igneum-genesis` and `igneum-hourly` and is not memory-hard.
|
||||
|
||||
## CLI
|
||||
|
||||
|
|
@ -101,56 +122,60 @@ The eight are pinned in `bind::tests::bound_vectors`. The pack vectors (96 per p
|
|||
cargo build --release
|
||||
./target/release/igneum-pow bench --seed igneum-genesis [--warps 20] [--closed-form] [--day 2026-10-03]
|
||||
./target/release/igneum-pow export --seed igneum-genesis --out <dir> [--closed-form]
|
||||
./target/release/igneum-pow export --epoch-hex <64 hex> --day-hex <hex> --out <dir> the chain's byte seeds
|
||||
./target/release/igneum-pow hash --seed igneum-genesis --nonce 4103
|
||||
./target/release/igneum-pow hash-bound --seed igneum-genesis --prehash <64 hex> --nonce <u64>
|
||||
./target/release/igneum-pow accept --seed igneum-census-2026-10-03/22 every candidate with its verdict
|
||||
./target/release/igneum-pow show --seed igneum-genesis the accepted program, one line per instruction
|
||||
```
|
||||
|
||||
The packs were regenerated on 4 October 2026 with exactly these commands:
|
||||
|
||||
```
|
||||
igneum-pow export --closed-form --seed igneum-genesis --out ../proto-cuda/packs/igneum-genesis
|
||||
igneum-pow export --closed-form --seed igneum-hourly --out ../proto-cuda/packs/igneum-hourly
|
||||
igneum-pow export --seed igneum-genesis --out ../proto-cuda/packs/igneum-genesis-mh
|
||||
igneum-pow export --epoch-hex edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07 \
|
||||
--day-hex 69676e65756d2d6461792ffa50000000000000 --out ../proto-cuda/packs/igneum-devnet-v4-epoch0
|
||||
```
|
||||
|
||||
The last is the chain's own derivation for devnet v4 epoch 0: the devnet genesis hash as the epoch seed and
|
||||
`bind::day_bytes(20730)` (2026-10-04) as the day bytes.
|
||||
|
||||
## Tests
|
||||
|
||||
`cargo test` (29 tests, 1 s after compile; the dev profile is optimised so the cache fill is quick):
|
||||
`cargo test --release` (39 tests, about 2 s after compile; the dev profile is optimised so the cache fill is quick):
|
||||
|
||||
| Check | Pack | Result |
|
||||
|---|---|---|
|
||||
| program.json instruction by instruction, op mix, loads per hash | igneum-genesis, igneum-genesis-mh, igneum-hourly | 3 x 64 match |
|
||||
| Mixer parameters (key, rot, mul, rc) | igneum-genesis-mh | match |
|
||||
| Cache head, last line, FNV-1a 64 `48c4f5bf24166b2e` | igneum-genesis-mh | match |
|
||||
| Dataset head (16), `[MASK]`, 64 sampled words | all three | match |
|
||||
| 96 hash vectors (3 warps x 32 lanes) | igneum-genesis-mh | 96/96 |
|
||||
| 96 hash vectors | igneum-genesis, igneum-hourly | 96/96 each |
|
||||
| kernel.cu, program.metal, kernel.cl, program.h byte-identical | all three | identical |
|
||||
| memhard.h, memhard.metal byte-identical | igneum-genesis-mh | identical |
|
||||
| program.json byte-identical (after the fix below) | all three | identical |
|
||||
| vectors.json, vectors.h byte-identical apart from the provenance string | all three | identical |
|
||||
| program.json instruction by instruction from `seed_bytes`, generator 2, attempt, program id, op mix, 128 loads, acceptance | all four | match |
|
||||
| Mixer parameters (key, rot, mul, rc) | igneum-genesis-mh, igneum-devnet-v4-epoch0 | match |
|
||||
| Cache head, last line, FNV-1a 64 (`48c4f5bf24166b2e` for day 2026-10-03, `448274a57f508cbc` for day bytes 20730) | the two memory-hard packs | match |
|
||||
| Dataset head (16), `[MASK]`, 64 sampled words | all four | match |
|
||||
| 96 hash vectors (3 warps x 32 lanes) | all four | 96/96 each |
|
||||
| kernel.cu, kernel_bound.cu, program.metal, program_bound.metal, kernel.cl, kernel_bound.cl, program.h, program.json byte-identical; 16 masked loads per kernel | all four | identical |
|
||||
| memhard.h, memhard.metal byte-identical | the two memory-hard packs | identical |
|
||||
| vectors.json, vectors.h byte-identical; no stale file in any pack directory | all four | identical |
|
||||
| bound vectors (8), bound warp == bound single, H and nonce_hi enter the hash | igneum-genesis-mh | pass |
|
||||
| the devnet pack equals `Epoch::from_seed_bytes(genesis, day_bytes(20730))` | igneum-devnet-v4-epoch0 | pass |
|
||||
| acceptance: instrumented interpreter == `hash_warp`, cyclic stale-load detection, injecting-write detection, rejection under 12.5 percent and distinct loads above 127 on 400 census seeds, version 1 programs mostly rejected | unit tests | pass |
|
||||
| generator: 16 loads, none at instruction 0, contract on 200 candidates; fresh sources; attempt words; program ids separate versions and attempts | unit tests | pass |
|
||||
|
||||
An independent `diff -r` of `igneum-pow export` output against the checked-in packs shows the same two lines
|
||||
only: the provenance string and the `"item"` line, plus the two bound files that only the Rust exporter writes
|
||||
(`program_bound.metal`, `kernel_bound.cu`).
|
||||
## Measured, Apple M5 Max, one core, release build
|
||||
|
||||
One deliberate difference: `proto-cuda/packs/igneum-genesis-mh/program.json` as written by the Swift is not
|
||||
valid JSON (main.swift line 1291 uses `jhex` inside the `"item"` string, so the cache line mask is quoted inside a
|
||||
quoted string). The Rust emitter writes `0x003fffff` bare; the test normalises that one line before comparing.
|
||||
A node must hand miners valid JSON, so the Rust side does not reproduce the defect.
|
||||
|
||||
## Measured, 3 October 2026, Apple M5 Max, one core, release build
|
||||
|
||||
| Step | Rust | Swift (MEMHARD.md) |
|
||||
| Step | 3 October 2026 (v1, 104 loads) | 4 October 2026 (v2, 128 loads, 4,096 items) |
|
||||
|---|---|---|
|
||||
| Cache fill, 256 MiB, 65,536 chains x 64 ChaCha12 blocks | 175 to 181 ms (5 quiet runs; 200 ms once with another build running) | 184.5 to 190.6 ms (C++ host reference 161.5) |
|
||||
| CPU verify per warp, igneum-genesis, 104 loads, 3,328 items, avg of 20 | 0.441 ms | 0.649 ms |
|
||||
| igneum-genesis/epoch1, 104 loads | 0.411 ms | 0.631 ms |
|
||||
| igneum-genesis/epoch2, 112 loads | 0.488 ms | 0.701 ms |
|
||||
| igneum-second-seed, 104 loads | 0.482 ms | 0.801 ms |
|
||||
| igneum-second-seed/epoch1, 144 loads, 4,608 items | 0.579 ms | 1.205 ms |
|
||||
| Cold single warps across the five seeds | 0.41 to 0.87 ms | 1.16 to 2.11 ms |
|
||||
| Closed form, igneum-genesis | 0.002 ms | 0.017 ms |
|
||||
| Cache fill, 256 MiB | 175 to 181 ms | 179 ms |
|
||||
| CPU verify per warp, igneum-genesis, avg of 20 | 0.441 ms (3,328 items) | 0.631 ms |
|
||||
| Cold single warps (bases 0, 4096, 1000000) | 0.41 to 0.87 ms | 0.67 to 0.81 ms |
|
||||
| Acceptance rule per candidate | | 1.3 to 3.4 ms |
|
||||
| Closed form, igneum-genesis | 0.002 ms | 0.002 ms |
|
||||
|
||||
The Rust verifier is 1.4x to 2.1x faster than the Swift one per warp; the registers are kept register-major
|
||||
(`r[reg][lane]`) so the lane loops vectorise, and the item derivation interleaves the 32 lanes round by round as
|
||||
the Swift does. The 10 ms gate holds with a margin of about 17x on the steady figure and 11x on the worst cold warp.
|
||||
The 10 ms gate holds with a margin of about 16x steady on this core. Every verified unit now derives exactly 4,096
|
||||
items, the design bound of spec section 1.11.
|
||||
|
||||
## Not done here
|
||||
|
||||
- No GPU. The vectors tie this crate to the Metal and CUDA results through the packs; nothing here runs a kernel.
|
||||
- `Epoch::from_seed_bytes` takes the epoch seed as bytes (the epoch block hash on devnet v0); the VDF output enters there.
|
||||
- The OpenCL pack has no bound kernel yet; `proto-opencl/host.c` builds its own from `kernel.cl` text at runtime.
|
||||
- No GPU. The vectors tie this crate to the Metal, CUDA and OpenCL results through the packs; nothing here runs a kernel.
|
||||
- The epoch seed enters at `Epoch::from_seed_bytes` (the epoch block hash on devnet); the VDF output replaces it there.
|
||||
- The dynamic acceptance test is specified on the closed-form dataset at 2^28 words; if the prototype mask ever changes, the rule's constant stays at 2^28.
|
||||
|
|
|
|||
459
igneum-pow/src/accept.rs
Normal file
459
igneum-pow/src/accept.rs
Normal file
|
|
@ -0,0 +1,459 @@
|
|||
//! Program acceptance (spec 01 section 1.4.6, adopted 4 October 2026 from the weak-program census of
|
||||
//! `docs/analysis/weak-program-census-2026-10-03.md`, section 6).
|
||||
//!
|
||||
//! A candidate program is accepted only if every test below holds. Every conforming implementation evaluates
|
||||
//! exactly these tests on exactly these inputs, so every node skips the same seeds.
|
||||
//!
|
||||
//! | Part | Test |
|
||||
//! |---|---|
|
||||
//! | (a) static | for every `load`, some instruction between the previous `load` from the same source register and this one, in cyclic order over the 64 instructions, writes that register |
|
||||
//! | (b) static | every register `r0..r7` is the destination of at least one `add`, `sub`, `xor`, `mad`, `shfl` or `load` |
|
||||
//! | (c) dynamic | the program is interpreted for [`ACCEPT_UNITS`] (64) units of 32 lanes at base nonces drawn from SplitMix64 seeded with `FNV-1a-64("igneum-accept/" \|\| seed words as little-endian bytes)`, each `low32(next()) AND NOT 31`, with init words equal to the seed words and the closed-form dataset `dataset_elem(idx, S[0], S[1])` at [`ACCEPT_DATASET_LOG2`] (2^28 words) in place of the memory-hard dataset. Over the 2,048 evaluations: no register has a bit equal in every final value; no load site (iteration, instruction) reads one address in all 32 lanes of any unit; fewer than [`MAX_SATURATED`] (164, 1 percent of 16,384) final register values are 0 or 2^32 - 1; every output bit's ones count is within [`BIAS_TOLERANCE`] (136, 6 sigma) of 1,024; the distinct masked addresses read by one lane in one evaluation, summed over the 2,048 evaluations, exceed [`MIN_DISTINCT_SUM`] (245,760, a mean above 120 of the 128 loads) |
|
||||
//!
|
||||
//! The dynamic test uses the closed form so that it is a pure function of the program (no cache, no day) and
|
||||
//! costs about a millisecond on one core. The census (section 7.3) checked on 100,000 programs that the
|
||||
//! closed-form verdict agrees with the memory-hard one on all but 39 threshold-edge cases.
|
||||
|
||||
use crate::generator::{Instr, Op, Program, INSTR_COUNT, ITERATIONS, LANES};
|
||||
use crate::seed::{fnv1a64, SplitMix64};
|
||||
use crate::verify::{dataset_elem, splitmix32};
|
||||
|
||||
/// Units (32-lane warps) the dynamic test interprets.
|
||||
pub const ACCEPT_UNITS: usize = 64;
|
||||
/// Hashes the dynamic test evaluates: 2,048.
|
||||
pub const ACCEPT_HASHES: usize = ACCEPT_UNITS * LANES;
|
||||
/// Domain tag of the base-nonce stream.
|
||||
pub const ACCEPT_TAG: &[u8] = b"igneum-accept/";
|
||||
/// log2 of the closed-form dataset the test addresses: the prototype's 2^28 words, MASK 0x0fffffff.
|
||||
pub const ACCEPT_DATASET_LOG2: u32 = 28;
|
||||
/// Final register values equal to 0 or 2^32 - 1 must number fewer than this (1 percent of 8 x 2,048).
|
||||
pub const MAX_SATURATED: u32 = 164;
|
||||
/// Every output bit's ones count must be within this of 1,024 (6 x sqrt(2048) / 2, rounded).
|
||||
pub const BIAS_TOLERANCE: u32 = 136;
|
||||
/// Distinct addresses per lane per evaluation, summed over 2,048 evaluations, must exceed this (mean above 120).
|
||||
pub const MIN_DISTINCT_SUM: u64 = 245_760;
|
||||
|
||||
/// Why a candidate was rejected. The verdict (accept or reject) is what consensus depends on; the reason is the
|
||||
/// first failing test in the order of the module table.
|
||||
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
|
||||
pub enum Reject {
|
||||
/// (a): instruction `instr` loads from `reg`, which no instruction wrote since the previous load from it.
|
||||
StaleLoadSource { instr: u8, reg: u8 },
|
||||
/// (b): no injecting op writes `reg`.
|
||||
NoInjectingWrite { reg: u8 },
|
||||
/// (c): `reg` has `bits` bits equal in all 2,048 final values.
|
||||
ConstantBit { reg: u8, bits: u8 },
|
||||
/// (c): the load at `instr` in `iteration` read one address in all 32 lanes of `unit`.
|
||||
LaneConstantSite { iteration: u8, instr: u8, unit: u8 },
|
||||
/// (c): `count` final register values were 0 or all ones.
|
||||
Saturated { count: u32 },
|
||||
/// (c): output bit `bit` was set in `ones` of 2,048 hashes.
|
||||
OutputBias { bit: u8, ones: u32 },
|
||||
/// (c): the distinct-address sum was `sum`.
|
||||
DistinctAddresses { sum: u64 },
|
||||
}
|
||||
|
||||
impl std::fmt::Display for Reject {
|
||||
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
|
||||
match self {
|
||||
Reject::StaleLoadSource { instr, reg } => {
|
||||
write!(f, "(a) load at instruction {instr} reads r{reg}, unwritten since the previous load from it")
|
||||
}
|
||||
Reject::NoInjectingWrite { reg } => write!(f, "(b) r{reg} has no add, sub, xor, mad, shfl or load write"),
|
||||
Reject::ConstantBit { reg, bits } => write!(f, "(c) r{reg} has {bits} nonce-independent bits"),
|
||||
Reject::LaneConstantSite { iteration, instr, unit } => {
|
||||
write!(f, "(c) load at iteration {iteration} instruction {instr} reads one address in all lanes of unit {unit}")
|
||||
}
|
||||
Reject::Saturated { count } => write!(f, "(c) {count} of 16384 final register values saturated (limit 163)"),
|
||||
Reject::OutputBias { bit, ones } => write!(f, "(c) output bit {bit} set in {ones} of 2048 hashes"),
|
||||
Reject::DistinctAddresses { sum } => {
|
||||
write!(f, "(c) distinct addresses {sum} over 2048 hashes (mean {:.2}, needs above 120)", *sum as f64 / 2048.0)
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// What the dynamic test measured on an accepted program.
|
||||
#[derive(Clone, Copy, Debug, Default, PartialEq, Eq)]
|
||||
pub struct AcceptReport {
|
||||
/// Distinct masked addresses per lane per evaluation, summed over the 2,048 evaluations.
|
||||
pub distinct_sum: u64,
|
||||
/// Final register values equal to 0 or all ones.
|
||||
pub saturated: u32,
|
||||
/// The largest `|ones - 1024|` over the 64 output bits.
|
||||
pub bias_max: u32,
|
||||
}
|
||||
|
||||
impl AcceptReport {
|
||||
/// Mean distinct addresses per hash (128 at most).
|
||||
pub fn distinct_mean(&self) -> f64 {
|
||||
self.distinct_sum as f64 / ACCEPT_HASHES as f64
|
||||
}
|
||||
}
|
||||
|
||||
/// Part (a): no load whose source is unwritten since the previous load from it, cyclically.
|
||||
fn check_stale_loads(instrs: &[Instr]) -> Result<(), Reject> {
|
||||
// `pending[r]`: a load has read r and nothing has written r since. Two passes over the list so the second
|
||||
// pass sees the state carried over the iteration boundary.
|
||||
let mut pending = [false; 8];
|
||||
for _pass in 0..2 {
|
||||
for (k, ins) in instrs.iter().enumerate() {
|
||||
if ins.op.is_load() && pending[ins.src as usize] {
|
||||
return Err(Reject::StaleLoadSource { instr: k as u8, reg: ins.src });
|
||||
}
|
||||
pending[ins.dst as usize] = false;
|
||||
if ins.op.is_load() {
|
||||
pending[ins.src as usize] = true;
|
||||
}
|
||||
}
|
||||
}
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// Part (b): every register has an injecting write.
|
||||
fn check_injecting_writes(instrs: &[Instr]) -> Result<(), Reject> {
|
||||
let mut injected = [false; 8];
|
||||
for ins in instrs {
|
||||
if ins.op.injects() {
|
||||
injected[ins.dst as usize] = true;
|
||||
}
|
||||
}
|
||||
for (reg, ok) in injected.iter().enumerate() {
|
||||
if !ok {
|
||||
return Err(Reject::NoInjectingWrite { reg: reg as u8 });
|
||||
}
|
||||
}
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// Parts (a) and (b).
|
||||
pub fn check_static(p: &Program) -> Result<(), Reject> {
|
||||
if p.instrs.len() != INSTR_COUNT {
|
||||
panic!("acceptance needs a {INSTR_COUNT}-instruction program");
|
||||
}
|
||||
check_stale_loads(&p.instrs)?;
|
||||
check_injecting_writes(&p.instrs)
|
||||
}
|
||||
|
||||
/// The 64 base nonces of the dynamic test for seed words `seed`.
|
||||
pub fn accept_base_nonces(seed: &[u32; 8]) -> [u32; ACCEPT_UNITS] {
|
||||
let mut b = Vec::with_capacity(ACCEPT_TAG.len() + 32);
|
||||
b.extend_from_slice(ACCEPT_TAG);
|
||||
for w in seed {
|
||||
b.extend_from_slice(&w.to_le_bytes());
|
||||
}
|
||||
let mut rng = SplitMix64::new(fnv1a64(&b));
|
||||
let mut out = [0u32; ACCEPT_UNITS];
|
||||
for o in out.iter_mut() {
|
||||
*o = (rng.next() as u32) & !31;
|
||||
}
|
||||
out
|
||||
}
|
||||
|
||||
#[inline(always)]
|
||||
fn mulhi32(a: u32, b: u32) -> u32 {
|
||||
((a as u64 * b as u64) >> 32) as u32
|
||||
}
|
||||
|
||||
/// Accumulators of the dynamic test over the 64 units.
|
||||
struct Acc {
|
||||
and_acc: [u32; 8],
|
||||
or_acc: [u32; 8],
|
||||
saturated: u32,
|
||||
bit_ones: [u32; 64],
|
||||
distinct_sum: u64,
|
||||
}
|
||||
|
||||
/// One unit of the dynamic test: the interpreter of `verify.rs` with the closed-form dataset, instrumented.
|
||||
/// Returns the first lane-constant load site, if any.
|
||||
fn run_unit(p: &Program, unit: usize, base: u32, acc: &mut Acc, lane_addrs: &mut [u32]) -> Result<(), Reject> {
|
||||
let seed = &p.seed;
|
||||
let mask: u32 = (1u32 << ACCEPT_DATASET_LOG2) - 1;
|
||||
let (d0, d1) = (seed[0], seed[1]);
|
||||
let loads = p.loads_per_hash();
|
||||
let mut r = [[0u32; LANES]; 8];
|
||||
for lane in 0..LANES {
|
||||
let nonce = base.wrapping_add(lane as u32);
|
||||
for i in 0..8 {
|
||||
let mut x = nonce ^ seed[i];
|
||||
x = x.wrapping_add(0x9e3779b9u32.wrapping_mul(i as u32 + 1));
|
||||
x = splitmix32(x);
|
||||
r[i][lane] = x ^ seed[(i + 1) & 7];
|
||||
}
|
||||
}
|
||||
let mut idx = [0u32; LANES];
|
||||
let mut nload = 0usize;
|
||||
for it in 0..ITERATIONS {
|
||||
let sel = r[0];
|
||||
for (k, ins) in p.instrs.iter().enumerate() {
|
||||
let d = ins.dst as usize;
|
||||
let a = ins.src as usize;
|
||||
match ins.op {
|
||||
Op::Add => {
|
||||
let (imm, imm2, bit) = (ins.imm, ins.imm2, ins.bit as u32);
|
||||
let src = r[a];
|
||||
for lane in 0..LANES {
|
||||
let s = (sel[lane] >> bit) & 1;
|
||||
let c = if s != 0 { imm2 } else { imm };
|
||||
r[d][lane] = r[d][lane].wrapping_add(src[lane]).wrapping_add(c);
|
||||
}
|
||||
}
|
||||
Op::Sub => {
|
||||
let src = r[a];
|
||||
for lane in 0..LANES {
|
||||
r[d][lane] = r[d][lane].wrapping_sub(src[lane]);
|
||||
}
|
||||
}
|
||||
Op::Mul => {
|
||||
let src = r[a];
|
||||
for lane in 0..LANES {
|
||||
r[d][lane] = r[d][lane].wrapping_mul(src[lane]);
|
||||
}
|
||||
}
|
||||
Op::MulHi => {
|
||||
let src = r[a];
|
||||
for lane in 0..LANES {
|
||||
r[d][lane] = mulhi32(r[d][lane], src[lane]);
|
||||
}
|
||||
}
|
||||
Op::Xor => {
|
||||
let src = r[a];
|
||||
for lane in 0..LANES {
|
||||
r[d][lane] ^= src[lane];
|
||||
}
|
||||
}
|
||||
Op::Or => {
|
||||
let src = r[a];
|
||||
for lane in 0..LANES {
|
||||
r[d][lane] |= src[lane];
|
||||
}
|
||||
}
|
||||
Op::Rotl => {
|
||||
let n = ins.rot;
|
||||
for lane in 0..LANES {
|
||||
r[d][lane] = r[d][lane].rotate_left(n);
|
||||
}
|
||||
}
|
||||
Op::Rotr => {
|
||||
let src = r[a];
|
||||
for lane in 0..LANES {
|
||||
r[d][lane] = r[d][lane].rotate_right(src[lane] & 31);
|
||||
}
|
||||
}
|
||||
Op::Mad => {
|
||||
let src = r[a];
|
||||
let src2 = r[ins.src2 as usize];
|
||||
for lane in 0..LANES {
|
||||
r[d][lane] = src[lane].wrapping_mul(src2[lane]).wrapping_add(r[d][lane]);
|
||||
}
|
||||
}
|
||||
Op::Shfl => {
|
||||
let src = r[a];
|
||||
let m = ins.mask as usize;
|
||||
for lane in 0..LANES {
|
||||
r[d][lane] ^= src[lane ^ m];
|
||||
}
|
||||
}
|
||||
Op::Load => {
|
||||
for lane in 0..LANES {
|
||||
idx[lane] = r[a][lane] & mask;
|
||||
}
|
||||
if idx.iter().all(|&x| x == idx[0]) {
|
||||
return Err(Reject::LaneConstantSite { iteration: it as u8, instr: k as u8, unit: unit as u8 });
|
||||
}
|
||||
for lane in 0..LANES {
|
||||
r[d][lane] ^= dataset_elem(idx[lane], d0, d1);
|
||||
lane_addrs[lane * loads + nload] = idx[lane];
|
||||
}
|
||||
nload += 1;
|
||||
}
|
||||
Op::WLoad => {
|
||||
let b = (r[a][0] & mask) & !31;
|
||||
for lane in 0..LANES {
|
||||
idx[lane] = b + lane as u32;
|
||||
r[d][lane] ^= dataset_elem(idx[lane], d0, d1);
|
||||
lane_addrs[lane * loads + nload] = idx[lane];
|
||||
}
|
||||
nload += 1;
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
for i in 0..8 {
|
||||
for lane in 0..LANES {
|
||||
let v = r[i][lane];
|
||||
acc.and_acc[i] &= v;
|
||||
acc.or_acc[i] |= v;
|
||||
acc.saturated += (v == 0 || v == u32::MAX) as u32;
|
||||
}
|
||||
}
|
||||
for lane in 0..LANES {
|
||||
let lo = r[0][lane] ^ r[1][lane].rotate_left(7) ^ r[2][lane].rotate_left(14) ^ r[3][lane].rotate_left(21);
|
||||
let hi = r[4][lane] ^ r[5][lane].rotate_left(9) ^ r[6][lane].rotate_left(18) ^ r[7][lane].rotate_left(27);
|
||||
let h = ((hi as u64) << 32) | lo as u64;
|
||||
for j in 0..64 {
|
||||
acc.bit_ones[j] += ((h >> j) & 1) as u32;
|
||||
}
|
||||
let sl = &mut lane_addrs[lane * loads..(lane + 1) * loads];
|
||||
sl.sort_unstable();
|
||||
let mut distinct = 0u64;
|
||||
for k in 0..loads {
|
||||
if k == 0 || sl[k] != sl[k - 1] {
|
||||
distinct += 1;
|
||||
}
|
||||
}
|
||||
acc.distinct_sum += distinct;
|
||||
}
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// Part (c).
|
||||
pub fn check_dynamic(p: &Program) -> Result<AcceptReport, Reject> {
|
||||
let loads = p.loads_per_hash();
|
||||
let mut acc = Acc { and_acc: [u32::MAX; 8], or_acc: [0; 8], saturated: 0, bit_ones: [0; 64], distinct_sum: 0 };
|
||||
let mut lane_addrs = vec![0u32; LANES * loads];
|
||||
for (unit, &base) in accept_base_nonces(&p.seed).iter().enumerate() {
|
||||
run_unit(p, unit, base, &mut acc, &mut lane_addrs)?;
|
||||
}
|
||||
for reg in 0..8 {
|
||||
let bits = (acc.and_acc[reg] | !acc.or_acc[reg]).count_ones();
|
||||
if bits != 0 {
|
||||
return Err(Reject::ConstantBit { reg: reg as u8, bits: bits as u8 });
|
||||
}
|
||||
}
|
||||
if acc.saturated >= MAX_SATURATED {
|
||||
return Err(Reject::Saturated { count: acc.saturated });
|
||||
}
|
||||
let half = (ACCEPT_HASHES / 2) as u32;
|
||||
let mut bias_max = 0u32;
|
||||
for (bit, &ones) in acc.bit_ones.iter().enumerate() {
|
||||
let d = ones.abs_diff(half);
|
||||
if d > BIAS_TOLERANCE {
|
||||
return Err(Reject::OutputBias { bit: bit as u8, ones });
|
||||
}
|
||||
bias_max = bias_max.max(d);
|
||||
}
|
||||
if acc.distinct_sum <= MIN_DISTINCT_SUM {
|
||||
return Err(Reject::DistinctAddresses { sum: acc.distinct_sum });
|
||||
}
|
||||
Ok(AcceptReport { distinct_sum: acc.distinct_sum, saturated: acc.saturated, bias_max })
|
||||
}
|
||||
|
||||
/// The whole rule: (a), (b), then (c).
|
||||
pub fn check(p: &Program) -> Result<AcceptReport, Reject> {
|
||||
check_static(p)?;
|
||||
check_dynamic(p)
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
use crate::generator::{candidate, generate, GeneratorConfig, generate_v1};
|
||||
use crate::verify::{DatasetMode, DatasetSource};
|
||||
|
||||
/// The instrumented interpreter agrees with `verify.rs` on the closed-form dataset keyed by the seed words.
|
||||
#[test]
|
||||
fn instrumented_interpreter_matches_verify() {
|
||||
for i in 0..20u32 {
|
||||
let s = format!("igneum-accept-test/{i}");
|
||||
let p = candidate(&s, s.as_bytes(), 0);
|
||||
let ds = DatasetSource::from_key(p.seed, DatasetMode::ClosedForm, ACCEPT_DATASET_LOG2);
|
||||
let bases = accept_base_nonces(&p.seed);
|
||||
let mut acc = Acc { and_acc: [u32::MAX; 8], or_acc: [0; 8], saturated: 0, bit_ones: [0; 64], distinct_sum: 0 };
|
||||
let mut la = vec![0u32; LANES * p.loads_per_hash()];
|
||||
let mut ones = [0u32; 64];
|
||||
let mut any = false;
|
||||
for (u, &b) in bases.iter().enumerate() {
|
||||
if run_unit(&p, u, b, &mut acc, &mut la).is_err() {
|
||||
continue;
|
||||
}
|
||||
any = true;
|
||||
let w = crate::verify::hash_warp(&p, b, &ds);
|
||||
for h in w {
|
||||
for j in 0..64 {
|
||||
ones[j] += ((h >> j) & 1) as u32;
|
||||
}
|
||||
}
|
||||
}
|
||||
if any {
|
||||
assert_eq!(acc.bit_ones, ones, "bit counts of {s} match the reference interpreter's hashes");
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn base_nonces_are_aligned_and_seed_dependent() {
|
||||
let a = accept_base_nonces(&[1, 2, 3, 4, 5, 6, 7, 8]);
|
||||
let b = accept_base_nonces(&[1, 2, 3, 4, 5, 6, 7, 9]);
|
||||
assert!(a.iter().all(|x| x & 31 == 0));
|
||||
assert_ne!(a, b);
|
||||
assert_eq!(a, accept_base_nonces(&[1, 2, 3, 4, 5, 6, 7, 8]));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn stale_load_detection_is_cyclic() {
|
||||
let mut p = candidate("igneum-genesis", b"igneum-genesis", 0);
|
||||
assert!(check_stale_loads(&p.instrs).is_ok(), "an accepted candidate has no stale load");
|
||||
// Make the last instruction a load from r3 and the first a load from r3 with no write between (wrap).
|
||||
let (first, last) = (0usize, INSTR_COUNT - 1);
|
||||
p.instrs[last].op = Op::Load;
|
||||
p.instrs[last].src = 3;
|
||||
p.instrs[last].dst = 4;
|
||||
p.instrs[first].op = Op::Load;
|
||||
p.instrs[first].src = 3;
|
||||
p.instrs[first].dst = 5;
|
||||
assert_eq!(check_stale_loads(&p.instrs), Err(Reject::StaleLoadSource { instr: 0, reg: 3 }));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn injecting_write_detection() {
|
||||
let mut p = candidate("igneum-genesis", b"igneum-genesis", 0);
|
||||
for ins in p.instrs.iter_mut() {
|
||||
if ins.dst == 6 && ins.op.injects() {
|
||||
ins.op = Op::Rotl;
|
||||
}
|
||||
}
|
||||
assert_eq!(check_injecting_writes(&p.instrs), Err(Reject::NoInjectingWrite { reg: 6 }));
|
||||
}
|
||||
|
||||
/// The census's measured rates: about 5 percent of candidates rejected, 128 distinct loads for the rest.
|
||||
#[test]
|
||||
fn rejection_rate_and_distinct_loads_on_a_sample() {
|
||||
let mut rejected = 0;
|
||||
let mut dsum = 0.0;
|
||||
let mut accepted = 0;
|
||||
for i in 0..400u32 {
|
||||
let s = format!("igneum-census-2026-10-03/{i}");
|
||||
match check(&candidate(&s, s.as_bytes(), 0)) {
|
||||
Ok(r) => {
|
||||
accepted += 1;
|
||||
dsum += r.distinct_mean();
|
||||
assert!(r.distinct_mean() > 120.0);
|
||||
}
|
||||
Err(_) => rejected += 1,
|
||||
}
|
||||
}
|
||||
assert!(rejected < 50, "{rejected} of 400 rejected");
|
||||
assert!(dsum / accepted as f64 > 127.0, "mean distinct {}", dsum / accepted as f64);
|
||||
}
|
||||
|
||||
/// The retired generator fails the rule on nearly every program (the census: 95 percent).
|
||||
#[test]
|
||||
fn v1_programs_are_mostly_rejected() {
|
||||
let mut rejected = 0;
|
||||
for i in 0..100u32 {
|
||||
let s = format!("igneum-census-2026-10-03/{i}");
|
||||
if check(&generate_v1(&s, &GeneratorConfig::default())).is_err() {
|
||||
rejected += 1;
|
||||
}
|
||||
}
|
||||
assert!(rejected > 80, "{rejected} of 100 rejected");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn generated_programs_pass() {
|
||||
for s in ["igneum-genesis", "igneum-hourly", "igneum-second-seed"] {
|
||||
assert!(check(&generate(s)).is_ok());
|
||||
}
|
||||
}
|
||||
}
|
||||
|
|
@ -18,7 +18,8 @@
|
|||
//! | Epoch seed bytes | the 32 bytes of the epoch block hash (devnet v0: the last selected-chain block below the epoch's start DAA score, genesis for epoch 0) | Section 1.12: `S_e = seed_words_from_bytes(program_seed_e)`; the VDF output replaces the block hash later without touching this crate |
|
||||
//!
|
||||
//! The packs' vectors (init words equal to the program seed) stay the conformance vectors for the generator,
|
||||
//! interpreter and dataset. The bound vectors are in `README.md` and in the tests below.
|
||||
//! interpreter and dataset. The bound vectors are in `README.md` and in the tests below (re-cut for generator
|
||||
//! version 2 on 4 October 2026).
|
||||
|
||||
use crate::generator::LANES;
|
||||
use crate::seed::seed_words_from_bytes;
|
||||
|
|
@ -157,7 +158,10 @@ mod tests {
|
|||
assert_eq!(&b[..13], b"igneum-block/");
|
||||
assert_eq!(&b[13..45], &prehash_b());
|
||||
assert_eq!(&b[45..49], &[0x02, 0x01, 0x00, 0x00]);
|
||||
assert_eq!(block_init_words(&prehash_b(), 0x0000_0102_0000_0007), block_init_words(&prehash_b(), 0x0000_0102_ffff_ffff));
|
||||
assert_eq!(
|
||||
block_init_words(&prehash_b(), 0x0000_0102_0000_0007),
|
||||
block_init_words(&prehash_b(), 0x0000_0102_ffff_ffff)
|
||||
);
|
||||
assert_ne!(block_init_words(&prehash_b(), 0), block_init_words(&prehash_b(), 1 << 32));
|
||||
}
|
||||
|
||||
|
|
@ -200,21 +204,22 @@ mod tests {
|
|||
assert_eq!(e.pow_bound(&a, 7), pow256_from_lane(t));
|
||||
}
|
||||
|
||||
/// The 8 bound vectors printed in README.md (seed igneum-genesis, day 2026-10-03, memory-hard, 2^28 words).
|
||||
/// The 8 bound vectors printed in README.md (seed igneum-genesis, day 2026-10-03, memory-hard, 2^28 words,
|
||||
/// generator v2 since 4 October 2026).
|
||||
#[test]
|
||||
fn bound_vectors() {
|
||||
let e = epoch();
|
||||
let a = prehash_a();
|
||||
let b = prehash_b();
|
||||
let cases: [(&[u8; 32], u64, u64); 8] = [
|
||||
(&a, 0, 0x2c619692d823263b),
|
||||
(&a, 1, 0x55d21ed545735ce8),
|
||||
(&a, 31, 0x862eebe7fbda564e),
|
||||
(&a, 4096, 0x37aadc51f95725df),
|
||||
(&a, 1 << 32, 0xb62e28b8a90e554f),
|
||||
(&b, 0, 0x9b2437118e087833),
|
||||
(&b, (1 << 32) | 5, 0x714ae31e369e0446),
|
||||
(&b, u64::MAX, 0x4ca4f84079025a13),
|
||||
(&a, 0, 0x746c567b090acf6a),
|
||||
(&a, 1, 0x45a619f860880c73),
|
||||
(&a, 31, 0x9aa495e43dedbfe6),
|
||||
(&a, 4096, 0x2e6ffd7624d3cba2),
|
||||
(&a, 1 << 32, 0x38a5cea1fb01431a),
|
||||
(&b, 0, 0x2a79c5e4797bf6aa),
|
||||
(&b, (1 << 32) | 5, 0xa243e0c61aa1b82e),
|
||||
(&b, u64::MAX, 0x9c7bbfbd064fe1a4),
|
||||
];
|
||||
for (h, nonce, want) in cases {
|
||||
assert_eq!(e.hash_bound(h, nonce), want, "H {} nonce {nonce}", hex(h));
|
||||
|
|
|
|||
|
|
@ -1,13 +1,15 @@
|
|||
//! Kernel source emitters. Each function here writes the same bytes as its namesake in
|
||||
//! `proto-metal/main.swift` (`generateMSL`, `memhardMSL`, `emitMemhardCore`, `generateCUDA`, `generateOpenCL`,
|
||||
//! `generateProgramHeader`, `generateMemhardHeader`, `generateVectorsHeader`, `generateProgramJSON`,
|
||||
//! `generateVectorsJSON`). The pack tests diff them against `proto-cuda/packs/*`.
|
||||
//! Kernel source emitters. Since 4 October 2026 (generator version 2) this crate is the source of every pack in
|
||||
//! `proto-cuda/packs/`; the pack tests diff the emitters against the checked-in files. Each function started as a
|
||||
//! byte-for-byte twin of its namesake in `proto-metal/main.swift` (`generateMSL`, `memhardMSL`, `emitMemhardCore`,
|
||||
//! `generateCUDA`, `generateOpenCL`, `generateProgramHeader`, `generateMemhardHeader`, `generateVectorsHeader`,
|
||||
//! `generateProgramJSON`, `generateVectorsJSON`); the kernel text is unchanged by version 2, and `program.json` and
|
||||
//! `program.h` carry the generator version, the attempt and the program id so no version 1 pack can be mistaken
|
||||
//! for a current one.
|
||||
//!
|
||||
//! One deliberate difference: `program_json` writes the cache line mask inside the `"item"` string as a bare
|
||||
//! `0x003fffff`. The Swift writes it quoted (`jhex`), which makes `igneum-genesis-mh/program.json` invalid
|
||||
//! JSON. A node must hand miners valid JSON, so the Rust side does not reproduce that defect.
|
||||
//! One deliberate difference from the Swift: `program_json` writes the cache line mask inside the `"item"` string
|
||||
//! as a bare `0x003fffff`. The Swift writes it quoted (`jhex`), which is not valid JSON.
|
||||
|
||||
use crate::generator::{Op, Program, INSTR_COUNT, ITERATIONS};
|
||||
use crate::generator::{Op, Program, GENERATOR_VERSION, INSTR_COUNT, ITERATIONS, LOAD_SLOTS};
|
||||
use crate::memhard::{
|
||||
MixParams, CACHE_LINES_PER_SEGMENT, CACHE_LINE_MASK, CACHE_LOG2_WORDS, CACHE_SEGMENTS, CACHE_SEGMENT_LOG2_LINES,
|
||||
CACHE_TAG, CACHE_WORDS, CHACHA_ROUNDS, CHACHA_SIGMA, ITEM_ROUNDS,
|
||||
|
|
@ -334,7 +336,11 @@ fn metal_program_impl(p: &Program, dataset_log2: u32, source: LoadSource, bound:
|
|||
pub const METAL_FILL: &str = "#include <metal_stdlib>\nusing namespace metal;\ninline uint ds_elem(uint i, uint d0, uint d1) {\n uint x = i ^ d0;\n x *= 0x9E3779B1u; x ^= x >> 15;\n x += d1;\n x *= 0x85EBCA77u; x ^= x >> 13;\n x *= 0xC2B2AE3Du; x ^= x >> 16;\n return x;\n}\nkernel void igneum_fill(device uint* dataset [[buffer(0)]],\n constant uint2& day [[buffer(1)]],\n uint gid [[thread_position_in_grid]]) {\n dataset[gid] = ds_elem(gid, day.x, day.y);\n}";
|
||||
|
||||
fn generated_by(seed: &str) -> String {
|
||||
format!("// Generated by proto-metal/igneum-bench --export-pack for seed \"{seed}\". Do not edit by hand.\n")
|
||||
format!("// Generated by igneum-pow export (generator v{GENERATOR_VERSION}) for seed \"{seed}\". Do not edit by hand.\n")
|
||||
}
|
||||
|
||||
fn hex_bytes(b: &[u8]) -> String {
|
||||
b.iter().map(|x| format!("{x:02x}")).collect()
|
||||
}
|
||||
|
||||
fn init_line(p: &Program, u: &str, i: usize) -> String {
|
||||
|
|
@ -516,11 +522,15 @@ pub fn cuda_kernel(p: &Program, memhard: Option<&MixParams>) -> String {
|
|||
pub fn cuda_kernel_bound(p: &Program, memhard: Option<&MixParams>) -> String {
|
||||
let mut s = String::with_capacity(9000);
|
||||
s.push_str(&generated_by(&p.seed_string));
|
||||
s.push_str("// Header-bound twin of igneum_hash in kernel.cu: the init words come from a kernel argument, not SEEDW.\n");
|
||||
s.push_str(
|
||||
"// Header-bound twin of igneum_hash in kernel.cu: the init words come from a kernel argument, not SEEDW.\n",
|
||||
);
|
||||
s.push_str("// Host declarations (also in program_bound.h if present):\n");
|
||||
s.push_str("// struct IgneumInitWords { uint32_t w[8]; };\n");
|
||||
s.push_str("// cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,\n");
|
||||
s.push_str("// IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps);\n");
|
||||
s.push_str(
|
||||
"// IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps);\n",
|
||||
);
|
||||
s.push_str("// cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);\n");
|
||||
s.push_str("#include <cuda_runtime.h>\n");
|
||||
s.push_str("#include <cstdint>\n");
|
||||
|
|
@ -560,7 +570,9 @@ pub fn cuda_kernel_bound(p: &Program, memhard: Option<&MixParams>) -> String {
|
|||
s.push_str(" out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;\n");
|
||||
s.push_str("}\n");
|
||||
s.push('\n');
|
||||
s.push_str("cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,\n");
|
||||
s.push_str(
|
||||
"cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,\n",
|
||||
);
|
||||
s.push_str(" IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps) {\n");
|
||||
s.push_str(" if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;\n");
|
||||
s.push_str(" uint32_t block = 32u * blockWarps;\n");
|
||||
|
|
@ -616,7 +628,9 @@ fn opencl_instr_lines(p: &Program) -> String {
|
|||
pub fn opencl_kernel_bound(p: &Program, memhard: Option<&MixParams>) -> String {
|
||||
let mut s = opencl_kernel(p, memhard);
|
||||
s.push('\n');
|
||||
s.push_str("// Header-bound variant (bind.rs): the init words come from initw, not SEEDW. Same body as igneum_hash.\n");
|
||||
s.push_str(
|
||||
"// Header-bound variant (bind.rs): the init words come from initw, not SEEDW. Same body as igneum_hash.\n",
|
||||
);
|
||||
s.push_str("IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw) {\n");
|
||||
s.push_str(" uint gid = (uint)get_global_id(0);\n");
|
||||
s.push_str(" uint lid = (uint)get_local_id(0);\n");
|
||||
|
|
@ -784,13 +798,10 @@ fn mask_for(dataset_log2: u32) -> u32 {
|
|||
const STDINT_BLOCK: &str = "#ifdef __cplusplus\n#include <cstdint>\n#else\n#include <stdint.h>\n#endif\n";
|
||||
|
||||
/// program.h (`generateProgramHeader`).
|
||||
pub fn program_header(
|
||||
p: &Program,
|
||||
day: &str,
|
||||
key: &[u32; 8],
|
||||
dataset_log2: u32,
|
||||
memhard: Option<&MixParams>,
|
||||
) -> String {
|
||||
pub fn program_header(p: &Program, day: &str, ds: &DatasetSource) -> String {
|
||||
let key = &ds.key;
|
||||
let dataset_log2 = ds.log2_words;
|
||||
let memhard = ds.memhard().map(|m| &m.params);
|
||||
let mask = mask_for(dataset_log2);
|
||||
let mut s = String::with_capacity(2600);
|
||||
s.push_str(&generated_by(&p.seed_string));
|
||||
|
|
@ -801,7 +812,12 @@ pub fn program_header(
|
|||
s.push_str("#ifndef IGNEUM_NO_CUDA\n#include <cuda_runtime.h>\n#endif\n");
|
||||
s.push('\n');
|
||||
s.push_str(&format!("#define IGNEUM_SEED_STRING {}\n", jstr(&p.seed_string)));
|
||||
s.push_str(&format!("#define IGNEUM_SEED_BYTES_HEX {}\n", jstr(&hex_bytes(&p.seed_bytes))));
|
||||
s.push_str(&format!("#define IGNEUM_GENERATOR {}\n", p.generator));
|
||||
s.push_str(&format!("#define IGNEUM_PROGRAM_ATTEMPT {}\n", p.attempt));
|
||||
s.push_str(&format!("#define IGNEUM_PROGRAM_ID {}\n", hex64(p.program_id())));
|
||||
s.push_str(&format!("#define IGNEUM_DAY_STRING {}\n", jstr(day)));
|
||||
s.push_str(&format!("#define IGNEUM_DAY_BYTES_HEX {}\n", jstr(&hex_bytes(&ds.key_bytes))));
|
||||
s.push_str(&format!("#define IGNEUM_DAY0 {}\n", hex(key[0])));
|
||||
s.push_str(&format!("#define IGNEUM_DAY1 {}\n", hex(key[1])));
|
||||
s.push_str(&format!("#define IGNEUM_DATASET_LOG2 {dataset_log2}\n"));
|
||||
|
|
@ -962,18 +978,27 @@ pub fn vectors_header(
|
|||
}
|
||||
|
||||
/// program.json (`generateProgramJSON`). Valid JSON (see the module note about the `"item"` line).
|
||||
pub fn program_json(p: &Program, day: &str, key: &[u32; 8], dataset_log2: u32, memhard: Option<&MixParams>) -> String {
|
||||
pub fn program_json(p: &Program, day: &str, ds: &DatasetSource) -> String {
|
||||
let key = &ds.key;
|
||||
let dataset_log2 = ds.log2_words;
|
||||
let memhard = ds.memhard().map(|m| &m.params);
|
||||
let mask = mask_for(dataset_log2);
|
||||
let mut s = String::with_capacity(13000);
|
||||
let mut s = String::with_capacity(14000);
|
||||
s.push_str("{\n");
|
||||
s.push_str(" \"format\": \"igneum-program-pack-2\",\n");
|
||||
s.push_str(" \"format\": \"igneum-program-pack-3\",\n");
|
||||
s.push_str(&format!(" \"generator\": {},\n", p.generator));
|
||||
s.push_str(&format!(" \"attempt\": {},\n", p.attempt));
|
||||
s.push_str(&format!(" \"program_id\": {},\n", jhex64(p.program_id())));
|
||||
s.push_str(" \"program_id_derivation\": \"FNV-1a 64 over 'igneum-program/' || generator_le32 || seed_words as little-endian bytes || attempt_le32\",\n");
|
||||
s.push_str(&format!(
|
||||
" \"dataset_mode\": {},\n",
|
||||
jstr(if memhard.is_some() { "memory-hard" } else { "closed-form" })
|
||||
));
|
||||
s.push_str(&format!(" \"seed\": {},\n", jstr(&p.seed_string)));
|
||||
s.push_str(&format!(" \"seed_bytes\": {},\n", jstr(&hex_bytes(&p.seed_bytes))));
|
||||
s.push_str(&format!(" \"seed_words\": [{}],\n", join_jhex(&p.seed)));
|
||||
s.push_str(" \"seed_derivation\": \"FNV-1a 64 over UTF-8 of seed, basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32\",\n");
|
||||
s.push_str(" \"seed_derivation\": \"seed_words = FNV-1a 64 over seed_bytes (attempt 0) or seed_bytes || attempt_le32 (attempt k >= 1), basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32\",\n");
|
||||
s.push_str(&format!(" \"generator_rule\": \"version {GENERATOR_VERSION}: exactly {LOAD_SLOTS} load slots drawn first from instructions 1..63 (partial Fisher-Yates), the other 48 ops from the ten non-load weights (sum 75); a load's source is drawn from the registers other than dst written by an earlier instruction and not read by a load since; the candidate must pass the acceptance rule of spec 01 section 1.4.6 (static: no cyclically stale load source, every register has an injecting write; dynamic: 64 units on the seed-keyed closed-form dataset with no constant register bit, no lane-constant load site, under 164 saturated final values, every output bit within 136 of 1024, distinct addresses above 245760), else the next attempt of the seed is tried\",\n"));
|
||||
s.push_str(" \"lanes\": 32,\n");
|
||||
s.push_str(" \"registers\": 8,\n");
|
||||
s.push_str(&format!(" \"iterations\": {ITERATIONS},\n"));
|
||||
|
|
@ -1010,16 +1035,15 @@ pub fn program_json(p: &Program, day: &str, key: &[u32; 8], dataset_log2: u32, m
|
|||
s.push_str(&format!(" \"bytes\": {},\n", 1u64 << (dataset_log2 as u64 + 2)));
|
||||
s.push_str(&format!(" \"mask\": {},\n", jhex(mask)));
|
||||
s.push_str(&format!(" \"day\": {},\n", jstr(day)));
|
||||
s.push_str(&format!(" \"day_words_from\": {},\n", jstr(&format!("day/{day}"))));
|
||||
s.push_str(&format!(" \"day_bytes\": {},\n", jstr(&hex_bytes(&ds.key_bytes))));
|
||||
s.push_str(" \"day_words_from\": \"seed_words_from_bytes(day_bytes)\",\n");
|
||||
s.push_str(&format!(" \"d0\": {},\n", jhex(key[0])));
|
||||
s.push_str(&format!(" \"d1\": {},\n", jhex(key[1])));
|
||||
if let Some(mp) = memhard {
|
||||
s.push_str(" \"mode\": \"memory-hard\",\n");
|
||||
s.push_str(" \"spec\": \"proto-metal/MEMHARD.md\",\n");
|
||||
s.push_str(&format!(" \"key\": [{}],\n", join_jhex(&mp.key)));
|
||||
s.push_str(
|
||||
" \"key_derivation\": \"the 8 words of seedWords(\\\"day/\\\" + day); d0, d1 are key[0], key[1]\",\n",
|
||||
);
|
||||
s.push_str(" \"key_derivation\": \"the 8 words of seed_words_from_bytes(day_bytes); d0, d1 are key[0], key[1]\",\n");
|
||||
s.push_str(&format!(
|
||||
" \"cache\": {{\"log2_words\": {CACHE_LOG2_WORDS}, \"bytes\": {}, \"line_words\": 16, \"segment_lines\": {CACHE_LINES_PER_SEGMENT}, \"segments\": {CACHE_SEGMENTS}, \"block\": \"ChaCha{CHACHA_ROUNDS} core + feed-forward, rotations 16 12 8 7\", \"sigma\": [{}], \"tag\": [{}], \"chain\": \"in_j = prev_line ^ (sigma[0..3] || key[0..7] || seg || j || tag[0..1]); line_j = block(in_j); prev_0 = 0\"}},\n",
|
||||
CACHE_WORDS as u64 * 4,
|
||||
|
|
@ -1162,11 +1186,11 @@ pub fn export_pack(epoch: &Epoch, day: &str, source: &str) -> Pack {
|
|||
}
|
||||
let is_mh = memhard.is_some();
|
||||
let mut files = vec![
|
||||
("program.json".to_string(), program_json(p, day, &ds.key, ds.log2_words, memhard)),
|
||||
("program.json".to_string(), program_json(p, day, ds)),
|
||||
("vectors.json".to_string(), vectors_json(p, day, ds.log2_words, &bases, &outs, &v, mask, source, is_mh)),
|
||||
("kernel.cu".to_string(), cuda_kernel(p, memhard)),
|
||||
("kernel.cl".to_string(), opencl_kernel(p, memhard)),
|
||||
("program.h".to_string(), program_header(p, day, &ds.key, ds.log2_words, memhard)),
|
||||
("program.h".to_string(), program_header(p, day, ds)),
|
||||
("vectors.h".to_string(), vectors_header(p, &bases, &outs, &v, mask, source, is_mh)),
|
||||
("program.metal".to_string(), metal_program(p, ds.log2_words, LoadSource::Stored)),
|
||||
// Header-bound kernels (3 October 2026, bind.rs): new files, the seven above are unchanged.
|
||||
|
|
|
|||
|
|
@ -1,7 +1,26 @@
|
|||
//! The program generator: 64 integer instructions over 8 x u32 lane registers, run for 8 iterations.
|
||||
//! Draw order, weights and the lever rules are those of `generateProgram` in `proto-metal/main.swift`.
|
||||
//!
|
||||
//! Generator version 2 (adopted 4 October 2026 from `docs/analysis/weak-program-census-2026-10-03.md`,
|
||||
//! spec 01 sections 1.4.2, 1.4.3 and 1.4.6):
|
||||
//!
|
||||
//! * G1, exact load count: every program has exactly [`LOAD_SLOTS`] (16) `load` instructions, drawn first as a
|
||||
//! uniform 16-subset of instruction slots 1..63 by a partial Fisher-Yates over the program stream. The other
|
||||
//! 48 ops come from the ten non-load families with the weights of [`NONLOAD_WEIGHTS`] (sum 75).
|
||||
//! * G2, fresh source: on a load slot the source register is drawn from `E`, the registers other than `dst` that
|
||||
//! an earlier instruction of this program has written and that no later load has read. A load's address is then
|
||||
//! a value produced in this iteration that no earlier load used, so no load of a hash repeats an earlier load's
|
||||
//! address, across the iteration boundary included. If `E` is empty the source is drawn as on an ALU slot and
|
||||
//! the acceptance rule of [`crate::accept`] rejects the program.
|
||||
//! * R, acceptance: a candidate must pass [`crate::accept::check`]. A rejected candidate is replaced by the next
|
||||
//! attempt, `seed_words_from_bytes(program_seed || k_le32)` for `k = 1, 2, ...` (attempt 0 is the bare seed),
|
||||
//! so every node derives the same program from the same seed.
|
||||
//!
|
||||
//! The retired version 1 generator (op rolled per instruction with a 25 percent load weight, no acceptance) is
|
||||
//! kept as [`generate_v1`] for the census tool and the lever measurements of `proto-metal/MEMHARD.md`. Its
|
||||
//! programs are not the lottery hash and no pack or vector of version 1 is current.
|
||||
|
||||
use crate::seed::{program_rng, seed_words};
|
||||
use crate::accept::{check, Reject};
|
||||
use crate::seed::{fnv1a64, program_rng, seed_words_from_bytes};
|
||||
|
||||
/// Iterations of the instruction list per hash.
|
||||
pub const ITERATIONS: usize = 8;
|
||||
|
|
@ -9,6 +28,15 @@ pub const ITERATIONS: usize = 8;
|
|||
pub const INSTR_COUNT: usize = 64;
|
||||
/// Lanes per verification unit (one SIMD group / warp).
|
||||
pub const LANES: usize = 32;
|
||||
/// The generator version written into every pack and program id. Version 1 programs never mix with these.
|
||||
pub const GENERATOR_VERSION: u32 = 2;
|
||||
/// Load instructions per program under version 2 (G1): 128 loads per hash, 4,096 items per 32-lane unit.
|
||||
pub const LOAD_SLOTS: usize = 16;
|
||||
/// Attempts before an implementation may treat the seed as a consensus fault (spec 01 section 1.4.6). At the
|
||||
/// measured 5.14 percent rejection rate the chance of 32 consecutive rejections is below 2^-136.
|
||||
pub const MAX_ATTEMPTS: u32 = 32;
|
||||
/// Domain tag of the program id.
|
||||
pub const PROGRAM_ID_TAG: &[u8] = b"igneum-program/";
|
||||
|
||||
#[derive(Clone, Copy, Debug, PartialEq, Eq, Hash)]
|
||||
pub enum Op {
|
||||
|
|
@ -23,7 +51,7 @@ pub enum Op {
|
|||
Mad,
|
||||
Shfl,
|
||||
Load,
|
||||
/// Warp-coalesced load (lever b). Never emitted unless `wide_frac > 0`.
|
||||
/// Warp-coalesced load (lever b of the version 1 generator). Never emitted by version 2.
|
||||
WLoad,
|
||||
}
|
||||
|
||||
|
|
@ -63,6 +91,16 @@ impl Op {
|
|||
_ => return None,
|
||||
})
|
||||
}
|
||||
|
||||
/// An injecting op: bijective in `dst` and bringing another register (or the dataset) in. The acceptance
|
||||
/// rule's part (b) requires one such write per register.
|
||||
pub fn injects(self) -> bool {
|
||||
matches!(self, Op::Add | Op::Sub | Op::Xor | Op::Mad | Op::Shfl | Op::Load | Op::WLoad)
|
||||
}
|
||||
|
||||
pub fn is_load(self) -> bool {
|
||||
matches!(self, Op::Load | Op::WLoad)
|
||||
}
|
||||
}
|
||||
|
||||
/// One instruction. Every field is drawn for every instruction whether the op uses it or not, so the
|
||||
|
|
@ -90,14 +128,23 @@ pub struct Instr {
|
|||
|
||||
#[derive(Clone, Debug, PartialEq, Eq)]
|
||||
pub struct Program {
|
||||
/// A label for packs and logs: the seed string, or whatever the caller named a byte seed.
|
||||
pub seed_string: String,
|
||||
/// The program seed bytes (`program_seed` of spec 01 section 1.12): the UTF-8 of a string seed, the 32-byte
|
||||
/// epoch seed on the chain. Attempt `k` of this seed is `seed_words_from_bytes(seed_bytes || k_le32)`.
|
||||
pub seed_bytes: Vec<u8>,
|
||||
/// The seed words of this attempt (what the program stream and the register init use).
|
||||
pub seed: [u32; 8],
|
||||
/// Generator version, [`GENERATOR_VERSION`] for every current program; 1 for the retired generator.
|
||||
pub generator: u32,
|
||||
/// Attempt index: 0 for the bare seed, `k` for the k-th re-derivation after rejections.
|
||||
pub attempt: u32,
|
||||
pub instrs: Vec<Instr>,
|
||||
}
|
||||
|
||||
impl Program {
|
||||
pub fn loads_per_hash(&self) -> usize {
|
||||
self.instrs.iter().filter(|i| i.op == Op::Load || i.op == Op::WLoad).count() * ITERATIONS
|
||||
self.instrs.iter().filter(|i| i.op.is_load()).count() * ITERATIONS
|
||||
}
|
||||
pub fn wide_loads_per_hash(&self) -> usize {
|
||||
self.instrs.iter().filter(|i| i.op == Op::WLoad).count() * ITERATIONS
|
||||
|
|
@ -109,7 +156,7 @@ impl Program {
|
|||
pub fn items_per_warp(&self) -> usize {
|
||||
(self.loads_per_hash() - self.wide_loads_per_hash()) * 32 + self.wide_loads_per_hash() * 2
|
||||
}
|
||||
/// Op histogram, count descending then name ascending, as the Swift prints it.
|
||||
/// Op histogram, count descending then name ascending.
|
||||
pub fn histogram(&self) -> Vec<(&'static str, usize)> {
|
||||
let mut counts: Vec<(&'static str, usize)> = Vec::new();
|
||||
for i in &self.instrs {
|
||||
|
|
@ -122,13 +169,45 @@ impl Program {
|
|||
counts.sort_by(|a, b| b.1.cmp(&a.1).then_with(|| a.0.cmp(b.0)));
|
||||
counts
|
||||
}
|
||||
/// "load=13 xor=13 ..." as written into program.h.
|
||||
/// "load=16 add=8 ..." as written into program.h.
|
||||
pub fn op_mix(&self) -> String {
|
||||
self.histogram().iter().map(|(n, c)| format!("{n}={c}")).collect::<Vec<_>>().join(" ")
|
||||
}
|
||||
/// The program id: FNV-1a 64 over `"igneum-program/" || generator_le32 || seed words as little-endian bytes
|
||||
/// || attempt_le32`. Written into every pack so a version 1 program, or another attempt of the same seed,
|
||||
/// can never be mistaken for this one.
|
||||
pub fn program_id(&self) -> u64 {
|
||||
program_id(self.generator, &self.seed, self.attempt)
|
||||
}
|
||||
}
|
||||
|
||||
/// Weights sum to 100. Loads are 25 percent so the kernel leans on memory.
|
||||
pub fn program_id(generator: u32, seed: &[u32; 8], attempt: u32) -> u64 {
|
||||
let mut b = Vec::with_capacity(PROGRAM_ID_TAG.len() + 4 + 32 + 4);
|
||||
b.extend_from_slice(PROGRAM_ID_TAG);
|
||||
b.extend_from_slice(&generator.to_le_bytes());
|
||||
for w in seed {
|
||||
b.extend_from_slice(&w.to_le_bytes());
|
||||
}
|
||||
b.extend_from_slice(&attempt.to_le_bytes());
|
||||
fnv1a64(&b)
|
||||
}
|
||||
|
||||
/// Weights of the ten non-load families under version 2, in draw order. Sum 75. The load family has no
|
||||
/// weight: its count is fixed by [`LOAD_SLOTS`].
|
||||
pub const NONLOAD_WEIGHTS: [(Op, u64); 10] = [
|
||||
(Op::Add, 12),
|
||||
(Op::Xor, 10),
|
||||
(Op::Mul, 8),
|
||||
(Op::Mad, 8),
|
||||
(Op::Shfl, 8),
|
||||
(Op::Rotl, 7),
|
||||
(Op::Sub, 6),
|
||||
(Op::MulHi, 6),
|
||||
(Op::Rotr, 6),
|
||||
(Op::Or, 4),
|
||||
];
|
||||
|
||||
/// Version 1 weights (retired). Sum 100, load at 25 percent.
|
||||
pub const OP_WEIGHTS: [(Op, u64); 11] = [
|
||||
(Op::Load, 25),
|
||||
(Op::Add, 12),
|
||||
|
|
@ -143,7 +222,167 @@ pub const OP_WEIGHTS: [(Op, u64); 11] = [
|
|||
(Op::Or, 4),
|
||||
];
|
||||
|
||||
/// Generator levers (MEMHARD.md section 2.4). The defaults reproduce the original generator exactly.
|
||||
/// The seed words of attempt `k` of a program seed: `seed_words_from_bytes(seed_bytes)` for `k = 0`,
|
||||
/// `seed_words_from_bytes(seed_bytes || k_le32)` otherwise.
|
||||
pub fn attempt_words(seed_bytes: &[u8], attempt: u32) -> [u32; 8] {
|
||||
if attempt == 0 {
|
||||
return seed_words_from_bytes(seed_bytes);
|
||||
}
|
||||
let mut b = Vec::with_capacity(seed_bytes.len() + 4);
|
||||
b.extend_from_slice(seed_bytes);
|
||||
b.extend_from_slice(&attempt.to_le_bytes());
|
||||
seed_words_from_bytes(&b)
|
||||
}
|
||||
|
||||
/// One version 2 candidate from its seed words, before the acceptance rule. Spec 01 section 1.4.3: 16 slot draws,
|
||||
/// then nine draws per instruction, 592 per program.
|
||||
pub fn candidate_from_words(seed_string: &str, seed_bytes: &[u8], seed: [u32; 8], attempt: u32) -> Program {
|
||||
let mut rng = program_rng(&seed);
|
||||
// (1) Load slots: a uniform 16-subset of 1..63 by partial Fisher-Yates. Instruction 0 is never a load.
|
||||
let mut p: [u8; INSTR_COUNT - 1] = [0; INSTR_COUNT - 1];
|
||||
for (i, slot) in p.iter_mut().enumerate() {
|
||||
*slot = (i + 1) as u8;
|
||||
}
|
||||
for i in 0..LOAD_SLOTS {
|
||||
let j = i + rng.below((INSTR_COUNT - 1 - i) as u64) as usize;
|
||||
p.swap(i, j);
|
||||
}
|
||||
let mut is_load = [false; INSTR_COUNT];
|
||||
for &slot in &p[..LOAD_SLOTS] {
|
||||
is_load[slot as usize] = true;
|
||||
}
|
||||
// (2) The instructions. `fresh[r]`: r was written by an earlier instruction and no load has read it since.
|
||||
let mut fresh = [false; 8];
|
||||
let mut instrs = Vec::with_capacity(INSTR_COUNT);
|
||||
for k in 0..INSTR_COUNT {
|
||||
let mut roll = rng.below(75);
|
||||
let mut op = Op::Add;
|
||||
for &(o, w) in &NONLOAD_WEIGHTS {
|
||||
if roll < w {
|
||||
op = o;
|
||||
break;
|
||||
}
|
||||
roll -= w;
|
||||
}
|
||||
if is_load[k] {
|
||||
op = Op::Load;
|
||||
}
|
||||
let dst = rng.below(8);
|
||||
let src = if op == Op::Load {
|
||||
let mut eligible = [0u64; 8];
|
||||
let mut n = 0usize;
|
||||
for r in 0..8u64 {
|
||||
if r != dst && fresh[r as usize] {
|
||||
eligible[n] = r;
|
||||
n += 1;
|
||||
}
|
||||
}
|
||||
if n == 0 {
|
||||
let a = rng.below(7);
|
||||
if a >= dst {
|
||||
a + 1
|
||||
} else {
|
||||
a
|
||||
}
|
||||
} else {
|
||||
eligible[rng.below(n as u64) as usize]
|
||||
}
|
||||
} else {
|
||||
let a = rng.below(7);
|
||||
if a >= dst {
|
||||
a + 1
|
||||
} else {
|
||||
a
|
||||
}
|
||||
};
|
||||
let b = rng.below(8);
|
||||
let imm = rng.next() as u32;
|
||||
let imm2 = rng.next() as u32;
|
||||
let rot = 1 + rng.below(31) as u32;
|
||||
let bit = rng.below(32);
|
||||
let mask = 1u8 << rng.below(5);
|
||||
if op == Op::Load {
|
||||
fresh[src as usize] = false;
|
||||
}
|
||||
fresh[dst as usize] = true;
|
||||
instrs.push(Instr { op, dst: dst as u8, src: src as u8, src2: b as u8, imm, imm2, rot, bit: bit as u8, mask });
|
||||
}
|
||||
Program {
|
||||
seed_string: seed_string.to_string(),
|
||||
seed_bytes: seed_bytes.to_vec(),
|
||||
seed,
|
||||
generator: GENERATOR_VERSION,
|
||||
attempt,
|
||||
instrs,
|
||||
}
|
||||
}
|
||||
|
||||
/// Candidate `attempt` of a program seed, before the acceptance rule.
|
||||
pub fn candidate(seed_string: &str, seed_bytes: &[u8], attempt: u32) -> Program {
|
||||
candidate_from_words(seed_string, seed_bytes, attempt_words(seed_bytes, attempt), attempt)
|
||||
}
|
||||
|
||||
/// Why no program could be derived from a seed.
|
||||
#[derive(Clone, Debug, PartialEq, Eq)]
|
||||
pub struct Exhausted {
|
||||
pub seed_string: String,
|
||||
pub attempts: u32,
|
||||
pub last: Reject,
|
||||
}
|
||||
|
||||
impl std::fmt::Display for Exhausted {
|
||||
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
|
||||
write!(f, "seed {:?}: {} consecutive candidates rejected, last: {}", self.seed_string, self.attempts, self.last)
|
||||
}
|
||||
}
|
||||
|
||||
impl std::error::Error for Exhausted {}
|
||||
|
||||
/// The program of a seed: the first accepted candidate over attempts `0, 1, 2, ...`, at most [`MAX_ATTEMPTS`].
|
||||
/// This is what the chain calls (`Epoch::from_seed_bytes`) with the 32-byte epoch seed, and what the packs call
|
||||
/// with the UTF-8 of a seed string.
|
||||
pub fn try_generate_from_seed_bytes(seed_string: &str, seed_bytes: &[u8]) -> Result<Program, Exhausted> {
|
||||
let mut last = None;
|
||||
for attempt in 0..MAX_ATTEMPTS {
|
||||
let p = candidate(seed_string, seed_bytes, attempt);
|
||||
match check(&p) {
|
||||
Ok(_) => return Ok(p),
|
||||
Err(r) => last = Some(r),
|
||||
}
|
||||
}
|
||||
Err(Exhausted { seed_string: seed_string.to_string(), attempts: MAX_ATTEMPTS, last: last.unwrap() })
|
||||
}
|
||||
|
||||
/// [`try_generate_from_seed_bytes`], treating exhaustion as the consensus fault it is.
|
||||
pub fn generate_from_seed_bytes(seed_string: &str, seed_bytes: &[u8]) -> Program {
|
||||
try_generate_from_seed_bytes(seed_string, seed_bytes).unwrap_or_else(|e| panic!("{e}"))
|
||||
}
|
||||
|
||||
/// The program of a seed string (its UTF-8 bytes are the program seed).
|
||||
pub fn generate(seed_string: &str) -> Program {
|
||||
generate_from_seed_bytes(seed_string, seed_string.as_bytes())
|
||||
}
|
||||
|
||||
/// Every candidate of a seed up to and including the accepted one, with each rejection. For reports and tests.
|
||||
pub fn attempts(seed_string: &str, seed_bytes: &[u8]) -> Vec<(Program, Result<(), Reject>)> {
|
||||
let mut out = Vec::new();
|
||||
for attempt in 0..MAX_ATTEMPTS {
|
||||
let p = candidate(seed_string, seed_bytes, attempt);
|
||||
let verdict = check(&p).map(|_| ());
|
||||
let accepted = verdict.is_ok();
|
||||
out.push((p, verdict));
|
||||
if accepted {
|
||||
break;
|
||||
}
|
||||
}
|
||||
out
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------------------------------------
|
||||
// Version 1 (retired 4 October 2026)
|
||||
// ---------------------------------------------------------------------------------------------------------
|
||||
|
||||
/// Version 1 levers (`proto-metal/MEMHARD.md` section 2.4). The defaults reproduce the version 1 generator exactly.
|
||||
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
|
||||
pub struct GeneratorConfig {
|
||||
/// Percent weight of the load op.
|
||||
|
|
@ -185,19 +424,9 @@ impl GeneratorConfig {
|
|||
}
|
||||
}
|
||||
|
||||
/// The default generator for a seed string.
|
||||
pub fn generate(seed_string: &str) -> Program {
|
||||
generate_with(seed_string, &GeneratorConfig::default())
|
||||
}
|
||||
|
||||
/// The generator with levers. `generateProgram` in the Swift, draw for draw.
|
||||
pub fn generate_with(seed_string: &str, cfg: &GeneratorConfig) -> Program {
|
||||
let seed = seed_words(seed_string);
|
||||
generate_from_words(seed_string, seed, cfg)
|
||||
}
|
||||
|
||||
/// The generator from already-derived seed words (what the chain will call once the VDF output is in).
|
||||
pub fn generate_from_words(seed_string: &str, seed: [u32; 8], cfg: &GeneratorConfig) -> Program {
|
||||
/// The retired version 1 generator from seed words: op rolled per instruction against the 11-family table,
|
||||
/// load count free, no acceptance rule. `generateProgramV1` in the Swift, draw for draw.
|
||||
pub fn generate_v1_from_words(seed_string: &str, seed: [u32; 8], cfg: &GeneratorConfig) -> Program {
|
||||
let mut rng = program_rng(&seed);
|
||||
let weights = cfg.weights();
|
||||
let mut instrs = Vec::with_capacity(INSTR_COUNT);
|
||||
|
|
@ -228,7 +457,19 @@ pub fn generate_from_words(seed_string: &str, seed: [u32; 8], cfg: &GeneratorCon
|
|||
}
|
||||
instrs.push(Instr { op, dst: dst as u8, src: a as u8, src2: b as u8, imm, imm2, rot, bit: bit as u8, mask });
|
||||
}
|
||||
Program { seed_string: seed_string.to_string(), seed, instrs }
|
||||
Program {
|
||||
seed_string: seed_string.to_string(),
|
||||
seed_bytes: seed_string.as_bytes().to_vec(),
|
||||
seed,
|
||||
generator: 1,
|
||||
attempt: 0,
|
||||
instrs,
|
||||
}
|
||||
}
|
||||
|
||||
/// The retired version 1 generator for a seed string.
|
||||
pub fn generate_v1(seed_string: &str, cfg: &GeneratorConfig) -> Program {
|
||||
generate_v1_from_words(seed_string, seed_words_from_bytes(seed_string.as_bytes()), cfg)
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
|
|
@ -238,6 +479,8 @@ mod tests {
|
|||
#[test]
|
||||
fn default_weights_unchanged() {
|
||||
assert_eq!(GeneratorConfig::default().weights(), OP_WEIGHTS.to_vec());
|
||||
assert_eq!(NONLOAD_WEIGHTS.iter().map(|w| w.1).sum::<u64>(), 75);
|
||||
assert_eq!(&OP_WEIGHTS[1..], &NONLOAD_WEIGHTS[..]);
|
||||
}
|
||||
|
||||
#[test]
|
||||
|
|
@ -262,15 +505,73 @@ mod tests {
|
|||
}
|
||||
|
||||
#[test]
|
||||
fn genesis_shape() {
|
||||
let p = generate("igneum-genesis");
|
||||
assert_eq!(p.instrs.len(), 64);
|
||||
fn v1_genesis_shape() {
|
||||
let p = generate_v1("igneum-genesis", &GeneratorConfig::default());
|
||||
assert_eq!(p.generator, 1);
|
||||
assert_eq!(p.loads_per_hash(), 104);
|
||||
assert_eq!(p.op_mix(), "load=13 xor=13 sub=7 shfl=6 add=5 mulhi=5 mad=4 rotr=4 mul=3 rotl=3 or=1");
|
||||
for i in &p.instrs {
|
||||
assert_ne!(i.dst, i.src);
|
||||
assert!((1..=31).contains(&i.rot));
|
||||
assert!(i.mask.is_power_of_two() && i.mask <= 16);
|
||||
}
|
||||
|
||||
/// Every version 2 candidate has 16 loads, none at instruction 0, and honours the generator contract.
|
||||
#[test]
|
||||
fn v2_shape_and_contract() {
|
||||
for i in 0..200u32 {
|
||||
let s = format!("igneum-shape/{i}");
|
||||
let p = candidate(&s, s.as_bytes(), 0);
|
||||
assert_eq!(p.generator, GENERATOR_VERSION);
|
||||
assert_eq!(p.instrs.len(), INSTR_COUNT);
|
||||
assert_eq!(p.loads_per_hash(), 8 * LOAD_SLOTS);
|
||||
assert_ne!(p.instrs[0].op, Op::Load);
|
||||
assert!(!p.has_wide());
|
||||
for ins in &p.instrs {
|
||||
assert_ne!(ins.dst, ins.src);
|
||||
assert!((1..=31).contains(&ins.rot));
|
||||
assert!(ins.mask.is_power_of_two() && ins.mask <= 16);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// The fresh-source rule by construction: unless E was empty, no load reads a register that an earlier load
|
||||
/// read without a write in between, cyclically.
|
||||
#[test]
|
||||
fn v2_loads_are_fresh_unless_fallback() {
|
||||
let mut fallbacks = 0;
|
||||
for i in 0..500u32 {
|
||||
let s = format!("igneum-fresh/{i}");
|
||||
let p = candidate(&s, s.as_bytes(), 0);
|
||||
let stale = crate::accept::check_static(&p).err();
|
||||
if let Some(Reject::StaleLoadSource { .. }) = stale {
|
||||
fallbacks += 1;
|
||||
}
|
||||
}
|
||||
// The census measured about 1.6 percent of candidates with a static repeat.
|
||||
assert!(fallbacks < 30, "{fallbacks} of 500 candidates with a stale load");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn attempt_words_differ_and_are_stable() {
|
||||
let a0 = attempt_words(b"igneum-genesis", 0);
|
||||
assert_eq!(a0, seed_words_from_bytes(b"igneum-genesis"));
|
||||
let a1 = attempt_words(b"igneum-genesis", 1);
|
||||
assert_ne!(a0, a1);
|
||||
assert_eq!(a1, seed_words_from_bytes(b"igneum-genesis\x01\x00\x00\x00"));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn program_id_separates_versions_and_attempts() {
|
||||
let p = generate("igneum-genesis");
|
||||
let v1 = generate_v1("igneum-genesis", &GeneratorConfig::default());
|
||||
assert_ne!(p.program_id(), v1.program_id());
|
||||
assert_ne!(program_id(2, &p.seed, 0), program_id(2, &p.seed, 1));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn generate_returns_an_accepted_program() {
|
||||
let p = generate("igneum-genesis");
|
||||
assert!(check(&p).is_ok());
|
||||
assert_eq!(p.seed, attempt_words(b"igneum-genesis", p.attempt));
|
||||
let again = generate_from_seed_bytes("other label", b"igneum-genesis");
|
||||
assert_eq!(again.instrs, p.instrs);
|
||||
assert_eq!(again.attempt, p.attempt);
|
||||
}
|
||||
}
|
||||
|
|
|
|||
|
|
@ -3,7 +3,9 @@
|
|||
//! The crate has five parts, each mirroring one section of the prototype:
|
||||
//!
|
||||
//! * [`seed`]: the 32-byte seed words from a string (FNV-1a 64, four salts) and the SplitMix64 stream.
|
||||
//! * [`generator`]: the 64-instruction program drawn from a seed.
|
||||
//! * [`generator`]: the 64-instruction program drawn from a seed (version 2: 16 load slots, fresh sources).
|
||||
//! * [`accept`]: the acceptance rule every candidate program must pass; a rejected candidate is replaced by the
|
||||
//! next attempt of the same seed.
|
||||
//! * [`memhard`]: the 256 MiB ChaCha12 cache and the 8-round dataset item derivation (`proto-metal/MEMHARD.md`).
|
||||
//! * [`verify`]: the 32-lane warp interpreter that computes the 64-bit hash on the CPU, deriving dataset
|
||||
//! words on demand from the cache (or from the closed form, for the old packs).
|
||||
|
|
@ -19,6 +21,7 @@
|
|||
#![allow(clippy::needless_range_loop, clippy::should_implement_trait, clippy::large_enum_variant)]
|
||||
#![allow(clippy::manual_slice_size_calculation, clippy::too_many_arguments)]
|
||||
|
||||
pub mod accept;
|
||||
pub mod bind;
|
||||
pub mod emit;
|
||||
pub mod generator;
|
||||
|
|
@ -26,8 +29,9 @@ pub mod memhard;
|
|||
pub mod seed;
|
||||
pub mod verify;
|
||||
|
||||
pub use generator::{generate, Instr, Op, Program};
|
||||
pub use bind::{block_init_words, day_bytes, pow256_from_lane, target64_from_le256};
|
||||
pub use accept::{check as accept_program, AcceptReport, Reject};
|
||||
pub use generator::{generate, generate_from_seed_bytes, Instr, Op, Program, GENERATOR_VERSION};
|
||||
pub use memhard::{Cache, MemhardCpu, MixParams};
|
||||
pub use seed::{fnv1a64, seed_words, SplitMix64};
|
||||
pub use bind::{block_init_words, day_bytes, pow256_from_lane, target64_from_le256};
|
||||
pub use verify::{hash_warp, interpret_warp_init, verify_block, DatasetMode, DatasetSource, Epoch};
|
||||
|
|
|
|||
|
|
@ -5,6 +5,8 @@
|
|||
//! igneum-pow hash --seed <s> --nonce <n> [--day <d>] [--closed-form] [--dataset-log2 28]
|
||||
//! igneum-pow hash-bound --seed <s> --prehash <64 hex> --nonce <u64> [--day <d>] [--closed-form] [--dataset-log2 28]
|
||||
//! [--epoch-hex <64 hex> --day-hex <hex>] byte seeds instead of strings (Epoch::from_seed_bytes)
|
||||
//! igneum-pow accept --seed <s> [--epoch-hex <64 hex>] every candidate of the seed with its verdict (spec 01 section 1.4.6)
|
||||
//! igneum-pow show --seed <s> [--epoch-hex <64 hex>] the accepted program, one instruction per line
|
||||
|
||||
use igneum_pow::emit::export_pack;
|
||||
use igneum_pow::memhard::Cache;
|
||||
|
|
@ -32,7 +34,9 @@ fn usage() -> ! {
|
|||
\x20 bench [--warps 20] fill the cache, then time the CPU verifier per 32-lane warp\n\
|
||||
\x20 export --out <dir> write the program pack (kernel.cu, kernel.cl, program.metal, memhard.h, ...)\n\
|
||||
\x20 hash --nonce <n> print the 64-bit hash of one nonce (pack form, init words = seed words)\n\
|
||||
\x20 hash-bound --prehash <64 hex> --nonce <u64> print the header-bound hash (bind.rs) of one 64-bit nonce"
|
||||
\x20 hash-bound --prehash <64 hex> --nonce <u64> print the header-bound hash (bind.rs) of one 64-bit nonce\n\
|
||||
\x20 accept every candidate of the seed (or --epoch-hex) with its acceptance verdict\n\
|
||||
\x20 show the accepted program, one instruction per line"
|
||||
);
|
||||
std::process::exit(2)
|
||||
}
|
||||
|
|
@ -78,6 +82,8 @@ fn main() {
|
|||
match a.cmd.as_str() {
|
||||
"bench" => bench(&a, mode),
|
||||
"export" => export(&a, mode),
|
||||
"accept" => accept(&a),
|
||||
"show" => show(&a),
|
||||
"hash" => {
|
||||
let e = Epoch::new(&a.seed, &a.day, mode, a.dataset_log2);
|
||||
println!("{:016x}", e.hash(a.nonce as u32));
|
||||
|
|
@ -95,10 +101,7 @@ fn main() {
|
|||
_ => Epoch::new(&a.seed, &a.day, mode, a.dataset_log2),
|
||||
};
|
||||
let init = igneum_pow::bind::block_init_words(&prehash, a.nonce);
|
||||
println!(
|
||||
"init words {}",
|
||||
init.iter().map(|w| format!("{w:08x}")).collect::<Vec<_>>().join(" ")
|
||||
);
|
||||
println!("init words {}", init.iter().map(|w| format!("{w:08x}")).collect::<Vec<_>>().join(" "));
|
||||
println!("{:016x}", e.hash_bound(&prehash, a.nonce));
|
||||
}
|
||||
_ => usage(),
|
||||
|
|
@ -154,28 +157,31 @@ fn bench(a: &Args, mode: DatasetMode) {
|
|||
fn export(a: &Args, mode: DatasetMode) {
|
||||
let out = a.out.clone().unwrap_or_else(|| usage());
|
||||
let t0 = Instant::now();
|
||||
let e = match (&a.epoch_hex, &a.day_hex) {
|
||||
// --epoch-hex / --day-hex: the chain's byte seeds; the day label then names the day bytes
|
||||
let (e, day_label) = match (&a.epoch_hex, &a.day_hex) {
|
||||
(Some(eh), Some(dh)) => {
|
||||
let eb = igneum_pow::bind::unhex(eh).unwrap_or_else(|| usage());
|
||||
let db = igneum_pow::bind::unhex(dh).unwrap_or_else(|| usage());
|
||||
Epoch::from_seed_bytes(&eb, &db, &format!("igneum-epoch/{eh}/day/{dh}"))
|
||||
(Epoch::from_seed_bytes(&eb, &db, &format!("igneum-epoch/{eh}/day/{dh}")), format!("bytes:{dh}"))
|
||||
}
|
||||
_ => Epoch::new(&a.seed, &a.day, mode, a.dataset_log2),
|
||||
_ => (Epoch::new(&a.seed, &a.day, mode, a.dataset_log2), a.day.clone()),
|
||||
};
|
||||
let build_ms = t0.elapsed().as_secs_f64() * 1e3;
|
||||
println!("igneum-pow export {out}");
|
||||
println!(
|
||||
"seed \"{}\", day \"{}\", dataset 2^{} words ({}), loads/hash {}, wide loads/hash {}; epoch built in {build_ms:.1} ms",
|
||||
a.seed,
|
||||
a.day,
|
||||
a.dataset_log2,
|
||||
mode.name(),
|
||||
e.program.loads_per_hash(),
|
||||
e.program.wide_loads_per_hash()
|
||||
"seed \"{}\", day \"{}\", dataset 2^{} words ({}), generator v{} attempt {} program id {:016x}, loads/hash {}; epoch built in {build_ms:.1} ms",
|
||||
e.program.seed_string,
|
||||
day_label,
|
||||
e.dataset.log2_words,
|
||||
e.dataset.mode().name(),
|
||||
e.program.generator,
|
||||
e.program.attempt,
|
||||
e.program.program_id(),
|
||||
e.program.loads_per_hash()
|
||||
);
|
||||
println!("op mix: {}", e.program.op_mix());
|
||||
let source = format!("igneum-pow (Rust) CPU interpreter, {} dataset", mode.name());
|
||||
let pack = export_pack(&e, &a.day, &source);
|
||||
let source = format!("igneum-pow (Rust) CPU interpreter, generator v{}, {} dataset", e.program.generator, e.dataset.mode().name());
|
||||
let pack = export_pack(&e, &day_label, &source);
|
||||
let dir = std::path::Path::new(&out);
|
||||
if let Err(err) = pack.write_to(dir) {
|
||||
eprintln!("FAIL: write error {err}");
|
||||
|
|
@ -187,8 +193,68 @@ fn export(a: &Args, mode: DatasetMode) {
|
|||
for (i, b) in pack.bases.iter().enumerate() {
|
||||
println!("vector warp base {b}: lane0 {:016x} lane31 {:016x}", pack.outs[i][0], pack.outs[i][31]);
|
||||
}
|
||||
if mode == DatasetMode::MemoryHard {
|
||||
if e.dataset.mode() == DatasetMode::MemoryHard {
|
||||
println!("cache FNV-1a 64 {:016x}", pack.vectors.cache_fnv);
|
||||
}
|
||||
println!("OVERALL: PASS (pack written)");
|
||||
}
|
||||
|
||||
fn seed_bytes_of(a: &Args) -> (String, Vec<u8>) {
|
||||
match &a.epoch_hex {
|
||||
Some(eh) => (format!("igneum-epoch/{eh}"), igneum_pow::bind::unhex(eh).unwrap_or_else(|| usage())),
|
||||
None => (a.seed.clone(), a.seed.as_bytes().to_vec()),
|
||||
}
|
||||
}
|
||||
|
||||
fn accept(a: &Args) {
|
||||
let (label, bytes) = seed_bytes_of(a);
|
||||
let t0 = Instant::now();
|
||||
let tries = igneum_pow::generator::attempts(&label, &bytes);
|
||||
let ms = t0.elapsed().as_secs_f64() * 1e3;
|
||||
for (p, verdict) in &tries {
|
||||
match verdict {
|
||||
Ok(()) => {
|
||||
let r = igneum_pow::accept::check(p).unwrap();
|
||||
println!(
|
||||
"attempt {}: ACCEPTED program id {:016x}, op mix {}, distinct {:.3} per hash, saturated {}, bias max {}",
|
||||
p.attempt,
|
||||
p.program_id(),
|
||||
p.op_mix(),
|
||||
r.distinct_mean(),
|
||||
r.saturated,
|
||||
r.bias_max
|
||||
);
|
||||
}
|
||||
Err(r) => println!("attempt {}: rejected, {r}", p.attempt),
|
||||
}
|
||||
}
|
||||
println!("{} candidates in {ms:.2} ms", tries.len());
|
||||
}
|
||||
|
||||
fn show(a: &Args) {
|
||||
let (label, bytes) = seed_bytes_of(a);
|
||||
let p = igneum_pow::generator::generate_from_seed_bytes(&label, &bytes);
|
||||
println!(
|
||||
"seed \"{}\" generator v{} attempt {} program id {:016x} seed words {}",
|
||||
p.seed_string,
|
||||
p.generator,
|
||||
p.attempt,
|
||||
p.program_id(),
|
||||
p.seed.iter().map(|w| format!("{w:08x}")).collect::<Vec<_>>().join(" ")
|
||||
);
|
||||
println!("op mix {} loads/hash {}", p.op_mix(), p.loads_per_hash());
|
||||
for (k, i) in p.instrs.iter().enumerate() {
|
||||
println!(
|
||||
"{k:2}: {:5} dst={} src={} src2={} imm={:#010x} imm2={:#010x} rot={} bit={} mask={}",
|
||||
i.op.name(),
|
||||
i.dst,
|
||||
i.src,
|
||||
i.src2,
|
||||
i.imm,
|
||||
i.imm2,
|
||||
i.rot,
|
||||
i.bit,
|
||||
i.mask
|
||||
);
|
||||
}
|
||||
}
|
||||
|
|
|
|||
|
|
@ -61,13 +61,18 @@ pub struct DatasetSource {
|
|||
pub mask: u32,
|
||||
/// The day key `K`; `d0, d1 = K[0], K[1]`.
|
||||
pub key: [u32; 8],
|
||||
/// The bytes `K` was derived from (`"day/<day>"` for a string day, `bind::day_bytes` on the chain), recorded
|
||||
/// in packs so any implementation can rebuild the key. Empty when the key was given directly.
|
||||
pub key_bytes: Vec<u8>,
|
||||
pub dataset: Dataset,
|
||||
}
|
||||
|
||||
impl DatasetSource {
|
||||
/// Build the source for a day. Memory-hard mode fills the 256 MiB cache on the calling thread.
|
||||
pub fn new(day: &str, mode: DatasetMode, log2_words: u32) -> Self {
|
||||
Self::from_key(day_key(day), mode, log2_words)
|
||||
let mut ds = Self::from_key(day_key(day), mode, log2_words);
|
||||
ds.key_bytes = format!("day/{day}").into_bytes();
|
||||
ds
|
||||
}
|
||||
|
||||
pub fn from_key(key: [u32; 8], mode: DatasetMode, log2_words: u32) -> Self {
|
||||
|
|
@ -77,7 +82,7 @@ impl DatasetSource {
|
|||
DatasetMode::ClosedForm => Dataset::ClosedForm { d0: key[0], d1: key[1] },
|
||||
DatasetMode::MemoryHard => Dataset::MemoryHard(MemhardCpu::new(key)),
|
||||
};
|
||||
Self { log2_words, mask, key, dataset }
|
||||
Self { log2_words, mask, key, key_bytes: Vec::new(), dataset }
|
||||
}
|
||||
|
||||
pub fn mode(&self) -> DatasetMode {
|
||||
|
|
@ -299,14 +304,16 @@ impl Epoch {
|
|||
Self::new(seed, day, DatasetMode::MemoryHard, DEFAULT_DATASET_LOG2)
|
||||
}
|
||||
|
||||
/// The chain's shape: program from `seed_words_from_bytes(epoch_seed)` (devnet v0: the 32-byte epoch block
|
||||
/// hash; later the VDF output) and the cache from `seed_words_from_bytes(day_bytes)` (`bind::day_bytes`).
|
||||
/// Memory-hard, 1 GiB dataset. `label` is only recorded in emitted packs.
|
||||
/// The chain's shape: program from the 32-byte epoch seed (devnet: the epoch block hash; later the VDF
|
||||
/// output) through the version 2 generator and its acceptance rule, and the cache from
|
||||
/// `seed_words_from_bytes(day_bytes)` (`bind::day_bytes`). Memory-hard, 1 GiB dataset. `label` is only
|
||||
/// recorded in emitted packs.
|
||||
pub fn from_seed_bytes(epoch_seed: &[u8], day_bytes: &[u8], label: &str) -> Self {
|
||||
let words = crate::seed::seed_words_from_bytes(epoch_seed);
|
||||
let program = crate::generator::generate_from_words(label, words, &crate::generator::GeneratorConfig::default());
|
||||
let program = crate::generator::generate_from_seed_bytes(label, epoch_seed);
|
||||
let key = crate::seed::seed_words_from_bytes(day_bytes);
|
||||
Self { program, dataset: DatasetSource::from_key(key, DatasetMode::MemoryHard, DEFAULT_DATASET_LOG2) }
|
||||
let mut dataset = DatasetSource::from_key(key, DatasetMode::MemoryHard, DEFAULT_DATASET_LOG2);
|
||||
dataset.key_bytes = day_bytes.to_vec();
|
||||
Self { program, dataset }
|
||||
}
|
||||
|
||||
/// The 32 hashes of the warp starting at `base_nonce`.
|
||||
|
|
@ -351,10 +358,11 @@ mod tests {
|
|||
|
||||
#[test]
|
||||
fn closed_form_genesis_vector_lane0() {
|
||||
// Generator v2 vectors (4 October 2026), proto-cuda/packs/igneum-genesis/vectors.json.
|
||||
let e = Epoch::new("igneum-genesis", "2026-10-03", DatasetMode::ClosedForm, 28);
|
||||
let w = e.hash_warp(0);
|
||||
assert_eq!(w[0], 0x2941e93c76cb1910);
|
||||
assert_eq!(w[31], 0x453388e1be04e25f);
|
||||
assert_eq!(w[0], 0x31e7555c3dfd007f);
|
||||
assert_eq!(w[31], 0xaab617183923ab2a);
|
||||
assert_eq!(e.hash(0), w[0]);
|
||||
assert_eq!(e.hash(31), w[31]);
|
||||
assert!(e.verify_block(0, u64::MAX));
|
||||
|
|
|
|||
|
|
@ -1,21 +1,25 @@
|
|||
//! Agreement with the Swift prototype through the checked-in packs under proto-cuda/packs/.
|
||||
//! The checked-in packs under proto-cuda/packs/ against this crate (their source since generator version 2,
|
||||
//! 4 October 2026): every program instruction by instruction from its seed bytes, the generator version, attempt
|
||||
//! and program id, the dataset and cache self-test words, the 96 hash vectors per pack, and every emitted file
|
||||
//! byte for byte. A pack that drifts from the emitters, or a generator change that moves a vector, fails here.
|
||||
//!
|
||||
//! igneum-genesis-mh: memory-hard dataset (cache FNV, head and last line, 64 sampled words, 96 hashes).
|
||||
//! igneum-genesis and igneum-hourly: closed-form dataset (head, last, 64 samples, 96 hashes each).
|
||||
//! All three: program.json instruction by instruction, and every emitted source file byte for byte.
|
||||
//! Packs: igneum-genesis-mh and igneum-devnet-v4-epoch0 (memory-hard; the latter from the devnet genesis hash as
|
||||
//! the epoch seed and the day bytes of 2026-10-04), igneum-genesis and igneum-hourly (closed-form dataset,
|
||||
//! interpreter regression only).
|
||||
|
||||
use igneum_pow::accept;
|
||||
use igneum_pow::emit::{
|
||||
cuda_kernel, cuda_memhard_header, export_pack, metal_memhard, metal_program, opencl_kernel, program_header,
|
||||
program_json, LoadSource,
|
||||
cuda_kernel, cuda_kernel_bound, cuda_memhard_header, export_pack, metal_memhard, metal_program,
|
||||
metal_program_bound, opencl_kernel, opencl_kernel_bound, program_header, program_json, LoadSource,
|
||||
};
|
||||
use igneum_pow::generator::{generate, Op};
|
||||
use igneum_pow::generator::{generate_from_seed_bytes, Op, GENERATOR_VERSION, LOAD_SLOTS};
|
||||
use igneum_pow::memhard::CACHE_WORDS;
|
||||
use igneum_pow::verify::{DatasetMode, Epoch};
|
||||
use igneum_pow::verify::{DatasetMode, DatasetSource, Epoch};
|
||||
use serde_json::Value;
|
||||
use std::path::PathBuf;
|
||||
use std::sync::OnceLock;
|
||||
|
||||
const DAY: &str = "2026-10-03";
|
||||
const PACKS: [&str; 4] = ["igneum-genesis-mh", "igneum-devnet-v4-epoch0", "igneum-genesis", "igneum-hourly"];
|
||||
|
||||
fn packs_dir() -> PathBuf {
|
||||
PathBuf::from(env!("CARGO_MANIFEST_DIR")).join("../proto-cuda/packs")
|
||||
|
|
@ -26,15 +30,8 @@ fn read(pack: &str, file: &str) -> String {
|
|||
std::fs::read_to_string(&p).unwrap_or_else(|e| panic!("read {}: {e}", p.display()))
|
||||
}
|
||||
|
||||
/// The Swift exporter writes the cache line mask quoted inside the "item" string, which is not valid JSON.
|
||||
/// Normalise that one defect so the file can be parsed and compared; the Rust emitter writes it bare.
|
||||
fn fix_swift_item_line(s: &str) -> String {
|
||||
s.replace("& \"0x003fffff\";", "& 0x003fffff;")
|
||||
}
|
||||
|
||||
fn json(pack: &str, file: &str) -> Value {
|
||||
let text = fix_swift_item_line(&read(pack, file));
|
||||
serde_json::from_str(&text).unwrap_or_else(|e| panic!("{pack}/{file}: {e}"))
|
||||
serde_json::from_str(&read(pack, file)).unwrap_or_else(|e| panic!("{pack}/{file}: {e}"))
|
||||
}
|
||||
|
||||
fn hex32(v: &Value) -> u32 {
|
||||
|
|
@ -43,24 +40,55 @@ fn hex32(v: &Value) -> u32 {
|
|||
fn hex64(v: &Value) -> u64 {
|
||||
u64::from_str_radix(v.as_str().unwrap().trim_start_matches("0x"), 16).unwrap()
|
||||
}
|
||||
|
||||
/// One memory-hard epoch shared by every test (the cache fill is 256 MiB and about 0.2 s).
|
||||
fn mh_epoch() -> &'static Epoch {
|
||||
static E: OnceLock<Epoch> = OnceLock::new();
|
||||
E.get_or_init(|| Epoch::new("igneum-genesis", DAY, DatasetMode::MemoryHard, 28))
|
||||
fn unhex(v: &Value) -> Vec<u8> {
|
||||
igneum_pow::bind::unhex(v.as_str().unwrap()).unwrap()
|
||||
}
|
||||
|
||||
fn closed_epoch(seed: &str) -> Epoch {
|
||||
Epoch::new(seed, DAY, DatasetMode::ClosedForm, 28)
|
||||
/// The epoch a pack describes, rebuilt from program.json alone: the program from `seed_bytes`, the dataset from
|
||||
/// `dataset.day_bytes` in the pack's mode and size. Memory-hard packs fill a 256 MiB cache (about 0.2 s), so
|
||||
/// each is built once.
|
||||
fn epoch(pack: &str) -> &'static Epoch {
|
||||
static E: OnceLock<Vec<(String, Epoch)>> = OnceLock::new();
|
||||
let all = E.get_or_init(|| {
|
||||
PACKS
|
||||
.iter()
|
||||
.map(|p| {
|
||||
let j = json(p, "program.json");
|
||||
let seed = j["seed"].as_str().unwrap();
|
||||
let seed_bytes = unhex(&j["seed_bytes"]);
|
||||
let day_bytes = unhex(&j["dataset"]["day_bytes"]);
|
||||
let mode = match j["dataset_mode"].as_str().unwrap() {
|
||||
"memory-hard" => DatasetMode::MemoryHard,
|
||||
_ => DatasetMode::ClosedForm,
|
||||
};
|
||||
let log2 = j["dataset"]["log2_words"].as_u64().unwrap() as u32;
|
||||
let program = generate_from_seed_bytes(seed, &seed_bytes);
|
||||
let mut dataset =
|
||||
DatasetSource::from_key(igneum_pow::seed::seed_words_from_bytes(&day_bytes), mode, log2);
|
||||
dataset.key_bytes = day_bytes;
|
||||
(p.to_string(), Epoch { program, dataset })
|
||||
})
|
||||
.collect()
|
||||
});
|
||||
&all.iter().find(|(n, _)| n == pack).unwrap().1
|
||||
}
|
||||
|
||||
fn day_label(pack: &str) -> String {
|
||||
json(pack, "program.json")["dataset"]["day"].as_str().unwrap().to_string()
|
||||
}
|
||||
|
||||
fn check_program_json(pack: &str) {
|
||||
let j = json(pack, "program.json");
|
||||
let seed = j["seed"].as_str().unwrap();
|
||||
let p = generate(seed);
|
||||
let p = &epoch(pack).program;
|
||||
assert_eq!(j["format"].as_str().unwrap(), "igneum-program-pack-3");
|
||||
assert_eq!(j["generator"].as_u64().unwrap() as u32, GENERATOR_VERSION, "{pack}: generator version");
|
||||
assert_eq!(j["attempt"].as_u64().unwrap() as u32, p.attempt, "{pack}: attempt");
|
||||
assert_eq!(hex64(&j["program_id"]), p.program_id(), "{pack}: program id");
|
||||
let sw: Vec<u32> = j["seed_words"].as_array().unwrap().iter().map(hex32).collect();
|
||||
assert_eq!(p.seed.to_vec(), sw, "{pack}: seed words");
|
||||
assert_eq!(p.loads_per_hash() as u64, j["loads_per_hash"].as_u64().unwrap(), "{pack}: loads per hash");
|
||||
assert_eq!(p.loads_per_hash(), 8 * LOAD_SLOTS);
|
||||
assert!(accept::check(p).is_ok(), "{pack}: the pack's program passes the acceptance rule");
|
||||
let instrs = j["instructions"].as_array().unwrap();
|
||||
assert_eq!(instrs.len(), p.instrs.len(), "{pack}: instruction count");
|
||||
for (k, (ins, ji)) in p.instrs.iter().zip(instrs).enumerate() {
|
||||
|
|
@ -75,7 +103,6 @@ fn check_program_json(pack: &str) {
|
|||
assert_eq!(ji["bit"].as_u64().unwrap(), ins.bit as u64, "{pack} #{k} bit");
|
||||
assert_eq!(ji["mask"].as_u64().unwrap(), ins.mask as u64, "{pack} #{k} mask");
|
||||
}
|
||||
// Op mix.
|
||||
let mix = j["op_mix"].as_object().unwrap();
|
||||
for (name, count) in p.histogram() {
|
||||
assert_eq!(mix[name].as_u64().unwrap() as usize, count, "{pack}: op_mix {name}");
|
||||
|
|
@ -85,45 +112,61 @@ fn check_program_json(pack: &str) {
|
|||
|
||||
#[test]
|
||||
fn program_json_matches_all_packs() {
|
||||
for pack in ["igneum-genesis", "igneum-genesis-mh", "igneum-hourly"] {
|
||||
for pack in PACKS {
|
||||
check_program_json(pack);
|
||||
}
|
||||
}
|
||||
|
||||
/// The genesis program of generator v2 (spec 01 section 1.4.3 vector).
|
||||
#[test]
|
||||
fn genesis_program_shape() {
|
||||
let p = &epoch("igneum-genesis-mh").program;
|
||||
assert_eq!(p.attempt, 0);
|
||||
assert_eq!(p.op_mix(), "load=16 add=8 shfl=8 xor=6 mad=5 mul=5 mulhi=5 sub=4 rotl=3 rotr=3 or=1");
|
||||
assert_eq!(p.program_id(), 0xbcc1248b10cc90f2);
|
||||
assert_eq!(&epoch("igneum-genesis").program, p, "closed-form and memory-hard packs share the program");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn mixer_params_match_pack() {
|
||||
let j = json("igneum-genesis-mh", "program.json");
|
||||
let mp = &mh_epoch().dataset.memhard().unwrap().params;
|
||||
let key: Vec<u32> = j["dataset"]["key"].as_array().unwrap().iter().map(hex32).collect();
|
||||
assert_eq!(mp.key.to_vec(), key);
|
||||
let rot: Vec<u32> =
|
||||
j["dataset"]["mixer"]["rot"].as_array().unwrap().iter().map(|v| v.as_u64().unwrap() as u32).collect();
|
||||
assert_eq!(mp.rot.to_vec(), rot);
|
||||
let mul: Vec<u32> = j["dataset"]["mixer"]["mul"].as_array().unwrap().iter().map(hex32).collect();
|
||||
assert_eq!(mp.mul.to_vec(), mul);
|
||||
let rc: Vec<u32> = j["dataset"]["mixer"]["rc"].as_array().unwrap().iter().map(hex32).collect();
|
||||
assert_eq!(mp.rc.to_vec(), rc);
|
||||
assert_eq!(hex32(&j["dataset"]["d0"]), mp.key[0]);
|
||||
assert_eq!(hex32(&j["dataset"]["d1"]), mp.key[1]);
|
||||
for pack in ["igneum-genesis-mh", "igneum-devnet-v4-epoch0"] {
|
||||
let j = json(pack, "program.json");
|
||||
let mp = &epoch(pack).dataset.memhard().unwrap().params;
|
||||
let key: Vec<u32> = j["dataset"]["key"].as_array().unwrap().iter().map(hex32).collect();
|
||||
assert_eq!(mp.key.to_vec(), key);
|
||||
let rot: Vec<u32> =
|
||||
j["dataset"]["mixer"]["rot"].as_array().unwrap().iter().map(|v| v.as_u64().unwrap() as u32).collect();
|
||||
assert_eq!(mp.rot.to_vec(), rot);
|
||||
let mul: Vec<u32> = j["dataset"]["mixer"]["mul"].as_array().unwrap().iter().map(hex32).collect();
|
||||
assert_eq!(mp.mul.to_vec(), mul);
|
||||
let rc: Vec<u32> = j["dataset"]["mixer"]["rc"].as_array().unwrap().iter().map(hex32).collect();
|
||||
assert_eq!(mp.rc.to_vec(), rc);
|
||||
assert_eq!(hex32(&j["dataset"]["d0"]), mp.key[0]);
|
||||
assert_eq!(hex32(&j["dataset"]["d1"]), mp.key[1]);
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn cache_matches_vectors() {
|
||||
let v = json("igneum-genesis-mh", "vectors.json");
|
||||
let m = mh_epoch().dataset.memhard().unwrap();
|
||||
let m = epoch("igneum-genesis-mh").dataset.memhard().unwrap();
|
||||
let w = m.cache.words();
|
||||
assert_eq!(w.len(), CACHE_WORDS);
|
||||
let head: Vec<u32> = v["cache_head"].as_array().unwrap().iter().map(hex32).collect();
|
||||
assert_eq!(&w[..16], &head[..]);
|
||||
let last: Vec<u32> = v["cache_last_line"].as_array().unwrap().iter().map(hex32).collect();
|
||||
assert_eq!(&w[CACHE_WORDS - 16..], &last[..]);
|
||||
assert_eq!(m.cache.fnv1a64(), 0x48c4f5bf24166b2e, "cache FNV-1a 64 (MEMHARD.md)");
|
||||
assert_eq!(m.cache.fnv1a64(), 0x48c4f5bf24166b2e, "cache FNV-1a 64 of day 2026-10-03 (MEMHARD.md), unchanged by v2");
|
||||
assert_eq!(m.cache.fnv1a64(), hex64(&v["cache_fnv1a64"]));
|
||||
let v = json("igneum-devnet-v4-epoch0", "vectors.json");
|
||||
let m = epoch("igneum-devnet-v4-epoch0").dataset.memhard().unwrap();
|
||||
assert_eq!(m.cache.fnv1a64(), hex64(&v["cache_fnv1a64"]));
|
||||
assert_eq!(m.cache.fnv1a64(), 0x448274a57f508cbc, "cache FNV-1a 64 of day bytes igneum-day/20730");
|
||||
}
|
||||
|
||||
fn check_dataset_words(pack: &str, e: &Epoch) {
|
||||
fn check_dataset_words(pack: &str) {
|
||||
let v = json(pack, "vectors.json");
|
||||
let ds = &e.dataset;
|
||||
let ds = &epoch(pack).dataset;
|
||||
let head: Vec<u32> = v["dataset_head"].as_array().unwrap().iter().map(hex32).collect();
|
||||
for (i, h) in head.iter().enumerate() {
|
||||
assert_eq!(ds.word(i as u32), *h, "{pack}: dataset[{i}]");
|
||||
|
|
@ -133,29 +176,23 @@ fn check_dataset_words(pack: &str, e: &Epoch) {
|
|||
assert_eq!(ds.word(last_index), hex32(&v["dataset_last"]), "{pack}: dataset[MASK]");
|
||||
let samples = v["dataset_samples"].as_array().unwrap();
|
||||
assert_eq!(samples.len(), 64);
|
||||
let mut n = 0;
|
||||
for s in samples {
|
||||
let idx = s["index"].as_u64().unwrap() as u32;
|
||||
assert_eq!(ds.word(idx), hex32(&s["value"]), "{pack}: dataset[{idx}]");
|
||||
n += 1;
|
||||
}
|
||||
assert_eq!(n, 64);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn dataset_words_match_memhard_pack() {
|
||||
check_dataset_words("igneum-genesis-mh", mh_epoch());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn dataset_words_match_closed_packs() {
|
||||
check_dataset_words("igneum-genesis", &closed_epoch("igneum-genesis"));
|
||||
check_dataset_words("igneum-hourly", &closed_epoch("igneum-hourly"));
|
||||
fn dataset_words_match_all_packs() {
|
||||
for pack in PACKS {
|
||||
check_dataset_words(pack);
|
||||
}
|
||||
}
|
||||
|
||||
/// Returns the number of hashes compared (3 warps x 32 lanes = 96).
|
||||
fn check_vectors(pack: &str, e: &Epoch) -> usize {
|
||||
fn check_vectors(pack: &str) -> usize {
|
||||
let v = json(pack, "vectors.json");
|
||||
let e = epoch(pack);
|
||||
assert_eq!(v["dataset_mode"].as_str().unwrap(), e.dataset.mode().name());
|
||||
assert_eq!(v["dataset_log2_words"].as_u64().unwrap() as u32, e.dataset.log2_words);
|
||||
let warps = v["warps"].as_array().unwrap();
|
||||
|
|
@ -178,24 +215,26 @@ fn check_vectors(pack: &str, e: &Epoch) -> usize {
|
|||
}
|
||||
|
||||
#[test]
|
||||
fn vectors_memhard_96() {
|
||||
assert_eq!(check_vectors("igneum-genesis-mh", mh_epoch()), 96);
|
||||
fn vectors_96_per_pack() {
|
||||
for pack in PACKS {
|
||||
assert_eq!(check_vectors(pack), 96, "{pack}");
|
||||
}
|
||||
}
|
||||
|
||||
/// The spec 01 section 1.17 vectors (igneum-genesis-mh, generator v2).
|
||||
#[test]
|
||||
fn vectors_closed_form_genesis_96() {
|
||||
assert_eq!(check_vectors("igneum-genesis", &closed_epoch("igneum-genesis")), 96);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn vectors_closed_form_hourly_96() {
|
||||
assert_eq!(check_vectors("igneum-hourly", &closed_epoch("igneum-hourly")), 96);
|
||||
fn spec_vectors_genesis_mh() {
|
||||
let e = epoch("igneum-genesis-mh");
|
||||
let w = e.hash_warp(0);
|
||||
assert_eq!(w[0], 0x42246ba99fc58e4f);
|
||||
assert_eq!(w[31], 0xb08446b1f2de7793);
|
||||
assert_eq!(e.hash_warp(4096)[0], 0x3d3903e310ca038f);
|
||||
assert_eq!(e.hash_warp(1_000_000)[0], 0xf218c1bd58e6dfe0);
|
||||
}
|
||||
|
||||
fn assert_same_text(pack: &str, file: &str, got: &str) {
|
||||
let want = read(pack, file);
|
||||
if got != want {
|
||||
// Find the first differing line for a readable failure.
|
||||
let (gl, wl): (Vec<&str>, Vec<&str>) = (got.lines().collect(), want.lines().collect());
|
||||
for i in 0..gl.len().max(wl.len()) {
|
||||
let g = gl.get(i).copied().unwrap_or("<eof>");
|
||||
|
|
@ -208,93 +247,106 @@ fn assert_same_text(pack: &str, file: &str, got: &str) {
|
|||
}
|
||||
}
|
||||
|
||||
fn check_sources(pack: &str, e: &Epoch) {
|
||||
fn check_sources(pack: &str) {
|
||||
let e = epoch(pack);
|
||||
let p = &e.program;
|
||||
let day = day_label(pack);
|
||||
let mp = e.dataset.memhard().map(|m| &m.params);
|
||||
assert_same_text(pack, "kernel.cu", &cuda_kernel(p, mp));
|
||||
assert_same_text(pack, "kernel_bound.cu", &cuda_kernel_bound(p, mp));
|
||||
assert_same_text(pack, "program.metal", &metal_program(p, e.dataset.log2_words, LoadSource::Stored));
|
||||
assert_same_text(pack, "program_bound.metal", &metal_program_bound(p, e.dataset.log2_words));
|
||||
assert_same_text(pack, "kernel.cl", &opencl_kernel(p, mp));
|
||||
assert_same_text(pack, "program.h", &program_header(p, DAY, &e.dataset.key, e.dataset.log2_words, mp));
|
||||
assert_same_text(pack, "kernel_bound.cl", &opencl_kernel_bound(p, mp));
|
||||
assert_same_text(pack, "program.h", &program_header(p, &day, &e.dataset));
|
||||
if let Some(mp) = mp {
|
||||
assert_same_text(pack, "memhard.h", &cuda_memhard_header(p, mp));
|
||||
assert_same_text(pack, "memhard.metal", &metal_memhard(mp));
|
||||
}
|
||||
// program.json: byte-identical after normalising the Swift quoting defect in the "item" line.
|
||||
let want = fix_swift_item_line(&read(pack, "program.json"));
|
||||
let got = program_json(p, DAY, &e.dataset.key, e.dataset.log2_words, mp);
|
||||
assert_eq!(got, want, "{pack}/program.json");
|
||||
let _: Value = serde_json::from_str(&got).expect("Rust program.json is valid JSON");
|
||||
let got = program_json(p, &day, &e.dataset);
|
||||
assert_same_text(pack, "program.json", &got);
|
||||
let _: Value = serde_json::from_str(&got).expect("program.json is valid JSON");
|
||||
// Every load in every emitted hash kernel has the masked form, and there are exactly 16 of them.
|
||||
for (file, load, masked) in [
|
||||
("kernel.cu", "ds[r", " & mask]"),
|
||||
("kernel_bound.cu", "ds[r", " & mask]"),
|
||||
("program.metal", "dataset[r", " & MASK]"),
|
||||
("program_bound.metal", "dataset[r", " & MASK]"),
|
||||
] {
|
||||
let text = read(pack, file);
|
||||
assert_eq!(text.matches(load).count(), LOAD_SLOTS, "{pack}/{file}: 16 loads");
|
||||
assert_eq!(text.matches(masked).count(), LOAD_SLOTS, "{pack}/{file}: 16 masked loads");
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn emitted_sources_match_memhard_pack() {
|
||||
check_sources("igneum-genesis-mh", mh_epoch());
|
||||
fn emitted_sources_match_all_packs() {
|
||||
for pack in PACKS {
|
||||
check_sources(pack);
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn emitted_sources_match_closed_packs() {
|
||||
check_sources("igneum-genesis", &closed_epoch("igneum-genesis"));
|
||||
check_sources("igneum-hourly", &closed_epoch("igneum-hourly"));
|
||||
}
|
||||
|
||||
/// The whole pack as `export` writes it: vectors.json and vectors.h match the Swift ones apart from the
|
||||
/// provenance string (the Swift adds "Metal GPU cross-check PASS", which the Rust side cannot claim).
|
||||
fn check_export(pack: &str, e: &Epoch) {
|
||||
let swift_v = json(pack, "vectors.json");
|
||||
let source = swift_v["source"].as_str().unwrap();
|
||||
let out = export_pack(e, DAY, source);
|
||||
/// The whole pack as `export` writes it: vectors.json and vectors.h match, and the file list is the full set.
|
||||
fn check_export(pack: &str) {
|
||||
let e = epoch(pack);
|
||||
let v = json(pack, "vectors.json");
|
||||
let source = v["source"].as_str().unwrap();
|
||||
let out = export_pack(e, &day_label(pack), source);
|
||||
let file = |name: &str| -> &str { &out.files.iter().find(|(n, _)| n == name).unwrap().1 };
|
||||
assert_same_text(pack, "vectors.json", file("vectors.json"));
|
||||
assert_same_text(pack, "vectors.h", file("vectors.h"));
|
||||
let expected: Vec<&str> = if e.dataset.mode() == DatasetMode::MemoryHard {
|
||||
vec![
|
||||
"program.json",
|
||||
"vectors.json",
|
||||
"kernel.cu",
|
||||
"kernel.cl",
|
||||
"program.h",
|
||||
"vectors.h",
|
||||
"program.metal",
|
||||
"program_bound.metal",
|
||||
"kernel_bound.cu",
|
||||
"kernel_bound.cl",
|
||||
"memhard.h",
|
||||
"memhard.metal",
|
||||
]
|
||||
} else {
|
||||
vec![
|
||||
"program.json",
|
||||
"vectors.json",
|
||||
"kernel.cu",
|
||||
"kernel.cl",
|
||||
"program.h",
|
||||
"vectors.h",
|
||||
"program.metal",
|
||||
"program_bound.metal",
|
||||
"kernel_bound.cu",
|
||||
"kernel_bound.cl",
|
||||
]
|
||||
};
|
||||
let mut expected = vec![
|
||||
"program.json",
|
||||
"vectors.json",
|
||||
"kernel.cu",
|
||||
"kernel.cl",
|
||||
"program.h",
|
||||
"vectors.h",
|
||||
"program.metal",
|
||||
"program_bound.metal",
|
||||
"kernel_bound.cu",
|
||||
"kernel_bound.cl",
|
||||
];
|
||||
if e.dataset.mode() == DatasetMode::MemoryHard {
|
||||
expected.extend(["memhard.h", "memhard.metal"]);
|
||||
}
|
||||
assert_eq!(out.files.iter().map(|(n, _)| n.as_str()).collect::<Vec<_>>(), expected);
|
||||
let mut on_disk: Vec<String> = std::fs::read_dir(packs_dir().join(pack))
|
||||
.unwrap()
|
||||
.map(|d| d.unwrap().file_name().to_string_lossy().to_string())
|
||||
.filter(|n| !n.starts_with('.'))
|
||||
.collect();
|
||||
on_disk.sort();
|
||||
let mut want: Vec<String> = expected.iter().map(|s| s.to_string()).collect();
|
||||
want.sort();
|
||||
assert_eq!(on_disk, want, "{pack}: no stale file in the pack directory");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn export_pack_matches_memhard_pack() {
|
||||
check_export("igneum-genesis-mh", mh_epoch());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn export_pack_matches_closed_packs() {
|
||||
check_export("igneum-genesis", &closed_epoch("igneum-genesis"));
|
||||
check_export("igneum-hourly", &closed_epoch("igneum-hourly"));
|
||||
fn export_pack_matches_all_packs() {
|
||||
for pack in PACKS {
|
||||
check_export(pack);
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn item_is_independent_of_dataset_size() {
|
||||
// MEMHARD.md 1.7: a smaller dataset is a prefix of items, so dataset[w] is the same at every size.
|
||||
let big = &mh_epoch().dataset;
|
||||
let big = &epoch("igneum-genesis-mh").dataset;
|
||||
let m = big.memhard().unwrap();
|
||||
for w in [0u32, 1, 15, 16, 17, 0x00ff_ffff, 0x03ff_ffff] {
|
||||
assert_eq!(big.word(w), m.word(w));
|
||||
}
|
||||
}
|
||||
|
||||
/// The devnet pack is the chain's own derivation: epoch seed = the devnet genesis hash, day bytes = `bind::day_bytes(20730)`.
|
||||
#[test]
|
||||
fn devnet_pack_is_the_chain_derivation() {
|
||||
let j = json("igneum-devnet-v4-epoch0", "program.json");
|
||||
let genesis = igneum_pow::bind::unhex("edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07").unwrap();
|
||||
assert_eq!(unhex(&j["seed_bytes"]), genesis);
|
||||
assert_eq!(unhex(&j["dataset"]["day_bytes"]), igneum_pow::bind::day_bytes(20_730).to_vec());
|
||||
let e = Epoch::from_seed_bytes(&genesis, &igneum_pow::bind::day_bytes(20_730), "devnet");
|
||||
assert_eq!(e.program.instrs, epoch("igneum-devnet-v4-epoch0").program.instrs);
|
||||
assert_eq!(e.hash_warp(0), epoch("igneum-devnet-v4-epoch0").hash_warp(0));
|
||||
}
|
||||
|
|
|
|||
|
|
@ -11,6 +11,16 @@ Status on 3 October 2026: the two closed-form packs ran on the RTX 5090 (96/96 v
|
|||
`../proto-metal/MEMHARD.md`) has passed only the clang emulation on the Mac; its 5090 and AMD runs are pending, and no
|
||||
NVIDIA figure for the memory-hard dataset exists yet.
|
||||
|
||||
Status on 4 October 2026, generator version 2: every pack was regenerated by `igneum-pow export` (the Rust crate is
|
||||
now the pack source; the Swift exporter is a cross-check) with the adopted generator rule (exactly 16 loads per
|
||||
program, fresh-source loads, the acceptance rule of spec 01 section 1.4.6). All four packs pass the clang emulation
|
||||
(`emu/emu.sh`, 96/96 standalone, 2 warps per block in batch) and Apple OpenCL on the M5 Max; Metal cross-checked the
|
||||
exported genesis pack and a 2,000-program fuzz. The version 1 vectors, including the 5090's 192 of 192 from
|
||||
3 October, are retired: the kernel text is unchanged, so those runs remain evidence that the ops agree on NVIDIA, but
|
||||
the first 5090 run on a version 2 pack is still owed (`docs/bench-log.md`, 4 October 2026 entry). `program.h` now
|
||||
carries `IGNEUM_GENERATOR 2`, `IGNEUM_PROGRAM_ATTEMPT` and `IGNEUM_PROGRAM_ID`; a worker built from a version 1 pack
|
||||
cannot serve a version 2 node.
|
||||
|
||||
## Layout
|
||||
|
||||
```
|
||||
|
|
@ -19,7 +29,7 @@ proto-cuda/
|
|||
build.sh Linux build (nvcc)
|
||||
build.bat Windows build (nvcc + Visual Studio Build Tools)
|
||||
CHECKLIST.md Metal/CUDA equivalence, op by op, and what was verified where
|
||||
packs/<seed>/ one program pack per seed, written by proto-metal/igneum-bench --export-pack
|
||||
packs/<seed>/ one program pack per seed, written by igneum-pow export (generator v2, 4 October 2026)
|
||||
kernel.cu the program as a CUDA kernel, plus fill kernel and host launch wrappers
|
||||
kernel.cl the same program as OpenCL C for proto-opencl (AMD and any other OpenCL device), built at runtime
|
||||
program.h seed, day words, dataset size, loads per hash, wrapper declarations (C99-safe: proto-opencl/host.c includes it too)
|
||||
|
|
@ -32,7 +42,7 @@ proto-cuda/
|
|||
emu/ CPU emulation shim: compile and check a pack with plain clang++/g++, no GPU
|
||||
```
|
||||
|
||||
Three packs are checked in. `igneum-genesis` (104 loads per hash) and `igneum-hourly` (128 loads per hash) use the
|
||||
Four packs are checked in, all generator version 2 with 128 loads per hash. `igneum-genesis` and `igneum-hourly` use the
|
||||
original closed-form dataset (`IGNEUM_DATASET_MODE 0`, implied when the macro is absent). `igneum-genesis-mh` is the
|
||||
same program as `igneum-genesis` over the memory-hard dataset (`IGNEUM_DATASET_MODE 1`): a 256 MiB cache of chained
|
||||
ChaCha12 blocks filled on the GPU from the day key, and every 64-byte dataset item derived from 8 dependent cache
|
||||
|
|
@ -163,16 +173,20 @@ It is not an nvcc build and says nothing about NVIDIA hardware.
|
|||
|
||||
## Regenerating a pack
|
||||
|
||||
On the Mac:
|
||||
Since 4 October 2026 the packs are written by the Rust crate, which is the normative implementation:
|
||||
|
||||
```
|
||||
cd proto-metal
|
||||
swiftc -O -o igneum-bench main.swift -framework Metal
|
||||
./igneum-bench --seed igneum-genesis --export-pack ../proto-cuda/packs/igneum-genesis-mh # memory-hard (default)
|
||||
./igneum-bench --closed-form --seed igneum-genesis --export-pack ../proto-cuda/packs/igneum-genesis
|
||||
cd igneum-pow && cargo build --release
|
||||
./target/release/igneum-pow export --seed igneum-genesis --out ../proto-cuda/packs/igneum-genesis-mh # memory-hard (default)
|
||||
./target/release/igneum-pow export --closed-form --seed igneum-genesis --out ../proto-cuda/packs/igneum-genesis
|
||||
./target/release/igneum-pow export --closed-form --seed igneum-hourly --out ../proto-cuda/packs/igneum-hourly
|
||||
./target/release/igneum-pow export --epoch-hex <32-byte epoch seed> --day-hex <day bytes> --out ../proto-cuda/packs/<name>
|
||||
cargo test --release # the packs must match the emitters byte for byte
|
||||
```
|
||||
|
||||
The exporter runs the CPU interpreter for the three vector warps, runs the Metal kernel for the same warps,
|
||||
and refuses to write anything unless all 96 outputs match. For a memory-hard pack it also refuses unless the GPU
|
||||
cache equals the CPU cache on every word and the sampled GPU dataset words equal the CPU derivation. `--day` and `--dataset-log2` change the dataset
|
||||
constants and are recorded in the pack.
|
||||
The fourth form is the chain's derivation (`igneum-devnet-v4-epoch0`: the devnet genesis hash and the day bytes of
|
||||
2026-10-04). The Metal cross-check is a separate step: `proto-metal/igneum-bench --seed <seed> --export-pack <scratch dir>`
|
||||
derives the same program with the Swift generator, runs the Metal kernel for the three vector warps and refuses to
|
||||
write unless all 96 outputs match its CPU interpreter; diff its `program.json` instruction list and `vectors.json`
|
||||
against the Rust pack (identical on every seed tried, `docs/bench-log.md`). `--day` and `--dataset-log2` change the
|
||||
dataset constants and are recorded in the pack.
|
||||
|
|
|
|||
278
proto-cuda/packs/igneum-devnet-v4-epoch0/kernel.cl
Normal file
278
proto-cuda/packs/igneum-devnet-v4-epoch0/kernel.cl
Normal file
|
|
@ -0,0 +1,278 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal).
|
||||
// Built from source at runtime by proto-opencl/host.c, which passes these defines:
|
||||
// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit)
|
||||
// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default)
|
||||
// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32
|
||||
// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition
|
||||
// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units;
|
||||
// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32.
|
||||
#ifndef IGNEUM_GROUP
|
||||
#define IGNEUM_GROUP 32
|
||||
#endif
|
||||
#ifndef IGNEUM_EXCHANGE
|
||||
#define IGNEUM_EXCHANGE 0
|
||||
#endif
|
||||
#ifdef __OPENCL_VERSION__
|
||||
#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1)))
|
||||
#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n]
|
||||
#if IGNEUM_EXCHANGE == 1
|
||||
#ifdef cl_khr_subgroups
|
||||
#pragma OPENCL EXTENSION cl_khr_subgroups : enable
|
||||
#endif
|
||||
#ifdef cl_khr_subgroup_shuffle
|
||||
#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable
|
||||
#endif
|
||||
#elif IGNEUM_EXCHANGE == 2
|
||||
#pragma OPENCL EXTENSION cl_intel_subgroups : enable
|
||||
#endif
|
||||
#else
|
||||
// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros.
|
||||
#include "emu_opencl.h"
|
||||
#endif
|
||||
|
||||
#if IGNEUM_EXCHANGE == 1
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m))
|
||||
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
|
||||
#elif IGNEUM_EXCHANGE == 2
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m))
|
||||
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
|
||||
#else
|
||||
// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per
|
||||
// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane
|
||||
// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's
|
||||
// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier.
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; }
|
||||
#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; }
|
||||
#endif
|
||||
|
||||
static inline uint splitmix32(uint x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32.
|
||||
static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); }
|
||||
// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x.
|
||||
static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); }
|
||||
static inline uint ds_elem(uint i, uint d0, uint d1) {
|
||||
uint x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
|
||||
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
|
||||
#define MH_CACHE_LINE_MASK 0x003fffffu
|
||||
#define MH_SEGMENT_LINES 64u
|
||||
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
|
||||
static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
|
||||
|
||||
// y = ChaCha12 core(x) + x
|
||||
static inline void mh_chacha_block(const uint* x, uint* y) {
|
||||
for (uint i = 0u; i < 16u; ++i) y[i] = x[i];
|
||||
for (uint r = 0u; r < 6u; ++r) {
|
||||
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
|
||||
}
|
||||
for (uint i = 0u; i < 16u; ++i) y[i] += x[i];
|
||||
}
|
||||
|
||||
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
|
||||
static inline void mh_cache_segment(__global uint* cache, uint seg) {
|
||||
uint prev[16]; uint x[16]; uint y[16];
|
||||
for (uint i = 0u; i < 16u; ++i) prev[i] = 0u;
|
||||
for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) {
|
||||
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
|
||||
x[4] = 0xceed56d7u ^ prev[4];
|
||||
x[5] = 0x9ba270d2u ^ prev[5];
|
||||
x[6] = 0x82caab2du ^ prev[6];
|
||||
x[7] = 0x81ebce0eu ^ prev[7];
|
||||
x[8] = 0x12b6ecf1u ^ prev[8];
|
||||
x[9] = 0xd0f3fd7cu ^ prev[9];
|
||||
x[10] = 0xd872eefeu ^ prev[10];
|
||||
x[11] = 0xc158c7bdu ^ prev[11];
|
||||
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
|
||||
mh_chacha_block(x, y);
|
||||
__global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
|
||||
}
|
||||
}
|
||||
|
||||
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
|
||||
static inline void mh_mixer(uint* s, uint rk) {
|
||||
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
|
||||
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
|
||||
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
|
||||
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
|
||||
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
|
||||
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
|
||||
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
|
||||
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
|
||||
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
|
||||
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
|
||||
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
|
||||
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
|
||||
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
|
||||
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
|
||||
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
|
||||
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
|
||||
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
|
||||
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
|
||||
}
|
||||
|
||||
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
|
||||
static inline void mh_item(__global const uint* cache, uint t, uint* s) {
|
||||
s[0] = 0xceed56d7u;
|
||||
s[1] = 0x9ba270d2u;
|
||||
s[2] = 0x82caab2du;
|
||||
s[3] = 0x81ebce0eu;
|
||||
s[4] = 0x12b6ecf1u;
|
||||
s[5] = 0xd0f3fd7cu;
|
||||
s[6] = 0xd872eefeu;
|
||||
s[7] = 0xc158c7bdu;
|
||||
s[8] = t * 0xf351d601u + 0xc6892460u;
|
||||
s[9] = t * 0xa3bb398fu + 0x25b7228au;
|
||||
s[10] = t * 0xb5a09e35u + 0xcd515004u;
|
||||
s[11] = t * 0x7509c9c1u + 0x2846527au;
|
||||
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
|
||||
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
|
||||
s[14] = t * 0xded91851u + 0x82961bacu;
|
||||
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
|
||||
for (uint r = 0u; r < 8u; ++r) {
|
||||
mh_mixer(s, 0x9E3779B9u * (r + 1u));
|
||||
__global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i];
|
||||
}
|
||||
mh_mixer(s, 0x9E3779B9u * 9u);
|
||||
}
|
||||
// dataset[w] without the dataset: derive item w >> 4 and take word w & 15.
|
||||
static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; }
|
||||
|
||||
// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item.
|
||||
// The same constants as memhard.h in this pack (one emitter, three dialects).
|
||||
__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) {
|
||||
uint seg = (uint)get_global_id(0);
|
||||
if (seg < nSegments) mh_cache_segment(cache, seg);
|
||||
}
|
||||
__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) {
|
||||
uint t = (uint)get_global_id(0);
|
||||
if (t < nItems) {
|
||||
uint s[16];
|
||||
mh_item(cache, t, s);
|
||||
__global uint* d = ds + ((ulong)t * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) d[i] = s[i];
|
||||
}
|
||||
}
|
||||
|
||||
// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the
|
||||
// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and
|
||||
// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all).
|
||||
IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) {
|
||||
uint gid = (uint)get_global_id(0);
|
||||
uint lid = (uint)get_local_id(0);
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
#if IGNEUM_EXCHANGE == 0
|
||||
IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
|
||||
uint xk = 0u;
|
||||
#else
|
||||
(void)lid;
|
||||
#endif
|
||||
{ uint x = nonce ^ 0x667d0fbdu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x7b8e5963u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
|
||||
{ uint x = nonce ^ 0x7b8e5963u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x31c67e5eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
|
||||
{ uint x = nonce ^ 0x31c67e5eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4529ddc6u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
|
||||
{ uint x = nonce ^ 0x4529ddc6u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef19d6d8u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
|
||||
{ uint x = nonce ^ 0xef19d6d8u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xaccf6211u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
|
||||
{ uint x = nonce ^ 0xaccf6211u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xda0aed32u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
|
||||
{ uint x = nonce ^ 0xda0aed32u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xabc6df31u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
|
||||
{ uint x = nonce ^ 0xabc6df31u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x667d0fbdu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r0, 4u); r2 = r2 ^ t_; } // 1 shfl
|
||||
r3 = r3 + r2 + ((((sel >> 10u) & 1u) != 0u) ? 0x642e66dbu : 0x2cccb6cau); // 2 add
|
||||
r0 = rotl_imm(r0, 19u); // 3 rotl
|
||||
r7 = rotr_var(r7, r6); // 4 rotr
|
||||
r7 = r7 + r4 + ((((sel >> 21u) & 1u) != 0u) ? 0xc1535555u : 0xee02465fu); // 5 add
|
||||
r1 = mul_hi(r1, r7); // 6 mulhi
|
||||
r4 = r4 ^ ds[r2 & mask]; // 7 load
|
||||
r7 = r7 ^ ds[r4 & mask]; // 8 load
|
||||
r0 = r0 ^ ds[r3 & mask]; // 9 load
|
||||
r5 = r5 ^ ds[r1 & mask]; // 10 load
|
||||
r1 = r1 ^ ds[r5 & mask]; // 11 load
|
||||
r3 = mul_hi(r3, r5); // 12 mulhi
|
||||
r1 = r1 ^ ds[r3 & mask]; // 13 load
|
||||
r0 = r0 - r3; // 14 sub
|
||||
r5 = r1 * r3 + r5; // 15 mad
|
||||
r6 = mul_hi(r6, r1); // 16 mulhi
|
||||
r5 = r5 + r2 + ((((sel >> 28u) & 1u) != 0u) ? 0x8b965b57u : 0x697b3d00u); // 17 add
|
||||
r0 = mul_hi(r0, r6); // 18 mulhi
|
||||
r5 = rotr_var(r5, r3); // 19 rotr
|
||||
r5 = mul_hi(r5, r2); // 20 mulhi
|
||||
r1 = r1 + r0 + ((((sel >> 1u) & 1u) != 0u) ? 0x6d7e8d05u : 0xebcf247au); // 21 add
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 22 add
|
||||
r1 = mul_hi(r1, r5); // 23 mulhi
|
||||
r2 = r2 - r5; // 24 sub
|
||||
r7 = r7 + r4 + ((((sel >> 2u) & 1u) != 0u) ? 0x699ef1bbu : 0x08ffa6c7u); // 25 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 26 shfl
|
||||
r7 = r7 + r1 + ((((sel >> 14u) & 1u) != 0u) ? 0xb4ead2fbu : 0xe60fea84u); // 27 add
|
||||
r3 = r3 + r1 + ((((sel >> 6u) & 1u) != 0u) ? 0x8f30d21du : 0x65c76dabu); // 28 add
|
||||
r2 = r2 ^ ds[r1 & mask]; // 29 load
|
||||
r5 = r5 ^ ds[r7 & mask]; // 30 load
|
||||
r2 = r2 ^ ds[r5 & mask]; // 31 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r1 = r1 ^ t_; } // 32 shfl
|
||||
r4 = r5 * r7 + r4; // 33 mad
|
||||
r4 = r4 + r2 + ((((sel >> 21u) & 1u) != 0u) ? 0xc7ce690cu : 0x0480debeu); // 34 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 8u); r3 = r3 ^ t_; } // 35 shfl
|
||||
r7 = r7 + r1 + ((((sel >> 2u) & 1u) != 0u) ? 0xe10c2c95u : 0xc53b542eu); // 36 add
|
||||
r5 = r5 ^ r7; // 37 xor
|
||||
r2 = r2 | r1; // 38 or
|
||||
r1 = mul_hi(r1, r0); // 39 mulhi
|
||||
r6 = rotl_imm(r6, 19u); // 40 rotl
|
||||
r4 = mul_hi(r4, r6); // 41 mulhi
|
||||
r6 = r6 - r0; // 42 sub
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r6 = r6 ^ t_; } // 43 shfl
|
||||
r4 = r4 ^ ds[r2 & mask]; // 44 load
|
||||
r1 = r1 ^ r3; // 45 xor
|
||||
r7 = r7 ^ ds[r0 & mask]; // 46 load
|
||||
r3 = r3 ^ ds[r1 & mask]; // 47 load
|
||||
r5 = r5 * r3; // 48 mul
|
||||
r1 = r1 - r5; // 49 sub
|
||||
r2 = rotl_imm(r2, 8u); // 50 rotl
|
||||
r1 = r1 + r5 + ((((sel >> 23u) & 1u) != 0u) ? 0x77b9bd43u : 0xa900fec4u); // 51 add
|
||||
r4 = r4 ^ ds[r7 & mask]; // 52 load
|
||||
r2 = r2 - r7; // 53 sub
|
||||
r4 = r4 ^ r0; // 54 xor
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 55 add
|
||||
r2 = r2 ^ ds[r4 & mask]; // 56 load
|
||||
r0 = r1 * r4 + r0; // 57 mad
|
||||
r3 = r3 ^ ds[r5 & mask]; // 58 load
|
||||
r5 = r5 | r6; // 59 or
|
||||
r6 = r5 * r7 + r6; // 60 mad
|
||||
r4 = rotl_imm(r4, 28u); // 61 rotl
|
||||
r5 = mul_hi(r5, r0); // 62 mulhi
|
||||
r3 = r3 ^ ds[r6 & mask]; // 63 load
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
|
||||
#if IGNEUM_EXCHANGE != 0
|
||||
// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the
|
||||
// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a
|
||||
// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md.
|
||||
IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) {
|
||||
if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); }
|
||||
}
|
||||
#endif
|
||||
164
proto-cuda/packs/igneum-devnet-v4-epoch0/kernel.cu
Normal file
164
proto-cuda/packs/igneum-devnet-v4-epoch0/kernel.cu
Normal file
|
|
@ -0,0 +1,164 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal).
|
||||
// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC.
|
||||
#include <cuda_runtime.h>
|
||||
#include <cstdint>
|
||||
#include "program.h"
|
||||
#include "memhard.h"
|
||||
|
||||
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31.
|
||||
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
|
||||
// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x.
|
||||
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) {
|
||||
uint32_t x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item.
|
||||
// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host.
|
||||
__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) {
|
||||
uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
if (seg < nSegments) mh_cache_segment(cache, seg);
|
||||
}
|
||||
__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
|
||||
uint32_t t = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
if (t < nItems) {
|
||||
uint32_t s[16];
|
||||
mh_item(cache, t, s);
|
||||
uint32_t* d = ds + (size_t)t * 16u;
|
||||
for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i];
|
||||
}
|
||||
}
|
||||
|
||||
// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every
|
||||
// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a
|
||||
// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid.
|
||||
__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) {
|
||||
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
uint32_t nonce = baseNonce + gid;
|
||||
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint32_t x = nonce ^ 0x667d0fbdu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x7b8e5963u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
|
||||
{ uint32_t x = nonce ^ 0x7b8e5963u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x31c67e5eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
|
||||
{ uint32_t x = nonce ^ 0x31c67e5eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4529ddc6u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
|
||||
{ uint32_t x = nonce ^ 0x4529ddc6u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef19d6d8u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
|
||||
{ uint32_t x = nonce ^ 0xef19d6d8u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xaccf6211u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
|
||||
{ uint32_t x = nonce ^ 0xaccf6211u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xda0aed32u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
|
||||
{ uint32_t x = nonce ^ 0xda0aed32u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xabc6df31u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
|
||||
{ uint32_t x = nonce ^ 0xabc6df31u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x667d0fbdu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
|
||||
|
||||
for (uint32_t it = 0u; it < 8u; ++it) {
|
||||
uint32_t sel = r0;
|
||||
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
|
||||
r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r0, 4); // 1 shfl
|
||||
r3 = r3 + r2 + ((((sel >> 10u) & 1u) != 0u) ? 0x642e66dbu : 0x2cccb6cau); // 2 add
|
||||
r0 = rotl_imm(r0, 19u); // 3 rotl
|
||||
r7 = rotr_var(r7, r6); // 4 rotr
|
||||
r7 = r7 + r4 + ((((sel >> 21u) & 1u) != 0u) ? 0xc1535555u : 0xee02465fu); // 5 add
|
||||
r1 = __umulhi(r1, r7); // 6 mulhi
|
||||
r4 = r4 ^ ds[r2 & mask]; // 7 load
|
||||
r7 = r7 ^ ds[r4 & mask]; // 8 load
|
||||
r0 = r0 ^ ds[r3 & mask]; // 9 load
|
||||
r5 = r5 ^ ds[r1 & mask]; // 10 load
|
||||
r1 = r1 ^ ds[r5 & mask]; // 11 load
|
||||
r3 = __umulhi(r3, r5); // 12 mulhi
|
||||
r1 = r1 ^ ds[r3 & mask]; // 13 load
|
||||
r0 = r0 - r3; // 14 sub
|
||||
r5 = r1 * r3 + r5; // 15 mad
|
||||
r6 = __umulhi(r6, r1); // 16 mulhi
|
||||
r5 = r5 + r2 + ((((sel >> 28u) & 1u) != 0u) ? 0x8b965b57u : 0x697b3d00u); // 17 add
|
||||
r0 = __umulhi(r0, r6); // 18 mulhi
|
||||
r5 = rotr_var(r5, r3); // 19 rotr
|
||||
r5 = __umulhi(r5, r2); // 20 mulhi
|
||||
r1 = r1 + r0 + ((((sel >> 1u) & 1u) != 0u) ? 0x6d7e8d05u : 0xebcf247au); // 21 add
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 22 add
|
||||
r1 = __umulhi(r1, r5); // 23 mulhi
|
||||
r2 = r2 - r5; // 24 sub
|
||||
r7 = r7 + r4 + ((((sel >> 2u) & 1u) != 0u) ? 0x699ef1bbu : 0x08ffa6c7u); // 25 add
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 26 shfl
|
||||
r7 = r7 + r1 + ((((sel >> 14u) & 1u) != 0u) ? 0xb4ead2fbu : 0xe60fea84u); // 27 add
|
||||
r3 = r3 + r1 + ((((sel >> 6u) & 1u) != 0u) ? 0x8f30d21du : 0x65c76dabu); // 28 add
|
||||
r2 = r2 ^ ds[r1 & mask]; // 29 load
|
||||
r5 = r5 ^ ds[r7 & mask]; // 30 load
|
||||
r2 = r2 ^ ds[r5 & mask]; // 31 load
|
||||
r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r7, 4); // 32 shfl
|
||||
r4 = r5 * r7 + r4; // 33 mad
|
||||
r4 = r4 + r2 + ((((sel >> 21u) & 1u) != 0u) ? 0xc7ce690cu : 0x0480debeu); // 34 add
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r7, 8); // 35 shfl
|
||||
r7 = r7 + r1 + ((((sel >> 2u) & 1u) != 0u) ? 0xe10c2c95u : 0xc53b542eu); // 36 add
|
||||
r5 = r5 ^ r7; // 37 xor
|
||||
r2 = r2 | r1; // 38 or
|
||||
r1 = __umulhi(r1, r0); // 39 mulhi
|
||||
r6 = rotl_imm(r6, 19u); // 40 rotl
|
||||
r4 = __umulhi(r4, r6); // 41 mulhi
|
||||
r6 = r6 - r0; // 42 sub
|
||||
r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 43 shfl
|
||||
r4 = r4 ^ ds[r2 & mask]; // 44 load
|
||||
r1 = r1 ^ r3; // 45 xor
|
||||
r7 = r7 ^ ds[r0 & mask]; // 46 load
|
||||
r3 = r3 ^ ds[r1 & mask]; // 47 load
|
||||
r5 = r5 * r3; // 48 mul
|
||||
r1 = r1 - r5; // 49 sub
|
||||
r2 = rotl_imm(r2, 8u); // 50 rotl
|
||||
r1 = r1 + r5 + ((((sel >> 23u) & 1u) != 0u) ? 0x77b9bd43u : 0xa900fec4u); // 51 add
|
||||
r4 = r4 ^ ds[r7 & mask]; // 52 load
|
||||
r2 = r2 - r7; // 53 sub
|
||||
r4 = r4 ^ r0; // 54 xor
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 55 add
|
||||
r2 = r2 ^ ds[r4 & mask]; // 56 load
|
||||
r0 = r1 * r4 + r0; // 57 mad
|
||||
r3 = r3 ^ ds[r5 & mask]; // 58 load
|
||||
r5 = r5 | r6; // 59 or
|
||||
r6 = r5 * r7 + r6; // 60 mad
|
||||
r4 = rotl_imm(r4, 28u); // 61 rotl
|
||||
r5 = __umulhi(r5, r0); // 62 mulhi
|
||||
r3 = r3 ^ ds[r6 & mask]; // 63 load
|
||||
}
|
||||
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
|
||||
}
|
||||
|
||||
// Host-side launch wrappers. Declared in program.h, called from host.cu.
|
||||
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) {
|
||||
if (nSegments == 0u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 256u;
|
||||
uint32_t grid = (nSegments + block - 1u) / block;
|
||||
igneum_cache_fill<<<grid, block>>>(cache, nSegments);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
|
||||
if (nItems == 0u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 256u;
|
||||
uint32_t grid = (nItems + block - 1u) / block;
|
||||
igneum_build<<<grid, block>>>(ds, cache, nItems);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
uint32_t nonces, uint32_t blockWarps) {
|
||||
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 32u * blockWarps;
|
||||
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
|
||||
igneum_hash<<<nonces / block, block>>>(ds, out, baseNonce, mask);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
|
||||
cudaFuncAttributes attr;
|
||||
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash);
|
||||
if (e != cudaSuccess) return e;
|
||||
*numRegs = attr.numRegs;
|
||||
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0);
|
||||
}
|
||||
372
proto-cuda/packs/igneum-devnet-v4-epoch0/kernel_bound.cl
Normal file
372
proto-cuda/packs/igneum-devnet-v4-epoch0/kernel_bound.cl
Normal file
|
|
@ -0,0 +1,372 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal).
|
||||
// Built from source at runtime by proto-opencl/host.c, which passes these defines:
|
||||
// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit)
|
||||
// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default)
|
||||
// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32
|
||||
// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition
|
||||
// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units;
|
||||
// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32.
|
||||
#ifndef IGNEUM_GROUP
|
||||
#define IGNEUM_GROUP 32
|
||||
#endif
|
||||
#ifndef IGNEUM_EXCHANGE
|
||||
#define IGNEUM_EXCHANGE 0
|
||||
#endif
|
||||
#ifdef __OPENCL_VERSION__
|
||||
#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1)))
|
||||
#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n]
|
||||
#if IGNEUM_EXCHANGE == 1
|
||||
#ifdef cl_khr_subgroups
|
||||
#pragma OPENCL EXTENSION cl_khr_subgroups : enable
|
||||
#endif
|
||||
#ifdef cl_khr_subgroup_shuffle
|
||||
#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable
|
||||
#endif
|
||||
#elif IGNEUM_EXCHANGE == 2
|
||||
#pragma OPENCL EXTENSION cl_intel_subgroups : enable
|
||||
#endif
|
||||
#else
|
||||
// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros.
|
||||
#include "emu_opencl.h"
|
||||
#endif
|
||||
|
||||
#if IGNEUM_EXCHANGE == 1
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m))
|
||||
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
|
||||
#elif IGNEUM_EXCHANGE == 2
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m))
|
||||
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
|
||||
#else
|
||||
// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per
|
||||
// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane
|
||||
// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's
|
||||
// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier.
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; }
|
||||
#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; }
|
||||
#endif
|
||||
|
||||
static inline uint splitmix32(uint x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32.
|
||||
static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); }
|
||||
// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x.
|
||||
static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); }
|
||||
static inline uint ds_elem(uint i, uint d0, uint d1) {
|
||||
uint x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
|
||||
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
|
||||
#define MH_CACHE_LINE_MASK 0x003fffffu
|
||||
#define MH_SEGMENT_LINES 64u
|
||||
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
|
||||
static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
|
||||
|
||||
// y = ChaCha12 core(x) + x
|
||||
static inline void mh_chacha_block(const uint* x, uint* y) {
|
||||
for (uint i = 0u; i < 16u; ++i) y[i] = x[i];
|
||||
for (uint r = 0u; r < 6u; ++r) {
|
||||
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
|
||||
}
|
||||
for (uint i = 0u; i < 16u; ++i) y[i] += x[i];
|
||||
}
|
||||
|
||||
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
|
||||
static inline void mh_cache_segment(__global uint* cache, uint seg) {
|
||||
uint prev[16]; uint x[16]; uint y[16];
|
||||
for (uint i = 0u; i < 16u; ++i) prev[i] = 0u;
|
||||
for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) {
|
||||
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
|
||||
x[4] = 0xceed56d7u ^ prev[4];
|
||||
x[5] = 0x9ba270d2u ^ prev[5];
|
||||
x[6] = 0x82caab2du ^ prev[6];
|
||||
x[7] = 0x81ebce0eu ^ prev[7];
|
||||
x[8] = 0x12b6ecf1u ^ prev[8];
|
||||
x[9] = 0xd0f3fd7cu ^ prev[9];
|
||||
x[10] = 0xd872eefeu ^ prev[10];
|
||||
x[11] = 0xc158c7bdu ^ prev[11];
|
||||
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
|
||||
mh_chacha_block(x, y);
|
||||
__global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
|
||||
}
|
||||
}
|
||||
|
||||
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
|
||||
static inline void mh_mixer(uint* s, uint rk) {
|
||||
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
|
||||
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
|
||||
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
|
||||
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
|
||||
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
|
||||
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
|
||||
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
|
||||
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
|
||||
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
|
||||
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
|
||||
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
|
||||
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
|
||||
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
|
||||
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
|
||||
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
|
||||
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
|
||||
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
|
||||
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
|
||||
}
|
||||
|
||||
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
|
||||
static inline void mh_item(__global const uint* cache, uint t, uint* s) {
|
||||
s[0] = 0xceed56d7u;
|
||||
s[1] = 0x9ba270d2u;
|
||||
s[2] = 0x82caab2du;
|
||||
s[3] = 0x81ebce0eu;
|
||||
s[4] = 0x12b6ecf1u;
|
||||
s[5] = 0xd0f3fd7cu;
|
||||
s[6] = 0xd872eefeu;
|
||||
s[7] = 0xc158c7bdu;
|
||||
s[8] = t * 0xf351d601u + 0xc6892460u;
|
||||
s[9] = t * 0xa3bb398fu + 0x25b7228au;
|
||||
s[10] = t * 0xb5a09e35u + 0xcd515004u;
|
||||
s[11] = t * 0x7509c9c1u + 0x2846527au;
|
||||
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
|
||||
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
|
||||
s[14] = t * 0xded91851u + 0x82961bacu;
|
||||
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
|
||||
for (uint r = 0u; r < 8u; ++r) {
|
||||
mh_mixer(s, 0x9E3779B9u * (r + 1u));
|
||||
__global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i];
|
||||
}
|
||||
mh_mixer(s, 0x9E3779B9u * 9u);
|
||||
}
|
||||
// dataset[w] without the dataset: derive item w >> 4 and take word w & 15.
|
||||
static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; }
|
||||
|
||||
// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item.
|
||||
// The same constants as memhard.h in this pack (one emitter, three dialects).
|
||||
__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) {
|
||||
uint seg = (uint)get_global_id(0);
|
||||
if (seg < nSegments) mh_cache_segment(cache, seg);
|
||||
}
|
||||
__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) {
|
||||
uint t = (uint)get_global_id(0);
|
||||
if (t < nItems) {
|
||||
uint s[16];
|
||||
mh_item(cache, t, s);
|
||||
__global uint* d = ds + ((ulong)t * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) d[i] = s[i];
|
||||
}
|
||||
}
|
||||
|
||||
// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the
|
||||
// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and
|
||||
// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all).
|
||||
IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) {
|
||||
uint gid = (uint)get_global_id(0);
|
||||
uint lid = (uint)get_local_id(0);
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
#if IGNEUM_EXCHANGE == 0
|
||||
IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
|
||||
uint xk = 0u;
|
||||
#else
|
||||
(void)lid;
|
||||
#endif
|
||||
{ uint x = nonce ^ 0x667d0fbdu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x7b8e5963u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
|
||||
{ uint x = nonce ^ 0x7b8e5963u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x31c67e5eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
|
||||
{ uint x = nonce ^ 0x31c67e5eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4529ddc6u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
|
||||
{ uint x = nonce ^ 0x4529ddc6u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef19d6d8u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
|
||||
{ uint x = nonce ^ 0xef19d6d8u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xaccf6211u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
|
||||
{ uint x = nonce ^ 0xaccf6211u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xda0aed32u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
|
||||
{ uint x = nonce ^ 0xda0aed32u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xabc6df31u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
|
||||
{ uint x = nonce ^ 0xabc6df31u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x667d0fbdu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r0, 4u); r2 = r2 ^ t_; } // 1 shfl
|
||||
r3 = r3 + r2 + ((((sel >> 10u) & 1u) != 0u) ? 0x642e66dbu : 0x2cccb6cau); // 2 add
|
||||
r0 = rotl_imm(r0, 19u); // 3 rotl
|
||||
r7 = rotr_var(r7, r6); // 4 rotr
|
||||
r7 = r7 + r4 + ((((sel >> 21u) & 1u) != 0u) ? 0xc1535555u : 0xee02465fu); // 5 add
|
||||
r1 = mul_hi(r1, r7); // 6 mulhi
|
||||
r4 = r4 ^ ds[r2 & mask]; // 7 load
|
||||
r7 = r7 ^ ds[r4 & mask]; // 8 load
|
||||
r0 = r0 ^ ds[r3 & mask]; // 9 load
|
||||
r5 = r5 ^ ds[r1 & mask]; // 10 load
|
||||
r1 = r1 ^ ds[r5 & mask]; // 11 load
|
||||
r3 = mul_hi(r3, r5); // 12 mulhi
|
||||
r1 = r1 ^ ds[r3 & mask]; // 13 load
|
||||
r0 = r0 - r3; // 14 sub
|
||||
r5 = r1 * r3 + r5; // 15 mad
|
||||
r6 = mul_hi(r6, r1); // 16 mulhi
|
||||
r5 = r5 + r2 + ((((sel >> 28u) & 1u) != 0u) ? 0x8b965b57u : 0x697b3d00u); // 17 add
|
||||
r0 = mul_hi(r0, r6); // 18 mulhi
|
||||
r5 = rotr_var(r5, r3); // 19 rotr
|
||||
r5 = mul_hi(r5, r2); // 20 mulhi
|
||||
r1 = r1 + r0 + ((((sel >> 1u) & 1u) != 0u) ? 0x6d7e8d05u : 0xebcf247au); // 21 add
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 22 add
|
||||
r1 = mul_hi(r1, r5); // 23 mulhi
|
||||
r2 = r2 - r5; // 24 sub
|
||||
r7 = r7 + r4 + ((((sel >> 2u) & 1u) != 0u) ? 0x699ef1bbu : 0x08ffa6c7u); // 25 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 26 shfl
|
||||
r7 = r7 + r1 + ((((sel >> 14u) & 1u) != 0u) ? 0xb4ead2fbu : 0xe60fea84u); // 27 add
|
||||
r3 = r3 + r1 + ((((sel >> 6u) & 1u) != 0u) ? 0x8f30d21du : 0x65c76dabu); // 28 add
|
||||
r2 = r2 ^ ds[r1 & mask]; // 29 load
|
||||
r5 = r5 ^ ds[r7 & mask]; // 30 load
|
||||
r2 = r2 ^ ds[r5 & mask]; // 31 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r1 = r1 ^ t_; } // 32 shfl
|
||||
r4 = r5 * r7 + r4; // 33 mad
|
||||
r4 = r4 + r2 + ((((sel >> 21u) & 1u) != 0u) ? 0xc7ce690cu : 0x0480debeu); // 34 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 8u); r3 = r3 ^ t_; } // 35 shfl
|
||||
r7 = r7 + r1 + ((((sel >> 2u) & 1u) != 0u) ? 0xe10c2c95u : 0xc53b542eu); // 36 add
|
||||
r5 = r5 ^ r7; // 37 xor
|
||||
r2 = r2 | r1; // 38 or
|
||||
r1 = mul_hi(r1, r0); // 39 mulhi
|
||||
r6 = rotl_imm(r6, 19u); // 40 rotl
|
||||
r4 = mul_hi(r4, r6); // 41 mulhi
|
||||
r6 = r6 - r0; // 42 sub
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r6 = r6 ^ t_; } // 43 shfl
|
||||
r4 = r4 ^ ds[r2 & mask]; // 44 load
|
||||
r1 = r1 ^ r3; // 45 xor
|
||||
r7 = r7 ^ ds[r0 & mask]; // 46 load
|
||||
r3 = r3 ^ ds[r1 & mask]; // 47 load
|
||||
r5 = r5 * r3; // 48 mul
|
||||
r1 = r1 - r5; // 49 sub
|
||||
r2 = rotl_imm(r2, 8u); // 50 rotl
|
||||
r1 = r1 + r5 + ((((sel >> 23u) & 1u) != 0u) ? 0x77b9bd43u : 0xa900fec4u); // 51 add
|
||||
r4 = r4 ^ ds[r7 & mask]; // 52 load
|
||||
r2 = r2 - r7; // 53 sub
|
||||
r4 = r4 ^ r0; // 54 xor
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 55 add
|
||||
r2 = r2 ^ ds[r4 & mask]; // 56 load
|
||||
r0 = r1 * r4 + r0; // 57 mad
|
||||
r3 = r3 ^ ds[r5 & mask]; // 58 load
|
||||
r5 = r5 | r6; // 59 or
|
||||
r6 = r5 * r7 + r6; // 60 mad
|
||||
r4 = rotl_imm(r4, 28u); // 61 rotl
|
||||
r5 = mul_hi(r5, r0); // 62 mulhi
|
||||
r3 = r3 ^ ds[r6 & mask]; // 63 load
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
|
||||
#if IGNEUM_EXCHANGE != 0
|
||||
// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the
|
||||
// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a
|
||||
// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md.
|
||||
IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) {
|
||||
if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); }
|
||||
}
|
||||
#endif
|
||||
|
||||
// Header-bound variant (bind.rs): the init words come from initw, not SEEDW. Same body as igneum_hash.
|
||||
IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw) {
|
||||
uint gid = (uint)get_global_id(0);
|
||||
uint lid = (uint)get_local_id(0);
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7];
|
||||
#if IGNEUM_EXCHANGE == 0
|
||||
IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
|
||||
uint xk = 0u;
|
||||
#else
|
||||
(void)lid;
|
||||
#endif
|
||||
{ uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; }
|
||||
{ uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; }
|
||||
{ uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; }
|
||||
{ uint x = nonce ^ iw3; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw4; }
|
||||
{ uint x = nonce ^ iw4; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw5; }
|
||||
{ uint x = nonce ^ iw5; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw6; }
|
||||
{ uint x = nonce ^ iw6; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw7; }
|
||||
{ uint x = nonce ^ iw7; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw0; }
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r0, 4u); r2 = r2 ^ t_; } // 1 shfl
|
||||
r3 = r3 + r2 + ((((sel >> 10u) & 1u) != 0u) ? 0x642e66dbu : 0x2cccb6cau); // 2 add
|
||||
r0 = rotl_imm(r0, 19u); // 3 rotl
|
||||
r7 = rotr_var(r7, r6); // 4 rotr
|
||||
r7 = r7 + r4 + ((((sel >> 21u) & 1u) != 0u) ? 0xc1535555u : 0xee02465fu); // 5 add
|
||||
r1 = mul_hi(r1, r7); // 6 mulhi
|
||||
r4 = r4 ^ ds[r2 & mask]; // 7 load
|
||||
r7 = r7 ^ ds[r4 & mask]; // 8 load
|
||||
r0 = r0 ^ ds[r3 & mask]; // 9 load
|
||||
r5 = r5 ^ ds[r1 & mask]; // 10 load
|
||||
r1 = r1 ^ ds[r5 & mask]; // 11 load
|
||||
r3 = mul_hi(r3, r5); // 12 mulhi
|
||||
r1 = r1 ^ ds[r3 & mask]; // 13 load
|
||||
r0 = r0 - r3; // 14 sub
|
||||
r5 = r1 * r3 + r5; // 15 mad
|
||||
r6 = mul_hi(r6, r1); // 16 mulhi
|
||||
r5 = r5 + r2 + ((((sel >> 28u) & 1u) != 0u) ? 0x8b965b57u : 0x697b3d00u); // 17 add
|
||||
r0 = mul_hi(r0, r6); // 18 mulhi
|
||||
r5 = rotr_var(r5, r3); // 19 rotr
|
||||
r5 = mul_hi(r5, r2); // 20 mulhi
|
||||
r1 = r1 + r0 + ((((sel >> 1u) & 1u) != 0u) ? 0x6d7e8d05u : 0xebcf247au); // 21 add
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 22 add
|
||||
r1 = mul_hi(r1, r5); // 23 mulhi
|
||||
r2 = r2 - r5; // 24 sub
|
||||
r7 = r7 + r4 + ((((sel >> 2u) & 1u) != 0u) ? 0x699ef1bbu : 0x08ffa6c7u); // 25 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 26 shfl
|
||||
r7 = r7 + r1 + ((((sel >> 14u) & 1u) != 0u) ? 0xb4ead2fbu : 0xe60fea84u); // 27 add
|
||||
r3 = r3 + r1 + ((((sel >> 6u) & 1u) != 0u) ? 0x8f30d21du : 0x65c76dabu); // 28 add
|
||||
r2 = r2 ^ ds[r1 & mask]; // 29 load
|
||||
r5 = r5 ^ ds[r7 & mask]; // 30 load
|
||||
r2 = r2 ^ ds[r5 & mask]; // 31 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 4u); r1 = r1 ^ t_; } // 32 shfl
|
||||
r4 = r5 * r7 + r4; // 33 mad
|
||||
r4 = r4 + r2 + ((((sel >> 21u) & 1u) != 0u) ? 0xc7ce690cu : 0x0480debeu); // 34 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r7, 8u); r3 = r3 ^ t_; } // 35 shfl
|
||||
r7 = r7 + r1 + ((((sel >> 2u) & 1u) != 0u) ? 0xe10c2c95u : 0xc53b542eu); // 36 add
|
||||
r5 = r5 ^ r7; // 37 xor
|
||||
r2 = r2 | r1; // 38 or
|
||||
r1 = mul_hi(r1, r0); // 39 mulhi
|
||||
r6 = rotl_imm(r6, 19u); // 40 rotl
|
||||
r4 = mul_hi(r4, r6); // 41 mulhi
|
||||
r6 = r6 - r0; // 42 sub
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r6 = r6 ^ t_; } // 43 shfl
|
||||
r4 = r4 ^ ds[r2 & mask]; // 44 load
|
||||
r1 = r1 ^ r3; // 45 xor
|
||||
r7 = r7 ^ ds[r0 & mask]; // 46 load
|
||||
r3 = r3 ^ ds[r1 & mask]; // 47 load
|
||||
r5 = r5 * r3; // 48 mul
|
||||
r1 = r1 - r5; // 49 sub
|
||||
r2 = rotl_imm(r2, 8u); // 50 rotl
|
||||
r1 = r1 + r5 + ((((sel >> 23u) & 1u) != 0u) ? 0x77b9bd43u : 0xa900fec4u); // 51 add
|
||||
r4 = r4 ^ ds[r7 & mask]; // 52 load
|
||||
r2 = r2 - r7; // 53 sub
|
||||
r4 = r4 ^ r0; // 54 xor
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 55 add
|
||||
r2 = r2 ^ ds[r4 & mask]; // 56 load
|
||||
r0 = r1 * r4 + r0; // 57 mad
|
||||
r3 = r3 ^ ds[r5 & mask]; // 58 load
|
||||
r5 = r5 | r6; // 59 or
|
||||
r6 = r5 * r7 + r6; // 60 mad
|
||||
r4 = rotl_imm(r4, 28u); // 61 rotl
|
||||
r5 = mul_hi(r5, r0); // 62 mulhi
|
||||
r3 = r3 ^ ds[r6 & mask]; // 63 load
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
123
proto-cuda/packs/igneum-devnet-v4-epoch0/kernel_bound.cu
Normal file
123
proto-cuda/packs/igneum-devnet-v4-epoch0/kernel_bound.cu
Normal file
|
|
@ -0,0 +1,123 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// Header-bound twin of igneum_hash in kernel.cu: the init words come from a kernel argument, not SEEDW.
|
||||
// Host declarations (also in program_bound.h if present):
|
||||
// struct IgneumInitWords { uint32_t w[8]; };
|
||||
// cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
// IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps);
|
||||
// cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
|
||||
#include <cuda_runtime.h>
|
||||
#include <cstdint>
|
||||
#include "program.h"
|
||||
|
||||
struct IgneumInitWords { uint32_t w[8]; };
|
||||
|
||||
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
|
||||
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
|
||||
__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw) {
|
||||
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
uint32_t nonce = baseNonce + gid;
|
||||
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint32_t x = nonce ^ iw.w[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw.w[1]; }
|
||||
{ uint32_t x = nonce ^ iw.w[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw.w[2]; }
|
||||
{ uint32_t x = nonce ^ iw.w[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw.w[3]; }
|
||||
{ uint32_t x = nonce ^ iw.w[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw.w[4]; }
|
||||
{ uint32_t x = nonce ^ iw.w[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw.w[5]; }
|
||||
{ uint32_t x = nonce ^ iw.w[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw.w[6]; }
|
||||
{ uint32_t x = nonce ^ iw.w[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw.w[7]; }
|
||||
{ uint32_t x = nonce ^ iw.w[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw.w[0]; }
|
||||
|
||||
for (uint32_t it = 0u; it < 8u; ++it) {
|
||||
uint32_t sel = r0;
|
||||
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
|
||||
r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r0, 4); // 1 shfl
|
||||
r3 = r3 + r2 + ((((sel >> 10u) & 1u) != 0u) ? 0x642e66dbu : 0x2cccb6cau); // 2 add
|
||||
r0 = rotl_imm(r0, 19u); // 3 rotl
|
||||
r7 = rotr_var(r7, r6); // 4 rotr
|
||||
r7 = r7 + r4 + ((((sel >> 21u) & 1u) != 0u) ? 0xc1535555u : 0xee02465fu); // 5 add
|
||||
r1 = __umulhi(r1, r7); // 6 mulhi
|
||||
r4 = r4 ^ ds[r2 & mask]; // 7 load
|
||||
r7 = r7 ^ ds[r4 & mask]; // 8 load
|
||||
r0 = r0 ^ ds[r3 & mask]; // 9 load
|
||||
r5 = r5 ^ ds[r1 & mask]; // 10 load
|
||||
r1 = r1 ^ ds[r5 & mask]; // 11 load
|
||||
r3 = __umulhi(r3, r5); // 12 mulhi
|
||||
r1 = r1 ^ ds[r3 & mask]; // 13 load
|
||||
r0 = r0 - r3; // 14 sub
|
||||
r5 = r1 * r3 + r5; // 15 mad
|
||||
r6 = __umulhi(r6, r1); // 16 mulhi
|
||||
r5 = r5 + r2 + ((((sel >> 28u) & 1u) != 0u) ? 0x8b965b57u : 0x697b3d00u); // 17 add
|
||||
r0 = __umulhi(r0, r6); // 18 mulhi
|
||||
r5 = rotr_var(r5, r3); // 19 rotr
|
||||
r5 = __umulhi(r5, r2); // 20 mulhi
|
||||
r1 = r1 + r0 + ((((sel >> 1u) & 1u) != 0u) ? 0x6d7e8d05u : 0xebcf247au); // 21 add
|
||||
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 22 add
|
||||
r1 = __umulhi(r1, r5); // 23 mulhi
|
||||
r2 = r2 - r5; // 24 sub
|
||||
r7 = r7 + r4 + ((((sel >> 2u) & 1u) != 0u) ? 0x699ef1bbu : 0x08ffa6c7u); // 25 add
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 26 shfl
|
||||
r7 = r7 + r1 + ((((sel >> 14u) & 1u) != 0u) ? 0xb4ead2fbu : 0xe60fea84u); // 27 add
|
||||
r3 = r3 + r1 + ((((sel >> 6u) & 1u) != 0u) ? 0x8f30d21du : 0x65c76dabu); // 28 add
|
||||
r2 = r2 ^ ds[r1 & mask]; // 29 load
|
||||
r5 = r5 ^ ds[r7 & mask]; // 30 load
|
||||
r2 = r2 ^ ds[r5 & mask]; // 31 load
|
||||
r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r7, 4); // 32 shfl
|
||||
r4 = r5 * r7 + r4; // 33 mad
|
||||
r4 = r4 + r2 + ((((sel >> 21u) & 1u) != 0u) ? 0xc7ce690cu : 0x0480debeu); // 34 add
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r7, 8); // 35 shfl
|
||||
r7 = r7 + r1 + ((((sel >> 2u) & 1u) != 0u) ? 0xe10c2c95u : 0xc53b542eu); // 36 add
|
||||
r5 = r5 ^ r7; // 37 xor
|
||||
r2 = r2 | r1; // 38 or
|
||||
r1 = __umulhi(r1, r0); // 39 mulhi
|
||||
r6 = rotl_imm(r6, 19u); // 40 rotl
|
||||
r4 = __umulhi(r4, r6); // 41 mulhi
|
||||
r6 = r6 - r0; // 42 sub
|
||||
r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 43 shfl
|
||||
r4 = r4 ^ ds[r2 & mask]; // 44 load
|
||||
r1 = r1 ^ r3; // 45 xor
|
||||
r7 = r7 ^ ds[r0 & mask]; // 46 load
|
||||
r3 = r3 ^ ds[r1 & mask]; // 47 load
|
||||
r5 = r5 * r3; // 48 mul
|
||||
r1 = r1 - r5; // 49 sub
|
||||
r2 = rotl_imm(r2, 8u); // 50 rotl
|
||||
r1 = r1 + r5 + ((((sel >> 23u) & 1u) != 0u) ? 0x77b9bd43u : 0xa900fec4u); // 51 add
|
||||
r4 = r4 ^ ds[r7 & mask]; // 52 load
|
||||
r2 = r2 - r7; // 53 sub
|
||||
r4 = r4 ^ r0; // 54 xor
|
||||
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 55 add
|
||||
r2 = r2 ^ ds[r4 & mask]; // 56 load
|
||||
r0 = r1 * r4 + r0; // 57 mad
|
||||
r3 = r3 ^ ds[r5 & mask]; // 58 load
|
||||
r5 = r5 | r6; // 59 or
|
||||
r6 = r5 * r7 + r6; // 60 mad
|
||||
r4 = rotl_imm(r4, 28u); // 61 rotl
|
||||
r5 = __umulhi(r5, r0); // 62 mulhi
|
||||
r3 = r3 ^ ds[r6 & mask]; // 63 load
|
||||
}
|
||||
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps) {
|
||||
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 32u * blockWarps;
|
||||
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
|
||||
igneum_hash_bound<<<nonces / block, block>>>(ds, out, baseNonce, mask, iw);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
|
||||
cudaFuncAttributes attr;
|
||||
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash_bound);
|
||||
if (e != cudaSuccess) return e;
|
||||
*numRegs = attr.numRegs;
|
||||
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash_bound, (int)(32u * blockWarps), 0);
|
||||
}
|
||||
108
proto-cuda/packs/igneum-devnet-v4-epoch0/memhard.h
Normal file
108
proto-cuda/packs/igneum-devnet-v4-epoch0/memhard.h
Normal file
|
|
@ -0,0 +1,108 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// Memory-hard dataset core, the same text that the Mac's Metal kernels and CPU verifier were checked against.
|
||||
// Included by kernel.cu (device), host.cu (host reference) and proto-opencl/host.c (C99 host reference).
|
||||
// See proto-metal/MEMHARD.md for the construction. kernel.cl carries the same text in OpenCL C.
|
||||
#pragma once
|
||||
#ifdef __cplusplus
|
||||
#include <cstdint>
|
||||
#else
|
||||
#include <stdint.h>
|
||||
#endif
|
||||
#if defined(__CUDACC__)
|
||||
#define IGNEUM_HD __host__ __device__ __forceinline__
|
||||
#elif defined(_MSC_VER) && !defined(__cplusplus)
|
||||
#define IGNEUM_HD static __inline
|
||||
#else
|
||||
#define IGNEUM_HD static inline
|
||||
#endif
|
||||
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
|
||||
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
|
||||
#define MH_CACHE_LINE_MASK 0x003fffffu
|
||||
#define MH_SEGMENT_LINES 64u
|
||||
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
|
||||
IGNEUM_HD uint32_t mh_rotl(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
|
||||
|
||||
// y = ChaCha12 core(x) + x
|
||||
IGNEUM_HD void mh_chacha_block(const uint32_t* x, uint32_t* y) {
|
||||
for (uint32_t i = 0u; i < 16u; ++i) y[i] = x[i];
|
||||
for (uint32_t r = 0u; r < 6u; ++r) {
|
||||
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
|
||||
}
|
||||
for (uint32_t i = 0u; i < 16u; ++i) y[i] += x[i];
|
||||
}
|
||||
|
||||
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
|
||||
IGNEUM_HD void mh_cache_segment(uint32_t* cache, uint32_t seg) {
|
||||
uint32_t prev[16]; uint32_t x[16]; uint32_t y[16];
|
||||
for (uint32_t i = 0u; i < 16u; ++i) prev[i] = 0u;
|
||||
for (uint32_t j = 0u; j < MH_SEGMENT_LINES; ++j) {
|
||||
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
|
||||
x[4] = 0xceed56d7u ^ prev[4];
|
||||
x[5] = 0x9ba270d2u ^ prev[5];
|
||||
x[6] = 0x82caab2du ^ prev[6];
|
||||
x[7] = 0x81ebce0eu ^ prev[7];
|
||||
x[8] = 0x12b6ecf1u ^ prev[8];
|
||||
x[9] = 0xd0f3fd7cu ^ prev[9];
|
||||
x[10] = 0xd872eefeu ^ prev[10];
|
||||
x[11] = 0xc158c7bdu ^ prev[11];
|
||||
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
|
||||
mh_chacha_block(x, y);
|
||||
uint32_t* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
|
||||
for (uint32_t i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
|
||||
}
|
||||
}
|
||||
|
||||
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
|
||||
IGNEUM_HD void mh_mixer(uint32_t* s, uint32_t rk) {
|
||||
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
|
||||
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
|
||||
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
|
||||
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
|
||||
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
|
||||
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
|
||||
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
|
||||
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
|
||||
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
|
||||
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
|
||||
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
|
||||
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
|
||||
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
|
||||
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
|
||||
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
|
||||
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
|
||||
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
|
||||
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
|
||||
}
|
||||
|
||||
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
|
||||
IGNEUM_HD void mh_item(const uint32_t* cache, uint32_t t, uint32_t* s) {
|
||||
s[0] = 0xceed56d7u;
|
||||
s[1] = 0x9ba270d2u;
|
||||
s[2] = 0x82caab2du;
|
||||
s[3] = 0x81ebce0eu;
|
||||
s[4] = 0x12b6ecf1u;
|
||||
s[5] = 0xd0f3fd7cu;
|
||||
s[6] = 0xd872eefeu;
|
||||
s[7] = 0xc158c7bdu;
|
||||
s[8] = t * 0xf351d601u + 0xc6892460u;
|
||||
s[9] = t * 0xa3bb398fu + 0x25b7228au;
|
||||
s[10] = t * 0xb5a09e35u + 0xcd515004u;
|
||||
s[11] = t * 0x7509c9c1u + 0x2846527au;
|
||||
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
|
||||
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
|
||||
s[14] = t * 0xded91851u + 0x82961bacu;
|
||||
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
|
||||
for (uint32_t r = 0u; r < 8u; ++r) {
|
||||
mh_mixer(s, 0x9E3779B9u * (r + 1u));
|
||||
const uint32_t* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
|
||||
for (uint32_t i = 0u; i < 16u; ++i) s[i] ^= line[i];
|
||||
}
|
||||
mh_mixer(s, 0x9E3779B9u * 9u);
|
||||
}
|
||||
// dataset[w] without the dataset: derive item w >> 4 and take word w & 15.
|
||||
IGNEUM_HD uint32_t mh_word(const uint32_t* cache, uint32_t w) { uint32_t s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; }
|
||||
106
proto-cuda/packs/igneum-devnet-v4-epoch0/memhard.metal
Normal file
106
proto-cuda/packs/igneum-devnet-v4-epoch0/memhard.metal
Normal file
|
|
@ -0,0 +1,106 @@
|
|||
#include <metal_stdlib>
|
||||
using namespace metal;
|
||||
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
|
||||
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
|
||||
#define MH_CACHE_LINE_MASK 0x003fffffu
|
||||
#define MH_SEGMENT_LINES 64u
|
||||
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
|
||||
inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
|
||||
|
||||
// y = ChaCha12 core(x) + x
|
||||
inline void mh_chacha_block(const thread uint* x, thread uint* y) {
|
||||
for (uint i = 0u; i < 16u; ++i) y[i] = x[i];
|
||||
for (uint r = 0u; r < 6u; ++r) {
|
||||
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
|
||||
}
|
||||
for (uint i = 0u; i < 16u; ++i) y[i] += x[i];
|
||||
}
|
||||
|
||||
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
|
||||
inline void mh_cache_segment(device uint* cache, uint seg) {
|
||||
uint prev[16]; uint x[16]; uint y[16];
|
||||
for (uint i = 0u; i < 16u; ++i) prev[i] = 0u;
|
||||
for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) {
|
||||
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
|
||||
x[4] = 0xceed56d7u ^ prev[4];
|
||||
x[5] = 0x9ba270d2u ^ prev[5];
|
||||
x[6] = 0x82caab2du ^ prev[6];
|
||||
x[7] = 0x81ebce0eu ^ prev[7];
|
||||
x[8] = 0x12b6ecf1u ^ prev[8];
|
||||
x[9] = 0xd0f3fd7cu ^ prev[9];
|
||||
x[10] = 0xd872eefeu ^ prev[10];
|
||||
x[11] = 0xc158c7bdu ^ prev[11];
|
||||
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
|
||||
mh_chacha_block(x, y);
|
||||
device uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
|
||||
}
|
||||
}
|
||||
|
||||
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
|
||||
inline void mh_mixer(thread uint* s, uint rk) {
|
||||
s[0] = (s[0] ^ (0xc6892460u + rk)) * 0xf351d601u;
|
||||
s[1] = (s[1] ^ (0x25b7228au + rk)) * 0xa3bb398fu;
|
||||
s[2] = (s[2] ^ (0xcd515004u + rk)) * 0xb5a09e35u;
|
||||
s[3] = (s[3] ^ (0x2846527au + rk)) * 0x7509c9c1u;
|
||||
s[4] = (s[4] ^ (0xa6324241u + rk)) * 0x6bbf31e9u;
|
||||
s[5] = (s[5] ^ (0x36e3ec53u + rk)) * 0xfc849a79u;
|
||||
s[6] = (s[6] ^ (0x82961bacu + rk)) * 0xded91851u;
|
||||
s[7] = (s[7] ^ (0x0f97ba7du + rk)) * 0x8d9113d1u;
|
||||
s[8] = (s[8] ^ (0xb6f921a9u + rk)) * 0x0ff15225u;
|
||||
s[9] = (s[9] ^ (0x3ada24e5u + rk)) * 0x3a5bdd41u;
|
||||
s[10] = (s[10] ^ (0xde20ab91u + rk)) * 0xab533435u;
|
||||
s[11] = (s[11] ^ (0x5378eeb2u + rk)) * 0xe1c55ad5u;
|
||||
s[12] = (s[12] ^ (0x7d161662u + rk)) * 0xe6d3bd0du;
|
||||
s[13] = (s[13] ^ (0x89353cc1u + rk)) * 0x9d9ffbbdu;
|
||||
s[14] = (s[14] ^ (0xb1aa03a2u + rk)) * 0xbb2a3cf3u;
|
||||
s[15] = (s[15] ^ (0x788acae6u + rk)) * 0x50a7c08du;
|
||||
MH_QR(s[0], s[4], s[8], s[12], 17u, 12u, 20u, 23u) MH_QR(s[1], s[5], s[9], s[13], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[2], s[6], s[10], s[14], 17u, 12u, 20u, 23u) MH_QR(s[3], s[7], s[11], s[15], 17u, 12u, 20u, 23u)
|
||||
MH_QR(s[0], s[5], s[10], s[15], 7u, 3u, 27u, 16u) MH_QR(s[1], s[6], s[11], s[12], 7u, 3u, 27u, 16u)
|
||||
MH_QR(s[2], s[7], s[8], s[13], 7u, 3u, 27u, 16u) MH_QR(s[3], s[4], s[9], s[14], 7u, 3u, 27u, 16u)
|
||||
}
|
||||
|
||||
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
|
||||
inline void mh_item(device const uint* cache, uint t, thread uint* s) {
|
||||
s[0] = 0xceed56d7u;
|
||||
s[1] = 0x9ba270d2u;
|
||||
s[2] = 0x82caab2du;
|
||||
s[3] = 0x81ebce0eu;
|
||||
s[4] = 0x12b6ecf1u;
|
||||
s[5] = 0xd0f3fd7cu;
|
||||
s[6] = 0xd872eefeu;
|
||||
s[7] = 0xc158c7bdu;
|
||||
s[8] = t * 0xf351d601u + 0xc6892460u;
|
||||
s[9] = t * 0xa3bb398fu + 0x25b7228au;
|
||||
s[10] = t * 0xb5a09e35u + 0xcd515004u;
|
||||
s[11] = t * 0x7509c9c1u + 0x2846527au;
|
||||
s[12] = t * 0x6bbf31e9u + 0xa6324241u;
|
||||
s[13] = t * 0xfc849a79u + 0x36e3ec53u;
|
||||
s[14] = t * 0xded91851u + 0x82961bacu;
|
||||
s[15] = t * 0x8d9113d1u + 0x0f97ba7du;
|
||||
for (uint r = 0u; r < 8u; ++r) {
|
||||
mh_mixer(s, 0x9E3779B9u * (r + 1u));
|
||||
device const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i];
|
||||
}
|
||||
mh_mixer(s, 0x9E3779B9u * 9u);
|
||||
}
|
||||
// dataset[w] without the dataset: derive item w >> 4 and take word w & 15.
|
||||
inline uint mh_word(device const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; }
|
||||
|
||||
// One thread per segment (2^16 threads).
|
||||
kernel void igneum_cache_fill(device uint* cache [[buffer(0)]], uint gid [[thread_position_in_grid]]) {
|
||||
mh_cache_segment(cache, gid);
|
||||
}
|
||||
// One thread per 64-byte item (dataset words / 16 threads).
|
||||
kernel void igneum_build(device const uint* cache [[buffer(0)]], device uint* dataset [[buffer(1)]],
|
||||
uint gid [[thread_position_in_grid]]) {
|
||||
uint s[16];
|
||||
mh_item(cache, gid, s);
|
||||
device uint* d = dataset + gid * 16u;
|
||||
for (uint i = 0u; i < 16u; ++i) d[i] = s[i];
|
||||
}
|
||||
51
proto-cuda/packs/igneum-devnet-v4-epoch0/program.h
Normal file
51
proto-cuda/packs/igneum-devnet-v4-epoch0/program.h
Normal file
|
|
@ -0,0 +1,51 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// Program metadata for host.cu plus the launch wrappers defined in kernel.cu.
|
||||
// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros.
|
||||
#pragma once
|
||||
#ifdef __cplusplus
|
||||
#include <cstdint>
|
||||
#else
|
||||
#include <stdint.h>
|
||||
#endif
|
||||
#ifndef IGNEUM_NO_CUDA
|
||||
#include <cuda_runtime.h>
|
||||
#endif
|
||||
|
||||
#define IGNEUM_SEED_STRING "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000"
|
||||
#define IGNEUM_SEED_BYTES_HEX "edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07"
|
||||
#define IGNEUM_GENERATOR 2
|
||||
#define IGNEUM_PROGRAM_ATTEMPT 0
|
||||
#define IGNEUM_PROGRAM_ID 0x4be132dd1f2ff270ull
|
||||
#define IGNEUM_DAY_STRING "bytes:69676e65756d2d6461792ffa50000000000000"
|
||||
#define IGNEUM_DAY_BYTES_HEX "69676e65756d2d6461792ffa50000000000000"
|
||||
#define IGNEUM_DAY0 0xceed56d7u
|
||||
#define IGNEUM_DAY1 0x9ba270d2u
|
||||
#define IGNEUM_DATASET_LOG2 28
|
||||
#define IGNEUM_MASK 0x0fffffffu
|
||||
#define IGNEUM_LANES 32
|
||||
#define IGNEUM_ITERATIONS 8
|
||||
#define IGNEUM_INSTR_COUNT 64
|
||||
#define IGNEUM_LOADS_PER_HASH 128
|
||||
#define IGNEUM_WIDE_LOADS_PER_HASH 0
|
||||
#define IGNEUM_OP_MIX "load=16 add=13 mulhi=9 shfl=5 sub=5 mad=4 rotl=4 xor=3 or=2 rotr=2 mul=1"
|
||||
// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h)
|
||||
#define IGNEUM_DATASET_MODE 1
|
||||
|
||||
#define IGNEUM_SEEDW_INIT { 0x667d0fbdu, 0x7b8e5963u, 0x31c67e5eu, 0x4529ddc6u, 0xef19d6d8u, 0xaccf6211u, 0xda0aed32u, 0xabc6df31u }
|
||||
#define IGNEUM_KEY_INIT { 0xceed56d7u, 0x9ba270d2u, 0x82caab2du, 0x81ebce0eu, 0x12b6ecf1u, 0xd0f3fd7cu, 0xd872eefeu, 0xc158c7bdu }
|
||||
#define IGNEUM_CACHE_LOG2_WORDS 26
|
||||
#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6
|
||||
#define IGNEUM_CACHE_SEGMENTS 65536u
|
||||
#define IGNEUM_ITEM_ROUNDS 8
|
||||
#define IGNEUM_MIX_ROT_INIT { 17u, 12u, 20u, 23u, 7u, 3u, 27u, 16u }
|
||||
#define IGNEUM_MIX_MUL_INIT { 0xf351d601u, 0xa3bb398fu, 0xb5a09e35u, 0x7509c9c1u, 0x6bbf31e9u, 0xfc849a79u, 0xded91851u, 0x8d9113d1u, 0x0ff15225u, 0x3a5bdd41u, 0xab533435u, 0xe1c55ad5u, 0xe6d3bd0du, 0x9d9ffbbdu, 0xbb2a3cf3u, 0x50a7c08du }
|
||||
#define IGNEUM_MIX_RC_INIT { 0xc6892460u, 0x25b7228au, 0xcd515004u, 0x2846527au, 0xa6324241u, 0x36e3ec53u, 0x82961bacu, 0x0f97ba7du, 0xb6f921a9u, 0x3ada24e5u, 0xde20ab91u, 0x5378eeb2u, 0x7d161662u, 0x89353cc1u, 0xb1aa03a2u, 0x788acae6u }
|
||||
|
||||
#ifndef IGNEUM_NO_CUDA
|
||||
// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError().
|
||||
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments);
|
||||
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems);
|
||||
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
uint32_t nonces, uint32_t blockWarps);
|
||||
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
|
||||
#endif
|
||||
121
proto-cuda/packs/igneum-devnet-v4-epoch0/program.json
Normal file
121
proto-cuda/packs/igneum-devnet-v4-epoch0/program.json
Normal file
|
|
@ -0,0 +1,121 @@
|
|||
{
|
||||
"format": "igneum-program-pack-3",
|
||||
"generator": 2,
|
||||
"attempt": 0,
|
||||
"program_id": "0x4be132dd1f2ff270",
|
||||
"program_id_derivation": "FNV-1a 64 over 'igneum-program/' || generator_le32 || seed_words as little-endian bytes || attempt_le32",
|
||||
"dataset_mode": "memory-hard",
|
||||
"seed": "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000",
|
||||
"seed_bytes": "edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07",
|
||||
"seed_words": ["0x667d0fbd", "0x7b8e5963", "0x31c67e5e", "0x4529ddc6", "0xef19d6d8", "0xaccf6211", "0xda0aed32", "0xabc6df31"],
|
||||
"seed_derivation": "seed_words = FNV-1a 64 over seed_bytes (attempt 0) or seed_bytes || attempt_le32 (attempt k >= 1), basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32",
|
||||
"generator_rule": "version 2: exactly 16 load slots drawn first from instructions 1..63 (partial Fisher-Yates), the other 48 ops from the ten non-load weights (sum 75); a load's source is drawn from the registers other than dst written by an earlier instruction and not read by a load since; the candidate must pass the acceptance rule of spec 01 section 1.4.6 (static: no cyclically stale load source, every register has an injecting write; dynamic: 64 units on the seed-keyed closed-form dataset with no constant register bit, no lane-constant load site, under 164 saturated final values, every output bit within 136 of 1024, distinct addresses above 245760), else the next attempt of the seed is tried",
|
||||
"lanes": 32,
|
||||
"registers": 8,
|
||||
"iterations": 8,
|
||||
"instruction_count": 64,
|
||||
"loads_per_hash": 128,
|
||||
"op_mix": {"load": 16, "add": 13, "mulhi": 9, "shfl": 5, "sub": 5, "mad": 4, "rotl": 4, "xor": 3, "or": 2, "rotr": 2, "mul": 1},
|
||||
"register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]",
|
||||
"splitmix32": "x ^= x>>16; x *= 0x7feb352d; x ^= x>>15; x *= 0x846ca68b; x ^= x>>16",
|
||||
"iteration": "sel = r0 sampled once at the top of each iteration, then all instructions in order",
|
||||
"output": "lo = r0 ^ rotl(r1,7) ^ rotl(r2,14) ^ rotl(r3,21); hi = r4 ^ rotl(r5,9) ^ rotl(r6,18) ^ rotl(r7,27); out = (hi << 32) | lo",
|
||||
"op_semantics": {
|
||||
"add": "dst = dst + src + (bit `bit` of sel ? imm2 : imm)",
|
||||
"sub": "dst = dst - src",
|
||||
"mul": "dst = dst * src (low 32)",
|
||||
"mulhi": "dst = high 32 bits of dst * src",
|
||||
"xor": "dst = dst ^ src",
|
||||
"or": "dst = dst | src",
|
||||
"rotl": "dst = rotl(dst, rot), rot in 1..31",
|
||||
"rotr": "dst = rotr(dst, src & 31)",
|
||||
"mad": "dst = src * src2 + dst",
|
||||
"shfl": "dst = dst ^ (src of lane (lane ^ mask)), mask in {1,2,4,8,16}, within the 32-lane warp",
|
||||
"load": "dst = dst ^ dataset[src & dataset.mask]",
|
||||
"wload": "base = (src of lane 0 & dataset.mask) & ~31; dst = dst ^ dataset[base + lane] (warp-coalesced 128-byte load, lever b, only when --wide-frac > 0)"
|
||||
},
|
||||
"dataset": {
|
||||
"log2_words": 28,
|
||||
"bytes": 1073741824,
|
||||
"mask": "0x0fffffff",
|
||||
"day": "bytes:69676e65756d2d6461792ffa50000000000000",
|
||||
"day_bytes": "69676e65756d2d6461792ffa50000000000000",
|
||||
"day_words_from": "seed_words_from_bytes(day_bytes)",
|
||||
"d0": "0xceed56d7",
|
||||
"d1": "0x9ba270d2",
|
||||
"mode": "memory-hard",
|
||||
"spec": "proto-metal/MEMHARD.md",
|
||||
"key": ["0xceed56d7", "0x9ba270d2", "0x82caab2d", "0x81ebce0e", "0x12b6ecf1", "0xd0f3fd7c", "0xd872eefe", "0xc158c7bd"],
|
||||
"key_derivation": "the 8 words of seed_words_from_bytes(day_bytes); d0, d1 are key[0], key[1]",
|
||||
"cache": {"log2_words": 26, "bytes": 268435456, "line_words": 16, "segment_lines": 64, "segments": 65536, "block": "ChaCha12 core + feed-forward, rotations 16 12 8 7", "sigma": ["0x61707865", "0x3320646e", "0x79622d32", "0x6b206574"], "tag": ["0x49676e65", "0x756d4d48"], "chain": "in_j = prev_line ^ (sigma[0..3] || key[0..7] || seg || j || tag[0..1]); line_j = block(in_j); prev_0 = 0"},
|
||||
"mixer": {"draw": "SplitMix64 seeded with key[0] | key[1] << 32: rot[0..7] = 1 + next() % 31, mul[0..15] = low32(next()) | 1, rc[0..15] = low32(next())", "rot": [17, 12, 20, 23, 7, 3, 27, 16], "mul": ["0xf351d601", "0xa3bb398f", "0xb5a09e35", "0x7509c9c1", "0x6bbf31e9", "0xfc849a79", "0xded91851", "0x8d9113d1", "0x0ff15225", "0x3a5bdd41", "0xab533435", "0xe1c55ad5", "0xe6d3bd0d", "0x9d9ffbbd", "0xbb2a3cf3", "0x50a7c08d"], "rc": ["0xc6892460", "0x25b7228a", "0xcd515004", "0x2846527a", "0xa6324241", "0x36e3ec53", "0x82961bac", "0x0f97ba7d", "0xb6f921a9", "0x3ada24e5", "0xde20ab91", "0x5378eeb2", "0x7d161662", "0x89353cc1", "0xb1aa03a2", "0x788acae6"], "round": "for i in 0..15: s[i] = (s[i] ^ (rc[i] + (r+1) * 0x9E3779B9)) * mul[i]; then quarter rounds on columns (0,4,8,12) (1,5,9,13) (2,6,10,14) (3,7,11,15) with rot[0..3] and diagonals (0,5,10,15) (1,6,11,12) (2,7,8,13) (3,4,9,14) with rot[4..7]", "quarter_round": "a += b; d ^= a; d = rotl(d, r1); c += d; b ^= c; b = rotl(b, r2); a += b; d ^= a; d = rotl(d, r3); c += d; b ^= c; b = rotl(b, r4)"},
|
||||
"item": "s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..7: s = M_r(s); line = s[0] & 0x003fffff; s[i] ^= cache[line * 16 + i]; then s = M_8(s); item(t) = s",
|
||||
"word": "dataset[w] = item(w >> 4)[w & 15]"
|
||||
},
|
||||
"instructions": [
|
||||
{"i": 0, "op": "add", "dst": 4, "src": 5, "src2": 7, "imm": "0xea86e152", "imm2": "0x5810667a", "rot": 27, "bit": 13, "mask": 2},
|
||||
{"i": 1, "op": "shfl", "dst": 2, "src": 0, "src2": 7, "imm": "0xe3c2f9cb", "imm2": "0xde0bea4e", "rot": 6, "bit": 9, "mask": 4},
|
||||
{"i": 2, "op": "add", "dst": 3, "src": 2, "src2": 0, "imm": "0x2cccb6ca", "imm2": "0x642e66db", "rot": 27, "bit": 10, "mask": 2},
|
||||
{"i": 3, "op": "rotl", "dst": 0, "src": 2, "src2": 1, "imm": "0xb59e83b2", "imm2": "0x19c26fb9", "rot": 19, "bit": 20, "mask": 4},
|
||||
{"i": 4, "op": "rotr", "dst": 7, "src": 6, "src2": 5, "imm": "0x6f055f55", "imm2": "0x550e4ea1", "rot": 31, "bit": 20, "mask": 4},
|
||||
{"i": 5, "op": "add", "dst": 7, "src": 4, "src2": 7, "imm": "0xee02465f", "imm2": "0xc1535555", "rot": 31, "bit": 21, "mask": 4},
|
||||
{"i": 6, "op": "mulhi", "dst": 1, "src": 7, "src2": 5, "imm": "0x7826a6a7", "imm2": "0x946f7818", "rot": 18, "bit": 21, "mask": 8},
|
||||
{"i": 7, "op": "load", "dst": 4, "src": 2, "src2": 0, "imm": "0x5d080878", "imm2": "0xdf885578", "rot": 5, "bit": 18, "mask": 4},
|
||||
{"i": 8, "op": "load", "dst": 7, "src": 4, "src2": 1, "imm": "0x875bbb36", "imm2": "0x594a838f", "rot": 24, "bit": 4, "mask": 2},
|
||||
{"i": 9, "op": "load", "dst": 0, "src": 3, "src2": 1, "imm": "0xe5e607c2", "imm2": "0xa5cd9f75", "rot": 26, "bit": 28, "mask": 4},
|
||||
{"i": 10, "op": "load", "dst": 5, "src": 1, "src2": 3, "imm": "0xea3f7b43", "imm2": "0x10c8d4e7", "rot": 5, "bit": 29, "mask": 8},
|
||||
{"i": 11, "op": "load", "dst": 1, "src": 5, "src2": 4, "imm": "0x4454980f", "imm2": "0xebd31581", "rot": 10, "bit": 28, "mask": 16},
|
||||
{"i": 12, "op": "mulhi", "dst": 3, "src": 5, "src2": 7, "imm": "0xbfd5615c", "imm2": "0xd8224ae2", "rot": 21, "bit": 12, "mask": 2},
|
||||
{"i": 13, "op": "load", "dst": 1, "src": 3, "src2": 5, "imm": "0x35c07cc5", "imm2": "0xe82db54f", "rot": 7, "bit": 6, "mask": 8},
|
||||
{"i": 14, "op": "sub", "dst": 0, "src": 3, "src2": 0, "imm": "0x587ee0f1", "imm2": "0xfd23eefd", "rot": 16, "bit": 21, "mask": 16},
|
||||
{"i": 15, "op": "mad", "dst": 5, "src": 1, "src2": 3, "imm": "0x6974dd29", "imm2": "0xc8148960", "rot": 3, "bit": 11, "mask": 2},
|
||||
{"i": 16, "op": "mulhi", "dst": 6, "src": 1, "src2": 7, "imm": "0x3072c3c6", "imm2": "0x55ee21f8", "rot": 26, "bit": 1, "mask": 1},
|
||||
{"i": 17, "op": "add", "dst": 5, "src": 2, "src2": 2, "imm": "0x697b3d00", "imm2": "0x8b965b57", "rot": 9, "bit": 28, "mask": 1},
|
||||
{"i": 18, "op": "mulhi", "dst": 0, "src": 6, "src2": 3, "imm": "0x2910cacb", "imm2": "0x6ac79431", "rot": 7, "bit": 6, "mask": 4},
|
||||
{"i": 19, "op": "rotr", "dst": 5, "src": 3, "src2": 0, "imm": "0xed96a94a", "imm2": "0x4c988c10", "rot": 24, "bit": 21, "mask": 8},
|
||||
{"i": 20, "op": "mulhi", "dst": 5, "src": 2, "src2": 7, "imm": "0x40d2fc76", "imm2": "0x2f7c7eca", "rot": 13, "bit": 8, "mask": 4},
|
||||
{"i": 21, "op": "add", "dst": 1, "src": 0, "src2": 5, "imm": "0xebcf247a", "imm2": "0x6d7e8d05", "rot": 31, "bit": 1, "mask": 16},
|
||||
{"i": 22, "op": "add", "dst": 7, "src": 5, "src2": 7, "imm": "0xf66e7017", "imm2": "0xb9e3577e", "rot": 9, "bit": 12, "mask": 16},
|
||||
{"i": 23, "op": "mulhi", "dst": 1, "src": 5, "src2": 7, "imm": "0xfdfe72fd", "imm2": "0x735eeb8d", "rot": 30, "bit": 25, "mask": 8},
|
||||
{"i": 24, "op": "sub", "dst": 2, "src": 5, "src2": 0, "imm": "0x32cf1258", "imm2": "0xd813deb6", "rot": 30, "bit": 6, "mask": 16},
|
||||
{"i": 25, "op": "add", "dst": 7, "src": 4, "src2": 2, "imm": "0x08ffa6c7", "imm2": "0x699ef1bb", "rot": 7, "bit": 2, "mask": 16},
|
||||
{"i": 26, "op": "shfl", "dst": 3, "src": 4, "src2": 5, "imm": "0x6f53c70d", "imm2": "0x3357513f", "rot": 3, "bit": 26, "mask": 2},
|
||||
{"i": 27, "op": "add", "dst": 7, "src": 1, "src2": 4, "imm": "0xe60fea84", "imm2": "0xb4ead2fb", "rot": 14, "bit": 14, "mask": 1},
|
||||
{"i": 28, "op": "add", "dst": 3, "src": 1, "src2": 0, "imm": "0x65c76dab", "imm2": "0x8f30d21d", "rot": 24, "bit": 6, "mask": 1},
|
||||
{"i": 29, "op": "load", "dst": 2, "src": 1, "src2": 2, "imm": "0x82fad9a6", "imm2": "0x8c6358db", "rot": 7, "bit": 31, "mask": 2},
|
||||
{"i": 30, "op": "load", "dst": 5, "src": 7, "src2": 3, "imm": "0x6e947ee0", "imm2": "0xaf9a2dda", "rot": 2, "bit": 30, "mask": 4},
|
||||
{"i": 31, "op": "load", "dst": 2, "src": 5, "src2": 1, "imm": "0x608bb7ce", "imm2": "0x4be663db", "rot": 19, "bit": 1, "mask": 16},
|
||||
{"i": 32, "op": "shfl", "dst": 1, "src": 7, "src2": 6, "imm": "0x88e52e20", "imm2": "0x77647269", "rot": 20, "bit": 25, "mask": 4},
|
||||
{"i": 33, "op": "mad", "dst": 4, "src": 5, "src2": 7, "imm": "0xb48420ae", "imm2": "0x3d1f2485", "rot": 14, "bit": 28, "mask": 1},
|
||||
{"i": 34, "op": "add", "dst": 4, "src": 2, "src2": 3, "imm": "0x0480debe", "imm2": "0xc7ce690c", "rot": 12, "bit": 21, "mask": 16},
|
||||
{"i": 35, "op": "shfl", "dst": 3, "src": 7, "src2": 5, "imm": "0xa73f59de", "imm2": "0x84b9e329", "rot": 21, "bit": 27, "mask": 8},
|
||||
{"i": 36, "op": "add", "dst": 7, "src": 1, "src2": 7, "imm": "0xc53b542e", "imm2": "0xe10c2c95", "rot": 21, "bit": 2, "mask": 4},
|
||||
{"i": 37, "op": "xor", "dst": 5, "src": 7, "src2": 4, "imm": "0x81cd7b0e", "imm2": "0x21a51823", "rot": 12, "bit": 10, "mask": 2},
|
||||
{"i": 38, "op": "or", "dst": 2, "src": 1, "src2": 3, "imm": "0x7894e657", "imm2": "0xf8e4b972", "rot": 18, "bit": 9, "mask": 4},
|
||||
{"i": 39, "op": "mulhi", "dst": 1, "src": 0, "src2": 4, "imm": "0xbee8421f", "imm2": "0x070888a8", "rot": 20, "bit": 28, "mask": 4},
|
||||
{"i": 40, "op": "rotl", "dst": 6, "src": 1, "src2": 4, "imm": "0x4609857a", "imm2": "0xaeecb156", "rot": 19, "bit": 21, "mask": 1},
|
||||
{"i": 41, "op": "mulhi", "dst": 4, "src": 6, "src2": 0, "imm": "0x1a84e1e9", "imm2": "0x9b26bb72", "rot": 28, "bit": 19, "mask": 1},
|
||||
{"i": 42, "op": "sub", "dst": 6, "src": 0, "src2": 6, "imm": "0xbf62908e", "imm2": "0xdf03ea88", "rot": 11, "bit": 27, "mask": 16},
|
||||
{"i": 43, "op": "shfl", "dst": 6, "src": 3, "src2": 4, "imm": "0xa966241c", "imm2": "0x9c639efa", "rot": 12, "bit": 4, "mask": 4},
|
||||
{"i": 44, "op": "load", "dst": 4, "src": 2, "src2": 3, "imm": "0x7b5b5474", "imm2": "0x45cfc5dd", "rot": 17, "bit": 18, "mask": 8},
|
||||
{"i": 45, "op": "xor", "dst": 1, "src": 3, "src2": 6, "imm": "0x006193d0", "imm2": "0xfc2acc3f", "rot": 25, "bit": 1, "mask": 1},
|
||||
{"i": 46, "op": "load", "dst": 7, "src": 0, "src2": 6, "imm": "0x4777c4f8", "imm2": "0x3cf0a02f", "rot": 11, "bit": 14, "mask": 4},
|
||||
{"i": 47, "op": "load", "dst": 3, "src": 1, "src2": 3, "imm": "0x10692532", "imm2": "0x1a292ea5", "rot": 20, "bit": 23, "mask": 1},
|
||||
{"i": 48, "op": "mul", "dst": 5, "src": 3, "src2": 7, "imm": "0xadf5bd13", "imm2": "0xb999de2e", "rot": 23, "bit": 10, "mask": 4},
|
||||
{"i": 49, "op": "sub", "dst": 1, "src": 5, "src2": 4, "imm": "0x68ff101e", "imm2": "0xbdaaf46a", "rot": 25, "bit": 23, "mask": 16},
|
||||
{"i": 50, "op": "rotl", "dst": 2, "src": 6, "src2": 4, "imm": "0x92d9a412", "imm2": "0x0daf96ea", "rot": 8, "bit": 4, "mask": 4},
|
||||
{"i": 51, "op": "add", "dst": 1, "src": 5, "src2": 0, "imm": "0xa900fec4", "imm2": "0x77b9bd43", "rot": 6, "bit": 23, "mask": 2},
|
||||
{"i": 52, "op": "load", "dst": 4, "src": 7, "src2": 1, "imm": "0x51392a72", "imm2": "0x99e8bb36", "rot": 11, "bit": 9, "mask": 1},
|
||||
{"i": 53, "op": "sub", "dst": 2, "src": 7, "src2": 2, "imm": "0x0ffe2ac7", "imm2": "0x030743df", "rot": 9, "bit": 30, "mask": 16},
|
||||
{"i": 54, "op": "xor", "dst": 4, "src": 0, "src2": 4, "imm": "0x9123ff15", "imm2": "0x10c329a7", "rot": 28, "bit": 15, "mask": 2},
|
||||
{"i": 55, "op": "add", "dst": 1, "src": 6, "src2": 5, "imm": "0xe09f54e9", "imm2": "0x83e825bf", "rot": 23, "bit": 14, "mask": 8},
|
||||
{"i": 56, "op": "load", "dst": 2, "src": 4, "src2": 3, "imm": "0x239de52c", "imm2": "0xf80bae18", "rot": 15, "bit": 24, "mask": 1},
|
||||
{"i": 57, "op": "mad", "dst": 0, "src": 1, "src2": 4, "imm": "0x31dede8e", "imm2": "0xd3f619e6", "rot": 29, "bit": 7, "mask": 2},
|
||||
{"i": 58, "op": "load", "dst": 3, "src": 5, "src2": 3, "imm": "0x206437d6", "imm2": "0x28d1c290", "rot": 17, "bit": 28, "mask": 4},
|
||||
{"i": 59, "op": "or", "dst": 5, "src": 6, "src2": 3, "imm": "0x8fffd674", "imm2": "0x0507903a", "rot": 26, "bit": 27, "mask": 2},
|
||||
{"i": 60, "op": "mad", "dst": 6, "src": 5, "src2": 7, "imm": "0xf572bdb9", "imm2": "0xeda2af31", "rot": 21, "bit": 8, "mask": 2},
|
||||
{"i": 61, "op": "rotl", "dst": 4, "src": 2, "src2": 7, "imm": "0x84f12ddf", "imm2": "0x81ef22e1", "rot": 28, "bit": 30, "mask": 1},
|
||||
{"i": 62, "op": "mulhi", "dst": 5, "src": 0, "src2": 6, "imm": "0xf6bb45ee", "imm2": "0x6bfb632d", "rot": 22, "bit": 0, "mask": 4},
|
||||
{"i": 63, "op": "load", "dst": 3, "src": 6, "src2": 6, "imm": "0x50ec702a", "imm2": "0xae6ee96e", "rot": 20, "bit": 25, "mask": 2}
|
||||
]
|
||||
}
|
||||
109
proto-cuda/packs/igneum-devnet-v4-epoch0/program.metal
Normal file
109
proto-cuda/packs/igneum-devnet-v4-epoch0/program.metal
Normal file
|
|
@ -0,0 +1,109 @@
|
|||
#include <metal_stdlib>
|
||||
using namespace metal;
|
||||
|
||||
#define MASK 0x0fffffffu
|
||||
constant uint SEEDW[8] = { 0x667d0fbdu, 0x7b8e5963u, 0x31c67e5eu, 0x4529ddc6u, 0xef19d6d8u, 0xaccf6211u, 0xda0aed32u, 0xabc6df31u };
|
||||
|
||||
inline uint splitmix32(uint x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31
|
||||
inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
inline uint ds_elem(uint i, uint d0, uint d1) {
|
||||
uint x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
kernel void igneum_hash(device const uint* dataset [[buffer(0)]],
|
||||
device ulong* out [[buffer(1)]],
|
||||
constant uint& baseNonce [[buffer(2)]],
|
||||
uint gid [[thread_position_in_grid]]) {
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint x = nonce ^ SEEDW[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ SEEDW[1]; }
|
||||
{ uint x = nonce ^ SEEDW[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ SEEDW[2]; }
|
||||
{ uint x = nonce ^ SEEDW[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ SEEDW[3]; }
|
||||
{ uint x = nonce ^ SEEDW[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ SEEDW[4]; }
|
||||
{ uint x = nonce ^ SEEDW[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ SEEDW[5]; }
|
||||
{ uint x = nonce ^ SEEDW[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ SEEDW[6]; }
|
||||
{ uint x = nonce ^ SEEDW[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ SEEDW[7]; }
|
||||
{ uint x = nonce ^ SEEDW[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ SEEDW[0]; }
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = r4 + r5 + select(0xea86e152u, 0x5810667au, ((sel >> 13u) & 1u) != 0u); // 0
|
||||
r2 = r2 ^ simd_shuffle_xor(r0, (ushort)4); // 1
|
||||
r3 = r3 + r2 + select(0x2cccb6cau, 0x642e66dbu, ((sel >> 10u) & 1u) != 0u); // 2
|
||||
r0 = rotl_imm(r0, 19u); // 3
|
||||
r7 = rotr_var(r7, r6); // 4
|
||||
r7 = r7 + r4 + select(0xee02465fu, 0xc1535555u, ((sel >> 21u) & 1u) != 0u); // 5
|
||||
r1 = mulhi(r1, r7); // 6
|
||||
r4 = r4 ^ dataset[r2 & MASK]; // 7
|
||||
r7 = r7 ^ dataset[r4 & MASK]; // 8
|
||||
r0 = r0 ^ dataset[r3 & MASK]; // 9
|
||||
r5 = r5 ^ dataset[r1 & MASK]; // 10
|
||||
r1 = r1 ^ dataset[r5 & MASK]; // 11
|
||||
r3 = mulhi(r3, r5); // 12
|
||||
r1 = r1 ^ dataset[r3 & MASK]; // 13
|
||||
r0 = r0 - r3; // 14
|
||||
r5 = r1 * r3 + r5; // 15
|
||||
r6 = mulhi(r6, r1); // 16
|
||||
r5 = r5 + r2 + select(0x697b3d00u, 0x8b965b57u, ((sel >> 28u) & 1u) != 0u); // 17
|
||||
r0 = mulhi(r0, r6); // 18
|
||||
r5 = rotr_var(r5, r3); // 19
|
||||
r5 = mulhi(r5, r2); // 20
|
||||
r1 = r1 + r0 + select(0xebcf247au, 0x6d7e8d05u, ((sel >> 1u) & 1u) != 0u); // 21
|
||||
r7 = r7 + r5 + select(0xf66e7017u, 0xb9e3577eu, ((sel >> 12u) & 1u) != 0u); // 22
|
||||
r1 = mulhi(r1, r5); // 23
|
||||
r2 = r2 - r5; // 24
|
||||
r7 = r7 + r4 + select(0x08ffa6c7u, 0x699ef1bbu, ((sel >> 2u) & 1u) != 0u); // 25
|
||||
r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 26
|
||||
r7 = r7 + r1 + select(0xe60fea84u, 0xb4ead2fbu, ((sel >> 14u) & 1u) != 0u); // 27
|
||||
r3 = r3 + r1 + select(0x65c76dabu, 0x8f30d21du, ((sel >> 6u) & 1u) != 0u); // 28
|
||||
r2 = r2 ^ dataset[r1 & MASK]; // 29
|
||||
r5 = r5 ^ dataset[r7 & MASK]; // 30
|
||||
r2 = r2 ^ dataset[r5 & MASK]; // 31
|
||||
r1 = r1 ^ simd_shuffle_xor(r7, (ushort)4); // 32
|
||||
r4 = r5 * r7 + r4; // 33
|
||||
r4 = r4 + r2 + select(0x0480debeu, 0xc7ce690cu, ((sel >> 21u) & 1u) != 0u); // 34
|
||||
r3 = r3 ^ simd_shuffle_xor(r7, (ushort)8); // 35
|
||||
r7 = r7 + r1 + select(0xc53b542eu, 0xe10c2c95u, ((sel >> 2u) & 1u) != 0u); // 36
|
||||
r5 = r5 ^ r7; // 37
|
||||
r2 = r2 | r1; // 38
|
||||
r1 = mulhi(r1, r0); // 39
|
||||
r6 = rotl_imm(r6, 19u); // 40
|
||||
r4 = mulhi(r4, r6); // 41
|
||||
r6 = r6 - r0; // 42
|
||||
r6 = r6 ^ simd_shuffle_xor(r3, (ushort)4); // 43
|
||||
r4 = r4 ^ dataset[r2 & MASK]; // 44
|
||||
r1 = r1 ^ r3; // 45
|
||||
r7 = r7 ^ dataset[r0 & MASK]; // 46
|
||||
r3 = r3 ^ dataset[r1 & MASK]; // 47
|
||||
r5 = r5 * r3; // 48
|
||||
r1 = r1 - r5; // 49
|
||||
r2 = rotl_imm(r2, 8u); // 50
|
||||
r1 = r1 + r5 + select(0xa900fec4u, 0x77b9bd43u, ((sel >> 23u) & 1u) != 0u); // 51
|
||||
r4 = r4 ^ dataset[r7 & MASK]; // 52
|
||||
r2 = r2 - r7; // 53
|
||||
r4 = r4 ^ r0; // 54
|
||||
r1 = r1 + r6 + select(0xe09f54e9u, 0x83e825bfu, ((sel >> 14u) & 1u) != 0u); // 55
|
||||
r2 = r2 ^ dataset[r4 & MASK]; // 56
|
||||
r0 = r1 * r4 + r0; // 57
|
||||
r3 = r3 ^ dataset[r5 & MASK]; // 58
|
||||
r5 = r5 | r6; // 59
|
||||
r6 = r5 * r7 + r6; // 60
|
||||
r4 = rotl_imm(r4, 28u); // 61
|
||||
r5 = mulhi(r5, r0); // 62
|
||||
r3 = r3 ^ dataset[r6 & MASK]; // 63
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
111
proto-cuda/packs/igneum-devnet-v4-epoch0/program_bound.metal
Normal file
111
proto-cuda/packs/igneum-devnet-v4-epoch0/program_bound.metal
Normal file
|
|
@ -0,0 +1,111 @@
|
|||
#include <metal_stdlib>
|
||||
using namespace metal;
|
||||
|
||||
#define MASK 0x0fffffffu
|
||||
constant uint SEEDW[8] = { 0x667d0fbdu, 0x7b8e5963u, 0x31c67e5eu, 0x4529ddc6u, 0xef19d6d8u, 0xaccf6211u, 0xda0aed32u, 0xabc6df31u };
|
||||
|
||||
inline uint splitmix32(uint x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31
|
||||
inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
inline uint ds_elem(uint i, uint d0, uint d1) {
|
||||
uint x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// Header-bound variant: the init words come from buffer 3 (bind.rs), not from SEEDW.
|
||||
kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]],
|
||||
device ulong* out [[buffer(1)]],
|
||||
constant uint& baseNonce [[buffer(2)]],
|
||||
constant uint* initw [[buffer(3)]],
|
||||
uint gid [[thread_position_in_grid]]) {
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint x = nonce ^ initw[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ initw[1]; }
|
||||
{ uint x = nonce ^ initw[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ initw[2]; }
|
||||
{ uint x = nonce ^ initw[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ initw[3]; }
|
||||
{ uint x = nonce ^ initw[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ initw[4]; }
|
||||
{ uint x = nonce ^ initw[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ initw[5]; }
|
||||
{ uint x = nonce ^ initw[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ initw[6]; }
|
||||
{ uint x = nonce ^ initw[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ initw[7]; }
|
||||
{ uint x = nonce ^ initw[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ initw[0]; }
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = r4 + r5 + select(0xea86e152u, 0x5810667au, ((sel >> 13u) & 1u) != 0u); // 0
|
||||
r2 = r2 ^ simd_shuffle_xor(r0, (ushort)4); // 1
|
||||
r3 = r3 + r2 + select(0x2cccb6cau, 0x642e66dbu, ((sel >> 10u) & 1u) != 0u); // 2
|
||||
r0 = rotl_imm(r0, 19u); // 3
|
||||
r7 = rotr_var(r7, r6); // 4
|
||||
r7 = r7 + r4 + select(0xee02465fu, 0xc1535555u, ((sel >> 21u) & 1u) != 0u); // 5
|
||||
r1 = mulhi(r1, r7); // 6
|
||||
r4 = r4 ^ dataset[r2 & MASK]; // 7
|
||||
r7 = r7 ^ dataset[r4 & MASK]; // 8
|
||||
r0 = r0 ^ dataset[r3 & MASK]; // 9
|
||||
r5 = r5 ^ dataset[r1 & MASK]; // 10
|
||||
r1 = r1 ^ dataset[r5 & MASK]; // 11
|
||||
r3 = mulhi(r3, r5); // 12
|
||||
r1 = r1 ^ dataset[r3 & MASK]; // 13
|
||||
r0 = r0 - r3; // 14
|
||||
r5 = r1 * r3 + r5; // 15
|
||||
r6 = mulhi(r6, r1); // 16
|
||||
r5 = r5 + r2 + select(0x697b3d00u, 0x8b965b57u, ((sel >> 28u) & 1u) != 0u); // 17
|
||||
r0 = mulhi(r0, r6); // 18
|
||||
r5 = rotr_var(r5, r3); // 19
|
||||
r5 = mulhi(r5, r2); // 20
|
||||
r1 = r1 + r0 + select(0xebcf247au, 0x6d7e8d05u, ((sel >> 1u) & 1u) != 0u); // 21
|
||||
r7 = r7 + r5 + select(0xf66e7017u, 0xb9e3577eu, ((sel >> 12u) & 1u) != 0u); // 22
|
||||
r1 = mulhi(r1, r5); // 23
|
||||
r2 = r2 - r5; // 24
|
||||
r7 = r7 + r4 + select(0x08ffa6c7u, 0x699ef1bbu, ((sel >> 2u) & 1u) != 0u); // 25
|
||||
r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 26
|
||||
r7 = r7 + r1 + select(0xe60fea84u, 0xb4ead2fbu, ((sel >> 14u) & 1u) != 0u); // 27
|
||||
r3 = r3 + r1 + select(0x65c76dabu, 0x8f30d21du, ((sel >> 6u) & 1u) != 0u); // 28
|
||||
r2 = r2 ^ dataset[r1 & MASK]; // 29
|
||||
r5 = r5 ^ dataset[r7 & MASK]; // 30
|
||||
r2 = r2 ^ dataset[r5 & MASK]; // 31
|
||||
r1 = r1 ^ simd_shuffle_xor(r7, (ushort)4); // 32
|
||||
r4 = r5 * r7 + r4; // 33
|
||||
r4 = r4 + r2 + select(0x0480debeu, 0xc7ce690cu, ((sel >> 21u) & 1u) != 0u); // 34
|
||||
r3 = r3 ^ simd_shuffle_xor(r7, (ushort)8); // 35
|
||||
r7 = r7 + r1 + select(0xc53b542eu, 0xe10c2c95u, ((sel >> 2u) & 1u) != 0u); // 36
|
||||
r5 = r5 ^ r7; // 37
|
||||
r2 = r2 | r1; // 38
|
||||
r1 = mulhi(r1, r0); // 39
|
||||
r6 = rotl_imm(r6, 19u); // 40
|
||||
r4 = mulhi(r4, r6); // 41
|
||||
r6 = r6 - r0; // 42
|
||||
r6 = r6 ^ simd_shuffle_xor(r3, (ushort)4); // 43
|
||||
r4 = r4 ^ dataset[r2 & MASK]; // 44
|
||||
r1 = r1 ^ r3; // 45
|
||||
r7 = r7 ^ dataset[r0 & MASK]; // 46
|
||||
r3 = r3 ^ dataset[r1 & MASK]; // 47
|
||||
r5 = r5 * r3; // 48
|
||||
r1 = r1 - r5; // 49
|
||||
r2 = rotl_imm(r2, 8u); // 50
|
||||
r1 = r1 + r5 + select(0xa900fec4u, 0x77b9bd43u, ((sel >> 23u) & 1u) != 0u); // 51
|
||||
r4 = r4 ^ dataset[r7 & MASK]; // 52
|
||||
r2 = r2 - r7; // 53
|
||||
r4 = r4 ^ r0; // 54
|
||||
r1 = r1 + r6 + select(0xe09f54e9u, 0x83e825bfu, ((sel >> 14u) & 1u) != 0u); // 55
|
||||
r2 = r2 ^ dataset[r4 & MASK]; // 56
|
||||
r0 = r1 * r4 + r0; // 57
|
||||
r3 = r3 ^ dataset[r5 & MASK]; // 58
|
||||
r5 = r5 | r6; // 59
|
||||
r6 = r5 * r7 + r6; // 60
|
||||
r4 = rotl_imm(r4, 28u); // 61
|
||||
r5 = mulhi(r5, r0); // 62
|
||||
r3 = r3 ^ dataset[r6 & MASK]; // 63
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
57
proto-cuda/packs/igneum-devnet-v4-epoch0/vectors.h
Normal file
57
proto-cuda/packs/igneum-devnet-v4-epoch0/vectors.h
Normal file
|
|
@ -0,0 +1,57 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
|
||||
// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset
|
||||
#pragma once
|
||||
#ifdef __cplusplus
|
||||
#include <cstdint>
|
||||
#else
|
||||
#include <stdint.h>
|
||||
#endif
|
||||
|
||||
#define IGNEUM_VEC_WARPS 3
|
||||
static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u };
|
||||
static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = {
|
||||
{ // base nonce 0
|
||||
0x285a83011e7ac3fcull, 0x5b6c4419a08ec70dull, 0x09b91d7b7ef0635aull, 0x1d7db2088b88e203ull, 0x364313aca5e0271bull, 0xb266d17bac4d2071ull, 0x7cfaea561ba1f79bull, 0x68907e8f5ace2dc2ull,
|
||||
0xfd61773b9d2444b1ull, 0x07f611eb83a9d282ull, 0x28ba6343cfb5fb02ull, 0xf64c793df2070ea4ull, 0x4c4b07824107ee4bull, 0x61f065a8548cbc0bull, 0x85beb9e842013806ull, 0x0908a7f6bfa688ccull,
|
||||
0xa802bdd4f062079full, 0x317cadd2d682e176ull, 0xba55bf9b76b2201full, 0x99fd2ea48ffa2982ull, 0x8ef077e8f7e5ea8cull, 0x86b97137129886c8ull, 0x97ec8b3ef390ea2cull, 0xaa31e9669ab72bfeull,
|
||||
0x8ccc96b63128e884ull, 0xf4bc0a18d985287eull, 0x532827f2e3531a1bull, 0x4d1479139bf17cc2ull, 0x0c41d6a96efd6f54ull, 0x61c5034c68f70786ull, 0x7fb2a60f5e211e67ull, 0x6f1136558c20e3d8ull
|
||||
},
|
||||
{ // base nonce 4096
|
||||
0xae0f5d1ac66966f6ull, 0xb335bec4d4590f48ull, 0xd74cb94a4a5dfa77ull, 0x4601059d3c526c70ull, 0xa81accb2016844a2ull, 0x2fdfee2240eb5b04ull, 0xc1209b4978bbf44cull, 0x9bd881612e285c5eull,
|
||||
0xf0accd0fee8bcb0full, 0x9b44d0702b454f27ull, 0xf6e6905cc5fa4355ull, 0x1180e75d94566648ull, 0xd5f41e4e108ab48dull, 0x3a7dbb2f2dff9f37ull, 0xaa2cb75a0022a1acull, 0x0bb58f47a32918d3ull,
|
||||
0x4b91e06435ab3c0cull, 0x2c94509d82bb9949ull, 0x9152ca4b8686422full, 0xdcebdd18080d6460ull, 0x93355a1421212c39ull, 0xcb434978558a7e56ull, 0x1f4c7d3f6827d7ecull, 0x9598835d863b9752ull,
|
||||
0x24b6d66a4b9d3b26ull, 0x6f73888e0bd77b1eull, 0xc4862e608103b587ull, 0xf423c5afc15c73feull, 0x781a3c1ca923598dull, 0xb43b71d378787341ull, 0xc72c42a31199c4a7ull, 0x0b580b0e42a0a42eull
|
||||
},
|
||||
{ // base nonce 1000000
|
||||
0xedb431758b34810cull, 0x27637ad21af42a3bull, 0x194b178a9203463cull, 0x9c650cd30a69bab8ull, 0xb2f52ad5d695a2e3ull, 0x771a9bc87187745full, 0xa1e7249f093bb167ull, 0x4fed709a7e2f19c4ull,
|
||||
0x4c3c778a686b5c37ull, 0x89f6e0d8b69f1a75ull, 0xfe2a3a1c7adf5c3bull, 0x0bf31eee62d6336full, 0x66a79fbaa66b4484ull, 0x436eac23e3d68eeaull, 0xadcb3040ba4c9313ull, 0x3919c3e0b09904aeull,
|
||||
0xa4b9d51a181a9c5cull, 0x9e8a465be8776ae5ull, 0x190910061da1eec1ull, 0x3d9f5f0ea1335bceull, 0x7c61fed158b1872bull, 0x8627027cd58d409cull, 0x098d036a81fa7ab8ull, 0xe25bbb24c3a135b9ull,
|
||||
0x59bab56f04f0fde2ull, 0x919621c583cd59beull, 0x95ca061dca251d5aull, 0xbda544905079a010ull, 0xd0cb1cf3fe8a6aa9ull, 0xb43efee3c24ee26aull, 0x1ce3a2494f68f693ull, 0x87cc652fe6d87bb8ull
|
||||
}
|
||||
};
|
||||
|
||||
// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455).
|
||||
static const uint32_t IGNEUM_DS_HEAD[16] = {
|
||||
0x3dd50b1fu, 0x48edec90u, 0x5119540eu, 0x657ac748u, 0x090717d5u, 0x5181125eu, 0x9ba8b16eu, 0x1843d8f4u,
|
||||
0x43f523fbu, 0xd7194d6fu, 0xae29ea68u, 0x182ecb6au, 0x8dfe0a23u, 0xba291b20u, 0x995fbedeu, 0x15b5fab6u
|
||||
};
|
||||
static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u;
|
||||
static const uint32_t IGNEUM_DS_LAST = 0xf7b7180eu;
|
||||
// 64 sampled dataset words (index, value) computed on the Mac.
|
||||
#define IGNEUM_DS_SAMPLES 64
|
||||
static const uint32_t IGNEUM_DS_SAMPLE_INDEX[IGNEUM_DS_SAMPLES] = {
|
||||
59471966u, 217795994u, 208353206u, 42483309u, 172547758u, 148076330u, 183853158u, 214389424u, 267488061u, 169781097u, 184093494u, 153880993u, 84977930u, 46426879u, 3093825u, 225364072u, 44593546u, 260713159u, 168250303u, 52384140u, 223401610u, 45554030u, 95410555u, 175039924u, 79171087u, 267580473u, 24168642u, 37981670u, 171551130u, 195559979u, 204611762u, 140997658u, 138925853u, 86637313u, 20736778u, 219665210u, 160430336u, 264654675u, 8013395u, 228945585u, 213884386u, 104419827u, 44185464u, 142737231u, 99284897u, 132475900u, 61861762u, 132056166u, 262388043u, 91878046u, 117353561u, 124768597u, 71352993u, 190698941u, 46055428u, 55281366u, 165145231u, 106810753u, 171985651u, 232085256u, 159510492u, 40072060u, 209107596u, 39023794u
|
||||
};
|
||||
static const uint32_t IGNEUM_DS_SAMPLE_VALUE[IGNEUM_DS_SAMPLES] = {
|
||||
0x300c5208u, 0x70d7f6afu, 0xbad7ce7cu, 0xa763a1cbu, 0x6a6cb69du, 0x7cc540a7u, 0x4d0a5a4fu, 0x6228f1d9u, 0x2ca22cbfu, 0xb7646a0du, 0x8c39fdaeu, 0xbfd11737u, 0x51d66d78u, 0xb3b3f153u, 0x75685180u, 0x4ad13cc3u, 0x1fff336cu, 0x683cb116u, 0xfcc9bddau, 0x9720e59bu, 0xc9ae54dfu, 0x09d5d5beu, 0xc3d0bbc5u, 0x005b5ca5u, 0x679d372du, 0x66bec2abu, 0x06bab271u, 0xe7e42a5du, 0x269b17e8u, 0x587e10f4u, 0x29698e73u, 0x6319943cu, 0x8ccf60abu, 0xd7cce064u, 0xce77c60eu, 0x3c3caef9u, 0x9369aa85u, 0x6d748b5eu, 0x85192c35u, 0x983de049u, 0xa9e017d5u, 0x99bc5f97u, 0xe17c3e02u, 0xc3b2634bu, 0x59b00e33u, 0x54e1bfd1u, 0x06e2d5f8u, 0xcf862800u, 0xeb4bfa4au, 0x2066805cu, 0x6032e02cu, 0x9fbd6780u, 0x53f8e649u, 0x50d4e5eeu, 0x3d239e6cu, 0x23d9566cu, 0xe3031b82u, 0x8489e44du, 0x5cb39f0bu, 0xa91c131cu, 0x20833a1bu, 0xc26997c6u, 0x904455bdu, 0x8402782au
|
||||
};
|
||||
// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^26 words.
|
||||
static const uint32_t IGNEUM_CACHE_HEAD[16] = {
|
||||
0xebd9055cu, 0x7eaf6b21u, 0x9b610855u, 0x134a8562u, 0xb10059bau, 0x7f459e58u, 0x8f3c7873u, 0x3523e609u,
|
||||
0x75b8f4cau, 0x750d31b4u, 0x0dae0782u, 0x26f10015u, 0x7ba87a41u, 0x0efd8543u, 0x691d6368u, 0x8929d967u
|
||||
};
|
||||
static const uint32_t IGNEUM_CACHE_LAST[16] = {
|
||||
0x51474edfu, 0xc3cfba93u, 0xf21454cfu, 0x9b79baacu, 0x4d4fce23u, 0xd701dfd5u, 0x37357ab3u, 0x1be693fau,
|
||||
0xa701cb3bu, 0x7467c620u, 0x428184e8u, 0xf4010df0u, 0xe33aa88fu, 0xc78d6d62u, 0xae9ba9c0u, 0xf4bba97eu
|
||||
};
|
||||
static const uint64_t IGNEUM_CACHE_FNV64 = 0x448274a57f508cbcull;
|
||||
36
proto-cuda/packs/igneum-devnet-v4-epoch0/vectors.json
Normal file
36
proto-cuda/packs/igneum-devnet-v4-epoch0/vectors.json
Normal file
|
|
@ -0,0 +1,36 @@
|
|||
{
|
||||
"seed": "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000",
|
||||
"day": "bytes:69676e65756d2d6461792ffa50000000000000",
|
||||
"dataset_mode": "memory-hard",
|
||||
"dataset_log2_words": 28,
|
||||
"mask": "0x0fffffff",
|
||||
"lanes": 32,
|
||||
"source": "igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset",
|
||||
"warps": [
|
||||
{"base_nonce": 0, "expected": [
|
||||
"0x285a83011e7ac3fc", "0x5b6c4419a08ec70d", "0x09b91d7b7ef0635a", "0x1d7db2088b88e203", "0x364313aca5e0271b", "0xb266d17bac4d2071", "0x7cfaea561ba1f79b", "0x68907e8f5ace2dc2",
|
||||
"0xfd61773b9d2444b1", "0x07f611eb83a9d282", "0x28ba6343cfb5fb02", "0xf64c793df2070ea4", "0x4c4b07824107ee4b", "0x61f065a8548cbc0b", "0x85beb9e842013806", "0x0908a7f6bfa688cc",
|
||||
"0xa802bdd4f062079f", "0x317cadd2d682e176", "0xba55bf9b76b2201f", "0x99fd2ea48ffa2982", "0x8ef077e8f7e5ea8c", "0x86b97137129886c8", "0x97ec8b3ef390ea2c", "0xaa31e9669ab72bfe",
|
||||
"0x8ccc96b63128e884", "0xf4bc0a18d985287e", "0x532827f2e3531a1b", "0x4d1479139bf17cc2", "0x0c41d6a96efd6f54", "0x61c5034c68f70786", "0x7fb2a60f5e211e67", "0x6f1136558c20e3d8"
|
||||
]},
|
||||
{"base_nonce": 4096, "expected": [
|
||||
"0xae0f5d1ac66966f6", "0xb335bec4d4590f48", "0xd74cb94a4a5dfa77", "0x4601059d3c526c70", "0xa81accb2016844a2", "0x2fdfee2240eb5b04", "0xc1209b4978bbf44c", "0x9bd881612e285c5e",
|
||||
"0xf0accd0fee8bcb0f", "0x9b44d0702b454f27", "0xf6e6905cc5fa4355", "0x1180e75d94566648", "0xd5f41e4e108ab48d", "0x3a7dbb2f2dff9f37", "0xaa2cb75a0022a1ac", "0x0bb58f47a32918d3",
|
||||
"0x4b91e06435ab3c0c", "0x2c94509d82bb9949", "0x9152ca4b8686422f", "0xdcebdd18080d6460", "0x93355a1421212c39", "0xcb434978558a7e56", "0x1f4c7d3f6827d7ec", "0x9598835d863b9752",
|
||||
"0x24b6d66a4b9d3b26", "0x6f73888e0bd77b1e", "0xc4862e608103b587", "0xf423c5afc15c73fe", "0x781a3c1ca923598d", "0xb43b71d378787341", "0xc72c42a31199c4a7", "0x0b580b0e42a0a42e"
|
||||
]},
|
||||
{"base_nonce": 1000000, "expected": [
|
||||
"0xedb431758b34810c", "0x27637ad21af42a3b", "0x194b178a9203463c", "0x9c650cd30a69bab8", "0xb2f52ad5d695a2e3", "0x771a9bc87187745f", "0xa1e7249f093bb167", "0x4fed709a7e2f19c4",
|
||||
"0x4c3c778a686b5c37", "0x89f6e0d8b69f1a75", "0xfe2a3a1c7adf5c3b", "0x0bf31eee62d6336f", "0x66a79fbaa66b4484", "0x436eac23e3d68eea", "0xadcb3040ba4c9313", "0x3919c3e0b09904ae",
|
||||
"0xa4b9d51a181a9c5c", "0x9e8a465be8776ae5", "0x190910061da1eec1", "0x3d9f5f0ea1335bce", "0x7c61fed158b1872b", "0x8627027cd58d409c", "0x098d036a81fa7ab8", "0xe25bbb24c3a135b9",
|
||||
"0x59bab56f04f0fde2", "0x919621c583cd59be", "0x95ca061dca251d5a", "0xbda544905079a010", "0xd0cb1cf3fe8a6aa9", "0xb43efee3c24ee26a", "0x1ce3a2494f68f693", "0x87cc652fe6d87bb8"
|
||||
]}
|
||||
],
|
||||
"dataset_head": ["0x3dd50b1f", "0x48edec90", "0x5119540e", "0x657ac748", "0x090717d5", "0x5181125e", "0x9ba8b16e", "0x1843d8f4", "0x43f523fb", "0xd7194d6f", "0xae29ea68", "0x182ecb6a", "0x8dfe0a23", "0xba291b20", "0x995fbede", "0x15b5fab6"],
|
||||
"dataset_last_index": 268435455,
|
||||
"dataset_last": "0xf7b7180e",
|
||||
"dataset_samples": [{"index": 59471966, "value": "0x300c5208"}, {"index": 217795994, "value": "0x70d7f6af"}, {"index": 208353206, "value": "0xbad7ce7c"}, {"index": 42483309, "value": "0xa763a1cb"}, {"index": 172547758, "value": "0x6a6cb69d"}, {"index": 148076330, "value": "0x7cc540a7"}, {"index": 183853158, "value": "0x4d0a5a4f"}, {"index": 214389424, "value": "0x6228f1d9"}, {"index": 267488061, "value": "0x2ca22cbf"}, {"index": 169781097, "value": "0xb7646a0d"}, {"index": 184093494, "value": "0x8c39fdae"}, {"index": 153880993, "value": "0xbfd11737"}, {"index": 84977930, "value": "0x51d66d78"}, {"index": 46426879, "value": "0xb3b3f153"}, {"index": 3093825, "value": "0x75685180"}, {"index": 225364072, "value": "0x4ad13cc3"}, {"index": 44593546, "value": "0x1fff336c"}, {"index": 260713159, "value": "0x683cb116"}, {"index": 168250303, "value": "0xfcc9bdda"}, {"index": 52384140, "value": "0x9720e59b"}, {"index": 223401610, "value": "0xc9ae54df"}, {"index": 45554030, "value": "0x09d5d5be"}, {"index": 95410555, "value": "0xc3d0bbc5"}, {"index": 175039924, "value": "0x005b5ca5"}, {"index": 79171087, "value": "0x679d372d"}, {"index": 267580473, "value": "0x66bec2ab"}, {"index": 24168642, "value": "0x06bab271"}, {"index": 37981670, "value": "0xe7e42a5d"}, {"index": 171551130, "value": "0x269b17e8"}, {"index": 195559979, "value": "0x587e10f4"}, {"index": 204611762, "value": "0x29698e73"}, {"index": 140997658, "value": "0x6319943c"}, {"index": 138925853, "value": "0x8ccf60ab"}, {"index": 86637313, "value": "0xd7cce064"}, {"index": 20736778, "value": "0xce77c60e"}, {"index": 219665210, "value": "0x3c3caef9"}, {"index": 160430336, "value": "0x9369aa85"}, {"index": 264654675, "value": "0x6d748b5e"}, {"index": 8013395, "value": "0x85192c35"}, {"index": 228945585, "value": "0x983de049"}, {"index": 213884386, "value": "0xa9e017d5"}, {"index": 104419827, "value": "0x99bc5f97"}, {"index": 44185464, "value": "0xe17c3e02"}, {"index": 142737231, "value": "0xc3b2634b"}, {"index": 99284897, "value": "0x59b00e33"}, {"index": 132475900, "value": "0x54e1bfd1"}, {"index": 61861762, "value": "0x06e2d5f8"}, {"index": 132056166, "value": "0xcf862800"}, {"index": 262388043, "value": "0xeb4bfa4a"}, {"index": 91878046, "value": "0x2066805c"}, {"index": 117353561, "value": "0x6032e02c"}, {"index": 124768597, "value": "0x9fbd6780"}, {"index": 71352993, "value": "0x53f8e649"}, {"index": 190698941, "value": "0x50d4e5ee"}, {"index": 46055428, "value": "0x3d239e6c"}, {"index": 55281366, "value": "0x23d9566c"}, {"index": 165145231, "value": "0xe3031b82"}, {"index": 106810753, "value": "0x8489e44d"}, {"index": 171985651, "value": "0x5cb39f0b"}, {"index": 232085256, "value": "0xa91c131c"}, {"index": 159510492, "value": "0x20833a1b"}, {"index": 40072060, "value": "0xc26997c6"}, {"index": 209107596, "value": "0x904455bd"}, {"index": 39023794, "value": "0x8402782a"}],
|
||||
"cache_head": ["0xebd9055c", "0x7eaf6b21", "0x9b610855", "0x134a8562", "0xb10059ba", "0x7f459e58", "0x8f3c7873", "0x3523e609", "0x75b8f4ca", "0x750d31b4", "0x0dae0782", "0x26f10015", "0x7ba87a41", "0x0efd8543", "0x691d6368", "0x8929d967"],
|
||||
"cache_last_line": ["0x51474edf", "0xc3cfba93", "0xf21454cf", "0x9b79baac", "0x4d4fce23", "0xd701dfd5", "0x37357ab3", "0x1be693fa", "0xa701cb3b", "0x7467c620", "0x428184e8", "0xf4010df0", "0xe33aa88f", "0xc78d6d62", "0xae9ba9c0", "0xf4bba97e"],
|
||||
"cache_fnv1a64": "0x448274a57f508cbc"
|
||||
}
|
||||
|
|
@ -1,4 +1,4 @@
|
|||
// Generated by proto-metal/igneum-bench --export-pack for seed "igneum-genesis". Do not edit by hand.
|
||||
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
|
||||
// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal).
|
||||
// Built from source at runtime by proto-opencl/host.c, which passes these defines:
|
||||
// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit)
|
||||
|
|
@ -198,70 +198,70 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out
|
|||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = rotl_imm(r4, 25u); // 0 rotl
|
||||
r0 = r0 - r5; // 1 sub
|
||||
r4 = r4 ^ ds[r3 & mask]; // 2 load
|
||||
r1 = rotl_imm(r1, 1u); // 3 rotl
|
||||
r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 4 add
|
||||
r5 = r5 ^ ds[r3 & mask]; // 5 load
|
||||
r5 = r5 - r7; // 6 sub
|
||||
r3 = r3 + r4 + ((((sel >> 26u) & 1u) != 0u) ? 0x5a069596u : 0x52f2dbf4u); // 7 add
|
||||
r0 = r0 ^ r4; // 8 xor
|
||||
r4 = r4 ^ r0; // 9 xor
|
||||
r2 = r2 ^ ds[r0 & mask]; // 10 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 16u); r4 = r4 ^ t_; } // 11 shfl
|
||||
r1 = r1 - r5; // 12 sub
|
||||
r2 = r2 ^ r1; // 13 xor
|
||||
r4 = r4 ^ ds[r5 & mask]; // 14 load
|
||||
r2 = r2 ^ ds[r4 & mask]; // 15 load
|
||||
r3 = r3 ^ ds[r0 & mask]; // 16 load
|
||||
r4 = r4 ^ r6; // 17 xor
|
||||
r2 = r4 * r6 + r2; // 18 mad
|
||||
r6 = rotr_var(r6, r1); // 19 rotr
|
||||
r3 = r3 ^ r4; // 20 xor
|
||||
r1 = r3 * r5 + r1; // 21 mad
|
||||
r7 = mul_hi(r7, r4); // 22 mulhi
|
||||
r5 = mul_hi(r5, r2); // 23 mulhi
|
||||
r0 = r0 ^ ds[r6 & mask]; // 24 load
|
||||
r5 = r5 * r6; // 25 mul
|
||||
r7 = r7 ^ ds[r1 & mask]; // 26 load
|
||||
r3 = rotr_var(r3, r1); // 27 rotr
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r1, 8u); r5 = r5 ^ t_; } // 28 shfl
|
||||
r7 = r7 ^ r5; // 29 xor
|
||||
r7 = rotl_imm(r7, 23u); // 30 rotl
|
||||
r2 = r2 - r0; // 31 sub
|
||||
r7 = r7 ^ r2; // 32 xor
|
||||
r2 = r2 ^ r6; // 33 xor
|
||||
r6 = r6 ^ ds[r1 & mask]; // 34 load
|
||||
r1 = r1 ^ r4; // 35 xor
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r0, 1u); r3 = r3 ^ t_; } // 36 shfl
|
||||
r2 = r2 * r6; // 37 mul
|
||||
r5 = r5 + r3 + ((((sel >> 12u) & 1u) != 0u) ? 0xa29f4338u : 0x71f30417u); // 38 add
|
||||
r7 = r7 ^ r6; // 39 xor
|
||||
r7 = r7 ^ r3; // 40 xor
|
||||
r3 = rotr_var(r3, r4); // 41 rotr
|
||||
r5 = r5 ^ r3; // 42 xor
|
||||
r3 = rotr_var(r3, r6); // 43 rotr
|
||||
r1 = r3 * r5 + r1; // 44 mad
|
||||
r7 = r7 + r3 + ((((sel >> 29u) & 1u) != 0u) ? 0xa907b90bu : 0xc1ae8d3bu); // 45 add
|
||||
r7 = r7 | r2; // 46 or
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 16u); r2 = r2 ^ t_; } // 47 shfl
|
||||
r1 = r1 ^ ds[r7 & mask]; // 48 load
|
||||
r5 = r5 - r1; // 49 sub
|
||||
r3 = r3 ^ ds[r1 & mask]; // 50 load
|
||||
r2 = r2 - r3; // 51 sub
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r2, 2u); r6 = r6 ^ t_; } // 52 shfl
|
||||
r2 = r2 - r0; // 53 sub
|
||||
r0 = r0 ^ ds[r3 & mask]; // 54 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r1, 16u); r2 = r2 ^ t_; } // 55 shfl
|
||||
r0 = r0 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0xf4689674u : 0x25955401u); // 56 add
|
||||
r2 = mul_hi(r2, r0); // 57 mulhi
|
||||
r4 = mul_hi(r4, r2); // 58 mulhi
|
||||
r2 = r6 * r7 + r2; // 59 mad
|
||||
r3 = r3 ^ r1; // 60 xor
|
||||
r4 = r4 * r2; // 61 mul
|
||||
r0 = r0 ^ ds[r2 & mask]; // 62 load
|
||||
r7 = mul_hi(r7, r0); // 63 mulhi
|
||||
r2 = r3 * r4 + r2; // 0 mad
|
||||
r2 = r1 * r1 + r2; // 1 mad
|
||||
r2 = r3 * r2 + r2; // 2 mad
|
||||
r3 = r3 ^ r5; // 3 xor
|
||||
r7 = r7 ^ ds[r2 & mask]; // 4 load
|
||||
r5 = r5 ^ ds[r7 & mask]; // 5 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r1 = r1 ^ t_; } // 6 shfl
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 8u); r7 = r7 ^ t_; } // 7 shfl
|
||||
r1 = mul_hi(r1, r5); // 8 mulhi
|
||||
r6 = rotr_var(r6, r3); // 9 rotr
|
||||
r3 = r3 | r4; // 10 or
|
||||
r4 = r4 ^ ds[r3 & mask]; // 11 load
|
||||
r0 = mul_hi(r0, r4); // 12 mulhi
|
||||
r5 = r5 + r1 + ((((sel >> 30u) & 1u) != 0u) ? 0xd3177981u : 0xc7934706u); // 13 add
|
||||
r0 = r0 ^ ds[r4 & mask]; // 14 load
|
||||
r2 = r2 - r4; // 15 sub
|
||||
r2 = r2 ^ ds[r0 & mask]; // 16 load
|
||||
r7 = r7 ^ ds[r2 & mask]; // 17 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 18 shfl
|
||||
r5 = r5 * r0; // 19 mul
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 20 shfl
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 16u); r2 = r2 ^ t_; } // 21 shfl
|
||||
r6 = mul_hi(r6, r2); // 22 mulhi
|
||||
r6 = r6 ^ ds[r1 & mask]; // 23 load
|
||||
r5 = r5 * r0; // 24 mul
|
||||
r5 = rotl_imm(r5, 19u); // 25 rotl
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 2u); r7 = r7 ^ t_; } // 26 shfl
|
||||
r0 = r0 ^ r5; // 27 xor
|
||||
r0 = r0 ^ r4; // 28 xor
|
||||
r3 = r3 - r0; // 29 sub
|
||||
r5 = r5 * r1; // 30 mul
|
||||
r7 = r7 ^ ds[r2 & mask]; // 31 load
|
||||
r1 = r1 ^ ds[r0 & mask]; // 32 load
|
||||
r5 = r5 ^ r6; // 33 xor
|
||||
r5 = r5 ^ ds[r1 & mask]; // 34 load
|
||||
r0 = mul_hi(r0, r5); // 35 mulhi
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r2, 4u); r5 = r5 ^ t_; } // 36 shfl
|
||||
r7 = r7 ^ ds[r0 & mask]; // 37 load
|
||||
r3 = r3 + r1 + ((((sel >> 27u) & 1u) != 0u) ? 0x230c005cu : 0x75ba2fadu); // 38 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 4u); r1 = r1 ^ t_; } // 39 shfl
|
||||
r2 = r2 ^ r5; // 40 xor
|
||||
r3 = r6 * r3 + r3; // 41 mad
|
||||
r6 = r6 - r7; // 42 sub
|
||||
r7 = r7 ^ r0; // 43 xor
|
||||
r1 = r1 ^ ds[r7 & mask]; // 44 load
|
||||
r2 = r2 * r3; // 45 mul
|
||||
r1 = mul_hi(r1, r5); // 46 mulhi
|
||||
r4 = r4 - r3; // 47 sub
|
||||
r2 = rotr_var(r2, r6); // 48 rotr
|
||||
r3 = r3 ^ ds[r5 & mask]; // 49 load
|
||||
r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 50 add
|
||||
r0 = r0 * r2; // 51 mul
|
||||
r0 = r0 + r2 + ((((sel >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); // 52 add
|
||||
r1 = r1 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0x77b1520du : 0x2bb965afu); // 53 add
|
||||
r7 = rotl_imm(r7, 14u); // 54 rotl
|
||||
r3 = r3 + r7 + ((((sel >> 1u) & 1u) != 0u) ? 0xa54c55a0u : 0x7b0fe07au); // 55 add
|
||||
r6 = r6 ^ ds[r7 & mask]; // 56 load
|
||||
r1 = rotr_var(r1, r5); // 57 rotr
|
||||
r5 = r5 ^ ds[r4 & mask]; // 58 load
|
||||
r6 = r6 ^ ds[r2 & mask]; // 59 load
|
||||
r3 = r5 * r0 + r3; // 60 mad
|
||||
r5 = r5 + r7 + ((((sel >> 31u) & 1u) != 0u) ? 0xad7493e7u : 0xaf9dd72du); // 61 add
|
||||
r4 = r4 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); // 62 add
|
||||
r5 = rotl_imm(r5, 19u); // 63 rotl
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
|
|
|
|||
|
|
@ -1,4 +1,4 @@
|
|||
// Generated by proto-metal/igneum-bench --export-pack for seed "igneum-genesis". Do not edit by hand.
|
||||
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
|
||||
// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal).
|
||||
// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC.
|
||||
#include <cuda_runtime.h>
|
||||
|
|
@ -59,70 +59,70 @@ __global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonc
|
|||
|
||||
for (uint32_t it = 0u; it < 8u; ++it) {
|
||||
uint32_t sel = r0;
|
||||
r4 = rotl_imm(r4, 25u); // 0 rotl
|
||||
r0 = r0 - r5; // 1 sub
|
||||
r4 = r4 ^ ds[r3 & mask]; // 2 load
|
||||
r1 = rotl_imm(r1, 1u); // 3 rotl
|
||||
r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 4 add
|
||||
r5 = r5 ^ ds[r3 & mask]; // 5 load
|
||||
r5 = r5 - r7; // 6 sub
|
||||
r3 = r3 + r4 + ((((sel >> 26u) & 1u) != 0u) ? 0x5a069596u : 0x52f2dbf4u); // 7 add
|
||||
r0 = r0 ^ r4; // 8 xor
|
||||
r4 = r4 ^ r0; // 9 xor
|
||||
r2 = r2 ^ ds[r0 & mask]; // 10 load
|
||||
r4 = r4 ^ __shfl_xor_sync(0xffffffffu, r6, 16); // 11 shfl
|
||||
r1 = r1 - r5; // 12 sub
|
||||
r2 = r2 ^ r1; // 13 xor
|
||||
r4 = r4 ^ ds[r5 & mask]; // 14 load
|
||||
r2 = r2 ^ ds[r4 & mask]; // 15 load
|
||||
r3 = r3 ^ ds[r0 & mask]; // 16 load
|
||||
r4 = r4 ^ r6; // 17 xor
|
||||
r2 = r4 * r6 + r2; // 18 mad
|
||||
r6 = rotr_var(r6, r1); // 19 rotr
|
||||
r3 = r3 ^ r4; // 20 xor
|
||||
r1 = r3 * r5 + r1; // 21 mad
|
||||
r7 = __umulhi(r7, r4); // 22 mulhi
|
||||
r5 = __umulhi(r5, r2); // 23 mulhi
|
||||
r0 = r0 ^ ds[r6 & mask]; // 24 load
|
||||
r5 = r5 * r6; // 25 mul
|
||||
r7 = r7 ^ ds[r1 & mask]; // 26 load
|
||||
r3 = rotr_var(r3, r1); // 27 rotr
|
||||
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r1, 8); // 28 shfl
|
||||
r7 = r7 ^ r5; // 29 xor
|
||||
r7 = rotl_imm(r7, 23u); // 30 rotl
|
||||
r2 = r2 - r0; // 31 sub
|
||||
r7 = r7 ^ r2; // 32 xor
|
||||
r2 = r2 ^ r6; // 33 xor
|
||||
r6 = r6 ^ ds[r1 & mask]; // 34 load
|
||||
r1 = r1 ^ r4; // 35 xor
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r0, 1); // 36 shfl
|
||||
r2 = r2 * r6; // 37 mul
|
||||
r5 = r5 + r3 + ((((sel >> 12u) & 1u) != 0u) ? 0xa29f4338u : 0x71f30417u); // 38 add
|
||||
r7 = r7 ^ r6; // 39 xor
|
||||
r7 = r7 ^ r3; // 40 xor
|
||||
r3 = rotr_var(r3, r4); // 41 rotr
|
||||
r5 = r5 ^ r3; // 42 xor
|
||||
r3 = rotr_var(r3, r6); // 43 rotr
|
||||
r1 = r3 * r5 + r1; // 44 mad
|
||||
r7 = r7 + r3 + ((((sel >> 29u) & 1u) != 0u) ? 0xa907b90bu : 0xc1ae8d3bu); // 45 add
|
||||
r7 = r7 | r2; // 46 or
|
||||
r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r3, 16); // 47 shfl
|
||||
r1 = r1 ^ ds[r7 & mask]; // 48 load
|
||||
r5 = r5 - r1; // 49 sub
|
||||
r3 = r3 ^ ds[r1 & mask]; // 50 load
|
||||
r2 = r2 - r3; // 51 sub
|
||||
r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r2, 2); // 52 shfl
|
||||
r2 = r2 - r0; // 53 sub
|
||||
r0 = r0 ^ ds[r3 & mask]; // 54 load
|
||||
r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r1, 16); // 55 shfl
|
||||
r0 = r0 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0xf4689674u : 0x25955401u); // 56 add
|
||||
r2 = __umulhi(r2, r0); // 57 mulhi
|
||||
r4 = __umulhi(r4, r2); // 58 mulhi
|
||||
r2 = r6 * r7 + r2; // 59 mad
|
||||
r3 = r3 ^ r1; // 60 xor
|
||||
r4 = r4 * r2; // 61 mul
|
||||
r0 = r0 ^ ds[r2 & mask]; // 62 load
|
||||
r7 = __umulhi(r7, r0); // 63 mulhi
|
||||
r2 = r3 * r4 + r2; // 0 mad
|
||||
r2 = r1 * r1 + r2; // 1 mad
|
||||
r2 = r3 * r2 + r2; // 2 mad
|
||||
r3 = r3 ^ r5; // 3 xor
|
||||
r7 = r7 ^ ds[r2 & mask]; // 4 load
|
||||
r5 = r5 ^ ds[r7 & mask]; // 5 load
|
||||
r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 6 shfl
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 8); // 7 shfl
|
||||
r1 = __umulhi(r1, r5); // 8 mulhi
|
||||
r6 = rotr_var(r6, r3); // 9 rotr
|
||||
r3 = r3 | r4; // 10 or
|
||||
r4 = r4 ^ ds[r3 & mask]; // 11 load
|
||||
r0 = __umulhi(r0, r4); // 12 mulhi
|
||||
r5 = r5 + r1 + ((((sel >> 30u) & 1u) != 0u) ? 0xd3177981u : 0xc7934706u); // 13 add
|
||||
r0 = r0 ^ ds[r4 & mask]; // 14 load
|
||||
r2 = r2 - r4; // 15 sub
|
||||
r2 = r2 ^ ds[r0 & mask]; // 16 load
|
||||
r7 = r7 ^ ds[r2 & mask]; // 17 load
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 18 shfl
|
||||
r5 = r5 * r0; // 19 mul
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 20 shfl
|
||||
r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r4, 16); // 21 shfl
|
||||
r6 = __umulhi(r6, r2); // 22 mulhi
|
||||
r6 = r6 ^ ds[r1 & mask]; // 23 load
|
||||
r5 = r5 * r0; // 24 mul
|
||||
r5 = rotl_imm(r5, 19u); // 25 rotl
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r6, 2); // 26 shfl
|
||||
r0 = r0 ^ r5; // 27 xor
|
||||
r0 = r0 ^ r4; // 28 xor
|
||||
r3 = r3 - r0; // 29 sub
|
||||
r5 = r5 * r1; // 30 mul
|
||||
r7 = r7 ^ ds[r2 & mask]; // 31 load
|
||||
r1 = r1 ^ ds[r0 & mask]; // 32 load
|
||||
r5 = r5 ^ r6; // 33 xor
|
||||
r5 = r5 ^ ds[r1 & mask]; // 34 load
|
||||
r0 = __umulhi(r0, r5); // 35 mulhi
|
||||
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r2, 4); // 36 shfl
|
||||
r7 = r7 ^ ds[r0 & mask]; // 37 load
|
||||
r3 = r3 + r1 + ((((sel >> 27u) & 1u) != 0u) ? 0x230c005cu : 0x75ba2fadu); // 38 add
|
||||
r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r5, 4); // 39 shfl
|
||||
r2 = r2 ^ r5; // 40 xor
|
||||
r3 = r6 * r3 + r3; // 41 mad
|
||||
r6 = r6 - r7; // 42 sub
|
||||
r7 = r7 ^ r0; // 43 xor
|
||||
r1 = r1 ^ ds[r7 & mask]; // 44 load
|
||||
r2 = r2 * r3; // 45 mul
|
||||
r1 = __umulhi(r1, r5); // 46 mulhi
|
||||
r4 = r4 - r3; // 47 sub
|
||||
r2 = rotr_var(r2, r6); // 48 rotr
|
||||
r3 = r3 ^ ds[r5 & mask]; // 49 load
|
||||
r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 50 add
|
||||
r0 = r0 * r2; // 51 mul
|
||||
r0 = r0 + r2 + ((((sel >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); // 52 add
|
||||
r1 = r1 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0x77b1520du : 0x2bb965afu); // 53 add
|
||||
r7 = rotl_imm(r7, 14u); // 54 rotl
|
||||
r3 = r3 + r7 + ((((sel >> 1u) & 1u) != 0u) ? 0xa54c55a0u : 0x7b0fe07au); // 55 add
|
||||
r6 = r6 ^ ds[r7 & mask]; // 56 load
|
||||
r1 = rotr_var(r1, r5); // 57 rotr
|
||||
r5 = r5 ^ ds[r4 & mask]; // 58 load
|
||||
r6 = r6 ^ ds[r2 & mask]; // 59 load
|
||||
r3 = r5 * r0 + r3; // 60 mad
|
||||
r5 = r5 + r7 + ((((sel >> 31u) & 1u) != 0u) ? 0xad7493e7u : 0xaf9dd72du); // 61 add
|
||||
r4 = r4 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); // 62 add
|
||||
r5 = rotl_imm(r5, 19u); // 63 rotl
|
||||
}
|
||||
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
|
|
|
|||
372
proto-cuda/packs/igneum-genesis-mh/kernel_bound.cl
Normal file
372
proto-cuda/packs/igneum-genesis-mh/kernel_bound.cl
Normal file
|
|
@ -0,0 +1,372 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
|
||||
// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal).
|
||||
// Built from source at runtime by proto-opencl/host.c, which passes these defines:
|
||||
// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit)
|
||||
// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default)
|
||||
// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32
|
||||
// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition
|
||||
// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units;
|
||||
// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32.
|
||||
#ifndef IGNEUM_GROUP
|
||||
#define IGNEUM_GROUP 32
|
||||
#endif
|
||||
#ifndef IGNEUM_EXCHANGE
|
||||
#define IGNEUM_EXCHANGE 0
|
||||
#endif
|
||||
#ifdef __OPENCL_VERSION__
|
||||
#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1)))
|
||||
#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n]
|
||||
#if IGNEUM_EXCHANGE == 1
|
||||
#ifdef cl_khr_subgroups
|
||||
#pragma OPENCL EXTENSION cl_khr_subgroups : enable
|
||||
#endif
|
||||
#ifdef cl_khr_subgroup_shuffle
|
||||
#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable
|
||||
#endif
|
||||
#elif IGNEUM_EXCHANGE == 2
|
||||
#pragma OPENCL EXTENSION cl_intel_subgroups : enable
|
||||
#endif
|
||||
#else
|
||||
// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros.
|
||||
#include "emu_opencl.h"
|
||||
#endif
|
||||
|
||||
#if IGNEUM_EXCHANGE == 1
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m))
|
||||
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
|
||||
#elif IGNEUM_EXCHANGE == 2
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m))
|
||||
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
|
||||
#else
|
||||
// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per
|
||||
// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane
|
||||
// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's
|
||||
// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier.
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; }
|
||||
#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; }
|
||||
#endif
|
||||
|
||||
static inline uint splitmix32(uint x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32.
|
||||
static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); }
|
||||
// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x.
|
||||
static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); }
|
||||
static inline uint ds_elem(uint i, uint d0, uint d1) {
|
||||
uint x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// Memory-hard dataset core (MEMHARD.md). Cache: 2^26 words in 2^16 segments of 64 chained ChaCha12 lines.
|
||||
// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.
|
||||
#define MH_CACHE_LINE_MASK 0x003fffffu
|
||||
#define MH_SEGMENT_LINES 64u
|
||||
#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }
|
||||
static inline uint mh_rotl(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31 at every call site
|
||||
|
||||
// y = ChaCha12 core(x) + x
|
||||
static inline void mh_chacha_block(const uint* x, uint* y) {
|
||||
for (uint i = 0u; i < 16u; ++i) y[i] = x[i];
|
||||
for (uint r = 0u; r < 6u; ++r) {
|
||||
MH_QR(y[0], y[4], y[8], y[12], 16u, 12u, 8u, 7u) MH_QR(y[1], y[5], y[9], y[13], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[6], y[10], y[14], 16u, 12u, 8u, 7u) MH_QR(y[3], y[7], y[11], y[15], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[0], y[5], y[10], y[15], 16u, 12u, 8u, 7u) MH_QR(y[1], y[6], y[11], y[12], 16u, 12u, 8u, 7u)
|
||||
MH_QR(y[2], y[7], y[8], y[13], 16u, 12u, 8u, 7u) MH_QR(y[3], y[4], y[9], y[14], 16u, 12u, 8u, 7u)
|
||||
}
|
||||
for (uint i = 0u; i < 16u; ++i) y[i] += x[i];
|
||||
}
|
||||
|
||||
// One cache segment: 64 chained lines written at cache[seg * 1024]. in_j = prev ^ (sigma || K || seg || j || tag), prev_0 = 0.
|
||||
static inline void mh_cache_segment(__global uint* cache, uint seg) {
|
||||
uint prev[16]; uint x[16]; uint y[16];
|
||||
for (uint i = 0u; i < 16u; ++i) prev[i] = 0u;
|
||||
for (uint j = 0u; j < MH_SEGMENT_LINES; ++j) {
|
||||
x[0] = 0x61707865u ^ prev[0]; x[1] = 0x3320646eu ^ prev[1]; x[2] = 0x79622d32u ^ prev[2]; x[3] = 0x6b206574u ^ prev[3];
|
||||
x[4] = 0x3067619fu ^ prev[4];
|
||||
x[5] = 0x3c269176u ^ prev[5];
|
||||
x[6] = 0x84a03b03u ^ prev[6];
|
||||
x[7] = 0xf8c63294u ^ prev[7];
|
||||
x[8] = 0xff977c5bu ^ prev[8];
|
||||
x[9] = 0xe60def3eu ^ prev[9];
|
||||
x[10] = 0x63630141u ^ prev[10];
|
||||
x[11] = 0xb8fbcb58u ^ prev[11];
|
||||
x[12] = seg ^ prev[12]; x[13] = j ^ prev[13]; x[14] = 0x49676e65u ^ prev[14]; x[15] = 0x756d4d48u ^ prev[15];
|
||||
mh_chacha_block(x, y);
|
||||
__global uint* line = cache + ((seg * MH_SEGMENT_LINES + j) * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) { line[i] = y[i]; prev[i] = y[i]; }
|
||||
}
|
||||
}
|
||||
|
||||
// M_r: per word (s ^ (RC + rk)) * MUL, then a column round and a diagonal round with the seed-drawn rotations.
|
||||
static inline void mh_mixer(uint* s, uint rk) {
|
||||
s[0] = (s[0] ^ (0xbab68293u + rk)) * 0x42146205u;
|
||||
s[1] = (s[1] ^ (0xcc162340u + rk)) * 0x52cbe0fbu;
|
||||
s[2] = (s[2] ^ (0x6ce151ccu + rk)) * 0x7ecf4a03u;
|
||||
s[3] = (s[3] ^ (0xe62b8997u + rk)) * 0x6728907fu;
|
||||
s[4] = (s[4] ^ (0xc9c80297u + rk)) * 0xd81d9751u;
|
||||
s[5] = (s[5] ^ (0xf74a1654u + rk)) * 0x132952c3u;
|
||||
s[6] = (s[6] ^ (0x3d704af5u + rk)) * 0xf60de277u;
|
||||
s[7] = (s[7] ^ (0x3cf522b7u + rk)) * 0x05358035u;
|
||||
s[8] = (s[8] ^ (0x2b9cac04u + rk)) * 0xbaf6499du;
|
||||
s[9] = (s[9] ^ (0xa880ac10u + rk)) * 0xe4db9667u;
|
||||
s[10] = (s[10] ^ (0x13e5dd1du + rk)) * 0x3e98f45du;
|
||||
s[11] = (s[11] ^ (0x6fc3e233u + rk)) * 0xd0004eddu;
|
||||
s[12] = (s[12] ^ (0x2d83eeacu + rk)) * 0x2691630du;
|
||||
s[13] = (s[13] ^ (0x9006e8bfu + rk)) * 0x9beb3bcfu;
|
||||
s[14] = (s[14] ^ (0x2c4b5362u + rk)) * 0xab310379u;
|
||||
s[15] = (s[15] ^ (0x31b49ee2u + rk)) * 0x99cfb423u;
|
||||
MH_QR(s[0], s[4], s[8], s[12], 20u, 20u, 19u, 4u) MH_QR(s[1], s[5], s[9], s[13], 20u, 20u, 19u, 4u)
|
||||
MH_QR(s[2], s[6], s[10], s[14], 20u, 20u, 19u, 4u) MH_QR(s[3], s[7], s[11], s[15], 20u, 20u, 19u, 4u)
|
||||
MH_QR(s[0], s[5], s[10], s[15], 26u, 3u, 3u, 27u) MH_QR(s[1], s[6], s[11], s[12], 26u, 3u, 3u, 27u)
|
||||
MH_QR(s[2], s[7], s[8], s[13], 26u, 3u, 3u, 27u) MH_QR(s[3], s[4], s[9], s[14], 26u, 3u, 3u, 27u)
|
||||
}
|
||||
|
||||
// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); 8 rounds of mixer + cache line s[0] & mask; final mixer.
|
||||
static inline void mh_item(__global const uint* cache, uint t, uint* s) {
|
||||
s[0] = 0x3067619fu;
|
||||
s[1] = 0x3c269176u;
|
||||
s[2] = 0x84a03b03u;
|
||||
s[3] = 0xf8c63294u;
|
||||
s[4] = 0xff977c5bu;
|
||||
s[5] = 0xe60def3eu;
|
||||
s[6] = 0x63630141u;
|
||||
s[7] = 0xb8fbcb58u;
|
||||
s[8] = t * 0x42146205u + 0xbab68293u;
|
||||
s[9] = t * 0x52cbe0fbu + 0xcc162340u;
|
||||
s[10] = t * 0x7ecf4a03u + 0x6ce151ccu;
|
||||
s[11] = t * 0x6728907fu + 0xe62b8997u;
|
||||
s[12] = t * 0xd81d9751u + 0xc9c80297u;
|
||||
s[13] = t * 0x132952c3u + 0xf74a1654u;
|
||||
s[14] = t * 0xf60de277u + 0x3d704af5u;
|
||||
s[15] = t * 0x05358035u + 0x3cf522b7u;
|
||||
for (uint r = 0u; r < 8u; ++r) {
|
||||
mh_mixer(s, 0x9E3779B9u * (r + 1u));
|
||||
__global const uint* line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) s[i] ^= line[i];
|
||||
}
|
||||
mh_mixer(s, 0x9E3779B9u * 9u);
|
||||
}
|
||||
// dataset[w] without the dataset: derive item w >> 4 and take word w & 15.
|
||||
static inline uint mh_word(__global const uint* cache, uint w) { uint s[16]; mh_item(cache, w >> 4u, s); return s[w & 15u]; }
|
||||
|
||||
// Memory-hard dataset (MEMHARD.md). One work-item per cache segment; one work-item per 64-byte dataset item.
|
||||
// The same constants as memhard.h in this pack (one emitter, three dialects).
|
||||
__kernel void igneum_cache_fill(__global uint* cache, uint nSegments) {
|
||||
uint seg = (uint)get_global_id(0);
|
||||
if (seg < nSegments) mh_cache_segment(cache, seg);
|
||||
}
|
||||
__kernel void igneum_build(__global uint* ds, __global const uint* cache, uint nItems) {
|
||||
uint t = (uint)get_global_id(0);
|
||||
if (t < nItems) {
|
||||
uint s[16];
|
||||
mh_item(cache, t, s);
|
||||
__global uint* d = ds + ((ulong)t * 16u);
|
||||
for (uint i = 0u; i < 16u; ++i) d[i] = s[i];
|
||||
}
|
||||
}
|
||||
|
||||
// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the
|
||||
// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and
|
||||
// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all).
|
||||
IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) {
|
||||
uint gid = (uint)get_global_id(0);
|
||||
uint lid = (uint)get_local_id(0);
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
#if IGNEUM_EXCHANGE == 0
|
||||
IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
|
||||
uint xk = 0u;
|
||||
#else
|
||||
(void)lid;
|
||||
#endif
|
||||
{ uint x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
|
||||
{ uint x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
|
||||
{ uint x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
|
||||
{ uint x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
|
||||
{ uint x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
|
||||
{ uint x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
|
||||
{ uint x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
|
||||
{ uint x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r2 = r3 * r4 + r2; // 0 mad
|
||||
r2 = r1 * r1 + r2; // 1 mad
|
||||
r2 = r3 * r2 + r2; // 2 mad
|
||||
r3 = r3 ^ r5; // 3 xor
|
||||
r7 = r7 ^ ds[r2 & mask]; // 4 load
|
||||
r5 = r5 ^ ds[r7 & mask]; // 5 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r1 = r1 ^ t_; } // 6 shfl
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 8u); r7 = r7 ^ t_; } // 7 shfl
|
||||
r1 = mul_hi(r1, r5); // 8 mulhi
|
||||
r6 = rotr_var(r6, r3); // 9 rotr
|
||||
r3 = r3 | r4; // 10 or
|
||||
r4 = r4 ^ ds[r3 & mask]; // 11 load
|
||||
r0 = mul_hi(r0, r4); // 12 mulhi
|
||||
r5 = r5 + r1 + ((((sel >> 30u) & 1u) != 0u) ? 0xd3177981u : 0xc7934706u); // 13 add
|
||||
r0 = r0 ^ ds[r4 & mask]; // 14 load
|
||||
r2 = r2 - r4; // 15 sub
|
||||
r2 = r2 ^ ds[r0 & mask]; // 16 load
|
||||
r7 = r7 ^ ds[r2 & mask]; // 17 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 18 shfl
|
||||
r5 = r5 * r0; // 19 mul
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 20 shfl
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 16u); r2 = r2 ^ t_; } // 21 shfl
|
||||
r6 = mul_hi(r6, r2); // 22 mulhi
|
||||
r6 = r6 ^ ds[r1 & mask]; // 23 load
|
||||
r5 = r5 * r0; // 24 mul
|
||||
r5 = rotl_imm(r5, 19u); // 25 rotl
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 2u); r7 = r7 ^ t_; } // 26 shfl
|
||||
r0 = r0 ^ r5; // 27 xor
|
||||
r0 = r0 ^ r4; // 28 xor
|
||||
r3 = r3 - r0; // 29 sub
|
||||
r5 = r5 * r1; // 30 mul
|
||||
r7 = r7 ^ ds[r2 & mask]; // 31 load
|
||||
r1 = r1 ^ ds[r0 & mask]; // 32 load
|
||||
r5 = r5 ^ r6; // 33 xor
|
||||
r5 = r5 ^ ds[r1 & mask]; // 34 load
|
||||
r0 = mul_hi(r0, r5); // 35 mulhi
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r2, 4u); r5 = r5 ^ t_; } // 36 shfl
|
||||
r7 = r7 ^ ds[r0 & mask]; // 37 load
|
||||
r3 = r3 + r1 + ((((sel >> 27u) & 1u) != 0u) ? 0x230c005cu : 0x75ba2fadu); // 38 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 4u); r1 = r1 ^ t_; } // 39 shfl
|
||||
r2 = r2 ^ r5; // 40 xor
|
||||
r3 = r6 * r3 + r3; // 41 mad
|
||||
r6 = r6 - r7; // 42 sub
|
||||
r7 = r7 ^ r0; // 43 xor
|
||||
r1 = r1 ^ ds[r7 & mask]; // 44 load
|
||||
r2 = r2 * r3; // 45 mul
|
||||
r1 = mul_hi(r1, r5); // 46 mulhi
|
||||
r4 = r4 - r3; // 47 sub
|
||||
r2 = rotr_var(r2, r6); // 48 rotr
|
||||
r3 = r3 ^ ds[r5 & mask]; // 49 load
|
||||
r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 50 add
|
||||
r0 = r0 * r2; // 51 mul
|
||||
r0 = r0 + r2 + ((((sel >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); // 52 add
|
||||
r1 = r1 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0x77b1520du : 0x2bb965afu); // 53 add
|
||||
r7 = rotl_imm(r7, 14u); // 54 rotl
|
||||
r3 = r3 + r7 + ((((sel >> 1u) & 1u) != 0u) ? 0xa54c55a0u : 0x7b0fe07au); // 55 add
|
||||
r6 = r6 ^ ds[r7 & mask]; // 56 load
|
||||
r1 = rotr_var(r1, r5); // 57 rotr
|
||||
r5 = r5 ^ ds[r4 & mask]; // 58 load
|
||||
r6 = r6 ^ ds[r2 & mask]; // 59 load
|
||||
r3 = r5 * r0 + r3; // 60 mad
|
||||
r5 = r5 + r7 + ((((sel >> 31u) & 1u) != 0u) ? 0xad7493e7u : 0xaf9dd72du); // 61 add
|
||||
r4 = r4 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); // 62 add
|
||||
r5 = rotl_imm(r5, 19u); // 63 rotl
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
|
||||
#if IGNEUM_EXCHANGE != 0
|
||||
// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the
|
||||
// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a
|
||||
// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md.
|
||||
IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) {
|
||||
if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); }
|
||||
}
|
||||
#endif
|
||||
|
||||
// Header-bound variant (bind.rs): the init words come from initw, not SEEDW. Same body as igneum_hash.
|
||||
IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw) {
|
||||
uint gid = (uint)get_global_id(0);
|
||||
uint lid = (uint)get_local_id(0);
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7];
|
||||
#if IGNEUM_EXCHANGE == 0
|
||||
IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
|
||||
uint xk = 0u;
|
||||
#else
|
||||
(void)lid;
|
||||
#endif
|
||||
{ uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; }
|
||||
{ uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; }
|
||||
{ uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; }
|
||||
{ uint x = nonce ^ iw3; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw4; }
|
||||
{ uint x = nonce ^ iw4; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw5; }
|
||||
{ uint x = nonce ^ iw5; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw6; }
|
||||
{ uint x = nonce ^ iw6; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw7; }
|
||||
{ uint x = nonce ^ iw7; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw0; }
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r2 = r3 * r4 + r2; // 0 mad
|
||||
r2 = r1 * r1 + r2; // 1 mad
|
||||
r2 = r3 * r2 + r2; // 2 mad
|
||||
r3 = r3 ^ r5; // 3 xor
|
||||
r7 = r7 ^ ds[r2 & mask]; // 4 load
|
||||
r5 = r5 ^ ds[r7 & mask]; // 5 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r1 = r1 ^ t_; } // 6 shfl
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 8u); r7 = r7 ^ t_; } // 7 shfl
|
||||
r1 = mul_hi(r1, r5); // 8 mulhi
|
||||
r6 = rotr_var(r6, r3); // 9 rotr
|
||||
r3 = r3 | r4; // 10 or
|
||||
r4 = r4 ^ ds[r3 & mask]; // 11 load
|
||||
r0 = mul_hi(r0, r4); // 12 mulhi
|
||||
r5 = r5 + r1 + ((((sel >> 30u) & 1u) != 0u) ? 0xd3177981u : 0xc7934706u); // 13 add
|
||||
r0 = r0 ^ ds[r4 & mask]; // 14 load
|
||||
r2 = r2 - r4; // 15 sub
|
||||
r2 = r2 ^ ds[r0 & mask]; // 16 load
|
||||
r7 = r7 ^ ds[r2 & mask]; // 17 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 18 shfl
|
||||
r5 = r5 * r0; // 19 mul
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 20 shfl
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 16u); r2 = r2 ^ t_; } // 21 shfl
|
||||
r6 = mul_hi(r6, r2); // 22 mulhi
|
||||
r6 = r6 ^ ds[r1 & mask]; // 23 load
|
||||
r5 = r5 * r0; // 24 mul
|
||||
r5 = rotl_imm(r5, 19u); // 25 rotl
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 2u); r7 = r7 ^ t_; } // 26 shfl
|
||||
r0 = r0 ^ r5; // 27 xor
|
||||
r0 = r0 ^ r4; // 28 xor
|
||||
r3 = r3 - r0; // 29 sub
|
||||
r5 = r5 * r1; // 30 mul
|
||||
r7 = r7 ^ ds[r2 & mask]; // 31 load
|
||||
r1 = r1 ^ ds[r0 & mask]; // 32 load
|
||||
r5 = r5 ^ r6; // 33 xor
|
||||
r5 = r5 ^ ds[r1 & mask]; // 34 load
|
||||
r0 = mul_hi(r0, r5); // 35 mulhi
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r2, 4u); r5 = r5 ^ t_; } // 36 shfl
|
||||
r7 = r7 ^ ds[r0 & mask]; // 37 load
|
||||
r3 = r3 + r1 + ((((sel >> 27u) & 1u) != 0u) ? 0x230c005cu : 0x75ba2fadu); // 38 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 4u); r1 = r1 ^ t_; } // 39 shfl
|
||||
r2 = r2 ^ r5; // 40 xor
|
||||
r3 = r6 * r3 + r3; // 41 mad
|
||||
r6 = r6 - r7; // 42 sub
|
||||
r7 = r7 ^ r0; // 43 xor
|
||||
r1 = r1 ^ ds[r7 & mask]; // 44 load
|
||||
r2 = r2 * r3; // 45 mul
|
||||
r1 = mul_hi(r1, r5); // 46 mulhi
|
||||
r4 = r4 - r3; // 47 sub
|
||||
r2 = rotr_var(r2, r6); // 48 rotr
|
||||
r3 = r3 ^ ds[r5 & mask]; // 49 load
|
||||
r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 50 add
|
||||
r0 = r0 * r2; // 51 mul
|
||||
r0 = r0 + r2 + ((((sel >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); // 52 add
|
||||
r1 = r1 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0x77b1520du : 0x2bb965afu); // 53 add
|
||||
r7 = rotl_imm(r7, 14u); // 54 rotl
|
||||
r3 = r3 + r7 + ((((sel >> 1u) & 1u) != 0u) ? 0xa54c55a0u : 0x7b0fe07au); // 55 add
|
||||
r6 = r6 ^ ds[r7 & mask]; // 56 load
|
||||
r1 = rotr_var(r1, r5); // 57 rotr
|
||||
r5 = r5 ^ ds[r4 & mask]; // 58 load
|
||||
r6 = r6 ^ ds[r2 & mask]; // 59 load
|
||||
r3 = r5 * r0 + r3; // 60 mad
|
||||
r5 = r5 + r7 + ((((sel >> 31u) & 1u) != 0u) ? 0xad7493e7u : 0xaf9dd72du); // 61 add
|
||||
r4 = r4 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); // 62 add
|
||||
r5 = rotl_imm(r5, 19u); // 63 rotl
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
123
proto-cuda/packs/igneum-genesis-mh/kernel_bound.cu
Normal file
123
proto-cuda/packs/igneum-genesis-mh/kernel_bound.cu
Normal file
|
|
@ -0,0 +1,123 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
|
||||
// Header-bound twin of igneum_hash in kernel.cu: the init words come from a kernel argument, not SEEDW.
|
||||
// Host declarations (also in program_bound.h if present):
|
||||
// struct IgneumInitWords { uint32_t w[8]; };
|
||||
// cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
// IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps);
|
||||
// cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
|
||||
#include <cuda_runtime.h>
|
||||
#include <cstdint>
|
||||
#include "program.h"
|
||||
|
||||
struct IgneumInitWords { uint32_t w[8]; };
|
||||
|
||||
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
|
||||
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
|
||||
__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw) {
|
||||
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
uint32_t nonce = baseNonce + gid;
|
||||
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint32_t x = nonce ^ iw.w[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw.w[1]; }
|
||||
{ uint32_t x = nonce ^ iw.w[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw.w[2]; }
|
||||
{ uint32_t x = nonce ^ iw.w[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw.w[3]; }
|
||||
{ uint32_t x = nonce ^ iw.w[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw.w[4]; }
|
||||
{ uint32_t x = nonce ^ iw.w[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw.w[5]; }
|
||||
{ uint32_t x = nonce ^ iw.w[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw.w[6]; }
|
||||
{ uint32_t x = nonce ^ iw.w[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw.w[7]; }
|
||||
{ uint32_t x = nonce ^ iw.w[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw.w[0]; }
|
||||
|
||||
for (uint32_t it = 0u; it < 8u; ++it) {
|
||||
uint32_t sel = r0;
|
||||
r2 = r3 * r4 + r2; // 0 mad
|
||||
r2 = r1 * r1 + r2; // 1 mad
|
||||
r2 = r3 * r2 + r2; // 2 mad
|
||||
r3 = r3 ^ r5; // 3 xor
|
||||
r7 = r7 ^ ds[r2 & mask]; // 4 load
|
||||
r5 = r5 ^ ds[r7 & mask]; // 5 load
|
||||
r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 6 shfl
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 8); // 7 shfl
|
||||
r1 = __umulhi(r1, r5); // 8 mulhi
|
||||
r6 = rotr_var(r6, r3); // 9 rotr
|
||||
r3 = r3 | r4; // 10 or
|
||||
r4 = r4 ^ ds[r3 & mask]; // 11 load
|
||||
r0 = __umulhi(r0, r4); // 12 mulhi
|
||||
r5 = r5 + r1 + ((((sel >> 30u) & 1u) != 0u) ? 0xd3177981u : 0xc7934706u); // 13 add
|
||||
r0 = r0 ^ ds[r4 & mask]; // 14 load
|
||||
r2 = r2 - r4; // 15 sub
|
||||
r2 = r2 ^ ds[r0 & mask]; // 16 load
|
||||
r7 = r7 ^ ds[r2 & mask]; // 17 load
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 18 shfl
|
||||
r5 = r5 * r0; // 19 mul
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 20 shfl
|
||||
r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r4, 16); // 21 shfl
|
||||
r6 = __umulhi(r6, r2); // 22 mulhi
|
||||
r6 = r6 ^ ds[r1 & mask]; // 23 load
|
||||
r5 = r5 * r0; // 24 mul
|
||||
r5 = rotl_imm(r5, 19u); // 25 rotl
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r6, 2); // 26 shfl
|
||||
r0 = r0 ^ r5; // 27 xor
|
||||
r0 = r0 ^ r4; // 28 xor
|
||||
r3 = r3 - r0; // 29 sub
|
||||
r5 = r5 * r1; // 30 mul
|
||||
r7 = r7 ^ ds[r2 & mask]; // 31 load
|
||||
r1 = r1 ^ ds[r0 & mask]; // 32 load
|
||||
r5 = r5 ^ r6; // 33 xor
|
||||
r5 = r5 ^ ds[r1 & mask]; // 34 load
|
||||
r0 = __umulhi(r0, r5); // 35 mulhi
|
||||
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r2, 4); // 36 shfl
|
||||
r7 = r7 ^ ds[r0 & mask]; // 37 load
|
||||
r3 = r3 + r1 + ((((sel >> 27u) & 1u) != 0u) ? 0x230c005cu : 0x75ba2fadu); // 38 add
|
||||
r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r5, 4); // 39 shfl
|
||||
r2 = r2 ^ r5; // 40 xor
|
||||
r3 = r6 * r3 + r3; // 41 mad
|
||||
r6 = r6 - r7; // 42 sub
|
||||
r7 = r7 ^ r0; // 43 xor
|
||||
r1 = r1 ^ ds[r7 & mask]; // 44 load
|
||||
r2 = r2 * r3; // 45 mul
|
||||
r1 = __umulhi(r1, r5); // 46 mulhi
|
||||
r4 = r4 - r3; // 47 sub
|
||||
r2 = rotr_var(r2, r6); // 48 rotr
|
||||
r3 = r3 ^ ds[r5 & mask]; // 49 load
|
||||
r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 50 add
|
||||
r0 = r0 * r2; // 51 mul
|
||||
r0 = r0 + r2 + ((((sel >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); // 52 add
|
||||
r1 = r1 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0x77b1520du : 0x2bb965afu); // 53 add
|
||||
r7 = rotl_imm(r7, 14u); // 54 rotl
|
||||
r3 = r3 + r7 + ((((sel >> 1u) & 1u) != 0u) ? 0xa54c55a0u : 0x7b0fe07au); // 55 add
|
||||
r6 = r6 ^ ds[r7 & mask]; // 56 load
|
||||
r1 = rotr_var(r1, r5); // 57 rotr
|
||||
r5 = r5 ^ ds[r4 & mask]; // 58 load
|
||||
r6 = r6 ^ ds[r2 & mask]; // 59 load
|
||||
r3 = r5 * r0 + r3; // 60 mad
|
||||
r5 = r5 + r7 + ((((sel >> 31u) & 1u) != 0u) ? 0xad7493e7u : 0xaf9dd72du); // 61 add
|
||||
r4 = r4 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); // 62 add
|
||||
r5 = rotl_imm(r5, 19u); // 63 rotl
|
||||
}
|
||||
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps) {
|
||||
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 32u * blockWarps;
|
||||
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
|
||||
igneum_hash_bound<<<nonces / block, block>>>(ds, out, baseNonce, mask, iw);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
|
||||
cudaFuncAttributes attr;
|
||||
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash_bound);
|
||||
if (e != cudaSuccess) return e;
|
||||
*numRegs = attr.numRegs;
|
||||
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash_bound, (int)(32u * blockWarps), 0);
|
||||
}
|
||||
|
|
@ -1,4 +1,4 @@
|
|||
// Generated by proto-metal/igneum-bench --export-pack for seed "igneum-genesis". Do not edit by hand.
|
||||
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
|
||||
// Memory-hard dataset core, the same text that the Mac's Metal kernels and CPU verifier were checked against.
|
||||
// Included by kernel.cu (device), host.cu (host reference) and proto-opencl/host.c (C99 host reference).
|
||||
// See proto-metal/MEMHARD.md for the construction. kernel.cl carries the same text in OpenCL C.
|
||||
|
|
|
|||
|
|
@ -1,4 +1,4 @@
|
|||
// Generated by proto-metal/igneum-bench --export-pack for seed "igneum-genesis". Do not edit by hand.
|
||||
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
|
||||
// Program metadata for host.cu plus the launch wrappers defined in kernel.cu.
|
||||
// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros.
|
||||
#pragma once
|
||||
|
|
@ -12,7 +12,12 @@
|
|||
#endif
|
||||
|
||||
#define IGNEUM_SEED_STRING "igneum-genesis"
|
||||
#define IGNEUM_SEED_BYTES_HEX "69676e65756d2d67656e65736973"
|
||||
#define IGNEUM_GENERATOR 2
|
||||
#define IGNEUM_PROGRAM_ATTEMPT 0
|
||||
#define IGNEUM_PROGRAM_ID 0xbcc1248b10cc90f2ull
|
||||
#define IGNEUM_DAY_STRING "2026-10-03"
|
||||
#define IGNEUM_DAY_BYTES_HEX "6461792f323032362d31302d3033"
|
||||
#define IGNEUM_DAY0 0x3067619fu
|
||||
#define IGNEUM_DAY1 0x3c269176u
|
||||
#define IGNEUM_DATASET_LOG2 28
|
||||
|
|
@ -20,9 +25,9 @@
|
|||
#define IGNEUM_LANES 32
|
||||
#define IGNEUM_ITERATIONS 8
|
||||
#define IGNEUM_INSTR_COUNT 64
|
||||
#define IGNEUM_LOADS_PER_HASH 104
|
||||
#define IGNEUM_LOADS_PER_HASH 128
|
||||
#define IGNEUM_WIDE_LOADS_PER_HASH 0
|
||||
#define IGNEUM_OP_MIX "load=13 xor=13 sub=7 shfl=6 add=5 mulhi=5 mad=4 rotr=4 mul=3 rotl=3 or=1"
|
||||
#define IGNEUM_OP_MIX "load=16 add=8 shfl=8 xor=6 mad=5 mul=5 mulhi=5 sub=4 rotl=3 rotr=3 or=1"
|
||||
// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h)
|
||||
#define IGNEUM_DATASET_MODE 1
|
||||
|
||||
|
|
|
|||
|
|
@ -1,15 +1,21 @@
|
|||
{
|
||||
"format": "igneum-program-pack-2",
|
||||
"format": "igneum-program-pack-3",
|
||||
"generator": 2,
|
||||
"attempt": 0,
|
||||
"program_id": "0xbcc1248b10cc90f2",
|
||||
"program_id_derivation": "FNV-1a 64 over 'igneum-program/' || generator_le32 || seed_words as little-endian bytes || attempt_le32",
|
||||
"dataset_mode": "memory-hard",
|
||||
"seed": "igneum-genesis",
|
||||
"seed_bytes": "69676e65756d2d67656e65736973",
|
||||
"seed_words": ["0x67a9a7be", "0x1a155b25", "0xfddfb732", "0x4b5af2e8", "0xc55caf33", "0xa27c13b7", "0x06628a48", "0x03852469"],
|
||||
"seed_derivation": "FNV-1a 64 over UTF-8 of seed, basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32",
|
||||
"seed_derivation": "seed_words = FNV-1a 64 over seed_bytes (attempt 0) or seed_bytes || attempt_le32 (attempt k >= 1), basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32",
|
||||
"generator_rule": "version 2: exactly 16 load slots drawn first from instructions 1..63 (partial Fisher-Yates), the other 48 ops from the ten non-load weights (sum 75); a load's source is drawn from the registers other than dst written by an earlier instruction and not read by a load since; the candidate must pass the acceptance rule of spec 01 section 1.4.6 (static: no cyclically stale load source, every register has an injecting write; dynamic: 64 units on the seed-keyed closed-form dataset with no constant register bit, no lane-constant load site, under 164 saturated final values, every output bit within 136 of 1024, distinct addresses above 245760), else the next attempt of the seed is tried",
|
||||
"lanes": 32,
|
||||
"registers": 8,
|
||||
"iterations": 8,
|
||||
"instruction_count": 64,
|
||||
"loads_per_hash": 104,
|
||||
"op_mix": {"load": 13, "xor": 13, "sub": 7, "shfl": 6, "add": 5, "mulhi": 5, "mad": 4, "rotr": 4, "mul": 3, "rotl": 3, "or": 1},
|
||||
"loads_per_hash": 128,
|
||||
"op_mix": {"load": 16, "add": 8, "shfl": 8, "xor": 6, "mad": 5, "mul": 5, "mulhi": 5, "sub": 4, "rotl": 3, "rotr": 3, "or": 1},
|
||||
"register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]",
|
||||
"splitmix32": "x ^= x>>16; x *= 0x7feb352d; x ^= x>>15; x *= 0x846ca68b; x ^= x>>16",
|
||||
"iteration": "sel = r0 sampled once at the top of each iteration, then all instructions in order",
|
||||
|
|
@ -33,82 +39,83 @@
|
|||
"bytes": 1073741824,
|
||||
"mask": "0x0fffffff",
|
||||
"day": "2026-10-03",
|
||||
"day_words_from": "day/2026-10-03",
|
||||
"day_bytes": "6461792f323032362d31302d3033",
|
||||
"day_words_from": "seed_words_from_bytes(day_bytes)",
|
||||
"d0": "0x3067619f",
|
||||
"d1": "0x3c269176",
|
||||
"mode": "memory-hard",
|
||||
"spec": "proto-metal/MEMHARD.md",
|
||||
"key": ["0x3067619f", "0x3c269176", "0x84a03b03", "0xf8c63294", "0xff977c5b", "0xe60def3e", "0x63630141", "0xb8fbcb58"],
|
||||
"key_derivation": "the 8 words of seedWords(\"day/\" + day); d0, d1 are key[0], key[1]",
|
||||
"key_derivation": "the 8 words of seed_words_from_bytes(day_bytes); d0, d1 are key[0], key[1]",
|
||||
"cache": {"log2_words": 26, "bytes": 268435456, "line_words": 16, "segment_lines": 64, "segments": 65536, "block": "ChaCha12 core + feed-forward, rotations 16 12 8 7", "sigma": ["0x61707865", "0x3320646e", "0x79622d32", "0x6b206574"], "tag": ["0x49676e65", "0x756d4d48"], "chain": "in_j = prev_line ^ (sigma[0..3] || key[0..7] || seg || j || tag[0..1]); line_j = block(in_j); prev_0 = 0"},
|
||||
"mixer": {"draw": "SplitMix64 seeded with key[0] | key[1] << 32: rot[0..7] = 1 + next() % 31, mul[0..15] = low32(next()) | 1, rc[0..15] = low32(next())", "rot": [20, 20, 19, 4, 26, 3, 3, 27], "mul": ["0x42146205", "0x52cbe0fb", "0x7ecf4a03", "0x6728907f", "0xd81d9751", "0x132952c3", "0xf60de277", "0x05358035", "0xbaf6499d", "0xe4db9667", "0x3e98f45d", "0xd0004edd", "0x2691630d", "0x9beb3bcf", "0xab310379", "0x99cfb423"], "rc": ["0xbab68293", "0xcc162340", "0x6ce151cc", "0xe62b8997", "0xc9c80297", "0xf74a1654", "0x3d704af5", "0x3cf522b7", "0x2b9cac04", "0xa880ac10", "0x13e5dd1d", "0x6fc3e233", "0x2d83eeac", "0x9006e8bf", "0x2c4b5362", "0x31b49ee2"], "round": "for i in 0..15: s[i] = (s[i] ^ (rc[i] + (r+1) * 0x9E3779B9)) * mul[i]; then quarter rounds on columns (0,4,8,12) (1,5,9,13) (2,6,10,14) (3,7,11,15) with rot[0..3] and diagonals (0,5,10,15) (1,6,11,12) (2,7,8,13) (3,4,9,14) with rot[4..7]", "quarter_round": "a += b; d ^= a; d = rotl(d, r1); c += d; b ^= c; b = rotl(b, r2); a += b; d ^= a; d = rotl(d, r3); c += d; b ^= c; b = rotl(b, r4)"},
|
||||
"item": "s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..7: s = M_r(s); line = s[0] & "0x003fffff"; s[i] ^= cache[line * 16 + i]; then s = M_8(s); item(t) = s",
|
||||
"item": "s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..7: s = M_r(s); line = s[0] & 0x003fffff; s[i] ^= cache[line * 16 + i]; then s = M_8(s); item(t) = s",
|
||||
"word": "dataset[w] = item(w >> 4)[w & 15]"
|
||||
},
|
||||
"instructions": [
|
||||
{"i": 0, "op": "rotl", "dst": 4, "src": 2, "src2": 7, "imm": "0x20699878", "imm2": "0x6f1a6170", "rot": 25, "bit": 7, "mask": 16},
|
||||
{"i": 1, "op": "sub", "dst": 0, "src": 5, "src2": 4, "imm": "0xf81a0b9d", "imm2": "0xf0505e88", "rot": 1, "bit": 4, "mask": 1},
|
||||
{"i": 2, "op": "load", "dst": 4, "src": 3, "src2": 6, "imm": "0xb1978a0b", "imm2": "0x2ca4e162", "rot": 10, "bit": 21, "mask": 1},
|
||||
{"i": 3, "op": "rotl", "dst": 1, "src": 6, "src2": 7, "imm": "0xc3bd2355", "imm2": "0xa8c5f27e", "rot": 1, "bit": 22, "mask": 1},
|
||||
{"i": 4, "op": "add", "dst": 2, "src": 3, "src2": 0, "imm": "0x61f0b51c", "imm2": "0x2735a174", "rot": 4, "bit": 26, "mask": 2},
|
||||
{"i": 5, "op": "load", "dst": 5, "src": 3, "src2": 3, "imm": "0x4d183796", "imm2": "0x679648a8", "rot": 4, "bit": 30, "mask": 4},
|
||||
{"i": 6, "op": "sub", "dst": 5, "src": 7, "src2": 7, "imm": "0x265677dc", "imm2": "0x9043323e", "rot": 30, "bit": 7, "mask": 4},
|
||||
{"i": 7, "op": "add", "dst": 3, "src": 4, "src2": 6, "imm": "0x52f2dbf4", "imm2": "0x5a069596", "rot": 5, "bit": 26, "mask": 2},
|
||||
{"i": 8, "op": "xor", "dst": 0, "src": 4, "src2": 1, "imm": "0x3303ec4b", "imm2": "0xfeca75be", "rot": 21, "bit": 7, "mask": 1},
|
||||
{"i": 9, "op": "xor", "dst": 4, "src": 0, "src2": 0, "imm": "0x5dc5959c", "imm2": "0x023f44a8", "rot": 11, "bit": 7, "mask": 4},
|
||||
{"i": 10, "op": "load", "dst": 2, "src": 0, "src2": 3, "imm": "0xde716173", "imm2": "0xc21e924d", "rot": 30, "bit": 15, "mask": 1},
|
||||
{"i": 11, "op": "shfl", "dst": 4, "src": 6, "src2": 1, "imm": "0x3cda6d48", "imm2": "0x0970145b", "rot": 12, "bit": 22, "mask": 16},
|
||||
{"i": 12, "op": "sub", "dst": 1, "src": 5, "src2": 4, "imm": "0x48866c15", "imm2": "0x4eaee50c", "rot": 30, "bit": 20, "mask": 16},
|
||||
{"i": 13, "op": "xor", "dst": 2, "src": 1, "src2": 5, "imm": "0x196d165c", "imm2": "0x0f730511", "rot": 11, "bit": 3, "mask": 4},
|
||||
{"i": 14, "op": "load", "dst": 4, "src": 5, "src2": 1, "imm": "0xc5c3b55d", "imm2": "0xec061424", "rot": 26, "bit": 27, "mask": 8},
|
||||
{"i": 15, "op": "load", "dst": 2, "src": 4, "src2": 1, "imm": "0x17c9c95b", "imm2": "0x306542fe", "rot": 27, "bit": 17, "mask": 16},
|
||||
{"i": 16, "op": "load", "dst": 3, "src": 0, "src2": 2, "imm": "0x590f9e11", "imm2": "0xa4c13036", "rot": 3, "bit": 28, "mask": 16},
|
||||
{"i": 17, "op": "xor", "dst": 4, "src": 6, "src2": 3, "imm": "0x9c4eeee9", "imm2": "0xf069b834", "rot": 11, "bit": 23, "mask": 8},
|
||||
{"i": 18, "op": "mad", "dst": 2, "src": 4, "src2": 6, "imm": "0xa242a28b", "imm2": "0xb8974bdf", "rot": 13, "bit": 30, "mask": 1},
|
||||
{"i": 19, "op": "rotr", "dst": 6, "src": 1, "src2": 0, "imm": "0xcd69ed50", "imm2": "0xadf52615", "rot": 30, "bit": 2, "mask": 2},
|
||||
{"i": 20, "op": "xor", "dst": 3, "src": 4, "src2": 0, "imm": "0xe3059a24", "imm2": "0x16b8dd86", "rot": 25, "bit": 31, "mask": 16},
|
||||
{"i": 21, "op": "mad", "dst": 1, "src": 3, "src2": 5, "imm": "0xc3ae8ae1", "imm2": "0x8e126e8a", "rot": 2, "bit": 28, "mask": 1},
|
||||
{"i": 22, "op": "mulhi", "dst": 7, "src": 4, "src2": 2, "imm": "0x6909af7a", "imm2": "0xa388b4b9", "rot": 14, "bit": 10, "mask": 2},
|
||||
{"i": 23, "op": "mulhi", "dst": 5, "src": 2, "src2": 2, "imm": "0xdf099cfb", "imm2": "0xd0133a01", "rot": 31, "bit": 3, "mask": 4},
|
||||
{"i": 24, "op": "load", "dst": 0, "src": 6, "src2": 3, "imm": "0x0a3056de", "imm2": "0x7f0c25c3", "rot": 27, "bit": 13, "mask": 8},
|
||||
{"i": 25, "op": "mul", "dst": 5, "src": 6, "src2": 7, "imm": "0x089f5404", "imm2": "0xbd066e1d", "rot": 7, "bit": 10, "mask": 4},
|
||||
{"i": 26, "op": "load", "dst": 7, "src": 1, "src2": 2, "imm": "0x3e26afea", "imm2": "0xba573970", "rot": 11, "bit": 15, "mask": 8},
|
||||
{"i": 27, "op": "rotr", "dst": 3, "src": 1, "src2": 5, "imm": "0xe2466d63", "imm2": "0x30db112b", "rot": 31, "bit": 15, "mask": 2},
|
||||
{"i": 28, "op": "shfl", "dst": 5, "src": 1, "src2": 3, "imm": "0xdc20297d", "imm2": "0xcb813557", "rot": 3, "bit": 21, "mask": 8},
|
||||
{"i": 29, "op": "xor", "dst": 7, "src": 5, "src2": 2, "imm": "0x6ede5c13", "imm2": "0x1f14267f", "rot": 22, "bit": 31, "mask": 1},
|
||||
{"i": 30, "op": "rotl", "dst": 7, "src": 5, "src2": 6, "imm": "0x24cdb54b", "imm2": "0xa44e008d", "rot": 23, "bit": 4, "mask": 16},
|
||||
{"i": 31, "op": "sub", "dst": 2, "src": 0, "src2": 2, "imm": "0xf94d9b65", "imm2": "0x3457264c", "rot": 11, "bit": 5, "mask": 2},
|
||||
{"i": 32, "op": "xor", "dst": 7, "src": 2, "src2": 3, "imm": "0x8d2046e5", "imm2": "0xf68893bb", "rot": 21, "bit": 2, "mask": 1},
|
||||
{"i": 33, "op": "xor", "dst": 2, "src": 6, "src2": 5, "imm": "0xe0ebc4ce", "imm2": "0x02773069", "rot": 21, "bit": 16, "mask": 1},
|
||||
{"i": 34, "op": "load", "dst": 6, "src": 1, "src2": 2, "imm": "0x3b2d2124", "imm2": "0x187a9128", "rot": 1, "bit": 9, "mask": 16},
|
||||
{"i": 35, "op": "xor", "dst": 1, "src": 4, "src2": 6, "imm": "0xe3f24158", "imm2": "0x5c64a589", "rot": 13, "bit": 0, "mask": 8},
|
||||
{"i": 36, "op": "shfl", "dst": 3, "src": 0, "src2": 3, "imm": "0x63578bc1", "imm2": "0xbf64a89f", "rot": 16, "bit": 26, "mask": 1},
|
||||
{"i": 37, "op": "mul", "dst": 2, "src": 6, "src2": 4, "imm": "0xfa624b69", "imm2": "0x0389cf85", "rot": 28, "bit": 4, "mask": 1},
|
||||
{"i": 38, "op": "add", "dst": 5, "src": 3, "src2": 5, "imm": "0x71f30417", "imm2": "0xa29f4338", "rot": 2, "bit": 12, "mask": 2},
|
||||
{"i": 39, "op": "xor", "dst": 7, "src": 6, "src2": 3, "imm": "0xfa3b845c", "imm2": "0x9717695d", "rot": 18, "bit": 24, "mask": 1},
|
||||
{"i": 40, "op": "xor", "dst": 7, "src": 3, "src2": 4, "imm": "0x8408dc67", "imm2": "0x8d389f9b", "rot": 21, "bit": 16, "mask": 16},
|
||||
{"i": 41, "op": "rotr", "dst": 3, "src": 4, "src2": 1, "imm": "0x4bbcad92", "imm2": "0xbc6375cd", "rot": 18, "bit": 4, "mask": 2},
|
||||
{"i": 42, "op": "xor", "dst": 5, "src": 3, "src2": 3, "imm": "0xbd31a2ea", "imm2": "0x742d28ea", "rot": 4, "bit": 12, "mask": 4},
|
||||
{"i": 43, "op": "rotr", "dst": 3, "src": 6, "src2": 5, "imm": "0x35c52d04", "imm2": "0x0f8e8621", "rot": 3, "bit": 22, "mask": 8},
|
||||
{"i": 44, "op": "mad", "dst": 1, "src": 3, "src2": 5, "imm": "0x3958f280", "imm2": "0x8713c7e1", "rot": 5, "bit": 23, "mask": 16},
|
||||
{"i": 45, "op": "add", "dst": 7, "src": 3, "src2": 7, "imm": "0xc1ae8d3b", "imm2": "0xa907b90b", "rot": 13, "bit": 29, "mask": 1},
|
||||
{"i": 46, "op": "or", "dst": 7, "src": 2, "src2": 2, "imm": "0xc4f18bec", "imm2": "0x8a3e2464", "rot": 30, "bit": 16, "mask": 2},
|
||||
{"i": 47, "op": "shfl", "dst": 2, "src": 3, "src2": 6, "imm": "0xf73ba9e3", "imm2": "0x028aa63c", "rot": 3, "bit": 20, "mask": 16},
|
||||
{"i": 48, "op": "load", "dst": 1, "src": 7, "src2": 5, "imm": "0x53be87d4", "imm2": "0x690a5729", "rot": 1, "bit": 17, "mask": 8},
|
||||
{"i": 49, "op": "sub", "dst": 5, "src": 1, "src2": 7, "imm": "0xecdc31a6", "imm2": "0xed6bd9f2", "rot": 1, "bit": 13, "mask": 8},
|
||||
{"i": 50, "op": "load", "dst": 3, "src": 1, "src2": 3, "imm": "0x55f91dbc", "imm2": "0xf026553d", "rot": 30, "bit": 27, "mask": 1},
|
||||
{"i": 51, "op": "sub", "dst": 2, "src": 3, "src2": 5, "imm": "0x6689eef2", "imm2": "0x24f3ae21", "rot": 6, "bit": 22, "mask": 8},
|
||||
{"i": 52, "op": "shfl", "dst": 6, "src": 2, "src2": 4, "imm": "0x8fd37aad", "imm2": "0x89ed5d67", "rot": 8, "bit": 16, "mask": 2},
|
||||
{"i": 53, "op": "sub", "dst": 2, "src": 0, "src2": 0, "imm": "0x963bb7e6", "imm2": "0x4738b84f", "rot": 18, "bit": 16, "mask": 4},
|
||||
{"i": 54, "op": "load", "dst": 0, "src": 3, "src2": 0, "imm": "0x838b5065", "imm2": "0x36360066", "rot": 3, "bit": 31, "mask": 4},
|
||||
{"i": 55, "op": "shfl", "dst": 2, "src": 1, "src2": 5, "imm": "0xb17aad78", "imm2": "0x8458f7ac", "rot": 5, "bit": 7, "mask": 16},
|
||||
{"i": 56, "op": "add", "dst": 0, "src": 6, "src2": 0, "imm": "0x25955401", "imm2": "0xf4689674", "rot": 14, "bit": 27, "mask": 4},
|
||||
{"i": 57, "op": "mulhi", "dst": 2, "src": 0, "src2": 0, "imm": "0xa4e8c86a", "imm2": "0x14bde8e1", "rot": 8, "bit": 9, "mask": 16},
|
||||
{"i": 58, "op": "mulhi", "dst": 4, "src": 2, "src2": 7, "imm": "0xa36b1f4f", "imm2": "0x372e072f", "rot": 31, "bit": 31, "mask": 1},
|
||||
{"i": 59, "op": "mad", "dst": 2, "src": 6, "src2": 7, "imm": "0xfcd3b17f", "imm2": "0x6a8583df", "rot": 24, "bit": 9, "mask": 2},
|
||||
{"i": 60, "op": "xor", "dst": 3, "src": 1, "src2": 0, "imm": "0x3bbea2ae", "imm2": "0x914e2f0a", "rot": 2, "bit": 15, "mask": 16},
|
||||
{"i": 61, "op": "mul", "dst": 4, "src": 2, "src2": 6, "imm": "0x50aaa099", "imm2": "0xd1b988a1", "rot": 27, "bit": 31, "mask": 4},
|
||||
{"i": 62, "op": "load", "dst": 0, "src": 2, "src2": 3, "imm": "0x4ccdbdea", "imm2": "0xf3a60b41", "rot": 1, "bit": 20, "mask": 2},
|
||||
{"i": 63, "op": "mulhi", "dst": 7, "src": 0, "src2": 7, "imm": "0x3e400372", "imm2": "0x92c6201f", "rot": 1, "bit": 13, "mask": 8}
|
||||
{"i": 0, "op": "mad", "dst": 2, "src": 3, "src2": 4, "imm": "0xbf7b174d", "imm2": "0x337b762e", "rot": 17, "bit": 2, "mask": 2},
|
||||
{"i": 1, "op": "mad", "dst": 2, "src": 1, "src2": 1, "imm": "0xdd04a5da", "imm2": "0x42da7657", "rot": 15, "bit": 30, "mask": 16},
|
||||
{"i": 2, "op": "mad", "dst": 2, "src": 3, "src2": 2, "imm": "0x734003fa", "imm2": "0x5bb67700", "rot": 3, "bit": 20, "mask": 1},
|
||||
{"i": 3, "op": "xor", "dst": 3, "src": 5, "src2": 5, "imm": "0xc55a1b1c", "imm2": "0xa19720f3", "rot": 7, "bit": 8, "mask": 1},
|
||||
{"i": 4, "op": "load", "dst": 7, "src": 2, "src2": 5, "imm": "0xad572dd7", "imm2": "0x9ceb3ea7", "rot": 18, "bit": 30, "mask": 2},
|
||||
{"i": 5, "op": "load", "dst": 5, "src": 7, "src2": 3, "imm": "0x769a53be", "imm2": "0x80f9067e", "rot": 12, "bit": 22, "mask": 1},
|
||||
{"i": 6, "op": "shfl", "dst": 1, "src": 4, "src2": 0, "imm": "0xd3613d88", "imm2": "0x262fb219", "rot": 10, "bit": 30, "mask": 8},
|
||||
{"i": 7, "op": "shfl", "dst": 7, "src": 3, "src2": 4, "imm": "0xce38e42f", "imm2": "0xb868b818", "rot": 11, "bit": 8, "mask": 8},
|
||||
{"i": 8, "op": "mulhi", "dst": 1, "src": 5, "src2": 2, "imm": "0xa5eebca5", "imm2": "0x5703a72b", "rot": 13, "bit": 13, "mask": 16},
|
||||
{"i": 9, "op": "rotr", "dst": 6, "src": 3, "src2": 4, "imm": "0x17a5a9c7", "imm2": "0xdcfb93a1", "rot": 20, "bit": 27, "mask": 2},
|
||||
{"i": 10, "op": "or", "dst": 3, "src": 4, "src2": 1, "imm": "0xccb7d785", "imm2": "0xc335364c", "rot": 14, "bit": 12, "mask": 4},
|
||||
{"i": 11, "op": "load", "dst": 4, "src": 3, "src2": 2, "imm": "0x88cb9af3", "imm2": "0x4e7dc10d", "rot": 24, "bit": 17, "mask": 4},
|
||||
{"i": 12, "op": "mulhi", "dst": 0, "src": 4, "src2": 4, "imm": "0x45374321", "imm2": "0x3cd91989", "rot": 11, "bit": 4, "mask": 2},
|
||||
{"i": 13, "op": "add", "dst": 5, "src": 1, "src2": 2, "imm": "0xc7934706", "imm2": "0xd3177981", "rot": 16, "bit": 30, "mask": 2},
|
||||
{"i": 14, "op": "load", "dst": 0, "src": 4, "src2": 3, "imm": "0xfd7f56bb", "imm2": "0x65e14f52", "rot": 13, "bit": 22, "mask": 2},
|
||||
{"i": 15, "op": "sub", "dst": 2, "src": 4, "src2": 4, "imm": "0x35a80b49", "imm2": "0x060f2d13", "rot": 16, "bit": 20, "mask": 16},
|
||||
{"i": 16, "op": "load", "dst": 2, "src": 0, "src2": 2, "imm": "0xae0a32c2", "imm2": "0x4c2a4cfe", "rot": 8, "bit": 31, "mask": 16},
|
||||
{"i": 17, "op": "load", "dst": 7, "src": 2, "src2": 6, "imm": "0x82a84cc3", "imm2": "0x21a38d68", "rot": 15, "bit": 21, "mask": 2},
|
||||
{"i": 18, "op": "shfl", "dst": 7, "src": 3, "src2": 3, "imm": "0xa3818806", "imm2": "0x8f66b5c8", "rot": 14, "bit": 6, "mask": 4},
|
||||
{"i": 19, "op": "mul", "dst": 5, "src": 0, "src2": 1, "imm": "0xa00de107", "imm2": "0x77bfcaa5", "rot": 3, "bit": 10, "mask": 2},
|
||||
{"i": 20, "op": "shfl", "dst": 3, "src": 4, "src2": 7, "imm": "0x1d2b8cab", "imm2": "0x80b4f9a2", "rot": 14, "bit": 25, "mask": 2},
|
||||
{"i": 21, "op": "shfl", "dst": 2, "src": 4, "src2": 5, "imm": "0x3ac915d2", "imm2": "0x5fba7bc2", "rot": 16, "bit": 1, "mask": 16},
|
||||
{"i": 22, "op": "mulhi", "dst": 6, "src": 2, "src2": 0, "imm": "0xdc3ec8fd", "imm2": "0x599e2fa3", "rot": 22, "bit": 3, "mask": 2},
|
||||
{"i": 23, "op": "load", "dst": 6, "src": 1, "src2": 5, "imm": "0x2a6b16d5", "imm2": "0xd73e396f", "rot": 28, "bit": 29, "mask": 2},
|
||||
{"i": 24, "op": "mul", "dst": 5, "src": 0, "src2": 7, "imm": "0x376d0223", "imm2": "0xe1c2169a", "rot": 4, "bit": 16, "mask": 16},
|
||||
{"i": 25, "op": "rotl", "dst": 5, "src": 7, "src2": 3, "imm": "0x78ad8c60", "imm2": "0x6f5b77d5", "rot": 19, "bit": 11, "mask": 16},
|
||||
{"i": 26, "op": "shfl", "dst": 7, "src": 6, "src2": 5, "imm": "0x93915b9f", "imm2": "0x1e61fb6b", "rot": 28, "bit": 23, "mask": 2},
|
||||
{"i": 27, "op": "xor", "dst": 0, "src": 5, "src2": 7, "imm": "0x6378fe15", "imm2": "0x66c78f42", "rot": 12, "bit": 31, "mask": 8},
|
||||
{"i": 28, "op": "xor", "dst": 0, "src": 4, "src2": 7, "imm": "0x20a57fda", "imm2": "0x088c848e", "rot": 16, "bit": 13, "mask": 4},
|
||||
{"i": 29, "op": "sub", "dst": 3, "src": 0, "src2": 2, "imm": "0x49d95fd5", "imm2": "0x1a5f946a", "rot": 6, "bit": 12, "mask": 1},
|
||||
{"i": 30, "op": "mul", "dst": 5, "src": 1, "src2": 7, "imm": "0x0a816217", "imm2": "0x405c4f73", "rot": 13, "bit": 27, "mask": 4},
|
||||
{"i": 31, "op": "load", "dst": 7, "src": 2, "src2": 2, "imm": "0x09ed045e", "imm2": "0xd69c4715", "rot": 5, "bit": 9, "mask": 2},
|
||||
{"i": 32, "op": "load", "dst": 1, "src": 0, "src2": 6, "imm": "0xeb79ea49", "imm2": "0xcc587f5a", "rot": 6, "bit": 8, "mask": 16},
|
||||
{"i": 33, "op": "xor", "dst": 5, "src": 6, "src2": 1, "imm": "0x3027401e", "imm2": "0x5f20c27e", "rot": 18, "bit": 9, "mask": 2},
|
||||
{"i": 34, "op": "load", "dst": 5, "src": 1, "src2": 3, "imm": "0x0e1cab07", "imm2": "0x09356c5b", "rot": 19, "bit": 31, "mask": 1},
|
||||
{"i": 35, "op": "mulhi", "dst": 0, "src": 5, "src2": 2, "imm": "0x90e31357", "imm2": "0xabd32484", "rot": 26, "bit": 5, "mask": 8},
|
||||
{"i": 36, "op": "shfl", "dst": 5, "src": 2, "src2": 5, "imm": "0xee9a955f", "imm2": "0x31b3faed", "rot": 8, "bit": 24, "mask": 4},
|
||||
{"i": 37, "op": "load", "dst": 7, "src": 0, "src2": 7, "imm": "0x3ba2f832", "imm2": "0x1160dcd3", "rot": 4, "bit": 29, "mask": 1},
|
||||
{"i": 38, "op": "add", "dst": 3, "src": 1, "src2": 7, "imm": "0x75ba2fad", "imm2": "0x230c005c", "rot": 4, "bit": 27, "mask": 1},
|
||||
{"i": 39, "op": "shfl", "dst": 1, "src": 5, "src2": 3, "imm": "0xdbf37e75", "imm2": "0xb5ac1969", "rot": 30, "bit": 13, "mask": 4},
|
||||
{"i": 40, "op": "xor", "dst": 2, "src": 5, "src2": 5, "imm": "0x47f136c5", "imm2": "0x06ce9153", "rot": 19, "bit": 10, "mask": 2},
|
||||
{"i": 41, "op": "mad", "dst": 3, "src": 6, "src2": 3, "imm": "0xce13eff8", "imm2": "0x04cc1d55", "rot": 3, "bit": 1, "mask": 4},
|
||||
{"i": 42, "op": "sub", "dst": 6, "src": 7, "src2": 1, "imm": "0x6a65ab71", "imm2": "0x8fbc1bcd", "rot": 4, "bit": 1, "mask": 8},
|
||||
{"i": 43, "op": "xor", "dst": 7, "src": 0, "src2": 7, "imm": "0xdaeb4928", "imm2": "0xc0423027", "rot": 24, "bit": 11, "mask": 8},
|
||||
{"i": 44, "op": "load", "dst": 1, "src": 7, "src2": 7, "imm": "0x778f01c9", "imm2": "0x28cedcea", "rot": 12, "bit": 4, "mask": 16},
|
||||
{"i": 45, "op": "mul", "dst": 2, "src": 3, "src2": 2, "imm": "0xf4264f1b", "imm2": "0x0f627d56", "rot": 5, "bit": 28, "mask": 8},
|
||||
{"i": 46, "op": "mulhi", "dst": 1, "src": 5, "src2": 1, "imm": "0xffb2147a", "imm2": "0xccde9b05", "rot": 13, "bit": 9, "mask": 2},
|
||||
{"i": 47, "op": "sub", "dst": 4, "src": 3, "src2": 5, "imm": "0x73b36234", "imm2": "0x3f5d5997", "rot": 7, "bit": 18, "mask": 2},
|
||||
{"i": 48, "op": "rotr", "dst": 2, "src": 6, "src2": 3, "imm": "0x3a4d9aa9", "imm2": "0x212bec7b", "rot": 4, "bit": 29, "mask": 16},
|
||||
{"i": 49, "op": "load", "dst": 3, "src": 5, "src2": 2, "imm": "0x626f11df", "imm2": "0x56cd5bfd", "rot": 7, "bit": 1, "mask": 1},
|
||||
{"i": 50, "op": "add", "dst": 1, "src": 5, "src2": 6, "imm": "0x81b8bc2c", "imm2": "0x1907970c", "rot": 28, "bit": 7, "mask": 4},
|
||||
{"i": 51, "op": "mul", "dst": 0, "src": 2, "src2": 2, "imm": "0xa8848b30", "imm2": "0xef6ac348", "rot": 9, "bit": 15, "mask": 8},
|
||||
{"i": 52, "op": "add", "dst": 0, "src": 2, "src2": 0, "imm": "0x4f92b968", "imm2": "0x699fd448", "rot": 22, "bit": 6, "mask": 4},
|
||||
{"i": 53, "op": "add", "dst": 1, "src": 0, "src2": 2, "imm": "0x2bb965af", "imm2": "0x77b1520d", "rot": 2, "bit": 12, "mask": 8},
|
||||
{"i": 54, "op": "rotl", "dst": 7, "src": 1, "src2": 0, "imm": "0x553e678b", "imm2": "0x3cc8eae0", "rot": 14, "bit": 20, "mask": 2},
|
||||
{"i": 55, "op": "add", "dst": 3, "src": 7, "src2": 2, "imm": "0x7b0fe07a", "imm2": "0xa54c55a0", "rot": 10, "bit": 1, "mask": 1},
|
||||
{"i": 56, "op": "load", "dst": 6, "src": 7, "src2": 4, "imm": "0x01eba9aa", "imm2": "0x2758c0f7", "rot": 14, "bit": 15, "mask": 4},
|
||||
{"i": 57, "op": "rotr", "dst": 1, "src": 5, "src2": 2, "imm": "0x1f5267b3", "imm2": "0x236f5a27", "rot": 2, "bit": 31, "mask": 16},
|
||||
{"i": 58, "op": "load", "dst": 5, "src": 4, "src2": 3, "imm": "0xa9954a9b", "imm2": "0x6a54d4e8", "rot": 11, "bit": 10, "mask": 16},
|
||||
{"i": 59, "op": "load", "dst": 6, "src": 2, "src2": 4, "imm": "0x9923ff88", "imm2": "0x9357254e", "rot": 16, "bit": 1, "mask": 16},
|
||||
{"i": 60, "op": "mad", "dst": 3, "src": 5, "src2": 0, "imm": "0xc0cc51a6", "imm2": "0x3fd7701b", "rot": 20, "bit": 1, "mask": 4},
|
||||
{"i": 61, "op": "add", "dst": 5, "src": 7, "src2": 7, "imm": "0xaf9dd72d", "imm2": "0xad7493e7", "rot": 7, "bit": 31, "mask": 16},
|
||||
{"i": 62, "op": "add", "dst": 4, "src": 6, "src2": 2, "imm": "0x89841d87", "imm2": "0x1e07c3d9", "rot": 6, "bit": 27, "mask": 1},
|
||||
{"i": 63, "op": "rotl", "dst": 5, "src": 4, "src2": 2, "imm": "0xae210f8d", "imm2": "0x8e499ba4", "rot": 19, "bit": 9, "mask": 1}
|
||||
]
|
||||
}
|
||||
|
|
|
|||
|
|
@ -38,70 +38,70 @@ kernel void igneum_hash(device const uint* dataset [[buffer(0)]],
|
|||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = rotl_imm(r4, 25u); // 0
|
||||
r0 = r0 - r5; // 1
|
||||
r4 = r4 ^ dataset[r3 & MASK]; // 2
|
||||
r1 = rotl_imm(r1, 1u); // 3
|
||||
r2 = r2 + r3 + select(0x61f0b51cu, 0x2735a174u, ((sel >> 26u) & 1u) != 0u); // 4
|
||||
r5 = r5 ^ dataset[r3 & MASK]; // 5
|
||||
r5 = r5 - r7; // 6
|
||||
r3 = r3 + r4 + select(0x52f2dbf4u, 0x5a069596u, ((sel >> 26u) & 1u) != 0u); // 7
|
||||
r0 = r0 ^ r4; // 8
|
||||
r4 = r4 ^ r0; // 9
|
||||
r2 = r2 ^ dataset[r0 & MASK]; // 10
|
||||
r4 = r4 ^ simd_shuffle_xor(r6, (ushort)16); // 11
|
||||
r1 = r1 - r5; // 12
|
||||
r2 = r2 ^ r1; // 13
|
||||
r4 = r4 ^ dataset[r5 & MASK]; // 14
|
||||
r2 = r2 ^ dataset[r4 & MASK]; // 15
|
||||
r3 = r3 ^ dataset[r0 & MASK]; // 16
|
||||
r4 = r4 ^ r6; // 17
|
||||
r2 = r4 * r6 + r2; // 18
|
||||
r6 = rotr_var(r6, r1); // 19
|
||||
r3 = r3 ^ r4; // 20
|
||||
r1 = r3 * r5 + r1; // 21
|
||||
r7 = mulhi(r7, r4); // 22
|
||||
r5 = mulhi(r5, r2); // 23
|
||||
r0 = r0 ^ dataset[r6 & MASK]; // 24
|
||||
r5 = r5 * r6; // 25
|
||||
r7 = r7 ^ dataset[r1 & MASK]; // 26
|
||||
r3 = rotr_var(r3, r1); // 27
|
||||
r5 = r5 ^ simd_shuffle_xor(r1, (ushort)8); // 28
|
||||
r7 = r7 ^ r5; // 29
|
||||
r7 = rotl_imm(r7, 23u); // 30
|
||||
r2 = r2 - r0; // 31
|
||||
r7 = r7 ^ r2; // 32
|
||||
r2 = r2 ^ r6; // 33
|
||||
r6 = r6 ^ dataset[r1 & MASK]; // 34
|
||||
r1 = r1 ^ r4; // 35
|
||||
r3 = r3 ^ simd_shuffle_xor(r0, (ushort)1); // 36
|
||||
r2 = r2 * r6; // 37
|
||||
r5 = r5 + r3 + select(0x71f30417u, 0xa29f4338u, ((sel >> 12u) & 1u) != 0u); // 38
|
||||
r7 = r7 ^ r6; // 39
|
||||
r7 = r7 ^ r3; // 40
|
||||
r3 = rotr_var(r3, r4); // 41
|
||||
r5 = r5 ^ r3; // 42
|
||||
r3 = rotr_var(r3, r6); // 43
|
||||
r1 = r3 * r5 + r1; // 44
|
||||
r7 = r7 + r3 + select(0xc1ae8d3bu, 0xa907b90bu, ((sel >> 29u) & 1u) != 0u); // 45
|
||||
r7 = r7 | r2; // 46
|
||||
r2 = r2 ^ simd_shuffle_xor(r3, (ushort)16); // 47
|
||||
r1 = r1 ^ dataset[r7 & MASK]; // 48
|
||||
r5 = r5 - r1; // 49
|
||||
r3 = r3 ^ dataset[r1 & MASK]; // 50
|
||||
r2 = r2 - r3; // 51
|
||||
r6 = r6 ^ simd_shuffle_xor(r2, (ushort)2); // 52
|
||||
r2 = r2 - r0; // 53
|
||||
r0 = r0 ^ dataset[r3 & MASK]; // 54
|
||||
r2 = r2 ^ simd_shuffle_xor(r1, (ushort)16); // 55
|
||||
r0 = r0 + r6 + select(0x25955401u, 0xf4689674u, ((sel >> 27u) & 1u) != 0u); // 56
|
||||
r2 = mulhi(r2, r0); // 57
|
||||
r4 = mulhi(r4, r2); // 58
|
||||
r2 = r6 * r7 + r2; // 59
|
||||
r3 = r3 ^ r1; // 60
|
||||
r4 = r4 * r2; // 61
|
||||
r0 = r0 ^ dataset[r2 & MASK]; // 62
|
||||
r7 = mulhi(r7, r0); // 63
|
||||
r2 = r3 * r4 + r2; // 0
|
||||
r2 = r1 * r1 + r2; // 1
|
||||
r2 = r3 * r2 + r2; // 2
|
||||
r3 = r3 ^ r5; // 3
|
||||
r7 = r7 ^ dataset[r2 & MASK]; // 4
|
||||
r5 = r5 ^ dataset[r7 & MASK]; // 5
|
||||
r1 = r1 ^ simd_shuffle_xor(r4, (ushort)8); // 6
|
||||
r7 = r7 ^ simd_shuffle_xor(r3, (ushort)8); // 7
|
||||
r1 = mulhi(r1, r5); // 8
|
||||
r6 = rotr_var(r6, r3); // 9
|
||||
r3 = r3 | r4; // 10
|
||||
r4 = r4 ^ dataset[r3 & MASK]; // 11
|
||||
r0 = mulhi(r0, r4); // 12
|
||||
r5 = r5 + r1 + select(0xc7934706u, 0xd3177981u, ((sel >> 30u) & 1u) != 0u); // 13
|
||||
r0 = r0 ^ dataset[r4 & MASK]; // 14
|
||||
r2 = r2 - r4; // 15
|
||||
r2 = r2 ^ dataset[r0 & MASK]; // 16
|
||||
r7 = r7 ^ dataset[r2 & MASK]; // 17
|
||||
r7 = r7 ^ simd_shuffle_xor(r3, (ushort)4); // 18
|
||||
r5 = r5 * r0; // 19
|
||||
r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 20
|
||||
r2 = r2 ^ simd_shuffle_xor(r4, (ushort)16); // 21
|
||||
r6 = mulhi(r6, r2); // 22
|
||||
r6 = r6 ^ dataset[r1 & MASK]; // 23
|
||||
r5 = r5 * r0; // 24
|
||||
r5 = rotl_imm(r5, 19u); // 25
|
||||
r7 = r7 ^ simd_shuffle_xor(r6, (ushort)2); // 26
|
||||
r0 = r0 ^ r5; // 27
|
||||
r0 = r0 ^ r4; // 28
|
||||
r3 = r3 - r0; // 29
|
||||
r5 = r5 * r1; // 30
|
||||
r7 = r7 ^ dataset[r2 & MASK]; // 31
|
||||
r1 = r1 ^ dataset[r0 & MASK]; // 32
|
||||
r5 = r5 ^ r6; // 33
|
||||
r5 = r5 ^ dataset[r1 & MASK]; // 34
|
||||
r0 = mulhi(r0, r5); // 35
|
||||
r5 = r5 ^ simd_shuffle_xor(r2, (ushort)4); // 36
|
||||
r7 = r7 ^ dataset[r0 & MASK]; // 37
|
||||
r3 = r3 + r1 + select(0x75ba2fadu, 0x230c005cu, ((sel >> 27u) & 1u) != 0u); // 38
|
||||
r1 = r1 ^ simd_shuffle_xor(r5, (ushort)4); // 39
|
||||
r2 = r2 ^ r5; // 40
|
||||
r3 = r6 * r3 + r3; // 41
|
||||
r6 = r6 - r7; // 42
|
||||
r7 = r7 ^ r0; // 43
|
||||
r1 = r1 ^ dataset[r7 & MASK]; // 44
|
||||
r2 = r2 * r3; // 45
|
||||
r1 = mulhi(r1, r5); // 46
|
||||
r4 = r4 - r3; // 47
|
||||
r2 = rotr_var(r2, r6); // 48
|
||||
r3 = r3 ^ dataset[r5 & MASK]; // 49
|
||||
r1 = r1 + r5 + select(0x81b8bc2cu, 0x1907970cu, ((sel >> 7u) & 1u) != 0u); // 50
|
||||
r0 = r0 * r2; // 51
|
||||
r0 = r0 + r2 + select(0x4f92b968u, 0x699fd448u, ((sel >> 6u) & 1u) != 0u); // 52
|
||||
r1 = r1 + r0 + select(0x2bb965afu, 0x77b1520du, ((sel >> 12u) & 1u) != 0u); // 53
|
||||
r7 = rotl_imm(r7, 14u); // 54
|
||||
r3 = r3 + r7 + select(0x7b0fe07au, 0xa54c55a0u, ((sel >> 1u) & 1u) != 0u); // 55
|
||||
r6 = r6 ^ dataset[r7 & MASK]; // 56
|
||||
r1 = rotr_var(r1, r5); // 57
|
||||
r5 = r5 ^ dataset[r4 & MASK]; // 58
|
||||
r6 = r6 ^ dataset[r2 & MASK]; // 59
|
||||
r3 = r5 * r0 + r3; // 60
|
||||
r5 = r5 + r7 + select(0xaf9dd72du, 0xad7493e7u, ((sel >> 31u) & 1u) != 0u); // 61
|
||||
r4 = r4 + r6 + select(0x89841d87u, 0x1e07c3d9u, ((sel >> 27u) & 1u) != 0u); // 62
|
||||
r5 = rotl_imm(r5, 19u); // 63
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
|
|
|
|||
111
proto-cuda/packs/igneum-genesis-mh/program_bound.metal
Normal file
111
proto-cuda/packs/igneum-genesis-mh/program_bound.metal
Normal file
|
|
@ -0,0 +1,111 @@
|
|||
#include <metal_stdlib>
|
||||
using namespace metal;
|
||||
|
||||
#define MASK 0x0fffffffu
|
||||
constant uint SEEDW[8] = { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u };
|
||||
|
||||
inline uint splitmix32(uint x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31
|
||||
inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
inline uint ds_elem(uint i, uint d0, uint d1) {
|
||||
uint x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// Header-bound variant: the init words come from buffer 3 (bind.rs), not from SEEDW.
|
||||
kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]],
|
||||
device ulong* out [[buffer(1)]],
|
||||
constant uint& baseNonce [[buffer(2)]],
|
||||
constant uint* initw [[buffer(3)]],
|
||||
uint gid [[thread_position_in_grid]]) {
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint x = nonce ^ initw[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ initw[1]; }
|
||||
{ uint x = nonce ^ initw[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ initw[2]; }
|
||||
{ uint x = nonce ^ initw[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ initw[3]; }
|
||||
{ uint x = nonce ^ initw[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ initw[4]; }
|
||||
{ uint x = nonce ^ initw[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ initw[5]; }
|
||||
{ uint x = nonce ^ initw[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ initw[6]; }
|
||||
{ uint x = nonce ^ initw[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ initw[7]; }
|
||||
{ uint x = nonce ^ initw[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ initw[0]; }
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r2 = r3 * r4 + r2; // 0
|
||||
r2 = r1 * r1 + r2; // 1
|
||||
r2 = r3 * r2 + r2; // 2
|
||||
r3 = r3 ^ r5; // 3
|
||||
r7 = r7 ^ dataset[r2 & MASK]; // 4
|
||||
r5 = r5 ^ dataset[r7 & MASK]; // 5
|
||||
r1 = r1 ^ simd_shuffle_xor(r4, (ushort)8); // 6
|
||||
r7 = r7 ^ simd_shuffle_xor(r3, (ushort)8); // 7
|
||||
r1 = mulhi(r1, r5); // 8
|
||||
r6 = rotr_var(r6, r3); // 9
|
||||
r3 = r3 | r4; // 10
|
||||
r4 = r4 ^ dataset[r3 & MASK]; // 11
|
||||
r0 = mulhi(r0, r4); // 12
|
||||
r5 = r5 + r1 + select(0xc7934706u, 0xd3177981u, ((sel >> 30u) & 1u) != 0u); // 13
|
||||
r0 = r0 ^ dataset[r4 & MASK]; // 14
|
||||
r2 = r2 - r4; // 15
|
||||
r2 = r2 ^ dataset[r0 & MASK]; // 16
|
||||
r7 = r7 ^ dataset[r2 & MASK]; // 17
|
||||
r7 = r7 ^ simd_shuffle_xor(r3, (ushort)4); // 18
|
||||
r5 = r5 * r0; // 19
|
||||
r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 20
|
||||
r2 = r2 ^ simd_shuffle_xor(r4, (ushort)16); // 21
|
||||
r6 = mulhi(r6, r2); // 22
|
||||
r6 = r6 ^ dataset[r1 & MASK]; // 23
|
||||
r5 = r5 * r0; // 24
|
||||
r5 = rotl_imm(r5, 19u); // 25
|
||||
r7 = r7 ^ simd_shuffle_xor(r6, (ushort)2); // 26
|
||||
r0 = r0 ^ r5; // 27
|
||||
r0 = r0 ^ r4; // 28
|
||||
r3 = r3 - r0; // 29
|
||||
r5 = r5 * r1; // 30
|
||||
r7 = r7 ^ dataset[r2 & MASK]; // 31
|
||||
r1 = r1 ^ dataset[r0 & MASK]; // 32
|
||||
r5 = r5 ^ r6; // 33
|
||||
r5 = r5 ^ dataset[r1 & MASK]; // 34
|
||||
r0 = mulhi(r0, r5); // 35
|
||||
r5 = r5 ^ simd_shuffle_xor(r2, (ushort)4); // 36
|
||||
r7 = r7 ^ dataset[r0 & MASK]; // 37
|
||||
r3 = r3 + r1 + select(0x75ba2fadu, 0x230c005cu, ((sel >> 27u) & 1u) != 0u); // 38
|
||||
r1 = r1 ^ simd_shuffle_xor(r5, (ushort)4); // 39
|
||||
r2 = r2 ^ r5; // 40
|
||||
r3 = r6 * r3 + r3; // 41
|
||||
r6 = r6 - r7; // 42
|
||||
r7 = r7 ^ r0; // 43
|
||||
r1 = r1 ^ dataset[r7 & MASK]; // 44
|
||||
r2 = r2 * r3; // 45
|
||||
r1 = mulhi(r1, r5); // 46
|
||||
r4 = r4 - r3; // 47
|
||||
r2 = rotr_var(r2, r6); // 48
|
||||
r3 = r3 ^ dataset[r5 & MASK]; // 49
|
||||
r1 = r1 + r5 + select(0x81b8bc2cu, 0x1907970cu, ((sel >> 7u) & 1u) != 0u); // 50
|
||||
r0 = r0 * r2; // 51
|
||||
r0 = r0 + r2 + select(0x4f92b968u, 0x699fd448u, ((sel >> 6u) & 1u) != 0u); // 52
|
||||
r1 = r1 + r0 + select(0x2bb965afu, 0x77b1520du, ((sel >> 12u) & 1u) != 0u); // 53
|
||||
r7 = rotl_imm(r7, 14u); // 54
|
||||
r3 = r3 + r7 + select(0x7b0fe07au, 0xa54c55a0u, ((sel >> 1u) & 1u) != 0u); // 55
|
||||
r6 = r6 ^ dataset[r7 & MASK]; // 56
|
||||
r1 = rotr_var(r1, r5); // 57
|
||||
r5 = r5 ^ dataset[r4 & MASK]; // 58
|
||||
r6 = r6 ^ dataset[r2 & MASK]; // 59
|
||||
r3 = r5 * r0 + r3; // 60
|
||||
r5 = r5 + r7 + select(0xaf9dd72du, 0xad7493e7u, ((sel >> 31u) & 1u) != 0u); // 61
|
||||
r4 = r4 + r6 + select(0x89841d87u, 0x1e07c3d9u, ((sel >> 27u) & 1u) != 0u); // 62
|
||||
r5 = rotl_imm(r5, 19u); // 63
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
|
|
@ -1,5 +1,5 @@
|
|||
// Generated by proto-metal/igneum-bench --export-pack for seed "igneum-genesis". Do not edit by hand.
|
||||
// Expected outputs: proto-metal CPU interpreter (cpuWarp, memory-hard dataset) on Apple M5 Max; Metal GPU cross-check PASS 3/3 warps
|
||||
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
|
||||
// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset
|
||||
#pragma once
|
||||
#ifdef __cplusplus
|
||||
#include <cstdint>
|
||||
|
|
@ -11,22 +11,22 @@
|
|||
static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u };
|
||||
static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = {
|
||||
{ // base nonce 0
|
||||
0x1fb0b3bbc1ac8279ull, 0x61533759ca995ac8ull, 0x97cd15ec31242001ull, 0x3c19e71e2dbff828ull, 0xc7aba60b6c2ab016ull, 0x8f13725d2a59bf84ull, 0x7b826b82fdec5a2full, 0x3131d7817418c04eull,
|
||||
0xd7c7e57c6aaac948ull, 0xb3b639431ebdd3d6ull, 0x400af448658a1d56ull, 0x17f5760b8cadb8fdull, 0x01ded5c4894c411cull, 0xd6b3fdbc57bdb128ull, 0x888efb103cecb983ull, 0xac73538353a340c4ull,
|
||||
0x3f9914ad9445f052ull, 0x8a860b28afd3dd42ull, 0xc2a122ef9c1fa330ull, 0x96c4e58663f82dceull, 0xbaae4b9d6a3c4320ull, 0xcd6c2dd653e07743ull, 0x7d3d00f0fb46325bull, 0x7636224050132b09ull,
|
||||
0xdc5d76c741cc60f4ull, 0xaef3031b0d4b701cull, 0x735987104a975f5full, 0xa58fce4010bf99ddull, 0xb16c863d2ffdc0ddull, 0x0337b56e39529c90ull, 0x802ac0bc1d0c0696ull, 0xfa052263a854f3deull
|
||||
0x42246ba99fc58e4full, 0x19561eec0db4f9f4ull, 0x11a4fb70ea7b688full, 0x8872bfadec960949ull, 0x085e6cf2ff6b8303ull, 0x780e0d76e504fe6cull, 0x7bb05a713f8283faull, 0xfef7beeffd5ab1d1ull,
|
||||
0x1b83afb12e78d65bull, 0x2576c5384a2e1caeull, 0x549d620a0735d6d8ull, 0x2eb4e32f32e8d5f1ull, 0x81bb414492f69586ull, 0x092d17324b465e01ull, 0x8bde350ee2354b5aull, 0xd64b47e9c9e0ec07ull,
|
||||
0xdbbf1b78b9979a3full, 0xbee382abde89c111ull, 0x6b598d99c16a8c70ull, 0x69d72da59fb3b6e9ull, 0x1e309d2a549632faull, 0x98fa8255ab65f005ull, 0x2f48ab1bb516110cull, 0x7d2af17cadd18beaull,
|
||||
0x547c5978bd005e02ull, 0x25ea2e21e88ea9d3ull, 0x1c575ec6e43efc58ull, 0x077b80cb079958b4ull, 0xa96eebd0c8634981ull, 0x44b18bdf19eb7838ull, 0xc3ccb9fe9ef5953cull, 0xb08446b1f2de7793ull
|
||||
},
|
||||
{ // base nonce 4096
|
||||
0xf219cf7ecf6ec450ull, 0xb575ab4b388c01f7ull, 0x09ff62fd27055f6dull, 0x0c8abc61759d3564ull, 0xa5909acbcc981201ull, 0x1272880e65dc626dull, 0xf49baed726dabd3aull, 0xd9bf25b9d7cd2c59ull,
|
||||
0xc1a453f0e2a0e26eull, 0x9c8bd1806912812full, 0xc7bc5603edeead7eull, 0x9a0218abf9fa1fd0ull, 0x01affc92769181b7ull, 0xe54fc07ff716c78aull, 0x70606c49b68b2fb3ull, 0x51a7e9c37930bcc3ull,
|
||||
0x47540e9d54ce6117ull, 0xd1dfe07ffafb6952ull, 0x701a455888ed0a2aull, 0x4baec8d5ffb99fadull, 0xc6dffc6807b38113ull, 0xdc9c3cabc5965b78ull, 0xa3e4e56eee3fe8d6ull, 0xf6decd4c32d0b018ull,
|
||||
0x85cd622043eb59b1ull, 0x4642d5079564dcdfull, 0x84e19ac8eefef8acull, 0x2f4414ad779564deull, 0xa288c93ca1c51b26ull, 0xa977d793f289fe1cull, 0xa1bc6f71f19384ffull, 0xd5b198435d9c40daull
|
||||
0x3d3903e310ca038full, 0x9e72b9a86ebb29e4ull, 0x1dca847362709effull, 0x4d0b021e131d6a4full, 0x88de07bd1539fa72ull, 0x042a436373b96ffeull, 0x3fcb1ef447b976fdull, 0x0237159348bd27fbull,
|
||||
0x282fb8b521df5e60ull, 0xc16d95f79fd99903ull, 0xd75a20a4d69df082ull, 0x809381bf5581fd22ull, 0xea6e9ed106945107ull, 0x3ef5fadeb7ac79d9ull, 0x75a4c882c4d02e44ull, 0x66b95d3105a88bf4ull,
|
||||
0xca7e020f5d415c29ull, 0x32dc62b7072cd5e4ull, 0x253fd43c6e7c2d57ull, 0x228235093e06d1afull, 0x1085653fe1d382a4ull, 0xe6bcd4d4693fd10dull, 0x065058e340b134a5ull, 0x4febdd41ea5e1913ull,
|
||||
0x8a9fba060f8c1d9full, 0x2be8ec07bf80f713ull, 0x476855e9a37e1b80ull, 0x5830538ad9e75e0aull, 0xb32a3ac396dc34bcull, 0xf521612e59ee414eull, 0xd5687c5ae184a180ull, 0x61c242509efdccddull
|
||||
},
|
||||
{ // base nonce 1000000
|
||||
0x8a6f7a32b06edb52ull, 0x362d2b789a0c0ebfull, 0x944ec8bd1d08c026ull, 0x6c40b67428eb7af3ull, 0x306e546583cb011eull, 0x27a55ab280344f21ull, 0xa6247ff1d8e72bceull, 0xf11780aa78365f00ull,
|
||||
0x2919860ba64f212aull, 0x5cfa0a33f4dd691dull, 0xce626a9cc0d2427dull, 0xf42f1047cdfbcb8bull, 0xd247b8ab5c35aeaaull, 0x47989e53b0251facull, 0xb55e4e5fa8df8217ull, 0x66b17338a66756caull,
|
||||
0x34b56d64da199729ull, 0x167a7c11aad8f00dull, 0x4292a627e69d80cbull, 0xce30b0820c123a4bull, 0xcd3b7367ba86adadull, 0xb0c1b22a10274c28ull, 0x7326d56bc7140e08ull, 0xea6490cba6bb2b4full,
|
||||
0xab57b3e93a1dae5cull, 0x2f3c7c5a624a8f0aull, 0x1f9a062062647b0bull, 0x83cf530898ba80f8ull, 0xa9f28306c97622baull, 0xe4212399129bddeeull, 0x8a2e1f8094e0a1f4ull, 0x7040299873672a35ull
|
||||
0xf218c1bd58e6dfe0ull, 0x8c5a362ee98971c2ull, 0x06394d22d03aa4eeull, 0x68166b204b0cf6ddull, 0x6e8e505080b2de7aull, 0xe922e3a6684d8c0aull, 0x89d0c35410c44bc9ull, 0x5a3df3edbf4d62a8ull,
|
||||
0xecf6b45e5a90ce00ull, 0x02221b7088fafc7full, 0x5d23dd2a7a4d3e16ull, 0x04e8e12e17a940f3ull, 0x4f671f11554578a6ull, 0x4e49f1b1117a97abull, 0xec0fb9eac4520ae9ull, 0x4ca8b3859d421f9bull,
|
||||
0xbf593793d02b0309ull, 0xcb11dc0a40088de6ull, 0xaa93a48f712fb2adull, 0xda314ca08359d2baull, 0xfba7742340db6445ull, 0xf0138acf64f37cb2ull, 0xfc307731b74f1d99ull, 0x660553bd0b06bd63ull,
|
||||
0xbecb2b151282f602ull, 0x373bead8040e3eb8ull, 0xbe39a4642591f3dfull, 0xb6c5a400b6586ff3ull, 0xd614efa47d7118e3ull, 0x1be1d5c73feceb6eull, 0xca6e33dd721350c4ull, 0x6c3b2c11adfbfcacull
|
||||
}
|
||||
};
|
||||
|
||||
|
|
|
|||
|
|
@ -5,25 +5,25 @@
|
|||
"dataset_log2_words": 28,
|
||||
"mask": "0x0fffffff",
|
||||
"lanes": 32,
|
||||
"source": "proto-metal CPU interpreter (cpuWarp, memory-hard dataset) on Apple M5 Max; Metal GPU cross-check PASS 3/3 warps",
|
||||
"source": "igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset",
|
||||
"warps": [
|
||||
{"base_nonce": 0, "expected": [
|
||||
"0x1fb0b3bbc1ac8279", "0x61533759ca995ac8", "0x97cd15ec31242001", "0x3c19e71e2dbff828", "0xc7aba60b6c2ab016", "0x8f13725d2a59bf84", "0x7b826b82fdec5a2f", "0x3131d7817418c04e",
|
||||
"0xd7c7e57c6aaac948", "0xb3b639431ebdd3d6", "0x400af448658a1d56", "0x17f5760b8cadb8fd", "0x01ded5c4894c411c", "0xd6b3fdbc57bdb128", "0x888efb103cecb983", "0xac73538353a340c4",
|
||||
"0x3f9914ad9445f052", "0x8a860b28afd3dd42", "0xc2a122ef9c1fa330", "0x96c4e58663f82dce", "0xbaae4b9d6a3c4320", "0xcd6c2dd653e07743", "0x7d3d00f0fb46325b", "0x7636224050132b09",
|
||||
"0xdc5d76c741cc60f4", "0xaef3031b0d4b701c", "0x735987104a975f5f", "0xa58fce4010bf99dd", "0xb16c863d2ffdc0dd", "0x0337b56e39529c90", "0x802ac0bc1d0c0696", "0xfa052263a854f3de"
|
||||
"0x42246ba99fc58e4f", "0x19561eec0db4f9f4", "0x11a4fb70ea7b688f", "0x8872bfadec960949", "0x085e6cf2ff6b8303", "0x780e0d76e504fe6c", "0x7bb05a713f8283fa", "0xfef7beeffd5ab1d1",
|
||||
"0x1b83afb12e78d65b", "0x2576c5384a2e1cae", "0x549d620a0735d6d8", "0x2eb4e32f32e8d5f1", "0x81bb414492f69586", "0x092d17324b465e01", "0x8bde350ee2354b5a", "0xd64b47e9c9e0ec07",
|
||||
"0xdbbf1b78b9979a3f", "0xbee382abde89c111", "0x6b598d99c16a8c70", "0x69d72da59fb3b6e9", "0x1e309d2a549632fa", "0x98fa8255ab65f005", "0x2f48ab1bb516110c", "0x7d2af17cadd18bea",
|
||||
"0x547c5978bd005e02", "0x25ea2e21e88ea9d3", "0x1c575ec6e43efc58", "0x077b80cb079958b4", "0xa96eebd0c8634981", "0x44b18bdf19eb7838", "0xc3ccb9fe9ef5953c", "0xb08446b1f2de7793"
|
||||
]},
|
||||
{"base_nonce": 4096, "expected": [
|
||||
"0xf219cf7ecf6ec450", "0xb575ab4b388c01f7", "0x09ff62fd27055f6d", "0x0c8abc61759d3564", "0xa5909acbcc981201", "0x1272880e65dc626d", "0xf49baed726dabd3a", "0xd9bf25b9d7cd2c59",
|
||||
"0xc1a453f0e2a0e26e", "0x9c8bd1806912812f", "0xc7bc5603edeead7e", "0x9a0218abf9fa1fd0", "0x01affc92769181b7", "0xe54fc07ff716c78a", "0x70606c49b68b2fb3", "0x51a7e9c37930bcc3",
|
||||
"0x47540e9d54ce6117", "0xd1dfe07ffafb6952", "0x701a455888ed0a2a", "0x4baec8d5ffb99fad", "0xc6dffc6807b38113", "0xdc9c3cabc5965b78", "0xa3e4e56eee3fe8d6", "0xf6decd4c32d0b018",
|
||||
"0x85cd622043eb59b1", "0x4642d5079564dcdf", "0x84e19ac8eefef8ac", "0x2f4414ad779564de", "0xa288c93ca1c51b26", "0xa977d793f289fe1c", "0xa1bc6f71f19384ff", "0xd5b198435d9c40da"
|
||||
"0x3d3903e310ca038f", "0x9e72b9a86ebb29e4", "0x1dca847362709eff", "0x4d0b021e131d6a4f", "0x88de07bd1539fa72", "0x042a436373b96ffe", "0x3fcb1ef447b976fd", "0x0237159348bd27fb",
|
||||
"0x282fb8b521df5e60", "0xc16d95f79fd99903", "0xd75a20a4d69df082", "0x809381bf5581fd22", "0xea6e9ed106945107", "0x3ef5fadeb7ac79d9", "0x75a4c882c4d02e44", "0x66b95d3105a88bf4",
|
||||
"0xca7e020f5d415c29", "0x32dc62b7072cd5e4", "0x253fd43c6e7c2d57", "0x228235093e06d1af", "0x1085653fe1d382a4", "0xe6bcd4d4693fd10d", "0x065058e340b134a5", "0x4febdd41ea5e1913",
|
||||
"0x8a9fba060f8c1d9f", "0x2be8ec07bf80f713", "0x476855e9a37e1b80", "0x5830538ad9e75e0a", "0xb32a3ac396dc34bc", "0xf521612e59ee414e", "0xd5687c5ae184a180", "0x61c242509efdccdd"
|
||||
]},
|
||||
{"base_nonce": 1000000, "expected": [
|
||||
"0x8a6f7a32b06edb52", "0x362d2b789a0c0ebf", "0x944ec8bd1d08c026", "0x6c40b67428eb7af3", "0x306e546583cb011e", "0x27a55ab280344f21", "0xa6247ff1d8e72bce", "0xf11780aa78365f00",
|
||||
"0x2919860ba64f212a", "0x5cfa0a33f4dd691d", "0xce626a9cc0d2427d", "0xf42f1047cdfbcb8b", "0xd247b8ab5c35aeaa", "0x47989e53b0251fac", "0xb55e4e5fa8df8217", "0x66b17338a66756ca",
|
||||
"0x34b56d64da199729", "0x167a7c11aad8f00d", "0x4292a627e69d80cb", "0xce30b0820c123a4b", "0xcd3b7367ba86adad", "0xb0c1b22a10274c28", "0x7326d56bc7140e08", "0xea6490cba6bb2b4f",
|
||||
"0xab57b3e93a1dae5c", "0x2f3c7c5a624a8f0a", "0x1f9a062062647b0b", "0x83cf530898ba80f8", "0xa9f28306c97622ba", "0xe4212399129bddee", "0x8a2e1f8094e0a1f4", "0x7040299873672a35"
|
||||
"0xf218c1bd58e6dfe0", "0x8c5a362ee98971c2", "0x06394d22d03aa4ee", "0x68166b204b0cf6dd", "0x6e8e505080b2de7a", "0xe922e3a6684d8c0a", "0x89d0c35410c44bc9", "0x5a3df3edbf4d62a8",
|
||||
"0xecf6b45e5a90ce00", "0x02221b7088fafc7f", "0x5d23dd2a7a4d3e16", "0x04e8e12e17a940f3", "0x4f671f11554578a6", "0x4e49f1b1117a97ab", "0xec0fb9eac4520ae9", "0x4ca8b3859d421f9b",
|
||||
"0xbf593793d02b0309", "0xcb11dc0a40088de6", "0xaa93a48f712fb2ad", "0xda314ca08359d2ba", "0xfba7742340db6445", "0xf0138acf64f37cb2", "0xfc307731b74f1d99", "0x660553bd0b06bd63",
|
||||
"0xbecb2b151282f602", "0x373bead8040e3eb8", "0xbe39a4642591f3df", "0xb6c5a400b6586ff3", "0xd614efa47d7118e3", "0x1be1d5c73feceb6e", "0xca6e33dd721350c4", "0x6c3b2c11adfbfcac"
|
||||
]}
|
||||
],
|
||||
"dataset_head": ["0xffc3cd94", "0x5920ccd8", "0x392f44bb", "0x5e57f67a", "0x2f2bc2a9", "0x620b0e36", "0xbdc09014", "0x436654bf", "0x311e0b48", "0x1abd93ad", "0x59cc7ce8", "0xee5247b2", "0x86171fe8", "0x6d874751", "0xc9f7728f", "0x7c2a435d"],
|
||||
|
|
|
|||
|
|
@ -1,4 +1,4 @@
|
|||
// Generated by proto-metal/igneum-bench --export-pack for seed "igneum-genesis". Do not edit by hand.
|
||||
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
|
||||
// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal).
|
||||
// Built from source at runtime by proto-opencl/host.c, which passes these defines:
|
||||
// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit)
|
||||
|
|
@ -96,70 +96,70 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out
|
|||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = rotl_imm(r4, 25u); // 0 rotl
|
||||
r0 = r0 - r5; // 1 sub
|
||||
r4 = r4 ^ ds[r3 & mask]; // 2 load
|
||||
r1 = rotl_imm(r1, 1u); // 3 rotl
|
||||
r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 4 add
|
||||
r5 = r5 ^ ds[r3 & mask]; // 5 load
|
||||
r5 = r5 - r7; // 6 sub
|
||||
r3 = r3 + r4 + ((((sel >> 26u) & 1u) != 0u) ? 0x5a069596u : 0x52f2dbf4u); // 7 add
|
||||
r0 = r0 ^ r4; // 8 xor
|
||||
r4 = r4 ^ r0; // 9 xor
|
||||
r2 = r2 ^ ds[r0 & mask]; // 10 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 16u); r4 = r4 ^ t_; } // 11 shfl
|
||||
r1 = r1 - r5; // 12 sub
|
||||
r2 = r2 ^ r1; // 13 xor
|
||||
r4 = r4 ^ ds[r5 & mask]; // 14 load
|
||||
r2 = r2 ^ ds[r4 & mask]; // 15 load
|
||||
r3 = r3 ^ ds[r0 & mask]; // 16 load
|
||||
r4 = r4 ^ r6; // 17 xor
|
||||
r2 = r4 * r6 + r2; // 18 mad
|
||||
r6 = rotr_var(r6, r1); // 19 rotr
|
||||
r3 = r3 ^ r4; // 20 xor
|
||||
r1 = r3 * r5 + r1; // 21 mad
|
||||
r7 = mul_hi(r7, r4); // 22 mulhi
|
||||
r5 = mul_hi(r5, r2); // 23 mulhi
|
||||
r0 = r0 ^ ds[r6 & mask]; // 24 load
|
||||
r5 = r5 * r6; // 25 mul
|
||||
r7 = r7 ^ ds[r1 & mask]; // 26 load
|
||||
r3 = rotr_var(r3, r1); // 27 rotr
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r1, 8u); r5 = r5 ^ t_; } // 28 shfl
|
||||
r7 = r7 ^ r5; // 29 xor
|
||||
r7 = rotl_imm(r7, 23u); // 30 rotl
|
||||
r2 = r2 - r0; // 31 sub
|
||||
r7 = r7 ^ r2; // 32 xor
|
||||
r2 = r2 ^ r6; // 33 xor
|
||||
r6 = r6 ^ ds[r1 & mask]; // 34 load
|
||||
r1 = r1 ^ r4; // 35 xor
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r0, 1u); r3 = r3 ^ t_; } // 36 shfl
|
||||
r2 = r2 * r6; // 37 mul
|
||||
r5 = r5 + r3 + ((((sel >> 12u) & 1u) != 0u) ? 0xa29f4338u : 0x71f30417u); // 38 add
|
||||
r7 = r7 ^ r6; // 39 xor
|
||||
r7 = r7 ^ r3; // 40 xor
|
||||
r3 = rotr_var(r3, r4); // 41 rotr
|
||||
r5 = r5 ^ r3; // 42 xor
|
||||
r3 = rotr_var(r3, r6); // 43 rotr
|
||||
r1 = r3 * r5 + r1; // 44 mad
|
||||
r7 = r7 + r3 + ((((sel >> 29u) & 1u) != 0u) ? 0xa907b90bu : 0xc1ae8d3bu); // 45 add
|
||||
r7 = r7 | r2; // 46 or
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 16u); r2 = r2 ^ t_; } // 47 shfl
|
||||
r1 = r1 ^ ds[r7 & mask]; // 48 load
|
||||
r5 = r5 - r1; // 49 sub
|
||||
r3 = r3 ^ ds[r1 & mask]; // 50 load
|
||||
r2 = r2 - r3; // 51 sub
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r2, 2u); r6 = r6 ^ t_; } // 52 shfl
|
||||
r2 = r2 - r0; // 53 sub
|
||||
r0 = r0 ^ ds[r3 & mask]; // 54 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r1, 16u); r2 = r2 ^ t_; } // 55 shfl
|
||||
r0 = r0 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0xf4689674u : 0x25955401u); // 56 add
|
||||
r2 = mul_hi(r2, r0); // 57 mulhi
|
||||
r4 = mul_hi(r4, r2); // 58 mulhi
|
||||
r2 = r6 * r7 + r2; // 59 mad
|
||||
r3 = r3 ^ r1; // 60 xor
|
||||
r4 = r4 * r2; // 61 mul
|
||||
r0 = r0 ^ ds[r2 & mask]; // 62 load
|
||||
r7 = mul_hi(r7, r0); // 63 mulhi
|
||||
r2 = r3 * r4 + r2; // 0 mad
|
||||
r2 = r1 * r1 + r2; // 1 mad
|
||||
r2 = r3 * r2 + r2; // 2 mad
|
||||
r3 = r3 ^ r5; // 3 xor
|
||||
r7 = r7 ^ ds[r2 & mask]; // 4 load
|
||||
r5 = r5 ^ ds[r7 & mask]; // 5 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r1 = r1 ^ t_; } // 6 shfl
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 8u); r7 = r7 ^ t_; } // 7 shfl
|
||||
r1 = mul_hi(r1, r5); // 8 mulhi
|
||||
r6 = rotr_var(r6, r3); // 9 rotr
|
||||
r3 = r3 | r4; // 10 or
|
||||
r4 = r4 ^ ds[r3 & mask]; // 11 load
|
||||
r0 = mul_hi(r0, r4); // 12 mulhi
|
||||
r5 = r5 + r1 + ((((sel >> 30u) & 1u) != 0u) ? 0xd3177981u : 0xc7934706u); // 13 add
|
||||
r0 = r0 ^ ds[r4 & mask]; // 14 load
|
||||
r2 = r2 - r4; // 15 sub
|
||||
r2 = r2 ^ ds[r0 & mask]; // 16 load
|
||||
r7 = r7 ^ ds[r2 & mask]; // 17 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 18 shfl
|
||||
r5 = r5 * r0; // 19 mul
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 20 shfl
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 16u); r2 = r2 ^ t_; } // 21 shfl
|
||||
r6 = mul_hi(r6, r2); // 22 mulhi
|
||||
r6 = r6 ^ ds[r1 & mask]; // 23 load
|
||||
r5 = r5 * r0; // 24 mul
|
||||
r5 = rotl_imm(r5, 19u); // 25 rotl
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 2u); r7 = r7 ^ t_; } // 26 shfl
|
||||
r0 = r0 ^ r5; // 27 xor
|
||||
r0 = r0 ^ r4; // 28 xor
|
||||
r3 = r3 - r0; // 29 sub
|
||||
r5 = r5 * r1; // 30 mul
|
||||
r7 = r7 ^ ds[r2 & mask]; // 31 load
|
||||
r1 = r1 ^ ds[r0 & mask]; // 32 load
|
||||
r5 = r5 ^ r6; // 33 xor
|
||||
r5 = r5 ^ ds[r1 & mask]; // 34 load
|
||||
r0 = mul_hi(r0, r5); // 35 mulhi
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r2, 4u); r5 = r5 ^ t_; } // 36 shfl
|
||||
r7 = r7 ^ ds[r0 & mask]; // 37 load
|
||||
r3 = r3 + r1 + ((((sel >> 27u) & 1u) != 0u) ? 0x230c005cu : 0x75ba2fadu); // 38 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 4u); r1 = r1 ^ t_; } // 39 shfl
|
||||
r2 = r2 ^ r5; // 40 xor
|
||||
r3 = r6 * r3 + r3; // 41 mad
|
||||
r6 = r6 - r7; // 42 sub
|
||||
r7 = r7 ^ r0; // 43 xor
|
||||
r1 = r1 ^ ds[r7 & mask]; // 44 load
|
||||
r2 = r2 * r3; // 45 mul
|
||||
r1 = mul_hi(r1, r5); // 46 mulhi
|
||||
r4 = r4 - r3; // 47 sub
|
||||
r2 = rotr_var(r2, r6); // 48 rotr
|
||||
r3 = r3 ^ ds[r5 & mask]; // 49 load
|
||||
r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 50 add
|
||||
r0 = r0 * r2; // 51 mul
|
||||
r0 = r0 + r2 + ((((sel >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); // 52 add
|
||||
r1 = r1 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0x77b1520du : 0x2bb965afu); // 53 add
|
||||
r7 = rotl_imm(r7, 14u); // 54 rotl
|
||||
r3 = r3 + r7 + ((((sel >> 1u) & 1u) != 0u) ? 0xa54c55a0u : 0x7b0fe07au); // 55 add
|
||||
r6 = r6 ^ ds[r7 & mask]; // 56 load
|
||||
r1 = rotr_var(r1, r5); // 57 rotr
|
||||
r5 = r5 ^ ds[r4 & mask]; // 58 load
|
||||
r6 = r6 ^ ds[r2 & mask]; // 59 load
|
||||
r3 = r5 * r0 + r3; // 60 mad
|
||||
r5 = r5 + r7 + ((((sel >> 31u) & 1u) != 0u) ? 0xad7493e7u : 0xaf9dd72du); // 61 add
|
||||
r4 = r4 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); // 62 add
|
||||
r5 = rotl_imm(r5, 19u); // 63 rotl
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
|
|
|
|||
|
|
@ -1,4 +1,4 @@
|
|||
// Generated by proto-metal/igneum-bench --export-pack for seed "igneum-genesis". Do not edit by hand.
|
||||
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
|
||||
// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal).
|
||||
// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC.
|
||||
#include <cuda_runtime.h>
|
||||
|
|
@ -48,70 +48,70 @@ __global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonc
|
|||
|
||||
for (uint32_t it = 0u; it < 8u; ++it) {
|
||||
uint32_t sel = r0;
|
||||
r4 = rotl_imm(r4, 25u); // 0 rotl
|
||||
r0 = r0 - r5; // 1 sub
|
||||
r4 = r4 ^ ds[r3 & mask]; // 2 load
|
||||
r1 = rotl_imm(r1, 1u); // 3 rotl
|
||||
r2 = r2 + r3 + ((((sel >> 26u) & 1u) != 0u) ? 0x2735a174u : 0x61f0b51cu); // 4 add
|
||||
r5 = r5 ^ ds[r3 & mask]; // 5 load
|
||||
r5 = r5 - r7; // 6 sub
|
||||
r3 = r3 + r4 + ((((sel >> 26u) & 1u) != 0u) ? 0x5a069596u : 0x52f2dbf4u); // 7 add
|
||||
r0 = r0 ^ r4; // 8 xor
|
||||
r4 = r4 ^ r0; // 9 xor
|
||||
r2 = r2 ^ ds[r0 & mask]; // 10 load
|
||||
r4 = r4 ^ __shfl_xor_sync(0xffffffffu, r6, 16); // 11 shfl
|
||||
r1 = r1 - r5; // 12 sub
|
||||
r2 = r2 ^ r1; // 13 xor
|
||||
r4 = r4 ^ ds[r5 & mask]; // 14 load
|
||||
r2 = r2 ^ ds[r4 & mask]; // 15 load
|
||||
r3 = r3 ^ ds[r0 & mask]; // 16 load
|
||||
r4 = r4 ^ r6; // 17 xor
|
||||
r2 = r4 * r6 + r2; // 18 mad
|
||||
r6 = rotr_var(r6, r1); // 19 rotr
|
||||
r3 = r3 ^ r4; // 20 xor
|
||||
r1 = r3 * r5 + r1; // 21 mad
|
||||
r7 = __umulhi(r7, r4); // 22 mulhi
|
||||
r5 = __umulhi(r5, r2); // 23 mulhi
|
||||
r0 = r0 ^ ds[r6 & mask]; // 24 load
|
||||
r5 = r5 * r6; // 25 mul
|
||||
r7 = r7 ^ ds[r1 & mask]; // 26 load
|
||||
r3 = rotr_var(r3, r1); // 27 rotr
|
||||
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r1, 8); // 28 shfl
|
||||
r7 = r7 ^ r5; // 29 xor
|
||||
r7 = rotl_imm(r7, 23u); // 30 rotl
|
||||
r2 = r2 - r0; // 31 sub
|
||||
r7 = r7 ^ r2; // 32 xor
|
||||
r2 = r2 ^ r6; // 33 xor
|
||||
r6 = r6 ^ ds[r1 & mask]; // 34 load
|
||||
r1 = r1 ^ r4; // 35 xor
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r0, 1); // 36 shfl
|
||||
r2 = r2 * r6; // 37 mul
|
||||
r5 = r5 + r3 + ((((sel >> 12u) & 1u) != 0u) ? 0xa29f4338u : 0x71f30417u); // 38 add
|
||||
r7 = r7 ^ r6; // 39 xor
|
||||
r7 = r7 ^ r3; // 40 xor
|
||||
r3 = rotr_var(r3, r4); // 41 rotr
|
||||
r5 = r5 ^ r3; // 42 xor
|
||||
r3 = rotr_var(r3, r6); // 43 rotr
|
||||
r1 = r3 * r5 + r1; // 44 mad
|
||||
r7 = r7 + r3 + ((((sel >> 29u) & 1u) != 0u) ? 0xa907b90bu : 0xc1ae8d3bu); // 45 add
|
||||
r7 = r7 | r2; // 46 or
|
||||
r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r3, 16); // 47 shfl
|
||||
r1 = r1 ^ ds[r7 & mask]; // 48 load
|
||||
r5 = r5 - r1; // 49 sub
|
||||
r3 = r3 ^ ds[r1 & mask]; // 50 load
|
||||
r2 = r2 - r3; // 51 sub
|
||||
r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r2, 2); // 52 shfl
|
||||
r2 = r2 - r0; // 53 sub
|
||||
r0 = r0 ^ ds[r3 & mask]; // 54 load
|
||||
r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r1, 16); // 55 shfl
|
||||
r0 = r0 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0xf4689674u : 0x25955401u); // 56 add
|
||||
r2 = __umulhi(r2, r0); // 57 mulhi
|
||||
r4 = __umulhi(r4, r2); // 58 mulhi
|
||||
r2 = r6 * r7 + r2; // 59 mad
|
||||
r3 = r3 ^ r1; // 60 xor
|
||||
r4 = r4 * r2; // 61 mul
|
||||
r0 = r0 ^ ds[r2 & mask]; // 62 load
|
||||
r7 = __umulhi(r7, r0); // 63 mulhi
|
||||
r2 = r3 * r4 + r2; // 0 mad
|
||||
r2 = r1 * r1 + r2; // 1 mad
|
||||
r2 = r3 * r2 + r2; // 2 mad
|
||||
r3 = r3 ^ r5; // 3 xor
|
||||
r7 = r7 ^ ds[r2 & mask]; // 4 load
|
||||
r5 = r5 ^ ds[r7 & mask]; // 5 load
|
||||
r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 6 shfl
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 8); // 7 shfl
|
||||
r1 = __umulhi(r1, r5); // 8 mulhi
|
||||
r6 = rotr_var(r6, r3); // 9 rotr
|
||||
r3 = r3 | r4; // 10 or
|
||||
r4 = r4 ^ ds[r3 & mask]; // 11 load
|
||||
r0 = __umulhi(r0, r4); // 12 mulhi
|
||||
r5 = r5 + r1 + ((((sel >> 30u) & 1u) != 0u) ? 0xd3177981u : 0xc7934706u); // 13 add
|
||||
r0 = r0 ^ ds[r4 & mask]; // 14 load
|
||||
r2 = r2 - r4; // 15 sub
|
||||
r2 = r2 ^ ds[r0 & mask]; // 16 load
|
||||
r7 = r7 ^ ds[r2 & mask]; // 17 load
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 18 shfl
|
||||
r5 = r5 * r0; // 19 mul
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 20 shfl
|
||||
r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r4, 16); // 21 shfl
|
||||
r6 = __umulhi(r6, r2); // 22 mulhi
|
||||
r6 = r6 ^ ds[r1 & mask]; // 23 load
|
||||
r5 = r5 * r0; // 24 mul
|
||||
r5 = rotl_imm(r5, 19u); // 25 rotl
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r6, 2); // 26 shfl
|
||||
r0 = r0 ^ r5; // 27 xor
|
||||
r0 = r0 ^ r4; // 28 xor
|
||||
r3 = r3 - r0; // 29 sub
|
||||
r5 = r5 * r1; // 30 mul
|
||||
r7 = r7 ^ ds[r2 & mask]; // 31 load
|
||||
r1 = r1 ^ ds[r0 & mask]; // 32 load
|
||||
r5 = r5 ^ r6; // 33 xor
|
||||
r5 = r5 ^ ds[r1 & mask]; // 34 load
|
||||
r0 = __umulhi(r0, r5); // 35 mulhi
|
||||
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r2, 4); // 36 shfl
|
||||
r7 = r7 ^ ds[r0 & mask]; // 37 load
|
||||
r3 = r3 + r1 + ((((sel >> 27u) & 1u) != 0u) ? 0x230c005cu : 0x75ba2fadu); // 38 add
|
||||
r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r5, 4); // 39 shfl
|
||||
r2 = r2 ^ r5; // 40 xor
|
||||
r3 = r6 * r3 + r3; // 41 mad
|
||||
r6 = r6 - r7; // 42 sub
|
||||
r7 = r7 ^ r0; // 43 xor
|
||||
r1 = r1 ^ ds[r7 & mask]; // 44 load
|
||||
r2 = r2 * r3; // 45 mul
|
||||
r1 = __umulhi(r1, r5); // 46 mulhi
|
||||
r4 = r4 - r3; // 47 sub
|
||||
r2 = rotr_var(r2, r6); // 48 rotr
|
||||
r3 = r3 ^ ds[r5 & mask]; // 49 load
|
||||
r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 50 add
|
||||
r0 = r0 * r2; // 51 mul
|
||||
r0 = r0 + r2 + ((((sel >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); // 52 add
|
||||
r1 = r1 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0x77b1520du : 0x2bb965afu); // 53 add
|
||||
r7 = rotl_imm(r7, 14u); // 54 rotl
|
||||
r3 = r3 + r7 + ((((sel >> 1u) & 1u) != 0u) ? 0xa54c55a0u : 0x7b0fe07au); // 55 add
|
||||
r6 = r6 ^ ds[r7 & mask]; // 56 load
|
||||
r1 = rotr_var(r1, r5); // 57 rotr
|
||||
r5 = r5 ^ ds[r4 & mask]; // 58 load
|
||||
r6 = r6 ^ ds[r2 & mask]; // 59 load
|
||||
r3 = r5 * r0 + r3; // 60 mad
|
||||
r5 = r5 + r7 + ((((sel >> 31u) & 1u) != 0u) ? 0xad7493e7u : 0xaf9dd72du); // 61 add
|
||||
r4 = r4 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); // 62 add
|
||||
r5 = rotl_imm(r5, 19u); // 63 rotl
|
||||
}
|
||||
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
|
|
|
|||
270
proto-cuda/packs/igneum-genesis/kernel_bound.cl
Normal file
270
proto-cuda/packs/igneum-genesis/kernel_bound.cl
Normal file
|
|
@ -0,0 +1,270 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
|
||||
// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal).
|
||||
// Built from source at runtime by proto-opencl/host.c, which passes these defines:
|
||||
// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit)
|
||||
// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default)
|
||||
// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32
|
||||
// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition
|
||||
// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units;
|
||||
// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32.
|
||||
#ifndef IGNEUM_GROUP
|
||||
#define IGNEUM_GROUP 32
|
||||
#endif
|
||||
#ifndef IGNEUM_EXCHANGE
|
||||
#define IGNEUM_EXCHANGE 0
|
||||
#endif
|
||||
#ifdef __OPENCL_VERSION__
|
||||
#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1)))
|
||||
#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n]
|
||||
#if IGNEUM_EXCHANGE == 1
|
||||
#ifdef cl_khr_subgroups
|
||||
#pragma OPENCL EXTENSION cl_khr_subgroups : enable
|
||||
#endif
|
||||
#ifdef cl_khr_subgroup_shuffle
|
||||
#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable
|
||||
#endif
|
||||
#elif IGNEUM_EXCHANGE == 2
|
||||
#pragma OPENCL EXTENSION cl_intel_subgroups : enable
|
||||
#endif
|
||||
#else
|
||||
// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros.
|
||||
#include "emu_opencl.h"
|
||||
#endif
|
||||
|
||||
#if IGNEUM_EXCHANGE == 1
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m))
|
||||
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
|
||||
#elif IGNEUM_EXCHANGE == 2
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m))
|
||||
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
|
||||
#else
|
||||
// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per
|
||||
// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane
|
||||
// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's
|
||||
// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier.
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; }
|
||||
#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; }
|
||||
#endif
|
||||
|
||||
static inline uint splitmix32(uint x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32.
|
||||
static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); }
|
||||
// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x.
|
||||
static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); }
|
||||
static inline uint ds_elem(uint i, uint d0, uint d1) {
|
||||
uint x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// dataset[i] = ds_elem(i, d0, d1) for i < n. Same closed form as the Metal igneum_fill kernel.
|
||||
__kernel void igneum_fill(__global uint* ds, uint n, uint d0, uint d1) {
|
||||
uint i = (uint)get_global_id(0);
|
||||
if (i < n) ds[i] = ds_elem(i, d0, d1);
|
||||
}
|
||||
|
||||
// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the
|
||||
// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and
|
||||
// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all).
|
||||
IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) {
|
||||
uint gid = (uint)get_global_id(0);
|
||||
uint lid = (uint)get_local_id(0);
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
#if IGNEUM_EXCHANGE == 0
|
||||
IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
|
||||
uint xk = 0u;
|
||||
#else
|
||||
(void)lid;
|
||||
#endif
|
||||
{ uint x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
|
||||
{ uint x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
|
||||
{ uint x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
|
||||
{ uint x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
|
||||
{ uint x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
|
||||
{ uint x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
|
||||
{ uint x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
|
||||
{ uint x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r2 = r3 * r4 + r2; // 0 mad
|
||||
r2 = r1 * r1 + r2; // 1 mad
|
||||
r2 = r3 * r2 + r2; // 2 mad
|
||||
r3 = r3 ^ r5; // 3 xor
|
||||
r7 = r7 ^ ds[r2 & mask]; // 4 load
|
||||
r5 = r5 ^ ds[r7 & mask]; // 5 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r1 = r1 ^ t_; } // 6 shfl
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 8u); r7 = r7 ^ t_; } // 7 shfl
|
||||
r1 = mul_hi(r1, r5); // 8 mulhi
|
||||
r6 = rotr_var(r6, r3); // 9 rotr
|
||||
r3 = r3 | r4; // 10 or
|
||||
r4 = r4 ^ ds[r3 & mask]; // 11 load
|
||||
r0 = mul_hi(r0, r4); // 12 mulhi
|
||||
r5 = r5 + r1 + ((((sel >> 30u) & 1u) != 0u) ? 0xd3177981u : 0xc7934706u); // 13 add
|
||||
r0 = r0 ^ ds[r4 & mask]; // 14 load
|
||||
r2 = r2 - r4; // 15 sub
|
||||
r2 = r2 ^ ds[r0 & mask]; // 16 load
|
||||
r7 = r7 ^ ds[r2 & mask]; // 17 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 18 shfl
|
||||
r5 = r5 * r0; // 19 mul
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 20 shfl
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 16u); r2 = r2 ^ t_; } // 21 shfl
|
||||
r6 = mul_hi(r6, r2); // 22 mulhi
|
||||
r6 = r6 ^ ds[r1 & mask]; // 23 load
|
||||
r5 = r5 * r0; // 24 mul
|
||||
r5 = rotl_imm(r5, 19u); // 25 rotl
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 2u); r7 = r7 ^ t_; } // 26 shfl
|
||||
r0 = r0 ^ r5; // 27 xor
|
||||
r0 = r0 ^ r4; // 28 xor
|
||||
r3 = r3 - r0; // 29 sub
|
||||
r5 = r5 * r1; // 30 mul
|
||||
r7 = r7 ^ ds[r2 & mask]; // 31 load
|
||||
r1 = r1 ^ ds[r0 & mask]; // 32 load
|
||||
r5 = r5 ^ r6; // 33 xor
|
||||
r5 = r5 ^ ds[r1 & mask]; // 34 load
|
||||
r0 = mul_hi(r0, r5); // 35 mulhi
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r2, 4u); r5 = r5 ^ t_; } // 36 shfl
|
||||
r7 = r7 ^ ds[r0 & mask]; // 37 load
|
||||
r3 = r3 + r1 + ((((sel >> 27u) & 1u) != 0u) ? 0x230c005cu : 0x75ba2fadu); // 38 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 4u); r1 = r1 ^ t_; } // 39 shfl
|
||||
r2 = r2 ^ r5; // 40 xor
|
||||
r3 = r6 * r3 + r3; // 41 mad
|
||||
r6 = r6 - r7; // 42 sub
|
||||
r7 = r7 ^ r0; // 43 xor
|
||||
r1 = r1 ^ ds[r7 & mask]; // 44 load
|
||||
r2 = r2 * r3; // 45 mul
|
||||
r1 = mul_hi(r1, r5); // 46 mulhi
|
||||
r4 = r4 - r3; // 47 sub
|
||||
r2 = rotr_var(r2, r6); // 48 rotr
|
||||
r3 = r3 ^ ds[r5 & mask]; // 49 load
|
||||
r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 50 add
|
||||
r0 = r0 * r2; // 51 mul
|
||||
r0 = r0 + r2 + ((((sel >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); // 52 add
|
||||
r1 = r1 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0x77b1520du : 0x2bb965afu); // 53 add
|
||||
r7 = rotl_imm(r7, 14u); // 54 rotl
|
||||
r3 = r3 + r7 + ((((sel >> 1u) & 1u) != 0u) ? 0xa54c55a0u : 0x7b0fe07au); // 55 add
|
||||
r6 = r6 ^ ds[r7 & mask]; // 56 load
|
||||
r1 = rotr_var(r1, r5); // 57 rotr
|
||||
r5 = r5 ^ ds[r4 & mask]; // 58 load
|
||||
r6 = r6 ^ ds[r2 & mask]; // 59 load
|
||||
r3 = r5 * r0 + r3; // 60 mad
|
||||
r5 = r5 + r7 + ((((sel >> 31u) & 1u) != 0u) ? 0xad7493e7u : 0xaf9dd72du); // 61 add
|
||||
r4 = r4 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); // 62 add
|
||||
r5 = rotl_imm(r5, 19u); // 63 rotl
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
|
||||
#if IGNEUM_EXCHANGE != 0
|
||||
// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the
|
||||
// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a
|
||||
// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md.
|
||||
IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) {
|
||||
if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); }
|
||||
}
|
||||
#endif
|
||||
|
||||
// Header-bound variant (bind.rs): the init words come from initw, not SEEDW. Same body as igneum_hash.
|
||||
IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw) {
|
||||
uint gid = (uint)get_global_id(0);
|
||||
uint lid = (uint)get_local_id(0);
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7];
|
||||
#if IGNEUM_EXCHANGE == 0
|
||||
IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
|
||||
uint xk = 0u;
|
||||
#else
|
||||
(void)lid;
|
||||
#endif
|
||||
{ uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; }
|
||||
{ uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; }
|
||||
{ uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; }
|
||||
{ uint x = nonce ^ iw3; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw4; }
|
||||
{ uint x = nonce ^ iw4; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw5; }
|
||||
{ uint x = nonce ^ iw5; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw6; }
|
||||
{ uint x = nonce ^ iw6; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw7; }
|
||||
{ uint x = nonce ^ iw7; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw0; }
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r2 = r3 * r4 + r2; // 0 mad
|
||||
r2 = r1 * r1 + r2; // 1 mad
|
||||
r2 = r3 * r2 + r2; // 2 mad
|
||||
r3 = r3 ^ r5; // 3 xor
|
||||
r7 = r7 ^ ds[r2 & mask]; // 4 load
|
||||
r5 = r5 ^ ds[r7 & mask]; // 5 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r1 = r1 ^ t_; } // 6 shfl
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 8u); r7 = r7 ^ t_; } // 7 shfl
|
||||
r1 = mul_hi(r1, r5); // 8 mulhi
|
||||
r6 = rotr_var(r6, r3); // 9 rotr
|
||||
r3 = r3 | r4; // 10 or
|
||||
r4 = r4 ^ ds[r3 & mask]; // 11 load
|
||||
r0 = mul_hi(r0, r4); // 12 mulhi
|
||||
r5 = r5 + r1 + ((((sel >> 30u) & 1u) != 0u) ? 0xd3177981u : 0xc7934706u); // 13 add
|
||||
r0 = r0 ^ ds[r4 & mask]; // 14 load
|
||||
r2 = r2 - r4; // 15 sub
|
||||
r2 = r2 ^ ds[r0 & mask]; // 16 load
|
||||
r7 = r7 ^ ds[r2 & mask]; // 17 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 4u); r7 = r7 ^ t_; } // 18 shfl
|
||||
r5 = r5 * r0; // 19 mul
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 20 shfl
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 16u); r2 = r2 ^ t_; } // 21 shfl
|
||||
r6 = mul_hi(r6, r2); // 22 mulhi
|
||||
r6 = r6 ^ ds[r1 & mask]; // 23 load
|
||||
r5 = r5 * r0; // 24 mul
|
||||
r5 = rotl_imm(r5, 19u); // 25 rotl
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 2u); r7 = r7 ^ t_; } // 26 shfl
|
||||
r0 = r0 ^ r5; // 27 xor
|
||||
r0 = r0 ^ r4; // 28 xor
|
||||
r3 = r3 - r0; // 29 sub
|
||||
r5 = r5 * r1; // 30 mul
|
||||
r7 = r7 ^ ds[r2 & mask]; // 31 load
|
||||
r1 = r1 ^ ds[r0 & mask]; // 32 load
|
||||
r5 = r5 ^ r6; // 33 xor
|
||||
r5 = r5 ^ ds[r1 & mask]; // 34 load
|
||||
r0 = mul_hi(r0, r5); // 35 mulhi
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r2, 4u); r5 = r5 ^ t_; } // 36 shfl
|
||||
r7 = r7 ^ ds[r0 & mask]; // 37 load
|
||||
r3 = r3 + r1 + ((((sel >> 27u) & 1u) != 0u) ? 0x230c005cu : 0x75ba2fadu); // 38 add
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 4u); r1 = r1 ^ t_; } // 39 shfl
|
||||
r2 = r2 ^ r5; // 40 xor
|
||||
r3 = r6 * r3 + r3; // 41 mad
|
||||
r6 = r6 - r7; // 42 sub
|
||||
r7 = r7 ^ r0; // 43 xor
|
||||
r1 = r1 ^ ds[r7 & mask]; // 44 load
|
||||
r2 = r2 * r3; // 45 mul
|
||||
r1 = mul_hi(r1, r5); // 46 mulhi
|
||||
r4 = r4 - r3; // 47 sub
|
||||
r2 = rotr_var(r2, r6); // 48 rotr
|
||||
r3 = r3 ^ ds[r5 & mask]; // 49 load
|
||||
r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 50 add
|
||||
r0 = r0 * r2; // 51 mul
|
||||
r0 = r0 + r2 + ((((sel >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); // 52 add
|
||||
r1 = r1 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0x77b1520du : 0x2bb965afu); // 53 add
|
||||
r7 = rotl_imm(r7, 14u); // 54 rotl
|
||||
r3 = r3 + r7 + ((((sel >> 1u) & 1u) != 0u) ? 0xa54c55a0u : 0x7b0fe07au); // 55 add
|
||||
r6 = r6 ^ ds[r7 & mask]; // 56 load
|
||||
r1 = rotr_var(r1, r5); // 57 rotr
|
||||
r5 = r5 ^ ds[r4 & mask]; // 58 load
|
||||
r6 = r6 ^ ds[r2 & mask]; // 59 load
|
||||
r3 = r5 * r0 + r3; // 60 mad
|
||||
r5 = r5 + r7 + ((((sel >> 31u) & 1u) != 0u) ? 0xad7493e7u : 0xaf9dd72du); // 61 add
|
||||
r4 = r4 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); // 62 add
|
||||
r5 = rotl_imm(r5, 19u); // 63 rotl
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
123
proto-cuda/packs/igneum-genesis/kernel_bound.cu
Normal file
123
proto-cuda/packs/igneum-genesis/kernel_bound.cu
Normal file
|
|
@ -0,0 +1,123 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
|
||||
// Header-bound twin of igneum_hash in kernel.cu: the init words come from a kernel argument, not SEEDW.
|
||||
// Host declarations (also in program_bound.h if present):
|
||||
// struct IgneumInitWords { uint32_t w[8]; };
|
||||
// cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
// IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps);
|
||||
// cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
|
||||
#include <cuda_runtime.h>
|
||||
#include <cstdint>
|
||||
#include "program.h"
|
||||
|
||||
struct IgneumInitWords { uint32_t w[8]; };
|
||||
|
||||
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
|
||||
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
|
||||
__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw) {
|
||||
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
uint32_t nonce = baseNonce + gid;
|
||||
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint32_t x = nonce ^ iw.w[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw.w[1]; }
|
||||
{ uint32_t x = nonce ^ iw.w[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw.w[2]; }
|
||||
{ uint32_t x = nonce ^ iw.w[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw.w[3]; }
|
||||
{ uint32_t x = nonce ^ iw.w[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw.w[4]; }
|
||||
{ uint32_t x = nonce ^ iw.w[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw.w[5]; }
|
||||
{ uint32_t x = nonce ^ iw.w[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw.w[6]; }
|
||||
{ uint32_t x = nonce ^ iw.w[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw.w[7]; }
|
||||
{ uint32_t x = nonce ^ iw.w[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw.w[0]; }
|
||||
|
||||
for (uint32_t it = 0u; it < 8u; ++it) {
|
||||
uint32_t sel = r0;
|
||||
r2 = r3 * r4 + r2; // 0 mad
|
||||
r2 = r1 * r1 + r2; // 1 mad
|
||||
r2 = r3 * r2 + r2; // 2 mad
|
||||
r3 = r3 ^ r5; // 3 xor
|
||||
r7 = r7 ^ ds[r2 & mask]; // 4 load
|
||||
r5 = r5 ^ ds[r7 & mask]; // 5 load
|
||||
r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 6 shfl
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 8); // 7 shfl
|
||||
r1 = __umulhi(r1, r5); // 8 mulhi
|
||||
r6 = rotr_var(r6, r3); // 9 rotr
|
||||
r3 = r3 | r4; // 10 or
|
||||
r4 = r4 ^ ds[r3 & mask]; // 11 load
|
||||
r0 = __umulhi(r0, r4); // 12 mulhi
|
||||
r5 = r5 + r1 + ((((sel >> 30u) & 1u) != 0u) ? 0xd3177981u : 0xc7934706u); // 13 add
|
||||
r0 = r0 ^ ds[r4 & mask]; // 14 load
|
||||
r2 = r2 - r4; // 15 sub
|
||||
r2 = r2 ^ ds[r0 & mask]; // 16 load
|
||||
r7 = r7 ^ ds[r2 & mask]; // 17 load
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 18 shfl
|
||||
r5 = r5 * r0; // 19 mul
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 20 shfl
|
||||
r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r4, 16); // 21 shfl
|
||||
r6 = __umulhi(r6, r2); // 22 mulhi
|
||||
r6 = r6 ^ ds[r1 & mask]; // 23 load
|
||||
r5 = r5 * r0; // 24 mul
|
||||
r5 = rotl_imm(r5, 19u); // 25 rotl
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r6, 2); // 26 shfl
|
||||
r0 = r0 ^ r5; // 27 xor
|
||||
r0 = r0 ^ r4; // 28 xor
|
||||
r3 = r3 - r0; // 29 sub
|
||||
r5 = r5 * r1; // 30 mul
|
||||
r7 = r7 ^ ds[r2 & mask]; // 31 load
|
||||
r1 = r1 ^ ds[r0 & mask]; // 32 load
|
||||
r5 = r5 ^ r6; // 33 xor
|
||||
r5 = r5 ^ ds[r1 & mask]; // 34 load
|
||||
r0 = __umulhi(r0, r5); // 35 mulhi
|
||||
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r2, 4); // 36 shfl
|
||||
r7 = r7 ^ ds[r0 & mask]; // 37 load
|
||||
r3 = r3 + r1 + ((((sel >> 27u) & 1u) != 0u) ? 0x230c005cu : 0x75ba2fadu); // 38 add
|
||||
r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r5, 4); // 39 shfl
|
||||
r2 = r2 ^ r5; // 40 xor
|
||||
r3 = r6 * r3 + r3; // 41 mad
|
||||
r6 = r6 - r7; // 42 sub
|
||||
r7 = r7 ^ r0; // 43 xor
|
||||
r1 = r1 ^ ds[r7 & mask]; // 44 load
|
||||
r2 = r2 * r3; // 45 mul
|
||||
r1 = __umulhi(r1, r5); // 46 mulhi
|
||||
r4 = r4 - r3; // 47 sub
|
||||
r2 = rotr_var(r2, r6); // 48 rotr
|
||||
r3 = r3 ^ ds[r5 & mask]; // 49 load
|
||||
r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 50 add
|
||||
r0 = r0 * r2; // 51 mul
|
||||
r0 = r0 + r2 + ((((sel >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); // 52 add
|
||||
r1 = r1 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0x77b1520du : 0x2bb965afu); // 53 add
|
||||
r7 = rotl_imm(r7, 14u); // 54 rotl
|
||||
r3 = r3 + r7 + ((((sel >> 1u) & 1u) != 0u) ? 0xa54c55a0u : 0x7b0fe07au); // 55 add
|
||||
r6 = r6 ^ ds[r7 & mask]; // 56 load
|
||||
r1 = rotr_var(r1, r5); // 57 rotr
|
||||
r5 = r5 ^ ds[r4 & mask]; // 58 load
|
||||
r6 = r6 ^ ds[r2 & mask]; // 59 load
|
||||
r3 = r5 * r0 + r3; // 60 mad
|
||||
r5 = r5 + r7 + ((((sel >> 31u) & 1u) != 0u) ? 0xad7493e7u : 0xaf9dd72du); // 61 add
|
||||
r4 = r4 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); // 62 add
|
||||
r5 = rotl_imm(r5, 19u); // 63 rotl
|
||||
}
|
||||
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps) {
|
||||
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 32u * blockWarps;
|
||||
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
|
||||
igneum_hash_bound<<<nonces / block, block>>>(ds, out, baseNonce, mask, iw);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
|
||||
cudaFuncAttributes attr;
|
||||
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash_bound);
|
||||
if (e != cudaSuccess) return e;
|
||||
*numRegs = attr.numRegs;
|
||||
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash_bound, (int)(32u * blockWarps), 0);
|
||||
}
|
||||
|
|
@ -1,4 +1,4 @@
|
|||
// Generated by proto-metal/igneum-bench --export-pack for seed "igneum-genesis". Do not edit by hand.
|
||||
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
|
||||
// Program metadata for host.cu plus the launch wrappers defined in kernel.cu.
|
||||
// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros.
|
||||
#pragma once
|
||||
|
|
@ -12,7 +12,12 @@
|
|||
#endif
|
||||
|
||||
#define IGNEUM_SEED_STRING "igneum-genesis"
|
||||
#define IGNEUM_SEED_BYTES_HEX "69676e65756d2d67656e65736973"
|
||||
#define IGNEUM_GENERATOR 2
|
||||
#define IGNEUM_PROGRAM_ATTEMPT 0
|
||||
#define IGNEUM_PROGRAM_ID 0xbcc1248b10cc90f2ull
|
||||
#define IGNEUM_DAY_STRING "2026-10-03"
|
||||
#define IGNEUM_DAY_BYTES_HEX "6461792f323032362d31302d3033"
|
||||
#define IGNEUM_DAY0 0x3067619fu
|
||||
#define IGNEUM_DAY1 0x3c269176u
|
||||
#define IGNEUM_DATASET_LOG2 28
|
||||
|
|
@ -20,9 +25,9 @@
|
|||
#define IGNEUM_LANES 32
|
||||
#define IGNEUM_ITERATIONS 8
|
||||
#define IGNEUM_INSTR_COUNT 64
|
||||
#define IGNEUM_LOADS_PER_HASH 104
|
||||
#define IGNEUM_LOADS_PER_HASH 128
|
||||
#define IGNEUM_WIDE_LOADS_PER_HASH 0
|
||||
#define IGNEUM_OP_MIX "load=13 xor=13 sub=7 shfl=6 add=5 mulhi=5 mad=4 rotr=4 mul=3 rotl=3 or=1"
|
||||
#define IGNEUM_OP_MIX "load=16 add=8 shfl=8 xor=6 mad=5 mul=5 mulhi=5 sub=4 rotl=3 rotr=3 or=1"
|
||||
// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h)
|
||||
#define IGNEUM_DATASET_MODE 0
|
||||
|
||||
|
|
|
|||
|
|
@ -1,15 +1,21 @@
|
|||
{
|
||||
"format": "igneum-program-pack-2",
|
||||
"format": "igneum-program-pack-3",
|
||||
"generator": 2,
|
||||
"attempt": 0,
|
||||
"program_id": "0xbcc1248b10cc90f2",
|
||||
"program_id_derivation": "FNV-1a 64 over 'igneum-program/' || generator_le32 || seed_words as little-endian bytes || attempt_le32",
|
||||
"dataset_mode": "closed-form",
|
||||
"seed": "igneum-genesis",
|
||||
"seed_bytes": "69676e65756d2d67656e65736973",
|
||||
"seed_words": ["0x67a9a7be", "0x1a155b25", "0xfddfb732", "0x4b5af2e8", "0xc55caf33", "0xa27c13b7", "0x06628a48", "0x03852469"],
|
||||
"seed_derivation": "FNV-1a 64 over UTF-8 of seed, basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32",
|
||||
"seed_derivation": "seed_words = FNV-1a 64 over seed_bytes (attempt 0) or seed_bytes || attempt_le32 (attempt k >= 1), basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32",
|
||||
"generator_rule": "version 2: exactly 16 load slots drawn first from instructions 1..63 (partial Fisher-Yates), the other 48 ops from the ten non-load weights (sum 75); a load's source is drawn from the registers other than dst written by an earlier instruction and not read by a load since; the candidate must pass the acceptance rule of spec 01 section 1.4.6 (static: no cyclically stale load source, every register has an injecting write; dynamic: 64 units on the seed-keyed closed-form dataset with no constant register bit, no lane-constant load site, under 164 saturated final values, every output bit within 136 of 1024, distinct addresses above 245760), else the next attempt of the seed is tried",
|
||||
"lanes": 32,
|
||||
"registers": 8,
|
||||
"iterations": 8,
|
||||
"instruction_count": 64,
|
||||
"loads_per_hash": 104,
|
||||
"op_mix": {"load": 13, "xor": 13, "sub": 7, "shfl": 6, "add": 5, "mulhi": 5, "mad": 4, "rotr": 4, "mul": 3, "rotl": 3, "or": 1},
|
||||
"loads_per_hash": 128,
|
||||
"op_mix": {"load": 16, "add": 8, "shfl": 8, "xor": 6, "mad": 5, "mul": 5, "mulhi": 5, "sub": 4, "rotl": 3, "rotr": 3, "or": 1},
|
||||
"register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]",
|
||||
"splitmix32": "x ^= x>>16; x *= 0x7feb352d; x ^= x>>15; x *= 0x846ca68b; x ^= x>>16",
|
||||
"iteration": "sel = r0 sampled once at the top of each iteration, then all instructions in order",
|
||||
|
|
@ -33,76 +39,77 @@
|
|||
"bytes": 1073741824,
|
||||
"mask": "0x0fffffff",
|
||||
"day": "2026-10-03",
|
||||
"day_words_from": "day/2026-10-03",
|
||||
"day_bytes": "6461792f323032362d31302d3033",
|
||||
"day_words_from": "seed_words_from_bytes(day_bytes)",
|
||||
"d0": "0x3067619f",
|
||||
"d1": "0x3c269176",
|
||||
"mode": "closed-form",
|
||||
"formula": "x = i ^ d0; x *= 0x9E3779B1; x ^= x>>15; x += d1; x *= 0x85EBCA77; x ^= x>>13; x *= 0xC2B2AE3D; x ^= x>>16 (all mod 2^32)"
|
||||
},
|
||||
"instructions": [
|
||||
{"i": 0, "op": "rotl", "dst": 4, "src": 2, "src2": 7, "imm": "0x20699878", "imm2": "0x6f1a6170", "rot": 25, "bit": 7, "mask": 16},
|
||||
{"i": 1, "op": "sub", "dst": 0, "src": 5, "src2": 4, "imm": "0xf81a0b9d", "imm2": "0xf0505e88", "rot": 1, "bit": 4, "mask": 1},
|
||||
{"i": 2, "op": "load", "dst": 4, "src": 3, "src2": 6, "imm": "0xb1978a0b", "imm2": "0x2ca4e162", "rot": 10, "bit": 21, "mask": 1},
|
||||
{"i": 3, "op": "rotl", "dst": 1, "src": 6, "src2": 7, "imm": "0xc3bd2355", "imm2": "0xa8c5f27e", "rot": 1, "bit": 22, "mask": 1},
|
||||
{"i": 4, "op": "add", "dst": 2, "src": 3, "src2": 0, "imm": "0x61f0b51c", "imm2": "0x2735a174", "rot": 4, "bit": 26, "mask": 2},
|
||||
{"i": 5, "op": "load", "dst": 5, "src": 3, "src2": 3, "imm": "0x4d183796", "imm2": "0x679648a8", "rot": 4, "bit": 30, "mask": 4},
|
||||
{"i": 6, "op": "sub", "dst": 5, "src": 7, "src2": 7, "imm": "0x265677dc", "imm2": "0x9043323e", "rot": 30, "bit": 7, "mask": 4},
|
||||
{"i": 7, "op": "add", "dst": 3, "src": 4, "src2": 6, "imm": "0x52f2dbf4", "imm2": "0x5a069596", "rot": 5, "bit": 26, "mask": 2},
|
||||
{"i": 8, "op": "xor", "dst": 0, "src": 4, "src2": 1, "imm": "0x3303ec4b", "imm2": "0xfeca75be", "rot": 21, "bit": 7, "mask": 1},
|
||||
{"i": 9, "op": "xor", "dst": 4, "src": 0, "src2": 0, "imm": "0x5dc5959c", "imm2": "0x023f44a8", "rot": 11, "bit": 7, "mask": 4},
|
||||
{"i": 10, "op": "load", "dst": 2, "src": 0, "src2": 3, "imm": "0xde716173", "imm2": "0xc21e924d", "rot": 30, "bit": 15, "mask": 1},
|
||||
{"i": 11, "op": "shfl", "dst": 4, "src": 6, "src2": 1, "imm": "0x3cda6d48", "imm2": "0x0970145b", "rot": 12, "bit": 22, "mask": 16},
|
||||
{"i": 12, "op": "sub", "dst": 1, "src": 5, "src2": 4, "imm": "0x48866c15", "imm2": "0x4eaee50c", "rot": 30, "bit": 20, "mask": 16},
|
||||
{"i": 13, "op": "xor", "dst": 2, "src": 1, "src2": 5, "imm": "0x196d165c", "imm2": "0x0f730511", "rot": 11, "bit": 3, "mask": 4},
|
||||
{"i": 14, "op": "load", "dst": 4, "src": 5, "src2": 1, "imm": "0xc5c3b55d", "imm2": "0xec061424", "rot": 26, "bit": 27, "mask": 8},
|
||||
{"i": 15, "op": "load", "dst": 2, "src": 4, "src2": 1, "imm": "0x17c9c95b", "imm2": "0x306542fe", "rot": 27, "bit": 17, "mask": 16},
|
||||
{"i": 16, "op": "load", "dst": 3, "src": 0, "src2": 2, "imm": "0x590f9e11", "imm2": "0xa4c13036", "rot": 3, "bit": 28, "mask": 16},
|
||||
{"i": 17, "op": "xor", "dst": 4, "src": 6, "src2": 3, "imm": "0x9c4eeee9", "imm2": "0xf069b834", "rot": 11, "bit": 23, "mask": 8},
|
||||
{"i": 18, "op": "mad", "dst": 2, "src": 4, "src2": 6, "imm": "0xa242a28b", "imm2": "0xb8974bdf", "rot": 13, "bit": 30, "mask": 1},
|
||||
{"i": 19, "op": "rotr", "dst": 6, "src": 1, "src2": 0, "imm": "0xcd69ed50", "imm2": "0xadf52615", "rot": 30, "bit": 2, "mask": 2},
|
||||
{"i": 20, "op": "xor", "dst": 3, "src": 4, "src2": 0, "imm": "0xe3059a24", "imm2": "0x16b8dd86", "rot": 25, "bit": 31, "mask": 16},
|
||||
{"i": 21, "op": "mad", "dst": 1, "src": 3, "src2": 5, "imm": "0xc3ae8ae1", "imm2": "0x8e126e8a", "rot": 2, "bit": 28, "mask": 1},
|
||||
{"i": 22, "op": "mulhi", "dst": 7, "src": 4, "src2": 2, "imm": "0x6909af7a", "imm2": "0xa388b4b9", "rot": 14, "bit": 10, "mask": 2},
|
||||
{"i": 23, "op": "mulhi", "dst": 5, "src": 2, "src2": 2, "imm": "0xdf099cfb", "imm2": "0xd0133a01", "rot": 31, "bit": 3, "mask": 4},
|
||||
{"i": 24, "op": "load", "dst": 0, "src": 6, "src2": 3, "imm": "0x0a3056de", "imm2": "0x7f0c25c3", "rot": 27, "bit": 13, "mask": 8},
|
||||
{"i": 25, "op": "mul", "dst": 5, "src": 6, "src2": 7, "imm": "0x089f5404", "imm2": "0xbd066e1d", "rot": 7, "bit": 10, "mask": 4},
|
||||
{"i": 26, "op": "load", "dst": 7, "src": 1, "src2": 2, "imm": "0x3e26afea", "imm2": "0xba573970", "rot": 11, "bit": 15, "mask": 8},
|
||||
{"i": 27, "op": "rotr", "dst": 3, "src": 1, "src2": 5, "imm": "0xe2466d63", "imm2": "0x30db112b", "rot": 31, "bit": 15, "mask": 2},
|
||||
{"i": 28, "op": "shfl", "dst": 5, "src": 1, "src2": 3, "imm": "0xdc20297d", "imm2": "0xcb813557", "rot": 3, "bit": 21, "mask": 8},
|
||||
{"i": 29, "op": "xor", "dst": 7, "src": 5, "src2": 2, "imm": "0x6ede5c13", "imm2": "0x1f14267f", "rot": 22, "bit": 31, "mask": 1},
|
||||
{"i": 30, "op": "rotl", "dst": 7, "src": 5, "src2": 6, "imm": "0x24cdb54b", "imm2": "0xa44e008d", "rot": 23, "bit": 4, "mask": 16},
|
||||
{"i": 31, "op": "sub", "dst": 2, "src": 0, "src2": 2, "imm": "0xf94d9b65", "imm2": "0x3457264c", "rot": 11, "bit": 5, "mask": 2},
|
||||
{"i": 32, "op": "xor", "dst": 7, "src": 2, "src2": 3, "imm": "0x8d2046e5", "imm2": "0xf68893bb", "rot": 21, "bit": 2, "mask": 1},
|
||||
{"i": 33, "op": "xor", "dst": 2, "src": 6, "src2": 5, "imm": "0xe0ebc4ce", "imm2": "0x02773069", "rot": 21, "bit": 16, "mask": 1},
|
||||
{"i": 34, "op": "load", "dst": 6, "src": 1, "src2": 2, "imm": "0x3b2d2124", "imm2": "0x187a9128", "rot": 1, "bit": 9, "mask": 16},
|
||||
{"i": 35, "op": "xor", "dst": 1, "src": 4, "src2": 6, "imm": "0xe3f24158", "imm2": "0x5c64a589", "rot": 13, "bit": 0, "mask": 8},
|
||||
{"i": 36, "op": "shfl", "dst": 3, "src": 0, "src2": 3, "imm": "0x63578bc1", "imm2": "0xbf64a89f", "rot": 16, "bit": 26, "mask": 1},
|
||||
{"i": 37, "op": "mul", "dst": 2, "src": 6, "src2": 4, "imm": "0xfa624b69", "imm2": "0x0389cf85", "rot": 28, "bit": 4, "mask": 1},
|
||||
{"i": 38, "op": "add", "dst": 5, "src": 3, "src2": 5, "imm": "0x71f30417", "imm2": "0xa29f4338", "rot": 2, "bit": 12, "mask": 2},
|
||||
{"i": 39, "op": "xor", "dst": 7, "src": 6, "src2": 3, "imm": "0xfa3b845c", "imm2": "0x9717695d", "rot": 18, "bit": 24, "mask": 1},
|
||||
{"i": 40, "op": "xor", "dst": 7, "src": 3, "src2": 4, "imm": "0x8408dc67", "imm2": "0x8d389f9b", "rot": 21, "bit": 16, "mask": 16},
|
||||
{"i": 41, "op": "rotr", "dst": 3, "src": 4, "src2": 1, "imm": "0x4bbcad92", "imm2": "0xbc6375cd", "rot": 18, "bit": 4, "mask": 2},
|
||||
{"i": 42, "op": "xor", "dst": 5, "src": 3, "src2": 3, "imm": "0xbd31a2ea", "imm2": "0x742d28ea", "rot": 4, "bit": 12, "mask": 4},
|
||||
{"i": 43, "op": "rotr", "dst": 3, "src": 6, "src2": 5, "imm": "0x35c52d04", "imm2": "0x0f8e8621", "rot": 3, "bit": 22, "mask": 8},
|
||||
{"i": 44, "op": "mad", "dst": 1, "src": 3, "src2": 5, "imm": "0x3958f280", "imm2": "0x8713c7e1", "rot": 5, "bit": 23, "mask": 16},
|
||||
{"i": 45, "op": "add", "dst": 7, "src": 3, "src2": 7, "imm": "0xc1ae8d3b", "imm2": "0xa907b90b", "rot": 13, "bit": 29, "mask": 1},
|
||||
{"i": 46, "op": "or", "dst": 7, "src": 2, "src2": 2, "imm": "0xc4f18bec", "imm2": "0x8a3e2464", "rot": 30, "bit": 16, "mask": 2},
|
||||
{"i": 47, "op": "shfl", "dst": 2, "src": 3, "src2": 6, "imm": "0xf73ba9e3", "imm2": "0x028aa63c", "rot": 3, "bit": 20, "mask": 16},
|
||||
{"i": 48, "op": "load", "dst": 1, "src": 7, "src2": 5, "imm": "0x53be87d4", "imm2": "0x690a5729", "rot": 1, "bit": 17, "mask": 8},
|
||||
{"i": 49, "op": "sub", "dst": 5, "src": 1, "src2": 7, "imm": "0xecdc31a6", "imm2": "0xed6bd9f2", "rot": 1, "bit": 13, "mask": 8},
|
||||
{"i": 50, "op": "load", "dst": 3, "src": 1, "src2": 3, "imm": "0x55f91dbc", "imm2": "0xf026553d", "rot": 30, "bit": 27, "mask": 1},
|
||||
{"i": 51, "op": "sub", "dst": 2, "src": 3, "src2": 5, "imm": "0x6689eef2", "imm2": "0x24f3ae21", "rot": 6, "bit": 22, "mask": 8},
|
||||
{"i": 52, "op": "shfl", "dst": 6, "src": 2, "src2": 4, "imm": "0x8fd37aad", "imm2": "0x89ed5d67", "rot": 8, "bit": 16, "mask": 2},
|
||||
{"i": 53, "op": "sub", "dst": 2, "src": 0, "src2": 0, "imm": "0x963bb7e6", "imm2": "0x4738b84f", "rot": 18, "bit": 16, "mask": 4},
|
||||
{"i": 54, "op": "load", "dst": 0, "src": 3, "src2": 0, "imm": "0x838b5065", "imm2": "0x36360066", "rot": 3, "bit": 31, "mask": 4},
|
||||
{"i": 55, "op": "shfl", "dst": 2, "src": 1, "src2": 5, "imm": "0xb17aad78", "imm2": "0x8458f7ac", "rot": 5, "bit": 7, "mask": 16},
|
||||
{"i": 56, "op": "add", "dst": 0, "src": 6, "src2": 0, "imm": "0x25955401", "imm2": "0xf4689674", "rot": 14, "bit": 27, "mask": 4},
|
||||
{"i": 57, "op": "mulhi", "dst": 2, "src": 0, "src2": 0, "imm": "0xa4e8c86a", "imm2": "0x14bde8e1", "rot": 8, "bit": 9, "mask": 16},
|
||||
{"i": 58, "op": "mulhi", "dst": 4, "src": 2, "src2": 7, "imm": "0xa36b1f4f", "imm2": "0x372e072f", "rot": 31, "bit": 31, "mask": 1},
|
||||
{"i": 59, "op": "mad", "dst": 2, "src": 6, "src2": 7, "imm": "0xfcd3b17f", "imm2": "0x6a8583df", "rot": 24, "bit": 9, "mask": 2},
|
||||
{"i": 60, "op": "xor", "dst": 3, "src": 1, "src2": 0, "imm": "0x3bbea2ae", "imm2": "0x914e2f0a", "rot": 2, "bit": 15, "mask": 16},
|
||||
{"i": 61, "op": "mul", "dst": 4, "src": 2, "src2": 6, "imm": "0x50aaa099", "imm2": "0xd1b988a1", "rot": 27, "bit": 31, "mask": 4},
|
||||
{"i": 62, "op": "load", "dst": 0, "src": 2, "src2": 3, "imm": "0x4ccdbdea", "imm2": "0xf3a60b41", "rot": 1, "bit": 20, "mask": 2},
|
||||
{"i": 63, "op": "mulhi", "dst": 7, "src": 0, "src2": 7, "imm": "0x3e400372", "imm2": "0x92c6201f", "rot": 1, "bit": 13, "mask": 8}
|
||||
{"i": 0, "op": "mad", "dst": 2, "src": 3, "src2": 4, "imm": "0xbf7b174d", "imm2": "0x337b762e", "rot": 17, "bit": 2, "mask": 2},
|
||||
{"i": 1, "op": "mad", "dst": 2, "src": 1, "src2": 1, "imm": "0xdd04a5da", "imm2": "0x42da7657", "rot": 15, "bit": 30, "mask": 16},
|
||||
{"i": 2, "op": "mad", "dst": 2, "src": 3, "src2": 2, "imm": "0x734003fa", "imm2": "0x5bb67700", "rot": 3, "bit": 20, "mask": 1},
|
||||
{"i": 3, "op": "xor", "dst": 3, "src": 5, "src2": 5, "imm": "0xc55a1b1c", "imm2": "0xa19720f3", "rot": 7, "bit": 8, "mask": 1},
|
||||
{"i": 4, "op": "load", "dst": 7, "src": 2, "src2": 5, "imm": "0xad572dd7", "imm2": "0x9ceb3ea7", "rot": 18, "bit": 30, "mask": 2},
|
||||
{"i": 5, "op": "load", "dst": 5, "src": 7, "src2": 3, "imm": "0x769a53be", "imm2": "0x80f9067e", "rot": 12, "bit": 22, "mask": 1},
|
||||
{"i": 6, "op": "shfl", "dst": 1, "src": 4, "src2": 0, "imm": "0xd3613d88", "imm2": "0x262fb219", "rot": 10, "bit": 30, "mask": 8},
|
||||
{"i": 7, "op": "shfl", "dst": 7, "src": 3, "src2": 4, "imm": "0xce38e42f", "imm2": "0xb868b818", "rot": 11, "bit": 8, "mask": 8},
|
||||
{"i": 8, "op": "mulhi", "dst": 1, "src": 5, "src2": 2, "imm": "0xa5eebca5", "imm2": "0x5703a72b", "rot": 13, "bit": 13, "mask": 16},
|
||||
{"i": 9, "op": "rotr", "dst": 6, "src": 3, "src2": 4, "imm": "0x17a5a9c7", "imm2": "0xdcfb93a1", "rot": 20, "bit": 27, "mask": 2},
|
||||
{"i": 10, "op": "or", "dst": 3, "src": 4, "src2": 1, "imm": "0xccb7d785", "imm2": "0xc335364c", "rot": 14, "bit": 12, "mask": 4},
|
||||
{"i": 11, "op": "load", "dst": 4, "src": 3, "src2": 2, "imm": "0x88cb9af3", "imm2": "0x4e7dc10d", "rot": 24, "bit": 17, "mask": 4},
|
||||
{"i": 12, "op": "mulhi", "dst": 0, "src": 4, "src2": 4, "imm": "0x45374321", "imm2": "0x3cd91989", "rot": 11, "bit": 4, "mask": 2},
|
||||
{"i": 13, "op": "add", "dst": 5, "src": 1, "src2": 2, "imm": "0xc7934706", "imm2": "0xd3177981", "rot": 16, "bit": 30, "mask": 2},
|
||||
{"i": 14, "op": "load", "dst": 0, "src": 4, "src2": 3, "imm": "0xfd7f56bb", "imm2": "0x65e14f52", "rot": 13, "bit": 22, "mask": 2},
|
||||
{"i": 15, "op": "sub", "dst": 2, "src": 4, "src2": 4, "imm": "0x35a80b49", "imm2": "0x060f2d13", "rot": 16, "bit": 20, "mask": 16},
|
||||
{"i": 16, "op": "load", "dst": 2, "src": 0, "src2": 2, "imm": "0xae0a32c2", "imm2": "0x4c2a4cfe", "rot": 8, "bit": 31, "mask": 16},
|
||||
{"i": 17, "op": "load", "dst": 7, "src": 2, "src2": 6, "imm": "0x82a84cc3", "imm2": "0x21a38d68", "rot": 15, "bit": 21, "mask": 2},
|
||||
{"i": 18, "op": "shfl", "dst": 7, "src": 3, "src2": 3, "imm": "0xa3818806", "imm2": "0x8f66b5c8", "rot": 14, "bit": 6, "mask": 4},
|
||||
{"i": 19, "op": "mul", "dst": 5, "src": 0, "src2": 1, "imm": "0xa00de107", "imm2": "0x77bfcaa5", "rot": 3, "bit": 10, "mask": 2},
|
||||
{"i": 20, "op": "shfl", "dst": 3, "src": 4, "src2": 7, "imm": "0x1d2b8cab", "imm2": "0x80b4f9a2", "rot": 14, "bit": 25, "mask": 2},
|
||||
{"i": 21, "op": "shfl", "dst": 2, "src": 4, "src2": 5, "imm": "0x3ac915d2", "imm2": "0x5fba7bc2", "rot": 16, "bit": 1, "mask": 16},
|
||||
{"i": 22, "op": "mulhi", "dst": 6, "src": 2, "src2": 0, "imm": "0xdc3ec8fd", "imm2": "0x599e2fa3", "rot": 22, "bit": 3, "mask": 2},
|
||||
{"i": 23, "op": "load", "dst": 6, "src": 1, "src2": 5, "imm": "0x2a6b16d5", "imm2": "0xd73e396f", "rot": 28, "bit": 29, "mask": 2},
|
||||
{"i": 24, "op": "mul", "dst": 5, "src": 0, "src2": 7, "imm": "0x376d0223", "imm2": "0xe1c2169a", "rot": 4, "bit": 16, "mask": 16},
|
||||
{"i": 25, "op": "rotl", "dst": 5, "src": 7, "src2": 3, "imm": "0x78ad8c60", "imm2": "0x6f5b77d5", "rot": 19, "bit": 11, "mask": 16},
|
||||
{"i": 26, "op": "shfl", "dst": 7, "src": 6, "src2": 5, "imm": "0x93915b9f", "imm2": "0x1e61fb6b", "rot": 28, "bit": 23, "mask": 2},
|
||||
{"i": 27, "op": "xor", "dst": 0, "src": 5, "src2": 7, "imm": "0x6378fe15", "imm2": "0x66c78f42", "rot": 12, "bit": 31, "mask": 8},
|
||||
{"i": 28, "op": "xor", "dst": 0, "src": 4, "src2": 7, "imm": "0x20a57fda", "imm2": "0x088c848e", "rot": 16, "bit": 13, "mask": 4},
|
||||
{"i": 29, "op": "sub", "dst": 3, "src": 0, "src2": 2, "imm": "0x49d95fd5", "imm2": "0x1a5f946a", "rot": 6, "bit": 12, "mask": 1},
|
||||
{"i": 30, "op": "mul", "dst": 5, "src": 1, "src2": 7, "imm": "0x0a816217", "imm2": "0x405c4f73", "rot": 13, "bit": 27, "mask": 4},
|
||||
{"i": 31, "op": "load", "dst": 7, "src": 2, "src2": 2, "imm": "0x09ed045e", "imm2": "0xd69c4715", "rot": 5, "bit": 9, "mask": 2},
|
||||
{"i": 32, "op": "load", "dst": 1, "src": 0, "src2": 6, "imm": "0xeb79ea49", "imm2": "0xcc587f5a", "rot": 6, "bit": 8, "mask": 16},
|
||||
{"i": 33, "op": "xor", "dst": 5, "src": 6, "src2": 1, "imm": "0x3027401e", "imm2": "0x5f20c27e", "rot": 18, "bit": 9, "mask": 2},
|
||||
{"i": 34, "op": "load", "dst": 5, "src": 1, "src2": 3, "imm": "0x0e1cab07", "imm2": "0x09356c5b", "rot": 19, "bit": 31, "mask": 1},
|
||||
{"i": 35, "op": "mulhi", "dst": 0, "src": 5, "src2": 2, "imm": "0x90e31357", "imm2": "0xabd32484", "rot": 26, "bit": 5, "mask": 8},
|
||||
{"i": 36, "op": "shfl", "dst": 5, "src": 2, "src2": 5, "imm": "0xee9a955f", "imm2": "0x31b3faed", "rot": 8, "bit": 24, "mask": 4},
|
||||
{"i": 37, "op": "load", "dst": 7, "src": 0, "src2": 7, "imm": "0x3ba2f832", "imm2": "0x1160dcd3", "rot": 4, "bit": 29, "mask": 1},
|
||||
{"i": 38, "op": "add", "dst": 3, "src": 1, "src2": 7, "imm": "0x75ba2fad", "imm2": "0x230c005c", "rot": 4, "bit": 27, "mask": 1},
|
||||
{"i": 39, "op": "shfl", "dst": 1, "src": 5, "src2": 3, "imm": "0xdbf37e75", "imm2": "0xb5ac1969", "rot": 30, "bit": 13, "mask": 4},
|
||||
{"i": 40, "op": "xor", "dst": 2, "src": 5, "src2": 5, "imm": "0x47f136c5", "imm2": "0x06ce9153", "rot": 19, "bit": 10, "mask": 2},
|
||||
{"i": 41, "op": "mad", "dst": 3, "src": 6, "src2": 3, "imm": "0xce13eff8", "imm2": "0x04cc1d55", "rot": 3, "bit": 1, "mask": 4},
|
||||
{"i": 42, "op": "sub", "dst": 6, "src": 7, "src2": 1, "imm": "0x6a65ab71", "imm2": "0x8fbc1bcd", "rot": 4, "bit": 1, "mask": 8},
|
||||
{"i": 43, "op": "xor", "dst": 7, "src": 0, "src2": 7, "imm": "0xdaeb4928", "imm2": "0xc0423027", "rot": 24, "bit": 11, "mask": 8},
|
||||
{"i": 44, "op": "load", "dst": 1, "src": 7, "src2": 7, "imm": "0x778f01c9", "imm2": "0x28cedcea", "rot": 12, "bit": 4, "mask": 16},
|
||||
{"i": 45, "op": "mul", "dst": 2, "src": 3, "src2": 2, "imm": "0xf4264f1b", "imm2": "0x0f627d56", "rot": 5, "bit": 28, "mask": 8},
|
||||
{"i": 46, "op": "mulhi", "dst": 1, "src": 5, "src2": 1, "imm": "0xffb2147a", "imm2": "0xccde9b05", "rot": 13, "bit": 9, "mask": 2},
|
||||
{"i": 47, "op": "sub", "dst": 4, "src": 3, "src2": 5, "imm": "0x73b36234", "imm2": "0x3f5d5997", "rot": 7, "bit": 18, "mask": 2},
|
||||
{"i": 48, "op": "rotr", "dst": 2, "src": 6, "src2": 3, "imm": "0x3a4d9aa9", "imm2": "0x212bec7b", "rot": 4, "bit": 29, "mask": 16},
|
||||
{"i": 49, "op": "load", "dst": 3, "src": 5, "src2": 2, "imm": "0x626f11df", "imm2": "0x56cd5bfd", "rot": 7, "bit": 1, "mask": 1},
|
||||
{"i": 50, "op": "add", "dst": 1, "src": 5, "src2": 6, "imm": "0x81b8bc2c", "imm2": "0x1907970c", "rot": 28, "bit": 7, "mask": 4},
|
||||
{"i": 51, "op": "mul", "dst": 0, "src": 2, "src2": 2, "imm": "0xa8848b30", "imm2": "0xef6ac348", "rot": 9, "bit": 15, "mask": 8},
|
||||
{"i": 52, "op": "add", "dst": 0, "src": 2, "src2": 0, "imm": "0x4f92b968", "imm2": "0x699fd448", "rot": 22, "bit": 6, "mask": 4},
|
||||
{"i": 53, "op": "add", "dst": 1, "src": 0, "src2": 2, "imm": "0x2bb965af", "imm2": "0x77b1520d", "rot": 2, "bit": 12, "mask": 8},
|
||||
{"i": 54, "op": "rotl", "dst": 7, "src": 1, "src2": 0, "imm": "0x553e678b", "imm2": "0x3cc8eae0", "rot": 14, "bit": 20, "mask": 2},
|
||||
{"i": 55, "op": "add", "dst": 3, "src": 7, "src2": 2, "imm": "0x7b0fe07a", "imm2": "0xa54c55a0", "rot": 10, "bit": 1, "mask": 1},
|
||||
{"i": 56, "op": "load", "dst": 6, "src": 7, "src2": 4, "imm": "0x01eba9aa", "imm2": "0x2758c0f7", "rot": 14, "bit": 15, "mask": 4},
|
||||
{"i": 57, "op": "rotr", "dst": 1, "src": 5, "src2": 2, "imm": "0x1f5267b3", "imm2": "0x236f5a27", "rot": 2, "bit": 31, "mask": 16},
|
||||
{"i": 58, "op": "load", "dst": 5, "src": 4, "src2": 3, "imm": "0xa9954a9b", "imm2": "0x6a54d4e8", "rot": 11, "bit": 10, "mask": 16},
|
||||
{"i": 59, "op": "load", "dst": 6, "src": 2, "src2": 4, "imm": "0x9923ff88", "imm2": "0x9357254e", "rot": 16, "bit": 1, "mask": 16},
|
||||
{"i": 60, "op": "mad", "dst": 3, "src": 5, "src2": 0, "imm": "0xc0cc51a6", "imm2": "0x3fd7701b", "rot": 20, "bit": 1, "mask": 4},
|
||||
{"i": 61, "op": "add", "dst": 5, "src": 7, "src2": 7, "imm": "0xaf9dd72d", "imm2": "0xad7493e7", "rot": 7, "bit": 31, "mask": 16},
|
||||
{"i": 62, "op": "add", "dst": 4, "src": 6, "src2": 2, "imm": "0x89841d87", "imm2": "0x1e07c3d9", "rot": 6, "bit": 27, "mask": 1},
|
||||
{"i": 63, "op": "rotl", "dst": 5, "src": 4, "src2": 2, "imm": "0xae210f8d", "imm2": "0x8e499ba4", "rot": 19, "bit": 9, "mask": 1}
|
||||
]
|
||||
}
|
||||
|
|
|
|||
|
|
@ -38,70 +38,70 @@ kernel void igneum_hash(device const uint* dataset [[buffer(0)]],
|
|||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r4 = rotl_imm(r4, 25u); // 0
|
||||
r0 = r0 - r5; // 1
|
||||
r4 = r4 ^ dataset[r3 & MASK]; // 2
|
||||
r1 = rotl_imm(r1, 1u); // 3
|
||||
r2 = r2 + r3 + select(0x61f0b51cu, 0x2735a174u, ((sel >> 26u) & 1u) != 0u); // 4
|
||||
r5 = r5 ^ dataset[r3 & MASK]; // 5
|
||||
r5 = r5 - r7; // 6
|
||||
r3 = r3 + r4 + select(0x52f2dbf4u, 0x5a069596u, ((sel >> 26u) & 1u) != 0u); // 7
|
||||
r0 = r0 ^ r4; // 8
|
||||
r4 = r4 ^ r0; // 9
|
||||
r2 = r2 ^ dataset[r0 & MASK]; // 10
|
||||
r4 = r4 ^ simd_shuffle_xor(r6, (ushort)16); // 11
|
||||
r1 = r1 - r5; // 12
|
||||
r2 = r2 ^ r1; // 13
|
||||
r4 = r4 ^ dataset[r5 & MASK]; // 14
|
||||
r2 = r2 ^ dataset[r4 & MASK]; // 15
|
||||
r3 = r3 ^ dataset[r0 & MASK]; // 16
|
||||
r4 = r4 ^ r6; // 17
|
||||
r2 = r4 * r6 + r2; // 18
|
||||
r6 = rotr_var(r6, r1); // 19
|
||||
r3 = r3 ^ r4; // 20
|
||||
r1 = r3 * r5 + r1; // 21
|
||||
r7 = mulhi(r7, r4); // 22
|
||||
r5 = mulhi(r5, r2); // 23
|
||||
r0 = r0 ^ dataset[r6 & MASK]; // 24
|
||||
r5 = r5 * r6; // 25
|
||||
r7 = r7 ^ dataset[r1 & MASK]; // 26
|
||||
r3 = rotr_var(r3, r1); // 27
|
||||
r5 = r5 ^ simd_shuffle_xor(r1, (ushort)8); // 28
|
||||
r7 = r7 ^ r5; // 29
|
||||
r7 = rotl_imm(r7, 23u); // 30
|
||||
r2 = r2 - r0; // 31
|
||||
r7 = r7 ^ r2; // 32
|
||||
r2 = r2 ^ r6; // 33
|
||||
r6 = r6 ^ dataset[r1 & MASK]; // 34
|
||||
r1 = r1 ^ r4; // 35
|
||||
r3 = r3 ^ simd_shuffle_xor(r0, (ushort)1); // 36
|
||||
r2 = r2 * r6; // 37
|
||||
r5 = r5 + r3 + select(0x71f30417u, 0xa29f4338u, ((sel >> 12u) & 1u) != 0u); // 38
|
||||
r7 = r7 ^ r6; // 39
|
||||
r7 = r7 ^ r3; // 40
|
||||
r3 = rotr_var(r3, r4); // 41
|
||||
r5 = r5 ^ r3; // 42
|
||||
r3 = rotr_var(r3, r6); // 43
|
||||
r1 = r3 * r5 + r1; // 44
|
||||
r7 = r7 + r3 + select(0xc1ae8d3bu, 0xa907b90bu, ((sel >> 29u) & 1u) != 0u); // 45
|
||||
r7 = r7 | r2; // 46
|
||||
r2 = r2 ^ simd_shuffle_xor(r3, (ushort)16); // 47
|
||||
r1 = r1 ^ dataset[r7 & MASK]; // 48
|
||||
r5 = r5 - r1; // 49
|
||||
r3 = r3 ^ dataset[r1 & MASK]; // 50
|
||||
r2 = r2 - r3; // 51
|
||||
r6 = r6 ^ simd_shuffle_xor(r2, (ushort)2); // 52
|
||||
r2 = r2 - r0; // 53
|
||||
r0 = r0 ^ dataset[r3 & MASK]; // 54
|
||||
r2 = r2 ^ simd_shuffle_xor(r1, (ushort)16); // 55
|
||||
r0 = r0 + r6 + select(0x25955401u, 0xf4689674u, ((sel >> 27u) & 1u) != 0u); // 56
|
||||
r2 = mulhi(r2, r0); // 57
|
||||
r4 = mulhi(r4, r2); // 58
|
||||
r2 = r6 * r7 + r2; // 59
|
||||
r3 = r3 ^ r1; // 60
|
||||
r4 = r4 * r2; // 61
|
||||
r0 = r0 ^ dataset[r2 & MASK]; // 62
|
||||
r7 = mulhi(r7, r0); // 63
|
||||
r2 = r3 * r4 + r2; // 0
|
||||
r2 = r1 * r1 + r2; // 1
|
||||
r2 = r3 * r2 + r2; // 2
|
||||
r3 = r3 ^ r5; // 3
|
||||
r7 = r7 ^ dataset[r2 & MASK]; // 4
|
||||
r5 = r5 ^ dataset[r7 & MASK]; // 5
|
||||
r1 = r1 ^ simd_shuffle_xor(r4, (ushort)8); // 6
|
||||
r7 = r7 ^ simd_shuffle_xor(r3, (ushort)8); // 7
|
||||
r1 = mulhi(r1, r5); // 8
|
||||
r6 = rotr_var(r6, r3); // 9
|
||||
r3 = r3 | r4; // 10
|
||||
r4 = r4 ^ dataset[r3 & MASK]; // 11
|
||||
r0 = mulhi(r0, r4); // 12
|
||||
r5 = r5 + r1 + select(0xc7934706u, 0xd3177981u, ((sel >> 30u) & 1u) != 0u); // 13
|
||||
r0 = r0 ^ dataset[r4 & MASK]; // 14
|
||||
r2 = r2 - r4; // 15
|
||||
r2 = r2 ^ dataset[r0 & MASK]; // 16
|
||||
r7 = r7 ^ dataset[r2 & MASK]; // 17
|
||||
r7 = r7 ^ simd_shuffle_xor(r3, (ushort)4); // 18
|
||||
r5 = r5 * r0; // 19
|
||||
r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 20
|
||||
r2 = r2 ^ simd_shuffle_xor(r4, (ushort)16); // 21
|
||||
r6 = mulhi(r6, r2); // 22
|
||||
r6 = r6 ^ dataset[r1 & MASK]; // 23
|
||||
r5 = r5 * r0; // 24
|
||||
r5 = rotl_imm(r5, 19u); // 25
|
||||
r7 = r7 ^ simd_shuffle_xor(r6, (ushort)2); // 26
|
||||
r0 = r0 ^ r5; // 27
|
||||
r0 = r0 ^ r4; // 28
|
||||
r3 = r3 - r0; // 29
|
||||
r5 = r5 * r1; // 30
|
||||
r7 = r7 ^ dataset[r2 & MASK]; // 31
|
||||
r1 = r1 ^ dataset[r0 & MASK]; // 32
|
||||
r5 = r5 ^ r6; // 33
|
||||
r5 = r5 ^ dataset[r1 & MASK]; // 34
|
||||
r0 = mulhi(r0, r5); // 35
|
||||
r5 = r5 ^ simd_shuffle_xor(r2, (ushort)4); // 36
|
||||
r7 = r7 ^ dataset[r0 & MASK]; // 37
|
||||
r3 = r3 + r1 + select(0x75ba2fadu, 0x230c005cu, ((sel >> 27u) & 1u) != 0u); // 38
|
||||
r1 = r1 ^ simd_shuffle_xor(r5, (ushort)4); // 39
|
||||
r2 = r2 ^ r5; // 40
|
||||
r3 = r6 * r3 + r3; // 41
|
||||
r6 = r6 - r7; // 42
|
||||
r7 = r7 ^ r0; // 43
|
||||
r1 = r1 ^ dataset[r7 & MASK]; // 44
|
||||
r2 = r2 * r3; // 45
|
||||
r1 = mulhi(r1, r5); // 46
|
||||
r4 = r4 - r3; // 47
|
||||
r2 = rotr_var(r2, r6); // 48
|
||||
r3 = r3 ^ dataset[r5 & MASK]; // 49
|
||||
r1 = r1 + r5 + select(0x81b8bc2cu, 0x1907970cu, ((sel >> 7u) & 1u) != 0u); // 50
|
||||
r0 = r0 * r2; // 51
|
||||
r0 = r0 + r2 + select(0x4f92b968u, 0x699fd448u, ((sel >> 6u) & 1u) != 0u); // 52
|
||||
r1 = r1 + r0 + select(0x2bb965afu, 0x77b1520du, ((sel >> 12u) & 1u) != 0u); // 53
|
||||
r7 = rotl_imm(r7, 14u); // 54
|
||||
r3 = r3 + r7 + select(0x7b0fe07au, 0xa54c55a0u, ((sel >> 1u) & 1u) != 0u); // 55
|
||||
r6 = r6 ^ dataset[r7 & MASK]; // 56
|
||||
r1 = rotr_var(r1, r5); // 57
|
||||
r5 = r5 ^ dataset[r4 & MASK]; // 58
|
||||
r6 = r6 ^ dataset[r2 & MASK]; // 59
|
||||
r3 = r5 * r0 + r3; // 60
|
||||
r5 = r5 + r7 + select(0xaf9dd72du, 0xad7493e7u, ((sel >> 31u) & 1u) != 0u); // 61
|
||||
r4 = r4 + r6 + select(0x89841d87u, 0x1e07c3d9u, ((sel >> 27u) & 1u) != 0u); // 62
|
||||
r5 = rotl_imm(r5, 19u); // 63
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
|
|
|
|||
111
proto-cuda/packs/igneum-genesis/program_bound.metal
Normal file
111
proto-cuda/packs/igneum-genesis/program_bound.metal
Normal file
|
|
@ -0,0 +1,111 @@
|
|||
#include <metal_stdlib>
|
||||
using namespace metal;
|
||||
|
||||
#define MASK 0x0fffffffu
|
||||
constant uint SEEDW[8] = { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u };
|
||||
|
||||
inline uint splitmix32(uint x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31
|
||||
inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
inline uint ds_elem(uint i, uint d0, uint d1) {
|
||||
uint x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// Header-bound variant: the init words come from buffer 3 (bind.rs), not from SEEDW.
|
||||
kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]],
|
||||
device ulong* out [[buffer(1)]],
|
||||
constant uint& baseNonce [[buffer(2)]],
|
||||
constant uint* initw [[buffer(3)]],
|
||||
uint gid [[thread_position_in_grid]]) {
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint x = nonce ^ initw[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ initw[1]; }
|
||||
{ uint x = nonce ^ initw[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ initw[2]; }
|
||||
{ uint x = nonce ^ initw[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ initw[3]; }
|
||||
{ uint x = nonce ^ initw[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ initw[4]; }
|
||||
{ uint x = nonce ^ initw[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ initw[5]; }
|
||||
{ uint x = nonce ^ initw[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ initw[6]; }
|
||||
{ uint x = nonce ^ initw[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ initw[7]; }
|
||||
{ uint x = nonce ^ initw[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ initw[0]; }
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r2 = r3 * r4 + r2; // 0
|
||||
r2 = r1 * r1 + r2; // 1
|
||||
r2 = r3 * r2 + r2; // 2
|
||||
r3 = r3 ^ r5; // 3
|
||||
r7 = r7 ^ dataset[r2 & MASK]; // 4
|
||||
r5 = r5 ^ dataset[r7 & MASK]; // 5
|
||||
r1 = r1 ^ simd_shuffle_xor(r4, (ushort)8); // 6
|
||||
r7 = r7 ^ simd_shuffle_xor(r3, (ushort)8); // 7
|
||||
r1 = mulhi(r1, r5); // 8
|
||||
r6 = rotr_var(r6, r3); // 9
|
||||
r3 = r3 | r4; // 10
|
||||
r4 = r4 ^ dataset[r3 & MASK]; // 11
|
||||
r0 = mulhi(r0, r4); // 12
|
||||
r5 = r5 + r1 + select(0xc7934706u, 0xd3177981u, ((sel >> 30u) & 1u) != 0u); // 13
|
||||
r0 = r0 ^ dataset[r4 & MASK]; // 14
|
||||
r2 = r2 - r4; // 15
|
||||
r2 = r2 ^ dataset[r0 & MASK]; // 16
|
||||
r7 = r7 ^ dataset[r2 & MASK]; // 17
|
||||
r7 = r7 ^ simd_shuffle_xor(r3, (ushort)4); // 18
|
||||
r5 = r5 * r0; // 19
|
||||
r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 20
|
||||
r2 = r2 ^ simd_shuffle_xor(r4, (ushort)16); // 21
|
||||
r6 = mulhi(r6, r2); // 22
|
||||
r6 = r6 ^ dataset[r1 & MASK]; // 23
|
||||
r5 = r5 * r0; // 24
|
||||
r5 = rotl_imm(r5, 19u); // 25
|
||||
r7 = r7 ^ simd_shuffle_xor(r6, (ushort)2); // 26
|
||||
r0 = r0 ^ r5; // 27
|
||||
r0 = r0 ^ r4; // 28
|
||||
r3 = r3 - r0; // 29
|
||||
r5 = r5 * r1; // 30
|
||||
r7 = r7 ^ dataset[r2 & MASK]; // 31
|
||||
r1 = r1 ^ dataset[r0 & MASK]; // 32
|
||||
r5 = r5 ^ r6; // 33
|
||||
r5 = r5 ^ dataset[r1 & MASK]; // 34
|
||||
r0 = mulhi(r0, r5); // 35
|
||||
r5 = r5 ^ simd_shuffle_xor(r2, (ushort)4); // 36
|
||||
r7 = r7 ^ dataset[r0 & MASK]; // 37
|
||||
r3 = r3 + r1 + select(0x75ba2fadu, 0x230c005cu, ((sel >> 27u) & 1u) != 0u); // 38
|
||||
r1 = r1 ^ simd_shuffle_xor(r5, (ushort)4); // 39
|
||||
r2 = r2 ^ r5; // 40
|
||||
r3 = r6 * r3 + r3; // 41
|
||||
r6 = r6 - r7; // 42
|
||||
r7 = r7 ^ r0; // 43
|
||||
r1 = r1 ^ dataset[r7 & MASK]; // 44
|
||||
r2 = r2 * r3; // 45
|
||||
r1 = mulhi(r1, r5); // 46
|
||||
r4 = r4 - r3; // 47
|
||||
r2 = rotr_var(r2, r6); // 48
|
||||
r3 = r3 ^ dataset[r5 & MASK]; // 49
|
||||
r1 = r1 + r5 + select(0x81b8bc2cu, 0x1907970cu, ((sel >> 7u) & 1u) != 0u); // 50
|
||||
r0 = r0 * r2; // 51
|
||||
r0 = r0 + r2 + select(0x4f92b968u, 0x699fd448u, ((sel >> 6u) & 1u) != 0u); // 52
|
||||
r1 = r1 + r0 + select(0x2bb965afu, 0x77b1520du, ((sel >> 12u) & 1u) != 0u); // 53
|
||||
r7 = rotl_imm(r7, 14u); // 54
|
||||
r3 = r3 + r7 + select(0x7b0fe07au, 0xa54c55a0u, ((sel >> 1u) & 1u) != 0u); // 55
|
||||
r6 = r6 ^ dataset[r7 & MASK]; // 56
|
||||
r1 = rotr_var(r1, r5); // 57
|
||||
r5 = r5 ^ dataset[r4 & MASK]; // 58
|
||||
r6 = r6 ^ dataset[r2 & MASK]; // 59
|
||||
r3 = r5 * r0 + r3; // 60
|
||||
r5 = r5 + r7 + select(0xaf9dd72du, 0xad7493e7u, ((sel >> 31u) & 1u) != 0u); // 61
|
||||
r4 = r4 + r6 + select(0x89841d87u, 0x1e07c3d9u, ((sel >> 27u) & 1u) != 0u); // 62
|
||||
r5 = rotl_imm(r5, 19u); // 63
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
|
|
@ -1,5 +1,5 @@
|
|||
// Generated by proto-metal/igneum-bench --export-pack for seed "igneum-genesis". Do not edit by hand.
|
||||
// Expected outputs: proto-metal CPU interpreter (cpuWarp, closed-form dataset) on Apple M5 Max; Metal GPU cross-check PASS 3/3 warps
|
||||
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
|
||||
// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v2, closed-form dataset
|
||||
#pragma once
|
||||
#ifdef __cplusplus
|
||||
#include <cstdint>
|
||||
|
|
@ -11,22 +11,22 @@
|
|||
static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u };
|
||||
static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = {
|
||||
{ // base nonce 0
|
||||
0x2941e93c76cb1910ull, 0xc0099c34df8280aeull, 0xe61da8a237797181ull, 0xfb944794eaa8ea1dull, 0x0093934e41befb07ull, 0xd4ba2031c0385915ull, 0x75a8a43902cee166ull, 0xb4ce2d6829aaa27cull,
|
||||
0x478095c2e9911e5dull, 0x595c8107b2792a66ull, 0x21b1b582752cc1d2ull, 0x16e746aa4d9532b8ull, 0x35129dbb17efe2e4ull, 0x926e0b679bf37640ull, 0xb2701dd71a189e19ull, 0x100e94f487022795ull,
|
||||
0x2b9bc5aa1452be87ull, 0x8dabbbe8b909130bull, 0x93a2cd4cba088297ull, 0x662c3e253432099eull, 0xf489da4c9c770867ull, 0xe37864f932e97adfull, 0x3ddf029de9b3c1edull, 0x9059da45130736acull,
|
||||
0x1a667325a17e0016ull, 0x74723d71ab69828aull, 0x121026c2f14795c1ull, 0x4989e1480c662cc8ull, 0xac3b8fd5307c0ab4ull, 0xdc3fc0bb843e84e8ull, 0xfa86df690d119639ull, 0x453388e1be04e25full
|
||||
0x31e7555c3dfd007full, 0x504db59fe5dbe7ecull, 0x40245e17aa00c1daull, 0x41072786c448a218ull, 0xad24c5fd54221ae1ull, 0xdb5a2d3c12002fbdull, 0x4c4f375ff4397f2bull, 0x168fd09d50861741ull,
|
||||
0x37e7ff0f8ad4e164ull, 0xb8f283fcf9ec7c45ull, 0xeae42085b3de9e1bull, 0xe9485b57f2bf3e7bull, 0xe9710ec888bad149ull, 0x2f79adf888179b91ull, 0xbedccd717aaed652ull, 0xbd748a9b3e1bc54full,
|
||||
0x41309662f79b53a0ull, 0x81e5230ebeefc1deull, 0xbcee43f20b08fa91ull, 0x0aedf8399a54dc49ull, 0x55e0cad163ce23eeull, 0x7500caccfdb72ebfull, 0x70e55ef021e6ae7bull, 0xc7ade69c9691176full,
|
||||
0x098eb370ff045d40ull, 0x8d79439fbf37aec8ull, 0xd318d93b86a2f612ull, 0x899bc50487532ae4ull, 0x47b8a0f002c5b1bdull, 0x52dafc738fb7f023ull, 0x4fa59d427c7ad89full, 0xaab617183923ab2aull
|
||||
},
|
||||
{ // base nonce 4096
|
||||
0x198f269c663acaf6ull, 0x6c125dcda3cade45ull, 0x0ae4eba3f2a0cdecull, 0xc112952633d9477cull, 0x2dfc52ebb7281f9eull, 0x70b0e04375823959ull, 0x699e78e051dfaa76ull, 0xd91ae88bc2eba03full,
|
||||
0x928b2939d0b23e68ull, 0x0e5feb811f4dd40bull, 0xd0e4ce74e6693b53ull, 0xbd8ac2d12ea0ea1bull, 0xa39bcd116eec4f88ull, 0xdd07f4f6a5a07514ull, 0xdc1cf8b4615b8842ull, 0xde0fb1599057c2ddull,
|
||||
0x56edb1995f7c341eull, 0x02782ce22dd25939ull, 0x150b52d7b4b87319ull, 0xd5c89a9a6c2786c3ull, 0x8aef759270e00bbaull, 0xf9f26409968ea1eaull, 0x2013ae93398863f0ull, 0x863877ca875900deull,
|
||||
0x6e05e03111bf02a3ull, 0x21ec57017818f221ull, 0x62b33a3456032a3bull, 0x5bc24dc2bf442511ull, 0xb97041874e099be7ull, 0x2d9f09fb5d11fa13ull, 0x43f3f83f764ea247ull, 0xa3ec25a3072b734cull
|
||||
0x7d4a15e723d8817full, 0x2bcdf701795dec45ull, 0x7717812a29875d92ull, 0x9e49e845ba8b23c5ull, 0x6963b6c7e297f4c4ull, 0xbebcda40625761aaull, 0x6083dd36d5abf8f3ull, 0x81206a07d9bbeaa3ull,
|
||||
0xd9f2a9013fd61b70ull, 0x5989d0df80bb696eull, 0x190a28c488cce037ull, 0x509d459608856838ull, 0x40f83aab28ed3347ull, 0x4042ded5f4ef3130ull, 0x64f2c36a1e9bba58ull, 0xbf0a2c66b322d325ull,
|
||||
0x35e9b6f18cd7bb5dull, 0x1cca51c37a7494c2ull, 0x46e46e297ad731cbull, 0x55d054ff7d0e95a9ull, 0xc0442156654b1e3bull, 0x16d42f3229fd7a15ull, 0x21217990d564f3dfull, 0x327fbd5f661f5e5dull,
|
||||
0x9156974deada88e4ull, 0x115951ae239cb392ull, 0x5ac3d4c4850de851ull, 0x2851edaffc3dc794ull, 0xf7d1aed104349d6dull, 0x87f3bca79811253full, 0x4f2d749c2fbd3546ull, 0x601afcfa08fa5865ull
|
||||
},
|
||||
{ // base nonce 1000000
|
||||
0xa63f6d32a9dc2bacull, 0x3263d16a85274409ull, 0xca7b34b1061437c7ull, 0xaf24083fbe9a33acull, 0x68c42544f8798a73ull, 0x5ba44c42278e0d0aull, 0x12b31f0da8ea79e3ull, 0xa4a6186f5f5b7280ull,
|
||||
0x1534d7e10b3f87aeull, 0xb04618c9238e7436ull, 0x28be2d38f9db17d7ull, 0xebed3a6006a491d1ull, 0x23d50e8d3518ee39ull, 0xd75fac65bc8d1975ull, 0x3fd5c0a165c0669bull, 0x882720cae6122082ull,
|
||||
0xc107a8305a1f009dull, 0xa58cc24784025f2cull, 0x7ee69d59a87a92ebull, 0x46009435f6f19482ull, 0xf50c2623a7f856acull, 0x1cc0ddf7b3dc20faull, 0xcbc7f0bb8cd2ec14ull, 0x119c2409ce2462baull,
|
||||
0x1b544f117b80334cull, 0xdb47d6cdd3cf93fcull, 0xdb8c149c9914ddaeull, 0xc6c5e4b5ea412748ull, 0x4437723f548837ffull, 0x1fc1eb1bda8e067cull, 0x786dfa4dc7a84237ull, 0x1fc0ac22d46fb45dull
|
||||
0x9ce07319b124a0d5ull, 0x4349cd39e0d740b5ull, 0xcb681f7c18702383ull, 0xefa90e1ccc2a2b0cull, 0x116c915bca40a85full, 0x0e8323325a63a231ull, 0x7fa7607212b9129bull, 0x5fcd9021b217d391ull,
|
||||
0x9441c1bcbda71988ull, 0xc3322820075d7d22ull, 0x169f475188a31b53ull, 0x9e36cadfb6db42dbull, 0x5a29514f523ca4b2ull, 0x37403b02111cae70ull, 0x2bd904fbe89dc528ull, 0x79d81ed80afd1d00ull,
|
||||
0xf2a50f3cf0516787ull, 0xb3a3a65396b871b3ull, 0xb33190a0b7d859f5ull, 0x9116c7f365b8435bull, 0xf40aba86a73d8ca7ull, 0xe6fe2a11c145fcf9ull, 0x34121d9f8bf51129ull, 0xe219b2070364f3f1ull,
|
||||
0x9ea4199ee8d9b61cull, 0x071e4094c8cb0c9cull, 0xb84fdc6e9e075071ull, 0xd749e28a6f0040b7ull, 0xec2f2075fd67decaull, 0x097ce985109872f9ull, 0x7091dd1d203b22feull, 0x0fcf30256ce1b4a1ull
|
||||
}
|
||||
};
|
||||
|
||||
|
|
|
|||
|
|
@ -5,25 +5,25 @@
|
|||
"dataset_log2_words": 28,
|
||||
"mask": "0x0fffffff",
|
||||
"lanes": 32,
|
||||
"source": "proto-metal CPU interpreter (cpuWarp, closed-form dataset) on Apple M5 Max; Metal GPU cross-check PASS 3/3 warps",
|
||||
"source": "igneum-pow (Rust) CPU interpreter, generator v2, closed-form dataset",
|
||||
"warps": [
|
||||
{"base_nonce": 0, "expected": [
|
||||
"0x2941e93c76cb1910", "0xc0099c34df8280ae", "0xe61da8a237797181", "0xfb944794eaa8ea1d", "0x0093934e41befb07", "0xd4ba2031c0385915", "0x75a8a43902cee166", "0xb4ce2d6829aaa27c",
|
||||
"0x478095c2e9911e5d", "0x595c8107b2792a66", "0x21b1b582752cc1d2", "0x16e746aa4d9532b8", "0x35129dbb17efe2e4", "0x926e0b679bf37640", "0xb2701dd71a189e19", "0x100e94f487022795",
|
||||
"0x2b9bc5aa1452be87", "0x8dabbbe8b909130b", "0x93a2cd4cba088297", "0x662c3e253432099e", "0xf489da4c9c770867", "0xe37864f932e97adf", "0x3ddf029de9b3c1ed", "0x9059da45130736ac",
|
||||
"0x1a667325a17e0016", "0x74723d71ab69828a", "0x121026c2f14795c1", "0x4989e1480c662cc8", "0xac3b8fd5307c0ab4", "0xdc3fc0bb843e84e8", "0xfa86df690d119639", "0x453388e1be04e25f"
|
||||
"0x31e7555c3dfd007f", "0x504db59fe5dbe7ec", "0x40245e17aa00c1da", "0x41072786c448a218", "0xad24c5fd54221ae1", "0xdb5a2d3c12002fbd", "0x4c4f375ff4397f2b", "0x168fd09d50861741",
|
||||
"0x37e7ff0f8ad4e164", "0xb8f283fcf9ec7c45", "0xeae42085b3de9e1b", "0xe9485b57f2bf3e7b", "0xe9710ec888bad149", "0x2f79adf888179b91", "0xbedccd717aaed652", "0xbd748a9b3e1bc54f",
|
||||
"0x41309662f79b53a0", "0x81e5230ebeefc1de", "0xbcee43f20b08fa91", "0x0aedf8399a54dc49", "0x55e0cad163ce23ee", "0x7500caccfdb72ebf", "0x70e55ef021e6ae7b", "0xc7ade69c9691176f",
|
||||
"0x098eb370ff045d40", "0x8d79439fbf37aec8", "0xd318d93b86a2f612", "0x899bc50487532ae4", "0x47b8a0f002c5b1bd", "0x52dafc738fb7f023", "0x4fa59d427c7ad89f", "0xaab617183923ab2a"
|
||||
]},
|
||||
{"base_nonce": 4096, "expected": [
|
||||
"0x198f269c663acaf6", "0x6c125dcda3cade45", "0x0ae4eba3f2a0cdec", "0xc112952633d9477c", "0x2dfc52ebb7281f9e", "0x70b0e04375823959", "0x699e78e051dfaa76", "0xd91ae88bc2eba03f",
|
||||
"0x928b2939d0b23e68", "0x0e5feb811f4dd40b", "0xd0e4ce74e6693b53", "0xbd8ac2d12ea0ea1b", "0xa39bcd116eec4f88", "0xdd07f4f6a5a07514", "0xdc1cf8b4615b8842", "0xde0fb1599057c2dd",
|
||||
"0x56edb1995f7c341e", "0x02782ce22dd25939", "0x150b52d7b4b87319", "0xd5c89a9a6c2786c3", "0x8aef759270e00bba", "0xf9f26409968ea1ea", "0x2013ae93398863f0", "0x863877ca875900de",
|
||||
"0x6e05e03111bf02a3", "0x21ec57017818f221", "0x62b33a3456032a3b", "0x5bc24dc2bf442511", "0xb97041874e099be7", "0x2d9f09fb5d11fa13", "0x43f3f83f764ea247", "0xa3ec25a3072b734c"
|
||||
"0x7d4a15e723d8817f", "0x2bcdf701795dec45", "0x7717812a29875d92", "0x9e49e845ba8b23c5", "0x6963b6c7e297f4c4", "0xbebcda40625761aa", "0x6083dd36d5abf8f3", "0x81206a07d9bbeaa3",
|
||||
"0xd9f2a9013fd61b70", "0x5989d0df80bb696e", "0x190a28c488cce037", "0x509d459608856838", "0x40f83aab28ed3347", "0x4042ded5f4ef3130", "0x64f2c36a1e9bba58", "0xbf0a2c66b322d325",
|
||||
"0x35e9b6f18cd7bb5d", "0x1cca51c37a7494c2", "0x46e46e297ad731cb", "0x55d054ff7d0e95a9", "0xc0442156654b1e3b", "0x16d42f3229fd7a15", "0x21217990d564f3df", "0x327fbd5f661f5e5d",
|
||||
"0x9156974deada88e4", "0x115951ae239cb392", "0x5ac3d4c4850de851", "0x2851edaffc3dc794", "0xf7d1aed104349d6d", "0x87f3bca79811253f", "0x4f2d749c2fbd3546", "0x601afcfa08fa5865"
|
||||
]},
|
||||
{"base_nonce": 1000000, "expected": [
|
||||
"0xa63f6d32a9dc2bac", "0x3263d16a85274409", "0xca7b34b1061437c7", "0xaf24083fbe9a33ac", "0x68c42544f8798a73", "0x5ba44c42278e0d0a", "0x12b31f0da8ea79e3", "0xa4a6186f5f5b7280",
|
||||
"0x1534d7e10b3f87ae", "0xb04618c9238e7436", "0x28be2d38f9db17d7", "0xebed3a6006a491d1", "0x23d50e8d3518ee39", "0xd75fac65bc8d1975", "0x3fd5c0a165c0669b", "0x882720cae6122082",
|
||||
"0xc107a8305a1f009d", "0xa58cc24784025f2c", "0x7ee69d59a87a92eb", "0x46009435f6f19482", "0xf50c2623a7f856ac", "0x1cc0ddf7b3dc20fa", "0xcbc7f0bb8cd2ec14", "0x119c2409ce2462ba",
|
||||
"0x1b544f117b80334c", "0xdb47d6cdd3cf93fc", "0xdb8c149c9914ddae", "0xc6c5e4b5ea412748", "0x4437723f548837ff", "0x1fc1eb1bda8e067c", "0x786dfa4dc7a84237", "0x1fc0ac22d46fb45d"
|
||||
"0x9ce07319b124a0d5", "0x4349cd39e0d740b5", "0xcb681f7c18702383", "0xefa90e1ccc2a2b0c", "0x116c915bca40a85f", "0x0e8323325a63a231", "0x7fa7607212b9129b", "0x5fcd9021b217d391",
|
||||
"0x9441c1bcbda71988", "0xc3322820075d7d22", "0x169f475188a31b53", "0x9e36cadfb6db42db", "0x5a29514f523ca4b2", "0x37403b02111cae70", "0x2bd904fbe89dc528", "0x79d81ed80afd1d00",
|
||||
"0xf2a50f3cf0516787", "0xb3a3a65396b871b3", "0xb33190a0b7d859f5", "0x9116c7f365b8435b", "0xf40aba86a73d8ca7", "0xe6fe2a11c145fcf9", "0x34121d9f8bf51129", "0xe219b2070364f3f1",
|
||||
"0x9ea4199ee8d9b61c", "0x071e4094c8cb0c9c", "0xb84fdc6e9e075071", "0xd749e28a6f0040b7", "0xec2f2075fd67deca", "0x097ce985109872f9", "0x7091dd1d203b22fe", "0x0fcf30256ce1b4a1"
|
||||
]}
|
||||
],
|
||||
"dataset_head": ["0x82174c0f", "0x577bdb9c", "0x111053bf", "0x2bb85514", "0xd83e190f", "0x8ccd9427", "0x3f69d5e4", "0xf3bbe1dc", "0xb3aa7e90", "0xc33f5f73", "0xb83a2b10", "0xc4d7c8ff", "0xefa6d1a8", "0x7029b116", "0x5e48fab0", "0x0ef66e40"],
|
||||
|
|
|
|||
|
|
@ -1,4 +1,4 @@
|
|||
// Generated by proto-metal/igneum-bench --export-pack for seed "igneum-hourly". Do not edit by hand.
|
||||
// Generated by igneum-pow export (generator v2) for seed "igneum-hourly". Do not edit by hand.
|
||||
// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal).
|
||||
// Built from source at runtime by proto-opencl/host.c, which passes these defines:
|
||||
// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit)
|
||||
|
|
@ -96,70 +96,70 @@ IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out
|
|||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r6 = r6 ^ ds[r5 & mask]; // 0 load
|
||||
r3 = r3 ^ r7; // 1 xor
|
||||
r6 = mul_hi(r6, r2); // 2 mulhi
|
||||
r1 = r1 + r0 + ((((sel >> 0u) & 1u) != 0u) ? 0x43f8f369u : 0x1eb46b1cu); // 3 add
|
||||
r3 = r3 ^ ds[r0 & mask]; // 4 load
|
||||
r5 = r5 | r7; // 5 or
|
||||
r4 = r4 ^ ds[r6 & mask]; // 6 load
|
||||
r4 = rotl_imm(r4, 21u); // 7 rotl
|
||||
r6 = r6 ^ ds[r3 & mask]; // 8 load
|
||||
r6 = r6 ^ ds[r1 & mask]; // 9 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 8u); r0 = r0 ^ t_; } // 10 shfl
|
||||
r2 = r2 ^ r3; // 11 xor
|
||||
r2 = r2 + r7 + ((((sel >> 19u) & 1u) != 0u) ? 0xc26c7c2au : 0x3a1ce85eu); // 12 add
|
||||
r4 = r4 ^ ds[r3 & mask]; // 13 load
|
||||
r7 = r7 ^ ds[r1 & mask]; // 14 load
|
||||
r2 = mul_hi(r2, r3); // 15 mulhi
|
||||
r5 = r5 ^ ds[r2 & mask]; // 16 load
|
||||
r5 = r5 ^ ds[r1 & mask]; // 17 load
|
||||
r4 = r4 ^ ds[r7 & mask]; // 18 load
|
||||
r2 = rotr_var(r2, r1); // 19 rotr
|
||||
r7 = r7 ^ ds[r0 & mask]; // 20 load
|
||||
r4 = r4 ^ ds[r6 & mask]; // 21 load
|
||||
r7 = rotl_imm(r7, 25u); // 22 rotl
|
||||
r3 = r3 + r5 + ((((sel >> 15u) & 1u) != 0u) ? 0x98ae0055u : 0x942b819bu); // 23 add
|
||||
r3 = mul_hi(r3, r4); // 24 mulhi
|
||||
r6 = r6 ^ r1; // 25 xor
|
||||
r1 = rotl_imm(r1, 31u); // 26 rotl
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 4u); r3 = r3 ^ t_; } // 27 shfl
|
||||
r6 = r6 - r5; // 28 sub
|
||||
r6 = rotr_var(r6, r3); // 29 rotr
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 4u); r0 = r0 ^ t_; } // 30 shfl
|
||||
r4 = rotl_imm(r4, 30u); // 31 rotl
|
||||
r2 = r2 - r1; // 32 sub
|
||||
r5 = r5 | r4; // 33 or
|
||||
r7 = r6 * r3 + r7; // 34 mad
|
||||
r5 = r5 * r0; // 35 mul
|
||||
r5 = r5 - r3; // 36 sub
|
||||
r2 = r2 + r7 + ((((sel >> 5u) & 1u) != 0u) ? 0x6b5970b5u : 0x473ecfd5u); // 37 add
|
||||
r2 = r2 ^ ds[r7 & mask]; // 38 load
|
||||
r2 = rotr_var(r2, r6); // 39 rotr
|
||||
r0 = r0 ^ r5; // 40 xor
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r5, 8u); r4 = r4 ^ t_; } // 41 shfl
|
||||
r1 = r1 * r6; // 42 mul
|
||||
r0 = r4 * r1 + r0; // 43 mad
|
||||
r1 = r1 + r4 + ((((sel >> 20u) & 1u) != 0u) ? 0x4a502c22u : 0x04e78f3bu); // 44 add
|
||||
r6 = r2 * r7 + r6; // 45 mad
|
||||
r1 = r1 ^ r0; // 46 xor
|
||||
r5 = r5 ^ ds[r7 & mask]; // 47 load
|
||||
r0 = r0 | r4; // 48 or
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 16u); r5 = r5 ^ t_; } // 49 shfl
|
||||
r7 = r7 + r0 + ((((sel >> 17u) & 1u) != 0u) ? 0x6fabf9ceu : 0x0a3df170u); // 50 add
|
||||
r6 = r6 ^ r3; // 51 xor
|
||||
r1 = r1 + r2 + ((((sel >> 21u) & 1u) != 0u) ? 0xad344ca0u : 0xc99bce6fu); // 52 add
|
||||
r6 = r6 * r7; // 53 mul
|
||||
r3 = mul_hi(r3, r4); // 54 mulhi
|
||||
r7 = r7 * r1; // 55 mul
|
||||
r7 = r7 ^ r1; // 56 xor
|
||||
r2 = r2 * r7; // 57 mul
|
||||
r2 = r2 ^ ds[r1 & mask]; // 58 load
|
||||
r7 = r4 * r5 + r7; // 59 mad
|
||||
r2 = r2 ^ ds[r7 & mask]; // 60 load
|
||||
r0 = r2 * r3 + r0; // 61 mad
|
||||
r1 = mul_hi(r1, r5); // 62 mulhi
|
||||
r7 = r7 + r3 + ((((sel >> 24u) & 1u) != 0u) ? 0x05e4fc1du : 0xc23e27c9u); // 63 add
|
||||
r5 = mul_hi(r5, r0); // 0 mulhi
|
||||
r6 = r6 ^ ds[r5 & mask]; // 1 load
|
||||
r0 = r0 ^ ds[r6 & mask]; // 2 load
|
||||
r5 = r5 ^ ds[r0 & mask]; // 3 load
|
||||
r3 = r3 - r0; // 4 sub
|
||||
r6 = r5 * r4 + r6; // 5 mad
|
||||
r7 = r7 * r2; // 6 mul
|
||||
r5 = r5 + r1 + ((((sel >> 1u) & 1u) != 0u) ? 0xc4382c99u : 0x99d561fcu); // 7 add
|
||||
r1 = r1 ^ ds[r7 & mask]; // 8 load
|
||||
r5 = r5 ^ ds[r1 & mask]; // 9 load
|
||||
r0 = r6 * r2 + r0; // 10 mad
|
||||
r6 = mul_hi(r6, r0); // 11 mulhi
|
||||
r2 = rotl_imm(r2, 25u); // 12 rotl
|
||||
r5 = r5 ^ ds[r0 & mask]; // 13 load
|
||||
r2 = mul_hi(r2, r1); // 14 mulhi
|
||||
r7 = r7 + r4 + ((((sel >> 17u) & 1u) != 0u) ? 0x4ab35569u : 0xa105b846u); // 15 add
|
||||
r1 = r1 ^ ds[r6 & mask]; // 16 load
|
||||
r7 = r7 ^ ds[r1 & mask]; // 17 load
|
||||
r2 = rotr_var(r2, r0); // 18 rotr
|
||||
r7 = r7 ^ r3; // 19 xor
|
||||
r1 = mul_hi(r1, r6); // 20 mulhi
|
||||
r3 = r3 ^ ds[r5 & mask]; // 21 load
|
||||
r6 = r6 ^ r3; // 22 xor
|
||||
r0 = r0 * r3; // 23 mul
|
||||
r4 = r4 + r0 + ((((sel >> 2u) & 1u) != 0u) ? 0x344a9ec0u : 0x3706948au); // 24 add
|
||||
r7 = rotl_imm(r7, 24u); // 25 rotl
|
||||
r3 = r3 ^ r4; // 26 xor
|
||||
r2 = r2 ^ r0; // 27 xor
|
||||
r0 = r0 ^ r7; // 28 xor
|
||||
r3 = r3 - r1; // 29 sub
|
||||
r5 = r5 ^ ds[r7 & mask]; // 30 load
|
||||
r0 = r0 ^ ds[r2 & mask]; // 31 load
|
||||
r0 = r1 * r7 + r0; // 32 mad
|
||||
r5 = r5 ^ ds[r4 & mask]; // 33 load
|
||||
r6 = r6 ^ r0; // 34 xor
|
||||
r1 = r1 ^ ds[r0 & mask]; // 35 load
|
||||
r1 = r0 * r2 + r1; // 36 mad
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r7 = r7 ^ t_; } // 37 shfl
|
||||
r6 = rotl_imm(r6, 21u); // 38 rotl
|
||||
r5 = r5 ^ r4; // 39 xor
|
||||
r5 = r1 * r1 + r5; // 40 mad
|
||||
r3 = r3 + r4 + ((((sel >> 17u) & 1u) != 0u) ? 0x7116ab79u : 0x2f1426d3u); // 41 add
|
||||
r7 = r7 - r1; // 42 sub
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 43 shfl
|
||||
r6 = rotr_var(r6, r3); // 44 rotr
|
||||
r1 = rotl_imm(r1, 18u); // 45 rotl
|
||||
r1 = r1 ^ ds[r3 & mask]; // 46 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 8u); r5 = r5 ^ t_; } // 47 shfl
|
||||
r0 = r0 ^ ds[r7 & mask]; // 48 load
|
||||
r2 = rotl_imm(r2, 27u); // 49 rotl
|
||||
r1 = r1 ^ r7; // 50 xor
|
||||
r6 = r6 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xbc8977b1u : 0xaa5a26d7u); // 51 add
|
||||
r0 = r0 ^ r5; // 52 xor
|
||||
r1 = rotr_var(r1, r5); // 53 rotr
|
||||
r6 = rotl_imm(r6, 24u); // 54 rotl
|
||||
r1 = r1 ^ r5; // 55 xor
|
||||
r6 = r6 ^ r7; // 56 xor
|
||||
r5 = r5 ^ r7; // 57 xor
|
||||
r2 = r2 ^ r5; // 58 xor
|
||||
r0 = r4 * r0 + r0; // 59 mad
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 8u); r7 = r7 ^ t_; } // 60 shfl
|
||||
r6 = r6 ^ ds[r5 & mask]; // 61 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 2u); r7 = r7 ^ t_; } // 62 shfl
|
||||
r5 = r5 | r2; // 63 or
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
|
|
|
|||
|
|
@ -1,4 +1,4 @@
|
|||
// Generated by proto-metal/igneum-bench --export-pack for seed "igneum-hourly". Do not edit by hand.
|
||||
// Generated by igneum-pow export (generator v2) for seed "igneum-hourly". Do not edit by hand.
|
||||
// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal).
|
||||
// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC.
|
||||
#include <cuda_runtime.h>
|
||||
|
|
@ -48,70 +48,70 @@ __global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonc
|
|||
|
||||
for (uint32_t it = 0u; it < 8u; ++it) {
|
||||
uint32_t sel = r0;
|
||||
r6 = r6 ^ ds[r5 & mask]; // 0 load
|
||||
r3 = r3 ^ r7; // 1 xor
|
||||
r6 = __umulhi(r6, r2); // 2 mulhi
|
||||
r1 = r1 + r0 + ((((sel >> 0u) & 1u) != 0u) ? 0x43f8f369u : 0x1eb46b1cu); // 3 add
|
||||
r3 = r3 ^ ds[r0 & mask]; // 4 load
|
||||
r5 = r5 | r7; // 5 or
|
||||
r4 = r4 ^ ds[r6 & mask]; // 6 load
|
||||
r4 = rotl_imm(r4, 21u); // 7 rotl
|
||||
r6 = r6 ^ ds[r3 & mask]; // 8 load
|
||||
r6 = r6 ^ ds[r1 & mask]; // 9 load
|
||||
r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r3, 8); // 10 shfl
|
||||
r2 = r2 ^ r3; // 11 xor
|
||||
r2 = r2 + r7 + ((((sel >> 19u) & 1u) != 0u) ? 0xc26c7c2au : 0x3a1ce85eu); // 12 add
|
||||
r4 = r4 ^ ds[r3 & mask]; // 13 load
|
||||
r7 = r7 ^ ds[r1 & mask]; // 14 load
|
||||
r2 = __umulhi(r2, r3); // 15 mulhi
|
||||
r5 = r5 ^ ds[r2 & mask]; // 16 load
|
||||
r5 = r5 ^ ds[r1 & mask]; // 17 load
|
||||
r4 = r4 ^ ds[r7 & mask]; // 18 load
|
||||
r2 = rotr_var(r2, r1); // 19 rotr
|
||||
r7 = r7 ^ ds[r0 & mask]; // 20 load
|
||||
r4 = r4 ^ ds[r6 & mask]; // 21 load
|
||||
r7 = rotl_imm(r7, 25u); // 22 rotl
|
||||
r3 = r3 + r5 + ((((sel >> 15u) & 1u) != 0u) ? 0x98ae0055u : 0x942b819bu); // 23 add
|
||||
r3 = __umulhi(r3, r4); // 24 mulhi
|
||||
r6 = r6 ^ r1; // 25 xor
|
||||
r1 = rotl_imm(r1, 31u); // 26 rotl
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r6, 4); // 27 shfl
|
||||
r6 = r6 - r5; // 28 sub
|
||||
r6 = rotr_var(r6, r3); // 29 rotr
|
||||
r0 = r0 ^ __shfl_xor_sync(0xffffffffu, r4, 4); // 30 shfl
|
||||
r4 = rotl_imm(r4, 30u); // 31 rotl
|
||||
r2 = r2 - r1; // 32 sub
|
||||
r5 = r5 | r4; // 33 or
|
||||
r7 = r6 * r3 + r7; // 34 mad
|
||||
r5 = r5 * r0; // 35 mul
|
||||
r5 = r5 - r3; // 36 sub
|
||||
r2 = r2 + r7 + ((((sel >> 5u) & 1u) != 0u) ? 0x6b5970b5u : 0x473ecfd5u); // 37 add
|
||||
r2 = r2 ^ ds[r7 & mask]; // 38 load
|
||||
r2 = rotr_var(r2, r6); // 39 rotr
|
||||
r0 = r0 ^ r5; // 40 xor
|
||||
r4 = r4 ^ __shfl_xor_sync(0xffffffffu, r5, 8); // 41 shfl
|
||||
r1 = r1 * r6; // 42 mul
|
||||
r0 = r4 * r1 + r0; // 43 mad
|
||||
r1 = r1 + r4 + ((((sel >> 20u) & 1u) != 0u) ? 0x4a502c22u : 0x04e78f3bu); // 44 add
|
||||
r6 = r2 * r7 + r6; // 45 mad
|
||||
r1 = r1 ^ r0; // 46 xor
|
||||
r5 = r5 ^ ds[r7 & mask]; // 47 load
|
||||
r0 = r0 | r4; // 48 or
|
||||
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r4, 16); // 49 shfl
|
||||
r7 = r7 + r0 + ((((sel >> 17u) & 1u) != 0u) ? 0x6fabf9ceu : 0x0a3df170u); // 50 add
|
||||
r6 = r6 ^ r3; // 51 xor
|
||||
r1 = r1 + r2 + ((((sel >> 21u) & 1u) != 0u) ? 0xad344ca0u : 0xc99bce6fu); // 52 add
|
||||
r6 = r6 * r7; // 53 mul
|
||||
r3 = __umulhi(r3, r4); // 54 mulhi
|
||||
r7 = r7 * r1; // 55 mul
|
||||
r7 = r7 ^ r1; // 56 xor
|
||||
r2 = r2 * r7; // 57 mul
|
||||
r2 = r2 ^ ds[r1 & mask]; // 58 load
|
||||
r7 = r4 * r5 + r7; // 59 mad
|
||||
r2 = r2 ^ ds[r7 & mask]; // 60 load
|
||||
r0 = r2 * r3 + r0; // 61 mad
|
||||
r1 = __umulhi(r1, r5); // 62 mulhi
|
||||
r7 = r7 + r3 + ((((sel >> 24u) & 1u) != 0u) ? 0x05e4fc1du : 0xc23e27c9u); // 63 add
|
||||
r5 = __umulhi(r5, r0); // 0 mulhi
|
||||
r6 = r6 ^ ds[r5 & mask]; // 1 load
|
||||
r0 = r0 ^ ds[r6 & mask]; // 2 load
|
||||
r5 = r5 ^ ds[r0 & mask]; // 3 load
|
||||
r3 = r3 - r0; // 4 sub
|
||||
r6 = r5 * r4 + r6; // 5 mad
|
||||
r7 = r7 * r2; // 6 mul
|
||||
r5 = r5 + r1 + ((((sel >> 1u) & 1u) != 0u) ? 0xc4382c99u : 0x99d561fcu); // 7 add
|
||||
r1 = r1 ^ ds[r7 & mask]; // 8 load
|
||||
r5 = r5 ^ ds[r1 & mask]; // 9 load
|
||||
r0 = r6 * r2 + r0; // 10 mad
|
||||
r6 = __umulhi(r6, r0); // 11 mulhi
|
||||
r2 = rotl_imm(r2, 25u); // 12 rotl
|
||||
r5 = r5 ^ ds[r0 & mask]; // 13 load
|
||||
r2 = __umulhi(r2, r1); // 14 mulhi
|
||||
r7 = r7 + r4 + ((((sel >> 17u) & 1u) != 0u) ? 0x4ab35569u : 0xa105b846u); // 15 add
|
||||
r1 = r1 ^ ds[r6 & mask]; // 16 load
|
||||
r7 = r7 ^ ds[r1 & mask]; // 17 load
|
||||
r2 = rotr_var(r2, r0); // 18 rotr
|
||||
r7 = r7 ^ r3; // 19 xor
|
||||
r1 = __umulhi(r1, r6); // 20 mulhi
|
||||
r3 = r3 ^ ds[r5 & mask]; // 21 load
|
||||
r6 = r6 ^ r3; // 22 xor
|
||||
r0 = r0 * r3; // 23 mul
|
||||
r4 = r4 + r0 + ((((sel >> 2u) & 1u) != 0u) ? 0x344a9ec0u : 0x3706948au); // 24 add
|
||||
r7 = rotl_imm(r7, 24u); // 25 rotl
|
||||
r3 = r3 ^ r4; // 26 xor
|
||||
r2 = r2 ^ r0; // 27 xor
|
||||
r0 = r0 ^ r7; // 28 xor
|
||||
r3 = r3 - r1; // 29 sub
|
||||
r5 = r5 ^ ds[r7 & mask]; // 30 load
|
||||
r0 = r0 ^ ds[r2 & mask]; // 31 load
|
||||
r0 = r1 * r7 + r0; // 32 mad
|
||||
r5 = r5 ^ ds[r4 & mask]; // 33 load
|
||||
r6 = r6 ^ r0; // 34 xor
|
||||
r1 = r1 ^ ds[r0 & mask]; // 35 load
|
||||
r1 = r0 * r2 + r1; // 36 mad
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 37 shfl
|
||||
r6 = rotl_imm(r6, 21u); // 38 rotl
|
||||
r5 = r5 ^ r4; // 39 xor
|
||||
r5 = r1 * r1 + r5; // 40 mad
|
||||
r3 = r3 + r4 + ((((sel >> 17u) & 1u) != 0u) ? 0x7116ab79u : 0x2f1426d3u); // 41 add
|
||||
r7 = r7 - r1; // 42 sub
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 43 shfl
|
||||
r6 = rotr_var(r6, r3); // 44 rotr
|
||||
r1 = rotl_imm(r1, 18u); // 45 rotl
|
||||
r1 = r1 ^ ds[r3 & mask]; // 46 load
|
||||
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r3, 8); // 47 shfl
|
||||
r0 = r0 ^ ds[r7 & mask]; // 48 load
|
||||
r2 = rotl_imm(r2, 27u); // 49 rotl
|
||||
r1 = r1 ^ r7; // 50 xor
|
||||
r6 = r6 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xbc8977b1u : 0xaa5a26d7u); // 51 add
|
||||
r0 = r0 ^ r5; // 52 xor
|
||||
r1 = rotr_var(r1, r5); // 53 rotr
|
||||
r6 = rotl_imm(r6, 24u); // 54 rotl
|
||||
r1 = r1 ^ r5; // 55 xor
|
||||
r6 = r6 ^ r7; // 56 xor
|
||||
r5 = r5 ^ r7; // 57 xor
|
||||
r2 = r2 ^ r5; // 58 xor
|
||||
r0 = r4 * r0 + r0; // 59 mad
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r6, 8); // 60 shfl
|
||||
r6 = r6 ^ ds[r5 & mask]; // 61 load
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 2); // 62 shfl
|
||||
r5 = r5 | r2; // 63 or
|
||||
}
|
||||
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
|
|
|
|||
270
proto-cuda/packs/igneum-hourly/kernel_bound.cl
Normal file
270
proto-cuda/packs/igneum-hourly/kernel_bound.cl
Normal file
|
|
@ -0,0 +1,270 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-hourly". Do not edit by hand.
|
||||
// OpenCL C twin of the Metal kernel for the same seed (see proto-opencl/README.md, WAVEFRONT.md and program.metal).
|
||||
// Built from source at runtime by proto-opencl/host.c, which passes these defines:
|
||||
// IGNEUM_GROUP work-group size of igneum_hash, a multiple of 32 (default 32: one work-group = one 32-lane unit)
|
||||
// IGNEUM_EXCHANGE 0 = local-memory exchange with a barrier (any device, any wave width; the default)
|
||||
// 1 = sub_group_shuffle_xor (cl_khr_subgroup_shuffle), only with IGNEUM_GROUP 32 and a sub-group size of exactly 32
|
||||
// 2 = intel_sub_group_shuffle_xor (cl_intel_subgroups), same condition
|
||||
// The verification unit is always 32 lanes. A 64-wide hardware wave (AMD GCN/CDNA, RDNA in wave64) runs two units;
|
||||
// the exchange masks are 1, 2, 4, 8, 16, so every partner lane lies inside the lane's own aligned run of 32.
|
||||
#ifndef IGNEUM_GROUP
|
||||
#define IGNEUM_GROUP 32
|
||||
#endif
|
||||
#ifndef IGNEUM_EXCHANGE
|
||||
#define IGNEUM_EXCHANGE 0
|
||||
#endif
|
||||
#ifdef __OPENCL_VERSION__
|
||||
#define IGNEUM_KERNEL_HASH __kernel __attribute__((reqd_work_group_size(IGNEUM_GROUP, 1, 1)))
|
||||
#define IGNEUM_LOCAL_WORDS(name, n) __local uint name[n]
|
||||
#if IGNEUM_EXCHANGE == 1
|
||||
#ifdef cl_khr_subgroups
|
||||
#pragma OPENCL EXTENSION cl_khr_subgroups : enable
|
||||
#endif
|
||||
#ifdef cl_khr_subgroup_shuffle
|
||||
#pragma OPENCL EXTENSION cl_khr_subgroup_shuffle : enable
|
||||
#endif
|
||||
#elif IGNEUM_EXCHANGE == 2
|
||||
#pragma OPENCL EXTENSION cl_intel_subgroups : enable
|
||||
#endif
|
||||
#else
|
||||
// Not an OpenCL compiler: proto-opencl/emu compiles this file as C++ and supplies the built-ins and these two macros.
|
||||
#include "emu_opencl.h"
|
||||
#endif
|
||||
|
||||
#if IGNEUM_EXCHANGE == 1
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) dst = sub_group_shuffle_xor((a), (uint)(m))
|
||||
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
|
||||
#elif IGNEUM_EXCHANGE == 2
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) dst = intel_sub_group_shuffle_xor((a), (uint)(m))
|
||||
#define IGNEUM_BCAST0(dst, a) dst = sub_group_broadcast((a), 0u)
|
||||
#else
|
||||
// Local-memory exchange. Two buffers of IGNEUM_GROUP words alternate (xk counts exchanges), so one barrier per
|
||||
// exchange is enough: a lane can only overwrite buffer b at exchange k+2 after passing barrier k+1, and every lane
|
||||
// reaches barrier k+1 only after its read of buffer b at exchange k. The partner lid ^ m stays inside the lane's
|
||||
// aligned run of 32 because m < 32. Control flow is uniform, so every work-item reaches every barrier.
|
||||
#define IGNEUM_SHFL_XOR(dst, a, m) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid ^ (uint)(m))]; xk += 1u; }
|
||||
#define IGNEUM_BCAST0(dst, a) { xch[(xk & 1u) * IGNEUM_GROUP + lid] = (a); barrier(CLK_LOCAL_MEM_FENCE); dst = xch[(xk & 1u) * IGNEUM_GROUP + (lid & ~31u)]; xk += 1u; }
|
||||
#endif
|
||||
|
||||
static inline uint splitmix32(uint x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
// n is a literal in 1..31 at every call site. OpenCL rotate() rotates left by n modulo 32.
|
||||
static inline uint rotl_imm(uint x, uint n) { return rotate(x, n); }
|
||||
// Right rotation by n modulo 32 as a left rotation by (32 - n) modulo 32; n == 0 gives x.
|
||||
static inline uint rotr_var(uint x, uint n) { return rotate(x, (0u - n) & 31u); }
|
||||
static inline uint ds_elem(uint i, uint d0, uint d1) {
|
||||
uint x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// dataset[i] = ds_elem(i, d0, d1) for i < n. Same closed form as the Metal igneum_fill kernel.
|
||||
__kernel void igneum_fill(__global uint* ds, uint n, uint d0, uint d1) {
|
||||
uint i = (uint)get_global_id(0);
|
||||
if (i < n) ds[i] = ds_elem(i, d0, d1);
|
||||
}
|
||||
|
||||
// One hash per work-item. IGNEUM_GROUP is a multiple of 32; lane = lid & 31 and every exchange stays inside the
|
||||
// lane's own aligned run of 32 work-items, exactly like simd_shuffle_xor inside a 32-wide Metal SIMD group and
|
||||
// __shfl_xor_sync inside a CUDA warp. Control flow is uniform (no branches at all).
|
||||
IGNEUM_KERNEL_HASH void igneum_hash(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask) {
|
||||
uint gid = (uint)get_global_id(0);
|
||||
uint lid = (uint)get_local_id(0);
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
#if IGNEUM_EXCHANGE == 0
|
||||
IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
|
||||
uint xk = 0u;
|
||||
#else
|
||||
(void)lid;
|
||||
#endif
|
||||
{ uint x = nonce ^ 0x6bdee811u; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x8f488bbeu; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
|
||||
{ uint x = nonce ^ 0x8f488bbeu; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xc5cdece7u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
|
||||
{ uint x = nonce ^ 0xc5cdece7u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x210af22du; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
|
||||
{ uint x = nonce ^ 0x210af22du; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0x2f687b65u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
|
||||
{ uint x = nonce ^ 0x2f687b65u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0x17471eeeu; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
|
||||
{ uint x = nonce ^ 0x17471eeeu; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xee16e284u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
|
||||
{ uint x = nonce ^ 0xee16e284u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xfc9eb8f9u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
|
||||
{ uint x = nonce ^ 0xfc9eb8f9u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x6bdee811u; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r5 = mul_hi(r5, r0); // 0 mulhi
|
||||
r6 = r6 ^ ds[r5 & mask]; // 1 load
|
||||
r0 = r0 ^ ds[r6 & mask]; // 2 load
|
||||
r5 = r5 ^ ds[r0 & mask]; // 3 load
|
||||
r3 = r3 - r0; // 4 sub
|
||||
r6 = r5 * r4 + r6; // 5 mad
|
||||
r7 = r7 * r2; // 6 mul
|
||||
r5 = r5 + r1 + ((((sel >> 1u) & 1u) != 0u) ? 0xc4382c99u : 0x99d561fcu); // 7 add
|
||||
r1 = r1 ^ ds[r7 & mask]; // 8 load
|
||||
r5 = r5 ^ ds[r1 & mask]; // 9 load
|
||||
r0 = r6 * r2 + r0; // 10 mad
|
||||
r6 = mul_hi(r6, r0); // 11 mulhi
|
||||
r2 = rotl_imm(r2, 25u); // 12 rotl
|
||||
r5 = r5 ^ ds[r0 & mask]; // 13 load
|
||||
r2 = mul_hi(r2, r1); // 14 mulhi
|
||||
r7 = r7 + r4 + ((((sel >> 17u) & 1u) != 0u) ? 0x4ab35569u : 0xa105b846u); // 15 add
|
||||
r1 = r1 ^ ds[r6 & mask]; // 16 load
|
||||
r7 = r7 ^ ds[r1 & mask]; // 17 load
|
||||
r2 = rotr_var(r2, r0); // 18 rotr
|
||||
r7 = r7 ^ r3; // 19 xor
|
||||
r1 = mul_hi(r1, r6); // 20 mulhi
|
||||
r3 = r3 ^ ds[r5 & mask]; // 21 load
|
||||
r6 = r6 ^ r3; // 22 xor
|
||||
r0 = r0 * r3; // 23 mul
|
||||
r4 = r4 + r0 + ((((sel >> 2u) & 1u) != 0u) ? 0x344a9ec0u : 0x3706948au); // 24 add
|
||||
r7 = rotl_imm(r7, 24u); // 25 rotl
|
||||
r3 = r3 ^ r4; // 26 xor
|
||||
r2 = r2 ^ r0; // 27 xor
|
||||
r0 = r0 ^ r7; // 28 xor
|
||||
r3 = r3 - r1; // 29 sub
|
||||
r5 = r5 ^ ds[r7 & mask]; // 30 load
|
||||
r0 = r0 ^ ds[r2 & mask]; // 31 load
|
||||
r0 = r1 * r7 + r0; // 32 mad
|
||||
r5 = r5 ^ ds[r4 & mask]; // 33 load
|
||||
r6 = r6 ^ r0; // 34 xor
|
||||
r1 = r1 ^ ds[r0 & mask]; // 35 load
|
||||
r1 = r0 * r2 + r1; // 36 mad
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r7 = r7 ^ t_; } // 37 shfl
|
||||
r6 = rotl_imm(r6, 21u); // 38 rotl
|
||||
r5 = r5 ^ r4; // 39 xor
|
||||
r5 = r1 * r1 + r5; // 40 mad
|
||||
r3 = r3 + r4 + ((((sel >> 17u) & 1u) != 0u) ? 0x7116ab79u : 0x2f1426d3u); // 41 add
|
||||
r7 = r7 - r1; // 42 sub
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 43 shfl
|
||||
r6 = rotr_var(r6, r3); // 44 rotr
|
||||
r1 = rotl_imm(r1, 18u); // 45 rotl
|
||||
r1 = r1 ^ ds[r3 & mask]; // 46 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 8u); r5 = r5 ^ t_; } // 47 shfl
|
||||
r0 = r0 ^ ds[r7 & mask]; // 48 load
|
||||
r2 = rotl_imm(r2, 27u); // 49 rotl
|
||||
r1 = r1 ^ r7; // 50 xor
|
||||
r6 = r6 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xbc8977b1u : 0xaa5a26d7u); // 51 add
|
||||
r0 = r0 ^ r5; // 52 xor
|
||||
r1 = rotr_var(r1, r5); // 53 rotr
|
||||
r6 = rotl_imm(r6, 24u); // 54 rotl
|
||||
r1 = r1 ^ r5; // 55 xor
|
||||
r6 = r6 ^ r7; // 56 xor
|
||||
r5 = r5 ^ r7; // 57 xor
|
||||
r2 = r2 ^ r5; // 58 xor
|
||||
r0 = r4 * r0 + r0; // 59 mad
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 8u); r7 = r7 ^ t_; } // 60 shfl
|
||||
r6 = r6 ^ ds[r5 & mask]; // 61 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 2u); r7 = r7 ^ t_; } // 62 shfl
|
||||
r5 = r5 | r2; // 63 or
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
|
||||
#if IGNEUM_EXCHANGE != 0
|
||||
// Reports the sub-group size this device uses for a work-group of IGNEUM_GROUP items. host.c runs it only when the
|
||||
// per-kernel query (clGetKernelSubGroupInfoKHR on igneum_hash) is unavailable; that query is preferred because a
|
||||
// compiler may pick a different wave width per kernel (RDNA: wave32 or wave64). See WAVEFRONT.md.
|
||||
IGNEUM_KERNEL_HASH void igneum_probe_subgroup(__global uint* out) {
|
||||
if (get_local_id(0) == 0u) { out[0] = get_sub_group_size(); out[1] = get_num_sub_groups(); }
|
||||
}
|
||||
#endif
|
||||
|
||||
// Header-bound variant (bind.rs): the init words come from initw, not SEEDW. Same body as igneum_hash.
|
||||
IGNEUM_KERNEL_HASH void igneum_hash_bound(__global const uint* ds, __global ulong* out, uint baseNonce, uint mask, __global const uint* initw) {
|
||||
uint gid = (uint)get_global_id(0);
|
||||
uint lid = (uint)get_local_id(0);
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
uint iw0 = initw[0], iw1 = initw[1], iw2 = initw[2], iw3 = initw[3], iw4 = initw[4], iw5 = initw[5], iw6 = initw[6], iw7 = initw[7];
|
||||
#if IGNEUM_EXCHANGE == 0
|
||||
IGNEUM_LOCAL_WORDS(xch, 2 * IGNEUM_GROUP);
|
||||
uint xk = 0u;
|
||||
#else
|
||||
(void)lid;
|
||||
#endif
|
||||
{ uint x = nonce ^ iw0; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw1; }
|
||||
{ uint x = nonce ^ iw1; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw2; }
|
||||
{ uint x = nonce ^ iw2; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw3; }
|
||||
{ uint x = nonce ^ iw3; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw4; }
|
||||
{ uint x = nonce ^ iw4; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw5; }
|
||||
{ uint x = nonce ^ iw5; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw6; }
|
||||
{ uint x = nonce ^ iw6; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw7; }
|
||||
{ uint x = nonce ^ iw7; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw0; }
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r5 = mul_hi(r5, r0); // 0 mulhi
|
||||
r6 = r6 ^ ds[r5 & mask]; // 1 load
|
||||
r0 = r0 ^ ds[r6 & mask]; // 2 load
|
||||
r5 = r5 ^ ds[r0 & mask]; // 3 load
|
||||
r3 = r3 - r0; // 4 sub
|
||||
r6 = r5 * r4 + r6; // 5 mad
|
||||
r7 = r7 * r2; // 6 mul
|
||||
r5 = r5 + r1 + ((((sel >> 1u) & 1u) != 0u) ? 0xc4382c99u : 0x99d561fcu); // 7 add
|
||||
r1 = r1 ^ ds[r7 & mask]; // 8 load
|
||||
r5 = r5 ^ ds[r1 & mask]; // 9 load
|
||||
r0 = r6 * r2 + r0; // 10 mad
|
||||
r6 = mul_hi(r6, r0); // 11 mulhi
|
||||
r2 = rotl_imm(r2, 25u); // 12 rotl
|
||||
r5 = r5 ^ ds[r0 & mask]; // 13 load
|
||||
r2 = mul_hi(r2, r1); // 14 mulhi
|
||||
r7 = r7 + r4 + ((((sel >> 17u) & 1u) != 0u) ? 0x4ab35569u : 0xa105b846u); // 15 add
|
||||
r1 = r1 ^ ds[r6 & mask]; // 16 load
|
||||
r7 = r7 ^ ds[r1 & mask]; // 17 load
|
||||
r2 = rotr_var(r2, r0); // 18 rotr
|
||||
r7 = r7 ^ r3; // 19 xor
|
||||
r1 = mul_hi(r1, r6); // 20 mulhi
|
||||
r3 = r3 ^ ds[r5 & mask]; // 21 load
|
||||
r6 = r6 ^ r3; // 22 xor
|
||||
r0 = r0 * r3; // 23 mul
|
||||
r4 = r4 + r0 + ((((sel >> 2u) & 1u) != 0u) ? 0x344a9ec0u : 0x3706948au); // 24 add
|
||||
r7 = rotl_imm(r7, 24u); // 25 rotl
|
||||
r3 = r3 ^ r4; // 26 xor
|
||||
r2 = r2 ^ r0; // 27 xor
|
||||
r0 = r0 ^ r7; // 28 xor
|
||||
r3 = r3 - r1; // 29 sub
|
||||
r5 = r5 ^ ds[r7 & mask]; // 30 load
|
||||
r0 = r0 ^ ds[r2 & mask]; // 31 load
|
||||
r0 = r1 * r7 + r0; // 32 mad
|
||||
r5 = r5 ^ ds[r4 & mask]; // 33 load
|
||||
r6 = r6 ^ r0; // 34 xor
|
||||
r1 = r1 ^ ds[r0 & mask]; // 35 load
|
||||
r1 = r0 * r2 + r1; // 36 mad
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 8u); r7 = r7 ^ t_; } // 37 shfl
|
||||
r6 = rotl_imm(r6, 21u); // 38 rotl
|
||||
r5 = r5 ^ r4; // 39 xor
|
||||
r5 = r1 * r1 + r5; // 40 mad
|
||||
r3 = r3 + r4 + ((((sel >> 17u) & 1u) != 0u) ? 0x7116ab79u : 0x2f1426d3u); // 41 add
|
||||
r7 = r7 - r1; // 42 sub
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r4, 2u); r3 = r3 ^ t_; } // 43 shfl
|
||||
r6 = rotr_var(r6, r3); // 44 rotr
|
||||
r1 = rotl_imm(r1, 18u); // 45 rotl
|
||||
r1 = r1 ^ ds[r3 & mask]; // 46 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 8u); r5 = r5 ^ t_; } // 47 shfl
|
||||
r0 = r0 ^ ds[r7 & mask]; // 48 load
|
||||
r2 = rotl_imm(r2, 27u); // 49 rotl
|
||||
r1 = r1 ^ r7; // 50 xor
|
||||
r6 = r6 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xbc8977b1u : 0xaa5a26d7u); // 51 add
|
||||
r0 = r0 ^ r5; // 52 xor
|
||||
r1 = rotr_var(r1, r5); // 53 rotr
|
||||
r6 = rotl_imm(r6, 24u); // 54 rotl
|
||||
r1 = r1 ^ r5; // 55 xor
|
||||
r6 = r6 ^ r7; // 56 xor
|
||||
r5 = r5 ^ r7; // 57 xor
|
||||
r2 = r2 ^ r5; // 58 xor
|
||||
r0 = r4 * r0 + r0; // 59 mad
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r6, 8u); r7 = r7 ^ t_; } // 60 shfl
|
||||
r6 = r6 ^ ds[r5 & mask]; // 61 load
|
||||
{ uint t_; IGNEUM_SHFL_XOR(t_, r3, 2u); r7 = r7 ^ t_; } // 62 shfl
|
||||
r5 = r5 | r2; // 63 or
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
123
proto-cuda/packs/igneum-hourly/kernel_bound.cu
Normal file
123
proto-cuda/packs/igneum-hourly/kernel_bound.cu
Normal file
|
|
@ -0,0 +1,123 @@
|
|||
// Generated by igneum-pow export (generator v2) for seed "igneum-hourly". Do not edit by hand.
|
||||
// Header-bound twin of igneum_hash in kernel.cu: the init words come from a kernel argument, not SEEDW.
|
||||
// Host declarations (also in program_bound.h if present):
|
||||
// struct IgneumInitWords { uint32_t w[8]; };
|
||||
// cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
// IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps);
|
||||
// cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
|
||||
#include <cuda_runtime.h>
|
||||
#include <cstdint>
|
||||
#include "program.h"
|
||||
|
||||
struct IgneumInitWords { uint32_t w[8]; };
|
||||
|
||||
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
|
||||
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
|
||||
__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw) {
|
||||
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
uint32_t nonce = baseNonce + gid;
|
||||
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint32_t x = nonce ^ iw.w[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw.w[1]; }
|
||||
{ uint32_t x = nonce ^ iw.w[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw.w[2]; }
|
||||
{ uint32_t x = nonce ^ iw.w[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw.w[3]; }
|
||||
{ uint32_t x = nonce ^ iw.w[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw.w[4]; }
|
||||
{ uint32_t x = nonce ^ iw.w[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw.w[5]; }
|
||||
{ uint32_t x = nonce ^ iw.w[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw.w[6]; }
|
||||
{ uint32_t x = nonce ^ iw.w[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw.w[7]; }
|
||||
{ uint32_t x = nonce ^ iw.w[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw.w[0]; }
|
||||
|
||||
for (uint32_t it = 0u; it < 8u; ++it) {
|
||||
uint32_t sel = r0;
|
||||
r5 = __umulhi(r5, r0); // 0 mulhi
|
||||
r6 = r6 ^ ds[r5 & mask]; // 1 load
|
||||
r0 = r0 ^ ds[r6 & mask]; // 2 load
|
||||
r5 = r5 ^ ds[r0 & mask]; // 3 load
|
||||
r3 = r3 - r0; // 4 sub
|
||||
r6 = r5 * r4 + r6; // 5 mad
|
||||
r7 = r7 * r2; // 6 mul
|
||||
r5 = r5 + r1 + ((((sel >> 1u) & 1u) != 0u) ? 0xc4382c99u : 0x99d561fcu); // 7 add
|
||||
r1 = r1 ^ ds[r7 & mask]; // 8 load
|
||||
r5 = r5 ^ ds[r1 & mask]; // 9 load
|
||||
r0 = r6 * r2 + r0; // 10 mad
|
||||
r6 = __umulhi(r6, r0); // 11 mulhi
|
||||
r2 = rotl_imm(r2, 25u); // 12 rotl
|
||||
r5 = r5 ^ ds[r0 & mask]; // 13 load
|
||||
r2 = __umulhi(r2, r1); // 14 mulhi
|
||||
r7 = r7 + r4 + ((((sel >> 17u) & 1u) != 0u) ? 0x4ab35569u : 0xa105b846u); // 15 add
|
||||
r1 = r1 ^ ds[r6 & mask]; // 16 load
|
||||
r7 = r7 ^ ds[r1 & mask]; // 17 load
|
||||
r2 = rotr_var(r2, r0); // 18 rotr
|
||||
r7 = r7 ^ r3; // 19 xor
|
||||
r1 = __umulhi(r1, r6); // 20 mulhi
|
||||
r3 = r3 ^ ds[r5 & mask]; // 21 load
|
||||
r6 = r6 ^ r3; // 22 xor
|
||||
r0 = r0 * r3; // 23 mul
|
||||
r4 = r4 + r0 + ((((sel >> 2u) & 1u) != 0u) ? 0x344a9ec0u : 0x3706948au); // 24 add
|
||||
r7 = rotl_imm(r7, 24u); // 25 rotl
|
||||
r3 = r3 ^ r4; // 26 xor
|
||||
r2 = r2 ^ r0; // 27 xor
|
||||
r0 = r0 ^ r7; // 28 xor
|
||||
r3 = r3 - r1; // 29 sub
|
||||
r5 = r5 ^ ds[r7 & mask]; // 30 load
|
||||
r0 = r0 ^ ds[r2 & mask]; // 31 load
|
||||
r0 = r1 * r7 + r0; // 32 mad
|
||||
r5 = r5 ^ ds[r4 & mask]; // 33 load
|
||||
r6 = r6 ^ r0; // 34 xor
|
||||
r1 = r1 ^ ds[r0 & mask]; // 35 load
|
||||
r1 = r0 * r2 + r1; // 36 mad
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 37 shfl
|
||||
r6 = rotl_imm(r6, 21u); // 38 rotl
|
||||
r5 = r5 ^ r4; // 39 xor
|
||||
r5 = r1 * r1 + r5; // 40 mad
|
||||
r3 = r3 + r4 + ((((sel >> 17u) & 1u) != 0u) ? 0x7116ab79u : 0x2f1426d3u); // 41 add
|
||||
r7 = r7 - r1; // 42 sub
|
||||
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 43 shfl
|
||||
r6 = rotr_var(r6, r3); // 44 rotr
|
||||
r1 = rotl_imm(r1, 18u); // 45 rotl
|
||||
r1 = r1 ^ ds[r3 & mask]; // 46 load
|
||||
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r3, 8); // 47 shfl
|
||||
r0 = r0 ^ ds[r7 & mask]; // 48 load
|
||||
r2 = rotl_imm(r2, 27u); // 49 rotl
|
||||
r1 = r1 ^ r7; // 50 xor
|
||||
r6 = r6 + r2 + ((((sel >> 16u) & 1u) != 0u) ? 0xbc8977b1u : 0xaa5a26d7u); // 51 add
|
||||
r0 = r0 ^ r5; // 52 xor
|
||||
r1 = rotr_var(r1, r5); // 53 rotr
|
||||
r6 = rotl_imm(r6, 24u); // 54 rotl
|
||||
r1 = r1 ^ r5; // 55 xor
|
||||
r6 = r6 ^ r7; // 56 xor
|
||||
r5 = r5 ^ r7; // 57 xor
|
||||
r2 = r2 ^ r5; // 58 xor
|
||||
r0 = r4 * r0 + r0; // 59 mad
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r6, 8); // 60 shfl
|
||||
r6 = r6 ^ ds[r5 & mask]; // 61 load
|
||||
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 2); // 62 shfl
|
||||
r5 = r5 | r2; // 63 or
|
||||
}
|
||||
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
|
||||
}
|
||||
|
||||
cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
|
||||
IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps) {
|
||||
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
|
||||
uint32_t block = 32u * blockWarps;
|
||||
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
|
||||
igneum_hash_bound<<<nonces / block, block>>>(ds, out, baseNonce, mask, iw);
|
||||
return cudaGetLastError();
|
||||
}
|
||||
|
||||
cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
|
||||
cudaFuncAttributes attr;
|
||||
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash_bound);
|
||||
if (e != cudaSuccess) return e;
|
||||
*numRegs = attr.numRegs;
|
||||
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash_bound, (int)(32u * blockWarps), 0);
|
||||
}
|
||||
|
|
@ -1,4 +1,4 @@
|
|||
// Generated by proto-metal/igneum-bench --export-pack for seed "igneum-hourly". Do not edit by hand.
|
||||
// Generated by igneum-pow export (generator v2) for seed "igneum-hourly". Do not edit by hand.
|
||||
// Program metadata for host.cu plus the launch wrappers defined in kernel.cu.
|
||||
// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros.
|
||||
#pragma once
|
||||
|
|
@ -12,7 +12,12 @@
|
|||
#endif
|
||||
|
||||
#define IGNEUM_SEED_STRING "igneum-hourly"
|
||||
#define IGNEUM_SEED_BYTES_HEX "69676e65756d2d686f75726c79"
|
||||
#define IGNEUM_GENERATOR 2
|
||||
#define IGNEUM_PROGRAM_ATTEMPT 0
|
||||
#define IGNEUM_PROGRAM_ID 0xa4c4d00961c855dfull
|
||||
#define IGNEUM_DAY_STRING "2026-10-03"
|
||||
#define IGNEUM_DAY_BYTES_HEX "6461792f323032362d31302d3033"
|
||||
#define IGNEUM_DAY0 0x3067619fu
|
||||
#define IGNEUM_DAY1 0x3c269176u
|
||||
#define IGNEUM_DATASET_LOG2 28
|
||||
|
|
@ -22,7 +27,7 @@
|
|||
#define IGNEUM_INSTR_COUNT 64
|
||||
#define IGNEUM_LOADS_PER_HASH 128
|
||||
#define IGNEUM_WIDE_LOADS_PER_HASH 0
|
||||
#define IGNEUM_OP_MIX "load=16 add=8 xor=7 mad=5 mul=5 mulhi=5 shfl=5 rotl=4 or=3 rotr=3 sub=3"
|
||||
#define IGNEUM_OP_MIX "load=16 xor=13 mad=6 rotl=6 add=5 shfl=5 mulhi=4 rotr=3 sub=3 mul=2 or=1"
|
||||
// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h)
|
||||
#define IGNEUM_DATASET_MODE 0
|
||||
|
||||
|
|
|
|||
|
|
@ -1,15 +1,21 @@
|
|||
{
|
||||
"format": "igneum-program-pack-2",
|
||||
"format": "igneum-program-pack-3",
|
||||
"generator": 2,
|
||||
"attempt": 0,
|
||||
"program_id": "0xa4c4d00961c855df",
|
||||
"program_id_derivation": "FNV-1a 64 over 'igneum-program/' || generator_le32 || seed_words as little-endian bytes || attempt_le32",
|
||||
"dataset_mode": "closed-form",
|
||||
"seed": "igneum-hourly",
|
||||
"seed_bytes": "69676e65756d2d686f75726c79",
|
||||
"seed_words": ["0x6bdee811", "0x8f488bbe", "0xc5cdece7", "0x210af22d", "0x2f687b65", "0x17471eee", "0xee16e284", "0xfc9eb8f9"],
|
||||
"seed_derivation": "FNV-1a 64 over UTF-8 of seed, basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32",
|
||||
"seed_derivation": "seed_words = FNV-1a 64 over seed_bytes (attempt 0) or seed_bytes || attempt_le32 (attempt k >= 1), basis ^ (salt * 0x9E3779B97F4A7C15) for salt 0..3, then h ^= h>>33; h *= 0xff51afd7ed558ccd; h ^= h>>33; words[2*salt] = low 32, words[2*salt+1] = high 32",
|
||||
"generator_rule": "version 2: exactly 16 load slots drawn first from instructions 1..63 (partial Fisher-Yates), the other 48 ops from the ten non-load weights (sum 75); a load's source is drawn from the registers other than dst written by an earlier instruction and not read by a load since; the candidate must pass the acceptance rule of spec 01 section 1.4.6 (static: no cyclically stale load source, every register has an injecting write; dynamic: 64 units on the seed-keyed closed-form dataset with no constant register bit, no lane-constant load site, under 164 saturated final values, every output bit within 136 of 1024, distinct addresses above 245760), else the next attempt of the seed is tried",
|
||||
"lanes": 32,
|
||||
"registers": 8,
|
||||
"iterations": 8,
|
||||
"instruction_count": 64,
|
||||
"loads_per_hash": 128,
|
||||
"op_mix": {"load": 16, "add": 8, "xor": 7, "mad": 5, "mul": 5, "mulhi": 5, "shfl": 5, "rotl": 4, "or": 3, "rotr": 3, "sub": 3},
|
||||
"op_mix": {"load": 16, "xor": 13, "mad": 6, "rotl": 6, "add": 5, "shfl": 5, "mulhi": 4, "rotr": 3, "sub": 3, "mul": 2, "or": 1},
|
||||
"register_init": "for i in 0..7: x = nonce ^ seed_words[i]; x += 0x9e3779b9 * (i+1) (mod 2^32); x = splitmix32(x); r[i] = x ^ seed_words[(i+1) & 7]",
|
||||
"splitmix32": "x ^= x>>16; x *= 0x7feb352d; x ^= x>>15; x *= 0x846ca68b; x ^= x>>16",
|
||||
"iteration": "sel = r0 sampled once at the top of each iteration, then all instructions in order",
|
||||
|
|
@ -33,76 +39,77 @@
|
|||
"bytes": 1073741824,
|
||||
"mask": "0x0fffffff",
|
||||
"day": "2026-10-03",
|
||||
"day_words_from": "day/2026-10-03",
|
||||
"day_bytes": "6461792f323032362d31302d3033",
|
||||
"day_words_from": "seed_words_from_bytes(day_bytes)",
|
||||
"d0": "0x3067619f",
|
||||
"d1": "0x3c269176",
|
||||
"mode": "closed-form",
|
||||
"formula": "x = i ^ d0; x *= 0x9E3779B1; x ^= x>>15; x += d1; x *= 0x85EBCA77; x ^= x>>13; x *= 0xC2B2AE3D; x ^= x>>16 (all mod 2^32)"
|
||||
},
|
||||
"instructions": [
|
||||
{"i": 0, "op": "load", "dst": 6, "src": 5, "src2": 1, "imm": "0xf9ca5168", "imm2": "0xd27a709a", "rot": 1, "bit": 24, "mask": 16},
|
||||
{"i": 1, "op": "xor", "dst": 3, "src": 7, "src2": 0, "imm": "0x4731bc24", "imm2": "0x8c57bc59", "rot": 23, "bit": 3, "mask": 16},
|
||||
{"i": 2, "op": "mulhi", "dst": 6, "src": 2, "src2": 7, "imm": "0xc05a41e5", "imm2": "0xb5d20e75", "rot": 3, "bit": 7, "mask": 1},
|
||||
{"i": 3, "op": "add", "dst": 1, "src": 0, "src2": 0, "imm": "0x1eb46b1c", "imm2": "0x43f8f369", "rot": 31, "bit": 0, "mask": 8},
|
||||
{"i": 4, "op": "load", "dst": 3, "src": 0, "src2": 0, "imm": "0x8b6ea7af", "imm2": "0xa02a6965", "rot": 10, "bit": 22, "mask": 1},
|
||||
{"i": 5, "op": "or", "dst": 5, "src": 7, "src2": 7, "imm": "0x440ba791", "imm2": "0xf46d731a", "rot": 20, "bit": 14, "mask": 4},
|
||||
{"i": 6, "op": "load", "dst": 4, "src": 6, "src2": 7, "imm": "0xc7fa7a14", "imm2": "0xf0426f08", "rot": 16, "bit": 11, "mask": 4},
|
||||
{"i": 7, "op": "rotl", "dst": 4, "src": 3, "src2": 3, "imm": "0x535a367f", "imm2": "0x1d053f10", "rot": 21, "bit": 0, "mask": 4},
|
||||
{"i": 8, "op": "load", "dst": 6, "src": 3, "src2": 2, "imm": "0x0407ee13", "imm2": "0x691717fb", "rot": 20, "bit": 27, "mask": 8},
|
||||
{"i": 9, "op": "load", "dst": 6, "src": 1, "src2": 1, "imm": "0x8735ec1e", "imm2": "0xea050d61", "rot": 12, "bit": 26, "mask": 4},
|
||||
{"i": 10, "op": "shfl", "dst": 0, "src": 3, "src2": 2, "imm": "0x9bdf14a9", "imm2": "0xba78183e", "rot": 11, "bit": 20, "mask": 8},
|
||||
{"i": 11, "op": "xor", "dst": 2, "src": 3, "src2": 7, "imm": "0xf2010dad", "imm2": "0x52b64f96", "rot": 8, "bit": 0, "mask": 4},
|
||||
{"i": 12, "op": "add", "dst": 2, "src": 7, "src2": 1, "imm": "0x3a1ce85e", "imm2": "0xc26c7c2a", "rot": 10, "bit": 19, "mask": 4},
|
||||
{"i": 13, "op": "load", "dst": 4, "src": 3, "src2": 1, "imm": "0x816ece48", "imm2": "0xa0c6535a", "rot": 18, "bit": 21, "mask": 1},
|
||||
{"i": 14, "op": "load", "dst": 7, "src": 1, "src2": 3, "imm": "0x1b89fffb", "imm2": "0xf0330d16", "rot": 2, "bit": 0, "mask": 16},
|
||||
{"i": 15, "op": "mulhi", "dst": 2, "src": 3, "src2": 2, "imm": "0xcc6a5993", "imm2": "0xb853c4da", "rot": 28, "bit": 23, "mask": 4},
|
||||
{"i": 16, "op": "load", "dst": 5, "src": 2, "src2": 1, "imm": "0x14489190", "imm2": "0x4c71c842", "rot": 29, "bit": 7, "mask": 1},
|
||||
{"i": 17, "op": "load", "dst": 5, "src": 1, "src2": 1, "imm": "0x3ecd1efe", "imm2": "0xe582d071", "rot": 25, "bit": 29, "mask": 2},
|
||||
{"i": 18, "op": "load", "dst": 4, "src": 7, "src2": 6, "imm": "0xccfb0ae5", "imm2": "0xee788287", "rot": 3, "bit": 17, "mask": 16},
|
||||
{"i": 19, "op": "rotr", "dst": 2, "src": 1, "src2": 5, "imm": "0xc9ac3074", "imm2": "0x4d5eaabf", "rot": 19, "bit": 29, "mask": 16},
|
||||
{"i": 20, "op": "load", "dst": 7, "src": 0, "src2": 0, "imm": "0x637ba816", "imm2": "0xdeb9f971", "rot": 25, "bit": 2, "mask": 1},
|
||||
{"i": 21, "op": "load", "dst": 4, "src": 6, "src2": 4, "imm": "0x59d5f45d", "imm2": "0x701a90ad", "rot": 9, "bit": 2, "mask": 1},
|
||||
{"i": 22, "op": "rotl", "dst": 7, "src": 3, "src2": 5, "imm": "0xc6e87564", "imm2": "0x7aee3663", "rot": 25, "bit": 23, "mask": 4},
|
||||
{"i": 23, "op": "add", "dst": 3, "src": 5, "src2": 4, "imm": "0x942b819b", "imm2": "0x98ae0055", "rot": 8, "bit": 15, "mask": 2},
|
||||
{"i": 24, "op": "mulhi", "dst": 3, "src": 4, "src2": 1, "imm": "0x4b3f70bb", "imm2": "0x1c24bab9", "rot": 13, "bit": 18, "mask": 1},
|
||||
{"i": 25, "op": "xor", "dst": 6, "src": 1, "src2": 7, "imm": "0x82815069", "imm2": "0xecdc8c4c", "rot": 3, "bit": 6, "mask": 4},
|
||||
{"i": 26, "op": "rotl", "dst": 1, "src": 0, "src2": 0, "imm": "0x33471012", "imm2": "0x9ce1a3e2", "rot": 31, "bit": 7, "mask": 8},
|
||||
{"i": 27, "op": "shfl", "dst": 3, "src": 6, "src2": 6, "imm": "0x3a18a1b5", "imm2": "0xc49b3103", "rot": 18, "bit": 1, "mask": 4},
|
||||
{"i": 28, "op": "sub", "dst": 6, "src": 5, "src2": 1, "imm": "0x79e2cc72", "imm2": "0x4f9dab61", "rot": 5, "bit": 9, "mask": 16},
|
||||
{"i": 29, "op": "rotr", "dst": 6, "src": 3, "src2": 6, "imm": "0x5bea26e5", "imm2": "0x2aa1df19", "rot": 18, "bit": 14, "mask": 16},
|
||||
{"i": 30, "op": "shfl", "dst": 0, "src": 4, "src2": 7, "imm": "0xb954881e", "imm2": "0xa6eeddea", "rot": 1, "bit": 29, "mask": 4},
|
||||
{"i": 31, "op": "rotl", "dst": 4, "src": 5, "src2": 4, "imm": "0x46a02201", "imm2": "0x67e4f5b6", "rot": 30, "bit": 27, "mask": 4},
|
||||
{"i": 32, "op": "sub", "dst": 2, "src": 1, "src2": 6, "imm": "0x2b8c1fcd", "imm2": "0xf1e88c21", "rot": 9, "bit": 18, "mask": 8},
|
||||
{"i": 33, "op": "or", "dst": 5, "src": 4, "src2": 2, "imm": "0xb3fcf9dd", "imm2": "0x948436e1", "rot": 21, "bit": 4, "mask": 2},
|
||||
{"i": 34, "op": "mad", "dst": 7, "src": 6, "src2": 3, "imm": "0xc076bcd6", "imm2": "0x224748b2", "rot": 29, "bit": 9, "mask": 2},
|
||||
{"i": 35, "op": "mul", "dst": 5, "src": 0, "src2": 2, "imm": "0xd7f50971", "imm2": "0x4e7ada3e", "rot": 14, "bit": 12, "mask": 2},
|
||||
{"i": 36, "op": "sub", "dst": 5, "src": 3, "src2": 6, "imm": "0x349ca5cf", "imm2": "0x82b0e280", "rot": 15, "bit": 27, "mask": 2},
|
||||
{"i": 37, "op": "add", "dst": 2, "src": 7, "src2": 5, "imm": "0x473ecfd5", "imm2": "0x6b5970b5", "rot": 25, "bit": 5, "mask": 1},
|
||||
{"i": 38, "op": "load", "dst": 2, "src": 7, "src2": 1, "imm": "0xefc24111", "imm2": "0x1be7ad5b", "rot": 6, "bit": 3, "mask": 2},
|
||||
{"i": 39, "op": "rotr", "dst": 2, "src": 6, "src2": 4, "imm": "0x2eade1b0", "imm2": "0xe8e01fb0", "rot": 13, "bit": 30, "mask": 4},
|
||||
{"i": 40, "op": "xor", "dst": 0, "src": 5, "src2": 6, "imm": "0xdaf9a7a2", "imm2": "0x89165d3f", "rot": 1, "bit": 3, "mask": 1},
|
||||
{"i": 41, "op": "shfl", "dst": 4, "src": 5, "src2": 6, "imm": "0xd100247c", "imm2": "0x9697c51f", "rot": 27, "bit": 17, "mask": 8},
|
||||
{"i": 42, "op": "mul", "dst": 1, "src": 6, "src2": 7, "imm": "0xebbb5184", "imm2": "0xf79998f9", "rot": 15, "bit": 16, "mask": 4},
|
||||
{"i": 43, "op": "mad", "dst": 0, "src": 4, "src2": 1, "imm": "0x83ee60eb", "imm2": "0xbc017371", "rot": 26, "bit": 21, "mask": 2},
|
||||
{"i": 44, "op": "add", "dst": 1, "src": 4, "src2": 3, "imm": "0x04e78f3b", "imm2": "0x4a502c22", "rot": 20, "bit": 20, "mask": 4},
|
||||
{"i": 45, "op": "mad", "dst": 6, "src": 2, "src2": 7, "imm": "0xba1891ac", "imm2": "0xe7992a86", "rot": 27, "bit": 25, "mask": 2},
|
||||
{"i": 46, "op": "xor", "dst": 1, "src": 0, "src2": 1, "imm": "0xbd2c89d9", "imm2": "0x0bd3d33b", "rot": 29, "bit": 15, "mask": 8},
|
||||
{"i": 47, "op": "load", "dst": 5, "src": 7, "src2": 5, "imm": "0x67776361", "imm2": "0x20a4c205", "rot": 17, "bit": 13, "mask": 8},
|
||||
{"i": 48, "op": "or", "dst": 0, "src": 4, "src2": 1, "imm": "0x6b0e9ba2", "imm2": "0xf52a7752", "rot": 7, "bit": 21, "mask": 16},
|
||||
{"i": 49, "op": "shfl", "dst": 5, "src": 4, "src2": 2, "imm": "0x03494dfd", "imm2": "0x9c326606", "rot": 21, "bit": 22, "mask": 16},
|
||||
{"i": 50, "op": "add", "dst": 7, "src": 0, "src2": 2, "imm": "0x0a3df170", "imm2": "0x6fabf9ce", "rot": 5, "bit": 17, "mask": 16},
|
||||
{"i": 51, "op": "xor", "dst": 6, "src": 3, "src2": 6, "imm": "0x0c22f139", "imm2": "0xe642b0be", "rot": 27, "bit": 24, "mask": 1},
|
||||
{"i": 52, "op": "add", "dst": 1, "src": 2, "src2": 6, "imm": "0xc99bce6f", "imm2": "0xad344ca0", "rot": 31, "bit": 21, "mask": 2},
|
||||
{"i": 53, "op": "mul", "dst": 6, "src": 7, "src2": 1, "imm": "0x85e422cd", "imm2": "0x95145330", "rot": 27, "bit": 21, "mask": 1},
|
||||
{"i": 54, "op": "mulhi", "dst": 3, "src": 4, "src2": 4, "imm": "0xf61f4490", "imm2": "0x2e3345ba", "rot": 5, "bit": 14, "mask": 2},
|
||||
{"i": 55, "op": "mul", "dst": 7, "src": 1, "src2": 1, "imm": "0x2266ef41", "imm2": "0x81c0e541", "rot": 1, "bit": 27, "mask": 16},
|
||||
{"i": 56, "op": "xor", "dst": 7, "src": 1, "src2": 3, "imm": "0x5799668a", "imm2": "0x381b80ef", "rot": 13, "bit": 17, "mask": 8},
|
||||
{"i": 57, "op": "mul", "dst": 2, "src": 7, "src2": 2, "imm": "0x221abae1", "imm2": "0x8bf02d19", "rot": 31, "bit": 2, "mask": 8},
|
||||
{"i": 58, "op": "load", "dst": 2, "src": 1, "src2": 5, "imm": "0xf27b8255", "imm2": "0x820332d2", "rot": 14, "bit": 23, "mask": 16},
|
||||
{"i": 59, "op": "mad", "dst": 7, "src": 4, "src2": 5, "imm": "0xb6772f7b", "imm2": "0x07904b83", "rot": 14, "bit": 0, "mask": 8},
|
||||
{"i": 60, "op": "load", "dst": 2, "src": 7, "src2": 5, "imm": "0x6760e166", "imm2": "0x2e5a2418", "rot": 30, "bit": 12, "mask": 4},
|
||||
{"i": 61, "op": "mad", "dst": 0, "src": 2, "src2": 3, "imm": "0x3a2d60b7", "imm2": "0x4ec18be1", "rot": 15, "bit": 22, "mask": 1},
|
||||
{"i": 62, "op": "mulhi", "dst": 1, "src": 5, "src2": 6, "imm": "0xf423297d", "imm2": "0xd87f149c", "rot": 28, "bit": 6, "mask": 4},
|
||||
{"i": 63, "op": "add", "dst": 7, "src": 3, "src2": 3, "imm": "0xc23e27c9", "imm2": "0x05e4fc1d", "rot": 26, "bit": 24, "mask": 4}
|
||||
{"i": 0, "op": "mulhi", "dst": 5, "src": 0, "src2": 6, "imm": "0x34bfa8e1", "imm2": "0x88326f37", "rot": 14, "bit": 21, "mask": 8},
|
||||
{"i": 1, "op": "load", "dst": 6, "src": 5, "src2": 1, "imm": "0xe3e34b7c", "imm2": "0x17326c18", "rot": 22, "bit": 9, "mask": 1},
|
||||
{"i": 2, "op": "load", "dst": 0, "src": 6, "src2": 3, "imm": "0x83fd3087", "imm2": "0x10053a30", "rot": 16, "bit": 5, "mask": 16},
|
||||
{"i": 3, "op": "load", "dst": 5, "src": 0, "src2": 5, "imm": "0xf51ec9c2", "imm2": "0x72ec1e9f", "rot": 7, "bit": 26, "mask": 4},
|
||||
{"i": 4, "op": "sub", "dst": 3, "src": 0, "src2": 4, "imm": "0x70e50f1b", "imm2": "0x8efb1547", "rot": 16, "bit": 8, "mask": 16},
|
||||
{"i": 5, "op": "mad", "dst": 6, "src": 5, "src2": 4, "imm": "0x41861fa8", "imm2": "0x5faa64f3", "rot": 28, "bit": 16, "mask": 1},
|
||||
{"i": 6, "op": "mul", "dst": 7, "src": 2, "src2": 6, "imm": "0x0166e539", "imm2": "0x0490cd12", "rot": 24, "bit": 27, "mask": 8},
|
||||
{"i": 7, "op": "add", "dst": 5, "src": 1, "src2": 6, "imm": "0x99d561fc", "imm2": "0xc4382c99", "rot": 15, "bit": 1, "mask": 1},
|
||||
{"i": 8, "op": "load", "dst": 1, "src": 7, "src2": 0, "imm": "0x0c762bba", "imm2": "0x6bb9bf02", "rot": 10, "bit": 30, "mask": 2},
|
||||
{"i": 9, "op": "load", "dst": 5, "src": 1, "src2": 2, "imm": "0xee89ffa2", "imm2": "0x213faec7", "rot": 25, "bit": 22, "mask": 2},
|
||||
{"i": 10, "op": "mad", "dst": 0, "src": 6, "src2": 2, "imm": "0xdf34c691", "imm2": "0x59b3cdd1", "rot": 20, "bit": 10, "mask": 16},
|
||||
{"i": 11, "op": "mulhi", "dst": 6, "src": 0, "src2": 4, "imm": "0x767c0a8a", "imm2": "0x98d7a4b1", "rot": 11, "bit": 26, "mask": 2},
|
||||
{"i": 12, "op": "rotl", "dst": 2, "src": 5, "src2": 7, "imm": "0x66093200", "imm2": "0x16ea25f3", "rot": 25, "bit": 22, "mask": 8},
|
||||
{"i": 13, "op": "load", "dst": 5, "src": 0, "src2": 2, "imm": "0xe21e7a04", "imm2": "0xb457ced2", "rot": 1, "bit": 26, "mask": 2},
|
||||
{"i": 14, "op": "mulhi", "dst": 2, "src": 1, "src2": 5, "imm": "0xb2187445", "imm2": "0xb3fe3369", "rot": 27, "bit": 2, "mask": 16},
|
||||
{"i": 15, "op": "add", "dst": 7, "src": 4, "src2": 5, "imm": "0xa105b846", "imm2": "0x4ab35569", "rot": 22, "bit": 17, "mask": 8},
|
||||
{"i": 16, "op": "load", "dst": 1, "src": 6, "src2": 4, "imm": "0xa130c86d", "imm2": "0x34f71fb6", "rot": 1, "bit": 7, "mask": 8},
|
||||
{"i": 17, "op": "load", "dst": 7, "src": 1, "src2": 2, "imm": "0xc36b4a6c", "imm2": "0x0125b78d", "rot": 22, "bit": 31, "mask": 4},
|
||||
{"i": 18, "op": "rotr", "dst": 2, "src": 0, "src2": 7, "imm": "0x69f531b8", "imm2": "0x9a6effe0", "rot": 3, "bit": 17, "mask": 16},
|
||||
{"i": 19, "op": "xor", "dst": 7, "src": 3, "src2": 4, "imm": "0x43d25d5c", "imm2": "0xd8b7f1fc", "rot": 19, "bit": 13, "mask": 1},
|
||||
{"i": 20, "op": "mulhi", "dst": 1, "src": 6, "src2": 7, "imm": "0x6c2d2701", "imm2": "0x5d2f262d", "rot": 4, "bit": 3, "mask": 16},
|
||||
{"i": 21, "op": "load", "dst": 3, "src": 5, "src2": 3, "imm": "0x0aadba5b", "imm2": "0xad70ed9c", "rot": 24, "bit": 21, "mask": 8},
|
||||
{"i": 22, "op": "xor", "dst": 6, "src": 3, "src2": 3, "imm": "0xbc9f5d87", "imm2": "0xe341ffd9", "rot": 2, "bit": 25, "mask": 4},
|
||||
{"i": 23, "op": "mul", "dst": 0, "src": 3, "src2": 6, "imm": "0xbe314cd5", "imm2": "0xb3bbc40f", "rot": 20, "bit": 12, "mask": 16},
|
||||
{"i": 24, "op": "add", "dst": 4, "src": 0, "src2": 1, "imm": "0x3706948a", "imm2": "0x344a9ec0", "rot": 8, "bit": 2, "mask": 4},
|
||||
{"i": 25, "op": "rotl", "dst": 7, "src": 5, "src2": 3, "imm": "0x461c978c", "imm2": "0x66acfaee", "rot": 24, "bit": 3, "mask": 16},
|
||||
{"i": 26, "op": "xor", "dst": 3, "src": 4, "src2": 6, "imm": "0x48146102", "imm2": "0x54d979b9", "rot": 14, "bit": 1, "mask": 8},
|
||||
{"i": 27, "op": "xor", "dst": 2, "src": 0, "src2": 6, "imm": "0x6e08250b", "imm2": "0x72a4e086", "rot": 22, "bit": 25, "mask": 2},
|
||||
{"i": 28, "op": "xor", "dst": 0, "src": 7, "src2": 0, "imm": "0x07a2677d", "imm2": "0x6a26e517", "rot": 2, "bit": 10, "mask": 4},
|
||||
{"i": 29, "op": "sub", "dst": 3, "src": 1, "src2": 4, "imm": "0xb86c4c0a", "imm2": "0xf672488c", "rot": 12, "bit": 22, "mask": 4},
|
||||
{"i": 30, "op": "load", "dst": 5, "src": 7, "src2": 2, "imm": "0xe3c23946", "imm2": "0x7ef55fb6", "rot": 24, "bit": 1, "mask": 2},
|
||||
{"i": 31, "op": "load", "dst": 0, "src": 2, "src2": 5, "imm": "0x3e10f7d6", "imm2": "0x523b0052", "rot": 13, "bit": 1, "mask": 16},
|
||||
{"i": 32, "op": "mad", "dst": 0, "src": 1, "src2": 7, "imm": "0xe6e8ec65", "imm2": "0x6543d173", "rot": 14, "bit": 18, "mask": 2},
|
||||
{"i": 33, "op": "load", "dst": 5, "src": 4, "src2": 5, "imm": "0xfbca73bd", "imm2": "0x1efa415a", "rot": 2, "bit": 30, "mask": 4},
|
||||
{"i": 34, "op": "xor", "dst": 6, "src": 0, "src2": 5, "imm": "0xc910e7c3", "imm2": "0x472fd126", "rot": 23, "bit": 0, "mask": 1},
|
||||
{"i": 35, "op": "load", "dst": 1, "src": 0, "src2": 2, "imm": "0xb107b4b8", "imm2": "0xfb4a884d", "rot": 22, "bit": 21, "mask": 16},
|
||||
{"i": 36, "op": "mad", "dst": 1, "src": 0, "src2": 2, "imm": "0x62977314", "imm2": "0x0b6fc4c1", "rot": 27, "bit": 27, "mask": 8},
|
||||
{"i": 37, "op": "shfl", "dst": 7, "src": 4, "src2": 2, "imm": "0xe67ec3cd", "imm2": "0x98d67acc", "rot": 30, "bit": 16, "mask": 8},
|
||||
{"i": 38, "op": "rotl", "dst": 6, "src": 5, "src2": 0, "imm": "0xc04dc7a1", "imm2": "0x3e3f5816", "rot": 21, "bit": 31, "mask": 2},
|
||||
{"i": 39, "op": "xor", "dst": 5, "src": 4, "src2": 4, "imm": "0xbe7d7afb", "imm2": "0x62491b26", "rot": 27, "bit": 31, "mask": 4},
|
||||
{"i": 40, "op": "mad", "dst": 5, "src": 1, "src2": 1, "imm": "0x802dc394", "imm2": "0x17c3fe97", "rot": 1, "bit": 25, "mask": 16},
|
||||
{"i": 41, "op": "add", "dst": 3, "src": 4, "src2": 0, "imm": "0x2f1426d3", "imm2": "0x7116ab79", "rot": 3, "bit": 17, "mask": 4},
|
||||
{"i": 42, "op": "sub", "dst": 7, "src": 1, "src2": 1, "imm": "0xce63ca87", "imm2": "0xa98afd6b", "rot": 21, "bit": 2, "mask": 4},
|
||||
{"i": 43, "op": "shfl", "dst": 3, "src": 4, "src2": 6, "imm": "0xc6bbf6a6", "imm2": "0x56b2ce6f", "rot": 12, "bit": 6, "mask": 2},
|
||||
{"i": 44, "op": "rotr", "dst": 6, "src": 3, "src2": 1, "imm": "0xc5175172", "imm2": "0x5d6bdd29", "rot": 31, "bit": 27, "mask": 4},
|
||||
{"i": 45, "op": "rotl", "dst": 1, "src": 0, "src2": 5, "imm": "0xb8a931a2", "imm2": "0xc4e87b4d", "rot": 18, "bit": 5, "mask": 2},
|
||||
{"i": 46, "op": "load", "dst": 1, "src": 3, "src2": 0, "imm": "0x30cccb06", "imm2": "0xccc19d89", "rot": 29, "bit": 18, "mask": 4},
|
||||
{"i": 47, "op": "shfl", "dst": 5, "src": 3, "src2": 5, "imm": "0xec2b1deb", "imm2": "0x9079e102", "rot": 7, "bit": 6, "mask": 8},
|
||||
{"i": 48, "op": "load", "dst": 0, "src": 7, "src2": 7, "imm": "0x1d1461ac", "imm2": "0x8ceaba8a", "rot": 29, "bit": 14, "mask": 2},
|
||||
{"i": 49, "op": "rotl", "dst": 2, "src": 4, "src2": 6, "imm": "0x751e4f44", "imm2": "0x201122ae", "rot": 27, "bit": 30, "mask": 4},
|
||||
{"i": 50, "op": "xor", "dst": 1, "src": 7, "src2": 1, "imm": "0x4579666a", "imm2": "0xfeb1d0fe", "rot": 8, "bit": 0, "mask": 16},
|
||||
{"i": 51, "op": "add", "dst": 6, "src": 2, "src2": 6, "imm": "0xaa5a26d7", "imm2": "0xbc8977b1", "rot": 15, "bit": 16, "mask": 8},
|
||||
{"i": 52, "op": "xor", "dst": 0, "src": 5, "src2": 3, "imm": "0x3cfc50c3", "imm2": "0xbbc53f24", "rot": 16, "bit": 26, "mask": 1},
|
||||
{"i": 53, "op": "rotr", "dst": 1, "src": 5, "src2": 7, "imm": "0x72226bff", "imm2": "0xa0c9e939", "rot": 11, "bit": 1, "mask": 1},
|
||||
{"i": 54, "op": "rotl", "dst": 6, "src": 7, "src2": 7, "imm": "0x2a4e680d", "imm2": "0x7c6e0c13", "rot": 24, "bit": 15, "mask": 8},
|
||||
{"i": 55, "op": "xor", "dst": 1, "src": 5, "src2": 2, "imm": "0x8212fe2a", "imm2": "0x1ab25f32", "rot": 22, "bit": 25, "mask": 4},
|
||||
{"i": 56, "op": "xor", "dst": 6, "src": 7, "src2": 2, "imm": "0x6b16e101", "imm2": "0xdc4555d5", "rot": 3, "bit": 18, "mask": 8},
|
||||
{"i": 57, "op": "xor", "dst": 5, "src": 7, "src2": 7, "imm": "0x043cd22c", "imm2": "0x6a4cff6d", "rot": 22, "bit": 3, "mask": 1},
|
||||
{"i": 58, "op": "xor", "dst": 2, "src": 5, "src2": 2, "imm": "0xd3639dfd", "imm2": "0xfa54efad", "rot": 16, "bit": 24, "mask": 8},
|
||||
{"i": 59, "op": "mad", "dst": 0, "src": 4, "src2": 0, "imm": "0xa97cf54a", "imm2": "0xbbf10afb", "rot": 6, "bit": 1, "mask": 16},
|
||||
{"i": 60, "op": "shfl", "dst": 7, "src": 6, "src2": 1, "imm": "0x6b9a98e1", "imm2": "0x9c87ac06", "rot": 25, "bit": 28, "mask": 8},
|
||||
{"i": 61, "op": "load", "dst": 6, "src": 5, "src2": 7, "imm": "0xa6c37405", "imm2": "0x5937486b", "rot": 12, "bit": 29, "mask": 16},
|
||||
{"i": 62, "op": "shfl", "dst": 7, "src": 3, "src2": 7, "imm": "0x3045ddfc", "imm2": "0xf9d73ef3", "rot": 25, "bit": 16, "mask": 2},
|
||||
{"i": 63, "op": "or", "dst": 5, "src": 2, "src2": 4, "imm": "0x34a864e2", "imm2": "0xb5c6eca6", "rot": 14, "bit": 24, "mask": 1}
|
||||
]
|
||||
}
|
||||
|
|
|
|||
|
|
@ -38,70 +38,70 @@ kernel void igneum_hash(device const uint* dataset [[buffer(0)]],
|
|||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r6 = r6 ^ dataset[r5 & MASK]; // 0
|
||||
r3 = r3 ^ r7; // 1
|
||||
r6 = mulhi(r6, r2); // 2
|
||||
r1 = r1 + r0 + select(0x1eb46b1cu, 0x43f8f369u, ((sel >> 0u) & 1u) != 0u); // 3
|
||||
r3 = r3 ^ dataset[r0 & MASK]; // 4
|
||||
r5 = r5 | r7; // 5
|
||||
r4 = r4 ^ dataset[r6 & MASK]; // 6
|
||||
r4 = rotl_imm(r4, 21u); // 7
|
||||
r6 = r6 ^ dataset[r3 & MASK]; // 8
|
||||
r6 = r6 ^ dataset[r1 & MASK]; // 9
|
||||
r0 = r0 ^ simd_shuffle_xor(r3, (ushort)8); // 10
|
||||
r2 = r2 ^ r3; // 11
|
||||
r2 = r2 + r7 + select(0x3a1ce85eu, 0xc26c7c2au, ((sel >> 19u) & 1u) != 0u); // 12
|
||||
r4 = r4 ^ dataset[r3 & MASK]; // 13
|
||||
r7 = r7 ^ dataset[r1 & MASK]; // 14
|
||||
r2 = mulhi(r2, r3); // 15
|
||||
r5 = r5 ^ dataset[r2 & MASK]; // 16
|
||||
r5 = r5 ^ dataset[r1 & MASK]; // 17
|
||||
r4 = r4 ^ dataset[r7 & MASK]; // 18
|
||||
r2 = rotr_var(r2, r1); // 19
|
||||
r7 = r7 ^ dataset[r0 & MASK]; // 20
|
||||
r4 = r4 ^ dataset[r6 & MASK]; // 21
|
||||
r7 = rotl_imm(r7, 25u); // 22
|
||||
r3 = r3 + r5 + select(0x942b819bu, 0x98ae0055u, ((sel >> 15u) & 1u) != 0u); // 23
|
||||
r3 = mulhi(r3, r4); // 24
|
||||
r6 = r6 ^ r1; // 25
|
||||
r1 = rotl_imm(r1, 31u); // 26
|
||||
r3 = r3 ^ simd_shuffle_xor(r6, (ushort)4); // 27
|
||||
r6 = r6 - r5; // 28
|
||||
r6 = rotr_var(r6, r3); // 29
|
||||
r0 = r0 ^ simd_shuffle_xor(r4, (ushort)4); // 30
|
||||
r4 = rotl_imm(r4, 30u); // 31
|
||||
r2 = r2 - r1; // 32
|
||||
r5 = r5 | r4; // 33
|
||||
r7 = r6 * r3 + r7; // 34
|
||||
r5 = r5 * r0; // 35
|
||||
r5 = r5 - r3; // 36
|
||||
r2 = r2 + r7 + select(0x473ecfd5u, 0x6b5970b5u, ((sel >> 5u) & 1u) != 0u); // 37
|
||||
r2 = r2 ^ dataset[r7 & MASK]; // 38
|
||||
r2 = rotr_var(r2, r6); // 39
|
||||
r0 = r0 ^ r5; // 40
|
||||
r4 = r4 ^ simd_shuffle_xor(r5, (ushort)8); // 41
|
||||
r1 = r1 * r6; // 42
|
||||
r0 = r4 * r1 + r0; // 43
|
||||
r1 = r1 + r4 + select(0x04e78f3bu, 0x4a502c22u, ((sel >> 20u) & 1u) != 0u); // 44
|
||||
r6 = r2 * r7 + r6; // 45
|
||||
r1 = r1 ^ r0; // 46
|
||||
r5 = r5 ^ dataset[r7 & MASK]; // 47
|
||||
r0 = r0 | r4; // 48
|
||||
r5 = r5 ^ simd_shuffle_xor(r4, (ushort)16); // 49
|
||||
r7 = r7 + r0 + select(0x0a3df170u, 0x6fabf9ceu, ((sel >> 17u) & 1u) != 0u); // 50
|
||||
r6 = r6 ^ r3; // 51
|
||||
r1 = r1 + r2 + select(0xc99bce6fu, 0xad344ca0u, ((sel >> 21u) & 1u) != 0u); // 52
|
||||
r6 = r6 * r7; // 53
|
||||
r3 = mulhi(r3, r4); // 54
|
||||
r7 = r7 * r1; // 55
|
||||
r7 = r7 ^ r1; // 56
|
||||
r2 = r2 * r7; // 57
|
||||
r2 = r2 ^ dataset[r1 & MASK]; // 58
|
||||
r7 = r4 * r5 + r7; // 59
|
||||
r2 = r2 ^ dataset[r7 & MASK]; // 60
|
||||
r0 = r2 * r3 + r0; // 61
|
||||
r1 = mulhi(r1, r5); // 62
|
||||
r7 = r7 + r3 + select(0xc23e27c9u, 0x05e4fc1du, ((sel >> 24u) & 1u) != 0u); // 63
|
||||
r5 = mulhi(r5, r0); // 0
|
||||
r6 = r6 ^ dataset[r5 & MASK]; // 1
|
||||
r0 = r0 ^ dataset[r6 & MASK]; // 2
|
||||
r5 = r5 ^ dataset[r0 & MASK]; // 3
|
||||
r3 = r3 - r0; // 4
|
||||
r6 = r5 * r4 + r6; // 5
|
||||
r7 = r7 * r2; // 6
|
||||
r5 = r5 + r1 + select(0x99d561fcu, 0xc4382c99u, ((sel >> 1u) & 1u) != 0u); // 7
|
||||
r1 = r1 ^ dataset[r7 & MASK]; // 8
|
||||
r5 = r5 ^ dataset[r1 & MASK]; // 9
|
||||
r0 = r6 * r2 + r0; // 10
|
||||
r6 = mulhi(r6, r0); // 11
|
||||
r2 = rotl_imm(r2, 25u); // 12
|
||||
r5 = r5 ^ dataset[r0 & MASK]; // 13
|
||||
r2 = mulhi(r2, r1); // 14
|
||||
r7 = r7 + r4 + select(0xa105b846u, 0x4ab35569u, ((sel >> 17u) & 1u) != 0u); // 15
|
||||
r1 = r1 ^ dataset[r6 & MASK]; // 16
|
||||
r7 = r7 ^ dataset[r1 & MASK]; // 17
|
||||
r2 = rotr_var(r2, r0); // 18
|
||||
r7 = r7 ^ r3; // 19
|
||||
r1 = mulhi(r1, r6); // 20
|
||||
r3 = r3 ^ dataset[r5 & MASK]; // 21
|
||||
r6 = r6 ^ r3; // 22
|
||||
r0 = r0 * r3; // 23
|
||||
r4 = r4 + r0 + select(0x3706948au, 0x344a9ec0u, ((sel >> 2u) & 1u) != 0u); // 24
|
||||
r7 = rotl_imm(r7, 24u); // 25
|
||||
r3 = r3 ^ r4; // 26
|
||||
r2 = r2 ^ r0; // 27
|
||||
r0 = r0 ^ r7; // 28
|
||||
r3 = r3 - r1; // 29
|
||||
r5 = r5 ^ dataset[r7 & MASK]; // 30
|
||||
r0 = r0 ^ dataset[r2 & MASK]; // 31
|
||||
r0 = r1 * r7 + r0; // 32
|
||||
r5 = r5 ^ dataset[r4 & MASK]; // 33
|
||||
r6 = r6 ^ r0; // 34
|
||||
r1 = r1 ^ dataset[r0 & MASK]; // 35
|
||||
r1 = r0 * r2 + r1; // 36
|
||||
r7 = r7 ^ simd_shuffle_xor(r4, (ushort)8); // 37
|
||||
r6 = rotl_imm(r6, 21u); // 38
|
||||
r5 = r5 ^ r4; // 39
|
||||
r5 = r1 * r1 + r5; // 40
|
||||
r3 = r3 + r4 + select(0x2f1426d3u, 0x7116ab79u, ((sel >> 17u) & 1u) != 0u); // 41
|
||||
r7 = r7 - r1; // 42
|
||||
r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 43
|
||||
r6 = rotr_var(r6, r3); // 44
|
||||
r1 = rotl_imm(r1, 18u); // 45
|
||||
r1 = r1 ^ dataset[r3 & MASK]; // 46
|
||||
r5 = r5 ^ simd_shuffle_xor(r3, (ushort)8); // 47
|
||||
r0 = r0 ^ dataset[r7 & MASK]; // 48
|
||||
r2 = rotl_imm(r2, 27u); // 49
|
||||
r1 = r1 ^ r7; // 50
|
||||
r6 = r6 + r2 + select(0xaa5a26d7u, 0xbc8977b1u, ((sel >> 16u) & 1u) != 0u); // 51
|
||||
r0 = r0 ^ r5; // 52
|
||||
r1 = rotr_var(r1, r5); // 53
|
||||
r6 = rotl_imm(r6, 24u); // 54
|
||||
r1 = r1 ^ r5; // 55
|
||||
r6 = r6 ^ r7; // 56
|
||||
r5 = r5 ^ r7; // 57
|
||||
r2 = r2 ^ r5; // 58
|
||||
r0 = r4 * r0 + r0; // 59
|
||||
r7 = r7 ^ simd_shuffle_xor(r6, (ushort)8); // 60
|
||||
r6 = r6 ^ dataset[r5 & MASK]; // 61
|
||||
r7 = r7 ^ simd_shuffle_xor(r3, (ushort)2); // 62
|
||||
r5 = r5 | r2; // 63
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
|
|
|
|||
111
proto-cuda/packs/igneum-hourly/program_bound.metal
Normal file
111
proto-cuda/packs/igneum-hourly/program_bound.metal
Normal file
|
|
@ -0,0 +1,111 @@
|
|||
#include <metal_stdlib>
|
||||
using namespace metal;
|
||||
|
||||
#define MASK 0x0fffffffu
|
||||
constant uint SEEDW[8] = { 0x6bdee811u, 0x8f488bbeu, 0xc5cdece7u, 0x210af22du, 0x2f687b65u, 0x17471eeeu, 0xee16e284u, 0xfc9eb8f9u };
|
||||
|
||||
inline uint splitmix32(uint x) {
|
||||
x ^= x >> 16; x *= 0x7feb352du;
|
||||
x ^= x >> 15; x *= 0x846ca68bu;
|
||||
x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31
|
||||
inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
|
||||
inline uint ds_elem(uint i, uint d0, uint d1) {
|
||||
uint x = i ^ d0;
|
||||
x *= 0x9E3779B1u; x ^= x >> 15;
|
||||
x += d1;
|
||||
x *= 0x85EBCA77u; x ^= x >> 13;
|
||||
x *= 0xC2B2AE3Du; x ^= x >> 16;
|
||||
return x;
|
||||
}
|
||||
|
||||
// Header-bound variant: the init words come from buffer 3 (bind.rs), not from SEEDW.
|
||||
kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]],
|
||||
device ulong* out [[buffer(1)]],
|
||||
constant uint& baseNonce [[buffer(2)]],
|
||||
constant uint* initw [[buffer(3)]],
|
||||
uint gid [[thread_position_in_grid]]) {
|
||||
uint nonce = baseNonce + gid;
|
||||
uint r0, r1, r2, r3, r4, r5, r6, r7;
|
||||
{ uint x = nonce ^ initw[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ initw[1]; }
|
||||
{ uint x = nonce ^ initw[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ initw[2]; }
|
||||
{ uint x = nonce ^ initw[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ initw[3]; }
|
||||
{ uint x = nonce ^ initw[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ initw[4]; }
|
||||
{ uint x = nonce ^ initw[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ initw[5]; }
|
||||
{ uint x = nonce ^ initw[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ initw[6]; }
|
||||
{ uint x = nonce ^ initw[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ initw[7]; }
|
||||
{ uint x = nonce ^ initw[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ initw[0]; }
|
||||
|
||||
for (uint it = 0u; it < 8u; ++it) {
|
||||
uint sel = r0;
|
||||
r5 = mulhi(r5, r0); // 0
|
||||
r6 = r6 ^ dataset[r5 & MASK]; // 1
|
||||
r0 = r0 ^ dataset[r6 & MASK]; // 2
|
||||
r5 = r5 ^ dataset[r0 & MASK]; // 3
|
||||
r3 = r3 - r0; // 4
|
||||
r6 = r5 * r4 + r6; // 5
|
||||
r7 = r7 * r2; // 6
|
||||
r5 = r5 + r1 + select(0x99d561fcu, 0xc4382c99u, ((sel >> 1u) & 1u) != 0u); // 7
|
||||
r1 = r1 ^ dataset[r7 & MASK]; // 8
|
||||
r5 = r5 ^ dataset[r1 & MASK]; // 9
|
||||
r0 = r6 * r2 + r0; // 10
|
||||
r6 = mulhi(r6, r0); // 11
|
||||
r2 = rotl_imm(r2, 25u); // 12
|
||||
r5 = r5 ^ dataset[r0 & MASK]; // 13
|
||||
r2 = mulhi(r2, r1); // 14
|
||||
r7 = r7 + r4 + select(0xa105b846u, 0x4ab35569u, ((sel >> 17u) & 1u) != 0u); // 15
|
||||
r1 = r1 ^ dataset[r6 & MASK]; // 16
|
||||
r7 = r7 ^ dataset[r1 & MASK]; // 17
|
||||
r2 = rotr_var(r2, r0); // 18
|
||||
r7 = r7 ^ r3; // 19
|
||||
r1 = mulhi(r1, r6); // 20
|
||||
r3 = r3 ^ dataset[r5 & MASK]; // 21
|
||||
r6 = r6 ^ r3; // 22
|
||||
r0 = r0 * r3; // 23
|
||||
r4 = r4 + r0 + select(0x3706948au, 0x344a9ec0u, ((sel >> 2u) & 1u) != 0u); // 24
|
||||
r7 = rotl_imm(r7, 24u); // 25
|
||||
r3 = r3 ^ r4; // 26
|
||||
r2 = r2 ^ r0; // 27
|
||||
r0 = r0 ^ r7; // 28
|
||||
r3 = r3 - r1; // 29
|
||||
r5 = r5 ^ dataset[r7 & MASK]; // 30
|
||||
r0 = r0 ^ dataset[r2 & MASK]; // 31
|
||||
r0 = r1 * r7 + r0; // 32
|
||||
r5 = r5 ^ dataset[r4 & MASK]; // 33
|
||||
r6 = r6 ^ r0; // 34
|
||||
r1 = r1 ^ dataset[r0 & MASK]; // 35
|
||||
r1 = r0 * r2 + r1; // 36
|
||||
r7 = r7 ^ simd_shuffle_xor(r4, (ushort)8); // 37
|
||||
r6 = rotl_imm(r6, 21u); // 38
|
||||
r5 = r5 ^ r4; // 39
|
||||
r5 = r1 * r1 + r5; // 40
|
||||
r3 = r3 + r4 + select(0x2f1426d3u, 0x7116ab79u, ((sel >> 17u) & 1u) != 0u); // 41
|
||||
r7 = r7 - r1; // 42
|
||||
r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 43
|
||||
r6 = rotr_var(r6, r3); // 44
|
||||
r1 = rotl_imm(r1, 18u); // 45
|
||||
r1 = r1 ^ dataset[r3 & MASK]; // 46
|
||||
r5 = r5 ^ simd_shuffle_xor(r3, (ushort)8); // 47
|
||||
r0 = r0 ^ dataset[r7 & MASK]; // 48
|
||||
r2 = rotl_imm(r2, 27u); // 49
|
||||
r1 = r1 ^ r7; // 50
|
||||
r6 = r6 + r2 + select(0xaa5a26d7u, 0xbc8977b1u, ((sel >> 16u) & 1u) != 0u); // 51
|
||||
r0 = r0 ^ r5; // 52
|
||||
r1 = rotr_var(r1, r5); // 53
|
||||
r6 = rotl_imm(r6, 24u); // 54
|
||||
r1 = r1 ^ r5; // 55
|
||||
r6 = r6 ^ r7; // 56
|
||||
r5 = r5 ^ r7; // 57
|
||||
r2 = r2 ^ r5; // 58
|
||||
r0 = r4 * r0 + r0; // 59
|
||||
r7 = r7 ^ simd_shuffle_xor(r6, (ushort)8); // 60
|
||||
r6 = r6 ^ dataset[r5 & MASK]; // 61
|
||||
r7 = r7 ^ simd_shuffle_xor(r3, (ushort)2); // 62
|
||||
r5 = r5 | r2; // 63
|
||||
}
|
||||
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
|
||||
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
|
||||
out[gid] = ((ulong)hi << 32) | (ulong)lo;
|
||||
}
|
||||
|
|
@ -1,5 +1,5 @@
|
|||
// Generated by proto-metal/igneum-bench --export-pack for seed "igneum-hourly". Do not edit by hand.
|
||||
// Expected outputs: proto-metal CPU interpreter (cpuWarp, closed-form dataset) on Apple M5 Max; Metal GPU cross-check PASS 3/3 warps
|
||||
// Generated by igneum-pow export (generator v2) for seed "igneum-hourly". Do not edit by hand.
|
||||
// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v2, closed-form dataset
|
||||
#pragma once
|
||||
#ifdef __cplusplus
|
||||
#include <cstdint>
|
||||
|
|
@ -11,22 +11,22 @@
|
|||
static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u };
|
||||
static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = {
|
||||
{ // base nonce 0
|
||||
0x787a737506455bbeull, 0xf93d248e364f7e1eull, 0x8a7ce1701b5d0489ull, 0xd4f5b2177c414f3full, 0xe88fdf6fb22aee5cull, 0xd68bc0baf30b41fcull, 0x87f6a81eb0a6ac69ull, 0xb3cd8b6e6e57b347ull,
|
||||
0x2eef97b916c626bfull, 0x4ded79caeecc1b81ull, 0xac5141672d56faaaull, 0xe5c20290edf36932ull, 0x8358bbb9c3fa668dull, 0x3394e6f0736ecd30ull, 0xef402d26f32ed7cfull, 0x0ce267a5c3e18a42ull,
|
||||
0xd92370857d7ec52bull, 0x4f8cd5815e18afbeull, 0x2fbf597253c4f35eull, 0x747180e71a09f68dull, 0x2238f847f0ed84dcull, 0xf9c277614e6b908dull, 0x39dbe66723750ad4ull, 0x8dc8b244ac675fbbull,
|
||||
0x02f535158b9d7c8eull, 0x71c6a2147fb4787aull, 0xcb3a0d709ec6edb6ull, 0x560cdae1ccc99cd7ull, 0xb8e12af183034f40ull, 0xe3088768bbf66638ull, 0x62b3109b60c3103dull, 0x0bb06c1609ff43d5ull
|
||||
0xed8c8c2a229082deull, 0x788bf8c95f117eeaull, 0x87eec7805f34f307ull, 0x020cc55434e23c37ull, 0x9053cd6b9650fad2ull, 0x38214f1197f8e18dull, 0xe11830340f9b1f9aull, 0x9c5d17d01f75bc80ull,
|
||||
0xe18df405d8dd5ca6ull, 0xfe65db408bedc853ull, 0xeade05871da382e1ull, 0x1e553492d1b7a9cbull, 0xaa5c278d24ea5b63ull, 0xf48beaf06cfb79ffull, 0x745e2d1aee8cbaf6ull, 0x9f7ee6e8f4760820ull,
|
||||
0xa6eed160478a5461ull, 0x49587ddce18631e7ull, 0xdfe6737f9ced9c0cull, 0x1fddcc0632181b10ull, 0x2416eb6b92c250f4ull, 0x401ed2c272df58b2ull, 0x39aababbc108f740ull, 0x645b9684389b26b4ull,
|
||||
0x710e62c990af3118ull, 0xa0ec7df1dbb7cf72ull, 0xc6924e913b427cc5ull, 0x3f944943d6bfabd9ull, 0x2c3bc5e317ff2b9cull, 0x156aec9a26a5e8acull, 0x8dbf6d080b1154a9ull, 0xdfd36edc4fe52cceull
|
||||
},
|
||||
{ // base nonce 4096
|
||||
0x37bb8728f884e6b8ull, 0x14be2ec14497f818ull, 0x2086e0b6e0d72b74ull, 0x50b1159a7a48b25aull, 0xc7e8af0d47ae9524ull, 0x11b6c5919a802722ull, 0x47191c35d09f473eull, 0x5a17776c88350cbcull,
|
||||
0xf8b63bb823a85140ull, 0x7d1536d3c3a77a4aull, 0x3fd4978fa97a3170ull, 0xb93a559eef3ca040ull, 0x3ab4c9b478dc4f70ull, 0xe383445b250ec1c5ull, 0x4b8c339a4059d0eaull, 0xb9edfbab431dc7dcull,
|
||||
0x67e73ea59e305c36ull, 0x85ad26bb831207a3ull, 0x7fa9ca66562cd371ull, 0x763b3ca6c6efdf75ull, 0x6f02716b104a47ceull, 0x033d54978e01388full, 0xed13f0a661b95976ull, 0x4cf586279408f177ull,
|
||||
0x339f105764d3a0dfull, 0xf5f5a600702c84e7ull, 0xbec42d884060c3ebull, 0x5e687d7d88c4dc86ull, 0xfe9202424271b313ull, 0xe9a93837f53818ddull, 0x2efe4cd7be3bc6ccull, 0x2718ca6898992e25ull
|
||||
0x48262c09eb666bb1ull, 0xfaae09746a766581ull, 0x9044c563fd782193ull, 0x38e1629374c81654ull, 0xc8c857db9859535dull, 0xfc48e1eb2395e951ull, 0x285a8bc2a1e83bccull, 0x968ada8c1dc2570bull,
|
||||
0x8035ace9b4fabe47ull, 0xedeb0b6bd0042439ull, 0xb37fdb28f075046dull, 0x3107cd1e294ddbdbull, 0xfd6048600aa346fbull, 0x838318fea1812cbeull, 0xd345b8c644f76facull, 0xf8c6982e60d73ccaull,
|
||||
0xbd406d8e3b1d93c6ull, 0xcea3caad3b50616dull, 0x30a5bf3ee3496035ull, 0xc1cc211ba9e892d1ull, 0x1e0e7863efb9c9e9ull, 0x37a3a64471cc0833ull, 0x0ceebf4ccb27a0feull, 0xa03cd804e19c1832ull,
|
||||
0x2fea21e917f689a6ull, 0x87e4826e41d3f7b9ull, 0x7a92c1436fc291eeull, 0xfc28387d8a5893d2ull, 0x911c0ceb9443ee66ull, 0xbf9fd4e3ebc60137ull, 0xe0a05dc0f781b49eull, 0xcbb408d3f084209dull
|
||||
},
|
||||
{ // base nonce 1000000
|
||||
0x3ea4d63013ee1447ull, 0x3920ed5486c12342ull, 0x6d71c6114a18c926ull, 0x76851391b25e9e27ull, 0x7552ab2e8ba173e4ull, 0x2ed7c0fe0ed382e7ull, 0x3965791801d021b2ull, 0x5638111a801de9f4ull,
|
||||
0xffdb8b13126c51efull, 0xec7513be841d1ee6ull, 0x5ee78112615a43a0ull, 0x447c921628b22b82ull, 0xa8683cc17b258e26ull, 0xc4161c3fb96be5f0ull, 0xf4bde10f52aef480ull, 0xfb9a90df739d4b7full,
|
||||
0x8b6a11be7a8067c6ull, 0x0c5e0b16f9dddd20ull, 0xad830d9f20c0cf41ull, 0x7b45608e60aa036cull, 0x871e0fc8a4c82071ull, 0x087bba8af99e5f89ull, 0xc9b353a02ccf042dull, 0xa8b5e4ca3e95f99eull,
|
||||
0x1f636af31962da8full, 0x4e3c7276699005ccull, 0xeb138ecf3dfdc723ull, 0x572fe00fa91ec1f0ull, 0x7eb7d1613efc2fb6ull, 0x4ec9b30dfb44aa0full, 0x8883fe27c6464e60ull, 0xca9d04210422e8b8ull
|
||||
0x5fdd97cdbef80859ull, 0xa6746d5b69e14862ull, 0xcf8de84cd68c481full, 0xc988b34d937f737bull, 0xaff10b57e563c5dfull, 0x204513cc06f71ebcull, 0xcfa09b6c2ddd7b1aull, 0x27fcc09db36e7ee1ull,
|
||||
0x84dfa831ccbd61d9ull, 0x0f24f4b0e2f26c9eull, 0x0d00c06606b83ac9ull, 0x43a1c3dc2d5261caull, 0xf888b28880490a49ull, 0xb31c71698c4bf8d3ull, 0x58c62b77fcfa2aa1ull, 0xbbc7af747001b89dull,
|
||||
0xa813af19ec466253ull, 0x237d24c85bd66221ull, 0x9feca7c7f10337e3ull, 0x23f45a840592f604ull, 0x61b395829f95b0bdull, 0x08241ec828e5cb7full, 0x57713b7b22caa588ull, 0x36db28ea3655a4c3ull,
|
||||
0x4f094dfebc197855ull, 0x75fd74c91fb5ed8dull, 0xcdcfca84af7e6235ull, 0x1436e2ebbda7fdc4ull, 0x0951a0f4dd09d993ull, 0x113c1fc5b35e8ed7ull, 0x5daf8ee6e0c45c56ull, 0xc94bcda16f821943ull
|
||||
}
|
||||
};
|
||||
|
||||
|
|
|
|||
|
|
@ -5,25 +5,25 @@
|
|||
"dataset_log2_words": 28,
|
||||
"mask": "0x0fffffff",
|
||||
"lanes": 32,
|
||||
"source": "proto-metal CPU interpreter (cpuWarp, closed-form dataset) on Apple M5 Max; Metal GPU cross-check PASS 3/3 warps",
|
||||
"source": "igneum-pow (Rust) CPU interpreter, generator v2, closed-form dataset",
|
||||
"warps": [
|
||||
{"base_nonce": 0, "expected": [
|
||||
"0x787a737506455bbe", "0xf93d248e364f7e1e", "0x8a7ce1701b5d0489", "0xd4f5b2177c414f3f", "0xe88fdf6fb22aee5c", "0xd68bc0baf30b41fc", "0x87f6a81eb0a6ac69", "0xb3cd8b6e6e57b347",
|
||||
"0x2eef97b916c626bf", "0x4ded79caeecc1b81", "0xac5141672d56faaa", "0xe5c20290edf36932", "0x8358bbb9c3fa668d", "0x3394e6f0736ecd30", "0xef402d26f32ed7cf", "0x0ce267a5c3e18a42",
|
||||
"0xd92370857d7ec52b", "0x4f8cd5815e18afbe", "0x2fbf597253c4f35e", "0x747180e71a09f68d", "0x2238f847f0ed84dc", "0xf9c277614e6b908d", "0x39dbe66723750ad4", "0x8dc8b244ac675fbb",
|
||||
"0x02f535158b9d7c8e", "0x71c6a2147fb4787a", "0xcb3a0d709ec6edb6", "0x560cdae1ccc99cd7", "0xb8e12af183034f40", "0xe3088768bbf66638", "0x62b3109b60c3103d", "0x0bb06c1609ff43d5"
|
||||
"0xed8c8c2a229082de", "0x788bf8c95f117eea", "0x87eec7805f34f307", "0x020cc55434e23c37", "0x9053cd6b9650fad2", "0x38214f1197f8e18d", "0xe11830340f9b1f9a", "0x9c5d17d01f75bc80",
|
||||
"0xe18df405d8dd5ca6", "0xfe65db408bedc853", "0xeade05871da382e1", "0x1e553492d1b7a9cb", "0xaa5c278d24ea5b63", "0xf48beaf06cfb79ff", "0x745e2d1aee8cbaf6", "0x9f7ee6e8f4760820",
|
||||
"0xa6eed160478a5461", "0x49587ddce18631e7", "0xdfe6737f9ced9c0c", "0x1fddcc0632181b10", "0x2416eb6b92c250f4", "0x401ed2c272df58b2", "0x39aababbc108f740", "0x645b9684389b26b4",
|
||||
"0x710e62c990af3118", "0xa0ec7df1dbb7cf72", "0xc6924e913b427cc5", "0x3f944943d6bfabd9", "0x2c3bc5e317ff2b9c", "0x156aec9a26a5e8ac", "0x8dbf6d080b1154a9", "0xdfd36edc4fe52cce"
|
||||
]},
|
||||
{"base_nonce": 4096, "expected": [
|
||||
"0x37bb8728f884e6b8", "0x14be2ec14497f818", "0x2086e0b6e0d72b74", "0x50b1159a7a48b25a", "0xc7e8af0d47ae9524", "0x11b6c5919a802722", "0x47191c35d09f473e", "0x5a17776c88350cbc",
|
||||
"0xf8b63bb823a85140", "0x7d1536d3c3a77a4a", "0x3fd4978fa97a3170", "0xb93a559eef3ca040", "0x3ab4c9b478dc4f70", "0xe383445b250ec1c5", "0x4b8c339a4059d0ea", "0xb9edfbab431dc7dc",
|
||||
"0x67e73ea59e305c36", "0x85ad26bb831207a3", "0x7fa9ca66562cd371", "0x763b3ca6c6efdf75", "0x6f02716b104a47ce", "0x033d54978e01388f", "0xed13f0a661b95976", "0x4cf586279408f177",
|
||||
"0x339f105764d3a0df", "0xf5f5a600702c84e7", "0xbec42d884060c3eb", "0x5e687d7d88c4dc86", "0xfe9202424271b313", "0xe9a93837f53818dd", "0x2efe4cd7be3bc6cc", "0x2718ca6898992e25"
|
||||
"0x48262c09eb666bb1", "0xfaae09746a766581", "0x9044c563fd782193", "0x38e1629374c81654", "0xc8c857db9859535d", "0xfc48e1eb2395e951", "0x285a8bc2a1e83bcc", "0x968ada8c1dc2570b",
|
||||
"0x8035ace9b4fabe47", "0xedeb0b6bd0042439", "0xb37fdb28f075046d", "0x3107cd1e294ddbdb", "0xfd6048600aa346fb", "0x838318fea1812cbe", "0xd345b8c644f76fac", "0xf8c6982e60d73cca",
|
||||
"0xbd406d8e3b1d93c6", "0xcea3caad3b50616d", "0x30a5bf3ee3496035", "0xc1cc211ba9e892d1", "0x1e0e7863efb9c9e9", "0x37a3a64471cc0833", "0x0ceebf4ccb27a0fe", "0xa03cd804e19c1832",
|
||||
"0x2fea21e917f689a6", "0x87e4826e41d3f7b9", "0x7a92c1436fc291ee", "0xfc28387d8a5893d2", "0x911c0ceb9443ee66", "0xbf9fd4e3ebc60137", "0xe0a05dc0f781b49e", "0xcbb408d3f084209d"
|
||||
]},
|
||||
{"base_nonce": 1000000, "expected": [
|
||||
"0x3ea4d63013ee1447", "0x3920ed5486c12342", "0x6d71c6114a18c926", "0x76851391b25e9e27", "0x7552ab2e8ba173e4", "0x2ed7c0fe0ed382e7", "0x3965791801d021b2", "0x5638111a801de9f4",
|
||||
"0xffdb8b13126c51ef", "0xec7513be841d1ee6", "0x5ee78112615a43a0", "0x447c921628b22b82", "0xa8683cc17b258e26", "0xc4161c3fb96be5f0", "0xf4bde10f52aef480", "0xfb9a90df739d4b7f",
|
||||
"0x8b6a11be7a8067c6", "0x0c5e0b16f9dddd20", "0xad830d9f20c0cf41", "0x7b45608e60aa036c", "0x871e0fc8a4c82071", "0x087bba8af99e5f89", "0xc9b353a02ccf042d", "0xa8b5e4ca3e95f99e",
|
||||
"0x1f636af31962da8f", "0x4e3c7276699005cc", "0xeb138ecf3dfdc723", "0x572fe00fa91ec1f0", "0x7eb7d1613efc2fb6", "0x4ec9b30dfb44aa0f", "0x8883fe27c6464e60", "0xca9d04210422e8b8"
|
||||
"0x5fdd97cdbef80859", "0xa6746d5b69e14862", "0xcf8de84cd68c481f", "0xc988b34d937f737b", "0xaff10b57e563c5df", "0x204513cc06f71ebc", "0xcfa09b6c2ddd7b1a", "0x27fcc09db36e7ee1",
|
||||
"0x84dfa831ccbd61d9", "0x0f24f4b0e2f26c9e", "0x0d00c06606b83ac9", "0x43a1c3dc2d5261ca", "0xf888b28880490a49", "0xb31c71698c4bf8d3", "0x58c62b77fcfa2aa1", "0xbbc7af747001b89d",
|
||||
"0xa813af19ec466253", "0x237d24c85bd66221", "0x9feca7c7f10337e3", "0x23f45a840592f604", "0x61b395829f95b0bd", "0x08241ec828e5cb7f", "0x57713b7b22caa588", "0x36db28ea3655a4c3",
|
||||
"0x4f094dfebc197855", "0x75fd74c91fb5ed8d", "0xcdcfca84af7e6235", "0x1436e2ebbda7fdc4", "0x0951a0f4dd09d993", "0x113c1fc5b35e8ed7", "0x5daf8ee6e0c45c56", "0xc94bcda16f821943"
|
||||
]}
|
||||
],
|
||||
"dataset_head": ["0x82174c0f", "0x577bdb9c", "0x111053bf", "0x2bb85514", "0xd83e190f", "0x8ccd9427", "0x3f69d5e4", "0xf3bbe1dc", "0xb3aa7e90", "0xc33f5f73", "0xb83a2b10", "0xc4d7c8ff", "0xefa6d1a8", "0x7029b116", "0x5e48fab0", "0x0ef66e40"],
|
||||
|
|
|
|||
|
|
@ -11,6 +11,13 @@ file are from the original closed-form dataset and reproduce with `--closed-form
|
|||
here was re-run on the new dataset with the same commands and passed; those results are in `MEMHARD.md` section 2.5.
|
||||
The shortcut of section 7 is answered there (section 2.2): the inline kernel is now 4.8x slower than the honest one.
|
||||
|
||||
On 4 October 2026 the generator became version 2 (exactly 16 loads per program, fresh-source loads, the acceptance
|
||||
rule of spec 01 section 1.4.6; `igneum-pow/src/{generator,accept}.rs`, ported into `main.swift` as
|
||||
`generateProgramV2`, `candidateProgram`, `acceptProgram`). Sections 1 to 7 below describe the version 1 generator and
|
||||
its vectors, which are retired; the loads-per-hash ranges they quote no longer occur. Section 9 holds the version 2
|
||||
run. The `--load-weight` and `--wide-frac` levers select the retired version 1 generator and exist only so
|
||||
`MEMHARD.md` section 2.4 reproduces.
|
||||
|
||||
The tests live in `main.swift` next to the bench and share its generator, MSL emitter, dataset build and CPU
|
||||
interpreter without modification. The CPU interpreter gained an optional trace hook (`cpuWarpTraced`) that the bench
|
||||
does not use; `cpuWarp` calls it with `nil`.
|
||||
|
|
@ -315,10 +322,47 @@ measurement is why it is not optional.
|
|||
edge-case programs. Then fix the hash as a specification with test vectors and replace the ad-hoc seed
|
||||
derivation (FNV-1a plus SplitMix) with a standard hash so the seed-to-program mapping is auditable.
|
||||
|
||||
## 9. Generator version 2 (4 October 2026)
|
||||
|
||||
Build: `swiftc -O -o igneum-bench main.swift -framework Metal` (7 s), same machine, idle (load 3).
|
||||
|
||||
Fuzz, the required run at a reduced size (the version 1 run was 10,000):
|
||||
|
||||
```
|
||||
./igneum-bench --fuzz 2000 --fuzz-seed igneum-fuzz-gen2-2026-10-04
|
||||
```
|
||||
|
||||
| Dataset | Programs | Pass | Fail |
|
||||
|---|---|---|---|
|
||||
| 2^24 words (64 MiB) | 647 | 647 | 0 |
|
||||
| 2^26 words (256 MiB) | 645 | 645 | 0 |
|
||||
| 2^28 words (1 GiB) | 708 | 708 | 0 |
|
||||
| all | 2000 | 2000 | 0 |
|
||||
|
||||
2,000 programs: pass 2,000, mismatch 0, compile failures 0, static mask failures 0, contract failures 0. 8,000 warps
|
||||
compared (256,000 hashes), memory-hard dataset. Loads per hash 128 to 128 (every program; version 1 ranged 40 to
|
||||
232). Op totals: load 32,000, add 15,350, xor 12,802, mad 10,380, mul 10,208, shfl 10,147, rotl 8,983, sub 7,704,
|
||||
rotr 7,632, mulhi 7,596, or 5,198. Compile min 15.7 ms, avg 21.8 ms, max 67.3 ms (all cold). GPU dispatch 1.7 s
|
||||
total, CPU interpreter 17.7 s total (the memory-hard dataset), wall 91.4 s. Every fuzzed program is the accepted
|
||||
attempt of its seed, so the acceptance rule ran 2,000 times inside this run as well. FUZZ: PASS.
|
||||
|
||||
Cross-implementation, the Swift generator against the Rust crate (`igneum-pow export` for the same seeds, closed
|
||||
form unless stated): instruction lists, seed words and all 96 vectors identical on `igneum-genesis`,
|
||||
`igneum-hourly`, and the three seeds whose attempt 0 is rejected and attempt 1 accepted
|
||||
(`igneum-census-2026-10-03/22`, `/37`, `/51`); `igneum-genesis` memory-hard: identical instruction list and 96
|
||||
vectors between the Swift export (Metal GPU cross-check PASS 3 of 3 warps, cache FNV `48c4f5bf24166b2e`) and the
|
||||
checked-in Rust pack `proto-cuda/packs/igneum-genesis-mh`. The Swift exporter is no longer the pack source; it is
|
||||
the Metal cross-check of the Rust packs (`proto-cuda/README.md`, "Regenerating a pack"), and its `program.json`
|
||||
still writes format `igneum-program-pack-2` without the generator fields.
|
||||
|
||||
Not re-run under version 2: sections 2 (edge, hand-built programs that bypass the generator), 3 (stats), 4
|
||||
(determinism), 5 (memcheck) and 7 (shortcut). Nothing in them depends on how the instruction list is drawn.
|
||||
|
||||
## Files
|
||||
|
||||
- `main.swift`: tests under `// MARK: - Hardening tests`, flags `--fuzz`, `--fuzz-seed`, `--edge`, `--stats`,
|
||||
`--determinism`, `--memcheck`, `--inline-dataset`.
|
||||
`--determinism`, `--memcheck`, `--inline-dataset`; generator version 2 and the acceptance rule under
|
||||
`// MARK: - Generator version 2 and the acceptance rule`.
|
||||
- `TESTS.md`: this file.
|
||||
- Raw logs of the runs quoted above were kept in the session scratchpad and are not checked in; every table
|
||||
is reproducible with the command above it.
|
||||
|
|
|
|||
|
|
@ -7,6 +7,13 @@ both Windows (Adrenalin driver) and Linux (ROCm, or Mesa), and checks every outp
|
|||
This is a test harness, not a miner. No pool, no network, no wallet, no mining protocol. It fills a dataset, checks the
|
||||
device against known answers, and times the kernel. Nothing here earns anything.
|
||||
|
||||
Status on 4 October 2026: every pack is now generator version 2 (`igneum-pow export`, spec 01 sections 1.4.2 to
|
||||
1.4.6; 128 loads per hash, program id in `program.h`). Apple OpenCL on the M5 Max gives 96/96 on all four packs, batch
|
||||
fingerprint `f2a95d5bb84d961e` at 2^13 for `igneum-genesis-mh` (`8e22ad069cb2a8c3` for `igneum-devnet-v4-epoch0`),
|
||||
identical to the emulator in the sub-group 32 and wave64 configurations; 27.9 Mhash/s at 1 GiB, down from 45.0 under
|
||||
the retired 80-distinct-load genesis program, exactly the 128-distinct-load projection of the census. The AMD gfx1036
|
||||
and RTX 5090 OpenCL runs of 3 October were on version 1 packs and are owed again.
|
||||
|
||||
Status on 3 October 2026: no AMD device has run this yet. Everything that could be proven without one has been
|
||||
(`WAVEFRONT.md`, "What was proven"): Apple's deprecated OpenCL 1.2 runtime on the M5 Max, pocl 7.2 on the Mac's CPU,
|
||||
and a CPU emulator with 32- and 64-wide sub-groups all give the Mac's 96/96 vectors and cache FNV for the memory-hard
|
||||
|
|
@ -21,13 +28,14 @@ proto-opencl/
|
|||
build.bat Windows (MSVC cl.exe + OpenCL.lib)
|
||||
WAVEFRONT.md wave32 vs wave64 on AMD, and why the kernel cannot tell the difference
|
||||
emu/ CPU emulator: compiles kernel.cl as C++ and runs it with a 32- or 64-wide sub-group
|
||||
../proto-cuda/packs/<seed>/kernel.cl the OpenCL C kernel of each pack, written by proto-metal/igneum-bench --export-pack
|
||||
../proto-cuda/packs/<seed>/kernel.cl the OpenCL C kernel of each pack, written by igneum-pow export (generator v2)
|
||||
```
|
||||
|
||||
The packs stay in `proto-cuda/packs/` because the three implementations share one `program.h`, `vectors.h` and
|
||||
`memhard.h` per seed; `kernel.cl` sits next to `kernel.cu` and `program.metal`. Three packs are checked in:
|
||||
`igneum-genesis-mh` (memory-hard dataset, the current construction), `igneum-genesis` and `igneum-hourly`
|
||||
(closed-form dataset, kept for comparison).
|
||||
`memhard.h` per seed; `kernel.cl` sits next to `kernel.cu` and `program.metal`. Four packs are checked in, all
|
||||
generator version 2: `igneum-genesis-mh` and `igneum-devnet-v4-epoch0` (memory-hard dataset, the current
|
||||
construction; the second is the devnet's own epoch 0 derivation), `igneum-genesis` and `igneum-hourly` (closed-form
|
||||
dataset, kept for comparison).
|
||||
|
||||
## What the OpenCL path does differently
|
||||
|
||||
|
|
@ -155,9 +163,9 @@ In order, the run prints:
|
|||
6. Six `verify warp ... PASS` lines: three standalone (bases 0, 4096, 1000000) and the same three read out of the
|
||||
warm-up batch.
|
||||
7. `batch fingerprint`: FNV-1a 64 of every output of the warm-up batch. At `--batch-log2 13` every implementation so far
|
||||
prints `f99fb375b3abeaf5` for igneum-genesis-mh (Apple OpenCL, pocl on both exchange paths, the emulator in seven
|
||||
configurations); an AMD run at `--batch-log2 13` must print the same value. At the default 2^24 the value is a new
|
||||
reference to compare AMD against the next machine.
|
||||
prints `f2a95d5bb84d961e` for the version 2 igneum-genesis-mh (Apple OpenCL, the emulator in the sub-group 32 and
|
||||
wave64 configurations; the version 1 value was `f99fb375b3abeaf5`); an AMD run at `--batch-log2 13` must print the
|
||||
same value. At the default 2^24 Apple OpenCL prints `25f96e7dce90bd4e` (`3cc4fbf90fa6366c` for the devnet pack).
|
||||
8. `timed:` with device and wall time, then the `rate` line: Mhash/s and GB/s useful.
|
||||
9. The summary table and `OVERALL: PASS`. Exit code 0 on PASS, 1 on FAIL, 2 on an OpenCL error or a build failure (the
|
||||
build log is printed in full).
|
||||
|
|
@ -174,6 +182,16 @@ odd, run it again and keep both.
|
|||
|
||||
## Results so far (no AMD)
|
||||
|
||||
Generator version 2 packs (4 October 2026):
|
||||
|
||||
| Machine, runtime | Exchange | igneum-genesis-mh | igneum-devnet-v4-epoch0 | closed-form packs | Rate at 1 GiB |
|
||||
|---|---|---|---|---|---|
|
||||
| Apple M5 Max, Apple OpenCL 1.2 | local memory | cache FNV = Rust, 96/96, fingerprint `f2a95d5bb84d961e` at 2^13 (also at `--group-warps 2`) | cache FNV `448274a57f508cbc`, 96/96, `8e22ad069cb2a8c3` | 96/96 and 96/96 | 27.9 Mhash/s, 14.3 GB/s useful, 128 loads per hash (every pack within 1 percent of this rate) |
|
||||
| CPU emulator (`emu/`), sub-group 32 and wave64 | both | 96/96, `f2a95d5bb84d961e` | 96/96, `8e22ad069cb2a8c3` | not run | none |
|
||||
|
||||
Version 1 packs (3 October 2026, retired vectors; the kernel text is unchanged, so the exchange and arithmetic
|
||||
findings stand):
|
||||
|
||||
| Machine, runtime | Exchange | igneum-genesis-mh | closed-form packs | Rate at 1 GiB |
|
||||
|---|---|---|---|---|
|
||||
| Apple M5 Max, Apple OpenCL 1.2 (deprecated runtime) | local memory (runtime lists no sub-group extension) | cache FNV = Mac, 96/96; also 96/96 at `--group-warps 2` and `4` | 96/96 and 96/96 | 45.03 Mhash/s, 18.73 GB/s useful (Metal on the same chip: 45.2). Apple OpenCL on the M5 Max, not an AMD number |
|
||||
|
|
@ -200,14 +218,14 @@ so configurations can be compared for every nonce, not only the vector warps. It
|
|||
## Regenerating a pack
|
||||
|
||||
```
|
||||
cd proto-metal
|
||||
swiftc -O -o igneum-bench main.swift -framework Metal
|
||||
./igneum-bench --seed igneum-genesis --export-pack ../proto-cuda/packs/igneum-genesis-mh
|
||||
./igneum-bench --closed-form --seed igneum-genesis --export-pack ../proto-cuda/packs/igneum-genesis
|
||||
./igneum-bench --closed-form --seed igneum-hourly --export-pack ../proto-cuda/packs/igneum-hourly
|
||||
cd igneum-pow && cargo build --release
|
||||
./target/release/igneum-pow export --seed igneum-genesis --out ../proto-cuda/packs/igneum-genesis-mh
|
||||
./target/release/igneum-pow export --closed-form --seed igneum-genesis --out ../proto-cuda/packs/igneum-genesis
|
||||
./target/release/igneum-pow export --closed-form --seed igneum-hourly --out ../proto-cuda/packs/igneum-hourly
|
||||
```
|
||||
|
||||
The exporter writes `kernel.cl` next to `kernel.cu` from the same instruction list, refuses to write unless the Metal
|
||||
GPU matches the CPU interpreter on all 96 vector outputs, and for memory-hard packs unless the GPU cache equals the
|
||||
CPU cache word for word. The memory-hard core is emitted once (`emitMemhardCore`) in three dialects (Metal, CUDA C++,
|
||||
OpenCL C), so the constants in `kernel.cl` are the literals of `memhard.h`.
|
||||
The Rust exporter (since 4 October 2026 the pack source) writes `kernel.cl` and `kernel_bound.cl` next to `kernel.cu`
|
||||
from the same instruction list. The memory-hard core is emitted once (`emit_memhard_core`) in three dialects (Metal,
|
||||
CUDA C++, OpenCL C), so the constants in `kernel.cl` are the literals of `memhard.h`. The Metal cross-check of a pack is
|
||||
`proto-metal/igneum-bench --seed <seed> --export-pack <scratch dir>` and a diff of its `vectors.json` (see
|
||||
`../proto-cuda/README.md`, "Regenerating a pack").
|
||||
|
|
|
|||
Loading…
Reference in a new issue