Merge branch 'ca3-derive' into ca3-coord

# Conflicts:
#	docs/analysis/chip-model-v3.md
This commit is contained in:
igneum-labs 2026-10-06 07:51:00 +00:00
commit 1c08438a12
38 changed files with 58874 additions and 18 deletions

View file

@ -332,3 +332,46 @@ measurements land.
check that they are the right order. check that they are the right order.
- The ALU budgets of the M5 Max and the 9070 XT, their power at the hash, and the verifier's cost at N = 100,000 - The ALU budgets of the M5 Max and the 9070 XT, their power at the hash, and the verifier's cost at N = 100,000
program ops are estimates; the program-length lever is a design item with its own measurements, not a result. program ops are estimates; the program-length lever is a design item with its own measurements, not a result.
## 6. The per-day derivation (item 2)
6 October 2026, Counter ASIC 3.0 item 2, worker `derive` (`docs/plans/counter-asic-3-derivation.md`; everything
PROPOSED, a prototype behind load class `dr736`). The fixed-shape mixer of section 2's rows is replaced by nine
straight-line programs of 736 instructions per item drawn from the day key stream (twelve two-register forms,
the chain rule, an acceptance test with the x8 mixer's counts as floors). The chip's cost per hash is still item
derivations; what changes is the fixed-function factor, because the chip must now execute an arbitrary program
of the day from a 12-form set over 16 registers (a sequencer: instruction store, register file, operand muxes, a
32-bit ALU with a multiplier and a rotator) instead of a wired pipeline of 72 mixer stages with the day's
constants in the wires. The counts are from the code (`memhard::mixer`: 144 ops per application as written, 128
with the round constants hoisted, 16 multiplies; the x8 item is 10,368 / 9,216 / 1,152), not the 130 of section 1;
the day program's floor is those counts, so the bare row cannot fall below x8's.
| Row | Derivation | Chip ops per hash | Chip rate at 50 T op/s | SRAM the chip holds | mm^2 / $ (N5 headline) | Bare gain against 136.1 MH/s | Allowance 1.2x (ProgPoW's claimed range, history 2.4 [S67] [S70]) | Allowance 1.5x (cautious upper bound, approximate) | The old 3x (the fixed shape's; does not apply) | Equal silicon, SRAM deducted, at 1.2x / 1.5x |
|---|---|---|---|---|---|---|---|---|---|---|
| x8 as shipped (section 2's v3 row, re-counted from the code with constants hoisted) | fixed mixer, 72 x 128 | 1,179,648 | 42.4 MH/s | 256 MiB | 128 / $46 | 0.31x | 0.37x | 0.47x | 0.93x | 0.31x / 0.39x |
| **dr736, the genesis day's draw** (9,992 chip ops, 1,461 multiplies per item; the floor is x8's 9,216) | the day program, 9 x 736 instructions | 1,278,976 | 39.1 MH/s | 256 MiB | 128 / $46 | 0.29x | 0.34x | 0.43x | 0.86x | 0.29x / 0.36x |
| dr736 at the floor (a day whose draw sits exactly on the acceptance floor) | the day program | 1,179,648 | 42.4 | 256 MiB | 128 / $46 | 0.31x | 0.37x | 0.47x | 0.93x | 0.31x / 0.39x |
| dr368, the fallback (the x4-equivalent count: 5,004 chip ops per item on the genesis day) | the day program, 9 x 368 | 640,512 | 78.1 | 256 MiB | 128 / $46 | 0.57x | 0.69x | 0.86x | 1.72x | 0.57x / 0.71x |
| dr736 at year 4 (cache 512 MiB) | the day program | 1,278,976 | 39.1 | 512 MiB | 255 / $111 | 0.29x | 0.34x | 0.43x | 0.86x | 0.23x / 0.28x |
Arithmetic, row dr736: 9,992 x 128 = 1,278,976; 50 x 10^12 / 1,278,976 = 39.1 x 10^6; 39.1 / 136.1 = 0.287; x 1.2
= 0.345; x 1.5 = 0.431; x 3 = 0.862; equal silicon (750 - 128) / 750 = 0.829, x 0.345 = 0.286, x 0.431 = 0.357.
The allowance argument, plainly: the 3x of section 1 was the credit for "a pipeline with no scheduling or
divergence", which a fixed dataflow earns because the chip wires the 72 applications and bakes the constants in;
with a program that changes daily the chip keeps no divergence (the GPU has none here either: the item function
is straight-line), the constants folded into an instruction store, and no warp scheduler or operand collector,
and it loses the wiring. That residual is what ProgPoW's audits priced at 1.1x to 1.2x for a conventional
compute chip (Rao: "conventional compute chips gain little on ProgPoW", history section 2.4); 1.5x is a cautious
upper bound of mine (approximate) for a chip that also drops the GPU's float and graphics area. The chain rule
(every instruction reads the register the previous one wrote) adds a cost the row does not credit: with no
intra-item parallelism a single engine completes one dependent instruction per cycle at best and must interleave
items to keep its multiplier busy, which is a register file per item in flight (RandomX's light-mode argument,
history 2.4). The measured costs that buy this: the verifier 4.88 ms per unit on one M5 Max core against x8's
2.06 (the derivation document's section 5.1), the Mac's daily build 29 ms against 22, the hash rate unchanged;
the 5090's build and compile are the PC 2 job, the 9070 XT's OWED.
What this does not settle: the rows are the same 50 T op/s budget and the same denominator as section 2 (their
margins apply); no chip has been priced for its instruction store or its register files per item in flight; the
random ARX programs have had no cryptanalysis (the item 3 brief should name them beside `M_r`); the 2019-class
core measurement (O-1.14) decides whether 736 or 368 is the length, and the derivation document's section 0
carries that verdict.

View file

@ -1961,3 +1961,84 @@ Ten distinct programs (seed strings `igneum-devnet-v4-epoch0`, `/epoch1` .. `/ep
Cache fill 1.95 ms GPU (192.4 ms one core), dataset build 20.8 ms GPU for 1 GiB. The devnet pack three times through `packbench --pack ../proto-cuda/packs/igneum-devnet-v4-epoch0 --batches 1 --batch-log2 20 --group 256` (the pack's two libraries, `memhard.metal` and `program.metal`): compile 79 ms, 1 ms, 1 ms (the system shader cache answers the identical source from the second run); cache fill 0.6 to 0.7 ms GPU, dataset build 20.7 to 20.8 ms GPU. Cache fill 1.95 ms GPU (192.4 ms one core), dataset build 20.8 ms GPU for 1 GiB. The devnet pack three times through `packbench --pack ../proto-cuda/packs/igneum-devnet-v4-epoch0 --batches 1 --batch-log2 20 --group 256` (the pack's two libraries, `memhard.metal` and `program.metal`): compile 79 ms, 1 ms, 1 ms (the system shader cache answers the identical source from the second run); cache fill 0.6 to 0.7 ms GPU, dataset build 20.7 to 20.8 ms GPU.
Reading: a fresh program compiles in about 18 ms on this card with the Metal compiler service warm, 79 ms for a pack with its dataset kernels, up to 1.8 s cold (the variant-racing entry's first seed), 0 to 444 ms at the fleet's live boundaries (M11). The hot table fill of layer 5 is 0.07 to 0.22 ms (ca2-cache). So the Mac's per-epoch compile-ahead is under 2 s without the race and about 38 s with it (M11: 34.0 / 34.9 / 37.8 s), and the race is the only item visible against the 600-s window in which the program is known (lead 1,200 s minus the 600-s VDF, fixed at every epoch length). PC cards, cited in the plan: RTX 5090 NVRTC 151 to 180 ms, prepare 0.5 to 1.0 s without the dataset (M11), race one round about 37 s; RX 9070 XT OpenCL compile NOT MEASURED at the current worker (owed: `host.c` times `clBuildProgram` only in the `prepare` path and no `prepared` line from gfx1201 is in any upload); Intel UHD build 3.0 to 6.4 s (M11). Floor by the rule (slowest compile-ahead under 10% of the epoch and inside the window, dataset excluded): 600 DAA s, carried by the race at 6.3% of 600 s; with the race off (M11 found base wins on both the 5090 and the Mac) the slowest measured row is the Intel iGPU at 1.1%. Consequences per tier and the difficulty-settle constraint (24% of a 600-s epoch in settle at the measured 144 s) are in the plan. Reading: a fresh program compiles in about 18 ms on this card with the Metal compiler service warm, 79 ms for a pack with its dataset kernels, up to 1.8 s cold (the variant-racing entry's first seed), 0 to 444 ms at the fleet's live boundaries (M11). The hot table fill of layer 5 is 0.07 to 0.22 ms (ca2-cache). So the Mac's per-epoch compile-ahead is under 2 s without the race and about 38 s with it (M11: 34.0 / 34.9 / 37.8 s), and the race is the only item visible against the 600-s window in which the program is known (lead 1,200 s minus the 600-s VDF, fixed at every epoch length). PC cards, cited in the plan: RTX 5090 NVRTC 151 to 180 ms, prepare 0.5 to 1.0 s without the dataset (M11), race one round about 37 s; RX 9070 XT OpenCL compile NOT MEASURED at the current worker (owed: `host.c` times `clBuildProgram` only in the `prepare` path and no `prepared` line from gfx1201 is in any upload); Intel UHD build 3.0 to 6.4 s (M11). Floor by the rule (slowest compile-ahead under 10% of the epoch and inside the window, dataset excluded): 600 DAA s, carried by the race at 6.3% of 600 s; with the race off (M11 found base wins on both the 5090 and the Mac) the slowest measured row is the Intel iGPU at 1.1%. Consequences per tier and the difficulty-settle constraint (24% of a 600-s epoch in settle at the measured 144 s) are in the plan.
## 6 October 2026, Counter ASIC 3.0 item 2: the per-day derivation
Branch `ca3-derive` (worker "derive", from ca3-coord 50df751; commits acb96ee and after), design, spec text and
the chip row in `docs/plans/counter-asic-3-derivation.md` and `docs/analysis/chip-model-v3.md` section 6.
Question (the plan's item 2): replace the fixed-shape mixer (the chip model's 3x fixed-function allowance, 0.31x
to 0.92x) with a random item-derivation program drawn per day from the day key stream (RandomX's SuperscalarHash
idea, `superscalar.cpp` read at upstream 7607fb2), keep the 8 dependent cache reads per item exactly, keep the op
count per item at or above x8's, and measure the verifier against the 10 ms gate, bit-exactness, the daily build
and the hash rate. Machine: Apple M5 Max, 64 GiB, Darwin 25.6.0; every timing row names its lock and load average.
The construction (class `dr736`, `igneum-pow/src/derive.rs`): nine straight-line programs of 736 instructions per
item (one before each cache read, one after the last), four draws per instruction from the mixer's own SplitMix64
stream after its 40 draws, twelve two-register forms (add, sub, xor, mul-lo by `c|1`, rotate-add, xor-rotate,
add-constant, xor-constant, the `M_r` form `(d ^ c) * odd`, `d * odd + c`, `d ^= c & b`, `d += c | b`), every
instruction reading the register the previous one wrote (the chain, `s[0]` first) and writing another, every form
a bijection on the state; the acceptance test rejects a register never written, fewer than 8 distinct rotations,
or a draw under the x8 mixer's counts from the code (72 x 128 = 9,216 chip ops, 72 x 144 = 10,368 as written,
1,152 multiplies; the coordinator's correction of the 130-per-application figure). The genesis day draws 6,624
instructions, 10,659 GPU ops, 9,992 chip ops, 1,461 multiplies per item; the verifier runs it with a word-major
interpreter over the 32 items of a load, dispatching on instruction pairs, no JIT.
**Verifier per 32-lane unit, one core (`with-lock.sh measure`, one session 07:42:20 to 07:42:33 UTC, load
average 4.91 / 4.53 / 5.34 at the start, 4.46 / 4.44 / 5.30 at the end; `igneum-pow bench --seed igneum-genesis
--day 2026-10-03 --class <c> --warps 50`, two rounds, then the devnet seeds once):**
| Class | ms per unit, avg of 50 (round 1 / 2) | Worst cold of three | Against x8 | Ops per item (GPU / chip / mul) |
|---|---|---|---|---|
| v2 | 0.598 / 0.594 | 0.697 | | 1,296 / 1,152 / 144 |
| x8 (mx8, class v3) | 2.061 / 2.063 | 2.179 | 1 | 10,368 / 9,216 / 1,152 |
| dr736 | 4.875 / 4.944 | 5.241 | 2.37x | 10,659 / 9,992 / 1,461 |
| x8, devnet seeds | 2.078 | 2.155 | | |
| dr736, devnet seeds | 4.872 | 5.241 | 2.34x | 10,701 / 10,083 / 1,362 |
| dr368 (half length, the x4-equivalent fallback) | 2.692 | 2.898 | 1.31x | 5,350 / 5,004 / 752 |
The v2 row reads the quiet nights' 0.60 (readwidth 0.604 to 0.626; 6.4a 0.607 to 0.611), so these are quiet-core
figures. The interpreter's cost split (`examples/derive_perf.rs`, a functional run under the run lock, load 3.9 to
4.9): 10.7 ns per instruction per 32-item batch cold, 7.18 with one dispatch per instruction, 4.98 with pair
dispatch; a uniform program (predictable dispatch) 3.4 to 4.3 ns, so about 1.5 ns is dispatch and 3.5 ns the
vector body (NEON, 1,180 `.4s` instructions in the binary).
**Bit-exactness (`with-lock.sh run`):** dr736-genesis on Metal (`packbench --batches 1 --batch-log2 24`) cache
FNV 48c4f5bf24166b2e PASS, dataset head and word [MASK] PASS, vectors 3/3 standalone and 3/3 in batch, fingerprint
2^24 50e3eaa779da4f1e, compile 784 ms cold; on Apple OpenCL (`igneum-bench-cl-dr736-genesis --bench-pack`) the
self-test PASS with the 64 samples and 96 of 96 lanes, fingerprint 50e3eaa779da4f1e (equal); dr736-devnet-epoch0
on Metal PASS, fingerprint 9553f6d5c667205a. Two compilers agree with the Rust interpreter on the derived dataset
and on 2^24 outputs.
**Daily 1 GiB build and hash rate, Metal (`with-lock.sh measure`, the same session, `packbench --batches 2
--batch-log2 22 --group 256`, three rounds):**
| Pack | Compile (1 / 2 / 3) | Build, GPU ms (1 / 2 / 3) | MH/s GPU (1 / 2 / 3) |
|---|---|---|---|
| mx8-genesis (x8, the control) | 80 / 1 / 1 ms | 31.3 / 22.1 / 22.1 | 27.155 / 27.076 / 27.123 |
| dr736-genesis | 751 / 1 / 1 ms | 28.9 / 29.0 / 29.1 | 27.125 / 27.129 / 27.063 |
**Chip model (chip-model-v3.md section 6):** 1,278,976 chip ops per hash on the genesis day, 39.1 MH/s at
50 T op/s, 0.29x bare (0.31x at the floor, x8's figure); the fixed-function allowance of the wired mixer (3x) no
longer applies to a chip that must run the day's program: at ProgPoW's claimed 1.2x the row reads 0.34x, at a
cautious 1.5x 0.43x, at the old 3x 0.86x; equal silicon 0.29x / 0.36x. dr368: 0.57x bare, 0.69x / 0.86x.
**Consequences per tier.** The verifier: no miner tier runs it; a node on any 2026 core verifies a block in 5 ms
(x8: 2.1), a pool core serves 205 shares per second (x8: 485; a 22,000-member pool at one share per 10 s needs
11 cores against 4.5), IBD over 108,000 headers is 8.8 min on one core (x8: 3.7); on a 2019-class laptop core
(2.5x, approximate, O-1.14 unmeasured) 736 reads about 12 ms, over the gate, and 368 about 6.7 ms, under it. The
build: the Mac pays 7 ms more per day (29 against 22 ms), nothing to any tier; the 5090 is the PC 2 job below;
the 9070 XT is OWED (PC 1 is the project lead's desk today; its x8 build was 72 to 77 ms); the integrated gfx1036 tier
already misses the per-prepare rule at x8 (epoch-length.md 6.1: 6.9 / 9.4 / 11.7 s prepares at x1, about 55 to
94 s at x8, approximate) and the day program leaves that need (per-day dataset reuse in the workers, 0.3.12) the
same in kind. The compile: the Metal item library is 0.75 to 0.8 s cold once a day and 1 ms from the shader cache;
the CUDA worker compiles memhard.h into every per-epoch kernel and every race variant, so the 5090's nvrtc line
is the number to read. The hash rate: unchanged within 0.3% on the Mac, as the hash kernel only loads. Packs grow
by about 550 KB (memhard.h 196 KB, program.json 156 KB): nothing to any tier.
**Go / no-go:** GO as reserve entry R0 (the PROPOSED text in the derivation document's section 6, not in
docs/spec); NO-GO for genesis-live at 736 instructions until the 2019-class core measurement lands under 10 ms;
the number that decides it is 4.88 ms per unit on one M5 Max core (pass) against about 12 ms on the approximate
laptop row (fail); dr368 passes both rows at 2.69 ms with the chip at 0.57x bare.
**RTX 5090 (PC 2, one job `relay/playbooks/ca3-derive-pc2.ps1`):** PENDING the proving agent's clear and the PC 2
lock; the rows are appended below when the closing report is read. **RX 9070 XT:** OWED.

View file

@ -0,0 +1,415 @@
# Counter ASIC 3.0 item 2: a random item-derivation program per day
6 October 2026. Worker `derive` (branch `ca3-derive`), under the brief of `docs/plans/counter-asic-3.md` item 2 and
the history audit's addition 2 (`docs/analysis/asic-resistance-history.md` section 4.3). Everything here is
PROPOSED: a prototype behind a load class (`dr736`), measured on the Mac and on PC 2, written as a reserve entry for
spec 1.13.2 (section 6) that lives in this file until the project lead's word. Nothing is published and no vector of class v2
or v3 moves (section 3.4).
What it does, in one line: the fixed-shape mixer `M_r` of spec 1.8.4, applied 72 times per item under class v3,
is replaced by nine straight-line programs of 736 instructions drawn once a day from the day key stream, so the chip
that holds the cache on die must run an arbitrary program instead of a wired pipeline; the 8 dependent cache reads
per item, the cache, the loads and the hash kernel are untouched.
## 0. The verdict first
| Gate | Number | Bar | Result |
|---|---|---|---|
| Verifier per 32-lane unit, one M5 Max core, `with-lock.sh measure`, load average 4.9 / 4.5 / 5.3 | 4.875 / 4.944 ms (two rounds of 50), worst cold unit 5.241; the devnet seeds 4.872 | 10 ms | passes, 5.1 ms of margin (x8 reads 2.061 / 2.063 in the same session: 2.37x) |
| The same on a 2019-class laptop core (2.5x, approximate, the design document's ratio; O-1.14 unmeasured) | about 12.2 ms steady, 13.1 worst cold | 10 ms | FAILS on the approximate row; the half-length class `dr368` (the x4-equivalent op count) reads 2.692 ms here, about 6.7 ms on that row, and passes |
| Bit-exact: Metal and Apple OpenCL against the Rust CPU interpreter, two packs | cache FNV, dataset head and word [MASK], 64 samples (OpenCL), 96 vector lanes, 2^24 fingerprint 50e3eaa779da4f1e (dr736-genesis, both compilers) and 9553f6d5c667205a (dr736-devnet-epoch0, Metal) | equal | passes on two compilers; CUDA (PC 2) section 5.4 |
| Daily 1 GiB build, M5 Max, Metal, measure lock | 29.0 / 29.1 / 28.9 ms GPU against mx8's 22.1 / 22.1 ms in the same session (+32%) | under 1 s on every discrete card | passes on the Mac; the 5090 section 5.4; the 9070 XT OWED (PC 1 is the project lead's desk today) |
| Hash rate, M5 Max, Metal, measure lock | dr736-genesis 27.06 to 27.13 MH/s GPU, mx8-genesis 27.08 to 27.16 | equal within noise | equal (0.3%): the hash kernel does not change |
| Chip model (section 7) | 1,278,976 chip ops per hash, 39.1 MH/s at 50 T op/s, 0.29x bare; 0.34x at a 1.2x allowance, 0.43x at 1.5x, 0.86x at the old 3x | under 1x | the allowance is the result: the 3x of the fixed shape no longer applies |
Go / no-go: GO as reserve entry R0 (section 6), NO-GO for genesis-live at 736 instructions until the 2019-class
core measurement (O-1.14) lands under 10 ms; the number that decides it is 4.88 ms per unit on one M5 Max core
(pass) against about 12 ms on the approximate laptop row (fail). The class that passes both rows today is `dr368`
(2.69 ms), at the x4-equivalent op count, with the chip row at 0.58x bare. The way to the 736 figure under the gate
on a laptop is the JIT (section 4.3), which is out of scope and named with its risk.
## 1. Why: what the chip model says the fixed shape is worth
`docs/analysis/chip-model-v3.md` section 2 prices the on-die-cache recompute chip at 50 T op/s: class v3 (x8) costs
it 1,198,080 integer ops per hash (72 mixer applications x 128 items x about 130 ops), 41.7 MH/s, 0.31x the 5090's
136.1 MH/s bare, and 0.92x with "the 3x fixed-function factor (approximate, from memory: 2x to 5x is the usual
credit for a pipeline with no scheduling or divergence)". That credit is the mixer's fixed shape: the chip unrolls
the 72 applications into a wired pipeline with the day's constants baked in, no instruction fetch, no register
file, no operand muxes. RandomX's answer (`vendor/RandomX/src/superscalar.cpp` is not in this tree; read on
6 October 2026 from the upstream repository at commit 7607fb2 into the session scratchpad; the design argument is
`doc/design.md`, history section 2.4) is SuperscalarHash: the dataset item derivation is itself a random program
drawn from the cache key, 8 programs of about 450 instructions per item, so a light-mode chip "becomes a CPU".
RandomX's generator (`generateSuperscalar`, lines 653 to 850) schedules for a superscalar x86 core: it picks a
decode-buffer configuration per cycle, selects a source register that is ready at the cycle and a destination
that is not the source and was not last written by the same op group (`selectDestination`, line 495: no
"xor r,r2; xor r,r2", no "ror r,C1; ror r,C2", no two multiplies in a row on one register), and then computes the
program's ASIC latency as the longest dependency chain (lines 810 to 824) and sets the address register to the
register with the highest one. The item init is `rl[0] = (itemNumber + 1) * superscalarMul0`, the other seven
registers `rl[0] ^ superscalarAdd_i`, then per cache access: run the program, XOR the mix block in, next address
from the address register (`dataset.cpp` `initDatasetItem`, lines 164 to 190).
What carries over here and what does not: the per-day program, the acceptance by construction, the dependency
chain and the address register idea carry over; the x86 port scheduling does not (our verifier interprets the
program for 32 items at once and our miners compile it for a GPU, so the schedule that matters is the GPU's), and
the latency bound RandomX relies on (the program's critical path against DRAM) is replaced by ours: the 8
dependent cache reads per item, which are untouched.
## 2. The design
### 2.1 The draw
One SplitMix64 stream seeded with `K[0] | (K[1] << 32)` (the day key, spec 1.8.1), the 40 draws of 1.8.4
(`ROT`, `MUL`, `RC`) first, exactly as today (the item init `s[8 + i] = t * MUL[i] + RC[i]` still uses them), then
the program: `DERIVE_PROGRAMS = 9` round programs (one before each of the 8 cache reads, one after the last) of
`derive_len` instructions each, four draws per instruction in a fixed order:
| Draw | Range | Sets |
|---|---|---|
| `below(100)` | op roll | the form, by the weight table of 2.2 (cumulative) |
| `below(15)` | destination roll | `d`: the 15 registers other than the chain `c`, in ascending order (roll >= c adds one) |
| `below(14)` | third-register roll | `b`: the 14 registers other than `d` and `c`, ascending (two skips); used by `andx` and `orx`, consumed by every form |
| `next()` | the immediate | `k = 1 + (low32 mod 31)` for the rotate forms; `imm = low32` for `addc`, `xorc`; `imm = low32 OR 1` for `mulc`, `mulc2`; consumed by every form |
The chain register `c` is `s[0]` (the address word) at the start of each round program and the destination of the
previous instruction after it. Every instruction reads `c` and writes `d != c`; the next instruction's chain is
`d`. So no two instructions of a program can run in parallel (the "every instruction consumes the newest result"
rule of the brief, SuperscalarHash's chain made strict), and no two consecutive instructions write one register,
which is what removes the mergeable pairs SuperscalarHash's `selectDestination` guards against. Four draws per
instruction whether the form uses them or not, so the stream position of every draw is fixed by its index and a
future change to one form's draw changes no other draw (the rule 1.13.1 follows for `epoch_len`). A rejected
candidate (2.4) is followed by the next: the stream continues, `attempt + 1`, as the program generator of 1.4.6.
Source: `igneum-pow/src/derive.rs` (`DeriveProgram::draw_candidate`, `draw`), `memhard.rs` (`MixParams::with_shape`).
### 2.2 The instruction set: twelve two-register forms, fixed at genesis
`c` the chain, `d` the destination, `b` the third register, `k` in 1..31, `i` a 32-bit constant (odd for the
multiplies). All arithmetic modulo 2^32, no division, no float, no data-dependent branch. Every form is a bijection
on the 16-word state (the old `d` enters through `+=`, `-=`, `^=`, an odd multiply, or a rotation of itself; `c`
and `b` are not written), so a program loses no entropy, the property `M_r` has.
| Form | Semantics | Weight (percent) | GPU ops | Chip ops | Multiply |
|---|---|---|---|---|---|
| `add` | `d += c` | 14 | 1 | 1 | |
| `sub` | `d -= c` | 10 | 1 | 1 | |
| `xor` | `d ^= c` | 14 | 1 | 1 | |
| `mul` | `d *= (c OR 1)` | 10 | 2 | 1 (the OR is a wire) | yes |
| `rot` | `d = rotl(d, k) + c` | 10 | 2 | 2 | |
| `xrot` | `d = rotl(d ^ c, k)` | 10 | 2 | 2 | |
| `addc` | `d += c + i` | 6 | 2 | 2 | |
| `xorc` | `d ^= c ^ i` | 6 | 2 | 2 | |
| `mulc` | `d = (d ^ c) * i` (the per-word form of `M_r`, the chain in place of the round constant) | 8 | 2 | 2 | yes |
| `mulc2` | `d = d * i + c` | 4 | 2 | 2 | yes |
| `andx` | `d ^= (c AND b)` | 4 | 2 | 2 | |
| `orx` | `d += (c OR b)` | 4 | 2 | 2 | |
Mean 1.62 GPU ops and 1.52 chip ops per instruction, 22% multiplies. `and` and `or` enter only as `andx` and
`orx` (a destructive `d &= c` would lose bits; the XOR and add of a conjunction keep `d` invertible). The forms are
the brief's set (add, sub, mul-lo, xor, rotate by 1..31, and, or, the ARX-multiply forms of `M_r`); `mulhi` is in
the lottery hash's families and bit-exact on the three vendors, and is left out of the derivation on purpose so
every form is one that the three compilers lower to a single integer instruction (section 3.3).
### 2.3 Length and the floors: the x8-equivalent op count
The coordinator's rule for item 2: the total op count per item equals or exceeds today's x8 count, so the chip's
budget row does not fall. The x8 mixer counted from the code (`memhard::mixer`): 16 x (xor, add, mul) + 8 quarter
rounds x 12 = 144 ops as written, 128 with the `RC[i] + rk` adds hoisted as constants (the chip and every compiler
do that), 16 multiplies; 72 applications per item = 10,368 ops as written, 9,216 hoisted, 1,152 multiplies
(`chip-model-v3.md` section 1 prices 130 per application from the spec text; the item 3 worker counted the same
144 / 128, coordinator's note of 6 October). The floors are those three. `derive_len = 736` gives 6,624
instructions per item, expected 10,731 GPU ops, 10,068 chip ops and 1,457 multiplies (the genesis day draws
10,659 / 9,992 / 1,461; the devnet day 10,701 / 10,083 / 1,362), 7 to 22 standard deviations above the floors, so a
rejection on a floor is a rare event and the test exists for the degenerate class. The class name is `dr736`; the
half-length `dr368` (the x4-equivalent) is measured beside it as the fallback with its floors scaled.
### 2.4 The acceptance test
`DeriveProgram::check`, on every candidate; a rejection draws the next attempt from the stream:
| Test | Rejects | Expected rate at 736 |
|---|---|---|
| every register written in every round program | a register no instruction of a round program writes (its init word would never enter the chain within that round) | 16 x 9 x (14/15)^736 = under 10^-20 |
| at least 8 distinct rotation amounts across the item's programs | all rotations equal (the brief's degenerate draw; the mixer's "ROT draw of eight equal values is possible and untested", spec 1.8.4) | about 1,300 rotate forms over 31 values: never |
| chip ops >= 9,216, GPU ops >= 10,368, multiplies >= 1,152 per item (scaled to the length) | an op draw under the x8 count | 22, 9 and 9 standard deviations below the mean: never |
| structural (asserted, hold by construction): `d != c`, `b` distinct from both, `k` in 1..31, odd multiplier constants | a generator bug | n/a |
The generator panics after 64 rejected candidates in a row (`MAX_ATTEMPTS`), which the rates above put beyond
any day the chain will see; the panic is the right failure (a node that cannot derive the day's program cannot
verify, and must say so rather than guess). The weights are fixed at genesis; only the order, the registers and
the constants are drawn, so the family mix of a program cannot be steered by the draw.
### 2.5 The item, with the program in place
Spec 1.8.5 under the derivation class (`derive_len` nonzero, `mixer_mult` unused):
```
s[i] = K[i] for i in 0..7
s[8 + i] = t * MUL[i] + RC[i] for i in 0..7
for r in 0..7:
s = P_r(s) round program r, the chain starting at s[0]
a = s[0] AND (2^(C - 4) - 1) cache line index, as today
s[i] = s[i] XOR cache[line a][i] for i in 0..15
s = P_8(s)
item(t) = s
```
The 8 dependent cache reads per item are exactly today's: the address of read `r` is `s[0]` after program `r`, and
`s[0]` depends on every earlier read through the chain (every round program writes every register, 2.4, and the
XOR of the line into all 16 words feeds the next program). The verifier's latency part (8 dependent misses per
item, overlapped across the up to 32 items of a load, spec 1.11) is unchanged, which the measurement shows: the
difference against x8 is the ALU part only (section 5.1).
## 3. The prototype
### 3.1 Where it lives
| Item | Where |
|---|---|
| `DOp`, `DInstr`, `DeriveProgram` (draw, check, counts, fingerprint), the SoA interpreter `run_round` (pair dispatch), the scalar reference `run_round_scalar`, the text forms `instr_text` and `instr_line` | `igneum-pow/src/derive.rs` |
| `Shape::derive_len`, `Shape::is_derived`, `MixParams::derive` (drawn after the 40 mixer draws), `derive_items` dispatching to `derive_items_program` | `igneum-pow/src/memhard.rs` |
| `LoadClass::derive_len`, `LoadClass::DR736`, `with_derive`, parse and name `dr<len>`, the program id (`derive/` + the length) | `igneum-pow/src/generator.rs` |
| `mh_round_0..8` and the program-driven `mh_item` in memhard.h, memhard.metal and kernel.cl; `IGNEUM_DERIVE_*` in program.h; `derive_len`, the op mix, the floors' counts and the nine programs (one line per instruction) in program.json | `igneum-pow/src/emit.rs` |
| `--class dr736` (or any `dr<len>`) on every command; the bench prints the program's counts | `igneum-pow/src/main.rs` |
| `examples/derive_perf.rs`: the interpreter's cost per instruction per batch, drawn program against uniform programs | `igneum-pow/examples/` |
| Packs `dr736-genesis` (seed igneum-genesis, day 2026-10-03, program id 72c1d8048aef9542, program fingerprint 463535d01511350d) and `dr736-devnet-epoch0` (the devnet epoch 0 and day seeds, 7f4a5ca0a3637820, 771868df4e64d6ab); generator 2 with the class in the id, the x4-record shape, not a class v3 pack | `proto-cuda/packs-ca3-derive/` |
| Tests: `src/derive.rs` (5), `tests/derive.rs` (7: by hand on a small cache, batches, v2 and v3 untouched, the stream and the class, determinism and the pack text, stats beside x8, the text forms against the scalar reference, the word path) | `igneum-pow` |
| The PC 2 job | `relay/playbooks/ca3-derive-pc2.ps1` (section 5.4) |
### 3.2 The verifier without a JIT: the word-major interpreter
The verifier derives up to 32 distinct items per load (one per lane of the unit, `MemhardCpu::fetch`). The
interpreter keeps the 32 item states word-major (`st[reg][lane]`, 2 KiB) and runs each instruction across the
whole batch in one straight loop the compiler vectorises (NEON `add.4s`, `mul.4s`, `ushl.4s` and so on: 1,180
such instructions in the example binary), so the dispatch is paid once per instruction per batch, not per item;
the cache reads of the batch are issued together after each round program, as the fixed-mixer loop does, so the 8
dependent misses of independent items overlap. The dispatch is on PAIRS of instructions (144 arms, one indirect
branch per two instructions): a drawn op sequence is random, the predictor misses most dispatches, and pairing
halves the misses per instruction. Measured with `examples/derive_perf.rs` (a functional run, load average 3.9 to
4.9): 10.7 ns per instruction per batch cold, 7.18 warm with single dispatch, 4.98 with pair dispatch; a uniform
program of one form (predictable dispatch) 3.4 to 4.3 ns, so the body is about 3.5 ns and the remaining dispatch
cost about 1.5 ns. 6,624 x 4.98 ns x 128 batches = 4.2 ms per unit of interpreter time; the measured 4.88 ms
includes the latency part and the transposes.
### 3.3 Bit-exactness on three vendors, by construction
Each form is one C statement on `uint` with `+`, `-`, `^`, `*`, `|`, `&` and the memhard core's `mh_rotl` (a
shift pair, `n` in 1..31 at every call site), the same text in Metal, CUDA C and OpenCL C, every operand a 32-bit
unsigned integer: the same argument as spec 1.14 for the lottery hash's families, which have run bit-exact on the
three vendors since 4 October. The program is emitted as nine functions of straight-line statements (196 KB of
memhard.h per day); NVRTC, the Metal compiler and the OpenCL compilers see no loop, no branch and no call inside a
round program. Measured: section 5.2 (Metal and Apple OpenCL), 5.4 (CUDA).
### 3.4 The v2 and v3 paths are untouched
`Shape::derive_len` is 0 and `LoadClass::derive_len` is 0 on `V2`, `MX4`, `MX8` and `V3_CLASS`; `MixParams::derive`
is `None`; `derive_items` takes the fixed-mixer loop as before; the emitter's text for a shape without a program is
unchanged. `cargo test -p igneum-pow` (commit acb96ee): 58 lib, 7 derive, 4 mixer, 19 packs (every pinned pack of
v2 and v3 regenerated and compared byte for byte), 7 scratch, all green.
## 4. What the chip keeps, and the fallbacks
### 4.1 The fixed-function allowance after the change
The chip of `chip-model-v3.md` now has to execute, per item, 6,624 instructions from a 12-form set over a
16-entry register file, with three operand fields and a constant per instruction, in an order and with operands
that change every day. That is a sequencer: an instruction store (6,624 x 9 bytes = 60 KB per day program, in
SRAM beside the cache mirror), a register file with three read ports and one write port, a 32-bit ALU with a
multiplier and a barrel rotator, operand muxes, and a program counter. A GPU streaming multiprocessor is the same
machine with a wider register file and a warp scheduler. What the chip keeps over the GPU: no warp scheduler, no
operand collector, no instruction cache hierarchy for a 1 MB kernel (the GPU's day program is about 60,000 SASS
instructions per thread; whether the 5090's instruction cache holds it is in the build time of 5.4), no graphics
or float units idle on the die, and the day's constants folded into the instruction store. What it loses: the
wired pipeline (the mixer's 72 applications as 72 stages with the constants in the wires, no fetch, no register
file, no crossbar), which is the thing the 3x credit paid for. The chain rule adds a second loss: with no
intra-item parallelism a single engine finishes one instruction per cycle at best and its multiplies serialise
unless it interleaves items, which costs a register file per item in flight (RandomX's light-mode argument,
history 2.4: a chip paying "760 cycles and 1,240 multiplies per item"). ProgPoW claimed 1.1x to 1.2x for exactly
this kind of chip ("conventional compute chips gain little on ProgPoW", Bob Rao's hardware audit, history 2.4,
[S70]; EIP-1057's own claim 1.1x to 1.2x, [S67]). So the allowance this document carries is 1.2x (ProgPoW's
claimed range, cited) with 1.5x as the cautious upper bound (approximate, mine), against the 3x of the fixed shape
(approximate, from memory, M16). Section 7 prices all of 1.0x, 1.2x, 1.5x, 2x and 3x.
### 4.2 What it does not change
The partial-store chip (item 1): a chip that stores the dataset and never derives items pays nothing for the
program; item 1's rows stand on their own. The cryptanalysis question (item 3) changes shape: instead of one
fixed `M_r` to attack, the attacker gets a fresh random ARX program every day, which is RandomX's bet and is
untested here; the day's programs can be audited by the same tools as random ARX ciphers, and the acceptance test
is the place to add a structural rule if one is found. The era draws and the cache growth are untouched.
### 4.3 The JIT, named as the fallback with its risk
A per-day JIT for the verifier (emit NEON or AVX2 code for the 32-lane batch, the chain row kept in registers
across instructions since `c` is always the row just written) would remove the 1.5 ns dispatch and about a third
of the 3.5 ns body (the chain row's loads), about 2.5 ms per unit on this core, the x8 figure. Its risk: a code
generator in the consensus path on two architectures (arm64, x86-64) whose output must equal the interpreter's
bit for bit, writable-executable memory in a node and in every pool verifier, a new attack surface the history's
lesson 9 says to audit before launch, and a second implementation per platform to keep in lockstep. Out of scope
for this item; it is the route to 736 under the gate on a laptop core if the measurement of O-1.14 confirms the
approximate row, and `dr368` is the route that needs no JIT.
## 5. Measurements
All under the locks of the brief; every row says its lock and the load average. Machine: Apple M5 Max, 64 GiB,
Darwin 25.6.0. Commit acb96ee (the code and the packs), packbench built from this worktree.
### 5.1 The verifier per unit, one M5 Max core (`with-lock.sh measure`, one session, 07:42:20 to 07:42:33 UTC, load average 4.91 / 4.53 / 5.34 at the start, 4.46 / 4.44 / 5.30 at the end)
Script `measure-ca3-derive.sh` (session scratchpad): `igneum-pow bench --seed igneum-genesis --day 2026-10-03
--class <v2|mx8|dr736> --warps 50`, two rounds, then the devnet seeds (`--epoch-hex edc4fa84...fb07 --day-hex
69676e65756d2d6461792ffa50000000000000`) and `dr368` once.
| Class | Verifier, ms per unit, avg of 50 (round 1 / round 2) | Worst cold unit of three | Against v2 | Against x8 | Items per unit | Ops per item (GPU / chip / multiplies) |
|---|---|---|---|---|---|---|
| v2 (igneum-genesis) | 0.598 / 0.594 | 0.697 | 1 | | 4,096 | 1,296 / 1,152 / 144 (9 x 144, hoisted 9 x 128) |
| x8, mx8 (class v3) | 2.061 / 2.063 | 2.179 | 3.46x | 1 | 4,096 | 10,368 / 9,216 / 1,152 |
| **dr736** | **4.875 / 4.944** | **5.241** | 8.2x | 2.37x | 4,096 | 10,659 / 9,992 / 1,461 |
| x8, the devnet seeds | 2.078 | 2.155 | | | 4,095 to 4,096 | |
| dr736, the devnet seeds | 4.872 | 5.241 | | 2.34x | 4,096 | 10,701 / 10,083 / 1,362 |
| dr368 (the x4-equivalent fallback) | 2.692 | 2.898 | 4.5x | 1.31x | 4,096 | 5,350 / 5,004 / 752 |
Reading: the v2 row reads 0.59 to 0.60, the quiet readwidth night's 0.604 to 0.626 and 6.4a's 0.607 to 0.611, so
this session is a quiet-core figure and no scaling applies. At the same chip-op count as x8 (plus 8%), the
interpreter costs 2.4x what the compiled mixer costs: 4.98 ns per instruction per batch (3.2), of which about
1.5 ns is dispatch and 3.5 ns the vector body, against the mixer's compiled straight-line loop over the same
batch. The latency part is the same in every row (the 8 dependent misses per item; the dr368 row at half the
instructions saves 2.2 ms of the 4.9, which puts the latency-and-transpose share at about 0.5 ms). The 10 ms gate
keeps 5.1 ms (worst cold 4.76 ms) at 736 on this core and 7.3 ms at 368.
### 5.1a Verification throughput per tier (the form of mixer-x4.md 6.5)
| Figure | x8 (this session) | dr736 | dr368 | Note |
|---|---|---|---|---|
| ms per unit, quiet M5 Max core (measured, this session) | 2.06 | 4.88 | 2.69 | |
| ms per unit, 2019-class laptop core (2.5x, approximate, O-1.14 unmeasured) | 5.2 | 12.2 | 6.7 | the figure that fixes the gate is a measurement, not this row |
| Shares per second per core (quiet M5 Max) | 485 | 205 | 372 | |
| Cores for a 22,000-member pool at one share per member per 10 s (2,200 shares per second) | 4.5 | 10.7 | 5.9 | a pool verifying two units at once would halve the dispatch share (3.2); a 64-lane interpreter is a follow-up, unmeasured |
| Node: worst cold single unit (per block) | 2.2 ms | 5.2 ms | 2.9 ms | a block's verification stays under the 1 s block time by 190x |
| IBD over 108,000 headers on one core | 3.7 min | 8.8 min | 4.8 min | laptop (approximate): 9.4 / 22 / 12 min |
| Margin left under the 10 ms gate (worst cold, this core) | 7.8 ms | 4.8 ms | 7.1 ms | on the laptop row (approximate): 4.8 / none (over by 2.2) / 3.3 ms |
Consequences per tier: every miner tier is untouched by the verifier (the miner never runs it); a pool operator
pays 2.4x the cores at 736 (11 cores for a 22,000-member pool against 4.5) or 1.3x at 368; a node on any 2026 core
verifies a block in 5 ms; a node on a 2019-class laptop core is the open question (O-1.14), and at 736 the
approximate row says it misses the gate, so 736 does not go genesis-live on an approximation, and 368 passes it.
### 5.2 Bit-exactness on the Mac (`with-lock.sh run`, 08:38 to 08:41 local)
`packbench --pack <dir> --batches 1 --batch-log2 24 --group 256` (Metal, built from this worktree) and
`igneum-bench-cl-dr736-genesis --bench-pack --pack <dir> --batches 1 --batch-log2 24` (Apple OpenCL, `proto-opencl/
build.sh` on the pack through a temporary link). Vectors are the Rust interpreter's.
| Pack | Harness | Cache FNV-1a 64 | Dataset head, word [MASK], 64 samples | Vectors | Fingerprint 2^24 | Compile | 1 GiB build, GPU ms (run lock, indicative) |
|---|---|---|---|---|---|---|---|
| dr736-genesis | Metal | 48c4f5bf24166b2e PASS | head and last PASS (packbench checks no samples) | 3/3 standalone, 3/3 in batch | 50e3eaa779da4f1e | 784 ms (276 on the second process, 1 ms once the shader cache has it) | 38.2 / 38.3 |
| dr736-genesis | Apple OpenCL | PASS (head, last line, FNV) | head, word [268435455], 64 samples PASS | 96 of 96 lanes | 50e3eaa779da4f1e | (in the 385 ms prepare) | 57 ms wall |
| dr736-devnet-epoch0 | Metal | 448274a57f508cbc PASS | PASS | 3/3, 3/3 | 9553f6d5c667205a | 763 ms | 37.4 |
Reading: the Rust interpreter, Metal and Apple OpenCL agree on the derived dataset (head, word [MASK], the 64
samples through OpenCL), on every vector lane and on the 2^24-output fingerprint of dr736-genesis across both
compilers; the devnet-seed pack agrees on Metal. The Apple OpenCL compile of a 6,624-statement item function
is inside a 385 ms prepare; the Metal compile is 0.75 to 0.8 s cold per day (the memhard library is the day's,
compiled once a day, not per epoch) and 1 ms from the shader cache.
### 5.3 The daily build and the hash rate on the M5 Max (`with-lock.sh measure`, the session of 5.1, three rounds)
`packbench --pack <dir> --batches 2 --batch-log2 22 --group 256`, mx8-genesis then dr736-genesis, three rounds.
| Pack | Compile (round 1 / 2 / 3) | 1 GiB build, GPU ms (round 1 / 2 / 3) | MH/s GPU (round 1 / 2 / 3) | Vectors, self-tests |
|---|---|---|---|---|
| mx8-genesis (class v3, the control) | 80 / 1 / 1 ms | 31.3 / 22.1 / 22.1 | 27.155 / 27.076 / 27.123 | 3/3 + 3/3, PASS |
| dr736-genesis | 751 / 1 / 1 ms | 28.9 / 29.0 / 29.1 | 27.125 / 27.129 / 27.063 | 3/3 + 3/3, PASS |
Reading: the build is 29 ms against 22 (+32%, +7 ms): the Mac's build was latency-bound at x1, x4 and x8 (21 to
22 ms at every multiplier, mixer-x4.md 6.4) and the serial chain per thread now shows (no intra-item ILP for the
compiler to schedule), still 34x under the 1 s bar. The hash rate is the x8 rate within 0.3% (27.06 to 27.13
against 27.08 to 27.16), as it must be: the hash kernel only loads. Consequences per tier: a daily build of 29 ms
on Apple silicon costs nothing to any tier; the 5090's figure is section 5.4, the 9070 XT's is OWED (PC 1 not
released today; its x8 build was 72 to 77 ms and arithmetic-bound at none of x1, x4, x8, so the chain's cost there
is the open number); the integrated tier is the one to watch (5.5).
### 5.4 RTX 5090, PC 2 (one job, `relay/playbooks/ca3-derive-pc2.ps1`)
PENDING at the time of writing: the job waits for `/tmp/igneum-devnet/pc2-ca3.clear` (the proving agent's
30-minute measurement) and the `pc2-ca3.lock`. The job downloads the packs zip itself (sha256
aadce58ca54f136d8c41dc15154b3ae107624808e7f061c266df228b200bc676: dr736-genesis, dr736-devnet-epoch0, mx8-genesis,
v2-genesis-mh), runs the installed app's igneum-worker-cuda.exe (NVRTC compiles each pack's own text) with the
NVIDIA card off in the app only under test (its key from settings.json), `--check` for the `nvrtc .. cache ..
dataset .. ms` line and the self-test, `--bench` at 2^24 for the fingerprint and the rate, block-warps 1 and 8.
The rows land in the bench-log entry when the closing report is read.
### 5.5 The build per tier, with the integrated tier
| Card | Build at x8 | Build with the day program | Source |
|---|---|---|---|
| M5 Max, Metal | 22.1 ms | 29.0 ms (+32%) | 5.3, measure lock |
| RTX 5090, CUDA | 23 ms | section 5.4 | the PC 2 job |
| RX 9070 XT, OpenCL | 72 to 77 ms | OWED (PC 1) | |
| Radeon integrated gfx1036 (PC 2), OpenCL, per prepare | about 55 to 94 s (approximate, mixer-x4.md 6.5: the iGPU's build is arithmetic-bound at x1 already, scaled x8 from 6.9 / 9.4 / 11.7 s) | about the same count of ops at a lower ILP: 55 to 120 s (approximate, unmeasured) | `docs/plans/epoch-length.md` 6.1 iGPU rows |
| gfx1036 beside WSL build jobs (PC 1) | about 7 to 17 min (approximate) | the same or worse (approximate) | epoch-length.md 6.1 |
| 8 GB-class discrete card (not owned, about a tenth of the 5090, approximate) | about 1 s | about 1 to 1.3 s (approximate) | scaled |
Consequences: nothing changes for a discrete card of any size on any vendor (the build is under a second), the
Mac row measured; the integrated tier already misses the per-prepare rule at x8 and needs the per-day dataset
reuse in the workers (0.3.12) or a restart per epoch, and the day program makes that need the same, not larger in
kind; the per-day compile of the item function (0.75 s Metal; NVRTC on the 5090 in 5.4) lands once a day in the
worker's day-cache build, not per epoch, unless the hash kernel's compile includes memhard.h, which on the CUDA
worker it does (kernel_bound.cu includes it, and the variant race compiles 17 variants): the 5.4 job's nvrtc line
is the number for that, and if it is large the fix is to compile the item function once per day into its own
module.
## 6. PROPOSED spec text for 1.13.2: reserve entry R0, `derive` (the per-day item-derivation program)
Not written into `docs/spec`; it lives here until the project lead's word. Named R0, ahead of R1 (mm8), because the reserve
is to be ordered by chip-unfriendliness (counter-asic-3.md item 6; mm8 last) and a derivation program is the most
chip-unfriendly entry the reserve can hold: it removes the fixed-function allowance of the recompute chip rather
than adding a family that chip can license.
> Reserve entry R0, `derive` (the per-day item-derivation program). Semantics: section 1.8.5 under
> `derive_len = 736`: the nine mixer slots of the item derivation (one before each of the 8 cache reads, one after
> the last) each run a straight-line program of 736 instructions drawn from the day key stream of 1.8.4 after its
> 40 draws, four draws per instruction (`below(100)` the form, `below(15)` the destination among the registers
> other than the chain, `below(14)` the third register among those other than the destination and the chain,
> `next()` the immediate: `1 + low32 mod 31` for the rotate forms, `low32` for `addc` and `xorc`, `low32 OR 1` for
> `mulc` and `mulc2`); the chain is `s[0]` at the start of each program and the previous destination after; the
> twelve forms and weights of `docs/plans/counter-asic-3-derivation.md` section 2.2, fixed; every form a
> bijection on the state; the 8 dependent cache reads, the mixer constants of the item init, the cache and the
> dataset mapping of 1.8.5 unchanged; `mixer_mult` unused under R0. Acceptance test, per candidate, the next
> attempt on rejection (the stream continues): every register written in every round program; at least 8 distinct
> rotation amounts; per item at least 9,216 operations with constants folded, 10,368 as written and 1,152
> multiplies (the x8 mixer's counts from `memhard::mixer`: 72 x 128, 72 x 144, 72 x 16; the floors scale with the
> length). Edge vectors, each a hand-built item run on every vendor: item 0, item 1, item 2^28 - 1 and item 2^32 - 1
> of the genesis day on a 2^16-word cache; a program whose first instruction is each of the twelve forms with
> `d = 15`, `c = 0`, `b = 14`, `k = 31`, `i = 0xffffffff` (odd for the multiplies) on the all-ones state and on the
> all-zero state (the wrap of every form); the day of the pinned pack `dr736-genesis` (program fingerprint
> 463535d01511350d, dataset head `vectors.json`, 2^24 fingerprint 50e3eaa779da4f1e) and of `dr736-devnet-epoch0`
> (771868df4e64d6ab, 9553f6d5c667205a). Unlock: at the start of era n = 2 (DAA 31,104,000), or earlier by the 90%
> signalling path of section 5.7, or at genesis if the verifier on a 2019-class core (O-1.14) reads under 10 ms per
> unit at 736, else at the length that does (368 measured at 2.69 ms on an M5 Max core); never by a release. The
> verifier procedure: the word-major interpreter of `igneum-pow/src/derive.rs` (32 item states per batch, pair
> dispatch), 4.88 ms per unit on one M5 Max core (section 5.1), no JIT; a JIT is the named fallback (4.3). Vendor
> paths: none needed; every form is a single 32-bit integer statement on all three compilers (3.3).
## 7. The chip model row
Written into `docs/analysis/chip-model-v3.md` section 6 in the form of its section 2. Ops per hash: 128 items x
9,992 chip ops (the genesis day's draw; the floor 9,216) = 1,278,976 (floor 1,179,648); chip rate at 50 T op/s =
39.1 MH/s (floor 42.4); bare against 136.1 MH/s = 0.287x (floor 0.31x, the x8 row's figure, as the floor is x8's
count). The allowance rows: 1.0x 0.29x; 1.2x (ProgPoW's claim) 0.34x; 1.5x (cautious upper bound, approximate)
0.43x; 2x 0.57x; 3x (the fixed shape's, which no longer applies) 0.86x. Equal silicon (x 0.829): 0.24 / 0.29 /
0.36 / 0.48 / 0.71. The x8 row read 0.92x at 3x and 0.76x at equal silicon; at the same 0.31x bare the day program
takes the chip from 0.92x to 0.34x to 0.43x, which is the margin the item was for. dr368 (the fallback): 639,488
chip ops per hash, 78.2 MH/s, 0.57x bare, 0.69x at 1.2x, 0.86x at 1.5x: under 1x, with less margin than x8 had at
3x and more than x4 had (1.84x).
## 8. What is unverified or owed
| Item | State |
|---|---|
| RTX 5090: the daily build, NVRTC compile per pack, the self-test and 2^24 fingerprints, the rate | PENDING the PC 2 job (section 5.4); the playbook and the zip are ready; the clear file is polled every 60 s |
| RX 9070 XT (PC 1) | OWED: PC 1 is the project lead's desk today; the same job shape runs there with `igneum-worker-opencl.exe --bench-pack` when released |
| The 2019-class laptop core (O-1.14) | unmeasured; the approximate row decides against 736 at genesis and for 368, and a measurement replaces it |
| Cryptanalysis of random ARX programs | none; item 3's brief should name the day program as a target beside `M_r` |
| The integrated tier's build with the day program | approximate (5.5); the gfx1036 measurement is a PC 2 OpenCL job, not run today (the one PC 2 job carries the 5090) |
| A 64-lane interpreter for pools (two units per batch) | unimplemented; it would cut the dispatch share for pool verifiers only |
| The NVRTC cost of memhard.h inside the per-epoch hash kernel compile and the variant race | the PC 2 job's nvrtc line; the fix, if large, is one module per day for the item function |

View file

@ -0,0 +1,51 @@
//! The derivation interpreter's cost per instruction per batch (Counter ASIC 3.0 item 2): the day's program against
//! a uniform program of the same length (every instruction one form, so the dispatch is predictable), which splits
//! the per-instruction cost into the dispatch and the vector body. A functional tool, not a bench-log number on its
//! own: run it under the measure lock and state the load average when a figure is recorded.
//! cargo run --release --example derive_perf [len] [reps]
use igneum_pow::derive::{run_round, DInstr, DOp, DeriveProgram, SoaState, DERIVE_LEN_X8, DERIVE_REGS, SOA_LANES};
use igneum_pow::seed::SplitMix64;
use std::time::Instant;
fn time(prog: &DeriveProgram, reps: usize) -> (f64, u64) {
let mut st: SoaState = [[0u32; SOA_LANES]; DERIVE_REGS];
let mut x = SplitMix64::new(5);
for r in 0..DERIVE_REGS {
for k in 0..SOA_LANES {
st[r][k] = x.next() as u32;
}
}
let t = Instant::now();
for _ in 0..reps {
for p in &prog.rounds {
run_round(p, &mut st);
}
}
let ns = t.elapsed().as_nanos() as f64 / (reps as f64 * prog.instr_count() as f64);
(ns, st[0][0] as u64 ^ st[15][31] as u64)
}
fn uniform(len: u32, op: DOp) -> DeriveProgram {
let mut p = DeriveProgram::draw_candidate(&mut SplitMix64::new(1), len, 0);
for ins in p.rounds.iter_mut().flatten() {
*ins = DInstr { op, rot: if op.has_rot() { 13 } else { 0 }, imm: if op.has_imm() { 0x9e37_79b9 } else { 0 }, ..*ins };
}
p
}
fn main() {
let a: Vec<String> = std::env::args().collect();
let len: u32 = a.get(1).and_then(|s| s.parse().ok()).unwrap_or(DERIVE_LEN_X8);
let reps: usize = a.get(2).and_then(|s| s.parse().ok()).unwrap_or(400);
let real = DeriveProgram::draw(&mut SplitMix64::new(0x3067619f3c269176), len);
let per_item = real.instr_count();
println!("len {len}: {per_item} instructions per item, {} GPU ops, {} chip ops, {} multiplies; batch of {SOA_LANES} lanes, {reps} reps", real.gpu_ops(), real.chip_ops(), real.muls());
for round in 0..2 {
let (ns, sink) = time(&real, reps);
println!("round {round}: drawn program {ns:.2} ns per instruction per batch ({:.1} us per batch, {:.2} ms per 128 batches) sink {sink:x}", ns * per_item as f64 / 1e3, ns * per_item as f64 * 128.0 / 1e6);
for op in [DOp::Add, DOp::Mul, DOp::XRot, DOp::MulC, DOp::AndX] {
let (ns, sink) = time(&uniform(len, op), reps);
println!("round {round}: uniform {:5} {ns:.2} ns per instruction per batch sink {sink:x}", op.name());
}
}
}

978
igneum-pow/src/derive.rs Normal file
View file

@ -0,0 +1,978 @@
//! The per-day item-derivation program (Counter ASIC 3.0 item 2, `docs/plans/counter-asic-3-derivation.md`):
//! RandomX's SuperscalarHash idea (`vendor/RandomX/src/superscalar.cpp`, read 6 October 2026 at commit 7607fb2)
//! rebuilt for a 16-word item on a GPU. In place of the fixed-shape mixer `M_r` of spec 01 section 1.8.4, each of
//! the nine mixer slots of an item (one before each of the 8 dependent cache reads, one after the last) runs a
//! straight-line program of [`DERIVE_LEN`] instructions drawn once a day from the day key stream, from a fixed
//! set of twelve two-register forms. The 8 dependent cache reads per item are untouched.
//!
//! Rules of the draw (the dependency chain of SuperscalarHash, made strict):
//! * every instruction reads the chain register `c`, the register the previous instruction wrote (`s[0]`, the
//! address word, at the start of each round program), and writes a register `d != c`, which becomes the chain;
//! so no two instructions of a program can run in parallel, and no two consecutive instructions write one
//! register (the "ror r,C1; ror r,C2" and "xor r,r2; xor r,r2" merges of SuperscalarHash's `selectDestination`
//! cannot arise);
//! * every form is a bijection on the 16-word state (the old `d` enters through `+=`, `-=`, `^=`, an odd multiply,
//! or a rotation of itself), so a program loses no entropy, the property `M_r` has;
//! * the forms are integer only, modulo 2^32, with rotations by 1..31: bit-exact on Metal, CUDA and OpenCL by the
//! same argument as the lottery hash's families (spec 01 section 1.14); no division, no float, no branch;
//! * four draws per instruction in a fixed order, so the stream position of every draw is fixed by the index.
//!
//! The acceptance test ([`DeriveProgram::check`]) rejects a degenerate draw and the next attempt is drawn from the
//! continuation of the stream, the rule the program generator uses (spec 01 section 1.4.6).
use crate::memhard::ITEM_ROUNDS;
use crate::seed::SplitMix64;
/// Registers of the item state (the item is 16 words).
pub const DERIVE_REGS: usize = 16;
/// Round programs per item: one before each cache read and one after the last (`ITEM_ROUNDS + 1`).
pub const DERIVE_PROGRAMS: usize = ITEM_ROUNDS + 1;
/// Instructions per round program for the x8-equivalent operation count (the candidate, class "dr736"): 9 x 736
/// = 6,624 instructions per item at a mean of 1.62 GPU operations (1.52 chip operations) each, about 10,730 GPU
/// operations, 10,070 chip operations and 1,460 multiplies per item. The x8 mixer, counted from the code
/// (`memhard::mixer`, 16 x (xor, add, mul) + 8 quarter rounds x 12 = 144 operations as written, 128 with the
/// `RC + rk` adds hoisted as constants, 16 multiplies; `chip-model-v3.md` section 1 prices 130 from the spec text):
/// 72 applications = 10,368 as written, 9,216 hoisted, 1,152 multiplies. The floors below are those three.
pub const DERIVE_LEN_X8: u32 = 736;
/// Draws per instruction: the op roll, the destination roll, the second-source roll and the immediate.
pub const DRAWS_PER_INSTR: u64 = 4;
/// The floor of chip operations per item (the x8 mixer with its constants hoisted: 72 x 128).
pub const OPS_FLOOR_X8: u64 = 9_216;
/// The floor of GPU operations per item (the x8 mixer as written: 72 x 144).
pub const GPU_OPS_FLOOR_X8: u64 = 10_368;
/// The floor of multiplies per item (the x8 mixer's 72 x 16).
pub const MULS_FLOOR_X8: u64 = 1_152;
/// Distinct rotation amounts an item's programs must use, at least.
pub const DISTINCT_ROTS_FLOOR: usize = 8;
/// Attempts before the generator gives up (never reached: see [`DeriveProgram::draw`]).
pub const MAX_ATTEMPTS: u32 = 64;
/// The twelve forms. `c` is the chain register (the previous destination), `d` the destination (`d != c`), `b` a
/// third register (`b != d`, `b != c`), `k` a rotation in 1..31, `i` a 32-bit constant (odd for `MulC`).
#[derive(Clone, Copy, Debug, PartialEq, Eq, Hash)]
#[repr(u8)]
pub enum DOp {
/// `d += c`
Add = 0,
/// `d -= c`
Sub = 1,
/// `d ^= c`
Xor = 2,
/// `d *= (c OR 1)`: the multiply-lo form, odd so it is a bijection on `d`
Mul = 3,
/// `d = rotl(d, k) + c`
Rot = 4,
/// `d = rotl(d ^ c, k)`
XRot = 5,
/// `d += c + i`
AddC = 6,
/// `d ^= c ^ i`
XorC = 7,
/// `d = (d ^ c) * i`, `i` odd: the per-word form of `M_r` with the chain in place of the round constant
MulC = 8,
/// `d = d * i + c`, `i` odd
MulC2 = 9,
/// `d ^= (c AND b)`
AndX = 10,
/// `d += (c OR b)`
OrX = 11,
}
/// The op weights in percent, in draw order (sum 100). Fixed at genesis; only the order, the registers and the
/// constants are drawn.
pub const DOP_WEIGHTS: [(DOp, u64); 12] = [
(DOp::Add, 14),
(DOp::Sub, 10),
(DOp::Xor, 14),
(DOp::Mul, 10),
(DOp::Rot, 10),
(DOp::XRot, 10),
(DOp::AddC, 6),
(DOp::XorC, 6),
(DOp::MulC, 8),
(DOp::MulC2, 4),
(DOp::AndX, 4),
(DOp::OrX, 4),
];
impl DOp {
pub fn from_u8(v: u8) -> Option<DOp> {
DOP_WEIGHTS.iter().map(|(o, _)| *o).find(|o| *o as u8 == v)
}
pub fn name(self) -> &'static str {
match self {
DOp::Add => "add",
DOp::Sub => "sub",
DOp::Xor => "xor",
DOp::Mul => "mul",
DOp::Rot => "rot",
DOp::XRot => "xrot",
DOp::AddC => "addc",
DOp::XorC => "xorc",
DOp::MulC => "mulc",
DOp::MulC2 => "mulc2",
DOp::AndX => "andx",
DOp::OrX => "orx",
}
}
/// Integer operations as a GPU executes the form (every `|`, `&`, `+`, `^`, `*`, rotate counts one).
pub fn gpu_ops(self) -> u64 {
match self {
DOp::Add | DOp::Sub | DOp::Xor => 1,
_ => 2,
}
}
/// Integer operations as the chip model counts them (`c OR 1` is a wire on a chip, so `Mul` is one multiply;
/// a constant folded into a chain value is still an add or an xor, so every other two-op form stays two).
pub fn chip_ops(self) -> u64 {
match self {
DOp::Add | DOp::Sub | DOp::Xor | DOp::Mul => 1,
_ => 2,
}
}
pub fn is_mul(self) -> bool {
matches!(self, DOp::Mul | DOp::MulC | DOp::MulC2)
}
pub fn has_rot(self) -> bool {
matches!(self, DOp::Rot | DOp::XRot)
}
pub fn has_third(self) -> bool {
matches!(self, DOp::AndX | DOp::OrX)
}
pub fn has_imm(self) -> bool {
matches!(self, DOp::AddC | DOp::XorC | DOp::MulC | DOp::MulC2)
}
/// The op of a roll in 0..99.
pub fn for_roll(roll: u64) -> DOp {
let mut acc = 0u64;
for (op, w) in DOP_WEIGHTS {
acc += w;
if roll < acc {
return op;
}
}
DOp::OrX
}
}
/// One instruction. `src` is the chain register (carried so the interpreter and the emitter need no state).
#[derive(Clone, Copy, Debug, PartialEq, Eq, Hash)]
pub struct DInstr {
pub op: DOp,
pub dst: u8,
pub src: u8,
/// The third register of `AndX` and `OrX`; 0 on every other form (drawn and unused).
pub src2: u8,
/// The rotation 1..31 of `Rot` and `XRot`, else 0.
pub rot: u8,
/// The constant of `AddC`, `XorC` (any), `MulC` and `MulC2` (odd); 0 on every other form.
pub imm: u32,
}
/// The nine round programs of an item for one day, with the attempt that passed the acceptance test.
#[derive(Clone, Debug, PartialEq, Eq)]
pub struct DeriveProgram {
pub len: u32,
pub attempt: u32,
pub rounds: Vec<Vec<DInstr>>,
}
/// Why a candidate was rejected.
#[derive(Clone, Debug, PartialEq, Eq)]
pub enum DeriveReject {
/// A register no instruction of round program `round` writes.
RegisterNeverWritten { round: usize, reg: u8 },
/// Fewer than [`DISTINCT_ROTS_FLOOR`] distinct rotation amounts over the item's programs.
RotationsDegenerate { distinct: usize },
/// Chip operations per item under [`OPS_FLOOR_X8`] scaled to the length.
OpsUnderFloor { ops: u64, floor: u64 },
/// GPU operations per item under [`GPU_OPS_FLOOR_X8`] scaled to the length.
GpuOpsUnderFloor { ops: u64, floor: u64 },
/// Multiplies per item under [`MULS_FLOOR_X8`] scaled to the length.
MulsUnderFloor { muls: u64, floor: u64 },
}
impl std::fmt::Display for DeriveReject {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
match self {
DeriveReject::RegisterNeverWritten { round, reg } => write!(f, "register {reg} never written in round program {round}"),
DeriveReject::RotationsDegenerate { distinct } => write!(f, "only {distinct} distinct rotation amounts"),
DeriveReject::OpsUnderFloor { ops, floor } => write!(f, "{ops} chip operations per item, floor {floor}"),
DeriveReject::GpuOpsUnderFloor { ops, floor } => write!(f, "{ops} GPU operations per item, floor {floor}"),
DeriveReject::MulsUnderFloor { muls, floor } => write!(f, "{muls} multiplies per item, floor {floor}"),
}
}
}
impl DeriveProgram {
/// Draw one candidate of `len` instructions per round program from `rng` (four draws per instruction).
pub fn draw_candidate(rng: &mut SplitMix64, len: u32, attempt: u32) -> DeriveProgram {
let mut rounds = Vec::with_capacity(DERIVE_PROGRAMS);
for _ in 0..DERIVE_PROGRAMS {
let mut prog = Vec::with_capacity(len as usize);
let mut chain = 0u8;
for _ in 0..len {
let op = DOp::for_roll(rng.below(100));
// the destination: the 15 registers other than the chain, in ascending order
let d_roll = rng.below((DERIVE_REGS - 1) as u64) as u8;
let dst = if d_roll >= chain { d_roll + 1 } else { d_roll };
// the third register: the 14 registers other than dst and the chain, in ascending order
let b_roll = rng.below((DERIVE_REGS - 2) as u64) as u8;
let (lo, hi) = if dst < chain { (dst, chain) } else { (chain, dst) };
let mut b = b_roll;
if b >= lo {
b += 1;
}
if b >= hi {
b += 1;
}
let x = rng.next() as u32;
let mut ins = DInstr { op, dst, src: chain, src2: 0, rot: 0, imm: 0 };
if op.has_third() {
ins.src2 = b;
}
if op.has_rot() {
ins.rot = 1 + (x % 31) as u8;
}
if op.has_imm() {
ins.imm = if matches!(op, DOp::MulC | DOp::MulC2) { x | 1 } else { x };
}
prog.push(ins);
chain = dst;
}
rounds.push(prog);
}
DeriveProgram { len, attempt, rounds }
}
/// Draw the program of a day: candidates from `rng` in turn until one passes [`DeriveProgram::check`].
/// Panics after [`MAX_ATTEMPTS`] (the floors sit more than 7 standard deviations under the expected counts, so
/// a rejection is a rare event and 64 in a row is not one that happens).
pub fn draw(rng: &mut SplitMix64, len: u32) -> DeriveProgram {
for attempt in 0..MAX_ATTEMPTS {
let p = Self::draw_candidate(rng, len, attempt);
if p.check().is_ok() {
return p;
}
}
panic!("derivation program: {MAX_ATTEMPTS} candidates rejected in a row");
}
/// The floors for this length (chip operations, GPU operations, multiplies): the x8 floors scaled by
/// `len / DERIVE_LEN_X8`, so a shorter class, measured as a fallback, has its own proportional floors.
pub fn floors(len: u32) -> (u64, u64, u64) {
let scale = |f: u64| f * len as u64 / DERIVE_LEN_X8 as u64;
(scale(OPS_FLOOR_X8), scale(GPU_OPS_FLOOR_X8), scale(MULS_FLOOR_X8))
}
/// The acceptance test: every register written in every round program; at least [`DISTINCT_ROTS_FLOOR`]
/// distinct rotation amounts; chip operations, GPU operations and multiplies per item at or above the floors
/// (the x8 mixer's counts from the code). The structural
/// rules (`dst != src`, the third register distinct, rotations in 1..31, odd multiplier constants) hold by
/// construction and are asserted.
pub fn check(&self) -> Result<(), DeriveReject> {
let mut rots = [false; 32];
for (r, prog) in self.rounds.iter().enumerate() {
let mut written = [false; DERIVE_REGS];
let mut chain = 0u8;
for ins in prog {
assert!(ins.src == chain && ins.dst != ins.src && (ins.dst as usize) < DERIVE_REGS, "chain rule");
if ins.op.has_third() {
assert!(ins.src2 != ins.dst && ins.src2 != ins.src && (ins.src2 as usize) < DERIVE_REGS, "third register");
}
if ins.op.has_rot() {
assert!((1..=31).contains(&ins.rot), "rotation");
rots[ins.rot as usize] = true;
}
if matches!(ins.op, DOp::MulC | DOp::MulC2) {
assert!(ins.imm & 1 == 1, "odd multiplier");
}
written[ins.dst as usize] = true;
chain = ins.dst;
}
if let Some(reg) = written.iter().position(|w| !w) {
return Err(DeriveReject::RegisterNeverWritten { round: r, reg: reg as u8 });
}
}
let distinct = rots.iter().filter(|r| **r).count();
if distinct < DISTINCT_ROTS_FLOOR {
return Err(DeriveReject::RotationsDegenerate { distinct });
}
let (ops_floor, gpu_floor, muls_floor) = Self::floors(self.len);
let ops = self.chip_ops();
if ops < ops_floor {
return Err(DeriveReject::OpsUnderFloor { ops, floor: ops_floor });
}
let gpu = self.gpu_ops();
if gpu < gpu_floor {
return Err(DeriveReject::GpuOpsUnderFloor { ops: gpu, floor: gpu_floor });
}
let muls = self.muls();
if muls < muls_floor {
return Err(DeriveReject::MulsUnderFloor { muls, floor: muls_floor });
}
Ok(())
}
pub fn instr_count(&self) -> u64 {
self.rounds.iter().map(|p| p.len() as u64).sum()
}
pub fn gpu_ops(&self) -> u64 {
self.rounds.iter().flatten().map(|i| i.op.gpu_ops()).sum()
}
pub fn chip_ops(&self) -> u64 {
self.rounds.iter().flatten().map(|i| i.op.chip_ops()).sum()
}
pub fn muls(&self) -> u64 {
self.rounds.iter().flatten().filter(|i| i.op.is_mul()).count() as u64
}
/// Count per op, in [`DOP_WEIGHTS`] order.
pub fn op_counts(&self) -> [u64; 12] {
let mut c = [0u64; 12];
for i in self.rounds.iter().flatten() {
c[i.op as usize] += 1;
}
c
}
/// "add=887 sub=..." in weight order.
pub fn op_mix(&self) -> String {
let c = self.op_counts();
DOP_WEIGHTS.iter().map(|(o, _)| format!("{}={}", o.name(), c[*o as usize])).collect::<Vec<_>>().join(" ")
}
/// FNV-1a 64 over the instruction stream (op, dst, src, src2, rot, imm as bytes): the program's fingerprint
/// for packs and logs.
pub fn fingerprint(&self) -> u64 {
let mut b = Vec::with_capacity(self.instr_count() as usize * 9);
for i in self.rounds.iter().flatten() {
b.push(i.op as u8);
b.push(i.dst);
b.push(i.src);
b.push(i.src2);
b.push(i.rot);
b.extend_from_slice(&i.imm.to_le_bytes());
}
crate::seed::fnv1a64(&b)
}
}
/// Lanes of the SoA interpreter: the verifier derives up to 32 distinct items per load (one per lane of the
/// unit), so each instruction runs across 32 item states at once and the dispatch is paid once per 32 items.
pub const SOA_LANES: usize = 32;
/// The item states of a batch, word-major: `st[reg][lane]`.
pub type SoaState = [[u32; SOA_LANES]; DERIVE_REGS];
#[inline(always)]
fn rotl(x: u32, n: u32) -> u32 {
x.rotate_left(n)
}
/// The twelve forms over a batch, one function each, every one a straight loop over the lanes the compiler
/// vectorises. The destination row and the source rows are distinct by the chain rule (`dst != src`, and the third
/// register distinct from both: asserted by [`DeriveProgram::check`] and checked here in debug builds), so the
/// rows are addressed through raw pointers rather than copied out of the state.
mod forms {
use super::{rotl, DInstr, SoaState, SOA_LANES};
#[inline(always)]
pub fn add(ins: &DInstr, st: &mut SoaState) {
debug_assert!(ins.dst != ins.src);
// SAFETY: dst, src (and src2) are distinct registers below DERIVE_REGS, so the rows do not alias and the
// pointers stay inside `st`.
unsafe {
let dp = st.as_mut_ptr().add(ins.dst as usize) as *mut u32;
let cp = st.as_ptr().add(ins.src as usize) as *const u32;
for k in 0..SOA_LANES {
let d = &mut *dp.add(k);
let c = &*cp.add(k);
*d = d.wrapping_add(*c);
}
}
}
#[inline(always)]
pub fn sub(ins: &DInstr, st: &mut SoaState) {
debug_assert!(ins.dst != ins.src);
// SAFETY: dst, src (and src2) are distinct registers below DERIVE_REGS, so the rows do not alias and the
// pointers stay inside `st`.
unsafe {
let dp = st.as_mut_ptr().add(ins.dst as usize) as *mut u32;
let cp = st.as_ptr().add(ins.src as usize) as *const u32;
for k in 0..SOA_LANES {
let d = &mut *dp.add(k);
let c = &*cp.add(k);
*d = d.wrapping_sub(*c);
}
}
}
#[inline(always)]
pub fn xor(ins: &DInstr, st: &mut SoaState) {
debug_assert!(ins.dst != ins.src);
// SAFETY: dst, src (and src2) are distinct registers below DERIVE_REGS, so the rows do not alias and the
// pointers stay inside `st`.
unsafe {
let dp = st.as_mut_ptr().add(ins.dst as usize) as *mut u32;
let cp = st.as_ptr().add(ins.src as usize) as *const u32;
for k in 0..SOA_LANES {
let d = &mut *dp.add(k);
let c = &*cp.add(k);
*d ^= *c;
}
}
}
#[inline(always)]
pub fn mul(ins: &DInstr, st: &mut SoaState) {
debug_assert!(ins.dst != ins.src);
// SAFETY: dst, src (and src2) are distinct registers below DERIVE_REGS, so the rows do not alias and the
// pointers stay inside `st`.
unsafe {
let dp = st.as_mut_ptr().add(ins.dst as usize) as *mut u32;
let cp = st.as_ptr().add(ins.src as usize) as *const u32;
for k in 0..SOA_LANES {
let d = &mut *dp.add(k);
let c = &*cp.add(k);
*d = d.wrapping_mul(*c | 1);
}
}
}
#[inline(always)]
pub fn rot(ins: &DInstr, st: &mut SoaState) {
debug_assert!(ins.dst != ins.src);
let r = ins.rot as u32;
// SAFETY: dst, src (and src2) are distinct registers below DERIVE_REGS, so the rows do not alias and the
// pointers stay inside `st`.
unsafe {
let dp = st.as_mut_ptr().add(ins.dst as usize) as *mut u32;
let cp = st.as_ptr().add(ins.src as usize) as *const u32;
for k in 0..SOA_LANES {
let d = &mut *dp.add(k);
let c = &*cp.add(k);
*d = rotl(*d, r).wrapping_add(*c);
}
}
}
#[inline(always)]
pub fn xrot(ins: &DInstr, st: &mut SoaState) {
debug_assert!(ins.dst != ins.src);
let r = ins.rot as u32;
// SAFETY: dst, src (and src2) are distinct registers below DERIVE_REGS, so the rows do not alias and the
// pointers stay inside `st`.
unsafe {
let dp = st.as_mut_ptr().add(ins.dst as usize) as *mut u32;
let cp = st.as_ptr().add(ins.src as usize) as *const u32;
for k in 0..SOA_LANES {
let d = &mut *dp.add(k);
let c = &*cp.add(k);
*d = rotl(*d ^ *c, r);
}
}
}
#[inline(always)]
pub fn addc(ins: &DInstr, st: &mut SoaState) {
debug_assert!(ins.dst != ins.src);
let i = ins.imm;
// SAFETY: dst, src (and src2) are distinct registers below DERIVE_REGS, so the rows do not alias and the
// pointers stay inside `st`.
unsafe {
let dp = st.as_mut_ptr().add(ins.dst as usize) as *mut u32;
let cp = st.as_ptr().add(ins.src as usize) as *const u32;
for k in 0..SOA_LANES {
let d = &mut *dp.add(k);
let c = &*cp.add(k);
*d = d.wrapping_add(c.wrapping_add(i));
}
}
}
#[inline(always)]
pub fn xorc(ins: &DInstr, st: &mut SoaState) {
debug_assert!(ins.dst != ins.src);
let i = ins.imm;
// SAFETY: dst, src (and src2) are distinct registers below DERIVE_REGS, so the rows do not alias and the
// pointers stay inside `st`.
unsafe {
let dp = st.as_mut_ptr().add(ins.dst as usize) as *mut u32;
let cp = st.as_ptr().add(ins.src as usize) as *const u32;
for k in 0..SOA_LANES {
let d = &mut *dp.add(k);
let c = &*cp.add(k);
*d ^= *c ^ i;
}
}
}
#[inline(always)]
pub fn mulc(ins: &DInstr, st: &mut SoaState) {
debug_assert!(ins.dst != ins.src);
let i = ins.imm;
// SAFETY: dst, src (and src2) are distinct registers below DERIVE_REGS, so the rows do not alias and the
// pointers stay inside `st`.
unsafe {
let dp = st.as_mut_ptr().add(ins.dst as usize) as *mut u32;
let cp = st.as_ptr().add(ins.src as usize) as *const u32;
for k in 0..SOA_LANES {
let d = &mut *dp.add(k);
let c = &*cp.add(k);
*d = (*d ^ *c).wrapping_mul(i);
}
}
}
#[inline(always)]
pub fn mulc2(ins: &DInstr, st: &mut SoaState) {
debug_assert!(ins.dst != ins.src);
let i = ins.imm;
// SAFETY: dst, src (and src2) are distinct registers below DERIVE_REGS, so the rows do not alias and the
// pointers stay inside `st`.
unsafe {
let dp = st.as_mut_ptr().add(ins.dst as usize) as *mut u32;
let cp = st.as_ptr().add(ins.src as usize) as *const u32;
for k in 0..SOA_LANES {
let d = &mut *dp.add(k);
let c = &*cp.add(k);
*d = d.wrapping_mul(i).wrapping_add(*c);
}
}
}
#[inline(always)]
pub fn andx(ins: &DInstr, st: &mut SoaState) {
debug_assert!(ins.dst != ins.src && ins.src2 != ins.dst && ins.src2 != ins.src);
// SAFETY: dst, src (and src2) are distinct registers below DERIVE_REGS, so the rows do not alias and the
// pointers stay inside `st`.
unsafe {
let dp = st.as_mut_ptr().add(ins.dst as usize) as *mut u32;
let cp = st.as_ptr().add(ins.src as usize) as *const u32;
let bp = st.as_ptr().add(ins.src2 as usize) as *const u32;
for k in 0..SOA_LANES {
let d = &mut *dp.add(k);
let c = &*cp.add(k);
let b = &*bp.add(k);
*d ^= *c & *b;
}
}
}
#[inline(always)]
pub fn orx(ins: &DInstr, st: &mut SoaState) {
debug_assert!(ins.dst != ins.src && ins.src2 != ins.dst && ins.src2 != ins.src);
// SAFETY: dst, src (and src2) are distinct registers below DERIVE_REGS, so the rows do not alias and the
// pointers stay inside `st`.
unsafe {
let dp = st.as_mut_ptr().add(ins.dst as usize) as *mut u32;
let cp = st.as_ptr().add(ins.src as usize) as *const u32;
let bp = st.as_ptr().add(ins.src2 as usize) as *const u32;
for k in 0..SOA_LANES {
let d = &mut *dp.add(k);
let c = &*cp.add(k);
let b = &*bp.add(k);
*d = d.wrapping_add(*c | *b);
}
}
}
}
/// One instruction over the batch (the single dispatch; [`run_round`] dispatches on pairs).
#[inline(always)]
pub fn run_instr(ins: &DInstr, st: &mut SoaState) {
match ins.op {
DOp::Add => forms::add(ins, st),
DOp::Sub => forms::sub(ins, st),
DOp::Xor => forms::xor(ins, st),
DOp::Mul => forms::mul(ins, st),
DOp::Rot => forms::rot(ins, st),
DOp::XRot => forms::xrot(ins, st),
DOp::AddC => forms::addc(ins, st),
DOp::XorC => forms::xorc(ins, st),
DOp::MulC => forms::mulc(ins, st),
DOp::MulC2 => forms::mulc2(ins, st),
DOp::AndX => forms::andx(ins, st),
DOp::OrX => forms::orx(ins, st),
}
}
/// Run one round program over the batch. The dispatch is on PAIRS of instructions (144 arms, one indirect branch
/// per two instructions): the op sequence of a drawn program is random, so the branch predictor misses most
/// dispatches, and the miss (about 3.7 ns of the 7.2 ns an instruction cost per batch on one M5 Max core, measured
/// with `examples/derive_perf.rs` on 6 October 2026, a functional run) is paid once per pair instead of once per
/// instruction. The result is bit for bit that of [`run_instr`] in sequence.
#[inline(never)]
pub fn run_round(prog: &[DInstr], st: &mut SoaState) {
let mut it = prog.chunks_exact(2);
for pair in &mut it {
let (a, b) = (&pair[0], &pair[1]);
match (a.op as u8) * 12 + b.op as u8 {
0 => { forms::add(a, st); forms::add(b, st); }
1 => { forms::add(a, st); forms::sub(b, st); }
2 => { forms::add(a, st); forms::xor(b, st); }
3 => { forms::add(a, st); forms::mul(b, st); }
4 => { forms::add(a, st); forms::rot(b, st); }
5 => { forms::add(a, st); forms::xrot(b, st); }
6 => { forms::add(a, st); forms::addc(b, st); }
7 => { forms::add(a, st); forms::xorc(b, st); }
8 => { forms::add(a, st); forms::mulc(b, st); }
9 => { forms::add(a, st); forms::mulc2(b, st); }
10 => { forms::add(a, st); forms::andx(b, st); }
11 => { forms::add(a, st); forms::orx(b, st); }
12 => { forms::sub(a, st); forms::add(b, st); }
13 => { forms::sub(a, st); forms::sub(b, st); }
14 => { forms::sub(a, st); forms::xor(b, st); }
15 => { forms::sub(a, st); forms::mul(b, st); }
16 => { forms::sub(a, st); forms::rot(b, st); }
17 => { forms::sub(a, st); forms::xrot(b, st); }
18 => { forms::sub(a, st); forms::addc(b, st); }
19 => { forms::sub(a, st); forms::xorc(b, st); }
20 => { forms::sub(a, st); forms::mulc(b, st); }
21 => { forms::sub(a, st); forms::mulc2(b, st); }
22 => { forms::sub(a, st); forms::andx(b, st); }
23 => { forms::sub(a, st); forms::orx(b, st); }
24 => { forms::xor(a, st); forms::add(b, st); }
25 => { forms::xor(a, st); forms::sub(b, st); }
26 => { forms::xor(a, st); forms::xor(b, st); }
27 => { forms::xor(a, st); forms::mul(b, st); }
28 => { forms::xor(a, st); forms::rot(b, st); }
29 => { forms::xor(a, st); forms::xrot(b, st); }
30 => { forms::xor(a, st); forms::addc(b, st); }
31 => { forms::xor(a, st); forms::xorc(b, st); }
32 => { forms::xor(a, st); forms::mulc(b, st); }
33 => { forms::xor(a, st); forms::mulc2(b, st); }
34 => { forms::xor(a, st); forms::andx(b, st); }
35 => { forms::xor(a, st); forms::orx(b, st); }
36 => { forms::mul(a, st); forms::add(b, st); }
37 => { forms::mul(a, st); forms::sub(b, st); }
38 => { forms::mul(a, st); forms::xor(b, st); }
39 => { forms::mul(a, st); forms::mul(b, st); }
40 => { forms::mul(a, st); forms::rot(b, st); }
41 => { forms::mul(a, st); forms::xrot(b, st); }
42 => { forms::mul(a, st); forms::addc(b, st); }
43 => { forms::mul(a, st); forms::xorc(b, st); }
44 => { forms::mul(a, st); forms::mulc(b, st); }
45 => { forms::mul(a, st); forms::mulc2(b, st); }
46 => { forms::mul(a, st); forms::andx(b, st); }
47 => { forms::mul(a, st); forms::orx(b, st); }
48 => { forms::rot(a, st); forms::add(b, st); }
49 => { forms::rot(a, st); forms::sub(b, st); }
50 => { forms::rot(a, st); forms::xor(b, st); }
51 => { forms::rot(a, st); forms::mul(b, st); }
52 => { forms::rot(a, st); forms::rot(b, st); }
53 => { forms::rot(a, st); forms::xrot(b, st); }
54 => { forms::rot(a, st); forms::addc(b, st); }
55 => { forms::rot(a, st); forms::xorc(b, st); }
56 => { forms::rot(a, st); forms::mulc(b, st); }
57 => { forms::rot(a, st); forms::mulc2(b, st); }
58 => { forms::rot(a, st); forms::andx(b, st); }
59 => { forms::rot(a, st); forms::orx(b, st); }
60 => { forms::xrot(a, st); forms::add(b, st); }
61 => { forms::xrot(a, st); forms::sub(b, st); }
62 => { forms::xrot(a, st); forms::xor(b, st); }
63 => { forms::xrot(a, st); forms::mul(b, st); }
64 => { forms::xrot(a, st); forms::rot(b, st); }
65 => { forms::xrot(a, st); forms::xrot(b, st); }
66 => { forms::xrot(a, st); forms::addc(b, st); }
67 => { forms::xrot(a, st); forms::xorc(b, st); }
68 => { forms::xrot(a, st); forms::mulc(b, st); }
69 => { forms::xrot(a, st); forms::mulc2(b, st); }
70 => { forms::xrot(a, st); forms::andx(b, st); }
71 => { forms::xrot(a, st); forms::orx(b, st); }
72 => { forms::addc(a, st); forms::add(b, st); }
73 => { forms::addc(a, st); forms::sub(b, st); }
74 => { forms::addc(a, st); forms::xor(b, st); }
75 => { forms::addc(a, st); forms::mul(b, st); }
76 => { forms::addc(a, st); forms::rot(b, st); }
77 => { forms::addc(a, st); forms::xrot(b, st); }
78 => { forms::addc(a, st); forms::addc(b, st); }
79 => { forms::addc(a, st); forms::xorc(b, st); }
80 => { forms::addc(a, st); forms::mulc(b, st); }
81 => { forms::addc(a, st); forms::mulc2(b, st); }
82 => { forms::addc(a, st); forms::andx(b, st); }
83 => { forms::addc(a, st); forms::orx(b, st); }
84 => { forms::xorc(a, st); forms::add(b, st); }
85 => { forms::xorc(a, st); forms::sub(b, st); }
86 => { forms::xorc(a, st); forms::xor(b, st); }
87 => { forms::xorc(a, st); forms::mul(b, st); }
88 => { forms::xorc(a, st); forms::rot(b, st); }
89 => { forms::xorc(a, st); forms::xrot(b, st); }
90 => { forms::xorc(a, st); forms::addc(b, st); }
91 => { forms::xorc(a, st); forms::xorc(b, st); }
92 => { forms::xorc(a, st); forms::mulc(b, st); }
93 => { forms::xorc(a, st); forms::mulc2(b, st); }
94 => { forms::xorc(a, st); forms::andx(b, st); }
95 => { forms::xorc(a, st); forms::orx(b, st); }
96 => { forms::mulc(a, st); forms::add(b, st); }
97 => { forms::mulc(a, st); forms::sub(b, st); }
98 => { forms::mulc(a, st); forms::xor(b, st); }
99 => { forms::mulc(a, st); forms::mul(b, st); }
100 => { forms::mulc(a, st); forms::rot(b, st); }
101 => { forms::mulc(a, st); forms::xrot(b, st); }
102 => { forms::mulc(a, st); forms::addc(b, st); }
103 => { forms::mulc(a, st); forms::xorc(b, st); }
104 => { forms::mulc(a, st); forms::mulc(b, st); }
105 => { forms::mulc(a, st); forms::mulc2(b, st); }
106 => { forms::mulc(a, st); forms::andx(b, st); }
107 => { forms::mulc(a, st); forms::orx(b, st); }
108 => { forms::mulc2(a, st); forms::add(b, st); }
109 => { forms::mulc2(a, st); forms::sub(b, st); }
110 => { forms::mulc2(a, st); forms::xor(b, st); }
111 => { forms::mulc2(a, st); forms::mul(b, st); }
112 => { forms::mulc2(a, st); forms::rot(b, st); }
113 => { forms::mulc2(a, st); forms::xrot(b, st); }
114 => { forms::mulc2(a, st); forms::addc(b, st); }
115 => { forms::mulc2(a, st); forms::xorc(b, st); }
116 => { forms::mulc2(a, st); forms::mulc(b, st); }
117 => { forms::mulc2(a, st); forms::mulc2(b, st); }
118 => { forms::mulc2(a, st); forms::andx(b, st); }
119 => { forms::mulc2(a, st); forms::orx(b, st); }
120 => { forms::andx(a, st); forms::add(b, st); }
121 => { forms::andx(a, st); forms::sub(b, st); }
122 => { forms::andx(a, st); forms::xor(b, st); }
123 => { forms::andx(a, st); forms::mul(b, st); }
124 => { forms::andx(a, st); forms::rot(b, st); }
125 => { forms::andx(a, st); forms::xrot(b, st); }
126 => { forms::andx(a, st); forms::addc(b, st); }
127 => { forms::andx(a, st); forms::xorc(b, st); }
128 => { forms::andx(a, st); forms::mulc(b, st); }
129 => { forms::andx(a, st); forms::mulc2(b, st); }
130 => { forms::andx(a, st); forms::andx(b, st); }
131 => { forms::andx(a, st); forms::orx(b, st); }
132 => { forms::orx(a, st); forms::add(b, st); }
133 => { forms::orx(a, st); forms::sub(b, st); }
134 => { forms::orx(a, st); forms::xor(b, st); }
135 => { forms::orx(a, st); forms::mul(b, st); }
136 => { forms::orx(a, st); forms::rot(b, st); }
137 => { forms::orx(a, st); forms::xrot(b, st); }
138 => { forms::orx(a, st); forms::addc(b, st); }
139 => { forms::orx(a, st); forms::xorc(b, st); }
140 => { forms::orx(a, st); forms::mulc(b, st); }
141 => { forms::orx(a, st); forms::mulc2(b, st); }
142 => { forms::orx(a, st); forms::andx(b, st); }
143 => { forms::orx(a, st); forms::orx(b, st); }
_ => unreachable!(),
}
}
for ins in it.remainder() {
run_instr(ins, st);
}
}
/// The scalar reference: one instruction on one 16-word state, the text the kernels carry (`emit.rs`,
/// `derive_instr_text`) restated in Rust. The tests pin the SoA interpreter against it.
pub fn run_round_scalar(prog: &[DInstr], s: &mut [u32; DERIVE_REGS]) {
for ins in prog {
let d = ins.dst as usize;
let c = s[ins.src as usize];
match ins.op {
DOp::Add => s[d] = s[d].wrapping_add(c),
DOp::Sub => s[d] = s[d].wrapping_sub(c),
DOp::Xor => s[d] ^= c,
DOp::Mul => s[d] = s[d].wrapping_mul(c | 1),
DOp::Rot => s[d] = rotl(s[d], ins.rot as u32).wrapping_add(c),
DOp::XRot => s[d] = rotl(s[d] ^ c, ins.rot as u32),
DOp::AddC => s[d] = s[d].wrapping_add(c.wrapping_add(ins.imm)),
DOp::XorC => s[d] ^= c ^ ins.imm,
DOp::MulC => s[d] = (s[d] ^ c).wrapping_mul(ins.imm),
DOp::MulC2 => s[d] = s[d].wrapping_mul(ins.imm).wrapping_add(c),
DOp::AndX => s[d] ^= c & s[ins.src2 as usize],
DOp::OrX => s[d] = s[d].wrapping_add(c | s[ins.src2 as usize]),
}
}
}
/// The source text of one instruction in the C-family dialects (the same text in Metal, CUDA C and OpenCL C:
/// `s` is the 16-word state, `mh_rotl` the rotate of the memhard core).
pub fn instr_text(ins: &DInstr) -> String {
let (d, c, b) = (ins.dst, ins.src, ins.src2);
match ins.op {
DOp::Add => format!("s[{d}] += s[{c}];"),
DOp::Sub => format!("s[{d}] -= s[{c}];"),
DOp::Xor => format!("s[{d}] ^= s[{c}];"),
DOp::Mul => format!("s[{d}] *= (s[{c}] | 1u);"),
DOp::Rot => format!("s[{d}] = mh_rotl(s[{d}], {}u) + s[{c}];", ins.rot),
DOp::XRot => format!("s[{d}] = mh_rotl(s[{d}] ^ s[{c}], {}u);", ins.rot),
DOp::AddC => format!("s[{d}] += s[{c}] + {:#010x}u;", ins.imm),
DOp::XorC => format!("s[{d}] ^= s[{c}] ^ {:#010x}u;", ins.imm),
DOp::MulC => format!("s[{d}] = (s[{d}] ^ s[{c}]) * {:#010x}u;", ins.imm),
DOp::MulC2 => format!("s[{d}] = s[{d}] * {:#010x}u + s[{c}];", ins.imm),
DOp::AndX => format!("s[{d}] ^= (s[{c}] & s[{b}]);"),
DOp::OrX => format!("s[{d}] += (s[{c}] | s[{b}]);"),
}
}
/// One instruction as a line of program.json: `"add d=3 c=0"`, `"mulc d=5 c=3 imm=0x..."`.
pub fn instr_line(ins: &DInstr) -> String {
let mut s = format!("{} d={} c={}", ins.op.name(), ins.dst, ins.src);
if ins.op.has_third() {
s.push_str(&format!(" b={}", ins.src2));
}
if ins.op.has_rot() {
s.push_str(&format!(" k={}", ins.rot));
}
if ins.op.has_imm() {
s.push_str(&format!(" imm={:#010x}", ins.imm));
}
s
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn weights_sum_and_ops() {
assert_eq!(DOP_WEIGHTS.iter().map(|(_, w)| w).sum::<u64>(), 100);
for (i, (op, _)) in DOP_WEIGHTS.iter().enumerate() {
assert_eq!(*op as usize, i);
assert_eq!(DOp::from_u8(i as u8), Some(*op));
}
assert_eq!(DOp::for_roll(0), DOp::Add);
assert_eq!(DOp::for_roll(13), DOp::Add);
assert_eq!(DOp::for_roll(14), DOp::Sub);
assert_eq!(DOp::for_roll(99), DOp::OrX);
// the expected chip operations per instruction, 1.52 (1.62 on a GPU), put 736 x 9 over the x8 floors
let mean: f64 = DOP_WEIGHTS.iter().map(|(o, w)| o.chip_ops() as f64 * *w as f64 / 100.0).sum();
assert!((mean - 1.52).abs() < 1e-9, "{mean}");
let gpu_mean: f64 = DOP_WEIGHTS.iter().map(|(o, w)| o.gpu_ops() as f64 * *w as f64 / 100.0).sum();
assert!((gpu_mean - 1.62).abs() < 1e-9, "{gpu_mean}");
assert!(9.0 * DERIVE_LEN_X8 as f64 * mean > OPS_FLOOR_X8 as f64);
assert!(9.0 * DERIVE_LEN_X8 as f64 * gpu_mean > GPU_OPS_FLOOR_X8 as f64);
assert_eq!(DeriveProgram::floors(DERIVE_LEN_X8), (9_216, 10_368, 1_152));
assert_eq!(DeriveProgram::floors(368), (4_608, 5_184, 576));
let mul_share: f64 = DOP_WEIGHTS.iter().filter(|(o, _)| o.is_mul()).map(|(_, w)| *w as f64 / 100.0).sum();
assert!(9.0 * DERIVE_LEN_X8 as f64 * mul_share > MULS_FLOOR_X8 as f64);
}
#[test]
fn draw_is_structural_and_accepted() {
let mut rng = SplitMix64::new(0x1234_5678_9abc_def0);
let p = DeriveProgram::draw(&mut rng, DERIVE_LEN_X8);
assert_eq!(p.attempt, 0, "the first candidate of this seed passes");
assert_eq!(p.rounds.len(), DERIVE_PROGRAMS);
assert_eq!(p.instr_count(), 9 * DERIVE_LEN_X8 as u64);
assert!(p.check().is_ok());
assert!(p.chip_ops() >= OPS_FLOOR_X8 && p.gpu_ops() >= GPU_OPS_FLOOR_X8 && p.muls() >= MULS_FLOOR_X8);
assert!(p.gpu_ops() > p.chip_ops());
// every instruction consumes the newest result
for prog in &p.rounds {
let mut chain = 0u8;
for ins in prog {
assert_eq!(ins.src, chain);
assert_ne!(ins.dst, chain);
chain = ins.dst;
}
}
// four draws per instruction: the same program again from the same seed, and a different one one draw on
let mut rng2 = SplitMix64::new(0x1234_5678_9abc_def0);
assert_eq!(DeriveProgram::draw(&mut rng2, DERIVE_LEN_X8), p);
let mut rng3 = SplitMix64::new(0x1234_5678_9abc_def0);
rng3.next();
assert_ne!(DeriveProgram::draw(&mut rng3, DERIVE_LEN_X8), p);
}
#[test]
fn soa_matches_scalar_and_is_a_bijection() {
let mut rng = SplitMix64::new(7);
let p = DeriveProgram::draw(&mut rng, 64);
let mut st: SoaState = [[0u32; SOA_LANES]; DERIVE_REGS];
let mut scalars = [[0u32; DERIVE_REGS]; SOA_LANES];
let mut x = SplitMix64::new(99);
for k in 0..SOA_LANES {
for r in 0..DERIVE_REGS {
let v = x.next() as u32;
st[r][k] = v;
scalars[k][r] = v;
}
}
let before = scalars;
for prog in &p.rounds {
run_round(prog, &mut st);
for k in 0..SOA_LANES {
run_round_scalar(prog, &mut scalars[k]);
}
}
for k in 0..SOA_LANES {
for r in 0..DERIVE_REGS {
assert_eq!(st[r][k], scalars[k][r], "lane {k} reg {r}");
}
}
// distinct inputs stay distinct (a bijection on the state, spot-checked: 32 lanes, no collision)
for a in 0..SOA_LANES {
for b in a + 1..SOA_LANES {
assert_ne!(scalars[a], scalars[b]);
assert_ne!(before[a], before[b]);
}
}
}
#[test]
fn acceptance_rejects_degenerate_draws() {
let mut rng = SplitMix64::new(3);
let mut p = DeriveProgram::draw(&mut rng, 64);
// a register never written: make every write of round 2 go to the chain's neighbour
let mut q = p.clone();
for ins in q.rounds[2].iter_mut() {
ins.dst = if ins.src == 1 { 2 } else { 1 };
}
let mut chain = 0u8;
for ins in q.rounds[2].iter_mut() {
ins.src = chain;
ins.dst = if chain == 1 { 2 } else { 1 };
chain = ins.dst;
}
assert!(matches!(q.check(), Err(DeriveReject::RegisterNeverWritten { round: 2, .. })));
// all rotations equal
let mut q = p.clone();
for ins in q.rounds.iter_mut().flatten() {
if ins.op.has_rot() {
ins.rot = 5;
}
}
assert!(matches!(q.check(), Err(DeriveReject::RotationsDegenerate { distinct: 1 })));
// every op an add apart from the xor-rotates (so the rotations stay distinct): under the ops floor
for ins in p.rounds.iter_mut().flatten() {
if ins.op != DOp::XRot {
ins.op = DOp::Add;
ins.rot = 0;
ins.imm = 0;
ins.src2 = 0;
}
}
assert!(matches!(p.check(), Err(DeriveReject::OpsUnderFloor { .. })), "{:?}", p.check());
// no multiplies at all but the ops floor met: under the multiply floor
let mut q = DeriveProgram::draw(&mut SplitMix64::new(11), 64);
for ins in q.rounds.iter_mut().flatten() {
if ins.op.is_mul() {
ins.op = DOp::AddC;
}
}
assert!(matches!(q.check(), Err(DeriveReject::MulsUnderFloor { .. })), "{:?}", q.check());
}
#[test]
fn text_forms() {
let i = DInstr { op: DOp::MulC, dst: 5, src: 3, src2: 0, rot: 0, imm: 0x9e37_79b9 };
assert_eq!(instr_text(&i), "s[5] = (s[5] ^ s[3]) * 0x9e3779b9u;");
assert_eq!(instr_line(&i), "mulc d=5 c=3 imm=0x9e3779b9");
let i = DInstr { op: DOp::XRot, dst: 0, src: 15, src2: 0, rot: 17, imm: 0 };
assert_eq!(instr_text(&i), "s[0] = mh_rotl(s[0] ^ s[15], 17u);");
let i = DInstr { op: DOp::AndX, dst: 2, src: 9, src2: 14, rot: 0, imm: 0 };
assert_eq!(instr_text(&i), "s[2] ^= (s[9] & s[14]);");
}
}

View file

@ -10,6 +10,7 @@
//! as a bare `0x003fffff`. The Swift writes it quoted (`jhex`), which is not valid JSON. //! as a bare `0x003fffff`. The Swift writes it quoted (`jhex`), which is not valid JSON.
use crate::generator::{EraParams, Instr, Op, Program, ProgramClass, GENERATOR_VERSION, INSTR_COUNT, ITERATIONS, LOAD_SLOTS}; use crate::generator::{EraParams, Instr, Op, Program, ProgramClass, GENERATOR_VERSION, INSTR_COUNT, ITERATIONS, LOAD_SLOTS};
use crate::derive::{instr_line as derive_instr_line, instr_text as derive_instr_text, DERIVE_PROGRAMS};
use crate::memhard::{ use crate::memhard::{
hot_key, hot_segments, hot_words, Layout, MixParams, Shape, CACHE_LINES_PER_SEGMENT, CACHE_SEGMENT_LOG2_LINES, CACHE_TAG, hot_key, hot_segments, hot_words, Layout, MixParams, Shape, CACHE_LINES_PER_SEGMENT, CACHE_SEGMENT_LOG2_LINES, CACHE_TAG,
CHACHA_ROUNDS, CHACHA_SIGMA, HOT_TAG, ITEM_ROUNDS, CHACHA_ROUNDS, CHACHA_SIGMA, HOT_TAG, ITEM_ROUNDS,
@ -195,6 +196,14 @@ fn class_header_lines(p: &Program) -> String {
", p.class.mixer_mult)); ", p.class.mixer_mult));
s.push_str(&format!("#define IGNEUM_CACHE_GROWTH {} // 1: cache words = 2^(26 + doublings(day)), doublings = floor(log2(1 + day / 1460)) s.push_str(&format!("#define IGNEUM_CACHE_GROWTH {} // 1: cache words = 2^(26 + doublings(day)), doublings = floor(log2(1 + day / 1460))
", p.class.growth as u8)); ", p.class.growth as u8));
}
if p.class.derive_len != 0 {
s.push_str("// Counter ASIC 3.0 item 2 (6 October 2026, docs/plans/counter-asic-3-derivation.md, a prototype, NOT class v3): the item
");
s.push_str("// derivation runs the day's drawn program (memhard.h: mh_round_0..8, IGNEUM_DERIVE_LEN instructions each) in place of the mixer.
");
s.push_str(&format!("#define IGNEUM_CLASS_DERIVE_LEN {}
", p.class.derive_len));
} }
s.push_str(&format!("#define IGNEUM_LOAD_SLOTS {} s.push_str(&format!("#define IGNEUM_LOAD_SLOTS {}
", p.class.load_slots)); ", p.class.load_slots));
@ -474,12 +483,18 @@ pub fn emit_memhard_core_layout(mp: &MixParams, dialect: CoreDialect, layout: La
CACHE_LINES_PER_SEGMENT, CACHE_LINES_PER_SEGMENT,
CHACHA_ROUNDS CHACHA_ROUNDS
)); ));
if m == 1 { if mp.derive.is_some() {
s.push_str("// Item: 8 rounds of (the day's round program + one 64-byte cache read), then round program 8. All parameters are literals.\n");
} else if m == 1 {
s.push_str("// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.\n"); s.push_str("// Item: 8 rounds of seed-parameterised mixer + one 64-byte cache read, then a final mixer. All parameters are literals.\n");
} else { } else {
s.push_str(&format!("// Item: 8 rounds of {m} x seed-parameterised mixer + one 64-byte cache read, then {m} x final mixer (class v3, mixer multiplier {m},\n")); s.push_str(&format!("// Item: 8 rounds of {m} x seed-parameterised mixer + one 64-byte cache read, then {m} x final mixer (class v3, mixer multiplier {m},\n"));
s.push_str("// docs/plans/mixer-x4.md: the round key of application j of round r is 0x9E3779B9 * (r * m + j + 1)). All parameters are literals.\n"); s.push_str("// docs/plans/mixer-x4.md: the round key of application j of round r is 0x9E3779B9 * (r * m + j + 1)). All parameters are literals.\n");
} }
if let Some(dp) = &mp.derive {
s.push_str(&format!("// Counter ASIC 3.0 item 2 (docs/plans/counter-asic-3-derivation.md): the mixer slots run the day's drawn program, {} instructions per round\n", dp.len));
s.push_str(&format!("// program (mh_round_0..{}), {} per item, drawn from the day key stream after the mixer constants (attempt {}, fingerprint {:016x}).\n", DERIVE_PROGRAMS - 1, dp.instr_count(), dp.attempt, dp.fingerprint()));
}
s.push_str(&format!("#define MH_CACHE_LINE_MASK {}\n", hex(cache_line_mask))); s.push_str(&format!("#define MH_CACHE_LINE_MASK {}\n", hex(cache_line_mask)));
s.push_str(&format!("#define MH_SEGMENT_LINES {}u\n", CACHE_LINES_PER_SEGMENT)); s.push_str(&format!("#define MH_SEGMENT_LINES {}u\n", CACHE_LINES_PER_SEGMENT));
s.push_str("#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }\n"); s.push_str("#define MH_QR(a, b, c, d, r1, r2, r3, r4) { a += b; d ^= a; d = mh_rotl(d, r1); c += d; b ^= c; b = mh_rotl(b, r2); a += b; d ^= a; d = mh_rotl(d, r3); c += d; b ^= c; b = mh_rotl(b, r4); }\n");
@ -559,6 +574,34 @@ pub fn emit_memhard_core_layout(mp: &MixParams, dialect: CoreDialect, layout: La
"// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); {ITEM_ROUNDS} rounds of {m} x mixer + cache line s[0] & mask; {m} x final mixer.\n" "// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); {ITEM_ROUNDS} rounds of {m} x mixer + cache line s[0] & mask; {m} x final mixer.\n"
)); ));
} }
if let Some(dp) = &mp.derive {
// the nine round programs as straight-line functions over the 16-word state (the same text in every dialect)
for (r, prog) in dp.rounds.iter().enumerate() {
s.push_str(&format!("// Round program {r}: {} instructions, chain rule (every instruction reads the register the previous one wrote; s[0] first).\n", prog.len()));
s.push_str(&format!("{fn_} void mh_round_{r}({lptr} s) {{\n"));
for ins in prog {
s.push_str(" ");
s.push_str(&derive_instr_text(ins));
s.push('\n');
}
s.push_str("}\n");
}
s.push_str(&format!("// Item t: 16 words. s = (K, t * MUL[i] + RC[i]); {ITEM_ROUNDS} rounds of (round program r, cache line s[0] & mask); round program {ITEM_ROUNDS}.\n"));
s.push_str(&format!("{fn_} void mh_item({cptr} cache, {u} t, {lptr} s) {{\n"));
for i in 0..8 {
s.push_str(&format!(" s[{i}] = {};\n", hex(k[i])));
}
for i in 0..8 {
s.push_str(&format!(" s[{}] = t * {} + {};\n", 8 + i, hex(mul[i]), hex(c[i])));
}
for r in 0..ITEM_ROUNDS {
s.push_str(&format!(" mh_round_{r}(s);\n"));
s.push_str(&format!(" {{ {cptr} line = cache + ((s[0] & MH_CACHE_LINE_MASK) * 16u); for ({u} i = 0u; i < 16u; ++i) s[i] ^= line[i]; }}\n"));
}
s.push_str(&format!(" mh_round_{ITEM_ROUNDS}(s);\n"));
s.push_str("}\n");
return finish_memhard_core(s, layout, u, fn_, cptr);
}
s.push_str(&format!("{fn_} void mh_item({cptr} cache, {u} t, {lptr} s) {{\n")); s.push_str(&format!("{fn_} void mh_item({cptr} cache, {u} t, {lptr} s) {{\n"));
for i in 0..8 { for i in 0..8 {
s.push_str(&format!(" s[{i}] = {};\n", hex(k[i]))); s.push_str(&format!(" s[{i}] = {};\n", hex(k[i])));
@ -584,6 +627,11 @@ pub fn emit_memhard_core_layout(mp: &MixParams, dialect: CoreDialect, layout: La
)); ));
} }
s.push_str("}\n"); s.push_str("}\n");
finish_memhard_core(s, layout, u, fn_, cptr)
}
/// The tail of the memhard core: `mh_word` (and the era layout helpers) after `mh_item`.
fn finish_memhard_core(mut s: String, layout: Layout, u: &str, fn_: &str, cptr: &str) -> String {
if layout.is_linear() { if layout.is_linear() {
s.push_str("// dataset[w] without the dataset: derive item w >> 4 and take word w & 15.\n"); s.push_str("// dataset[w] without the dataset: derive item w >> 4 and take word w & 15.\n");
s.push_str(&format!( s.push_str(&format!(
@ -1486,6 +1534,15 @@ pub fn program_header(p: &Program, day: &str, ds: &DatasetSource) -> String {
if mp.shape.mixer_mult != 1 { if mp.shape.mixer_mult != 1 {
s.push_str(&format!("#define IGNEUM_MIXER_MULT {} // mixer applications per round and after the last read (class v3, docs/plans/mixer-x4.md)\n", mp.shape.mixer_mult)); s.push_str(&format!("#define IGNEUM_MIXER_MULT {} // mixer applications per round and after the last read (class v3, docs/plans/mixer-x4.md)\n", mp.shape.mixer_mult));
} }
if let Some(dp) = &mp.derive {
s.push_str(&format!("#define IGNEUM_DERIVE_LEN {} // instructions per round program of the day's item-derivation program (Counter ASIC 3.0 item 2; memhard.h mh_round_0..8)\n", dp.len));
s.push_str(&format!("#define IGNEUM_DERIVE_ATTEMPT {}\n", dp.attempt));
s.push_str(&format!("#define IGNEUM_DERIVE_FINGERPRINT {}\n", hex64(dp.fingerprint())));
s.push_str(&format!("#define IGNEUM_DERIVE_INSTRS_PER_ITEM {}\n", dp.instr_count()));
s.push_str(&format!("#define IGNEUM_DERIVE_GPU_OPS_PER_ITEM {}\n", dp.gpu_ops()));
s.push_str(&format!("#define IGNEUM_DERIVE_CHIP_OPS_PER_ITEM {}\n", dp.chip_ops()));
s.push_str(&format!("#define IGNEUM_DERIVE_MULS_PER_ITEM {}\n", dp.muls()));
}
s.push_str(&format!( s.push_str(&format!(
"#define IGNEUM_MIX_ROT_INIT {{ {} }}\n", "#define IGNEUM_MIX_ROT_INIT {{ {} }}\n",
mp.rot.iter().map(|r| format!("{r}u")).collect::<Vec<_>>().join(", ") mp.rot.iter().map(|r| format!("{r}u")).collect::<Vec<_>>().join(", ")
@ -1694,6 +1751,10 @@ pub fn program_json(p: &Program, day: &str, ds: &DatasetSource) -> String {
if !p.class.is_v2() { if !p.class.is_v2() {
let c = p.width_counts(); let c = p.width_counts();
s.push_str(&format!(" \"load_class\": {},\n", jstr(&p.class.name()))); s.push_str(&format!(" \"load_class\": {},\n", jstr(&p.class.name())));
if p.class.derive_len != 0 {
s.push_str(&format!(" \"derive_len\": {},\n", p.class.derive_len));
s.push_str(" \"derive\": \"Counter ASIC 3.0 item 2 (6 October 2026, docs/plans/counter-asic-3-derivation.md; a prototype, not class v3): the nine mixer slots of the item derivation each run a straight-line program of derive_len instructions drawn from the day key stream after the 40 mixer draws, four draws per instruction (op roll below(100), destination roll below(15), third-register roll below(14), the immediate next()); every instruction reads the register the previous one wrote (s[0] first) and writes another; twelve forms, each a bijection on the state; the 8 dependent cache reads per item unchanged; the acceptance test of derive.rs (every register written per round program, 8 distinct rotations, the x8 mixer's operation and multiply counts as floors) rejects a draw and the next attempt continues the stream\",\n");
}
if p.class.mixer_mult != 1 || p.class.growth { if p.class.mixer_mult != 1 || p.class.growth {
s.push_str(&format!(" \"mixer_mult\": {},\n", p.class.mixer_mult)); s.push_str(&format!(" \"mixer_mult\": {},\n", p.class.mixer_mult));
s.push_str(&format!(" \"cache_growth\": {},\n", p.class.growth)); s.push_str(&format!(" \"cache_growth\": {},\n", p.class.growth));
@ -1801,7 +1862,24 @@ pub fn program_json(p: &Program, day: &str, ds: &DatasetSource) -> String {
join_jhex(&mp.rc) join_jhex(&mp.rc)
)); ));
// The Swift writes jhex(cacheLineMask) here, which breaks the JSON. We write the bare literal. // The Swift writes jhex(cacheLineMask) here, which breaks the JSON. We write the bare literal.
if shape.mixer_mult == 1 { if let Some(dp) = &mp.derive {
s.push_str(&format!(
" \"derive_len\": {},\n \"derive_attempt\": {},\n \"derive_fingerprint\": {},\n \"derive_op_mix\": {},\n \"derive_instrs_per_item\": {},\n \"derive_gpu_ops_per_item\": {},\n \"derive_chip_ops_per_item\": {},\n \"derive_muls_per_item\": {},\n",
dp.len, dp.attempt, jhex64(dp.fingerprint()), jstr(&dp.op_mix()), dp.instr_count(), dp.gpu_ops(), dp.chip_ops(), dp.muls()
));
s.push_str(&format!(
" \"item\": \"s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..{}: s = P_r(s); line = s[0] & 0x{:08x}; s[i] ^= cache[line * 16 + i]; then s = P_{ITEM_ROUNDS}(s); item(t) = s; P_r is round program r below (c = the register the previous instruction wrote, s[0] first): add d += c; sub d -= c; xor d ^= c; mul d *= (c | 1); rot d = rotl(d, k) + c; xrot d = rotl(d ^ c, k); addc d += c + imm; xorc d ^= c ^ imm; mulc d = (d ^ c) * imm; mulc2 d = d * imm + c; andx d ^= (c & b); orx d += (c | b)\",\n",
ITEM_ROUNDS - 1,
shape.cache_line_mask()
));
s.push_str(" \"programs\": [\n");
for (r, prog) in dp.rounds.iter().enumerate() {
s.push_str(" [");
s.push_str(&prog.iter().map(|i| jstr(&derive_instr_line(i))).collect::<Vec<_>>().join(", "));
s.push_str(if r + 1 < dp.rounds.len() { "],\n" } else { "]\n" });
}
s.push_str(" ],\n");
} else if shape.mixer_mult == 1 {
s.push_str(&format!( s.push_str(&format!(
" \"item\": \"s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..{}: s = M_r(s); line = s[0] & 0x{:08x}; s[i] ^= cache[line * 16 + i]; then s = M_{ITEM_ROUNDS}(s); item(t) = s\",\n", " \"item\": \"s[0..7] = key; s[8+i] = t * mul[i] + rc[i] for i in 0..7; for r in 0..{}: s = M_r(s); line = s[0] & 0x{:08x}; s[i] ^= cache[line * 16 + i]; then s = M_{ITEM_ROUNDS}(s); item(t) = s\",\n",
ITEM_ROUNDS - 1, ITEM_ROUNDS - 1,

View file

@ -216,6 +216,10 @@ pub struct LoadClass {
/// Hot table (`docs/plans/hot-table.md`, measured 5 October 2026 and not adopted): `Some(HotClass { mb, k, added })` /// Hot table (`docs/plans/hot-table.md`, measured 5 October 2026 and not adopted): `Some(HotClass { mb, k, added })`
/// turns `k` load slots into reads of an `mb` MiB epoch table. `None` for every other class, class v3 included. /// turns `k` load slots into reads of an `mb` MiB epoch table. `None` for every other class, class v3 included.
pub hot: Option<HotClass>, pub hot: Option<HotClass>,
/// Counter ASIC 3.0 item 2 (`crate::derive`, `docs/plans/counter-asic-3-derivation.md`, 6 October 2026, a
/// prototype behind the class): instructions per round program of the per-day item-derivation program that
/// replaces the fixed mixer when non-zero (the mixer multiplier is then unused and 1). 0 for every other class.
pub derive_len: u16,
} }
/// The parameters one era draws from its seed `E_n` (`docs/plans/era-layout.md` section 1.1, the proposed text of /// The parameters one era draws from its seed `E_n` (`docs/plans/era-layout.md` section 1.1, the proposed text of
@ -365,13 +369,13 @@ impl LoadClass {
impl LoadClass { impl LoadClass {
/// Generator version 2 as adopted on 4 October 2026: 16 loads of one word. The lottery hash. /// Generator version 2 as adopted on 4 October 2026: 16 loads of one word. The lottery hash.
pub const V2: LoadClass = pub const V2: LoadClass =
LoadClass { mix: [100, 0, 0], load_slots: LOAD_SLOTS as u8, scratch: None, scratch_kb: 0, mixer_mult: 1, growth: false, era: None, hot: None }; LoadClass { mix: [100, 0, 0], load_slots: LOAD_SLOTS as u8, scratch: None, scratch_kb: 0, mixer_mult: 1, growth: false, era: None, hot: None, derive_len: 0 };
/// The construction decided for program class v3 on 5 October 2026 (Counter ASIC 2.0, `docs/plans/mixer-x4.md`): /// The construction decided for program class v3 on 5 October 2026 (Counter ASIC 2.0, `docs/plans/mixer-x4.md`):
/// version 2 loads (16 slots of one word, no scratch, no width roll, so the program stream is version 2's), the /// version 2 loads (16 slots of one word, no scratch, no width roll, so the program stream is version 2's), the
/// mixer applied 4 times per round, and the cache growth rule. Name "mx4". /// mixer applied 4 times per round, and the cache growth rule. Name "mx4".
pub const MX4: LoadClass = pub const MX4: LoadClass =
LoadClass { mix: [100, 0, 0], load_slots: LOAD_SLOTS as u8, scratch: None, scratch_kb: 0, mixer_mult: 4, growth: true, era: None, hot: None }; LoadClass { mix: [100, 0, 0], load_slots: LOAD_SLOTS as u8, scratch: None, scratch_kb: 0, mixer_mult: 4, growth: true, era: None, hot: None, derive_len: 0 };
/// The era class over `base` (`docs/plans/era-layout.md`): the parameters drawn by [`era_draw`]; when `allowed` /// The era class over `base` (`docs/plans/era-layout.md`): the parameters drawn by [`era_draw`]; when `allowed`
/// has more than one width the drawn width becomes the class mix (every load that width), otherwise the base /// has more than one width the drawn width becomes the class mix (every load that width), otherwise the base
@ -406,6 +410,24 @@ impl LoadClass {
pub const MX8: LoadClass = pub const MX8: LoadClass =
LoadClass { mixer_mult: 8, ..LoadClass::MX4 }; LoadClass { mixer_mult: 8, ..LoadClass::MX4 };
/// Counter ASIC 3.0 item 2 (6 October 2026, `docs/plans/counter-asic-3-derivation.md`): version 2 loads and the
/// growth rule of class v3, with the per-day derivation program of `crate::derive` (736 instructions per round
/// program, the x8-equivalent operation count) in place of the mixer. Name "dr736". A prototype class; not v3.
pub const DR736: LoadClass =
LoadClass { mixer_mult: 1, derive_len: crate::derive::DERIVE_LEN_X8 as u16, ..LoadClass::MX4 };
/// This class with a derivation program of `len` instructions per round program (0: the fixed mixer); the
/// mixer multiplier is set to 1 under a program, since no mixer is applied.
pub fn with_derive(self, len: u16) -> LoadClass {
assert!(len == 0 || (32..=4096).contains(&len), "derivation program length must be 0 or 32..=4096");
LoadClass { derive_len: len, mixer_mult: if len != 0 { 1 } else { self.mixer_mult }, ..self }
}
/// Whether the item derivation is the per-day program (Counter ASIC 3.0 item 2).
pub fn is_derived(&self) -> bool {
self.derive_len != 0
}
/// A fixed width (1, 4 or 16 words) with `load_slots` loads per program. /// A fixed width (1, 4 or 16 words) with `load_slots` loads per program.
pub fn fixed(width_words: u8, load_slots: u8) -> LoadClass { pub fn fixed(width_words: u8, load_slots: u8) -> LoadClass {
let mut mix = [0u8; 3]; let mut mix = [0u8; 3];
@ -505,6 +527,16 @@ impl LoadClass {
if s == "mx8" { if s == "mx8" {
return Some(LoadClass::MX8); return Some(LoadClass::MX8);
} }
// Counter ASIC 3.0 item 2: "dr<len>" is the derivation class on the v3 loads and growth rule
if let Some(digits) = s.strip_prefix("dr") {
if !digits.is_empty() && digits.bytes().all(|b| b.is_ascii_digit()) {
let len: u16 = digits.parse().ok()?;
if len == 0 || !(32..=4096).contains(&len) {
return None;
}
return Some(LoadClass::MX4.with_derive(len));
}
}
// the mixer suffix: "...m<mult>" then an optional "g" // the mixer suffix: "...m<mult>" then an optional "g"
let (s, growth) = match s.strip_suffix('g') { let (s, growth) = match s.strip_suffix('g') {
Some(base) if base.rsplit_once('m').map(|(_, d)| !d.is_empty() && d.bytes().all(|b| b.is_ascii_digit())).unwrap_or(false) => (base, true), Some(base) if base.rsplit_once('m').map(|(_, d)| !d.is_empty() && d.bytes().all(|b| b.is_ascii_digit())).unwrap_or(false) => (base, true),
@ -627,6 +659,11 @@ impl LoadClass {
if *self == LoadClass::MX8 { if *self == LoadClass::MX8 {
return "mx8".to_string(); return "mx8".to_string();
} }
if self.derive_len != 0 {
// the derivation class: "dr<len>" on the v3 loads; any other base keeps its name with the suffix
let base = LoadClass { derive_len: 0, mixer_mult: 1, ..*self }.base_name();
return if base == "v2m1g" || base == "v2" { format!("dr{}", self.derive_len) } else { format!("{base}dr{}", self.derive_len) };
}
let loads = LoadClass { mixer_mult: 1, growth: false, ..*self }; let loads = LoadClass { mixer_mult: 1, growth: false, ..*self };
let base = if loads.is_v2() { let base = if loads.is_v2() {
"v2".to_string() "v2".to_string()
@ -892,6 +929,11 @@ pub fn program_id_class(generator: u32, seed: &[u32; 8], attempt: u32, class: &L
b.push(class.mixer_mult); b.push(class.mixer_mult);
b.push(class.growth as u8); b.push(class.growth as u8);
} }
if class.derive_len != 0 {
// Counter ASIC 3.0 item 2: the derivation program's length is part of the construction
b.extend_from_slice(b"derive/");
b.extend_from_slice(&class.derive_len.to_le_bytes());
}
if let Some(e) = class.era { if let Some(e) = class.era {
b.extend_from_slice(b"era/"); b.extend_from_slice(b"era/");
b.extend_from_slice(&e.id_bytes()); b.extend_from_slice(&e.id_bytes());

View file

@ -7,6 +7,7 @@
//! * [`accept`]: the acceptance rule every candidate program must pass; a rejected candidate is replaced by the //! * [`accept`]: the acceptance rule every candidate program must pass; a rejected candidate is replaced by the
//! next attempt of the same seed. //! next attempt of the same seed.
//! * [`memhard`]: the 256 MiB ChaCha12 cache and the 8-round dataset item derivation (`proto-metal/MEMHARD.md`). //! * [`memhard`]: the 256 MiB ChaCha12 cache and the 8-round dataset item derivation (`proto-metal/MEMHARD.md`).
//! * [`derive`]: the per-day item-derivation program of Counter ASIC 3.0 item 2 (a prototype behind a load class).
//! * [`verify`]: the 32-lane warp interpreter that computes the 64-bit hash on the CPU, deriving dataset //! * [`verify`]: the 32-lane warp interpreter that computes the 64-bit hash on the CPU, deriving dataset
//! words on demand from the cache (or from the closed form, for the old packs). //! words on demand from the cache (or from the closed form, for the old packs).
//! * [`emit`]: the Metal, CUDA and OpenCL kernel text for a program, byte-identical to the Swift exporter. //! * [`emit`]: the Metal, CUDA and OpenCL kernel text for a program, byte-identical to the Swift exporter.
@ -23,6 +24,7 @@
pub mod accept; pub mod accept;
pub mod bind; pub mod bind;
pub mod derive;
pub mod emit; pub mod emit;
pub mod generator; pub mod generator;
pub mod memhard; pub mod memhard;

View file

@ -93,7 +93,7 @@ fn usage() -> ! {
\x20 hash-bound --prehash <64 hex> --nonce <u64> print the header-bound hash (bind.rs) of one 64-bit nonce\n\ \x20 hash-bound --prehash <64 hex> --nonce <u64> print the header-bound hash (bind.rs) of one 64-bit nonce\n\
\x20 accept every candidate of the seed (or --epoch-hex) with its acceptance verdict\n\ \x20 accept every candidate of the seed (or --epoch-hex) with its acceptance verdict\n\
\x20 show the accepted program, one instruction per line\n\ \x20 show the accepted program, one instruction per line\n\
\x20 --class C load class: v2 (default), mx4 (class v3: mixer x4, cache growth), w4, w16, w64, w64x4, p4,p16,p64[xN], <class>m<mult>[g]\n\ \x20 --class C load class: v2 (default), mx4, mx8 (class v3: mixer x8, cache growth), dr<len> (Counter ASIC 3.0 item 2: the per-day derivation program, dr736 = the x8-equivalent), w4, w16, w64, w64x4, p4,p16,p64[xN], <class>m<mult>[g]\n\
\x20 --days N days since genesis for the cache growth rule of a class with it (default 0: the 2^26-word cache)\n\ \x20 --days N days since genesis for the cache growth rule of a class with it (default 0: the 2^26-word cache)\n\
\x20 --program-class v2|v3 the program class of the seam (v3 = generator 3 on V3_CLASS, the chain's own derivation; --era-hex records the era seed)\n\ \x20 --program-class v2|v3 the program class of the seam (v3 = generator 3 on V3_CLASS, the chain's own derivation; --era-hex records the era seed)\n\
\x20 --era E era layout over --class: igneum-era-test/<n> or <n>:<64 hex> (the 32-byte era seed E_n)\n\ \x20 --era E era layout over --class: igneum-era-test/<n> or <n>:<64 hex> (the 32-byte era seed E_n)\n\
@ -274,6 +274,20 @@ fn bench(a: &Args, mode: DatasetMode) {
shape.cache_log2_words, shape.cache_log2_words,
e.program.op_mix() e.program.op_mix()
); );
if let Some(dp) = e.dataset.memhard().and_then(|m| m.params.derive.as_ref()) {
// Counter ASIC 3.0 item 2: the day's derivation program
println!(
"derivation program: {} instructions per round program, {} per item ({} GPU ops, {} chip ops, {} multiplies per item), attempt {}, fingerprint {:016x}, op mix {}",
dp.len,
dp.instr_count(),
dp.gpu_ops(),
dp.chip_ops(),
dp.muls(),
dp.attempt,
dp.fingerprint(),
dp.op_mix()
);
}
let bases = [0u32, 4096, 1_000_000]; let bases = [0u32, 4096, 1_000_000];
for &b in &bases { for &b in &bases {
let t = Instant::now(); let t = Instant::now();

View file

@ -9,6 +9,7 @@
//! `m` applications with distinct round keys, the 8 dependent cache reads unchanged; the cache doubles when the //! `m` applications with distinct round keys, the 8 dependent cache reads unchanged; the cache doubles when the
//! dataset doubles ([`growth_doublings`]). [`Shape::V2`] (`m = 1`, 2^26 words) is version 2 bit for bit. //! dataset doubles ([`growth_doublings`]). [`Shape::V2`] (`m = 1`, 2^26 words) is version 2 bit for bit.
use crate::derive::{run_round, DeriveProgram, SoaState, DERIVE_REGS, SOA_LANES};
use crate::generator::LoadClass; use crate::generator::LoadClass;
use crate::seed::{day_key, fnv1a64_words, SplitMix64}; use crate::seed::{day_key, fnv1a64_words, SplitMix64};
@ -45,11 +46,15 @@ pub struct Shape {
pub mixer_mult: u32, pub mixer_mult: u32,
/// The cache is 2^cache_log2_words words (26 at genesis; 27 and 28 after the dataset doublings of 1.13.3). /// The cache is 2^cache_log2_words words (26 at genesis; 27 and 28 after the dataset doublings of 1.13.3).
pub cache_log2_words: u32, pub cache_log2_words: u32,
/// Counter ASIC 3.0 item 2 (`crate::derive`): instructions per round program of the per-day derivation
/// program, which replaces the `mixer_mult` applications of `M_r` in every mixer slot when non-zero. 0 for
/// version 2 and class v3 (the fixed mixer).
pub derive_len: u32,
} }
impl Shape { impl Shape {
/// Version 2: one mixer application per round, a 2^26-word cache. /// Version 2: one mixer application per round, a 2^26-word cache.
pub const V2: Shape = Shape { mixer_mult: 1, cache_log2_words: CACHE_LOG2_WORDS as u32 }; pub const V2: Shape = Shape { mixer_mult: 1, cache_log2_words: CACHE_LOG2_WORDS as u32, derive_len: 0 };
/// The shape of a load class on day 0 of the chain (and on every day for a class without the growth rule). /// The shape of a load class on day 0 of the chain (and on every day for a class without the growth rule).
pub fn for_class(class: &LoadClass) -> Shape { pub fn for_class(class: &LoadClass) -> Shape {
@ -62,6 +67,7 @@ impl Shape {
Shape { Shape {
mixer_mult: class.mixer_mult(), mixer_mult: class.mixer_mult(),
cache_log2_words: if class.growth { cache_log2_words(days_since_genesis) } else { CACHE_LOG2_WORDS as u32 }, cache_log2_words: if class.growth { cache_log2_words(days_since_genesis) } else { CACHE_LOG2_WORDS as u32 },
derive_len: class.derive_len as u32,
} }
} }
@ -83,9 +89,21 @@ impl Shape {
pub fn log2_segments(&self) -> u32 { pub fn log2_segments(&self) -> u32 {
self.cache_log2_words - 4 - CACHE_SEGMENT_LOG2_LINES as u32 self.cache_log2_words - 4 - CACHE_SEGMENT_LOG2_LINES as u32
} }
/// Mixer applications per item: `(ITEM_ROUNDS + 1) x m`. /// Mixer applications per item: `(ITEM_ROUNDS + 1) x m` (0 under a derivation program, which has no mixer).
pub fn mixers_per_item(&self) -> u32 { pub fn mixers_per_item(&self) -> u32 {
(ITEM_ROUNDS as u32 + 1) * self.mixer_mult if self.is_derived() {
0
} else {
(ITEM_ROUNDS as u32 + 1) * self.mixer_mult
}
}
/// Whether the item derivation is the per-day program of `crate::derive` (Counter ASIC 3.0 item 2).
pub fn is_derived(&self) -> bool {
self.derive_len != 0
}
/// Instructions per item under a derivation program: `(ITEM_ROUNDS + 1) x derive_len`.
pub fn derive_instrs_per_item(&self) -> u32 {
(ITEM_ROUNDS as u32 + 1) * self.derive_len
} }
} }
@ -169,7 +187,10 @@ pub fn chacha_block(x: &[u32; 16]) -> [u32; 16] {
} }
/// Mixer parameters drawn from the day key, plus the [`Shape`] the mixer is applied under. Draw order: /// Mixer parameters drawn from the day key, plus the [`Shape`] the mixer is applied under. Draw order:
/// ROT[0..7] (1..31), MUL[0..15] (odd), RC[0..15]. The shape is not drawn: it is the class's. /// ROT[0..7] (1..31), MUL[0..15] (odd), RC[0..15]. The shape is not drawn: it is the class's. Under a shape with
/// a derivation program (Counter ASIC 3.0 item 2) the same stream continues after the 40 draws with the program's
/// draws (`DeriveProgram::draw`); the mixer constants are still drawn (the item init uses MUL and RC) and the
/// mixer itself is not applied.
#[derive(Clone, Debug, PartialEq, Eq)] #[derive(Clone, Debug, PartialEq, Eq)]
pub struct MixParams { pub struct MixParams {
pub key: [u32; 8], pub key: [u32; 8],
@ -177,6 +198,8 @@ pub struct MixParams {
pub mul: [u32; 16], pub mul: [u32; 16],
pub rc: [u32; 16], pub rc: [u32; 16],
pub shape: Shape, pub shape: Shape,
/// The per-day derivation program when `shape.derive_len != 0`, else `None`.
pub derive: Option<DeriveProgram>,
} }
impl MixParams { impl MixParams {
@ -198,7 +221,8 @@ impl MixParams {
for c in rc.iter_mut() { for c in rc.iter_mut() {
*c = rng.next() as u32; *c = rng.next() as u32;
} }
Self { key, rot, mul, rc, shape } let derive = if shape.is_derived() { Some(DeriveProgram::draw(&mut rng, shape.derive_len)) } else { None };
Self { key, rot, mul, rc, shape, derive }
} }
/// Parameters for a day string: the key is `seed_words("day/" + day)`. /// Parameters for a day string: the key is `seed_words("day/" + day)`.
pub fn for_day(day: &str) -> Self { pub fn for_day(day: &str) -> Self {
@ -350,7 +374,7 @@ impl Cache {
/// are the smaller cache's segments word for word. /// are the smaller cache's segments word for word.
pub fn fill_log2(key: [u32; 8], log2_words: u32) -> Cache { pub fn fill_log2(key: [u32; 8], log2_words: u32) -> Cache {
assert!((10..=30).contains(&log2_words), "cache log2 words must be in 10..=30"); assert!((10..=30).contains(&log2_words), "cache log2 words must be in 10..=30");
let shape = Shape { mixer_mult: 1, cache_log2_words: log2_words }; let shape = Shape { mixer_mult: 1, cache_log2_words: log2_words, derive_len: 0 };
let mut words = vec![0u32; shape.cache_words()]; let mut words = vec![0u32; shape.cache_words()];
for seg in 0..shape.cache_segments() { for seg in 0..shape.cache_segments() {
Self::fill_segment(&mut words, seg, &key); Self::fill_segment(&mut words, seg, &key);
@ -481,6 +505,9 @@ impl HotTable {
/// (`mp.shape.mixer_mult`) round `r` applies `M` with keys `round_key(r m + j)` for `j = 0 .. m - 1` before its /// (`mp.shape.mixer_mult`) round `r` applies `M` with keys `round_key(r m + j)` for `j = 0 .. m - 1` before its
/// one cache read; the final mixer applies `M` with keys `round_key(8 m + j)`. `m = 1` is version 2. /// one cache read; the final mixer applies `M` with keys `round_key(8 m + j)`. `m = 1` is version 2.
pub fn derive_items(ts: &[u32], mp: &MixParams, cache: &Cache, out: &mut [[u32; 16]]) { pub fn derive_items(ts: &[u32], mp: &MixParams, cache: &Cache, out: &mut [[u32; 16]]) {
if let Some(prog) = &mp.derive {
return derive_items_program(ts, mp, prog, cache, out);
}
// The item loop lives in its own function, one instance per cache size the growth rule can reach with the line // The item loop lives in its own function, one instance per cache size the growth rule can reach with the line
// mask a constant, never inlined into the callers. Inlined into `MemhardCpu::fetch` it ran at 1.33 ms per unit // mask a constant, never inlined into the callers. Inlined into `MemhardCpu::fetch` it ran at 1.33 ms per unit
// against 0.61 out of line (the version 2 verifier, bisected on one core under the measure lock, 5 October 2026, // against 0.61 out of line (the version 2 verifier, bisected on one core under the measure lock, 5 October 2026,
@ -535,6 +562,44 @@ fn derive_items_mask<const LINE_MASK: u32>(ts: &[u32], mp: &MixParams, cache: &C
} }
} }
/// [`derive_items`] under a per-day derivation program (Counter ASIC 3.0 item 2, `crate::derive`): the same
/// init and the same 8 dependent cache reads, with round program `r` in place of the `m` mixer applications of
/// round `r` and program 8 in place of the final applications. The states are kept word-major
/// (`st[reg][lane]`) so every instruction runs across the batch's items in one vectorised loop and the
/// interpreter's dispatch is paid once per instruction per batch of up to 32 items, not once per item. The
/// cache reads of the batch are issued together, as in the fixed-mixer loop, so the 8 dependent misses of
/// independent items overlap in the memory system.
#[inline(never)]
pub fn derive_items_program(ts: &[u32], mp: &MixParams, prog: &DeriveProgram, cache: &Cache, out: &mut [[u32; 16]]) {
let n = ts.len();
debug_assert!(out.len() >= n && n <= SOA_LANES);
assert_eq!(prog.rounds.len(), ITEM_ROUNDS + 1);
let mut st: SoaState = [[0u32; SOA_LANES]; DERIVE_REGS];
for k in 0..n {
let t = ts[k];
for i in 0..8 {
st[i][k] = mp.key[i];
st[8 + i][k] = t.wrapping_mul(mp.mul[i]).wrapping_add(mp.rc[i]);
}
}
let mask = cache.line_mask();
for r in 0..ITEM_ROUNDS {
run_round(&prog.rounds[r], &mut st);
for k in 0..n {
let line = cache.line(st[0][k] & mask);
for i in 0..16 {
st[i][k] ^= line[i];
}
}
}
run_round(&prog.rounds[ITEM_ROUNDS], &mut st);
for k in 0..n {
for i in 0..16 {
out[k][i] = st[i][k];
}
}
}
/// One dataset item, 16 words. /// One dataset item, 16 words.
pub fn derive_item(t: u32, mp: &MixParams, cache: &Cache) -> [u32; 16] { pub fn derive_item(t: u32, mp: &MixParams, cache: &Cache) -> [u32; 16] {
let mut out = [[0u32; 16]; 1]; let mut out = [[0u32; 16]; 1];
@ -786,7 +851,7 @@ mod tests {
let v2 = Shape::for_class_day(&LoadClass::V2, 100_000); let v2 = Shape::for_class_day(&LoadClass::V2, 100_000);
assert_eq!(v2, Shape::V2); assert_eq!(v2, Shape::V2);
let v3 = Shape::for_class_day(&LoadClass::MX4, 0); let v3 = Shape::for_class_day(&LoadClass::MX4, 0);
assert_eq!(v3, Shape { mixer_mult: 4, cache_log2_words: 26 }); assert_eq!(v3, Shape { mixer_mult: 4, cache_log2_words: 26, derive_len: 0 });
assert_eq!(Shape::for_class_day(&LoadClass::MX4, 1_460).cache_log2_words, 27); assert_eq!(Shape::for_class_day(&LoadClass::MX4, 1_460).cache_log2_words, 27);
assert_eq!(v3.mixers_per_item(), 36); assert_eq!(v3.mixers_per_item(), 36);
assert_eq!(Shape::V2.mixers_per_item(), 9); assert_eq!(Shape::V2.mixers_per_item(), 9);
@ -807,7 +872,7 @@ mod tests {
assert_eq!(small.segments(), 64); assert_eq!(small.segments(), 64);
assert_eq!(small.line_mask(), 4095); assert_eq!(small.line_mask(), 4095);
for m in [1u32, 2, 4] { for m in [1u32, 2, 4] {
let mp = MixParams::with_shape(key, Shape { mixer_mult: m, cache_log2_words: 16 }); let mp = MixParams::with_shape(key, Shape { mixer_mult: m, cache_log2_words: 16, derive_len: 0 });
for t in [0u32, 1, 12_345, u32::MAX] { for t in [0u32, 1, 12_345, u32::MAX] {
let got = derive_item(t, &mp, &small); let got = derive_item(t, &mp, &small);
let mut s = [0u32; 16]; let mut s = [0u32; 16];
@ -830,8 +895,8 @@ mod tests {
assert_eq!(got, s, "m {m} t {t}"); assert_eq!(got, s, "m {m} t {t}");
} }
} }
let v2 = MixParams::with_shape(key, Shape { mixer_mult: 1, cache_log2_words: 16 }); let v2 = MixParams::with_shape(key, Shape { mixer_mult: 1, cache_log2_words: 16, derive_len: 0 });
let v3 = MixParams::with_shape(key, Shape { mixer_mult: 4, cache_log2_words: 16 }); let v3 = MixParams::with_shape(key, Shape { mixer_mult: 4, cache_log2_words: 16, derive_len: 0 });
assert_ne!(derive_item(0, &v2, &small), derive_item(0, &v3, &small)); assert_ne!(derive_item(0, &v2, &small), derive_item(0, &v3, &small));
assert_eq!(round_key_mult(0, 0, 1), round_key(0)); assert_eq!(round_key_mult(0, 0, 1), round_key(0));
assert_eq!(round_key_mult(8, 0, 1), round_key(8)); assert_eq!(round_key_mult(8, 0, 1), round_key(8));

275
igneum-pow/tests/derive.rs Normal file
View file

@ -0,0 +1,275 @@
//! The per-day item-derivation program (Counter ASIC 3.0 item 2, `docs/plans/counter-asic-3-derivation.md`): the
//! soundness runs on the CPU.
//!
//! 1. By hand: the derived item of class dr736 restated with the scalar reference on a small cache equals the
//! verifier's batched (SoA) derivation, at the item boundary and the index wrap.
//! 2. The v2 and v3 paths are untouched: the derivation field is 0 on both, their items are the fixed-mixer items.
//! 3. Determinism and the stream: two epochs agree on every vector and file; the program is the continuation of the
//! mixer's stream (the first draw of the program follows the 40 mixer draws); a different day draws a different
//! program; the class is in the program id and the name.
//! 4. Stats beside x8: bit balance and single-bit avalanche of the derived items and of the hash, on the same seeds.
//! 5. A program whose text is the kernels': every instruction's C text evaluated by hand on one state matches the
//! scalar reference (the text forms are what Metal, CUDA and OpenCL compile).
use igneum_pow::derive::{instr_text, run_round_scalar, DOp, DeriveProgram, DERIVE_LEN_X8, DERIVE_PROGRAMS};
use igneum_pow::emit::export_pack;
use igneum_pow::generator::{generate_from_seed_bytes_class, LoadClass, V3_CLASS};
use igneum_pow::memhard::{derive_item, derive_items, mixer, round_key, Cache, MixParams, Shape, ITEM_ROUNDS};
use igneum_pow::seed::{day_key, SplitMix64};
use igneum_pow::verify::{DatasetMode, DatasetSource, Epoch};
const DAY: &str = "2026-10-03";
/// The item of a derivation class restated by hand with the scalar reference.
fn item_by_hand(t: u32, mp: &MixParams, cache: &Cache) -> [u32; 16] {
let prog = mp.derive.as_ref().expect("a derivation class");
let mut s = [0u32; 16];
s[..8].copy_from_slice(&mp.key);
for i in 0..8 {
s[8 + i] = t.wrapping_mul(mp.mul[i]).wrapping_add(mp.rc[i]);
}
for r in 0..ITEM_ROUNDS {
run_round_scalar(&prog.rounds[r], &mut s);
let line = cache.line(s[0]);
for i in 0..16 {
s[i] ^= line[i];
}
}
run_round_scalar(&prog.rounds[ITEM_ROUNDS], &mut s);
s
}
#[test]
fn derived_item_by_hand_and_in_batches() {
let key = day_key(DAY);
let cache = Cache::fill_log2(key, 16);
let shape = Shape { mixer_mult: 1, cache_log2_words: 16, derive_len: DERIVE_LEN_X8 };
let mp = MixParams::with_shape(key, shape);
let prog = mp.derive.as_ref().unwrap();
assert_eq!(prog.rounds.len(), DERIVE_PROGRAMS);
assert!(prog.check().is_ok());
assert_eq!(shape.derive_instrs_per_item(), 9 * DERIVE_LEN_X8);
assert_eq!(shape.mixers_per_item(), 0);
for t in [0u32, 1, 2, 15, 16, 17, 12_345, (1 << 28) - 1, u32::MAX - 1, u32::MAX] {
assert_eq!(derive_item(t, &mp, &cache), item_by_hand(t, &mp, &cache), "t {t}");
}
// a batch of 32 distinct items against one at a time, and a short batch
let ts: Vec<u32> = (0..32).map(|k| k * 7_919 + 3).collect();
let mut out = [[0u32; 16]; 32];
derive_items(&ts, &mp, &cache, &mut out);
for (k, &t) in ts.iter().enumerate() {
assert_eq!(out[k], item_by_hand(t, &mp, &cache), "batch slot {k}");
}
let mut out5 = [[0u32; 16]; 5];
derive_items(&ts[..5], &mp, &cache, &mut out5);
assert_eq!(&out5[..], &out[..5]);
// the fixed mixer of the same key gives other items
let v3 = MixParams::with_shape(key, Shape { mixer_mult: 8, cache_log2_words: 16, derive_len: 0 });
assert!(v3.derive.is_none());
assert_ne!(derive_item(0, &v3, &cache), derive_item(0, &mp, &cache));
}
#[test]
fn v2_and_v3_are_untouched() {
assert_eq!(LoadClass::V2.derive_len, 0);
assert_eq!(V3_CLASS.derive_len, 0);
assert_eq!(LoadClass::MX8.derive_len, 0);
assert_eq!(Shape::V2.derive_len, 0);
assert!(!Shape::for_class(&V3_CLASS).is_derived());
let key = day_key(DAY);
let cache = Cache::fill_log2(key, 16);
// the version 2 item restated by hand (the mixer_mult_by_hand test of memhard.rs, m = 1)
let v2 = MixParams::with_shape(key, Shape { mixer_mult: 1, cache_log2_words: 16, derive_len: 0 });
let t = 12_345u32;
let mut s = [0u32; 16];
s[..8].copy_from_slice(&key);
for i in 0..8 {
s[8 + i] = t.wrapping_mul(v2.mul[i]).wrapping_add(v2.rc[i]);
}
for r in 0..8usize {
mixer(&mut s, round_key(r), &v2);
let line = cache.line(s[0]);
for i in 0..16 {
s[i] ^= line[i];
}
}
mixer(&mut s, round_key(8), &v2);
assert_eq!(derive_item(t, &v2, &cache), s);
// the mixer constants of the derivation class are the v2 draws (the stream continues after them)
let dr = MixParams::with_shape(key, Shape { mixer_mult: 1, cache_log2_words: 16, derive_len: DERIVE_LEN_X8 });
assert_eq!((dr.rot, dr.mul, dr.rc), (v2.rot, v2.mul, v2.rc));
}
#[test]
fn stream_class_name_and_id() {
let key = day_key(DAY);
// the program is the continuation of the mixer stream: 40 draws, then the program
let mut rng = SplitMix64::new(key[0] as u64 | ((key[1] as u64) << 32));
for _ in 0..40 {
rng.next();
}
let expect = DeriveProgram::draw(&mut rng, DERIVE_LEN_X8);
let mp = MixParams::with_shape(key, Shape { mixer_mult: 1, cache_log2_words: 26, derive_len: DERIVE_LEN_X8 });
assert_eq!(mp.derive.as_ref().unwrap(), &expect);
// another day, another program; another length, another program
let other = MixParams::with_shape(day_key("2026-10-04"), Shape { mixer_mult: 1, cache_log2_words: 26, derive_len: DERIVE_LEN_X8 });
assert_ne!(other.derive.as_ref().unwrap().fingerprint(), expect.fingerprint());
let short = MixParams::with_shape(key, Shape { mixer_mult: 1, cache_log2_words: 26, derive_len: 368 });
assert_eq!(short.derive.as_ref().unwrap().instr_count(), 9 * 368);
// the class: name, parse, id, and the v2 program stream (v2 loads, no width roll)
let c = LoadClass::DR736;
assert_eq!(c.name(), "dr736");
assert_eq!(LoadClass::parse("dr736"), Some(c));
assert_eq!(LoadClass::parse("dr368"), Some(LoadClass::MX4.with_derive(368)));
assert_eq!(LoadClass::parse("dr0"), None);
assert_eq!(LoadClass::parse("dr5000"), None);
assert!(c.v2_loads() && !c.takes_width_roll() && c.growth && c.mixer_mult == 1 && c.is_derived());
let p = generate_from_seed_bytes_class("igneum-genesis", b"igneum-genesis", c);
let v2 = generate_from_seed_bytes_class("igneum-genesis", b"igneum-genesis", LoadClass::V2);
let mx8 = generate_from_seed_bytes_class("igneum-genesis", b"igneum-genesis", LoadClass::MX8);
assert_eq!(p.instrs, v2.instrs, "the v2 program of the seed");
assert_ne!(p.program_id(), v2.program_id());
assert_ne!(p.program_id(), mx8.program_id());
assert_ne!(
generate_from_seed_bytes_class("igneum-genesis", b"igneum-genesis", LoadClass::MX4.with_derive(368)).program_id(),
p.program_id()
);
}
#[test]
fn determinism_and_pack_text() {
let a = Epoch::new_class("igneum-genesis", DAY, DatasetMode::MemoryHard, 20, LoadClass::DR736);
let b = Epoch::new_class("igneum-genesis", DAY, DatasetMode::MemoryHard, 20, LoadClass::DR736);
let pa = export_pack(&a, DAY, "test");
let pb = export_pack(&b, DAY, "test");
assert_eq!(pa.outs, pb.outs);
assert_eq!(pa.files, pb.files);
let names: Vec<&str> = pa.files.iter().map(|(n, _)| n.as_str()).collect();
assert!(names.contains(&"memhard.h") && names.contains(&"memhard.metal"));
for (name, text) in &pa.files {
if name == "memhard.h" || name == "memhard.metal" || name == "kernel.cl" {
for r in 0..DERIVE_PROGRAMS {
assert!(text.contains(&format!("mh_round_{r}(")), "{name} carries round program {r}");
}
assert!(!text.contains("mh_mixer(s,"), "{name}: no mixer application under a derivation program");
let dp = a.dataset.memhard().unwrap().params.derive.as_ref().unwrap();
// every instruction's text appears, in order, inside the round functions
let first = instr_text(&dp.rounds[0][0]);
assert!(text.contains(&first), "{name} carries the first instruction {first}");
}
if name == "program.h" {
assert!(text.contains("#define IGNEUM_LOAD_CLASS \"dr736\""));
assert!(text.contains("#define IGNEUM_DERIVE_LEN 736"));
assert!(text.contains("#define IGNEUM_CLASS_DERIVE_LEN 736"));
assert!(!text.contains("#define IGNEUM_MIXER_MULT"));
assert!(text.contains("#define IGNEUM_GENERATOR 2"));
}
if name == "program.json" {
assert!(text.contains("\"derive_len\": 736"));
assert!(text.contains("\"programs\": ["));
serde_json::from_str::<serde_json::Value>(text).expect("valid JSON");
}
}
// the vectors are the interpreter's
assert_eq!(pa.outs[0], a.hash_warp(0));
}
/// Bit balance and single-bit avalanche of the derived items (t flipped one bit) and of the hash, beside x8.
#[test]
fn stats_beside_x8() {
let key = day_key(DAY);
let cache = Cache::fill_log2(key, 18);
let dr = MixParams::with_shape(key, Shape { mixer_mult: 1, cache_log2_words: 18, derive_len: DERIVE_LEN_X8 });
let x8 = MixParams::with_shape(key, Shape { mixer_mult: 8, cache_log2_words: 18, derive_len: 0 });
for (label, mp) in [("dr736", &dr), ("x8", &x8)] {
let n = 2048u32;
let mut ones = [0u32; 512];
let mut flips = 0u64;
let mut flip_n = 0u64;
let mut seen = std::collections::HashSet::new();
for t in 0..n {
let a = derive_item(t * 2_654_435_761, mp, &cache);
assert!(seen.insert(a), "{label}: duplicate item");
for i in 0..16 {
for b in 0..32 {
ones[i * 32 + b] += (a[i] >> b) & 1;
}
}
let bit = t % 32;
let c = derive_item((t * 2_654_435_761) ^ (1 << bit), mp, &cache);
for i in 0..16 {
flips += (a[i] ^ c[i]).count_ones() as u64;
}
flip_n += 512;
}
let avalanche = flips as f64 / flip_n as f64 * 100.0;
let worst_z = ones.iter().map(|&o| ((o as f64 - n as f64 / 2.0) / (n as f64 / 4.0).sqrt()).abs()).fold(0.0, f64::max);
println!("{label}: avalanche {avalanche:.2} percent, worst bit z {worst_z:.2}");
assert!((48.5..=51.5).contains(&avalanche), "{label}: avalanche {avalanche}");
assert!(worst_z < 4.5, "{label}: worst bit z {worst_z}");
}
// the hash on the same seeds: avalanche across a nonce flip within the unit
let e = Epoch::new_class("igneum-genesis", DAY, DatasetMode::MemoryHard, 20, LoadClass::DR736);
let w0 = e.hash_warp(0);
let w1 = e.hash_warp(32);
let mut d = 0u64;
for l in 0..32 {
d += (w0[l] ^ w1[l]).count_ones() as u64;
}
let av = d as f64 / (32.0 * 64.0) * 100.0;
println!("hash unit 0 against unit 1: {av:.2} percent of bits differ");
assert!((44.0..=56.0).contains(&av));
}
/// The kernels' text forms: every form's C text, read back into the scalar reference's arithmetic by hand.
#[test]
fn text_forms_match_scalar_reference() {
let mut rng = SplitMix64::new(42);
let p = DeriveProgram::draw_candidate(&mut rng, 64, 0);
let mut s = [0u32; 16];
for (i, v) in s.iter_mut().enumerate() {
*v = 0x9e37_79b9u32.wrapping_mul(i as u32 + 1);
}
let mut seen = [false; 12];
for ins in p.rounds.iter().flatten() {
seen[ins.op as usize] = true;
let before = s;
run_round_scalar(std::slice::from_ref(ins), &mut s);
let (d, c, b) = (ins.dst as usize, ins.src as usize, ins.src2 as usize);
let (dv, cv, bv) = (before[d], before[c], before[b]);
let want = match ins.op {
DOp::Add => dv.wrapping_add(cv),
DOp::Sub => dv.wrapping_sub(cv),
DOp::Xor => dv ^ cv,
DOp::Mul => dv.wrapping_mul(cv | 1),
DOp::Rot => dv.rotate_left(ins.rot as u32).wrapping_add(cv),
DOp::XRot => (dv ^ cv).rotate_left(ins.rot as u32),
DOp::AddC => dv.wrapping_add(cv.wrapping_add(ins.imm)),
DOp::XorC => dv ^ cv ^ ins.imm,
DOp::MulC => (dv ^ cv).wrapping_mul(ins.imm),
DOp::MulC2 => dv.wrapping_mul(ins.imm).wrapping_add(cv),
DOp::AndX => dv ^ (cv & bv),
DOp::OrX => dv.wrapping_add(cv | bv),
};
assert_eq!(s[d], want, "{}", instr_text(ins));
for i in 0..16 {
if i != d {
assert_eq!(s[i], before[i], "only the destination changes: {}", instr_text(ins));
}
}
assert!(instr_text(ins).starts_with(&format!("s[{d}]")));
}
assert!(seen.iter().all(|s| *s), "64 x 9 draws cover every form");
}
/// The dataset source of the class on a day: the verifier's `word` path derives through the program.
#[test]
fn dataset_source_word_path() {
let ds = DatasetSource::new_shape(DAY, DatasetMode::MemoryHard, 20, Shape { mixer_mult: 1, cache_log2_words: 16, derive_len: DERIVE_LEN_X8 });
let m = ds.memhard().unwrap();
let item = derive_item(3, &m.params, &m.cache);
for j in 0..16u32 {
assert_eq!(ds.word(3 * 16 + j), item[j as usize]);
}
assert_eq!(ds.word((1 << 20) + 5), ds.word(5), "the mask");
}

View file

@ -181,7 +181,7 @@ fn edge_items_every_multiplier() {
let key = day_key(DAY); let key = day_key(DAY);
let cache = Cache::fill_log2(key, 14); let cache = Cache::fill_log2(key, 14);
for m in [1u32, 2, 4, 8] { for m in [1u32, 2, 4, 8] {
let mp = MixParams::with_shape(key, Shape { mixer_mult: m, cache_log2_words: 14 }); let mp = MixParams::with_shape(key, Shape { mixer_mult: m, cache_log2_words: 14, derive_len: 0 });
let by_hand = |t: u32| -> [u32; 16] { let by_hand = |t: u32| -> [u32; 16] {
let mut s = [0u32; 16]; let mut s = [0u32; 16];
s[..8].copy_from_slice(&key); s[..8].copy_from_slice(&key);

View file

@ -369,7 +369,7 @@ fn v3_packs_are_the_v2_seeds_under_mixer_x8() {
assert_eq!(e3.program.program_id(), igneum_pow::generator::program_id(GENERATOR_VERSION_V3, &e3.program.seed, e3.program.attempt)); assert_eq!(e3.program.program_id(), igneum_pow::generator::program_id(GENERATOR_VERSION_V3, &e3.program.seed, e3.program.attempt));
let m3 = e3.dataset.memhard().unwrap(); let m3 = e3.dataset.memhard().unwrap();
let m2 = e2.dataset.memhard().unwrap(); let m2 = e2.dataset.memhard().unwrap();
assert_eq!(m3.shape(), Shape { mixer_mult: 8, cache_log2_words: 26 }); assert_eq!(m3.shape(), Shape { mixer_mult: 8, cache_log2_words: 26, derive_len: 0 });
assert_eq!(m3.cache.fnv1a64(), m2.cache.fnv1a64(), "{v3}: the same cache as v2 on day 0"); assert_eq!(m3.cache.fnv1a64(), m2.cache.fnv1a64(), "{v3}: the same cache as v2 on day 0");
assert_eq!(m3.params.rot, m2.params.rot); assert_eq!(m3.params.rot, m2.params.rot);
assert_eq!(e3.dataset.log2_words, 28); assert_eq!(e3.dataset.log2_words, 28);
@ -405,7 +405,7 @@ fn v3_packs_are_the_v2_seeds_under_mixer_x8() {
assert_eq!(j["load_class"].as_str().unwrap(), "mx4"); assert_eq!(j["load_class"].as_str().unwrap(), "mx4");
assert_eq!(e4.program.class, LoadClass::MX4); assert_eq!(e4.program.class, LoadClass::MX4);
assert_eq!(e4.program.instrs, epoch(v2).program.instrs); assert_eq!(e4.program.instrs, epoch(v2).program.instrs);
assert_eq!(e4.dataset.memhard().unwrap().shape(), Shape { mixer_mult: 4, cache_log2_words: 26 }); assert_eq!(e4.dataset.memhard().unwrap().shape(), Shape { mixer_mult: 4, cache_log2_words: 26, derive_len: 0 });
assert!(read(x4, "memhard.h").contains("j < 4u; ++j) mh_mixer(s, 0x9E3779B9u * (r * 4u + j + 1u))")); assert!(read(x4, "memhard.h").contains("j < 4u; ++j) mh_mixer(s, 0x9E3779B9u * (r * 4u + j + 1u))"));
} }
} }

File diff suppressed because it is too large Load diff

View file

@ -0,0 +1,164 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal).
// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC.
#include <cuda_runtime.h>
#include <cstdint>
#include "program.h"
#include "memhard.h"
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
x ^= x >> 16; x *= 0x7feb352du;
x ^= x >> 15; x *= 0x846ca68bu;
x ^= x >> 16;
return x;
}
// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31.
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x.
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) {
uint32_t x = i ^ d0;
x *= 0x9E3779B1u; x ^= x >> 15;
x += d1;
x *= 0x85EBCA77u; x ^= x >> 13;
x *= 0xC2B2AE3Du; x ^= x >> 16;
return x;
}
// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item.
// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host.
__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) {
uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x;
if (seg < nSegments) mh_cache_segment(cache, seg);
}
__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
uint32_t t = blockIdx.x * blockDim.x + threadIdx.x;
if (t < nItems) {
uint32_t s[16];
mh_item(cache, t, s);
uint32_t* d = ds + (size_t)t * 16u;
for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i];
}
}
// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every
// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a
// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid.
__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) {
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
uint32_t nonce = baseNonce + gid;
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
{ uint32_t x = nonce ^ 0x667d0fbdu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x7b8e5963u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
{ uint32_t x = nonce ^ 0x7b8e5963u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0x31c67e5eu; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
{ uint32_t x = nonce ^ 0x31c67e5eu; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4529ddc6u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
{ uint32_t x = nonce ^ 0x4529ddc6u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xef19d6d8u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
{ uint32_t x = nonce ^ 0xef19d6d8u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xaccf6211u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
{ uint32_t x = nonce ^ 0xaccf6211u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0xda0aed32u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
{ uint32_t x = nonce ^ 0xda0aed32u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0xabc6df31u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
{ uint32_t x = nonce ^ 0xabc6df31u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x667d0fbdu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
for (uint32_t it = 0u; it < 8u; ++it) {
uint32_t sel = r0;
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r0, 4); // 1 shfl
r3 = r3 + r2 + ((((sel >> 10u) & 1u) != 0u) ? 0x642e66dbu : 0x2cccb6cau); // 2 add
r0 = rotl_imm(r0, 19u); // 3 rotl
r7 = rotr_var(r7, r6); // 4 rotr
r7 = r7 + r4 + ((((sel >> 21u) & 1u) != 0u) ? 0xc1535555u : 0xee02465fu); // 5 add
r1 = __umulhi(r1, r7); // 6 mulhi
r4 = r4 ^ ds[r2 & mask]; // 7 load
r7 = r7 ^ ds[r4 & mask]; // 8 load
r0 = r0 ^ ds[r3 & mask]; // 9 load
r5 = r5 ^ ds[r1 & mask]; // 10 load
r1 = r1 ^ ds[r5 & mask]; // 11 load
r3 = __umulhi(r3, r5); // 12 mulhi
r1 = r1 ^ ds[r3 & mask]; // 13 load
r0 = r0 - r3; // 14 sub
r5 = r1 * r3 + r5; // 15 mad
r6 = __umulhi(r6, r1); // 16 mulhi
r5 = r5 + r2 + ((((sel >> 28u) & 1u) != 0u) ? 0x8b965b57u : 0x697b3d00u); // 17 add
r0 = __umulhi(r0, r6); // 18 mulhi
r5 = rotr_var(r5, r3); // 19 rotr
r5 = __umulhi(r5, r2); // 20 mulhi
r1 = r1 + r0 + ((((sel >> 1u) & 1u) != 0u) ? 0x6d7e8d05u : 0xebcf247au); // 21 add
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 22 add
r1 = __umulhi(r1, r5); // 23 mulhi
r2 = r2 - r5; // 24 sub
r7 = r7 + r4 + ((((sel >> 2u) & 1u) != 0u) ? 0x699ef1bbu : 0x08ffa6c7u); // 25 add
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 26 shfl
r7 = r7 + r1 + ((((sel >> 14u) & 1u) != 0u) ? 0xb4ead2fbu : 0xe60fea84u); // 27 add
r3 = r3 + r1 + ((((sel >> 6u) & 1u) != 0u) ? 0x8f30d21du : 0x65c76dabu); // 28 add
r2 = r2 ^ ds[r1 & mask]; // 29 load
r5 = r5 ^ ds[r7 & mask]; // 30 load
r2 = r2 ^ ds[r5 & mask]; // 31 load
r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r7, 4); // 32 shfl
r4 = r5 * r7 + r4; // 33 mad
r4 = r4 + r2 + ((((sel >> 21u) & 1u) != 0u) ? 0xc7ce690cu : 0x0480debeu); // 34 add
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r7, 8); // 35 shfl
r7 = r7 + r1 + ((((sel >> 2u) & 1u) != 0u) ? 0xe10c2c95u : 0xc53b542eu); // 36 add
r5 = r5 ^ r7; // 37 xor
r2 = r2 | r1; // 38 or
r1 = __umulhi(r1, r0); // 39 mulhi
r6 = rotl_imm(r6, 19u); // 40 rotl
r4 = __umulhi(r4, r6); // 41 mulhi
r6 = r6 - r0; // 42 sub
r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 43 shfl
r4 = r4 ^ ds[r2 & mask]; // 44 load
r1 = r1 ^ r3; // 45 xor
r7 = r7 ^ ds[r0 & mask]; // 46 load
r3 = r3 ^ ds[r1 & mask]; // 47 load
r5 = r5 * r3; // 48 mul
r1 = r1 - r5; // 49 sub
r2 = rotl_imm(r2, 8u); // 50 rotl
r1 = r1 + r5 + ((((sel >> 23u) & 1u) != 0u) ? 0x77b9bd43u : 0xa900fec4u); // 51 add
r4 = r4 ^ ds[r7 & mask]; // 52 load
r2 = r2 - r7; // 53 sub
r4 = r4 ^ r0; // 54 xor
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 55 add
r2 = r2 ^ ds[r4 & mask]; // 56 load
r0 = r1 * r4 + r0; // 57 mad
r3 = r3 ^ ds[r5 & mask]; // 58 load
r5 = r5 | r6; // 59 or
r6 = r5 * r7 + r6; // 60 mad
r4 = rotl_imm(r4, 28u); // 61 rotl
r5 = __umulhi(r5, r0); // 62 mulhi
r3 = r3 ^ ds[r6 & mask]; // 63 load
}
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
}
// Host-side launch wrappers. Declared in program.h, called from host.cu.
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) {
if (nSegments == 0u) return cudaErrorInvalidValue;
uint32_t block = 256u;
uint32_t grid = (nSegments + block - 1u) / block;
igneum_cache_fill<<<grid, block>>>(cache, nSegments);
return cudaGetLastError();
}
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
if (nItems == 0u) return cudaErrorInvalidValue;
uint32_t block = 256u;
uint32_t grid = (nItems + block - 1u) / block;
igneum_build<<<grid, block>>>(ds, cache, nItems);
return cudaGetLastError();
}
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
uint32_t nonces, uint32_t blockWarps) {
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
uint32_t block = 32u * blockWarps;
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
igneum_hash<<<nonces / block, block>>>(ds, out, baseNonce, mask);
return cudaGetLastError();
}
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
cudaFuncAttributes attr;
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash);
if (e != cudaSuccess) return e;
*numRegs = attr.numRegs;
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0);
}

File diff suppressed because it is too large Load diff

View file

@ -0,0 +1,123 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
// Header-bound twin of igneum_hash in kernel.cu: the init words come from a kernel argument, not SEEDW.
// Host declarations (also in program_bound.h if present):
// struct IgneumInitWords { uint32_t w[8]; };
// cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
// IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps);
// cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
#include <cuda_runtime.h>
#include <cstdint>
#include "program.h"
struct IgneumInitWords { uint32_t w[8]; };
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
x ^= x >> 16; x *= 0x7feb352du;
x ^= x >> 15; x *= 0x846ca68bu;
x ^= x >> 16;
return x;
}
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw) {
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
uint32_t nonce = baseNonce + gid;
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
{ uint32_t x = nonce ^ iw.w[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw.w[1]; }
{ uint32_t x = nonce ^ iw.w[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw.w[2]; }
{ uint32_t x = nonce ^ iw.w[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw.w[3]; }
{ uint32_t x = nonce ^ iw.w[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw.w[4]; }
{ uint32_t x = nonce ^ iw.w[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw.w[5]; }
{ uint32_t x = nonce ^ iw.w[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw.w[6]; }
{ uint32_t x = nonce ^ iw.w[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw.w[7]; }
{ uint32_t x = nonce ^ iw.w[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw.w[0]; }
for (uint32_t it = 0u; it < 8u; ++it) {
uint32_t sel = r0;
r4 = r4 + r5 + ((((sel >> 13u) & 1u) != 0u) ? 0x5810667au : 0xea86e152u); // 0 add
r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r0, 4); // 1 shfl
r3 = r3 + r2 + ((((sel >> 10u) & 1u) != 0u) ? 0x642e66dbu : 0x2cccb6cau); // 2 add
r0 = rotl_imm(r0, 19u); // 3 rotl
r7 = rotr_var(r7, r6); // 4 rotr
r7 = r7 + r4 + ((((sel >> 21u) & 1u) != 0u) ? 0xc1535555u : 0xee02465fu); // 5 add
r1 = __umulhi(r1, r7); // 6 mulhi
r4 = r4 ^ ds[r2 & mask]; // 7 load
r7 = r7 ^ ds[r4 & mask]; // 8 load
r0 = r0 ^ ds[r3 & mask]; // 9 load
r5 = r5 ^ ds[r1 & mask]; // 10 load
r1 = r1 ^ ds[r5 & mask]; // 11 load
r3 = __umulhi(r3, r5); // 12 mulhi
r1 = r1 ^ ds[r3 & mask]; // 13 load
r0 = r0 - r3; // 14 sub
r5 = r1 * r3 + r5; // 15 mad
r6 = __umulhi(r6, r1); // 16 mulhi
r5 = r5 + r2 + ((((sel >> 28u) & 1u) != 0u) ? 0x8b965b57u : 0x697b3d00u); // 17 add
r0 = __umulhi(r0, r6); // 18 mulhi
r5 = rotr_var(r5, r3); // 19 rotr
r5 = __umulhi(r5, r2); // 20 mulhi
r1 = r1 + r0 + ((((sel >> 1u) & 1u) != 0u) ? 0x6d7e8d05u : 0xebcf247au); // 21 add
r7 = r7 + r5 + ((((sel >> 12u) & 1u) != 0u) ? 0xb9e3577eu : 0xf66e7017u); // 22 add
r1 = __umulhi(r1, r5); // 23 mulhi
r2 = r2 - r5; // 24 sub
r7 = r7 + r4 + ((((sel >> 2u) & 1u) != 0u) ? 0x699ef1bbu : 0x08ffa6c7u); // 25 add
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 26 shfl
r7 = r7 + r1 + ((((sel >> 14u) & 1u) != 0u) ? 0xb4ead2fbu : 0xe60fea84u); // 27 add
r3 = r3 + r1 + ((((sel >> 6u) & 1u) != 0u) ? 0x8f30d21du : 0x65c76dabu); // 28 add
r2 = r2 ^ ds[r1 & mask]; // 29 load
r5 = r5 ^ ds[r7 & mask]; // 30 load
r2 = r2 ^ ds[r5 & mask]; // 31 load
r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r7, 4); // 32 shfl
r4 = r5 * r7 + r4; // 33 mad
r4 = r4 + r2 + ((((sel >> 21u) & 1u) != 0u) ? 0xc7ce690cu : 0x0480debeu); // 34 add
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r7, 8); // 35 shfl
r7 = r7 + r1 + ((((sel >> 2u) & 1u) != 0u) ? 0xe10c2c95u : 0xc53b542eu); // 36 add
r5 = r5 ^ r7; // 37 xor
r2 = r2 | r1; // 38 or
r1 = __umulhi(r1, r0); // 39 mulhi
r6 = rotl_imm(r6, 19u); // 40 rotl
r4 = __umulhi(r4, r6); // 41 mulhi
r6 = r6 - r0; // 42 sub
r6 = r6 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 43 shfl
r4 = r4 ^ ds[r2 & mask]; // 44 load
r1 = r1 ^ r3; // 45 xor
r7 = r7 ^ ds[r0 & mask]; // 46 load
r3 = r3 ^ ds[r1 & mask]; // 47 load
r5 = r5 * r3; // 48 mul
r1 = r1 - r5; // 49 sub
r2 = rotl_imm(r2, 8u); // 50 rotl
r1 = r1 + r5 + ((((sel >> 23u) & 1u) != 0u) ? 0x77b9bd43u : 0xa900fec4u); // 51 add
r4 = r4 ^ ds[r7 & mask]; // 52 load
r2 = r2 - r7; // 53 sub
r4 = r4 ^ r0; // 54 xor
r1 = r1 + r6 + ((((sel >> 14u) & 1u) != 0u) ? 0x83e825bfu : 0xe09f54e9u); // 55 add
r2 = r2 ^ ds[r4 & mask]; // 56 load
r0 = r1 * r4 + r0; // 57 mad
r3 = r3 ^ ds[r5 & mask]; // 58 load
r5 = r5 | r6; // 59 or
r6 = r5 * r7 + r6; // 60 mad
r4 = rotl_imm(r4, 28u); // 61 rotl
r5 = __umulhi(r5, r0); // 62 mulhi
r3 = r3 ^ ds[r6 & mask]; // 63 load
}
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
}
cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps) {
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
uint32_t block = 32u * blockWarps;
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
igneum_hash_bound<<<nonces / block, block>>>(ds, out, baseNonce, mask, iw);
return cudaGetLastError();
}
cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
cudaFuncAttributes attr;
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash_bound);
if (e != cudaSuccess) return e;
*numRegs = attr.numRegs;
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash_bound, (int)(32u * blockWarps), 0);
}

File diff suppressed because it is too large Load diff

File diff suppressed because it is too large Load diff

View file

@ -0,0 +1,72 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
// Program metadata for host.cu plus the launch wrappers defined in kernel.cu.
// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros.
#pragma once
#ifdef __cplusplus
#include <cstdint>
#else
#include <stdint.h>
#endif
#ifndef IGNEUM_NO_CUDA
#include <cuda_runtime.h>
#endif
#define IGNEUM_SEED_STRING "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000"
#define IGNEUM_SEED_BYTES_HEX "edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07"
#define IGNEUM_GENERATOR 2
#define IGNEUM_PROGRAM_ATTEMPT 0
#define IGNEUM_PROGRAM_ID 0x7f4a5ca0a3637820ull
#define IGNEUM_DAY_STRING "bytes:69676e65756d2d6461792ffa50000000000000"
#define IGNEUM_DAY_BYTES_HEX "69676e65756d2d6461792ffa50000000000000"
#define IGNEUM_DAY0 0xceed56d7u
#define IGNEUM_DAY1 0x9ba270d2u
#define IGNEUM_DATASET_LOG2 28
#define IGNEUM_MASK 0x0fffffffu
#define IGNEUM_LANES 32
#define IGNEUM_ITERATIONS 8
#define IGNEUM_INSTR_COUNT 64
#define IGNEUM_LOADS_PER_HASH 128
#define IGNEUM_WIDE_LOADS_PER_HASH 0
#define IGNEUM_OP_MIX "load=16 add=13 mulhi=9 shfl=5 sub=5 mad=4 rotl=4 xor=3 or=2 rotr=2 mul=1"
// Class v3 construction (Counter ASIC 2.0, 5 October 2026, docs/plans/mixer-x4.md): version 2 loads; the dataset item
// derivation applies the mixer IGNEUM_MIXER_MULT times per round (memhard.h), and the cache follows the growth rule.
#define IGNEUM_LOAD_CLASS "dr736"
#define IGNEUM_CLASS_MIXER_MULT 1
#define IGNEUM_CACHE_GROWTH 1 // 1: cache words = 2^(26 + doublings(day)), doublings = floor(log2(1 + day / 1460))
// Counter ASIC 3.0 item 2 (6 October 2026, docs/plans/counter-asic-3-derivation.md, a prototype, NOT class v3): the item
// derivation runs the day's drawn program (memhard.h: mh_round_0..8, IGNEUM_DERIVE_LEN instructions each) in place of the mixer.
#define IGNEUM_CLASS_DERIVE_LEN 736
#define IGNEUM_LOAD_SLOTS 16
#define IGNEUM_LOAD_MIX { 100, 0, 0 }
#define IGNEUM_LOAD_WIDTH_COUNTS { 16, 0, 0 } // loads of 4, 16, 64 bytes per program
#define IGNEUM_BYTES_PER_HASH 512
#define IGNEUM_FOLD_ROT 11
#define IGNEUM_FOLD_MUL 0x9e3779b1u
// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h)
#define IGNEUM_DATASET_MODE 1
#define IGNEUM_SEEDW_INIT { 0x667d0fbdu, 0x7b8e5963u, 0x31c67e5eu, 0x4529ddc6u, 0xef19d6d8u, 0xaccf6211u, 0xda0aed32u, 0xabc6df31u }
#define IGNEUM_KEY_INIT { 0xceed56d7u, 0x9ba270d2u, 0x82caab2du, 0x81ebce0eu, 0x12b6ecf1u, 0xd0f3fd7cu, 0xd872eefeu, 0xc158c7bdu }
#define IGNEUM_CACHE_LOG2_WORDS 26
#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6
#define IGNEUM_CACHE_SEGMENTS 65536u
#define IGNEUM_ITEM_ROUNDS 8
#define IGNEUM_DERIVE_LEN 736 // instructions per round program of the day's item-derivation program (Counter ASIC 3.0 item 2; memhard.h mh_round_0..8)
#define IGNEUM_DERIVE_ATTEMPT 0
#define IGNEUM_DERIVE_FINGERPRINT 0x771868df4e64d6abull
#define IGNEUM_DERIVE_INSTRS_PER_ITEM 6624
#define IGNEUM_DERIVE_GPU_OPS_PER_ITEM 10708
#define IGNEUM_DERIVE_CHIP_OPS_PER_ITEM 10072
#define IGNEUM_DERIVE_MULS_PER_ITEM 1475
#define IGNEUM_MIX_ROT_INIT { 17u, 12u, 20u, 23u, 7u, 3u, 27u, 16u }
#define IGNEUM_MIX_MUL_INIT { 0xf351d601u, 0xa3bb398fu, 0xb5a09e35u, 0x7509c9c1u, 0x6bbf31e9u, 0xfc849a79u, 0xded91851u, 0x8d9113d1u, 0x0ff15225u, 0x3a5bdd41u, 0xab533435u, 0xe1c55ad5u, 0xe6d3bd0du, 0x9d9ffbbdu, 0xbb2a3cf3u, 0x50a7c08du }
#define IGNEUM_MIX_RC_INIT { 0xc6892460u, 0x25b7228au, 0xcd515004u, 0x2846527au, 0xa6324241u, 0x36e3ec53u, 0x82961bacu, 0x0f97ba7du, 0xb6f921a9u, 0x3ada24e5u, 0xde20ab91u, 0x5378eeb2u, 0x7d161662u, 0x89353cc1u, 0xb1aa03a2u, 0x788acae6u }
#ifndef IGNEUM_NO_CUDA
// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError().
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments);
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems);
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
uint32_t nonces, uint32_t blockWarps);
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
#endif

File diff suppressed because one or more lines are too long

View file

@ -0,0 +1,109 @@
#include <metal_stdlib>
using namespace metal;
#define MASK 0x0fffffffu
constant uint SEEDW[8] = { 0x667d0fbdu, 0x7b8e5963u, 0x31c67e5eu, 0x4529ddc6u, 0xef19d6d8u, 0xaccf6211u, 0xda0aed32u, 0xabc6df31u };
inline uint splitmix32(uint x) {
x ^= x >> 16; x *= 0x7feb352du;
x ^= x >> 15; x *= 0x846ca68bu;
x ^= x >> 16;
return x;
}
inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31
inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
inline uint ds_elem(uint i, uint d0, uint d1) {
uint x = i ^ d0;
x *= 0x9E3779B1u; x ^= x >> 15;
x += d1;
x *= 0x85EBCA77u; x ^= x >> 13;
x *= 0xC2B2AE3Du; x ^= x >> 16;
return x;
}
kernel void igneum_hash(device const uint* dataset [[buffer(0)]],
device ulong* out [[buffer(1)]],
constant uint& baseNonce [[buffer(2)]],
uint gid [[thread_position_in_grid]]) {
uint nonce = baseNonce + gid;
uint r0, r1, r2, r3, r4, r5, r6, r7;
{ uint x = nonce ^ SEEDW[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ SEEDW[1]; }
{ uint x = nonce ^ SEEDW[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ SEEDW[2]; }
{ uint x = nonce ^ SEEDW[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ SEEDW[3]; }
{ uint x = nonce ^ SEEDW[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ SEEDW[4]; }
{ uint x = nonce ^ SEEDW[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ SEEDW[5]; }
{ uint x = nonce ^ SEEDW[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ SEEDW[6]; }
{ uint x = nonce ^ SEEDW[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ SEEDW[7]; }
{ uint x = nonce ^ SEEDW[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ SEEDW[0]; }
for (uint it = 0u; it < 8u; ++it) {
uint sel = r0;
r4 = r4 + r5 + select(0xea86e152u, 0x5810667au, ((sel >> 13u) & 1u) != 0u); // 0
r2 = r2 ^ simd_shuffle_xor(r0, (ushort)4); // 1
r3 = r3 + r2 + select(0x2cccb6cau, 0x642e66dbu, ((sel >> 10u) & 1u) != 0u); // 2
r0 = rotl_imm(r0, 19u); // 3
r7 = rotr_var(r7, r6); // 4
r7 = r7 + r4 + select(0xee02465fu, 0xc1535555u, ((sel >> 21u) & 1u) != 0u); // 5
r1 = mulhi(r1, r7); // 6
r4 = r4 ^ dataset[r2 & MASK]; // 7
r7 = r7 ^ dataset[r4 & MASK]; // 8
r0 = r0 ^ dataset[r3 & MASK]; // 9
r5 = r5 ^ dataset[r1 & MASK]; // 10
r1 = r1 ^ dataset[r5 & MASK]; // 11
r3 = mulhi(r3, r5); // 12
r1 = r1 ^ dataset[r3 & MASK]; // 13
r0 = r0 - r3; // 14
r5 = r1 * r3 + r5; // 15
r6 = mulhi(r6, r1); // 16
r5 = r5 + r2 + select(0x697b3d00u, 0x8b965b57u, ((sel >> 28u) & 1u) != 0u); // 17
r0 = mulhi(r0, r6); // 18
r5 = rotr_var(r5, r3); // 19
r5 = mulhi(r5, r2); // 20
r1 = r1 + r0 + select(0xebcf247au, 0x6d7e8d05u, ((sel >> 1u) & 1u) != 0u); // 21
r7 = r7 + r5 + select(0xf66e7017u, 0xb9e3577eu, ((sel >> 12u) & 1u) != 0u); // 22
r1 = mulhi(r1, r5); // 23
r2 = r2 - r5; // 24
r7 = r7 + r4 + select(0x08ffa6c7u, 0x699ef1bbu, ((sel >> 2u) & 1u) != 0u); // 25
r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 26
r7 = r7 + r1 + select(0xe60fea84u, 0xb4ead2fbu, ((sel >> 14u) & 1u) != 0u); // 27
r3 = r3 + r1 + select(0x65c76dabu, 0x8f30d21du, ((sel >> 6u) & 1u) != 0u); // 28
r2 = r2 ^ dataset[r1 & MASK]; // 29
r5 = r5 ^ dataset[r7 & MASK]; // 30
r2 = r2 ^ dataset[r5 & MASK]; // 31
r1 = r1 ^ simd_shuffle_xor(r7, (ushort)4); // 32
r4 = r5 * r7 + r4; // 33
r4 = r4 + r2 + select(0x0480debeu, 0xc7ce690cu, ((sel >> 21u) & 1u) != 0u); // 34
r3 = r3 ^ simd_shuffle_xor(r7, (ushort)8); // 35
r7 = r7 + r1 + select(0xc53b542eu, 0xe10c2c95u, ((sel >> 2u) & 1u) != 0u); // 36
r5 = r5 ^ r7; // 37
r2 = r2 | r1; // 38
r1 = mulhi(r1, r0); // 39
r6 = rotl_imm(r6, 19u); // 40
r4 = mulhi(r4, r6); // 41
r6 = r6 - r0; // 42
r6 = r6 ^ simd_shuffle_xor(r3, (ushort)4); // 43
r4 = r4 ^ dataset[r2 & MASK]; // 44
r1 = r1 ^ r3; // 45
r7 = r7 ^ dataset[r0 & MASK]; // 46
r3 = r3 ^ dataset[r1 & MASK]; // 47
r5 = r5 * r3; // 48
r1 = r1 - r5; // 49
r2 = rotl_imm(r2, 8u); // 50
r1 = r1 + r5 + select(0xa900fec4u, 0x77b9bd43u, ((sel >> 23u) & 1u) != 0u); // 51
r4 = r4 ^ dataset[r7 & MASK]; // 52
r2 = r2 - r7; // 53
r4 = r4 ^ r0; // 54
r1 = r1 + r6 + select(0xe09f54e9u, 0x83e825bfu, ((sel >> 14u) & 1u) != 0u); // 55
r2 = r2 ^ dataset[r4 & MASK]; // 56
r0 = r1 * r4 + r0; // 57
r3 = r3 ^ dataset[r5 & MASK]; // 58
r5 = r5 | r6; // 59
r6 = r5 * r7 + r6; // 60
r4 = rotl_imm(r4, 28u); // 61
r5 = mulhi(r5, r0); // 62
r3 = r3 ^ dataset[r6 & MASK]; // 63
}
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((ulong)hi << 32) | (ulong)lo;
}

View file

@ -0,0 +1,111 @@
#include <metal_stdlib>
using namespace metal;
#define MASK 0x0fffffffu
constant uint SEEDW[8] = { 0x667d0fbdu, 0x7b8e5963u, 0x31c67e5eu, 0x4529ddc6u, 0xef19d6d8u, 0xaccf6211u, 0xda0aed32u, 0xabc6df31u };
inline uint splitmix32(uint x) {
x ^= x >> 16; x *= 0x7feb352du;
x ^= x >> 15; x *= 0x846ca68bu;
x ^= x >> 16;
return x;
}
inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31
inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
inline uint ds_elem(uint i, uint d0, uint d1) {
uint x = i ^ d0;
x *= 0x9E3779B1u; x ^= x >> 15;
x += d1;
x *= 0x85EBCA77u; x ^= x >> 13;
x *= 0xC2B2AE3Du; x ^= x >> 16;
return x;
}
// Header-bound variant: the init words come from buffer 3 (bind.rs), not from SEEDW.
kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]],
device ulong* out [[buffer(1)]],
constant uint& baseNonce [[buffer(2)]],
constant uint* initw [[buffer(3)]],
uint gid [[thread_position_in_grid]]) {
uint nonce = baseNonce + gid;
uint r0, r1, r2, r3, r4, r5, r6, r7;
{ uint x = nonce ^ initw[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ initw[1]; }
{ uint x = nonce ^ initw[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ initw[2]; }
{ uint x = nonce ^ initw[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ initw[3]; }
{ uint x = nonce ^ initw[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ initw[4]; }
{ uint x = nonce ^ initw[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ initw[5]; }
{ uint x = nonce ^ initw[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ initw[6]; }
{ uint x = nonce ^ initw[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ initw[7]; }
{ uint x = nonce ^ initw[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ initw[0]; }
for (uint it = 0u; it < 8u; ++it) {
uint sel = r0;
r4 = r4 + r5 + select(0xea86e152u, 0x5810667au, ((sel >> 13u) & 1u) != 0u); // 0
r2 = r2 ^ simd_shuffle_xor(r0, (ushort)4); // 1
r3 = r3 + r2 + select(0x2cccb6cau, 0x642e66dbu, ((sel >> 10u) & 1u) != 0u); // 2
r0 = rotl_imm(r0, 19u); // 3
r7 = rotr_var(r7, r6); // 4
r7 = r7 + r4 + select(0xee02465fu, 0xc1535555u, ((sel >> 21u) & 1u) != 0u); // 5
r1 = mulhi(r1, r7); // 6
r4 = r4 ^ dataset[r2 & MASK]; // 7
r7 = r7 ^ dataset[r4 & MASK]; // 8
r0 = r0 ^ dataset[r3 & MASK]; // 9
r5 = r5 ^ dataset[r1 & MASK]; // 10
r1 = r1 ^ dataset[r5 & MASK]; // 11
r3 = mulhi(r3, r5); // 12
r1 = r1 ^ dataset[r3 & MASK]; // 13
r0 = r0 - r3; // 14
r5 = r1 * r3 + r5; // 15
r6 = mulhi(r6, r1); // 16
r5 = r5 + r2 + select(0x697b3d00u, 0x8b965b57u, ((sel >> 28u) & 1u) != 0u); // 17
r0 = mulhi(r0, r6); // 18
r5 = rotr_var(r5, r3); // 19
r5 = mulhi(r5, r2); // 20
r1 = r1 + r0 + select(0xebcf247au, 0x6d7e8d05u, ((sel >> 1u) & 1u) != 0u); // 21
r7 = r7 + r5 + select(0xf66e7017u, 0xb9e3577eu, ((sel >> 12u) & 1u) != 0u); // 22
r1 = mulhi(r1, r5); // 23
r2 = r2 - r5; // 24
r7 = r7 + r4 + select(0x08ffa6c7u, 0x699ef1bbu, ((sel >> 2u) & 1u) != 0u); // 25
r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 26
r7 = r7 + r1 + select(0xe60fea84u, 0xb4ead2fbu, ((sel >> 14u) & 1u) != 0u); // 27
r3 = r3 + r1 + select(0x65c76dabu, 0x8f30d21du, ((sel >> 6u) & 1u) != 0u); // 28
r2 = r2 ^ dataset[r1 & MASK]; // 29
r5 = r5 ^ dataset[r7 & MASK]; // 30
r2 = r2 ^ dataset[r5 & MASK]; // 31
r1 = r1 ^ simd_shuffle_xor(r7, (ushort)4); // 32
r4 = r5 * r7 + r4; // 33
r4 = r4 + r2 + select(0x0480debeu, 0xc7ce690cu, ((sel >> 21u) & 1u) != 0u); // 34
r3 = r3 ^ simd_shuffle_xor(r7, (ushort)8); // 35
r7 = r7 + r1 + select(0xc53b542eu, 0xe10c2c95u, ((sel >> 2u) & 1u) != 0u); // 36
r5 = r5 ^ r7; // 37
r2 = r2 | r1; // 38
r1 = mulhi(r1, r0); // 39
r6 = rotl_imm(r6, 19u); // 40
r4 = mulhi(r4, r6); // 41
r6 = r6 - r0; // 42
r6 = r6 ^ simd_shuffle_xor(r3, (ushort)4); // 43
r4 = r4 ^ dataset[r2 & MASK]; // 44
r1 = r1 ^ r3; // 45
r7 = r7 ^ dataset[r0 & MASK]; // 46
r3 = r3 ^ dataset[r1 & MASK]; // 47
r5 = r5 * r3; // 48
r1 = r1 - r5; // 49
r2 = rotl_imm(r2, 8u); // 50
r1 = r1 + r5 + select(0xa900fec4u, 0x77b9bd43u, ((sel >> 23u) & 1u) != 0u); // 51
r4 = r4 ^ dataset[r7 & MASK]; // 52
r2 = r2 - r7; // 53
r4 = r4 ^ r0; // 54
r1 = r1 + r6 + select(0xe09f54e9u, 0x83e825bfu, ((sel >> 14u) & 1u) != 0u); // 55
r2 = r2 ^ dataset[r4 & MASK]; // 56
r0 = r1 * r4 + r0; // 57
r3 = r3 ^ dataset[r5 & MASK]; // 58
r5 = r5 | r6; // 59
r6 = r5 * r7 + r6; // 60
r4 = rotl_imm(r4, 28u); // 61
r5 = mulhi(r5, r0); // 62
r3 = r3 ^ dataset[r6 & MASK]; // 63
}
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((ulong)hi << 32) | (ulong)lo;
}

View file

@ -0,0 +1,57 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000". Do not edit by hand.
// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset
#pragma once
#ifdef __cplusplus
#include <cstdint>
#else
#include <stdint.h>
#endif
#define IGNEUM_VEC_WARPS 3
static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u };
static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = {
{ // base nonce 0
0x435cd45db22fc83bull, 0xeb08b0eec4635b92ull, 0x5dac8870aed5c678ull, 0x2779f39382882c13ull, 0x465e11a6d64c87f4ull, 0x1a379f4f699ca9aeull, 0x03c6a8eb7ff849ebull, 0xbc1b878d6c277748ull,
0x69adeb1f2d29c0ebull, 0x375894dbc9868b50ull, 0x79d2d8527b76b932ull, 0x4dcefb4f0a2ec3fcull, 0xe5ad9fb702dd5966ull, 0xa296d3e34edcfbc6ull, 0xc5c839c600924f03ull, 0xdc7228f16532bfbeull,
0x81f595f907c9cb04ull, 0x1152850f76d4e90aull, 0x25446f27b34b11a9ull, 0xee986312f9e4e522ull, 0x3c5baa9645d0c99full, 0x144aff659162feb9ull, 0x497c9200b23e8e07ull, 0xc92501e7f6157dfeull,
0x8caf96daa4ca385aull, 0x00062ab96f31e514ull, 0x9508e98517444c91ull, 0x8a3f948cd7fb9797ull, 0xab156a88a1b77e41ull, 0x2408a16a50352fa9ull, 0x2237f2014a735ac3ull, 0x7da1a30d3b416f06ull
},
{ // base nonce 4096
0xb868dcaeacd8e012ull, 0xe88143fdca4bfab2ull, 0x9ee48b8d044a1dbeull, 0xc62e9314832773f3ull, 0x998beb89868141d7ull, 0xb730b261ff57cc38ull, 0xb886beb356bc26edull, 0x7472e83c4b9f208cull,
0x3c436e168b67d002ull, 0xb3c4c537b341c70full, 0xeea62a85a5c4d571ull, 0x31fb743150950537ull, 0x23d68caf4fb30f60ull, 0x7fa71a956c2cdab0ull, 0x2aaac765ae72764full, 0x5f109809ed36b7fbull,
0xebde9ca62f1a205cull, 0xa684f0ab53357506ull, 0x5b68d0c5246705a2ull, 0x923549e6f6aad924ull, 0x056ad95d6c1b0cffull, 0x96879d74b28d43ddull, 0xc7cb03983950f903ull, 0x09391d92779b4915ull,
0x92c061ed4bf9de47ull, 0xee84d54a38b36d07ull, 0x78da496aa8c9afd7ull, 0x0afe22f66ce1fc6eull, 0xf9576b3c7dc747c2ull, 0x16bfb648d41c69a2ull, 0xbf8f0934f1b1c327ull, 0x3781f5d303ba67f8ull
},
{ // base nonce 1000000
0xcd1ed6c453bf0f8cull, 0x19a6d570facc3c28ull, 0x99a0491a592d2705ull, 0x84fe73aa729eeaebull, 0x05d6496a724a48bcull, 0x427d4a53e3cd224bull, 0x566c3ffcc3dcfd37ull, 0x8a7edbb1da3510c0ull,
0x845f1b6d3c02f911ull, 0x9dc5b77664c6fb53ull, 0x7dc5df7872141de5ull, 0x281dd1300d03f493ull, 0xe56337a62077fb83ull, 0xbf3f0ceccb20a1e3ull, 0x61dd42e89af96505ull, 0x012b28dac792c2ffull,
0x703e2d57a7228db9ull, 0x93617feef48f6fa0ull, 0xa9b76a0cc401f839ull, 0xdb06176d832ccce9ull, 0xf1c45184eef7624bull, 0x2b6dc1cd2eae9856ull, 0xebb01b8100e4b19cull, 0x7f9305e00616dfe9ull,
0xdccda8f0b817fc71ull, 0x3a927aa1fd8d9bc8ull, 0xa4d82662a5fc6d5bull, 0x126c846254679b8bull, 0x8d792378f29a8a2bull, 0xa79da48e070f183full, 0xdfc380765068d276ull, 0x84aee1d7f9c8b2a6ull
}
};
// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455).
static const uint32_t IGNEUM_DS_HEAD[16] = {
0x0da55599u, 0x16ebf912u, 0x26ebba50u, 0x644a8670u, 0xb86da339u, 0xfdea8ce8u, 0xd8d38a59u, 0x5f930123u,
0xc99284a8u, 0xfa9fa963u, 0xfd990357u, 0x8d1aad8cu, 0xc5c03c95u, 0x82af7120u, 0x9fbf6800u, 0x913cc5e8u
};
static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u;
static const uint32_t IGNEUM_DS_LAST = 0x7c357826u;
// 64 sampled dataset words (index, value) computed on the Mac.
#define IGNEUM_DS_SAMPLES 64
static const uint32_t IGNEUM_DS_SAMPLE_INDEX[IGNEUM_DS_SAMPLES] = {
59471966u, 217795994u, 208353206u, 42483309u, 172547758u, 148076330u, 183853158u, 214389424u, 267488061u, 169781097u, 184093494u, 153880993u, 84977930u, 46426879u, 3093825u, 225364072u, 44593546u, 260713159u, 168250303u, 52384140u, 223401610u, 45554030u, 95410555u, 175039924u, 79171087u, 267580473u, 24168642u, 37981670u, 171551130u, 195559979u, 204611762u, 140997658u, 138925853u, 86637313u, 20736778u, 219665210u, 160430336u, 264654675u, 8013395u, 228945585u, 213884386u, 104419827u, 44185464u, 142737231u, 99284897u, 132475900u, 61861762u, 132056166u, 262388043u, 91878046u, 117353561u, 124768597u, 71352993u, 190698941u, 46055428u, 55281366u, 165145231u, 106810753u, 171985651u, 232085256u, 159510492u, 40072060u, 209107596u, 39023794u
};
static const uint32_t IGNEUM_DS_SAMPLE_VALUE[IGNEUM_DS_SAMPLES] = {
0xf4b0010au, 0xc82a3ca4u, 0xfc72371bu, 0x4c31b63cu, 0x0f0d86b4u, 0x0860e638u, 0x6016ffc1u, 0xf1cac2c4u, 0xb520efedu, 0x9eac76d6u, 0xca4ccfe1u, 0xa07e7108u, 0xc1c60eacu, 0xfdd31f2du, 0x2225e701u, 0x71f7fde6u, 0xb52aed34u, 0x75a3dc19u, 0xfe023726u, 0x308b9e1du, 0x6003222cu, 0xa5cc70a2u, 0x22c2940fu, 0xdbcd5db2u, 0xeed44d67u, 0xfca48f4cu, 0x2744926eu, 0x24f03506u, 0xa90cad7du, 0x47db755eu, 0x2051b7a9u, 0x18365d2du, 0x69c0d7e2u, 0x0f1f31c6u, 0x912e45bfu, 0x03e4035eu, 0xa30eb890u, 0xf65796a0u, 0x0b97d1e6u, 0xb0014ea8u, 0xa02bb1dfu, 0x87476bf2u, 0xc2f22b36u, 0x2f0452a5u, 0xb6d0c4c3u, 0xc6a765e2u, 0xf2b8cbd5u, 0x928e50adu, 0x85dbe732u, 0x2b9ee10cu, 0x2a1c9124u, 0x7f9a88e2u, 0x92da37eau, 0x519214cbu, 0x7869844eu, 0x19062ee1u, 0x045f0735u, 0x09f5e114u, 0x6082c8d4u, 0x6bc68181u, 0x700987d5u, 0x3bca401du, 0x7829f7aau, 0xe325295cu
};
// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^26 words.
static const uint32_t IGNEUM_CACHE_HEAD[16] = {
0xebd9055cu, 0x7eaf6b21u, 0x9b610855u, 0x134a8562u, 0xb10059bau, 0x7f459e58u, 0x8f3c7873u, 0x3523e609u,
0x75b8f4cau, 0x750d31b4u, 0x0dae0782u, 0x26f10015u, 0x7ba87a41u, 0x0efd8543u, 0x691d6368u, 0x8929d967u
};
static const uint32_t IGNEUM_CACHE_LAST[16] = {
0x51474edfu, 0xc3cfba93u, 0xf21454cfu, 0x9b79baacu, 0x4d4fce23u, 0xd701dfd5u, 0x37357ab3u, 0x1be693fau,
0xa701cb3bu, 0x7467c620u, 0x428184e8u, 0xf4010df0u, 0xe33aa88fu, 0xc78d6d62u, 0xae9ba9c0u, 0xf4bba97eu
};
static const uint64_t IGNEUM_CACHE_FNV64 = 0x448274a57f508cbcull;

View file

@ -0,0 +1,36 @@
{
"seed": "igneum-epoch/edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07/day/69676e65756d2d6461792ffa50000000000000",
"day": "bytes:69676e65756d2d6461792ffa50000000000000",
"dataset_mode": "memory-hard",
"dataset_log2_words": 28,
"mask": "0x0fffffff",
"lanes": 32,
"source": "igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset",
"warps": [
{"base_nonce": 0, "expected": [
"0x435cd45db22fc83b", "0xeb08b0eec4635b92", "0x5dac8870aed5c678", "0x2779f39382882c13", "0x465e11a6d64c87f4", "0x1a379f4f699ca9ae", "0x03c6a8eb7ff849eb", "0xbc1b878d6c277748",
"0x69adeb1f2d29c0eb", "0x375894dbc9868b50", "0x79d2d8527b76b932", "0x4dcefb4f0a2ec3fc", "0xe5ad9fb702dd5966", "0xa296d3e34edcfbc6", "0xc5c839c600924f03", "0xdc7228f16532bfbe",
"0x81f595f907c9cb04", "0x1152850f76d4e90a", "0x25446f27b34b11a9", "0xee986312f9e4e522", "0x3c5baa9645d0c99f", "0x144aff659162feb9", "0x497c9200b23e8e07", "0xc92501e7f6157dfe",
"0x8caf96daa4ca385a", "0x00062ab96f31e514", "0x9508e98517444c91", "0x8a3f948cd7fb9797", "0xab156a88a1b77e41", "0x2408a16a50352fa9", "0x2237f2014a735ac3", "0x7da1a30d3b416f06"
]},
{"base_nonce": 4096, "expected": [
"0xb868dcaeacd8e012", "0xe88143fdca4bfab2", "0x9ee48b8d044a1dbe", "0xc62e9314832773f3", "0x998beb89868141d7", "0xb730b261ff57cc38", "0xb886beb356bc26ed", "0x7472e83c4b9f208c",
"0x3c436e168b67d002", "0xb3c4c537b341c70f", "0xeea62a85a5c4d571", "0x31fb743150950537", "0x23d68caf4fb30f60", "0x7fa71a956c2cdab0", "0x2aaac765ae72764f", "0x5f109809ed36b7fb",
"0xebde9ca62f1a205c", "0xa684f0ab53357506", "0x5b68d0c5246705a2", "0x923549e6f6aad924", "0x056ad95d6c1b0cff", "0x96879d74b28d43dd", "0xc7cb03983950f903", "0x09391d92779b4915",
"0x92c061ed4bf9de47", "0xee84d54a38b36d07", "0x78da496aa8c9afd7", "0x0afe22f66ce1fc6e", "0xf9576b3c7dc747c2", "0x16bfb648d41c69a2", "0xbf8f0934f1b1c327", "0x3781f5d303ba67f8"
]},
{"base_nonce": 1000000, "expected": [
"0xcd1ed6c453bf0f8c", "0x19a6d570facc3c28", "0x99a0491a592d2705", "0x84fe73aa729eeaeb", "0x05d6496a724a48bc", "0x427d4a53e3cd224b", "0x566c3ffcc3dcfd37", "0x8a7edbb1da3510c0",
"0x845f1b6d3c02f911", "0x9dc5b77664c6fb53", "0x7dc5df7872141de5", "0x281dd1300d03f493", "0xe56337a62077fb83", "0xbf3f0ceccb20a1e3", "0x61dd42e89af96505", "0x012b28dac792c2ff",
"0x703e2d57a7228db9", "0x93617feef48f6fa0", "0xa9b76a0cc401f839", "0xdb06176d832ccce9", "0xf1c45184eef7624b", "0x2b6dc1cd2eae9856", "0xebb01b8100e4b19c", "0x7f9305e00616dfe9",
"0xdccda8f0b817fc71", "0x3a927aa1fd8d9bc8", "0xa4d82662a5fc6d5b", "0x126c846254679b8b", "0x8d792378f29a8a2b", "0xa79da48e070f183f", "0xdfc380765068d276", "0x84aee1d7f9c8b2a6"
]}
],
"dataset_head": ["0x0da55599", "0x16ebf912", "0x26ebba50", "0x644a8670", "0xb86da339", "0xfdea8ce8", "0xd8d38a59", "0x5f930123", "0xc99284a8", "0xfa9fa963", "0xfd990357", "0x8d1aad8c", "0xc5c03c95", "0x82af7120", "0x9fbf6800", "0x913cc5e8"],
"dataset_last_index": 268435455,
"dataset_last": "0x7c357826",
"dataset_samples": [{"index": 59471966, "value": "0xf4b0010a"}, {"index": 217795994, "value": "0xc82a3ca4"}, {"index": 208353206, "value": "0xfc72371b"}, {"index": 42483309, "value": "0x4c31b63c"}, {"index": 172547758, "value": "0x0f0d86b4"}, {"index": 148076330, "value": "0x0860e638"}, {"index": 183853158, "value": "0x6016ffc1"}, {"index": 214389424, "value": "0xf1cac2c4"}, {"index": 267488061, "value": "0xb520efed"}, {"index": 169781097, "value": "0x9eac76d6"}, {"index": 184093494, "value": "0xca4ccfe1"}, {"index": 153880993, "value": "0xa07e7108"}, {"index": 84977930, "value": "0xc1c60eac"}, {"index": 46426879, "value": "0xfdd31f2d"}, {"index": 3093825, "value": "0x2225e701"}, {"index": 225364072, "value": "0x71f7fde6"}, {"index": 44593546, "value": "0xb52aed34"}, {"index": 260713159, "value": "0x75a3dc19"}, {"index": 168250303, "value": "0xfe023726"}, {"index": 52384140, "value": "0x308b9e1d"}, {"index": 223401610, "value": "0x6003222c"}, {"index": 45554030, "value": "0xa5cc70a2"}, {"index": 95410555, "value": "0x22c2940f"}, {"index": 175039924, "value": "0xdbcd5db2"}, {"index": 79171087, "value": "0xeed44d67"}, {"index": 267580473, "value": "0xfca48f4c"}, {"index": 24168642, "value": "0x2744926e"}, {"index": 37981670, "value": "0x24f03506"}, {"index": 171551130, "value": "0xa90cad7d"}, {"index": 195559979, "value": "0x47db755e"}, {"index": 204611762, "value": "0x2051b7a9"}, {"index": 140997658, "value": "0x18365d2d"}, {"index": 138925853, "value": "0x69c0d7e2"}, {"index": 86637313, "value": "0x0f1f31c6"}, {"index": 20736778, "value": "0x912e45bf"}, {"index": 219665210, "value": "0x03e4035e"}, {"index": 160430336, "value": "0xa30eb890"}, {"index": 264654675, "value": "0xf65796a0"}, {"index": 8013395, "value": "0x0b97d1e6"}, {"index": 228945585, "value": "0xb0014ea8"}, {"index": 213884386, "value": "0xa02bb1df"}, {"index": 104419827, "value": "0x87476bf2"}, {"index": 44185464, "value": "0xc2f22b36"}, {"index": 142737231, "value": "0x2f0452a5"}, {"index": 99284897, "value": "0xb6d0c4c3"}, {"index": 132475900, "value": "0xc6a765e2"}, {"index": 61861762, "value": "0xf2b8cbd5"}, {"index": 132056166, "value": "0x928e50ad"}, {"index": 262388043, "value": "0x85dbe732"}, {"index": 91878046, "value": "0x2b9ee10c"}, {"index": 117353561, "value": "0x2a1c9124"}, {"index": 124768597, "value": "0x7f9a88e2"}, {"index": 71352993, "value": "0x92da37ea"}, {"index": 190698941, "value": "0x519214cb"}, {"index": 46055428, "value": "0x7869844e"}, {"index": 55281366, "value": "0x19062ee1"}, {"index": 165145231, "value": "0x045f0735"}, {"index": 106810753, "value": "0x09f5e114"}, {"index": 171985651, "value": "0x6082c8d4"}, {"index": 232085256, "value": "0x6bc68181"}, {"index": 159510492, "value": "0x700987d5"}, {"index": 40072060, "value": "0x3bca401d"}, {"index": 209107596, "value": "0x7829f7aa"}, {"index": 39023794, "value": "0xe325295c"}],
"cache_head": ["0xebd9055c", "0x7eaf6b21", "0x9b610855", "0x134a8562", "0xb10059ba", "0x7f459e58", "0x8f3c7873", "0x3523e609", "0x75b8f4ca", "0x750d31b4", "0x0dae0782", "0x26f10015", "0x7ba87a41", "0x0efd8543", "0x691d6368", "0x8929d967"],
"cache_last_line": ["0x51474edf", "0xc3cfba93", "0xf21454cf", "0x9b79baac", "0x4d4fce23", "0xd701dfd5", "0x37357ab3", "0x1be693fa", "0xa701cb3b", "0x7467c620", "0x428184e8", "0xf4010df0", "0xe33aa88f", "0xc78d6d62", "0xae9ba9c0", "0xf4bba97e"],
"cache_fnv1a64": "0x448274a57f508cbc"
}

File diff suppressed because it is too large Load diff

View file

@ -0,0 +1,164 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
// Bit-exact twin of the Metal kernel for the same seed (see proto-cuda/CHECKLIST.md and program.metal).
// Compiled ahead of time by nvcc together with proto-cuda/host.cu. No NVRTC.
#include <cuda_runtime.h>
#include <cstdint>
#include "program.h"
#include "memhard.h"
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
x ^= x >> 16; x *= 0x7feb352du;
x ^= x >> 15; x *= 0x846ca68bu;
x ^= x >> 16;
return x;
}
// n is a literal in 1..31 at every call site, so both shift amounts are in 1..31.
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
// n is masked to 0..31; the second shift amount is masked too, so n == 0 gives x.
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
__device__ __forceinline__ uint32_t ds_elem(uint32_t i, uint32_t d0, uint32_t d1) {
uint32_t x = i ^ d0;
x *= 0x9E3779B1u; x ^= x >> 15;
x += d1;
x *= 0x85EBCA77u; x ^= x >> 13;
x *= 0xC2B2AE3Du; x ^= x >> 16;
return x;
}
// Memory-hard dataset (MEMHARD.md). One thread per cache segment; one thread per 64-byte dataset item.
// The core functions (mh_cache_segment, mh_item) are in memhard.h and are also compiled for the host.
__global__ void igneum_cache_fill(uint32_t* cache, uint32_t nSegments) {
uint32_t seg = blockIdx.x * blockDim.x + threadIdx.x;
if (seg < nSegments) mh_cache_segment(cache, seg);
}
__global__ void igneum_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
uint32_t t = blockIdx.x * blockDim.x + threadIdx.x;
if (t < nItems) {
uint32_t s[16];
mh_item(cache, t, s);
uint32_t* d = ds + (size_t)t * 16u;
for (uint32_t i = 0u; i < 16u; ++i) d[i] = s[i];
}
}
// One hash per thread. blockDim.x is a multiple of 32; lane = threadIdx.x & 31 and every
// __shfl_xor_sync stays inside the lane's own warp, exactly like simd_shuffle_xor inside a
// 32-wide Metal SIMD group. Control flow is uniform, so the full 0xffffffff member mask is valid.
__global__ void igneum_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask) {
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
uint32_t nonce = baseNonce + gid;
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
{ uint32_t x = nonce ^ 0x67a9a7beu; x += 0x9e3779b9u; x = splitmix32(x); r0 = x ^ 0x1a155b25u; } // SEEDW[0], 0x9e3779b9u * 1u, SEEDW[1]
{ uint32_t x = nonce ^ 0x1a155b25u; x += 0x3c6ef372u; x = splitmix32(x); r1 = x ^ 0xfddfb732u; } // SEEDW[1], 0x9e3779b9u * 2u, SEEDW[2]
{ uint32_t x = nonce ^ 0xfddfb732u; x += 0xdaa66d2bu; x = splitmix32(x); r2 = x ^ 0x4b5af2e8u; } // SEEDW[2], 0x9e3779b9u * 3u, SEEDW[3]
{ uint32_t x = nonce ^ 0x4b5af2e8u; x += 0x78dde6e4u; x = splitmix32(x); r3 = x ^ 0xc55caf33u; } // SEEDW[3], 0x9e3779b9u * 4u, SEEDW[4]
{ uint32_t x = nonce ^ 0xc55caf33u; x += 0x1715609du; x = splitmix32(x); r4 = x ^ 0xa27c13b7u; } // SEEDW[4], 0x9e3779b9u * 5u, SEEDW[5]
{ uint32_t x = nonce ^ 0xa27c13b7u; x += 0xb54cda56u; x = splitmix32(x); r5 = x ^ 0x06628a48u; } // SEEDW[5], 0x9e3779b9u * 6u, SEEDW[6]
{ uint32_t x = nonce ^ 0x06628a48u; x += 0x5384540fu; x = splitmix32(x); r6 = x ^ 0x03852469u; } // SEEDW[6], 0x9e3779b9u * 7u, SEEDW[7]
{ uint32_t x = nonce ^ 0x03852469u; x += 0xf1bbcdc8u; x = splitmix32(x); r7 = x ^ 0x67a9a7beu; } // SEEDW[7], 0x9e3779b9u * 8u, SEEDW[0]
for (uint32_t it = 0u; it < 8u; ++it) {
uint32_t sel = r0;
r2 = r3 * r4 + r2; // 0 mad
r2 = r1 * r1 + r2; // 1 mad
r2 = r3 * r2 + r2; // 2 mad
r3 = r3 ^ r5; // 3 xor
r7 = r7 ^ ds[r2 & mask]; // 4 load
r5 = r5 ^ ds[r7 & mask]; // 5 load
r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 6 shfl
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 8); // 7 shfl
r1 = __umulhi(r1, r5); // 8 mulhi
r6 = rotr_var(r6, r3); // 9 rotr
r3 = r3 | r4; // 10 or
r4 = r4 ^ ds[r3 & mask]; // 11 load
r0 = __umulhi(r0, r4); // 12 mulhi
r5 = r5 + r1 + ((((sel >> 30u) & 1u) != 0u) ? 0xd3177981u : 0xc7934706u); // 13 add
r0 = r0 ^ ds[r4 & mask]; // 14 load
r2 = r2 - r4; // 15 sub
r2 = r2 ^ ds[r0 & mask]; // 16 load
r7 = r7 ^ ds[r2 & mask]; // 17 load
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 18 shfl
r5 = r5 * r0; // 19 mul
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 20 shfl
r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r4, 16); // 21 shfl
r6 = __umulhi(r6, r2); // 22 mulhi
r6 = r6 ^ ds[r1 & mask]; // 23 load
r5 = r5 * r0; // 24 mul
r5 = rotl_imm(r5, 19u); // 25 rotl
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r6, 2); // 26 shfl
r0 = r0 ^ r5; // 27 xor
r0 = r0 ^ r4; // 28 xor
r3 = r3 - r0; // 29 sub
r5 = r5 * r1; // 30 mul
r7 = r7 ^ ds[r2 & mask]; // 31 load
r1 = r1 ^ ds[r0 & mask]; // 32 load
r5 = r5 ^ r6; // 33 xor
r5 = r5 ^ ds[r1 & mask]; // 34 load
r0 = __umulhi(r0, r5); // 35 mulhi
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r2, 4); // 36 shfl
r7 = r7 ^ ds[r0 & mask]; // 37 load
r3 = r3 + r1 + ((((sel >> 27u) & 1u) != 0u) ? 0x230c005cu : 0x75ba2fadu); // 38 add
r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r5, 4); // 39 shfl
r2 = r2 ^ r5; // 40 xor
r3 = r6 * r3 + r3; // 41 mad
r6 = r6 - r7; // 42 sub
r7 = r7 ^ r0; // 43 xor
r1 = r1 ^ ds[r7 & mask]; // 44 load
r2 = r2 * r3; // 45 mul
r1 = __umulhi(r1, r5); // 46 mulhi
r4 = r4 - r3; // 47 sub
r2 = rotr_var(r2, r6); // 48 rotr
r3 = r3 ^ ds[r5 & mask]; // 49 load
r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 50 add
r0 = r0 * r2; // 51 mul
r0 = r0 + r2 + ((((sel >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); // 52 add
r1 = r1 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0x77b1520du : 0x2bb965afu); // 53 add
r7 = rotl_imm(r7, 14u); // 54 rotl
r3 = r3 + r7 + ((((sel >> 1u) & 1u) != 0u) ? 0xa54c55a0u : 0x7b0fe07au); // 55 add
r6 = r6 ^ ds[r7 & mask]; // 56 load
r1 = rotr_var(r1, r5); // 57 rotr
r5 = r5 ^ ds[r4 & mask]; // 58 load
r6 = r6 ^ ds[r2 & mask]; // 59 load
r3 = r5 * r0 + r3; // 60 mad
r5 = r5 + r7 + ((((sel >> 31u) & 1u) != 0u) ? 0xad7493e7u : 0xaf9dd72du); // 61 add
r4 = r4 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); // 62 add
r5 = rotl_imm(r5, 19u); // 63 rotl
}
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
}
// Host-side launch wrappers. Declared in program.h, called from host.cu.
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments) {
if (nSegments == 0u) return cudaErrorInvalidValue;
uint32_t block = 256u;
uint32_t grid = (nSegments + block - 1u) / block;
igneum_cache_fill<<<grid, block>>>(cache, nSegments);
return cudaGetLastError();
}
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems) {
if (nItems == 0u) return cudaErrorInvalidValue;
uint32_t block = 256u;
uint32_t grid = (nItems + block - 1u) / block;
igneum_build<<<grid, block>>>(ds, cache, nItems);
return cudaGetLastError();
}
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
uint32_t nonces, uint32_t blockWarps) {
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
uint32_t block = 32u * blockWarps;
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
igneum_hash<<<nonces / block, block>>>(ds, out, baseNonce, mask);
return cudaGetLastError();
}
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
cudaFuncAttributes attr;
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash);
if (e != cudaSuccess) return e;
*numRegs = attr.numRegs;
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash, (int)(32u * blockWarps), 0);
}

File diff suppressed because it is too large Load diff

View file

@ -0,0 +1,123 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
// Header-bound twin of igneum_hash in kernel.cu: the init words come from a kernel argument, not SEEDW.
// Host declarations (also in program_bound.h if present):
// struct IgneumInitWords { uint32_t w[8]; };
// cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
// IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps);
// cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
#include <cuda_runtime.h>
#include <cstdint>
#include "program.h"
struct IgneumInitWords { uint32_t w[8]; };
__device__ __forceinline__ uint32_t splitmix32(uint32_t x) {
x ^= x >> 16; x *= 0x7feb352du;
x ^= x >> 15; x *= 0x846ca68bu;
x ^= x >> 16;
return x;
}
__device__ __forceinline__ uint32_t rotl_imm(uint32_t x, uint32_t n) { return (x << n) | (x >> (32u - n)); }
__device__ __forceinline__ uint32_t rotr_var(uint32_t x, uint32_t n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
__global__ void igneum_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask, IgneumInitWords iw) {
uint32_t gid = blockIdx.x * blockDim.x + threadIdx.x;
uint32_t nonce = baseNonce + gid;
uint32_t r0, r1, r2, r3, r4, r5, r6, r7;
{ uint32_t x = nonce ^ iw.w[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ iw.w[1]; }
{ uint32_t x = nonce ^ iw.w[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ iw.w[2]; }
{ uint32_t x = nonce ^ iw.w[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ iw.w[3]; }
{ uint32_t x = nonce ^ iw.w[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ iw.w[4]; }
{ uint32_t x = nonce ^ iw.w[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ iw.w[5]; }
{ uint32_t x = nonce ^ iw.w[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ iw.w[6]; }
{ uint32_t x = nonce ^ iw.w[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ iw.w[7]; }
{ uint32_t x = nonce ^ iw.w[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ iw.w[0]; }
for (uint32_t it = 0u; it < 8u; ++it) {
uint32_t sel = r0;
r2 = r3 * r4 + r2; // 0 mad
r2 = r1 * r1 + r2; // 1 mad
r2 = r3 * r2 + r2; // 2 mad
r3 = r3 ^ r5; // 3 xor
r7 = r7 ^ ds[r2 & mask]; // 4 load
r5 = r5 ^ ds[r7 & mask]; // 5 load
r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r4, 8); // 6 shfl
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 8); // 7 shfl
r1 = __umulhi(r1, r5); // 8 mulhi
r6 = rotr_var(r6, r3); // 9 rotr
r3 = r3 | r4; // 10 or
r4 = r4 ^ ds[r3 & mask]; // 11 load
r0 = __umulhi(r0, r4); // 12 mulhi
r5 = r5 + r1 + ((((sel >> 30u) & 1u) != 0u) ? 0xd3177981u : 0xc7934706u); // 13 add
r0 = r0 ^ ds[r4 & mask]; // 14 load
r2 = r2 - r4; // 15 sub
r2 = r2 ^ ds[r0 & mask]; // 16 load
r7 = r7 ^ ds[r2 & mask]; // 17 load
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r3, 4); // 18 shfl
r5 = r5 * r0; // 19 mul
r3 = r3 ^ __shfl_xor_sync(0xffffffffu, r4, 2); // 20 shfl
r2 = r2 ^ __shfl_xor_sync(0xffffffffu, r4, 16); // 21 shfl
r6 = __umulhi(r6, r2); // 22 mulhi
r6 = r6 ^ ds[r1 & mask]; // 23 load
r5 = r5 * r0; // 24 mul
r5 = rotl_imm(r5, 19u); // 25 rotl
r7 = r7 ^ __shfl_xor_sync(0xffffffffu, r6, 2); // 26 shfl
r0 = r0 ^ r5; // 27 xor
r0 = r0 ^ r4; // 28 xor
r3 = r3 - r0; // 29 sub
r5 = r5 * r1; // 30 mul
r7 = r7 ^ ds[r2 & mask]; // 31 load
r1 = r1 ^ ds[r0 & mask]; // 32 load
r5 = r5 ^ r6; // 33 xor
r5 = r5 ^ ds[r1 & mask]; // 34 load
r0 = __umulhi(r0, r5); // 35 mulhi
r5 = r5 ^ __shfl_xor_sync(0xffffffffu, r2, 4); // 36 shfl
r7 = r7 ^ ds[r0 & mask]; // 37 load
r3 = r3 + r1 + ((((sel >> 27u) & 1u) != 0u) ? 0x230c005cu : 0x75ba2fadu); // 38 add
r1 = r1 ^ __shfl_xor_sync(0xffffffffu, r5, 4); // 39 shfl
r2 = r2 ^ r5; // 40 xor
r3 = r6 * r3 + r3; // 41 mad
r6 = r6 - r7; // 42 sub
r7 = r7 ^ r0; // 43 xor
r1 = r1 ^ ds[r7 & mask]; // 44 load
r2 = r2 * r3; // 45 mul
r1 = __umulhi(r1, r5); // 46 mulhi
r4 = r4 - r3; // 47 sub
r2 = rotr_var(r2, r6); // 48 rotr
r3 = r3 ^ ds[r5 & mask]; // 49 load
r1 = r1 + r5 + ((((sel >> 7u) & 1u) != 0u) ? 0x1907970cu : 0x81b8bc2cu); // 50 add
r0 = r0 * r2; // 51 mul
r0 = r0 + r2 + ((((sel >> 6u) & 1u) != 0u) ? 0x699fd448u : 0x4f92b968u); // 52 add
r1 = r1 + r0 + ((((sel >> 12u) & 1u) != 0u) ? 0x77b1520du : 0x2bb965afu); // 53 add
r7 = rotl_imm(r7, 14u); // 54 rotl
r3 = r3 + r7 + ((((sel >> 1u) & 1u) != 0u) ? 0xa54c55a0u : 0x7b0fe07au); // 55 add
r6 = r6 ^ ds[r7 & mask]; // 56 load
r1 = rotr_var(r1, r5); // 57 rotr
r5 = r5 ^ ds[r4 & mask]; // 58 load
r6 = r6 ^ ds[r2 & mask]; // 59 load
r3 = r5 * r0 + r3; // 60 mad
r5 = r5 + r7 + ((((sel >> 31u) & 1u) != 0u) ? 0xad7493e7u : 0xaf9dd72du); // 61 add
r4 = r4 + r6 + ((((sel >> 27u) & 1u) != 0u) ? 0x1e07c3d9u : 0x89841d87u); // 62 add
r5 = rotl_imm(r5, 19u); // 63 rotl
}
uint32_t lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint32_t hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((uint64_t)hi << 32) | (uint64_t)lo;
}
cudaError_t igneum_launch_hash_bound(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
IgneumInitWords iw, uint32_t nonces, uint32_t blockWarps) {
if (blockWarps == 0u || blockWarps > 32u) return cudaErrorInvalidValue;
uint32_t block = 32u * blockWarps;
if (nonces == 0u || (nonces % block) != 0u) return cudaErrorInvalidValue;
igneum_hash_bound<<<nonces / block, block>>>(ds, out, baseNonce, mask, iw);
return cudaGetLastError();
}
cudaError_t igneum_hash_bound_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps) {
cudaFuncAttributes attr;
cudaError_t e = cudaFuncGetAttributes(&attr, igneum_hash_bound);
if (e != cudaSuccess) return e;
*numRegs = attr.numRegs;
return cudaOccupancyMaxActiveBlocksPerMultiprocessor(blocksPerSM, igneum_hash_bound, (int)(32u * blockWarps), 0);
}

File diff suppressed because it is too large Load diff

File diff suppressed because it is too large Load diff

View file

@ -0,0 +1,72 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
// Program metadata for host.cu plus the launch wrappers defined in kernel.cu.
// Also included by proto-opencl/host.c (C99), which defines IGNEUM_NO_CUDA first and reads only the macros.
#pragma once
#ifdef __cplusplus
#include <cstdint>
#else
#include <stdint.h>
#endif
#ifndef IGNEUM_NO_CUDA
#include <cuda_runtime.h>
#endif
#define IGNEUM_SEED_STRING "igneum-genesis"
#define IGNEUM_SEED_BYTES_HEX "69676e65756d2d67656e65736973"
#define IGNEUM_GENERATOR 2
#define IGNEUM_PROGRAM_ATTEMPT 0
#define IGNEUM_PROGRAM_ID 0x72c1d8048aef9542ull
#define IGNEUM_DAY_STRING "2026-10-03"
#define IGNEUM_DAY_BYTES_HEX "6461792f323032362d31302d3033"
#define IGNEUM_DAY0 0x3067619fu
#define IGNEUM_DAY1 0x3c269176u
#define IGNEUM_DATASET_LOG2 28
#define IGNEUM_MASK 0x0fffffffu
#define IGNEUM_LANES 32
#define IGNEUM_ITERATIONS 8
#define IGNEUM_INSTR_COUNT 64
#define IGNEUM_LOADS_PER_HASH 128
#define IGNEUM_WIDE_LOADS_PER_HASH 0
#define IGNEUM_OP_MIX "load=16 add=8 shfl=8 xor=6 mad=5 mul=5 mulhi=5 sub=4 rotl=3 rotr=3 or=1"
// Class v3 construction (Counter ASIC 2.0, 5 October 2026, docs/plans/mixer-x4.md): version 2 loads; the dataset item
// derivation applies the mixer IGNEUM_MIXER_MULT times per round (memhard.h), and the cache follows the growth rule.
#define IGNEUM_LOAD_CLASS "dr736"
#define IGNEUM_CLASS_MIXER_MULT 1
#define IGNEUM_CACHE_GROWTH 1 // 1: cache words = 2^(26 + doublings(day)), doublings = floor(log2(1 + day / 1460))
// Counter ASIC 3.0 item 2 (6 October 2026, docs/plans/counter-asic-3-derivation.md, a prototype, NOT class v3): the item
// derivation runs the day's drawn program (memhard.h: mh_round_0..8, IGNEUM_DERIVE_LEN instructions each) in place of the mixer.
#define IGNEUM_CLASS_DERIVE_LEN 736
#define IGNEUM_LOAD_SLOTS 16
#define IGNEUM_LOAD_MIX { 100, 0, 0 }
#define IGNEUM_LOAD_WIDTH_COUNTS { 16, 0, 0 } // loads of 4, 16, 64 bytes per program
#define IGNEUM_BYTES_PER_HASH 512
#define IGNEUM_FOLD_ROT 11
#define IGNEUM_FOLD_MUL 0x9e3779b1u
// 0 = closed-form dataset (ds_elem), 1 = memory-hard cache construction (MEMHARD.md, memhard.h)
#define IGNEUM_DATASET_MODE 1
#define IGNEUM_SEEDW_INIT { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u }
#define IGNEUM_KEY_INIT { 0x3067619fu, 0x3c269176u, 0x84a03b03u, 0xf8c63294u, 0xff977c5bu, 0xe60def3eu, 0x63630141u, 0xb8fbcb58u }
#define IGNEUM_CACHE_LOG2_WORDS 26
#define IGNEUM_CACHE_SEGMENT_LOG2_LINES 6
#define IGNEUM_CACHE_SEGMENTS 65536u
#define IGNEUM_ITEM_ROUNDS 8
#define IGNEUM_DERIVE_LEN 736 // instructions per round program of the day's item-derivation program (Counter ASIC 3.0 item 2; memhard.h mh_round_0..8)
#define IGNEUM_DERIVE_ATTEMPT 0
#define IGNEUM_DERIVE_FINGERPRINT 0x463535d01511350dull
#define IGNEUM_DERIVE_INSTRS_PER_ITEM 6624
#define IGNEUM_DERIVE_GPU_OPS_PER_ITEM 10659
#define IGNEUM_DERIVE_CHIP_OPS_PER_ITEM 9992
#define IGNEUM_DERIVE_MULS_PER_ITEM 1461
#define IGNEUM_MIX_ROT_INIT { 20u, 20u, 19u, 4u, 26u, 3u, 3u, 27u }
#define IGNEUM_MIX_MUL_INIT { 0x42146205u, 0x52cbe0fbu, 0x7ecf4a03u, 0x6728907fu, 0xd81d9751u, 0x132952c3u, 0xf60de277u, 0x05358035u, 0xbaf6499du, 0xe4db9667u, 0x3e98f45du, 0xd0004eddu, 0x2691630du, 0x9beb3bcfu, 0xab310379u, 0x99cfb423u }
#define IGNEUM_MIX_RC_INIT { 0xbab68293u, 0xcc162340u, 0x6ce151ccu, 0xe62b8997u, 0xc9c80297u, 0xf74a1654u, 0x3d704af5u, 0x3cf522b7u, 0x2b9cac04u, 0xa880ac10u, 0x13e5dd1du, 0x6fc3e233u, 0x2d83eeacu, 0x9006e8bfu, 0x2c4b5362u, 0x31b49ee2u }
#ifndef IGNEUM_NO_CUDA
// Defined in kernel.cu. All launch on the default stream and return cudaGetLastError().
cudaError_t igneum_launch_cache_fill(uint32_t* cache, uint32_t nSegments);
cudaError_t igneum_launch_build(uint32_t* ds, const uint32_t* cache, uint32_t nItems);
cudaError_t igneum_launch_hash(const uint32_t* ds, uint64_t* out, uint32_t baseNonce, uint32_t mask,
uint32_t nonces, uint32_t blockWarps);
cudaError_t igneum_hash_info(int* numRegs, int* blocksPerSM, uint32_t blockWarps);
#endif

File diff suppressed because one or more lines are too long

View file

@ -0,0 +1,109 @@
#include <metal_stdlib>
using namespace metal;
#define MASK 0x0fffffffu
constant uint SEEDW[8] = { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u };
inline uint splitmix32(uint x) {
x ^= x >> 16; x *= 0x7feb352du;
x ^= x >> 15; x *= 0x846ca68bu;
x ^= x >> 16;
return x;
}
inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31
inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
inline uint ds_elem(uint i, uint d0, uint d1) {
uint x = i ^ d0;
x *= 0x9E3779B1u; x ^= x >> 15;
x += d1;
x *= 0x85EBCA77u; x ^= x >> 13;
x *= 0xC2B2AE3Du; x ^= x >> 16;
return x;
}
kernel void igneum_hash(device const uint* dataset [[buffer(0)]],
device ulong* out [[buffer(1)]],
constant uint& baseNonce [[buffer(2)]],
uint gid [[thread_position_in_grid]]) {
uint nonce = baseNonce + gid;
uint r0, r1, r2, r3, r4, r5, r6, r7;
{ uint x = nonce ^ SEEDW[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ SEEDW[1]; }
{ uint x = nonce ^ SEEDW[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ SEEDW[2]; }
{ uint x = nonce ^ SEEDW[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ SEEDW[3]; }
{ uint x = nonce ^ SEEDW[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ SEEDW[4]; }
{ uint x = nonce ^ SEEDW[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ SEEDW[5]; }
{ uint x = nonce ^ SEEDW[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ SEEDW[6]; }
{ uint x = nonce ^ SEEDW[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ SEEDW[7]; }
{ uint x = nonce ^ SEEDW[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ SEEDW[0]; }
for (uint it = 0u; it < 8u; ++it) {
uint sel = r0;
r2 = r3 * r4 + r2; // 0
r2 = r1 * r1 + r2; // 1
r2 = r3 * r2 + r2; // 2
r3 = r3 ^ r5; // 3
r7 = r7 ^ dataset[r2 & MASK]; // 4
r5 = r5 ^ dataset[r7 & MASK]; // 5
r1 = r1 ^ simd_shuffle_xor(r4, (ushort)8); // 6
r7 = r7 ^ simd_shuffle_xor(r3, (ushort)8); // 7
r1 = mulhi(r1, r5); // 8
r6 = rotr_var(r6, r3); // 9
r3 = r3 | r4; // 10
r4 = r4 ^ dataset[r3 & MASK]; // 11
r0 = mulhi(r0, r4); // 12
r5 = r5 + r1 + select(0xc7934706u, 0xd3177981u, ((sel >> 30u) & 1u) != 0u); // 13
r0 = r0 ^ dataset[r4 & MASK]; // 14
r2 = r2 - r4; // 15
r2 = r2 ^ dataset[r0 & MASK]; // 16
r7 = r7 ^ dataset[r2 & MASK]; // 17
r7 = r7 ^ simd_shuffle_xor(r3, (ushort)4); // 18
r5 = r5 * r0; // 19
r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 20
r2 = r2 ^ simd_shuffle_xor(r4, (ushort)16); // 21
r6 = mulhi(r6, r2); // 22
r6 = r6 ^ dataset[r1 & MASK]; // 23
r5 = r5 * r0; // 24
r5 = rotl_imm(r5, 19u); // 25
r7 = r7 ^ simd_shuffle_xor(r6, (ushort)2); // 26
r0 = r0 ^ r5; // 27
r0 = r0 ^ r4; // 28
r3 = r3 - r0; // 29
r5 = r5 * r1; // 30
r7 = r7 ^ dataset[r2 & MASK]; // 31
r1 = r1 ^ dataset[r0 & MASK]; // 32
r5 = r5 ^ r6; // 33
r5 = r5 ^ dataset[r1 & MASK]; // 34
r0 = mulhi(r0, r5); // 35
r5 = r5 ^ simd_shuffle_xor(r2, (ushort)4); // 36
r7 = r7 ^ dataset[r0 & MASK]; // 37
r3 = r3 + r1 + select(0x75ba2fadu, 0x230c005cu, ((sel >> 27u) & 1u) != 0u); // 38
r1 = r1 ^ simd_shuffle_xor(r5, (ushort)4); // 39
r2 = r2 ^ r5; // 40
r3 = r6 * r3 + r3; // 41
r6 = r6 - r7; // 42
r7 = r7 ^ r0; // 43
r1 = r1 ^ dataset[r7 & MASK]; // 44
r2 = r2 * r3; // 45
r1 = mulhi(r1, r5); // 46
r4 = r4 - r3; // 47
r2 = rotr_var(r2, r6); // 48
r3 = r3 ^ dataset[r5 & MASK]; // 49
r1 = r1 + r5 + select(0x81b8bc2cu, 0x1907970cu, ((sel >> 7u) & 1u) != 0u); // 50
r0 = r0 * r2; // 51
r0 = r0 + r2 + select(0x4f92b968u, 0x699fd448u, ((sel >> 6u) & 1u) != 0u); // 52
r1 = r1 + r0 + select(0x2bb965afu, 0x77b1520du, ((sel >> 12u) & 1u) != 0u); // 53
r7 = rotl_imm(r7, 14u); // 54
r3 = r3 + r7 + select(0x7b0fe07au, 0xa54c55a0u, ((sel >> 1u) & 1u) != 0u); // 55
r6 = r6 ^ dataset[r7 & MASK]; // 56
r1 = rotr_var(r1, r5); // 57
r5 = r5 ^ dataset[r4 & MASK]; // 58
r6 = r6 ^ dataset[r2 & MASK]; // 59
r3 = r5 * r0 + r3; // 60
r5 = r5 + r7 + select(0xaf9dd72du, 0xad7493e7u, ((sel >> 31u) & 1u) != 0u); // 61
r4 = r4 + r6 + select(0x89841d87u, 0x1e07c3d9u, ((sel >> 27u) & 1u) != 0u); // 62
r5 = rotl_imm(r5, 19u); // 63
}
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((ulong)hi << 32) | (ulong)lo;
}

View file

@ -0,0 +1,111 @@
#include <metal_stdlib>
using namespace metal;
#define MASK 0x0fffffffu
constant uint SEEDW[8] = { 0x67a9a7beu, 0x1a155b25u, 0xfddfb732u, 0x4b5af2e8u, 0xc55caf33u, 0xa27c13b7u, 0x06628a48u, 0x03852469u };
inline uint splitmix32(uint x) {
x ^= x >> 16; x *= 0x7feb352du;
x ^= x >> 15; x *= 0x846ca68bu;
x ^= x >> 16;
return x;
}
inline uint rotl_imm(uint x, uint n) { return (x << n) | (x >> (32u - n)); } // n in 1..31
inline uint rotr_var(uint x, uint n) { n &= 31u; return (x >> n) | (x << ((32u - n) & 31u)); }
inline uint ds_elem(uint i, uint d0, uint d1) {
uint x = i ^ d0;
x *= 0x9E3779B1u; x ^= x >> 15;
x += d1;
x *= 0x85EBCA77u; x ^= x >> 13;
x *= 0xC2B2AE3Du; x ^= x >> 16;
return x;
}
// Header-bound variant: the init words come from buffer 3 (bind.rs), not from SEEDW.
kernel void igneum_hash_bound(device const uint* dataset [[buffer(0)]],
device ulong* out [[buffer(1)]],
constant uint& baseNonce [[buffer(2)]],
constant uint* initw [[buffer(3)]],
uint gid [[thread_position_in_grid]]) {
uint nonce = baseNonce + gid;
uint r0, r1, r2, r3, r4, r5, r6, r7;
{ uint x = nonce ^ initw[0]; x += 0x9e3779b9u * 1u; x = splitmix32(x); r0 = x ^ initw[1]; }
{ uint x = nonce ^ initw[1]; x += 0x9e3779b9u * 2u; x = splitmix32(x); r1 = x ^ initw[2]; }
{ uint x = nonce ^ initw[2]; x += 0x9e3779b9u * 3u; x = splitmix32(x); r2 = x ^ initw[3]; }
{ uint x = nonce ^ initw[3]; x += 0x9e3779b9u * 4u; x = splitmix32(x); r3 = x ^ initw[4]; }
{ uint x = nonce ^ initw[4]; x += 0x9e3779b9u * 5u; x = splitmix32(x); r4 = x ^ initw[5]; }
{ uint x = nonce ^ initw[5]; x += 0x9e3779b9u * 6u; x = splitmix32(x); r5 = x ^ initw[6]; }
{ uint x = nonce ^ initw[6]; x += 0x9e3779b9u * 7u; x = splitmix32(x); r6 = x ^ initw[7]; }
{ uint x = nonce ^ initw[7]; x += 0x9e3779b9u * 8u; x = splitmix32(x); r7 = x ^ initw[0]; }
for (uint it = 0u; it < 8u; ++it) {
uint sel = r0;
r2 = r3 * r4 + r2; // 0
r2 = r1 * r1 + r2; // 1
r2 = r3 * r2 + r2; // 2
r3 = r3 ^ r5; // 3
r7 = r7 ^ dataset[r2 & MASK]; // 4
r5 = r5 ^ dataset[r7 & MASK]; // 5
r1 = r1 ^ simd_shuffle_xor(r4, (ushort)8); // 6
r7 = r7 ^ simd_shuffle_xor(r3, (ushort)8); // 7
r1 = mulhi(r1, r5); // 8
r6 = rotr_var(r6, r3); // 9
r3 = r3 | r4; // 10
r4 = r4 ^ dataset[r3 & MASK]; // 11
r0 = mulhi(r0, r4); // 12
r5 = r5 + r1 + select(0xc7934706u, 0xd3177981u, ((sel >> 30u) & 1u) != 0u); // 13
r0 = r0 ^ dataset[r4 & MASK]; // 14
r2 = r2 - r4; // 15
r2 = r2 ^ dataset[r0 & MASK]; // 16
r7 = r7 ^ dataset[r2 & MASK]; // 17
r7 = r7 ^ simd_shuffle_xor(r3, (ushort)4); // 18
r5 = r5 * r0; // 19
r3 = r3 ^ simd_shuffle_xor(r4, (ushort)2); // 20
r2 = r2 ^ simd_shuffle_xor(r4, (ushort)16); // 21
r6 = mulhi(r6, r2); // 22
r6 = r6 ^ dataset[r1 & MASK]; // 23
r5 = r5 * r0; // 24
r5 = rotl_imm(r5, 19u); // 25
r7 = r7 ^ simd_shuffle_xor(r6, (ushort)2); // 26
r0 = r0 ^ r5; // 27
r0 = r0 ^ r4; // 28
r3 = r3 - r0; // 29
r5 = r5 * r1; // 30
r7 = r7 ^ dataset[r2 & MASK]; // 31
r1 = r1 ^ dataset[r0 & MASK]; // 32
r5 = r5 ^ r6; // 33
r5 = r5 ^ dataset[r1 & MASK]; // 34
r0 = mulhi(r0, r5); // 35
r5 = r5 ^ simd_shuffle_xor(r2, (ushort)4); // 36
r7 = r7 ^ dataset[r0 & MASK]; // 37
r3 = r3 + r1 + select(0x75ba2fadu, 0x230c005cu, ((sel >> 27u) & 1u) != 0u); // 38
r1 = r1 ^ simd_shuffle_xor(r5, (ushort)4); // 39
r2 = r2 ^ r5; // 40
r3 = r6 * r3 + r3; // 41
r6 = r6 - r7; // 42
r7 = r7 ^ r0; // 43
r1 = r1 ^ dataset[r7 & MASK]; // 44
r2 = r2 * r3; // 45
r1 = mulhi(r1, r5); // 46
r4 = r4 - r3; // 47
r2 = rotr_var(r2, r6); // 48
r3 = r3 ^ dataset[r5 & MASK]; // 49
r1 = r1 + r5 + select(0x81b8bc2cu, 0x1907970cu, ((sel >> 7u) & 1u) != 0u); // 50
r0 = r0 * r2; // 51
r0 = r0 + r2 + select(0x4f92b968u, 0x699fd448u, ((sel >> 6u) & 1u) != 0u); // 52
r1 = r1 + r0 + select(0x2bb965afu, 0x77b1520du, ((sel >> 12u) & 1u) != 0u); // 53
r7 = rotl_imm(r7, 14u); // 54
r3 = r3 + r7 + select(0x7b0fe07au, 0xa54c55a0u, ((sel >> 1u) & 1u) != 0u); // 55
r6 = r6 ^ dataset[r7 & MASK]; // 56
r1 = rotr_var(r1, r5); // 57
r5 = r5 ^ dataset[r4 & MASK]; // 58
r6 = r6 ^ dataset[r2 & MASK]; // 59
r3 = r5 * r0 + r3; // 60
r5 = r5 + r7 + select(0xaf9dd72du, 0xad7493e7u, ((sel >> 31u) & 1u) != 0u); // 61
r4 = r4 + r6 + select(0x89841d87u, 0x1e07c3d9u, ((sel >> 27u) & 1u) != 0u); // 62
r5 = rotl_imm(r5, 19u); // 63
}
uint lo = r0 ^ rotl_imm(r1, 7u) ^ rotl_imm(r2, 14u) ^ rotl_imm(r3, 21u);
uint hi = r4 ^ rotl_imm(r5, 9u) ^ rotl_imm(r6, 18u) ^ rotl_imm(r7, 27u);
out[gid] = ((ulong)hi << 32) | (ulong)lo;
}

View file

@ -0,0 +1,57 @@
// Generated by igneum-pow export (generator v2) for seed "igneum-genesis". Do not edit by hand.
// Expected outputs: igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset
#pragma once
#ifdef __cplusplus
#include <cstdint>
#else
#include <stdint.h>
#endif
#define IGNEUM_VEC_WARPS 3
static const uint32_t IGNEUM_VEC_BASE[IGNEUM_VEC_WARPS] = { 0u, 4096u, 1000000u };
static const uint64_t IGNEUM_VEC_OUT[IGNEUM_VEC_WARPS][32] = {
{ // base nonce 0
0xe23d389f3eea0c83ull, 0xdb20ad461b5a2b40ull, 0x19f7ba3b1f71682cull, 0x85f636966fc9aca0ull, 0x6487ed96bdc27c42ull, 0xc490c3313b8029a3ull, 0xe361eba9844902d4ull, 0xf08e800494c84491ull,
0xa8bad93be02c1cacull, 0xe4e1f2c75c891797ull, 0x63b050548aa5d9b3ull, 0x7b8851cd9a064497ull, 0x9873c6fdd34355abull, 0xbe51e127ed02c404ull, 0x1aa1ee5f4c909067ull, 0x422684f3c86c5058ull,
0x42c3e7bc0f3578a3ull, 0x3e34aef37a263942ull, 0xcf0a9f5f13bd905cull, 0x44ba872dba90ac37ull, 0xbd590fe07552ad9bull, 0x62b32d43d066161eull, 0x213423a10a13d5d1ull, 0x6bca949b59253f84ull,
0xfdb3dba0dcb2951bull, 0x1f871d6c8a19d696ull, 0x413da5a7e4e5f500ull, 0xa448622cdd32bb6dull, 0x5193eba1f8804e2aull, 0x567869c0cfdb19acull, 0x2b2b619c048918dcull, 0x6605db059b381bd9ull
},
{ // base nonce 4096
0xfdb4b214da8ce292ull, 0xa04a87297fc6a9c0ull, 0x968cecee7a3bcd99ull, 0xa41939b15019ad50ull, 0xffc694d1b04e4da2ull, 0xc262e3973d436fdaull, 0x665a2ce13954b603ull, 0x6ece51e9d9d5d921ull,
0x5ff85c609da62800ull, 0xde9a8e850dd6c1e9ull, 0x39e876e19ae70187ull, 0xdb55c10b53c52bf2ull, 0x610845800cd92811ull, 0x696316e0c89c9d52ull, 0x3f18086d783e7d7aull, 0xb0d29c8dfb421fd5ull,
0x580543d5977c1138ull, 0xaba172c0bc028b29ull, 0x43ffd7726859240dull, 0xf04dcb8e1b09dcafull, 0xa1ebc5eb5873c4fbull, 0x4141b328723483e0ull, 0xba9fad0940aa4905ull, 0xed0671189b176bd0ull,
0x7ed7083b52f320fdull, 0x30c86e3033463bf7ull, 0x8169974907e0f4a3ull, 0x89fd7aad0e41bf03ull, 0x1aab38280fa3ee63ull, 0x6f525aeaa4c8a186ull, 0x5bd2a274c3905d8bull, 0xdb698d03437d74f7ull
},
{ // base nonce 1000000
0x534671b1bf5cea36ull, 0xe3f81c66bf210aa3ull, 0xb79ffc3d3cca6731ull, 0x5df44d8e9e3ac008ull, 0x445792668a534b0dull, 0x293f8c07ed96f576ull, 0x5b2e23716649e117ull, 0x55aa3dcc43378f3full,
0xfec021749867a3a8ull, 0x3f06d1c7cabc290full, 0x875d22f8b0811415ull, 0x0cd09e82747f0eefull, 0x2bcf3b7cd81a64c8ull, 0x23c180008f4a8677ull, 0x8bae18c724c5810aull, 0xb6a3fe36acf8b5e3ull,
0xc145ffcaa45352baull, 0xf1a0f8edaaf82d2dull, 0x205ba1313b3852b1ull, 0x6922192747873b0bull, 0x9937304a65a44a7dull, 0xca589d2fe74a7e01ull, 0xbecb0ddd75ba8dddull, 0x752e223db126522aull,
0x92f919e76899e682ull, 0x6c9bf07fddaae5efull, 0xfa7a39ed604236c0ull, 0x1478fb9d0050bc69ull, 0x9e078959178a2797ull, 0x7e6311fd8ff88b13ull, 0x911c33694145c818ull, 0x733b123353f13d21ull
}
};
// Dataset self-test: dataset[0..15] and dataset[IGNEUM_MASK] (268435455).
static const uint32_t IGNEUM_DS_HEAD[16] = {
0xe73af5cau, 0x45d77effu, 0xb640f499u, 0x351c2ca1u, 0xfed2d16fu, 0x1199bd1du, 0x1db2bbecu, 0x4b1769deu,
0x45b535eau, 0x414c389cu, 0xb8cada9cu, 0x32a3eb0fu, 0x405305cau, 0xddbd43e4u, 0xbdd280acu, 0x70358961u
};
static const uint32_t IGNEUM_DS_LAST_INDEX = 268435455u;
static const uint32_t IGNEUM_DS_LAST = 0x7cdbf6b5u;
// 64 sampled dataset words (index, value) computed on the Mac.
#define IGNEUM_DS_SAMPLES 64
static const uint32_t IGNEUM_DS_SAMPLE_INDEX[IGNEUM_DS_SAMPLES] = {
59471966u, 217795994u, 208353206u, 42483309u, 172547758u, 148076330u, 183853158u, 214389424u, 267488061u, 169781097u, 184093494u, 153880993u, 84977930u, 46426879u, 3093825u, 225364072u, 44593546u, 260713159u, 168250303u, 52384140u, 223401610u, 45554030u, 95410555u, 175039924u, 79171087u, 267580473u, 24168642u, 37981670u, 171551130u, 195559979u, 204611762u, 140997658u, 138925853u, 86637313u, 20736778u, 219665210u, 160430336u, 264654675u, 8013395u, 228945585u, 213884386u, 104419827u, 44185464u, 142737231u, 99284897u, 132475900u, 61861762u, 132056166u, 262388043u, 91878046u, 117353561u, 124768597u, 71352993u, 190698941u, 46055428u, 55281366u, 165145231u, 106810753u, 171985651u, 232085256u, 159510492u, 40072060u, 209107596u, 39023794u
};
static const uint32_t IGNEUM_DS_SAMPLE_VALUE[IGNEUM_DS_SAMPLES] = {
0x0c799a6du, 0xa21ecab8u, 0xcc71d9a3u, 0xb0d7af9cu, 0x1925c929u, 0x457795ffu, 0x4718201au, 0x985cde9eu, 0x4bb9b17eu, 0x677407ddu, 0x6093de87u, 0x2a754135u, 0xbc58f70du, 0xd805eb4bu, 0xf453a94du, 0x5a28b420u, 0xe92fb02bu, 0x8a48d35au, 0x2024c448u, 0x48ac5a95u, 0x8fa7880cu, 0xfff38e1bu, 0x4e98033au, 0xb6a33ee4u, 0xdfb0f94fu, 0x80d99b65u, 0x7d1c8bdeu, 0x06f4c2c4u, 0x05e3527cu, 0x8690f1beu, 0xa392261au, 0xac2293ffu, 0x0d067d65u, 0x7133a61eu, 0x1b115d00u, 0x566e466fu, 0x40e4c461u, 0x39b8ff78u, 0x580c8cafu, 0xbdc4761bu, 0x18e3d94cu, 0x7d742821u, 0xaeb2e213u, 0xbf49ac05u, 0xa607dc56u, 0x225cafb1u, 0xb4f81cf3u, 0x68afbfa5u, 0xe6d35b11u, 0x86050c01u, 0x799c3487u, 0xa90362eau, 0x051a2bbeu, 0xced2c729u, 0xe9429d0bu, 0x5bcdc16cu, 0x0af50eb5u, 0x1493ced7u, 0xcd80ab7eu, 0x4b67919au, 0x633f820au, 0x5fd0275eu, 0x6d20c837u, 0x200a638eu
};
// Cache self-test (memory-hard mode): cache[0..15], the last 16 words, and FNV-1a 64 over all 2^26 words.
static const uint32_t IGNEUM_CACHE_HEAD[16] = {
0x355a86d2u, 0x7957db1cu, 0xd21772afu, 0x6fc1e09bu, 0xd55ce61du, 0x6e6a278bu, 0xd3f543ceu, 0x223d8e82u,
0x143ab337u, 0x2e9f05bdu, 0x2eb389bfu, 0x0c6e449eu, 0x5cfa4222u, 0xba6560feu, 0x8e3e1aa4u, 0xdbcc1d53u
};
static const uint32_t IGNEUM_CACHE_LAST[16] = {
0x41190d91u, 0xbd277957u, 0x22ddbb49u, 0x6986f207u, 0xdf69a4d6u, 0x26401a3au, 0x818230fbu, 0xc417122du,
0x3597b211u, 0xb553ce55u, 0xcf39cc0du, 0x3b7fc43au, 0x3fd43b00u, 0x67e1c80eu, 0xffa7ea7du, 0xca2960abu
};
static const uint64_t IGNEUM_CACHE_FNV64 = 0x48c4f5bf24166b2eull;

View file

@ -0,0 +1,36 @@
{
"seed": "igneum-genesis",
"day": "2026-10-03",
"dataset_mode": "memory-hard",
"dataset_log2_words": 28,
"mask": "0x0fffffff",
"lanes": 32,
"source": "igneum-pow (Rust) CPU interpreter, generator v2, memory-hard dataset",
"warps": [
{"base_nonce": 0, "expected": [
"0xe23d389f3eea0c83", "0xdb20ad461b5a2b40", "0x19f7ba3b1f71682c", "0x85f636966fc9aca0", "0x6487ed96bdc27c42", "0xc490c3313b8029a3", "0xe361eba9844902d4", "0xf08e800494c84491",
"0xa8bad93be02c1cac", "0xe4e1f2c75c891797", "0x63b050548aa5d9b3", "0x7b8851cd9a064497", "0x9873c6fdd34355ab", "0xbe51e127ed02c404", "0x1aa1ee5f4c909067", "0x422684f3c86c5058",
"0x42c3e7bc0f3578a3", "0x3e34aef37a263942", "0xcf0a9f5f13bd905c", "0x44ba872dba90ac37", "0xbd590fe07552ad9b", "0x62b32d43d066161e", "0x213423a10a13d5d1", "0x6bca949b59253f84",
"0xfdb3dba0dcb2951b", "0x1f871d6c8a19d696", "0x413da5a7e4e5f500", "0xa448622cdd32bb6d", "0x5193eba1f8804e2a", "0x567869c0cfdb19ac", "0x2b2b619c048918dc", "0x6605db059b381bd9"
]},
{"base_nonce": 4096, "expected": [
"0xfdb4b214da8ce292", "0xa04a87297fc6a9c0", "0x968cecee7a3bcd99", "0xa41939b15019ad50", "0xffc694d1b04e4da2", "0xc262e3973d436fda", "0x665a2ce13954b603", "0x6ece51e9d9d5d921",
"0x5ff85c609da62800", "0xde9a8e850dd6c1e9", "0x39e876e19ae70187", "0xdb55c10b53c52bf2", "0x610845800cd92811", "0x696316e0c89c9d52", "0x3f18086d783e7d7a", "0xb0d29c8dfb421fd5",
"0x580543d5977c1138", "0xaba172c0bc028b29", "0x43ffd7726859240d", "0xf04dcb8e1b09dcaf", "0xa1ebc5eb5873c4fb", "0x4141b328723483e0", "0xba9fad0940aa4905", "0xed0671189b176bd0",
"0x7ed7083b52f320fd", "0x30c86e3033463bf7", "0x8169974907e0f4a3", "0x89fd7aad0e41bf03", "0x1aab38280fa3ee63", "0x6f525aeaa4c8a186", "0x5bd2a274c3905d8b", "0xdb698d03437d74f7"
]},
{"base_nonce": 1000000, "expected": [
"0x534671b1bf5cea36", "0xe3f81c66bf210aa3", "0xb79ffc3d3cca6731", "0x5df44d8e9e3ac008", "0x445792668a534b0d", "0x293f8c07ed96f576", "0x5b2e23716649e117", "0x55aa3dcc43378f3f",
"0xfec021749867a3a8", "0x3f06d1c7cabc290f", "0x875d22f8b0811415", "0x0cd09e82747f0eef", "0x2bcf3b7cd81a64c8", "0x23c180008f4a8677", "0x8bae18c724c5810a", "0xb6a3fe36acf8b5e3",
"0xc145ffcaa45352ba", "0xf1a0f8edaaf82d2d", "0x205ba1313b3852b1", "0x6922192747873b0b", "0x9937304a65a44a7d", "0xca589d2fe74a7e01", "0xbecb0ddd75ba8ddd", "0x752e223db126522a",
"0x92f919e76899e682", "0x6c9bf07fddaae5ef", "0xfa7a39ed604236c0", "0x1478fb9d0050bc69", "0x9e078959178a2797", "0x7e6311fd8ff88b13", "0x911c33694145c818", "0x733b123353f13d21"
]}
],
"dataset_head": ["0xe73af5ca", "0x45d77eff", "0xb640f499", "0x351c2ca1", "0xfed2d16f", "0x1199bd1d", "0x1db2bbec", "0x4b1769de", "0x45b535ea", "0x414c389c", "0xb8cada9c", "0x32a3eb0f", "0x405305ca", "0xddbd43e4", "0xbdd280ac", "0x70358961"],
"dataset_last_index": 268435455,
"dataset_last": "0x7cdbf6b5",
"dataset_samples": [{"index": 59471966, "value": "0x0c799a6d"}, {"index": 217795994, "value": "0xa21ecab8"}, {"index": 208353206, "value": "0xcc71d9a3"}, {"index": 42483309, "value": "0xb0d7af9c"}, {"index": 172547758, "value": "0x1925c929"}, {"index": 148076330, "value": "0x457795ff"}, {"index": 183853158, "value": "0x4718201a"}, {"index": 214389424, "value": "0x985cde9e"}, {"index": 267488061, "value": "0x4bb9b17e"}, {"index": 169781097, "value": "0x677407dd"}, {"index": 184093494, "value": "0x6093de87"}, {"index": 153880993, "value": "0x2a754135"}, {"index": 84977930, "value": "0xbc58f70d"}, {"index": 46426879, "value": "0xd805eb4b"}, {"index": 3093825, "value": "0xf453a94d"}, {"index": 225364072, "value": "0x5a28b420"}, {"index": 44593546, "value": "0xe92fb02b"}, {"index": 260713159, "value": "0x8a48d35a"}, {"index": 168250303, "value": "0x2024c448"}, {"index": 52384140, "value": "0x48ac5a95"}, {"index": 223401610, "value": "0x8fa7880c"}, {"index": 45554030, "value": "0xfff38e1b"}, {"index": 95410555, "value": "0x4e98033a"}, {"index": 175039924, "value": "0xb6a33ee4"}, {"index": 79171087, "value": "0xdfb0f94f"}, {"index": 267580473, "value": "0x80d99b65"}, {"index": 24168642, "value": "0x7d1c8bde"}, {"index": 37981670, "value": "0x06f4c2c4"}, {"index": 171551130, "value": "0x05e3527c"}, {"index": 195559979, "value": "0x8690f1be"}, {"index": 204611762, "value": "0xa392261a"}, {"index": 140997658, "value": "0xac2293ff"}, {"index": 138925853, "value": "0x0d067d65"}, {"index": 86637313, "value": "0x7133a61e"}, {"index": 20736778, "value": "0x1b115d00"}, {"index": 219665210, "value": "0x566e466f"}, {"index": 160430336, "value": "0x40e4c461"}, {"index": 264654675, "value": "0x39b8ff78"}, {"index": 8013395, "value": "0x580c8caf"}, {"index": 228945585, "value": "0xbdc4761b"}, {"index": 213884386, "value": "0x18e3d94c"}, {"index": 104419827, "value": "0x7d742821"}, {"index": 44185464, "value": "0xaeb2e213"}, {"index": 142737231, "value": "0xbf49ac05"}, {"index": 99284897, "value": "0xa607dc56"}, {"index": 132475900, "value": "0x225cafb1"}, {"index": 61861762, "value": "0xb4f81cf3"}, {"index": 132056166, "value": "0x68afbfa5"}, {"index": 262388043, "value": "0xe6d35b11"}, {"index": 91878046, "value": "0x86050c01"}, {"index": 117353561, "value": "0x799c3487"}, {"index": 124768597, "value": "0xa90362ea"}, {"index": 71352993, "value": "0x051a2bbe"}, {"index": 190698941, "value": "0xced2c729"}, {"index": 46055428, "value": "0xe9429d0b"}, {"index": 55281366, "value": "0x5bcdc16c"}, {"index": 165145231, "value": "0x0af50eb5"}, {"index": 106810753, "value": "0x1493ced7"}, {"index": 171985651, "value": "0xcd80ab7e"}, {"index": 232085256, "value": "0x4b67919a"}, {"index": 159510492, "value": "0x633f820a"}, {"index": 40072060, "value": "0x5fd0275e"}, {"index": 209107596, "value": "0x6d20c837"}, {"index": 39023794, "value": "0x200a638e"}],
"cache_head": ["0x355a86d2", "0x7957db1c", "0xd21772af", "0x6fc1e09b", "0xd55ce61d", "0x6e6a278b", "0xd3f543ce", "0x223d8e82", "0x143ab337", "0x2e9f05bd", "0x2eb389bf", "0x0c6e449e", "0x5cfa4222", "0xba6560fe", "0x8e3e1aa4", "0xdbcc1d53"],
"cache_last_line": ["0x41190d91", "0xbd277957", "0x22ddbb49", "0x6986f207", "0xdf69a4d6", "0x26401a3a", "0x818230fb", "0xc417122d", "0x3597b211", "0xb553ce55", "0xcf39cc0d", "0x3b7fc43a", "0x3fd43b00", "0x67e1c80e", "0xffa7ea7d", "0xca2960ab"],
"cache_fnv1a64": "0x48c4f5bf24166b2e"
}

View file

@ -0,0 +1,118 @@
# Igneum run job: Counter ASIC 3.0 item 2, the per-day item-derivation program (docs/plans/counter-asic-3-derivation.md),
# on PC 2's RTX 5090 (machine 1ccfe586), 6 October 2026. ONE job carries everything (the PC 2 rule of the ca3 brief):
# it downloads the packs zip itself from the downloads host (sha256 checked), finds the installed app's
# igneum-worker-cuda.exe (NVRTC compiles each pack's own kernel text, so no new worker build is needed), switches the
# NVIDIA card off in the app ONLY while the packs run (the card key from the app's settings.json, never /api/state;
# restored after with the settings it had), and never quits, restarts or updates the installed app. The numbers this
# job is for: per pack, the worker's own `nvrtc .. cache .. dataset .. ms` line (the daily 1 GiB build on the 5090),
# the self-test (bit-exactness: cache FNV, dataset head, word [MASK], 64 samples, 96 vector lanes against the Rust CPU
# interpreter), the 2^24 fingerprint at base nonce 0 (against the Mac: dr736-genesis 50e3eaa779da4f1e,
# dr736-devnet-epoch0 9553f6d5c667205a, mx8-genesis 7c28cfb06c5c65a9, v2-genesis-mh 25f96e7dce90bd4e) and the hash rate.
# Packs: dr736-genesis, dr736-devnet-epoch0 (the derivation class), mx8-genesis (the x8 control), v2-genesis-mh (the v2
# control). Every result line starts with RESULT. The placeholder __DL_BASE__ is substituted at publish time; the
# downloads token never enters the repository.
$ErrorActionPreference = 'Continue'
function Say([string] $m) { Write-Host ("[" + (Get-Date -Format 'HH:mm:ss') + "] " + $m) }
function Stamp { (Get-Date).ToUniversalTime().ToString('yyyy-MM-ddTHH:mm:ssZ') }
# everything lives in this job's own folder (the wiped-jobs-folder rule of tools/ci/kit-path-check.sh: nothing of
# another job's is reached)
$work = Join-Path $env:IGNEUM_JOB_DIR 'ca3-derive'
New-Item -ItemType Directory -Force -Path $work | Out-Null
$zipUrl = '__DL_BASE__/igneum-ca3-derive-packs.zip'
$zipSha = 'aadce58ca54f136d8c41dc15154b3ae107624808e7f061c266df228b200bc676'
$zip = Join-Path $work 'igneum-ca3-derive-packs.zip'
Write-Output "RESULT job ca3-derive-pc2 start $(Stamp) machine $env:COMPUTERNAME"
# ---- the packs: downloaded by this job, sha256 checked, expanded fresh ----
try {
[Net.ServicePointManager]::SecurityProtocol = [Net.SecurityProtocolType]::Tls12
Invoke-WebRequest -Uri $zipUrl -OutFile $zip -TimeoutSec 120 -Headers @{ 'Cache-Control' = 'no-cache' }
} catch { Write-Output ("RESULT error download: " + $_.Exception.Message); exit 2 }
$got = (Get-FileHash -Algorithm SHA256 $zip).Hash.ToLower()
if ($got -ne $zipSha) { Write-Output "RESULT error zip sha256 $got expected $zipSha"; exit 2 }
Write-Output "RESULT zip sha256 $got size $((Get-Item $zip).Length) OK"
$packs = Join-Path $work 'packs-ca3-derive'
if (Test-Path $packs) { Remove-Item -Recurse -Force $packs }
Expand-Archive -Path $zip -DestinationPath $work -Force
$packList = @('v2-genesis-mh', 'mx8-genesis', 'dr736-genesis', 'dr736-devnet-epoch0')
foreach ($pk in $packList) {
$d = Join-Path $packs $pk
if (-not (Test-Path (Join-Path $d 'memhard.h'))) { Write-Output "RESULT error pack $pk missing after extract"; exit 2 }
Write-Output ("RESULT pack $pk memhard.h sha256 " + (Get-FileHash -Algorithm SHA256 (Join-Path $d 'memhard.h')).Hash.ToLower() + " program.h sha256 " + (Get-FileHash -Algorithm SHA256 (Join-Path $d 'program.h')).Hash.ToLower())
}
# ---- the worker: the installed app's igneum-worker-cuda.exe, run in place (its NVRTC DLLs sit beside it) ----
$inst = @("$env:LOCALAPPDATA\Programs\Igneum Miner", "$env:ProgramFiles\Igneum Miner") | Where-Object { Test-Path (Join-Path $_ 'igneum-worker-cuda.exe') } | Select-Object -First 1
if (-not $inst) { Write-Output 'RESULT error no installed igneum-worker-cuda.exe'; exit 2 }
$cuda = Join-Path $inst 'igneum-worker-cuda.exe'
Write-Output "RESULT worker-cuda $cuda sha256 $((Get-FileHash -Algorithm SHA256 $cuda).Hash.ToLower()) nvrtc_dlls $((Get-ChildItem $inst -Filter 'nvrtc*.dll').Count)"
$help = (& $cuda --help 2>&1 | Out-String)
$hasBench = $help -match '--bench'
Write-Output "RESULT worker-cuda has --bench: $hasBench"
if (-not $hasBench) { Write-Output 'RESULT note the installed worker has no --bench; --check gives the build time and the self-test, no fingerprint and no hash rate' }
# ---- the app: the NVIDIA card's key and settings from settings.json; switched off through POST api/cards only ----
$appDir = $env:IGNEUM_APP_DIR
if (-not $appDir) { $appDir = Join-Path $env:LOCALAPPDATA 'igneum\app' }
$urlFile = Join-Path $appDir 'app.url'
$url = $null
if (Test-Path $urlFile) { $url = (Get-Content -LiteralPath $urlFile -Raw).Trim() }
$sj = Join-Path $appDir 'settings.json'
$cardKey = $null; $cardPref = $null
if (Test-Path $sj) {
try {
$settings = Get-Content -LiteralPath $sj -Raw | ConvertFrom-Json
if ($settings.cards) {
foreach ($p in $settings.cards.PSObject.Properties) { if ($p.Name -like 'nvidia:*') { $cardKey = $p.Name; $cardPref = $p.Value; break } }
}
} catch { Say ("settings.json: " + $_.Exception.Message) }
}
if ($cardKey) {
Write-Output ("RESULT card " + $cardKey + " enabled=" + $cardPref.enabled + " identities=" + $cardPref.identities + " power_pct=" + $cardPref.power_pct + " (settings.json)")
} else {
Write-Output 'RESULT card none in settings.json (no nvidia:* entry); the app keeps mining on the card and the numbers carry that load'
}
$cardOff = $false
if ($cardKey -and $url) {
$ident = 1; if ($cardPref.identities) { $ident = [int]$cardPref.identities }
$pp = 0; if ($cardPref.power_pct) { $pp = [int]$cardPref.power_pct }
$body = @{ cards = @(@{ key = $cardKey; enabled = $false; identities = $ident; power_pct = $pp }) } | ConvertTo-Json -Depth 5
try { Invoke-RestMethod -Uri ($url + 'api/cards') -Method POST -Body $body -ContentType 'application/json' -TimeoutSec 10 | Out-Null; $cardOff = $true; Say 'card off requested' } catch { Write-Output ("RESULT error api/cards off: " + $_.Exception.Message) }
# wait for the worker process to go, up to 90 s (the process list, not /api/state)
$t = 0
while ($t -lt 90) {
Start-Sleep -Seconds 5; $t += 5
$w = Get-Process -Name 'igneum-worker-cuda' -ErrorAction SilentlyContinue
if (-not $w) { break }
}
$w = Get-Process -Name 'igneum-worker-cuda' -ErrorAction SilentlyContinue
Write-Output ("RESULT card-off " + $cardKey + " after " + $t + " s, worker processes left " + (($w | Measure-Object).Count))
Start-Sleep -Seconds 5
}
& nvidia-smi --query-gpu=name,driver_version,power.limit,clocks.sm,clocks.mem,memory.used,temperature.gpu --format=csv,noheader 2>&1 | ForEach-Object { "RESULT gpu-before $_" }
# ---- the runs: --check for the build line and the self-test, --bench for the fingerprint and the rate ----
foreach ($pk in $packList) {
$d = Join-Path $packs $pk
Write-Output "RESULT run $pk check start $(Stamp)"
& $cuda --check --pack $d 2>&1 | ForEach-Object { "RESULT check $pk $_" }
Write-Output "RESULT run $pk check exit $LASTEXITCODE"
if ($hasBench) {
Write-Output "RESULT run $pk bench start $(Stamp)"
& $cuda --bench --pack $d --batches 5 --batch-log2 24 --block-warps 1 2>&1 | ForEach-Object { "RESULT bench $pk $_" }
Write-Output "RESULT run $pk bench exit $LASTEXITCODE"
& $cuda --bench --pack $d --batches 5 --batch-log2 24 --block-warps 8 2>&1 | Where-Object { $_ -match '^RESULT|error|FAIL|dataset' } | ForEach-Object { "RESULT bench8 $pk $_" }
}
}
& nvidia-smi --query-gpu=power.draw,clocks.sm,clocks.mem,memory.used,temperature.gpu --format=csv,noheader 2>&1 | ForEach-Object { "RESULT gpu-after $_" }
# ---- restore the card with the settings it had ----
if ($cardOff) {
$ident = 1; if ($cardPref.identities) { $ident = [int]$cardPref.identities }
$pp = 0; if ($cardPref.power_pct) { $pp = [int]$cardPref.power_pct }
$en = $true; if ($null -ne $cardPref.enabled) { $en = [bool]$cardPref.enabled }
$body = @{ cards = @(@{ key = $cardKey; enabled = $en; identities = $ident; power_pct = $pp }) } | ConvertTo-Json -Depth 5
try { Invoke-RestMethod -Uri ($url + 'api/cards') -Method POST -Body $body -ContentType 'application/json' -TimeoutSec 10 | Out-Null; Write-Output ("RESULT card restored " + $cardKey + " enabled=" + $en) } catch { Write-Output ("RESULT error card restore " + $cardKey + ": " + $_.Exception.Message) }
}
Write-Output "RESULT job ca3-derive-pc2 end $(Stamp)"
exit 0