61 KiB
Igneum protocol specification, section 1: the lottery hash
Spec version 0.2, 4 October 2026 (0.1 on 3 October 2026). Status of this section: Measured for the construction as implemented (Implemented values with test vectors on three GPU vendors and a CPU reference); Designed for the header binding, the day key, the dataset growth and the era schedule; Open where marked. Version 0.2 adopts generator version 2 (sections 1.4.2, 1.4.3 and 1.4.6: exact load count, fresh-source loads, program acceptance) from the weak-program census of docs/analysis/weak-program-census-2026-10-03.md; every vector of version 0.1 is retired and re-cut (section 1.17).
Normative implementation: igneum-pow/src/{seed,generator,accept,memhard,verify}.rs. Where this text and that code disagree, the code and its test vectors win until this text is corrected (section 0.4). The Swift prototype proto-metal/main.swift carries the same generator and acceptance rule and is bit-exact with the crate on every pack (docs/bench-log.md, entries "igneum-pow: Rust crate bit-exact with proto-metal" and "generator version 2").
Every parameter marked "prototype value, to be fixed at gate 1" is carried by the implementation today, is part of the test vectors, and is confirmed or replaced by the named measurement before the hash is frozen (section 1.16).
1.1 What the hash is for
The lottery hash decides who produces the next block. It MUST be:
- Deterministic and bit-exact on every conforming implementation: GPU kernels on any vendor, the CPU verifier, and any future implementation (section 1.14).
- Cheap to verify on one CPU core without the dataset: one 32-lane unit of work in under 10 ms (Target, Measured at 0.63 ms steady and 0.67 to 0.81 ms cold under generator version 2, section 1.11).
- Bound by random access to a dataset larger than any on-chip cache, so that computing dataset words is slower than loading them (Measured 4.8x slower on Apple, section 1.8.4; not measured on NVIDIA or AMD).
- Unknowable until shortly before it is needed, so that a miner cannot grind the seed (section 4).
It does not need to be a general-purpose cryptographic hash (preimage, collision). It needs to be a fair lottery: no shortcut cheaper than honest evaluation and no bias a miner can exploit. No analysis of either property exists yet (docs/fud-ledger.md, entries M6 and M7; section 6).
1.2 Notation
All arithmetic is on unsigned 32-bit integers modulo 2^32 unless a 64-bit type is named. rotl(x, n) and rotr(x, n) rotate by n in 0..31. mulhi(a, b) is bits 32..63 of the 64-bit product. low32(v) is bits 0..31 of a 64-bit v. || is byte concatenation. Indices are zero based. Words are hashed into FNV as little-endian bytes (section 0.6).
1.3 Seed words and the SplitMix64 stream
Implemented, prototype value, to be fixed at gate 1 (ledger M7 asks for a standard hash in place of FNV-1a plus SplitMix so the seed-to-program mapping is auditable; the measurement that fixes it is the weak-program census of section 1.16, which must be re-run on whichever derivation is chosen).
1.3.1 seed_words_from_bytes(b) -> [u32; 8]
For salt in 0..3:
basis = 0xcbf29ce484222325 XOR (salt * 0x9E3779B97F4A7C15) (64-bit, wrapping)
h = basis
for each byte x of b: h = (h XOR x) * 0x100000001b3 (64-bit, wrapping)
h = h XOR (h >> 33); h = h * 0xff51afd7ed558ccd; h = h XOR (h >> 33)
words[2 * salt] = low32(h)
words[2 * salt + 1] = h >> 32
seed_words(s) is seed_words_from_bytes over the UTF-8 bytes of the string s. This function is the boundary at which the chain's bytes enter the hash: the epoch program seed from section 4 and the day key bytes (section 1.12) both go through it. In the prototype the input is a string ("igneum-genesis", "day/2026-10-03").
Test vector: seed_words("igneum-genesis") = 67a9a7be 1a155b25 fddfb732 4b5af2e8 c55caf33 a27c13b7 06628a48 03852469 (proto-cuda/packs/igneum-genesis-mh/program.json, seed_words).
1.3.2 SplitMix64
State s (64-bit). next():
s = s + 0x9E3779B97F4A7C15
z = s
z = (z XOR (z >> 30)) * 0xBF58476D1CE4E5B9
z = (z XOR (z >> 27)) * 0x94D049BB133111EB
return z XOR (z >> 31)
below(n) = next() mod n (modulo, not rejection sampling; the bias at n <= 100 is under 2^-57 and is part of the definition).
1.3.3 The program stream
For seed words w[0..7]: lo = w[0] | (w[1] << 32), hi = w[2] | (w[3] << 32), and the generator's SplitMix64 state starts at lo XOR (hi * 0x9E3779B97F4A7C15) (64-bit wrapping multiply). Words w[4..7] are not used by the generator; all eight are used by the register initialisation (section 1.6).
1.4 The generator
Implemented (igneum-pow/src/generator.rs). A program is a list of INSTR_COUNT instructions over 8 lane registers r0..r7, executed ITERATIONS times per hash.
| Parameter | Value | Label |
|---|---|---|
| Registers per lane | 8 x u32 | prototype value, to be fixed at gate 1 (fixed by the ASIC-gain target and the register-pressure measurement of 1.16) |
| Instructions per program | 64, of which exactly 16 are load (section 1.4.2) |
prototype value, to be fixed at gate 1 (fixed by the CPU-verify measurement on a 2019-class core); the load count is Definition since 4 October 2026 |
| Iterations per hash | 8 | prototype value, to be fixed at gate 1 (same measurement) |
| Lanes per unit of work | 32 | Definition. Not a tuning parameter (section 1.9) |
| Shuffle masks | {1, 2, 4, 8, 16} | Definition, follows from 32 lanes |
| Rotate immediates | 1..31 | Definition. 0 is never emitted and MUST NOT be relied on (proto-metal/TESTS.md section 2) |
1.4.1 Instruction set
Eleven families. dst, src, src2 name registers; src != dst always; src2 may equal either.
| Op | Semantics (per lane) | Operands used |
|---|---|---|
add |
dst = dst + src + (bit bit of sel ? imm2 : imm) |
dst, src, imm, imm2, bit |
sub |
dst = dst - src |
dst, src |
mul |
dst = low32(dst * src) |
dst, src |
mulhi |
dst = mulhi(dst, src) |
dst, src |
xor |
dst = dst XOR src |
dst, src |
or |
dst = dst OR src |
dst, src |
rotl |
dst = rotl(dst, rot), rot in 1..31 |
dst, rot |
rotr |
dst = rotr(dst, src AND 31) |
dst, src |
mad |
dst = low32(src * src2) + dst |
dst, src, src2 |
shfl |
dst = dst XOR src_of_lane(lane XOR mask) |
dst, src, mask |
load |
dst = dst XOR dataset[src AND MASK] |
dst, src |
sel is the value of r0 sampled once at the top of each iteration, before instruction 0, and held for all 64 instructions of that iteration. This is the per-hash nonce-dependent select: the immediates an add uses depend on the lane's own state, so no two nonces run the same constant sequence (ProgPoW-style data-dependent path, integer only).
A twelfth family wload (warp-coalesced 128-byte load) exists in the code as lever (b) and is never emitted at the default configuration (wide_frac = 0). It is NOT part of the lottery hash. It is retained only so the measurement in proto-metal/MEMHARD.md section 2.4 stays reproducible, and the recommendation there is not to adopt it.
1.4.2 Op weights and the load count
Implemented (generator version 2, 4 October 2026, igneum-pow/src/generator.rs; ledger M5 Fixed). Every program contains exactly 16 load instructions (LOAD_SLOTS, Definition), so every hash performs 128 loads and a 32-lane unit derives at most 4,096 dataset items, the bound of section 1.11. The other 48 instructions are drawn from the ten non-load families with these weights (sum 75, prototype value, to be fixed at gate 1), applied in this order:
| Op | Weight |
|---|---|
| add | 12 |
| xor | 10 |
| mul | 8 |
| mad | 8 |
| shfl | 8 |
| rotl | 7 |
| sub | 6 |
| mulhi | 6 |
| rotr | 6 |
| or | 4 |
The load weight 25 of version 1 is retired; it survives only as the ratio 16 of 64. Why the count is fixed and not merely expected: under version 1 the static count ran 24 to 256 loads per hash over 100,000 programs and the GPU rate tracked the number of distinct addresses, 24 to 200, because 19.9 percent of all loads re-read an address the same hash had already read (census sections 3 and 5). Fixing the static count alone would not fix the memory work; the fresh-source rule of 1.4.3 fixes the distinct count by construction, and 1.4.6 rejects the few programs where it cannot.
The version 1 lever measurement (load_weight = 17, proto-metal/MEMHARD.md section 2.4) is kept in the code as generate_v1 for reproduction only; its programs are not the lottery hash.
1.4.3 Draw order
From the program stream of 1.3.3, in this order, whether or not an op uses a value.
(1) Load slots. Let p[0..62] = 1..63 (instruction 0 is never a load: nothing is fresh before it). For i in 0..15 draw j = i + below(63 - i) and swap p[i] and p[j]. The load slots are p[0..15], a uniform 16-subset of 1..63.
(2) For each instruction k in 0..63, nine draws:
roll = below(75); op = first entry of the table of 1.4.2 whose cumulative weight exceeds roll
(on a load slot the roll is drawn and ignored and op = load)
dst = below(8)
src: on an ALU slot a = below(7); src = a + (a >= dst)
on a load slot E = the registers other than dst, in register order, that an earlier instruction of this
program has written and that no later load has used as its source (E is empty before
instruction 0); if E is not empty, a = below(|E|) and src = E[a];
if E is empty, a = below(7) and src = a + (a >= dst), and 1.4.6 (a) rejects the program
b = below(8) (src2)
imm = low32(next())
imm2 = low32(next())
rot = 1 + below(31)
bit = below(32)
mask = 1 << below(5)
16 slot draws plus 64 x 9: 592 draws per program. A program is fully determined by its eight seed words. A load's source holds a value written in the same iteration that no earlier load has read, so no load repeats the address of an earlier load of the same hash, across the iteration boundary included (the first form of the rule, with every register eligible at instruction 0, left the wrap open and failed 1.4.6 (a) on 36 percent of programs; census section 7.1).
Test vector: for seed igneum-genesis (attempt 0, program id bcc1248b10cc90f2, section 1.4.6) the first eight instructions are (proto-cuda/packs/igneum-genesis-mh/program.json):
0: mad dst=2 src=3 src2=4 imm=0xbf7b174d imm2=0x337b762e rot=17 bit=2 mask=2
1: mad dst=2 src=1 src2=1 imm=0xdd04a5da imm2=0x42da7657 rot=15 bit=30 mask=16
2: mad dst=2 src=3 src2=2 imm=0x734003fa imm2=0x5bb67700 rot=3 bit=20 mask=1
3: xor dst=3 src=5 src2=5 imm=0xc55a1b1c imm2=0xa19720f3 rot=7 bit=8 mask=1
4: load dst=7 src=2 src2=5 imm=0xad572dd7 imm2=0x9ceb3ea7 rot=18 bit=30 mask=2
5: load dst=5 src=7 src2=3 imm=0x769a53be imm2=0x80f9067e rot=12 bit=22 mask=1
6: shfl dst=1 src=4 src2=0 imm=0xd3613d88 imm2=0x262fb219 rot=10 bit=30 mask=8
7: shfl dst=7 src=3 src2=4 imm=0xce38e42f imm2=0xb868b818 rot=11 bit=8 mask=8
The whole program has op mix load=16 add=8 shfl=8 xor=6 mad=5 mul=5 mulhi=5 sub=4 rotl=3 rotr=3 or=1, 128 loads per hash. Instruction 4 reads r2, written by instructions 0 to 2; instruction 5 reads r7, written by instruction 4.
1.4.4 Generator contract
Every emitted instruction satisfies: rot in 1..31, mask in {1, 2, 4, 8, 16}, src != dst; every program has exactly 16 load instructions and none at instruction 0. This held on every instruction of 10,200 fuzzed version 1 programs (TESTS.md section 1) and of 2,000 fuzzed version 2 programs (TESTS.md section 9, 4 October 2026). A kernel emitter MAY rely on it; an interpreter MUST NOT accept a program that violates it.
1.4.5 Encoding
A program is transmitted as the seed bytes, never as instructions. A node hands a miner the pack it emits itself (igneum-pow/src/emit.rs: kernel.cu, kernel.cl, program.metal, program.h, program.json, memhard.h, vectors.*, the three *_bound kernels), and a miner MAY regenerate everything from the seed bytes by the procedure of 1.4.6. program.json (format igneum-program-pack-3) is the interchange form; its field names are those of Instr in generator.rs, and it carries generator (2), attempt, program_id and seed_bytes. program.h carries the same as IGNEUM_GENERATOR, IGNEUM_PROGRAM_ATTEMPT, IGNEUM_PROGRAM_ID and IGNEUM_SEED_BYTES_HEX. An implementation MUST refuse a pack whose generator version is not its own. Program class v3 (Counter ASIC 2.0, 5 October 2026, activated by the height switch program_class_v3_activation_daa from the first epoch whose start score is at or above it, docs/plans/counter-asic-2-rollout.md) writes generator 3, and every pack of it carries IGNEUM_PROGRAM_CLASS (v3) and IGNEUM_ERA_SEED_HEX (the 32-byte era seed of section 1.13.1, or its devnet stand-in) beside IGNEUM_GENERATOR; the serve protocol's prepare and job lines carry class=v3 era=<hex> for v3 epochs and nothing for v2 ones. A worker MUST refuse a pack whose class or era seed does not match the line it was prepared for (igneum-pow/src/packcheck.rs, verify_pack_dir_chain; proto-cuda/nvrtc/packfile.h), and a pack of a generator other than 2 or 3.
1.4.6 Program acceptance
Implemented (igneum-pow/src/accept.rs, proto-metal/main.swift; ledger M6 Fixed). A candidate program is accepted only if all of the following hold, and every conforming implementation MUST evaluate them identically.
(a) For every load, some instruction between the previous load from the same source register and this one, in cyclic order over the 64 instructions, writes that register.
(b) Every register r0..r7 is the destination of at least one add, sub, xor, mad, shfl or load.
(c) The program is interpreted (section 1.7) for 64 units at base nonces low32(next()) AND NOT 31 from a SplitMix64 stream seeded with FNV-1a-64("igneum-accept/" || seed words as little-endian bytes), with init words I equal to the seed words and dataset words dataset_elem(idx, S[0], S[1]) of verify.rs (the six-operation closed form of the version 0.1 packs) at 2^28 words (idx = src AND 0x0fffffff, a constant of this rule whatever the live dataset size) in place of the memory-hard dataset. Over those 2,048 evaluations: no register has a bit equal in every final value; no load site (iteration, instruction) reads one address in all 32 lanes of any unit; the number of final register values equal to 0 or 2^32 - 1 is below 164 (1 percent of 16,384); every output bit's ones count is within 136 of 1,024 (6 sigma); and the number of distinct masked dataset addresses read by one lane in one evaluation, summed over the 2,048 evaluations, exceeds 245,760 (a mean above 120 of the 128 loads).
Attempts. Attempt 0 of a program seed b (the 32-byte epoch seed, or the UTF-8 of a seed string) is the candidate drawn from seed_words_from_bytes(b). If it fails, attempt k = 1, 2, ... is drawn from seed_words_from_bytes(b || k_le32); the first accepted candidate is the program of the epoch. Measured rejection rate under this generator: 5.14 percent over 100,000 seeds (census section 7) and the 20,000-seed confirmation of docs/bench-log.md (4 October 2026), so the probability that 32 consecutive candidates fail is below 2^-136, and an implementation MAY treat 32 consecutive failures as a consensus fault (MAX_ATTEMPTS).
Program id. FNV-1a-64("igneum-program/" || generator_le32 || seed words as little-endian bytes || attempt_le32) with generator = 2 under class v2 and generator = 3 under class v3, written into every pack. Two implementations that agree on the id agree on the generator version, the seed words and the attempt.
Why the closed form: the test is then a pure function of the program (no cache, no day), costs 1.3 to 3.4 ms on one core, and the census checked on 100,000 programs that its verdict agrees with the memory-hard dataset's on all but 39 threshold-edge cases (section 7.3). What the three parts catch: (a) the empty-list fallback of 1.4.3; (b) registers that saturate to all ones (2.4 percent of candidates); (c) zero-absorbing register sets, lane-constant load sites, output bias and value-level address repeats (2.1 percent). Not in the rule, and why: a contraction as the last write (80 percent of programs) and the or count are too common and (c) already catches the cases that matter; the load critical path is a hash-rate question, not a weakness.
Test vectors for the rule (igneum-pow accept --seed ...):
| Seed | Attempt 0 | Attempt 1 |
|---|---|---|
igneum-genesis |
accepted, program id bcc1248b10cc90f2, 128.000 distinct addresses per hash, 0 saturated, bias max 54 |
|
igneum-hourly |
accepted, a4c4d00961c855df |
|
igneum-census-2026-10-03/22 |
rejected, (b) r7 has no injecting write | accepted, 22ed0609d079f4cf |
igneum-census-2026-10-03/37 |
rejected, (c) 245,230 distinct addresses (mean 119.74) | accepted, 947705cc4eb1df0a |
igneum-census-2026-10-03/51 |
rejected, (b) r4 has no injecting write | accepted, 9869afcc028bf9f1 |
The Rust crate and the Swift prototype derive identical instruction lists and identical 96-vector sets on all five seeds (docs/bench-log.md, 4 October 2026).
1.5 Fixed memory footprint
Implemented for the prototype size; Designed for genesis.
| Quantity | Prototype (packs, vectors) | Genesis | Label |
|---|---|---|---|
| Dataset words | 2^28 (1 GiB) | 2^29 (2 GiB) | Prototype: Implemented. Genesis: Designed (design document, "Which cards mine?") |
| MASK | 0x0fffffff | 0x1fffffff | as above |
| Dataset item | 16 words (64 bytes) | same | Implemented |
| Cache | 2^26 words (256 MiB) | same | Implemented, prototype value, to be fixed at gate 1 (fixed by the shortcut-ratio measurement on the RTX 5090 and an AMD discrete card: the cache must exceed the largest on-chip cache of any card that mines, and 96 MiB of L2 on the 5090 is the figure to beat) |
A load reads one 4-byte word at src AND MASK. Every load in every emitted kernel has exactly this form; the static check in TESTS.md section 5 is part of conformance (section 1.15). Because item values do not depend on the dataset size (section 1.8.5), the 1 GiB vectors remain valid for words below 2^28 at any larger size.
Growth beyond genesis is in section 1.13; under program class v3 the cache follows the dataset's doublings (1.13.3) and the verifier holds 256 MiB, then 512 MiB from year 4 and 1 GiB from year 12. Beside the schedule the latency ladder (docs/design/latency-ladder.md) carries one cache rung (genesis forward-compatibility, 7 October 2026, docs/design/genesis-forward.md section 3): LatencyLadder::cache_rung, 512 MiB, entered by the same rule as a shadow rung (90 percent of blue blocks with both ladder bits set in each of seven consecutive windows, one decision per seven windows, one rung, never back), because a consumer last-level cache at the cache size gives that card's owners a 2 to 3x shortcut (chip-model-v3; consumer LLC 96 to 128 MB today, datacentre 256 MB). Its gate is the rung's admissible flag in the genesis list, false until the cold verify with the larger cache on the reference core with its sibling loaded is measured under 10 ms, the day-cache build on a 2019-class core under twice today's, and the 8 GB tier still holds dataset, cache and the prover footprint; the engine's consumption of the larger cache is owed with that measurement, and the rule never enters the rung before the flag is set.
1.6 Register initialisation
Implemented (verify.rs, interpret_warp). For lane nonce n (32-bit) and init words I[0..7]:
splitmix32(x): x ^= x >> 16; x *= 0x7feb352d; x ^= x >> 15; x *= 0x846ca68b; x ^= x >> 16
for i in 0..7:
x = n XOR I[i]
x = x + 0x9e3779b9 * (i + 1)
x = splitmix32(x)
r[i] = x XOR I[(i + 1) AND 7]
In the prototype and in every pack, I is the program's own seed words (the same eight words that drive the generator). On the chain the hash MUST also commit to the block being mined, which the prototype does not do. The binding is Designed, proposed here, Open until gate 1 (item O-1.9 in section 6):
- The header nonce is 64 bits (rusty-kaspa
Header.nonce). Its low 32 bits are the lane noncen. Its high 32 bits and the 256-bit pre-PoW header hashH(section 2, fork point a5) form the init words:I = seed_words_from_bytes("igneum-block/" || H || nonce_hi_le32). Iis a kernel argument, not a compile-time constant. The program (from the epoch seed) is compiled once per epoch;Ichanges per block template.- The packs' vectors, where
Iequals the program seed, remain the conformance vectors for the generator, the interpreter and the dataset. A second vector set withIderived from a header is produced when section 2 fixes the header hash.
1.7 Execution of one hash
r = init(n, I)
repeat ITERATIONS (8) times:
sel = r0
for each instruction in order: apply it (section 1.4.1)
lo = r0 XOR rotl(r1, 7) XOR rotl(r2, 14) XOR rotl(r3, 21)
hi = r4 XOR rotl(r5, 9) XOR rotl(r6, 18) XOR rotl(r7, 27)
hash = (hi << 32) | lo (64 bits)
The output folding rotations (7, 14, 21; 9, 18, 27) are Implemented, prototype value, to be fixed at gate 1 (the stats run of TESTS.md section 3 is the check; a different fold must pass the same run).
shfl makes the 32 lanes of a unit interdependent: the hash of one nonce is defined only as a member of its aligned group of 32 (section 1.9).
1.8 The memory-hard dataset
Implemented (memhard.rs), construction and measurements in proto-metal/MEMHARD.md. Every constant below is a prototype value, to be fixed at gate 1, unless marked Definition. What fixes them is the shortcut-ratio and time-memory trade-off measurement on NVIDIA and AMD discrete cards (section 1.16 items 2 and 3) and an external review of the primitives (ledger M7).
1.8.1 Day key
K[0..7] = seed_words_from_bytes(day_bytes). In the prototype day_bytes is the UTF-8 of "day/" + day with day an ISO date; on the chain see section 1.12. For day/2026-10-03: K = 3067619f 3c269176 84a03b03 f8c63294 ff977c5b e60def3e 63630141 b8fbcb58.
1.8.2 Block function B (ChaCha12 with feed-forward)
Standard ChaCha quarter round with rotations (16, 12, 8, 7), six double rounds (columns then diagonals, on the 16-word state), then y[i] = y[i] + x[i]. Twelve rounds, no key schedule beyond the input block. CHACHA_ROUNDS = 12 is a prototype value; sigma = (0x61707865, 0x3320646e, 0x79622d32, 0x6b206574) is a Definition.
1.8.3 Cache fill
The cache is 2^26 words = 2^22 lines of 16 words, in 2^16 segments of 64 lines. Segment s, line j is at word offset (s * 64 + j) * 16. Each segment is a chain:
tag = (0x49676e65, 0x756d4d48) "Igne", "umMH"
prev = 0^16
for j in 0..63:
in = prev XOR (sigma[0..3] || K[0..7] || s || j || tag[0..1]) 16 words
line = B(in)
cache[segment s, line j] = line
prev = line
Line j costs j + 1 block evaluations from nothing, 32.5 on average. The 65,536 segments are independent (one GPU thread each). The fill is a once-per-day cost.
| Fill time | Value | Source |
|---|---|---|
| One M5 Max core, Rust | 175 to 181 ms | Measured, docs/bench-log.md, igneum-pow entry |
| One M5 Max core, Swift | 184.5 to 190.6 ms | Measured, same log, memory-hard dataset entry |
| RTX 5090, one host thread | 223 ms | Measured, same log, "RTX 5090, memory-hard dataset" |
| M5 Max GPU | 0.6 to 2.1 ms | Measured, memory-hard dataset entry (variance not isolated) |
| RTX 5090 GPU | 0.67 ms | Measured, "RTX 5090, memory-hard dataset" |
1.8.4 Mixer parameters and the mixer M_r
One SplitMix64 stream seeded with K[0] | (K[1] << 32), drawn in this order: ROT[0..7] = 1 + below(31) (eight draws), MUL[0..15] = low32(next()) OR 1 (sixteen draws, always odd so each multiply is a bijection), RC[0..15] = low32(next()) (sixteen draws).
M_r(s) on a 16-word state with round key rk = (r + 1) * 0x9E3779B9:
for i in 0..15: s[i] = (s[i] XOR (RC[i] + rk)) * MUL[i]
QR(s0, s4, s8, s12; ROT0..3) QR(s1, s5, s9, s13; ROT0..3) QR(s2, s6, s10, s14; ROT0..3) QR(s3, s7, s11, s15; ROT0..3)
QR(s0, s5, s10, s15; ROT4..7) QR(s1, s6, s11, s12; ROT4..7) QR(s2, s7, s8, s13; ROT4..7) QR(s3, s4, s9, s14; ROT4..7)
where QR(a, b, c, d; r1, r2, r3, r4) is the ChaCha quarter round with those four rotations. About 130 integer operations. This is the prototype's stand-in for RandomX's SuperscalarHash: fixed shape, seed-drawn constants. It has had no cryptanalysis (MEMHARD.md section 3, unproven item 3; a ROT draw of eight equal values is possible and untested).
Test vector, day 2026-10-03 (proto-cuda/packs/igneum-genesis-mh/program.h):
ROT = 20 20 19 4 26 3 3 27
MUL = 42146205 52cbe0fb 7ecf4a03 6728907f d81d9751 132952c3 f60de277 05358035
baf6499d e4db9667 3e98f45d d0004edd 2691630d 9beb3bcf ab310379 99cfb423
RC = bab68293 cc162340 6ce151cc e62b8997 c9c80297 f74a1654 3d704af5 3cf522b7
2b9cac04 a880ac10 13e5dd1d 6fc3e233 2d83eeac 9006e8bf 2c4b5362 31b49ee2
1.8.5 Item derivation and dataset mapping
Item t (16 words):
s[i] = K[i] for i in 0..7
s[8 + i] = t * MUL[i] + RC[i] for i in 0..7
for r in 0..7:
s = M_r(s)
a = s[0] AND 0x003fffff cache line index, 2^22 lines
s[i] = s[i] XOR cache[line a][i] for i in 0..15
s = M_8(s)
item(t) = s
Eight dependent cache reads (ITEM_ROUNDS = 8, prototype value): the address of read r depends on every earlier read. Nine mixer applications under program class v2.
Program class v3 (Counter ASIC 2.0, decided 5 October 2026, delegated; the project lead confirms for the public testnet genesis) applies the mixer m = 8 times per round with distinct round keys (LoadClass::mixer_mult; docs/plans/mixer-x4.md section 2), the eight dependent reads unchanged:
for r in 0..7:
for j in 0..m-1:
s = M(s, rk = (r * m + j + 1) * 0x9E3779B9)
a = s[0] AND (2^(C - 4) - 1) cache line index, 2^(C - 4) lines of a 2^C-word cache (C from 1.13.3)
s[i] = s[i] XOR cache[line a][i] for i in 0..15
for j in 0..m-1:
s = M(s, rk = (8 * m + j + 1) * 0x9E3779B9)
Under m = 1 the keys are (r + 1) * 0x9E3779B9 and 9 * 0x9E3779B9, the class v2 text exactly; the 9 m keys are the first 9 m values of the sequence k * 0x9E3779B9, all distinct. Why m = 8: the recompute attacker's cost is operations per item (docs/analysis/m16-recompute-attacker-2026-10-05.md); the honest miner pays the mixer once a day in the dataset build, which stays latency-bound (RTX 5090 23 to 25 ms, RX 9070 XT 72 to 77 ms, M5 Max 21 ms at m = 1, 4 and 8, measured 5 October 2026); the verifier pays m per item it derives: 0.61 ms per warp at m = 1, 1.24 at 4, 2.08 at 8 on one loaded M5 Max core (measured 5 October 2026, docs/plans/mixer-x4.md 6.4a), inside the 10 ms gate. The on-die-cache recompute chip's gain against the RTX 5090 falls from 2.4x (m = 1) to 0.92x with a 3x fixed-function factor at m = 8 (docs/analysis/chip-model-v3.md, approximate factor). Under class v3 an item's value also depends on C through the line mask, so the items change on the day the cache doubles (1.13.3); the emitted mh_item carries the m loop only for m > 1, so every class v2 pack keeps its text. Class v3 also draws the dataset layout and the load windows per era (docs/plans/era-layout.md; the strided windowed load address and the interleaved mapping mh_addr, one text form in the three dialects). dataset[w] = item(w >> 4)[w AND 15]. A dataset of 2^D words is the prefix of items 0 .. 2^(D-4) - 1, so an item has the same value at every dataset size.
What the construction buys (Measured, MEMHARD.md section 2.2, M5 Max, seed igneum-genesis, 1 GiB):
| Kernel | Mhash/s | Ratio to honest |
|---|---|---|
| Honest (loads from the dataset buffer) | 45.2 | 1 |
| Inline, closed-form dataset (the prototype before this construction) | 5,014 | 111x faster |
| Inline, memory-hard (recomputes every word, never reads the dataset) | 9.49 | 0.21 (4.8x slower) |
| Inline, memory-hard, against a 256 MiB honest dataset | 9.48 vs 94.8 | 0.10 |
The inline kernel is bound by 104 x 8 = 832 dependent 64-byte cache reads per hash against 104 independent 4-byte reads for the honest kernel. Not measured: the same ratio on NVIDIA or AMD, partial-storage trade-offs between the two points, a smarter attacker kernel (hoisting the t-only part of round 0, caching hot lines), and a census of distinct cache lines touched per hash (MEMHARD.md section 3, items 1, 2, 4, 5). Analytically a hash touches up to 832 of 4,194,304 lines and a warp's working set is 26,624 lines.
Test vectors, day 2026-10-03 (vectors.json of the igneum-genesis-mh pack):
cache line 0 (segment 0, line 0):
355a86d2 7957db1c d21772af 6fc1e09b d55ce61d 6e6a278b d3f543ce 223d8e82
143ab337 2e9f05bd 2eb389bf 0c6e449e 5cfa4222 ba6560fe 8e3e1aa4 dbcc1d53
cache line 4194303 (segment 65535, line 63):
41190d91 bd277957 22ddbb49 6986f207 df69a4d6 26401a3a 818230fb c417122d
3597b211 b553ce55 cf39cc0d 3b7fc43a 3fd43b00 67e1c80e ffa7ea7d ca2960ab
cache fingerprint, FNV-1a 64 over all 2^26 words as little-endian bytes: 48c4f5bf24166b2e
dataset words 0..15 (item 0):
ffc3cd94 5920ccd8 392f44bb 5e57f67a 2f2bc2a9 620b0e36 bdc09014 436654bf
311e0b48 1abd93ad 59cc7ce8 ee5247b2 86171fe8 6d874751 c9f7728f 7c2a435d
dataset[0x0fffffff] = a33ada72
dataset[59471966] = e8b73d94 dataset[217795994] = 337028b5
dataset[3093825] = e3922dca dataset[267580473] = 26b5f1d8
The cache fingerprint has been reproduced by: the Swift CPU fill, the Rust fill, the Metal GPU fill, the clang emulation of the CUDA fill, the RTX 5090 (CUDA and NVIDIA OpenCL), Apple OpenCL, pocl, and the AMD gfx1036 (docs/bench-log.md, the five entries dated 3 October 2026 for memory-hard, igneum-pow, proto-opencl, gfx1036 and NVIDIA OpenCL).
1.9 The 32-lane unit of work
Definition. The hash is defined over an aligned group of 32 consecutive nonces g .. g + 31 with g AND 31 = 0 (wrapping modulo 2^32 at the top of the range). shfl exchanges registers within that group: lane l reads from lane l XOR mask, mask < 32, so the exchange never leaves the group. The verification unit is the group: hash(n) is computed by evaluating the group n AND ~31 and taking lane n AND 31 (verify.rs, Epoch::hash).
Exchange semantics on hardware: the group MUST be realised so that the exchange is exactly lane XOR mask over the 32 logical lanes, independent of the hardware wave width. The rule proto-opencl/WAVEFRONT.md fixes and kernel.cl implements:
| Path | When permitted |
|---|---|
Hardware sub-group or warp shuffle (simd_shuffle_xor, __shfl_xor_sync, sub_group_shuffle_xor) |
Only when the work-group is exactly 32 items and the device reports a sub-group size of exactly 32 for that kernel and work-group |
Local-memory exchange (each lane writes its register to shared memory, one barrier, reads slot lid XOR mask) |
Always permitted. REQUIRED whenever the sub-group size is not exactly 32 or cannot be queried per kernel |
The wave64 rule: on a device whose native wave is 64 lanes (AMD GCN, CDNA, RDNA compiled as wave64; figures approximate per WAVEFRONT.md), a wave holds two logical units. A sub-group shuffle would still compute lane XOR mask correctly (the emulator rows in WAVEFRONT.md show this), but the specification does not permit it because the lane-to-work-item mapping is not guaranteed and because a broadcast would read the wrong half. Such devices MUST use the local-memory exchange. Both paths produced the identical batch fingerprint f99fb375b3abeaf5 (FNV-1a 64 over the 2^13 outputs at base nonce 0, pack igneum-genesis-mh) on Apple OpenCL, pocl with sub-group shuffles, pocl with local memory, and seven emulator configurations including wave64 (docs/bench-log.md, proto-opencl entry); the AMD gfx1036 and the RTX 5090 via OpenCL printed 98af644e993239e2 at 2^24 outputs, identical to each other (gfx1036 and NVIDIA OpenCL entries). Cost of the local-memory path against the warp shuffle on the 5090: about 4%, approximate (NVIDIA OpenCL entry).
1.10 Target comparison
The hash is 64 bits. A block is valid for the lottery when hash <= target64. The mapping between the chain's 256-bit difficulty target (rusty-kaspa Uint256, fork point a3) and target64, and the block-level computation that pruning proofs use (calc_level_from_pow, fork map a3, assumes a uniform 256-bit output), are forward references to section 2 and are Open (section 6, O-2.4). Candidate: target64 = target256 >> 192 and level = leading_zeros(hash) capped at 64.
1.11 CPU verification procedure
Implemented (verify.rs, memhard.rs). A verifier holds the program for the epoch, the mixer parameters and the 256 MiB cache for the day. It never holds the dataset. To verify a block with lane nonce n:
- Evaluate the group
g = n AND ~31with the interpreter of section 1.7, register-major (r[reg][lane]) so lane loops vectorise. - At each
load, gather the 32 masked indices, deduplicate by item (idx >> 4), derive the distinct items with all chains interleaved round by round (all mixers for roundr, then all cache-line XORs for roundr), and hand each lane its word. Interleaving lets the eight dependent misses of each item overlap across up to 32 items; without it the verifier pays about 8 x 100 ns of DRAM latency per item in series. - Fold and compare lane
n AND 31againsttarget64.
The verifier does at most 128 x 32 = 4,096 item derivations per unit, the design bound (design document, Lottery seeds item 3), because every program has exactly 16 loads (1.4.2) and an accepted program reads at least 120 and typically 128 distinct addresses per hash (1.4.6). Under version 1 the count ran from 3,328 to 4,608 items.
| Verifier | ms per 32-lane unit, steady (avg of 20) | Worst cold single unit | Source |
|---|---|---|---|
| Rust, one M5 Max performance core, generator v2, 128 loads, 4,096 items | 0.631 | 0.67 to 0.81 across the three vector units | Measured 4 October 2026, igneum-pow/README.md |
| Rust, version 1, 104 loads, 3,328 items (retired) | 0.441 | 0.41 to 0.87 across five seeds | Measured, docs/bench-log.md, igneum-pow entry |
| Rust, version 1, 144 loads, 4,608 items (retired) | 0.579 | same | |
| Swift, 104 loads | 0.649 | 1.16 to 2.11 | Measured, memory-hard dataset entry |
| Swift, 144 loads | 1.205 | same | |
| Closed-form dataset (not memory-hard, for scale) | 0.002 (Rust), 0.017 (Swift) | same entries |
The 10 ms gate (Target) is met with a margin of about 16x steady and 12x worst-cold on this core. Not measured: a 2019-class laptop core (design document, "Three experiments before gate 3"), which is what fixes the gate.
1.12 Schedules: epoch, day, era
All times are DAA seconds since genesis (section 0.6). At 1 block per second one DAA second is about one block; the schedules below are written in DAA seconds so that block-rate steps (section 2) do not move them.
| Clock | Length | What changes | Label |
|---|---|---|---|
| Epoch | epoch_len(d) DAA s, base 3,600; the ladder 600, 900, 1,200, 1,800, 2,400, 3,600, 4,800, 7,200; set by 90% miner signal at a day boundary (sections 1.13.1 and 5.7) |
The program: new seed words from the VDF of section 4, new kernel | Designed (Counter ASIC 2.0 layer 9, 5 October 2026, docs/plans/epoch-length.md); 3,600 stays the value on every network until a signal moves it, and stays the prototype value to be fixed at gate 2 by the difficulty-tracking measurement (fork map c1: the hash rate steps by program, 35 to 48 Mhash/s across seeds on the M5 Max, so the DAA window must track within an epoch) |
| Day | 86,400 DAA s | The day key, hence the cache and the dataset | Designed |
| Era | 15,552,000 DAA s (180 days) | Era parameters and one instruction-family unlock, section 1.13 | Designed; the length is a prototype value (the design says "every 6 months") |
Epoch (d, e) covers DAA scores [86,400 d + L e, 86,400 d + L (e + 1)) with L = epoch_len(d) and e in 0 .. 86,400 / L; every ladder step divides 86,400, so day boundaries are epoch boundaries, and the epoch is identified by its start score s = 86,400 d + L e. At the base L = 3,600 this is [3,600 e, 3,600 (e + 1)) and nothing below differs from the earlier text. The epoch of a block is the epoch of its own DAA score, and epoch_len(d) is a function of the blue blocks of the signalling window that closed at least 2 days before day d (section 1.13.1), which are in the header's past, so "which program was this block mined under" is a function of the header alone once the seed is known. T_epoch and the 1,200-s lead of section 4.3 are genesis constants and do not follow epoch_len: the program of every epoch is known 600 s before it starts on the reference core at every length. The program for epoch e is generate_from_seed_bytes(program_seed_e): attempt 0 is drawn from S_e = seed_words_from_bytes(program_seed_e), and a rejected attempt is replaced as 1.4.6 says; program_seed_e is the 32-byte VDF output of section 4.3.
Implementation note (devnet, 3 October 2026, docs/fork-divergence.md "Epoch seed"): until the VDF of section 4 is in the node, program_seed_e is the hash of the last selected-chain block whose DAA score is below 3,600 e - 600. The 600-DAA-score lead stands in for section 4.3's 20-minute lead: the program of epoch e is knowable about 10 minutes before it starts, every block template reports it (pow_epoch.next_epoch_seed), and a GPU worker compiles it in the background and swaps at the boundary with no pause (serve protocol prepare, proto-metal/main.swift, proto-cuda/host.cu, proto-opencl/host.c). Measured across boundaries on a short-epoch test network in docs/bench-log.md (hot-swap entry). The program schedule is a protocol constant; a miner that cannot compile ahead sees the same seed at the same time as everyone else, only later.
Day d covers DAA scores [86,400 d, 86,400 (d + 1)). The design document names a day seed and does not say how it is derived. Proposed (Designed, Open, O-1.10): day_bytes = "igneum-day/" || d_le64 || program_seed of the first epoch of day d, so the day key is as unpredictable as the epoch seed and known 20 minutes before the day starts (section 4.5), which is enough for a 0.2 s CPU cache fill or a 2 ms GPU one plus a 13 to 30 ms GPU dataset build (Measured, section 1.8.3 and docs/bench-log.md RTX 5090 memory-hard entry: 13.4 ms for 1 GiB).
Era n covers DAA scores [15,552,000 n, 15,552,000 (n + 1)).
1.13 Era parameter draw, instruction-family reserve, dataset growth
Designed at the level of a sentence in the design document ("a new instruction mix and memory pattern drawn from chain state, plus one instruction family unlocked from a reserve fixed at genesis"; "the dataset grows on a genesis-fixed schedule"). No draw procedure, reserve list or growth rule exists in code (section 0.3). This section proposes them so that they can be reviewed; everything here is Open (section 6, O-1.11 to O-1.13) until gate 1 fixes it. The rule that is not open: nothing in this section is ever changed by a human release. The draw and the unlock are functions of genesis constants and chain state.
1.13.1 Era seed and draw
The era seed E_n is the 32-byte output of the 1-hour VDF of section 4.4 (in the node from era_vdf_activation_daa, 7 October 2026, era VDF lane: the delay over the hash of the cut block's day under the genesis scheme byte; before the activation, and on every network today, the devnet stand-in below). One SplitMix64 stream seeded from seed_words_from_bytes("igneum-era/" || n_le64 || E_n) words 0 and 1, drawn in a fixed order, sets the era parameters within genesis-fixed bounds:
| Parameter | Base (era 0) | Draw | Bound |
|---|---|---|---|
| Op weights for the ten non-load ops | table 1.4.2 | each perturbed by below(2 * B + 1) - B points, then renormalised by largest remainder to 100 minus the load weight |
B = 2 points, proposed |
| Load count | 16 of 64 | not drawn | fixed, so every era is equally memory-bound |
| Output fold rotations | (7, 14, 21), (9, 18, 27) | each 1 + below(31) |
1..31 |
| Mixer round count | 8 | not drawn | fixed, so the verify budget holds |
Mixer applications per round mixer_mult |
8 (class v3, LoadClass::MX8, decided x8 at 22:05 UTC on 5 October 2026; 1 under class v2; corrected 6 October 2026, the row had said 4) |
not drawn | fixed at genesis (Counter ASIC 2.0, 5 October 2026: the M16 recompute chip's only measured lever; section 1.8.5 carries the form; docs/plans/mixer-x4.md) |
Epoch length epoch_len |
3,600 DAA s | one draw of the era stream consumed and not used (the value is set by signal) | the ladder 600, 900, 1,200, 1,800, 2,400, 3,600, 4,800, 7,200 (layer 9, docs/plans/epoch-length.md) |
epoch_len is the one era-table parameter set by miners rather than by the draw: 90% of blue blocks over a 7-day window carrying the same ladder index (3 bits of the header version, encoding Open in section 5.8) sets that length from the first day boundary at least 2 days after the window closes (section 5.7). It is not a code upgrade: the rule, the ladder and the window are genesis constants, and the chain carries no release. The era stream consumes its draw so that a future draw of this parameter changes no other parameter's value. The threat it answers is a per-program hard datapath (an FPGA fleet: 42 to 160 minutes per compile on a mid-size part, PRflow, FPT 2019, hours on large parts; at 600 s nothing it compiles ever runs); it does not answer a programmable chip, which the other layers answer. The floor 600 is set by the slowest compile-ahead measured (the Metal variant race, 38 s on the M5 Max, 6.3% of a 600-s epoch and inside the 600-s seed window; docs/plans/epoch-length.md section 6).
The table layout and the working-set window (Counter ASIC 2.0 layers 4 and 8, decided IN on 5 October 2026, delegated: the six-era hash-rate spread is 1.3% on the RTX 5090, 3.2% on the RX 9070 XT and 0.8% on the M5 Max, under the 5% rule; docs/plans/era-layout.md) are drawn under program class v3 by a second stream S seeded with words 0 and 1 of seed_words_from_bytes("igneum-era/" || E_n) (the index is not in the preimage: E_n commits to n through the VDF input), seven draws in this order whether or not a value is used:
W = allowed[below(|allowed|)]: the width in words of every dataset load of the era, from the genesis-fixed setallowed; the set is{1}(4 bytes, the read-width decision of 5 October 2026), so the draw is consumed and the width pinned.M = low32(next()) OR 1: the stride multiplier, odd, sox -> x * Mis a bijection.R = 1 + below(31): the stride rotation.- to 7.
r_i = next()foriin 0..3: the interleave draws. Withb = log2(W)andfree = 4 - b,c = [b, ..., 15]; foriin0..free:j = i + (r_i mod (16 - b - i)), swapc[i]andc[j]; the interleave ispos = [0, ..., b - 1] ++ sort(c[0..free]), four ascending bit positions below 16.
The era parameters are (W, M, R, pos). Dataset mapping under class v3: word w holds word j(w) of item t(w), where bit i of j(w) is bit pos[i] of w and t(w) is w with bits pos[0..3] removed; with pos = [0, 1, 2, 3] this is dataset[w] = item(w >> 4)[w AND 15] byte for byte; an item keeps its value at every dataset size of at least 2^16 words, and the W words of one aligned load lie in one item, so the 4,096-item verifier bound of 1.11 holds. Load address under class v3, for a load site with window draws (k_off, o) and a dataset of 2^D words: k = min(k_off, D - 26), y = rotl(x * M, R), idx = ((y AND (MASK >> k)) OR ((o AND (2^k - 1)) << (D - k))) AND MASK (uniform on the window, branch-free, three operations before the mask), one text form in Metal, CUDA and OpenCL. The window draws per instruction (layer 8), after the nine draws of 1.4.3: k_off = below(3) (the dataset, a half or a quarter) and o = low32(next()) AND (2^k_off - 1), used only on a load slot, so a class v3 program takes 720 draws; the window never goes below 2^26 words (256 MiB, above the largest on-chip cache in the benchmark) nor above the dataset, and sixteen sites with drawn offsets cover the dataset with high probability (a windows-union census over 300 programs: the SRAM mirror a chip would need is the whole dataset in every hour). The acceptance rule of 1.4.6 is unchanged in its tests and mirrors this address at its constant D = 28. Devnet stand-in for E_n below the activation of the VDF of 4.4 (the VDF is in the node since 7 October 2026, behind era_vdf_activation_daa, never until the project lead sets it per network): era 0 the genesis block hash; era n >= 1 the hash of the last selected-chain block whose DAA score is below 15,552,000 n - 7,200, the same block the VDF reads its input from once active. The attack pass's F7 harness showed the stand-in grindable with one block of hash at no delay (1 of 6 cuts) and the VDF closing it (docs/analysis/era-vdf-2026-10-07.md). What the interleave buys and does not: a chip that hard-wires one layout reads the wrong 15 words with every word once the era draws another; a chip whose address decoder can permute its address lines pays nothing (stated in the plan). The stride is a bijection with no cryptanalysis yet (Open).
"Memory pattern" in the design document is read here as the item-address pattern (the cache line index word, s[0] in 1.8.5, and the XOR-all-sixteen rule); the proposal is to leave it fixed at era 0 and let the unlocked families change the kernel instead, because every change to the item derivation changes the verify time and must be re-measured.
1.13.2 Instruction-family reserve
At genesis the generator carries the eleven families of 1.4.1 live and a reserve list of further families in a fixed order. At the start of era n >= 1, reserve family n becomes live with weight W_new taken proportionally from the live non-load families. A family may enter the reserve only if it is integer-exact and has passed the cross-vendor conformance of section 1.15 on every vendor in the benchmark (Metal, CUDA, OpenCL on NVIDIA and AMD), with its own edge-case vectors, before genesis. Candidate families, all integer ALU operations present on Apple, NVIDIA and AMD: variable left shift and logical right shift by src AND 31; bit-field extract with an immediate offset and width; andn (dst = dst AND NOT src); byte permute of dst by an immediate selector; population count and count-leading-zeros folded into dst by add; a three-register select (dst = bit of src2 ? src : dst); a second shuffle form (lane + delta mod 32). The order and W_new are Open, except the first entry, decided 5 October 2026 (Counter ASIC 2.0 layer 7, delegated; the project lead confirms for the public testnet genesis; docs/analysis/int8-matrix-family.md):
Reserve family R1,
mm8(integer matrix). Semantics: section 2.2 ofdocs/analysis/int8-matrix-family.md, uint8 operands fromsrcandsrc2in the m8n8k16 fragment layout, one int32 element of C per lane selected by the immediatebit, added intodstmodulo 2^32. Weight at unlockW_new = 4points, taken proportionally from the ten live non-load families (the load weight and count are untouched). Edge vectors, each a hand-built unit run on every vendor: all bytes 0xFF in A and B (C = 1,040,400 everywhere); all bytes 0x80 (C = 262,144); A all zero (C = 0);dst= 0xFFFFFFFF with a nonzero C (the wrap); alternating 0x00 and 0xFF by lane;bit= 0 and 1 on the same fragments. Unlock: at the start of era n = 4 (DAA 62,208,000), or earlier by the 90% signalling path of section 5.7; never by a release. Native paths: PTXmma.sync.u8(sm_75+), AMD WMMAi32_16x16x16_iu8(RDNA 3 and 4), Metal 4mpp::tensor_ops::matmul2d(uchar x uchar -> int); the per-lanedot4form is emulation on Apple (1.6x per op unsigned, measured 5 October 2026) and is not the reserved form.
A vendor that can only emulate. A family enters the reserve when it is bit-exact on every vendor of 1.15. A vendor that reaches the result only by emulation (no instruction or library path) does not block entry if the measured penalty of the emulation on that vendor, on the family's own probe (a dependent chain of the op against the same vendor's integer ALU chain), is at most 8x per op, AND the family's weight at unlock keeps the emulating vendor's hash-rate loss under 5% on the memory-hard hash, checked on the vendor's card with the family live. A family whose emulation exceeds either bound stays out of the reserve until the vendor ships a path.
A family that is not in the genesis reserve can only be added by the upgrade path of section 5.7.
1.13.3 Dataset growth
Designed: 2 GiB at genesis plus 0.5 GiB per year (design document, "Which cards mine?", labelled approximate there for the card-lifetime consequence, not for the schedule). Proposed rule: the dataset for day d has N_d items where
N_d = floor((2 GiB + 0.5 GiB * (86,400 d / 31,536,000)) / 64 bytes)
evaluated in integers (bytes), with one year = 31,536,000 DAA seconds. The dataset grows by about 23 KiB per day on this average and is recomputed with the day key. Decided 5 October 2026 (Counter ASIC 2.0 layer 6, delegated; the project lead confirms for the public testnet genesis): the cache grows with the dataset, doubling when the dataset doubles: cache_log2_words(d) = 26 + growth_doublings(d), growth_doublings(d) = floor(log2(1 + d / 1,460)) for day d since genesis (doublings at years 4, 12, 28 and 60), so the cache is 256 MiB at genesis, 512 MiB from year 4, 1 GiB from year 12; the verifier's one-core fill is 0.2, 0.4 and 0.8 s at those steps (0.2 s per 256 MiB, section 1.12), under 1 s at every step of the schedule. The dataset steps to the next power of two on the same doublings (option (b) below, recommended to the project lead with the card-lifetime consequences in docs/analysis/card-lifetime-2026-10-05.md: a 4 GB card mines to year 4, an 8 GB card to year 12, a 12 GB card to year 28 with the cache freed after the daily build). Why the cache grows at all: an SRAM mirror of a flat 256 MiB cache is about 128 mm^2 and $46 of silicon at N5 by shipped cache-die density (AMD V-Cache, 64 MB on 41 mm^2 at 7 nm; docs/analysis/sram-mirror.md), so the cache size never prices a chip out; its job is to stay above any GPU's on-die cache (96 MB on the RTX 5090, 128 MB on GB202), which a flat 256 MiB loses within the decade. The GPU frees the cache after the daily dataset build; the hash never reads it. Two consequences remain Open:
- Index mapping.
src AND MASKrequires a power-of-two size. For a non-power-of-twoN_dthe proposed mapping isidx = (src * N_words) >> 32computed in 64 bits (a multiply-shift range reduction; uniform to within 2^-32, branch-free, integer only). AtN_words = 2^28this givessrc >> 4, notsrc AND MASK, so adopting it changes the 1 GiB vectors; gate 1 chooses between (a) the multiply-shift mapping with new vectors, or (b) power-of-two sizes only, growing in steps (2 GiB, 4 GiB) on the same schedule's average, which keepsAND MASKand means a 4 GiB card lasts until the 4 GiB step instead of fading. - The item index
tis 32 bits, so the construction as written tops out at 2^32 items = 256 GiB, which the schedule reaches after 508 years. No action needed.
1.14 Determinism requirements
A conforming implementation MUST:
- Use only integer arithmetic. No floating point anywhere, including in index computation and in the mixer (floating point rounds differently per vendor and would split the chain; design document, hostile review table row 2).
- Mask or range-reduce every dataset index exactly as 1.5 and 1.13.3 state, and never read outside the dataset. Every load in emitted source MUST have the single form
dataset[rN AND MASK](or the adopted range reduction), checkable by text search (igneum-pow/tests/packs.rs: 16 of 16 loads masked in every emitted kernel of every pack;TESTS.mdsection 5 for the version 1 run). - Implement
rotrbysrc AND 31androtlby an immediate in 1..31; a rotate by 0 or 32 through the immediate path is undefined and MUST NOT occur. - Compute
mulhias the exact high 32 bits of the 64-bit product (__umulhi,mulhi,mul_hi). - Wrap on overflow everywhere (add, sub, mul, mad, the SplitMix and FNV state).
- Realise the exchange over exactly 32 logical lanes as 1.9 requires, independent of the hardware wave width.
- Produce the same output for the same (program, day key, dataset size, nonce group) on every run:
TESTS.mdsection 4 (5 runs and 3 compiles, one cold, fingerprint933787e8cfefccb7closed-form;62a4f0eb018df273memory-hard,MEMHARD.mdsection 2.5).
1.15 Conformance procedure for a miner implementation
A miner, kernel emitter or verifier conforms when all of the following pass. Each is a command that exists today; the AMD discrete-card rows are pending. Every vector below is of generator version 2 (4 October 2026); the version 0.1 vectors are retired and MUST NOT be used.
- Cache: fill the 256 MiB cache for day
2026-10-03and reproduce FNV-1a 6448c4f5bf24166b2e, cache line 0 and cache line 4,194,303 of section 1.8.5, word for word (unchanged by version 2). For the devnet day bytes"igneum-day/" || 20730_le64:448274a57f508cbc. - Dataset self-test at 1 GiB: words 0..15, word
0x0fffffff, and the 64 sampled words ofvectors.json(four are quoted in 1.8.5). - Generator and acceptance: derive the five seeds of the 1.4.6 table to the same attempt and program id, and the
igneum-genesisinstruction list of 1.4.3. - Vectors: all 96 outputs of the three units at base nonces 0, 4,096 and 1,000,000 for pack
igneum-genesis-mh, standalone (one unit per launch) and in batch (many units per launch, at least two units per work-group or block). Section 1.17 lists them. The same forigneum-devnet-v4-epoch0, the chain's own derivation (epoch seed = the devnet genesis hash, day bytes of 2026-10-04). - Batch fingerprint: FNV-1a 64 over the 2^13 outputs at base nonce 0 =
f2a95d5bb84d961e(igneum-genesis-mh;8e22ad069cb2a8c3for igneum-devnet-v4-epoch0), and over the 2^24 outputs =25f96e7dce90bd4e(igneum-genesis-mh;3cc4fbf90fa6366cdevnet). - Fuzz: at least 200 random programs of the current generator (
--fuzz 200or the Rust equivalent when it exists), 4 units each at base nonces drawn from the full 32-bit range including wraps past 2^32, at 64 MiB, 256 MiB and 1 GiB, zero mismatches against the CPU interpreter, zero compile failures, generator contract asserted on every instruction. Reference: 2,000 of 2,000 under version 2 (TESTS.mdsection 9), 10,200 of 10,200 under version 1. - Edge: the 14 hand-built programs of
TESTS.mdsection 2 (rotates by 0 and 31 through the register path,mulhiat the extremes, every shuffle mask, loads at index 0 and at MASK through in-range and out-of-range registers, wraparound on add, sub, mul, mad, zero loads, 64 loads), 128 of 128 lanes each. These are hand-built and bypass the generator; they test the interpreter and the kernels, not the rule. - Static mask check on every emitted kernel (1.14 item 2).
- Exchange rule: on any device whose sub-group size is not exactly 32, or cannot be queried per kernel, the local-memory path is used and the run says so.
Measured conformance to date (docs/bench-log.md), version 2 vectors: Apple Metal natively (3 of 3 units on the exported igneum-genesis-mh, fuzz 2,000 of 2,000, and the Swift generator identical to the Rust one on the five seeds of 1.4.6), Apple OpenCL (96 of 96 on all four packs, fingerprints as in item 5), the clang CUDA emulation (96 of 96 on all four packs, 2 warps per block), the clang OpenCL emulation (96 of 96 on the two memory-hard packs in sub-group 32 and wave64 configurations), and the Rust CPU reference. Not yet run on version 2 vectors: the RTX 5090 (CUDA and NVIDIA OpenCL), AMD gfx1036, pocl; the version 1 runs on those devices (3 October 2026) stand as evidence that the kernel text, which version 2 did not change, agrees across vendors. A discrete AMD card has not run anything (ledger M8).
1.16 Parameters marked "prototype value, to be fixed at gate 1" and what fixes them
| Parameter | Prototype value | Measurement or decision that fixes it |
|---|---|---|
| Seed derivation (FNV-1a plus SplitMix64) | section 1.3 | Replace with a standard hash (ledger M7); re-run the weak-program census and re-cut every vector on the result (the attempt derivation and the program id of 1.4.6 go through the same function) |
| Instructions per program, iterations | 64 x 8 | CPU verify on a 2019-class laptop core under 10 ms with the memory-hard dataset; register-pressure and occupancy on the three vendors |
| Registers per lane | 8 | same |
| Op weights, including the 25% load weight | table 1.4.2 | Fixed 4 October 2026 by the weak-program census (docs/analysis/weak-program-census-2026-10-03.md): exact load count 16 (1.4.2), fresh-source rule (1.4.3), acceptance rule (1.4.6). The ten non-load weights remain a prototype value for the era draw bounds of 1.13.1 |
| Output fold rotations | (7, 14, 21), (9, 18, 27) | Stats run of TESTS.md section 3 on the chosen fold |
| Cache size, lines per segment, ChaCha rounds | 256 MiB, 64, 12 | Shortcut ratio and time-memory curve on an RTX 5090 and an AMD discrete card; external review of the chained-block cache |
| Item rounds, mixer shape | 8, section 1.8.4 | Same measurement; external review of the mixer; weak-key check on ROT |
| Dataset size at genesis | 1 GiB in packs, 2 GiB designed | Hash-rate and shortcut ratio at 2 GiB on the three vendors; the growth and index-mapping decision of 1.13.3 |
| Epoch length | 3,600 DAA s | Difficulty tracking across program steps on the devnet (gate 2, fork map c1) |
Era length, draw bounds, reserve list, W_new |
section 1.13 | Design review; each reserve family's own conformance run |
| Header binding of the init words | section 1.6 | Fixed when section 2 fixes the pre-PoW header hash; new vector set |
1.17 Test vectors
Pack proto-cuda/packs/igneum-genesis-mh/ (seed igneum-genesis, generator version 2, attempt 0, program id bcc1248b10cc90f2, day 2026-10-03, memory-hard, 2^28 words, MASK 0x0fffffff, 32 lanes). Produced by igneum-pow export (the Rust CPU interpreter) on 4 October 2026, reproduced by the Swift CPU interpreter and cross-checked by Metal, Apple OpenCL and the two clang emulations the same day (section 1.15). The version 0.1 vectors of 3 October 2026 (lane 0 1fb0b3bbc1ac8279) are retired: they belong to a generator that no longer exists in the protocol. The two closed-form packs (igneum-genesis, igneum-hourly) are regression vectors for the interpreter only and are not the lottery hash.
Unit at base nonce 0, lanes 0..31:
42246ba99fc58e4f 19561eec0db4f9f4 11a4fb70ea7b688f 8872bfadec960949
085e6cf2ff6b8303 780e0d76e504fe6c 7bb05a713f8283fa fef7beeffd5ab1d1
1b83afb12e78d65b 2576c5384a2e1cae 549d620a0735d6d8 2eb4e32f32e8d5f1
81bb414492f69586 092d17324b465e01 8bde350ee2354b5a d64b47e9c9e0ec07
dbbf1b78b9979a3f bee382abde89c111 6b598d99c16a8c70 69d72da59fb3b6e9
1e309d2a549632fa 98fa8255ab65f005 2f48ab1bb516110c 7d2af17cadd18bea
547c5978bd005e02 25ea2e21e88ea9d3 1c575ec6e43efc58 077b80cb079958b4
a96eebd0c8634981 44b18bdf19eb7838 c3ccb9fe9ef5953c b08446b1f2de7793
Unit at base nonce 4,096: lane 0 3d3903e310ca038f, lane 1 9e72b9a86ebb29e4, lane 31 61c242509efdccdd. Unit at base nonce 1,000,000: lane 0 f218c1bd58e6dfe0, lane 1 8c5a362ee98971c2, lane 31 6c3b2c11adfbfcac. The remaining 58 values are in vectors.json and vectors.h of the pack.
Pack proto-cuda/packs/igneum-devnet-v4-epoch0/: epoch seed bytes edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07 (the devnet genesis hash), day bytes 69676e65756d2d6461792ffa50000000000000 ("igneum-day/" || 20730_le64, 2026-10-04), attempt 0, program id 4be132dd1f2ff270, cache FNV-1a 64 448274a57f508cbc, dataset words 0 and 1 3dd50b1f 48edec90, word [MASK] f7b7180e; unit at base nonce 0: lane 0 285a83011e7ac3fc, lane 31 6f1136558c20e3d8.
Header-bound vectors (section 1.6 rule, seed igneum-genesis, day 2026-10-03): igneum-pow/README.md, eight values, for example H = 32 zero bytes and nonce 0 give 746c567b090acf6a.
Program class v3 vectors (5 October 2026, proto-cuda/packs-ca2-mixer/ and proto-cuda/packs-ca2-era/; generator 3, mixer x8, the cache growth rule, the era draw inside the class): pack mx8-genesis (seed igneum-genesis, no era, program id e323b9dcaf283a6f, batch fingerprint 7c28cfb06c5c65a9) and pack mx8-devnet-epoch0 (the devnet epoch seed and day of 1.17 with the era stand-in E_0 = the devnet genesis hash inside the class, program id 73bcbfe8ccf988f1, unit 0 lane 0 d424577fce4a7a60, batch fingerprint 90f794dd556f7a3b over 2^24 outputs at base nonce 0), reproduced by the Rust interpreter, Metal and Apple OpenCL on 5 October 2026 (3/3 standalone, 3/3 in batch, 96 of 96 lanes each); the six era packs era-0 to era-5 (the same epoch seed and day, era test seeds 0 to 5, program id 73bcbfe8ccf988f1; 2^24 fingerprints 8e8e070db4eea52d, 891c01b8563bb47e, e54279fed2831b5d, 77e0ba8abbd0ae62, d898d8f4f2e7684b, a6927db380f7efb2). All seven fingerprints are equal on the RTX 5090 (CUDA/NVRTC), the RX 9070 XT (AMD OpenCL) and the M5 Max (Metal), self-test PASS on every pack, and 1,024 random nonces per card on era-0 and mx8-devnet-epoch0 re-hash to the same value on the Rust verifier (1,024 of 1,024 each): job run-ca2-era-pc1b-20261005, 5 October 2026, docs/bench-log.md. The class v2 vectors above stand unchanged (the v2 exports are byte-identical on the class v3 crate, igneum-pow/tests/packs.rs).
Cache, dataset and mixer vectors: section 1.8.4 and 1.8.5. Seed words: section 1.3.1. Generator: section 1.4.3. Acceptance: section 1.4.6. Batch fingerprints: section 1.15 item 5.