547 lines
60 KiB
Markdown
547 lines
60 KiB
Markdown
# Igneum protocol specification, section 1: the lottery hash
|
|
|
|
Spec version 0.2, 4 October 2026 (0.1 on 3 October 2026). Status of this section: Measured for the construction as implemented (Implemented values with test vectors on three GPU vendors and a CPU reference); Designed for the header binding, the day key, the dataset growth and the era schedule; Open where marked. Version 0.2 adopts generator version 2 (sections 1.4.2, 1.4.3 and 1.4.6: exact load count, fresh-source loads, program acceptance) from the weak-program census of `docs/analysis/weak-program-census-2026-10-03.md`; every vector of version 0.1 is retired and re-cut (section 1.17).
|
|
|
|
Normative implementation: `igneum-pow/src/{seed,generator,accept,memhard,verify}.rs`. Where this text and that code disagree, the code and its test vectors win until this text is corrected (section 0.4). The Swift prototype `proto-metal/main.swift` carries the same generator and acceptance rule and is bit-exact with the crate on every pack (`docs/bench-log.md`, entries "igneum-pow: Rust crate bit-exact with proto-metal" and "generator version 2").
|
|
|
|
Every parameter marked "prototype value, to be fixed at gate 1" is carried by the implementation today, is part of the test vectors, and is confirmed or replaced by the named measurement before the hash is frozen (section 1.16).
|
|
|
|
## 1.1 What the hash is for
|
|
|
|
The lottery hash decides who produces the next block. It MUST be:
|
|
|
|
1. Deterministic and bit-exact on every conforming implementation: GPU kernels on any vendor, the CPU verifier, and any future implementation (section 1.14).
|
|
2. Cheap to verify on one CPU core without the dataset: one 32-lane unit of work in under 10 ms (Target, Measured at 0.63 ms steady and 0.67 to 0.81 ms cold under generator version 2, section 1.11).
|
|
3. Bound by random access to a dataset larger than any on-chip cache, so that computing dataset words is slower than loading them (Measured 4.8x slower on Apple, section 1.8.4; not measured on NVIDIA or AMD).
|
|
4. Unknowable until shortly before it is needed, so that a miner cannot grind the seed (section 4).
|
|
|
|
It does not need to be a general-purpose cryptographic hash (preimage, collision). It needs to be a fair lottery: no shortcut cheaper than honest evaluation and no bias a miner can exploit. No analysis of either property exists yet (`docs/fud-ledger.md`, entries M6 and M7; section 6).
|
|
|
|
## 1.2 Notation
|
|
|
|
All arithmetic is on unsigned 32-bit integers modulo 2^32 unless a 64-bit type is named. `rotl(x, n)` and `rotr(x, n)` rotate by `n` in 0..31. `mulhi(a, b)` is bits 32..63 of the 64-bit product. `low32(v)` is bits 0..31 of a 64-bit `v`. `||` is byte concatenation. Indices are zero based. Words are hashed into FNV as little-endian bytes (section 0.6).
|
|
|
|
## 1.3 Seed words and the SplitMix64 stream
|
|
|
|
Implemented, prototype value, to be fixed at gate 1 (ledger M7 asks for a standard hash in place of FNV-1a plus SplitMix so the seed-to-program mapping is auditable; the measurement that fixes it is the weak-program census of section 1.16, which must be re-run on whichever derivation is chosen).
|
|
|
|
### 1.3.1 `seed_words_from_bytes(b) -> [u32; 8]`
|
|
|
|
For `salt` in 0..3:
|
|
|
|
```
|
|
basis = 0xcbf29ce484222325 XOR (salt * 0x9E3779B97F4A7C15) (64-bit, wrapping)
|
|
h = basis
|
|
for each byte x of b: h = (h XOR x) * 0x100000001b3 (64-bit, wrapping)
|
|
h = h XOR (h >> 33); h = h * 0xff51afd7ed558ccd; h = h XOR (h >> 33)
|
|
words[2 * salt] = low32(h)
|
|
words[2 * salt + 1] = h >> 32
|
|
```
|
|
|
|
`seed_words(s)` is `seed_words_from_bytes` over the UTF-8 bytes of the string `s`. This function is the boundary at which the chain's bytes enter the hash: the epoch program seed from section 4 and the day key bytes (section 1.12) both go through it. In the prototype the input is a string ("igneum-genesis", "day/2026-10-03").
|
|
|
|
Test vector: `seed_words("igneum-genesis")` = `67a9a7be 1a155b25 fddfb732 4b5af2e8 c55caf33 a27c13b7 06628a48 03852469` (`proto-cuda/packs/igneum-genesis-mh/program.json`, `seed_words`).
|
|
|
|
### 1.3.2 SplitMix64
|
|
|
|
State `s` (64-bit). `next()`:
|
|
|
|
```
|
|
s = s + 0x9E3779B97F4A7C15
|
|
z = s
|
|
z = (z XOR (z >> 30)) * 0xBF58476D1CE4E5B9
|
|
z = (z XOR (z >> 27)) * 0x94D049BB133111EB
|
|
return z XOR (z >> 31)
|
|
```
|
|
|
|
`below(n)` = `next() mod n` (modulo, not rejection sampling; the bias at n <= 100 is under 2^-57 and is part of the definition).
|
|
|
|
### 1.3.3 The program stream
|
|
|
|
For seed words `w[0..7]`: `lo = w[0] | (w[1] << 32)`, `hi = w[2] | (w[3] << 32)`, and the generator's SplitMix64 state starts at `lo XOR (hi * 0x9E3779B97F4A7C15)` (64-bit wrapping multiply). Words `w[4..7]` are not used by the generator; all eight are used by the register initialisation (section 1.6).
|
|
|
|
## 1.4 The generator
|
|
|
|
Implemented (`igneum-pow/src/generator.rs`). A program is a list of `INSTR_COUNT` instructions over 8 lane registers `r0..r7`, executed `ITERATIONS` times per hash.
|
|
|
|
| Parameter | Value | Label |
|
|
|---|---|---|
|
|
| Registers per lane | 8 x u32 | prototype value, to be fixed at gate 1 (fixed by the ASIC-gain target and the register-pressure measurement of 1.16) |
|
|
| Instructions per program | 64, of which exactly 16 are `load` (section 1.4.2) | prototype value, to be fixed at gate 1 (fixed by the CPU-verify measurement on a 2019-class core); the load count is Definition since 4 October 2026 |
|
|
| Iterations per hash | 8 | prototype value, to be fixed at gate 1 (same measurement) |
|
|
| Lanes per unit of work | 32 | Definition. Not a tuning parameter (section 1.9) |
|
|
| Shuffle masks | {1, 2, 4, 8, 16} | Definition, follows from 32 lanes |
|
|
| Rotate immediates | 1..31 | Definition. 0 is never emitted and MUST NOT be relied on (`proto-metal/TESTS.md` section 2) |
|
|
|
|
### 1.4.1 Instruction set
|
|
|
|
Eleven families. `dst`, `src`, `src2` name registers; `src != dst` always; `src2` may equal either.
|
|
|
|
| Op | Semantics (per lane) | Operands used |
|
|
|---|---|---|
|
|
| `add` | `dst = dst + src + (bit `bit` of sel ? imm2 : imm)` | dst, src, imm, imm2, bit |
|
|
| `sub` | `dst = dst - src` | dst, src |
|
|
| `mul` | `dst = low32(dst * src)` | dst, src |
|
|
| `mulhi` | `dst = mulhi(dst, src)` | dst, src |
|
|
| `xor` | `dst = dst XOR src` | dst, src |
|
|
| `or` | `dst = dst OR src` | dst, src |
|
|
| `rotl` | `dst = rotl(dst, rot)`, rot in 1..31 | dst, rot |
|
|
| `rotr` | `dst = rotr(dst, src AND 31)` | dst, src |
|
|
| `mad` | `dst = low32(src * src2) + dst` | dst, src, src2 |
|
|
| `shfl` | `dst = dst XOR src_of_lane(lane XOR mask)` | dst, src, mask |
|
|
| `load` | `dst = dst XOR dataset[src AND MASK]` | dst, src |
|
|
|
|
`sel` is the value of `r0` sampled once at the top of each iteration, before instruction 0, and held for all 64 instructions of that iteration. This is the per-hash nonce-dependent select: the immediates an `add` uses depend on the lane's own state, so no two nonces run the same constant sequence (ProgPoW-style data-dependent path, integer only).
|
|
|
|
A twelfth family `wload` (warp-coalesced 128-byte load) exists in the code as lever (b) and is never emitted at the default configuration (`wide_frac = 0`). It is NOT part of the lottery hash. It is retained only so the measurement in `proto-metal/MEMHARD.md` section 2.4 stays reproducible, and the recommendation there is not to adopt it.
|
|
|
|
### 1.4.2 Op weights and the load count
|
|
|
|
Implemented (generator version 2, 4 October 2026, `igneum-pow/src/generator.rs`; ledger M5 Fixed). Every program contains exactly 16 `load` instructions (`LOAD_SLOTS`, Definition), so every hash performs 128 loads and a 32-lane unit derives at most 4,096 dataset items, the bound of section 1.11. The other 48 instructions are drawn from the ten non-load families with these weights (sum 75, prototype value, to be fixed at gate 1), applied in this order:
|
|
|
|
| Op | Weight |
|
|
|---|---|
|
|
| add | 12 |
|
|
| xor | 10 |
|
|
| mul | 8 |
|
|
| mad | 8 |
|
|
| shfl | 8 |
|
|
| rotl | 7 |
|
|
| sub | 6 |
|
|
| mulhi | 6 |
|
|
| rotr | 6 |
|
|
| or | 4 |
|
|
|
|
The load weight 25 of version 1 is retired; it survives only as the ratio 16 of 64. Why the count is fixed and not merely expected: under version 1 the static count ran 24 to 256 loads per hash over 100,000 programs and the GPU rate tracked the number of distinct addresses, 24 to 200, because 19.9 percent of all loads re-read an address the same hash had already read (census sections 3 and 5). Fixing the static count alone would not fix the memory work; the fresh-source rule of 1.4.3 fixes the distinct count by construction, and 1.4.6 rejects the few programs where it cannot.
|
|
|
|
The version 1 lever measurement (`load_weight = 17`, `proto-metal/MEMHARD.md` section 2.4) is kept in the code as `generate_v1` for reproduction only; its programs are not the lottery hash.
|
|
|
|
### 1.4.3 Draw order
|
|
|
|
From the program stream of 1.3.3, in this order, whether or not an op uses a value.
|
|
|
|
(1) Load slots. Let `p[0..62] = 1..63` (instruction 0 is never a load: nothing is fresh before it). For `i` in 0..15 draw `j = i + below(63 - i)` and swap `p[i]` and `p[j]`. The load slots are `p[0..15]`, a uniform 16-subset of 1..63.
|
|
|
|
(2) For each instruction `k` in 0..63, nine draws:
|
|
|
|
```
|
|
roll = below(75); op = first entry of the table of 1.4.2 whose cumulative weight exceeds roll
|
|
(on a load slot the roll is drawn and ignored and op = load)
|
|
dst = below(8)
|
|
src: on an ALU slot a = below(7); src = a + (a >= dst)
|
|
on a load slot E = the registers other than dst, in register order, that an earlier instruction of this
|
|
program has written and that no later load has used as its source (E is empty before
|
|
instruction 0); if E is not empty, a = below(|E|) and src = E[a];
|
|
if E is empty, a = below(7) and src = a + (a >= dst), and 1.4.6 (a) rejects the program
|
|
b = below(8) (src2)
|
|
imm = low32(next())
|
|
imm2 = low32(next())
|
|
rot = 1 + below(31)
|
|
bit = below(32)
|
|
mask = 1 << below(5)
|
|
```
|
|
|
|
16 slot draws plus 64 x 9: 592 draws per program. A program is fully determined by its eight seed words. A load's source holds a value written in the same iteration that no earlier load has read, so no load repeats the address of an earlier load of the same hash, across the iteration boundary included (the first form of the rule, with every register eligible at instruction 0, left the wrap open and failed 1.4.6 (a) on 36 percent of programs; census section 7.1).
|
|
|
|
Test vector: for seed `igneum-genesis` (attempt 0, program id `bcc1248b10cc90f2`, section 1.4.6) the first eight instructions are (`proto-cuda/packs/igneum-genesis-mh/program.json`):
|
|
|
|
```
|
|
0: mad dst=2 src=3 src2=4 imm=0xbf7b174d imm2=0x337b762e rot=17 bit=2 mask=2
|
|
1: mad dst=2 src=1 src2=1 imm=0xdd04a5da imm2=0x42da7657 rot=15 bit=30 mask=16
|
|
2: mad dst=2 src=3 src2=2 imm=0x734003fa imm2=0x5bb67700 rot=3 bit=20 mask=1
|
|
3: xor dst=3 src=5 src2=5 imm=0xc55a1b1c imm2=0xa19720f3 rot=7 bit=8 mask=1
|
|
4: load dst=7 src=2 src2=5 imm=0xad572dd7 imm2=0x9ceb3ea7 rot=18 bit=30 mask=2
|
|
5: load dst=5 src=7 src2=3 imm=0x769a53be imm2=0x80f9067e rot=12 bit=22 mask=1
|
|
6: shfl dst=1 src=4 src2=0 imm=0xd3613d88 imm2=0x262fb219 rot=10 bit=30 mask=8
|
|
7: shfl dst=7 src=3 src2=4 imm=0xce38e42f imm2=0xb868b818 rot=11 bit=8 mask=8
|
|
```
|
|
|
|
The whole program has op mix `load=16 add=8 shfl=8 xor=6 mad=5 mul=5 mulhi=5 sub=4 rotl=3 rotr=3 or=1`, 128 loads per hash. Instruction 4 reads r2, written by instructions 0 to 2; instruction 5 reads r7, written by instruction 4.
|
|
|
|
### 1.4.4 Generator contract
|
|
|
|
Every emitted instruction satisfies: `rot` in 1..31, `mask` in {1, 2, 4, 8, 16}, `src != dst`; every program has exactly 16 `load` instructions and none at instruction 0. This held on every instruction of 10,200 fuzzed version 1 programs (`TESTS.md` section 1) and of 2,000 fuzzed version 2 programs (`TESTS.md` section 9, 4 October 2026). A kernel emitter MAY rely on it; an interpreter MUST NOT accept a program that violates it.
|
|
|
|
### 1.4.5 Encoding
|
|
|
|
A program is transmitted as the seed bytes, never as instructions. A node hands a miner the pack it emits itself (`igneum-pow/src/emit.rs`: `kernel.cu`, `kernel.cl`, `program.metal`, `program.h`, `program.json`, `memhard.h`, `vectors.*`, the three `*_bound` kernels), and a miner MAY regenerate everything from the seed bytes by the procedure of 1.4.6. `program.json` (format `igneum-program-pack-3`) is the interchange form; its field names are those of `Instr` in `generator.rs`, and it carries `generator` (2), `attempt`, `program_id` and `seed_bytes`. `program.h` carries the same as `IGNEUM_GENERATOR`, `IGNEUM_PROGRAM_ATTEMPT`, `IGNEUM_PROGRAM_ID` and `IGNEUM_SEED_BYTES_HEX`. An implementation MUST refuse a pack whose generator version is not its own. Program class v3 (Counter ASIC 2.0, 5 October 2026, activated by the height switch `program_class_v3_activation_daa` from the first epoch whose start score is at or above it, `docs/plans/counter-asic-2-rollout.md`) writes `generator` 3, and every pack of it carries `IGNEUM_PROGRAM_CLASS` (`v3`) and `IGNEUM_ERA_SEED_HEX` (the 32-byte era seed of section 1.13.1, or its devnet stand-in) beside `IGNEUM_GENERATOR`; the serve protocol's `prepare` and `job` lines carry `class=v3 era=<hex>` for v3 epochs and nothing for v2 ones. A worker MUST refuse a pack whose class or era seed does not match the line it was prepared for (`igneum-pow/src/packcheck.rs`, `verify_pack_dir_chain`; `proto-cuda/nvrtc/packfile.h`), and a pack of a generator other than 2 or 3.
|
|
|
|
### 1.4.6 Program acceptance
|
|
|
|
Implemented (`igneum-pow/src/accept.rs`, `proto-metal/main.swift`; ledger M6 Fixed). A candidate program is accepted only if all of the following hold, and every conforming implementation MUST evaluate them identically.
|
|
|
|
(a) For every `load`, some instruction between the previous `load` from the same source register and this one, in cyclic order over the 64 instructions, writes that register.
|
|
|
|
(b) Every register `r0..r7` is the destination of at least one `add`, `sub`, `xor`, `mad`, `shfl` or `load`.
|
|
|
|
(c) The program is interpreted (section 1.7) for 64 units at base nonces `low32(next()) AND NOT 31` from a SplitMix64 stream seeded with `FNV-1a-64("igneum-accept/" || seed words as little-endian bytes)`, with init words `I` equal to the seed words and dataset words `dataset_elem(idx, S[0], S[1])` of `verify.rs` (the six-operation closed form of the version 0.1 packs) at 2^28 words (`idx = src AND 0x0fffffff`, a constant of this rule whatever the live dataset size) in place of the memory-hard dataset. Over those 2,048 evaluations: no register has a bit equal in every final value; no `load` site (iteration, instruction) reads one address in all 32 lanes of any unit; the number of final register values equal to 0 or 2^32 - 1 is below 164 (1 percent of 16,384); every output bit's ones count is within 136 of 1,024 (6 sigma); and the number of distinct masked dataset addresses read by one lane in one evaluation, summed over the 2,048 evaluations, exceeds 245,760 (a mean above 120 of the 128 loads).
|
|
|
|
Attempts. Attempt 0 of a program seed `b` (the 32-byte epoch seed, or the UTF-8 of a seed string) is the candidate drawn from `seed_words_from_bytes(b)`. If it fails, attempt `k = 1, 2, ...` is drawn from `seed_words_from_bytes(b || k_le32)`; the first accepted candidate is the program of the epoch. Measured rejection rate under this generator: 5.14 percent over 100,000 seeds (census section 7) and the 20,000-seed confirmation of `docs/bench-log.md` (4 October 2026), so the probability that 32 consecutive candidates fail is below 2^-136, and an implementation MAY treat 32 consecutive failures as a consensus fault (`MAX_ATTEMPTS`).
|
|
|
|
Program id. `FNV-1a-64("igneum-program/" || generator_le32 || seed words as little-endian bytes || attempt_le32)` with `generator = 2` under class v2 and `generator = 3` under class v3, written into every pack. Two implementations that agree on the id agree on the generator version, the seed words and the attempt.
|
|
|
|
Why the closed form: the test is then a pure function of the program (no cache, no day), costs 1.3 to 3.4 ms on one core, and the census checked on 100,000 programs that its verdict agrees with the memory-hard dataset's on all but 39 threshold-edge cases (section 7.3). What the three parts catch: (a) the empty-list fallback of 1.4.3; (b) registers that saturate to all ones (2.4 percent of candidates); (c) zero-absorbing register sets, lane-constant load sites, output bias and value-level address repeats (2.1 percent). Not in the rule, and why: a contraction as the last write (80 percent of programs) and the `or` count are too common and (c) already catches the cases that matter; the load critical path is a hash-rate question, not a weakness.
|
|
|
|
Test vectors for the rule (`igneum-pow accept --seed ...`):
|
|
|
|
| Seed | Attempt 0 | Attempt 1 |
|
|
|---|---|---|
|
|
| `igneum-genesis` | accepted, program id `bcc1248b10cc90f2`, 128.000 distinct addresses per hash, 0 saturated, bias max 54 | |
|
|
| `igneum-hourly` | accepted, `a4c4d00961c855df` | |
|
|
| `igneum-census-2026-10-03/22` | rejected, (b) r7 has no injecting write | accepted, `22ed0609d079f4cf` |
|
|
| `igneum-census-2026-10-03/37` | rejected, (c) 245,230 distinct addresses (mean 119.74) | accepted, `947705cc4eb1df0a` |
|
|
| `igneum-census-2026-10-03/51` | rejected, (b) r4 has no injecting write | accepted, `9869afcc028bf9f1` |
|
|
|
|
The Rust crate and the Swift prototype derive identical instruction lists and identical 96-vector sets on all five seeds (`docs/bench-log.md`, 4 October 2026).
|
|
|
|
## 1.5 Fixed memory footprint
|
|
|
|
Implemented for the prototype size; Designed for genesis.
|
|
|
|
| Quantity | Prototype (packs, vectors) | Genesis | Label |
|
|
|---|---|---|---|
|
|
| Dataset words | 2^28 (1 GiB) | 2^29 (2 GiB) | Prototype: Implemented. Genesis: Designed (design document, "Which cards mine?") |
|
|
| MASK | 0x0fffffff | 0x1fffffff | as above |
|
|
| Dataset item | 16 words (64 bytes) | same | Implemented |
|
|
| Cache | 2^26 words (256 MiB) | same | Implemented, prototype value, to be fixed at gate 1 (fixed by the shortcut-ratio measurement on the RTX 5090 and an AMD discrete card: the cache must exceed the largest on-chip cache of any card that mines, and 96 MiB of L2 on the 5090 is the figure to beat) |
|
|
|
|
A load reads one 4-byte word at `src AND MASK`. Every load in every emitted kernel has exactly this form; the static check in `TESTS.md` section 5 is part of conformance (section 1.15). Because item values do not depend on the dataset size (section 1.8.5), the 1 GiB vectors remain valid for words below 2^28 at any larger size.
|
|
|
|
Growth beyond genesis is in section 1.13; under program class v3 the cache follows the dataset's doublings (1.13.3) and the verifier holds 256 MiB, then 512 MiB from year 4 and 1 GiB from year 12.
|
|
|
|
## 1.6 Register initialisation
|
|
|
|
Implemented (`verify.rs`, `interpret_warp`). For lane nonce `n` (32-bit) and init words `I[0..7]`:
|
|
|
|
```
|
|
splitmix32(x): x ^= x >> 16; x *= 0x7feb352d; x ^= x >> 15; x *= 0x846ca68b; x ^= x >> 16
|
|
for i in 0..7:
|
|
x = n XOR I[i]
|
|
x = x + 0x9e3779b9 * (i + 1)
|
|
x = splitmix32(x)
|
|
r[i] = x XOR I[(i + 1) AND 7]
|
|
```
|
|
|
|
In the prototype and in every pack, `I` is the program's own seed words (the same eight words that drive the generator). On the chain the hash MUST also commit to the block being mined, which the prototype does not do. The binding is Designed, proposed here, Open until gate 1 (item O-1.9 in section 6):
|
|
|
|
- The header nonce is 64 bits (rusty-kaspa `Header.nonce`). Its low 32 bits are the lane nonce `n`. Its high 32 bits and the 256-bit pre-PoW header hash `H` (section 2, fork point a5) form the init words: `I = seed_words_from_bytes("igneum-block/" || H || nonce_hi_le32)`.
|
|
- `I` is a kernel argument, not a compile-time constant. The program (from the epoch seed) is compiled once per epoch; `I` changes per block template.
|
|
- The packs' vectors, where `I` equals the program seed, remain the conformance vectors for the generator, the interpreter and the dataset. A second vector set with `I` derived from a header is produced when section 2 fixes the header hash.
|
|
|
|
## 1.7 Execution of one hash
|
|
|
|
```
|
|
r = init(n, I)
|
|
repeat ITERATIONS (8) times:
|
|
sel = r0
|
|
for each instruction in order: apply it (section 1.4.1)
|
|
lo = r0 XOR rotl(r1, 7) XOR rotl(r2, 14) XOR rotl(r3, 21)
|
|
hi = r4 XOR rotl(r5, 9) XOR rotl(r6, 18) XOR rotl(r7, 27)
|
|
hash = (hi << 32) | lo (64 bits)
|
|
```
|
|
|
|
The output folding rotations (7, 14, 21; 9, 18, 27) are Implemented, prototype value, to be fixed at gate 1 (the stats run of `TESTS.md` section 3 is the check; a different fold must pass the same run).
|
|
|
|
`shfl` makes the 32 lanes of a unit interdependent: the hash of one nonce is defined only as a member of its aligned group of 32 (section 1.9).
|
|
|
|
## 1.8 The memory-hard dataset
|
|
|
|
Implemented (`memhard.rs`), construction and measurements in `proto-metal/MEMHARD.md`. Every constant below is a prototype value, to be fixed at gate 1, unless marked Definition. What fixes them is the shortcut-ratio and time-memory trade-off measurement on NVIDIA and AMD discrete cards (section 1.16 items 2 and 3) and an external review of the primitives (ledger M7).
|
|
|
|
### 1.8.1 Day key
|
|
|
|
`K[0..7] = seed_words_from_bytes(day_bytes)`. In the prototype `day_bytes` is the UTF-8 of `"day/" + day` with `day` an ISO date; on the chain see section 1.12. For `day/2026-10-03`: `K = 3067619f 3c269176 84a03b03 f8c63294 ff977c5b e60def3e 63630141 b8fbcb58`.
|
|
|
|
### 1.8.2 Block function B (ChaCha12 with feed-forward)
|
|
|
|
Standard ChaCha quarter round with rotations (16, 12, 8, 7), six double rounds (columns then diagonals, on the 16-word state), then `y[i] = y[i] + x[i]`. Twelve rounds, no key schedule beyond the input block. `CHACHA_ROUNDS = 12` is a prototype value; `sigma = (0x61707865, 0x3320646e, 0x79622d32, 0x6b206574)` is a Definition.
|
|
|
|
### 1.8.3 Cache fill
|
|
|
|
The cache is 2^26 words = 2^22 lines of 16 words, in 2^16 segments of 64 lines. Segment `s`, line `j` is at word offset `(s * 64 + j) * 16`. Each segment is a chain:
|
|
|
|
```
|
|
tag = (0x49676e65, 0x756d4d48) "Igne", "umMH"
|
|
prev = 0^16
|
|
for j in 0..63:
|
|
in = prev XOR (sigma[0..3] || K[0..7] || s || j || tag[0..1]) 16 words
|
|
line = B(in)
|
|
cache[segment s, line j] = line
|
|
prev = line
|
|
```
|
|
|
|
Line `j` costs `j + 1` block evaluations from nothing, 32.5 on average. The 65,536 segments are independent (one GPU thread each). The fill is a once-per-day cost.
|
|
|
|
| Fill time | Value | Source |
|
|
|---|---|---|
|
|
| One M5 Max core, Rust | 175 to 181 ms | Measured, `docs/bench-log.md`, igneum-pow entry |
|
|
| One M5 Max core, Swift | 184.5 to 190.6 ms | Measured, same log, memory-hard dataset entry |
|
|
| RTX 5090, one host thread | 223 ms | Measured, same log, "RTX 5090, memory-hard dataset" |
|
|
| M5 Max GPU | 0.6 to 2.1 ms | Measured, memory-hard dataset entry (variance not isolated) |
|
|
| RTX 5090 GPU | 0.67 ms | Measured, "RTX 5090, memory-hard dataset" |
|
|
|
|
### 1.8.4 Mixer parameters and the mixer M_r
|
|
|
|
One SplitMix64 stream seeded with `K[0] | (K[1] << 32)`, drawn in this order: `ROT[0..7] = 1 + below(31)` (eight draws), `MUL[0..15] = low32(next()) OR 1` (sixteen draws, always odd so each multiply is a bijection), `RC[0..15] = low32(next())` (sixteen draws).
|
|
|
|
`M_r(s)` on a 16-word state with round key `rk = (r + 1) * 0x9E3779B9`:
|
|
|
|
```
|
|
for i in 0..15: s[i] = (s[i] XOR (RC[i] + rk)) * MUL[i]
|
|
QR(s0, s4, s8, s12; ROT0..3) QR(s1, s5, s9, s13; ROT0..3) QR(s2, s6, s10, s14; ROT0..3) QR(s3, s7, s11, s15; ROT0..3)
|
|
QR(s0, s5, s10, s15; ROT4..7) QR(s1, s6, s11, s12; ROT4..7) QR(s2, s7, s8, s13; ROT4..7) QR(s3, s4, s9, s14; ROT4..7)
|
|
```
|
|
|
|
where `QR(a, b, c, d; r1, r2, r3, r4)` is the ChaCha quarter round with those four rotations. About 130 integer operations. This is the prototype's stand-in for RandomX's SuperscalarHash: fixed shape, seed-drawn constants. It has had no cryptanalysis (`MEMHARD.md` section 3, unproven item 3; a `ROT` draw of eight equal values is possible and untested).
|
|
|
|
Test vector, day `2026-10-03` (`proto-cuda/packs/igneum-genesis-mh/program.h`):
|
|
|
|
```
|
|
ROT = 20 20 19 4 26 3 3 27
|
|
MUL = 42146205 52cbe0fb 7ecf4a03 6728907f d81d9751 132952c3 f60de277 05358035
|
|
baf6499d e4db9667 3e98f45d d0004edd 2691630d 9beb3bcf ab310379 99cfb423
|
|
RC = bab68293 cc162340 6ce151cc e62b8997 c9c80297 f74a1654 3d704af5 3cf522b7
|
|
2b9cac04 a880ac10 13e5dd1d 6fc3e233 2d83eeac 9006e8bf 2c4b5362 31b49ee2
|
|
```
|
|
|
|
### 1.8.5 Item derivation and dataset mapping
|
|
|
|
Item `t` (16 words):
|
|
|
|
```
|
|
s[i] = K[i] for i in 0..7
|
|
s[8 + i] = t * MUL[i] + RC[i] for i in 0..7
|
|
for r in 0..7:
|
|
s = M_r(s)
|
|
a = s[0] AND 0x003fffff cache line index, 2^22 lines
|
|
s[i] = s[i] XOR cache[line a][i] for i in 0..15
|
|
s = M_8(s)
|
|
item(t) = s
|
|
```
|
|
|
|
Eight dependent cache reads (`ITEM_ROUNDS = 8`, prototype value): the address of read `r` depends on every earlier read. Nine mixer applications under program class v2.
|
|
|
|
Program class v3 (Counter ASIC 2.0, decided 5 October 2026, delegated; the project lead confirms for the public testnet genesis) applies the mixer `m = 8` times per round with distinct round keys (`LoadClass::mixer_mult`; `docs/plans/mixer-x4.md` section 2), the eight dependent reads unchanged:
|
|
|
|
```
|
|
for r in 0..7:
|
|
for j in 0..m-1:
|
|
s = M(s, rk = (r * m + j + 1) * 0x9E3779B9)
|
|
a = s[0] AND (2^(C - 4) - 1) cache line index, 2^(C - 4) lines of a 2^C-word cache (C from 1.13.3)
|
|
s[i] = s[i] XOR cache[line a][i] for i in 0..15
|
|
for j in 0..m-1:
|
|
s = M(s, rk = (8 * m + j + 1) * 0x9E3779B9)
|
|
```
|
|
|
|
Under `m = 1` the keys are `(r + 1) * 0x9E3779B9` and `9 * 0x9E3779B9`, the class v2 text exactly; the `9 m` keys are the first `9 m` values of the sequence `k * 0x9E3779B9`, all distinct. Why `m = 8`: the recompute attacker's cost is operations per item (`docs/analysis/m16-recompute-attacker-2026-10-05.md`); the honest miner pays the mixer once a day in the dataset build, which stays latency-bound (RTX 5090 23 to 25 ms, RX 9070 XT 72 to 77 ms, M5 Max 21 ms at `m` = 1, 4 and 8, measured 5 October 2026); the verifier pays `m` per item it derives: 0.61 ms per warp at `m = 1`, 1.24 at 4, 2.08 at 8 on one loaded M5 Max core (measured 5 October 2026, `docs/plans/mixer-x4.md` 6.4a), inside the 10 ms gate. The on-die-cache recompute chip's gain against the RTX 5090 falls from 2.4x (`m = 1`) to 0.92x with a 3x fixed-function factor at `m = 8` (`docs/analysis/chip-model-v3.md`, approximate factor). Under class v3 an item's value also depends on `C` through the line mask, so the items change on the day the cache doubles (1.13.3); the emitted `mh_item` carries the `m` loop only for `m > 1`, so every class v2 pack keeps its text. Class v3 also draws the dataset layout and the load windows per era (`docs/plans/era-layout.md`; the strided windowed load address and the interleaved mapping `mh_addr`, one text form in the three dialects). `dataset[w] = item(w >> 4)[w AND 15]`. A dataset of 2^D words is the prefix of items `0 .. 2^(D-4) - 1`, so an item has the same value at every dataset size.
|
|
|
|
What the construction buys (Measured, `MEMHARD.md` section 2.2, M5 Max, seed igneum-genesis, 1 GiB):
|
|
|
|
| Kernel | Mhash/s | Ratio to honest |
|
|
|---|---|---|
|
|
| Honest (loads from the dataset buffer) | 45.2 | 1 |
|
|
| Inline, closed-form dataset (the prototype before this construction) | 5,014 | 111x faster |
|
|
| Inline, memory-hard (recomputes every word, never reads the dataset) | 9.49 | 0.21 (4.8x slower) |
|
|
| Inline, memory-hard, against a 256 MiB honest dataset | 9.48 vs 94.8 | 0.10 |
|
|
|
|
The inline kernel is bound by 104 x 8 = 832 dependent 64-byte cache reads per hash against 104 independent 4-byte reads for the honest kernel. Not measured: the same ratio on NVIDIA or AMD, partial-storage trade-offs between the two points, a smarter attacker kernel (hoisting the `t`-only part of round 0, caching hot lines), and a census of distinct cache lines touched per hash (`MEMHARD.md` section 3, items 1, 2, 4, 5). Analytically a hash touches up to 832 of 4,194,304 lines and a warp's working set is 26,624 lines.
|
|
|
|
Test vectors, day `2026-10-03` (`vectors.json` of the `igneum-genesis-mh` pack):
|
|
|
|
```
|
|
cache line 0 (segment 0, line 0):
|
|
355a86d2 7957db1c d21772af 6fc1e09b d55ce61d 6e6a278b d3f543ce 223d8e82
|
|
143ab337 2e9f05bd 2eb389bf 0c6e449e 5cfa4222 ba6560fe 8e3e1aa4 dbcc1d53
|
|
cache line 4194303 (segment 65535, line 63):
|
|
41190d91 bd277957 22ddbb49 6986f207 df69a4d6 26401a3a 818230fb c417122d
|
|
3597b211 b553ce55 cf39cc0d 3b7fc43a 3fd43b00 67e1c80e ffa7ea7d ca2960ab
|
|
cache fingerprint, FNV-1a 64 over all 2^26 words as little-endian bytes: 48c4f5bf24166b2e
|
|
dataset words 0..15 (item 0):
|
|
ffc3cd94 5920ccd8 392f44bb 5e57f67a 2f2bc2a9 620b0e36 bdc09014 436654bf
|
|
311e0b48 1abd93ad 59cc7ce8 ee5247b2 86171fe8 6d874751 c9f7728f 7c2a435d
|
|
dataset[0x0fffffff] = a33ada72
|
|
dataset[59471966] = e8b73d94 dataset[217795994] = 337028b5
|
|
dataset[3093825] = e3922dca dataset[267580473] = 26b5f1d8
|
|
```
|
|
|
|
The cache fingerprint has been reproduced by: the Swift CPU fill, the Rust fill, the Metal GPU fill, the clang emulation of the CUDA fill, the RTX 5090 (CUDA and NVIDIA OpenCL), Apple OpenCL, pocl, and the AMD gfx1036 (`docs/bench-log.md`, the five entries dated 3 October 2026 for memory-hard, igneum-pow, proto-opencl, gfx1036 and NVIDIA OpenCL).
|
|
|
|
## 1.9 The 32-lane unit of work
|
|
|
|
Definition. The hash is defined over an aligned group of 32 consecutive nonces `g .. g + 31` with `g AND 31 = 0` (wrapping modulo 2^32 at the top of the range). `shfl` exchanges registers within that group: lane `l` reads from lane `l XOR mask`, `mask < 32`, so the exchange never leaves the group. The verification unit is the group: `hash(n)` is computed by evaluating the group `n AND ~31` and taking lane `n AND 31` (`verify.rs`, `Epoch::hash`).
|
|
|
|
Exchange semantics on hardware: the group MUST be realised so that the exchange is exactly `lane XOR mask` over the 32 logical lanes, independent of the hardware wave width. The rule `proto-opencl/WAVEFRONT.md` fixes and `kernel.cl` implements:
|
|
|
|
| Path | When permitted |
|
|
|---|---|
|
|
| Hardware sub-group or warp shuffle (`simd_shuffle_xor`, `__shfl_xor_sync`, `sub_group_shuffle_xor`) | Only when the work-group is exactly 32 items and the device reports a sub-group size of exactly 32 for that kernel and work-group |
|
|
| Local-memory exchange (each lane writes its register to shared memory, one barrier, reads slot `lid XOR mask`) | Always permitted. REQUIRED whenever the sub-group size is not exactly 32 or cannot be queried per kernel |
|
|
|
|
The wave64 rule: on a device whose native wave is 64 lanes (AMD GCN, CDNA, RDNA compiled as wave64; figures approximate per `WAVEFRONT.md`), a wave holds two logical units. A sub-group shuffle would still compute `lane XOR mask` correctly (the emulator rows in `WAVEFRONT.md` show this), but the specification does not permit it because the lane-to-work-item mapping is not guaranteed and because a broadcast would read the wrong half. Such devices MUST use the local-memory exchange. Both paths produced the identical batch fingerprint `f99fb375b3abeaf5` (FNV-1a 64 over the 2^13 outputs at base nonce 0, pack igneum-genesis-mh) on Apple OpenCL, pocl with sub-group shuffles, pocl with local memory, and seven emulator configurations including wave64 (`docs/bench-log.md`, proto-opencl entry); the AMD gfx1036 and the RTX 5090 via OpenCL printed `98af644e993239e2` at 2^24 outputs, identical to each other (gfx1036 and NVIDIA OpenCL entries). Cost of the local-memory path against the warp shuffle on the 5090: about 4%, approximate (NVIDIA OpenCL entry).
|
|
|
|
## 1.10 Target comparison
|
|
|
|
The hash is 64 bits. A block is valid for the lottery when `hash <= target64`. The mapping between the chain's 256-bit difficulty target (rusty-kaspa `Uint256`, fork point a3) and `target64`, and the block-level computation that pruning proofs use (`calc_level_from_pow`, fork map a3, assumes a uniform 256-bit output), are forward references to section 2 and are Open (section 6, O-2.4). Candidate: `target64 = target256 >> 192` and `level = leading_zeros(hash)` capped at 64.
|
|
|
|
## 1.11 CPU verification procedure
|
|
|
|
Implemented (`verify.rs`, `memhard.rs`). A verifier holds the program for the epoch, the mixer parameters and the 256 MiB cache for the day. It never holds the dataset. To verify a block with lane nonce `n`:
|
|
|
|
1. Evaluate the group `g = n AND ~31` with the interpreter of section 1.7, register-major (`r[reg][lane]`) so lane loops vectorise.
|
|
2. At each `load`, gather the 32 masked indices, deduplicate by item (`idx >> 4`), derive the distinct items with all chains interleaved round by round (all mixers for round `r`, then all cache-line XORs for round `r`), and hand each lane its word. Interleaving lets the eight dependent misses of each item overlap across up to 32 items; without it the verifier pays about 8 x 100 ns of DRAM latency per item in series.
|
|
3. Fold and compare lane `n AND 31` against `target64`.
|
|
|
|
The verifier does at most 128 x 32 = 4,096 item derivations per unit, the design bound (design document, Lottery seeds item 3), because every program has exactly 16 loads (1.4.2) and an accepted program reads at least 120 and typically 128 distinct addresses per hash (1.4.6). Under version 1 the count ran from 3,328 to 4,608 items.
|
|
|
|
| Verifier | ms per 32-lane unit, steady (avg of 20) | Worst cold single unit | Source |
|
|
|---|---|---|---|
|
|
| Rust, one M5 Max performance core, generator v2, 128 loads, 4,096 items | 0.631 | 0.67 to 0.81 across the three vector units | Measured 4 October 2026, `igneum-pow/README.md` |
|
|
| Rust, version 1, 104 loads, 3,328 items (retired) | 0.441 | 0.41 to 0.87 across five seeds | Measured, `docs/bench-log.md`, igneum-pow entry |
|
|
| Rust, version 1, 144 loads, 4,608 items (retired) | 0.579 | | same |
|
|
| Swift, 104 loads | 0.649 | 1.16 to 2.11 | Measured, memory-hard dataset entry |
|
|
| Swift, 144 loads | 1.205 | | same |
|
|
| Closed-form dataset (not memory-hard, for scale) | 0.002 (Rust), 0.017 (Swift) | | same entries |
|
|
|
|
The 10 ms gate (Target) is met with a margin of about 16x steady and 12x worst-cold on this core. Not measured: a 2019-class laptop core (design document, "Three experiments before gate 3"), which is what fixes the gate.
|
|
|
|
## 1.12 Schedules: epoch, day, era
|
|
|
|
All times are DAA seconds since genesis (section 0.6). At 1 block per second one DAA second is about one block; the schedules below are written in DAA seconds so that block-rate steps (section 2) do not move them.
|
|
|
|
| Clock | Length | What changes | Label |
|
|
|---|---|---|---|
|
|
| Epoch | `epoch_len(d)` DAA s, base 3,600; the ladder 600, 900, 1,200, 1,800, 2,400, 3,600, 4,800, 7,200; set by 90% miner signal at a day boundary (sections 1.13.1 and 5.7) | The program: new seed words from the VDF of section 4, new kernel | Designed (Counter ASIC 2.0 layer 9, 5 October 2026, `docs/plans/epoch-length.md`); 3,600 stays the value on every network until a signal moves it, and stays the prototype value to be fixed at gate 2 by the difficulty-tracking measurement (fork map c1: the hash rate steps by program, 35 to 48 Mhash/s across seeds on the M5 Max, so the DAA window must track within an epoch) |
|
|
| Day | 86,400 DAA s | The day key, hence the cache and the dataset | Designed |
|
|
| Era | 15,552,000 DAA s (180 days) | Era parameters and one instruction-family unlock, section 1.13 | Designed; the length is a prototype value (the design says "every 6 months") |
|
|
|
|
Epoch `(d, e)` covers DAA scores `[86,400 d + L e, 86,400 d + L (e + 1))` with `L = epoch_len(d)` and `e` in `0 .. 86,400 / L`; every ladder step divides 86,400, so day boundaries are epoch boundaries, and the epoch is identified by its start score `s = 86,400 d + L e`. At the base `L = 3,600` this is `[3,600 e, 3,600 (e + 1))` and nothing below differs from the earlier text. The epoch of a block is the epoch of its own DAA score, and `epoch_len(d)` is a function of the blue blocks of the signalling window that closed at least 2 days before day `d` (section 1.13.1), which are in the header's past, so "which program was this block mined under" is a function of the header alone once the seed is known. `T_epoch` and the 1,200-s lead of section 4.3 are genesis constants and do not follow `epoch_len`: the program of every epoch is known 600 s before it starts on the reference core at every length. The program for epoch `e` is `generate_from_seed_bytes(program_seed_e)`: attempt 0 is drawn from `S_e = seed_words_from_bytes(program_seed_e)`, and a rejected attempt is replaced as 1.4.6 says; `program_seed_e` is the 32-byte VDF output of section 4.3.
|
|
|
|
Implementation note (devnet, 3 October 2026, `docs/fork-divergence.md` "Epoch seed"): until the VDF of section 4 is in the node, `program_seed_e` is the hash of the last selected-chain block whose DAA score is below `3,600 e - 600`. The 600-DAA-score lead stands in for section 4.3's 20-minute lead: the program of epoch `e` is knowable about 10 minutes before it starts, every block template reports it (`pow_epoch.next_epoch_seed`), and a GPU worker compiles it in the background and swaps at the boundary with no pause (serve protocol `prepare`, `proto-metal/main.swift`, `proto-cuda/host.cu`, `proto-opencl/host.c`). Measured across boundaries on a short-epoch test network in `docs/bench-log.md` (hot-swap entry). The program schedule is a protocol constant; a miner that cannot compile ahead sees the same seed at the same time as everyone else, only later.
|
|
|
|
Day `d` covers DAA scores `[86,400 d, 86,400 (d + 1))`. The design document names a day seed and does not say how it is derived. Proposed (Designed, Open, O-1.10): `day_bytes = "igneum-day/" || d_le64 || program_seed of the first epoch of day d`, so the day key is as unpredictable as the epoch seed and known 20 minutes before the day starts (section 4.5), which is enough for a 0.2 s CPU cache fill or a 2 ms GPU one plus a 13 to 30 ms GPU dataset build (Measured, section 1.8.3 and `docs/bench-log.md` RTX 5090 memory-hard entry: 13.4 ms for 1 GiB).
|
|
|
|
Era `n` covers DAA scores `[15,552,000 n, 15,552,000 (n + 1))`.
|
|
|
|
## 1.13 Era parameter draw, instruction-family reserve, dataset growth
|
|
|
|
Designed at the level of a sentence in the design document ("a new instruction mix and memory pattern drawn from chain state, plus one instruction family unlocked from a reserve fixed at genesis"; "the dataset grows on a genesis-fixed schedule"). No draw procedure, reserve list or growth rule exists in code (section 0.3). This section proposes them so that they can be reviewed; everything here is Open (section 6, O-1.11 to O-1.13) until gate 1 fixes it. The rule that is not open: nothing in this section is ever changed by a human release. The draw and the unlock are functions of genesis constants and chain state.
|
|
|
|
### 1.13.1 Era seed and draw
|
|
|
|
The era seed `E_n` is the 32-byte output of the 1-hour VDF of section 4.4. One SplitMix64 stream seeded from `seed_words_from_bytes("igneum-era/" || n_le64 || E_n)` words 0 and 1, drawn in a fixed order, sets the era parameters within genesis-fixed bounds:
|
|
|
|
| Parameter | Base (era 0) | Draw | Bound |
|
|
|---|---|---|---|
|
|
| Op weights for the ten non-load ops | table 1.4.2 | each perturbed by `below(2 * B + 1) - B` points, then renormalised by largest remainder to 100 minus the load weight | B = 2 points, proposed |
|
|
| Load count | 16 of 64 | not drawn | fixed, so every era is equally memory-bound |
|
|
| Output fold rotations | (7, 14, 21), (9, 18, 27) | each `1 + below(31)` | 1..31 |
|
|
| Mixer round count | 8 | not drawn | fixed, so the verify budget holds |
|
|
| Mixer applications per round `mixer_mult` | 8 (class v3, `LoadClass::MX8`, decided x8 at 22:05 UTC on 5 October 2026; 1 under class v2; corrected 6 October 2026, the row had said 4) | not drawn | fixed at genesis (Counter ASIC 2.0, 5 October 2026: the M16 recompute chip's only measured lever; section 1.8.5 carries the form; `docs/plans/mixer-x4.md`) |
|
|
| Epoch length `epoch_len` | 3,600 DAA s | one draw of the era stream consumed and not used (the value is set by signal) | the ladder 600, 900, 1,200, 1,800, 2,400, 3,600, 4,800, 7,200 (layer 9, `docs/plans/epoch-length.md`) |
|
|
|
|
`epoch_len` is the one era-table parameter set by miners rather than by the draw: 90% of blue blocks over a 7-day window carrying the same ladder index (3 bits of the header version, encoding Open in section 5.8) sets that length from the first day boundary at least 2 days after the window closes (section 5.7). It is not a code upgrade: the rule, the ladder and the window are genesis constants, and the chain carries no release. The era stream consumes its draw so that a future draw of this parameter changes no other parameter's value. The threat it answers is a per-program hard datapath (an FPGA fleet: 42 to 160 minutes per compile on a mid-size part, PRflow, FPT 2019, hours on large parts; at 600 s nothing it compiles ever runs); it does not answer a programmable chip, which the other layers answer. The floor 600 is set by the slowest compile-ahead measured (the Metal variant race, 38 s on the M5 Max, 6.3% of a 600-s epoch and inside the 600-s seed window; `docs/plans/epoch-length.md` section 6).
|
|
|
|
The table layout and the working-set window (Counter ASIC 2.0 layers 4 and 8, decided IN on 5 October 2026, delegated: the six-era hash-rate spread is 1.3% on the RTX 5090, 3.2% on the RX 9070 XT and 0.8% on the M5 Max, under the 5% rule; `docs/plans/era-layout.md`) are drawn under program class v3 by a second stream `S` seeded with words 0 and 1 of `seed_words_from_bytes("igneum-era/" || E_n)` (the index is not in the preimage: `E_n` commits to `n` through the VDF input), seven draws in this order whether or not a value is used:
|
|
|
|
1. `W = allowed[below(|allowed|)]`: the width in words of every dataset load of the era, from the genesis-fixed set `allowed`; the set is `{1}` (4 bytes, the read-width decision of 5 October 2026), so the draw is consumed and the width pinned.
|
|
2. `M = low32(next()) OR 1`: the stride multiplier, odd, so `x -> x * M` is a bijection.
|
|
3. `R = 1 + below(31)`: the stride rotation.
|
|
4. to 7. `r_i = next()` for `i` in 0..3: the interleave draws. With `b = log2(W)` and `free = 4 - b`, `c = [b, ..., 15]`; for `i` in `0..free`: `j = i + (r_i mod (16 - b - i))`, swap `c[i]` and `c[j]`; the interleave is `pos = [0, ..., b - 1] ++ sort(c[0..free])`, four ascending bit positions below 16.
|
|
|
|
The era parameters are `(W, M, R, pos)`. Dataset mapping under class v3: word `w` holds word `j(w)` of item `t(w)`, where bit `i` of `j(w)` is bit `pos[i]` of `w` and `t(w)` is `w` with bits `pos[0..3]` removed; with `pos = [0, 1, 2, 3]` this is `dataset[w] = item(w >> 4)[w AND 15]` byte for byte; an item keeps its value at every dataset size of at least 2^16 words, and the `W` words of one aligned load lie in one item, so the 4,096-item verifier bound of 1.11 holds. Load address under class v3, for a load site with window draws `(k_off, o)` and a dataset of `2^D` words: `k = min(k_off, D - 26)`, `y = rotl(x * M, R)`, `idx = ((y AND (MASK >> k)) OR ((o AND (2^k - 1)) << (D - k))) AND MASK` (uniform on the window, branch-free, three operations before the mask), one text form in Metal, CUDA and OpenCL. The window draws per instruction (layer 8), after the nine draws of 1.4.3: `k_off = below(3)` (the dataset, a half or a quarter) and `o = low32(next()) AND (2^k_off - 1)`, used only on a load slot, so a class v3 program takes 720 draws; the window never goes below 2^26 words (256 MiB, above the largest on-chip cache in the benchmark) nor above the dataset, and sixteen sites with drawn offsets cover the dataset with high probability (a windows-union census over 300 programs: the SRAM mirror a chip would need is the whole dataset in every hour). The acceptance rule of 1.4.6 is unchanged in its tests and mirrors this address at its constant `D = 28`. Devnet stand-in for `E_n` until the VDF of 4.4 is in the node: era 0 the genesis block hash; era `n >= 1` the hash of the last selected-chain block whose DAA score is below `15,552,000 n - 7,200`. What the interleave buys and does not: a chip that hard-wires one layout reads the wrong 15 words with every word once the era draws another; a chip whose address decoder can permute its address lines pays nothing (stated in the plan). The stride is a bijection with no cryptanalysis yet (Open).
|
|
|
|
"Memory pattern" in the design document is read here as the item-address pattern (the cache line index word, `s[0]` in 1.8.5, and the XOR-all-sixteen rule); the proposal is to leave it fixed at era 0 and let the unlocked families change the kernel instead, because every change to the item derivation changes the verify time and must be re-measured.
|
|
|
|
### 1.13.2 Instruction-family reserve
|
|
|
|
At genesis the generator carries the eleven families of 1.4.1 live and a reserve list of further families in a fixed order. At the start of era `n >= 1`, reserve family `n` becomes live with weight `W_new` taken proportionally from the live non-load families. A family may enter the reserve only if it is integer-exact and has passed the cross-vendor conformance of section 1.15 on every vendor in the benchmark (Metal, CUDA, OpenCL on NVIDIA and AMD), with its own edge-case vectors, before genesis. Candidate families, all integer ALU operations present on Apple, NVIDIA and AMD: variable left shift and logical right shift by `src AND 31`; bit-field extract with an immediate offset and width; `andn` (`dst = dst AND NOT src`); byte permute of `dst` by an immediate selector; population count and count-leading-zeros folded into `dst` by add; a three-register select (`dst = bit of src2 ? src : dst`); a second shuffle form (`lane + delta mod 32`). The order and `W_new` are Open, except the first entry, decided 5 October 2026 (Counter ASIC 2.0 layer 7, delegated; the project lead confirms for the public testnet genesis; `docs/analysis/int8-matrix-family.md`):
|
|
|
|
> Reserve family R1, `mm8` (integer matrix). Semantics: section 2.2 of `docs/analysis/int8-matrix-family.md`, uint8 operands from `src` and `src2` in the m8n8k16 fragment layout, one int32 element of C per lane selected by the immediate `bit`, added into `dst` modulo 2^32. Weight at unlock `W_new = 4` points, taken proportionally from the ten live non-load families (the load weight and count are untouched). Edge vectors, each a hand-built unit run on every vendor: all bytes 0xFF in A and B (C = 1,040,400 everywhere); all bytes 0x80 (C = 262,144); A all zero (C = 0); `dst` = 0xFFFFFFFF with a nonzero C (the wrap); alternating 0x00 and 0xFF by lane; `bit` = 0 and 1 on the same fragments. Unlock: at the start of era n = 4 (DAA 62,208,000), or earlier by the 90% signalling path of section 5.7; never by a release. Native paths: PTX `mma.sync` `.u8` (sm_75+), AMD WMMA `i32_16x16x16_iu8` (RDNA 3 and 4), Metal 4 `mpp::tensor_ops::matmul2d` (`uchar x uchar -> int`); the per-lane `dot4` form is emulation on Apple (1.6x per op unsigned, measured 5 October 2026) and is not the reserved form.
|
|
|
|
A vendor that can only emulate. A family enters the reserve when it is bit-exact on every vendor of 1.15. A vendor that reaches the result only by emulation (no instruction or library path) does not block entry if the measured penalty of the emulation on that vendor, on the family's own probe (a dependent chain of the op against the same vendor's integer ALU chain), is at most 8x per op, AND the family's weight at unlock keeps the emulating vendor's hash-rate loss under 5% on the memory-hard hash, checked on the vendor's card with the family live. A family whose emulation exceeds either bound stays out of the reserve until the vendor ships a path.
|
|
|
|
A family that is not in the genesis reserve can only be added by the upgrade path of section 5.7.
|
|
|
|
### 1.13.3 Dataset growth
|
|
|
|
Designed: 2 GiB at genesis plus 0.5 GiB per year (design document, "Which cards mine?", labelled approximate there for the card-lifetime consequence, not for the schedule). Proposed rule: the dataset for day `d` has `N_d` items where
|
|
|
|
```
|
|
N_d = floor((2 GiB + 0.5 GiB * (86,400 d / 31,536,000)) / 64 bytes)
|
|
```
|
|
|
|
evaluated in integers (bytes), with one year = 31,536,000 DAA seconds. The dataset grows by about 23 KiB per day on this average and is recomputed with the day key. Decided 5 October 2026 (Counter ASIC 2.0 layer 6, delegated; the project lead confirms for the public testnet genesis): the cache grows with the dataset, doubling when the dataset doubles: `cache_log2_words(d) = 26 + growth_doublings(d)`, `growth_doublings(d) = floor(log2(1 + d / 1,460))` for day `d` since genesis (doublings at years 4, 12, 28 and 60), so the cache is 256 MiB at genesis, 512 MiB from year 4, 1 GiB from year 12; the verifier's one-core fill is 0.2, 0.4 and 0.8 s at those steps (0.2 s per 256 MiB, section 1.12), under 1 s at every step of the schedule. The dataset steps to the next power of two on the same doublings (option (b) below, recommended to the project lead with the card-lifetime consequences in `docs/analysis/card-lifetime-2026-10-05.md`: a 4 GB card mines to year 4, an 8 GB card to year 12, a 12 GB card to year 28 with the cache freed after the daily build). Why the cache grows at all: an SRAM mirror of a flat 256 MiB cache is about 128 mm^2 and $46 of silicon at N5 by shipped cache-die density (AMD V-Cache, 64 MB on 41 mm^2 at 7 nm; `docs/analysis/sram-mirror.md`), so the cache size never prices a chip out; its job is to stay above any GPU's on-die cache (96 MB on the RTX 5090, 128 MB on GB202), which a flat 256 MiB loses within the decade. The GPU frees the cache after the daily dataset build; the hash never reads it. Two consequences remain Open:
|
|
|
|
- Index mapping. `src AND MASK` requires a power-of-two size. For a non-power-of-two `N_d` the proposed mapping is `idx = (src * N_words) >> 32` computed in 64 bits (a multiply-shift range reduction; uniform to within 2^-32, branch-free, integer only). At `N_words = 2^28` this gives `src >> 4`, not `src AND MASK`, so adopting it changes the 1 GiB vectors; gate 1 chooses between (a) the multiply-shift mapping with new vectors, or (b) power-of-two sizes only, growing in steps (2 GiB, 4 GiB) on the same schedule's average, which keeps `AND MASK` and means a 4 GiB card lasts until the 4 GiB step instead of fading.
|
|
- The item index `t` is 32 bits, so the construction as written tops out at 2^32 items = 256 GiB, which the schedule reaches after 508 years. No action needed.
|
|
|
|
## 1.14 Determinism requirements
|
|
|
|
A conforming implementation MUST:
|
|
|
|
1. Use only integer arithmetic. No floating point anywhere, including in index computation and in the mixer (floating point rounds differently per vendor and would split the chain; design document, hostile review table row 2).
|
|
2. Mask or range-reduce every dataset index exactly as 1.5 and 1.13.3 state, and never read outside the dataset. Every load in emitted source MUST have the single form `dataset[rN AND MASK]` (or the adopted range reduction), checkable by text search (`igneum-pow/tests/packs.rs`: 16 of 16 loads masked in every emitted kernel of every pack; `TESTS.md` section 5 for the version 1 run).
|
|
3. Implement `rotr` by `src AND 31` and `rotl` by an immediate in 1..31; a rotate by 0 or 32 through the immediate path is undefined and MUST NOT occur.
|
|
4. Compute `mulhi` as the exact high 32 bits of the 64-bit product (`__umulhi`, `mulhi`, `mul_hi`).
|
|
5. Wrap on overflow everywhere (add, sub, mul, mad, the SplitMix and FNV state).
|
|
6. Realise the exchange over exactly 32 logical lanes as 1.9 requires, independent of the hardware wave width.
|
|
7. Produce the same output for the same (program, day key, dataset size, nonce group) on every run: `TESTS.md` section 4 (5 runs and 3 compiles, one cold, fingerprint `933787e8cfefccb7` closed-form; `62a4f0eb018df273` memory-hard, `MEMHARD.md` section 2.5).
|
|
|
|
## 1.15 Conformance procedure for a miner implementation
|
|
|
|
A miner, kernel emitter or verifier conforms when all of the following pass. Each is a command that exists today; the AMD discrete-card rows are pending. Every vector below is of generator version 2 (4 October 2026); the version 0.1 vectors are retired and MUST NOT be used.
|
|
|
|
1. Cache: fill the 256 MiB cache for day `2026-10-03` and reproduce FNV-1a 64 `48c4f5bf24166b2e`, cache line 0 and cache line 4,194,303 of section 1.8.5, word for word (unchanged by version 2). For the devnet day bytes `"igneum-day/" || 20730_le64`: `448274a57f508cbc`.
|
|
2. Dataset self-test at 1 GiB: words 0..15, word `0x0fffffff`, and the 64 sampled words of `vectors.json` (four are quoted in 1.8.5).
|
|
3. Generator and acceptance: derive the five seeds of the 1.4.6 table to the same attempt and program id, and the `igneum-genesis` instruction list of 1.4.3.
|
|
4. Vectors: all 96 outputs of the three units at base nonces 0, 4,096 and 1,000,000 for pack `igneum-genesis-mh`, standalone (one unit per launch) and in batch (many units per launch, at least two units per work-group or block). Section 1.17 lists them. The same for `igneum-devnet-v4-epoch0`, the chain's own derivation (epoch seed = the devnet genesis hash, day bytes of 2026-10-04).
|
|
5. Batch fingerprint: FNV-1a 64 over the 2^13 outputs at base nonce 0 = `f2a95d5bb84d961e` (igneum-genesis-mh; `8e22ad069cb2a8c3` for igneum-devnet-v4-epoch0), and over the 2^24 outputs = `25f96e7dce90bd4e` (igneum-genesis-mh; `3cc4fbf90fa6366c` devnet).
|
|
6. Fuzz: at least 200 random programs of the current generator (`--fuzz 200` or the Rust equivalent when it exists), 4 units each at base nonces drawn from the full 32-bit range including wraps past 2^32, at 64 MiB, 256 MiB and 1 GiB, zero mismatches against the CPU interpreter, zero compile failures, generator contract asserted on every instruction. Reference: 2,000 of 2,000 under version 2 (`TESTS.md` section 9), 10,200 of 10,200 under version 1.
|
|
7. Edge: the 14 hand-built programs of `TESTS.md` section 2 (rotates by 0 and 31 through the register path, `mulhi` at the extremes, every shuffle mask, loads at index 0 and at MASK through in-range and out-of-range registers, wraparound on add, sub, mul, mad, zero loads, 64 loads), 128 of 128 lanes each. These are hand-built and bypass the generator; they test the interpreter and the kernels, not the rule.
|
|
8. Static mask check on every emitted kernel (1.14 item 2).
|
|
9. Exchange rule: on any device whose sub-group size is not exactly 32, or cannot be queried per kernel, the local-memory path is used and the run says so.
|
|
|
|
Measured conformance to date (`docs/bench-log.md`), version 2 vectors: Apple Metal natively (3 of 3 units on the exported `igneum-genesis-mh`, fuzz 2,000 of 2,000, and the Swift generator identical to the Rust one on the five seeds of 1.4.6), Apple OpenCL (96 of 96 on all four packs, fingerprints as in item 5), the clang CUDA emulation (96 of 96 on all four packs, 2 warps per block), the clang OpenCL emulation (96 of 96 on the two memory-hard packs in sub-group 32 and wave64 configurations), and the Rust CPU reference. Not yet run on version 2 vectors: the RTX 5090 (CUDA and NVIDIA OpenCL), AMD gfx1036, pocl; the version 1 runs on those devices (3 October 2026) stand as evidence that the kernel text, which version 2 did not change, agrees across vendors. A discrete AMD card has not run anything (ledger M8).
|
|
|
|
## 1.16 Parameters marked "prototype value, to be fixed at gate 1" and what fixes them
|
|
|
|
| Parameter | Prototype value | Measurement or decision that fixes it |
|
|
|---|---|---|
|
|
| Seed derivation (FNV-1a plus SplitMix64) | section 1.3 | Replace with a standard hash (ledger M7); re-run the weak-program census and re-cut every vector on the result (the attempt derivation and the program id of 1.4.6 go through the same function) |
|
|
| Instructions per program, iterations | 64 x 8 | CPU verify on a 2019-class laptop core under 10 ms with the memory-hard dataset; register-pressure and occupancy on the three vendors |
|
|
| Registers per lane | 8 | same |
|
|
| Op weights, including the 25% load weight | table 1.4.2 | Fixed 4 October 2026 by the weak-program census (`docs/analysis/weak-program-census-2026-10-03.md`): exact load count 16 (1.4.2), fresh-source rule (1.4.3), acceptance rule (1.4.6). The ten non-load weights remain a prototype value for the era draw bounds of 1.13.1 |
|
|
| Output fold rotations | (7, 14, 21), (9, 18, 27) | Stats run of `TESTS.md` section 3 on the chosen fold |
|
|
| Cache size, lines per segment, ChaCha rounds | 256 MiB, 64, 12 | Shortcut ratio and time-memory curve on an RTX 5090 and an AMD discrete card; external review of the chained-block cache |
|
|
| Item rounds, mixer shape | 8, section 1.8.4 | Same measurement; external review of the mixer; weak-key check on `ROT` |
|
|
| Dataset size at genesis | 1 GiB in packs, 2 GiB designed | Hash-rate and shortcut ratio at 2 GiB on the three vendors; the growth and index-mapping decision of 1.13.3 |
|
|
| Epoch length | 3,600 DAA s | Difficulty tracking across program steps on the devnet (gate 2, fork map c1) |
|
|
| Era length, draw bounds, reserve list, `W_new` | section 1.13 | Design review; each reserve family's own conformance run |
|
|
| Header binding of the init words | section 1.6 | Fixed when section 2 fixes the pre-PoW header hash; new vector set |
|
|
|
|
## 1.17 Test vectors
|
|
|
|
Pack `proto-cuda/packs/igneum-genesis-mh/` (seed `igneum-genesis`, generator version 2, attempt 0, program id `bcc1248b10cc90f2`, day `2026-10-03`, memory-hard, 2^28 words, MASK `0x0fffffff`, 32 lanes). Produced by `igneum-pow export` (the Rust CPU interpreter) on 4 October 2026, reproduced by the Swift CPU interpreter and cross-checked by Metal, Apple OpenCL and the two clang emulations the same day (section 1.15). The version 0.1 vectors of 3 October 2026 (lane 0 `1fb0b3bbc1ac8279`) are retired: they belong to a generator that no longer exists in the protocol. The two closed-form packs (`igneum-genesis`, `igneum-hourly`) are regression vectors for the interpreter only and are not the lottery hash.
|
|
|
|
Unit at base nonce 0, lanes 0..31:
|
|
|
|
```
|
|
42246ba99fc58e4f 19561eec0db4f9f4 11a4fb70ea7b688f 8872bfadec960949
|
|
085e6cf2ff6b8303 780e0d76e504fe6c 7bb05a713f8283fa fef7beeffd5ab1d1
|
|
1b83afb12e78d65b 2576c5384a2e1cae 549d620a0735d6d8 2eb4e32f32e8d5f1
|
|
81bb414492f69586 092d17324b465e01 8bde350ee2354b5a d64b47e9c9e0ec07
|
|
dbbf1b78b9979a3f bee382abde89c111 6b598d99c16a8c70 69d72da59fb3b6e9
|
|
1e309d2a549632fa 98fa8255ab65f005 2f48ab1bb516110c 7d2af17cadd18bea
|
|
547c5978bd005e02 25ea2e21e88ea9d3 1c575ec6e43efc58 077b80cb079958b4
|
|
a96eebd0c8634981 44b18bdf19eb7838 c3ccb9fe9ef5953c b08446b1f2de7793
|
|
```
|
|
|
|
Unit at base nonce 4,096: lane 0 `3d3903e310ca038f`, lane 1 `9e72b9a86ebb29e4`, lane 31 `61c242509efdccdd`. Unit at base nonce 1,000,000: lane 0 `f218c1bd58e6dfe0`, lane 1 `8c5a362ee98971c2`, lane 31 `6c3b2c11adfbfcac`. The remaining 58 values are in `vectors.json` and `vectors.h` of the pack.
|
|
|
|
Pack `proto-cuda/packs/igneum-devnet-v4-epoch0/`: epoch seed bytes `edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07` (the devnet genesis hash), day bytes `69676e65756d2d6461792ffa50000000000000` (`"igneum-day/" || 20730_le64`, 2026-10-04), attempt 0, program id `4be132dd1f2ff270`, cache FNV-1a 64 `448274a57f508cbc`, dataset words 0 and 1 `3dd50b1f 48edec90`, word `[MASK]` `f7b7180e`; unit at base nonce 0: lane 0 `285a83011e7ac3fc`, lane 31 `6f1136558c20e3d8`.
|
|
|
|
Header-bound vectors (section 1.6 rule, seed `igneum-genesis`, day `2026-10-03`): `igneum-pow/README.md`, eight values, for example H = 32 zero bytes and nonce 0 give `746c567b090acf6a`.
|
|
|
|
Program class v3 vectors (5 October 2026, `proto-cuda/packs-ca2-mixer/` and `proto-cuda/packs-ca2-era/`; generator 3, mixer x8, the cache growth rule, the era draw inside the class): pack `mx8-genesis` (seed `igneum-genesis`, no era, program id `e323b9dcaf283a6f`, batch fingerprint `7c28cfb06c5c65a9`) and pack `mx8-devnet-epoch0` (the devnet epoch seed and day of 1.17 with the era stand-in E_0 = the devnet genesis hash inside the class, program id `73bcbfe8ccf988f1`, unit 0 lane 0 `d424577fce4a7a60`, batch fingerprint `90f794dd556f7a3b` over 2^24 outputs at base nonce 0), reproduced by the Rust interpreter, Metal and Apple OpenCL on 5 October 2026 (3/3 standalone, 3/3 in batch, 96 of 96 lanes each); the six era packs `era-0` to `era-5` (the same epoch seed and day, era test seeds 0 to 5, program id `73bcbfe8ccf988f1`; 2^24 fingerprints `8e8e070db4eea52d`, `891c01b8563bb47e`, `e54279fed2831b5d`, `77e0ba8abbd0ae62`, `d898d8f4f2e7684b`, `a6927db380f7efb2`). All seven fingerprints are equal on the RTX 5090 (CUDA/NVRTC), the RX 9070 XT (AMD OpenCL) and the M5 Max (Metal), self-test PASS on every pack, and 1,024 random nonces per card on `era-0` and `mx8-devnet-epoch0` re-hash to the same value on the Rust verifier (1,024 of 1,024 each): job `run-ca2-era-pc1b-20261005`, 5 October 2026, `docs/bench-log.md`. The class v2 vectors above stand unchanged (the v2 exports are byte-identical on the class v3 crate, `igneum-pow/tests/packs.rs`).
|
|
|
|
Cache, dataset and mixer vectors: section 1.8.4 and 1.8.5. Seed words: section 1.3.1. Generator: section 1.4.3. Acceptance: section 1.4.6. Batch fingerprints: section 1.15 item 5.
|