From f69ec95f8fc5cd71b04d0563c9be2a29ed6d918c Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Sat, 3 Oct 2026 17:45:08 +0000 Subject: [PATCH] Spec 0.1: lottery hash, consensus delta, finality v2 with the quorum floor, seeds and VDF, fees, open items docs/spec/01 to 06 and the README index, alongside the existing overview. Section 1 is the normative definition of the lottery hash with test vectors copied from the igneum-genesis-mh pack (cache fingerprint 48c4f5bf24166b2e, 96 hashes, mixer constants) and every prototype value marked with the measurement that fixes it at gate 1. Sections 2 to 5 are the design as decided on 3 October 2026, labelled Designed. Section 6 lists 60 open items with the experiment or decision that closes each and its gate. Co-Authored-By: Claude Fable 5.1 --- docs/spec/00-overview.md | 78 +++++ docs/spec/01-lottery-hash.md | 466 +++++++++++++++++++++++++++++ docs/spec/02-consensus.md | 104 +++++++ docs/spec/03-finality.md | 120 ++++++++ docs/spec/04-seeds-and-vdf.md | 106 +++++++ docs/spec/05-fees-and-economics.md | 84 ++++++ docs/spec/06-open-items.md | 105 +++++++ docs/spec/README.md | 17 ++ 8 files changed, 1080 insertions(+) create mode 100644 docs/spec/00-overview.md create mode 100644 docs/spec/01-lottery-hash.md create mode 100644 docs/spec/02-consensus.md create mode 100644 docs/spec/03-finality.md create mode 100644 docs/spec/04-seeds-and-vdf.md create mode 100644 docs/spec/05-fees-and-economics.md create mode 100644 docs/spec/06-open-items.md create mode 100644 docs/spec/README.md diff --git a/docs/spec/00-overview.md b/docs/spec/00-overview.md new file mode 100644 index 00000000..dd29178e --- /dev/null +++ b/docs/spec/00-overview.md @@ -0,0 +1,78 @@ +# Igneum protocol specification, section 0: overview + +Spec version 0.1, 3 October 2026. Status of this section: Designed. + +## 0.1 Scope + +This specification defines the Igneum layer 1 so that a stranger can reimplement a node, a miner and a verifier from it and reach the same bits. It covers: + +| Section | File | What it fixes | +|---|---|---| +| 1 | `01-lottery-hash.md` | The random-program GPU lottery hash, its memory-hard dataset, the CPU verifier, the epoch, day and era schedules, test vectors and conformance | +| 2 | `02-consensus.md` | The ordering layer as a delta on rusty-kaspa: block rate, GHOSTDAG parameters, difficulty, emission, header changes, duplicate inclusion | +| 3 | `03-finality.md` | Sustained-mining finality, version 2, with the quorum floor | +| 4 | `04-seeds-and-vdf.md` | The epoch and era seed pipeline through the class-group VDF | +| 5 | `05-fees-and-economics.md` | Fees, splits, burns, the development fund, signalling thresholds | +| 6 | `06-open-items.md` | Every parameter or rule that is unmeasured, unreviewed or marked open, with the experiment that closes it | + +Out of scope for version 0.1: the zkEVM execution layer, the chunked proving protocol, the proof format carried in headers, the external job market, the P2P wire format, the RPC surface, the miner client and pool protocol. Where a later section depends on one of these, the dependency is named as a forward reference and the behaviour the dependency must provide is stated. + +## 0.2 Normative language + +MUST, MUST NOT, SHOULD, SHOULD NOT and MAY are used as in RFC 2119. A sentence without one of these words is descriptive and carries no conformance weight. + +Numbers come in four kinds and every number in this specification is labelled as one of them, in the sentence or in the table column that holds it: + +| Label | Meaning | Example | +|---|---|---| +| Measured | Produced on a named machine on a named date by a named command. The bench-log entry is cited | 0.441 ms per warp CPU verify, `docs/bench-log.md`, entry "igneum-pow: Rust crate bit-exact with proto-metal" | +| Implemented | Fixed by code that exists and has test vectors, whether or not the value is final | 64 instructions, 8 iterations, `igneum-pow/src/generator.rs` | +| Designed | Fixed by a decision in the design document and not yet in code | Hard cap 4 billion, halving every two years | +| Target | A number the design aims at and has not measured | A chip gains under 2x over a GPU | + +"Prototype value, to be fixed at gate 1" marks a parameter that the implementation carries today and that a named measurement will replace or confirm before the lottery hash is frozen. Every such parameter is listed at the end of section 1 and in section 6. + +## 0.3 What is implemented and measured, what is designed + +| Component | State on 3 October 2026 | Where | +|---|---|---| +| Lottery hash: generator, interpreter, memory-hard dataset, kernel emitters (Metal, CUDA, OpenCL) | Implemented in Rust (`igneum-pow`) and Swift (`proto-metal`), bit-exact with each other and with four GPU compiler paths on three vendors | `igneum-pow/`, `proto-metal/`, `proto-cuda/packs/`, `proto-opencl/` | +| CPU verification under 10 ms per warp | Measured: 0.41 to 0.58 ms steady, 0.87 ms worst cold, one M5 Max core | `docs/bench-log.md`, igneum-pow entry | +| Memory hardness of the dataset | Measured on Apple only: recomputing items is 4.8x slower than loading them. Not measured on NVIDIA or AMD | `proto-metal/MEMHARD.md` section 2.2 | +| Class-group VDF for the epoch seed | Implemented as a prototype, measured on one core; not reviewed by a second cryptographer | `proto-vdf/` | +| Finality rule version 2 | Designed, simulated with latency, partitions and eclipses (no DAG in the model). Not implemented | `sim/results_v2.md`, design document section "Finality rule, version 2" | +| rusty-kaspa fork | Base built and a 3-node devnet run at Kaspa's own parameters. No fork point implemented | `docs/fork-map.md`, bench-log devnet entry | +| Emission, fees, signalling | Designed | Design document | +| Era draw, instruction-family reserve, dataset growth | Designed at the level of a sentence. No draw procedure or reserve list exists | Design document, section 1.13 here | +| zkEVM, proving protocol, job market | Designed at the level of the design document. Nothing measured | Out of scope | + +## 0.4 Versioning of the specification + +The specification carries a version of the form MAJOR.MINOR. + +- MINOR increments when text is clarified, a test vector is added, or a parameter marked "prototype value" is fixed without changing any existing test vector. +- MAJOR increments when any consensus-relevant behaviour changes: a test vector in section 1 changes, a rule in sections 2 to 5 changes, or a parameter that was Designed becomes Implemented with a different value. +- Version 1.0 is the version frozen at gate 1 for section 1 and at gate 3 for section 3. Until 1.0 any section may change at any MAJOR step. + +Each section file carries its own status line (Measured, Designed or Open) and the index `README.md` collects them. The git history of `docs/spec/` is the change log; no separate change log is kept. + +Consensus rules in a running chain change only through the upgrade path of section 5.7: new code activates when 90% of blocks in a signalling window carry the signal. The specification is updated to describe the activated rules, not the other way round. + +## 0.5 How to submit a break + +A break is a reimplementation that disagrees with a test vector, an attack that beats a stated bound, a measurement that contradicts a Measured number, or an argument that a Designed rule fails its stated property. + +1. Check `docs/fud-ledger.md`. It holds every criticism received so far with its status. If the break is already there as Open, the entry names the experiment that will settle it and you can go straight to that. +2. If it is not there, write it the way the ledger writes entries: the claim in one paragraph, the evidence (a command and its output, a proof sketch, a reference), and the section and rule it breaks. +3. Submit it by the route named in the ledger's "Submitting a criticism" paragraph. On 3 October 2026 the repository is private and the public route does not yet exist; the ledger says so and the project has committed to opening one before the litepaper is shared (ledger entry X7). +4. Every submission that is not already in the ledger is added to it with credit if wanted, including submissions that turn out to be wrong, with the reason. + +A break that reproduces is a MAJOR version change here and a status change in the ledger. Nothing in the ledger is deleted. + +## 0.6 Conventions used in every section + +- All arithmetic in section 1 is on unsigned 32-bit integers modulo 2^32 unless stated. There is no floating point anywhere in consensus. +- Byte order for hashing memory is little-endian: a 32-bit word w is hashed as the four bytes `w & 0xff, (w >> 8) & 0xff, (w >> 16) & 0xff, w >> 24`. +- Time in sections 2 to 5 is DAA time: difficulty-adjusted seconds derived from DAA score, never wall-clock. One block per second at launch, so one DAA second is about one block. +- Sizes: KiB, MiB, GiB are powers of two. GB in a quoted hardware figure means what the vendor meant. +- The chain's hash function for everything that is not the lottery hash is the one rusty-kaspa uses at the forked commit (BLAKE2b-based `Hash` in `crypto/hashes`), until section 2 says otherwise. diff --git a/docs/spec/01-lottery-hash.md b/docs/spec/01-lottery-hash.md new file mode 100644 index 00000000..ffab7670 --- /dev/null +++ b/docs/spec/01-lottery-hash.md @@ -0,0 +1,466 @@ +# Igneum protocol specification, section 1: the lottery hash + +Spec version 0.1, 3 October 2026. Status of this section: Measured for the construction as implemented (Implemented values with test vectors on three GPU vendors and a CPU reference); Designed for the header binding, the day key, the dataset growth and the era schedule; Open where marked. + +Normative implementation: `igneum-pow/src/{seed,generator,memhard,verify}.rs`. Where this text and that code disagree, the code and its test vectors win until this text is corrected (section 0.4). The Swift prototype `proto-metal/main.swift` is bit-exact with the crate on every pack (`docs/bench-log.md`, entry "igneum-pow: Rust crate bit-exact with proto-metal"). + +Every parameter marked "prototype value, to be fixed at gate 1" is carried by the implementation today, is part of the test vectors, and is confirmed or replaced by the named measurement before the hash is frozen (section 1.16). + +## 1.1 What the hash is for + +The lottery hash decides who produces the next block. It MUST be: + +1. Deterministic and bit-exact on every conforming implementation: GPU kernels on any vendor, the CPU verifier, and any future implementation (section 1.14). +2. Cheap to verify on one CPU core without the dataset: one 32-lane unit of work in under 10 ms (Target, Measured at 0.41 to 0.58 ms, section 1.11). +3. Bound by random access to a dataset larger than any on-chip cache, so that computing dataset words is slower than loading them (Measured 4.8x slower on Apple, section 1.8.4; not measured on NVIDIA or AMD). +4. Unknowable until shortly before it is needed, so that a miner cannot grind the seed (section 4). + +It does not need to be a general-purpose cryptographic hash (preimage, collision). It needs to be a fair lottery: no shortcut cheaper than honest evaluation and no bias a miner can exploit. No analysis of either property exists yet (`docs/fud-ledger.md`, entries M6 and M7; section 6). + +## 1.2 Notation + +All arithmetic is on unsigned 32-bit integers modulo 2^32 unless a 64-bit type is named. `rotl(x, n)` and `rotr(x, n)` rotate by `n` in 0..31. `mulhi(a, b)` is bits 32..63 of the 64-bit product. `low32(v)` is bits 0..31 of a 64-bit `v`. `||` is byte concatenation. Indices are zero based. Words are hashed into FNV as little-endian bytes (section 0.6). + +## 1.3 Seed words and the SplitMix64 stream + +Implemented, prototype value, to be fixed at gate 1 (ledger M7 asks for a standard hash in place of FNV-1a plus SplitMix so the seed-to-program mapping is auditable; the measurement that fixes it is the weak-program census of section 1.16, which must be re-run on whichever derivation is chosen). + +### 1.3.1 `seed_words_from_bytes(b) -> [u32; 8]` + +For `salt` in 0..3: + +``` +basis = 0xcbf29ce484222325 XOR (salt * 0x9E3779B97F4A7C15) (64-bit, wrapping) +h = basis +for each byte x of b: h = (h XOR x) * 0x100000001b3 (64-bit, wrapping) +h = h XOR (h >> 33); h = h * 0xff51afd7ed558ccd; h = h XOR (h >> 33) +words[2 * salt] = low32(h) +words[2 * salt + 1] = h >> 32 +``` + +`seed_words(s)` is `seed_words_from_bytes` over the UTF-8 bytes of the string `s`. This function is the boundary at which the chain's bytes enter the hash: the epoch program seed from section 4 and the day key bytes (section 1.12) both go through it. In the prototype the input is a string ("igneum-genesis", "day/2026-10-03"). + +Test vector: `seed_words("igneum-genesis")` = `67a9a7be 1a155b25 fddfb732 4b5af2e8 c55caf33 a27c13b7 06628a48 03852469` (`proto-cuda/packs/igneum-genesis-mh/program.json`, `seed_words`). + +### 1.3.2 SplitMix64 + +State `s` (64-bit). `next()`: + +``` +s = s + 0x9E3779B97F4A7C15 +z = s +z = (z XOR (z >> 30)) * 0xBF58476D1CE4E5B9 +z = (z XOR (z >> 27)) * 0x94D049BB133111EB +return z XOR (z >> 31) +``` + +`below(n)` = `next() mod n` (modulo, not rejection sampling; the bias at n <= 100 is under 2^-57 and is part of the definition). + +### 1.3.3 The program stream + +For seed words `w[0..7]`: `lo = w[0] | (w[1] << 32)`, `hi = w[2] | (w[3] << 32)`, and the generator's SplitMix64 state starts at `lo XOR (hi * 0x9E3779B97F4A7C15)` (64-bit wrapping multiply). Words `w[4..7]` are not used by the generator; all eight are used by the register initialisation (section 1.6). + +## 1.4 The generator + +Implemented (`igneum-pow/src/generator.rs`). A program is a list of `INSTR_COUNT` instructions over 8 lane registers `r0..r7`, executed `ITERATIONS` times per hash. + +| Parameter | Value | Label | +|---|---|---| +| Registers per lane | 8 x u32 | prototype value, to be fixed at gate 1 (fixed by the ASIC-gain target and the register-pressure measurement of 1.16) | +| Instructions per program | 64 | prototype value, to be fixed at gate 1 (fixed by the CPU-verify measurement on a 2019-class core and the load-count rule) | +| Iterations per hash | 8 | prototype value, to be fixed at gate 1 (same measurement) | +| Lanes per unit of work | 32 | Definition. Not a tuning parameter (section 1.9) | +| Shuffle masks | {1, 2, 4, 8, 16} | Definition, follows from 32 lanes | +| Rotate immediates | 1..31 | Definition. 0 is never emitted and MUST NOT be relied on (`proto-metal/TESTS.md` section 2) | + +### 1.4.1 Instruction set + +Eleven families. `dst`, `src`, `src2` name registers; `src != dst` always; `src2` may equal either. + +| Op | Semantics (per lane) | Operands used | +|---|---|---| +| `add` | `dst = dst + src + (bit `bit` of sel ? imm2 : imm)` | dst, src, imm, imm2, bit | +| `sub` | `dst = dst - src` | dst, src | +| `mul` | `dst = low32(dst * src)` | dst, src | +| `mulhi` | `dst = mulhi(dst, src)` | dst, src | +| `xor` | `dst = dst XOR src` | dst, src | +| `or` | `dst = dst OR src` | dst, src | +| `rotl` | `dst = rotl(dst, rot)`, rot in 1..31 | dst, rot | +| `rotr` | `dst = rotr(dst, src AND 31)` | dst, src | +| `mad` | `dst = low32(src * src2) + dst` | dst, src, src2 | +| `shfl` | `dst = dst XOR src_of_lane(lane XOR mask)` | dst, src, mask | +| `load` | `dst = dst XOR dataset[src AND MASK]` | dst, src | + +`sel` is the value of `r0` sampled once at the top of each iteration, before instruction 0, and held for all 64 instructions of that iteration. This is the per-hash nonce-dependent select: the immediates an `add` uses depend on the lane's own state, so no two nonces run the same constant sequence (ProgPoW-style data-dependent path, integer only). + +A twelfth family `wload` (warp-coalesced 128-byte load) exists in the code as lever (b) and is never emitted at the default configuration (`wide_frac = 0`). It is NOT part of the lottery hash. It is retained only so the measurement in `proto-metal/MEMHARD.md` section 2.4 stays reproducible, and the recommendation there is not to adopt it. + +### 1.4.2 Op weights + +Implemented, prototype value, to be fixed at gate 1. Weights sum to 100 and are applied in this order: + +| Op | Weight | +|---|---| +| load | 25 | +| add | 12 | +| xor | 10 | +| mul | 8 | +| mad | 8 | +| shfl | 8 | +| rotl | 7 | +| sub | 6 | +| mulhi | 6 | +| rotr | 6 | +| or | 4 | + +What fixes them: (1) the load count per program is not fixed today, only its expectation (16 loads per 64 instructions, 128 per hash); over 10,000 programs it ranged 40 to 232 loads per hash (`proto-metal/TESTS.md` section 1) and the GPU rate scales with it (104 loads: 228 Mhash/s, 128 loads: 185 Mhash/s on the RTX 5090, `docs/bench-log.md` RTX 5090 sweep entry). Ledger M5. The gate 1 rule MUST fix the load count per program exactly (candidate: draw exactly 16 load positions, then draw the other 48 ops from the remaining weights), so every epoch is equally memory-bound and difficulty does not step on the hour. (2) The lever (a) measurement (`load_weight = 17`, `MEMHARD.md` section 2.4) is the fallback if a slower verifier core ever threatens the 10 ms gate: it cut CPU time to 0.46 to 0.51 ms per warp on the Swift verifier and left the kernel memory bound. + +### 1.4.3 Draw order + +For each of the 64 instructions, in order, draw from the program stream of 1.3.3 exactly these values in exactly this order, whether or not the op uses them: + +``` +roll = below(100); op = first entry of the weight table whose cumulative weight exceeds roll +dst = below(8) +a = below(7); if a >= dst then a = a + 1 (src, never equal to dst) +b = below(8) (src2) +imm = low32(next()) +imm2 = low32(next()) +rot = 1 + below(31) +bit = below(32) +mask = 1 << below(5) +``` + +Nine draws per instruction, 576 per program. A program is fully determined by its eight seed words. + +Test vector: for seed `igneum-genesis` the first eight instructions are (`proto-cuda/packs/igneum-genesis-mh/program.json`): + +``` +0: rotl dst=4 src=2 src2=7 imm=0x20699878 imm2=0x6f1a6170 rot=25 bit=7 mask=16 +1: sub dst=0 src=5 src2=4 imm=0xf81a0b9d imm2=0xf0505e88 rot=1 bit=4 mask=1 +2: load dst=4 src=3 src2=6 imm=0xb1978a0b imm2=0x2ca4e162 rot=10 bit=21 mask=1 +3: rotl dst=1 src=6 src2=7 imm=0xc3bd2355 imm2=0xa8c5f27e rot=1 bit=22 mask=1 +4: add dst=2 src=3 src2=0 imm=0x61f0b51c imm2=0x2735a174 rot=4 bit=26 mask=2 +5: load dst=5 src=3 src2=3 imm=0x4d183796 imm2=0x679648a8 rot=4 bit=30 mask=4 +6: sub dst=5 src=7 src2=7 imm=0x265677dc imm2=0x9043323e rot=30 bit=7 mask=4 +7: add dst=3 src=4 src2=6 imm=0x52f2dbf4 imm2=0x5a069596 rot=5 bit=26 mask=2 +``` + +The whole program has op mix `load=13 xor=13 sub=7 shfl=6 add=5 mulhi=5 mad=4 rotr=4 mul=3 rotl=3 or=1`, 104 loads per hash. + +### 1.4.4 Generator contract + +Every emitted instruction satisfies: `rot` in 1..31, `mask` in {1, 2, 4, 8, 16}, `src != dst`. This held on every instruction of 10,200 fuzzed programs (`TESTS.md` section 1). A kernel emitter MAY rely on it; an interpreter MUST NOT accept a program that violates it. + +### 1.4.5 Encoding + +A program is transmitted as the seed words, never as instructions. A node hands a miner the pack it emits itself (`igneum-pow/src/emit.rs`: `kernel.cu`, `kernel.cl`, `program.metal`, `program.h`, `program.json`, `memhard.h`, `vectors.*`), and a miner MAY regenerate everything from the seed. `program.json` is the interchange form; its field names are those of `Instr` in `generator.rs`. + +## 1.5 Fixed memory footprint + +Implemented for the prototype size; Designed for genesis. + +| Quantity | Prototype (packs, vectors) | Genesis | Label | +|---|---|---|---| +| Dataset words | 2^28 (1 GiB) | 2^29 (2 GiB) | Prototype: Implemented. Genesis: Designed (design document, "Which cards mine?") | +| MASK | 0x0fffffff | 0x1fffffff | as above | +| Dataset item | 16 words (64 bytes) | same | Implemented | +| Cache | 2^26 words (256 MiB) | same | Implemented, prototype value, to be fixed at gate 1 (fixed by the shortcut-ratio measurement on the RTX 5090 and an AMD discrete card: the cache must exceed the largest on-chip cache of any card that mines, and 96 MiB of L2 on the 5090 is the figure to beat) | + +A load reads one 4-byte word at `src AND MASK`. Every load in every emitted kernel has exactly this form; the static check in `TESTS.md` section 5 is part of conformance (section 1.15). Because item values do not depend on the dataset size (section 1.8.5), the 1 GiB vectors remain valid for words below 2^28 at any larger size. + +Growth beyond genesis is in section 1.13. + +## 1.6 Register initialisation + +Implemented (`verify.rs`, `interpret_warp`). For lane nonce `n` (32-bit) and init words `I[0..7]`: + +``` +splitmix32(x): x ^= x >> 16; x *= 0x7feb352d; x ^= x >> 15; x *= 0x846ca68b; x ^= x >> 16 +for i in 0..7: + x = n XOR I[i] + x = x + 0x9e3779b9 * (i + 1) + x = splitmix32(x) + r[i] = x XOR I[(i + 1) AND 7] +``` + +In the prototype and in every pack, `I` is the program's own seed words (the same eight words that drive the generator). On the chain the hash MUST also commit to the block being mined, which the prototype does not do. The binding is Designed, proposed here, Open until gate 1 (item O-1.9 in section 6): + +- The header nonce is 64 bits (rusty-kaspa `Header.nonce`). Its low 32 bits are the lane nonce `n`. Its high 32 bits and the 256-bit pre-PoW header hash `H` (section 2, fork point a5) form the init words: `I = seed_words_from_bytes("igneum-block/" || H || nonce_hi_le32)`. +- `I` is a kernel argument, not a compile-time constant. The program (from the epoch seed) is compiled once per epoch; `I` changes per block template. +- The packs' vectors, where `I` equals the program seed, remain the conformance vectors for the generator, the interpreter and the dataset. A second vector set with `I` derived from a header is produced when section 2 fixes the header hash. + +## 1.7 Execution of one hash + +``` +r = init(n, I) +repeat ITERATIONS (8) times: + sel = r0 + for each instruction in order: apply it (section 1.4.1) +lo = r0 XOR rotl(r1, 7) XOR rotl(r2, 14) XOR rotl(r3, 21) +hi = r4 XOR rotl(r5, 9) XOR rotl(r6, 18) XOR rotl(r7, 27) +hash = (hi << 32) | lo (64 bits) +``` + +The output folding rotations (7, 14, 21; 9, 18, 27) are Implemented, prototype value, to be fixed at gate 1 (the stats run of `TESTS.md` section 3 is the check; a different fold must pass the same run). + +`shfl` makes the 32 lanes of a unit interdependent: the hash of one nonce is defined only as a member of its aligned group of 32 (section 1.9). + +## 1.8 The memory-hard dataset + +Implemented (`memhard.rs`), construction and measurements in `proto-metal/MEMHARD.md`. Every constant below is a prototype value, to be fixed at gate 1, unless marked Definition. What fixes them is the shortcut-ratio and time-memory trade-off measurement on NVIDIA and AMD discrete cards (section 1.16 items 2 and 3) and an external review of the primitives (ledger M7). + +### 1.8.1 Day key + +`K[0..7] = seed_words_from_bytes(day_bytes)`. In the prototype `day_bytes` is the UTF-8 of `"day/" + day` with `day` an ISO date; on the chain see section 1.12. For `day/2026-10-03`: `K = 3067619f 3c269176 84a03b03 f8c63294 ff977c5b e60def3e 63630141 b8fbcb58`. + +### 1.8.2 Block function B (ChaCha12 with feed-forward) + +Standard ChaCha quarter round with rotations (16, 12, 8, 7), six double rounds (columns then diagonals, on the 16-word state), then `y[i] = y[i] + x[i]`. Twelve rounds, no key schedule beyond the input block. `CHACHA_ROUNDS = 12` is a prototype value; `sigma = (0x61707865, 0x3320646e, 0x79622d32, 0x6b206574)` is a Definition. + +### 1.8.3 Cache fill + +The cache is 2^26 words = 2^22 lines of 16 words, in 2^16 segments of 64 lines. Segment `s`, line `j` is at word offset `(s * 64 + j) * 16`. Each segment is a chain: + +``` +tag = (0x49676e65, 0x756d4d48) "Igne", "umMH" +prev = 0^16 +for j in 0..63: + in = prev XOR (sigma[0..3] || K[0..7] || s || j || tag[0..1]) 16 words + line = B(in) + cache[segment s, line j] = line + prev = line +``` + +Line `j` costs `j + 1` block evaluations from nothing, 32.5 on average. The 65,536 segments are independent (one GPU thread each). The fill is a once-per-day cost. + +| Fill time | Value | Source | +|---|---|---| +| One M5 Max core, Rust | 175 to 181 ms | Measured, `docs/bench-log.md`, igneum-pow entry | +| One M5 Max core, Swift | 184.5 to 190.6 ms | Measured, same log, memory-hard dataset entry | +| RTX 5090, one host thread | 223 ms | Measured, same log, "RTX 5090, memory-hard dataset" | +| M5 Max GPU | 0.6 to 2.1 ms | Measured, memory-hard dataset entry (variance not isolated) | +| RTX 5090 GPU | 0.67 ms | Measured, "RTX 5090, memory-hard dataset" | + +### 1.8.4 Mixer parameters and the mixer M_r + +One SplitMix64 stream seeded with `K[0] | (K[1] << 32)`, drawn in this order: `ROT[0..7] = 1 + below(31)` (eight draws), `MUL[0..15] = low32(next()) OR 1` (sixteen draws, always odd so each multiply is a bijection), `RC[0..15] = low32(next())` (sixteen draws). + +`M_r(s)` on a 16-word state with round key `rk = (r + 1) * 0x9E3779B9`: + +``` +for i in 0..15: s[i] = (s[i] XOR (RC[i] + rk)) * MUL[i] +QR(s0, s4, s8, s12; ROT0..3) QR(s1, s5, s9, s13; ROT0..3) QR(s2, s6, s10, s14; ROT0..3) QR(s3, s7, s11, s15; ROT0..3) +QR(s0, s5, s10, s15; ROT4..7) QR(s1, s6, s11, s12; ROT4..7) QR(s2, s7, s8, s13; ROT4..7) QR(s3, s4, s9, s14; ROT4..7) +``` + +where `QR(a, b, c, d; r1, r2, r3, r4)` is the ChaCha quarter round with those four rotations. About 130 integer operations. This is the prototype's stand-in for RandomX's SuperscalarHash: fixed shape, seed-drawn constants. It has had no cryptanalysis (`MEMHARD.md` section 3, unproven item 3; a `ROT` draw of eight equal values is possible and untested). + +Test vector, day `2026-10-03` (`proto-cuda/packs/igneum-genesis-mh/program.h`): + +``` +ROT = 20 20 19 4 26 3 3 27 +MUL = 42146205 52cbe0fb 7ecf4a03 6728907f d81d9751 132952c3 f60de277 05358035 + baf6499d e4db9667 3e98f45d d0004edd 2691630d 9beb3bcf ab310379 99cfb423 +RC = bab68293 cc162340 6ce151cc e62b8997 c9c80297 f74a1654 3d704af5 3cf522b7 + 2b9cac04 a880ac10 13e5dd1d 6fc3e233 2d83eeac 9006e8bf 2c4b5362 31b49ee2 +``` + +### 1.8.5 Item derivation and dataset mapping + +Item `t` (16 words): + +``` +s[i] = K[i] for i in 0..7 +s[8 + i] = t * MUL[i] + RC[i] for i in 0..7 +for r in 0..7: + s = M_r(s) + a = s[0] AND 0x003fffff cache line index, 2^22 lines + s[i] = s[i] XOR cache[line a][i] for i in 0..15 +s = M_8(s) +item(t) = s +``` + +Eight dependent cache reads (`ITEM_ROUNDS = 8`, prototype value): the address of read `r` depends on every earlier read. Nine mixer applications. `dataset[w] = item(w >> 4)[w AND 15]`. A dataset of 2^D words is the prefix of items `0 .. 2^(D-4) - 1`, so an item has the same value at every dataset size. + +What the construction buys (Measured, `MEMHARD.md` section 2.2, M5 Max, seed igneum-genesis, 1 GiB): + +| Kernel | Mhash/s | Ratio to honest | +|---|---|---| +| Honest (loads from the dataset buffer) | 45.2 | 1 | +| Inline, closed-form dataset (the prototype before this construction) | 5,014 | 111x faster | +| Inline, memory-hard (recomputes every word, never reads the dataset) | 9.49 | 0.21 (4.8x slower) | +| Inline, memory-hard, against a 256 MiB honest dataset | 9.48 vs 94.8 | 0.10 | + +The inline kernel is bound by 104 x 8 = 832 dependent 64-byte cache reads per hash against 104 independent 4-byte reads for the honest kernel. Not measured: the same ratio on NVIDIA or AMD, partial-storage trade-offs between the two points, a smarter attacker kernel (hoisting the `t`-only part of round 0, caching hot lines), and a census of distinct cache lines touched per hash (`MEMHARD.md` section 3, items 1, 2, 4, 5). Analytically a hash touches up to 832 of 4,194,304 lines and a warp's working set is 26,624 lines. + +Test vectors, day `2026-10-03` (`vectors.json` of the `igneum-genesis-mh` pack): + +``` +cache line 0 (segment 0, line 0): + 355a86d2 7957db1c d21772af 6fc1e09b d55ce61d 6e6a278b d3f543ce 223d8e82 + 143ab337 2e9f05bd 2eb389bf 0c6e449e 5cfa4222 ba6560fe 8e3e1aa4 dbcc1d53 +cache line 4194303 (segment 65535, line 63): + 41190d91 bd277957 22ddbb49 6986f207 df69a4d6 26401a3a 818230fb c417122d + 3597b211 b553ce55 cf39cc0d 3b7fc43a 3fd43b00 67e1c80e ffa7ea7d ca2960ab +cache fingerprint, FNV-1a 64 over all 2^26 words as little-endian bytes: 48c4f5bf24166b2e +dataset words 0..15 (item 0): + ffc3cd94 5920ccd8 392f44bb 5e57f67a 2f2bc2a9 620b0e36 bdc09014 436654bf + 311e0b48 1abd93ad 59cc7ce8 ee5247b2 86171fe8 6d874751 c9f7728f 7c2a435d +dataset[0x0fffffff] = a33ada72 +dataset[59471966] = e8b73d94 dataset[217795994] = 337028b5 +dataset[3093825] = e3922dca dataset[267580473] = 26b5f1d8 +``` + +The cache fingerprint has been reproduced by: the Swift CPU fill, the Rust fill, the Metal GPU fill, the clang emulation of the CUDA fill, the RTX 5090 (CUDA and NVIDIA OpenCL), Apple OpenCL, pocl, and the AMD gfx1036 (`docs/bench-log.md`, the five entries dated 3 October 2026 for memory-hard, igneum-pow, proto-opencl, gfx1036 and NVIDIA OpenCL). + +## 1.9 The 32-lane unit of work + +Definition. The hash is defined over an aligned group of 32 consecutive nonces `g .. g + 31` with `g AND 31 = 0` (wrapping modulo 2^32 at the top of the range). `shfl` exchanges registers within that group: lane `l` reads from lane `l XOR mask`, `mask < 32`, so the exchange never leaves the group. The verification unit is the group: `hash(n)` is computed by evaluating the group `n AND ~31` and taking lane `n AND 31` (`verify.rs`, `Epoch::hash`). + +Exchange semantics on hardware: the group MUST be realised so that the exchange is exactly `lane XOR mask` over the 32 logical lanes, independent of the hardware wave width. The rule `proto-opencl/WAVEFRONT.md` fixes and `kernel.cl` implements: + +| Path | When permitted | +|---|---| +| Hardware sub-group or warp shuffle (`simd_shuffle_xor`, `__shfl_xor_sync`, `sub_group_shuffle_xor`) | Only when the work-group is exactly 32 items and the device reports a sub-group size of exactly 32 for that kernel and work-group | +| Local-memory exchange (each lane writes its register to shared memory, one barrier, reads slot `lid XOR mask`) | Always permitted. REQUIRED whenever the sub-group size is not exactly 32 or cannot be queried per kernel | + +The wave64 rule: on a device whose native wave is 64 lanes (AMD GCN, CDNA, RDNA compiled as wave64; figures approximate per `WAVEFRONT.md`), a wave holds two logical units. A sub-group shuffle would still compute `lane XOR mask` correctly (the emulator rows in `WAVEFRONT.md` show this), but the specification does not permit it because the lane-to-work-item mapping is not guaranteed and because a broadcast would read the wrong half. Such devices MUST use the local-memory exchange. Both paths produced the identical batch fingerprint `f99fb375b3abeaf5` (FNV-1a 64 over the 2^13 outputs at base nonce 0, pack igneum-genesis-mh) on Apple OpenCL, pocl with sub-group shuffles, pocl with local memory, and seven emulator configurations including wave64 (`docs/bench-log.md`, proto-opencl entry); the AMD gfx1036 and the RTX 5090 via OpenCL printed `98af644e993239e2` at 2^24 outputs, identical to each other (gfx1036 and NVIDIA OpenCL entries). Cost of the local-memory path against the warp shuffle on the 5090: about 4%, approximate (NVIDIA OpenCL entry). + +## 1.10 Target comparison + +The hash is 64 bits. A block is valid for the lottery when `hash <= target64`. The mapping between the chain's 256-bit difficulty target (rusty-kaspa `Uint256`, fork point a3) and `target64`, and the block-level computation that pruning proofs use (`calc_level_from_pow`, fork map a3, assumes a uniform 256-bit output), are forward references to section 2 and are Open (section 6, O-2.4). Candidate: `target64 = target256 >> 192` and `level = leading_zeros(hash)` capped at 64. + +## 1.11 CPU verification procedure + +Implemented (`verify.rs`, `memhard.rs`). A verifier holds the program for the epoch, the mixer parameters and the 256 MiB cache for the day. It never holds the dataset. To verify a block with lane nonce `n`: + +1. Evaluate the group `g = n AND ~31` with the interpreter of section 1.7, register-major (`r[reg][lane]`) so lane loops vectorise. +2. At each `load`, gather the 32 masked indices, deduplicate by item (`idx >> 4`), derive the distinct items with all chains interleaved round by round (all mixers for round `r`, then all cache-line XORs for round `r`), and hand each lane its word. Interleaving lets the eight dependent misses of each item overlap across up to 32 items; without it the verifier pays about 8 x 100 ns of DRAM latency per item in series. +3. Fold and compare lane `n AND 31` against `target64`. + +The verifier does at most 104 x 32 = 3,328 item derivations for a 104-load program (fewer with duplicates); the design bound is 4,096 items per unit (design document, Lottery seeds item 3), which a 128-load program meets exactly and a 144-load program (4,608 items) exceeds. The load-count rule of 1.4.2 is what will enforce the bound. + +| Verifier | ms per 32-lane unit, steady (avg of 20) | Worst cold single unit | Source | +|---|---|---|---| +| Rust, one M5 Max performance core, 104 loads | 0.441 | 0.41 to 0.87 across five seeds | Measured, `docs/bench-log.md`, igneum-pow entry | +| Rust, 144 loads, 4,608 items | 0.579 | | same | +| Swift, 104 loads | 0.649 | 1.16 to 2.11 | Measured, memory-hard dataset entry | +| Swift, 144 loads | 1.205 | | same | +| Closed-form dataset (not memory-hard, for scale) | 0.002 (Rust), 0.017 (Swift) | | same entries | + +The 10 ms gate (Target) is met with a margin of about 17x steady and 11x worst-cold on this core. Not measured: a 2019-class laptop core (design document, "Three experiments before gate 3"), which is what fixes the gate. + +## 1.12 Schedules: epoch, day, era + +All times are DAA seconds since genesis (section 0.6). At 1 block per second one DAA second is about one block; the schedules below are written in DAA seconds so that block-rate steps (section 2) do not move them. + +| Clock | Length | What changes | Label | +|---|---|---|---| +| Epoch | 3,600 DAA s | The program: new seed words from the VDF of section 4, new kernel | Designed (design document, "Always evolving, on three clocks"); the epoch length is a prototype value, to be fixed at gate 2 by the difficulty-tracking measurement (fork map c1: the hash rate steps by program, 35 to 48 Mhash/s across seeds on the M5 Max, so the DAA window must track within an epoch) | +| Day | 86,400 DAA s | The day key, hence the cache and the dataset | Designed | +| Era | 15,552,000 DAA s (180 days) | Era parameters and one instruction-family unlock, section 1.13 | Designed; the length is a prototype value (the design says "every 6 months") | + +Epoch `e` covers DAA scores `[3,600 e, 3,600 (e + 1))`. The epoch of a block is the epoch of its own DAA score, so "which program was this block mined under" is a function of the header alone once the seed is known. The program for epoch `e` is `generate_from_words(S_e)` with `S_e = seed_words_from_bytes(program_seed_e)` and `program_seed_e` the 32-byte VDF output of section 4.3. + +Day `d` covers DAA scores `[86,400 d, 86,400 (d + 1))`. The design document names a day seed and does not say how it is derived. Proposed (Designed, Open, O-1.10): `day_bytes = "igneum-day/" || d_le64 || program_seed of the first epoch of day d`, so the day key is as unpredictable as the epoch seed and known 20 minutes before the day starts (section 4.5), which is enough for a 0.2 s CPU cache fill or a 2 ms GPU one plus a 13 to 30 ms GPU dataset build (Measured, section 1.8.3 and `docs/bench-log.md` RTX 5090 memory-hard entry: 13.4 ms for 1 GiB). + +Era `n` covers DAA scores `[15,552,000 n, 15,552,000 (n + 1))`. + +## 1.13 Era parameter draw, instruction-family reserve, dataset growth + +Designed at the level of a sentence in the design document ("a new instruction mix and memory pattern drawn from chain state, plus one instruction family unlocked from a reserve fixed at genesis"; "the dataset grows on a genesis-fixed schedule"). No draw procedure, reserve list or growth rule exists in code (section 0.3). This section proposes them so that they can be reviewed; everything here is Open (section 6, O-1.11 to O-1.13) until gate 1 fixes it. The rule that is not open: nothing in this section is ever changed by a human release. The draw and the unlock are functions of genesis constants and chain state. + +### 1.13.1 Era seed and draw + +The era seed `E_n` is the 32-byte output of the 1-hour VDF of section 4.4. One SplitMix64 stream seeded from `seed_words_from_bytes("igneum-era/" || n_le64 || E_n)` words 0 and 1, drawn in a fixed order, sets the era parameters within genesis-fixed bounds: + +| Parameter | Base (era 0) | Draw | Bound | +|---|---|---|---| +| Op weights for the ten non-load ops | table 1.4.2 | each perturbed by `below(2 * B + 1) - B` points, then renormalised by largest remainder to 100 minus the load weight | B = 2 points, proposed | +| Load weight | 25 | not drawn | fixed, so every era is equally memory-bound | +| Output fold rotations | (7, 14, 21), (9, 18, 27) | each `1 + below(31)` | 1..31 | +| Mixer round count | 8 | not drawn | fixed, so the verify budget holds | + +"Memory pattern" in the design document is read here as the item-address pattern (the cache line index word, `s[0]` in 1.8.5, and the XOR-all-sixteen rule); the proposal is to leave it fixed at era 0 and let the unlocked families change the kernel instead, because every change to the item derivation changes the verify time and must be re-measured. + +### 1.13.2 Instruction-family reserve + +At genesis the generator carries the eleven families of 1.4.1 live and a reserve list of further families in a fixed order. At the start of era `n >= 1`, reserve family `n` becomes live with weight `W_new` taken proportionally from the live non-load families. A family may enter the reserve only if it is integer-exact and has passed the cross-vendor conformance of section 1.15 on every vendor in the benchmark (Metal, CUDA, OpenCL on NVIDIA and AMD), with its own edge-case vectors, before genesis. Candidate families, all integer ALU operations present on Apple, NVIDIA and AMD: variable left shift and logical right shift by `src AND 31`; bit-field extract with an immediate offset and width; `andn` (`dst = dst AND NOT src`); byte permute of `dst` by an immediate selector; population count and count-leading-zeros folded into `dst` by add; a three-register select (`dst = bit of src2 ? src : dst`); a second shuffle form (`lane + delta mod 32`). The order and `W_new` are Open. A family that is not in the genesis reserve can only be added by the upgrade path of section 5.7. + +### 1.13.3 Dataset growth + +Designed: 2 GiB at genesis plus 0.5 GiB per year (design document, "Which cards mine?", labelled approximate there for the card-lifetime consequence, not for the schedule). Proposed rule: the dataset for day `d` has `N_d` items where + +``` +N_d = floor((2 GiB + 0.5 GiB * (86,400 d / 31,536,000)) / 64 bytes) +``` + +evaluated in integers (bytes), with one year = 31,536,000 DAA seconds. The dataset grows by about 23 KiB per day and is recomputed with the day key. Two consequences are Open: + +- Index mapping. `src AND MASK` requires a power-of-two size. For a non-power-of-two `N_d` the proposed mapping is `idx = (src * N_words) >> 32` computed in 64 bits (a multiply-shift range reduction; uniform to within 2^-32, branch-free, integer only). At `N_words = 2^28` this gives `src >> 4`, not `src AND MASK`, so adopting it changes the 1 GiB vectors; gate 1 chooses between (a) the multiply-shift mapping with new vectors, or (b) power-of-two sizes only, growing in steps (2 GiB, 4 GiB) on the same schedule's average, which keeps `AND MASK` and means a 4 GiB card lasts until the 4 GiB step instead of fading. +- The item index `t` is 32 bits, so the construction as written tops out at 2^32 items = 256 GiB, which the schedule reaches after 508 years. No action needed. + +## 1.14 Determinism requirements + +A conforming implementation MUST: + +1. Use only integer arithmetic. No floating point anywhere, including in index computation and in the mixer (floating point rounds differently per vendor and would split the chain; design document, hostile review table row 2). +2. Mask or range-reduce every dataset index exactly as 1.5 and 1.13.3 state, and never read outside the dataset. Every load in emitted source MUST have the single form `dataset[rN AND MASK]` (or the adopted range reduction), checkable by text search (`TESTS.md` section 5: 13 of 13 loads at three sizes, 416 of 416 indices exceeded MASK before masking). +3. Implement `rotr` by `src AND 31` and `rotl` by an immediate in 1..31; a rotate by 0 or 32 through the immediate path is undefined and MUST NOT occur. +4. Compute `mulhi` as the exact high 32 bits of the 64-bit product (`__umulhi`, `mulhi`, `mul_hi`). +5. Wrap on overflow everywhere (add, sub, mul, mad, the SplitMix and FNV state). +6. Realise the exchange over exactly 32 logical lanes as 1.9 requires, independent of the hardware wave width. +7. Produce the same output for the same (program, day key, dataset size, nonce group) on every run: `TESTS.md` section 4 (5 runs and 3 compiles, one cold, fingerprint `933787e8cfefccb7` closed-form; `62a4f0eb018df273` memory-hard, `MEMHARD.md` section 2.5). + +## 1.15 Conformance procedure for a miner implementation + +A miner, kernel emitter or verifier conforms when all of the following pass. Each is a command that exists today; the AMD discrete-card rows are pending. + +1. Cache: fill the 256 MiB cache for day `2026-10-03` and reproduce FNV-1a 64 `48c4f5bf24166b2e`, cache line 0 and cache line 4,194,303 of section 1.8.5, word for word. +2. Dataset self-test at 1 GiB: words 0..15, word `0x0fffffff`, and the 64 sampled words of `vectors.json` (four are quoted in 1.8.5). +3. Vectors: all 96 outputs of the three units at base nonces 0, 4,096 and 1,000,000 for pack `igneum-genesis-mh`, standalone (one unit per launch) and in batch (many units per launch, at least two units per work-group or block). Section 1.17 lists them. +4. Batch fingerprint: FNV-1a 64 over the 2^13 outputs at base nonce 0 = `f99fb375b3abeaf5`, and over the 2^24 outputs = `98af644e993239e2`. +5. Fuzz: at least 200 random programs (`--fuzz 200` or the Rust equivalent when it exists), 4 units each at base nonces drawn from the full 32-bit range including wraps past 2^32, at 64 MiB, 256 MiB and 1 GiB, zero mismatches against the CPU interpreter, zero compile failures, generator contract asserted on every instruction. Reference: 200 of 200 and 10,000 of 10,000 (`TESTS.md` section 1; re-run on the memory-hard dataset, `MEMHARD.md` section 2.5). +6. Edge: the 14 hand-built programs of `TESTS.md` section 2 (rotates by 0 and 31 through the register path, `mulhi` at the extremes, every shuffle mask, loads at index 0 and at MASK through in-range and out-of-range registers, wraparound on add, sub, mul, mad, zero loads, 64 loads), 128 of 128 lanes each. +7. Static mask check on every emitted kernel (1.14 item 2). +8. Exchange rule: on any device whose sub-group size is not exactly 32, or cannot be queried per kernel, the local-memory path is used and the run says so. + +Measured conformance to date (`docs/bench-log.md`): Apple Metal, Apple OpenCL, pocl (CPU), the clang emulators, NVIDIA RTX 5090 via CUDA and via NVIDIA OpenCL, AMD gfx1036 via AMD OpenCL, and the Rust and Swift CPU references all pass items 1 to 4; items 5 and 6 have run on Metal and the CPU references only. A discrete AMD card has not run anything (ledger M8). + +## 1.16 Parameters marked "prototype value, to be fixed at gate 1" and what fixes them + +| Parameter | Prototype value | Measurement or decision that fixes it | +|---|---|---| +| Seed derivation (FNV-1a plus SplitMix64) | section 1.3 | Replace with a standard hash (ledger M7); re-run the weak-program census on the result | +| Instructions per program, iterations | 64 x 8 | CPU verify on a 2019-class laptop core under 10 ms with the memory-hard dataset; register-pressure and occupancy on the three vendors | +| Registers per lane | 8 | same | +| Op weights, including the 25% load weight | table 1.4.2 | Weak-program census of at least 10^5 programs (bias, avalanche, distinct load addresses, `or` saturation, nonce-independent registers), then the rejection rule; and the exact-load-count rule (ledger M5) | +| Output fold rotations | (7, 14, 21), (9, 18, 27) | Stats run of `TESTS.md` section 3 on the chosen fold | +| Cache size, lines per segment, ChaCha rounds | 256 MiB, 64, 12 | Shortcut ratio and time-memory curve on an RTX 5090 and an AMD discrete card; external review of the chained-block cache | +| Item rounds, mixer shape | 8, section 1.8.4 | Same measurement; external review of the mixer; weak-key check on `ROT` | +| Dataset size at genesis | 1 GiB in packs, 2 GiB designed | Hash-rate and shortcut ratio at 2 GiB on the three vendors; the growth and index-mapping decision of 1.13.3 | +| Epoch length | 3,600 DAA s | Difficulty tracking across program steps on the devnet (gate 2, fork map c1) | +| Era length, draw bounds, reserve list, `W_new` | section 1.13 | Design review; each reserve family's own conformance run | +| Header binding of the init words | section 1.6 | Fixed when section 2 fixes the pre-PoW header hash; new vector set | + +## 1.17 Test vectors + +Pack `proto-cuda/packs/igneum-genesis-mh/` (seed `igneum-genesis`, day `2026-10-03`, memory-hard, 2^28 words, MASK `0x0fffffff`, 32 lanes). Produced by the Swift CPU interpreter, cross-checked by Metal before the pack was written, then reproduced by every implementation listed in 1.15. The two closed-form packs (`igneum-genesis`, `igneum-hourly`) are regression vectors for the interpreter only and are not the lottery hash. + +Unit at base nonce 0, lanes 0..31: + +``` +1fb0b3bbc1ac8279 61533759ca995ac8 97cd15ec31242001 3c19e71e2dbff828 +c7aba60b6c2ab016 8f13725d2a59bf84 7b826b82fdec5a2f 3131d7817418c04e +d7c7e57c6aaac948 b3b639431ebdd3d6 400af448658a1d56 17f5760b8cadb8fd +01ded5c4894c411c d6b3fdbc57bdb128 888efb103cecb983 ac73538353a340c4 +3f9914ad9445f052 8a860b28afd3dd42 c2a122ef9c1fa330 96c4e58663f82dce +baae4b9d6a3c4320 cd6c2dd653e07743 7d3d00f0fb46325b 7636224050132b09 +dc5d76c741cc60f4 aef3031b0d4b701c 735987104a975f5f a58fce4010bf99dd +b16c863d2ffdc0dd 0337b56e39529c90 802ac0bc1d0c0696 fa052263a854f3de +``` + +Unit at base nonce 4,096: lane 0 `f219cf7ecf6ec450`, lane 1 `b575ab4b388c01f7`, lane 31 `d5b198435d9c40da`. Unit at base nonce 1,000,000: lane 0 `8a6f7a32b06edb52`, lane 1 `362d2b789a0c0ebf`, lane 31 `7040299873672a35`. The remaining 58 values are in `vectors.json` and `vectors.h` of the pack. + +Cache, dataset and mixer vectors: section 1.8.4 and 1.8.5. Seed words: section 1.3.1. Generator: section 1.4.3. Batch fingerprints: section 1.15 item 4. diff --git a/docs/spec/02-consensus.md b/docs/spec/02-consensus.md new file mode 100644 index 00000000..ffd425c1 --- /dev/null +++ b/docs/spec/02-consensus.md @@ -0,0 +1,104 @@ +# Igneum protocol specification, section 2: the ordering layer + +Spec version 0.1, 3 October 2026. Status of this section: Designed. The base (rusty-kaspa at commit `01b532e8b553523216471682649693af92f0fd16`, v2.1.0) is built and has run a 3-node devnet at Kaspa's own parameters (`docs/bench-log.md`, entry "rusty-kaspa base build and 3-node devnet"). No fork point is implemented. `docs/fork-divergence.md` does not exist yet; when it does, it records what was changed against this section. + +This section is written as a delta on rusty-kaspa. Everything not named here is Kaspa's rule at the forked commit. File and line references are those of `docs/fork-map.md`, which must be re-checked after any vendor update. + +## 2.1 Block rate and GHOSTDAG parameters + +| Parameter | Launch value | Kaspa source | Label | +|---|---|---|---| +| Target block rate | 1 block per DAA second | `consensus/core/src/config/bps.rs`, `Bps::<1>` (fork map f1) | Designed | +| GHOSTDAG k | 18 | Kaspa's k table at 1 BPS (fork map f1) | Designed, taken from Kaspa's table (delta 0.01, network delay bound 5 s, `constants.rs:13 to 16`) | +| Max parents | 10 | same | Designed | +| Mergeset size limit | 180 | same | Designed | +| Merge depth | 3,600 DAA s (3,600 blocks at 1 BPS) | `MERGE_DEPTH_DURATION` (fork map e1) | Designed, unchanged from Kaspa | +| Finality depth (backstop only) | 43,200 DAA s | `FINALITY_DURATION` | Designed; live finality is section 3 | +| Pruning depth | 108,000 DAA s | `PRUNING_DURATION` | Designed; MUST stay above the longest checkpoint gap and MUST NOT pass the latest certified checkpoint (section 3, F3) | +| Coinbase maturity | 100 blocks | `bps.rs:119 to 121` | Designed | +| Block-rate steps | 1, then 4, then 10 per second | design document, decisions table | Designed. Each step is a planned fork with its own test campaign (as Kaspa's Crescendo), taken only once proving lag holds under 60 s. Each step re-derives k, max parents and mergeset limit from Kaspa's table; the DAA-second schedules of sections 1, 3, 4 and 5 do not move | + +Every Kaspa fork activation (`ForkActivation`) is `always()` on Igneum: there is no history to replay (fork map f2). Own genesis, network prefix, ports and seeders (`network.rs:42 to 60, 238 to 252`). + +## 2.2 Proof of work (fork points a1 to a5) + +| Fork point | Kaspa today | Igneum | +|---|---|---| +| a1, a2 | `PowHash` cSHAKE256 then `KHeavyHash` matrix | Both deleted. The lottery hash of section 1 replaces them. `kaspa_pow::State` carries the `igneum_pow::Epoch` for the header's epoch (program by epoch of the header's DAA score, dataset by day of the header's DAA score, section 1.12) | +| a3 | `pow <= target` on a `Uint256` | Section 1.10: `hash64 <= target64`. The mapping from the 256-bit target and the block level for pruning proofs are Open (O-2.4) | +| a4 | Header validated in isolation | Header validation gains a dependency on chain state: the epoch seed (section 4). The seed source is named in the header (section 2.4) so the dependency is checkable from the header plus its past. During IBD and pruning-proof validation (`pruning_proof/validate.rs:192`) the node MUST be able to derive the program for any header from headers and certificates alone; this is the hardest fork point (fork map, risk High) and is Open until the devnet proves it (O-2.5) | +| a5 | Pre-PoW hash over 12 fields | Adds `vote_key_hash`, `seed_source` and `proof_ref` (section 2.4) so all three are committed by the nonce. Every header-hash test vector changes, genesis included | + +What the nonce commits to: the pre-PoW header hash `H` enters the register initialisation (section 1.6, Open O-1.9). The VDF requirement that the seed checkpoint commits to the full block hash including the nonce (`proto-vdf/README.md`) is satisfied by Kaspa's `Header.hash` covering the nonce. + +## 2.3 Difficulty adjustment (fork points c1, c2) + +Kaspa's sampled DAA (KIP-4) is kept as the retarget: average target of a sampled window, `new_target = avg x measured / expected`, clamped to `MAX_DIFFICULTY_TARGET` = 2^255 - 1, minimum window 150 samples (`difficulty.rs:97 to 198`). + +| Parameter | Value | Label | +|---|---|---| +| Window duration | 2,641 DAA s (Kaspa's `DIFFICULTY_WINDOW_DURATION`) | Designed, prototype value, to be fixed at gate 2. The design document says "about 2,600 blocks, roughly 45 minutes" | +| Sample interval | 4 s, 661 samples | Kaspa's constants, kept | +| Keyed to | DAA score, never wall-clock | Designed | + +What fixes the window: the hash rate steps at every epoch because programs differ in cost (35 to 48 Mhash/s across seeds on the M5 Max, `docs/bench-log.md` first-run entry; 228 vs 185 Mhash/s for 104 vs 128 loads on the RTX 5090). With a 44-minute window a 30% step means roughly half an epoch at the wrong block rate (fork map c1). Two remedies, decided at gate 2 by a simpa run (fork map c2): shorten the window (candidate 1,800 s), or equalise cost per program through the exact-load-count rule of section 1.4.2, which this specification prefers because it removes the step rather than tracking it. + +The 30-day vote-weight window of section 3 is a new reader of DAA scores, not a retarget change. + +## 2.4 Header (fork point d) + +Kaspa's `Header` (version, parents, hash_merkle_root, accepted_id_merkle_root, utxo_commitment, timestamp, bits, nonce, daa_score, blue_score, blue_work, pruning_point) plus: + +| Field | Size | Meaning | Label | +|---|---|---|---| +| `vote_key_hash` | 32 bytes | Hash (the chain's BLAKE2b-based `Hash`) of the producer's BLS12-381 G1 compressed public key (48 bytes). The first block that uses a key reveals the key itself in the coinbase payload. Vote weight accrues to this key (section 3, W1) | Designed. The design document says a 32-byte hash; `docs/fork-map.md` row d says a 48-byte key. This specification takes the hash: 32 bytes x 86,400 blocks/day = 2.8 MB/day of header growth against 4.1 MB with the key | +| `seed_source` | 32 bytes | Hash of the checkpoint block from which the header's epoch program seed was derived (section 4.3) | Designed, proposed here, Open (O-4.3) | +| `proof_ref` | 32 bytes | Hash of the highest aggregated block proof the producer knows | Forward reference. The proof format, what "highest" means and the validation rule belong to the chunked proving protocol, which is out of scope for 0.1. Until that protocol is specified the field is all zeros and unvalidated | + +Also edited (fork map d): p2p `p2p.proto:76 to 89`, `convert/header.rs:13, 45`; RPC `rpc.proto:25`, `model/header.rs:85, 105, 248`; genesis headers; the headers store (serde, re-sync needed); header mass. + +Certificates (section 3, C3) travel in the block body, not the header: every block carries the highest certificate its producer knows, and a block whose selected chain does not pass through every certified checkpoint in its past is invalid (post-PoW validation, fork map e2). + +## 2.5 Emission (fork points b1, b2, b3) + +Designed (design document, "The token" and "Difficulty, block timing and proving cadence"). No pre-deflationary phase, no month table, no treasury. + +| Parameter | Value | Label | +|---|---|---| +| Hard cap | 4,000,000,000 IGN, approached and never reached | Designed | +| Year | 31,536,000 DAA s (365 days) | Designed, prototype value: the design document says "1 billion a year"; the day count is this specification's choice | +| Emission in the first two years | 1,000,000,000 IGN per year | Designed | +| Halving interval | 63,072,000 DAA s (2 years), for ever | Designed | +| Launch ramp | linear from 10% at genesis to 100% at DAA second 2,592,000 (30 days) | Designed | +| Split | 80% block producer, 20% proving pool | Designed | +| Base unit | Open (O-2.6): the EVM layer implies 10^18 per IGN; `docs/fork-map.md` b2 wrote `cap_sompi = 4e9 x 1e8`. One of the two is chosen before any coinbase code is written | + +Emission per DAA second at DAA score `t`: + +``` +E(t) = ramp(t) * floor(10^9 * UNIT / 31,536,000) >> floor(t / 63,072,000) +ramp(t) = min(1, 1/10 + 9/10 * t / 2,592,000) evaluated in integers as a rational with denominator 25,920,000 +``` + +With `UNIT = 10^18` the pre-ramp rate is about 31.7098 IGN per DAA second (Designed). The geometric series sums to 4 x 10^9 IGN; the ramp withholds 0.45 x 30/365 x 10^9 = about 37 million IGN that are never minted, and integer floors withhold a negligible further amount, so the cap is a strict bound. + +Emission is keyed to DAA score, not to blocks or timestamps: more blocks never means more coins and miner-chosen timestamps cannot mint (hostile review table, "Emission per wall-clock second invites timestamp games"). Payment is to blue blocks only, through the merging block's coinbase as in Kaspa (`coinbase.rs:97 to 142`): the coinbase of block B pays, for each blue block M in B's mergeset, `E(daa_score(B))` split 80% to M's miner and 20% to the proving pool. Red blocks are unpaid. How the DAA-score increment is apportioned among the blues of one mergeset at higher block rates is Open (O-2.7); at 1 BPS the mergeset is usually one block. + +The 20% proving share is paid to the prover set recorded for the proven block, which is known 20 to 60 s after the block (design document, "Proving lag"). The coinbase payload format gains prover outputs (fork map b3, risk High). The rule that the pool is paid as a fixed amount per block divided among shards by consensus proving cost is in section 5.3. + +Fees are not in the coinbase; they are in the execution layer (section 5). + +## 2.6 Duplicate inclusion for the EVM layer + +Designed (hostile review table, "Parallel blocks include the same transaction"). Blocks carry transactions only and make no claim about state. The execution layer orders transactions by the GHOSTDAG ordered sequence (the selected chain's mergeset order, Kaspa's), and: + +1. The first copy of a transaction (by transaction hash) in the ordered sequence executes and pays its fees. +2. Every later copy is deduplicated before execution and before proving. It executes nothing and pays nothing. +3. A later copy still occupies block space (mass) in the block that includes it, so the including miner bears the cost. +4. Transactions from one account are subject to the EVM nonce rule over the same ordered sequence: a transaction whose nonce is not the account's next nonce at its position is skipped, not failed, and pays nothing. This is Kaspa's "skip conflicting spends" pattern applied to the EVM. + +Block number, timestamp, blockhash, coinbase and prevrandao over the ordered sequence are out of scope for 0.1 and Open (ledger P5, O-2.8). + +## 2.7 What the devnet showed and did not show + +Measured (`docs/bench-log.md`, devnet entry): three kaspad nodes at Kaspa's devnet parameters (10 BPS, k 124) held identical block counts, DAA scores and sink hashes at 18 of 19 ten-second samples under a 27 MH/s CPU miner; the DAA raised difficulty from genesis bits at block 6,018 and the block rate fell from 56 to 62 blocks/s toward the 10 BPS target. GHOSTDAG k was not exercised (a single serial miner never produced parallel blocks); a second miner is the next step. Nothing in that run used an Igneum parameter. diff --git a/docs/spec/03-finality.md b/docs/spec/03-finality.md new file mode 100644 index 00000000..e103ace8 --- /dev/null +++ b/docs/spec/03-finality.md @@ -0,0 +1,120 @@ +# Igneum protocol specification, section 3: sustained-mining finality, version 2 + +Spec version 0.1, 3 October 2026. Status of this section: Designed. Simulated at the checkpoint level with latency, partitions and eclipses but without a DAG (`sim/finality_v2.py`, `sim/results_v2.md`, `docs/bench-log.md` entry "sim/finality_v2.py"). Not implemented. Not externally reviewed. Gate 3 is the external review that tries to break it. + +The rule is exactly the design document's "Finality rule, version 2, after the second hostile review" with the quorum floor added on 3 October 2026 after the second simulation. CLAUDE.md's FINALITY RULE V2 paragraph is the short form. Every quantity is in blue score or DAA seconds, never wall-clock (section 0.6). + +Two statements frame everything below. Finality is miner-only and self-contained: no stake, no bond, no other chain, no committee that is not the set of recent miners. Equivocation costs history, not coins, because there are no coins to slash. + +## 3.1 Weight + +- **W1.** Every block header names a vote key by `vote_key_hash` (section 2.4): the hash of a BLS12-381 G1 compressed public key. The first block that uses a key reveals the key in its coinbase payload. A header whose `vote_key_hash` has never been revealed is valid; the key simply cannot vote until it is revealed. +- **W2.** The weight of key k at checkpoint block C is the number of blue blocks in C's past whose header names k and whose DAA score is in `(daa(C) - 2,592,000, daa(C)]` (the trailing 30 days). No damping, no cap, no floor. Blocks are the only thing in proof of work that cannot be forged, so weight is counted in blocks. The damped rule of version 1 was removed because its 2x cap was defeated by splitting into free keys (`sim/results.md` table C; `docs/bench-log.md` finality_sim entry). +- **W3.** Dust: a key with fewer than 100 blocks in the window is not a voter and is in no denominator. Measured consequence (`sim/results_v2.md` A and G): every honest key in a 1,000-key Pareto network clears 100 blocks by day 20 from zero history and the smallest new keys need up to 25 days after a doubling; a 9x renter's victims lose about one point to dust (B). +- **W4.** Steady state: weight equals hashrate. Simulated correlation 1.00000 at day 60, Gini equal to three figures, weight-to-hash ratio within 0.82 to 1.12 for every key (`sim/results_v2.md` A). +- **W5.** Key succession: a message signed by the old key naming a new key, included in any block, moves the old key's weight history to the new key once. The old key is dead thereafter: its later blocks earn nothing and its votes are invalid. + +The headline arithmetic (W2 with constant hashrate): an attacker with share a of hashrate for t days holds weight share `(t/30) x a/(1+a)`, verified by simulation to 0.04 points for a = 1 and 2 (`sim/results_v2.md` B). + +| Attacker hashrate relative to honest (a) | Reaches 1/3 (can veto) | Reaches 2/3 (locks alone) | +|---|---|---| +| 1x (50% of blocks) | day 20.0 | never (ceiling 50%) | +| 2x (67%) | day 15.0 | day 30.0 | +| 4x (80%) | day 12.5 | day 25.0 | +| 9x (90%) | day 11.1 | day 22.2 | +| unbounded (100% of blocks, honest miners gone) | day 10 | day 20 | + +All in public, on the hashrate charts. 51% never reaches 2/3 while honest miners keep mining. + +## 3.2 Checkpoints + +- **C1.** Checkpoint i is the selected-chain block at blue score 30 i. It is determined once the virtual's blue score reaches 30 i + d. d = 60 at 1 block per second is a placeholder (ledger F7): the gate 3 devnet records the reorg-depth distribution and sets d so that a vote split at one index is rare and self-heals at the next. d scales with block rate. A lock lands about 90 to 120 s after a transaction (Designed; simulated lock latency after the checkpoint block is median 2.5 s, p99 4.6 s at a 2-s inter-region delay, `sim/results_v2.md` A). +- **C2.** A vote is a BLS signature over `(chain_id, i, hash(C_i))` under a fixed domain-separation tag. Votes gossip as their own message type. +- **C3.** A lock certificate for index i is an aggregate BLS signature over one checkpoint block hash with a bitmap of signers, whose signed weight meets Q3. Every block carries the highest certificate its producer knows. A block whose selected chain does not pass through every certified checkpoint in its past is invalid (section 2.4). +- **C4.** A node holding a certificate for index i rejects any other certificate for index i and publishes the pair as evidence (section 3.6). +- **C5.** No certificate may form in the chain's first 3,600 DAA seconds (design document). See 3.8 for the proposed first-month rule. + +## 3.3 Quorum + +- **Q1.** Presence window P = 240 checkpoint indices (2 hours of blue score at 1 BPS). +- **Q2.** Participation of key k at index i, "cert reading": the number of indices j in `[i - 240, i - 1]` for which a certificate containing k's vote is in the past of C_i, divided by 240, capped at 1. An index with no certificate in C_i's past credits nobody. A key whose first block is fewer than 240 indices old counts 1. This reading is objective (every node computes it from certificates in C_i's past) and self-healing; the "seen" reading is rejected as not objective and the "frozen" reading as total with extra steps (`sim/results_v2.md`, "Recommended parameters"). +- **Q3.** Active weight at i is the sum over voters of weight x participation. Total weight at i is the sum of weight over all keys above dust. A certificate for index i locks when the unscaled weight of its signers is at least **2/3 of active weight** AND at least **56.7% of total weight** (the floor 0.85 x 2/3 = 17/30). Both tests use weights and participation computed at C_i, so any node can verify a certificate from C_i's past. +- **Q4.** Certificate grace: an aggregator closes a certificate at the later of quorum time and `t0 + grace`, where grace MUST be at least 3x the worst honest one-way network delay (15 s in the simulation at a 2-s delay; at a 5-s delay the slowest region already lost 1.7 points of participation to a 15-s grace, `sim/results_v2.md` A). The value is Open (O-3.4) and does not affect validity, only which votes a certificate carries. + +### 3.3.1 Why the floor, from the simulation + +Both denominators alone fail (`sim/results_v2.md`, seed 7, seed 11 agrees on B and E): + +| Scenario | Active alone (cert reading, P = 240) | Total alone | Active + floor 0.85 | +|---|---|---|---| +| E: 50/50 honest partition, no attacker, 150 min | 90 conflicting locks, first at 60 min (240 with instant DAA retarget, first at 30 min) | 0 | 0 | +| E: 60/40 honest partition, 360 min, with retarget | 624 conflicts | 0 | 0; majority side locks from minute 13 | +| E: 33/33/34, 360 min, with retarget | 1,198 conflicts | 0 | 0 | +| F2: 34% attacker poisons one eclipsed 20% pool, 1 / 2 / 4 h | 10 / 65 / 174 conflicts, first at 49 min | 0 | 0 (floor 0.80 gave 10 / 65 / 174, because the eclipsed side's 54% clears 53.3%) | +| C: 34% of weight silent, keeps mining | 0 min to first lock, 0 stalls | never (8,666 stalls in 3 days, and for the whole window) | 0 min | +| C: 40% silent | 13 min, 26 stalls | never | 13 min, 138 intermittent stalls in 3 days | +| C: 45% silent | 20 min | never | never, for as long as they stay silent | +| D: 35% churn (stops mining and signing) | 2 min | 41 h | 2 min, 98 intermittent stalls in 3 days | +| D: 50% churn | 31 min | 10.1 days | 4.1 days (floor 0.80: 2.2 days) | + +The mechanism active alone fails by: a side that cannot see the other side's votes sees stalls, stalls shrink its denominator by 1/240 per index, and once the denominator has fallen to 1.5x the side's own weight the side certifies alone, after `240 x (1 - 1.5 s) / s` slots for a side of share s (confirmed within 5% across the sweep): 60 min for a 50% side, 120 for 40%, 180 for 33%. The floor makes the second test of Q3 bind before that point: no honest side of any tested split holds 56.7% of total. + +What the floor costs: liveness ends between 40% and 45% of weight silent (total alone: 34%; active alone: above 55%), 50% churn stalls 4.1 days, and at 40% silent or 35% churn the margin is one pool outage thin. The safety bound stays at 1/3 of weight for every event tested; the liveness bound moves from 1/3 (total) to about 42% silent. + +## 3.4 Sortition + +- **S1.** At launch every voter signs every checkpoint. A private VRF, keyed to the vote key and the checkpoint, selects 8 aggregators per checkpoint who publish certificates; anyone MAY aggregate and publish. Aggregator failure costs latency, not safety: the simulation's lock latency (median 2.5 s) assumes one aggregator per region and no failures, so it is a lower bound (`sim/results_v2.md`, "cannot tell us"). +- **S2.** If the number of voters above dust exceeds 8,192 at a checkpoint, the genesis rules switch, at that checkpoint and without a release, to Algorand-style binomial sub-user sortition with an expected 4,000 sub-users per checkpoint and both Q3 thresholds applied to expected sampled weight. The exact VRF construction, the binomial sampling procedure and the threshold on sampled weight are Open (O-3.5). + +Ledger F3 (participation grinding through the bitmap: an aggregator or block producer that drops a rival's votes from certificates lowers that rival's participation) is Open (O-3.3). Candidate fixes: a certificate MUST include every valid vote the aggregator received above a size bound; or participation also counts votes carried in blocks as transactions; or participation falls only on indices where the key's vote appears in no certificate and no block. + +## 3.5 Fork choice + +- **F1.** Candidate tips are tips whose selected chain passes through the highest certified checkpoint the node holds and every lower certified checkpoint. +- **F2.** Among candidates, GHOSTDAG selects by accumulated blue work under Kaspa's merge-depth bound of 3,600 DAA s (section 2.1). A certified checkpoint removes other tips from candidacy; it does not change how blue work is computed. +- **F3.** The pruning point never advances past the latest certified checkpoint. `virtual_finality_point` returns the latest certified checkpoint when it is newer than the depth-based finality point (fork map e2). +- **F4.** There is no hidden-block penalty in consensus. It was removed in review round 2 because it breaks DAG determinism and amplifies eclipse attacks. First-seen MAY break ties in a node's own block template only. +- **F5.** A node started with a configured trusted certificate follows it. A node started cold selects the DAG with the most accumulated blue work, then follows certificates found in it. A private DAG that out-works the public one over the window is a public 51% event lasting weeks. + +What a node does when it holds two valid certificates at one index after a partition heals is not modelled and not defined (`sim/results_v2.md`, "cannot tell us"; O-3.6). C4 says it publishes the pair. The proposal for gate 3: both certificates are evidence against every key that signed both; the node re-evaluates both against Q3 with those keys' weight struck, and if exactly one still locks it follows that one; if neither or both still lock, F2 decides among the two checkpoint blocks' descendants and the index is treated as uncertified. + +## 3.6 Equivocation evidence + +Two votes by one key for different checkpoint blocks at one index are equivocation. The evidence (the two votes) is a transaction that any block MAY include. On inclusion: the key's weight is zero for the rest of the current window and its blocks earn no weight for the next 2,592,000 DAA s. There is no coin penalty. In the simulation, evidence is detected only at the heal and the penalty is forward-looking only; the model does not revoke the conflicting certificates (`sim/results_v2.md`, "cannot tell us"), which is why 3.5's post-heal proposal strikes the equivocators' weight retroactively for the re-evaluation. + +## 3.7 Residual risks, stated + +1. Safety holds with under one third of window weight under hostile keys. Reaching a third takes ten days of 100% hashrate or twenty days of 51%, in public. Beyond a third, two locks can coexist under a partition (E: a 34% equivocator across a 50/50 split breaks every variant, 33% + 34% = 67% per side) and equivocation costs history, not coins. +2. Liveness pauses after a sudden loss of more than about 42% of weight from signing (3.3.1). During a pause the chain is proof-of-work only in practice, so the node ships a flag that tells exchanges to credit nothing until the next lock (3.9). +3. Pools hold their hashers' votes. Vote concentration equals pool concentration, as on Bitcoin, and is public. Stratum v2 job declaration changes transaction choice, not the vote key (ledger F10, G6). +4. A 51% owner who drives half the honest miners away and holds for 30 days owns finality thereafter. Same as Bitcoin, with a month's warning. +5. The delay function (section 4) is a new dependency. The evaluator ships in every node. +6. The dust threshold excludes solo miners under 100 blocks a month from voting, not from rewards. +7. New honest miners are under-weighted for the whole window: after an overnight doubling the new cohort holds t/60 of weight on day t and the old cohort can lock without a single new signature for 19 days (`sim/results_v2.md` G). A doubling and a 1x renter are the same event to the rule, by design. +8. The simulation has no DAG: a partition side's checkpoint block is "the block at blue score 30 i in that view" and conflict counts are index collisions, not reorg depths. Real GHOSTDAG merge under the 3,600-s bound, DAA lag (of the order of an hour, approximate), VRF aggregator noise, uptime (the 0.978 resting participation is a guess), regional silent sets and the cost of keys are all outside the model. + +## 3.8 The first month + +W2 gives the chain no weight at genesis. From zero history the honest network holds 3.3% of a full window on day 1, 33.4% on day 10 and 100% on day 30; 905 of 1,000 keys are under dust on day 1 and all are above it by day 20 (`sim/results_v2.md` A). Ledger F1 shows the consequence: with a one-day honest head start an attacker producing 75% of blocks from day 2 crosses two thirds on day 9 of the chain's life, and the 30-day emission ramp lowers the prize without removing the attack. + +Two rules are on the table: + +| Rule | Status | +|---|---| +| C5 as written: no certificate in the first 3,600 DAA s | Designed (design document) | +| No certificate may form until the window holds 30 days of history: before DAA second 2,592,000 the chain runs plain GHOSTDAG under the merge-depth bound, exchanges are told so, and the node flag of 3.9 is false | Proposed (ledger F1), Open (O-3.1). This specification recommends it: the first month has no weight to defend with, so a lock in that month is a lock by whoever showed up, and the merge-depth bound already bounds a reorg to one hour | + +Gate 3 decides, with the launch-month simulation that does not yet exist. Either way the litepaper states the rule. + +## 3.9 Exchange confirmation guidance + +Designed. The node exposes `finality_active` (true when a certificate has formed within the last 240 indices and the first-month rule of 3.8 has passed) and `last_certified` (index, block hash, DAA score). + +| Situation | Guidance | +|---|---| +| `finality_active` and the deposit's block is in the past of `last_certified` | Credit. Expected wait about 90 to 120 s after inclusion | +| `finality_active`, deposit not yet covered | Wait for the next certificate; do not count blocks | +| `finality_active` false (first month, or a pause under 3.7 item 2) | Treat the chain as proof of work with a one-hour merge-depth bound. Credit nothing below 3,600 DAA s of depth, and for amounts that matter apply the operator's own hashrate judgement, as for any young proof-of-work chain. Listings are not sought before launch in any case (ledger X8) | +| Two certificates at one index observed (C4 evidence) | Suspend credits until the node's view resolves under 3.5 | + +A certified checkpoint overrides the heaviest chain, so fresh hashrate cannot reorganise past `last_certified`; two thirds of 30-day weight can, and 3.1 says what that costs. diff --git a/docs/spec/04-seeds-and-vdf.md b/docs/spec/04-seeds-and-vdf.md new file mode 100644 index 00000000..31204de1 --- /dev/null +++ b/docs/spec/04-seeds-and-vdf.md @@ -0,0 +1,106 @@ +# Igneum protocol specification, section 4: epoch and era seeds through the class-group VDF + +Spec version 0.1, 3 October 2026. Status of this section: Measured for the VDF primitive on one machine (`proto-vdf/`, `docs/bench-log.md` entry "proto-vdf, Wesolowski VDF"); Designed for the pipeline; Open where marked. The class-group code has not been reviewed by a second cryptographer (`proto-vdf/README.md`, open item 1). + +## 4.1 Why a delay + +The hourly program is generated from a seed, the seed comes from a checkpoint, and the checkpoint commits to the blocks before it. The miner who finds the last block before a checkpoint can compute the program that block implies, benchmark it on its own fleet, and withhold the block if the program is bad for it. The review priced this at roughly 130 to 1 for a 30% miner; the prototype's own model gives 6 to 15 to 1 on the burned block for 10% to 40% miners (`proto-vdf/README.md`, grinding table, Measured by Monte Carlo over 2,000,000 epochs against the analytic model, agreement 0.03 blocks). Either way the sign is positive: with no delay, grinding pays and favours the largest miner. + +A miner has about 2 s to decide whether to publish a block under a 1 block/s DAG (approximate, from the block rate). The delay only has to exceed that by a margin no hardware advantage can close. With the delay, withholding has the same expected program as publishing and only burns the block: gain 0 in every row of the grinding table. + +## 4.2 The primitive + +Wesolowski VDF (Efficient Verifiable Delay Functions, EUROCRYPT 2019) in the class group of an imaginary quadratic field with a prime discriminant derived from the input. Chia's construction (`vendor/chiavdf`, commit 7e62ce14, 29 Sep 2026): no trusted setup, group order unknown to everyone, a fresh discriminant per input so nothing can be precomputed. NUDUPL and NUCOMP ported from chiavdf's `qfb_nudupl` and `qfb_nucomp` with the Lehmer partial xgcd; the textbook Cohen 5.4.7 composition and plain duplication are kept as oracles and agree on 15,000 random cases (Measured, `proto-vdf/README.md`). + +``` +input (32 bytes) + D = -HashPrime(tag_D || input), 1024 bits, |D| prime, D = 1 mod 8 + x = (2, 1, (1 - D) / 8), the generator form of Cl(D) + y = x^(2^T) T sequential squarings + l = HashPrime(tag_l || x || y || T), 256 bits Fiat-Shamir challenge + pi = x^floor(2^T / l) + output = SHA-256(tag_out || input || T || y) +proof = (T, y, pi) +verify: rederive D and x, recompute l, check pi^l * x^(2^T mod l) == y, recompute output +``` + +Tags (Implemented in `proto-vdf/src/seed.rs` for the epoch path): `tag_D = "igneum-epoch-discriminant"`, `tag_l = "igneum-vdf-challenge"`, `tag_out = "igneum-program-seed"`. The era path uses `"igneum-era-discriminant"` and `"igneum-era-seed"` (Designed, not yet in code). Prime |D| kills the 2-torsion, the low-order element Wesolowski must exclude. Whether to mirror chiavdf's byte layout for HashPrime exactly, and whether the proof should carry D (the verifier otherwise pays a 17 ms average prime search), are Open (O-4.5). + +| Parameter | Value | Label | +|---|---|---| +| Group | Class group, 1024-bit prime discriminant from the input | Designed (production choice) | +| Fiat-Shamir prime | 256 bits | Designed (Chia uses 264; 2x the 128-bit level) | +| Proof | (T, y, pi), 516 bytes with the prototype serialisation (sign byte plus fixed-width a and b) | Measured. The serialisation should become chiavdf's compact encoding before any wire format is frozen (Open, O-4.5) | +| Proof plan | 12-bit digits, at most 65,536 checkpoints (17 MB) during evaluation | Implemented; proving costs 12 to 13% of evaluation single-threaded and parallelises over residue classes | + +Measured on the Apple M5 Max, one core, rustc 1.69.0, GMP 6.3.0, 3 October 2026 (`docs/bench-log.md`, proto-vdf entry): + +| Group | Squarings/s | T for 600 s | T for 3,600 s | Verify | Proof bytes | +|---|---|---|---|---|---| +| Class group, 1024-bit D | 163,000 | 98.0 million | 588 million | 4.5 ms (12.6 ms including deriving D) | 516 | +| Class group, 2048-bit D | 83,500 | 50.1 million | 301 million | 8.0 ms | 1,028 | +| RSA-2048 (trusted-setup stand-in, timing only, never for production) | 1,257,000 | 754 million | 4.53 billion | 1.4 ms | 512 | + +Full-length run, class 1024: T = 97,126,043, evaluation 585.4 s at 165,900 sq/s, proof 9.1 s on 12 threads (56.8 s on one), verify 4.47 ms. Determinism: the same checkpoint hash gave the same seed and identical proof bytes in two processes at T = 1,000,000; a wrong checkpoint, a flipped seed bit and T + 1 are all rejected. + +## 4.3 The epoch seed pipeline + +Designed. Epoch e is the DAA-score interval `[3,600 e, 3,600 (e + 1))` (section 1.12). + +1. **Seed checkpoint.** `C(e)` is the highest-index checkpoint (section 3, C1) whose checkpoint block has DAA score at most `3,600 e - 1,200`: the latest checkpoint at least 20 minutes of DAA time before the epoch starts. The checkpoint block hash is Kaspa's full header hash, which covers the nonce, as the grinding defence requires (`proto-vdf/README.md`: a hash that covers only the body would let a miner start the VDF while still searching nonces). +2. **Evaluation.** `input = hash(C(e))`, `T = T_epoch`, run 4.2. `program_seed_e` is the 32-byte output; `proof_e` is (T, y, pi). +3. **Program.** `S_e = seed_words_from_bytes(program_seed_e)` (section 1.3.1), program = `generate_from_words(S_e)` (section 1.4). The 1,200-s lead is 2x the reference evaluation time, so a core half as fast as the reference still finishes before the epoch (4.6). +4. **Header.** Every header carries `seed_source = hash(C(e))` for its own epoch (section 2.4). A header is valid under the lottery only if its `seed_source` is a block on its own selected chain at the blue score of checkpoint index `i(C(e))`, and its PoW verifies under the program derived from that block. The proof is not in the header: a node verifies `proof_e` once per epoch (4.5 ms) and caches `S_e`. + +Determinism of step 1 is the point of the rule: which block is "the checkpoint at blue score 30 i" is a function of the header's own past, so two nodes validating the same header derive the same program, and a header mined under a reorged-away checkpoint names a block that is not on its chain and is invalid. Whether `C(e)` must be certified (section 3) or merely be the selected-chain block at that blue score in the header's past is Open (O-4.3): requiring certification couples mining to finality liveness (a stall longer than the lead would stop the program from being derivable), which the design document accepts ("an epoch cannot start without a valid proof") and this specification argues against, because the chain is meant to keep running on plain GHOSTDAG through a finality pause (section 3.7 item 2). The proposal: the selected-chain block at that blue score, certified or not, deep enough (1,200 DAA s plus d) that a reorg across it is a merge-depth-scale event. + +## 4.4 The era seed pipeline + +Designed (design document: "The era draw applies a one-hour delay function to the hash of all blue blocks in the day ending at the last certified checkpoint one epoch before the boundary"). Era n is the DAA-score interval `[15,552,000 n, 15,552,000 (n + 1))` (section 1.12). + +1. `C_era(n)` is the highest-index checkpoint whose block has DAA score at most `15,552,000 n - 7,200` (2 hours of lead, 2x the 1-hour evaluation). +2. `input = Hash(chain_id || n_le64 || h_1 || h_2 || ... || h_m)` where `h_1..h_m` are the hashes of the blue blocks in `C_era(n)`'s past with DAA score in `(daa(C_era(n)) - 86,400, daa(C_era(n))]`, in ascending (blue score, hash) order, and `Hash` is the chain's BLAKE2b-based hash. This is a function of `C_era(n)`'s past, so it is as deterministic as 4.3 step 1. +3. Run 4.2 with `T = T_era = 6 x T_epoch` and the era tags. `E_n` is the output. +4. The era draw of section 1.13.1 consumes `E_n` as its only randomness (`proto-vdf/README.md` open item 6: check that nothing else enters the draw). + +## 4.5 Who evaluates, who verifies + +Every mining node evaluates both VDFs itself: one CPU core for 10 minutes each hour (17% of one core) and once for an hour each era, plus 17 MB of RAM if it also proves. The GPU is untouched. The proof exists for nodes that did not evaluate (light clients, syncing nodes, late starters): 516 bytes and 4.5 ms per epoch. Proofs gossip as their own message type and are also retrievable by `seed_source` from any peer. + +There is no timelord role and no race: unlike Chia, the chain does not wait for the VDF; it uses a VDF output that was fixed 20 minutes (epoch) or 2 hours (era) earlier. "No node evaluates" means no miner is running a CPU, which means nobody is mining. + +## 4.6 T from a reference core rate, fixed at genesis + +Designed (`proto-vdf/README.md`, parameter recommendation). Before genesis, run `vdf bench` (or chiavdf's `vdf_bench`) on every devnet node type that will mine, take the fastest honest single-core NUDUPL rate observed as `r_ref` (squarings per second), and fix at genesis: + +| Constant | Definition | On the M5 Max core (r_ref = 163,000) | +|---|---|---| +| `T_epoch` | 600 x r_ref | 98.0 million | +| `T_era` | 3,600 x r_ref | 588 million | + +Choosing the fastest honest core, not the median, keeps the stated 10 minutes an upper bound for honest nodes and leaves the attacker margin intact. T MUST NOT be derived from on-chain timing, which is manipulable. T is a genesis constant; hardware will get faster over the years and the margin will erode slowly from 300x, which a fixed T covers for decades, and the upgrade path of section 5.7 exists if it is ever needed. + +| Attacker evaluator speed vs reference | Epoch delay | Era delay | Beats the 2-s window | +|---|---|---|---| +| 1x | 600 s | 3,600 s | no, margin 300x | +| 2x | 300 s | 1,800 s | no, 150x | +| 10x | 60 s | 360 s | no, 30x | +| 100x | 6 s | 36 s | no, 3x | +| 300x | 2 s | 12 s | epoch yes, era no | + +Chia's and the Ethereum Foundation's VDF hardware efforts targeted single to low double-digit speedups over CPUs (approximate, from memory, `proto-vdf/README.md`). chiavdf's AVX-512 IFMA assembly path is faster than the prototype's 163,000 sq/s and the x86 devnet nodes will set r_ref higher than this Mac (open item 2 there). + +## 4.7 Fallback for a node without the seed at an epoch boundary + +Open (O-4.4). A node that has not finished evaluating `program_seed_e` when epoch e starts cannot build the program, so it cannot mine the new program or validate new blocks' PoW until it has `y`. It can still receive and relay. Two options: + +| Option | Rule | For | Against | +|---|---|---|---| +| A (proposed by `proto-vdf/README.md`) | The previous epoch's program stays valid for the first N blocks of epoch e, and each header names the program seed it mined under | A slow node keeps mining through the boundary | Two programs are valid at once for N blocks; a miner may pick the better of the two, which reintroduces a small choice the VDF was meant to remove; N is a new parameter | +| B (this specification's preference) | No consensus fallback. The 1,200-s lead is 2x the reference evaluation; a node slower than that takes `(y, pi)` from any peer and verifies in 4.5 ms; `seed_source` in every header tells it which proof to ask for | No second valid program, no new parameter; the proof is 516 bytes | A node isolated from all peers and slower than 2x the reference loses up to the remainder of its own evaluation. Honest nodes evaluate at r_ref or take the proof from peers, so the case is a slow node that is also isolated | + +Decision at gate 3 with the devnet, where the time to first block after an epoch boundary is measured on the slowest node type. + +## 4.8 Open items from the prototype + +Carried into section 6: external review of `classgroup.rs` against chiavdf (O-4.1); reference core choice (O-4.2); whether `C(e)` must be certified and the header field (O-4.3); the fallback (O-4.4); HashPrime layout, carrying D in the proof, compact form encoding (O-4.5); reduction only when `a` exceeds 8 limbs as chiavdf does, a further speedup to port (O-4.6); no fuzzing of `deserialize` on hostile bytes beyond validity checks, no measurement on NVIDIA or AMD hosts' CPUs (O-4.7). diff --git a/docs/spec/05-fees-and-economics.md b/docs/spec/05-fees-and-economics.md new file mode 100644 index 00000000..ec0937e2 --- /dev/null +++ b/docs/spec/05-fees-and-economics.md @@ -0,0 +1,84 @@ +# Igneum protocol specification, section 5: fees, splits, burns, the development fund, signalling + +Spec version 0.1, 3 October 2026. Status of this section: Designed (design document, "The token", "Finality rule, version 2: Fees", decisions table "What do I get for being early?"). Nothing here is implemented or measured. Emission itself is section 2.5. + +Two rules frame everything: there is no stake anywhere in consensus, and not one unit of emission goes to a treasury, a fund, a founder or a stake. Development is funded from fees, under miner signalling. + +## 5.1 Base fee + +Designed. Every transaction pays a base fee in both gas dimensions: + +| Dimension | What it meters | Who sets it | +|---|---|---| +| Execution gas | EVM execution, Ethereum's rule | Ethereum's EIP-1559-style base fee over the ordered sequence | +| Proving-cost gas | Proving cycles the transaction will cost the provers | A second base fee adjusted per block from the unproven backlog, smoothed over the difficulty window (section 2.3) so cards do not flip between hashing and proving every block | + +The base fee in both dimensions is **burned in full**. A miner cannot stuff blocks with its own transactions for free; wash gas loses its whole base fee (ledger E3). The proving-cost budget per block is a consensus constant set from measured prover throughput (phase 2 gate: one shard on a 12 GB card in about 20 s, Target, unmeasured, ledger P1), so a transaction that is cheap to run and brutal to prove cannot stall the provers for everyone. How the node folds the proving-cost dimension into the quoted gas price so `eth_estimateGas` keeps working is an execution-layer matter and Open (ledger P5, O-2.8). + +## 5.2 Priority fee split + +Designed. The priority fee (tip) of every executed transaction splits: + +| Share | Recipient | Rule | +|---|---|---| +| 65% | Block producer and provers | The producer of the block in which the first copy executed (section 2.6) and the provers of that block, in the proportion the proving protocol defines (forward reference) | +| 15% | Developer | Attributed per call frame by gas consumed, to the developer address registered for the called contract at deployment. A frame in an unregistered contract sends its share to the burn | +| 15% | Burn | | +| 5% | Development fund | Section 5.5 | + +Per-call-frame attribution: a transaction that calls contract A, which calls B, which calls C, pays 15% of the tip in proportion to the gas each frame consumed, to A's, B's and C's registrations respectively. A frame's own gas excludes the gas of frames it called. + +Factory rule: a contract deployed by another contract (CREATE or CREATE2 from a contract) inherits the deploying contract's registration unless it re-registers in the same transaction. A contract deployed from an externally owned account registers the address the deployment names, or none. + +Self-dealing: a developer who also mines the including block collects 65% plus 15% and loses 20% of the tip plus the whole base fee, so wash gas is a guaranteed loss; it can still inflate a "gas earned" figure on a leaderboard, so no explorer ranking should use raw developer share without a self-dealing filter (ledger E3; not a consensus rule). + +## 5.3 Proving pool + +Designed. The 20% emission share (section 2.5) and the provers' part of the 65% tip share are paid per block as a fixed amount for that block, divided among the block's shards by consensus proving cost, so a stuffed block earns no more than an honest one. Shards are claimed with a small bond, slashed on a bad or late proof, and a withheld shard reopens with a rising fee until someone proves it (design document, security model table). First-come claiming favours the lowest-latency datacentre (ledger P8, C9); sortition of shards by a VRF keyed to the block, weighted by past proving, with open claiming after a timeout, is the scheduled replacement and is Open (O-5.1). Bond size and timeout are Open (O-5.6). An unproven block delays only its proof; execution and the 30-s lock do not wait for it (ledger P9). + +## 5.4 External job market + +Designed. At launch external proving jobs are paid on the customer's chain, in the customer's currency, to a payout contract keyed by miner address, because Igneum cannot yet see Ethereum; the customer chain's own bond and slashing apply (design document, "The first six months"). When the job market settles on Igneum, which needs the proof bridge (phase 2 consensus proof, ledger P4, E7), every job fee paid in IGN splits: + +| Share | Recipient | +|---|---| +| 85% | The prover that won the job | +| 10% | Burn | +| 5% | Development fund | + +The burn applies only to IGN-settled jobs (ledger P10). The timing of the switch and the bridge's trust model are Open (O-5.2). + +## 5.5 Development fund + +Designed. A contract, not an account. It receives 5% of every priority fee (5.2) and 5% of every IGN-settled external job fee (5.4), and never a unit of emission. Spending needs miner signalling: a proposal names a recipient and an amount; blocks signal for it; it passes when at least **60% of blue blocks** over a signalling window of **1,209,600 DAA s (two weeks)** carry the signal (the BIP 9 model). Miners hold the purse without filling it; users and rollups fill it. Before there is usage the founder's company carries development from the official client's 1% dev fee, its pool and its proving business (ledger E5: the homepage's "not one coin to a founder" is true of emission and false of the dev fee, and must say so). The fund contract's upgrade and key policy is published before launch (ledger G4, O-5.4). + +## 5.6 No stake, no treasury + +- No consensus rule reads a balance. No bond, deposit or stake affects block validity, vote weight, fork choice or finality. The prover shard bond of 5.3 is an execution-layer escrow, not a consensus weight. +- No unit of emission goes anywhere but the block producer (80%) and the proving pool (20%). +- The team holds no consensus key and no purse (hostile review table, "A team that controls a treasury"). + +## 5.7 Upgrade signalling + +Designed. Consensus rules change only through new code that activates when at least **90% of blue blocks** in a signalling window carry the upgrade's signal. The window length, the activation delay after the window closes, and the proposal identifier format in the block (candidate: a version-bits field in the header `version`, Kaspa's field) are Open (O-5.3). The specification is updated to describe activated rules, not the other way round (section 0.4). Automatic changes (epoch programs, day keys, era draws, the sortition switch at 8,192 voters, dataset growth) are not upgrades and need no signal. + +Emergency path: a soundness bug in the proof system cannot wait for a signalling window. The rule that contains it is execution-layer: a full node executes natively and rejects any proof whose claimed state root differs from its own execution, so a forged proof is a light-client problem and not a chain split (ledger P7). Writing the fix is human work; activating it is this section's signal. + +## 5.8 Signalling encoding + +Designed, Open (O-5.3). Each block carries a bitfield; bit b set means "this block signals for proposal b". Proposals are registered by a transaction naming the bit, the kind (fund spend or upgrade), the threshold (60% or 90%) and the window start in DAA score. The count is over blue blocks with DAA score in the window. A proposal that fails may be re-registered. + +## 5.9 Parameters in this section + +| Parameter | Value | Label | +|---|---|---| +| Base fee burn | 100%, both dimensions | Designed | +| Priority fee split | 65 / 15 / 15 / 5 | Designed | +| External job split (IGN-settled) | 85 / 10 / 5 | Designed | +| Fund signalling threshold | 60% of blue blocks | Designed | +| Fund signalling window | 1,209,600 DAA s | Designed (design document: "over two weeks") | +| Upgrade signalling threshold | 90% of blue blocks | Designed | +| Upgrade window, activation delay | not set | Open (O-5.3) | +| Shard bond, claim timeout | not set | Open (O-5.6) | +| Proving-cost budget per block | from the phase 2 measurement | Target | +| Emission to any treasury | 0 | Designed | diff --git a/docs/spec/06-open-items.md b/docs/spec/06-open-items.md new file mode 100644 index 00000000..25540a53 --- /dev/null +++ b/docs/spec/06-open-items.md @@ -0,0 +1,105 @@ +# Igneum protocol specification, section 6: open items + +Spec version 0.1, 3 October 2026. Status of this section: Open by definition. Every parameter or rule in sections 1 to 5 that is unmeasured, unreviewed or marked open is listed here once, with the experiment or decision that closes it and the gate it belongs to. Sources: the Open entries of `docs/fud-ledger.md`, the "unproven" and "not demonstrated" lists of `proto-metal/MEMHARD.md` and `proto-metal/TESTS.md`, the open items of `proto-vdf/README.md`, the "cannot tell us" sections of `sim/results.md` and `sim/results_v2.md`, and the notes of `docs/fork-map.md`. + +Gates: gate 1 = the lottery hash frozen (section 1 at version 1.0; cryptographer). Gate 2 = the fork runs a devnet at Igneum parameters (consensus engineer). Gate 3 = finality and seeds reviewed externally (section 3 at 1.0; cryptographer). Gate 4 = 1,000 independent miners for 30 days (miner-community lead). Phase 2 = the proving benchmark, out of scope for 0.1 but named where a parameter here waits on it. Items that are not a gate's business are marked "decision" with an owner. + +An item closes when its measurement is in `docs/bench-log.md` or its decision is in the design document, and the section it belongs to is updated at a MAJOR or MINOR step (section 0.4). + +## 6.1 Section 1, the lottery hash + +| Id | Item | What closes it | Gate | +|---|---|---|---| +| O-1.1 | Seed derivation is FNV-1a 64 plus SplitMix64, ad hoc (ledger M7) | Replace with a standard hash (candidate: BLAKE2b or SHA-256 into the eight words); re-run the weak-program census of O-1.3 on the result; new vectors | 1 | +| O-1.2 | Load count per program is not fixed, only its expectation; 40 to 232 loads per hash over 10,000 programs; hash rate scales with it (ledger M5) | Exact-load-count rule in the generator (draw exactly 16 load positions), re-fuzz 10,000 programs, confirm the loads-per-hash column is constant | 1 | +| O-1.3 | Weak programs: 3 seeds measured for bias out of the population; nothing rejects a program whose `or` chain saturates a register or whose load addresses collapse (ledger M6) | Census of at least 10^5 programs on the CPU interpreter: bit bias and avalanche on 2^12 nonces, distinct load addresses per hash, registers saturated by `or`, nonce-independent registers; a rejection rule in the generator; a bound on the advantage of a miner grinding across candidate blocks | 1 | +| O-1.4 | No cryptanalysis of the hash as a lottery: the init is a bijection of the nonce per register; the body is add-rotate-xor-multiply with `or`; no analysis of shortcuts or seed influence (`TESTS.md` section 8 item 1, ledger M7) | External review scoped to the fair-lottery properties of section 1.1 (no shortcut cheaper than honest evaluation, no exploitable bias), not general hash security | 1 | +| O-1.5 | Memory hardness is measured on Apple only: shortcut ratio 4.8x slower on the M5 Max; not measured on NVIDIA or AMD (`MEMHARD.md` section 3 item 1) | Run `--inline-dataset` (the CUDA and OpenCL equivalents) on the RTX 5090 and on an AMD discrete card; the ratio must stay above 1 on every card that mines | 1 | +| O-1.6 | Time-memory trade-offs between the two measured points, and a smarter attacker kernel (hoisting the t-only part of round 0, caching hot lines), were not attempted (`MEMHARD.md` item 2) | Draw the curve: store fraction f of the dataset, recompute the rest, for f in {0, 1/8, 1/4, 1/2, 1}; write the hoisting attacker and measure it | 1 | +| O-1.7 | Mixer M_r and the chained-block cache have had no cryptanalysis; weak `ROT` draws (all equal) possible and untested (`MEMHARD.md` item 3) | External review of section 1.8; a weak-key check on the mixer draw with a rejection rule if needed; decide whether to adopt RandomX's SuperscalarHash design instead | 1 | +| O-1.8 | Distinct cache lines touched per hash not measured; cache-resident shortcut ruled out by argument (832 of 4,194,304 lines per hash, 26,624 per unit) not by census (`MEMHARD.md` item 5); inlining-the-cache cost is an estimate (item 4) | Census over 2^20 nonces; write the cache-inlining attacker kernel and measure it | 1 | +| O-1.9 | The hash does not commit to the block header; the packs use the program seed as the init words (section 1.6) | Fix the pre-PoW header hash in section 2 (fork point a5), adopt the proposed `I = seed_words_from_bytes("igneum-block/" || H || nonce_hi)`, produce a header-bound vector set | 1, with 2 | +| O-1.10 | The day key derivation is not in the design document (section 1.12) | Decision: adopt the proposed `"igneum-day/" || d || program_seed of the day's first epoch`, or another; owner cryptographer | 1 | +| O-1.11 | Era parameter draw: no procedure exists; bounds proposed in section 1.13.1 | Design review of the bounds; every drawn configuration must pass the fuzz and stats runs of section 1.15; decide what "memory pattern" means for the draw | 1 | +| O-1.12 | Instruction-family reserve: no list exists; candidates in section 1.13.2 | Each candidate family passes cross-vendor conformance with its own edge vectors; fix the order and `W_new` at genesis | 1 | +| O-1.13 | Dataset growth: 2 GiB plus 0.5 GiB per year is Designed; the index mapping for non-power-of-two sizes changes the vectors (section 1.13.3) | Decision between the multiply-shift mapping (new vectors) and power-of-two steps (`AND MASK` kept); measure hash rate and shortcut ratio at 2 GiB on three vendors | 1 | +| O-1.14 | CPU verify measured on one M5 Max core only (Rust 0.41 to 0.58 ms); the 10 ms gate is for a 2019-class laptop core (ledger M9, design document "Three experiments") | Measure the Rust verifier on a 2019-class core at 104 and 144 loads with the memory-hard cache; if over 10 ms, apply lever (a) | 1 | +| O-1.15 | Cross-vendor conformance: AMD is one integrated gfx1036 chip; no discrete AMD card has run anything; Intel unmentioned; the fuzz and edge sets have run on Metal and the CPU references only (ledger M8) | Run the seven commands of `proto-opencl/README.md` on an AMD discrete card; run the 10,200-program fuzz set and the 14 edge programs on NVIDIA and AMD; Intel Arc after | 1 | +| O-1.16 | Runtime kernel compilation on real rigs (NVRTC, ROCm, mixed-generation cards, every hour) is unmeasured; the 18 to 52 ms compile figures are Metal on one Mac (ledger M11) | Miner client prototype on a multi-card rig; compile time and failure rate per hour | 4 | +| O-1.17 | ASIC gain under 2x is a Target, not a measurement (ledger M1) | Public benchmark with a leaderboard by card model and a standing bounty for any chip design beating a GPU by more than 2x, January 2027; the claim stays a target until the bounty has gone unclaimed for years | phase 2, standing | +| O-1.18 | The 64-bit output and the 256-bit target space (section 1.10) | See O-2.4 | 2 | +| O-1.19 | Epoch length 3,600 DAA s against difficulty tracking (section 1.12) | See O-2.1 | 2 | +| O-1.20 | The GPU cache-fill time varied 0.6 to 2.1 ms across runs with identical code; cause not isolated (`MEMHARD.md` item 7) | Isolate (GPU clock state suspected); no consensus consequence | none, note | + +## 6.2 Section 2, the ordering layer + +| Id | Item | What closes it | Gate | +|---|---|---|---| +| O-2.1 | DAA window against per-epoch hash-rate steps: 2,641 s Kaspa window, candidate 1,800 s, or equalise cost per program (fork map c1, c2) | simpa run with programs stepping 30% every 3,600 blocks at both windows, with and without O-1.2's exact load count | 2 | +| O-2.2 | GHOSTDAG k = 18 at 1 BPS is Kaspa's table value; k was never exercised on the devnet (single serial miner) | Two-miner devnet at Igneum parameters; measure parallel-block rate and red-block rate | 2 | +| O-2.3 | Pruning depth must stay above the longest checkpoint gap and never pass the latest lock (section 2.1, F3) | Define the bound once the first-month rule (O-3.1) and the stall behaviour (section 3.3.1) fix the longest gap; implement in `block_depth.rs` and the virtual processor | 2, with 3 | +| O-2.4 | Mapping of the 64-bit hash into the 256-bit target and the block level for pruning proofs (fork map a3) | Decision: `target64 = target256 >> 192`, `level = leading_zeros(hash64)` capped at 64, or widen the output; check `calc_level_from_pow` callers | 2 | +| O-2.5 | Header validation gains a chain-state dependency (the epoch seed) and must stay deterministic from headers plus certificates during IBD and pruning-proof validation (fork map a4, risk High) | Implement `seed_source` validation; sync a node from genesis and from a pruning point on the devnet | 2 | +| O-2.6 | Base unit: 10^18 (EVM) or 10^8 (fork map b2) | Decision, owner execution engineer with consensus engineer | 2 | +| O-2.7 | Apportioning the DAA-score increment among several blue blocks in one mergeset at higher block rates (section 2.5) | Define before the first block-rate step; at 1 BPS the mergeset is usually one block | 2, before the 4 BPS step | +| O-2.8 | EVM semantics on a DAG: block number, timestamp, blockhash, coinbase, prevrandao over the ordered sequence; folding the proving-cost dimension into the quoted gas price (ledger P5) | Execution-layer specification; prevrandao from the epoch VDF is the candidate | decision, execution engineer | +| O-2.9 | `docs/fork-divergence.md` does not exist | Write it as the fork is made | 2 | +| O-2.10 | Block-rate steps to 4 and 10 BPS each re-derive k, parents and mergeset limit and are hard forks (hostile review table) | Each step has its own test campaign, as Kaspa's Crescendo | later, consensus engineer | + +## 6.3 Section 3, finality + +| Id | Item | What closes it | Gate | +|---|---|---|---| +| O-3.1 | The first month: no weight exists; an attacker producing 75% of blocks from day 2 crosses 2/3 on day 9 (ledger F1). C5 says first hour; this specification proposes first 30 days | Launch-month simulation (not yet in `sim/`); decision; litepaper states the rule either way | 3 | +| O-3.2 | Checkpoint depth d = 60 is a placeholder (ledger F7) | Devnet with regional latency records the reorg-depth distribution at each block rate; set d so a vote split at one index is rare | 3 | +| O-3.3 | Participation grinding through the certificate bitmap (ledger F3): aggregators and block producers can shape rivals' participation | Choose among the candidate fixes of section 3.4; simulate with a hostile aggregator | 3 | +| O-3.4 | Certificate grace value; must be at least 3x the worst honest one-way delay (section 3.3, Q4) | Measure one-way delays on the devnet across regions; set grace | 3 | +| O-3.5 | VRF construction for aggregator selection and the binomial sub-user sortition above 8,192 voters (S1, S2) | Specify (candidate: BLS-based VRF on the vote key, Algorand's binomial sampling); simulate the threshold on sampled weight | 3 | +| O-3.6 | What a node does with two valid certificates at one index after a partition heals; post-heal fork choice is unmodelled (`sim/results_v2.md`, "cannot tell us") | Adopt or replace the proposal in section 3.5; devnet partition-and-heal test | 3 | +| O-3.7 | The presence window is an eclipse vector (ledger F2); the floor closes it in the model for 1, 2 and 4 h at a 34% attacker (F2) but the model grants the attacker the eclipse for free | Devnet with a single-node eclipse recording whether conflicting locks appear; choice between the 2-hour window and a 7-day silent-key rule | 3 | +| O-3.8 | The simulation has no DAG: conflict counts are index collisions; red blocks, merge under the 3,600-s bound and the finality overlay's effect on GHOSTDAG's guarantees are unmodelled (ledger C4, F8) | Devnet runs with the finality module on and off; a churn and adversary simulation driven by real pool-hashrate traces from mid-cap GPU coins (design document, "Three experiments") | 3 | +| O-3.9 | Model assumptions that move the numbers: perfect or instant DAA retarget (real lag of the order of an hour, approximate), uptime 97% / 99.5% is a guess, silent sets random by key not by pool or region, keys are free, VRF noise absent (`sim/results.md` and `results_v2.md`) | Re-run `finality_v2.py` with a DAA lag model, a top-pool silent set and a regional silent set; price keys through the P2P layer | 3 | +| O-3.10 | Equivocation evidence is detected only at the heal and is forward-looking; certificates signed by equivocators are not revoked in the model | Decide revocation (section 3.5 proposal) and simulate | 3 | +| O-3.11 | Key succession (W5): replay protection and what happens if both keys mine after the message | Specify the message (old key, new key, DAA score, signature) and that blocks naming the old key after inclusion earn nothing | 3 | +| O-3.12 | Vote message and certificate wire formats, aggregation rules, bitmap size at 10^4 keys | Specify with the P2P layer | 3 | +| O-3.13 | `finality_active` flag semantics and exchange guidance (section 3.9) are Designed and untested | Devnet stall test; exchange guidance reviewed by an operator | 3, 4 | + +## 6.4 Section 4, seeds and the VDF + +| Id | Item | What closes it | Gate | +|---|---|---|---| +| O-4.1 | `proto-vdf/src/classgroup.rs` (NUDUPL, NUCOMP, Lehmer partial xgcd) agrees with the textbook algorithms on 15,000 cases but has not been read by a second cryptographer against `vendor/chiavdf` | External review | 3 | +| O-4.2 | Reference core rate r_ref: 163,000 sq/s is this Mac with this code; x86 AVX-512 IFMA and chiavdf's assembly path differ | Run `vdf bench` on every devnet node type that will mine; fix `T_epoch = 600 r_ref`, `T_era = 3,600 r_ref` at genesis | 3 | +| O-4.3 | Whether the seed checkpoint must be certified or may be the uncertified selected-chain block at that blue score; the `seed_source` header field (section 4.3) | Decision at gate 3 with the devnet; this specification proposes uncertified-allowed so a finality pause does not stop mining | 3 | +| O-4.4 | Fallback for a node without the seed at an epoch boundary: option A (previous program valid for N blocks) or B (none) (section 4.7) | Measure time to first block after a boundary on the slowest devnet node type; decide | 3 | +| O-4.5 | HashPrime byte layout (mirror chiavdf or not); carry D in the proof so the verifier skips the 17 ms prime search; compact form encoding before the wire format freezes | Decision, cryptographer | 3 | +| O-4.6 | Reduction runs every squaring; chiavdf reduces only when `a` exceeds 8 limbs | Port; re-measure r_ref | 3 | +| O-4.7 | No fuzzing of `deserialize` on hostile bytes beyond validity checks; no VDF rate measured on NVIDIA or AMD hosts' CPUs | Fuzz the deserializer; bench on the devnet hosts (same run as O-4.2) | 3 | +| O-4.8 | The era draw must take the VDF output as its only randomness (`proto-vdf/README.md` item 6) | Check against section 1.13.1 once O-1.11 is fixed | 1, with 3 | +| O-4.9 | Grinding model assumptions: advantage uniform on 0 to 15% per program (the review's measured range), top-quartile keep rule, one block burned per candidate | Re-run `vdf grind` with the advantage distribution from the O-1.3 census | 3 | + +## 6.5 Section 5, fees and economics + +| Id | Item | What closes it | Gate | +|---|---|---|---| +| O-5.1 | First-come shard claiming lets the lowest-latency datacentre take every shard (ledger P8, C9) | Specify sortition of shards (VRF keyed to the block, weighted by past proving, open claiming after a timeout); measure on the phase 4 devnet | decision, cryptographer; measure in 4 | +| O-5.2 | External job settlement: on the customer's chain at launch, in IGN with a 10% burn once the proof bridge exists; the bridge's trust model at launch (committee-attested or proof-verified) is undecided (ledger P10, E7) | Decision before phase 4; the litepaper stops presenting the proof bridge as a genesis feature until then | decision, execution engineer | +| O-5.3 | Signalling encoding: version bits or a separate field, window length and activation delay for upgrades, proposal registration format (section 5.7, 5.8) | Specify with the consensus engineer; BIP 9 is the model | 2 | +| O-5.4 | The development fund contract's upgrade and key policy; every genesis contract's key policy (ledger G4) | Publish before launch | 4 | +| O-5.5 | The proving-cost gas dimension: the metering table per opcode and precompile, and the per-block budget from the phase 2 measurement | Phase 2 benchmark; execution-layer specification | phase 2 | +| O-5.6 | Shard bond size and claim timeout (ledger P9) | Set on the phase 4 devnet | 4 | +| O-5.7 | The provers' proportion of the 65% tip share (section 5.2) | Defined by the chunked proving protocol | phase 2 | +| O-5.8 | Developer registration format at deployment and the re-registration transaction (section 5.2) | Execution-layer specification | decision, execution engineer | + +## 6.6 Count + +| Section | Open items | +|---|---| +| 1 | 20 | +| 2 | 10 | +| 3 | 13 | +| 4 | 9 | +| 5 | 8 | +| Total | 60 | + +By gate (an item shared between two gates is counted at the earlier one): 16 belong to gate 1, 11 to gate 2, 21 to gate 3, 4 to gate 4, 3 to phase 2, 4 are decisions with a named owner outside a gate, and 1 (O-1.20) is a note with no consensus consequence. diff --git a/docs/spec/README.md b/docs/spec/README.md new file mode 100644 index 00000000..442ee6c8 --- /dev/null +++ b/docs/spec/README.md @@ -0,0 +1,17 @@ +# Igneum protocol specification, version 0.1 + +3 October 2026. Written so that a stranger can reimplement a node, a miner and a verifier and reach the same bits. Conventions, number labels (Measured, Implemented, Designed, Target) and the break-submission route are in section 0. + +| Section | File | Status | One line | +|---|---|---|---| +| 0 | `00-overview.md` | Designed | Scope, normative language, what is implemented and what is designed, versioning, how to submit a break | +| 1 | `01-lottery-hash.md` | Measured (construction, vectors on three GPU vendors and two CPU references); Designed (header binding, day key, growth, era schedule); Open where marked | Seed words, generator, memory-hard cache and items, 32-lane unit and the wave64 rule, CPU verifier, epoch, day and era schedules, test vectors, determinism, conformance, the prototype-value list | +| 2 | `02-consensus.md` | Designed | The ordering layer as a delta on rusty-kaspa: 1 BPS and the steps, k, merge depth, DAA, header fields, emission, duplicate inclusion | +| 3 | `03-finality.md` | Designed (simulated without a DAG, not implemented, not reviewed) | Weight W1 to W5, checkpoints C1 to C5, quorum Q1 to Q4 with the 56.7% floor and the simulation that justifies it, sortition, fork choice, equivocation, residual risks, the first month, exchange guidance | +| 4 | `04-seeds-and-vdf.md` | Measured (primitive, one machine); Designed (pipeline); Open (fallback) | Class-group Wesolowski VDF, T from a genesis reference rate, 20-minute and 2-hour leads, proof format, who evaluates, fallback options | +| 5 | `05-fees-and-economics.md` | Designed | Base fee burned, 65/15/15/5 with per-frame attribution and the factory rule, proving pool, job market, development fund at 60%, upgrades at 90%, no stake, no treasury | +| 6 | `06-open-items.md` | Open | 60 items, each with the experiment or decision that closes it and its gate | + +Prototype values (carried by the implementation today, part of the test vectors, confirmed or replaced at gate 1): seed derivation; 64 instructions x 8 iterations; 8 registers; the op weights including the 25% load weight; the output fold rotations; the 256 MiB cache, 64 lines per segment and ChaCha12; 8 item rounds and the mixer shape; the 1 GiB pack dataset against the 2 GiB genesis size; the 3,600 DAA s epoch; the era length, draw bounds and reserve list; the header binding of the init words. The full table with what fixes each is section 1.16. + +Normative code: `igneum-pow/`. Measurements: `docs/bench-log.md`. Criticisms and their status: `docs/fud-ledger.md`. Simulations: `sim/`. Prototypes: `proto-metal/`, `proto-cuda/`, `proto-opencl/`, `proto-vdf/`.