Merge d1-record 1254833b into master (gate: green on 36a0387c, recorded by tools/ci/pre-push.sh; landed on the box mirror)
This commit is contained in:
commit
0d94049bb2
2 changed files with 148 additions and 2 deletions
|
|
@ -41,6 +41,22 @@ Owner: coordinator (Counter ASIC 3.0) with the hash lane and the node lane.
|
|||
- [ ] Central measure served: cost per accepted unit of work = (annualised hardware + power + hosting, failures, fees) / annual accepted work.
|
||||
- Pass: an independent operator reproduces the baseline within declared tolerances from the served kit alone.
|
||||
|
||||
### D1 record, 8 October 2026 (the node lane; FROZEN 23:17 BST by the shipper's decision, candidate (c))
|
||||
|
||||
The frozen, reproducible baseline is the class v6 object commit e8773ff5 on the fork branch `class-v6-node-review` (both node mirrors; its tree merges the live devnet-4 literal ef0f2ed8), the generator tree class-v6 1a938abe4 (generator 6, `ProgramClass::V6`), built against the parent `ca3-v4-node` at 2ceb09be9 (key-succession-pin, off master 07e5fa680). Candidate (b) 83bb8ea6 on `class-v6-node-b` (the accept-fix tree abdecf2e6, fingerprint 6cdd922a4bc7f1bc…) is the named fallback; candidate (a) 3bd95052 was never eligible. The signing object F0 is program id 2a1d6caab4c24564 drawn at (epoch af89be5d…, era edc4fa84…, day 20730, the node1 state abb58003… with root 1c583d35…), read equal from the served pair's engine (`igneum-miner program-id`, 21:41Z) and the freeze CLI at 1a938abe4 (`igneum-pow show`, build-5, 21:48 BST); F0 signed and on master at 900cf82a. Every value here is pinned by digest in the kit the node serves (`packaging/pow-freeze.txt` line `5f4d6dc6… 1a938abe class-v6-review`, `packaging/d1-freeze.txt`, embedded at build time and printed at start as `D1 baseline: ...`; served on `igneum_getManifest` from node 2.0.1).
|
||||
|
||||
| Component | Digest | Source | Label |
|
||||
|---|---|---|---|
|
||||
| Generator | 5f4d6dc6199294db89042171004e6420c1e5791e716d4b39b73d375818c10b6f | igneum-pow `src/*.rs` at class-v6 1a938abe4 (the hash lane's final tree, 21:24 BST; generator 6, `ProgramClass::V6`, with the class v6 rules: the 64-register window, the index fold, the re-weight table rw1, the lossy band at base, the family gate 6.13; `CLASS_V6_FLAG_NOWIN` defined and off; no counter-asic-4 code path); the fingerprint the served pair prints at start (`IGNEUM_POW_FINGERPRINT=5f4d6dc6…`), as `consensus/pow/build.rs` computes it | Measured (the census sheet: fold, rw, foldrw, win, all PASS; rw2 FAIL and not carried) |
|
||||
| Verifier | shard 0x2b1a81cb413236cf063077b46ed3111628f6c41036bcf6e23ee4cbbf5679ef7a, aggregator 0x474678f35f7545db28055d5e5bbc308231d84a5a072202087a2a8d5b09123896 | the embedded program ids of the served release (the ELF manifest pinned 5 October 2026), in consensus from block zero on igneum-devnet-4 by `verifier_in_consensus` | Measured (0.67 to 0.71 s a shard proof on one EPYC core; the re-pinned guest afd952e89 with new ids waits for the next genesis or a key-succession field, the coordinator's ruling of 8 October 19:00) |
|
||||
| Dataset policy | 70a6c703787d5d75cdbc486b34acf3eb901f4188bae3e948c94274d5f4ac4bba | the object's `class_v6_dataset_steps` ([(0, 4,096 MiB)]: the 4 GiB floor from genesis, no later step) and `class_v6_family_flags` (0xf: the window, the fold, rw1, the lossy band) as the consensus digest arm hashes them; items = floor(bytes / 64), the ds55 mapping taking any count | Designed (main's default of 8 October 2026, 20:00 BST); the energy row Measured on a locked 5090 (+8 percent at 4 GiB, inside the 10 percent budget) |
|
||||
| Compiler config | 1.99.0-b940084d7;cargo-5f94df478;sccache-0.18.0;zig-0.17.0;profile.release:lto=thin,strip=true,overflow-checks=true;target:x86_64-unknown-linux-gnu | `rust-toolchain.toml` channel 1.99.0, the workspace release profile, the box's sccache and zig; the gate artefact native, the seed and rig pairs cross-built at glibc 2.35 and 2.31 | Measured (the build-remote read-back on igneum-build-2) |
|
||||
| Harness | v5-fasttime 92bf6a7f (records cross-v6-617cb441-3.{json,log}) | the fast-time crossing on the object's base 617cb441: v4 at e0 and e1, rung 1 by signal, v5 by its floor, v6 at byte 8 from its floor, the template at every boundary, the stale node's refusals, the restart step, the cold restart on four nodes (SUMMARY PASS 17:30:44 BST); the finality attack lines (split50 v3 known-failed, v4 pause, v4 recovery, split70) on bee41b5e's pair in sim/results_v2.md "Rule v4" | Measured; the five-together case named with its gap: proof queues and seed transitions not exercised in the same network as the crossing and the finality attack, owed to the next harness commit (the fast-time lane's morning item) |
|
||||
|
||||
Gates on the object, as read back on build-1 for e8773ff5 (the d1c chain, 20:5x to 21:3x BST): the six crate suites at gate priority (kaspa-consensus-core 184, kaspa-consensus 143, igneum-exec 74, igneum-miner 30, kaspa-p2p-flows 38, kaspa-pow 19 with the igneum-pow feature; 0 failed), the pair placed on build-1 and build-2 as `/srv/artefacts/200-e8773ff5/node-lane` (igneumd 5d7dfb311f50e5df…, igneum-miner 2ee4893de1735a6e…, the commit string read back), the canary set (the devnet-4 digest be5f4068… on two empty nodes, eth_chainId 0x1171, the class signal line "object version 6", the devnet-3 override refused, the cross-dial refused); the census re-read on the chain-seed draw (PASS: 2a1d6caab4c24564 accepted at attempt 0, min site ratio 0.99988, worst free bit 3.06 sigma, no site over 6 sigma, 256 of 256 seeds at r 0.129; the window trace on the signing program 1.0026 per site and 1.02x of its foldrw control on the item share; master f518e344); research lane D's family gate 6.13 (PASS under both era inputs, fddf641fd and its amendment); the same-work read on context 02 (node, CPU reference and CUDA equal on every nonce of 1,048,576 per phase, the day switch seen; the pool reader's rerun the one open cell at the freeze, a reader-install fault, not the engine); the five chain-paired packs under `/srv/artefacts/packs` (hl-v6-all-cs 0x2a1d6caab4c24564 and the four beside it) with the CPU reference b77c61d8…. Every floor never on every compiled network, so no live digest moves with the landing. The spec: `docs/spec/01-lottery-hash.md` at version 0.3 (section 1.0 the frozen object, 1.4 generator 6's rules with rw1, 1.8.5 W = 4 and the ds55 mapping, 1.13.3 the dataset policy, 1.16 the acceptance rule, the per-program rule as served, the fixed parameters and the excluded items).
|
||||
|
||||
The classes read today and closed in the node (release-0.3.25-node, then release-2.0.0-node): the cold-restart deadlock (a sink-age guard no paused chain could satisfy; f8da7515), the genesis sink (acaf08b0), the short-captures snapshot (acaf08b0; the mid-epoch resume rule in node 2.0.0), the class v5 genesis stream (6e04f7fc), the miner's exec-rpc default per network (f41f48a7), the old-object blocks under every kept branch after a stall (the fresh start; rule 20's addition: a stall stops every miner at once), the 8-decimal emission literal in an 18-decimal object (4cdcc488, the re-cut). Pass line for an independent operator: the served kit (the pair at the sha above, `igneum-miner program-id` beside the freeze CLI's `igneum-pow show` on the conformance seeds, the manifest's five digests) reproduces this table from the kit alone.
|
||||
|
||||
## D2. Two architectural experiments, not twenty knobs
|
||||
|
||||
Owner: class v6 invention lane (a) and the multi-family adversary lane (b).
|
||||
|
|
|
|||
|
|
@ -1,11 +1,25 @@
|
|||
# Igneum protocol specification, section 1: the lottery hash
|
||||
|
||||
Spec version 0.2, 4 October 2026 (0.1 on 3 October 2026). Status of this section: Measured for the construction as implemented (Implemented values with test vectors on three GPU vendors and a CPU reference); Designed for the header binding, the day key, the dataset growth and the era schedule; Open where marked. Version 0.2 adopts generator version 2 (sections 1.4.2, 1.4.3 and 1.4.6: exact load count, fresh-source loads, program acceptance) from the weak-program census of `docs/analysis/weak-program-census-2026-10-03.md`; every vector of version 0.1 is retired and re-cut (section 1.17).
|
||||
Spec version 0.3, 8 October 2026 (0.2 on 4 October 2026, 0.1 on 3 October 2026). Version 0.3 freezes the Igneum 2.0 Deliverable 1 object by digest (section 1.0: generator, verifier, dataset policy, compiler configuration, harness), adopts generator version 6 for epochs from the class v6 floor (1.4.8; version 0.2's generator 5 stands below it), fixes the dataset policy (1.8.5, 1.13.3), replaces the prototype table with the parameters the freeze fixes (1.16) and names the class v6 vectors as Open with their clock (1.17). Status of this section: Measured for the construction as implemented (Implemented values with test vectors on three GPU vendors and a CPU reference); Designed for the header binding, the day key, the dataset growth and the era schedule; Open where marked. Version 0.2 adopts generator version 2 (sections 1.4.2, 1.4.3 and 1.4.6: exact load count, fresh-source loads, program acceptance) from the weak-program census of `docs/analysis/weak-program-census-2026-10-03.md`; every vector of version 0.1 is retired and re-cut (section 1.17).
|
||||
|
||||
Normative implementation: `igneum-pow/src/{seed,generator,accept,memhard,verify}.rs`. Where this text and that code disagree, the code and its test vectors win until this text is corrected (section 0.4). The Swift prototype `proto-metal/main.swift` carries the same generator and acceptance rule and is bit-exact with the crate on every pack (`docs/bench-log.md`, entries "igneum-pow: Rust crate bit-exact with proto-metal" and "generator version 2").
|
||||
|
||||
Every parameter marked "prototype value, to be fixed at gate 1" is carried by the implementation today, is part of the test vectors, and is confirmed or replaced by the named measurement before the hash is frozen (section 1.16).
|
||||
|
||||
## 1.0 The frozen object (Igneum 2.0 Deliverable 1, 8 October 2026)
|
||||
|
||||
Spec version 0.3 freezes the object the served kit is built from, by digest, so the served spec and the served kit are one thing. Every digest below is the sha256 the build embeds (`packaging/pow-freeze.txt` for the generator, `packaging/d1-freeze.txt` for the rest; `consensus/pow/build.rs` refuses a build whose linked tree matches no listed freeze, and the daemon prints the five at start as `D1 baseline: ...`). A value marked Measured carries the row and date that measured it; Designed is the construction as written in `docs/design/class-v6-rotating-family.md`; Open is named with the clock that closes it.
|
||||
|
||||
| Component | Digest (sha256) | Source | Label |
|
||||
|---|---|---|---|
|
||||
| Generator | 5f4d6dc6199294db89042171004e6420c1e5791e716d4b39b73d375818c10b6f | igneum-pow `src/*.rs` tree at the class-v6 freeze commit 1a938abe4 (the hash lane's final tree, 21:24 BST; the fingerprint as `consensus/pow/build.rs` computes it, `find src -name '*.rs' \| sort \| xargs sha256sum \| sha256sum`, printed by the served pair at start); generator 6 (`ProgramClass::V6`) carrying the class v6 rules | Measured (the census lane's PASS on the chain-seed draw 2a1d6caab4c24564, 8 October 2026, 21:5x UK, and the signing program's window trace 1.0026 per site; research lane D's family gate 6.13 PASS) |
|
||||
| Verifier | 2b1a81cb413236cf063077b46ed3111628f6c41036bcf6e23ee4cbbf5679ef7a, 474678f35f7545db28055d5e5bbc308231d84a5a072202087a2a8d5b09123896 | the embedded shard program id and aggregator id (`proving/igneum-prove/elf/manifest.json`, pinned 5 October 2026), pinned in consensus by `verifier_in_consensus` and printed by the daemon's "Proving: consensus proof verification from DAA score 0" line | Measured (the enforced-proving lane: 0.67 to 0.71 s a shard proof on one EPYC core, 8 October 2026) |
|
||||
| Dataset policy | 70a6c703787d5d75cdbc486b34acf3eb901f4188bae3e948c94274d5f4ac4bba | the object's `class_v6_dataset_steps` and `class_v6_family_flags` as the consensus digest arm hashes them (section 1.13.3) | Designed (the schedule), Measured (the 5.5 GiB form: 136.56 MH/s against 141.48 at 1 GiB on a rented 5090 at stock, the hash lane, 8 October 2026) |
|
||||
| Compiler config | 1.99.0-b940084d7;cargo-5f94df478;sccache-0.18.0;zig-0.17.0;profile.release:lto=thin,strip=true,overflow-checks=true;target:x86_64-unknown-linux-gnu | `rust-toolchain.toml` channel 1.99.0 (rustc b940084d7 2026-09-28), the release profile of the workspace `Cargo.toml`, the box's sccache and zig; the gate artefact is the native target, the seed and rig builds cross with zig at glibc 2.35 and 2.31 | Measured (the build-remote read-back on igneum-build-2, 8 October 2026) |
|
||||
| Harness | v5-fasttime 92bf6a7f (the harness commit on the box mirror; its records cross-v6-617cb441-3.{json,log}) | the fast-time harness the class v6 crossing ran on: v4 at e0 and e1, rung 1 by signal, v5 at byte 6 by its floor, v6 at byte 8 from its floor one epoch after, the template at every boundary, the stale node's refusals, the restart step, the cold restart on four nodes; named by the record as the five-together case tonight, with the gap stated: proof queues and seed transitions not exercised in the same network as the crossing and the finality attack, owed to the next harness commit | Measured (SUMMARY PASS cross-v6-617cb441-3 at 17:30:44 BST, 8 October 2026; the finality attack lines on bee41b5e's pair in sim/results_v2.md) |
|
||||
|
||||
The object the digests describe: the fork branch `class-v6-node-review` at e8773ff5 (FROZEN 23:17 BST, 8 October 2026; the sha the served kit prints; the fallback `class-v6-node-b` 83bb8ea6 named beside it) (the class v6 fields on `Params`, section 1.13.3; the floor `program_class_v6_activation_daa` is never on every compiled network until its own cut, so no live digest moves with this landing). The served kit's miner carries `igneum-miner program-id`, which prints the engine's attempt and program id for an epoch seed, era, day and state stream through the same path the template prepare takes; read beside the freeze CLI's `igneum-pow show` on the same inputs, a differing id is the stop (the pairing read-back, rule 19).
|
||||
|
||||
## 1.1 What the hash is for
|
||||
|
||||
The lottery hash decides who produces the next block. It MUST be:
|
||||
|
|
@ -392,6 +406,49 @@ One row per Rust `pub const` of igneum-pow that class v5 adds; the gate's spec-c
|
|||
| `LAST_RESORT_SCAN` | 256 | `igneum-pow/src/generator.rs`, the verified last resort |
|
||||
| `GENERATOR_VERSION_V5` | 5 | `igneum-pow/src/generator.rs` |
|
||||
|
||||
### 1.4.8 Generator version 6, the class v6 freeze (version 0.3)
|
||||
|
||||
Implemented (`igneum-pow/src/generator.rs`, generator version 6 at the class-v6 freeze; generator 5's construction, section 1.4 of version 0.2, stands for every epoch before the class v6 floor). A program is a list of `INSTR_COUNT` instructions over the lane registers, executed `ITERATIONS` times per hash.
|
||||
|
||||
| Parameter | Value | Label |
|
||||
|---|---|---|
|
||||
| Registers per lane | 64 x u32, the 64-register window (the window is the core shape the floor lanes measured: k 0.78 at the lock against 0.56 with 8 registers; `docs/design/class-v6-rotating-family.md` 10.0c, the edge table 10.0e) | Designed; Measured on the chip side (the synthesised core row, floor lane 2, 8 October 2026); the GPU side Open until the frozen generator's pack is measured on the 5090, 5070 Ti and M5 Max (the hash lane, its own clock) |
|
||||
| Instructions per program | 64, of which exactly 16 are `load` (section 1.4.2) | Definition since 4 October 2026 |
|
||||
| Iterations per hash | 8 | Measured (the CPU-verify bound, section 1.11, held at version 0.2's figure) |
|
||||
| Lanes per unit of work | 32 | Definition (section 1.9) |
|
||||
| Shuffle masks | {1, 2, 4, 8, 16} | Definition, follows from 32 lanes |
|
||||
| Rotate immediates | 1..31 | Definition |
|
||||
| Read width W | 4 words (16 bytes), pinned at genesis, not drawn (section 1.8.5) | Measured (floor lane 3 and lane 5: 8 words not free on the 5090 at stock, +4.8 percent; 16 words killed, 8 October 2026) |
|
||||
| The index fold | every `load` folds the product's low bits before the stride rotation, so no era's R lands a biased product bit on an address bit (a ring-A design rule of layer 1; every era's draw passes through it) | Measured (the census lane, 8 October 2026: hl-v6-fold PASS, min site ratio 0.99986, no biased free bit, 256 of 256 chain-shaped seeds at r 0.701 with 0 exhausted, F8 16 of 16 within 1.0455x at 2^22) |
|
||||
| Op-mix weights | the re-weight table rw1 as exported in hl-v6-all, in draw order: add 16, xor 14, mul 4, mad 12, shfl 4, rotl 11, sub 10, mulhi 2, rotr 10, or 0 (sum 83; `or` never drawn); the lossy families (`mul`, `mulhi`, `or`) held at their base and the era's B = 4 points on the injecting families only (`add`, `sub`, `xor`, `mad`, `shfl`, `rotl`, `rotr`); the ring-A rule `or + mul + mulhi` at most the table's base plus B | Measured (the census lane, 8 October 2026: hl-v6-rw PASS r 0.117, 256 of 256, F8 at most 1.15x, the 6-sigma bucket flags on 4 of 16 seeds equal to the class v5 control's 5 of 16; hl-v6-foldrw PASS, F8 15 of 16 at most 1.0932x; hl-v6-all PASS 256 of 256 with 0 exhausted; lane D's acceptance at both widths; the census lane's own table rw2 FAIL on F8 at 1.3111x and not carried) |
|
||||
| The family bank | the admissible entries are the object's `class_v6_family_flags` bits: 0 the 64-register window, 1 the index fold, 2 the re-weight table (rw1), 3 the lossy band at base; an entry whose bit is off is not drawable by any era | Designed (section 10 of the design document; the flags in the consensus digest); Measured per entry (the census sheet: fold, rw, foldrw, win, all PASS, 8 October 2026 16:3x to 17:55 BST; the window's own address stream through the interpreter's trace on all 32 sites, per-site distinct ratio 1.0026 to 1.0027; the live-dataset F8 point at 2^24 for the two window packs owed) |
|
||||
|
||||
Sections 1.4.1 to 1.4.6 (the instruction set, the load count, the draw order, the generator contract, the encoding, program acceptance) stand as version 0.2 wrote them for generator 5 with the two changes above (the register count in the draw order's register fields; the fold in `load`'s address computation); the frozen tree is the normative text where they differ.
|
||||
|
||||
### 1.8.5 Item derivation and dataset mapping (version 0.3)
|
||||
|
||||
Item derivation stands as version 0.2 wrote it (the 16-word item from the day key, the mixer rounds and the cache lines). The dataset mapping and the read width change:
|
||||
|
||||
- Read width. One `load` reads W = 4 words (16 bytes) of the item at the computed index: pinned at genesis, never drawn. Measured: 8 words cost +4.8 percent on the 5090 at stock and 16 words were killed on four rented cards (floor lanes 3 and 5, 8 October 2026); the width is the only wire lever on the SRAM die (66x at 4 bytes against 44x at 16 at zero shadow), so pinning it at 4 words holds the chip's per-joule edge where the floor lanes measured it. Designed value, Measured consequence.
|
||||
- Index mapping. `idx = (src * N_words) >> 32` computed in 64 bits (the multiply-shift range reduction; uniform to within 2^-32, branch-free, integer only) takes any `N_words`, so the dataset is not a power of two: 5.5 GiB is 92,274,688 items of 64 bytes (1,476,395,008 words). The exporter refuses an item count off the 2^16 grid, and every step of the schedule lies on it. Measured: the 5.5 GiB form at 136.56 MH/s against 141.48 at 1 GiB on a rented 5090 at stock, fingerprint 23ced07a4d28b465 on the pinned class v3 program, the worker's self-test PASS (the hash lane, 8 October 2026).
|
||||
- Cache. 2^26 words until the dataset reaches 4 GiB, then doubling as section 1.13.3 of version 0.2 says. Designed.
|
||||
|
||||
### 1.13.3 The dataset policy (version 0.3)
|
||||
|
||||
Designed (layer 2 of `docs/design/class-v6-rotating-family.md`, section 3), in the consensus digest through the object's fields. The day's item count is the larger of the floor schedule at the chain's DAA height and the state rule, under a genesis ceiling:
|
||||
|
||||
```
|
||||
floor_bytes(daa) = the highest step (height, MiB) with height <= daa, in bytes; 0 below the first step
|
||||
state_bytes = records_at_era_cut * 64 (CLASS_V6_STATE_BYTES_PER_RECORD)
|
||||
bytes = min(max(floor_bytes(daa), state_bytes), ceiling_bytes) (no ceiling when 0)
|
||||
items = bytes / 64 (CLASS_V6_ITEM_BYTES)
|
||||
```
|
||||
|
||||
- The steps (`class_v6_dataset_steps`, up to four `(daa, MiB)` pairs; `(never, 0)` is no step): 5.5 GiB at the class v6 epoch, 8.5 GiB two years on, 11.5 GiB four years on, the fixed schedule's 16 GiB whenever the state brings it forward (section 3.4 of the design document, priced per tier there: each step retires about a quarter of today's measured consumer cards by count; the 8 GB tier holds 5.5 GiB at 72 to 76 percent of the card). Designed.
|
||||
- The state rule is read at the era cut (the records of the execution state the class v5 dataset is keyed on, section 2a of `docs/design/class-v5-stored-state.md`), so a chain whose state outgrows the schedule steps the dataset up by state, never down. Designed.
|
||||
- The family flags (`class_v6_family_flags`) sit beside the steps in the digest arm (the floor, then every set step as height and MiB, then the flags; entered only when the floor is set). Designed.
|
||||
- `igneum-pow`'s side: the mapping of section 1.8.5 takes the item count; the generator reads the live steps and flags through `kaspa_consensus_core::igneum::class_v6_dataset_items_at` and `class_v6_family_admissible`. Implemented on the node; the generator's read of the flags lands with the frozen tree.
|
||||
|
||||
## 1.5 Fixed memory footprint
|
||||
|
||||
Implemented for the prototype size; Designed for genesis.
|
||||
|
|
@ -593,6 +650,30 @@ and continues as 1.8.5 writes. Every item carries a leaf (the first form of the
|
|||
|
||||
The refresh (proof of following). The leaves of epoch `e` are derived from the state after epoch `e`'s seed block, so the dataset is rebuilt every epoch from the chain's own state at that cut (the refresh cadence is the epoch; the design names one hour and one epoch as the candidates and the epoch is the one shipped). A miner must therefore hold a following executor, or ask a node's RPC for each epoch's stream; a pool can serve its workers the leaves (64 bytes per item, the dataset's size per refresh) but cannot serve a dataset to a party that does not fetch it each epoch. The verifier on a node without the state (a node synced from a pruning proof) holds the state root per class v5 epoch as a witness beside the epoch seed header in the proof (`EpochSeedHeader.stateRoot`, section 10) and trusts the headers under it as trusted data until its own executor passes the cut. Measured cost (design section 7): dataset build +0.24 ms on an RTX 4090, the verifier +0.20 ms per warp on one core (+0.28 cold), hash rate equal within 0.02 percent, no change in watts.
|
||||
|
||||
### 1.8.7 Item derivation and dataset mapping under the freeze (version 0.3; amends 1.8.5)
|
||||
|
||||
Item derivation stands as version 0.2 wrote it (the 16-word item from the day key, the mixer rounds and the cache lines). The dataset mapping and the read width change:
|
||||
|
||||
- Read width. One `load` reads W = 4 words (16 bytes) of the item at the computed index: pinned at genesis, never drawn. Measured: 8 words cost +4.8 percent on the 5090 at stock and 16 words were killed on four rented cards (floor lanes 3 and 5, 8 October 2026); the width is the only wire lever on the SRAM die (66x at 4 bytes against 44x at 16 at zero shadow), so pinning it at 4 words holds the chip's per-joule edge where the floor lanes measured it. Designed value, Measured consequence.
|
||||
- Index mapping. `idx = (src * N_words) >> 32` computed in 64 bits (the multiply-shift range reduction; uniform to within 2^-32, branch-free, integer only) takes any `N_words`, so the dataset is not a power of two: 5.5 GiB is 92,274,688 items of 64 bytes (1,476,395,008 words). The exporter refuses an item count off the 2^16 grid, and every step of the schedule lies on it. Measured: the 5.5 GiB form at 136.56 MH/s against 141.48 at 1 GiB on a rented 5090 at stock, fingerprint 23ced07a4d28b465 on the pinned class v3 program, the worker's self-test PASS (the hash lane, 8 October 2026).
|
||||
- Cache. 2^26 words until the dataset reaches 4 GiB, then doubling as section 1.13.3 of version 0.2 says. Designed.
|
||||
|
||||
### 1.13.3 The dataset policy (version 0.3)
|
||||
|
||||
Designed (layer 2 of `docs/design/class-v6-rotating-family.md`, section 3), in the consensus digest through the object's fields. The day's item count is the larger of the floor schedule at the chain's DAA height and the state rule, under a genesis ceiling:
|
||||
|
||||
```
|
||||
floor_bytes(daa) = the highest step (height, MiB) with height <= daa, in bytes; 0 below the first step
|
||||
state_bytes = records_at_era_cut * 64 (CLASS_V6_STATE_BYTES_PER_RECORD)
|
||||
bytes = min(max(floor_bytes(daa), state_bytes), ceiling_bytes) (no ceiling when 0)
|
||||
items = bytes / 64 (CLASS_V6_ITEM_BYTES)
|
||||
```
|
||||
|
||||
- The steps (`class_v6_dataset_steps`, up to four `(daa, MiB)` pairs; `(never, 0)` is no step): 5.5 GiB at the class v6 epoch, 8.5 GiB two years on, 11.5 GiB four years on, the fixed schedule's 16 GiB whenever the state brings it forward (section 3.4 of the design document, priced per tier there: each step retires about a quarter of today's measured consumer cards by count; the 8 GB tier holds 5.5 GiB at 72 to 76 percent of the card). Designed.
|
||||
- The state rule is read at the era cut (the records of the execution state the class v5 dataset is keyed on, section 2a of `docs/design/class-v5-stored-state.md`), so a chain whose state outgrows the schedule steps the dataset up by state, never down. Designed.
|
||||
- The family flags (`class_v6_family_flags`) sit beside the steps in the digest arm (the floor, then every set step as height and MiB, then the flags; entered only when the floor is set). Designed.
|
||||
- `igneum-pow`'s side: the mapping of section 1.8.5 takes the item count; the generator reads the live steps and flags through `kaspa_consensus_core::igneum::class_v6_dataset_items_at` and `class_v6_family_admissible`. Implemented on the node; the generator's read of the flags lands with the frozen tree.
|
||||
|
||||
## 1.9 The 32-lane unit of work
|
||||
|
||||
Definition. The hash is defined over an aligned group of 32 consecutive nonces `g .. g + 31` with `g AND 31 = 0` (wrapping modulo 2^32 at the top of the range). `shfl` exchanges registers within that group: lane `l` reads from lane `l XOR mask`, `mask < 32`, so the exchange never leaves the group. The verification unit is the group: `hash(n)` is computed by evaluating the group `n AND ~31` and taking lane `n AND 31` (`verify.rs`, `Epoch::hash`).
|
||||
|
|
@ -702,6 +783,22 @@ evaluated in integers (bytes), with one year = 31,536,000 DAA seconds. The datas
|
|||
- Index mapping. `src AND MASK` requires a power-of-two size. For a non-power-of-two `N_d` the proposed mapping is `idx = (src * N_words) >> 32` computed in 64 bits (a multiply-shift range reduction; uniform to within 2^-32, branch-free, integer only). At `N_words = 2^28` this gives `src >> 4`, not `src AND MASK`, so adopting it changes the 1 GiB vectors; gate 1 chooses between (a) the multiply-shift mapping with new vectors, or (b) power-of-two sizes only, growing in steps (2 GiB, 4 GiB) on the same schedule's average, which keeps `AND MASK` and means a 4 GiB card lasts until the 4 GiB step instead of fading.
|
||||
- The item index `t` is 32 bits, so the construction as written tops out at 2^32 items = 256 GiB, which the schedule reaches after 508 years. No action needed.
|
||||
|
||||
### 1.13.4 The dataset policy under the freeze (version 0.3; amends 1.13.3)
|
||||
|
||||
Designed (layer 2 of `docs/design/class-v6-rotating-family.md`, section 3), in the consensus digest through the object's fields. The day's item count is the larger of the floor schedule at the chain's DAA height and the state rule, under a genesis ceiling:
|
||||
|
||||
```
|
||||
floor_bytes(daa) = the highest step (height, MiB) with height <= daa, in bytes; 0 below the first step
|
||||
state_bytes = records_at_era_cut * 64 (CLASS_V6_STATE_BYTES_PER_RECORD)
|
||||
bytes = min(max(floor_bytes(daa), state_bytes), ceiling_bytes) (no ceiling when 0)
|
||||
items = bytes / 64 (CLASS_V6_ITEM_BYTES)
|
||||
```
|
||||
|
||||
- The steps (`class_v6_dataset_steps`, up to four `(daa, MiB)` pairs; `(never, 0)` is no step): 5.5 GiB at the class v6 epoch, 8.5 GiB two years on, 11.5 GiB four years on, the fixed schedule's 16 GiB whenever the state brings it forward (section 3.4 of the design document, priced per tier there: each step retires about a quarter of today's measured consumer cards by count; the 8 GB tier holds 5.5 GiB at 72 to 76 percent of the card). Designed.
|
||||
- The state rule is read at the era cut (the records of the execution state the class v5 dataset is keyed on, section 2a of `docs/design/class-v5-stored-state.md`), so a chain whose state outgrows the schedule steps the dataset up by state, never down. Designed.
|
||||
- The family flags (`class_v6_family_flags`) sit beside the steps in the digest arm (the floor, then every set step as height and MiB, then the flags; entered only when the floor is set). Designed.
|
||||
- `igneum-pow`'s side: the mapping of section 1.8.5 takes the item count; the generator reads the live steps and flags through `kaspa_consensus_core::igneum::class_v6_dataset_items_at` and `class_v6_family_admissible`. Implemented on the node; the generator's read of the flags lands with the frozen tree.
|
||||
|
||||
## 1.14 Determinism requirements
|
||||
|
||||
A conforming implementation MUST:
|
||||
|
|
@ -730,7 +827,36 @@ A miner, kernel emitter or verifier conforms when all of the following pass. Eac
|
|||
|
||||
Measured conformance to date (`docs/bench-log.md`), version 2 vectors: Apple Metal natively (3 of 3 units on the exported `igneum-genesis-mh`, fuzz 2,000 of 2,000, and the Swift generator identical to the Rust one on the five seeds of 1.4.6), Apple OpenCL (96 of 96 on all four packs, fingerprints as in item 5), the clang CUDA emulation (96 of 96 on all four packs, 2 warps per block), the clang OpenCL emulation (96 of 96 on the two memory-hard packs in sub-group 32 and wave64 configurations), and the Rust CPU reference. Not yet run on version 2 vectors: the RTX 5090 (CUDA and NVIDIA OpenCL), AMD gfx1036, pocl; the version 1 runs on those devices (3 October 2026) stand as evidence that the kernel text, which version 2 did not change, agrees across vendors. A discrete AMD card has not run anything (ledger M8).
|
||||
|
||||
## 1.16 Parameters marked "prototype value, to be fixed at gate 1" and what fixes them
|
||||
## 1.16 Parameters fixed by the freeze (version 0.3)
|
||||
|
||||
### 1.16.1 The one acceptance rule (Designed; the plan's rule above the kept list)
|
||||
|
||||
A change to the hash passes only when the best adversarial implementation becomes worse relative to the best practical GPU implementation, within pre-agreed cost (10 percent of the GPU's rate per joule at the 1,300 MHz lock) and verification limits (the verifier's bound of 1.2 unchanged).
|
||||
|
||||
### 1.16.2 The per-program rule as served, the class v6 freeze (Measured on hl-v6-fold, hl-v6-rw, hl-v6-foldrw, hl-v6-win and hl-v6-all by the census lane, 8 October 2026, 16:3x to 17:55 BST, harness class-v6-census-fold 5b3486f0)
|
||||
|
||||
Every drawn program of an epoch is accepted or redrawn by one rule on the closed-form dataset, run before the first hash: (a) to (c') as class v4 sub-version 3 states them (the dataflow rule, the shared-operand rule, the lane-constant test, the total draw); (c'') the distinct-index ratio at every load site over the acceptance's units at or above 0.98; (c''') the hot-item floor, the minimum site ratio at or above 0.995 over 2^20 indices; the bucket bound, no 64-item bucket of any site's window above 6 sigma of its uniform expectation; the free-bit check, the one-count of every address bit of every site within 6 sigma of n / 2 (the index fold is the construction that keeps the era's rotation from landing a biased product bit on an address bit; the check reads that it held); and the draw census, over 256 chain-shaped seeds the redraw rate r at or below 0.2 with no seed exhausted (measured 0.117 on the re-weight table and 0.701 to 0.716 on the fold and window packs, mean attempt 0.13, max 2, (c''') refusals 0.34 percent). Beyond the per-program rule the freeze carries the F8 read on the live state at 2^22 reads per seed: the top 0.1 percent of items of every site within 1.2x of the window model (measured: 15 of 16 seeds PASS at most 1.0932x on hl-v6-foldrw; the uniform control flags 5 of 16 on the same statistic; the census lane's own table rw2 read 1.3111x on a sample seed and is not carried). The window packs' own address stream through the interpreter's trace reads a per-site distinct ratio of 1.0026 to 1.0027 on all 32 sites with the item share within 0.97 to 1.04x of each pack's control; their live-dataset F8 point at 2^24 is owed (the mirror interprets eight registers).
|
||||
|
||||
### 1.16.3 The parameters the freeze fixes
|
||||
|
||||
| Parameter | Fixed value | Fixed by | Label |
|
||||
|---|---|---|---|
|
||||
| Registers per lane | 64, the window with the full chain (reg64c) | the floor lanes' core row (k 0.78 at the lock), the census lane's hl-v6-win and hl-v6-all PASS | Designed; the GPU side Open (the frozen pack's measurement on the 5090, 5070 Ti and M5 Max, the hash lane's clock) |
|
||||
| Read width W | 4 words (16 bytes) | floor lanes 3 and 5, 8 October 2026 (the hl read-width b run: W = 8 at +4.8 percent energy per hash at stock; W = 16 and w64 dead at stock and at the knee, w64 bandwidth-bound at a 5090 share of 0.58; W = 32 out with W = 16) | Measured |
|
||||
| Index fold | on, every era | the census lane's hl-v6-fold PASS (min site ratio 0.99986, no biased free bit) | Measured |
|
||||
| Op-mix table | rw1, sum 83 (section 1.4); the lossy families at base, B = 4 on the injecting families | the census lane's hl-v6-rw, foldrw and all PASS; lane D's 0 of 3,000 eras exhausted | Measured |
|
||||
| State leaves | class v5's, the dataset keyed by the execution state after the epoch's reference block | class v5 (`docs/design/class-v5-stored-state.md`) | Measured (Devnet 3's crossing at 68,400 and the Igneum 2.0 devnet from genesis, 8 October 2026) |
|
||||
| Dataset at genesis | 4 GiB (4,096 MiB, 67,108,864 items of 64 bytes) from genesis, the floor; the execution state's bytes above it as the state grows (class v5's leaves, `CLASS_V6_STATE_BYTES_PER_RECORD` a record), the ceiling the miner's own flag | the research lane's 10.0u (5.5 GiB at +14.3 percent energy per hash on a locked 5090, over the 10 percent budget; 4 GiB at +8 percent at the knee, inside it); main's word of 8 October 2026 | Designed (main's default of 8 October 2026, 20:00 BST); the energy row Measured (the research lane: +8 percent energy per hash at 4 GiB on a locked 5090, inside the 10 percent budget; 5.5 GiB at +14.3 percent, outside it) |
|
||||
| Dataset floor schedule | `class_v6_dataset_steps` = [(0, 4,096 MiB)]: one step, the 4 GiB floor from DAA 0 and no later step; the later rungs of lane A's priced schedule enter by a cut that moves the digest (the v6 floor is never on every compiled network tonight, so the policy is carried by the object and outside its digest) | lane A's priced schedule, section 3.4 of the design document, and 10.0u | Designed |
|
||||
| Iterations, instruction count, lanes | 8, 64 (16 loads), 32 | version 0.2 | Measured / Definition |
|
||||
| The era draw's range | as section 1.13.1 of version 0.2 | layer 1 | Designed |
|
||||
| Test vectors | section 1.17 | the frozen generator's `igneum-pow show` on the conformance seeds | Open: re-cut by 01:00 BST 9 October 2026 |
|
||||
|
||||
### 1.16.4 The excluded knobs, kept as regression controls in the D1 kit, each with the row that killed it
|
||||
|
||||
W = 8 (hl read-width b at stock: W = 4's rate at +4.8 percent energy per hash; W = 4 pins). W = 16 and w64 (dead at stock and at the knee, the read-width b run; w64 bandwidth-bound, 5090 share 0.58). W = 32 (the same read-width family, out with W = 16). Long programs (knob 3, sh1024x27: 1.58x energy per hash on the 5090 and the 4090 at stock, the 1p5x bench). SM gating and SM-sparse (occupancy worth 1 to 4 percent at stock and nothing at the lock, floor lane SM-sparse; `--sm-sparse auto` ships off by default, on in the Efficiency and Balanced tiers at its measured 1 to 2.5 percent). Memory clock (no lever at the lock, the Ember rows; approximate). Select trees (rejected in the invention lane's KEEP/KILL table, floor lane 5). Sealed classes, random epoch lengths, per-tier scoring, VRF draws (the plan's decision register p. 7, excluded by rule, no GPU row). Mixed FP32 (the mixed-fp32 lane: deterministic FP32 costs the cards 15 to 26 percent energy per hash against the 10 percent budget, four fifths of it the integer masking that keeps the FP unit deterministic, and the chip's edge grows to 3.0x to 3.2x; `docs/analysis/class-v6/mixed-fp32.md`; KILL 8 October 16:3x). Connected state (cs64, the connected-state lane: an edge of 1.10x against the 1.25x gate; KILL; its 5090 row informational only). The census lane's weight table rw2 (F8 1.3111x on a sample seed against the 1.2x gate; the k lane's rw1 is what the all pack carries).
|
||||
|
||||
### 1.16.5 Version 0.2's prototype table, superseded by 1.16.3 and kept for the record
|
||||
|
||||
| Parameter | Prototype value | Measurement or decision that fixes it |
|
||||
|---|---|---|
|
||||
|
|
@ -748,6 +874,10 @@ Measured conformance to date (`docs/bench-log.md`), version 2 vectors: Apple Met
|
|||
|
||||
## 1.17 Test vectors
|
||||
|
||||
Open at version 0.3 (clock: 01:00 BST, 9 October 2026). The vectors of version 0.2 are generator 5's and stand for every epoch before the class v6 floor; the class v6 vectors are re-cut from the frozen tree's `igneum-pow show` on the conformance seeds the moment the class-v6 freeze commit is named (the hash lane's merge of class-v6-fold and reg64-v5, 21:00 to 23:00 BST), with `igneum-miner program-id` read beside each as the pairing check. Until then this section's version 0.2 text stands for generator 5 and a class v6 implementation conforms by matching the frozen crate on the seeds of section 1.15.
|
||||
|
||||
### 1.17.1 Version 0.2's vectors (generator 5; they stand for every epoch below the class v6 floor)
|
||||
|
||||
Pack `proto-cuda/packs/igneum-genesis-mh/` (seed `igneum-genesis`, generator version 2, attempt 0, program id `bcc1248b10cc90f2`, day `2026-10-03`, memory-hard, 2^28 words, MASK `0x0fffffff`, 32 lanes). Produced by `igneum-pow export` (the Rust CPU interpreter) on 4 October 2026, reproduced by the Swift CPU interpreter and cross-checked by Metal, Apple OpenCL and the two clang emulations the same day (section 1.15). The version 0.1 vectors of 3 October 2026 (lane 0 `1fb0b3bbc1ac8279`) are retired: they belong to a generator that no longer exists in the protocol. The two closed-form packs (`igneum-genesis`, `igneum-hourly`) are regression vectors for the interpreter only and are not the lottery hash.
|
||||
|
||||
Unit at base nonce 0, lanes 0..31:
|
||||
|
|
|
|||
Loading…
Reference in a new issue