igneum/igneum-pow
igneum-labs ab530e0e18 igneum-pow: OpenCL bound kernel in the pack, byte-seed options on the CLI
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-03 19:01:13 +00:00
..
src igneum-pow: OpenCL bound kernel in the pack, byte-seed options on the CLI 2026-10-03 19:01:13 +00:00
tests igneum-pow: OpenCL bound kernel in the pack, byte-seed options on the CLI 2026-10-03 19:01:13 +00:00
.gitignore igneum-pow: Rust crate bit-exact with proto-metal (seed, generator, memhard, verifier, emitters) 2026-10-03 16:53:18 +00:00
Cargo.lock igneum-pow: Rust crate bit-exact with proto-metal (seed, generator, memhard, verifier, emitters) 2026-10-03 16:53:18 +00:00
Cargo.toml igneum-pow: Rust crate bit-exact with proto-metal (seed, generator, memhard, verifier, emitters) 2026-10-03 16:53:18 +00:00
README.md igneum-pow: header binding (init words from the pre-PoW hash and nonce), bound kernels, 8 bound vectors 2026-10-03 18:39:34 +00:00
rustfmt.toml igneum-pow: Rust crate bit-exact with proto-metal (seed, generator, memhard, verifier, emitters) 2026-10-03 16:53:18 +00:00

igneum-pow

The Igneum lottery hash in Rust, bit-exact with the Swift prototype in proto-metal/main.swift. This is the crate the rusty-kaspa fork will call (docs/fork-map.md, rows a1 to a3) so a node written in Rust can verify any block and hand miners the kernel source for the epoch. No dependency outside the standard library; serde_json is a dev-dependency for reading the packs in the tests.

Date: 3 October 2026. Toolchain: rustc 1.99.0 via rustup (the Homebrew 1.69 on PATH is too old; use ~/.cargo/bin/cargo).

Modules

Module What it is Swift namesake
seed 32-byte seed words from a string (FNV-1a 64, four salts, finalised); seed_words_from_bytes is the boundary where the chain will feed the VDF output; SplitMix64 seedWords, SplitMix64
generator the 64-instruction program for a seed (op, dst, src, src2, imm, imm2, rot, bit, mask); levers load_weight and wide_frac generateProgram, GeneratorConfig
memhard 256 MiB cache (2^16 chains of 64 ChaCha12 blocks), mixer parameters, 8-round item derivation with the 32 lanes interleaved, MemhardCpu::fetch cpuFillCache, MixParams, deriveItems, MemhardCPU
verify the 32-lane warp interpreter, DatasetMode::{ClosedForm, MemoryHard}, Epoch, hash_warp, verify_block cpuWarp, DatasetSource
emit Metal, CUDA and OpenCL source, program.h, memhard.h, vectors.h, program.json, vectors.json, export_pack; since 3 October 2026 also the header-bound kernels program_bound.metal and kernel_bound.cu generateMSL, memhardMSL, emitMemhardCore, generateCUDA, generateOpenCL, exportPack
bind header binding (spec 01 section 1.6, O-1.9): init words from "igneum-block/" || H || nonce_hi_le32, bound hash API on Epoch, the 256-bit pow mapping, the interim day seed bytes none (new)

The API the fork calls

use igneum_pow::{Epoch, DatasetMode};

// Once per epoch and day: generates the program and fills the 256 MiB cache (about 0.18 s on one core).
let epoch = Epoch::memory_hard("igneum-genesis", "2026-10-03");

let h: u64 = epoch.hash(nonce);                 // one nonce (computes its aligned 32-nonce warp)
let w: [u64; 32] = epoch.hash_warp(base_nonce); // one warp
let ok: bool = epoch.verify_block(nonce, target_u64);

// Miner programs for the epoch, byte-identical to the Swift exporter.
let pack = igneum_pow::emit::export_pack(&epoch, "2026-10-03", "igneum node");
pack.write_to(std::path::Path::new("out"))?;   // kernel.cu, kernel.cl, program.metal, memhard.h, ...

The header-bound form (what the chain uses)

The pack form above initialises the lane registers from the program's own seed words, so one nonce has one hash per epoch whatever block is mined. On the chain the init words commit to the block (spec 01 section 1.6, open item O-1.9, implemented 3 October 2026 in src/bind.rs):

H        = header hash with the nonce field zeroed, every other field as mined
           (rusty-kaspa hash_override_nonce_time(header, 0, header.timestamp))
nonce    = 64 bits; lane nonce n = low 32 bits; nonce_hi = high 32 bits
I        = seed_words_from_bytes("igneum-block/" || H || nonce_hi_le32)      (49 bytes in)
hash     = interpret(program, I, n)                                          (section 1.7 of the spec)
pow256   = hash in the top 64 bits, low 192 bits zero (little-endian bytes 24..32)
valid    = pow256 <= target256, which is exactly hash <= target256 >> 192

Choices the spec left open and how they were fixed: H keeps the timestamp (Kaspa zeroes it in the pre-PoW hash and absorbs it in cSHAKE afterwards; the lane hash has no afterwards, so a nonce would otherwise be reusable across timestamps). The pow value puts the lane in the top 64 bits with zero low bits so a GPU worker and the node compare the same 64-bit numbers. The interim day seed is "igneum-day/" || day_le64 with day = timestamp_ms / 86,400,000 (O-1.10's proposal needs the VDF schedule). The epoch seed bytes are the 32 bytes of the epoch block hash (devnet v0).

use igneum_pow::{bind, Epoch};

let epoch = Epoch::from_seed_bytes(epoch_hash.as_bytes(), &bind::day_bytes(day), "label");
let lane: u64 = epoch.hash_bound(&prehash, nonce);               // one 64-bit nonce
let warp: [u64; 32] = epoch.hash_warp_bound(&prehash, nonce);    // its aligned 32-nonce group
let init = bind::block_init_words(&prehash, nonce);              // what a GPU kernel takes as its argument
let same = epoch.hash_warp_init(&init, nonce as u32 & !31);     // == warp
let pow: [u8; 32] = epoch.pow_bound(&prehash, nonce);
let ok = epoch.verify_block_bound(&prehash, nonce, bind::target64_from_le256(&target_le));

On the GPU the init words are a kernel argument: igneum_hash_bound in program_bound.metal takes constant uint* initw [[buffer(3)]], and kernel_bound.cu takes IgneumInitWords iw by value. Both are emitted into every pack next to the unchanged igneum_hash and differ from it only in the kernel name, the argument and the eight init lines (initw[i] / iw.w[i] instead of SEEDW[i]). The lane nonce stays baseNonce + gid.

Bound vectors (seed igneum-genesis, day 2026-10-03, memory-hard, 2^28 words; igneum-pow hash-bound):

H nonce hash_bound
32 zero bytes 0 2c619692d823263b
32 zero bytes 1 55d21ed545735ce8
32 zero bytes 31 862eebe7fbda564e
32 zero bytes 4096 37aadc51f95725df
32 zero bytes 4294967296 (1 << 32) b62e28b8a90e554f
bytes 00 01 02 .. 1f 0 9b2437118e087833
bytes 00 01 02 .. 1f 4294967301 ((1 << 32) + 5) 714ae31e369e0446
bytes 00 01 02 .. 1f 18446744073709551615 (u64::MAX) 4ca4f84079025a13

Init words for H = 32 zero bytes, nonce 0: 595a8f8a 37647e95 faadade1 cbbcf2a4 54f7cc13 f6851b5e 8c68ca04 7991ea9c. The eight are pinned in bind::tests::bound_vectors. The pack vectors (96 per pack) are unchanged.

Epoch is Send + Sync; build one and share it. DatasetMode::ClosedForm reproduces the two old packs (igneum-genesis, igneum-hourly) and is not memory-hard. The hash is 64 bits; the fork maps it into its 256-bit target space in consensus/pow/src/lib.rs.

CLI

cargo build --release
./target/release/igneum-pow bench  --seed igneum-genesis [--warps 20] [--closed-form] [--day 2026-10-03]
./target/release/igneum-pow export --seed igneum-genesis --out <dir> [--closed-form]
./target/release/igneum-pow hash   --seed igneum-genesis --nonce 4103
./target/release/igneum-pow hash-bound --seed igneum-genesis --prehash <64 hex> --nonce <u64>

Tests

cargo test (29 tests, 1 s after compile; the dev profile is optimised so the cache fill is quick):

Check Pack Result
program.json instruction by instruction, op mix, loads per hash igneum-genesis, igneum-genesis-mh, igneum-hourly 3 x 64 match
Mixer parameters (key, rot, mul, rc) igneum-genesis-mh match
Cache head, last line, FNV-1a 64 48c4f5bf24166b2e igneum-genesis-mh match
Dataset head (16), [MASK], 64 sampled words all three match
96 hash vectors (3 warps x 32 lanes) igneum-genesis-mh 96/96
96 hash vectors igneum-genesis, igneum-hourly 96/96 each
kernel.cu, program.metal, kernel.cl, program.h byte-identical all three identical
memhard.h, memhard.metal byte-identical igneum-genesis-mh identical
program.json byte-identical (after the fix below) all three identical
vectors.json, vectors.h byte-identical apart from the provenance string all three identical
bound vectors (8), bound warp == bound single, H and nonce_hi enter the hash igneum-genesis-mh pass

An independent diff -r of igneum-pow export output against the checked-in packs shows the same two lines only: the provenance string and the "item" line, plus the two bound files that only the Rust exporter writes (program_bound.metal, kernel_bound.cu).

One deliberate difference: proto-cuda/packs/igneum-genesis-mh/program.json as written by the Swift is not valid JSON (main.swift line 1291 uses jhex inside the "item" string, so the cache line mask is quoted inside a quoted string). The Rust emitter writes 0x003fffff bare; the test normalises that one line before comparing. A node must hand miners valid JSON, so the Rust side does not reproduce the defect.

Measured, 3 October 2026, Apple M5 Max, one core, release build

Step Rust Swift (MEMHARD.md)
Cache fill, 256 MiB, 65,536 chains x 64 ChaCha12 blocks 175 to 181 ms (5 quiet runs; 200 ms once with another build running) 184.5 to 190.6 ms (C++ host reference 161.5)
CPU verify per warp, igneum-genesis, 104 loads, 3,328 items, avg of 20 0.441 ms 0.649 ms
igneum-genesis/epoch1, 104 loads 0.411 ms 0.631 ms
igneum-genesis/epoch2, 112 loads 0.488 ms 0.701 ms
igneum-second-seed, 104 loads 0.482 ms 0.801 ms
igneum-second-seed/epoch1, 144 loads, 4,608 items 0.579 ms 1.205 ms
Cold single warps across the five seeds 0.41 to 0.87 ms 1.16 to 2.11 ms
Closed form, igneum-genesis 0.002 ms 0.017 ms

The Rust verifier is 1.4x to 2.1x faster than the Swift one per warp; the registers are kept register-major (r[reg][lane]) so the lane loops vectorise, and the item derivation interleaves the 32 lanes round by round as the Swift does. The 10 ms gate holds with a margin of about 17x on the steady figure and 11x on the worst cold warp.

Not done here

  • No GPU. The vectors tie this crate to the Metal and CUDA results through the packs; nothing here runs a kernel.
  • Epoch::from_seed_bytes takes the epoch seed as bytes (the epoch block hash on devnet v0); the VDF output enters there.
  • The OpenCL pack has no bound kernel yet; proto-opencl/host.c builds its own from kernel.cl text at runtime.