docs/analysis: binding review of template, nonce, work and result (Igneum 2.0 register row): the chain of bindings from the code, the twelve reuse paths with their cost and rule, the measured lines, five open questions (epoch-seed VDF not in the node line, the day keyed on the timestamp)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-08 17:11:56 +00:00
parent 33af70442e
commit c2459ae08f

View file

@ -0,0 +1,145 @@
# Binding review: template, nonce, work and result (October 2026)
Igneum 2.0 register row (plan p. 10, missed pins): "Cryptographic review of the template, nonce, work and result
binding: no expensive intermediate reused across cheap winning attempts." Adversarial seat, 8 October 2026 (evening,
UK), with the hash lane for the generator's side. Sources: the crate (`igneum-pow` at the mirror's master 494fb981),
the fork's node line (`release-2.0.0-node` 9fc9f42a on the box mirror; the Mac's vendor copy is the 7 October master
and is not used where the two differ), spec 01, 02, 04 and 09, and the attack pass of 7 to 8 October
(`docs/analysis/attack-pass-2026-10.md`, rows F3, F8, F9, F1 and their records). Every number is either from a
record named beside it or marked Designed. No consensus code is changed by this document.
The question in one line: for every quantity that is expensive to compute, which inputs is it a function of, and
can a miner hold it fixed while it varies something cheap that still produces distinct winning attempts? A "winning
attempt" is a 64-bit lane hash at or below the target for a header the chain accepts.
## 1. What the hash commits to, from the code
### 1.1 The chain of bindings
| Step | Function of | Where | Cost |
|---|---|---|---|
| The pre-PoW hash `H` | every header field but the nonce: version, parents (every level), `hash_merkle_root`, `accepted_id_merkle_root`, `utxo_commitment`, timestamp, bits, `daa_score`, `blue_score`, `blue_work`, `pruning_point`, `vote_key_hash`; the nonce field zeroed | fork `consensus/core/src/hashing/header.rs:7` (`hash_override_nonce_time(header, 0, header.timestamp)`), `consensus/pow/src/igneum.rs:115` (`header_prehash`) | one BLAKE2b over about 200 bytes per template |
| The init words `I[0..7]` | `seed_words_from_bytes("igneum-block/" || H || nonce_hi_le32)` | `igneum-pow/src/bind.rs:48` (`block_init_bytes`, 49 bytes), `:57` | four FNV-1a-64 passes over 49 bytes per `(H, nonce_hi)`: about 200 integer ops |
| The lane registers at start | `r[i] = splitmix32((n XOR I[i]) + 0x9e3779b9 (i + 1)) XOR I[(i + 1) AND 7]` for lane nonce `n` (the low 32 bits of the header nonce) | `verify.rs:570` (`interpret_warp_core`); spec 1.6 | 8 splitmix32 per lane |
| The program | the epoch seed: on the node line, the hash of the last selected-chain block below `epoch x 3,600 - 600` on the header's own selected-parent chain (`pow_epoch_seed_score`, `epoch_seed`), attempt by attempt through acceptance (spec 1.4.6.6); the era seed from the era VDF (spec 4.4) | fork `consensus/core/src/igneum.rs:343`, `consensus/src/pipeline/header_processor/processor.rs:365` (release-2.0.0-node); `generator.rs`, `accept.rs` | once per epoch per node: the draw plus acceptance, about 1.8 core-s per candidate at (c'') and (c''') (attack pass, F9 lane (d)), 3.2 candidates per seed |
| The dataset | the day key `seed_words_from_bytes("igneum-day/" || day_le64)` with `day = header.timestamp_ms / 86,400,000` (class v5 adds the epoch's state leaves) | `bind.rs:62`, fork `pre_ghostdag_validation.rs:147` (`day: day_index(header.timestamp)`); spec 1.8 | once per day per node: the 256 MiB cache fill (175 ms on one core) and, for a miner, the 1 to 2 GiB dataset |
| One lane hash | the program applied to the 32 lanes of the aligned group `n AND NOT 31`, 8 iterations of 64 base instructions plus the 256-instruction shadow block 27 times, 16 loads per iteration reading `dataset[load_index(x)]`, the output fold over `r0..r7` | `verify.rs:555` to `:654` | 55,296 instructions and 128 dependent dataset reads per lane; the warp is the unit |
| The verdict | `hash_bound(H, nonce) <= target64(bits)`, `target64` the top 64 bits of the 256-bit target | `bind.rs:128`, `igneum.rs:120`, `:173` | one 64-bit compare |
So a winning attempt is a function of `(H, nonce_hi, n, program(e), dataset(day), target)`, and every one of `H`,
`e` (through `daa_score` and the parents, both in `H`) and `day` (through the timestamp, in `H`) is under proof of
work. The expensive quantities are three: the program (per epoch), the dataset (per day), and the hash itself (per
warp). The cheap variables are the nonce's two halves and anything in the template the miner may change.
### 1.2 What is in the template that a miner may change, and what each change costs
| Field | Who sets it | Freedom | What changes downstream |
|---|---|---|---|
| nonce, low 32 bits (`n`) | the worker | 2^32 per `I`, iterated 32 at a time | the lane registers only; the warp's 32 lanes are hashed together |
| nonce, high 32 bits | the miner (`main.rs:565`, a random `hi`, incremented when the lane space wraps) | 2^32 per template | `I`, hence every register of every lane: a new warp space |
| timestamp | the miner, within Kaspa's bounds (10 s future tolerance, above the sampled past median) | about 10 s of milliseconds, 10,000 values | `H`, hence `I`; and the DAY at a day boundary (1.4) |
| the coinbase and transactions | the miner (mode A or C) or the pool | unbounded | `hash_merkle_root`, hence `H` |
| parents | the miner, among known tips within the merge bound | a few | `H`, `daa_score`, `blue_work`, `blue_score`, and the epoch seed if the selected parent's chain differs below the seed score |
| `vote_key_hash` | the member's client (spec 9.4.1 item 1) | one per key the miner holds | `H` |
| bits, `daa_score`, `blue_score`, `blue_work`, `pruning_point`, `accepted_id_merkle_root`, `utxo_commitment` | the node from the parents | none | |
Every one of these is an "extra nonce": it changes `H`, so `I`, so every register from the first instruction. None
of them is free of the hash: a template change costs the whole warp again.
## 2. The nonce's place in the chain and the warp as the unit
The lane nonce enters once, at register initialisation, XORed with each init word before splitmix32 (`verify.rs:570`).
Nothing later reads the nonce. So the hash of lane `n` is the composition `fold(program^8(init(n, I)))` and the only
state shared across lanes is what `shfl` moves (lane `l` reads lane `l XOR mask`), which couples the 32 lanes of the
aligned group into one computation. The consequences:
- A warp's 32 hashes are 32 attempts for the price of 32 lanes of work; no lane's result is reusable for another
nonce because every register of every lane starts from the nonce.
- The high half of the nonce is one `I` per 2^32 lanes. A miner that changes `nonce_hi` pays one FNV over 49 bytes
(about 200 ops, `bind.rs:48`) and gets a fresh warp space; it does not need to, since 2^32 lane nonces at 141.8
MH/s on an RTX 5090 (attack pass F9, the card row) is 30 s of work per `I`, 30 templates' worth at 1 block per
second. The crate's own test `bind::bound_hash_properties` pins that two `nonce_hi` values give different warps
under one `H` and that the lane hash depends on `H` (the fork's engine test pins the timestamp, `igneum.rs:237`).
- The pool's share check recomputes `hash_bound(job.prehash, nonce)` for the claimed nonce (`pool/src/verify.rs:52`),
so a share binds the whole template through `H`; a nonce under another template "hashes differently and is
refused as a wrong hash" (the crate's test `the same nonce on another template`, `verify.rs:101`).
## 3. What the 256 loads and the shadow block compute from
### 3.1 The loads
Each of the 16 load sites per iteration (128 per hash) reads `dataset[load_index(era, site, x)]` where `x` is the
source register's value in that lane at that instruction (`verify.rs:759`): `idx = x AND MASK` without an era, or
`rotl(x M, R)` masked into the site's window under an era (`load_index`, `:18`). The address is therefore a function
of the lane's register state, which from the first instruction is a function of `(I, n)` and, after the first load,
of dataset words too. The attack pass measured the part that is a function of `(I, n)` alone (F9, sub-row (c), taint
analysis `init_determined_sites`): on the devnet epoch-0 program the loads at instructions 7, 8, 9, 10 and 31 of
iteration 0; over 2,000 seeds 1 to 7 sites per program, median 3, always in iteration 0 only. Everything after the
first dataset word is data-dependent.
The 256 loads of a hash's "working set" are the 128 reads of its own lane plus the reads the warp's other lanes
make, and they are distinct per hash by construction: the acceptance rule refuses a program whose site reads fewer
than 120 distinct addresses per hash on average (`MIN_DISTINCT_SUM`, spec 1.4.6.4), whose 32 lanes compute one
address at any site (`LaneConstantSite`), whose site's distinct-index ratio over 2^20 evaluations falls under 0.98
(class v4, (c'')) or 0.995 (class v5, (c''')).
### 3.2 The shadow block
The shadow block is 256 ALU instructions from the ten non-load families (`generator.rs:413`), applied 27 times after
instruction 63 of every iteration with the iteration's `sel` (`r0` at the iteration's start; `verify.rs:609`,
`:631`). It has no loads. Its input is the lane's eight registers at that point and `sel`; its output is the
registers for the next iteration. It is therefore a per-lane function of the lane's state, with `shfl` coupling
lanes inside it as in the base block. Nothing in it is shared across nonces, iterations or templates: the same block
applied 27 times to a different register state each time. The attack pass's F1 row measured what an honest compiler
can remove from it (at most 4.688 percent of its instructions on the frozen class v5 tip at 10^5 programs, two
programs at 5.078 percent in the 2.2 x 10^5 partial, both at the honest-compiler parity the ruling AP-F1-1 names), so
the block's cost per application is fixed to within that margin for every implementer.
## 4. Where a result could be reused, each with its cost and the rule or the finding
The seven places an expensive intermediate could be shared across cheap attempts, in the order a miner would try
them. "Rule" names what forbids the reuse or makes it worthless; "cost" is what the reuse would save against what it
costs.
| # | Reuse | What is shared | Across what | Cost to the attacker | Rule or finding |
|---|---|---|---|---|---|
| R1 | The dataset | the day's 1 to 2 GiB and its 256 MiB cache | every hash of the day, every miner | nothing: this is the design (the memory-hard dataset is the shared intermediate on purpose) | Intended. The question is whether a SMALLER shared structure serves: the hot set (R5) and partial storage (R6) |
| R2 | The compiled program | the epoch's kernel | every hash of the epoch, every miner | nothing: intended; one compile per epoch per card | Intended. A miner that grinds the epoch seed to get a favourable program is R7 |
| R3 | The init words `I` | one FNV result | 2^32 lane nonces | nothing to gain: `I` costs about 200 ops and a warp costs 1.8 million lane-instructions (F9's count) | No rule needed; the ratio is 10^-4 |
| R4 | The register prefix before the first load | the first instructions of iteration 0, a function of `(I, n)` | nothing: `n` is per lane, so there is no second attempt with the same prefix | zero reuse: each lane's prefix is its own | Bound by construction (spec 1.6: the nonce enters every register) |
| R5 | A hot set of items | a few MB that serves a large share of reads | every hash of the epoch | the attack pass's F8 and F9 (b): at the frozen tip no program's top 0.1 percent of items carries more than 0.23 percent of reads (256 seeds, 245 within 1.2x of the window model), hot-set verdict clear on every seed; the exemplar that reads 3.35 percent (class v4 seed 100767) is refused by (c''') | AP-F8-1: the residue is a quarter-bit bucket concentration at one narrow-window site, named by mechanism, no cacheable hot set; the per-site largest-bucket bound is a next class's item |
| R6 | A partial dataset or cache | a stored fraction `f` of the cache lines, the rest recomputed | every hash | F3: every cache line costs exactly `j + 1` block evaluations from nothing (0 of 1,024 lines under the bound), the storage-against-recompute curve at f = 1/8 is 3.17 ops per read optimal against the 3.5 the funding note assumed, and a partial-cache chip is worse than the full mirror at every f below 1 | F3 PASS; the chip edge is the latency ladder's business, not a binding hole |
| R7 | The epoch seed | the program of epoch `e + 1` | every hash of that epoch | a miner who produces the seed block (the last selected-chain block below `3,600 (e + 1) - 600`) picks among the candidate blocks it can produce; each candidate costs a full block's work, and the 600-score lead gives every miner the same 10 minutes to compile; under the devnet stand-in there is no VDF on the epoch seed (spec 4.3 Designed, the era VDF Implemented) | Spec 4.3 (the 1,200-s lead and the 600-s VDF) is the rule; on the node line the lead is 600 and the VDF is not in the node. The grind's value is bounded by the program-to-program variance of hash rate: the attack pass's F8 and F1 rows put every accepted program within a few percent of the mean on the card (the per-program spread is the F10 ladder's reading), so a candidate block burned to pick a 2 percent better program is a bad trade at any hashrate under 50 percent. Finding, minor: the VDF of 4.3 is the stated defence and is not implemented; the lead alone does not stop the pick. Routed to the Counter ASIC lane's D1 list as a 2.0 item (6) |
| R8 | The day | the dataset of day `d` | every hash of the day | the day is keyed on the header TIMESTAMP (`pre_ghostdag_validation.rs:147`), which the miner sets within about 10 s; at a day boundary a miner may hash on either day's dataset for about 10 s, and a verifier must hold both caches across the boundary (the engine's `note_chain_day`, `:139`, builds the next in the background) | No gain: both datasets are full datasets. Observation: spec 1.12 keys the day on DAA score and the node keys it on the timestamp; one of the two texts moves (6) |
| R9 | The attempt chain | the acceptance work of attempts 0 to k-1 | every node and miner of the epoch | nothing to a miner: attempts are deterministic from the seed; a miner cannot choose the attempt. A verifier's cost is the row below | No rule needed |
| R10 | The warp | one hash computation | 32 nonces | intended: the 32 lanes are 32 attempts for 32 lanes of work, coupled by `shfl` so no lane is computable alone | Intended (spec 1.9). The verifier dedupes items within a warp (`verify.rs`, the 4,096 derivations bound); a miner that finds warps whose loads coincide is F9 (c): one coalesced pair per 3,400 tries is worth 0.2 percent of one warp and costs 2.4 hashes of search; five sites fully coalesced would be worth 43 percent and has probability 2^-115 per draw |
| R11 | The template | one `H` | 2^64 nonces | intended; a template change is a new `H` and costs nothing but one BLAKE2b; it buys nothing, since the nonce space under one `H` is 2^64 | No rule needed |
| R12 | The vote key | one `vote_key_hash` | every block of the key | a miner with many keys hashes under one `H` per key; keys are free (spec 03 W6) and weight is blocks, so more keys is more templates, not more hashes per hash | No gain; the pool design document (pool-vote-key-commitment.md) covers the pool case |
The cheapest reuse that survives as a gain is R10's coalesced pair at 0.2 percent of one warp per 2.4 hashes of
search, which is a loss, and R5's quarter-bit bucket, which is under the window model's own spread. No path was found
by which an intermediate computed once serves a second winning attempt at less than the honest cost of that attempt.
## 5. Known-failed tests and measured lines
| Test | Known-pass | Known-fail | Where | Result |
|---|---|---|---|---|
| The lane hash depends on `H` and on `nonce_hi` | two headers differ in one field: different warps | the same `(H, nonce_hi)`: the same warp, bit for bit | `igneum-pow/src/bind.rs`, `bound_hash_properties`, `bound_vectors`; the fork's `igneum.rs:200` smoke test (timestamp bound, `:237`) | see the run line below |
| A nonce on another template is refused | the pool's check against its own job | the same nonce on another template: `wrong_hash` | `pool/src/verify.rs:101`, `the same nonce on another template hashes differently` | in the pool crate's suite (pool.md section 4) |
| A stateless hasher is wrong on every class v5 item | the dataset built with the leaves | the dataset built without them: every item differs | `igneum-pow/src/state.rs:200`, `a_stateless_hasher_is_wrong_on_every_item` | see the run line below |
| No cache line costs less than `j + 1` | the real chain | a flattened chain (the F3 harness's known-fail) | `docs/analysis/attack-pass/f3-cache.md`, the two firings | PASS, 7 October |
| No hot set on the frozen tip | the window-model control | a planted hot site (F8's known-fail) | `f8-uniform.md` sections (d) and (e) | PASS, 61 of 64 and 245 of 256 within 1.2x |
| Init-determined sites are in iteration 0 only | the taint trace | a program with a load whose address never absorbs a word (F9's planted case) | `f9-grind.md` sub-row (c) | 1 to 7 sites, iteration 0 only, 2,000 seeds |
Run line (box 2 through `tools/build-remote.sh`, `cargo test --release -p igneum-pow -- bind::
state::a_stateless_hasher_is_wrong_on_every_item`): 8 October 2026, 19:0x UK, box 2 (igneum-build-2) through `tools/build-remote.sh --box 2`, the crate at the mirror's master 494fb981: `bind::tests::bound_hash_properties` ok, `bound_vectors` ok, `init_bytes_layout` ok, `day_bytes_layout` ok, `pow256_and_target64` ok, `closed_form_bound_also_works` ok (6 passed, 0 failed, 0.43 s); `state::tests::a_stateless_hasher_is_wrong_on_every_item` ok (1 passed, 0 failed, 2.15 s). The pool crate's `the same nonce on another template hashes differently` and the fork engine's timestamp-bound test are in their crates' suites (pool.md section 4; `igneum.rs:200`), not re-run here.
## 6. Open questions, named for main
| Id | Question | What closes it |
|---|---|---|
| B1 | The epoch-seed VDF (spec 4.3, 600 s of delay on the seed block's hash) is Designed and not in the node line; the era VDF is. Without it the seed-block producer picks among the candidates it can afford to burn (R7). Is the epoch VDF a 2.0 item, or does the measured program-to-program hash-rate spread (the F10 ladder's reading) close it as not worth a block? | A ruling, with the F10 spread as the number |
| B2 | The day is keyed on the header timestamp in the node (`day_index(header.timestamp)`) and on DAA score in spec 1.12 (R8). Which text moves? The timestamp key makes the day boundary a miner's choice within 10 s and costs the verifier two caches across it; the DAA key makes it a function of the past alone | A ruling; a one-line spec or code change |
| B3 | `seed_source` and `proof_ref` (spec 2.4, fork point a5) are not in the node line's header; the epoch seed is bound through `daa_score` and the parents instead, which is sufficient for binding but leaves the spec's field table ahead of the code | The spec's table marked to the code, or the fields added when the chunked proving protocol needs `proof_ref` |
| B4 | A verifier syncing from genesis derives every epoch's program (the draw plus acceptance, about 6 core-s per epoch at 3.2 candidates): about 15 core-hours per year of chain on one core, parallel across epochs | The IBD cost measured on the node line with the memo (`epoch_seed_memo`) and a per-epoch program cache; a number, not a rule |
| B5 | R10's coalesced-pair search (F9 (c)) is measured as a loss on the RTX 5090; it is not measured on a card whose load completes per lane rather than per warp (an AMD wave64 under the local-memory exchange) | One run of the F9 `grind` per-warp table on a 9070 XT |