igneum/docs/design/class-v5-stored-state.md

76 KiB

Class v5, proof of stored state: the daily dataset derives from the chain's own execution state

7 October 2026, morning UK. Horizon lane 8's first pick (docs/analysis/horizon/new-pow.md section 6, scheme C "sd1", measured on rented RTX 4090s on 6 October: hash rate and watts unchanged, build +1.4 ms, verifier +0.11 to 0.21 ms per unit). Branch class-v5 (repo) and class-v5-node (fork, from release-0.3.18-node e69e8a39). Status: implemented behind program_class_v5_activation_daa (never until set); not in docs/spec until adopted; nothing here touches the devnet.

0. Progress (kept current for the coordinator)

Time (UK) State
08:05 Lane started. Read: CLAUDE.md, new-pow.md whole, proto-newpow/state-dataset (README, sd.h), the latency ladder page, igneum-pow (generator, memhard, verify, emit, main), the fork's class signal (consensus/core/src/igneum.rs, processes/class_signal.rs, the header processor's seed path), the engine (consensus/pow/src/igneum.rs), the exec layer (state, snapshot, service, records), the miner's day handling, the class v4 signal harness, the ladder harness and bench script, spec 01 1.8.5 and 1.12, spec 04 4.3
08:20 Fleet agent asked for one quiet 4090; answered: a fresh RunPod 4090 pod from 10:00 UK to 12:00 UK
08:33 The canonical day stream written in the fork (igneum/exec/src/day_stream.rs) and run against node 1's exec snapshot on igneum-build-1: 93 records, 8,619 bytes, root recomputed and equal (section 7)
08:36 This page committed (ad25d2c3); the fork's day stream (f32641d1)
08:4x The coordinator relays the project lead's widening to proof of following; section 2a written (the refresh per window from the epoch's seed block, the pool and farm bound with the numbers: the floor is the state size, 6 KB today, 45 MB per member per hour is the WAN line at 10,000 members, a LAN farm is never bounded)
08:42 Section 2a committed (3cbe8c0f)
08:56 FIRST GREEN: igneum-pow class v5 (c6bbde78): blake2b, state leaves, the leaf XOR behind Shape.state, V5_CLASS, generator 5, the emitters' leaf buffer, the CLI's --state; 66 unit and 39 integration tests on igneum-build-1, the known-failed case (a stateless hasher) first
09:10 The fork builds (class-v5-node): ProgramClass::V5, program_class_v5_activation_daa (never, in the digest once set), the byte-5 tally and the one-step rule, the executor's per-epoch state captures and igneum_getPowStateLeaves, the engine's DayStateProvider and DayStateUnavailable, the miner's RPC provider with --exec-rpc and --freeze-state, the template's v5 fields (37 to 40); kaspa-pow 16 tests green; the harness infra/fast-time/class-v5-signal.mjs and the 4090 bench committed (e528cb05)
09:30 Fork suites on igneum-build-1: kaspa-pow with the feature 16 passed (the engine's known-failed case class_v5_refuses_without_state_and_refreshes_the_leaves_per_epoch among them), igneum-exec 28 passed (a_tampered_stream_is_refused, epoch_state_captures_follow_the_cut_and_unwind), kaspa-consensus class_signal 1 passed, kaspa-consensus-core 122 of 123: the one failure is fast_time_60x_file_is_the_devnet_at_60x demanding the emission field of the 0.3.18 fork's OverrideParams in the repo's 60x file, which master's file lacks as well (the economy lane's field, not this lane's; the v5 field is in the file and read)
09:33 The pinned packs exported on the box from the devnet's IGSD1 stream (proto-cuda/packs-ca3-v5/, 4d10ac56); the packs test 20 passed: the v4 control pack agrees with the v5 epoch on 0 of 32 lanes, the hash kernel text equal, the build kernel and leaves the difference
09:37 The verifier measured (section 7): v5 adds 0.28 ms cold and 0.20 ms on average over v4 on the loaded reference core; both under 10 ms
09:37 The harness's known-failed case started on the box (tools/class-v5/harness-remote.sh failed-case -- --signal 5,5,4 --expect flip)
09:48 The known-failed case FAILED as it must: no flip at 7,519 bps of byte 5 (two signalling nodes of three plus the stateless node), 665 blocks, 0 rejected, four sinks equal; the harness's own two faults fixed (e0c555cb)
10:00 The first flip run read a stranger's chain (two nodes of the known-failed case were still bound to its ports); the harness now kills its own previous pids and fails on a bind error, with a port base per case (6c983cba); the miner's engine panicked on a state refusal while the executor had not yet captured the seed block: it now says the refusal once and asks again every two seconds (bdc3bb34); the flip re-run started 10:24
10:23 The fleet's 4090 pod (cv5-2) measured (section 7): rate equal within 0.01 percent, build +0.24 ms, watts a warm-up drift with the classes interleaved, every vector and sample bit-exact
10:38 The flip re-run (--signal 5,5,5 --expect flip --stale 2, ports 29830 and up): the chain flipped v4 to v5 at epoch 8 by signal on 3 of 3 nodes (Program class v5 by miner signal: epoch 8 ... weakest 10000 bps), the first epoch with seven full windows after v4; and node 0's templates timed out from the flip on: the engine's provider took the executor's state lock from inside the template path while the follower held it and waited on consensus (a lock-order inversion), and the miners' stream fetches timed out the same way. Fix: the follower publishes each finalised epoch's stream (its cut passed, the stream rebuilt to its root) into a separate cache the provider and the RPC read without the state lock; the harness points the stateless node's miner at its own node's exec port. The 60x profile's day boundary was crossed inside the earlier runs (days 1244001 to 1244003 in the summaries)
11:17 to 11:30 The flip run on the lock-free stream cache: the flip at epoch 8 on 3 of 3 with equal roots, and no block mined under v5: the miner's stream reader skipped 14 characters past the streamHex key, which is 15, so every fetched stream began at x (3237e842)
11:31 to 11:46 A run that stalled at DAA 65 with the box at load 97 (other lanes' builds): all four CPU miners found no block for twelve minutes after the difficulty had adjusted up; no panic, no OOM; not the lane's
11:47 to 11:59 The flip run on the quiet box: the chain did everything the design says (the flip, four v5 epochs, a root per epoch, the stale miner off on 52 of 52, the stateless node stopped at 480); six FAILED CHECKs were the harness's own scope (it compared the stateless node's sink with the others, counted the stale miner's rejections as a fault, and read the flip epoch's stream after the executor had dropped it: captures are kept two epochs); the checks fixed
12:00 to 12:11 PASS on every check (section 10)
12:12 to 12:23 The no-flip case PASS (section 10); both branches pushed; master merged for the build-2 test route (d1285cef)
18:5x to 19:07 Resumed on the project lead's first-priority order (build class v5 now, on the 0.3.23 line, byte 6): the frozen sub-version 3 (ca3-v4-amend 017e7037) merged into class-v5 and release-0.3.23-node into class-v5-node; the draw's and the acceptance rules' class keys set the state flag aside (the acceptance key had judged a v5 candidate under the pre-amendment rule and the draw diverged at instruction 0; fixed, 70 unit tests green on build-2 at 18:04Z); class v5 pinned to object byte 6 counted exactly (byte 7 is v4 sub-version 3); the gap list as section 13; the FIRST CLASS V5 PACK proto-cuda/packs-ca3-v5/v5-dn3-epoch0 (Devnet 3's genesis 4020cb43... as epoch and era seed, day 20,733, the genesis state's 11 leaves under root 7e37a9fb...), exported on build-1 at 18:06:39 to 18:06:41Z, program id e5a4ac5978462156, its Metal fingerprint on the M5 Max under the measure lock at 18:07:04 to 18:07:09Z: 82b19cbde8557ea5, the three vector warps bit-exact, the build 29.8 ms GPU; the Metal pack bench takes leaves.bin
19:2x to 19:32 The AP-F4-1 mixer-draw rule and the AP-F1-1 shadow rule written behind the v5 class with their known-failed cases (2f9bf135; the suite running on build-2); the M5 Max rows measured (section 7: +0.23 ms per warp); the node side handed to the node lane, the kits to the v5-kits lane, the families to the attack-pass lane, a 5090 or 4090 fingerprint run of the first pack asked of the fleet lane
13:2x The coordinator's correction: the hot-set bound is NOT satisfied by sub-version 1 (F8's re-gate: nine of the first 30 seeds over 1.2x, p31 at 29.3x, 0.19 percent hot-set programs; saturated values pass saturation-preserving writers); v5 inherits the gap by merge and takes sub-version 2 by merge when it lands; the AP-F4-1 and AP-F1-1 rules come in the next round; the lane stops here
12:5x to 13:21 The amended class v4 taken in: ca3-v4-amend 8c728ca3 merged into class-v5 (4e737543) and release-0.3.20-node into class-v5-node (699db5a2); the source rule's class key sets the state flag aside so v5 and its rungs draw under it (a unit test pins the v5 chain draw equal to the amended v4's with no lossy-sourced load, the unamended v3 stream the known-failed case); generator 5 keeps program_id(5, seed, attempt) with no sub-version (the hash lane's reading: the suffix separates two streams inside generator 4); class v5 is object byte 6 above the amended v4's 5 (the fork, the harness, the default byte); igneum-pow 67 unit and 39 integration tests on igneum-build-2, the packs test with the re-exported control pack, the flip case PASS again with bytes 6,6,6 (the flip at epoch 8 on 3 of 3, 183 v5 blocks, the stale miner off on 59 of 59, the stateless node stopped at 480, equal roots)
10:4x The Counter ASIC coordinator relays the project lead's ruling of option A on AP-F8-1: the injecting-source load draw lands in class v4 itself on the hash lane's branch ca3-v4-amend (0.3.19); class v5 takes it by merging that branch once its commit is up, not by a second implementation; the page's section 11 row stands as the record of the bound and its gate
19:4x The flip-stale harness re-run (base 30500) read FAIL on v5_ids_equal_the_cli_v5_id, sinks_agree, block_counts_agree: the fork binaries on build-1 date from 18:56 UK, before AP-F4-1 (19:27) and AP-F1-1 (19:37), so the miners' v5 ids differ from the CLI's at redrawn epochs (9 and 10 differ, 11 equal). Binary skew; the fork rebuilds after (c''') and the case runs again on the matched pair. The harness itself: the 60x file now carries other lanes' override fields (sig_scheme, sig_scheme_activation_daa, finality_succession_activation_daa, latency_ladder_cache_rung) that the fork's OverrideParams (deny_unknown_fields) refuses; the harness drops the fields the fork source does not name and says so in its log (1c02465c)
19:5x Main's order through the coordinator: class v5's acceptance carries the fix for the residual hot-set class (adv-accept's seed 100767, section 14), the cheapest fix first (the raised per-site floor) against the per-site hot-item test, chosen by the lower clean rejection rate. (c''') in at ab6f980b: MIN_DISTINCT_RATIO_V5 = 0.995, known-failed first on the exemplar, the census test v5_hot_census behind the number
20:0x The v5-fasttime lane's exception, owned by the node lane: class_signal.rs decide() with the v5 object set and both floors at 0 resolves the epochs under the window to class v3 regardless of the v4 floor (e0:v3 e1:v3 e2:v4 on the 60x profile); on Devnet 3 (v4 from genesis) enabling v5 would re-read day one as v3. This lane's harness never saw it (its v4 floor is 120, v3 at 60). Gates the crossing object, not the hash side

1. The claim, in one paragraph

Under class v5 the day's dataset is built from the execution state at a reference block of the chain, so an item cannot be derived without that state and every hash proves the mining operation holds the chain. The hash kernel is class v4's byte for byte (the base program, the 16 loads, the latency-shadow block, the era draw and the ladder rung are v4's); what changes is one line of the item derivation (s[i] ^= leaf(t)[i] before the first mixer) and the daily build. The cost is measured and small (new-pow 5.2: rate and watts equal within noise, +1.4 ms on the build, +0.11 to 0.21 ms per verifier unit with the leaves in RAM). What it does not do: it moves nothing against the dataset-storing chip (the f = 1 chip stores whatever the items are; new-pow 5.3), and it does not claim to. What it removes: the f = 0 recompute chip as a category, and the pool miner who needs nothing but the day key.

2. Which state, and which block

Item Rule Why
The day the devnet's day index d = header.timestamp_ms / day_ms (docs/fork-divergence.md, kaspa_pow::igneum::day_index); the spec's DAA-day form (spec 01 1.12) takes the same rule with the DAA cut in place of the timestamp cut the class v5 dataset is the day's dataset with one more input, so it is keyed by the same day as the cache
The cut cut(d) = d x day_ms - lead_ms, lead_ms = day_ms / 24: one hour on the devnet (86,400,000 / 24), one minute on the fast-time profile (1,440,000 / 24 = 60,000 ms = 60 DAA, the profile's merge depth) the state must be known before the day starts and deep enough that a reorg across it is a merge-depth-scale event (spec 04 4.3, O-4.3: "deep enough that a reorg across it is a merge-depth event"); the epoch seed's lead (600 DAA) is a tenth of the merge depth and was chosen for a 10-minute VDF; the day's reference block is read by every block of the day, so it takes the whole merge depth
The reference block C_d walking down the header's selected chain from its selected parent, the first block whose header timestamp is below cut(d); genesis when none (a chain younger than the cut) a function of the header's own past, as the epoch seed block is (HeaderProcessor::epoch_seed): two nodes validating the same header derive the same C_d; a header on another chain derives its own C_d, which is the epoch seed's acceptance (spec 04 4.3 item 4). Kaspa's timestamp rule (above the past median, at most 2 min ahead) bounds the walk's non-monotone blips to seconds, so C_d is fixed within a minute of the cut and the day starts 59 minutes later
The state the execution state AFTER chain block C_d: the record's state_root (ChainBlockRecord.state_root, an output of execution, never a header field, design 2.2), written R_d every node that executes holds it; the exec snapshot ring already clones the state per chain block (ExecState.snapshots), so capturing the day's state is one more clone kept for the day
Certified or not the selected-chain block, certified or not mining never waits on finality (spec 04 4.3's argument, section 3.7 item 2): a finality pause (6 October, 18:42Z) does not stop the day's dataset. The lock usually lands within a minute of C_d anyway, so in practice R_d is finalised an hour before the day starts
The canonical stream DayStream (igneum/exec/src/day_stream.rs): accounts ascending, only those the state root counts (EIP-161); after each account its non-zero storage slots ascending; then its code in 64-byte chunks. Records: account `0x01
The leaves `D[i] = Blake2b-512("igneum-sd1/"
The leaf of item t leaf(t) = D[t mod n] EVERY item takes a leaf: a dataset where only the first n items carried state would let a stateless hasher be right on (1 - n / items)^128 of its hashes, 92.6 percent at today's 93 records against 2^24 items; with every item keyed, a stateless hasher is wrong on every item (the known-failed case, section 4). Items t and t + n share a leaf and differ by t in the init (s[8..15] = t x MUL + RC), as every item already does
Too much state when n > items (2^(D - 4), 2^24 at 1 GiB, 2^25 at the designed 2 GiB): the records are ordered by `Blake2b-256("igneum-sd1-sample/"
The item class v4's derivation with s[i] ^= leaf(t)[i] for i in 0..16 after the init and before the first mixer (spec 01 1.8.5, prototype sd.h mh_item_sd) the prototype's line, bit-exact on 1,024 items and 128 lanes on the 4090; the leaf enters a keyed chained derivation whose read addresses are mixer state, so no bias on 128 dependent reads is reachable below a block's cost (the MTP lesson, asic-resistance-history.md 2.3)
The program class v4's draw with generator 5 stamped (GENERATOR_VERSION_V5, V5_CLASS = { state: true, ..V4_CLASS }, program_id(5, seed, attempt)), the ladder's rung riding as it does for v4 (v5_class_at(reps)) a v5 pack and a v4 pack of one seed never share an id, so a worker on the wrong class refuses the pair (the id check of M28)
The kernel text unchanged for the hash; the build kernel gains a leaf buffer: igneum_build(ds, cache, leaves, nLeaves, nItems) with mh_item_sd(cache, leaves + 16 (t mod nLeaves), t, s) in the three dialects the prototype measured +1.43 ms for the one coalesced 64-byte read per item with a 1 GiB leaf array; at today's 6 KB of leaves the read is cache-resident and cheaper

Under proof of following (section 2a) the day keys the cache alone and the leaves refresh per window from C_w; the paragraph below describes the daily rule as first designed, which 2a replaces for the leaves. What moves with the day: at cut(d) every node's executor captures the day's state; from then on D for day d is computable, and the miner's existing one-lead-ahead prepare of the next day's pair (igneum-miner, day_eta_s <= epoch_lead: 10 minutes before the day on the devnet) builds the next dataset with it. The switch at the boundary is the day switch the miner already makes (a prepared pair swaps with no pause); the build itself is 32 ms on a 4090 (31.98 ms measured, new-pow 5.2) against the 1.8 ms cache fill, so even an unprepared worker loses under 50 ms of hashing at the boundary. No hash-rate dip by construction: the state is known 60 minutes before the day, the pair 10 minutes before.

2a. Proof of following (the project lead's widening, 7 October 2026, 08:4x UK, through the coordinator)

The daily dataset of section 2 proves the miner holds the chain once a day. Proof of following makes the leaves refresh every WINDOW from the chain's latest reference block, so a miner must keep executing and following the chain to keep mining: a machine that stops following builds the previous window's dataset at the refresh and is wrong on every hash from then on.

2a.1 The refresh rule

Item Rule Why
The window Params::pow_state_window_daa, a genesis parameter in DAA seconds; 3,600 on the devnet (one hour, which at the base epoch length is one epoch; the two candidates coincide there and the parameter keeps the refresh where it is when the epoch ladder of spec 01 1.12 moves epoch_len); 60 on the fast-time profile; in the digest once program_class_v5_activation_daa is set (the 0.3.15 rule) a refresh per hour is 32 ms of rebuild per hour on a 4090 (section 2a.3), so the cadence costs nothing; one epoch is the natural unit because the miner already swaps programs there
The window index w = daa_score / W (the block's own DAA score, as the epoch index is) a block's window is a function of its header
The reference block C_w the last selected-chain block below DAA score W x w - W / 6, walking down from the header's selected parent; genesis when none. With W = epoch_blocks and the devnet lead of 600 this IS the epoch's seed block (HeaderProcessor::epoch_seed), so the node pays no second walk and the state's block is the block the epoch seed is already taken from a function of the header's own past; the lead (W / 6, 10 minutes at the devnet window) is the epoch seed's lead and the miner's existing prepare horizon
The state the execution state after C_w: R_w = records[C_w].state_root, the stream DayStream of section 2 taken at C_w the executor's ring already clones the state per chain block; the capture keeps the one at C_w for the window and the next
Finalised or not the selected-chain block at the cut, certified or not; in practice the lock lands within a minute of C_w, ten minutes before the window starts, so the block is finalised when mined on; by rule mining never waits on finality (spec 04 4.3, section 3.7 item 2) "the latest finalised checkpoint" as a RULE couples mining to finality liveness: a pause longer than the lead would stop every miner, which the 6 October pause (18:42Z, 20 minutes) would have done; the cut rule gives the same block in every normal hour and keeps mining through a pause
The leaves `D_w[i] = Blake2b-512("igneum-sd1/"
The dataset of (day d, window w) the day's 256 MiB cache (the day key, unchanged) with D_w folded into every item: a rebuild of the dataset per window, never of the cache the cache is a function of the day key alone and stays; the build is the measured 32 ms on a 4090 (section 2a.3)
How a miner learns it the template's powEpoch carries stateBlock (the hash of C_w) and nextStateBlock once the next window's cut has passed (one lead before the boundary, with nextEpochSeed); the miner fetches the stream for that block from its node (igneum_getPowStateLeaves [block hash]), checks nothing when the node is its own (its node executed it), and prepares the next pair; the pack carries IGNEUM_STATE_BLOCK_HEX, IGNEUM_STATE_ROOT_HEX and leaves.bin the existing next-epoch prepare: with W = epoch_blocks the refresh and the program swap are one swap
The grace at the boundary none in validation (a block's window is its DAA score's; a block mined with the previous window's leaves after the boundary is invalid); the grace is the lead: the next window's leaves are knowable ten minutes before it, the miner prepares the pair then, and a prepared worker swaps with no pause (the 2.0 hot-swap path). An unprepared worker rebuilds at the boundary: 32 ms of hashing lost on a 4090, under 0.001 percent of the window no hash-rate dip by construction, as the epoch swap has none today
A miner whose node is behind the node serves no template for a window whose C_w it has not executed (section 5's refusal), so the miner's hash stops at the boundary rather than mining invalid blocks a node ten minutes behind the chain is already a node without a useful template

2a.2 What a pool can and cannot centralise, and the farm attack, with the numbers

The attack: a farm rents one node for ten thousand cards, rebuilds the dataset once per window on that node and ships it. The bound the design can give, honestly:

Fact Number Source
The dataset the cards hash over 1 GiB (2^24 items; 2 GiB designed) spec 01 1.13.3
What a card needs to rebuild it itself the day key (32 B, public) and the window's leaves D_w (64 B x min(n, items)) section 2
D_w at today's devnet state 5,952 B (93 records) igneum-build-1, 08:33 UK
D_w at a mainnet state ten times larger 60 KB arithmetic
D_w at the sample cap 1 GiB at the 1 GiB dataset, 2 GiB at 2 GiB section 2
Shipping the built DATASET to 10,000 cards per hour 1 GiB x 10,000 / 3,600 s = 2.98 GB/s (24 Gbit/s); 2 GiB: 5.96 GB/s (48 Gbit/s) arithmetic
Shipping the LEAVES instead, today 5,952 B x 10,000 / 3,600 s = 16.5 KB/s arithmetic
Where the leaves become too large to ship inside the window over a WAN pool link of 1 Gbit/s to 10,000 members 1 Gbit/s x 3,600 s / 10,000 = 45 MB per member per window, which is n above about 700,000 records (7,500x today's devnet state; approximate) arithmetic
Over a farm's LAN at 100 Gbit/s 4.5 GB per member per window: never, under the 2 GiB cap arithmetic

So the design target "a dataset too large to ship inside the window" is NOT reachable by construction while the state is small, and this page says so rather than claiming it: everything the lottery derives is derived from a seed and the state, so the only bytes a central node cannot compress away are the state's, and the state is the floor. A farm with one node ships D_w (6 KB today) to its cards and the chain cannot tell a card that derived the leaves from a card that received them, exactly as it cannot tell a pool member from a solo miner today. What proof of following does force, per machine and per window: holding the current window's state (or its leaves), a rebuild from it, and knowledge of the chain's reference block inside the lead; a machine cut off from the chain for one window stops producing valid blocks at the next refresh, where under the daily rule it kept mining until midnight. The bound bites on the WAN above about 700,000 state records (the leaves pass 45 MB per member per hour at 1 Gbit/s for 10,000 members) and on a LAN never; a used chain crosses the WAN line early (a chain with a million accounts does), a young chain does not.

What a pool can centralise: templates (as today), the day key (public), the stream or the leaves (a 6 KB to 2 GiB delivery per member per window), the built dataset (1 to 2 GiB per member per window). What it cannot: the rebuild (per machine, 32 ms on a 4090 per window) and the clock (a member that misses a window's delivery is wrong on every hash until it gets it). What the rule "members must hold and refresh the state themselves; the pool serves templates only" means in practice: it is a rule a pool can state and a member can follow, and nothing on the chain enforces it below the WAN line; above the line the link enforces it. The pool protocol (spec 09) gains the per-window state-leaves message and the page records that a pool serving leaves is serving the state, which is the design's "every mining operation holds the chain", not every card.

2a.3 The cost per refresh

Machine The rebuild (the dataset from the cache and the leaves) The leaf pass (n Blake2b-512) The stream check (to_db and the root, only when the leaves are not from the machine's own node) Per window at W = one hour
RTX 4090 31.98 ms resident (the prototype's second-pass build, new-pow 5.2), about 76 ms chunked over PCIe at the 1 GiB leaf array; at today's 6 KB of leaves the resident figure, cache-resident reads 93 hashes, under 0.1 ms on the host 0.1 ms at today's state 0.0009 percent of the hour (32 ms)
RTX 5090 13.4 ms for 1 GiB (the memory-hard build, docs/bench-log.md) plus the leaf read, approximate 15 ms the same the same 0.0004 percent
A laptop (M5 Max, Metal) the 1 GiB build on the M5 Max is the GPU item build of spec 01 1.12, 13 to 30 ms class, approximate until measured under the measure lock the same the same under 0.001 percent
A CPU miner (the harness's --engine igneum-pow) nothing: the CPU verifier derives items lazily; the refresh is the leaves in RAM (6 KB) the same the same nothing
A node (the verifier) nothing: lazy derivation; the leaf array swaps per window one pass per window its own execution one capture per window from the ring

2a.4 Hostile review of the refresh

Question Answer
A miner that stops following the known-failed case (section 4): it builds the window's dataset from the previous window's leaves, every item is wrong, every block is rejected from the first block of the new window; the harness runs a miner frozen on its first leaves and counts its accepted blocks after the first refresh (0)
A miner that follows with a one-window delay the same: a window's blocks need that window's leaves; there is no grace in validation
The lead is too short for a slow node the lead is ten minutes at the devnet window, the epoch seed's lead; a node that executes ten minutes behind the chain cannot serve a template today either
A reorg across C_w 600 DAA deep, the epoch seed's own exposure (spec 04 4.3 item 4, O-4.3), and with W = epoch_blocks the same block: a reorg that moves the epoch seed moves the leaves with it; every miner prepares again
A finality pause no coupling (2a.1); the reference block is the selected-chain block at the cut
The window parameter moved by a file a genesis parameter in the digest once v5 is set; a later move is a class change under the 95 percent rule
The farm with one node bounded by the state size on the WAN and unbounded on the LAN, with the numbers of 2a.2; stated, not hidden
Timestamp or DAA grinding at the cut the cut is a DAA score, as the epoch seed's; a producer chooses at most whether its own block is C_w (one bit between two honest states)
Two windows of leaves in RAM on the node and the worker the current and the next: 12 KB today, at most 4 GiB at the sample cap; the worker frees the leaves after the build (1,803 MiB resident while hashing, measured)

3. A miner with a pruned node, and what a chip must hold

Who What they hold How they get it
A full node that executes the day's DayStream captured by its own executor at cut(d), kept for the current and the next day, persisted beside the exec snapshot (day-stream-<d>.bin) so a restart inside the day does not replay nothing new: the follower passes cut(d) an hour before the day
A node that synced from a pruning proof no block bodies below its pruning point, so no execution of its own: it starts from a peer's exec snapshot already (--igneum-exec-snapshot=<file or peer>, the 6 October path); the snapshot's tip is above C_d for the current day in every normal case, so the day's state is BELOW the tip and the node cannot rewind to it. It takes the day's stream from the same peer, checks it (section 5) and executes forward from the snapshot as today the snapshot wire gains the current and next day's streams (ExecSnapshot version 2, day_streams; owed, section 11); until then the operator exports them with igneum_getPowDayState from an executing node and hands the file over with the snapshot
A GPU worker (Ember, HiveOS, a rig) the D array for the day: 64 x min(n, items) bytes, 5,952 bytes at today's devnet state (93 records); the pack's leaves.bin beside kernel.cu from its node over loopback with the pack, as the day key arrives today (prepare <epoch> <day> <pack_dir>); a worker without the leaf buffer builds nothing and the self-test refuses to serve (M28's shape)
A pool miner the same D array, or the built dataset from the pool, once a day (pool protocol, spec 09: a daily day-state fetch, owed); today a pool miner needs only the day key, which is the change the scheme makes on purpose
A chip the leaves (64 B x n) beside the cache, or the built dataset; a chip that derives items on the fly (the f = 0 recompute chip, chip-model-v3 5.4, 0.31x per chip) must now hold the day's state as well, which at a used chain's size is DRAM, so it becomes the f = 1 chip; the f = 1 chip is unchanged (5.1x on GDDR7 in the model) the same daily delivery as a miner, from a node it runs or a pool; "the chip holds the chain" is the design's claim, with the pool caveat stated

What a miner without the chain cannot do (the known-failed case, section 4): build any item of the day. A stateless hasher that fills the leaves with zeros, or with yesterday's leaves, or with leaves of another root, derives a dataset in which every item is wrong, so every hash it computes is wrong and every block it submits is rejected by every node that holds the state. There is no partial credit: leaf(t) = D[t mod n] keys every item.

Sizes:

State Records n Stream bytes Leaves D (64 B x min(n, items)) Verifier RAM beyond the cache Daily leaf pass (Blake2b-512 per record, one core) Source
The devnet today (node 1's exec snapshot, tip 159,357, root 0x1c583d35...) 93 (83 accounts, 0 storage slots, 10 code chunks of the registry) 8,619 5,952 B 6 KB under 0.1 ms (93 hashes at about 100 ns, approximate; measured in section 7) igneum-build-1, 08:33 UK, igneum-day-stream on node 1's exec-snapshot.bin
A mainnet state ten times larger 930 86 KB 60 KB 60 KB under 1 ms arithmetic
A used chain, approximate: 10 M accounts, 50 M storage slots, 1 M code chunks 61 M 5.1 GB capped at items: 1 GiB at the 1 GiB dataset (2^24 x 64 B), 2 GiB at the designed 2 GiB (2^25) 1 to 2 GiB (the sample) 61 M hashes, about 6 s on one core, 0.2 s on 32 threads (the prototype's 100 ns per leaf, new-pow 5.2) arithmetic on the prototype's rate
The chip's extra DRAM as the leaves column the f = 0 chip becomes f = 1

Per tier, what the sizes mean: a home miner with an 8, 12, 16, 24 or 32 GB card sees no change in device memory while hashing (the leaves are freed after the build, measured 1,803 MiB in every mode on the 4090) and at most one extra D buffer during the build (6 KB today, 1 GiB at the sample cap, which the chunked build keeps at 64 MiB); a rig shares one stream across its cards; a pool user takes a 6 KB daily fetch today and at most 1 to 2 GiB on a used chain; a solo miner runs a node that executes, which it already does for templates; a node operator holds the day's stream beside the 256 MiB cache (6 KB today, at most the sample cap) and pays one leaf pass per day; a light client gains nothing it can check without a full node (section 5); a prover and a rollup customer: nothing.

4. The known-failed case, first

Every gate here fires on a failed case before it is trusted (CLAUDE.md, 4 October 2026 rule).

Gate Known-failed case What it must say
igneum-pow state::tests::a_stateless_hasher_is_wrong_on_every_item the same seed, day and program, the dataset built with zero leaves, with the leaves of another root, and with the leaves of a stream one record short 0 of 32 lanes of the pinned warp agree with the stated hashes; 0 of 64 sampled items agree
igneum-pow packs test, the pinned v5 pack the v4 pack's vectors against the v5 program the program id differs; every vector differs; the kernel text is equal but for the build kernel and the class lines
Fork kaspa-consensus-core program_class_v5_rule 9,499 bps in one of the seven windows; byte 5 while program_class_v5_activation_daa is never; v5 signalled below v4 no flip; no flip (the byte counts as v4 only); v4's answer
Fork kaspa-pow with the feature a v5 epoch asked of an engine with no day-state provider PowEngineError::DayStateUnavailable, no build, the header not marked invalid
Fork igneum-exec day_stream::a_tampered_stream_is_refused a changed byte, a dropped record, storage before its account, a stale root check() errs on each
The fast-time harness (section 10) --signal 5,5,4 --expect flip FAIL with the eight flip checks named; and the node run with --evm-disable refuses every v5 template and mines no block after the flip while the chain carries on

5. The verifier on a node without the state (the trusted-data path)

A verifier derives up to 4,096 items per unit lazily (2.06 ms on an M5 Max core for class v3, 8.8 ms cold with the sibling loaded on the box's core 40 for class v4 at rung 0, the ladder page section 5) and under class v5 reads D[t mod n] for each: the measured +0.11 to 0.21 ms per unit with the array in RAM (new-pow 5.2 rows (i)). So the verifier needs D, which needs the stream, which needs the state at C_d. Three ways a node can have it:

Path Trust Cost per unit Who
Own execution none beyond its own code +0.11 to 0.21 ms every full node
A stream from a peer or an operator, checked the stream rebuilds to R_d (DayStream::check, 0.1 ms at today's state, seconds at a used chain's), and R_d is the state_root of C_d's record in the node's pinned exec snapshot (the pin the node already trusts) or in a verified segment record whose post_root covers C_d (proving v1) the same +0.11 to 0.21 ms a proof-synced node, a node whose executor is behind, a verifier box that runs no EVM
Openings against R_d per unit (a header-only verifier) 128 leaves per lane x 25 levels x 32 B = 102 KB per lane, 3.4 MiB per unit before dedup (new-pow 5.2) the openings' bytes and about 3,200 keccaks NOT offered: a light client keeps relying on certificates (spec 10) and new-pow's rank 6 RPC is the availability spot check, not PoW verification

A node that holds none of the three refuses the header with a retryable error (RuleError::PowCacheQueueFull's class: the node's own state, not the peer's fault, no strike in pow_guard), logs class v5 needs the day's state (day d, reference block C_d): this node cannot validate class v5 blocks until its executor passes the cut or a checked day stream is loaded, and serves no template (GetBlockTemplate errs the same line): nodes that cannot build the dataset refuse to mine. A node started with --evm-disable says so at start and is in this state for good under class v5.

IBD under class v5: headers of history need each day's state. The rule is the one the exec restart already set: a node executes from its snapshot forward, trusts the snapshot by its pin and the proofs, and validates PoW on the headers it has state for; headers below the snapshot's tip are the snapshot's business (the 0.3.14 exec_restart_state_root and exec_restart_trust_daa shape); a proof-synced node validates the proof's headers by the proof. The class-signal witness owed for v4 (counter-asic-3-node.md 6.3) covers v5's tally the same way; a day-state witness (the R_d per day beside the epoch seed in the proof format) is the matching spec item (new-pow 4.3, rank 1), owed.

6. The signal and the switch

The same machinery as class v4 and the ladder (counter-asic-3-node.md section 6; docs/design/latency-ladder.md section 3): one walk of the seven windows, one predicate per object.

Item Rule
The carrier the object version in bits 8 to 13 of the header version; CLASS_SIGNAL_V5 = 6 (the amended class v4 of the 0.3.20 node line is 5, the 6 October stream was 4); a byte at or above 6 signals v5 and, as before, the amended v4 (a later object contains the earlier); accepted by every node since 0.3.15 (only the low byte is checked), so no header rule changes
The switch Params::program_class_v5_activation_daa, u64::MAX never, the floor once set (rounded up to the epoch boundary as v3's and v4's); in the digest ONLY once set (the 0.3.15 rule), so a binary carrying the field peers with one that does not until a file sets it; --override-params-file field program_class_v5_activation_daa
Active class_v5_enabled() = program_class_v5_activation_daa() != never; with it never, a byte 5 on the chain counts as a v4 signal and nothing else moves
The window program_class_v4_signal_window_daa, shared: one window for every object's tally (86,400 on the devnet, 120 on the fast-time profile)
The decision for epoch e v5 when the floor says so; else when epoch e - 1 was v5 (monotone); else when epoch e - 1 was v4 and each of the seven consecutive windows ending at the seed block S_e is full and has at least 9,500 bps of its blue blocks at byte 5 or above; else the v4 rule's answer. One step per epoch: a chain goes v3, v4, v5 through three decisions at the earliest
This node's byte IGNEUM_CLASS_SIGNAL on devnet and simnet; the default is 6 when class_v5_enabled() and 5 (the amended v4's) otherwise, so a live file where v5 is never writes exactly what the 0.3.20 line writes
What a node reports PowEpochInfo fields 37 to 40: programClassV5ActivationDaa (optional, absent = never), programClassV5SignalBps, programClassV5SignalWeakestBps, programClassV5SignalEpoch; the daemon prints the floor and window lines for v5 beside v4's, and Program class v5 by miner signal: epoch E (...) once per flip
The miner takes the class from the template as before; a v5 pair needs the day stream, fetched from its node's exec RPC (igneum_getPowDayState [day], --exec-rpc on the miner, default the node's host at the exec port) when the engine's provider asks
The day-state provider kaspa_pow::igneum::install_day_state_provider(Arc<dyn DayStateProvider>): the daemon installs the executor's; the miner installs the RPC client's; an engine without one refuses v5 with DayStateUnavailable
The engine keys day caches on (day, dataset_class); v5's dataset class is v5 (its own DatasetSource carrying the leaves), so the flip costs one 256 MiB cache build per node and per miner at the boundary, prepared one lead ahead as the v4 flip was

7. Measured (igneum-build-1, 7 October 2026) and owed (the 4090 at 10:00 UK)

Row Value Where
The devnet's state (node 1's exec snapshot, tip 159,357) 93 records, 8,619 bytes; stream built in 0.1 ms; rebuilt and root-checked in 0.1 ms igneum-day-stream, 08:33 UK, core 40 at nice 19
Verifier per unit, class v4 against class v5 (one 32-lane warp, core 40 at 3.72 GHz, nice 19, the ladder's method; the box at load 70 from other lanes' builds, so every figure is a loaded-box figure: the ladder's rung-0 row read 5.14 alone and 8.77 loaded at load 25) v4: cold 8.11 ms alone, 8.74 ms with the sibling loaded; average of 50: 7.83 and 8.63. v5 over the devnet's 93 leaves: cold 8.71 alone, 9.03 loaded; average of 50: 8.54 and 8.83. So v5 adds 0.28 ms cold and 0.20 ms on average with the sibling loaded (+3.2 and +2.3 percent), inside the prototype's +0.11 to 0.21 ms row, and both classes stay under the 10 ms gate on the loaded core at load 70 tools/class-v5/verify-bench-remote.sh, 09:37 UK, raw lines under docs/design/class-v5-bench/20261007T083716Z/
One M5 Max core (the Mac under the measure lock, nice 19, the box rule's own order: class v5 against class v4 interleaved three times, 50 warps each, 19:31 UK, Mac load 8 to 9 from other agents) class v5 2.535, 2.649, 2.398 ms per warp (average of 50; cold 2.50, 2.99, 2.26); class v4 2.263, 2.224, 2.393 (cold 2.29, 2.25, 2.31): v5 adds 0.23 ms per warp on the average (+10 percent), the order's +0.2 ms; both classes at a quarter of the 10 ms gate on this core, and the 2019-class core by the 2.5x rule reads about 6.3 ms for v5 (O-1.14 still owed). Per tier: a node on any core from a 2019 laptop up keeps its margin; a miner's CPU engine (the harness's) pays the same 10 percent per verified warp and nothing per hash igneum-pow bench --program-class v5 --state <stream> --warps 50 on the Mac's own build (rustup 1.99.0; the sub-version 3 igneum-pow at 2f9bf135), raw lines kept in the lane's scratch
RTX 4090 (the fleet's pod cv5-2, driver 595.91.07, nvcc 12.8, 450 W limit, the card alone; proto-newpow/class-v5/bench.cu on the two pinned packs, 10:23 to 10:30 UK): hash rate, v4 against v5 v4 63.067, 63.066, 63.064, 63.065 MH/s; v5 63.058, 63.058, 63.052, 63.059 MH/s over 10 batches of 2^24 each (GPU event time): v5 is 0.01 percent under v4, equal within the batch noise; the hash kernel is the same text docs/design/class-v5-bench/pod-4090-20261007/run.log and run60.log
4090: watts the four 20-s windows read 269.4 (v4), 275.8 (v5), 275.9 (v4), 284.2 (v5) W and the four 60-s windows in the order v5, v4, v5, v4 read 284.5, 294.3, 295.8, 297.1 W: a monotone warm-up drift over the eight minutes of the run with the classes interleaved inside it (v4's last window is the highest), so the class moves nothing a window can read; the prototype's control-against-sd1 rows on 6 October read 205 to 208 W on another 4090 at the same 63.08 MH/s (this pod's card draws 4.3 to 4.7 microjoules per hash against that box's 3.27: another host, another power state, the same rate) the same logs
4090: dataset build with the devnet's leaves (93 x 64 B) v4 30.56 to 30.58 ms, v5 30.79 to 30.81 ms: +0.24 ms (+0.8 percent) for the cache-resident leaf read; the prototype's +1.43 ms was with a 1 GiB leaf array the same logs
4090: bit-exactness the v5 pack's dataset head, last word and 64 sampled words and its three vector warps (96 lanes) agree with the CPU derivation of the crate on the card, every pass; the v4 pack likewise; the fingerprints differ (3d2e8245cc084d07 against e438ca2fc66b5cd4) the same logs

Per tier, what the 4090 rows mean: every NVIDIA card from 8 to 32 GB keeps its hash rate and its watts under class v5 (the kernel is class v4's text; the only new work is one 64-byte read per item in the build, +0.24 ms per epoch at today's state, under 2 ms at the sample cap on this card by the prototype's 1 GiB row); device memory during the build gains the leaf buffer (6 KB today, at most 1 GiB resident or 64 MiB chunked at the cap) and nothing while hashing; AMD and Apple take the same build kernel in their dialects (emitted, unmeasured: owed with the worker hosts' leaf buffer); a rig's cards share one stream; a pool user takes a 6 KB fetch per epoch today; a solo miner's node holds the stream it already executes; a node operator pays 0.2 to 0.3 ms more per verified warp. What is done about the owed rows: the worker hosts' leaf buffer is the next item (section 11), the AMD and Apple rows ride on it.

8. Hostile review

Question Answer
A miner served stale state (yesterday's stream, a fork's stream, a tampered stream) every item it builds is wrong, every block it submits is rejected by every node that holds the day's state, and it learns it from the first rejection; a worker that takes the stream from its own node checks nothing (its node executed it), a worker or pool miner that takes it from elsewhere runs DayStream::check against the R_d its node reports (igneum_getPowDayState returns the root with the block)
A pool serving wrong state to its miners the pool loses every share it pays for; the pool's own node rejects its miners' solutions; there is no way to profit from a wrong dataset, only to waste the pool's hash
A state the verifier cannot reach (a header whose C_d the node never executed: a fork deeper than the lead) the engine refuses the header with the retryable error and says which block it lacks; a fork that deep is a merge-depth-scale reorg, which the follower re-executes when the chain adopts it ("a deep reorg never resets execution", 6 October rule), after which the header validates; a chain the node never adopts is a chain it never needs the state of
The executor is behind the cut when the day starts (a slow node, a node that restarted) it refuses to validate and to mine v5 headers until it passes cut(d) and says so once per epoch at warn, then at debug (the class signal's first_time shape); with the lead at one hour this is a node an hour behind the chain, which is already a node that serves no useful template
A node without an executor (--evm-disable) states at start that it cannot validate or mine class v5; after the flip it relays headers it cannot check as it relays blocks whose bodies it has not fetched, and accepts nothing it cannot validate (the same as a node without the day's cache slot today: BuildQueueFull)
The era draw's interaction the era draws the layout (t(w), j(w), the stride and the window of each load) and nothing about the leaves; the leaf XOR sits before the first mixer of item t, after the layout has named t, so every era's dataset of a day is built from one leaf array and the hash kernel of every era is unchanged; the era draws 8 and 9 stay consumed and unused as the ladder left them
The ladder's interaction a rung changes the shadow pass count of the base program and nothing in the derivation; a v5 program at rung r is v5_class_at(reps_r), generator 5, its own id
State grinding (a producer writes state to shape leaves) the leaf is keyed by R_d, which depends on every record, and enters a chained keyed derivation whose read addresses are mixer state; the producer of a block near the cut chooses at most whether its own block is C_d (one bit, between two honest states: a timestamp just under or just over the cut); no bias on 128 dependent reads is reachable below a block's cost (new-pow 4.3, the cryptographer's review)
Compressible or empty state the digests are full-entropy whatever the records; an empty chain has the registry (93 records today) and the dataset is keyed by those; graceful, as new-pow 3.3 says
The day boundary is not an epoch boundary on the devnet true today for the day key as well (the day is by timestamp, the epoch by DAA); the miner's prepare of (epoch, day + 1) one lead ahead is the existing path and v5 adds the leaves to it; under the spec's DAA day (1.12) day boundaries are epoch boundaries and nothing here changes
A finality pause no coupling: C_d is the selected-chain block, certified or not
A reorg across C_d one hour deep: a merge-depth-scale event; the executor unwinds and re-executes (the ring), the day's state is recaptured from the new chain's C_d, and the day's dataset changes; every miner prepares again; the epoch seed carries the same acceptance for a 600-DAA reorg
The serialisation splits the chain on a day boundary (the M20 and DAA 198,000 class) the stream round-trips to the root on every node (check is run by the capturing node too, so a node whose serialisation disagrees with its own state refuses to serve the day and says so rather than mining on it); the fast-time harness crosses a day boundary with a non-trivial state before Devnet 2; the Devnet 2 gate crosses a real day boundary
The chip unchanged against the f = 1 chip; the f = 0 chip must hold the state and becomes f = 1; stated, not claimed otherwise
The pool caveat a pool ships D once a day; "every miner holds the chain" is "every mining operation holds the state"; stated in new-pow 3.3 and here
A byte 5 on the live devnet before any file sets the switch counts as v4, moves nothing; the node's default byte stays 4 until a file enables v5, so nothing is even written

9. What moves the digest and the wire

Change Digest Wire
The binary carrying the field, nothing set no (the 0.3.15 rule) header version unchanged (the byte stays 4); PowEpochInfo gains fields 37 to 40 with defaults; igneum_getPowDayState on the exec RPC; an old miner ignores them
The file setting program_class_v5_activation_daa yes, once, at publish (bundled into one cut, miners first, hands last) the byte 5 may appear; every 0.3.15+ node accepts it
The flip no the program id of every epoch from the flip (program_id(5, ...)), the pack's IGNEUM_PROGRAM_CLASS "v5", leaves.bin in the pack, the job line's class=v5 token (the three hosts strip class= and era= from its tail, so the line's shape is unchanged)

10. The fast-time gate (the class v4 rehearsal's shape)

infra/fast-time/class-v5-signal.mjs: three nodes and three CPU miners on override-60x.json, v3 from DAA 60 and v4 from DAA 120 (the floors), the v5 floor at --floor (default never), the window 120 DAA, each node's byte by IGNEUM_CLASS_SIGNAL (--signal a,b,c), each node's exec RPC on its own loopback port and each miner's --exec-rpc pointing at its node; a fourth node with --evm-disable and its own miner joins with --stateless. The three cases and the known-failed case, as the v4 gate's: --signal 5,5,4 --expect no-flip; --signal 5,5,5 --expect flip (the flip at the first epoch whose seed block has seven full windows below it, every miner's v5 id equal to the CLI's --program-class v5 id and unequal to the same seed's v4 id; the stateless node's miner accepted 0 blocks after the flip; the chain's reject count 0); --signal 4,4,4 --floor <daa> --expect floor; and the known-failed case --signal 5,5,4 --expect flip, which must FAIL. The day boundary: the profile's day is 24 minutes and the cut one minute before it; a run of 10 epochs (10 minutes) crosses a boundary when started inside the right 14 minutes, and the harness reports whether it did (day_boundary_crossed, the day index of every epoch and the stream root each node served for each day; the roots must agree across nodes). Results (igneum-build-1, 7 October 2026; tools/class-v5/harness-remote.sh, summaries and logs under docs/design/class-v5-harness/):

Case Run (UK) SUMMARY
The known-failed case first (--signal 5,5,4 --expect flip) 09:37 to 09:48 FAIL as it must: no flip at 7,519 bps of byte 5 (two signalling nodes of three plus the stateless node), 665 blocks, 0 rejected, four sinks equal, eleven flip checks named FAILED CHECK, exit 1 (failed-case.json)
The flip with the stale miner and the stateless node (--signal 5,5,5 --expect flip --stale 2) 12:00 to 12:11 (the sixth run; the five before it found the harness's own faults and two of the fork's, section 0). CORRECTION 20:3x UK (the v5-fasttime lane): the two "accepted 0 after" counts of this row and the 13:21 row were parse misses (the harness read the miner's line clock as HH:MM:SS where it is epoch seconds), so they proved nothing; the stale miner's 66 rejections by its own node and the stateless node's chain stopped at 480 with 14 refusal lines are real and are what those two checks rest on until the re-run with the fixed parser (section 0) PASS on every check: Program class v5 by miner signal: epoch 8 (... weakest 10000 bps, 420 of 420 blue blocks) on 3 of 3 nodes, the first epoch with seven full windows after v4; the template class v5 from epoch 8 (DAA 480) at 486 s wall; 481 / 185 blocks across the boundary; the three signalling nodes at one sink, 665/665/665; the miners' v5 program ids equal to the CLI's --program-class v5 --state id on epochs 9, 10 and 11 and unequal to the same seeds' v4 ids; a different state root per v5 epoch (828e6d26..., 2ae82ead..., 3a94e581...: the refresh per window), the same root on every executing node; the stale miner (its first stream kept for every later epoch) accepted 0 after the first refresh and rejected on 66 of 66 by its own node; the stateless node (--evm-disable) accepted 0 after the flip, 14 refusal lines, its chain stopped at 480 (flip-stale.json)
Two of three signal v5 (--signal 5,5,4 --expect no-flip --epochs 11) 12:12 to 12:23 (the first start refused at once: port 30140 was a stranger's, the new bind check fired) PASS: no v5 epoch over epochs 0 to 11, the byte-5 share at the sink 6,274 to 9,000 bps epoch by epoch (two signalling nodes of three plus the stateless node at byte 5, the third at byte 4), on the chain 505 blocks at byte 5 and 158 at byte 4 (7,605 bps), no signal line on any node, 664 blocks, 0 rejected, four sinks equal at 663 (no-flip.json)

11. What is implemented, tested, owed

Implemented on class-v5 and class-v5-node (this page's time line in section 0 names each commit): the design; igneum_pow::blake2b, igneum_pow::state (leaves, sample, the leaf XOR in memhard::derive_items), V5_CLASS, ProgramClass::V5, generator 5, the emitters' build kernel with the leaf buffer, the CLI's --program-class v5 --state <stream>; the fork's DayStream, the day-state capture in the executor, the provider in kaspa_pow, program_class_v5_activation_daa, the v5 signal rule and tally, the RPC fields, the daemon lines, the miner's fetch; the tests with the known-failed cases first; the pinned v5 packs; the harness.

Owed, named here so nobody looks for them: an acceptance-rule bound on a program's hot-set share (the Counter ASIC lane's item for v5, 09:5x UK, on main's word: the attack-pass lane's F8 found that on a class v4 program the top 0.1 percent of items take 0.52 percent of reads, 4.05x a uniform map, one item 153x the mean, most of it class v4's designed per-site windows of layer 8; the v5 rule rejects a draw whose windows and offsets coincide into a hot set beyond the window model's tail, with the rejection's cost stated in accepted programs per 64 seeds under the 2.0 bounds; the bound's definition and number come from the hash lane's window-model analysis on branch ca3-v4-uniform and the gate is F8's 64-seed census, tools/attack/f8-uniform on branch attack-pass phase E, the top 0.1 percent within the bound on every seed, re-gated by the attack-pass lane; nothing of it touches v4, which is on the vote; the bound landed 10:3x UK from docs/analysis/ca3-v4-uniform.md 095f84a7: the fault is a lossy-sourced load, a load whose source register was last written by an entropy-losing op (or, mul, mulhi), whose saturated value recurs at (3/4)^32 per read and lands on one item through the era map; 96.6 percent of today's class v4 programs carry one; the bound H = W_0.1 from the 16 window draws (0.115 to 0.251 percent) plus the sum over load sites of h(last writer), h = 0.30 percent for or, 4.5 for an or chain, 0.067 mul, 0.049 mulhi, 0 for an injecting op or a rotate, with H at or under 1.2 x W_0.1, equal to the static rule "every load's source was last written by an injecting op or a rotate"; as a rejection it would cost 96.6 percent of candidates, so the form for v5 is a generator draw: a load's source drawn from the registers whose last writer injects, no attempts lost, rule (a) unchanged; its worth to a chip today, at most 1.005x on about half the hours, 1.048x on 5 percent, 1.067x at the ceiling; the project lead ruled option A at 15:2x UK by the coordinator's relay: the draw lands in class v4 itself (branch ca3-v4-amend, the hash lane, 0.3.19, a new program stream and vectors, the amended class with its own generator stamp), and class v5 inherited it by merging that branch into class-v5 (4e737543, 13:0x UK). NOT satisfied yet (the coordinator's correction, 13:2x UK): the attack-pass lane's F8 re-gate on sub-version 1 has nine of the first 30 seeds over 1.2x (p31 at 29.3x) and 0.19 percent hot-set programs, because saturated values pass through saturation-preserving writers; class v5 draws under the amended v4 by merge and so inherits the gap, and takes the fix the same way when the Counter ASIC lane lands sub-version 2 inside generator 4 (byte 6 stays v5's); gate as above, open); beside it, the mixer-draw acceptance rule of the attack-pass lane's AP-F4-1 (docs/analysis/attack-pass/f4-weakday.md sections 6.7, 7 and 8 on branch attack-pass, relayed 10:0x UK): the day key's MUL block has a tail of cheap multipliers (low NAF weight), a tail of a sum with no weak class behind it; against the DSP-bound datapath F4 passes (0 of 2^28 days over 1.1x), on the LUT-adder metric 5,476 of 2^24 days are over 1.1x (the worst 1.173x; the worst calendar day in the first 100 years is chain day 29,337 at 1.121x), worth at most 12.1 percent more hash rate that day to a per-day LUT-recompute FPGA and nothing to a stored-dataset FPGA or any chip; the v5 rule for the day's mixer draw: reject a MUL block with NAF sum under 163 (M1 cost under 211) and redraw from the next stream values, NAF weight at least 4 per word, at least 4 distinct ROT amounts, total rejection 6.1e-4 per day, the first calendar redraw at day 22,633 (5.2 years in), no devnet or testnet pack changes; gate: F4's harness (the 2^24-day census under both metrics, the worst day's M1 cost at or above 211 after the rule), re-gated by the attack-pass lane against this branch; it is the mixer-draw side of memhard.rs (MixParams::with_shape), independent of the leaves, and lands behind the v5 class; third, the attack-pass lane's AP-F1-1 shadow-block rule (docs/analysis/attack-pass/f1-shadow.md, harness tools/attack/f1-shadow, relayed 10:5x UK): over 100,000 class v4 programs the shadow block's peephole-removable instructions (a register written twice from one source with no write between: xor-cancel, sum-cancel, rotate merges, or-idempotence) average 0.62 percent of the 256 per pass, at most 5.078 percent (13 of 256, one program in 100,000 over 5 percent), nothing crossing a pass; clang -O3 removes the same from the honest kernel, so it is a bound on the shadow's useful work, not a chip shortcut; the v5 rule: the generator refuses a shadow block whose honest-compiler simplification exceeds a fraction and redraws from the next stream values; the fraction, set here from the census histogram (0.5-percent bins from 0: 55,595; 20,442; 11,790; 9,729; 1,447; 613; 256; 103; 17; 7; 1; 0): 3.0 percent, which rejects 384 of 100,000 draws (the bins from 3.0 up: 256 + 103 + 17 + 7 + 1, about 4e-3, under one attempt lost per 250 seeds, so the per-program spread and the acceptance rate of the 2.0 rule stay inside their bounds) and caps the removable share at 3 percent of the block where 5 percent would cap it at 5 for 1e-5 of draws; gate: F1's harness on 64 seeds of the v5 stream, every block under 3 percent, re-gated by the attack-pass lane; it is the shadow side of the v4 draw that v5 inherits from ca3-v4-amend and lands in generator.rs beside the shadow draw, behind the v5 class; the exec snapshot wire carrying the day streams (version 2); the GPU worker hosts' leaf buffer (proto-cuda/nvrtc, proto-opencl, proto-metal: the kernel text carries it, the hosts must upload it); the pool protocol's daily fetch (spec 09); the day-state witness in the pruning-proof format (with the class-signal witness); the spec text (01 1.8.5 the leaf line, 1.12 the cut and the reference block, 10 the witness); the 2019-class core row (O-1.14); Devnet 2 across a real day boundary.

12. The litepaper paragraph and the ledger row

Litepaper, Mining, after "The work that waits can grow": "The dataset is the chain. Class v5 builds each day's dataset from the chain's own execution state at a block one hour before the day: every account, every storage slot and every byte of contract code, hashed under the day's state root and folded into every item. A card that does not hold the state cannot build the dataset, and a card that builds it wrong is wrong on every hash. The hash itself does not change, and neither does its speed: measured on an RTX 4090 on 6 October 2026, 63.08 against 63.09 million hashes a second at 207 W, the daily build 1.4 ms longer, a node's check 0.1 to 0.2 ms longer per block. What it buys is a floor under what mining means: a chip that recomputes items instead of storing them must now hold the state too, and a pool miner who today needs nothing but the date needs the state. What it does not buy is stated as plainly: a pool can ship the state to its miners once a day, so the claim is that every mining operation holds the chain, not every card; and against a chip that stores the whole dataset it changes nothing, which is why it sits beside the latency-shadow work, not in its place. It ships switched off, behind the 95 percent class signal with a floor height."

Ledger row M35 (in docs/fud-ledger.md): the scheme, the measured cost, the pool caveat, the stateless known-failed case, the owed items.

13. The design-to-code gap (for the coordinator, 7 October 2026, 19:0x UK)

Every item of this page that has no code, no test or no measurement yet, with the owner and the hours (agent hours; the project lead's rule). "Done" rows are listed first so the gap reads against them.

Item State Owner Hours
The generator and verifier (igneum-pow): V5_CLASS, generator 5, blake2b, the state leaves and the leaf XOR, the emitters' leaf buffer, --state, the pinned packs, the known-failed tests done; on the frozen sub-version 3 base (ca3-v4-amend 017e7037 merged 18:5x UK, the source rule keyed with the state flag aside, 67 unit and 39 integration tests on build-2 before the merge, the suite re-running on it now) this lane 0
The +0.2 ms per warp cost against the 10 ms gate measured on igneum-build-1's reference core (section 7: +0.28 cold, +0.20 average, both classes under 10 ms loaded); NOT yet on one M5 Max core as the order asks (the Mac rule: one Metal or macOS run at a time under the lock) this lane 1
The node side: the state commitment into the derivation, the per-epoch capture, the lock-free stream cache, the provider, the refusal, igneum_getPowStateLeaves, the signal rule and the floor, the RPC fields, the miner's fetch done on the fork; rebased onto release-0.3.23-node at 19:0x UK (v5 pinned to object byte 6, counted exactly; byte 7 is v4 sub-version 3); the suites on build-2 owed on the rebased fork this lane 1
The stateless and stale-chip rule as a test, known-failed first done: a_stateless_hasher_is_wrong_on_every_item (igneum-pow), class_v5_refuses_without_state_and_refreshes_the_leaves_per_epoch (kaspa-pow), the harness's stale miner (66 of 66 rejected) and stateless node this lane 0
The digest-compat rule for the new field done: program_class_v5_activation_daa enters the digest only once set (the 0.3.15 rule; the params test covers the override file); the 60x file carries it at never this lane 0
The hot-set cache rule (AP-F8-1, 1.067x bound) inherited by merge from sub-version 3 (the dataflow freshness rule, the shared-operand rule, the 256 cap and the last resort); the bound itself is the attack-pass lane's F8 gate on the v5 stream attack-pass lane (a3832b1c3b274b310) for the gate; this lane for the v5 pack export it reads 2
The weak-day FPGA rule (AP-F4-1, NAF sum at least 163, at least 4 distinct rotations) NO code, NO test: recorded in section 11 with the rule text; lands in memhard.rs MixParams::with_shape behind the v5 class (the mixer draw redrawn from the next stream values), with the known-failed case (the worst calendar day 29,337 at 1.121x redrawn) and the per-day rejection count (6.1e-4) this lane 3
The shadow-redundancy rule (AP-F1-1, 3.0 percent) NO code, NO test: recorded in section 11 with the fraction; lands in generator.rs beside the shadow draw behind the v5 class (a block whose peephole-removable share exceeds 3.0 percent redrawn), with the known-failed case (the census's 5.078 percent block redrawn) this lane 3
The attack-pass families on v5 (F1, F4, F8's 64-seed gate at 2^24 with the 1.2x line, F9's exhaustion count) NOT run on v5; the harnesses exist on branch attack-pass for v4 attack-pass lane; this lane hands it the v5 stream (--program-class v5 --state) and the pinned pack 4 (its)
The kit for every platform: Metal the kernel text is emitted (the leaf buffer on igneum_build); the Metal pack bench takes leaves.bin as of 19:0x UK (this commit); the one-click worker's Metal host (proto-metal/main.swift) does NOT yet upload the leaves this lane (packbench), the worker lane for main.swift 2
The kit: CUDA the kernel text emitted and measured on the 4090 through proto-newpow/class-v5/bench.cu (section 7); the one-click NVRTC host (proto-cuda/nvrtc) does NOT yet upload the leaves worker lane 2
The kit: OpenCL (AMD), Intel the kernel text emitted; no host, no run worker lane (host), fleet or PC 1 (the 9070 XT run) 3
Fingerprints equal across platforms, G1 on the 5090 and the Mac, AMD and Intel after the 4090 pack fingerprints are in (section 7); the Mac's Metal fingerprint is the clock reading below; the 5090 (PC 2) and AMD rows owed this lane (Mac), the fleet and PC lanes (5090, 9070 XT) 2
The crossing on Devnet 3 by height after its gate NOT started: the floor in the Devnet 3 override, the Devnet 2 style gate (zero rejected across the flip, exec roots agreeing, every node holding the epoch's stream); the harness's flip case is the rehearsal (PASS twice) node lane (a283f5f0d364ceef0) owns the fork, the shipper (ae892a8b0f78fe31c) the cut; this lane the gate's v5 checks 3
Miner cost rows per card (5090, M5 Max, 9070 XT, 4070) the 4090 row only (rate equal within 0.01 percent, build +0.24 ms); the four named cards owed, each labelled measured with the date fleet (5090, 4070), this lane (M5 Max under the lock), PC 1 (9070 XT) 3
The exec snapshot wire carrying the day streams (a proof-synced node's trusted-data path) DONE by the node lane (19:3x UK): fork branch class-v5-node-wire 7737ebd9 on both box mirrors, igneum/exec only, SNAPSHOT_VERSION kept at 1 with epoch_streams as the last field and a legacy fallback (a 0.3.23 snapshot loads on a pre-field node and the other way round through the sweep); the writer carries the two newest captures, each checked to rebuild to its root; the loader rebuilds each, compares its root with the loaded state's record at that block, publishes the final one to ExecState.stream_cache at once and holds the next until its cut, refuses a tampered or foreign stream and installs the rest; tests snapshot::tests::the_legacy_shape_is_the_struct_without_the_streams and service::class_v5_tests::a_snapshot_carries_the_epoch_streams_and_a_loader_serves_them, known-failed first; exec suite 43 of 43 on build-2 at 18:33:09Z node lane 0
The pool protocol's per-epoch state fetch NO code pool lane 2
The day-state witness in the pruning-proof format DONE by the node lane (19:5x UK): class-v5-node-wire 6d827c5d on both mirrors. The witness is the state root after each class v5 epoch's seed block, EpochSeedHeader.stateRoot = 3 beside the epoch seed header (empty before class v5, so the old wire parses both ways), PruningProofStateRoots on the proof metadata; the prover reads each class v5 seed block's root from the installed provider (DayStateProvider::state_root(block), one addition to the kaspa-pow trait, default decodes the stream) or from a witness it took; the verifier's ProofSeeds::from_witnesses_with_roots refuses a class v5 epoch whose seed header has no root (the known-failed shape) and a root for an epoch with no header; check_header on DayStateUnavailable for a class v5 header accepts it as trusted data when the proof carries the epoch's witness, else refuses naming the missing witness; the importer installs the (seed block, root) pairs through install_day_state_witnesses; the executor's install_epoch_streams holds a stream for a block it has no record of to that witness (section 5's rule, both conditions). Suites on build-2 on the commit: kaspa-consensus 122, igneum-exec 44, kaspa-p2p-flows 37, kaspa-p2p-lib 20, kaspa-pow 18 node lane 0
The acceptance bound on a program's hot-set share (section 14, main's order 19:5x UK) (c''') at ab6f980b: the per-site distinct-index floor raised to 0.995 under the state flag, known-failed first on seed 100767; the census (clean rejection rate, attempts histogram) running on box 2; the number goes to the coordinator and here when it lands this lane 1
The class walk under the v5 object (class_signal.rs decide(): epochs under the window resolve to v3 regardless of the v4 floor once v5 is enabled; the v5-fasttime lane's reading 20:0x UK) OPEN, the node lane's: a known-failed test with v4 from genesis and the v5 object set must read v4 at epoch 0; this lane's harness gains the same case (v4 floor 0, v3 never) once the fork has the fix node lane (fork), this lane (harness case) 2
The spec text (01 1.8.5 the leaf line, 1.12 the cut, 10 the witness) NO text this lane 2
The litepaper paragraph and ledger M35 done (the litepaper's measured numbers are the 4090's) this lane 0

Sum of the gap: about 40 agent hours, 17 of them this lane's (the two draw rules, the M5 Max rows, the spec text, the gate's v5 checks), the rest the node, worker, pool, fleet and attack-pass lanes'.

14. The acceptance rule's class v5 part (c'''): the residual hot-set class (main's order, 7 October 2026, 19:5x UK)

The class the in-house attack pass attributed tonight (adv-accept, docs/analysis/cryptanalysis/report-acceptance-rule.md on branch adv-accept): a value-level constant from a lineage-fresh writer that neither (a') nor (c'') reaches. The exemplar is class v4 sub-version 3's seed 100767 of the f8 label space (epoch bytes the words of igneum-attack-f8/program/100767, era bytes of igneum-attack-f8/era/100767, attempt 2, program id 9d68e6286fc817d4, class mx8-era763e5847+sh256x27). It passes every part of the sub-version 3 rule: its load site 6 (instruction 23, source r6, a quarter window) reads a (c'') ratio of 0.9919 on the closed form and 0.9920 on the live day at the rule's own 2^20 sample, 0.012 above the 0.98 floor; live at 2^24 sequential nonces the site sends 3.35 percent of its reads to the top 0.1 percent of items (the multiples of 2^19: rotl(x * stride, rot) of a mad value that is near zero in a value class; r6 is zero in 0.0126 percent of evaluations, 1,677 reads of index 0 in 10^6 nonces), the top 0.1 percent of items taking 2.05x the window model's share, where the chip model's gate is 1.2x.

Two fixes were on the table, with the rule that the one with the lower rejection rate on clean seeds wins if it reaches the class: (i) a per-site hot-item test at live scale (the top items' share per site against the window model, which needs a sample far above 2^20: at 2^20 evaluations the hot set is 272,544 word indices in 17,034 items, 0.13 reads per index, so no index-multiplicity statistic reads it; the item-share statistic at 2^20 reads 2.05x only with the 2^24 run's resolution); (ii) the per-site distinct-index floor of (c'') raised from 0.98 to the bottom of the clean seeds' spread. The distinct count is the same statistic as the collision excess the hot set produces: 3.35 percent of a site's reads landing on 0.1 percent of its window's items is about 0.8 percent of its evaluations repeating an index at 2^20, which is exactly the 0.992 the ratio reads, while the model's spread at that sample is 0.0001 (the collision count is near Poisson with mean n^2 / 2W = 8,192 on the quarter window, so a clean site reads 0.9997 to 1.0001; adv-accept's census of sub-version 3 clean seeds reads 0.9960 at the minimum site, 0.9990 at p1, 1.0000 at the median).

The rule taken: MIN_DISTINCT_RATIO_V5 = 0.995 (igneum-pow accept.rs, ab6f980b), applied by check_indices_v5 on the ratio pass's own run when the candidate's class has the state flag: site_index_stats keeps every site's sorted word indices once (the distinct count, the colliding pairs, the most read index with its count), distinct_ratio_on names a site under 0.98 as (c'') first, then hot_item_pass names a site under 0.995 as (c''') Reject::HotItemSite with the ratio and the most read index. No class v4 verdict moves (the key is the state flag), no new sample is drawn, and the per-candidate cost is the sort (c'') already paid plus one linear pass. The exemplar is refused at attempt 2 under class v5; the class v5 draw moves to the next attempt; the class v4 draw still lands on it. The known-failed test class_v5_hot_set_rule_known_failed_seed_100767 holds all of that and the genesis draws of both classes clearing the floor. The census test v5_hot_census (ignored; run by hand on a box) reads over 4,600 seeds of the f8 label space the class v4 accepted programs' minimum-site spread, the count under floors 0.98, 0.99, 0.995, 0.998 and 0.999, the attempts histogram of the class v4 and class v5 draws and the seeds whose v5 draw lands on another attempt: the clean rejection rate and the retried-draw cost of the number. Its reading goes in this section when it lands.

What the floor does not reach, named: a hot set whose excess collisions at 2^20 stay under 0.5 percent of a site's evaluations (a top-0.1-percent item share under about 1.3x); adv-accept's lowest-ratio list between 0.995 and 0.998 reads 1.29x to 1.63x live on the f8 gate but no hot set by X_f >= f, and the floor rejects those too (cheap: one more attempt). A chip that holds 0.1 percent of a window's items serves at most the excess the floor allows, 0.5 percent of a site's reads, 1/16 of that over the program: under the 1.067x bound of AP-F8-1 by a wide margin. The per-site hot-item test at live scale stays owed only if the census reads the floor's clean rejection rate over 1 percent.