diff --git a/docs/bench-log.md b/docs/bench-log.md index 713986c01..8b718d35e 100644 --- a/docs/bench-log.md +++ b/docs/bench-log.md @@ -1147,3 +1147,45 @@ IGNEUMD= IGNEUM_FAST_TIME=1 IGNEUM_HARNESS_BASE_PORT=<29500|29600> IGNEU | Slope from 1,000 blocks to the end | 30.7 MB per 1,000 blocks | 30.2 MB per 1,000 blocks | Reading. Before: one cache build per epoch roll (25 rolls in 1,529 blocks, plus the two days) and the RSS steps with them up to five chunks: `KEEP = 4` plus one evicted 256 MiB chunk the allocator keeps and never returns to the OS (vmmap at 500, 1,000 and 1,500 blocks: 1.3 G resident in the region vmmap labels IOAccelerator, which held exactly the cache chunks; the malloc zones hold 9 to 22 MB). It stays at five: the before node is flat at 1,34x to 1,37x from 514 blocks on. After: one cache for the first day (319 MB at 510 blocks), a second when the fast-time day rolled at 01:36 UTC (both runs; the day cache is per day by design, `KEEP_DAYS = 3`), 2 of 2 builds in 1,526 blocks against 27 of 27 before; on the devnet profile (24-hour day, 1-hour epochs) that is 256 MiB flat, 512 MiB around midnight UTC, 768 MiB worst case, against 256 MiB per hour up to 1.3 GB before. So per 1,000 blocks: before, 1,300 MB in the first 500 blocks (the four caches plus the kept chunk) then 31 MB; after, 30 MB, plus 256 MiB once per day. The 30 MB per 1,000 blocks is the same on both builds and is not the PoW cache: at 1,526 blocks the after node's non-cache footprint is 584.0 M physical minus 2 x 256 MiB = 72 MB against 23 MB at 0 blocks, and the malloc zones account for 22 MB of it, so most of it is in large `vm_allocate` regions, which on this node means the consensus database's write buffers and block cache and the consensus in-memory caches filling toward their fixed sizes (rusty-kaspa sizes them in entries for mainnet), not the execution layer (1,526 chain-block records on an empty chain are about 2 to 3 MB, approximate) and not the finality key maps (1 voter here). That is a reading, not a measurement: it needs a longer run to see the plateau (the live node's own figures fit it: 1,081 to 1,193 MB over 52 minutes with no epoch roll is 36 MB per 1,000 blocks). The live app node's 2,258 MB at 4 h 14 min is more than the before build can hold on its own (five chunks plus 30 MB per 1,000 blocks gives about 1.7 GB at 15,000 blocks); the app runs a GPU worker beside the node (Metal buffers and a kernel per epoch), which this harness does not cover, so that node wants its own `vmmap -summary`. Result JSON `docs/benchmarks/memory-floods-2026-10-04/{before,after}-s8-steady.json`, vmmap summaries under `.../vmmap/`. + +## 4 October 2026 (night), round-4 consensus items F23, F24, G12, X18 and M31: unit tests and fast-time 3-node runs against the control (consensus engineer) + +Branches: node fork `fud-consensus` (worktree `vendor/igneum-node-fud`, from `finality-fixes` 6aa69a45, no remote; commits 9f738e2e the four items, ae9df8a3 pending certificates at fresh determinations, b755f43d merge of `fud-memory`, then M31 and the un-determination rule), main repo `fud-consensus`. Machine shared with the red team, the release builds and the memory runs all night (load average 100 to 146 until about 01:00 BST, under 20 after); every number here is a count, a lock or an index, not a timing. Runner `tools/finality-attacks/fud.mjs` (scenarios `digest`, `ban`, `reorg`; fast-time 60x profile with `finality_v3_activation_daa` 0 merged by the BigInt-safe `overrideParams`, ports 29400+, suffix 940, `/tmp/igneum-fin-fud`, node logs kept per scenario); the control is the shipping finality-fixes build `vendor/igneum-node/target-finality/release/igneumd` driven by the same miner. Raw result files in `docs/benchmarks/round4-consensus-2026-10-04/`. + +**Builds.** `with-lock.sh build nice -n 19 cargo build --release -j 4 -p kaspad -p igneum-miner --features kaspad/igneum-pow` on a target directory cloned from `target-finality` with `cp -Rc` (APFS clonefile, 1 min, no disk): 17 min 53 s the first time under load 140, 8 min 58 s the second. Unit tests in the release profile (`cargo test --release -j 4 -p --lib`): consensus-core `config::params::tests` and `igneum` 20 of 20 (new: `consensus_digest_covers_every_consensus_field_and_nothing_else`, `env_pow_schedule_is_devnet_and_simnet_only`), consensus `processes::finality` 6 of 6 (new: `ban_is_decided_by_the_carrying_block_so_nodes_agree_on_every_voter_list` with three `TestConsensus` nodes on one chain, `reorg_past_an_unlocked_checkpoint_re_determines_it_and_verifies_the_pending_certificate`, `a_locked_checkpoint_pins_the_chain_and_a_certificate_against_it_conflicts`). After M31 and the un-determination rule: FINAL_TESTS. + +**X18, the params digest** (`fud.mjs digest`: n1 listens on the shared override, n0 dials it with `finality.weight_window` 121 instead of 120, then n2 dials with the shared override). + +| Build | n0's digest | n1's digest | mismatch lines n0 / n1 / n2 | peers on n1 after 25 s, then after n2 dialled | n2 connected | +|---|---|---|---|---|---| +| fud-consensus (9f738e2e and the final pass) | 4bf763ba... | 7a40cc3b... | 1 to 2 / 2 / 0 | 0, then 1 | after 1 s | +| control, finality-fixes 6aa69a45 | none printed | none printed | 0 / 0 / 0 | 1, then 2 | after 1 s | + +The listener's line: `Refusing peer 127.0.0.1:...: consensus params digest mismatch, local 7a40cc3b... remote 4bf763ba... (the peer's override file, environment or build differs)`; the dialler sees the reject message with both digests. Two lines on the listener per pass because the dialler redials once within 25 s. The control connects the mismatched node and says nothing. + +**F23, the ban decided by the carrier** (`fud.mjs ban`: six voters at 1/6 of 1 block/s, two per node; `a0` on n0 equivocates once at index 9 (`vmine --equivocate-at 9`, the second vote reaches n0 over RPC only); P2 cut when n0's next index reaches 9 (252 to 259 s) and healed 45 s later, so n2 learns the evidence from the carrier block after the heal; 480 s). + +| Build, pass | EQUIVOCATION lines n0 / n1 / n2 (carried by block) | refused "names N voters" | CONFLICTING | indices with 5 voters, per node | voter counts agree / differ (indices with lines on 2+ nodes) | disagreeing locked indices | max locked | +|---|---|---|---|---|---|---|---| +| fud-consensus, first pass (9f738e2e) | 2 (1) / 1 (1) / 1 (1) | 0 / 0 / 0 | 0 / 0 / 0 | 10..13 on all three | 11 / 0 | 0 | 14 / 14 / 14 | +| fud-consensus, final pass (ae9df8a3) | 2 (1) / 1 (1) / 1 (1) | 0 / 0 / 0 | 0 / 0 / 0 | 10..12 on all three | 10 / 0 | 0 | 15 / 15 / 15 | +| control, finality-fixes | 2 (0) / 1 (0) / 1 (0) | 2 / 0 / 0 | 0 / 0 / 0 | 10..13 on all three | 9 / 0 | 0 | 14 / 14 / 14 | + +On the new build n2's one EQUIVOCATION line is the "carried by block" variety (it never saw the vote), and the stripped range is the same on the node that detected over RPC, the node that saw the carrier at once and the node that saw it 45 s late. The control refused two of the other nodes' certificates on n0 with "names N voters, this node counts M"; the red team's stock s1 (two keys equivocating at every index) gave 9 / 3 / 4 refusals on the same build. Locks still agreed on the control because each node could build its own certificate from the votes it held; the refusal is the defect, the disagreement would follow on a network where one node depends on another's certificate. + +**F24, re-determination after a deep reorg** (`fud.mjs reorg`: n0 holds `q0`, `q1` at 0.15 each, n1 and n2 hold `p0` to `p3` at 0.175 each, so the n1/n2 side has 70% of the weight; 230 s warm, P0 cut 180 s, healed, 150 s heal window). + +| Build, pass | n0 determined on its own chain during the split | re-determined lines on n0 | pending kept / verified on n0 | CONFLICTING | refused "is for X, this node's checkpoint is Y" (pre-F24 wording) | indices the majority locked that n0 did not | disagreeing locked indices | max locked at the end | +|---|---|---|---|---|---|---|---|---| +| fud-consensus, first pass (9f738e2e) | 1 (index 8) | 2 | 4 / not re-read at fresh determinations (the gap fixed in ae9df8a3) | 0 / 0 / 0 | 0 | 8 and 9 (no certificate was verified at them) | 0 | 15 / 15 / 15 | +| fud-consensus, final pass (ae9df8a3) | 2 (8, 9) | 2 | 2 / 2 | 0 / 0 / 0 | 0 | none (the majority locked 10 and 11 during the split, not 8 and 9: 70% nominal is Poisson noise away from the floor) | 0 | 16 / 16 / 16 | +| control, finality-fixes, two passes | 2 then 1 | 0 | 0 / 0 | 0 / 0 / 0 | 3 (second pass) | 8 and 9 (second pass): the permanent hole | 0 | 17 then 15 | +| fud-consensus after the un-determination rule | RERUN_REORG_ROW | + +What the final pass found: n0's index 9 was re-determined at a sink of blue score 263 to a block of blue score 263, below the index's target 270, because the majority chain became the sink by blue work (its difficulty drifted less than n0's during the split) before it had reached index 9's depth; a record never names a block below its target, so the rule now un-determines an index the new chain has not reached and determines it again when it has (fork commit after ae9df8a3, unit test `a_shallower_sink_un_determines_the_indices_it_cannot_reach`). The control's first pass logged no CONFLICTING because that build's refusal used other words ("certificate at index 8 is for X, this node's checkpoint is Y"), counted in the second pass. + +**The red team's own reproductions** (`tools/finality-attacks/redteam/rtfin.mjs`, fin-attacks miner at 6 blocks/s, run on the fud-consensus build ae9df8a3 through the main worktree's env-aware harness lib on ports 29550+): RT_RESULTS. + +**G12** is covered by the digest run (the environment's schedule is part of the digest, so a devnet node with `IGNEUM_POW_EPOCH_BLOCKS` set cannot connect to one without it) and by `env_pow_schedule_is_devnet_and_simnet_only`; no mainnet node was started tonight. **M31** is covered by `largest_coinbase_fits_on_every_network`; no simnet network was started tonight (the red team's `tools/exec-attacks/net.sh` is the run that would show templates on simnet without an override). + +**Uncertain.** (1) Every run is fast time (W = 120 DAA, ban 120, depth 20) on three nodes with 100-ms links; the mainnet values are 30 days, 30 days and 60 blocks. (2) The ban run shows one equivocation at one index; the red team's s1 (equivocation at every index, two keys) was not re-run on the new build tonight except through RT_RESULTS_NOTE. (3) The digest-less allowance on devnet and simnet is deliberate for the rollout and is a hole until removed. (4) The reorg run's final pass had the majority lock no index during the first 60 s of the split, so the re-determination at 8 and 9 was exercised, the pending-certificate path only at 10 and 11; the first pass exercised the opposite. (5) The un-determination rule has a unit test and RERUN_REORG_NOTE. diff --git a/docs/fud-ledger.md b/docs/fud-ledger.md index 15357f37b..18f8a6f89 100644 --- a/docs/fud-ledger.md +++ b/docs/fud-ledger.md @@ -1649,9 +1649,11 @@ Evidence: the files above. ### G12. The PoW schedule comes from the environment on every network, including mainnet "Your mainnet gate refuses the override file. It does not refuse `IGNEUM_POW_EPOCH_BLOCKS`. A node without a file installs the schedule from the environment and `Params.pow_epoch_blocks` is never consulted." -Status: Open (4 October 2026). +Status: Fix built, pending rollout (4 October 2026, night; node branch `fud-consensus`, local worktree `vendor/igneum-node-fud`, no remote; main repo branch `fud-consensus`). Was: Open (4 October 2026). -Answer: Correct. `daemon.rs:339-340` installs the file's schedule only when the file names one; otherwise `PowSchedule::from_env()` (`consensus/core/src/igneum.rs:137-145`) installs on first read; the difficulty manager takes the global (`services.rs:113`); `daemon.rs:316-319` gates the file only. Fix: delete the environment fallback and install `Params.pow_epoch_blocks` from the network params on every start. Review id R4.1.1. +The fix: `PowSchedule::from_env()` is gone. `PowSchedule::from_env_for(network, base)` applies the three variables on devnet and simnet only and returns `None` elsewhere; the daemon calls `Params::apply_env_pow_schedule()` after the override file, prints "PoW schedule from the environment (...)" when applied and "Ignoring IGNEUM_POW_... on igneum-mainnet: the environment never sets a consensus parameter outside devnet and simnet" when not, then installs the network's schedule from `Params` on every network (`install_pow_schedule`), so `Params.pow_epoch_blocks` is what runs; the lazy fallback in `pow_schedule()` installs the devnet constants, never the environment. The miner takes all three schedule values from the template (`pow_epoch.day_ms` joined the two epoch fields in `PowEpochInfo`, the RPC model and the gRPC proto), so `IGNEUM_POW_DAY_MS` has no reader left in the miner (R4.1.9's day split is closed with it). The effective schedule is part of the params digest of X18, so a devnet node with the variable set cannot connect to one without it. Unit test `env_pow_schedule_is_devnet_and_simnet_only` (consensus-core, `config::params::tests`): the variable moves the devnet and simnet schedule and digest, leaves mainnet's and testnet's untouched, and the caller can tell "ignored" from "nothing set". Spec 2.8 and 8.7 state the rule. + +Answer (as found): Correct. `daemon.rs:339-340` installs the file's schedule only when the file names one; otherwise `PowSchedule::from_env()` (`consensus/core/src/igneum.rs:137-145`) installs on first read; the difficulty manager takes the global (`services.rs:113`); `daemon.rs:316-319` gates the file only. Fix: delete the environment fallback and install `Params.pow_epoch_blocks` from the network params on every start. Review id R4.1.1. Evidence: the files above. Experiment: start a node with `IGNEUM_POW_EPOCH_BLOCKS=60` and no file; its template's `epoch_blocks` must be the network's value. @@ -1676,18 +1678,22 @@ Evidence: `git ls-files | xargs grep -lF ` counts, `git log -S`. Experime ### X18. Two nodes with two override files connect, and only some mismatches fork "Your handshake compares the network name and nothing else. A PoW or difficulty mismatch forks and bans; a `finality` mismatch is a WARN; `rollout-v2.sh` throws the finality block away when it writes the file; the app rewrites the packaged file on every start." -Status: Open (4 October 2026). +Status: Fix built, pending rollout (4 October 2026, night; node branch `fud-consensus`, main repo branch `fud-consensus`). Was: Open (4 October 2026). -Answer: Correct. `protocol/flows/src/flow_context.rs:833` compares `network`; the version message has no params digest and no genesis hash. `pre_pow_validation.rs:37` and `pow_guard.rs:25-42` fork and ban on PoW and difficulty fields; `processes/finality.rs:669-688` only warns on a finality mismatch; `infra/cloud-devnet/rollout-v2.sh:20,35` rewrites the file as two fields; `app/igneum-app/src/engine.rs:742-760` rewrites `override-params.json` each start. Fix: a digest of the effective consensus params plus the genesis hash in the version message, refused on mismatch; the finality WARN becomes a refusal with the reason; `rollout-v2.sh` merges rather than replaces. Review ids R4.1.2, R4.1.12. +The fix: `Params::consensus_digest()` (BLAKE2b-256, domain `IgneumParamsDigest`, every consensus field in a fixed tagged order: genesis, difficulty, mass and lane limits, blockrate, crescendo, the ten finality fields, the PoW schedule, the three activation heights; not the dead `timestamp_deviation_tolerance`, not seeders or ports; spec 2.8 lists it). The version message carries it (`paramsDigest`, field 11); `initialize_connection` refuses a peer whose digest differs with `ProtocolError::ParamsDigestMismatch` and one WARN naming both digests before any flow is registered, so a finality-only mismatch, which used to connect and WARN "names N voters" after the fact, never connects. A peer with no digest (an older build) is refused on mainnet and testnet and let in with a WARN on devnet and simnet while the devnet rolls (an allowance to remove afterwards). The node prints its digest at start. Unit test `consensus_digest_covers_every_consensus_field_and_nothing_else`. Measured (`docs/bench-log.md`, "round-4 consensus items", digest run; `tools/finality-attacks/fud.mjs digest`, fast time, ports 29400+): a listener on the shared fast-time override and a dialler whose finality block differs by one DAA second of window: the dialler was refused at the handshake on both sides ("Refusing peer ...: consensus params digest mismatch, local 4bf7... remote 7a40..." on the listener, the reject message on the dialler), 0 peers after 25 s on both; a third node on the shared override connected in 1 s. Control on the finality-fixes build 6aa69a45: the mismatched dialler connected (1 peer, no line). `rollout-v2.sh` still rewrites the file as two fields and the app still rewrites the packaged file (R4.1.12): with the digest both now fail loudly at the handshake instead of forking; the merge fix for the script is not done tonight. + +Answer (as found): Correct. `protocol/flows/src/flow_context.rs:833` compares `network`; the version message has no params digest and no genesis hash. `pre_pow_validation.rs:37` and `pow_guard.rs:25-42` fork and ban on PoW and difficulty fields; `processes/finality.rs:669-688` only warns on a finality mismatch; `infra/cloud-devnet/rollout-v2.sh:20,35` rewrites the file as two fields; `app/igneum-app/src/engine.rs:742-760` rewrites `override-params.json` each start. Fix: a digest of the effective consensus params plus the genesis hash in the version message, refused on mismatch; the finality WARN becomes a refusal with the reason; `rollout-v2.sh` merges rather than replaces. Review ids R4.1.2, R4.1.12. Evidence: the files above. Experiment: two nodes on different files; the handshake must fail with the field named. ### F23. The equivocation ban is node-local, so honest nodes refuse each other's certificates "Evidence detected from an RPC vote stamps the sink's DAA; evidence carried in a block stamps the carrier's DAA. Two honest nodes hold different `until` for the same key, their voter lists differ by one at every checkpoint between the two expiries, and `voter_count` refuses the other's certificate for good." -Status: Open (4 October 2026). +Status: Fix built, pending rollout (4 October 2026, night; node branch `fud-consensus`, main repo branch `fud-consensus`). Was: Open (4 October 2026); reproduced by the red team the same evening on the stock s1 scenario (n0 kept the ban until DAA 726 against 239 on n1 and n2, 9 and 3 to 4 certificates refused "names N voters"). -Answer: Correct. `ingest_evidence` (`processes/finality.rs:600-612`); `ingest_certificate` (`:680-688`). Certificate validity is not a function of the DAG, the same defect R3.9 found in the execution veto. Fix: stamp every ban with the DAA of the block that carries the evidence; evidence seen by RPC is only acted on once carried. Review id R4.1.3. +The fix: the node-local `stripped` map is gone. Evidence is kept as `EvidenceRecord` (the two votes, the carriers with their DAA scores) and the ban at a checkpoint C is a function of C's past (`bans_at`): the key is stripped at C when some carrier lies in C's past and `daa(C) < daa(lowest carrier in C's past) + ban`. Evidence detected over RPC or gossip strips nothing until a block carries it; the node puts it in its next templates. The weight table cache stays ban-free and `voters_at` applies the checkpoint's own bans, so a certificate built before the carrier existed verifies on a node that saw the evidence later, and a node that saw it over RPC counts the same voters as one that saw it in the block. Evidence records are bounded (4,096; dropped once the ban ended two windows below the sink or never carried within one ban of being seen; 16 carriers per record), the same for vote and certificate carriers. The persisted state is layout 2; a layout-1 blob is read and converted on start, so no devnet node loses its locks on the upgrade. Unit test `ban_is_decided_by_the_carrying_block_so_nodes_agree_on_every_voter_list` (three `TestConsensus` nodes on one chain: the voter list agrees on all three at every checkpoint, the key is a voter before the carrier and after the ban and nowhere in between, the third node verifies the first two's certificates at every locked index). Measured (`docs/bench-log.md`, "round-4 consensus items", ban run; `fud.mjs ban`, 480 s, six voters, one equivocation at index 9 by a voter on n0 over RPC, n2 cut off 45 s around it and healed, so it saw the evidence late from the carrier block): on the `fud-consensus` build every node names 5 voters at the same indices (10 to 12 in the final pass, 10 to 13 in the first), 0 certificates refused "names N voters", 0 conflicting certificates, 0 locked indices disagreeing, voter counts agree at every index with lines on two or more nodes, locks continue to index 14 or 15 on all three. Control on the finality-fixes build: n0 (the RPC detector) refused 2 certificates "names N voters, this node counts M", the rest agreed because each node built its own; the red team's stock s1 scenario the same evening gave 9 / 3 / 4 refusals. Spec 3.6 and the 3.10 row state the rule. + +Answer (as found): Correct. `ingest_evidence` (`processes/finality.rs:600-612`); `ingest_certificate` (`:680-688`). Certificate validity is not a function of the DAG, the same defect R3.9 found in the execution veto. Fix: stamp every ban with the DAA of the block that carries the evidence; evidence seen by RPC is only acted on once carried. Review id R4.1.3. Evidence: the files above. Experiment: `tools/finality-attacks` with one equivocation detected on node A by RPC and on node B from the carrying block 30 DAA later; count certificates refused with "names N voters" between the two expiries; after the fix, zero. @@ -1696,9 +1702,11 @@ Red-team run, 4 October 2026 (evening, the 0.3.4 finality-fixes build with rule ### F24. A checkpoint determination is never revisited "After a reorg deeper than `checkpoint_depth`, the node's record for that index names a block off its chain. Every certificate the network forms for that index is refused as conflicting, with no equivocation anywhere, and the node voted for a block that is not on its chain." -Status: Open (4 October 2026); extends F7 and C4. +Status: Fix built, pending rollout (4 October 2026, night; node branch `fud-consensus`, main repo branch `fud-consensus`). Was: Open (4 October 2026); extends F7 and C4. -Answer: Correct. `on_virtual_changed` (`processes/finality.rs:404-440`) inserts once and advances `next_index`; `:661-666` refuses any certificate whose checkpoint is not the node's record. Fix: re-determine an unlocked index when the virtual's chain at that blue score changes; refuse only a certificate that conflicts with a lock. Review id R4.1.4. +The fix: after every virtual change, every unlocked checkpoint record whose block is no longer a chain ancestor of the sink is determined again on the new chain ("re-determined" log line); the certificate held over the old block is dropped, the fold clock restarts. A certificate over a block other than the node's determination at an unlocked index (or at an index not yet determined, up to 64 ahead) is no longer logged CONFLICTING and discarded: it is held pending (4 per index) and verified when a re-determination names its block; a block that cannot be the index's checkpoint on any chain (C1 as a function of the DAG: blue score under the target, or the selected parent's not) is refused outright. CONFLICTING now means what 3.11.4 says: a certificate against a LOCK. The certificate verification itself is unchanged; a locked index is never re-determined (fork choice keeps the chain through it). The node's own keys do not re-vote at a re-determined index (that would be equivocation); the lock there comes from the network's certificate, which is the point. Unit tests `reorg_past_an_unlocked_checkpoint_re_determines_it_and_verifies_the_pending_certificate` and `a_locked_checkpoint_pins_the_chain_and_a_certificate_against_it_conflicts`. Measured (`docs/bench-log.md`, "round-4 consensus items", reorg run; `fud.mjs reorg`, n0 with 30% of the weight cut off for 180 s while the 70% side kept locking, then healed): on the `fud-consensus` build n0 re-determined its own 1 to 2 split indices on the majority chain, verified the pending certificates at determination (2 in the final pass), logged 0 CONFLICTING and 0 refusals, and ended with every index the majority locked locked on the same block (0 disagreeing). Control on the finality-fixes build: n0 refused 3 certificates "is for X, this node's checkpoint is Y" and never locked indices 8 and 9 that the majority locked (the permanent hole of this entry), 0 disagreeing because the hole is not a lock. The final pass also found that a chain which becomes the sink at a lower blue score than the old one (more blue work per block after the split's difficulty drift) re-determined index 9 at a sink below its depth, to a block below the target: fixed the same night (an index the new chain has not reached is un-determined and determined again when it has; unit test `a_shallower_sink_un_determines_the_indices_it_cannot_reach`), RERUN_REORG. Spec 3.2 C1 and C4, and the 3.10 rows, state the rule. + +Answer (as found): Correct. `on_virtual_changed` (`processes/finality.rs:404-440`) inserts once and advances `next_index`; `:661-666` refuses any certificate whose checkpoint is not the node's record. Fix: re-determine an unlocked index when the virtual's chain at that blue score changes; refuse only a certificate that conflicts with a lock. Review id R4.1.4. Evidence: the files above. Experiment: a 30 s cut on a 3-node devnet at d = 20; the losing side must accept the network's certificate at that index with no CONFLICTING line. @@ -1707,7 +1715,7 @@ Red-team run, 5 October 2026, 00:56 (the 0.3.4 finality-fixes build, rule v3 on, ### F25. The fast-time harnesses cannot start a node, and the timestamp probe tests the old rule "Both attack harnesses rebuild each node's override with `JSON.parse` and `JSON.stringify` of `infra/fast-time/override-60x.json`. That file now carries two `u64::MAX` sentinels (`difficulty_v2_activation_daa`, `proving_v0_activation_daa`); a JavaScript number cannot hold them, the round-trip writes `18446744073709552000`, and `igneumd` refuses the file as a floating point where a u64 is expected. Every `--fast-time` run of `tools/finality-attacks` and `tools/harness` fails at the first node. Scenario 2 of `tools/harness` still probes the 132 s future bound and reports FAIL against the 10 s rule the node has carried since the timestamp fix." -Status: Open (4 October 2026, red-team run). Tooling, low severity: no consensus effect, but every fast-time attack run is blind until it is fixed. +Status: Fixed in both harnesses (4 October 2026, night: `tools/finality-attacks/lib/net.mjs` kept the sentinels as BigInt through the merge since the v3 runner of the evening; `tools/harness/lib/net.mjs` got the same reviver on branch `fud-memory`, merged into `fud-consensus`); the stale timestamp criterion of scenario 2 is not touched. Was: Open (4 October 2026, red-team run). Tooling, low severity: no consensus effect, but every fast-time attack run is blind until it is fixed. Answer: Correct, measured. The red-team run's first scenario errored on it (`docs/review/redteam-2026-10-04.md`, "Tooling defect"); `tools/proving-v0/run.mjs` already edits the file as text for this reason. Scenario 2 live probe: past floor pmt+1, future flip between +130.00 and +130.01 s of the probe's own offsets, every stamp from +10 s rejected; the node is right, the criterion is stale. Smallest fix: in `tools/finality-attacks/lib/net.mjs` and `tools/harness/lib/net.mjs` `overrideParams`, drop the two sentinel fields before `stringify` (absent means never) or splice the extra fields into the file text; in `tools/harness/scenarios/s2-timestamp.mjs`, probe `max(pmt + 1, parent - 10 s)` and the +10 s bound. Also stale: `tools/exec-attacks/scenario3_pgas.mjs` waits for an over-budget transaction to be included and skipped with `BlockProvingBudget`; since F-exec-B the mempool refuses it with the metered pgas, so the check should accept `ProvingGasAboveBlockLimit` from the pool (`docs/review/redteam-2026-10-04.md` row 28). And `tools/finality-attacks` scenario 5 compares the burster's share of the weight window with its share of the whole run, which only agree when the run is shorter than the window (row 17). @@ -1815,7 +1823,9 @@ Evidence: the files above. ### M30. A block or transaction flood grows the 0.3.4 node by hundreds of megabytes in a minute "On the 3 October ordering-layer node (no execution layer) the resource-exhaustion scenario grew RSS by 4, 11 and 14 MB and the 50x block flood by 30 MB. On the 0.3.4 build, same harness, same scenarios, same 60 s: template flood +6 MB, submit flood +269 MB, mempool flood +270 MB, block flood 302 to 1,082 MB on both nodes (567 MB at 10 s, 824 MB at 20 s). The harness calls it a pass because its bound is baseline + 512 MB; a peer that keeps going is not bounded by the harness." -Status: Open (4 October 2026, red-team run). Serious: a single peer at 50 blocks/s or 500 transactions/s is the devnet's own fast-miner event, and a node that grows 13 MB/s under it runs out of memory in minutes on the 2 to 4 GB cloud nodes. +Status: Fix built, pending rollout (4 October 2026, night; node branch `fud-memory`, merged into `fud-consensus`; main repo branch `fud-memory`, merged into `fud-consensus`). Was: Open (4 October 2026, red-team run). Serious: a single peer at 50 blocks/s or 500 transactions/s is the devnet's own fast-miner event, and a node that grows 13 MB/s under it runs out of memory in minutes on the 2 to 4 GB cloud nodes. + +The cause, measured (not the one guessed below): every RSS step in both floods is one `PoW cache built` line, 256 MiB each. The lottery engine (`consensus/pow/src/igneum.rs`, `IgneumEngine`) keyed its resident entries by `(epoch seed, day)` and built a full 256 MiB cache per entry, `KEEP = 4`, although the cache depends on the day seed alone (`igneum_pow::Epoch::from_seed_bytes`: cache from the day bytes, program from the epoch seed). The floods ran on the 60x profile, where an epoch rolls every 60 DAA, so the 50x block flood rolled it every 10 to 20 s and paid a cache each time; the 3 October run was on the devnet profile (3,600-DAA epochs, no roll in 60 s), which is what differed, not the execution layer. The s6 figures were cumulative from one starting RSS: the mempool flood's "+270 MB" was the submit flood's growth carried forward (its own cost is 1 MB), and 197 chain blocks cost under 1 MB in `ExecState.records`. On the live devnet the same engine costs 256 MiB per hourly epoch roll up to `KEEP`, which is the steady-state growth the coordinator saw on the 0.3.4 app node (1,081 MB at 27 min, 2,258 MB at 4 h 14 min, about 270 MB an hour). The fix: caches keyed by day, `KEEP_DAYS = 3` (3 x 256 MiB resident, plus at most 2 in-flight builds), programs keyed by `(epoch seed, day)` at a few KB each (`KEEP = 8`, LRU); an epoch roll on the same day builds no cache; node and miner compute the identical hash through `EpochRef` (unit test `epoch_rolls_share_the_day_cache`). No consensus rule changed. Measured (`docs/bench-log.md`, "ledger M30", two-node fast-time floods, before on the shipping finality-fixes build, after on `fud-memory` 796f758d): s6 submit flood +263 / +257 MB with 1 cache build each, after +3 / +2 MB and 0 builds; s7 50x block flood 302 to 1,085 MB with 3 builds, after 302 to 318 MB and 0 builds; template and mempool floods unchanged at +6 and +1 MB. Steady state, measured on two nodes at 1 block/s with no flood for 1,500 blocks on the 60x profile (`tools/harness/scenarios/s8-steady.mjs`, bench-log "ledger M30", steady-state paragraph): the shipping build reached 1,342 MB by block 514 after 9 cache builds (one per epoch roll; five 256 MiB chunks, the fifth an evicted one the allocator keeps) and 1,371 MB at 1,529 blocks; the fixed build held 319 MB at 510 blocks with 1 build and 603 MB at 1,526 blocks with 2 (the fast-time day rolled once), so on the devnet profile it is 256 MiB flat, 512 MiB around midnight UTC, 768 MiB worst case. Both builds then climb 30 MB per 1,000 blocks (30.7 before, 30.2 after), which is not the PoW cache (72 MB of non-cache footprint at 1,526 blocks against 23 MB at 0; the reading, not a measurement, is the consensus database's buffers and rusty-kaspa's entry-sized caches filling); the live node's 36 MB per 1,000 blocks between epoch rolls fits it. The app node's 2,258 MB at 4 h 14 min exceeds what this node build can reach (about 1.7 GB at 15,000 blocks) and includes its GPU worker, which needs its own `vmmap -summary`. Still growing by design and not bounded tonight: `ExecState.records` (1 to 2 KB per chain block, approximate, 100 to 170 MB a day at 1 block/s; a window must cover the proving sortition window and the by-number RPC history), `SNAPSHOT_RING = 64` state clones scaling with state size, the finality key registry (grows with distinct vote keys; votes, certificates, locks and evidence are trimmed). The harness now reports per-load RSS deltas and cache-build counts and takes `IGNEUM_HARNESS_BASE_PORT` and `IGNEUM_HARNESS_TMP`. Answer: Measured, cause not yet isolated. What changed between the two builds is the execution layer, which this node links: `igneum/exec/src/service.rs` keeps every `ChainBlockRecord` in `ExecState.records` (`:304`, `:419`, pushed and never truncated) plus `tx_index` and `inclusions` maps per transaction, and the mempool flood's 14,998 rejected transactions cost 270 MB, so rejected transactions are retained somewhere too. Smallest fix: bound `ExecState.records` to the record window plus the pruning depth and drop `tx_index`/`inclusions` entries with them; discard a rejected transaction's bytes at rejection; then re-run `tools/harness` s6 and s7 and require growth under 50 MB, the 3 October figure. @@ -1824,7 +1834,9 @@ Evidence: `/tmp/igneum-redteam-ord/results/s6-exhaustion.json` and `s7-flood.jso ### M31. The 0.3.4 node cannot produce a block template on mainnet, testnet or simnet parameters "`getBlockTemplate` on a `--simnet` node from the finality-fixes build answers every call with `Coinbase payload is above max length (204). Try to shorten the extra data.` and the network never makes a block. The coinbase of this build carries the vote-key reveal, the proof-record section and the finality section; only `DEVNET_PARAMS` was raised to `MAX_COINBASE_PAYLOAD_LEN_WITH_FINALITY` (16,384). `MAINNET_PARAMS`, `TESTNET_PARAMS` and `SIMNET_PARAMS` still carry Kaspa's 204 (`consensus/core/src/config/params.rs:705, 766, 828` against `:900`)." -Status: Open (4 October 2026, red-team run). Serious for anything that is not the devnet: a mainnet or testnet genesis on these parameters cannot be mined by a voting miner at all; harmless on the live devnet, whose parameters carry the raise. +Status: Fix built, pending rollout (4 October 2026, night; node branch `fud-consensus`). Was: Open (4 October 2026, red-team run). Serious for anything that is not the devnet: a mainnet or testnet genesis on these parameters cannot be mined by a voting miner at all; harmless on the live devnet, whose parameters carry the raise. + +The fix: `max_coinbase_payload_len` is `MAX_COINBASE_PAYLOAD_LEN_WITH_FINALITY` (16,384) on mainnet, testnet and simnet as on devnet, and the template builder fits the finality section into what the payload has left (`finality::encode_section_within`, certificates first, then evidence, then votes; the RPC computes the budget from the fixed part, the script, the node version, the miner's extra data and the record section), so a coinbase can never exceed the limit on any network whatever the voter count. The worst case did not fit even on devnet before: 8 certificates over 8,192 voters, 8 pieces of evidence and 48 votes come to about 30 KB, so the per-item bounds alone were not a bound. Unit test `largest_coinbase_fits_on_every_network` (consensus-core): the fixed part, the version, a key reveal, a full record section (8 records of 274 bytes) and the cut finality section fit under every network's limit, the cut keeps every certificate and every piece of evidence and at least 8 votes, the uncut section would not fit, and a budget under one certificate yields an empty section. The exec-attacks network (`tools/exec-attacks/net.sh`, `--simnet`) can drop its override once this build ships; not re-run tonight. Answer: Measured on the execution-layer attack network (`tools/exec-attacks/net.sh` runs `--simnet` with no override): three nodes up, 0 blocks, every template refused with that line (`docs/review/redteam-2026-10-04.md` row 27). Smallest fix: set `max_coinbase_payload_len: MAX_COINBASE_PAYLOAD_LEN_WITH_FINALITY` on the three other networks, and add a unit test that builds a coinbase with a key reveal, the maximum record section and a full certificate and checks it under every network's limit. The red-team run worked around it with `{"max_coinbase_payload_len": 16384}` in an override file.