Round-4 consensus items: final runs on 977db931 (reorg-final2, the red team's f23 / f24b / f24c), the M31 and un-determination entries, result files

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-josh 2026-10-05 03:17:49 +01:00
parent 6332f7aa92
commit 0420c887e2
8 changed files with 117 additions and 5 deletions

View file

@ -1152,7 +1152,7 @@ Reading. Before: one cache build per epoch roll (25 rolls in 1,529 blocks, plus
Branches: node fork `fud-consensus` (worktree `vendor/igneum-node-fud`, from `finality-fixes` 6aa69a45, no remote; commits 9f738e2e the four items, ae9df8a3 pending certificates at fresh determinations, b755f43d merge of `fud-memory`, then M31 and the un-determination rule), main repo `fud-consensus`. Machine shared with the red team, the release builds and the memory runs all night (load average 100 to 146 until about 01:00 BST, under 20 after); every number here is a count, a lock or an index, not a timing. Runner `tools/finality-attacks/fud.mjs` (scenarios `digest`, `ban`, `reorg`; fast-time 60x profile with `finality_v3_activation_daa` 0 merged by the BigInt-safe `overrideParams`, ports 29400+, suffix 940, `/tmp/igneum-fin-fud`, node logs kept per scenario); the control is the shipping finality-fixes build `vendor/igneum-node/target-finality/release/igneumd` driven by the same miner. Raw result files in `docs/benchmarks/round4-consensus-2026-10-04/`.
**Builds.** `with-lock.sh build nice -n 19 cargo build --release -j 4 -p kaspad -p igneum-miner --features kaspad/igneum-pow` on a target directory cloned from `target-finality` with `cp -Rc` (APFS clonefile, 1 min, no disk): 17 min 53 s the first time under load 140, 8 min 58 s the second. Unit tests in the release profile (`cargo test --release -j 4 -p <crate> --lib`): consensus-core `config::params::tests` and `igneum` 20 of 20 (new: `consensus_digest_covers_every_consensus_field_and_nothing_else`, `env_pow_schedule_is_devnet_and_simnet_only`), consensus `processes::finality` 6 of 6 (new: `ban_is_decided_by_the_carrying_block_so_nodes_agree_on_every_voter_list` with three `TestConsensus` nodes on one chain, `reorg_past_an_unlocked_checkpoint_re_determines_it_and_verifies_the_pending_certificate`, `a_locked_checkpoint_pins_the_chain_and_a_certificate_against_it_conflicts`). After M31 and the un-determination rule: FINAL_TESTS.
**Builds.** `with-lock.sh build nice -n 19 cargo build --release -j 4 -p kaspad -p igneum-miner --features kaspad/igneum-pow` on a target directory cloned from `target-finality` with `cp -Rc` (APFS clonefile, 1 min, no disk): 17 min 53 s the first time under load 140, 8 min 58 s the second. Unit tests in the release profile (`cargo test --release -j 4 -p <crate> --lib`): consensus-core `config::params::tests` and `igneum` 20 of 20 (new: `consensus_digest_covers_every_consensus_field_and_nothing_else`, `env_pow_schedule_is_devnet_and_simnet_only`), consensus `processes::finality` 6 of 6 (new: `ban_is_decided_by_the_carrying_block_so_nodes_agree_on_every_voter_list` with three `TestConsensus` nodes on one chain, `reorg_past_an_unlocked_checkpoint_re_determines_it_and_verifies_the_pending_certificate`, `a_locked_checkpoint_pins_the_chain_and_a_certificate_against_it_conflicts`). After M31 and the un-determination rule (fork 977db931, with the `fud-memory` merge): consensus-core params and igneum tests 28 of 28 then params 12 of 12 (new: `largest_coinbase_fits_on_every_network`), consensus `processes::finality` 7 of 7 (new: `a_shallower_sink_un_determines_the_indices_it_cannot_reach`), kaspa-pow with `igneum-pow` 12 of 12; the second rebuild took 15 min 18 s under load 110 to 134.
**X18, the params digest** (`fud.mjs digest`: n1 listens on the shared override, n0 dials it with `finality.weight_window` 121 instead of 120, then n2 dials with the shared override).
@ -1180,12 +1180,18 @@ On the new build n2's one EQUIVOCATION line is the "carried by block" variety (i
| fud-consensus, first pass (9f738e2e) | 1 (index 8) | 2 | 4 / not re-read at fresh determinations (the gap fixed in ae9df8a3) | 0 / 0 / 0 | 0 | 8 and 9 (no certificate was verified at them) | 0 | 15 / 15 / 15 |
| fud-consensus, final pass (ae9df8a3) | 2 (8, 9) | 2 | 2 / 2 | 0 / 0 / 0 | 0 | none (the majority locked 10 and 11 during the split, not 8 and 9: 70% nominal is Poisson noise away from the floor) | 0 | 16 / 16 / 16 |
| control, finality-fixes, two passes | 2 then 1 | 0 | 0 / 0 | 0 / 0 / 0 | 3 (second pass) | 8 and 9 (second pass): the permanent hole | 0 | 17 then 15 |
| fud-consensus after the un-determination rule | RERUN_REORG_ROW |
| fud-consensus after the un-determination rule (977db931, `reorg-final2`) | 1 (index 9) | 1 | 1 / 1 | 0 / 0 / 0 | 0 | none (the majority locked nothing during the split this time: 8 at the cut, 8 at the heal, 16 at the end on all three) | 0 | 16 / 16 / 16; no record below its target |
What the final pass found: n0's index 9 was re-determined at a sink of blue score 263 to a block of blue score 263, below the index's target 270, because the majority chain became the sink by blue work (its difficulty drifted less than n0's during the split) before it had reached index 9's depth; a record never names a block below its target, so the rule now un-determines an index the new chain has not reached and determines it again when it has (fork commit after ae9df8a3, unit test `a_shallower_sink_un_determines_the_indices_it_cannot_reach`). The control's first pass logged no CONFLICTING because that build's refusal used other words ("certificate at index 8 is for X, this node's checkpoint is Y"), counted in the second pass.
**The red team's own reproductions** (`tools/finality-attacks/redteam/rtfin.mjs`, fin-attacks miner at 6 blocks/s, run on the fud-consensus build ae9df8a3 through the main worktree's env-aware harness lib on ports 29550+): RT_RESULTS.
**The red team's own reproductions** (`tools/finality-attacks/redteam/rtfin.mjs`, fin-attacks miner at 6 blocks/s, run on the fud-consensus build ae9df8a3 through the main worktree's env-aware harness lib on ports 29550+; result files in `docs/benchmarks/round4-consensus-2026-10-04/redteam-repro/`):
| Scenario | Build | Result |
|---|---|---|
| f23: two keys equivocating at every index, four honest voters, three nodes, 105 s (the evening's 9 / 3 / 4 refusals) | ae9df8a3 | PASS: equivocation detections 14 / 7 / 7, voter-count refusals 0 / 0 / 0, CONFLICTING 0 / 0 / 0, disagreeing locked indices 0, max locked 61 / 61 / 61 |
| f24c: 3/3 split, 16 s cut (96 DAA at 6 blocks/s), heal | ae9df8a3 | PASS: n1 determined 31..32 during the cut; after the heal 0 refusals, 0 CONFLICTING, 0 stuck indices, 0 disagreeing, max locked 61 / 61 |
| f24b: 4/2 split, 24 s cut, heal | ae9df8a3 and 977db931 | FAIL on both, outside F24: 24 s at 6 blocks/s is 144 DAA, longer than the 120-DAA window and the 60-DAA merge depth, so the chains never merge and each side locks its own chain alone ("100.0% of total, 100.0% of the table frozen at lock 39" on n1): the partition longer than a window of spec 3.7 item 9 (F21), mis-scaled by the scenario's assumption of 1 DAA a second. The 1,668 and 2,025 "PoW rejected" lines are the INFO line of `pre_ghostdag_validation.rs:158` for nonce-1 blocks, which `skip_proof_of_work` then accepts; no block was refused for them |
**G12** is covered by the digest run (the environment's schedule is part of the digest, so a devnet node with `IGNEUM_POW_EPOCH_BLOCKS` set cannot connect to one without it) and by `env_pow_schedule_is_devnet_and_simnet_only`; no mainnet node was started tonight. **M31** is covered by `largest_coinbase_fits_on_every_network`; no simnet network was started tonight (the red team's `tools/exec-attacks/net.sh` is the run that would show templates on simnet without an override).
**Uncertain.** (1) Every run is fast time (W = 120 DAA, ban 120, depth 20) on three nodes with 100-ms links; the mainnet values are 30 days, 30 days and 60 blocks. (2) The ban run shows one equivocation at one index; the red team's s1 (equivocation at every index, two keys) was not re-run on the new build tonight except through RT_RESULTS_NOTE. (3) The digest-less allowance on devnet and simnet is deliberate for the rollout and is a hole until removed. (4) The reorg run's final pass had the majority lock no index during the first 60 s of the split, so the re-determination at 8 and 9 was exercised, the pending-certificate path only at 10 and 11; the first pass exercised the opposite. (5) The un-determination rule has a unit test and RERUN_REORG_NOTE.
**Uncertain.** (1) Every run is fast time (W = 120 DAA, ban 120, depth 20) on three nodes with 100-ms links; the mainnet values are 30 days, 30 days and 60 blocks. (2) The ban run shows one equivocation at one index; the red team's s1 (equivocation at every index, two keys) was re-run on the new build only through the red team's f23 above (0 refusals where the evening had 9 / 3 / 4). (3) The digest-less allowance on devnet and simnet is deliberate for the rollout and is a hole until removed. (4) The reorg run's final pass had the majority lock no index during the first 60 s of the split, so the re-determination at 8 and 9 was exercised, the pending-certificate path only at 10 and 11; the first pass exercised the opposite. (5) The un-determination rule has a unit test and one network pass (`reorg-final2`) in which the shallow-sink case did not recur, so the rule is exercised by the test, not by a run; the case needs a split whose difficulty drifts enough for blue work to overtake blue score, which happened once in four runs. (6) The red team's f24b is a window-length partition at 6 blocks/s, so it measures F21's stated limit, not F24; a 4/2 cut under 20 s at that rate would be the F24 case.

View file

@ -0,0 +1,8 @@
[
{
"scenario": "F23 (equivocation ban node-local; honest nodes refuse each other's certs)",
"expected": "ban expiry identical across honest nodes; 0 voter-count refusals; 0 CONFLICTING; 0 disagreeing locked indices",
"observed": "equiv detections 14/7/7; stripped-until per node - | - | - (distinct values 0); voter-count-mismatch refusals 0/0/0; CONFLICTING 0/0/0; disagreeing locked indices 0; maxLocked 61/61/61",
"pass": true
}
]

View file

@ -0,0 +1,8 @@
[
{
"scenario": "F24b (deep reorg under merge depth; determination never revisited)",
"expected": "after a reorg deeper than checkpoint_depth the losing node re-determines the moved indices and accepts the network certificates; 0 false CONFLICTING, 0 stuck indices, 0 disagreeing locks",
"observed": "n1 determined 32..33 during the 24s cut; after heal: cert-for-other-block refusals 0/26, CONFLICTING 7/3, equivocation 0/0, PoW-rejected 1668/2025, sinks equal false, disagreeing locked indices 9, n1 indices stuck unlocked that n0 locked [33], maxLocked 54/43",
"pass": false
}
]

View file

@ -0,0 +1,8 @@
[
{
"scenario": "F24c (3/3 split, 16 s cut under merge depth; determination never revisited)",
"expected": "after a reorg deeper than checkpoint_depth the losing node re-determines the moved indices and accepts the network certificates; 0 false CONFLICTING, 0 stuck indices, 0 disagreeing locks",
"observed": "n1 determined 31..32 during the 16s cut; after heal: cert-for-other-block refusals 0/0, CONFLICTING 0/0, equivocation 0/0, PoW-rejected 1935/1930, sinks equal false, disagreeing locked indices 0, n1 indices stuck unlocked that n0 locked [], maxLocked 61/61",
"pass": true
}
]

View file

@ -0,0 +1,44 @@
### digest-final: n1 on the shared fast-time override, n0 with finality.weight_window 121 (one DAA second more), n2 on the shared override
| measure | n0 (mismatched) | n1 (listener) | n2 (matching) |
|---|---|---|---|
| params digest printed at start | 4bf763ba5b78c6ac88463932f522679186cb641f631671eb25fc17b01071aea9 | 7a40cc3b90c7726b9813113448510bd0597856baf9f0e6f245e3cf5cac494852 | 7a40cc3b90c7726b9813113448510bd0597856baf9f0e6f245e3cf5cac494852 |
| "consensus params digest mismatch" lines | 2 | 2 | 0 |
| peers after 25 s (n0, n1) and after n2 dialled (n1) | 0 | 0 then 1 | connected after 1 s |
n0's line: `WARN ] P2P, got reject message: consensus params digest mismatch - local: 7a40cc3b90c7726b9813113448510bd0597856baf9f0e6f245e3cf5cac494852, remote: 4bf763ba5b78c6ac88463932f522679186cb641f631671eb25fc17b01071aea9: the peer's override file, environment or build differs from peer: 127.0.0.1:29411`
### ban-final: 480 s, 1 blocks/s in all, 6 voters, delay 100 ms, a0 equivocates once at index 9; P2 cut at 252 s, healed at 297 s; n0 detected the equivocation at 288 s
| measure | n0 (saw it over RPC) | n1 (from the block at once) | n2 (from the block, after the heal) |
|---|---|---|---|
| EQUIVOCATION lines (of which "carried by block") | 2 (1) | 1 (1) | 1 (1) |
| certificates refused "names N voters, this node counts M" | 0 | 0 | 0 |
| CONFLICTING certificate lines | 0 | 0 | 0 |
| indices whose certificates name 5 voters (a0 stripped) | 10..12 (3) | 10..12 (3) | 10..12 (3) |
| max locked index at the end | 15 | 15 | 15 |
indices with certificate lines on at least two nodes: voter counts agree at 10, differ at 0; locked indices disagreeing across the three nodes: 0
### reorg-final: warm 230 s, split 180 s (n0 alone with 30% of the weight), heal window 150 s, 1 blocks/s in all, delay 100 ms
| measure | n0 (cut off, 30%) | n1 (70% side) | n2 (70% side) |
|---|---|---|---|
| max locked index at the cut | 7 | 7 | 7 |
| max locked index at the heal | 7 | 11 | 11 |
| max locked index at the end | 16 | 16 | 16 |
| "re-determined" lines | 2 | 0 | 0 |
| certificates kept pending | 2 | 0 | 0 |
| CONFLICTING certificate lines | 0 | 0 | 0 |
| certificates refused over another block, pre-F24 wording | 0 | 0 | 0 |
| pending certificates verified at determination (did not verify) | 2 (0) | 0 (0) | 0 (0) |
| indices n1 locked that this node did not lock | none | | |
n0 determined 2 checkpoint(s) on its own chain during the split (indices 8, 9); after the heal n0 holds the same locked block as n1 at 0 of them; locked indices disagreeing across the three nodes: 0; n0 reconnected 10 s after the gate reopened
Finality: checkpoint 8 re-determined: block a806eb0744d0e12c09c1de08bd6afc2445cd4c27edff2f78d50ed296c33fbfb7 (blue score 240, daa 239), was 46cfb6b02eb5728cba01cd729d87463fb3bd4603abbbf13921b4067b3f1b5895: the selected chain moved past it
Finality: checkpoint 9 re-determined: block 8c1b51dd691441543fcb76809c67058d1c161b432ad091d949e1879270225174 (blue score 263, daa 262), was 88138fd98a72ac2ec4a5778f97c3f0dd913a2dfc70335ff6fb487649e8ffc422: the selected chain moved past it
[PASS] digest-final
[PASS] ban-final
[FAIL] reorg-final

View file

@ -0,0 +1,20 @@
### reorg-final2: warm 230 s, split 180 s (n0 alone with 30% of the weight), heal window 150 s, 1 blocks/s in all, delay 100 ms
| measure | n0 (cut off, 30%) | n1 (70% side) | n2 (70% side) |
|---|---|---|---|
| max locked index at the cut | 8 | 8 | 8 |
| max locked index at the heal | 8 | 8 | 8 |
| max locked index at the end | 16 | 16 | 16 |
| "re-determined" lines | 1 | 0 | 0 |
| certificates kept pending | 1 | 0 | 0 |
| CONFLICTING certificate lines | 0 | 0 | 0 |
| certificates refused over another block, pre-F24 wording | 0 | 0 | 0 |
| pending certificates verified at determination (did not verify) | 1 (0) | 0 (0) | 0 (0) |
| indices n1 locked that this node did not lock | none | | |
| records whose block is below the index's target blue score | none | | |
n0 determined 1 checkpoint(s) on its own chain during the split (indices 9); after the heal n0 holds the same locked block as n1 at 0 of them; locked indices disagreeing across the three nodes: 0; n0 reconnected 10 s after the gate reopened
Finality: checkpoint 9 re-determined: block dd8f57d7a5bc1a80eab4d87704310396c614d7cdb5ea21595694b9453b21d1bf (blue score 270, daa 269), was c3379892186e5e96e679bdfb3ec6991a9cb41ebe953e84baf3f1cf0ba07388d6: the selected chain moved past it
[PASS] reorg-final2

View file

@ -0,0 +1,18 @@
### reorg-old2: warm 230 s, split 180 s (n0 alone with 30% of the weight), heal window 150 s, 1 blocks/s in all, delay 100 ms
| measure | n0 (cut off, 30%) | n1 (70% side) | n2 (70% side) |
|---|---|---|---|
| max locked index at the cut | 7 | 7 | 7 |
| max locked index at the heal | 7 | 11 | 11 |
| max locked index at the end | 15 | 15 | 15 |
| "re-determined" lines | 0 | 0 | 0 |
| certificates kept pending | 0 | 0 | 0 |
| CONFLICTING certificate lines | 0 | 0 | 0 |
| certificates refused over another block, pre-F24 wording | 3 | 0 | 0 |
| pending certificates verified at determination (did not verify) | 0 (0) | 0 (0) | 0 (0) |
| indices n1 locked that this node did not lock | 8 9 | | |
n0 determined 1 checkpoint(s) on its own chain during the split (indices 8); after the heal n0 holds the same locked block as n1 at 0 of them; locked indices disagreeing across the three nodes: 0; n0 reconnected 40 s after the gate reopened
[FAIL] reorg-old2

View file

@ -1704,7 +1704,7 @@ Red-team run, 4 October 2026 (evening, the 0.3.4 finality-fixes build with rule
Status: Fix built, pending rollout (4 October 2026, night; node branch `fud-consensus`, main repo branch `fud-consensus`). Was: Open (4 October 2026); extends F7 and C4.
The fix: after every virtual change, every unlocked checkpoint record whose block is no longer a chain ancestor of the sink is determined again on the new chain ("re-determined" log line); the certificate held over the old block is dropped, the fold clock restarts. A certificate over a block other than the node's determination at an unlocked index (or at an index not yet determined, up to 64 ahead) is no longer logged CONFLICTING and discarded: it is held pending (4 per index) and verified when a re-determination names its block; a block that cannot be the index's checkpoint on any chain (C1 as a function of the DAG: blue score under the target, or the selected parent's not) is refused outright. CONFLICTING now means what 3.11.4 says: a certificate against a LOCK. The certificate verification itself is unchanged; a locked index is never re-determined (fork choice keeps the chain through it). The node's own keys do not re-vote at a re-determined index (that would be equivocation); the lock there comes from the network's certificate, which is the point. Unit tests `reorg_past_an_unlocked_checkpoint_re_determines_it_and_verifies_the_pending_certificate` and `a_locked_checkpoint_pins_the_chain_and_a_certificate_against_it_conflicts`. Measured (`docs/bench-log.md`, "round-4 consensus items", reorg run; `fud.mjs reorg`, n0 with 30% of the weight cut off for 180 s while the 70% side kept locking, then healed): on the `fud-consensus` build n0 re-determined its own 1 to 2 split indices on the majority chain, verified the pending certificates at determination (2 in the final pass), logged 0 CONFLICTING and 0 refusals, and ended with every index the majority locked locked on the same block (0 disagreeing). Control on the finality-fixes build: n0 refused 3 certificates "is for X, this node's checkpoint is Y" and never locked indices 8 and 9 that the majority locked (the permanent hole of this entry), 0 disagreeing because the hole is not a lock. The final pass also found that a chain which becomes the sink at a lower blue score than the old one (more blue work per block after the split's difficulty drift) re-determined index 9 at a sink below its depth, to a block below the target: fixed the same night (an index the new chain has not reached is un-determined and determined again when it has; unit test `a_shallower_sink_un_determines_the_indices_it_cannot_reach`), RERUN_REORG. Spec 3.2 C1 and C4, and the 3.10 rows, state the rule.
The fix: after every virtual change, every unlocked checkpoint record whose block is no longer a chain ancestor of the sink is determined again on the new chain ("re-determined" log line); the certificate held over the old block is dropped, the fold clock restarts. A certificate over a block other than the node's determination at an unlocked index (or at an index not yet determined, up to 64 ahead) is no longer logged CONFLICTING and discarded: it is held pending (4 per index) and verified when a re-determination names its block; a block that cannot be the index's checkpoint on any chain (C1 as a function of the DAG: blue score under the target, or the selected parent's not) is refused outright. CONFLICTING now means what 3.11.4 says: a certificate against a LOCK. The certificate verification itself is unchanged; a locked index is never re-determined (fork choice keeps the chain through it). The node's own keys do not re-vote at a re-determined index (that would be equivocation); the lock there comes from the network's certificate, which is the point. Unit tests `reorg_past_an_unlocked_checkpoint_re_determines_it_and_verifies_the_pending_certificate` and `a_locked_checkpoint_pins_the_chain_and_a_certificate_against_it_conflicts`. Measured (`docs/bench-log.md`, "round-4 consensus items", reorg run; `fud.mjs reorg`, n0 with 30% of the weight cut off for 180 s while the 70% side kept locking, then healed): on the `fud-consensus` build n0 re-determined its own 1 to 2 split indices on the majority chain, verified the pending certificates at determination (2 in the final pass), logged 0 CONFLICTING and 0 refusals, and ended with every index the majority locked locked on the same block (0 disagreeing). Control on the finality-fixes build: n0 refused 3 certificates "is for X, this node's checkpoint is Y" and never locked indices 8 and 9 that the majority locked (the permanent hole of this entry), 0 disagreeing because the hole is not a lock. The final pass also found that a chain which becomes the sink at a lower blue score than the old one (more blue work per block after the split's difficulty drift) re-determined index 9 at a sink below its depth, to a block below the target: fixed the same night (an index the new chain has not reached is un-determined and determined again when it has; unit test `a_shallower_sink_un_determines_the_indices_it_cannot_reach`) and re-run (`reorg-final2`, fork 977db931): the majority locked nothing during the split this time (Poisson again), n0 re-determined its one split index, verified 1 pending certificate at determination, 0 CONFLICTING, 0 refused, 0 disagreeing, no record below its target, every index the majority locked locked on n0 too. The red team's own reproductions on the new build (`tools/finality-attacks/redteam/rtfin.mjs`, fin-attacks miner at 6 blocks/s): `f24c` (3/3 split, 16 s cut) PASS with 0 refusals, 0 CONFLICTING, 0 stuck indices, 0 disagreeing; `f24b` (4/2 split, 24 s cut) FAIL on both builds for a reason outside F24: at 6 blocks/s the DAA score advances 6 a second, so the 24-s cut is 144 DAA, longer than the 120-DAA weight window and the 60-DAA merge depth; the two chains cannot merge at the heal and each side's table holds only its own keys (n1's lone locks read "100.0% of total, 100.0% of the table frozen at lock 39"), which is the partition longer than a window that spec 3.7 item 9 states and F21 conceded, not a reorg. The scenario's "under merge depth" assumes 1 DAA a second. Spec 3.2 C1 and C4, and the 3.10 rows, state the rule.
Answer (as found): Correct. `on_virtual_changed` (`processes/finality.rs:404-440`) inserts once and advances `next_index`; `:661-666` refuses any certificate whose checkpoint is not the node's record. Fix: re-determine an unlocked index when the virtual's chain at that blue score changes; refuse only a certificate that conflicts with a lock. Review id R4.1.4.