diff --git a/docs/bench-log.md b/docs/bench-log.md index d367ccef4..e6d054cf9 100644 --- a/docs/bench-log.md +++ b/docs/bench-log.md @@ -1457,8 +1457,9 @@ Owner: the consensus engineer and cryptographer agent, worktrees `igneum-wt-c4` | v2 130 s, W 240 | same | v2, 240, 130 s | 2 (10, 11) | 114 s | 0 | apart, n0 locked 16 alone at 293 s | 0 / 1 / 1 | 0 | the sync gap again (n0 reconnected after v2's 210-s bound) | | v2 130 s, W 240 | c4 fix with the sync hook | v2, 240, 130 s | 1 (index 12) | 84 s | 3 (12, 13, 14 within 2 s of the first B block; 11 re-determined) | B, all three, A's split tip abandoned | 0 | 0 | PASS | | on 130 s, W 240 | same | v3, 240, 130 s | 0 (Poisson: 52 blue blocks, the index fell just short) | 84 s | n1 1, n2 2 (B's nodes adopted A's post-heal certificates and moved before IBD) | A, all three | 0 | 0 | the mirror case; not the C4 shape | +| on 140 s, W 240, addPeer at the heal | same | v3, 240, 140 s | 2 (11, 12, first at 12 s) | 3 s (the harness now dials through `addPeer`; the address goes as `{ip, port}`) | 1 (12 by certificate; 11 verified on the new chain) | B, all three, A's split tip abandoned | 0 | 0 | PASS | -Reading. With the consensus fix and the sync hook, a node on the heavier chain that receives a certificate for a chain it has never seen fetches that chain, verifies the certificate at its own block, locks it, moves its sink to the lighter certified chain and re-determines its own records onto it (the v2 W 240 row: 0 conflicts, 0 disagreements, every node on B's chain, which is the spec's F1 and the design's Fork choice items 1 to 4). The same code holds under rule v3 (unit test, the mirror case in the last row); a v3 row with B certifying during the split is Poisson-limited at these rates and is the run still owed (`SPLIT=140`). The fix does not and cannot cover a partition that outlasts the bound before the certificate arrives (rows 2, 3 and 6): there the node has already locked alone and 3.11.4 keeps that lock, the late certificate is CONFLICTING for the operator. On the live devnet (W 7,200 DAA, two hours) the bound is two hours after a side's last lock, so every partition under that heals by certificate. Raw: `scratchpad c4-results-*.md`, node logs `c4-*-n0.log`. +Reading. With the consensus fix and the sync hook, a node on the heavier chain that receives a certificate for a chain it has never seen fetches that chain, verifies the certificate at its own block, locks it, moves its sink to the lighter certified chain and re-determines its own records onto it (the v2 W 240 row: 0 conflicts, 0 disagreements, every node on B's chain, which is the spec's F1 and the design's Fork choice items 1 to 4). The same holds under rule v3 with the frozen table on (the last row: B certified 11 and 12 during a 140-s split, n0 reconnected 3 s after the heal once the harness dialled through `addPeer`, adopted 12 by certificate and ended on B's chain with the other two, 0 conflicts, 0 disagreements). The fix does not and cannot cover a partition that outlasts the bound before the certificate arrives (rows 2, 3 and 6): there the node has already locked alone and 3.11.4 keeps that lock, the late certificate is CONFLICTING for the operator. On the live devnet (W 7,200 DAA, two hours) the bound is two hours after a side's last lock, so every partition under that heals by certificate. Raw: `scratchpad c4-results-*.md`, node logs `c4-*-n0.log`. ## 5 October 2026 (evening), FUD ledger sweep round 6 diff --git a/docs/fud-ledger.md b/docs/fud-ledger.md index 2003f9718..365cf7072 100644 --- a/docs/fud-ledger.md +++ b/docs/fud-ledger.md @@ -620,7 +620,7 @@ Evidence: design doc Finality v2, Fork choice items 1 to 4; `sim/results.md` fin Sweep (5 October 2026, evening): the module-on against module-off comparison of O-3.8, run on the fast-time harness with the live node line (`tools/finality-attacks/c4.mjs`, fork 2b6d23ef, 3 nodes, 100-ms proxied links; raw tables in `docs/bench-log.md`, "FUD ledger sweep round 6", C4). The scenario separates weight from work: side B (n1, n2, four keys) holds 70% of the weight table and side A (n0, two keys) 30% when the link is cut; from the cut A mines at 0.6 blocks/s and B at 0.4, so A's chain is the heavier one by blue work while only B can certify under rule v3 (A holds 30% of the frozen table). Module off (`min_daa` never, so no certificate can form, fork choice bare GHOSTDAG): after a 150-s split the three nodes converged on A's heavier chain within 36 s of the heal, B's nodes re-determined their two split-time checkpoints onto it (F24), 0 conflicts. Module on (rule v3 from checkpoint DAA 0), 90-s split, n0 back on the link 6 s after the heal, A's chain at about 58 DAA of its own time, well inside the 120-DAA frozen table: during the split A locked nothing and B locked indices 7 and 8 on its own blocks, as designed; after the heal n0 did not switch. Its log: B's certificates for 8 and 9 arrived and were "kept pending until the chain decides (no lock at this index)" (the F24 path), n0's chain never changed because GHOSTDAG prefers its heavier tip and nothing in the node turns a verified certificate over an off-chain block into a fork-choice constraint, and one window after n0's last lock (index 7 at DAA 209, so from DAA 329) the frozen table no longer applied on A's chain ("no frozen table (no lock on this chain inside the window)"), A's two keys were 100% of A's own window table (B's post-cut blocks are red there and earn nothing), and n0 locked 10, 11 and 12 alone; B's certificates for 10 and 11 then logged CONFLICTING on n0, and B's nodes kept their certified chain. End state: sinks apart, 2 locked indices disagreeing across the nodes, a finality fork from a 96-s honest partition with no attacker and the frozen table intact at the heal; the same shape with a 150-s split (the table expired at the heal) and under rule v2 (the control: 1 conflict, sinks apart). So the answer to the critic is sharper than conceded: the overlay is specified to override blue work (spec 3.5, "GHOSTDAG among tips through all certified checkpoints") but the shipped node applies a certificate only to a block on its own chain, holds the rest pending a reorg that GHOSTDAG alone never produces, and after one window the heavier side certifies its own chain. Two honest views never reconcile. What closes it: a verified certificate over a block the node does not have on its selected chain must verify against the weight table at THAT block (its signers' weight there) and, when valid, constrain fork choice to tips through it, forcing the reorg (a certificate-driven reorg, bounded by the finality depth), with the node's own unlocked records re-determined on the new chain (F24); until then the exchange guidance of 3.9 (a node partitioned for more than a minute treats its locks as proof of work until it has seen the network's certificates agree with its own) is the only protection, and the 3.11.7 row for this case ("a certificate over a chain the node is not on") is missing. On the live devnet the window is 7,200 DAA (two hours) and the cliff is two hours after a side's last lock; a miner who joins with more hashrate than the weight table credits is the realistic work-majority side. The trace-driven adversary of O-3.8 is still owed. Spec rows: 3.5, 3.11.4, 3.11.7; node: `processes/finality.rs` (`ingest_certificate`'s pending branch, `fork_choice_lock`). Decision owner: the project lead (gate 3; a rule change to the node's fork choice). -Fix (5 October 2026, night): the certificate-driven reorg, built, unit-tested and measured; fork branch `c4-fix` on release-0.3.6 (a24ab01a), main branch `c4-fix`. The cause in the code: `processes/finality.rs` `ingest_certificate` verified a certificate only over the node's own determination and sent every other block to `hold_pending`; `fork_choice_lock` reads `state.locks`, which only `evaluate` filled over the node's own chain; so a certified block off the chain never became a lock. A second cause only the harness showed: the block relay (`protocol/flows/src/v10/blockrelay/flow.rs`) skips a relayed block lighter than the virtual's merge-depth root, and the certified chain is the lighter one by construction, so the heavier side never even received it. The fix: `ingest_off_chain` verifies a certificate against the voter table at its own block (canonical list, aggregate BLS, 2/3 of active and of total there, the frozen table under v3, the first-month gate), checks the block lies on the chain through the node's nearest locks (else CONFLICTING, 3.11.4, no lock withdrawn), locks the index on that block and asks the virtual processor to resolve (`VirtualStateProcessingMessage::Resolve`), so the sink search keeps only tips through it, whatever the blue work and whatever the merge depth (finality outranks merge depth; the depth-based finality point still bounds it, Kaspa's pruning safety, logged once); pending certificates over blocks the node lacks are retried on every virtual change; `evaluate` locks the node's own determination only on the chain through its locks; and while a pending certificate names a block the node lacks (`finality_wants_blocks`), the relay takes the lighter block, which orphans, falls out of range and triggers IBD of the certified chain. Not gated on v3: the live devnet's rule v2 took the same pending path (unit test `the_certificate_driven_reorg_holds_under_rule_v2`). Spec 3.5 carries the rule in one paragraph, 3.2 C4 and the 3.10 rows C4 and F1/F2 the implementation. Measured (bench-log "the C4 fix", `c4.mjs` with `WINDOW=240` so a 130-s split plus the 84-s p2p reconnect stays inside the window; the sweep's 120-DAA framing crosses F21's bound before any certificate can arrive once the real reconnect time is counted): under rule v2 side B locked index 12 during the split, n0 took B's first relayed block through the hook, locked 12, 13 and 14 by certificate within 2 s, re-determined 11, and all three nodes ended on B's certified chain with 0 CONFLICTING and 0 disagreeing locked indices (was: sinks apart, 5 conflicts on each node, 2 disagreeing); the module-off control is unchanged (heavier chain, 0 conflicts); the mirror case under v3 (B certified nothing, A certified after the heal) had B's nodes adopt A's certificates and move before IBD. Unit tests on PC 2: `kaspa-consensus` 97 passed, `kaspa-consensus-core` 101 passed. Still owed: a v3 harness row with B certifying during the split (Poisson at 0.4 blocks/s over 130 s; `SPLIT=140` queued), the live-devnet partition test of O-3.6, and the trace-driven adversary of O-3.8. What the fix does not cover, by design: a partition that outlasts the bound before the certificate arrives (the side has locked alone, 3.11.4 keeps it, the late certificate is CONFLICTING for the operator), which on the devnet means over two hours. Rollout: a consensus-behaviour change in the node with no params-digest change; a mixed fleet disagrees only in the state the old node already got wrong (an old node holds the certificate pending and stays on its heavier chain while new nodes move), and converges once every node is new; ship in the next node release with every node restarted on it. +Fix (5 October 2026, night): the certificate-driven reorg, built, unit-tested and measured; fork branch `c4-fix` on release-0.3.6 (a24ab01a), main branch `c4-fix`. The cause in the code: `processes/finality.rs` `ingest_certificate` verified a certificate only over the node's own determination and sent every other block to `hold_pending`; `fork_choice_lock` reads `state.locks`, which only `evaluate` filled over the node's own chain; so a certified block off the chain never became a lock. A second cause only the harness showed: the block relay (`protocol/flows/src/v10/blockrelay/flow.rs`) skips a relayed block lighter than the virtual's merge-depth root, and the certified chain is the lighter one by construction, so the heavier side never even received it. The fix: `ingest_off_chain` verifies a certificate against the voter table at its own block (canonical list, aggregate BLS, 2/3 of active and of total there, the frozen table under v3, the first-month gate), checks the block lies on the chain through the node's nearest locks (else CONFLICTING, 3.11.4, no lock withdrawn), locks the index on that block and asks the virtual processor to resolve (`VirtualStateProcessingMessage::Resolve`), so the sink search keeps only tips through it, whatever the blue work and whatever the merge depth (finality outranks merge depth; the depth-based finality point still bounds it, Kaspa's pruning safety, logged once); pending certificates over blocks the node lacks are retried on every virtual change; `evaluate` locks the node's own determination only on the chain through its locks; and while a pending certificate names a block the node lacks (`finality_wants_blocks`), the relay takes the lighter block, which orphans, falls out of range and triggers IBD of the certified chain. Not gated on v3: the live devnet's rule v2 took the same pending path (unit test `the_certificate_driven_reorg_holds_under_rule_v2`). Spec 3.5 carries the rule in one paragraph, 3.2 C4 and the 3.10 rows C4 and F1/F2 the implementation. Measured (bench-log "the C4 fix", `c4.mjs` with `WINDOW=240` so a 130-s split plus the 84-s p2p reconnect stays inside the window; the sweep's 120-DAA framing crosses F21's bound before any certificate can arrive once the real reconnect time is counted): under rule v2 side B locked index 12 during the split, n0 took B's first relayed block through the hook, locked 12, 13 and 14 by certificate within 2 s, re-determined 11, and all three nodes ended on B's certified chain with 0 CONFLICTING and 0 disagreeing locked indices (was: sinks apart, 5 conflicts on each node, 2 disagreeing); the module-off control is unchanged (heavier chain, 0 conflicts); under rule v3 (frozen table on, 140-s split, n0 reconnected 3 s after the heal) B locked 11 and 12 during the split, n0 adopted 12 by certificate and verified 11 on the new chain, all three nodes on B's chain, 0 CONFLICTING, 0 disagreeing; the mirror case (B certified nothing, A certified after the heal) had B's nodes adopt A's certificates and move before IBD. Unit tests on PC 2: `kaspa-consensus` 97 passed, `kaspa-consensus-core` 101 passed. Still owed: the live-devnet partition test of O-3.6 and the trace-driven adversary of O-3.8. What the fix does not cover, by design: a partition that outlasts the bound before the certificate arrives (the side has locked alone, 3.11.4 keeps it, the late certificate is CONFLICTING for the operator), which on the devnet means over two hours. Rollout: a consensus-behaviour change in the node with no params-digest change; a mixed fleet disagrees only in the state the old node already got wrong (an old node holds the certificate pending and stays on its heavier chain while new nodes move), and converges once every node is new; ship in the next node release with every node restarted on it. ### C5. vs Ethereum: you compare inclusion to finality "'Included in about one second, against twelve on Ethereum.' Inclusion in a DAG is not confirmation. Ethereum's twelve seconds is a slot, its finality is about thirteen minutes, and you compare your two-minute lock to that as if a two-minute lock by a pool committee were the same thing." diff --git a/docs/spec/03-finality.md b/docs/spec/03-finality.md index bbb6c88cb..00bf47335 100644 --- a/docs/spec/03-finality.md +++ b/docs/spec/03-finality.md @@ -282,7 +282,7 @@ Each guarantee, the scenario that tests it, and the measured result. Bench-log c | Acquired keys decay as the window moves (3.11.5) | results K, keys worth 20% and 40% bought, 30% hashrate, signing and silent, 30 days, five seeds | share follows b (1 - t/30) + 0.3 t/30 within 0.6 points at every sampled day in every seed; keys worth 20% rise to 30% on day 30 and never reach 1/3; keys worth 40% hold the veto from day 1 to day 19 or 20 (formula 20) and end at 30%; withholding its votes, the 40% buyer stalls 63,307 to 68,716 of 86,400 checkpoints in 30 days (the pause lasts until it has decayed below one third) and the 20% buyer 304 to 1,045; 0 conflicts (K, floor 2/3) | | Signing stops while mining continues, 1, 6, 24 h | results J and L1 | 34% and above: every checkpoint stalled for the whole silence; 33%: 665 to 727 of 720 in 6 h; 32%: 35 to 221; 30%: 0 to 40; first lock after resume 0 min at every weight; 0 conflicts (J, L1) | | Seeds during a pause (3.11.6) | devnet epoch boundary through a forced pause | not yet run (O-4.3 implementation) | -| A certificate over a chain the node is not on: the certificate-driven reorg (3.5) | fast-time harness `tools/finality-attacks/c4.mjs`, weight against work, 130-s split, 240-DAA window, rule v2 (the live devnet's) | the work-majority node fetched the certified chain, locked 12, 13 and 14 by certificate within 2 s of the first block, re-determined 11, and all three nodes ended on the certified chain with 0 conflicting certificates and 0 disagreeing locks; the module-off control took the heavier chain (bench-log "the C4 fix", 5 October 2026 night); the same under rule v3 is unit-tested, the harness row with the certifying side locking during the split is still owed | +| A certificate over a chain the node is not on: the certificate-driven reorg (3.5) | fast-time harness `tools/finality-attacks/c4.mjs`, weight against work, 130-s split, 240-DAA window, rule v2 (the live devnet's) | the work-majority node fetched the certified chain, locked 12, 13 and 14 by certificate within 2 s of the first block, re-determined 11, and all three nodes ended on the certified chain with 0 conflicting certificates and 0 disagreeing locks; the same under rule v3 with a 140-s split: the certifying side locked 11 and 12 during the split, the work-majority node adopted 12 by certificate 3 s after the heal and every node ended on the certified chain, 0 conflicting, 0 disagreeing; the module-off control took the heavier chain (bench-log "the C4 fix", 5 October 2026 night) | | Two certificates at one index: no lock withdrawn (3.11.4) | devnet with a forced double certificate | not yet run (O-3.17) | | `T` under the block reading (3.11.3) | O-3.3 re-run | not yet run (O-3.18) | | Participation credited only for the chain's checkpoint (3.11.1) | results C with an adversary voting for private blocks | not yet run (O-3.19) |