From 1cf83109a0254d315791064c0ef3c88dda4811bb Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Tue, 6 Oct 2026 13:21:19 +0000 Subject: [PATCH] bench-log: the finality route, the certificate echo below the window (fork fin-route-0313 5a339733), before and after rates Co-Authored-By: Claude Fable 5.1 --- docs/bench-log.md | 39 +++++++++++++++++++++++++++++++++++++++ 1 file changed, 39 insertions(+) diff --git a/docs/bench-log.md b/docs/bench-log.md index a78eade86..ec88c115a 100644 --- a/docs/bench-log.md +++ b/docs/bench-log.md @@ -2223,3 +2223,42 @@ Run b (`segments-pc2-pv1b`, 07:20Z to 07:51Z) claimed nothing in 88 passes: the | The rule as shipped (no switch): fresh refused while the previous segment is pending (known-failed), accepted after it is unproven | 22 passed | 166.2 s | | `--fresh-rule 0`: fresh accepted while the previous segment is pending, `freshAdmissible` true, still refused after a proven one, the second offer a duplicate ("segment already paid") | 23 passed | 139.9 s | + +## 6 October 2026, 12:25 to 13:20Z, the finality route: why 26 fresh nodes lost the seed every checkpoint (fork `fin-route-0313` 5a339733 on 83089544; release engineer) + +The fleet agent's finding (12:25Z): every rented node logged `P2P, route error: incoming route capacity for message type IgneumFinality has +been reached (peer: 188.245.5.161:26611)` every 20 to 60 s and reconnected at the checkpoint cadence (every 30 s); on a Vast box the seed +is the only peer, so each drop cost the node its only peer until the next dial. + +**The cause is an echo, not the burst.** A certificate for an index below a node's window (`next_index` minus `KEEP_CHECKPOINTS` 2,000: +trimmed history) finds no record, goes through the off-chain path (`ingest_off_chain`), is LOCKED, pushed to gossip and sent to every peer, +trimmed again on the next pass, and comes back from every peer that held it. The seed's journal (`igneumd-v4`, 12:40 to 12:47Z): + +| Line shape | Count in 7 min | +|---|---| +| `Finality: checkpoint N LOCKED by certificate: block ... is off this node's selected chain (not determined here yet)` | 13,354 (index 2954: 2,811; 2956: 2,799; 2957: 2,790; 2955: 2,778; 3897: 1,234; 1464: 942; the seed's next index was 6,127) | +| `route error: incoming route capacity for message type IgneumFinality` (the seed dropping ITS peers) | 13 | +| the real work (determined, received, LOCKED, folded, replaced by a heavier one) | 13 + 13 + 13 + 8 + 20 | + +A fresh node on the Mac against the seed only (the 0.3.12 binary 83089544, 300 s, `kaspa_p2p_flows=debug`): 11,700 `Finality relay: +certificate` lines, every one `new=false`, 15 distinct indices, 240 per second at the peak (2,530 per 10 s), 203 votes; no route error on +the Mac (it drains 240/s with a 256-deep route) and one connection, where the fleet's slower boxes filled the route and lost the peer. + +**The fix (four changes, 5a339733):** `ingest_certificate` ignores an index below `keep_from` (counted, debug: the echo stops at its source +once the seed runs it); the router's overflow policy for `IgneumFinality` is `Drop` with a counted warn once per 10 s per peer, never a +disconnect; the finality route is subscribed with 4,096 (a checkpoint's worst case is `MAX_VOTES_PER_BLOCK` 48 votes on each of 30 blocks +plus the certificates); the relay flow skips votes while IBD runs (counted, said once per 30 s; certificates still go in and land pending). +No consensus change, no digest change. Tests: the overflow-policy table (p2p 33 of 33), the flows crate (19 of 19), a certificate below +the window submitted twice (ignored, no gossip, counter 2; an index inside goes the normal way) with the finality tests (12 of 12). + +**After, on the fixed binary against the still-unfixed seed** (203ae727, same run, 13:15:31 to 13:20:31Z): 63,628 certificates received +(the seed's echo had grown to 3,032 per 10 s at the peak as more fleet nodes joined), 0 route errors, 0 drops, 4 connections kept (the seed +and three peers learned from it), 168 votes skipped during IBD. The receiver side of the fix holds under a storm five times the morning's; +the source side (the guard) cannot show on the seed until 0.3.13 runs there, and the fresh node's own guard never fires during IBD (its +window starts at genesis), which is correct. Harness s7 on the fixed binary (`--quick --live-only`): PASS, 192 blocks accepted in 60 s under +a 50 blocks/s flood from one peer, honest template p50/p95/max 0.4/0.6/1.4 ms, rss 306 to 321 MB. + +**Per tier:** a home miner joining today sees the warning and the peers=0 flicker every checkpoint until the seed runs 0.3.13; a rig the +same once; a pool user nothing; a fleet operator gets a node that keeps its only peer, and a seed that stops amplifying old certificates to +every peer (13,354 lines of work it did not need in seven minutes). Owed: the fleet agent's synced-node reading; a receiver-side limit on +certificates per index per minute as a second belt once the seed is fixed; the formatter's reflow of `finality.rs` (taken out of the commit).