release 0.3.13: merge fin-route (the bench-log entry on the finality route echo)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-06 13:21:35 +00:00
commit 4ee5fc0a31

View file

@ -2223,3 +2223,42 @@ Run b (`segments-pc2-pv1b`, 07:20Z to 07:51Z) claimed nothing in 88 passes: the
| The rule as shipped (no switch): fresh refused while the previous segment is pending (known-failed), accepted after it is unproven | 22 passed | 166.2 s |
| `--fresh-rule 0`: fresh accepted while the previous segment is pending, `freshAdmissible` true, still refused after a proven one, the second offer a duplicate ("segment already paid") | 23 passed | 139.9 s |
## 6 October 2026, 12:25 to 13:20Z, the finality route: why 26 fresh nodes lost the seed every checkpoint (fork `fin-route-0313` 5a339733 on 83089544; release engineer)
The fleet agent's finding (12:25Z): every rented node logged `P2P, route error: incoming route capacity for message type IgneumFinality has
been reached (peer: 188.245.5.161:26611)` every 20 to 60 s and reconnected at the checkpoint cadence (every 30 s); on a Vast box the seed
is the only peer, so each drop cost the node its only peer until the next dial.
**The cause is an echo, not the burst.** A certificate for an index below a node's window (`next_index` minus `KEEP_CHECKPOINTS` 2,000:
trimmed history) finds no record, goes through the off-chain path (`ingest_off_chain`), is LOCKED, pushed to gossip and sent to every peer,
trimmed again on the next pass, and comes back from every peer that held it. The seed's journal (`igneumd-v4`, 12:40 to 12:47Z):
| Line shape | Count in 7 min |
|---|---|
| `Finality: checkpoint N LOCKED by certificate: block <hash> ... is off this node's selected chain (not determined here yet)` | 13,354 (index 2954: 2,811; 2956: 2,799; 2957: 2,790; 2955: 2,778; 3897: 1,234; 1464: 942; the seed's next index was 6,127) |
| `route error: incoming route capacity for message type IgneumFinality` (the seed dropping ITS peers) | 13 |
| the real work (determined, received, LOCKED, folded, replaced by a heavier one) | 13 + 13 + 13 + 8 + 20 |
A fresh node on the Mac against the seed only (the 0.3.12 binary 83089544, 300 s, `kaspa_p2p_flows=debug`): 11,700 `Finality relay:
certificate` lines, every one `new=false`, 15 distinct indices, 240 per second at the peak (2,530 per 10 s), 203 votes; no route error on
the Mac (it drains 240/s with a 256-deep route) and one connection, where the fleet's slower boxes filled the route and lost the peer.
**The fix (four changes, 5a339733):** `ingest_certificate` ignores an index below `keep_from` (counted, debug: the echo stops at its source
once the seed runs it); the router's overflow policy for `IgneumFinality` is `Drop` with a counted warn once per 10 s per peer, never a
disconnect; the finality route is subscribed with 4,096 (a checkpoint's worst case is `MAX_VOTES_PER_BLOCK` 48 votes on each of 30 blocks
plus the certificates); the relay flow skips votes while IBD runs (counted, said once per 30 s; certificates still go in and land pending).
No consensus change, no digest change. Tests: the overflow-policy table (p2p 33 of 33), the flows crate (19 of 19), a certificate below
the window submitted twice (ignored, no gossip, counter 2; an index inside goes the normal way) with the finality tests (12 of 12).
**After, on the fixed binary against the still-unfixed seed** (203ae727, same run, 13:15:31 to 13:20:31Z): 63,628 certificates received
(the seed's echo had grown to 3,032 per 10 s at the peak as more fleet nodes joined), 0 route errors, 0 drops, 4 connections kept (the seed
and three peers learned from it), 168 votes skipped during IBD. The receiver side of the fix holds under a storm five times the morning's;
the source side (the guard) cannot show on the seed until 0.3.13 runs there, and the fresh node's own guard never fires during IBD (its
window starts at genesis), which is correct. Harness s7 on the fixed binary (`--quick --live-only`): PASS, 192 blocks accepted in 60 s under
a 50 blocks/s flood from one peer, honest template p50/p95/max 0.4/0.6/1.4 ms, rss 306 to 321 MB.
**Per tier:** a home miner joining today sees the warning and the peers=0 flicker every checkpoint until the seed runs 0.3.13; a rig the
same once; a pool user nothing; a fleet operator gets a node that keeps its only peer, and a seed that stops amplifying old certificates to
every peer (13,354 lines of work it did not need in seven minutes). Owed: the fleet agent's synced-node reading; a receiver-side limit on
certificates per index per minute as a second belt once the seed is fixed; the formatter's reflow of `finality.rs` (taken out of the commit).