docs: N9 fixed in dc141409; 7.7.5 the oracle at never and the retention store owed; the 0.3.19 renumbering

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-07 08:37:56 +00:00
parent a70f7b8eff
commit 69588aa155
2 changed files with 6 additions and 2 deletions

View file

@ -2097,7 +2097,7 @@ The e69e8a39 canary passed its headers proof (132,7xx headers, one session) and
Per tier: every fresh or wiped 0.3.18 install (a home miner's reinstall, a new pool, a new hand) would have stalled at the first record-carrying block of the chain for ever, alive and never synced; nodes upgraded in place would have carried their state until the first relayed record whose proof fails verification.
Fix on release-0.3.18-node: `ProofOracle::active_from` carries the switch (the sink takes it from `Params::proving_consensus_verify_daa` through `proof_sink`); the body rule, the IBD fetch and (through the rule) the relay retry apply to blocks at or above it only, so at never the node is 0.3.17 on this path, and with the switch set the history below it is not demanded. Known failed first on the pure gate (`the_carried_proof_rule_applies_from_the_switch_only`: the canary's block under never is not under the rule; the testnet's switch at genesis puts every block under it). What a switch set from genesis needs for a fresh join after the pool's retention is the testnet lane's go item: peers must serve the proofs of the pruning window. Commit hash: in the status log once green; the fleet's re-run of the c18-1 wipe on the new sha is the gate.
Fix on release-0.3.18-node: `ProofOracle::active_from` carries the switch (the sink takes it from `Params::proving_consensus_verify_daa` through `proof_sink`); the body rule, the IBD fetch and (through the rule) the relay retry apply to blocks at or above it only, so at never the node is 0.3.17 on this path, and with the switch set the history below it is not demanded. Known failed first on the pure gate (`the_carried_proof_rule_applies_from_the_switch_only`: the canary's block under never is not under the rule; the testnet's switch at genesis puts every block under it). What a switch set from genesis needs for a fresh join after the pool's retention is the testnet lane's go item: peers must serve the proofs of the pruning window. Fixed in dc141409 on release-0.3.19-node (09:36 UK; igneum-exec 27, consensus-core 123 with the gate's known-failed test, consensus 111 plus the two moved targets, p2p-flows 37, the kaspad and testing-integration checks clean on igneum-build-1); the fleet's re-run of the c18-1 wipe on it is the gate. The retention store (every node keeps the pruning window's proofs, serves them, drops below the pruning point; the rule above the pruning point only) is the second half, for a switch set from genesis.
## Status updates, 5 October 2026 (ledger sweep, night of 4 to 5 October)

View file

@ -329,7 +329,7 @@ Known failed first: the test holds the state as the catch-up does, asserts the t
### 7.7.2 The 0.3.18 node tree (7 October 2026, 06:30 to 07:09 UK)
The shipper's 0.3.18 node is ca3-v4-0318 merged into release-0.3.18-node, final at e69e8a39 (07:53 UK: the third merge, carrying c6a62e00 and 500ddd66, igneum-exec 26, the consensus crate 111 plus 1 plus 1, p2p-flows 37, the three checks clean); before it ae17ad00 (07:14 UK: 1cf43254 plus the second merge carrying 4f7c56a0 and the branch's copy of the test move; `cargo test -p kaspa-consensus` 111 plus 1 plus 1, p2p-flows 37, the kaspad, rpc-service and testing-integration checks clean). The first merge: ca3-v4-0318 (8220c944) into release-0.3.18-node (12153428: the ladder 1591ee1d, exec-sync 40fd8d8c, finality d8d1de1a, tail-emission 996b592f, kaspad's igneum-pow default, cc78a21f, the canary fix 90aaf38e), pushed as release-0.3.18-node = 1cf43254. Six conflicts: five taken as the release side (the header-version reading with the ladder's `header_signals_active` gate in igneum.rs, pre_ghostdag_validation.rs and ibd/flow.rs, where 0318 carried the same canary fix without the ladder; the shipper's two test fixes), the sixth the relay flow, where the pipelined `handle_block` now carries the release's 0.3.16 `IgneumProofMissing` retry arm ahead of the N6 arm. On igneum-build-1 at 1cf43254: cargo check -p kaspad and -p kaspa-testing-integration --tests clean, kaspa-consensus lib 110 twice, kaspa-consensus-core 123, kaspa-p2p-flows 37, kaspa-p2p-lib 20.
Renumbered 09:3x UK: the feature tree is 0.3.19 (mirror branch release-0.3.19-node, e69e8a39 at the rename, then dc141409), 0.3.18 is an app-only cut on the 0.3.17 tree, decimals 0.3.20. The shipper's feature-tree node is ca3-v4-0318 merged into release-0.3.18-node, final at e69e8a39 (07:53 UK: the third merge, carrying c6a62e00 and 500ddd66, igneum-exec 26, the consensus crate 111 plus 1 plus 1, p2p-flows 37, the three checks clean); before it ae17ad00 (07:14 UK: 1cf43254 plus the second merge carrying 4f7c56a0 and the branch's copy of the test move; `cargo test -p kaspa-consensus` 111 plus 1 plus 1, p2p-flows 37, the kaspad, rpc-service and testing-integration checks clean). The first merge: ca3-v4-0318 (8220c944) into release-0.3.18-node (12153428: the ladder 1591ee1d, exec-sync 40fd8d8c, finality d8d1de1a, tail-emission 996b592f, kaspad's igneum-pow default, cc78a21f, the canary fix 90aaf38e), pushed as release-0.3.18-node = 1cf43254. Six conflicts: five taken as the release side (the header-version reading with the ladder's `header_signals_active` gate in igneum.rs, pre_ghostdag_validation.rs and ibd/flow.rs, where 0318 carried the same canary fix without the ladder; the shipper's two test fixes), the sixth the relay flow, where the pipelined `handle_block` now carries the release's 0.3.16 `IgneumProofMissing` retry arm ahead of the N6 arm. On igneum-build-1 at 1cf43254: cargo check -p kaspad and -p kaspa-testing-integration --tests clean, kaspa-consensus lib 110 twice, kaspa-consensus-core 123, kaspa-p2p-flows 37, kaspa-p2p-lib 20.
The merge surfaced a test-order race both parents carry: the class-signal and ladder tables are process-wide statics, two tests install and restore them (`block_template_uses_current_block_version`, `cheap_checks_run_before_the_pow_engine`), and any block-building test of the same binary inside that window reads `WrongBlockVersion(1026 or 32770, 2)`; 1 of 112 then 4 of 112 under a box load of 178, 2 of 112 on the quiet box, and the 0.3.17 suite's "two load flakes that pass alone" were the same reading. Fix in 1cf43254 (and on ca3-v4-0318): the two installing tests are their own test targets (`consensus/tests/igneum_installed_signals.rs`, `consensus/tests/igneum_order_tests.rs`), one process each, and `INSTALL_TEST_LOCK` orders installers that share a binary; `cargo test -p kaspa-consensus --lib` no longer covers them, the crate without `--lib` does.
@ -351,6 +351,10 @@ The pool window measured get_block_template at 1 to 1.6 s on pool-1 and its solo
So on a young chain the RPC path is under a millisecond and the pool's second is chain-length work. The one per-request walk that scales with the chain is the live class-signal tally in `get_pow_epoch_info`: seven windows of 86,400 DAA on the devnet, which on pool-1's 250,000-DAA chain walks the whole chain and its mergesets on every template (250,000 header reads at a few microseconds each is the 1 to 1.6 s; the epoch-seed walk is at most one epoch and is memoised too). The reading is from the code and the arithmetic until the pool's stage line names it: 500ddd66's memo already serves the tally for 10 s per sink (in the 0.3.18 pin e69e8a39), and fd7de1b4 (the 0.3.19 line, 08:11 UK) refreshes it off the request path (the held tally served always, a refresh thread when stale, only a process's first request walks). The before is the pool's 1,632 ms; the after is the pool's stage line on the next canary. The 5,000-checkpoint join bench (7.7.1's asked size, before and after) was stopped at 07:00 UK because its rayon pool put the box at load 178 under the shipper's suites; it resumes when the 0.3.18 build chain has the box to itself.
### 7.7.5 The proof oracle at never (ledger N9; the 0.3.19 canary's stall, 7 October 2026)
The e69e8a39 canary on c18-1 passed its headers proof and stalled in the block stage at block 2,728: the daemon installed the proof oracle whatever `proving_consensus_verify_daa` said, the IBD flow then demanded the proof bytes of every record-carrying block from the syncer, and the serve flow answers from a peer's in-memory pool (records drop at 600 chain blocks, proof bytes live as long as the process), so a block from months ago was served by nobody; 56 asks of pool-1 and the hub, 110 failed IBDs, every peer 0.3.17, which installs no oracle and so never asked. Fix dc141409 on release-0.3.19-node: `ProofOracle::active_from` carries the switch, and the body rule, the IBD fetch and the relay retry apply from it only (known-failed test on the pure gate). Owed: the retention store for a switch set from genesis (the testnet): proof bytes persisted by the carrying block's DAA, kept for the pruning window plus the record window, served from the store, dropped below the pruning point; the rule above the pruning point only (the pipeline already skips trusted bodies); a retention test across a pruning move and a harness case (a fresh join after the window has passed) on the testnet lane's fast-time network with provers.
### 7.8 The stale-block nuisance guard (ledger N6)
The record (the shipper, 6 October 2026, 20:04 to 21:03Z): after the roll-back, four nodes whose datadirs held blocks mined under the older rule re-advertised their sink to the hub on every reconnect (the ready step sends an inv of the sink); the hub answered `wrong block version: got 1026 but expected 2, disconnecting`, the sender re-dialled in 30 s and advertised it again, and the hub's connection count climbed from 310 to 551 in an hour with no block mined. It stopped when the fleet wiped the datadirs.