Counter ASIC 3.0 node plan 7.7.6: the fresh-join cost measured on the dc141409 canary (weight tables 68 per checkpoint from the 64-table cache clear; the 1,024-table fix)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
a7adf116c6
commit
d296b6e33c
1 changed files with 13 additions and 0 deletions
|
|
@ -372,6 +372,19 @@ The pool window measured get_block_template at 1 to 1.6 s on pool-1 and its solo
|
|||
|
||||
So on a young chain the RPC path is under a millisecond and the pool's second is chain-length work. The one per-request walk that scales with the chain is the live class-signal tally in `get_pow_epoch_info`: seven windows of 86,400 DAA on the devnet, which on pool-1's 250,000-DAA chain walks the whole chain and its mergesets on every template (250,000 header reads at a few microseconds each is the 1 to 1.6 s; the epoch-seed walk is at most one epoch and is memoised too). The reading is from the code and the arithmetic until the pool's stage line names it: 500ddd66's memo already serves the tally for 10 s per sink (in the 0.3.18 pin e69e8a39), and fd7de1b4 (the 0.3.19 line, 08:11 UK) refreshes it off the request path (the held tally served always, a refresh thread when stale, only a process's first request walks). The before is the pool's 1,632 ms; the after is the pool's stage line on the next canary. The 5,000-checkpoint join bench (7.7.1's asked size, before and after) was stopped at 07:00 UK because its rayon pool put the box at load 178 under the shipper's suites; it resumes when the 0.3.18 build chain has the box to itself.
|
||||
|
||||
### 7.7.6 The fresh-join cost, measured on the dc141409 canary (c18-1, 7 October 2026, 08:51 to 10:29Z)
|
||||
|
||||
The 0.3.20 tree's first fresh join of the live devnet (RTX 3070 pod, wiped datadir): started 08:51:50Z, synced 10:29:19Z at 141,357 blocks, 97.5 minutes wall; one headers-proof IBD and five relay catch-ups of a second each; no proof asks, no idle drops, the app's poller 4,217 reads with 0 errors. The IBD-end line (8.7 above): "finality time: bodies 16,592,238 ms over 141,699 blocks, virtual 1,061,858 ms over 36,183 changes, weight tables 968,243 ms over 359,709 tables (3,832,285,840 blocks walked), signatures 711,416 ms over 439,307, persists 40,537 ms over 3,324".
|
||||
|
||||
| Item | Canary (c18-1) | Bench, after row (box, 16 cores, 2,000 checkpoints) | Reading |
|
||||
|---|---|---|---|
|
||||
| Weight tables | 359,709 tables, 68 per checkpoint, 10,650 blocks walked each, 968 s | 2,335 tables, 1.2 per checkpoint, 18 s | the cache held 64 tables and cleared itself past that; one IBD batch of 2,000 bodies spans about 66 checkpoints, so every batch threw the set away and walked the window again |
|
||||
| Signatures | 711 s over 439,307 | 29 s over 136,339 | 1.6 ms a signature against 0.2 ms: the pod's CPU, parallel pre-verification on fewer cores |
|
||||
| Virtual | 1,062 s over 36,183 changes | 206 s over 67,221 | 29 ms a change against 3 ms |
|
||||
| Bodies | 16,592 s over 141,699 (117 ms a block, summed over the body workers, 2.8x the wall) | 179 s over 67,221 (2.7 ms) | thread time, not CPU: the body workers queue on the finality state lock while one of them walks a table or verifies a certificate |
|
||||
|
||||
Fix on release-0.3.20-node: `WEIGHT_TABLE_CACHE` = 1,024 tables with oldest-first eviction, never a clear (`FinalityManager::weights_at`); gated by a unit test that fills the cache past the bound and reads the newest table back without a walk (the known-failed case is the old rule, which leaves one table after each clear). Expected on the next fresh join: the tables fall to about one per checkpoint (5,400 tables, about 15 s on the pod), the bodies thread time with them since the lock is held for the walk; the wall is then the bodies' own cost plus the signatures, about 60 to 70 minutes on the pod, approximate until the 0.3.20 canary's line.
|
||||
|
||||
### 7.7.5 The proof oracle at never (ledger N9; the 0.3.19 canary's stall, 7 October 2026)
|
||||
|
||||
The e69e8a39 canary on c18-1 passed its headers proof and stalled in the block stage at block 2,728: the daemon installed the proof oracle whatever `proving_consensus_verify_daa` said, the IBD flow then demanded the proof bytes of every record-carrying block from the syncer, and the serve flow answers from a peer's in-memory pool (records drop at 600 chain blocks, proof bytes live as long as the process), so a block from months ago was served by nobody; 56 asks of pool-1 and the hub, 110 failed IBDs, every peer 0.3.17, which installs no oracle and so never asked. Fix dc141409 on release-0.3.19-node: `ProofOracle::active_from` carries the switch, and the body rule, the IBD fetch and the relay retry apply from it only (known-failed test on the pure gate). The retention store landed as aea0ca5c on release-0.3.19-node (10:04 UK): the proof archive under the exec db dir's proofs/, one file per proof by the carrying block's DAA, the index rebuilt at start, read after the pool by the oracle's held check and the serve flow, dropped below the pruning point in the chain-facts refresh; the rule applies above the pruning point only (the pipeline skips trusted bodies). Known failed first in its unit test (served by a second instance after the pool holds nothing; gone once the pruning point passes the carrier). The harness hook is e5f993d4 (`--prover`, `--join-after`, per-node exec ports on the join harness); the late-join case runs on the testnet lane's fast-time network with provers, the pre-archive binary its known-failed side.
|
||||
|
|
|
|||
Loading…
Reference in a new issue