release 0.3.20 plan: PC 1's template path, the two app faults, the identity finding withdrawn, the weight-table cache

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-07 10:42:30 +00:00
parent a34600448b
commit 82046e8d69

View file

@ -316,3 +316,10 @@ c18-1 (RTX 3070 pod, wiped datadir, from 08:51:50Z; the Mac's reboot cut the for
**Fleet kit rule (main, 10:3xZ):** one prover identity per box; the kit never ships an identity file; box-prover generates its own at first start from the box's label and keeps it in the registry row; a shipped or duplicated identity is refused at start with a line. Migration on the shared boxes (9e4ba6b0 on three, faa34a1a on one) one at a time, prover only; prover identity keys are not vote keys (no weight, no signal), so the lock rule does not apply.
**Main (10:4xZ): miner-reliability rides 0.3.20's app** with ember-heat and ui-ota (ee09ae8b, docs only, goes with its three); no app-only follow-up and no renumber. One exception: if the reliability lane's watchdog-clock and export-lock changes land small and green before the node line is ready, main decides an app-only 0.3.20 then (the feature node would become 0.3.21). PC 1 today: the Arc held off and the packs-ahead restart on 0.3.19.
## 31. Rows from PC 1 on 0.3.19 (11:3x to 11:4x UK)
- The template path on the 0.3.17 node: on PC 1 (24 identities, 3 cards x 8) every getBlockTemplate times out at 5 s on every identity ("template fetch timed out (5 s) for identity N", 59 + 43 + 28 lines in two minutes), no STATUS line; the 0.3.19 watchdog treats it as the miner's heartbeat ("waiting for the node to answer block templates") so the cards wait rather than fault. The node logs no per-request time on this tree (the timers are 500ddd66, 0.3.20). Mitigation today: identities down to 2 per card by the project lead's tap (no signed job can set identities: a script may not POST /api/cards, and the runner's --cards-off only toggles enabled; a runner option for identities is a small jobrun.rs change, main's word). The fix is the 0.3.20 node on PC 1 first (the memoised tally 500ddd66; fd7de1b4 the off-path refresh is 0.3.21 material unless main moves it).
- Two app faults for the reliability lane's rules, both 0.3.20's app (main): (1) engine.rs:3032, the runner's --stop-miners hold is not released while a following job runs (the resume fires only when job_hold is set and no job holds the miners), so a read-only watch job kept both cards off for three minutes; (2) orphan igneum-miner.exe processes the app no longer tracks (two alive under --stop-miners with their rows at pid 0, one after) hammer the node's template RPC beside the tracked miners, part of why PC 1's template calls ran past 5 s. Sizing of the orphan-kill for an app-only 0.3.20: below.
- The fleet's prover-identity finding was withdrawn (no box shares a key; the "shared" hashes were a carrier block's record list); no migration. The kit rule stands on another ground: `igneum-miner key-hash <label>` is a pure function of the label, so the 0.3.20 kit seeds the identity from the box, keeps the hash on the registry row, refuses a duplicate.
- The node lane's IBD-end read (ca3-v4-node d49c71d0): the weight-table cache (64 tables, cleared past that) made every 2,000-body IBD batch walk the window again (359,709 tables for 5,335 checkpoints, 968 s); fix on release-0.3.20-node: WEIGHT_TABLE_CACHE = 1,024, oldest-first eviction, never a clear, a unit test; bodies 16,592 s is thread time behind the finality state lock, not CPU; the next fresh join on the pod should read 60 to 70 minutes.