docs/plans/pool.md 10.5d: the 21:30Z minute on the pool pair as it went (nodes by hand at 21:39Z and 21:42Z, daemons NODE NOT ANSWERING until the 21:50Z restart, the gate read moves to daa 21600)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-07 21:57:13 +00:00
parent 2ea2402ef9
commit c7d002138d

View file

@ -535,7 +535,7 @@ this process's run, and say which is which.
| The sweep | the 0.3.23 node 2720d8d2 moves Devnet 3's digest to ba75bf6f; a one-box-at-a-time sweep split the network (build-1's node at 83eb50cd read 27 "digest mismatch, local ba75bf6f" rejects between 20:19Z and 20:37Z); every Devnet 3 node restarts inside one minute at the fleet's clock, the member nodes among them, by the fleet lane on its boxes; the daemons stay (the pairing is unchanged on 2720d8d2) | a daemon row owed if a daemon's gRPC connection does not survive its node's restart (STATUS must return to "node ok" within a minute of templates) |
| The prepare stall | at the epoch-3 boundary (19:53:14Z) dn3-pool-b's miner wrote the checked pack, sent the prepare, and the worker never said prepared: 54 minutes with every job held, nothing sent, nothing rejected, 0 MH/s, until a hand restart. Two halves: the pack root `packs/dn3-prepare` was relative to the miner's working directory and handed as such to the worker, a separate process that never found it (the node lane's finding and fix, `absolute_pack_root` on class-v5-node-wire); and pool mode had no guard on an unanswered prepare and did not honour `--exit-on-seed-change` at a seed move the way the solo path does | fork 4789fbf7 (on v5-object-0323's tip c8f9b383, for the 0.3.24 line): a seed move in pool mode is logged as SEED CHANGE; a worker that cannot prepare exits 42 for the launcher under the flag; a prepare unanswered for 120 s is re-sent once with the pack written again, unanswered twice it is a fault (exit 42, or the worker replaced and the pair prepared again); tests `a_prepare_that_never_answers_is_resent_once_then_a_fault`, `a_seed_move_in_pool_mode_exits_for_the_launcher_or_hot_swaps`. Operators until then: an absolute `--prepare-packs` path or miner and worker from one directory. Register row MF-15 |
| The two rejected counters | the daemon's STATUS rejected=235,016 beside the member's zero rejected lines in the hour: the daemon's counter is the ledger's life (persisted in the state file across its eras: the 12-job-window era's unknown_job and the 45db7613 era's wrong_hash), the member's is its run | 4a123b89: STATUS prints `rejected_by_code(ledger)=` and `rejected_by_code(this run)=`; `/api/stats` carries `shares.rejected_by_code`, `shares.rejected_by_code_this_run` and a `counters_since` note; `refuse_share` takes the code |
| The daemon across its node's restart | `GrpcClient::connect` leaves the client's reconnect flag off, so the daemon's three node connections (templates and submits, the chain walker, the open pool's seed-block lookup) died with the node and never came back: a daemon running on across the 21:30Z sweep minute (member nodes restart on the 2720d8d2 hands pair, daemons a357c581 untouched) reads NODE NOT ANSWERING until a hand restarts it; the fleet's rule for the minute (restart a daemon that stays there past a minute) is that hand | from the next daemon every node connection is made with `connect_reconnecting` (reconnect on; requests fail while the node is away and succeed once it answers; the template feed re-subscribes beside it); the subscription's own re-subscribe was already there |
| The daemon across its node's restart | `GrpcClient::connect` leaves the client's reconnect flag off, so the daemon's three node connections (templates and submits, the chain walker, the open pool's seed-block lookup) died with the node and never came back: a daemon running on across the 21:30Z sweep minute (member nodes restart on the 2720d8d2 hands pair, daemons a357c581 untouched) reads NODE NOT ANSWERING until a hand restarts it; the fleet's rule for the minute (restart a daemon that stays there past a minute) is that hand. The minute as it went (the fleet lane, UTC): the member nodes did not restart at 21:30:00Z (the puller's "no box-dn3.sh environment" fault on every MINE=0 box) and were moved by hand, pool-a's node synced on 2720d8d2 at 21:39:35Z, pool-b's at 21:42:29Z; the daemons read NODE NOT ANSWERING from their last template (pool-a 21:30:23Z, pool-b 21:42:12Z) and stayed there; the fleet's first restart passed an empty command, the real one landed 21:50:49Z (pool-a) and 21:51:12Z (pool-b), same a357c581 from /proc; after it "NO TEMPLATE YET ... members=0" because the move's kill file had taken the member miners too, so the loops restarted after; the daa 18000 boundary (about 21:53Z) passed with the loops down and MF-15's gate read moves to daa 21600 (about 22:53Z) | from the next daemon every node connection is made with `connect_reconnecting` (reconnect on; requests fail while the node is away and succeed once it answers; the template feed re-subscribes beside it); the subscription's own re-subscribe was already there |
### 10.6 Open after this round