diff --git a/docs/plans/pool.md b/docs/plans/pool.md index 832e00fb2..c303f21e1 100644 --- a/docs/plans/pool.md +++ b/docs/plans/pool.md @@ -535,6 +535,7 @@ this process's run, and say which is which. | The sweep | the 0.3.23 node 2720d8d2 moves Devnet 3's digest to ba75bf6f; a one-box-at-a-time sweep split the network (build-1's node at 83eb50cd read 27 "digest mismatch, local ba75bf6f" rejects between 20:19Z and 20:37Z); every Devnet 3 node restarts inside one minute at the fleet's clock, the member nodes among them, by the fleet lane on its boxes; the daemons stay (the pairing is unchanged on 2720d8d2) | a daemon row owed if a daemon's gRPC connection does not survive its node's restart (STATUS must return to "node ok" within a minute of templates) | | The prepare stall | at the epoch-3 boundary (19:53:14Z) dn3-pool-b's miner wrote the checked pack, sent the prepare, and the worker never said prepared: 54 minutes with every job held, nothing sent, nothing rejected, 0 MH/s, until a hand restart. Two halves: the pack root `packs/dn3-prepare` was relative to the miner's working directory and handed as such to the worker, a separate process that never found it (the node lane's finding and fix, `absolute_pack_root` on class-v5-node-wire); and pool mode had no guard on an unanswered prepare and did not honour `--exit-on-seed-change` at a seed move the way the solo path does | fork 4789fbf7 (on v5-object-0323's tip c8f9b383, for the 0.3.24 line): a seed move in pool mode is logged as SEED CHANGE; a worker that cannot prepare exits 42 for the launcher under the flag; a prepare unanswered for 120 s is re-sent once with the pack written again, unanswered twice it is a fault (exit 42, or the worker replaced and the pair prepared again); tests `a_prepare_that_never_answers_is_resent_once_then_a_fault`, `a_seed_move_in_pool_mode_exits_for_the_launcher_or_hot_swaps`. Operators until then: an absolute `--prepare-packs` path or miner and worker from one directory. Register row MF-15 | | The two rejected counters | the daemon's STATUS rejected=235,016 beside the member's zero rejected lines in the hour: the daemon's counter is the ledger's life (persisted in the state file across its eras: the 12-job-window era's unknown_job and the 45db7613 era's wrong_hash), the member's is its run | 4a123b89: STATUS prints `rejected_by_code(ledger)=` and `rejected_by_code(this run)=`; `/api/stats` carries `shares.rejected_by_code`, `shares.rejected_by_code_this_run` and a `counters_since` note; `refuse_share` takes the code | +| The daemon across its node's restart | `GrpcClient::connect` leaves the client's reconnect flag off, so the daemon's three node connections (templates and submits, the chain walker, the open pool's seed-block lookup) died with the node and never came back: a daemon running on across the 21:30Z sweep minute (member nodes restart on the 2720d8d2 hands pair, daemons a357c581 untouched) reads NODE NOT ANSWERING until a hand restarts it; the fleet's rule for the minute (restart a daemon that stays there past a minute) is that hand | from the next daemon every node connection is made with `connect_reconnecting` (reconnect on; requests fail while the node is away and succeed once it answers; the template feed re-subscribes beside it); the subscription's own re-subscribe was already there | ### 10.6 Open after this round diff --git a/pool/src/main.rs b/pool/src/main.rs index 9595e0644..713bc5e24 100644 --- a/pool/src/main.rs +++ b/pool/src/main.rs @@ -39,6 +39,14 @@ use std::sync::atomic::AtomicU64; use std::sync::{Arc, Mutex}; use std::time::{Duration, Instant}; + +/// A node connection that comes back by itself after the node restarts: `GrpcClient::connect` sets the client's +/// reconnect flag off, so a request after the node's restart failed for ever; with it on, requests fail while the +/// node is away and succeed again once it answers (the template feed's own subscription re-subscribes beside it). +pub async fn connect_reconnecting(url: String) -> kaspa_grpc_client::error::Result { + GrpcClient::connect_with_args(kaspa_rpc_core::notify::mode::NotificationMode::Direct, url, None, true, None, false, None, Default::default()).await +} + #[tokio::main] async fn main() { let args: Vec = std::env::args().skip(1).collect(); @@ -93,10 +101,12 @@ async fn main() { eprintln!("--evm-rpc: {e}"); std::process::exit(1) }); - // nodes + // nodes: every connection reconnects on its own (the gRPC client's reconnect flag, off in `GrpcClient::connect`); + // 7 October 2026, the Devnet 3 sweep minute: the member nodes restart on a new hands pair while the daemons run on, + // and a daemon whose connection died with the old node would read NODE NOT ANSWERING until a hand restarted it let mut clients = Vec::new(); for url in &cfg.nodes { - match tokio::time::timeout(Duration::from_secs(10), GrpcClient::connect(url.clone())).await { + match tokio::time::timeout(Duration::from_secs(10), connect_reconnecting(url.clone())).await { Ok(Ok(c)) => clients.push(Arc::new(c)), Ok(Err(e)) => { eprintln!("node {url}: {e}"); @@ -109,7 +119,7 @@ async fn main() { } } let node = clients.remove(0); - let walker = match tokio::time::timeout(Duration::from_secs(10), GrpcClient::connect(cfg.nodes[0].clone())).await { + let walker = match tokio::time::timeout(Duration::from_secs(10), connect_reconnecting(cfg.nodes[0].clone())).await { Ok(Ok(c)) => Arc::new(c), Ok(Err(e)) => { eprintln!("node {} (second connection): {e}", cfg.nodes[0]); diff --git a/pool/src/node.rs b/pool/src/node.rs index 97df4190a..f14c7e0f1 100644 --- a/pool/src/node.rs +++ b/pool/src/node.rs @@ -248,7 +248,7 @@ async fn issue_job(pool: &Arc, member: &Arc, mut raw: RpcRawBlock, } async fn subscribe(url: &str) -> Result<(GrpcClient, async_channel::Receiver), String> { - let client = tokio::time::timeout(Duration::from_secs(5), GrpcClient::connect(url.to_string())).await.map_err(|_| "connect timed out".to_string())?.map_err(|e| e.to_string())?; + let client = tokio::time::timeout(Duration::from_secs(5), crate::connect_reconnecting(url.to_string())).await.map_err(|_| "connect timed out".to_string())?.map_err(|e| e.to_string())?; tokio::time::timeout(Duration::from_secs(5), client.start_notify(ListenerId::default(), NewBlockTemplateScope {}.into())) .await .map_err(|_| "start_notify timed out".to_string())?