Pool daemon: every node connection reconnects by itself after the node restarts (connect_reconnecting: the gRPC client reconnect flag, off in GrpcClient::connect; 7 October 2026, the Devnet 3 sweep minute with daemons running on across their nodes restart); pool.md 10.5d row
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
1c9c7d7e56
commit
2ea2402ef9
3 changed files with 15 additions and 4 deletions
|
|
@ -535,6 +535,7 @@ this process's run, and say which is which.
|
|||
| The sweep | the 0.3.23 node 2720d8d2 moves Devnet 3's digest to ba75bf6f; a one-box-at-a-time sweep split the network (build-1's node at 83eb50cd read 27 "digest mismatch, local ba75bf6f" rejects between 20:19Z and 20:37Z); every Devnet 3 node restarts inside one minute at the fleet's clock, the member nodes among them, by the fleet lane on its boxes; the daemons stay (the pairing is unchanged on 2720d8d2) | a daemon row owed if a daemon's gRPC connection does not survive its node's restart (STATUS must return to "node ok" within a minute of templates) |
|
||||
| The prepare stall | at the epoch-3 boundary (19:53:14Z) dn3-pool-b's miner wrote the checked pack, sent the prepare, and the worker never said prepared: 54 minutes with every job held, nothing sent, nothing rejected, 0 MH/s, until a hand restart. Two halves: the pack root `packs/dn3-prepare` was relative to the miner's working directory and handed as such to the worker, a separate process that never found it (the node lane's finding and fix, `absolute_pack_root` on class-v5-node-wire); and pool mode had no guard on an unanswered prepare and did not honour `--exit-on-seed-change` at a seed move the way the solo path does | fork 4789fbf7 (on v5-object-0323's tip c8f9b383, for the 0.3.24 line): a seed move in pool mode is logged as SEED CHANGE; a worker that cannot prepare exits 42 for the launcher under the flag; a prepare unanswered for 120 s is re-sent once with the pack written again, unanswered twice it is a fault (exit 42, or the worker replaced and the pair prepared again); tests `a_prepare_that_never_answers_is_resent_once_then_a_fault`, `a_seed_move_in_pool_mode_exits_for_the_launcher_or_hot_swaps`. Operators until then: an absolute `--prepare-packs` path or miner and worker from one directory. Register row MF-15 |
|
||||
| The two rejected counters | the daemon's STATUS rejected=235,016 beside the member's zero rejected lines in the hour: the daemon's counter is the ledger's life (persisted in the state file across its eras: the 12-job-window era's unknown_job and the 45db7613 era's wrong_hash), the member's is its run | 4a123b89: STATUS prints `rejected_by_code(ledger)=` and `rejected_by_code(this run)=`; `/api/stats` carries `shares.rejected_by_code`, `shares.rejected_by_code_this_run` and a `counters_since` note; `refuse_share` takes the code |
|
||||
| The daemon across its node's restart | `GrpcClient::connect` leaves the client's reconnect flag off, so the daemon's three node connections (templates and submits, the chain walker, the open pool's seed-block lookup) died with the node and never came back: a daemon running on across the 21:30Z sweep minute (member nodes restart on the 2720d8d2 hands pair, daemons a357c581 untouched) reads NODE NOT ANSWERING until a hand restarts it; the fleet's rule for the minute (restart a daemon that stays there past a minute) is that hand | from the next daemon every node connection is made with `connect_reconnecting` (reconnect on; requests fail while the node is away and succeed once it answers; the template feed re-subscribes beside it); the subscription's own re-subscribe was already there |
|
||||
|
||||
### 10.6 Open after this round
|
||||
|
||||
|
|
|
|||
|
|
@ -39,6 +39,14 @@ use std::sync::atomic::AtomicU64;
|
|||
use std::sync::{Arc, Mutex};
|
||||
use std::time::{Duration, Instant};
|
||||
|
||||
|
||||
/// A node connection that comes back by itself after the node restarts: `GrpcClient::connect` sets the client's
|
||||
/// reconnect flag off, so a request after the node's restart failed for ever; with it on, requests fail while the
|
||||
/// node is away and succeed again once it answers (the template feed's own subscription re-subscribes beside it).
|
||||
pub async fn connect_reconnecting(url: String) -> kaspa_grpc_client::error::Result<GrpcClient> {
|
||||
GrpcClient::connect_with_args(kaspa_rpc_core::notify::mode::NotificationMode::Direct, url, None, true, None, false, None, Default::default()).await
|
||||
}
|
||||
|
||||
#[tokio::main]
|
||||
async fn main() {
|
||||
let args: Vec<String> = std::env::args().skip(1).collect();
|
||||
|
|
@ -93,10 +101,12 @@ async fn main() {
|
|||
eprintln!("--evm-rpc: {e}");
|
||||
std::process::exit(1)
|
||||
});
|
||||
// nodes
|
||||
// nodes: every connection reconnects on its own (the gRPC client's reconnect flag, off in `GrpcClient::connect`);
|
||||
// 7 October 2026, the Devnet 3 sweep minute: the member nodes restart on a new hands pair while the daemons run on,
|
||||
// and a daemon whose connection died with the old node would read NODE NOT ANSWERING until a hand restarted it
|
||||
let mut clients = Vec::new();
|
||||
for url in &cfg.nodes {
|
||||
match tokio::time::timeout(Duration::from_secs(10), GrpcClient::connect(url.clone())).await {
|
||||
match tokio::time::timeout(Duration::from_secs(10), connect_reconnecting(url.clone())).await {
|
||||
Ok(Ok(c)) => clients.push(Arc::new(c)),
|
||||
Ok(Err(e)) => {
|
||||
eprintln!("node {url}: {e}");
|
||||
|
|
@ -109,7 +119,7 @@ async fn main() {
|
|||
}
|
||||
}
|
||||
let node = clients.remove(0);
|
||||
let walker = match tokio::time::timeout(Duration::from_secs(10), GrpcClient::connect(cfg.nodes[0].clone())).await {
|
||||
let walker = match tokio::time::timeout(Duration::from_secs(10), connect_reconnecting(cfg.nodes[0].clone())).await {
|
||||
Ok(Ok(c)) => Arc::new(c),
|
||||
Ok(Err(e)) => {
|
||||
eprintln!("node {} (second connection): {e}", cfg.nodes[0]);
|
||||
|
|
|
|||
|
|
@ -248,7 +248,7 @@ async fn issue_job(pool: &Arc<Pool>, member: &Arc<Member>, mut raw: RpcRawBlock,
|
|||
}
|
||||
|
||||
async fn subscribe(url: &str) -> Result<(GrpcClient, async_channel::Receiver<Notification>), String> {
|
||||
let client = tokio::time::timeout(Duration::from_secs(5), GrpcClient::connect(url.to_string())).await.map_err(|_| "connect timed out".to_string())?.map_err(|e| e.to_string())?;
|
||||
let client = tokio::time::timeout(Duration::from_secs(5), crate::connect_reconnecting(url.to_string())).await.map_err(|_| "connect timed out".to_string())?.map_err(|e| e.to_string())?;
|
||||
tokio::time::timeout(Duration::from_secs(5), client.start_notify(ListenerId::default(), NewBlockTemplateScope {}.into()))
|
||||
.await
|
||||
.map_err(|_| "start_notify timed out".to_string())?
|
||||
|
|
|
|||
Loading…
Reference in a new issue