diff --git a/docs/plans/pool.md b/docs/plans/pool.md index 4c28d914..732241fb 100644 --- a/docs/plans/pool.md +++ b/docs/plans/pool.md @@ -289,3 +289,40 @@ Consequences per tier: a pool operator needs a box whose memory serves the 1 GiB verifies about 100 shares a second per core at 9.8 ms; at one share per 10 s per member that is 1,000 members per core, so the verifier is not the limit once vardiff holds); a member on any card is unaffected by any of this; the chain walk now costs the node at most 600 `get_block` calls per 5 s on its own connection. + +### 9.3 The re-run's close (06:33:33Z) and the member rows from its logs + +After 2f6c0358 (06:20:34Z to 06:33:33Z, 13 minutes, four members): 207 blocks found, 144 confirmed, 69 orphaned (13 +inherited pending at the swap, 56 on its own jobs: a 27 percent residual at 1 bps), 20,404 shares accepted, 0 rejected, +0 mismatches, 0 refused, every STATUS `node ok`, the gRPC connection churn gone (ids #3 to #5 once each in a minute; the +193 connections in 150 s were under f808b3f3's pruning-point walk), re-subscribe lines 2 (one per daemon start: the +re-subscribe loop was never the churn). Before it, f808b3f3 had worked through the walk by about 06:14Z: 245 blocks, +69 confirmed, 163 orphaned. First CONFIRMED line 06:24:32Z, 4.80 IGN to the finder. Logs: `~/Desktop/fleet/pool-run-0607/`. + +The residual orphans are template latency: the daemon's STATUS reads `last template 1 s ago` and the solo miner on the +same node reads `template_ms=1632`; pool-1's node answers a template request in 1 to 1.6 s although its mining manager +answers a per-member request from its cache by a coinbase rewrite (`modify_block_template`) and its prewarm builds in +0.6 ms. At 1 bps a template that arrives 1.6 s old loses about a quarter of the blocks found on it, pool or solo. That +latency is the node's RPC path, a row for the node lane (measure `get_block_template` round trips on a devnet node under +a miner and a pool; the suspect is the RPC service's `pow_epoch` derivation per request). 82 "RPC request timeout" lines +after the 2f6c0358 swap are the gRPC client's own 5 s request timeout on that path; `--template-timeout-s` cannot +lengthen it. + +From pm-1's log (1.09 GB): 4,342,515 lines are `worker: error N epoch seed mismatch: this worker holds epoch eea66ce7...`, +one per re-queued job, for the 40 s the worker compiled the new epoch's pack (NVRTC 40,327 ms) and again after every +reconnect, because a session respawned the worker (a second process beside the first, the same compile again) and +the member fed jobs for a pair the worker did not hold yet. Vardiff from the same log: shift 10 to 11 (idle easing +while the worker compiled, which consumed the sized first correction), then 11 down to 0 one step per 30 s over 5.5 +minutes at about 190 shares a second, the load that put the share check at 10 to 11 ms on pool-1's cores. + +Fixed in the member (fork commit after b8070476): one worker process per run, jobs held for a pair whose prepare is +pending and never re-prepared once held (`WorkerMemory`, test `jobs_are_held_while_a_pair_is_being_prepared`), worker +error lines summarised (one per class, then a count per 1,000), `set_target` applied to the current work at once. +Fixed in the pool (`vardiff.rs`): the idle easing no longer consumes the sized first correction (test +`an_idle_easing_does_not_consume_the_sized_first_correction`). + +Consequences per tier: a member on a 3070-class card loses 40 s of hash at every epoch roll to the NVRTC compile unless +its pack is prepared ahead (the pool's `seeds` line carries `next_epoch_seed` for that, the member's prepare-ahead is +the next member row); a member on a reconnect now keeps its worker and its compiled pair; a pool operator's verifier +sees the sized correction within 10 s of a member's first shares (one share per 10 s per member from then on) instead of +5 minutes of a 190-share-a-second flood; the 10-member load figure is still owed and now has a clean form to run in. diff --git a/pool/src/vardiff.rs b/pool/src/vardiff.rs index 7811cee1..4f86908d 100644 --- a/pool/src/vardiff.rs +++ b/pool/src/vardiff.rs @@ -79,8 +79,11 @@ impl Vardiff { return None; } if n == 0 { - // nothing yet after an interval: wait up to three intervals, then one easier step - if (now_s - self.started_s) < 3.0 * self.interval_s { + // nothing yet after an interval: wait up to three intervals, then one easier step. The sized + // correction stays owed (first_done stays false): 7 October 2026, pm-1's worker compiled its pack for + // 40 s, the idle easing consumed the first correction, and the measured 190 shares a second then + // walked down one step per 30 s for 5.5 minutes, saturating the verifier + if now_s - self.last_change_s.max(self.started_s) < 3.0 * self.interval_s { return None; } self.shift = (self.shift + 1).min(cap); @@ -92,8 +95,8 @@ impl Vardiff { } else if per_interval < 0.66 { self.shift = (self.shift + steps).min(cap); } + self.first_done = true; } - self.first_done = true; } else { if since_change < self.min_change_s { return None; @@ -226,4 +229,23 @@ mod tests { } assert_eq!(v.retarget(t, 1 << 20), Some(2), "floor at min_shift"); } + + /// 7 October 2026, pm-1: a member idle for four intervals (its worker compiling) gets eased one step, and the + /// first measured rate must still size the jump (190 shares a second at shift 11 is 11 steps away from one per + /// 10 s; the sized step takes 8 at once), not walk down one step per 30 s. + #[test] + fn an_idle_easing_does_not_consume_the_sized_first_correction() { + let t64 = 1u64 << 36; + let mut v = Vardiff::new(10, 0, 20, 10.0, 0.0); + assert_eq!(v.retarget(31.0, t64), Some(11), "idle three intervals: one easier step"); + assert_eq!(v.retarget(40.0, t64), None, "still idle, nothing within the next three intervals"); + // the worker is ready: 190 shares a second for 10 s + let mut t = 40.0; + for _ in 0..1900 { + t += 10.0 / 1900.0; + v.on_share(t); + } + let s = v.retarget(t + 0.1, t64).expect("the first measured rate sizes the jump"); + assert!(s <= 3, "sized first correction from 11 took at most 8 steps at once, got shift {s}"); + } }