igneum/docs/analysis/hashrate-decay-2026-10-03.md
igneum-labs e70e3a58fc Hash-rate decay on the RTX 5090: diagnosis, Metal reproduction, proposed Seeder fix
No per-job growth in any worker or in the miner's memory. The STATUS rates are cumulative averages
(a fast first interval decays by construction), and the miner's Seeder walks the selected chain from
the sink to the epoch start on every memo miss (one getBlock per block, up to 3,600), a gap between
jobs that grew 0.10 s to 0.33 s across epoch 2 on the PC and reset at the epoch boundary while the
difficulty held. Reproduced on the Metal worker (churn on, off, on). One versus eight workers at two
fixed difficulties: 4% and 1% constant cost, no decay. Fix as a unified diff, not applied.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-03 23:31:23 +00:00

33 KiB

Per-identity hash rate "decay" on the RTX 5090: diagnosis and fix

3 October 2026, miner-community-lead. Source data: the uploaded logs of the project lead's PC (node tools/logs.mjs nvidia-DESKTOP-KMCV30N-1-20261003-222331 --all and the other identities, the launcher log igneum-DESKTOP-KMCV30N-20261003-222331), the three serve loops, the miner's worker mode, and five runs of the Metal worker on the Mac against private test nodes (ports 27500 and up, /tmp/igneum-decay-test). Figures from the logs are exact; the two labelled approximate are from memory.

1. Finding in one paragraph

There is no per-job growth in any worker or in the miner's memory. Two separate things produce the picture the project lead saw. First, the STATUS line's two rates are cumulative averages since the miner started (hashes_total / elapsed and hashes_total / gpu_ms_total in mine_worker), so a fast first interval decays as 1/t by construction; nvidia-1 was alone on the card for its first seconds and every later interval ran at a flat 17.8 MH/s wall, while the last-started identity, nvidia-8, shows the mirror image, a cumulative figure that rises. Second, inside an epoch the miner's time between jobs grows with the DAA position: Seeder::seeds_for memoises the epoch seed by (epoch, sink), the sink changes with every block, and every memo miss walks the selected chain from the sink back to the epoch's first block with one getBlock RPC per block (1,100 to 3,600 calls at the end of epoch 2). On the PC the gap between jobs grew from 0.10 s to 0.33 s per 0.8 s job over the 29 minutes of epoch 2, the card's total fell from 197 to about 165 MH/s, and at the 22:57 epoch boundary (walk back to 1) the gap fell to 2% while the difficulty did not move (84.5M to 83.0M). The fix is a per-block memo so a job costs as many getBlock calls as the sink moved (section 6).

2. What the Windows logs actually say

The miner prints hash=A MH/s wall (B MH/s inside jobs) with A = hashes_total / elapsed and B = hashes_total / gpu_ms_total, both since start. Consecutive STATUS lines give the per-interval rates: dH / dt and dH / dG with H = A x t and G = H / B. Every table below is that calculation (rates.py in the bench-log entry).

2.1 The segment the project lead quoted: 22:57 to 23:04 UTC, after the epoch-3 restart (DAA 10,801 on)

nvidia-1, jobs of 2^24 nonces, STATUS every 30 s:

t (s) jobs jobs in interval wall MH/s (cumulative, printed) inside MH/s (cumulative, printed) wall MH/s (this interval) inside MH/s (this interval)
30 40 40 22.18 22.70 22.18 22.70
60 72 32 20.03 20.31 17.84 17.96
91 104 32 19.28 19.51 18.30 17.97
121 136 32 18.89 19.10 17.55 17.86
181 200 32 18.54 18.73 17.76 17.98
241 264 32 18.35 18.57 17.79 18.15
302 328 32 18.23 18.49 17.68 18.31
362 392 32 18.16 18.46 17.59 18.24
392 424 32 18.13 18.45 17.63 18.33

Every interval after the first has exactly 32 jobs and 17.5 to 18.4 MH/s. The printed figure falls only because the first 30 s had 40 jobs: the eight workers printed ready between 22:57:37.3 and 22:57:42.0 (nvidia-1 first, nvidia-8 last, 4.7 s apart), so nvidia-1 had the card to itself or to a few for its first seconds. nvidia-8, started last, prints a cumulative rate that rises: 16.80, 17.29, 17.44, 17.53, 17.58, 17.62, 17.65, 17.66, 17.68, 17.69, 17.70, 17.71 MH/s over the same 7 minutes, with 32 jobs per interval throughout. Eight identities at 17.8 MH/s wall are 142 MH/s on the card, 2% under the 145 MH/s inside jobs. The 197 to 142 MH/s step at 22:57 is the epoch-3 program (a different instruction mix; the bench log records 228 vs 185 MH/s for 104 vs 128 loads on this card), not a decay.

2.2 The 29-minute segment before it: 22:23 to 22:57 UTC, epoch 2 (DAA 8,300 to 10,800)

nvidia-1, every third STATUS line (90 s intervals), from the --all uploads deduplicated by timestamp:

t (s) DAA at the time (node line) walk length if the memo misses (DAA minus 7,200) jobs in 90 s wall MH/s (interval) inside MH/s (interval) wall s per job GPU s per job gap s per job
121 8,474 1,274 44 24.58 28.69 0.683 0.585 0.098
303 8,661 1,461 44 23.98 29.51 0.700 0.569 0.131
485 8,916 1,716 44 23.97 31.02 0.700 0.541 0.159
667 9,157 1,957 42 23.32 34.06 0.720 0.493 0.227
849 9,433 2,233 44 24.16 32.97 0.695 0.509 0.186
1,032 9,645 2,445 42 23.49 33.42 0.714 0.502 0.212
1,215 9,834 2,634 37 21.11 32.01 0.795 0.524 0.271
1,397 10,041 2,841 35 19.94 29.80 0.842 0.563 0.279
1,579 10,292 3,092 37 21.06 33.25 0.797 0.505 0.292
1,762 10,442 3,242 37 20.18 33.55 0.832 0.500 0.332
1,945 10,513 3,313 38 20.63 34.74 0.813 0.483 0.330

Three things move together: the gap between jobs grows 3.4x, the walk length grows 2.6x (3.3x from 22:23), and the GPU time per job falls 17% (the other seven identities idle more, so this one's dispatches wait less). Summed over the eight identities the card delivered 197 MH/s at 22:27 and about 165 MH/s at 22:55 (the launcher's cumulative total shows 197.6 falling to 183.4). None of this is the GPU: the inside-jobs rate rose. At 22:57 the epoch changed, the launcher rebuilt the workers, the walk length went back to 1, and the gap went back to 2% (table 2.1) while the difficulty stayed at 83 to 85M. Difficulty is therefore not the driver: it rose 57M to 85M across epoch 2 and the gap tracked the DAA position, and it did not fall at 22:57 when the gap did. Jobs are a fixed 2^24 nonces (job_nonces: 1 << 24 in the miner), so a job does not get longer when blocks get rarer; a harder target only means fewer found lines, each of which costs the miner one CPU hash to re-check (0.1 per job on the PC, nothing).

Time-slicing of eight contexts is real but constant: with all eight busy, each gets one eighth of the card and the sum is the card's rate (142 MH/s in table 2.1, flat for 7 minutes). The only way the sum falls is the card going idle, which is what a growing CPU-side gap in every worker does.

2.3 The 22:03 and 22:09 runs

22:03 (four CUDA identities): 87 to 104 MH/s inside jobs but 3.2 to 4.1 MH/s wall and template_age of 6 to 7 s: the miner spent 6 s between 0.17 s jobs. 22:09 (eight identities, epoch 1, DAA 5,659 on): 24.2 to 25.1 MH/s wall per identity for 3 minutes with the gap at 4 to 8%. Both are consistent with the walk (epoch 1 started at DAA 3,600, so the walk was 2,000 blocks at 22:03 with a node still syncing).

3. Code audit: what is allocated per job, and what is freed

3.1 proto-cuda/host.cu, runServe (HEAD, the binary the project lead ran, built by the launcher at 22:09:57)

Allocation When Size Freed
gCache (cudaMalloc in setupCache) once 256 MiB at exit
dDs dataset once 2^28 words = 1 GiB at exit
dOut once 2^22 x 8 = 32 MiB at exit
hOut (std::vector<uint64_t>) once 32 MiB host at exit
line, f (8 std::string), prehash, epochSeed, daySeed vectors per job under 1 KiB end of the loop iteration (scope)
IgneumInitWords, 49-byte buffer per dispatch stack immediately
cudaMemcpy device to host per dispatch 32 MiB, 4 per job not an allocation; 180 MB/s per identity, 1.4 GB/s for eight, flat
CPU scan of hOut per dispatch 4M compares flat

No cudaMalloc, no container growth, no re-upload of the dataset, no module reload per job. The working-tree version (the hot-swap agent's) adds CudaPair (at most two resident, the old one released after the first job on the new pair, releasePair frees dataset, cache and both modules) and PrepareTask (deleted after the load); still nothing per job. cudaDeviceSynchronize per dispatch under the default cudaDeviceScheduleAuto spins the host thread when the process holds fewer contexts than the machine has cores, which is always true here (one context per process): that is the 6.2% CPU per worker the project lead saw (one of 16 threads). The hot-swap working tree sets cudaSetDeviceFlags(cudaDeviceScheduleBlockingSync) before the context is created (host.cu, main), which is the right call and the right place; the thread then sleeps on the dispatch. It does not change the hash rate directly, but eight spinning threads plus eight OpenCL workers plus the node plus sixteen miner processes contend for the PC's cores, and the walk is made of sixteen thousand small RPCs per minute that need those cores, so it shortens the gap too.

3.2 proto-opencl/host.c, runServe

Once: dOut (32 MiB), dInit (32 bytes), hOut, the dataset and cache (one ServePair, two at most during a prepare). Per dispatch: a 32-byte clEnqueueWriteBuffer, five clSetKernelArg, one launch1D whose event is released, clFinish (HEAD) or clWaitForEvents (working tree), a blocking clEnqueueReadBuffer of chunk x 8 bytes. Per job: stack buffers. Nothing grows. clFinish on the AMD runtime can spin a core (the working tree's comment says so and switches to the event wait, which some runtimes still spin on; only HWiNFO or Task Manager on the PC can tell for this driver). The AMD identities run 0.1 to 0.2 MH/s each with template_age of 15 to 21 s: a 2^24 job takes 15 s on the integrated GPU, so they contribute nothing to the hash rate and should probably not run at all while the 5090 is mining.

3.3 proto-metal/main.swift, runServe

Once: outBuf (32 MiB shared). Per pair: one compiled pipeline and one 1 GiB dataset buffer in ServeStore (at most two of each; prune after the first job on a new pair). Per dispatch: one command buffer and encoder, released by ARC when the loop iteration ends; waitUntilCompleted blocks without spinning. Per job: the split strings and unhex arrays. The Mac runs below confirm the worker's RSS is flat within every run (46 to 58 MiB, the figure depending on the program).

3.4 vendor/igneum-node/igneum/miner/src/main.rs, mine_worker

Per job: one job line String, a Vec<&str> per worker line, raw.clone() for a submit, and seeder.seeds_for. State that persists: hashes_total, gpu_ms_total, jobs, found, extra (counters); Seeder::memo (HashMap<(u64, Hash), Hash>, cleared at 10,000 entries, under 1 MiB); Voter::proposed_at (pruned to 200 entries) and Voter::lock_latencies_ms (one f64 per lock seen, one every 30 s, unbounded but 23 KiB a day; the Windows build predates the voter). No log buffering: every line is written as printed. The one per-job cost that grows is time, not memory: seeds_for.

// main.rs, Seeder::seeds_for (HEAD and the Windows build 745d41ef alike)
let sink = client.get_block_dag_info().await.expect("dag info").sink;      // one RPC per job
if let Some(seed) = self.memo.get(&(epoch, sink)) { return ... }           // hit only while the sink is unchanged
let epoch_start = epoch * POW_EPOCH_BLOCKS;
let mut cur = sink;
let seed = loop {                                                          // miss: one getBlock per block
    if cur == self.genesis { break self.genesis; }                         //   from the sink down to the
    let b = client.get_block(cur, false).await.expect("get block");        //   epoch's first block
    if b.header.daa_score < epoch_start { break cur; }
    cur = b.verbose_data.expect("verbose data").selected_parent_hash;
};

On the PC the sink changed 2.2 times a second (the node line: "network 2.18 blocks/s") against 11 jobs a second across the eight CUDA identities, so most jobs hit the memo and roughly one in five missed; at 3,300 blocks per miss and the 0.33 s average gap, a getBlock round trip cost about 0.5 ms (approximate, derived, not measured on the PC). The node pays for every one of those calls too (verbose data on each).

4. Mac reproduction with the Metal worker

Machine: Apple M5 Max, 64 GiB, the HEAD proto-metal/main.swift built into the scratchpad (swiftc -O, byte-identical in size to proto-metal/igneum-bench), the difficulty worktree's igneumd and igneum-miner (the pair proven on override-genesis networks; the Seeder and the worker protocol are the same code as HEAD), everything at nice -n 19. The Mac was carrying other agents' builds (load average 40 to 147 during R1, 4 to 25 later) and, during R1, another agent's Metal worker on the same GPU, so absolute rates are not comparable between runs; the per-interval shape within a run is what each run is for.

4.1 R1, control: one identity, epoch 0 (no walk), devnet genesis bits 0x1d100000, 308 s

t (s) jobs in interval wall MH/s (printed, cumulative) wall MH/s (interval) inside MH/s (interval) gap
31 46 25.18 25.18 25.78 2.3%
62 26 19.58 13.96 14.02 0.4%
92 25 17.69 13.64 13.96 2.3%
123 28 17.07 15.29 15.25 0%
153 26 16.52 14.07 14.39 2.2%
184 26 16.11 14.07 14.30 1.6%
215 26 15.82 14.12 14.27 1.0%
246 26 15.60 14.11 14.24 0.9%
277 26 15.43 14.06 14.33 1.9%
308 26 15.31 14.40 14.42 0.1%

Another agent's Metal worker started on the GPU 25 s into this run. The printed cumulative rate then "decays" from 25.2 to 15.3 over five minutes while every interval after the first sits at 14.0 to 14.4 MH/s: the same shape as nvidia-1 in table 2.1, from the same arithmetic. The gap is 0 to 2% throughout (epoch 0: seeds_for returns genesis without a walk). Worker RSS 56.8 to 56.9 MiB, miner RSS 268 MiB, flat. The run ended at 308 s when every process of this session was SIGTERMed at 22:21:37 UTC (the credits outage), and was not repeated because R2 to R6 cover the same ground.

4.2 R2, the walk reproduced: one identity in epoch 1, sink churn on, off, on (900 s)

Recipe: a fresh skip_proof_of_work node (difficulty worktree igneumd, override {"skip_proof_of_work": true, "genesis_bits": 487587840}), pumped to DAA 4,000 through getBlockTemplate + submitBlock with timestamps one second apart ending at now (so the difficulty rule saw 1 block/s and held at 76.8M, 0.11 founds per job, the PC's regime); then one Metal identity for 900 s while a pump added one block a second for 300 s (every job sees a new sink, so every job walks), nothing for 300 s (the sink moves only on the identity's own blocks, one per 8 s), and one block a second again for 300 s. Walk length when a job misses: DAA minus 3,600, from 400 at the start to about 1,000 at the end. Load average 2 to 9 for the first 660 s, 12 to 14 after (another agent's build).

t (s) jobs in interval wall MH/s (printed, cumulative) inside MH/s (printed, cumulative) wall MH/s (interval) inside MH/s (interval) gap template age s
30 44 24.27 26.94 24.27 26.94 10% 0.72
60 43 24.13 27.08 23.94 27.22 12% 0.73
91 43 23.99 27.14 24.18 27.26 11% 0.74
121 43 23.89 27.15 23.14 27.18 15% 0.62
152 42 23.77 27.17 23.84 27.25 13% 0.62
182 41 23.60 27.18 22.60 27.23 17% 0.62
212 41 23.46 27.18 22.26 27.18 18% 0.76
243 41 23.36 27.17 23.19 27.10 14% 0.79
273 42 23.37 27.17 23.41 27.17 14% 0.74
303 43 23.40 27.17 23.31 27.17 14% 0.79
334 48 23.68 27.14 26.93 26.88 -0% 0.62
364 48 23.92 27.13 26.32 27.03 3% 0.61
394 49 24.16 27.16 26.65 27.49 3% 0.61
425 48 24.34 27.19 27.32 27.54 1% 0.72
455 48 24.49 27.22 26.50 27.61 4% 0.72
485 47 24.59 27.24 25.87 27.53 6% 0.62
516 48 24.70 27.26 26.92 27.55 2% 0.61
546 47 24.79 27.28 26.33 27.61 5% 0.61
576 48 24.87 27.30 25.87 27.65 6% 0.61
606 45 24.87 27.31 24.63 27.50 10% 0.73
636 43 24.83 27.32 23.96 27.53 13% 0.74
667 43 24.77 27.33 23.82 27.55 14% 0.74
697 39 24.65 27.34 21.98 27.59 20% 0.84
727 33 24.38 27.34 18.00 27.34 34% 0.85
758 37 24.22 27.35 20.82 27.63 25% 0.86
788 36 24.06 27.35 19.95 27.35 27% 0.98
818 33 23.85 27.36 18.24 27.71 34% 0.87
848 37 23.74 27.36 20.61 27.36 25% 0.87
879 36 23.60 27.37 20.18 27.70 27% 0.89

The inside-jobs rate never moves (27.0 to 27.7 MH/s in every one of the 29 intervals). The wall rate is 22.3 to 24.3 MH/s with churn (gap 12 to 18% at a 400 to 700 walk), 25.9 to 27.3 MH/s without it (gap 0 to 4%), and 18.0 to 21.0 MH/s with churn again at a 700 to 1,000 walk (gap 25 to 35%). The miner process's own CPU went from 0 to 1% in the quiet phase to 11 to 21% in the last phase (the walk is the miner parsing a thousand getBlock answers per job); the worker's stayed at 0.4 to 0.9% and its RSS at 56 to 57 MiB. This is the shared miner logic, not CUDA, not the card, not the node's template path (template_age moves with the gap because it is measured from the template fetch, which precedes the walk).

5. The time-slicing hypothesis: one worker versus eight, at two fixed difficulties

the project lead's Task Manager reading (GPU memory flat at 14.7 GB, the card 99% busy, each CUDA worker at 6.2% CPU) and the hypothesis that eight contexts time-slicing one card with "longer jobs as blocks get rarer" explain the decay. A job is a fixed 2^24 nonces, so its length does not depend on the target, but the four runs below test the hypothesis as stated: one Metal worker and eight, each at a fixed low difficulty (2^25, 0.25 founds per job) and a fixed high one (2^31, 0.004 founds per job), 10 minutes each, on fresh networks under Kaspa's sampled rule, which holds the genesis bits for the first 600 blocks (no run made more than 242). Epoch 0 throughout, so no walk. The 6.2% CPU per CUDA worker is cudaDeviceSynchronize spinning (section 3.1); the Metal worker sleeps in waitUntilCompleted, so its CPU is the per-dispatch scan only.

5.1 R3, one worker, difficulty 2^25 (33.5M), 600 s: 30.51 MH/s wall, 30.72 inside, 1,091 jobs, 224 blocks

t (s) jobs in interval wall MH/s (printed, cumulative) inside MH/s (printed, cumulative) wall MH/s (interval) inside MH/s (interval) gap template age s
30 56 30.95 31.18 30.95 31.18 1% 0.54
91 56 31.14 31.25 32.09 31.23 -3% 0.55
151 56 31.10 31.18 31.19 31.26 0% 0.53
212 56 31.09 31.17 31.47 30.94 -2% 0.56
272 55 30.92 31.01 29.88 30.52 2% 0.56
333 53 30.60 30.80 28.73 29.77 4% 0.55
394 55 30.44 30.69 30.03 30.57 2% 0.55
455 56 30.48 30.71 31.42 30.85 -2% 0.54
515 56 30.53 30.74 30.40 30.90 2% 0.58
576 55 30.50 30.71 30.79 30.53 -1% 0.54

5.2 R4, eight workers, difficulty 2^25, 600 s: 29.38 MH/s wall summed, 29.45 inside summed, 1,053 jobs, 242 blocks

Identity 1 (3.52 MH/s wall, 3.53 inside) and identity 8 (3.46 and 3.47); the eight ran 3.32 to 4.38 MH/s each:

t (s) jobs in interval wall MH/s (printed, cumulative) inside MH/s (printed, cumulative) wall MH/s (interval) inside MH/s (interval) gap template age s
34 7 3.42 3.44 3.42 3.44 1% 4.69
134 7 3.51 3.52 3.58 3.58 0% 4.54
233 7 3.52 3.53 3.49 3.53 1% 4.64
334 7 3.52 3.52 3.58 3.52 -2% 4.85
434 7 3.52 3.52 3.64 3.52 -3% 4.62
534 7 3.52 3.53 3.50 3.69 5% 4.38
600 7 3.52 3.53 3.51 3.53 1% 4.76
t (s) jobs in interval wall MH/s (printed, cumulative) inside MH/s (printed, cumulative) wall MH/s (interval) inside MH/s (interval) gap template age s
34 7 3.41 3.43 3.41 3.43 1% 4.99
136 7 3.46 3.46 3.45 3.43 -1% 4.72
238 7 3.45 3.46 3.36 3.46 3% 4.83
339 7 3.46 3.47 3.38 3.57 5% 4.92
441 7 3.46 3.46 3.53 3.46 -2% 4.71
542 7 3.47 3.47 3.57 3.47 -3% 5.11
577 7 3.46 3.47 3.33 3.47 4% 4.90

Eight contexts cost 4% of the card against one (29.4 vs 30.5 MH/s) and the cost is constant: 7 jobs per 100 s per identity in every interval, gap 0 to 1%, no identity's rate moves. Each worker used 0.0 to 0.6% CPU and 46 to 57 MiB; each miner 0% and 268 MiB. A job takes 4.4 to 5.1 s wall per identity (eight times R3's 0.55 s) because the card is shared, which is time-slicing doing exactly what it should.

5.3 R5, one worker, difficulty 2^31 (2.15G), 600 s: 37.01 MH/s wall, 37.75 inside, 1,324 jobs, 4 blocks

t (s) jobs in interval wall MH/s (printed, cumulative) inside MH/s (printed, cumulative) wall MH/s (interval) inside MH/s (interval) gap template age s
30 65 36.18 36.87 36.18 36.87 2% 0.45
90 70 37.50 37.89 38.61 38.96 1% 0.43
151 67 37.67 38.05 36.76 37.73 3% 0.44
212 66 37.34 37.89 37.22 37.53 1% 0.45
272 68 37.43 37.96 37.30 38.12 2% 0.45
332 66 37.29 37.89 36.44 37.59 3% 0.45
393 65 37.14 37.83 36.07 37.46 4% 0.45
454 66 37.06 37.77 37.27 37.49 1% 0.44
514 67 37.06 37.77 36.70 37.77 3% 0.46
575 66 37.02 37.75 37.59 37.57 -0% 0.44

Blocks are 50 times rarer than in R3 and the job is the same 0.45 s; the rate is flat for ten minutes with a 1 to 3% gap. (R5's network has a different genesis hash from R3's, so a different epoch-0 program: 112 loads per hash against R3's; the 37 versus 30.5 MH/s is the program, as the bench log's per-seed spread predicts.)

5.4 R6, eight workers, difficulty 2^31, 600 s

36.67 MH/s wall summed, 36.75 inside summed, 1,314 jobs, 5 blocks; the eight ran 4.29 to 5.55 MH/s each. Identity 1 (4.62 MH/s wall, 4.63 inside):

t (s) jobs in interval wall MH/s (printed, cumulative) inside MH/s (printed, cumulative) wall MH/s (interval) inside MH/s (interval) gap template age s
32 8 4.24 4.27 4.24 4.27 1% 4.11
130 9 4.53 4.54 4.70 4.69 -0% 3.46
227 9 4.58 4.59 4.60 4.59 -0% 3.65
325 9 4.60 4.62 4.80 4.71 -2% 3.74
422 9 4.62 4.63 4.92 4.75 -4% 3.44
520 9 4.61 4.62 4.52 4.62 2% 3.53
585 9 4.62 4.63 4.73 4.63 -2% 3.56

Eight contexts cost 1% against one (36.7 vs 37.0 MH/s), 9 jobs per 100 s per identity in every interval, the gap within the noise of 30-second intervals holding 2 to 3 jobs (plus or minus 4%), CPU 0.0 to 0.6% per worker and 0% per miner, RSS flat (46 to 57 MiB per worker, 268 MiB per miner). Load average 17 to 33 during this run (other agents' builds), which did not move the rate.

5.5 Verdict on the hypothesis

Rarer blocks do not lengthen jobs and do not lower the rate (R3 against R5, one worker; R4 against R6, eight). Eight workers on one card lose a constant 4% to context switching and nothing over time (R4, R6). What the Task Manager reading does show is the CPU side: eight threads spinning in cudaDeviceSynchronize, which the hot-swap branch's cudaDeviceScheduleBlockingSync removes, and which matters because the sixteen miner processes and the node need those cores for the seed walk until the walk is fixed. The difficulty rising through the evening (9M to 83M) tracked the DAA position in the epoch, which is what the gap tracks; the epoch boundary at 22:57, where the difficulty held and the gap collapsed, separates the two.

6. The fix, as a unified diff (not applied)

Memoise per block instead of per sink. A block's epoch seed is a function of the block alone (its selected parent chain is fixed), so the walk from a new sink can stop at the first block already seen: a job then costs as many getBlock calls as the sink advanced since the last job (one or two at 2 blocks/s) instead of the DAA position in the epoch. The map is per epoch and the previous epoch's map is kept for the boundary. The second hunk prints per-interval rates in STATUS next to the cumulative ones and the number of seed-walk calls, so the launcher and the logs show what the card is doing now. Against vendor/igneum-node/igneum/miner/src/main.rs at the working tree of 3 October 2026 (git apply --check passes, and cargo check -p igneum-miner on a scratch worktree with the patch applied finishes with no new warnings; also saved as docs/analysis/hashrate-decay-2026-10-03.patch).

--- a/igneum/miner/src/main.rs
+++ b/igneum/miner/src/main.rs
@@ -366,8 +366,15 @@
 /// selected parent, which is the selected parent of any block built on the current template).
 struct Seeder {
     genesis: Hash,
-    /// (epoch index, sink the walk started from) -> seed
-    memo: HashMap<(u64, Hash), Hash>,
+    /// epoch index -> (selected-chain block -> that epoch's seed) for every block a walk has visited. A block's seed
+    /// is a function of the block alone (its selected parent chain is fixed), so a walk from a new sink stops at the
+    /// first block already in the map: a job costs as many `get_block` calls as the sink advanced, not the DAA
+    /// position in the epoch (the old `(epoch, sink)` key missed on every new block and walked back to the epoch's
+    /// first block, up to 3,600 calls, measured on the RTX 5090 on 3 October 2026 as a gap between jobs that grew
+    /// from 0.10 s to 0.33 s across epoch 2; docs/analysis/hashrate-decay-2026-10-03.md).
+    memo: HashMap<u64, HashMap<Hash, Hash>>,
+    /// `get_block` calls made by seed walks, for the STATUS line
+    walk_calls: u64,
 }
 
 impl Seeder {
@@ -392,7 +399,7 @@
         if genesis != DEVNET_PARAMS.genesis.hash {
             println!("{} node genesis {} differs from the shared devnet genesis: a private test network", now(), genesis);
         }
-        Self { genesis, memo: HashMap::new() }
+        Self { genesis, memo: HashMap::new(), walk_calls: 0 }
     }
 
     async fn seeds_for(&mut self, client: &GrpcClient, header: &Header) -> EpochSeeds {
@@ -402,25 +409,32 @@
             return EpochSeeds { epoch: self.genesis, day };
         }
         let sink = client.get_block_dag_info().await.expect("dag info").sink;
-        if let Some(seed) = self.memo.get(&(epoch, sink)) {
-            return EpochSeeds { epoch: *seed, day };
-        }
+        // This epoch's map and the previous one stay (a template can still land in the previous epoch at the
+        // boundary); older ones go. A map holds at most one entry per selected-chain block of its epoch (3,600).
+        self.memo.retain(|e, _| *e + 1 >= epoch);
+        let known = self.memo.entry(epoch).or_default();
         let epoch_start = epoch * POW_EPOCH_BLOCKS;
         let mut cur = sink;
+        let mut path = Vec::new();
         let seed = loop {
+            if let Some(seed) = known.get(&cur) {
+                break *seed;
+            }
             if cur == self.genesis {
                 break self.genesis;
             }
             let b = client.get_block(cur, false).await.expect("get block");
+            self.walk_calls += 1;
             if b.header.daa_score < epoch_start {
                 break cur;
             }
+            path.push(cur);
             cur = b.verbose_data.expect("verbose data").selected_parent_hash;
         };
-        if self.memo.len() > 10_000 {
-            self.memo.clear();
+        for h in path {
+            known.insert(h, seed);
         }
-        self.memo.insert((epoch, sink), seed);
+        known.insert(sink, seed);
         EpochSeeds { epoch: seed, day }
     }
 }
@@ -707,6 +721,9 @@
     let mut gpu_ms_total = 0.0f64;
     let mut jobs = 0u64;
     let mut last_report = Instant::now();
+    // The previous STATUS line's counters, for the per-interval rates (the cumulative ones decay by construction
+    // after a fast first interval, which read as a hash-rate decay on 3 October 2026)
+    let (mut last_hashes, mut last_gpu_ms, mut last_jobs, mut last_walk) = (0u64, 0.0f64, 0u64, 0u64);
     let mut last_seeds: Option<EpochSeeds> = None;
     let mut rng = rand::thread_rng();
     while start.elapsed() < Duration::from_secs(o.secs) {
@@ -813,8 +830,9 @@
         }
         if last_report.elapsed() > Duration::from_secs(o.status_secs) {
             let el = start.elapsed().as_secs_f64();
+            let (dt, dh, dg) = (last_report.elapsed().as_secs_f64(), (hashes_total - last_hashes) as f64, gpu_ms_total - last_gpu_ms);
             println!(
-                "{} STATUS '{}' [worker]: {:.0}s jobs={} accepted={} rejected={} mismatched={} extra={} rate={:.2} blocks/s hash={:.2} MH/s wall ({:.2} MH/s inside jobs) template_age={:.2}s synced={}",
+                "{} STATUS '{}' [worker]: {:.0}s jobs={} accepted={} rejected={} mismatched={} extra={} rate={:.2} blocks/s hash={:.2} MH/s wall ({:.2} MH/s inside jobs) now={:.2} MH/s wall ({:.2} MH/s inside jobs, {} jobs, seed walk {} calls) template_age={:.2}s synced={}",
                 now(),
                 o.vote_label,
                 el,
@@ -826,9 +844,14 @@
                 found as f64 / el,
                 hashes_total as f64 / el / 1e6,
                 if gpu_ms_total > 0.0 { hashes_total as f64 / gpu_ms_total / 1e3 } else { 0.0 },
+                if dt > 0.0 { dh / dt / 1e6 } else { 0.0 },
+                if dg > 0.0 { dh / dg / 1e3 } else { 0.0 },
+                jobs - last_jobs,
+                seeder.walk_calls - last_walk,
                 template_at.elapsed().as_secs_f64(),
                 synced
             );
+            (last_hashes, last_gpu_ms, last_jobs, last_walk) = (hashes_total, gpu_ms_total, jobs, seeder.walk_calls);
             last_report = Instant::now();
         }
     }

The test that proves it

  1. Unit: Seeder::seeds_for with a mocked client that counts get_block calls: after one job at DAA 3600 + n, a second job whose sink is one block ahead must make exactly 1 get_block call (today: n + 1), and a job on a sink from a different tip of the same height must walk only to the common ancestor.
  2. Network: the R2 recipe above (a skip_proof_of_work node pumped to DAA 4,000 with one-second timestamps, one Metal identity, pumped blocks at 1/s for 300 s, none for 300 s, 1/s for 300 s). Before the fix the wall-versus-inside gap and walk= in STATUS move with the churn phases; after the fix walk per job is at most the blocks the sink advanced (1 to 2) in every phase and the gap stays within 2%. On the PC the acceptance figure is table 2.2 flattening: the gap per job no better than 0.10 s at DAA 10,800 (3,600 into an epoch) with eight identities.

7. Two things for the project lead to check on the PC

  1. Dedicated GPU memory over time. Task Manager, Performance, GPU 0, "Dedicated GPU memory usage", or nvidia-smi --query-gpu=timestamp,memory.used,utilization.gpu,temperature.gpu,clocks.sm,clocks.mem,power.draw,clocks_throttle_reasons.active --format=csv -l 10 > gpu.csv. Expected: flat from the moment the eighth worker prints ready (each CUDA worker holds 1 GiB dataset + 256 MiB cache + 32 MiB output + the context, about 1.6 GiB; eight of them 13 to 15 GiB of the 32 GiB, which matches the flat 14.7 GB the project lead read). A line that climbs while the hash rate falls would mean a leak in the worker; none exists in the code and the Mac RSS traces are flat.
  2. GDDR7 memory junction temperature and the throttle reasons. HWiNFO64, Sensors, under the GPU: "GPU Memory Junction Temperature", "GPU Thermal Limit", "GPU Power Limit", "GPU Reliability Voltage Limit" (each "Yes" or "No"), "GPU Effective Clock" and "GPU Memory Clock"; or the clocks_throttle_reasons.active column above (0x0000000000000004 is SW power cap, 0x0000000000000020 SW thermal slowdown, 0x0000000000000040 HW thermal slowdown, 0x0000000000000080 HW power brake). Approximate thresholds, from memory: the core starts to pull clocks around 83 C (the project lead's 55 to 59 C is far below it); GDDR6X on the previous generations throttles from about 95 C junction and the hard limit is 105 C; GDDR7 figures are not published, so treat anything above 90 C junction as the zone to watch and a "Yes" on any limit row as the signal. The decisive sign of throttling is "GPU Effective Clock" falling while utilization stays at 99%; in the logs it would appear as the inside-jobs per-interval rate falling, which it never does (it rose). A flat line on both rows means the card was not the cause, which is what the logs already say.