No per-job growth in any worker or in the miner's memory. The STATUS rates are cumulative averages (a fast first interval decays by construction), and the miner's Seeder walks the selected chain from the sink to the epoch start on every memo miss (one getBlock per block, up to 3,600), a gap between jobs that grew 0.10 s to 0.33 s across epoch 2 on the PC and reset at the epoch boundary while the difficulty held. Reproduced on the Metal worker (churn on, off, on). One versus eight workers at two fixed difficulties: 4% and 1% constant cost, no decay. Fix as a unified diff, not applied. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
33 KiB
Per-identity hash rate "decay" on the RTX 5090: diagnosis and fix
3 October 2026, miner-community-lead. Source data: the uploaded logs of the project lead's PC (node tools/logs.mjs nvidia-DESKTOP-KMCV30N-1-20261003-222331 --all and the other identities, the launcher log
igneum-DESKTOP-KMCV30N-20261003-222331), the three serve loops, the miner's worker mode, and five runs of the
Metal worker on the Mac against private test nodes (ports 27500 and up, /tmp/igneum-decay-test). Figures
from the logs are exact; the two labelled approximate are from memory.
1. Finding in one paragraph
There is no per-job growth in any worker or in the miner's memory. Two separate things produce the picture
the project lead saw. First, the STATUS line's two rates are cumulative averages since the miner started
(hashes_total / elapsed and hashes_total / gpu_ms_total in mine_worker), so a fast first interval
decays as 1/t by construction; nvidia-1 was alone on the card for its first seconds and every later interval
ran at a flat 17.8 MH/s wall, while the last-started identity, nvidia-8, shows the mirror image, a cumulative
figure that rises. Second, inside an epoch the miner's time between jobs grows with the DAA position:
Seeder::seeds_for memoises the epoch seed by (epoch, sink), the sink changes with every block, and every
memo miss walks the selected chain from the sink back to the epoch's first block with one getBlock RPC per
block (1,100 to 3,600 calls at the end of epoch 2). On the PC the gap between jobs grew from 0.10 s to
0.33 s per 0.8 s job over the 29 minutes of epoch 2, the card's total fell from 197 to about 165 MH/s, and
at the 22:57 epoch boundary (walk back to 1) the gap fell to 2% while the difficulty did not move (84.5M to
83.0M). The fix is a per-block memo so a job costs as many getBlock calls as the sink moved (section 6).
2. What the Windows logs actually say
The miner prints hash=A MH/s wall (B MH/s inside jobs) with A = hashes_total / elapsed and
B = hashes_total / gpu_ms_total, both since start. Consecutive STATUS lines give the per-interval rates:
dH / dt and dH / dG with H = A x t and G = H / B. Every table below is that calculation
(rates.py in the bench-log entry).
2.1 The segment the project lead quoted: 22:57 to 23:04 UTC, after the epoch-3 restart (DAA 10,801 on)
nvidia-1, jobs of 2^24 nonces, STATUS every 30 s:
| t (s) | jobs | jobs in interval | wall MH/s (cumulative, printed) | inside MH/s (cumulative, printed) | wall MH/s (this interval) | inside MH/s (this interval) |
|---|---|---|---|---|---|---|
| 30 | 40 | 40 | 22.18 | 22.70 | 22.18 | 22.70 |
| 60 | 72 | 32 | 20.03 | 20.31 | 17.84 | 17.96 |
| 91 | 104 | 32 | 19.28 | 19.51 | 18.30 | 17.97 |
| 121 | 136 | 32 | 18.89 | 19.10 | 17.55 | 17.86 |
| 181 | 200 | 32 | 18.54 | 18.73 | 17.76 | 17.98 |
| 241 | 264 | 32 | 18.35 | 18.57 | 17.79 | 18.15 |
| 302 | 328 | 32 | 18.23 | 18.49 | 17.68 | 18.31 |
| 362 | 392 | 32 | 18.16 | 18.46 | 17.59 | 18.24 |
| 392 | 424 | 32 | 18.13 | 18.45 | 17.63 | 18.33 |
Every interval after the first has exactly 32 jobs and 17.5 to 18.4 MH/s. The printed figure falls only
because the first 30 s had 40 jobs: the eight workers printed ready between 22:57:37.3 and 22:57:42.0
(nvidia-1 first, nvidia-8 last, 4.7 s apart), so nvidia-1 had the card to itself or to a few for its first
seconds. nvidia-8, started last, prints a cumulative rate that rises: 16.80, 17.29, 17.44, 17.53, 17.58,
17.62, 17.65, 17.66, 17.68, 17.69, 17.70, 17.71 MH/s over the same 7 minutes, with 32 jobs per interval
throughout. Eight identities at 17.8 MH/s wall are 142 MH/s on the card, 2% under the 145 MH/s inside jobs.
The 197 to 142 MH/s step at 22:57 is the epoch-3 program (a different instruction mix; the bench log records
228 vs 185 MH/s for 104 vs 128 loads on this card), not a decay.
2.2 The 29-minute segment before it: 22:23 to 22:57 UTC, epoch 2 (DAA 8,300 to 10,800)
nvidia-1, every third STATUS line (90 s intervals), from the --all uploads deduplicated by timestamp:
| t (s) | DAA at the time (node line) | walk length if the memo misses (DAA minus 7,200) | jobs in 90 s | wall MH/s (interval) | inside MH/s (interval) | wall s per job | GPU s per job | gap s per job |
|---|---|---|---|---|---|---|---|---|
| 121 | 8,474 | 1,274 | 44 | 24.58 | 28.69 | 0.683 | 0.585 | 0.098 |
| 303 | 8,661 | 1,461 | 44 | 23.98 | 29.51 | 0.700 | 0.569 | 0.131 |
| 485 | 8,916 | 1,716 | 44 | 23.97 | 31.02 | 0.700 | 0.541 | 0.159 |
| 667 | 9,157 | 1,957 | 42 | 23.32 | 34.06 | 0.720 | 0.493 | 0.227 |
| 849 | 9,433 | 2,233 | 44 | 24.16 | 32.97 | 0.695 | 0.509 | 0.186 |
| 1,032 | 9,645 | 2,445 | 42 | 23.49 | 33.42 | 0.714 | 0.502 | 0.212 |
| 1,215 | 9,834 | 2,634 | 37 | 21.11 | 32.01 | 0.795 | 0.524 | 0.271 |
| 1,397 | 10,041 | 2,841 | 35 | 19.94 | 29.80 | 0.842 | 0.563 | 0.279 |
| 1,579 | 10,292 | 3,092 | 37 | 21.06 | 33.25 | 0.797 | 0.505 | 0.292 |
| 1,762 | 10,442 | 3,242 | 37 | 20.18 | 33.55 | 0.832 | 0.500 | 0.332 |
| 1,945 | 10,513 | 3,313 | 38 | 20.63 | 34.74 | 0.813 | 0.483 | 0.330 |
Three things move together: the gap between jobs grows 3.4x, the walk length grows 2.6x (3.3x from 22:23),
and the GPU time per job falls 17% (the other seven identities idle more, so this one's dispatches wait
less). Summed over the eight identities the card delivered 197 MH/s at 22:27 and about 165 MH/s at 22:55
(the launcher's cumulative total shows 197.6 falling to 183.4). None of this is the GPU: the inside-jobs rate
rose. At 22:57 the epoch changed, the launcher rebuilt the workers, the walk length went back to 1, and the
gap went back to 2% (table 2.1) while the difficulty stayed at 83 to 85M. Difficulty is therefore not the
driver: it rose 57M to 85M across epoch 2 and the gap tracked the DAA position, and it did not fall at 22:57
when the gap did. Jobs are a fixed 2^24 nonces (job_nonces: 1 << 24 in the miner), so a job does not get
longer when blocks get rarer; a harder target only means fewer found lines, each of which costs the miner
one CPU hash to re-check (0.1 per job on the PC, nothing).
Time-slicing of eight contexts is real but constant: with all eight busy, each gets one eighth of the card and the sum is the card's rate (142 MH/s in table 2.1, flat for 7 minutes). The only way the sum falls is the card going idle, which is what a growing CPU-side gap in every worker does.
2.3 The 22:03 and 22:09 runs
22:03 (four CUDA identities): 87 to 104 MH/s inside jobs but 3.2 to 4.1 MH/s wall and template_age of
6 to 7 s: the miner spent 6 s between 0.17 s jobs. 22:09 (eight identities, epoch 1, DAA 5,659 on):
24.2 to 25.1 MH/s wall per identity for 3 minutes with the gap at 4 to 8%. Both are consistent with the walk
(epoch 1 started at DAA 3,600, so the walk was 2,000 blocks at 22:03 with a node still syncing).
3. Code audit: what is allocated per job, and what is freed
3.1 proto-cuda/host.cu, runServe (HEAD, the binary the project lead ran, built by the launcher at 22:09:57)
| Allocation | When | Size | Freed |
|---|---|---|---|
gCache (cudaMalloc in setupCache) |
once | 256 MiB | at exit |
dDs dataset |
once | 2^28 words = 1 GiB | at exit |
dOut |
once | 2^22 x 8 = 32 MiB | at exit |
hOut (std::vector<uint64_t>) |
once | 32 MiB host | at exit |
line, f (8 std::string), prehash, epochSeed, daySeed vectors |
per job | under 1 KiB | end of the loop iteration (scope) |
IgneumInitWords, 49-byte buffer |
per dispatch | stack | immediately |
cudaMemcpy device to host |
per dispatch | 32 MiB, 4 per job | not an allocation; 180 MB/s per identity, 1.4 GB/s for eight, flat |
CPU scan of hOut |
per dispatch | 4M compares | flat |
No cudaMalloc, no container growth, no re-upload of the dataset, no module reload per job. The working-tree
version (the hot-swap agent's) adds CudaPair (at most two resident, the old one released after the first
job on the new pair, releasePair frees dataset, cache and both modules) and PrepareTask (deleted after
the load); still nothing per job. cudaDeviceSynchronize per dispatch under the default
cudaDeviceScheduleAuto spins the host thread when the process holds fewer contexts than the machine has
cores, which is always true here (one context per process): that is the 6.2% CPU per worker the project lead saw (one
of 16 threads). The hot-swap working tree sets cudaSetDeviceFlags(cudaDeviceScheduleBlockingSync) before
the context is created (host.cu, main), which is the right call and the right place; the thread then sleeps
on the dispatch. It does not change the hash rate directly, but eight spinning threads plus eight OpenCL
workers plus the node plus sixteen miner processes contend for the PC's cores, and the walk is made of
sixteen thousand small RPCs per minute that need those cores, so it shortens the gap too.
3.2 proto-opencl/host.c, runServe
Once: dOut (32 MiB), dInit (32 bytes), hOut, the dataset and cache (one ServePair, two at most
during a prepare). Per dispatch: a 32-byte clEnqueueWriteBuffer, five clSetKernelArg, one
launch1D whose event is released, clFinish (HEAD) or clWaitForEvents (working tree), a blocking
clEnqueueReadBuffer of chunk x 8 bytes. Per job: stack buffers. Nothing grows. clFinish on the AMD
runtime can spin a core (the working tree's comment says so and switches to the event wait, which some
runtimes still spin on; only HWiNFO or Task Manager on the PC can tell for this driver). The AMD identities
run 0.1 to 0.2 MH/s each with template_age of 15 to 21 s: a 2^24 job takes 15 s on the integrated GPU,
so they contribute nothing to the hash rate and should probably not run at all while the 5090 is mining.
3.3 proto-metal/main.swift, runServe
Once: outBuf (32 MiB shared). Per pair: one compiled pipeline and one 1 GiB dataset buffer in ServeStore
(at most two of each; prune after the first job on a new pair). Per dispatch: one command buffer and
encoder, released by ARC when the loop iteration ends; waitUntilCompleted blocks without spinning.
Per job: the split strings and unhex arrays. The Mac runs below confirm the worker's RSS is flat
within every run (46 to 58 MiB, the figure depending on the program).
3.4 vendor/igneum-node/igneum/miner/src/main.rs, mine_worker
Per job: one job line String, a Vec<&str> per worker line, raw.clone() for a submit, and
seeder.seeds_for. State that persists: hashes_total, gpu_ms_total, jobs, found, extra
(counters); Seeder::memo (HashMap<(u64, Hash), Hash>, cleared at 10,000 entries, under 1 MiB);
Voter::proposed_at (pruned to 200 entries) and Voter::lock_latencies_ms (one f64 per lock seen, one
every 30 s, unbounded but 23 KiB a day; the Windows build predates the voter). No log buffering: every line
is written as printed. The one per-job cost that grows is time, not memory: seeds_for.
// main.rs, Seeder::seeds_for (HEAD and the Windows build 745d41ef alike)
let sink = client.get_block_dag_info().await.expect("dag info").sink; // one RPC per job
if let Some(seed) = self.memo.get(&(epoch, sink)) { return ... } // hit only while the sink is unchanged
let epoch_start = epoch * POW_EPOCH_BLOCKS;
let mut cur = sink;
let seed = loop { // miss: one getBlock per block
if cur == self.genesis { break self.genesis; } // from the sink down to the
let b = client.get_block(cur, false).await.expect("get block"); // epoch's first block
if b.header.daa_score < epoch_start { break cur; }
cur = b.verbose_data.expect("verbose data").selected_parent_hash;
};
On the PC the sink changed 2.2 times a second (the node line: "network 2.18 blocks/s") against 11 jobs a
second across the eight CUDA identities, so most jobs hit the memo and roughly one in five missed; at
3,300 blocks per miss and the 0.33 s average gap, a getBlock round trip cost about 0.5 ms (approximate,
derived, not measured on the PC). The node pays for every one of those calls too (verbose data on each).
4. Mac reproduction with the Metal worker
Machine: Apple M5 Max, 64 GiB, the HEAD proto-metal/main.swift built into the scratchpad
(swiftc -O, byte-identical in size to proto-metal/igneum-bench), the difficulty worktree's igneumd
and igneum-miner (the pair proven on override-genesis networks; the Seeder and the worker protocol are
the same code as HEAD), everything at nice -n 19. The Mac was carrying other agents' builds (load average
40 to 147 during R1, 4 to 25 later) and, during R1, another agent's Metal worker on the same GPU, so absolute
rates are not comparable between runs; the per-interval shape within a run is what each run is for.
4.1 R1, control: one identity, epoch 0 (no walk), devnet genesis bits 0x1d100000, 308 s
| t (s) | jobs in interval | wall MH/s (printed, cumulative) | wall MH/s (interval) | inside MH/s (interval) | gap |
|---|---|---|---|---|---|
| 31 | 46 | 25.18 | 25.18 | 25.78 | 2.3% |
| 62 | 26 | 19.58 | 13.96 | 14.02 | 0.4% |
| 92 | 25 | 17.69 | 13.64 | 13.96 | 2.3% |
| 123 | 28 | 17.07 | 15.29 | 15.25 | 0% |
| 153 | 26 | 16.52 | 14.07 | 14.39 | 2.2% |
| 184 | 26 | 16.11 | 14.07 | 14.30 | 1.6% |
| 215 | 26 | 15.82 | 14.12 | 14.27 | 1.0% |
| 246 | 26 | 15.60 | 14.11 | 14.24 | 0.9% |
| 277 | 26 | 15.43 | 14.06 | 14.33 | 1.9% |
| 308 | 26 | 15.31 | 14.40 | 14.42 | 0.1% |
Another agent's Metal worker started on the GPU 25 s into this run. The printed cumulative rate then
"decays" from 25.2 to 15.3 over five minutes while every interval after the first sits at 14.0 to 14.4 MH/s:
the same shape as nvidia-1 in table 2.1, from the same arithmetic. The gap is 0 to 2% throughout (epoch 0:
seeds_for returns genesis without a walk). Worker RSS 56.8 to 56.9 MiB, miner RSS 268 MiB, flat. The run
ended at 308 s when every process of this session was SIGTERMed at 22:21:37 UTC (the credits outage), and
was not repeated because R2 to R6 cover the same ground.
4.2 R2, the walk reproduced: one identity in epoch 1, sink churn on, off, on (900 s)
Recipe: a fresh skip_proof_of_work node (difficulty worktree igneumd, override {"skip_proof_of_work": true, "genesis_bits": 487587840}), pumped to DAA 4,000 through getBlockTemplate + submitBlock with timestamps one
second apart ending at now (so the difficulty rule saw 1 block/s and held at 76.8M, 0.11 founds per job, the PC's
regime); then one Metal identity for 900 s while a pump added one block a second for 300 s (every job sees a new
sink, so every job walks), nothing for 300 s (the sink moves only on the identity's own blocks, one per 8 s), and
one block a second again for 300 s. Walk length when a job misses: DAA minus 3,600, from 400 at the start to
about 1,000 at the end. Load average 2 to 9 for the first 660 s, 12 to 14 after (another agent's build).
| t (s) | jobs in interval | wall MH/s (printed, cumulative) | inside MH/s (printed, cumulative) | wall MH/s (interval) | inside MH/s (interval) | gap | template age s |
|---|---|---|---|---|---|---|---|
| 30 | 44 | 24.27 | 26.94 | 24.27 | 26.94 | 10% | 0.72 |
| 60 | 43 | 24.13 | 27.08 | 23.94 | 27.22 | 12% | 0.73 |
| 91 | 43 | 23.99 | 27.14 | 24.18 | 27.26 | 11% | 0.74 |
| 121 | 43 | 23.89 | 27.15 | 23.14 | 27.18 | 15% | 0.62 |
| 152 | 42 | 23.77 | 27.17 | 23.84 | 27.25 | 13% | 0.62 |
| 182 | 41 | 23.60 | 27.18 | 22.60 | 27.23 | 17% | 0.62 |
| 212 | 41 | 23.46 | 27.18 | 22.26 | 27.18 | 18% | 0.76 |
| 243 | 41 | 23.36 | 27.17 | 23.19 | 27.10 | 14% | 0.79 |
| 273 | 42 | 23.37 | 27.17 | 23.41 | 27.17 | 14% | 0.74 |
| 303 | 43 | 23.40 | 27.17 | 23.31 | 27.17 | 14% | 0.79 |
| 334 | 48 | 23.68 | 27.14 | 26.93 | 26.88 | -0% | 0.62 |
| 364 | 48 | 23.92 | 27.13 | 26.32 | 27.03 | 3% | 0.61 |
| 394 | 49 | 24.16 | 27.16 | 26.65 | 27.49 | 3% | 0.61 |
| 425 | 48 | 24.34 | 27.19 | 27.32 | 27.54 | 1% | 0.72 |
| 455 | 48 | 24.49 | 27.22 | 26.50 | 27.61 | 4% | 0.72 |
| 485 | 47 | 24.59 | 27.24 | 25.87 | 27.53 | 6% | 0.62 |
| 516 | 48 | 24.70 | 27.26 | 26.92 | 27.55 | 2% | 0.61 |
| 546 | 47 | 24.79 | 27.28 | 26.33 | 27.61 | 5% | 0.61 |
| 576 | 48 | 24.87 | 27.30 | 25.87 | 27.65 | 6% | 0.61 |
| 606 | 45 | 24.87 | 27.31 | 24.63 | 27.50 | 10% | 0.73 |
| 636 | 43 | 24.83 | 27.32 | 23.96 | 27.53 | 13% | 0.74 |
| 667 | 43 | 24.77 | 27.33 | 23.82 | 27.55 | 14% | 0.74 |
| 697 | 39 | 24.65 | 27.34 | 21.98 | 27.59 | 20% | 0.84 |
| 727 | 33 | 24.38 | 27.34 | 18.00 | 27.34 | 34% | 0.85 |
| 758 | 37 | 24.22 | 27.35 | 20.82 | 27.63 | 25% | 0.86 |
| 788 | 36 | 24.06 | 27.35 | 19.95 | 27.35 | 27% | 0.98 |
| 818 | 33 | 23.85 | 27.36 | 18.24 | 27.71 | 34% | 0.87 |
| 848 | 37 | 23.74 | 27.36 | 20.61 | 27.36 | 25% | 0.87 |
| 879 | 36 | 23.60 | 27.37 | 20.18 | 27.70 | 27% | 0.89 |
The inside-jobs rate never moves (27.0 to 27.7 MH/s in every one of the 29 intervals). The wall rate is
22.3 to 24.3 MH/s with churn (gap 12 to 18% at a 400 to 700 walk), 25.9 to 27.3 MH/s without it (gap 0 to 4%),
and 18.0 to 21.0 MH/s with churn again at a 700 to 1,000 walk (gap 25 to 35%). The miner process's own CPU
went from 0 to 1% in the quiet phase to 11 to 21% in the last phase (the walk is the miner parsing a thousand
getBlock answers per job); the worker's stayed at 0.4 to 0.9% and its RSS at 56 to 57 MiB. This is the
shared miner logic, not CUDA, not the card, not the node's template path (template_age moves with the gap
because it is measured from the template fetch, which precedes the walk).
5. The time-slicing hypothesis: one worker versus eight, at two fixed difficulties
the project lead's Task Manager reading (GPU memory flat at 14.7 GB, the card 99% busy, each CUDA worker at 6.2% CPU) and the
hypothesis that eight contexts time-slicing one card with "longer jobs as blocks get rarer" explain the decay. A
job is a fixed 2^24 nonces, so its length does not depend on the target, but the four runs below test the
hypothesis as stated: one Metal worker and eight, each at a fixed low difficulty (2^25, 0.25 founds per job)
and a fixed high one (2^31, 0.004 founds per job), 10 minutes each, on fresh networks under Kaspa's sampled
rule, which holds the genesis bits for the first 600 blocks (no run made more than 242). Epoch 0 throughout, so
no walk. The 6.2% CPU per CUDA worker is cudaDeviceSynchronize spinning (section 3.1); the Metal worker
sleeps in waitUntilCompleted, so its CPU is the per-dispatch scan only.
5.1 R3, one worker, difficulty 2^25 (33.5M), 600 s: 30.51 MH/s wall, 30.72 inside, 1,091 jobs, 224 blocks
| t (s) | jobs in interval | wall MH/s (printed, cumulative) | inside MH/s (printed, cumulative) | wall MH/s (interval) | inside MH/s (interval) | gap | template age s |
|---|---|---|---|---|---|---|---|
| 30 | 56 | 30.95 | 31.18 | 30.95 | 31.18 | 1% | 0.54 |
| 91 | 56 | 31.14 | 31.25 | 32.09 | 31.23 | -3% | 0.55 |
| 151 | 56 | 31.10 | 31.18 | 31.19 | 31.26 | 0% | 0.53 |
| 212 | 56 | 31.09 | 31.17 | 31.47 | 30.94 | -2% | 0.56 |
| 272 | 55 | 30.92 | 31.01 | 29.88 | 30.52 | 2% | 0.56 |
| 333 | 53 | 30.60 | 30.80 | 28.73 | 29.77 | 4% | 0.55 |
| 394 | 55 | 30.44 | 30.69 | 30.03 | 30.57 | 2% | 0.55 |
| 455 | 56 | 30.48 | 30.71 | 31.42 | 30.85 | -2% | 0.54 |
| 515 | 56 | 30.53 | 30.74 | 30.40 | 30.90 | 2% | 0.58 |
| 576 | 55 | 30.50 | 30.71 | 30.79 | 30.53 | -1% | 0.54 |
5.2 R4, eight workers, difficulty 2^25, 600 s: 29.38 MH/s wall summed, 29.45 inside summed, 1,053 jobs, 242 blocks
Identity 1 (3.52 MH/s wall, 3.53 inside) and identity 8 (3.46 and 3.47); the eight ran 3.32 to 4.38 MH/s each:
| t (s) | jobs in interval | wall MH/s (printed, cumulative) | inside MH/s (printed, cumulative) | wall MH/s (interval) | inside MH/s (interval) | gap | template age s |
|---|---|---|---|---|---|---|---|
| 34 | 7 | 3.42 | 3.44 | 3.42 | 3.44 | 1% | 4.69 |
| 134 | 7 | 3.51 | 3.52 | 3.58 | 3.58 | 0% | 4.54 |
| 233 | 7 | 3.52 | 3.53 | 3.49 | 3.53 | 1% | 4.64 |
| 334 | 7 | 3.52 | 3.52 | 3.58 | 3.52 | -2% | 4.85 |
| 434 | 7 | 3.52 | 3.52 | 3.64 | 3.52 | -3% | 4.62 |
| 534 | 7 | 3.52 | 3.53 | 3.50 | 3.69 | 5% | 4.38 |
| 600 | 7 | 3.52 | 3.53 | 3.51 | 3.53 | 1% | 4.76 |
| t (s) | jobs in interval | wall MH/s (printed, cumulative) | inside MH/s (printed, cumulative) | wall MH/s (interval) | inside MH/s (interval) | gap | template age s |
|---|---|---|---|---|---|---|---|
| 34 | 7 | 3.41 | 3.43 | 3.41 | 3.43 | 1% | 4.99 |
| 136 | 7 | 3.46 | 3.46 | 3.45 | 3.43 | -1% | 4.72 |
| 238 | 7 | 3.45 | 3.46 | 3.36 | 3.46 | 3% | 4.83 |
| 339 | 7 | 3.46 | 3.47 | 3.38 | 3.57 | 5% | 4.92 |
| 441 | 7 | 3.46 | 3.46 | 3.53 | 3.46 | -2% | 4.71 |
| 542 | 7 | 3.47 | 3.47 | 3.57 | 3.47 | -3% | 5.11 |
| 577 | 7 | 3.46 | 3.47 | 3.33 | 3.47 | 4% | 4.90 |
Eight contexts cost 4% of the card against one (29.4 vs 30.5 MH/s) and the cost is constant: 7 jobs per 100 s per identity in every interval, gap 0 to 1%, no identity's rate moves. Each worker used 0.0 to 0.6% CPU and 46 to 57 MiB; each miner 0% and 268 MiB. A job takes 4.4 to 5.1 s wall per identity (eight times R3's 0.55 s) because the card is shared, which is time-slicing doing exactly what it should.
5.3 R5, one worker, difficulty 2^31 (2.15G), 600 s: 37.01 MH/s wall, 37.75 inside, 1,324 jobs, 4 blocks
| t (s) | jobs in interval | wall MH/s (printed, cumulative) | inside MH/s (printed, cumulative) | wall MH/s (interval) | inside MH/s (interval) | gap | template age s |
|---|---|---|---|---|---|---|---|
| 30 | 65 | 36.18 | 36.87 | 36.18 | 36.87 | 2% | 0.45 |
| 90 | 70 | 37.50 | 37.89 | 38.61 | 38.96 | 1% | 0.43 |
| 151 | 67 | 37.67 | 38.05 | 36.76 | 37.73 | 3% | 0.44 |
| 212 | 66 | 37.34 | 37.89 | 37.22 | 37.53 | 1% | 0.45 |
| 272 | 68 | 37.43 | 37.96 | 37.30 | 38.12 | 2% | 0.45 |
| 332 | 66 | 37.29 | 37.89 | 36.44 | 37.59 | 3% | 0.45 |
| 393 | 65 | 37.14 | 37.83 | 36.07 | 37.46 | 4% | 0.45 |
| 454 | 66 | 37.06 | 37.77 | 37.27 | 37.49 | 1% | 0.44 |
| 514 | 67 | 37.06 | 37.77 | 36.70 | 37.77 | 3% | 0.46 |
| 575 | 66 | 37.02 | 37.75 | 37.59 | 37.57 | -0% | 0.44 |
Blocks are 50 times rarer than in R3 and the job is the same 0.45 s; the rate is flat for ten minutes with a 1 to 3% gap. (R5's network has a different genesis hash from R3's, so a different epoch-0 program: 112 loads per hash against R3's; the 37 versus 30.5 MH/s is the program, as the bench log's per-seed spread predicts.)
5.4 R6, eight workers, difficulty 2^31, 600 s
36.67 MH/s wall summed, 36.75 inside summed, 1,314 jobs, 5 blocks; the eight ran 4.29 to 5.55 MH/s each. Identity 1 (4.62 MH/s wall, 4.63 inside):
| t (s) | jobs in interval | wall MH/s (printed, cumulative) | inside MH/s (printed, cumulative) | wall MH/s (interval) | inside MH/s (interval) | gap | template age s |
|---|---|---|---|---|---|---|---|
| 32 | 8 | 4.24 | 4.27 | 4.24 | 4.27 | 1% | 4.11 |
| 130 | 9 | 4.53 | 4.54 | 4.70 | 4.69 | -0% | 3.46 |
| 227 | 9 | 4.58 | 4.59 | 4.60 | 4.59 | -0% | 3.65 |
| 325 | 9 | 4.60 | 4.62 | 4.80 | 4.71 | -2% | 3.74 |
| 422 | 9 | 4.62 | 4.63 | 4.92 | 4.75 | -4% | 3.44 |
| 520 | 9 | 4.61 | 4.62 | 4.52 | 4.62 | 2% | 3.53 |
| 585 | 9 | 4.62 | 4.63 | 4.73 | 4.63 | -2% | 3.56 |
Eight contexts cost 1% against one (36.7 vs 37.0 MH/s), 9 jobs per 100 s per identity in every interval, the gap within the noise of 30-second intervals holding 2 to 3 jobs (plus or minus 4%), CPU 0.0 to 0.6% per worker and 0% per miner, RSS flat (46 to 57 MiB per worker, 268 MiB per miner). Load average 17 to 33 during this run (other agents' builds), which did not move the rate.
5.5 Verdict on the hypothesis
Rarer blocks do not lengthen jobs and do not lower the rate (R3 against R5, one worker; R4 against R6, eight).
Eight workers on one card lose a constant 4% to context switching and nothing over time (R4, R6). What the
Task Manager reading does show is the CPU side: eight threads spinning in cudaDeviceSynchronize, which the
hot-swap branch's cudaDeviceScheduleBlockingSync removes, and which matters because the sixteen miner
processes and the node need those cores for the seed walk until the walk is fixed. The difficulty rising
through the evening (9M to 83M) tracked the DAA position in the epoch, which is what the gap tracks; the
epoch boundary at 22:57, where the difficulty held and the gap collapsed, separates the two.
6. The fix, as a unified diff (not applied)
Memoise per block instead of per sink. A block's epoch seed is a function of the block alone (its selected
parent chain is fixed), so the walk from a new sink can stop at the first block already seen: a job then
costs as many getBlock calls as the sink advanced since the last job (one or two at 2 blocks/s) instead of
the DAA position in the epoch. The map is per epoch and the previous epoch's map is kept for the boundary.
The second hunk prints per-interval rates in STATUS next to the cumulative ones and the number of seed-walk
calls, so the launcher and the logs show what the card is doing now. Against
vendor/igneum-node/igneum/miner/src/main.rs at the working tree of 3 October 2026 (git apply --check
passes, and cargo check -p igneum-miner on a scratch worktree with the patch applied finishes with no new
warnings; also saved as docs/analysis/hashrate-decay-2026-10-03.patch).
--- a/igneum/miner/src/main.rs
+++ b/igneum/miner/src/main.rs
@@ -366,8 +366,15 @@
/// selected parent, which is the selected parent of any block built on the current template).
struct Seeder {
genesis: Hash,
- /// (epoch index, sink the walk started from) -> seed
- memo: HashMap<(u64, Hash), Hash>,
+ /// epoch index -> (selected-chain block -> that epoch's seed) for every block a walk has visited. A block's seed
+ /// is a function of the block alone (its selected parent chain is fixed), so a walk from a new sink stops at the
+ /// first block already in the map: a job costs as many `get_block` calls as the sink advanced, not the DAA
+ /// position in the epoch (the old `(epoch, sink)` key missed on every new block and walked back to the epoch's
+ /// first block, up to 3,600 calls, measured on the RTX 5090 on 3 October 2026 as a gap between jobs that grew
+ /// from 0.10 s to 0.33 s across epoch 2; docs/analysis/hashrate-decay-2026-10-03.md).
+ memo: HashMap<u64, HashMap<Hash, Hash>>,
+ /// `get_block` calls made by seed walks, for the STATUS line
+ walk_calls: u64,
}
impl Seeder {
@@ -392,7 +399,7 @@
if genesis != DEVNET_PARAMS.genesis.hash {
println!("{} node genesis {} differs from the shared devnet genesis: a private test network", now(), genesis);
}
- Self { genesis, memo: HashMap::new() }
+ Self { genesis, memo: HashMap::new(), walk_calls: 0 }
}
async fn seeds_for(&mut self, client: &GrpcClient, header: &Header) -> EpochSeeds {
@@ -402,25 +409,32 @@
return EpochSeeds { epoch: self.genesis, day };
}
let sink = client.get_block_dag_info().await.expect("dag info").sink;
- if let Some(seed) = self.memo.get(&(epoch, sink)) {
- return EpochSeeds { epoch: *seed, day };
- }
+ // This epoch's map and the previous one stay (a template can still land in the previous epoch at the
+ // boundary); older ones go. A map holds at most one entry per selected-chain block of its epoch (3,600).
+ self.memo.retain(|e, _| *e + 1 >= epoch);
+ let known = self.memo.entry(epoch).or_default();
let epoch_start = epoch * POW_EPOCH_BLOCKS;
let mut cur = sink;
+ let mut path = Vec::new();
let seed = loop {
+ if let Some(seed) = known.get(&cur) {
+ break *seed;
+ }
if cur == self.genesis {
break self.genesis;
}
let b = client.get_block(cur, false).await.expect("get block");
+ self.walk_calls += 1;
if b.header.daa_score < epoch_start {
break cur;
}
+ path.push(cur);
cur = b.verbose_data.expect("verbose data").selected_parent_hash;
};
- if self.memo.len() > 10_000 {
- self.memo.clear();
+ for h in path {
+ known.insert(h, seed);
}
- self.memo.insert((epoch, sink), seed);
+ known.insert(sink, seed);
EpochSeeds { epoch: seed, day }
}
}
@@ -707,6 +721,9 @@
let mut gpu_ms_total = 0.0f64;
let mut jobs = 0u64;
let mut last_report = Instant::now();
+ // The previous STATUS line's counters, for the per-interval rates (the cumulative ones decay by construction
+ // after a fast first interval, which read as a hash-rate decay on 3 October 2026)
+ let (mut last_hashes, mut last_gpu_ms, mut last_jobs, mut last_walk) = (0u64, 0.0f64, 0u64, 0u64);
let mut last_seeds: Option<EpochSeeds> = None;
let mut rng = rand::thread_rng();
while start.elapsed() < Duration::from_secs(o.secs) {
@@ -813,8 +830,9 @@
}
if last_report.elapsed() > Duration::from_secs(o.status_secs) {
let el = start.elapsed().as_secs_f64();
+ let (dt, dh, dg) = (last_report.elapsed().as_secs_f64(), (hashes_total - last_hashes) as f64, gpu_ms_total - last_gpu_ms);
println!(
- "{} STATUS '{}' [worker]: {:.0}s jobs={} accepted={} rejected={} mismatched={} extra={} rate={:.2} blocks/s hash={:.2} MH/s wall ({:.2} MH/s inside jobs) template_age={:.2}s synced={}",
+ "{} STATUS '{}' [worker]: {:.0}s jobs={} accepted={} rejected={} mismatched={} extra={} rate={:.2} blocks/s hash={:.2} MH/s wall ({:.2} MH/s inside jobs) now={:.2} MH/s wall ({:.2} MH/s inside jobs, {} jobs, seed walk {} calls) template_age={:.2}s synced={}",
now(),
o.vote_label,
el,
@@ -826,9 +844,14 @@
found as f64 / el,
hashes_total as f64 / el / 1e6,
if gpu_ms_total > 0.0 { hashes_total as f64 / gpu_ms_total / 1e3 } else { 0.0 },
+ if dt > 0.0 { dh / dt / 1e6 } else { 0.0 },
+ if dg > 0.0 { dh / dg / 1e3 } else { 0.0 },
+ jobs - last_jobs,
+ seeder.walk_calls - last_walk,
template_at.elapsed().as_secs_f64(),
synced
);
+ (last_hashes, last_gpu_ms, last_jobs, last_walk) = (hashes_total, gpu_ms_total, jobs, seeder.walk_calls);
last_report = Instant::now();
}
}
The test that proves it
- Unit:
Seeder::seeds_forwith a mocked client that countsget_blockcalls: after one job at DAA3600 + n, a second job whose sink is one block ahead must make exactly 1get_blockcall (today:n + 1), and a job on a sink from a different tip of the same height must walk only to the common ancestor. - Network: the R2 recipe above (a
skip_proof_of_worknode pumped to DAA 4,000 with one-second timestamps, one Metal identity, pumped blocks at 1/s for 300 s, none for 300 s, 1/s for 300 s). Before the fix the wall-versus-inside gap andwalk=in STATUS move with the churn phases; after the fixwalkper job is at most the blocks the sink advanced (1 to 2) in every phase and the gap stays within 2%. On the PC the acceptance figure is table 2.2 flattening: the gap per job no better than 0.10 s at DAA 10,800 (3,600 into an epoch) with eight identities.
7. Two things for the project lead to check on the PC
- Dedicated GPU memory over time. Task Manager, Performance, GPU 0, "Dedicated GPU memory usage", or
nvidia-smi --query-gpu=timestamp,memory.used,utilization.gpu,temperature.gpu,clocks.sm,clocks.mem,power.draw,clocks_throttle_reasons.active --format=csv -l 10 > gpu.csv. Expected: flat from the moment the eighth worker printsready(each CUDA worker holds 1 GiB dataset + 256 MiB cache + 32 MiB output + the context, about 1.6 GiB; eight of them 13 to 15 GiB of the 32 GiB, which matches the flat 14.7 GB the project lead read). A line that climbs while the hash rate falls would mean a leak in the worker; none exists in the code and the Mac RSS traces are flat. - GDDR7 memory junction temperature and the throttle reasons. HWiNFO64, Sensors, under the GPU: "GPU
Memory Junction Temperature", "GPU Thermal Limit", "GPU Power Limit", "GPU Reliability Voltage Limit"
(each "Yes" or "No"), "GPU Effective Clock" and "GPU Memory Clock"; or the
clocks_throttle_reasons.activecolumn above (0x0000000000000004is SW power cap,0x0000000000000020SW thermal slowdown,0x0000000000000040HW thermal slowdown,0x0000000000000080HW power brake). Approximate thresholds, from memory: the core starts to pull clocks around 83 C (the project lead's 55 to 59 C is far below it); GDDR6X on the previous generations throttles from about 95 C junction and the hard limit is 105 C; GDDR7 figures are not published, so treat anything above 90 C junction as the zone to watch and a "Yes" on any limit row as the signal. The decisive sign of throttling is "GPU Effective Clock" falling while utilization stays at 99%; in the logs it would appear as the inside-jobs per-interval rate falling, which it never does (it rose). A flat line on both rows means the card was not the cause, which is what the logs already say.