Workers answer a job they have no pair for with a need line; the miner prepares the current pair (PC 2 stuck on the previous epoch)

Root cause from the uploads (bench-log entry): the app exported the pack while its node was in IBD inside the previous
epoch, the OpenCL worker started after the boundary with no next epoch within lead, so no prepare was ever sent and
every job was a seed mismatch; the CUDA worker on the same PC had swapped correctly. Both workers now print
need <epoch> <day> before the error; the devnet-v4 miner (3bfe346f) prepares the current pair on a need line or three
mismatches, exits 42 for a worker without prepare support, and restarts a ready worker that completes no job for
60 s with jobs queued. Package rebuilt with the guarded miner (ship build on dc749905), payload inputs published.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-04 13:32:58 +00:00
parent bdd6f602f5
commit 4d907d8a14
5 changed files with 41 additions and 0 deletions

View file

@ -834,3 +834,22 @@ Open: the gap itself. Zero blocks for 78 minutes followed by every missed block
time is also what a stalled observer node delivering its own catch-up looks like; the lag metric now makes either
case visible on the page within 30 s, and the self-check covers the dead-subscription case. Node logs for 11:57
to 13:18 UTC would settle which it was.
## 4 October 2026, PC 2 at the 14:20 boundary: a worker stuck on the previous epoch (root cause from the uploads)
Run `win-1ccfe586-20261004-132055` (the Igneum Miner app, package 0.3.0 workers). Sequence from the node and miner
uploads: the app reinstalled and its node restarted at 14:20:29 BST in IBD from DAA 17,881, inside epoch 4 (seed
57ac7663...); the app exported `packs\devnet` from that node's template at once, so both workers started on 57ac. The
boundary at DAA 18,000 passed about a minute later. The CUDA miner's first templates still carried `next_epoch_seed`
c23e65dd... within lead: `PREPARE sent` at 14:20:57, `prepared` 0.9 s later (NVRTC 151 ms, cache 68, dataset 113,
self-test 511 ms), `switched to the prepared pair` at 14:21:28, then 60 MH/s with 8 identities and 101 accepted blocks
in 271 s. The OpenCL worker reported ready 43 s after the CUDA one (14:21:39); by then every template was on c23e as
the current pair and no next epoch was within lead, so the miner never sent a `prepare`, and the worker answered 514
jobs in a row with `epoch seed mismatch` (one every 0.5 s, the miner's error back-off) for the rest of the run. The
message text "holds prepared epoch 57ac..." is host.c's wording for a pack read at run time, which is why the stuck
worker looked like a wrong prediction: PC 2's node announced the same next epoch (c23e) as the chain. No node on this
PC predicted a different epoch, and the 57ac pack was simply the previous epoch's. Fixes: devnet-v4 miner 3bfe346f
(prepare the current pair after a `need` line or three mismatches; exit 42 without prepare support; restart a ready
worker with jobs queued and no job done for 60 s), workers emit `need <epoch> <day>` before the error. Not measured
here: the swap time of the forced prepare on the PC; the OpenCL run's "2 jobs in 154 s" were the two jobs before the
first mismatch and are not a rate.

View file

@ -22,6 +22,7 @@
// found <job_id> <nonce u64> <hash_hex 16>
// done <job_id> <hashes> <ms>
// error <job_id> <text>
// need <epoch_seed_hex> <day_seed_hex> before the mismatch error: the pair this worker lacks (the miner prepares it)
// prepared <epoch_seed_hex> <day_seed_hex> <ms> ... | prepare-failed <epoch_seed_hex> <day_seed_hex> <text>
// info ...
// The first pack comes from --pack <dir> (igneum-miner export-pack writes it; the launcher passes it). Every pack is
@ -706,10 +707,12 @@ static int runServe(Ctx& c, const Options& o, Pair* cur) {
old = cur; cur = prepared; prepared = nullptr; switched = true;
info(fmt("switched to the prepared pair epoch %.16s day %s in %.2f ms", cur->epochHex.c_str(), cur->dayHex.c_str(), wallMs() - t0));
} else if (std::memcmp(sw, cur->sw, 32) != 0) {
emit(fmt("need %s %s", f[6].c_str(), f[7].c_str())); // the miner prepares this pair (4 October 2026)
emit(fmt("error %s epoch seed mismatch: this worker holds epoch %.16s (seed words %08x %08x ...)%s, the job's epoch seed %.16s gives %08x %08x ...; send prepare with a pack directory",
jobId.c_str(), cur->epochHex.c_str(), cur->sw[0], cur->sw[1], prepared ? " plus one prepared pair" : "", f[6].c_str(), sw[0], sw[1]));
continue;
} else {
emit(fmt("need %s %s", f[6].c_str(), f[7].c_str()));
emit(fmt("error %s day seed mismatch: this worker's cache is for key %08x %08x ..., the job's day seed %s gives %08x %08x ...; send prepare with a pack directory",
jobId.c_str(), cur->kw[0], cur->kw[1], f[7].c_str(), kw[0], kw[1]));
continue;

View file

@ -8,6 +8,7 @@ Devnet v4 is a new chain from genesis. The node keeps its database in %LOCALAPPD
Payout: block rewards go to an EVM address. PAYOUT_EVM at the top of the bat sets one for the whole PC; left empty, the launcher derives one address per GPU vendor from this PC's name (the same on every run; it is printed in the launcher log and the status block). Finality voting is on: every identity signs every checkpoint (VOTE=0 switches it off).
The hourly program change: every hour the lottery program changes. About ten minutes before the boundary the node announces the next program; the miner writes its pack (proto-cuda\packs\prepare\<epoch>-<day>) and tells the worker, which compiles it on the card in the background, builds its cache and dataset, self-tests it and keeps it ready; at the boundary the worker swaps with no pause (the dashboard says "the next hourly program is compiled and resident"). If that ever fails, the CUDA worker finds the pack by itself when the first job on the new program arrives and compiles it then (a pause of a few seconds); and if a worker keeps erroring on the new seeds for 90 s, the launcher re-exports the pack and restarts that card's miner. --exit-on-seed-change stays on as the last fallback: a worker that cannot prepare at all exits with code 42 at the boundary and the launcher restarts it on a fresh pack.
Worker faults (added after the first field run on 4 October 2026, when an integrated AMD card's OpenCL runtime stopped running kernels after 600 s and kept answering every call with success): the OpenCL worker now treats any OpenCL error in a job as fatal, checks that each dispatch really completed, that it did not finish 20x faster per nonce than before and that its output changed, and exits with code 3 so the miner restarts it; every 200 jobs it prints a stats line with the live event and buffer counts. The miner itself kills and restarts a worker whose jobs finish in under 1/20 of the mean time per hash or whose rate jumps over 10x, drops the fake numbers, and prints "WORKER FAULT ..."; the dashboard then shows "worker fault" and "restarting" for that card instead of a rate, and the status block counts faults.
A worker started on a pack of an epoch that has just ended (the node was still catching up when the pack was exported) used to answer every job with a seed mismatch for the rest of its run. Now it asks for the pair it lacks ("need"), the miner writes that pack and sends it a prepare, and mining starts within seconds; a worker that cannot prepare is restarted on a fresh pack (exit 42). A worker that says ready but completes no job for 60 s while jobs are queued is restarted as well.
If a prebuilt worker never reports ready (150 s), the launcher says so in the events; when a CUDA Toolkit and Visual Studio happen to be installed it builds a worker from source with them instead, otherwise look at the miner log named in the event (the worker prints the reason: a missing DLL, a compile error, a driver too old for CUDA 12).
A node crash is restarted after a random 5 to 60 s; the miners are restarted once the node is back and synced (their connection dies with the node). The PC is held awake while the window runs; for the night also set Sleep to Never in Settings > System > Power.
Logs land next to the bat (igneum-<stamp>.log for the launcher, node-<stamp>.log for the node, nvidia-<stamp>.log and amd-<stamp>.log per miner process; the worker's ready, info and self-test lines are in the miner log) and are uploaded every 60 s with upload-log.bat under the labels igneum-<PC>, nodelog-<PC>, nvidia-<PC>, amd-<PC>.

View file

@ -101,6 +101,22 @@ and this is what they print:
If the fault returns on the gfx1036, the first `worker fault` line names which guard fired and that is the clue to
the runtime's failure mode; please send the amd log around it.
## The seed mismatch guard (third field run, PC 2 at the 14:20 boundary)
What happened: the app reinstalled at 14:20 and exported `packs\devnet` while its own node was still in IBD inside the
previous epoch (DAA 17,881, seed 57ac...); the boundary at 18,000 passed a minute later. The CUDA miner saw the next
seed within lead, sent `prepare` 2 s after start and swapped at 14:21:28 (fine). The OpenCL worker started 43 s later,
when the templates were already on c23e... with no next epoch within lead, so no `prepare` was ever sent and every job
was answered `error N epoch seed mismatch` for the rest of the run. PC 2's node never predicted a different epoch.
Now: a worker that lacks the job's pair prints `need <epoch hex> <day hex>` before that error; the miner, on a `need`
line or three mismatches in a row, prints `WORKER FAULT seed mismatch: ... preparing the current pair epoch ... for it`,
writes the pack under `packs\prepare\<epoch16>-<day>` and sends `prepare`; the worker answers `prepared ... self-test
PASS` and the next job switches (`info switched to the prepared pair`). A worker without prepare support gets `WORKER
FAULT seed mismatch, restarting on epoch ...` and the miner exits 42 (the launcher re-exports the pack). A ready
worker with jobs queued and no `done` for 60 s gets `WORKER FAULT no job completed for N s ...; restarting the worker`.
Expect at most one `WORKER FAULT seed mismatch` per worker start, and then mining; `faults=` on STATUS counts them.
## What to send back
The launcher log, the two miner logs (they are uploaded every minute as well) and, for the bench log, the three

View file

@ -1250,10 +1250,12 @@ static int runServe(Device* dv, const DeviceInfo* di, const Options* o) {
old = cur; cur = prepared; prepared = NULL; switched = 1;
printf("info switched to the prepared pair epoch %.16s day %s in %.2f ms\n", cur->epochHex, cur->dayHex, wallMs() - t0); fflush(stdout);
} else if (memcmp(sw, cur->sw, 32) != 0) {
printf("need %s %s\n", f[6], f[7]); /* the miner prepares this pair (4 October 2026) */
printf("error %s epoch seed mismatch: this worker holds %s%s (seed words %08x %08x ...)%s, the job's epoch seed %.16s gives %08x %08x ...; send prepare with a pack directory, or run igneum-miner export-pack and rebuild\n",
jobId, cur->epochHex[0] ? "prepared epoch " : "pack \"" IGNEUM_SEED_STRING "\"", cur->epochHex[0] ? cur->epochHex : "", cur->sw[0], cur->sw[1], prepared ? " plus one prepared pair" : "", f[6], sw[0], sw[1]);
fflush(stdout); continue;
} else {
printf("need %s %s\n", f[6], f[7]);
printf("error %s day seed mismatch: this worker's cache is for key %08x %08x ..., the job's day seed %s gives %08x %08x ...; send prepare with a pack directory, or run igneum-miner export-pack and rebuild\n",
jobId, cur->kw[0], cur->kw[1], f[7], kw[0], kw[1]);
fflush(stdout); continue;