Workers answer a job they have no pair for with a need line; the miner prepares the current pair (PC 2 stuck on the previous epoch)
Root cause from the uploads (bench-log entry): the app exported the pack while its node was in IBD inside the previous epoch, the OpenCL worker started after the boundary with no next epoch within lead, so no prepare was ever sent and every job was a seed mismatch; the CUDA worker on the same PC had swapped correctly. Both workers now print need <epoch> <day> before the error; the devnet-v4 miner (3bfe346f) prepares the current pair on a need line or three mismatches, exits 42 for a worker without prepare support, and restarts a ready worker that completes no job for 60 s with jobs queued. Package rebuilt with the guarded miner (ship build on dc749905), payload inputs published. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
45fd566bb1
commit
8bbbeb0e6c
5 changed files with 41 additions and 0 deletions
|
|
@ -834,3 +834,22 @@ Open: the gap itself. Zero blocks for 78 minutes followed by every missed block
|
|||
time is also what a stalled observer node delivering its own catch-up looks like; the lag metric now makes either
|
||||
case visible on the page within 30 s, and the self-check covers the dead-subscription case. Node logs for 11:57
|
||||
to 13:18 UTC would settle which it was.
|
||||
|
||||
## 4 October 2026, PC 2 at the 14:20 boundary: a worker stuck on the previous epoch (root cause from the uploads)
|
||||
|
||||
Run `win-1ccfe586-20261004-132055` (the Igneum Miner app, package 0.3.0 workers). Sequence from the node and miner
|
||||
uploads: the app reinstalled and its node restarted at 14:20:29 BST in IBD from DAA 17,881, inside epoch 4 (seed
|
||||
57ac7663...); the app exported `packs\devnet` from that node's template at once, so both workers started on 57ac. The
|
||||
boundary at DAA 18,000 passed about a minute later. The CUDA miner's first templates still carried `next_epoch_seed`
|
||||
c23e65dd... within lead: `PREPARE sent` at 14:20:57, `prepared` 0.9 s later (NVRTC 151 ms, cache 68, dataset 113,
|
||||
self-test 511 ms), `switched to the prepared pair` at 14:21:28, then 60 MH/s with 8 identities and 101 accepted blocks
|
||||
in 271 s. The OpenCL worker reported ready 43 s after the CUDA one (14:21:39); by then every template was on c23e as
|
||||
the current pair and no next epoch was within lead, so the miner never sent a `prepare`, and the worker answered 514
|
||||
jobs in a row with `epoch seed mismatch` (one every 0.5 s, the miner's error back-off) for the rest of the run. The
|
||||
message text "holds prepared epoch 57ac..." is host.c's wording for a pack read at run time, which is why the stuck
|
||||
worker looked like a wrong prediction: PC 2's node announced the same next epoch (c23e) as the chain. No node on this
|
||||
PC predicted a different epoch, and the 57ac pack was simply the previous epoch's. Fixes: devnet-v4 miner 3bfe346f
|
||||
(prepare the current pair after a `need` line or three mismatches; exit 42 without prepare support; restart a ready
|
||||
worker with jobs queued and no job done for 60 s), workers emit `need <epoch> <day>` before the error. Not measured
|
||||
here: the swap time of the forced prepare on the PC; the OpenCL run's "2 jobs in 154 s" were the two jobs before the
|
||||
first mismatch and are not a rate.
|
||||
|
|
|
|||
|
|
@ -22,6 +22,7 @@
|
|||
// found <job_id> <nonce u64> <hash_hex 16>
|
||||
// done <job_id> <hashes> <ms>
|
||||
// error <job_id> <text>
|
||||
// need <epoch_seed_hex> <day_seed_hex> before the mismatch error: the pair this worker lacks (the miner prepares it)
|
||||
// prepared <epoch_seed_hex> <day_seed_hex> <ms> ... | prepare-failed <epoch_seed_hex> <day_seed_hex> <text>
|
||||
// info ...
|
||||
// The first pack comes from --pack <dir> (igneum-miner export-pack writes it; the launcher passes it). Every pack is
|
||||
|
|
@ -706,10 +707,12 @@ static int runServe(Ctx& c, const Options& o, Pair* cur) {
|
|||
old = cur; cur = prepared; prepared = nullptr; switched = true;
|
||||
info(fmt("switched to the prepared pair epoch %.16s day %s in %.2f ms", cur->epochHex.c_str(), cur->dayHex.c_str(), wallMs() - t0));
|
||||
} else if (std::memcmp(sw, cur->sw, 32) != 0) {
|
||||
emit(fmt("need %s %s", f[6].c_str(), f[7].c_str())); // the miner prepares this pair (4 October 2026)
|
||||
emit(fmt("error %s epoch seed mismatch: this worker holds epoch %.16s (seed words %08x %08x ...)%s, the job's epoch seed %.16s gives %08x %08x ...; send prepare with a pack directory",
|
||||
jobId.c_str(), cur->epochHex.c_str(), cur->sw[0], cur->sw[1], prepared ? " plus one prepared pair" : "", f[6].c_str(), sw[0], sw[1]));
|
||||
continue;
|
||||
} else {
|
||||
emit(fmt("need %s %s", f[6].c_str(), f[7].c_str()));
|
||||
emit(fmt("error %s day seed mismatch: this worker's cache is for key %08x %08x ..., the job's day seed %s gives %08x %08x ...; send prepare with a pack directory",
|
||||
jobId.c_str(), cur->kw[0], cur->kw[1], f[7].c_str(), kw[0], kw[1]));
|
||||
continue;
|
||||
|
|
|
|||
|
|
@ -8,6 +8,7 @@ Devnet v4 is a new chain from genesis. The node keeps its database in %LOCALAPPD
|
|||
Payout: block rewards go to an EVM address. PAYOUT_EVM at the top of the bat sets one for the whole PC; left empty, the launcher derives one address per GPU vendor from this PC's name (the same on every run; it is printed in the launcher log and the status block). Finality voting is on: every identity signs every checkpoint (VOTE=0 switches it off).
|
||||
The hourly program change: every hour the lottery program changes. About ten minutes before the boundary the node announces the next program; the miner writes its pack (proto-cuda\packs\prepare\<epoch>-<day>) and tells the worker, which compiles it on the card in the background, builds its cache and dataset, self-tests it and keeps it ready; at the boundary the worker swaps with no pause (the dashboard says "the next hourly program is compiled and resident"). If that ever fails, the CUDA worker finds the pack by itself when the first job on the new program arrives and compiles it then (a pause of a few seconds); and if a worker keeps erroring on the new seeds for 90 s, the launcher re-exports the pack and restarts that card's miner. --exit-on-seed-change stays on as the last fallback: a worker that cannot prepare at all exits with code 42 at the boundary and the launcher restarts it on a fresh pack.
|
||||
Worker faults (added after the first field run on 4 October 2026, when an integrated AMD card's OpenCL runtime stopped running kernels after 600 s and kept answering every call with success): the OpenCL worker now treats any OpenCL error in a job as fatal, checks that each dispatch really completed, that it did not finish 20x faster per nonce than before and that its output changed, and exits with code 3 so the miner restarts it; every 200 jobs it prints a stats line with the live event and buffer counts. The miner itself kills and restarts a worker whose jobs finish in under 1/20 of the mean time per hash or whose rate jumps over 10x, drops the fake numbers, and prints "WORKER FAULT ..."; the dashboard then shows "worker fault" and "restarting" for that card instead of a rate, and the status block counts faults.
|
||||
A worker started on a pack of an epoch that has just ended (the node was still catching up when the pack was exported) used to answer every job with a seed mismatch for the rest of its run. Now it asks for the pair it lacks ("need"), the miner writes that pack and sends it a prepare, and mining starts within seconds; a worker that cannot prepare is restarted on a fresh pack (exit 42). A worker that says ready but completes no job for 60 s while jobs are queued is restarted as well.
|
||||
If a prebuilt worker never reports ready (150 s), the launcher says so in the events; when a CUDA Toolkit and Visual Studio happen to be installed it builds a worker from source with them instead, otherwise look at the miner log named in the event (the worker prints the reason: a missing DLL, a compile error, a driver too old for CUDA 12).
|
||||
A node crash is restarted after a random 5 to 60 s; the miners are restarted once the node is back and synced (their connection dies with the node). The PC is held awake while the window runs; for the night also set Sleep to Never in Settings > System > Power.
|
||||
Logs land next to the bat (igneum-<stamp>.log for the launcher, node-<stamp>.log for the node, nvidia-<stamp>.log and amd-<stamp>.log per miner process; the worker's ready, info and self-test lines are in the miner log) and are uploaded every 60 s with upload-log.bat under the labels igneum-<PC>, nodelog-<PC>, nvidia-<PC>, amd-<PC>.
|
||||
|
|
|
|||
|
|
@ -101,6 +101,22 @@ and this is what they print:
|
|||
If the fault returns on the gfx1036, the first `worker fault` line names which guard fired and that is the clue to
|
||||
the runtime's failure mode; please send the amd log around it.
|
||||
|
||||
## The seed mismatch guard (third field run, PC 2 at the 14:20 boundary)
|
||||
|
||||
What happened: the app reinstalled at 14:20 and exported `packs\devnet` while its own node was still in IBD inside the
|
||||
previous epoch (DAA 17,881, seed 57ac...); the boundary at 18,000 passed a minute later. The CUDA miner saw the next
|
||||
seed within lead, sent `prepare` 2 s after start and swapped at 14:21:28 (fine). The OpenCL worker started 43 s later,
|
||||
when the templates were already on c23e... with no next epoch within lead, so no `prepare` was ever sent and every job
|
||||
was answered `error N epoch seed mismatch` for the rest of the run. PC 2's node never predicted a different epoch.
|
||||
|
||||
Now: a worker that lacks the job's pair prints `need <epoch hex> <day hex>` before that error; the miner, on a `need`
|
||||
line or three mismatches in a row, prints `WORKER FAULT seed mismatch: ... preparing the current pair epoch ... for it`,
|
||||
writes the pack under `packs\prepare\<epoch16>-<day>` and sends `prepare`; the worker answers `prepared ... self-test
|
||||
PASS` and the next job switches (`info switched to the prepared pair`). A worker without prepare support gets `WORKER
|
||||
FAULT seed mismatch, restarting on epoch ...` and the miner exits 42 (the launcher re-exports the pack). A ready
|
||||
worker with jobs queued and no `done` for 60 s gets `WORKER FAULT no job completed for N s ...; restarting the worker`.
|
||||
Expect at most one `WORKER FAULT seed mismatch` per worker start, and then mining; `faults=` on STATUS counts them.
|
||||
|
||||
## What to send back
|
||||
|
||||
The launcher log, the two miner logs (they are uploaded every minute as well) and, for the bench log, the three
|
||||
|
|
|
|||
|
|
@ -1250,10 +1250,12 @@ static int runServe(Device* dv, const DeviceInfo* di, const Options* o) {
|
|||
old = cur; cur = prepared; prepared = NULL; switched = 1;
|
||||
printf("info switched to the prepared pair epoch %.16s day %s in %.2f ms\n", cur->epochHex, cur->dayHex, wallMs() - t0); fflush(stdout);
|
||||
} else if (memcmp(sw, cur->sw, 32) != 0) {
|
||||
printf("need %s %s\n", f[6], f[7]); /* the miner prepares this pair (4 October 2026) */
|
||||
printf("error %s epoch seed mismatch: this worker holds %s%s (seed words %08x %08x ...)%s, the job's epoch seed %.16s gives %08x %08x ...; send prepare with a pack directory, or run igneum-miner export-pack and rebuild\n",
|
||||
jobId, cur->epochHex[0] ? "prepared epoch " : "pack \"" IGNEUM_SEED_STRING "\"", cur->epochHex[0] ? cur->epochHex : "", cur->sw[0], cur->sw[1], prepared ? " plus one prepared pair" : "", f[6], sw[0], sw[1]);
|
||||
fflush(stdout); continue;
|
||||
} else {
|
||||
printf("need %s %s\n", f[6], f[7]);
|
||||
printf("error %s day seed mismatch: this worker's cache is for key %08x %08x ..., the job's day seed %s gives %08x %08x ...; send prepare with a pack directory, or run igneum-miner export-pack and rebuild\n",
|
||||
jobId, cur->kw[0], cur->kw[1], f[7], kw[0], kw[1]);
|
||||
fflush(stdout); continue;
|
||||
|
|
|
|||
Loading…
Reference in a new issue