igneum/proto-cuda/windows-app/TEST.md
igneum-labs 4d907d8a14 Workers answer a job they have no pair for with a need line; the miner prepares the current pair (PC 2 stuck on the previous epoch)
Root cause from the uploads (bench-log entry): the app exported the pack while its node was in IBD inside the previous
epoch, the OpenCL worker started after the boundary with no next epoch within lead, so no prepare was ever sent and
every job was a seed mismatch; the CUDA worker on the same PC had swapped correctly. Both workers now print
need <epoch> <day> before the error; the devnet-v4 miner (3bfe346f) prepares the current pair on a need line or three
mismatches, exits 42 for a worker without prepare support, and restarts a ready worker that completes no job for
60 s with jobs queued. Package rebuilt with the guarded miner (ship build on dc749905), payload inputs published.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 13:32:58 +00:00

9.3 KiB

First run of package 0.3.0 (prebuilt workers): what to look for

Written 4 October 2026 on the Mac, where the workers could only be checked without an NVIDIA GPU (the CUDA worker's host logic through a CPU emulation of the driver API and NVRTC, the OpenCL worker through Apple's OpenCL runtime). The real NVRTC compile and the real driver run happen for the first time on the RTX 5090 PC. This page says what a good run prints and what to send back if it does not.

Before you start

Extract the zip to a fresh short folder (C:\igneum-v4). Nothing to install. These files must sit together in it: igneum-worker-cuda.exe, nvrtc64_120_0.dll, nvrtc-builtins64_128.dll, igneum-worker-opencl.exe, igneum-miner.exe, igneumd.exe. The NVIDIA driver must be the one that already mines (617.14 on 4 October 2026 is fine; the driver API needs a driver for CUDA 12.x).

What the dashboard should show

  1. In the launcher log (igneum-.log) right after "GPUs found": prebuilt workers: igneum-worker-cuda.exe with nvrtc64_120_0.dll, igneum-worker-opencl.exe and workers to start: nvidia (prebuilt worker, NVRTC compiles on the card (driver only)), amd (prebuilt worker, the OpenCL driver compiles (driver only)). No "tools:" line: nothing is looked for when the prebuilt workers are there.
  2. After the node is synced: the phase line starting the prebuilt RTX 5090 worker, then the card appears with prebuilt worker, NVRTC compiles on the card (driver only) next to its name.
  3. Within about 10 s the event RTX 5090 worker ready (prebuilt worker, NVRTC compiles on the card (driver only)), hot swap at the hour boundary. The AMD card the same with "the OpenCL driver compiles".
  4. Then mining in green and a hash rate about the one the nvcc-built worker gave (229 MH/s on the 5090 with 8 identities on 3 October 2026; the kernel text is the same, only the compiler changed, so a large gap either way is worth a note).
  5. About ten minutes before the hour boundary: RTX 5090: the next hourly program is compiled and resident (hot swap ready). At the boundary no pause, no restart event; the miner log says switched to the prepared pair.

What the miner log should show (nvidia-.log)

The worker's own lines, in this order:

worker: info igneum-worker-cuda 1.0 (4 October 2026): device 0 NVIDIA_GeForce_RTX_5090 (sm_120, 170 SMs), driver 12.x from nvcuda.dll, NVRTC 12.8 from C:\igneum-v4\nvrtc64_120_0.dll, target sm_120 (the device's architecture, listed by NVRTC)
worker: ready cuda NVIDIA_GeForce_RTX_5090 pack igneum-epoch/<64 hex>/day/<hex> dataset-log2 28 batch 4194304 regs <N> prepare 1 path nvrtc 12.8 driver 12.x arch sm_120 worker 1.0 (4 October 2026)  (prepare support: yes, hot swap at the boundary)
worker: info first pack proto-cuda\packs\devnet: nvrtc <ms> cache <ms> dataset <ms> check <ms> ms; self-test PASS (cache head, last line and FNV-1a 64 <16 hex>; dataset head, word [268435455] and 64 samples; 96 of 96 vector lanes)

The self-test is the proof that NVRTC compiled the same program nvcc did: the 96 vector lanes are the outputs the Rust CPU interpreter put in the pack, and the cache FNV covers all 2^26 words of the cache. regs should be close to what the nvcc build reported for the same program. Every found is still re-checked by the miner's CPU code (WORKER MISMATCH in the log would be the alarm; expect none).

At the hourly prepare:

PREPARE sent for epoch seed ... (N DAA blocks before the boundary ...)
worker: info prepare started for epoch <16 hex> day <hex> from C:\igneum-v4\proto-cuda\packs\prepare\<epoch>-<day> (NVRTC sm_120 in the background)
worker: prepared <epoch hex> <day hex> <ms> nvrtc <ms> cache <ms> dataset <ms> check <ms> ms; self-test PASS (...) resident 2 programs 2 datasets

For the AMD worker (amd-.log): ready opencl <device> platform AMD_Accelerated_Parallel_Processing pack ... prepare 1 path prebuilt-generic, then info first pack ..\proto-cuda\packs\devnet: cache <ms> dataset <ms> check <ms> ms (...); self-test PASS (...).

Timings worth writing down (none measured yet on NVIDIA)

From the info first pack and prepared lines: the NVRTC compile time (nvrtc <ms>, expected a few seconds for the two files), the cache fill, the dataset build and the self-test time (the self-test reads 256 MiB back over PCIe and hashes it on the CPU, so a few hundred ms). These go into docs/bench-log.md with the date and the driver version.

If it fails

  • error 0 nvrtc64_120_0.dll (and nvrtc-builtins64_128.dll) must sit next to ...: the zip was not extracted whole.
  • error 0 the CUDA driver library (nvcuda.dll) is not installed or the driver library lacks ...: NVIDIA driver missing or too old. Update the driver.
  • error 0 nvrtcCompileProgram kernel.cu for sm_120: NVRTC_ERROR_COMPILATION: <log>: the compiler rejected the pack's text. The log is on that line (the first 600 characters). Send the whole miner log: this is the one thing the Mac could not try, and the fix is in the stub headers the worker hands NVRTC (proto-cuda/nvrtc/worker.cpp, STUB_CUDA_RUNTIME and STUB_CSTDINT).
  • error 0 pack ...: self-test FAIL (...): the compiled program ran but gave other numbers than the CPU reference. Send the miner log; the line says which check failed and the first differing lane.
  • error 0 cuModuleLoadData kernel.cu (sm_120): ...: the driver refused the cubin (driver older than the NVRTC target). The prebuilt worker picks the card's own architecture by itself (CUDA_ARCH is for the nvcc path only); WORKER_ARCH=compute_90 at the top of the bat makes it emit PTX for the driver to compile instead. Send the log.
  • The dashboard's the prebuilt RTX 5090 worker is not ready after 150 s: with CUDA Toolkit and Visual Studio still installed on this PC the launcher then builds the old worker with nvcc and mining continues on it (the card line changes to "worker built here with nvcc + MSVC"); the miner log of the first attempt holds the reason.
  • FORCE_BUILD=1 at the top of the bat forces the old nvcc / cl.exe path for a comparison run.

The fault guard (second field run)

PC 2's gfx1036 ran the OpenCL worker correctly for 577 s, then every job "completed" in half a millisecond with no hash: the runtime answered every call with success without running the kernel. Three guards now sit on that path, and this is what they print:

  • the worker (amd-.log, from the worker): error <job> worker fault: <what was seen>; exiting 3 so the miner restarts the worker, then the miner's worker exited (code Some(3)); restarting it in N s. The "what was seen" is one of: an OpenCL call failing (name and code), the dispatch event not CL_COMPLETE, a chunk 20x faster per nonce than the running mean, or the output buffer unchanged since the previous dispatch. Every 200 jobs the worker prints info stats jobs N ... events created X released Y live Z; buffers created ... live ...: live should stay at 0 for events and 4 for buffers (cache, dataset, out, init words; 6 while a prepared pair is resident).
  • the miner (same log): WORKER FAULT job ... reported done in ... ms, Nx faster than the running mean ...; restarting the worker (fault 1) or the 10x interval form; the STATUS line carries faults=N and its rates are rolled back to the last report, so the dashboard never shows the fake GH/s again.
  • the launcher: the card's cell says worker fault in red and restarting where the rate was, the events list <card> worker fault, restarting: ..., and the 30-s status block says worker faults N.

If the fault returns on the gfx1036, the first worker fault line names which guard fired and that is the clue to the runtime's failure mode; please send the amd log around it.

The seed mismatch guard (third field run, PC 2 at the 14:20 boundary)

What happened: the app reinstalled at 14:20 and exported packs\devnet while its own node was still in IBD inside the previous epoch (DAA 17,881, seed 57ac...); the boundary at 18,000 passed a minute later. The CUDA miner saw the next seed within lead, sent prepare 2 s after start and swapped at 14:21:28 (fine). The OpenCL worker started 43 s later, when the templates were already on c23e... with no next epoch within lead, so no prepare was ever sent and every job was answered error N epoch seed mismatch for the rest of the run. PC 2's node never predicted a different epoch.

Now: a worker that lacks the job's pair prints need <epoch hex> <day hex> before that error; the miner, on a need line or three mismatches in a row, prints WORKER FAULT seed mismatch: ... preparing the current pair epoch ... for it, writes the pack under packs\prepare\<epoch16>-<day> and sends prepare; the worker answers prepared ... self-test PASS and the next job switches (info switched to the prepared pair). A worker without prepare support gets WORKER FAULT seed mismatch, restarting on epoch ... and the miner exits 42 (the launcher re-exports the pack). A ready worker with jobs queued and no done for 60 s gets WORKER FAULT no job completed for N s ...; restarting the worker. Expect at most one WORKER FAULT seed mismatch per worker start, and then mining; faults= on STATUS counts them.

What to send back

The launcher log, the two miner logs (they are uploaded every minute as well) and, for the bench log, the three timing lines above plus nvidia-smi's driver version.