igneum/proto-cuda/windows-app/TEST.md
igneum-labs 4714b510f2 OpenCL worker fault guard, miner-side fault state in the launcher, NVRTC annotation rule in the emulation
After PC 2's gfx1036 (prebuilt-generic path) completed 2,000 jobs a second with no hash from 600 s on: every OpenCL
call in host.c's job path is now fatal on error (exit 3, the miner restarts the worker), the dispatch event must read
CL_COMPLETE, a chunk 20x faster per nonce than the running mean or an output buffer unchanged since the previous
dispatch is a fault, and a stats line every 200 jobs carries the live event and buffer counts (a leak over 16 events
or 12 buffers is fatal too). IGNEUM_FAULT_TEST=N exercises the detectors on a healthy device (verified on Apple
OpenCL: the stale-output guard fires on the chunk after the injected fault). The launcher shows 'worker fault' and
'restarting' for the card on the miner's WORKER FAULT line and drops the last rate. The emulation's NVRTC stand-in
now applies NVRTC's execution-space rule (program.h(46) igneum_launch_* declarations are host code unless
-default-device or -DIGNEUM_NO_CUDA), which is what the RTX 5090 reported; test.sh checks the rejection.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 10:42:36 +00:00

7.8 KiB

First run of package 0.3.0 (prebuilt workers): what to look for

Written 4 October 2026 on the Mac, where the workers could only be checked without an NVIDIA GPU (the CUDA worker's host logic through a CPU emulation of the driver API and NVRTC, the OpenCL worker through Apple's OpenCL runtime). The real NVRTC compile and the real driver run happen for the first time on the RTX 5090 PC. This page says what a good run prints and what to send back if it does not.

Before you start

Extract the zip to a fresh short folder (C:\igneum-v4). Nothing to install. These files must sit together in it: igneum-worker-cuda.exe, nvrtc64_120_0.dll, nvrtc-builtins64_128.dll, igneum-worker-opencl.exe, igneum-miner.exe, igneumd.exe. The NVIDIA driver must be the one that already mines (617.14 on 4 October 2026 is fine; the driver API needs a driver for CUDA 12.x).

What the dashboard should show

  1. In the launcher log (igneum-.log) right after "GPUs found": prebuilt workers: igneum-worker-cuda.exe with nvrtc64_120_0.dll, igneum-worker-opencl.exe and workers to start: nvidia (prebuilt worker, NVRTC compiles on the card (driver only)), amd (prebuilt worker, the OpenCL driver compiles (driver only)). No "tools:" line: nothing is looked for when the prebuilt workers are there.
  2. After the node is synced: the phase line starting the prebuilt RTX 5090 worker, then the card appears with prebuilt worker, NVRTC compiles on the card (driver only) next to its name.
  3. Within about 10 s the event RTX 5090 worker ready (prebuilt worker, NVRTC compiles on the card (driver only)), hot swap at the hour boundary. The AMD card the same with "the OpenCL driver compiles".
  4. Then mining in green and a hash rate about the one the nvcc-built worker gave (229 MH/s on the 5090 with 8 identities on 3 October 2026; the kernel text is the same, only the compiler changed, so a large gap either way is worth a note).
  5. About ten minutes before the hour boundary: RTX 5090: the next hourly program is compiled and resident (hot swap ready). At the boundary no pause, no restart event; the miner log says switched to the prepared pair.

What the miner log should show (nvidia-.log)

The worker's own lines, in this order:

worker: info igneum-worker-cuda 1.0 (4 October 2026): device 0 NVIDIA_GeForce_RTX_5090 (sm_120, 170 SMs), driver 12.x from nvcuda.dll, NVRTC 12.8 from C:\igneum-v4\nvrtc64_120_0.dll, target sm_120 (the device's architecture, listed by NVRTC)
worker: ready cuda NVIDIA_GeForce_RTX_5090 pack igneum-epoch/<64 hex>/day/<hex> dataset-log2 28 batch 4194304 regs <N> prepare 1 path nvrtc 12.8 driver 12.x arch sm_120 worker 1.0 (4 October 2026)  (prepare support: yes, hot swap at the boundary)
worker: info first pack proto-cuda\packs\devnet: nvrtc <ms> cache <ms> dataset <ms> check <ms> ms; self-test PASS (cache head, last line and FNV-1a 64 <16 hex>; dataset head, word [268435455] and 64 samples; 96 of 96 vector lanes)

The self-test is the proof that NVRTC compiled the same program nvcc did: the 96 vector lanes are the outputs the Rust CPU interpreter put in the pack, and the cache FNV covers all 2^26 words of the cache. regs should be close to what the nvcc build reported for the same program. Every found is still re-checked by the miner's CPU code (WORKER MISMATCH in the log would be the alarm; expect none).

At the hourly prepare:

PREPARE sent for epoch seed ... (N DAA blocks before the boundary ...)
worker: info prepare started for epoch <16 hex> day <hex> from C:\igneum-v4\proto-cuda\packs\prepare\<epoch>-<day> (NVRTC sm_120 in the background)
worker: prepared <epoch hex> <day hex> <ms> nvrtc <ms> cache <ms> dataset <ms> check <ms> ms; self-test PASS (...) resident 2 programs 2 datasets

For the AMD worker (amd-.log): ready opencl <device> platform AMD_Accelerated_Parallel_Processing pack ... prepare 1 path prebuilt-generic, then info first pack ..\proto-cuda\packs\devnet: cache <ms> dataset <ms> check <ms> ms (...); self-test PASS (...).

Timings worth writing down (none measured yet on NVIDIA)

From the info first pack and prepared lines: the NVRTC compile time (nvrtc <ms>, expected a few seconds for the two files), the cache fill, the dataset build and the self-test time (the self-test reads 256 MiB back over PCIe and hashes it on the CPU, so a few hundred ms). These go into docs/bench-log.md with the date and the driver version.

If it fails

  • error 0 nvrtc64_120_0.dll (and nvrtc-builtins64_128.dll) must sit next to ...: the zip was not extracted whole.
  • error 0 the CUDA driver library (nvcuda.dll) is not installed or the driver library lacks ...: NVIDIA driver missing or too old. Update the driver.
  • error 0 nvrtcCompileProgram kernel.cu for sm_120: NVRTC_ERROR_COMPILATION: <log>: the compiler rejected the pack's text. The log is on that line (the first 600 characters). Send the whole miner log: this is the one thing the Mac could not try, and the fix is in the stub headers the worker hands NVRTC (proto-cuda/nvrtc/worker.cpp, STUB_CUDA_RUNTIME and STUB_CSTDINT).
  • error 0 pack ...: self-test FAIL (...): the compiled program ran but gave other numbers than the CPU reference. Send the miner log; the line says which check failed and the first differing lane.
  • error 0 cuModuleLoadData kernel.cu (sm_120): ...: the driver refused the cubin (driver older than the NVRTC target). The prebuilt worker picks the card's own architecture by itself (CUDA_ARCH is for the nvcc path only); WORKER_ARCH=compute_90 at the top of the bat makes it emit PTX for the driver to compile instead. Send the log.
  • The dashboard's the prebuilt RTX 5090 worker is not ready after 150 s: with CUDA Toolkit and Visual Studio still installed on this PC the launcher then builds the old worker with nvcc and mining continues on it (the card line changes to "worker built here with nvcc + MSVC"); the miner log of the first attempt holds the reason.
  • FORCE_BUILD=1 at the top of the bat forces the old nvcc / cl.exe path for a comparison run.

The fault guard (second field run)

PC 2's gfx1036 ran the OpenCL worker correctly for 577 s, then every job "completed" in half a millisecond with no hash: the runtime answered every call with success without running the kernel. Three guards now sit on that path, and this is what they print:

  • the worker (amd-.log, from the worker): error <job> worker fault: <what was seen>; exiting 3 so the miner restarts the worker, then the miner's worker exited (code Some(3)); restarting it in N s. The "what was seen" is one of: an OpenCL call failing (name and code), the dispatch event not CL_COMPLETE, a chunk 20x faster per nonce than the running mean, or the output buffer unchanged since the previous dispatch. Every 200 jobs the worker prints info stats jobs N ... events created X released Y live Z; buffers created ... live ...: live should stay at 0 for events and 4 for buffers (cache, dataset, out, init words; 6 while a prepared pair is resident).
  • the miner (same log): WORKER FAULT job ... reported done in ... ms, Nx faster than the running mean ...; restarting the worker (fault 1) or the 10x interval form; the STATUS line carries faults=N and its rates are rolled back to the last report, so the dashboard never shows the fake GH/s again.
  • the launcher: the card's cell says worker fault in red and restarting where the rate was, the events list <card> worker fault, restarting: ..., and the 30-s status block says worker faults N.

If the fault returns on the gfx1036, the first worker fault line names which guard fired and that is the clue to the runtime's failure mode; please send the amd log around it.

What to send back

The launcher log, the two miner logs (they are uploaded every minute as well) and, for the bench log, the three timing lines above plus nvidia-smi's driver version.