the project lead, 4 Oct 2026 evening: "we need to make our miner better than anything else can be". The compile-ahead pipeline
built one kernel per program; it now builds a catalogue (unroll 2 or 8; the dataset load path __ldg, __ldcg, __ldcs
on NVIDIA; a register budget by --maxrregcount or __launch_bounds__, max_total_threads_per_threadgroup on Apple;
2, 4, 8 warps per block or 64 to 256 threads per threadgroup; combinations), self-tests each against the pack's
vector warps (NVIDIA) or the base kernel over 2^16 nonces (Metal), bit for bit or out, and times each for about two
seconds with the job loop paused (one mutex, mining resumes between variants). Base is the pack's text as shipped,
always first, never discarded; a race has a budget (default 120 s against the 600-DAA lead) and keeps the best so
far when it runs out, so the swap is never delayed. One line per race: variants, MH/s each, winner, gain, time.
A tuning file (--tuning, IGNEUM_TUNING_FILE) pins a variant or orders the candidates per card model.
NVIDIA: textual rewrites on the pack's own kernel_bound.cu with exact anchors from igneum-pow's emitter, so the
pack format, the miner and igneum-pow are untouched and old packs race. --race --pack <dir> runs the race alone.
Under IGNEUM_EMU the race is off (the stand-in checks the handed-over text is the pack's). Metal: the MSL hooks in
generateMSL, a lock-protected kernel slot per program, a pair compiled inline races after its first job,
--race-test --seed --day runs the race alone with a table. OpenCL is not raced yet. docs/design/miner-tuning.md.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Root cause from the uploads (bench-log entry): the app exported the pack while its node was in IBD inside the previous
epoch, the OpenCL worker started after the boundary with no next epoch within lead, so no prepare was ever sent and
every job was a seed mismatch; the CUDA worker on the same PC had swapped correctly. Both workers now print
need <epoch> <day> before the error; the devnet-v4 miner (3bfe346f) prepares the current pair on a need line or three
mismatches, exits 42 for a worker without prepare support, and restarts a ready worker that completes no job for
60 s with jobs queued. Package rebuilt with the guarded miner (ship build on dc749905), payload inputs published.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
proto-cuda/nvrtc/worker.cpp serves igneum-miner's worker protocol with nothing installed but the NVIDIA driver:
nvcuda.dll and the redistributable nvrtc64_120_0.dll are opened with LoadLibrary (cuda_api.h), the pack's kernel.cu
and kernel_bound.cu are handed to NVRTC byte for byte up to the host launch wrappers with program.h and memhard.h as
named headers, every pack is self-tested against its vectors.h (cache head, last line, FNV-1a 64, dataset head, last
word, 64 samples, 96 vector lanes) before it serves a job, prepare runs on a thread for the hourly swap, and a job on
seeds without a pair makes the worker find the miner's pack by seeds.txt and build it. packfile.h (C99) reads a pack
directory. fetch-redist.sh verifies and stages the NVIDIA 12.8.93 redistributables and the Khronos headers
(THIRD-PARTY.md records URLs, hashes and the EULA clause); build-windows.sh cross-compiles with mingw (static,
KERNEL32 + Universal CRT only). emu/: the two libraries as host functions, the pack kernels on host threads, the NVRTC
source compared with the pack files; test.sh PASS on the Mac (two real packs, swap, self-heal, 17 sampled hashes equal
igneum-pow hash-bound). Not run here: the real NVRTC compile and driver load (RTX 5090 PC).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>