From 27bcb8149426a949557b2d63f0338070ad74e07f Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Sun, 4 Oct 2026 19:08:45 +0000 Subject: [PATCH 1/5] Variant racing: the NVRTC and Metal workers compile several kernel variants at every prepare and keep the fastest for the hour the project lead, 4 Oct 2026 evening: "we need to make our miner better than anything else can be". The compile-ahead pipeline built one kernel per program; it now builds a catalogue (unroll 2 or 8; the dataset load path __ldg, __ldcg, __ldcs on NVIDIA; a register budget by --maxrregcount or __launch_bounds__, max_total_threads_per_threadgroup on Apple; 2, 4, 8 warps per block or 64 to 256 threads per threadgroup; combinations), self-tests each against the pack's vector warps (NVIDIA) or the base kernel over 2^16 nonces (Metal), bit for bit or out, and times each for about two seconds with the job loop paused (one mutex, mining resumes between variants). Base is the pack's text as shipped, always first, never discarded; a race has a budget (default 120 s against the 600-DAA lead) and keeps the best so far when it runs out, so the swap is never delayed. One line per race: variants, MH/s each, winner, gain, time. A tuning file (--tuning, IGNEUM_TUNING_FILE) pins a variant or orders the candidates per card model. NVIDIA: textual rewrites on the pack's own kernel_bound.cu with exact anchors from igneum-pow's emitter, so the pack format, the miner and igneum-pow are untouched and old packs race. --race --pack