Commit graph

6 commits

Author SHA1 Message Date
igneum-labs
382039ea5d Variant race: bench-log entry (Metal on the M5 Max, 14 variants, two programs, serve check), the base-only short-circuit, the PC 1 zip's hash in the plan
Under emulation (or --race with no other name) the race no longer times base alone: emu/test.sh passes again
(9 source checks, the serve protocol with prepare, swap and self-heal, 17 sampled hashes equal to igneum-pow).
Mac table: g256 (256 threads per threadgroup) +17.3% and +21.2% over the shipped 32-thread groups on two
programs, with the live app's worker sharing the GPU; the serve check found the unfair mutex (fixed in 123ee1a).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 20:37:27 +00:00
igneum-labs
123ee1a2ee Variant race: a 150 ms pause between timed windows so the job loop gets the card (the mutex is not fair; a queued job waited the whole race, 36 s, in the first serve check)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 19:18:04 +00:00
igneum-labs
27bcb81494 Variant racing: the NVRTC and Metal workers compile several kernel variants at every prepare and keep the fastest for the hour
the project lead, 4 Oct 2026 evening: "we need to make our miner better than anything else can be". The compile-ahead pipeline
built one kernel per program; it now builds a catalogue (unroll 2 or 8; the dataset load path __ldg, __ldcg, __ldcs
on NVIDIA; a register budget by --maxrregcount or __launch_bounds__, max_total_threads_per_threadgroup on Apple;
2, 4, 8 warps per block or 64 to 256 threads per threadgroup; combinations), self-tests each against the pack's
vector warps (NVIDIA) or the base kernel over 2^16 nonces (Metal), bit for bit or out, and times each for about two
seconds with the job loop paused (one mutex, mining resumes between variants). Base is the pack's text as shipped,
always first, never discarded; a race has a budget (default 120 s against the 600-DAA lead) and keeps the best so
far when it runs out, so the swap is never delayed. One line per race: variants, MH/s each, winner, gain, time.
A tuning file (--tuning, IGNEUM_TUNING_FILE) pins a variant or orders the candidates per card model.

NVIDIA: textual rewrites on the pack's own kernel_bound.cu with exact anchors from igneum-pow's emitter, so the
pack format, the miner and igneum-pow are untouched and old packs race. --race --pack <dir> runs the race alone.
Under IGNEUM_EMU the race is off (the stand-in checks the handed-over text is the pack's). Metal: the MSL hooks in
generateMSL, a lock-protected kernel slot per program, a pair compiled inline races after its first job,
--race-test --seed --day runs the race alone with a table. OpenCL is not raced yet. docs/design/miner-tuning.md.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 19:08:45 +00:00
igneum-labs
4d907d8a14 Workers answer a job they have no pair for with a need line; the miner prepares the current pair (PC 2 stuck on the previous epoch)
Root cause from the uploads (bench-log entry): the app exported the pack while its node was in IBD inside the previous
epoch, the OpenCL worker started after the boundary with no next epoch within lead, so no prepare was ever sent and
every job was a seed mismatch; the CUDA worker on the same PC had swapped correctly. Both workers now print
need <epoch> <day> before the error; the devnet-v4 miner (3bfe346f) prepares the current pair on a need line or three
mismatches, exits 42 for a worker without prepare support, and restarts a ready worker that completes no job for
60 s with jobs queued. Package rebuilt with the guarded miner (ship build on dc749905), payload inputs published.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 13:32:58 +00:00
igneum-labs
734dbd134f NVRTC worker: -default-device (NVRTC rejects unannotated pack helpers as host code; found on the first real card)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 10:12:35 +00:00
igneum-labs
f535abb59d One-click CUDA worker: driver API + NVRTC loaded at run time, pack read at run time, Mac emulation check
proto-cuda/nvrtc/worker.cpp serves igneum-miner's worker protocol with nothing installed but the NVIDIA driver:
nvcuda.dll and the redistributable nvrtc64_120_0.dll are opened with LoadLibrary (cuda_api.h), the pack's kernel.cu
and kernel_bound.cu are handed to NVRTC byte for byte up to the host launch wrappers with program.h and memhard.h as
named headers, every pack is self-tested against its vectors.h (cache head, last line, FNV-1a 64, dataset head, last
word, 64 samples, 96 vector lanes) before it serves a job, prepare runs on a thread for the hourly swap, and a job on
seeds without a pair makes the worker find the miner's pack by seeds.txt and build it. packfile.h (C99) reads a pack
directory. fetch-redist.sh verifies and stages the NVIDIA 12.8.93 redistributables and the Khronos headers
(THIRD-PARTY.md records URLs, hashes and the EULA clause); build-windows.sh cross-compiles with mingw (static,
KERNEL32 + Universal CRT only). emu/: the two libraries as host functions, the pack kernels on host threads, the NVRTC
source compared with the pack files; test.sh PASS on the Mac (two real packs, swap, self-heal, 17 sampled hashes equal
igneum-pow hash-bound). Not run here: the real NVRTC compile and driver load (RTX 5090 PC).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-04 09:55:05 +00:00