7.5 KiB
igneum-bench-cuda (proto-cuda)
The NVIDIA twin of proto-metal. It runs the same random-program proof-of-work kernels on a CUDA GPU and checks
them bit for bit against results produced on the Mac.
This is a test harness, not a miner. No pool, no network, no wallet, no mining protocol. It fills a dataset, checks the GPU against known answers, and times the kernel. Nothing here earns anything.
Status on 3 October 2026: packs exported and cross-checked on the Mac, CUDA run on the RTX 5090 pending. No NVIDIA hash rate has been measured. Any figure you see for NVIDIA in this repo before the 5090 run is wrong.
Layout
proto-cuda/
host.cu host program: device info, dataset fill, self-test, vectors, bench, size sweep
build.sh Linux build (nvcc)
build.bat Windows build (nvcc + Visual Studio Build Tools)
CHECKLIST.md Metal/CUDA equivalence, op by op, and what was verified where
packs/<seed>/ one program pack per seed, written by proto-metal/igneum-bench --export-pack
kernel.cu the program as a CUDA kernel, plus fill kernel and host launch wrappers
program.h seed, day words, dataset size, loads per hash, wrapper declarations
vectors.h expected outputs for 3 warps (96 x 64-bit) and dataset self-test values
program.json the instruction list and all constants, for any other implementation
vectors.json the same vectors as JSON
program.metal the Metal source the Mac ran, for diffing by eye
emu/ CPU emulation shim: compile and check a pack with plain clang++/g++, no GPU
Two packs are checked in: igneum-genesis (104 loads per hash) and igneum-hourly (128 loads per hash).
Both were cross-checked on the Mac's Metal GPU before being written.
Prerequisites
Linux
- An NVIDIA driver recent enough for the toolkit. For CUDA 12.8 that is the R570 series or newer (approximate,
from memory;
nvidia-smiprints the driver's maximum supported CUDA version in its header). - CUDA Toolkit 12.8 or newer. Blackwell (
sm_120, RTX 50 series) is not known to older toolkits. - A host compiler the toolkit supports (gcc 11 to 13 for 12.8, approximate).
Windows
- The same driver requirement.
- CUDA Toolkit 12.8 or newer. Tick the Visual Studio integration in the installer.
- Visual Studio 2022 Build Tools with the "Desktop development with C++" workload. nvcc needs
cl.exe. - Run
build.batfrom an "x64 Native Tools Command Prompt for VS 2022" socl.exeis on PATH.
Build
Linux:
cd proto-cuda
./build.sh # pack igneum-genesis, -arch=sm_120
./build.sh igneum-hourly # the second pack
./build.sh igneum-genesis native # if sm_120 is refused, let nvcc pick the installed GPU
Windows (x64 Native Tools Command Prompt):
cd proto-cuda
build.bat
build.bat igneum-hourly
build.bat igneum-genesis native
Both scripts run this one command (paths adjusted for the pack):
nvcc -O3 -std=c++17 -arch=sm_120 -I packs/igneum-genesis -o igneum-bench-cuda-igneum-genesis host.cu packs/igneum-genesis/kernel.cu
Notes
-arch=sm_120is Blackwell.-arch=native(CUDA 11.6 or newer) compiles for whatever GPU is in the machine and is the fallback if the toolkit is too old to knowsm_120(which means it is too old for a 5090 anyway: upgrade).- If nvcc on Windows refuses the Visual Studio version, add
-allow-unsupported-compilerto the nvcc line. - The host code is plain C++17 and the CUDA runtime API. No NVRTC, no third-party libraries, no JSON parser. The kernel is compiled ahead of time from the pack.
Run
./igneum-bench-cuda-igneum-genesis # 1 GiB dataset, 5 batches x 2^24 nonces, vectors checked
./igneum-bench-cuda-igneum-genesis --sweep # 4, 64, 256, 512, 1024 MiB in sequence (the Mac's sweep)
./igneum-bench-cuda-igneum-genesis --block-warps 4 # 4 warps per block instead of 1 (still bit-exact)
./igneum-bench-cuda-igneum-hourly # the second program
On Windows the binaries are igneum-bench-cuda-igneum-genesis.exe and so on.
Flags: --dataset-mib N (power of two, default 1024), --sweep, --batch-log2 B (default 24),
--batches N (default 5), --block-warps W (default 1, mirrors the Metal run's one SIMD group per threadgroup),
--device D.
What it prints, in order:
- GPU name, SM count, memory, clocks, L2, warp size, driver and runtime versions, registers per thread and resident warps per SM for the kernel.
- Dataset fill time (twice; the Mac saw a first-touch cost on the first fill) and write GB/s.
- Dataset self-test: 16 head words and word
[MASK]against values from the Mac, 64 random words against the host formula. - Vectors: 3 warps (base nonces 0, 4096, 1000000), each run standalone as one 32-thread block, then again read out of the warm-up batch so the bench configuration itself is checked. PASS or FAIL per warp, with the first differing lane printed on FAIL.
- Timing: 5 batches of 2^24 hashes after a warm-up batch, GPU event time and wall time, Mhash/s, hashes/s, GB/s useful (loads per hash x 4 bytes x hashes/s, the same definition as the Mac's table).
- A summary table in Markdown and
OVERALL: PASSorFAIL. Exit code 0 on PASS, 1 on FAIL, 2 on a CUDA error.
Vectors are only checked when the dataset is the pack's size (1024 MiB), because the outputs depend on the address mask. At other sizes the table says "skipped (not pack size)" and only the dataset self-test counts.
What PASS means
- The CUDA fill kernel produced the same dataset as the Mac's closed-form function (sampled, not every word).
- For 96 nonces spread across the nonce space, the RTX 5090 produced the same 64-bit outputs as the Mac's CPU interpreter, which had itself matched the Mac's Metal GPU. The random program, the register init, the warp shuffles, the multiply-high and rotates, and the dataset addressing all agree between Apple and NVIDIA.
- With
--block-warps Wthe in-batch check passing shows that packing W warps per block changed nothing.
A FAIL with a small number of differing lanes points at a shuffle; a FAIL in every lane points at an arithmetic op or the dataset. Send the whole printout either way.
Sending results back
Copy the printed header lines (GPU, CUDA versions, kernel line, program line) and the summary table into
docs/bench-log.md under a dated heading, together with the output of nvcc --version and the driver version
from nvidia-smi. Keep the full stdout as well. Run both packs and the sweep so the log has the same shape as
the Mac's entry. Do not edit the numbers; if a run looks odd, run it again and log both.
Checking a pack without a GPU
emu/emu.sh <pack> [flags] compiles host.cu and the pack's kernel.cu as plain C++17 against a shim
cuda_runtime.h and runs the kernels on host threads (32 per warp, a barrier inside __shfl_xor_sync). Use
small batches (--batch-log2 13 --batches 1). Only PASS/FAIL matters; the rates it prints are noise.
This is how the CUDA text was checked on the Mac on 3 October 2026 (both packs PASS, see CHECKLIST.md).
It is not an nvcc build and says nothing about NVIDIA hardware.
Regenerating a pack
On the Mac:
cd proto-metal
swiftc -O -o igneum-bench main.swift -framework Metal
./igneum-bench --seed igneum-genesis --export-pack ../proto-cuda/packs/igneum-genesis
The exporter runs the CPU interpreter for the three vector warps, runs the Metal kernel for the same warps,
and refuses to write anything unless all 96 outputs match. --day and --dataset-log2 change the dataset
constants and are recorded in the pack.