igneum-common.ps1 uses igneum-worker-cuda.exe (with nvrtc64_*_0.dll next to it) and igneum-worker-opencl.exe when they are in the folder, passes the exported pack with --pack, and only surveys the toolchain (Find-Toolchain) when a worker is missing or FORCE_BUILD=1; the dashboard and the status block name the path per card; a prebuilt worker that is not ready after 150 s falls back to the build path once when a toolchain exists; 90 s of seed mismatch errors re-export the pack and restart the vendor; WORKER_ARCH overrides the NVRTC target. make-package.sh ships the two exes, the two NVRTC DLLs, the licence texts, THIRD-PARTY.md and TEST.md (what the first RTX 5090 run should print and what to send back). README.txt, the bats, proto-cuda/README.md (nvrtc/ section), WINDOWS-MINER.md and the bench log updated with what the Mac measured. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
227 lines
15 KiB
Markdown
227 lines
15 KiB
Markdown
# igneum-bench-cuda (proto-cuda)
|
|
|
|
The NVIDIA twin of `proto-metal`. It runs the same random-program proof-of-work kernels on a CUDA GPU and checks
|
|
them bit for bit against results produced on the Mac.
|
|
|
|
This is a test harness, not a miner. No pool, no network, no wallet, no mining protocol. It fills a dataset,
|
|
checks the GPU against known answers, and times the kernel. Nothing here earns anything.
|
|
|
|
Status on 3 October 2026: the two closed-form packs ran on the RTX 5090 (96/96 vectors PASS each, see
|
|
`docs/bench-log.md`). The memory-hard pack `igneum-genesis-mh` (added later the same day, construction in
|
|
`../proto-metal/MEMHARD.md`) has passed only the clang emulation on the Mac; its 5090 and AMD runs are pending, and no
|
|
NVIDIA figure for the memory-hard dataset exists yet.
|
|
|
|
Status on 4 October 2026, generator version 2: every pack was regenerated by `igneum-pow export` (the Rust crate is
|
|
now the pack source; the Swift exporter is a cross-check) with the adopted generator rule (exactly 16 loads per
|
|
program, fresh-source loads, the acceptance rule of spec 01 section 1.4.6). All four packs pass the clang emulation
|
|
(`emu/emu.sh`, 96/96 standalone, 2 warps per block in batch) and Apple OpenCL on the M5 Max; Metal cross-checked the
|
|
exported genesis pack and a 2,000-program fuzz. The version 1 vectors, including the 5090's 192 of 192 from
|
|
3 October, are retired: the kernel text is unchanged, so those runs remain evidence that the ops agree on NVIDIA, but
|
|
the first 5090 run on a version 2 pack is still owed (`docs/bench-log.md`, 4 October 2026 entry). `program.h` now
|
|
carries `IGNEUM_GENERATOR 2`, `IGNEUM_PROGRAM_ATTEMPT` and `IGNEUM_PROGRAM_ID`; a worker built from a version 1 pack
|
|
cannot serve a version 2 node.
|
|
|
|
## The one-click worker (nvrtc/, 4 October 2026)
|
|
|
|
`nvrtc/worker.cpp` is the worker the Windows package ships as `igneum-worker-cuda.exe`: nothing to install but the
|
|
NVIDIA driver. It opens `nvcuda.dll` (the driver API, in every driver) and `nvrtc64_120_0.dll` (NVIDIA's runtime
|
|
compiler, redistributable, shipped next to the exe with `nvrtc-builtins64_128.dll`; `nvrtc/THIRD-PARTY.md`) with
|
|
`LoadLibrary`/`GetProcAddress` (`nvrtc/cuda_api.h`), reads a pack directory at run time (`nvrtc/packfile.h`: program.h,
|
|
seeds.txt, vectors.h), hands NVRTC the pack's `kernel.cu` and `kernel_bound.cu` up to the host launch wrappers with
|
|
`program.h` and `memhard.h` as named headers, byte for byte, loads the cubin through the driver, fills the cache, builds
|
|
the dataset and self-tests both plus the three vector warps against `vectors.h` before it serves a job. It speaks the
|
|
same `--serve` protocol as `host.cu` (jobs, `prepare` in the background for the hourly swap, `found`/`done`); a job on
|
|
seeds it has no pair for makes it look for the miner's pack by `seeds.txt` and build it in the foreground. `host.cu`
|
|
stays the ahead-of-time harness and the launcher's fallback when the prebuilt worker is missing and a toolkit exists.
|
|
|
|
```
|
|
nvrtc/
|
|
worker.cpp the worker (C++17, driver API + NVRTC loaded at run time, cross-compiled with mingw)
|
|
cuda_api.h the function-pointer table for the two libraries
|
|
packfile.h C99 reader for a pack directory: sizes, seeds, vectors; the self-test verdict (shared with proto-opencl/host.c --pack)
|
|
fetch-redist.sh downloads and verifies cuda.h, nvrtc.h, the two NVRTC DLLs and the Khronos OpenCL headers into nvrtc/redist/ (gitignored)
|
|
build-windows.sh cross-compiles igneum-worker-cuda.exe and proto-opencl/igneum-worker-opencl.exe (static, KERNEL32 + Universal CRT only)
|
|
THIRD-PARTY.md every third-party file, its URL, version, SHA-256 and licence clause
|
|
emu/ the Mac check: emu_backend.cpp (driver API and NVRTC as host functions, the pack's kernels on host threads,
|
|
the NVRTC source compared byte for byte with the pack), test.sh, serve-check.sh (shared with proto-opencl/test-generic.sh)
|
|
```
|
|
|
|
What was proven on the Mac (4 October 2026): `emu/test.sh` builds the worker against two real packs (the devnet epoch
|
|
0 pack and a second epoch and day exported by `igneum-pow`), runs `--check` (self-test PASS, 96 of 96 lanes) and a
|
|
`--serve` session: 64 + 64 + 32 found on pack A including the nonces either side of the 32-bit boundary, pack B
|
|
prepared in the background with its self-test PASS, the swap, the self-heal rebuild of pack A, 17 sampled hashes equal
|
|
to `igneum-pow hash-bound`, and the NVRTC source check PASS for all four files (kernel.cu 7,385 of 8,893 bytes with
|
|
the 1,508-byte host wrapper tail dropped, kernel_bound.cu 6,003 of 6,887, program.h and memhard.h identical). The exe
|
|
cross-compiles and links with mingw. Not proven here: that the real NVRTC accepts the text with the two stub headers
|
|
and that the driver runs the cubin; that is the RTX 5090 run (`windows-app/TEST.md`).
|
|
|
|
## Layout
|
|
|
|
```
|
|
proto-cuda/
|
|
host.cu host program: device info, dataset fill, self-test, vectors, bench, size sweep
|
|
build.sh Linux build (nvcc)
|
|
build.bat Windows build (nvcc + Visual Studio Build Tools)
|
|
CHECKLIST.md Metal/CUDA equivalence, op by op, and what was verified where
|
|
packs/<seed>/ one program pack per seed, written by igneum-pow export (generator v2, 4 October 2026)
|
|
kernel.cu the program as a CUDA kernel, plus fill kernel and host launch wrappers
|
|
kernel.cl the same program as OpenCL C for proto-opencl (AMD and any other OpenCL device), built at runtime
|
|
program.h seed, day words, dataset size, loads per hash, wrapper declarations (C99-safe: proto-opencl/host.c includes it too)
|
|
vectors.h expected outputs for 3 warps (96 x 64-bit) and dataset self-test values
|
|
program.json the instruction list and all constants, for any other implementation
|
|
vectors.json the same vectors as JSON
|
|
program.metal the Metal source the Mac ran, for diffing by eye
|
|
memhard.h memory-hard packs only: the cache fill and item derivation core, compiled for device and host
|
|
memhard.metal memory-hard packs only: the Metal cache-fill and build kernels the Mac ran
|
|
emu/ CPU emulation shim: compile and check a pack with plain clang++/g++, no GPU
|
|
nvrtc/ the one-click worker (driver API + NVRTC at run time), its Mac emulation and the third-party record
|
|
```
|
|
|
|
Four packs are checked in, all generator version 2 with 128 loads per hash. `igneum-genesis` and `igneum-hourly` use the
|
|
original closed-form dataset (`IGNEUM_DATASET_MODE 0`, implied when the macro is absent). `igneum-genesis-mh` is the
|
|
same program as `igneum-genesis` over the memory-hard dataset (`IGNEUM_DATASET_MODE 1`): a 256 MiB cache of chained
|
|
ChaCha12 blocks filled on the GPU from the day key, and every 64-byte dataset item derived from 8 dependent cache
|
|
reads through a seed-parameterised mixer (`../proto-metal/MEMHARD.md`). The hash kernel text is identical in both
|
|
packs; only the dataset contents differ, so the 96 expected outputs differ. All three were cross-checked on the
|
|
Mac's Metal GPU before being written.
|
|
|
|
## Prerequisites
|
|
|
|
Linux
|
|
- An NVIDIA driver recent enough for the toolkit. For CUDA 12.8 that is the R570 series or newer (approximate,
|
|
from memory; `nvidia-smi` prints the driver's maximum supported CUDA version in its header).
|
|
- CUDA Toolkit 12.8 or newer. Blackwell (`sm_120`, RTX 50 series) is not known to older toolkits.
|
|
- A host compiler the toolkit supports (gcc 11 to 13 for 12.8, approximate).
|
|
|
|
Windows
|
|
- The same driver requirement.
|
|
- CUDA Toolkit 12.8 or newer. Tick the Visual Studio integration in the installer.
|
|
- Visual Studio 2022 Build Tools with the "Desktop development with C++" workload. nvcc needs `cl.exe`.
|
|
- Run `build.bat` from an "x64 Native Tools Command Prompt for VS 2022" so `cl.exe` is on PATH.
|
|
|
|
## Build
|
|
|
|
Linux:
|
|
|
|
```
|
|
cd proto-cuda
|
|
./build.sh # pack igneum-genesis, -arch=sm_120
|
|
./build.sh igneum-hourly # the second pack
|
|
./build.sh igneum-genesis-mh # the memory-hard pack (needs 256 MiB more device memory for the cache)
|
|
./build.sh igneum-genesis native # if sm_120 is refused, let nvcc pick the installed GPU
|
|
```
|
|
|
|
Windows (x64 Native Tools Command Prompt):
|
|
|
|
```
|
|
cd proto-cuda
|
|
build.bat
|
|
build.bat igneum-hourly
|
|
build.bat igneum-genesis-mh
|
|
build.bat igneum-genesis native
|
|
```
|
|
|
|
Both scripts run this one command (paths adjusted for the pack):
|
|
|
|
```
|
|
nvcc -O3 -std=c++17 -arch=sm_120 -I packs/igneum-genesis -o igneum-bench-cuda-igneum-genesis host.cu packs/igneum-genesis/kernel.cu
|
|
```
|
|
|
|
Notes
|
|
- `-arch=sm_120` is Blackwell. `-arch=native` (CUDA 11.6 or newer) compiles for whatever GPU is in the
|
|
machine and is the fallback if the toolkit is too old to know `sm_120` (which means it is too old for a 5090
|
|
anyway: upgrade).
|
|
- If nvcc on Windows refuses the Visual Studio version, add `-allow-unsupported-compiler` to the nvcc line.
|
|
- The host code is plain C++17 and the CUDA runtime API. No NVRTC, no third-party libraries, no JSON parser.
|
|
The kernel is compiled ahead of time from the pack.
|
|
|
|
## Run
|
|
|
|
```
|
|
./igneum-bench-cuda-igneum-genesis # 1 GiB dataset, 5 batches x 2^24 nonces, vectors checked
|
|
./igneum-bench-cuda-igneum-genesis --sweep # 4, 64, 256, 512, 1024 MiB in sequence (the Mac's sweep)
|
|
./igneum-bench-cuda-igneum-genesis --block-warps 4 # 4 warps per block instead of 1 (still bit-exact)
|
|
./igneum-bench-cuda-igneum-hourly # the second program
|
|
./igneum-bench-cuda-igneum-genesis-mh # memory-hard dataset: cache fill + build, cache check, vectors, bench
|
|
./igneum-bench-cuda-igneum-genesis-mh --sweep # the same sweep over the memory-hard dataset
|
|
```
|
|
|
|
On Windows the binaries are `igneum-bench-cuda-igneum-genesis.exe` and so on.
|
|
|
|
Flags: `--dataset-mib N` (power of two, default 1024), `--sweep`, `--batch-log2 B` (default 24),
|
|
`--batches N` (default 5), `--block-warps W` (default 1, mirrors the Metal run's one SIMD group per threadgroup),
|
|
`--device D`.
|
|
|
|
What it prints, in order:
|
|
1. GPU name, SM count, memory, clocks, L2, warp size, driver and runtime versions, registers per thread and
|
|
resident warps per SM for the kernel, and which dataset construction the pack uses.
|
|
2. Memory-hard packs only: cache fill time on the GPU (twice), cache fill time on the host (one thread, the same
|
|
`memhard.h` text), then the cache check: every one of the 2^26 words GPU versus host, the host FNV-1a 64 against
|
|
the Mac's, and the head and last line against the Mac's. A FAIL here stops nothing but fails OVERALL.
|
|
3. Dataset fill time (closed form, twice, with write GB/s) or dataset build time (memory-hard, twice, with items/s
|
|
and cache-line reads/s).
|
|
4. Dataset self-test: 16 head words and word `[MASK]` against values from the Mac, 64 random words against the
|
|
host formula (closed form) or the host derivation from the host cache (memory-hard), and the Mac's 64 sampled
|
|
words (memory-hard packs; those inside the current dataset size).
|
|
5. Vectors: 3 warps (base nonces 0, 4096, 1000000), each run standalone as one 32-thread block, then again
|
|
read out of the warm-up batch so the bench configuration itself is checked. PASS or FAIL per warp, with the
|
|
first differing lane printed on FAIL.
|
|
6. Timing: 5 batches of 2^24 hashes after a warm-up batch, GPU event time and wall time, Mhash/s, hashes/s,
|
|
GB/s useful (loads per hash x 4 bytes x hashes/s, the same definition as the Mac's table).
|
|
7. A summary table in Markdown and `OVERALL: PASS` or `FAIL`. Exit code 0 on PASS, 1 on FAIL, 2 on a CUDA error.
|
|
|
|
Vectors are only checked when the dataset is the pack's size (1024 MiB), because the outputs depend on the
|
|
address mask. At other sizes the table says "skipped (not pack size)" and only the dataset self-test counts.
|
|
|
|
## What PASS means
|
|
|
|
- Closed-form packs: the CUDA fill kernel produced the same dataset as the Mac's closed-form function (sampled, not
|
|
every word).
|
|
- Memory-hard pack: the CUDA cache-fill kernel produced, word for word, the same 256 MiB cache as the host and as the
|
|
Mac (FNV-1a 64, head, last line), and the CUDA build kernel produced the same dataset words as the host derivation
|
|
and the Mac's samples (sampled, not every word). The whole chain from day key to dataset word agrees across three
|
|
compilers (Apple Metal, host C++, NVIDIA CUDA).
|
|
- For 96 nonces spread across the nonce space, the RTX 5090 produced the same 64-bit outputs as the Mac's CPU
|
|
interpreter, which had itself matched the Mac's Metal GPU. The random program, the register init, the warp
|
|
shuffles, the multiply-high and rotates, and the dataset addressing all agree between Apple and NVIDIA.
|
|
- With `--block-warps W` the in-batch check passing shows that packing W warps per block changed nothing.
|
|
|
|
A FAIL with a small number of differing lanes points at a shuffle; a FAIL in every lane points at an arithmetic
|
|
op or the dataset. Send the whole printout either way.
|
|
|
|
## Sending results back
|
|
|
|
Copy the printed header lines (GPU, CUDA versions, kernel line, program line) and the summary table into
|
|
`docs/bench-log.md` under a dated heading, together with the output of `nvcc --version` and the driver version
|
|
from `nvidia-smi`. Keep the full stdout as well. Run both packs and the sweep so the log has the same shape as
|
|
the Mac's entry. Do not edit the numbers; if a run looks odd, run it again and log both.
|
|
|
|
## Checking a pack without a GPU
|
|
|
|
`emu/emu.sh <pack> [flags]` compiles `host.cu` and the pack's `kernel.cu` as plain C++17 against a shim
|
|
`cuda_runtime.h` and runs the kernels on host threads (32 per warp, a barrier inside `__shfl_xor_sync`). Use
|
|
small batches (`--batch-log2 13 --batches 1`). Only PASS/FAIL matters; the rates it prints are noise.
|
|
This is how the CUDA text was checked on the Mac on 3 October 2026 (all three packs PASS, see `CHECKLIST.md`). The
|
|
memory-hard pack builds the 1 GiB dataset on host threads, which takes about a second on the M5 Max; the cache fill
|
|
on one host thread took 161 ms.
|
|
It is not an nvcc build and says nothing about NVIDIA hardware.
|
|
|
|
## Regenerating a pack
|
|
|
|
Since 4 October 2026 the packs are written by the Rust crate, which is the normative implementation:
|
|
|
|
```
|
|
cd igneum-pow && cargo build --release
|
|
./target/release/igneum-pow export --seed igneum-genesis --out ../proto-cuda/packs/igneum-genesis-mh # memory-hard (default)
|
|
./target/release/igneum-pow export --closed-form --seed igneum-genesis --out ../proto-cuda/packs/igneum-genesis
|
|
./target/release/igneum-pow export --closed-form --seed igneum-hourly --out ../proto-cuda/packs/igneum-hourly
|
|
./target/release/igneum-pow export --epoch-hex <32-byte epoch seed> --day-hex <day bytes> --out ../proto-cuda/packs/<name>
|
|
cargo test --release # the packs must match the emitters byte for byte
|
|
```
|
|
|
|
The fourth form is the chain's derivation (`igneum-devnet-v4-epoch0`: the devnet genesis hash and the day bytes of
|
|
2026-10-04). The Metal cross-check is a separate step: `proto-metal/igneum-bench --seed <seed> --export-pack <scratch dir>`
|
|
derives the same program with the Swift generator, runs the Metal kernel for the three vector warps and refuses to
|
|
write unless all 96 outputs match its CPU interpreter; diff its `program.json` instruction list and `vectors.json`
|
|
against the Rust pack (identical on every seed tried, `docs/bench-log.md`). `--day` and `--dataset-log2` change the
|
|
dataset constants and are recorded in the pack.
|