igneum/proto-cuda/README.md
igneum-labs 660f0eb16c proto-opencl: OpenCL path for AMD, proven on Apple OpenCL, pocl and a wave64 CPU emulator
Exporter writes kernel.cl next to kernel.cu (same instruction list; memory-hard core emitted in a third, OpenCL C
dialect with the same literals as memhard.h). Pack headers are now C99-safe so a plain C host can include them.

proto-opencl/host.c: C99 + OpenCL 1.2 API, device list, runtime build, cache fill and FNV check, dataset build and
self-test, 3 vector warps standalone and in batch, bench and sweep as host.cu, whole-batch fingerprint. The 32-lane
exchange is sub_group_shuffle_xor only when the queried sub-group size for a 32-item work-group is exactly 32;
otherwise a local-memory exchange with one barrier per exchange, so wave64 hardware cannot change the hash
(WAVEFRONT.md). build.sh (macOS, Linux), build.bat (MSVC), README with the exact AMD-rig commands.

Proven without AMD silicon: Apple OpenCL 1.2 on the M5 Max 96/96 on all three packs (45.0 Mhash/s at 1 GiB, Apple
number, not AMD); pocl 7.2 CPU device 96/96 on both exchange paths including the real sub_group_shuffle_xor text;
CPU emulator 7 configurations incl. 64-wide sub-groups, identical fingerprint f99fb375b3abeaf5 everywhere.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-03 16:52:24 +00:00

178 lines
10 KiB
Markdown

# igneum-bench-cuda (proto-cuda)
The NVIDIA twin of `proto-metal`. It runs the same random-program proof-of-work kernels on a CUDA GPU and checks
them bit for bit against results produced on the Mac.
This is a test harness, not a miner. No pool, no network, no wallet, no mining protocol. It fills a dataset,
checks the GPU against known answers, and times the kernel. Nothing here earns anything.
Status on 3 October 2026: the two closed-form packs ran on the RTX 5090 (96/96 vectors PASS each, see
`docs/bench-log.md`). The memory-hard pack `igneum-genesis-mh` (added later the same day, construction in
`../proto-metal/MEMHARD.md`) has passed only the clang emulation on the Mac; its 5090 and AMD runs are pending, and no
NVIDIA figure for the memory-hard dataset exists yet.
## Layout
```
proto-cuda/
host.cu host program: device info, dataset fill, self-test, vectors, bench, size sweep
build.sh Linux build (nvcc)
build.bat Windows build (nvcc + Visual Studio Build Tools)
CHECKLIST.md Metal/CUDA equivalence, op by op, and what was verified where
packs/<seed>/ one program pack per seed, written by proto-metal/igneum-bench --export-pack
kernel.cu the program as a CUDA kernel, plus fill kernel and host launch wrappers
kernel.cl the same program as OpenCL C for proto-opencl (AMD and any other OpenCL device), built at runtime
program.h seed, day words, dataset size, loads per hash, wrapper declarations (C99-safe: proto-opencl/host.c includes it too)
vectors.h expected outputs for 3 warps (96 x 64-bit) and dataset self-test values
program.json the instruction list and all constants, for any other implementation
vectors.json the same vectors as JSON
program.metal the Metal source the Mac ran, for diffing by eye
memhard.h memory-hard packs only: the cache fill and item derivation core, compiled for device and host
memhard.metal memory-hard packs only: the Metal cache-fill and build kernels the Mac ran
emu/ CPU emulation shim: compile and check a pack with plain clang++/g++, no GPU
```
Three packs are checked in. `igneum-genesis` (104 loads per hash) and `igneum-hourly` (128 loads per hash) use the
original closed-form dataset (`IGNEUM_DATASET_MODE 0`, implied when the macro is absent). `igneum-genesis-mh` is the
same program as `igneum-genesis` over the memory-hard dataset (`IGNEUM_DATASET_MODE 1`): a 256 MiB cache of chained
ChaCha12 blocks filled on the GPU from the day key, and every 64-byte dataset item derived from 8 dependent cache
reads through a seed-parameterised mixer (`../proto-metal/MEMHARD.md`). The hash kernel text is identical in both
packs; only the dataset contents differ, so the 96 expected outputs differ. All three were cross-checked on the
Mac's Metal GPU before being written.
## Prerequisites
Linux
- An NVIDIA driver recent enough for the toolkit. For CUDA 12.8 that is the R570 series or newer (approximate,
from memory; `nvidia-smi` prints the driver's maximum supported CUDA version in its header).
- CUDA Toolkit 12.8 or newer. Blackwell (`sm_120`, RTX 50 series) is not known to older toolkits.
- A host compiler the toolkit supports (gcc 11 to 13 for 12.8, approximate).
Windows
- The same driver requirement.
- CUDA Toolkit 12.8 or newer. Tick the Visual Studio integration in the installer.
- Visual Studio 2022 Build Tools with the "Desktop development with C++" workload. nvcc needs `cl.exe`.
- Run `build.bat` from an "x64 Native Tools Command Prompt for VS 2022" so `cl.exe` is on PATH.
## Build
Linux:
```
cd proto-cuda
./build.sh # pack igneum-genesis, -arch=sm_120
./build.sh igneum-hourly # the second pack
./build.sh igneum-genesis-mh # the memory-hard pack (needs 256 MiB more device memory for the cache)
./build.sh igneum-genesis native # if sm_120 is refused, let nvcc pick the installed GPU
```
Windows (x64 Native Tools Command Prompt):
```
cd proto-cuda
build.bat
build.bat igneum-hourly
build.bat igneum-genesis-mh
build.bat igneum-genesis native
```
Both scripts run this one command (paths adjusted for the pack):
```
nvcc -O3 -std=c++17 -arch=sm_120 -I packs/igneum-genesis -o igneum-bench-cuda-igneum-genesis host.cu packs/igneum-genesis/kernel.cu
```
Notes
- `-arch=sm_120` is Blackwell. `-arch=native` (CUDA 11.6 or newer) compiles for whatever GPU is in the
machine and is the fallback if the toolkit is too old to know `sm_120` (which means it is too old for a 5090
anyway: upgrade).
- If nvcc on Windows refuses the Visual Studio version, add `-allow-unsupported-compiler` to the nvcc line.
- The host code is plain C++17 and the CUDA runtime API. No NVRTC, no third-party libraries, no JSON parser.
The kernel is compiled ahead of time from the pack.
## Run
```
./igneum-bench-cuda-igneum-genesis # 1 GiB dataset, 5 batches x 2^24 nonces, vectors checked
./igneum-bench-cuda-igneum-genesis --sweep # 4, 64, 256, 512, 1024 MiB in sequence (the Mac's sweep)
./igneum-bench-cuda-igneum-genesis --block-warps 4 # 4 warps per block instead of 1 (still bit-exact)
./igneum-bench-cuda-igneum-hourly # the second program
./igneum-bench-cuda-igneum-genesis-mh # memory-hard dataset: cache fill + build, cache check, vectors, bench
./igneum-bench-cuda-igneum-genesis-mh --sweep # the same sweep over the memory-hard dataset
```
On Windows the binaries are `igneum-bench-cuda-igneum-genesis.exe` and so on.
Flags: `--dataset-mib N` (power of two, default 1024), `--sweep`, `--batch-log2 B` (default 24),
`--batches N` (default 5), `--block-warps W` (default 1, mirrors the Metal run's one SIMD group per threadgroup),
`--device D`.
What it prints, in order:
1. GPU name, SM count, memory, clocks, L2, warp size, driver and runtime versions, registers per thread and
resident warps per SM for the kernel, and which dataset construction the pack uses.
2. Memory-hard packs only: cache fill time on the GPU (twice), cache fill time on the host (one thread, the same
`memhard.h` text), then the cache check: every one of the 2^26 words GPU versus host, the host FNV-1a 64 against
the Mac's, and the head and last line against the Mac's. A FAIL here stops nothing but fails OVERALL.
3. Dataset fill time (closed form, twice, with write GB/s) or dataset build time (memory-hard, twice, with items/s
and cache-line reads/s).
4. Dataset self-test: 16 head words and word `[MASK]` against values from the Mac, 64 random words against the
host formula (closed form) or the host derivation from the host cache (memory-hard), and the Mac's 64 sampled
words (memory-hard packs; those inside the current dataset size).
5. Vectors: 3 warps (base nonces 0, 4096, 1000000), each run standalone as one 32-thread block, then again
read out of the warm-up batch so the bench configuration itself is checked. PASS or FAIL per warp, with the
first differing lane printed on FAIL.
6. Timing: 5 batches of 2^24 hashes after a warm-up batch, GPU event time and wall time, Mhash/s, hashes/s,
GB/s useful (loads per hash x 4 bytes x hashes/s, the same definition as the Mac's table).
7. A summary table in Markdown and `OVERALL: PASS` or `FAIL`. Exit code 0 on PASS, 1 on FAIL, 2 on a CUDA error.
Vectors are only checked when the dataset is the pack's size (1024 MiB), because the outputs depend on the
address mask. At other sizes the table says "skipped (not pack size)" and only the dataset self-test counts.
## What PASS means
- Closed-form packs: the CUDA fill kernel produced the same dataset as the Mac's closed-form function (sampled, not
every word).
- Memory-hard pack: the CUDA cache-fill kernel produced, word for word, the same 256 MiB cache as the host and as the
Mac (FNV-1a 64, head, last line), and the CUDA build kernel produced the same dataset words as the host derivation
and the Mac's samples (sampled, not every word). The whole chain from day key to dataset word agrees across three
compilers (Apple Metal, host C++, NVIDIA CUDA).
- For 96 nonces spread across the nonce space, the RTX 5090 produced the same 64-bit outputs as the Mac's CPU
interpreter, which had itself matched the Mac's Metal GPU. The random program, the register init, the warp
shuffles, the multiply-high and rotates, and the dataset addressing all agree between Apple and NVIDIA.
- With `--block-warps W` the in-batch check passing shows that packing W warps per block changed nothing.
A FAIL with a small number of differing lanes points at a shuffle; a FAIL in every lane points at an arithmetic
op or the dataset. Send the whole printout either way.
## Sending results back
Copy the printed header lines (GPU, CUDA versions, kernel line, program line) and the summary table into
`docs/bench-log.md` under a dated heading, together with the output of `nvcc --version` and the driver version
from `nvidia-smi`. Keep the full stdout as well. Run both packs and the sweep so the log has the same shape as
the Mac's entry. Do not edit the numbers; if a run looks odd, run it again and log both.
## Checking a pack without a GPU
`emu/emu.sh <pack> [flags]` compiles `host.cu` and the pack's `kernel.cu` as plain C++17 against a shim
`cuda_runtime.h` and runs the kernels on host threads (32 per warp, a barrier inside `__shfl_xor_sync`). Use
small batches (`--batch-log2 13 --batches 1`). Only PASS/FAIL matters; the rates it prints are noise.
This is how the CUDA text was checked on the Mac on 3 October 2026 (all three packs PASS, see `CHECKLIST.md`). The
memory-hard pack builds the 1 GiB dataset on host threads, which takes about a second on the M5 Max; the cache fill
on one host thread took 161 ms.
It is not an nvcc build and says nothing about NVIDIA hardware.
## Regenerating a pack
On the Mac:
```
cd proto-metal
swiftc -O -o igneum-bench main.swift -framework Metal
./igneum-bench --seed igneum-genesis --export-pack ../proto-cuda/packs/igneum-genesis-mh # memory-hard (default)
./igneum-bench --closed-form --seed igneum-genesis --export-pack ../proto-cuda/packs/igneum-genesis
```
The exporter runs the CPU interpreter for the three vector warps, runs the Metal kernel for the same warps,
and refuses to write anything unless all 96 outputs match. For a memory-hard pack it also refuses unless the GPU
cache equals the CPU cache on every word and the sampled GPU dataset words equal the CPU derivation. `--day` and `--dataset-log2` change the dataset
constants and are recorded in the pack.