Exporter writes kernel.cl next to kernel.cu (same instruction list; memory-hard core emitted in a third, OpenCL C dialect with the same literals as memhard.h). Pack headers are now C99-safe so a plain C host can include them. proto-opencl/host.c: C99 + OpenCL 1.2 API, device list, runtime build, cache fill and FNV check, dataset build and self-test, 3 vector warps standalone and in batch, bench and sweep as host.cu, whole-batch fingerprint. The 32-lane exchange is sub_group_shuffle_xor only when the queried sub-group size for a 32-item work-group is exactly 32; otherwise a local-memory exchange with one barrier per exchange, so wave64 hardware cannot change the hash (WAVEFRONT.md). build.sh (macOS, Linux), build.bat (MSVC), README with the exact AMD-rig commands. Proven without AMD silicon: Apple OpenCL 1.2 on the M5 Max 96/96 on all three packs (45.0 Mhash/s at 1 GiB, Apple number, not AMD); pocl 7.2 CPU device 96/96 on both exchange paths including the real sub_group_shuffle_xor text; CPU emulator 7 configurations incl. 64-wide sub-groups, identical fingerprint f99fb375b3abeaf5 everywhere. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
178 lines
10 KiB
Markdown
178 lines
10 KiB
Markdown
# igneum-bench-cuda (proto-cuda)
|
|
|
|
The NVIDIA twin of `proto-metal`. It runs the same random-program proof-of-work kernels on a CUDA GPU and checks
|
|
them bit for bit against results produced on the Mac.
|
|
|
|
This is a test harness, not a miner. No pool, no network, no wallet, no mining protocol. It fills a dataset,
|
|
checks the GPU against known answers, and times the kernel. Nothing here earns anything.
|
|
|
|
Status on 3 October 2026: the two closed-form packs ran on the RTX 5090 (96/96 vectors PASS each, see
|
|
`docs/bench-log.md`). The memory-hard pack `igneum-genesis-mh` (added later the same day, construction in
|
|
`../proto-metal/MEMHARD.md`) has passed only the clang emulation on the Mac; its 5090 and AMD runs are pending, and no
|
|
NVIDIA figure for the memory-hard dataset exists yet.
|
|
|
|
## Layout
|
|
|
|
```
|
|
proto-cuda/
|
|
host.cu host program: device info, dataset fill, self-test, vectors, bench, size sweep
|
|
build.sh Linux build (nvcc)
|
|
build.bat Windows build (nvcc + Visual Studio Build Tools)
|
|
CHECKLIST.md Metal/CUDA equivalence, op by op, and what was verified where
|
|
packs/<seed>/ one program pack per seed, written by proto-metal/igneum-bench --export-pack
|
|
kernel.cu the program as a CUDA kernel, plus fill kernel and host launch wrappers
|
|
kernel.cl the same program as OpenCL C for proto-opencl (AMD and any other OpenCL device), built at runtime
|
|
program.h seed, day words, dataset size, loads per hash, wrapper declarations (C99-safe: proto-opencl/host.c includes it too)
|
|
vectors.h expected outputs for 3 warps (96 x 64-bit) and dataset self-test values
|
|
program.json the instruction list and all constants, for any other implementation
|
|
vectors.json the same vectors as JSON
|
|
program.metal the Metal source the Mac ran, for diffing by eye
|
|
memhard.h memory-hard packs only: the cache fill and item derivation core, compiled for device and host
|
|
memhard.metal memory-hard packs only: the Metal cache-fill and build kernels the Mac ran
|
|
emu/ CPU emulation shim: compile and check a pack with plain clang++/g++, no GPU
|
|
```
|
|
|
|
Three packs are checked in. `igneum-genesis` (104 loads per hash) and `igneum-hourly` (128 loads per hash) use the
|
|
original closed-form dataset (`IGNEUM_DATASET_MODE 0`, implied when the macro is absent). `igneum-genesis-mh` is the
|
|
same program as `igneum-genesis` over the memory-hard dataset (`IGNEUM_DATASET_MODE 1`): a 256 MiB cache of chained
|
|
ChaCha12 blocks filled on the GPU from the day key, and every 64-byte dataset item derived from 8 dependent cache
|
|
reads through a seed-parameterised mixer (`../proto-metal/MEMHARD.md`). The hash kernel text is identical in both
|
|
packs; only the dataset contents differ, so the 96 expected outputs differ. All three were cross-checked on the
|
|
Mac's Metal GPU before being written.
|
|
|
|
## Prerequisites
|
|
|
|
Linux
|
|
- An NVIDIA driver recent enough for the toolkit. For CUDA 12.8 that is the R570 series or newer (approximate,
|
|
from memory; `nvidia-smi` prints the driver's maximum supported CUDA version in its header).
|
|
- CUDA Toolkit 12.8 or newer. Blackwell (`sm_120`, RTX 50 series) is not known to older toolkits.
|
|
- A host compiler the toolkit supports (gcc 11 to 13 for 12.8, approximate).
|
|
|
|
Windows
|
|
- The same driver requirement.
|
|
- CUDA Toolkit 12.8 or newer. Tick the Visual Studio integration in the installer.
|
|
- Visual Studio 2022 Build Tools with the "Desktop development with C++" workload. nvcc needs `cl.exe`.
|
|
- Run `build.bat` from an "x64 Native Tools Command Prompt for VS 2022" so `cl.exe` is on PATH.
|
|
|
|
## Build
|
|
|
|
Linux:
|
|
|
|
```
|
|
cd proto-cuda
|
|
./build.sh # pack igneum-genesis, -arch=sm_120
|
|
./build.sh igneum-hourly # the second pack
|
|
./build.sh igneum-genesis-mh # the memory-hard pack (needs 256 MiB more device memory for the cache)
|
|
./build.sh igneum-genesis native # if sm_120 is refused, let nvcc pick the installed GPU
|
|
```
|
|
|
|
Windows (x64 Native Tools Command Prompt):
|
|
|
|
```
|
|
cd proto-cuda
|
|
build.bat
|
|
build.bat igneum-hourly
|
|
build.bat igneum-genesis-mh
|
|
build.bat igneum-genesis native
|
|
```
|
|
|
|
Both scripts run this one command (paths adjusted for the pack):
|
|
|
|
```
|
|
nvcc -O3 -std=c++17 -arch=sm_120 -I packs/igneum-genesis -o igneum-bench-cuda-igneum-genesis host.cu packs/igneum-genesis/kernel.cu
|
|
```
|
|
|
|
Notes
|
|
- `-arch=sm_120` is Blackwell. `-arch=native` (CUDA 11.6 or newer) compiles for whatever GPU is in the
|
|
machine and is the fallback if the toolkit is too old to know `sm_120` (which means it is too old for a 5090
|
|
anyway: upgrade).
|
|
- If nvcc on Windows refuses the Visual Studio version, add `-allow-unsupported-compiler` to the nvcc line.
|
|
- The host code is plain C++17 and the CUDA runtime API. No NVRTC, no third-party libraries, no JSON parser.
|
|
The kernel is compiled ahead of time from the pack.
|
|
|
|
## Run
|
|
|
|
```
|
|
./igneum-bench-cuda-igneum-genesis # 1 GiB dataset, 5 batches x 2^24 nonces, vectors checked
|
|
./igneum-bench-cuda-igneum-genesis --sweep # 4, 64, 256, 512, 1024 MiB in sequence (the Mac's sweep)
|
|
./igneum-bench-cuda-igneum-genesis --block-warps 4 # 4 warps per block instead of 1 (still bit-exact)
|
|
./igneum-bench-cuda-igneum-hourly # the second program
|
|
./igneum-bench-cuda-igneum-genesis-mh # memory-hard dataset: cache fill + build, cache check, vectors, bench
|
|
./igneum-bench-cuda-igneum-genesis-mh --sweep # the same sweep over the memory-hard dataset
|
|
```
|
|
|
|
On Windows the binaries are `igneum-bench-cuda-igneum-genesis.exe` and so on.
|
|
|
|
Flags: `--dataset-mib N` (power of two, default 1024), `--sweep`, `--batch-log2 B` (default 24),
|
|
`--batches N` (default 5), `--block-warps W` (default 1, mirrors the Metal run's one SIMD group per threadgroup),
|
|
`--device D`.
|
|
|
|
What it prints, in order:
|
|
1. GPU name, SM count, memory, clocks, L2, warp size, driver and runtime versions, registers per thread and
|
|
resident warps per SM for the kernel, and which dataset construction the pack uses.
|
|
2. Memory-hard packs only: cache fill time on the GPU (twice), cache fill time on the host (one thread, the same
|
|
`memhard.h` text), then the cache check: every one of the 2^26 words GPU versus host, the host FNV-1a 64 against
|
|
the Mac's, and the head and last line against the Mac's. A FAIL here stops nothing but fails OVERALL.
|
|
3. Dataset fill time (closed form, twice, with write GB/s) or dataset build time (memory-hard, twice, with items/s
|
|
and cache-line reads/s).
|
|
4. Dataset self-test: 16 head words and word `[MASK]` against values from the Mac, 64 random words against the
|
|
host formula (closed form) or the host derivation from the host cache (memory-hard), and the Mac's 64 sampled
|
|
words (memory-hard packs; those inside the current dataset size).
|
|
5. Vectors: 3 warps (base nonces 0, 4096, 1000000), each run standalone as one 32-thread block, then again
|
|
read out of the warm-up batch so the bench configuration itself is checked. PASS or FAIL per warp, with the
|
|
first differing lane printed on FAIL.
|
|
6. Timing: 5 batches of 2^24 hashes after a warm-up batch, GPU event time and wall time, Mhash/s, hashes/s,
|
|
GB/s useful (loads per hash x 4 bytes x hashes/s, the same definition as the Mac's table).
|
|
7. A summary table in Markdown and `OVERALL: PASS` or `FAIL`. Exit code 0 on PASS, 1 on FAIL, 2 on a CUDA error.
|
|
|
|
Vectors are only checked when the dataset is the pack's size (1024 MiB), because the outputs depend on the
|
|
address mask. At other sizes the table says "skipped (not pack size)" and only the dataset self-test counts.
|
|
|
|
## What PASS means
|
|
|
|
- Closed-form packs: the CUDA fill kernel produced the same dataset as the Mac's closed-form function (sampled, not
|
|
every word).
|
|
- Memory-hard pack: the CUDA cache-fill kernel produced, word for word, the same 256 MiB cache as the host and as the
|
|
Mac (FNV-1a 64, head, last line), and the CUDA build kernel produced the same dataset words as the host derivation
|
|
and the Mac's samples (sampled, not every word). The whole chain from day key to dataset word agrees across three
|
|
compilers (Apple Metal, host C++, NVIDIA CUDA).
|
|
- For 96 nonces spread across the nonce space, the RTX 5090 produced the same 64-bit outputs as the Mac's CPU
|
|
interpreter, which had itself matched the Mac's Metal GPU. The random program, the register init, the warp
|
|
shuffles, the multiply-high and rotates, and the dataset addressing all agree between Apple and NVIDIA.
|
|
- With `--block-warps W` the in-batch check passing shows that packing W warps per block changed nothing.
|
|
|
|
A FAIL with a small number of differing lanes points at a shuffle; a FAIL in every lane points at an arithmetic
|
|
op or the dataset. Send the whole printout either way.
|
|
|
|
## Sending results back
|
|
|
|
Copy the printed header lines (GPU, CUDA versions, kernel line, program line) and the summary table into
|
|
`docs/bench-log.md` under a dated heading, together with the output of `nvcc --version` and the driver version
|
|
from `nvidia-smi`. Keep the full stdout as well. Run both packs and the sweep so the log has the same shape as
|
|
the Mac's entry. Do not edit the numbers; if a run looks odd, run it again and log both.
|
|
|
|
## Checking a pack without a GPU
|
|
|
|
`emu/emu.sh <pack> [flags]` compiles `host.cu` and the pack's `kernel.cu` as plain C++17 against a shim
|
|
`cuda_runtime.h` and runs the kernels on host threads (32 per warp, a barrier inside `__shfl_xor_sync`). Use
|
|
small batches (`--batch-log2 13 --batches 1`). Only PASS/FAIL matters; the rates it prints are noise.
|
|
This is how the CUDA text was checked on the Mac on 3 October 2026 (all three packs PASS, see `CHECKLIST.md`). The
|
|
memory-hard pack builds the 1 GiB dataset on host threads, which takes about a second on the M5 Max; the cache fill
|
|
on one host thread took 161 ms.
|
|
It is not an nvcc build and says nothing about NVIDIA hardware.
|
|
|
|
## Regenerating a pack
|
|
|
|
On the Mac:
|
|
|
|
```
|
|
cd proto-metal
|
|
swiftc -O -o igneum-bench main.swift -framework Metal
|
|
./igneum-bench --seed igneum-genesis --export-pack ../proto-cuda/packs/igneum-genesis-mh # memory-hard (default)
|
|
./igneum-bench --closed-form --seed igneum-genesis --export-pack ../proto-cuda/packs/igneum-genesis
|
|
```
|
|
|
|
The exporter runs the CPU interpreter for the three vector warps, runs the Metal kernel for the same warps,
|
|
and refuses to write anything unless all 96 outputs match. For a memory-hard pack it also refuses unless the GPU
|
|
cache equals the CPU cache on every word and the sampled GPU dataset words equal the CPU derivation. `--day` and `--dataset-log2` change the dataset
|
|
constants and are recorded in the pack.
|