Exporter writes kernel.cl next to kernel.cu (same instruction list; memory-hard core emitted in a third, OpenCL C dialect with the same literals as memhard.h). Pack headers are now C99-safe so a plain C host can include them. proto-opencl/host.c: C99 + OpenCL 1.2 API, device list, runtime build, cache fill and FNV check, dataset build and self-test, 3 vector warps standalone and in batch, bench and sweep as host.cu, whole-batch fingerprint. The 32-lane exchange is sub_group_shuffle_xor only when the queried sub-group size for a 32-item work-group is exactly 32; otherwise a local-memory exchange with one barrier per exchange, so wave64 hardware cannot change the hash (WAVEFRONT.md). build.sh (macOS, Linux), build.bat (MSVC), README with the exact AMD-rig commands. Proven without AMD silicon: Apple OpenCL 1.2 on the M5 Max 96/96 on all three packs (45.0 Mhash/s at 1 GiB, Apple number, not AMD); pocl 7.2 CPU device 96/96 on both exchange paths including the real sub_group_shuffle_xor text; CPU emulator 7 configurations incl. 64-wide sub-groups, identical fingerprint f99fb375b3abeaf5 everywhere. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
213 lines
13 KiB
Markdown
213 lines
13 KiB
Markdown
# igneum-bench-cl (proto-opencl)
|
|
|
|
The portable third path for Igneum's random-program proof-of-work kernels, after Apple Metal (`proto-metal`) and
|
|
NVIDIA CUDA (`proto-cuda`). It runs the same program packs through OpenCL 1.2, which is what an AMD card exposes on
|
|
both Windows (Adrenalin driver) and Linux (ROCm, or Mesa), and checks every output bit for bit against the Mac.
|
|
|
|
This is a test harness, not a miner. No pool, no network, no wallet, no mining protocol. It fills a dataset, checks the
|
|
device against known answers, and times the kernel. Nothing here earns anything.
|
|
|
|
Status on 3 October 2026: no AMD device has run this yet. Everything that could be proven without one has been
|
|
(`WAVEFRONT.md`, "What was proven"): Apple's deprecated OpenCL 1.2 runtime on the M5 Max, pocl 7.2 on the Mac's CPU,
|
|
and a CPU emulator with 32- and 64-wide sub-groups all give the Mac's 96/96 vectors and cache FNV for the memory-hard
|
|
pack. The AMD run itself is the next step and the commands for it are below.
|
|
|
|
## Layout
|
|
|
|
```
|
|
proto-opencl/
|
|
host.c C99 host: device list, runtime kernel build, cache + dataset fill, self-tests, vectors, bench, sweep
|
|
build.sh macOS (-framework OpenCL, or the Khronos ICD loader) and Linux (-lOpenCL)
|
|
build.bat Windows (MSVC cl.exe + OpenCL.lib)
|
|
WAVEFRONT.md wave32 vs wave64 on AMD, and why the kernel cannot tell the difference
|
|
emu/ CPU emulator: compiles kernel.cl as C++ and runs it with a 32- or 64-wide sub-group
|
|
../proto-cuda/packs/<seed>/kernel.cl the OpenCL C kernel of each pack, written by proto-metal/igneum-bench --export-pack
|
|
```
|
|
|
|
The packs stay in `proto-cuda/packs/` because the three implementations share one `program.h`, `vectors.h` and
|
|
`memhard.h` per seed; `kernel.cl` sits next to `kernel.cu` and `program.metal`. Three packs are checked in:
|
|
`igneum-genesis-mh` (memory-hard dataset, the current construction), `igneum-genesis` and `igneum-hourly`
|
|
(closed-form dataset, kept for comparison).
|
|
|
|
## What the OpenCL path does differently
|
|
|
|
| Item | CUDA (`host.cu`) | OpenCL (`host.c`) |
|
|
|---|---|---|
|
|
| Kernel compile | ahead of time by nvcc | at runtime by the driver from `kernel.cl`, with `-D IGNEUM_GROUP=32W -D IGNEUM_EXCHANGE=M` |
|
|
| 32-lane exchange | `__shfl_xor_sync` | `sub_group_shuffle_xor` only when the device's sub-group size is exactly 32 for a 32-item work-group; otherwise `__local` memory and a barrier (`WAVEFRONT.md`) |
|
|
| Timing | cudaEvent | event profiling (`CL_PROFILING_COMMAND_START/END`); wall time on Apple, whose OpenCL timestamps are unusable |
|
|
| Device choice | `--device D` | `--list`, then `--device D` by flat index over all platforms; default is the first GPU |
|
|
| Language | C++17 | C99, so MSVC builds it without nvcc or any C++ toolchain |
|
|
|
|
Everything else (dataset fill or build, cache check against the Mac's FNV, dataset self-test, 3 vector warps standalone
|
|
and inside the warm-up batch, 5 timed batches of 2^24, the sweep, the summary table, exit codes) mirrors `host.cu`
|
|
line for line. Rates have the same definition: Mhash/s, and GB/s useful = loads per hash x 4 bytes x hashes/s.
|
|
|
|
## Build
|
|
|
|
Linux (AMD ROCm, AMD Adrenalin for Linux, or Mesa rusticl; any of them registers with the system ICD loader):
|
|
|
|
```
|
|
sudo apt install ocl-icd-opencl-dev opencl-headers clinfo # Debian or Ubuntu; Fedora: ocl-icd-devel opencl-headers clinfo
|
|
clinfo | head -40 # the AMD device must appear here before anything else matters
|
|
cd proto-opencl
|
|
./build.sh igneum-genesis-mh # -> ./igneum-bench-cl-igneum-genesis-mh
|
|
./build.sh igneum-genesis # closed-form packs
|
|
./build.sh igneum-hourly
|
|
```
|
|
|
|
The one command behind the script, for a pack `P`:
|
|
|
|
```
|
|
cc -std=c99 -O2 -Wall -Wextra -I ../proto-cuda/packs/P -DIGNEUM_KERNEL_PATH='"../proto-cuda/packs/P/kernel.cl"' -o igneum-bench-cl-P host.c -lOpenCL -ldl
|
|
```
|
|
|
|
Windows (AMD Adrenalin driver installed; it provides `OpenCL.dll` and the AMD ICD, nothing else is needed to run):
|
|
|
|
1. Install Visual Studio 2022 or 2026 Build Tools with "Desktop development with C++". No CUDA, no nvcc.
|
|
2. Get the OpenCL headers and the import library `OpenCL.lib`, any one of these:
|
|
`vcpkg install opencl:x64-windows` then `set OPENCL_SDK=C:\vcpkg\installed\x64-windows`;
|
|
or the Khronos OpenCL-SDK release zip (GitHub KhronosGroup/OpenCL-SDK) unpacked anywhere, `set OPENCL_SDK=<that dir>`;
|
|
or an installed CUDA Toolkit, which ships both under `%CUDA_PATH%` (the script finds it on its own).
|
|
3. From an "x64 Native Tools Command Prompt":
|
|
|
|
```
|
|
cd proto-opencl
|
|
build.bat igneum-genesis-mh
|
|
build.bat igneum-genesis
|
|
build.bat igneum-hourly
|
|
```
|
|
|
|
The one command behind the script:
|
|
|
|
```
|
|
cl /nologo /O2 /W3 /std:c11 /I "%OPENCL_SDK%\include" /I "..\proto-cuda\packs\P" /DIGNEUM_KERNEL_PATH="\"../proto-cuda/packs/P/kernel.cl\"" /Fe:igneum-bench-cl-P.exe host.c /link "%OPENCL_SDK%\lib\OpenCL.lib"
|
|
```
|
|
|
|
macOS (correctness check only; Apple deprecated OpenCL in macOS 10.14 and the runtime is 1.2):
|
|
|
|
```
|
|
./build.sh igneum-genesis-mh # Apple OpenCL.framework
|
|
./build.sh igneum-genesis-mh khr # Khronos ICD loader from Homebrew, for pocl: brew install opencl-headers opencl-icd-loader pocl
|
|
OCL_ICD_VENDORS=/opt/homebrew/etc/OpenCL/vendors SDKROOT=$(xcrun --show-sdk-path) ./igneum-bench-cl-igneum-genesis-mh-khr --list
|
|
```
|
|
|
|
(`SDKROOT` is needed because pocl links its kernels with the system linker and otherwise cannot find `-lSystem`.)
|
|
|
|
## Run on the AMD rig
|
|
|
|
Run from `proto-opencl/` so the default kernel path resolves, or pass `--kernel`. Exactly this, in this order, and paste
|
|
the whole stdout back:
|
|
|
|
Windows:
|
|
|
|
```
|
|
cd proto-opencl
|
|
igneum-bench-cl-igneum-genesis-mh.exe --list
|
|
igneum-bench-cl-igneum-genesis-mh.exe > amd-mh-auto.txt
|
|
igneum-bench-cl-igneum-genesis-mh.exe --exchange local > amd-mh-local.txt
|
|
igneum-bench-cl-igneum-genesis-mh.exe --group-warps 2 > amd-mh-gw2.txt
|
|
igneum-bench-cl-igneum-genesis-mh.exe --sweep > amd-mh-sweep.txt
|
|
igneum-bench-cl-igneum-genesis.exe > amd-genesis.txt
|
|
igneum-bench-cl-igneum-hourly.exe > amd-hourly.txt
|
|
```
|
|
|
|
Linux:
|
|
|
|
```
|
|
cd proto-opencl
|
|
./igneum-bench-cl-igneum-genesis-mh --list
|
|
./igneum-bench-cl-igneum-genesis-mh | tee amd-mh-auto.txt
|
|
./igneum-bench-cl-igneum-genesis-mh --exchange local | tee amd-mh-local.txt
|
|
./igneum-bench-cl-igneum-genesis-mh --group-warps 2 | tee amd-mh-gw2.txt
|
|
./igneum-bench-cl-igneum-genesis-mh --sweep | tee amd-mh-sweep.txt
|
|
./igneum-bench-cl-igneum-genesis | tee amd-genesis.txt
|
|
./igneum-bench-cl-igneum-hourly | tee amd-hourly.txt
|
|
```
|
|
|
|
If `--list` shows more than one device (an iGPU, a CPU runtime, an NVIDIA card in the same box), add `--device D`
|
|
with the AMD card's index to every line. If the auto run says `exchange: sub_group_shuffle_xor`, also run
|
|
`--exchange subgroup` once so the log has an explicit sub-group run; if it says local memory, the `--exchange local`
|
|
run is a repeat and that is fine.
|
|
|
|
Flags: `--list`, `--device D`, `--dataset-mib N` (power of two, default 1024), `--sweep` (4, 64, 256, 512, 1024 MiB),
|
|
`--batch-log2 B` (default 24), `--batches N` (default 5), `--group-warps W` (32-lane units per work-group, 1..8,
|
|
default 1), `--exchange auto|local|subgroup`, `--kernel path`, `--build-opts "..."` (appended to clBuildProgram),
|
|
`--time event|wall`.
|
|
|
|
## What PASS looks like
|
|
|
|
In order, the run prints:
|
|
|
|
1. Every platform and device, with vendor, driver, OpenCL C version, compute units, memory, local memory, the
|
|
sub-group extension it lists, and (AMD) the wavefront width or (NVIDIA) the warp size. The chosen device is starred.
|
|
2. `build options:` and `exchange:`. The exchange line is the one that matters on AMD; it says which path was taken and
|
|
why, for example `local-memory exchange with barrier (the sub-group size for a 32-item work-group is not 32; queried
|
|
sub-group size 64)` on a wave64 card, or `sub_group_shuffle_xor (cl_khr_subgroup_shuffle), sub-group size 32 for a
|
|
32-item work-group` on a wave32 one. Both are correct by construction (`WAVEFRONT.md`).
|
|
3. `kernel:` (work-group limit and local memory of `igneum_hash`), `program:`, `seed words:`, `day`.
|
|
4. Memory-hard packs: cache fill time on the device (twice) and on one host thread, then `cache check: PASS (device ==
|
|
host all 67108864 words PASS, host FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head 16 vs Mac PASS,
|
|
last line vs Mac PASS)`. The FNV is the same on every machine that has ever run this construction for day
|
|
2026-10-03.
|
|
5. `dataset build` (or `dataset fill`) times, then `dataset self-test: PASS (head 16 vs Mac PASS, element [MASK] vs Mac
|
|
PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)`.
|
|
6. Six `verify warp ... PASS` lines: three standalone (bases 0, 4096, 1000000) and the same three read out of the
|
|
warm-up batch.
|
|
7. `batch fingerprint`: FNV-1a 64 of every output of the warm-up batch. At `--batch-log2 13` every implementation so far
|
|
prints `f99fb375b3abeaf5` for igneum-genesis-mh (Apple OpenCL, pocl on both exchange paths, the emulator in seven
|
|
configurations); an AMD run at `--batch-log2 13` must print the same value. At the default 2^24 the value is a new
|
|
reference to compare AMD against the next machine.
|
|
8. `timed:` with device and wall time, then the `rate` line: Mhash/s and GB/s useful.
|
|
9. The summary table and `OVERALL: PASS`. Exit code 0 on PASS, 1 on FAIL, 2 on an OpenCL error or a build failure (the
|
|
build log is printed in full).
|
|
|
|
A FAIL with a few lanes differing points at the exchange; a FAIL in every lane at an arithmetic op or the dataset; a
|
|
cache FAIL at the ChaCha fill. Send the whole printout either way, including the `--list` output and, on Linux,
|
|
`clinfo` and the ROCm or Mesa version, or on Windows the Adrenalin driver version from the AMD Software panel.
|
|
|
|
## What to paste back
|
|
|
|
All seven text files above, unedited. From them the bench log gets: device name and driver string, the `exchange:`
|
|
line, the cache check line, the 96/96 vector result per pack, the rate at 1 GiB, and the sweep table. If a run looks
|
|
odd, run it again and keep both.
|
|
|
|
## Results so far (no AMD)
|
|
|
|
| Machine, runtime | Exchange | igneum-genesis-mh | closed-form packs | Rate at 1 GiB |
|
|
|---|---|---|---|---|
|
|
| Apple M5 Max, Apple OpenCL 1.2 (deprecated runtime) | local memory (runtime lists no sub-group extension) | cache FNV = Mac, 96/96; also 96/96 at `--group-warps 2` and `4` | 96/96 and 96/96 | 45.03 Mhash/s, 18.73 GB/s useful (Metal on the same chip: 45.2). Apple OpenCL on the M5 Max, not an AMD number |
|
|
| Apple M5 Max, pocl 7.2 CPU device, OpenCL 3.0, LLVM 23 | `sub_group_shuffle_xor` (auto and `--exchange subgroup`; pocl's `clGetKernelSubGroupInfoKHR` fails with CL_INVALID_OPERATION, the probe kernel reports 32), and `--exchange local` | cache FNV = Mac, 96/96 on both paths, same batch fingerprint as Apple and the emulator at 2^13 | not run | CPU, not meaningful |
|
|
| CPU emulator (`emu/`), 7 configurations incl. sub-group 64 | both | 96/96 in every configuration, identical batch fingerprint `f99fb375b3abeaf5` | not run | none |
|
|
|
|
Sweep on Apple OpenCL (3 batches, wall time): 4 MiB 573.7, 64 MiB 178.9, 256 MiB 94.3, 512 MiB 68.8, 1024 MiB 45.0 Mhash/s.
|
|
Full tables in `docs/bench-log.md`.
|
|
|
|
## Checking a pack without a GPU
|
|
|
|
```
|
|
emu/emu.sh igneum-genesis-mh 0 32 --sg 32 # local-memory exchange, work-group 32, 32-wide sub-group
|
|
emu/emu.sh igneum-genesis-mh 0 64 --sg 64 # local-memory exchange on a wave64 model, two units per work-group
|
|
emu/emu.sh igneum-genesis-mh 1 32 --sg 32 # sub_group_shuffle_xor, 32-wide sub-group
|
|
emu/emu.sh igneum-genesis-mh 1 64 --sg 64 # sub_group_shuffle_xor over a 64-wide sub-group (two units in one wave)
|
|
```
|
|
|
|
The emulator compiles `kernel.cl` unchanged as C++ (`emu/emu_opencl.h` supplies the OpenCL built-ins; the kernel's own
|
|
prelude includes it when no OpenCL compiler is present) and runs work-groups on host threads with a barrier inside
|
|
every exchange. It prints the cache check, dataset self-test, vectors and a fingerprint of the whole 2^13-hash batch,
|
|
so configurations can be compared for every nonce, not only the vector warps. It says nothing about any GPU.
|
|
|
|
## Regenerating a pack
|
|
|
|
```
|
|
cd proto-metal
|
|
swiftc -O -o igneum-bench main.swift -framework Metal
|
|
./igneum-bench --seed igneum-genesis --export-pack ../proto-cuda/packs/igneum-genesis-mh
|
|
./igneum-bench --closed-form --seed igneum-genesis --export-pack ../proto-cuda/packs/igneum-genesis
|
|
./igneum-bench --closed-form --seed igneum-hourly --export-pack ../proto-cuda/packs/igneum-hourly
|
|
```
|
|
|
|
The exporter writes `kernel.cl` next to `kernel.cu` from the same instruction list, refuses to write unless the Metal
|
|
GPU matches the CPU interpreter on all 96 vector outputs, and for memory-hard packs unless the GPU cache equals the
|
|
CPU cache word for word. The memory-hard core is emitted once (`emitMemhardCore`) in three dialects (Metal, CUDA C++,
|
|
OpenCL C), so the constants in `kernel.cl` are the literals of `memhard.h`.
|