igneum/proto-opencl/README.md
igneum-labs 660f0eb16c proto-opencl: OpenCL path for AMD, proven on Apple OpenCL, pocl and a wave64 CPU emulator
Exporter writes kernel.cl next to kernel.cu (same instruction list; memory-hard core emitted in a third, OpenCL C
dialect with the same literals as memhard.h). Pack headers are now C99-safe so a plain C host can include them.

proto-opencl/host.c: C99 + OpenCL 1.2 API, device list, runtime build, cache fill and FNV check, dataset build and
self-test, 3 vector warps standalone and in batch, bench and sweep as host.cu, whole-batch fingerprint. The 32-lane
exchange is sub_group_shuffle_xor only when the queried sub-group size for a 32-item work-group is exactly 32;
otherwise a local-memory exchange with one barrier per exchange, so wave64 hardware cannot change the hash
(WAVEFRONT.md). build.sh (macOS, Linux), build.bat (MSVC), README with the exact AMD-rig commands.

Proven without AMD silicon: Apple OpenCL 1.2 on the M5 Max 96/96 on all three packs (45.0 Mhash/s at 1 GiB, Apple
number, not AMD); pocl 7.2 CPU device 96/96 on both exchange paths including the real sub_group_shuffle_xor text;
CPU emulator 7 configurations incl. 64-wide sub-groups, identical fingerprint f99fb375b3abeaf5 everywhere.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-03 16:52:24 +00:00

213 lines
13 KiB
Markdown

# igneum-bench-cl (proto-opencl)
The portable third path for Igneum's random-program proof-of-work kernels, after Apple Metal (`proto-metal`) and
NVIDIA CUDA (`proto-cuda`). It runs the same program packs through OpenCL 1.2, which is what an AMD card exposes on
both Windows (Adrenalin driver) and Linux (ROCm, or Mesa), and checks every output bit for bit against the Mac.
This is a test harness, not a miner. No pool, no network, no wallet, no mining protocol. It fills a dataset, checks the
device against known answers, and times the kernel. Nothing here earns anything.
Status on 3 October 2026: no AMD device has run this yet. Everything that could be proven without one has been
(`WAVEFRONT.md`, "What was proven"): Apple's deprecated OpenCL 1.2 runtime on the M5 Max, pocl 7.2 on the Mac's CPU,
and a CPU emulator with 32- and 64-wide sub-groups all give the Mac's 96/96 vectors and cache FNV for the memory-hard
pack. The AMD run itself is the next step and the commands for it are below.
## Layout
```
proto-opencl/
host.c C99 host: device list, runtime kernel build, cache + dataset fill, self-tests, vectors, bench, sweep
build.sh macOS (-framework OpenCL, or the Khronos ICD loader) and Linux (-lOpenCL)
build.bat Windows (MSVC cl.exe + OpenCL.lib)
WAVEFRONT.md wave32 vs wave64 on AMD, and why the kernel cannot tell the difference
emu/ CPU emulator: compiles kernel.cl as C++ and runs it with a 32- or 64-wide sub-group
../proto-cuda/packs/<seed>/kernel.cl the OpenCL C kernel of each pack, written by proto-metal/igneum-bench --export-pack
```
The packs stay in `proto-cuda/packs/` because the three implementations share one `program.h`, `vectors.h` and
`memhard.h` per seed; `kernel.cl` sits next to `kernel.cu` and `program.metal`. Three packs are checked in:
`igneum-genesis-mh` (memory-hard dataset, the current construction), `igneum-genesis` and `igneum-hourly`
(closed-form dataset, kept for comparison).
## What the OpenCL path does differently
| Item | CUDA (`host.cu`) | OpenCL (`host.c`) |
|---|---|---|
| Kernel compile | ahead of time by nvcc | at runtime by the driver from `kernel.cl`, with `-D IGNEUM_GROUP=32W -D IGNEUM_EXCHANGE=M` |
| 32-lane exchange | `__shfl_xor_sync` | `sub_group_shuffle_xor` only when the device's sub-group size is exactly 32 for a 32-item work-group; otherwise `__local` memory and a barrier (`WAVEFRONT.md`) |
| Timing | cudaEvent | event profiling (`CL_PROFILING_COMMAND_START/END`); wall time on Apple, whose OpenCL timestamps are unusable |
| Device choice | `--device D` | `--list`, then `--device D` by flat index over all platforms; default is the first GPU |
| Language | C++17 | C99, so MSVC builds it without nvcc or any C++ toolchain |
Everything else (dataset fill or build, cache check against the Mac's FNV, dataset self-test, 3 vector warps standalone
and inside the warm-up batch, 5 timed batches of 2^24, the sweep, the summary table, exit codes) mirrors `host.cu`
line for line. Rates have the same definition: Mhash/s, and GB/s useful = loads per hash x 4 bytes x hashes/s.
## Build
Linux (AMD ROCm, AMD Adrenalin for Linux, or Mesa rusticl; any of them registers with the system ICD loader):
```
sudo apt install ocl-icd-opencl-dev opencl-headers clinfo # Debian or Ubuntu; Fedora: ocl-icd-devel opencl-headers clinfo
clinfo | head -40 # the AMD device must appear here before anything else matters
cd proto-opencl
./build.sh igneum-genesis-mh # -> ./igneum-bench-cl-igneum-genesis-mh
./build.sh igneum-genesis # closed-form packs
./build.sh igneum-hourly
```
The one command behind the script, for a pack `P`:
```
cc -std=c99 -O2 -Wall -Wextra -I ../proto-cuda/packs/P -DIGNEUM_KERNEL_PATH='"../proto-cuda/packs/P/kernel.cl"' -o igneum-bench-cl-P host.c -lOpenCL -ldl
```
Windows (AMD Adrenalin driver installed; it provides `OpenCL.dll` and the AMD ICD, nothing else is needed to run):
1. Install Visual Studio 2022 or 2026 Build Tools with "Desktop development with C++". No CUDA, no nvcc.
2. Get the OpenCL headers and the import library `OpenCL.lib`, any one of these:
`vcpkg install opencl:x64-windows` then `set OPENCL_SDK=C:\vcpkg\installed\x64-windows`;
or the Khronos OpenCL-SDK release zip (GitHub KhronosGroup/OpenCL-SDK) unpacked anywhere, `set OPENCL_SDK=<that dir>`;
or an installed CUDA Toolkit, which ships both under `%CUDA_PATH%` (the script finds it on its own).
3. From an "x64 Native Tools Command Prompt":
```
cd proto-opencl
build.bat igneum-genesis-mh
build.bat igneum-genesis
build.bat igneum-hourly
```
The one command behind the script:
```
cl /nologo /O2 /W3 /std:c11 /I "%OPENCL_SDK%\include" /I "..\proto-cuda\packs\P" /DIGNEUM_KERNEL_PATH="\"../proto-cuda/packs/P/kernel.cl\"" /Fe:igneum-bench-cl-P.exe host.c /link "%OPENCL_SDK%\lib\OpenCL.lib"
```
macOS (correctness check only; Apple deprecated OpenCL in macOS 10.14 and the runtime is 1.2):
```
./build.sh igneum-genesis-mh # Apple OpenCL.framework
./build.sh igneum-genesis-mh khr # Khronos ICD loader from Homebrew, for pocl: brew install opencl-headers opencl-icd-loader pocl
OCL_ICD_VENDORS=/opt/homebrew/etc/OpenCL/vendors SDKROOT=$(xcrun --show-sdk-path) ./igneum-bench-cl-igneum-genesis-mh-khr --list
```
(`SDKROOT` is needed because pocl links its kernels with the system linker and otherwise cannot find `-lSystem`.)
## Run on the AMD rig
Run from `proto-opencl/` so the default kernel path resolves, or pass `--kernel`. Exactly this, in this order, and paste
the whole stdout back:
Windows:
```
cd proto-opencl
igneum-bench-cl-igneum-genesis-mh.exe --list
igneum-bench-cl-igneum-genesis-mh.exe > amd-mh-auto.txt
igneum-bench-cl-igneum-genesis-mh.exe --exchange local > amd-mh-local.txt
igneum-bench-cl-igneum-genesis-mh.exe --group-warps 2 > amd-mh-gw2.txt
igneum-bench-cl-igneum-genesis-mh.exe --sweep > amd-mh-sweep.txt
igneum-bench-cl-igneum-genesis.exe > amd-genesis.txt
igneum-bench-cl-igneum-hourly.exe > amd-hourly.txt
```
Linux:
```
cd proto-opencl
./igneum-bench-cl-igneum-genesis-mh --list
./igneum-bench-cl-igneum-genesis-mh | tee amd-mh-auto.txt
./igneum-bench-cl-igneum-genesis-mh --exchange local | tee amd-mh-local.txt
./igneum-bench-cl-igneum-genesis-mh --group-warps 2 | tee amd-mh-gw2.txt
./igneum-bench-cl-igneum-genesis-mh --sweep | tee amd-mh-sweep.txt
./igneum-bench-cl-igneum-genesis | tee amd-genesis.txt
./igneum-bench-cl-igneum-hourly | tee amd-hourly.txt
```
If `--list` shows more than one device (an iGPU, a CPU runtime, an NVIDIA card in the same box), add `--device D`
with the AMD card's index to every line. If the auto run says `exchange: sub_group_shuffle_xor`, also run
`--exchange subgroup` once so the log has an explicit sub-group run; if it says local memory, the `--exchange local`
run is a repeat and that is fine.
Flags: `--list`, `--device D`, `--dataset-mib N` (power of two, default 1024), `--sweep` (4, 64, 256, 512, 1024 MiB),
`--batch-log2 B` (default 24), `--batches N` (default 5), `--group-warps W` (32-lane units per work-group, 1..8,
default 1), `--exchange auto|local|subgroup`, `--kernel path`, `--build-opts "..."` (appended to clBuildProgram),
`--time event|wall`.
## What PASS looks like
In order, the run prints:
1. Every platform and device, with vendor, driver, OpenCL C version, compute units, memory, local memory, the
sub-group extension it lists, and (AMD) the wavefront width or (NVIDIA) the warp size. The chosen device is starred.
2. `build options:` and `exchange:`. The exchange line is the one that matters on AMD; it says which path was taken and
why, for example `local-memory exchange with barrier (the sub-group size for a 32-item work-group is not 32; queried
sub-group size 64)` on a wave64 card, or `sub_group_shuffle_xor (cl_khr_subgroup_shuffle), sub-group size 32 for a
32-item work-group` on a wave32 one. Both are correct by construction (`WAVEFRONT.md`).
3. `kernel:` (work-group limit and local memory of `igneum_hash`), `program:`, `seed words:`, `day`.
4. Memory-hard packs: cache fill time on the device (twice) and on one host thread, then `cache check: PASS (device ==
host all 67108864 words PASS, host FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head 16 vs Mac PASS,
last line vs Mac PASS)`. The FNV is the same on every machine that has ever run this construction for day
2026-10-03.
5. `dataset build` (or `dataset fill`) times, then `dataset self-test: PASS (head 16 vs Mac PASS, element [MASK] vs Mac
PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)`.
6. Six `verify warp ... PASS` lines: three standalone (bases 0, 4096, 1000000) and the same three read out of the
warm-up batch.
7. `batch fingerprint`: FNV-1a 64 of every output of the warm-up batch. At `--batch-log2 13` every implementation so far
prints `f99fb375b3abeaf5` for igneum-genesis-mh (Apple OpenCL, pocl on both exchange paths, the emulator in seven
configurations); an AMD run at `--batch-log2 13` must print the same value. At the default 2^24 the value is a new
reference to compare AMD against the next machine.
8. `timed:` with device and wall time, then the `rate` line: Mhash/s and GB/s useful.
9. The summary table and `OVERALL: PASS`. Exit code 0 on PASS, 1 on FAIL, 2 on an OpenCL error or a build failure (the
build log is printed in full).
A FAIL with a few lanes differing points at the exchange; a FAIL in every lane at an arithmetic op or the dataset; a
cache FAIL at the ChaCha fill. Send the whole printout either way, including the `--list` output and, on Linux,
`clinfo` and the ROCm or Mesa version, or on Windows the Adrenalin driver version from the AMD Software panel.
## What to paste back
All seven text files above, unedited. From them the bench log gets: device name and driver string, the `exchange:`
line, the cache check line, the 96/96 vector result per pack, the rate at 1 GiB, and the sweep table. If a run looks
odd, run it again and keep both.
## Results so far (no AMD)
| Machine, runtime | Exchange | igneum-genesis-mh | closed-form packs | Rate at 1 GiB |
|---|---|---|---|---|
| Apple M5 Max, Apple OpenCL 1.2 (deprecated runtime) | local memory (runtime lists no sub-group extension) | cache FNV = Mac, 96/96; also 96/96 at `--group-warps 2` and `4` | 96/96 and 96/96 | 45.03 Mhash/s, 18.73 GB/s useful (Metal on the same chip: 45.2). Apple OpenCL on the M5 Max, not an AMD number |
| Apple M5 Max, pocl 7.2 CPU device, OpenCL 3.0, LLVM 23 | `sub_group_shuffle_xor` (auto and `--exchange subgroup`; pocl's `clGetKernelSubGroupInfoKHR` fails with CL_INVALID_OPERATION, the probe kernel reports 32), and `--exchange local` | cache FNV = Mac, 96/96 on both paths, same batch fingerprint as Apple and the emulator at 2^13 | not run | CPU, not meaningful |
| CPU emulator (`emu/`), 7 configurations incl. sub-group 64 | both | 96/96 in every configuration, identical batch fingerprint `f99fb375b3abeaf5` | not run | none |
Sweep on Apple OpenCL (3 batches, wall time): 4 MiB 573.7, 64 MiB 178.9, 256 MiB 94.3, 512 MiB 68.8, 1024 MiB 45.0 Mhash/s.
Full tables in `docs/bench-log.md`.
## Checking a pack without a GPU
```
emu/emu.sh igneum-genesis-mh 0 32 --sg 32 # local-memory exchange, work-group 32, 32-wide sub-group
emu/emu.sh igneum-genesis-mh 0 64 --sg 64 # local-memory exchange on a wave64 model, two units per work-group
emu/emu.sh igneum-genesis-mh 1 32 --sg 32 # sub_group_shuffle_xor, 32-wide sub-group
emu/emu.sh igneum-genesis-mh 1 64 --sg 64 # sub_group_shuffle_xor over a 64-wide sub-group (two units in one wave)
```
The emulator compiles `kernel.cl` unchanged as C++ (`emu/emu_opencl.h` supplies the OpenCL built-ins; the kernel's own
prelude includes it when no OpenCL compiler is present) and runs work-groups on host threads with a barrier inside
every exchange. It prints the cache check, dataset self-test, vectors and a fingerprint of the whole 2^13-hash batch,
so configurations can be compared for every nonce, not only the vector warps. It says nothing about any GPU.
## Regenerating a pack
```
cd proto-metal
swiftc -O -o igneum-bench main.swift -framework Metal
./igneum-bench --seed igneum-genesis --export-pack ../proto-cuda/packs/igneum-genesis-mh
./igneum-bench --closed-form --seed igneum-genesis --export-pack ../proto-cuda/packs/igneum-genesis
./igneum-bench --closed-form --seed igneum-hourly --export-pack ../proto-cuda/packs/igneum-hourly
```
The exporter writes `kernel.cl` next to `kernel.cu` from the same instruction list, refuses to write unless the Metal
GPU matches the CPU interpreter on all 96 vector outputs, and for memory-hard packs unless the GPU cache equals the
CPU cache word for word. The memory-hard core is emitted once (`emitMemhardCore`) in three dialects (Metal, CUDA C++,
OpenCL C), so the constants in `kernel.cl` are the literals of `memhard.h`.