Exporter writes kernel.cl next to kernel.cu (same instruction list; memory-hard core emitted in a third, OpenCL C dialect with the same literals as memhard.h). Pack headers are now C99-safe so a plain C host can include them. proto-opencl/host.c: C99 + OpenCL 1.2 API, device list, runtime build, cache fill and FNV check, dataset build and self-test, 3 vector warps standalone and in batch, bench and sweep as host.cu, whole-batch fingerprint. The 32-lane exchange is sub_group_shuffle_xor only when the queried sub-group size for a 32-item work-group is exactly 32; otherwise a local-memory exchange with one barrier per exchange, so wave64 hardware cannot change the hash (WAVEFRONT.md). build.sh (macOS, Linux), build.bat (MSVC), README with the exact AMD-rig commands. Proven without AMD silicon: Apple OpenCL 1.2 on the M5 Max 96/96 on all three packs (45.0 Mhash/s at 1 GiB, Apple number, not AMD); pocl 7.2 CPU device 96/96 on both exchange paths including the real sub_group_shuffle_xor text; CPU emulator 7 configurations incl. 64-wide sub-groups, identical fingerprint f99fb375b3abeaf5 everywhere. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|---|---|---|
| .. | ||
| emu | ||
| .gitignore | ||
| build.bat | ||
| build.sh | ||
| host.c | ||
| README.md | ||
| WAVEFRONT.md | ||
igneum-bench-cl (proto-opencl)
The portable third path for Igneum's random-program proof-of-work kernels, after Apple Metal (proto-metal) and
NVIDIA CUDA (proto-cuda). It runs the same program packs through OpenCL 1.2, which is what an AMD card exposes on
both Windows (Adrenalin driver) and Linux (ROCm, or Mesa), and checks every output bit for bit against the Mac.
This is a test harness, not a miner. No pool, no network, no wallet, no mining protocol. It fills a dataset, checks the device against known answers, and times the kernel. Nothing here earns anything.
Status on 3 October 2026: no AMD device has run this yet. Everything that could be proven without one has been
(WAVEFRONT.md, "What was proven"): Apple's deprecated OpenCL 1.2 runtime on the M5 Max, pocl 7.2 on the Mac's CPU,
and a CPU emulator with 32- and 64-wide sub-groups all give the Mac's 96/96 vectors and cache FNV for the memory-hard
pack. The AMD run itself is the next step and the commands for it are below.
Layout
proto-opencl/
host.c C99 host: device list, runtime kernel build, cache + dataset fill, self-tests, vectors, bench, sweep
build.sh macOS (-framework OpenCL, or the Khronos ICD loader) and Linux (-lOpenCL)
build.bat Windows (MSVC cl.exe + OpenCL.lib)
WAVEFRONT.md wave32 vs wave64 on AMD, and why the kernel cannot tell the difference
emu/ CPU emulator: compiles kernel.cl as C++ and runs it with a 32- or 64-wide sub-group
../proto-cuda/packs/<seed>/kernel.cl the OpenCL C kernel of each pack, written by proto-metal/igneum-bench --export-pack
The packs stay in proto-cuda/packs/ because the three implementations share one program.h, vectors.h and
memhard.h per seed; kernel.cl sits next to kernel.cu and program.metal. Three packs are checked in:
igneum-genesis-mh (memory-hard dataset, the current construction), igneum-genesis and igneum-hourly
(closed-form dataset, kept for comparison).
What the OpenCL path does differently
| Item | CUDA (host.cu) |
OpenCL (host.c) |
|---|---|---|
| Kernel compile | ahead of time by nvcc | at runtime by the driver from kernel.cl, with -D IGNEUM_GROUP=32W -D IGNEUM_EXCHANGE=M |
| 32-lane exchange | __shfl_xor_sync |
sub_group_shuffle_xor only when the device's sub-group size is exactly 32 for a 32-item work-group; otherwise __local memory and a barrier (WAVEFRONT.md) |
| Timing | cudaEvent | event profiling (CL_PROFILING_COMMAND_START/END); wall time on Apple, whose OpenCL timestamps are unusable |
| Device choice | --device D |
--list, then --device D by flat index over all platforms; default is the first GPU |
| Language | C++17 | C99, so MSVC builds it without nvcc or any C++ toolchain |
Everything else (dataset fill or build, cache check against the Mac's FNV, dataset self-test, 3 vector warps standalone
and inside the warm-up batch, 5 timed batches of 2^24, the sweep, the summary table, exit codes) mirrors host.cu
line for line. Rates have the same definition: Mhash/s, and GB/s useful = loads per hash x 4 bytes x hashes/s.
Build
Linux (AMD ROCm, AMD Adrenalin for Linux, or Mesa rusticl; any of them registers with the system ICD loader):
sudo apt install ocl-icd-opencl-dev opencl-headers clinfo # Debian or Ubuntu; Fedora: ocl-icd-devel opencl-headers clinfo
clinfo | head -40 # the AMD device must appear here before anything else matters
cd proto-opencl
./build.sh igneum-genesis-mh # -> ./igneum-bench-cl-igneum-genesis-mh
./build.sh igneum-genesis # closed-form packs
./build.sh igneum-hourly
The one command behind the script, for a pack P:
cc -std=c99 -O2 -Wall -Wextra -I ../proto-cuda/packs/P -DIGNEUM_KERNEL_PATH='"../proto-cuda/packs/P/kernel.cl"' -o igneum-bench-cl-P host.c -lOpenCL -ldl
Windows (AMD Adrenalin driver installed; it provides OpenCL.dll and the AMD ICD, nothing else is needed to run):
- Install Visual Studio 2022 or 2026 Build Tools with "Desktop development with C++". No CUDA, no nvcc.
- Get the OpenCL headers and the import library
OpenCL.lib, any one of these:vcpkg install opencl:x64-windowsthenset OPENCL_SDK=C:\vcpkg\installed\x64-windows; or the Khronos OpenCL-SDK release zip (GitHub KhronosGroup/OpenCL-SDK) unpacked anywhere,set OPENCL_SDK=<that dir>; or an installed CUDA Toolkit, which ships both under%CUDA_PATH%(the script finds it on its own). - From an "x64 Native Tools Command Prompt":
cd proto-opencl
build.bat igneum-genesis-mh
build.bat igneum-genesis
build.bat igneum-hourly
The one command behind the script:
cl /nologo /O2 /W3 /std:c11 /I "%OPENCL_SDK%\include" /I "..\proto-cuda\packs\P" /DIGNEUM_KERNEL_PATH="\"../proto-cuda/packs/P/kernel.cl\"" /Fe:igneum-bench-cl-P.exe host.c /link "%OPENCL_SDK%\lib\OpenCL.lib"
macOS (correctness check only; Apple deprecated OpenCL in macOS 10.14 and the runtime is 1.2):
./build.sh igneum-genesis-mh # Apple OpenCL.framework
./build.sh igneum-genesis-mh khr # Khronos ICD loader from Homebrew, for pocl: brew install opencl-headers opencl-icd-loader pocl
OCL_ICD_VENDORS=/opt/homebrew/etc/OpenCL/vendors SDKROOT=$(xcrun --show-sdk-path) ./igneum-bench-cl-igneum-genesis-mh-khr --list
(SDKROOT is needed because pocl links its kernels with the system linker and otherwise cannot find -lSystem.)
Run on the AMD rig
Run from proto-opencl/ so the default kernel path resolves, or pass --kernel. Exactly this, in this order, and paste
the whole stdout back:
Windows:
cd proto-opencl
igneum-bench-cl-igneum-genesis-mh.exe --list
igneum-bench-cl-igneum-genesis-mh.exe > amd-mh-auto.txt
igneum-bench-cl-igneum-genesis-mh.exe --exchange local > amd-mh-local.txt
igneum-bench-cl-igneum-genesis-mh.exe --group-warps 2 > amd-mh-gw2.txt
igneum-bench-cl-igneum-genesis-mh.exe --sweep > amd-mh-sweep.txt
igneum-bench-cl-igneum-genesis.exe > amd-genesis.txt
igneum-bench-cl-igneum-hourly.exe > amd-hourly.txt
Linux:
cd proto-opencl
./igneum-bench-cl-igneum-genesis-mh --list
./igneum-bench-cl-igneum-genesis-mh | tee amd-mh-auto.txt
./igneum-bench-cl-igneum-genesis-mh --exchange local | tee amd-mh-local.txt
./igneum-bench-cl-igneum-genesis-mh --group-warps 2 | tee amd-mh-gw2.txt
./igneum-bench-cl-igneum-genesis-mh --sweep | tee amd-mh-sweep.txt
./igneum-bench-cl-igneum-genesis | tee amd-genesis.txt
./igneum-bench-cl-igneum-hourly | tee amd-hourly.txt
If --list shows more than one device (an iGPU, a CPU runtime, an NVIDIA card in the same box), add --device D
with the AMD card's index to every line. If the auto run says exchange: sub_group_shuffle_xor, also run
--exchange subgroup once so the log has an explicit sub-group run; if it says local memory, the --exchange local
run is a repeat and that is fine.
Flags: --list, --device D, --dataset-mib N (power of two, default 1024), --sweep (4, 64, 256, 512, 1024 MiB),
--batch-log2 B (default 24), --batches N (default 5), --group-warps W (32-lane units per work-group, 1..8,
default 1), --exchange auto|local|subgroup, --kernel path, --build-opts "..." (appended to clBuildProgram),
--time event|wall.
What PASS looks like
In order, the run prints:
- Every platform and device, with vendor, driver, OpenCL C version, compute units, memory, local memory, the sub-group extension it lists, and (AMD) the wavefront width or (NVIDIA) the warp size. The chosen device is starred.
build options:andexchange:. The exchange line is the one that matters on AMD; it says which path was taken and why, for examplelocal-memory exchange with barrier (the sub-group size for a 32-item work-group is not 32; queried sub-group size 64)on a wave64 card, orsub_group_shuffle_xor (cl_khr_subgroup_shuffle), sub-group size 32 for a 32-item work-groupon a wave32 one. Both are correct by construction (WAVEFRONT.md).kernel:(work-group limit and local memory ofigneum_hash),program:,seed words:,day.- Memory-hard packs: cache fill time on the device (twice) and on one host thread, then
cache check: PASS (device == host all 67108864 words PASS, host FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head 16 vs Mac PASS, last line vs Mac PASS). The FNV is the same on every machine that has ever run this construction for day 2026-10-03. dataset build(ordataset fill) times, thendataset self-test: PASS (head 16 vs Mac PASS, element [MASK] vs Mac PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS).- Six
verify warp ... PASSlines: three standalone (bases 0, 4096, 1000000) and the same three read out of the warm-up batch. batch fingerprint: FNV-1a 64 of every output of the warm-up batch. At--batch-log2 13every implementation so far printsf99fb375b3abeaf5for igneum-genesis-mh (Apple OpenCL, pocl on both exchange paths, the emulator in seven configurations); an AMD run at--batch-log2 13must print the same value. At the default 2^24 the value is a new reference to compare AMD against the next machine.timed:with device and wall time, then therateline: Mhash/s and GB/s useful.- The summary table and
OVERALL: PASS. Exit code 0 on PASS, 1 on FAIL, 2 on an OpenCL error or a build failure (the build log is printed in full).
A FAIL with a few lanes differing points at the exchange; a FAIL in every lane at an arithmetic op or the dataset; a
cache FAIL at the ChaCha fill. Send the whole printout either way, including the --list output and, on Linux,
clinfo and the ROCm or Mesa version, or on Windows the Adrenalin driver version from the AMD Software panel.
What to paste back
All seven text files above, unedited. From them the bench log gets: device name and driver string, the exchange:
line, the cache check line, the 96/96 vector result per pack, the rate at 1 GiB, and the sweep table. If a run looks
odd, run it again and keep both.
Results so far (no AMD)
| Machine, runtime | Exchange | igneum-genesis-mh | closed-form packs | Rate at 1 GiB |
|---|---|---|---|---|
| Apple M5 Max, Apple OpenCL 1.2 (deprecated runtime) | local memory (runtime lists no sub-group extension) | cache FNV = Mac, 96/96; also 96/96 at --group-warps 2 and 4 |
96/96 and 96/96 | 45.03 Mhash/s, 18.73 GB/s useful (Metal on the same chip: 45.2). Apple OpenCL on the M5 Max, not an AMD number |
| Apple M5 Max, pocl 7.2 CPU device, OpenCL 3.0, LLVM 23 | sub_group_shuffle_xor (auto and --exchange subgroup; pocl's clGetKernelSubGroupInfoKHR fails with CL_INVALID_OPERATION, the probe kernel reports 32), and --exchange local |
cache FNV = Mac, 96/96 on both paths, same batch fingerprint as Apple and the emulator at 2^13 | not run | CPU, not meaningful |
CPU emulator (emu/), 7 configurations incl. sub-group 64 |
both | 96/96 in every configuration, identical batch fingerprint f99fb375b3abeaf5 |
not run | none |
Sweep on Apple OpenCL (3 batches, wall time): 4 MiB 573.7, 64 MiB 178.9, 256 MiB 94.3, 512 MiB 68.8, 1024 MiB 45.0 Mhash/s.
Full tables in docs/bench-log.md.
Checking a pack without a GPU
emu/emu.sh igneum-genesis-mh 0 32 --sg 32 # local-memory exchange, work-group 32, 32-wide sub-group
emu/emu.sh igneum-genesis-mh 0 64 --sg 64 # local-memory exchange on a wave64 model, two units per work-group
emu/emu.sh igneum-genesis-mh 1 32 --sg 32 # sub_group_shuffle_xor, 32-wide sub-group
emu/emu.sh igneum-genesis-mh 1 64 --sg 64 # sub_group_shuffle_xor over a 64-wide sub-group (two units in one wave)
The emulator compiles kernel.cl unchanged as C++ (emu/emu_opencl.h supplies the OpenCL built-ins; the kernel's own
prelude includes it when no OpenCL compiler is present) and runs work-groups on host threads with a barrier inside
every exchange. It prints the cache check, dataset self-test, vectors and a fingerprint of the whole 2^13-hash batch,
so configurations can be compared for every nonce, not only the vector warps. It says nothing about any GPU.
Regenerating a pack
cd proto-metal
swiftc -O -o igneum-bench main.swift -framework Metal
./igneum-bench --seed igneum-genesis --export-pack ../proto-cuda/packs/igneum-genesis-mh
./igneum-bench --closed-form --seed igneum-genesis --export-pack ../proto-cuda/packs/igneum-genesis
./igneum-bench --closed-form --seed igneum-hourly --export-pack ../proto-cuda/packs/igneum-hourly
The exporter writes kernel.cl next to kernel.cu from the same instruction list, refuses to write unless the Metal
GPU matches the CPU interpreter on all 96 vector outputs, and for memory-hard packs unless the GPU cache equals the
CPU cache word for word. The memory-hard core is emitted once (emitMemhardCore) in three dialects (Metal, CUDA C++,
OpenCL C), so the constants in kernel.cl are the literals of memhard.h.