igneum/proto-opencl
igneum-labs 3c2e7aad81 GPU workers on the real hash: serve protocol in Metal, CUDA and OpenCL, Windows mining package, devnet v1 log
proto-metal --serve compiles igneum_hash_bound at runtime and mines jobs from stdin (init words in buffer 3, program
and dataset cached per seed); proto-cuda and proto-opencl host serve modes from the pack's kernel_bound.cu / .cl with a
seed guard; --vendor device filter for OpenCL. windows-miner/: START-MINING.bat + start-mining.ps1 (GPU and tool
detection, pack export, cached builds, MINERS identities per vendor, status every 30 s, uploads every 60 s, rebuild on
seed change, Ctrl+C summary), README.txt, make-package.sh (igneum-mine-test.zip with a cross-compiled miner).
docs: fork-divergence devnet v1, bench-log entry with the CPU, Metal and overnight numbers.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-03 19:18:19 +00:00
..
emu proto-opencl: OpenCL path for AMD, proven on Apple OpenCL, pocl and a wave64 CPU emulator 2026-10-03 16:52:24 +00:00
.gitignore proto-opencl: OpenCL path for AMD, proven on Apple OpenCL, pocl and a wave64 CPU emulator 2026-10-03 16:52:24 +00:00
build.bat OpenCL build.bat: library detection without parentheses in blocks (Program Files (x86) broke the parser) 2026-10-03 17:04:39 +00:00
build.sh proto-opencl: OpenCL path for AMD, proven on Apple OpenCL, pocl and a wave64 CPU emulator 2026-10-03 16:52:24 +00:00
host.c GPU workers on the real hash: serve protocol in Metal, CUDA and OpenCL, Windows mining package, devnet v1 log 2026-10-03 19:18:19 +00:00
README.md proto-opencl: OpenCL path for AMD, proven on Apple OpenCL, pocl and a wave64 CPU emulator 2026-10-03 16:52:24 +00:00
WAVEFRONT.md proto-opencl: OpenCL path for AMD, proven on Apple OpenCL, pocl and a wave64 CPU emulator 2026-10-03 16:52:24 +00:00

igneum-bench-cl (proto-opencl)

The portable third path for Igneum's random-program proof-of-work kernels, after Apple Metal (proto-metal) and NVIDIA CUDA (proto-cuda). It runs the same program packs through OpenCL 1.2, which is what an AMD card exposes on both Windows (Adrenalin driver) and Linux (ROCm, or Mesa), and checks every output bit for bit against the Mac.

This is a test harness, not a miner. No pool, no network, no wallet, no mining protocol. It fills a dataset, checks the device against known answers, and times the kernel. Nothing here earns anything.

Status on 3 October 2026: no AMD device has run this yet. Everything that could be proven without one has been (WAVEFRONT.md, "What was proven"): Apple's deprecated OpenCL 1.2 runtime on the M5 Max, pocl 7.2 on the Mac's CPU, and a CPU emulator with 32- and 64-wide sub-groups all give the Mac's 96/96 vectors and cache FNV for the memory-hard pack. The AMD run itself is the next step and the commands for it are below.

Layout

proto-opencl/
  host.c          C99 host: device list, runtime kernel build, cache + dataset fill, self-tests, vectors, bench, sweep
  build.sh        macOS (-framework OpenCL, or the Khronos ICD loader) and Linux (-lOpenCL)
  build.bat       Windows (MSVC cl.exe + OpenCL.lib)
  WAVEFRONT.md    wave32 vs wave64 on AMD, and why the kernel cannot tell the difference
  emu/            CPU emulator: compiles kernel.cl as C++ and runs it with a 32- or 64-wide sub-group
../proto-cuda/packs/<seed>/kernel.cl   the OpenCL C kernel of each pack, written by proto-metal/igneum-bench --export-pack

The packs stay in proto-cuda/packs/ because the three implementations share one program.h, vectors.h and memhard.h per seed; kernel.cl sits next to kernel.cu and program.metal. Three packs are checked in: igneum-genesis-mh (memory-hard dataset, the current construction), igneum-genesis and igneum-hourly (closed-form dataset, kept for comparison).

What the OpenCL path does differently

Item CUDA (host.cu) OpenCL (host.c)
Kernel compile ahead of time by nvcc at runtime by the driver from kernel.cl, with -D IGNEUM_GROUP=32W -D IGNEUM_EXCHANGE=M
32-lane exchange __shfl_xor_sync sub_group_shuffle_xor only when the device's sub-group size is exactly 32 for a 32-item work-group; otherwise __local memory and a barrier (WAVEFRONT.md)
Timing cudaEvent event profiling (CL_PROFILING_COMMAND_START/END); wall time on Apple, whose OpenCL timestamps are unusable
Device choice --device D --list, then --device D by flat index over all platforms; default is the first GPU
Language C++17 C99, so MSVC builds it without nvcc or any C++ toolchain

Everything else (dataset fill or build, cache check against the Mac's FNV, dataset self-test, 3 vector warps standalone and inside the warm-up batch, 5 timed batches of 2^24, the sweep, the summary table, exit codes) mirrors host.cu line for line. Rates have the same definition: Mhash/s, and GB/s useful = loads per hash x 4 bytes x hashes/s.

Build

Linux (AMD ROCm, AMD Adrenalin for Linux, or Mesa rusticl; any of them registers with the system ICD loader):

sudo apt install ocl-icd-opencl-dev opencl-headers clinfo     # Debian or Ubuntu; Fedora: ocl-icd-devel opencl-headers clinfo
clinfo | head -40                                            # the AMD device must appear here before anything else matters
cd proto-opencl
./build.sh igneum-genesis-mh                                  # -> ./igneum-bench-cl-igneum-genesis-mh
./build.sh igneum-genesis                                     # closed-form packs
./build.sh igneum-hourly

The one command behind the script, for a pack P:

cc -std=c99 -O2 -Wall -Wextra -I ../proto-cuda/packs/P -DIGNEUM_KERNEL_PATH='"../proto-cuda/packs/P/kernel.cl"' -o igneum-bench-cl-P host.c -lOpenCL -ldl

Windows (AMD Adrenalin driver installed; it provides OpenCL.dll and the AMD ICD, nothing else is needed to run):

  1. Install Visual Studio 2022 or 2026 Build Tools with "Desktop development with C++". No CUDA, no nvcc.
  2. Get the OpenCL headers and the import library OpenCL.lib, any one of these: vcpkg install opencl:x64-windows then set OPENCL_SDK=C:\vcpkg\installed\x64-windows; or the Khronos OpenCL-SDK release zip (GitHub KhronosGroup/OpenCL-SDK) unpacked anywhere, set OPENCL_SDK=<that dir>; or an installed CUDA Toolkit, which ships both under %CUDA_PATH% (the script finds it on its own).
  3. From an "x64 Native Tools Command Prompt":
cd proto-opencl
build.bat igneum-genesis-mh
build.bat igneum-genesis
build.bat igneum-hourly

The one command behind the script:

cl /nologo /O2 /W3 /std:c11 /I "%OPENCL_SDK%\include" /I "..\proto-cuda\packs\P" /DIGNEUM_KERNEL_PATH="\"../proto-cuda/packs/P/kernel.cl\"" /Fe:igneum-bench-cl-P.exe host.c /link "%OPENCL_SDK%\lib\OpenCL.lib"

macOS (correctness check only; Apple deprecated OpenCL in macOS 10.14 and the runtime is 1.2):

./build.sh igneum-genesis-mh            # Apple OpenCL.framework
./build.sh igneum-genesis-mh khr        # Khronos ICD loader from Homebrew, for pocl: brew install opencl-headers opencl-icd-loader pocl
OCL_ICD_VENDORS=/opt/homebrew/etc/OpenCL/vendors SDKROOT=$(xcrun --show-sdk-path) ./igneum-bench-cl-igneum-genesis-mh-khr --list

(SDKROOT is needed because pocl links its kernels with the system linker and otherwise cannot find -lSystem.)

Run on the AMD rig

Run from proto-opencl/ so the default kernel path resolves, or pass --kernel. Exactly this, in this order, and paste the whole stdout back:

Windows:

cd proto-opencl
igneum-bench-cl-igneum-genesis-mh.exe --list
igneum-bench-cl-igneum-genesis-mh.exe                          > amd-mh-auto.txt
igneum-bench-cl-igneum-genesis-mh.exe --exchange local         > amd-mh-local.txt
igneum-bench-cl-igneum-genesis-mh.exe --group-warps 2          > amd-mh-gw2.txt
igneum-bench-cl-igneum-genesis-mh.exe --sweep                  > amd-mh-sweep.txt
igneum-bench-cl-igneum-genesis.exe                             > amd-genesis.txt
igneum-bench-cl-igneum-hourly.exe                              > amd-hourly.txt

Linux:

cd proto-opencl
./igneum-bench-cl-igneum-genesis-mh --list
./igneum-bench-cl-igneum-genesis-mh                            | tee amd-mh-auto.txt
./igneum-bench-cl-igneum-genesis-mh --exchange local           | tee amd-mh-local.txt
./igneum-bench-cl-igneum-genesis-mh --group-warps 2            | tee amd-mh-gw2.txt
./igneum-bench-cl-igneum-genesis-mh --sweep                    | tee amd-mh-sweep.txt
./igneum-bench-cl-igneum-genesis                               | tee amd-genesis.txt
./igneum-bench-cl-igneum-hourly                                | tee amd-hourly.txt

If --list shows more than one device (an iGPU, a CPU runtime, an NVIDIA card in the same box), add --device D with the AMD card's index to every line. If the auto run says exchange: sub_group_shuffle_xor, also run --exchange subgroup once so the log has an explicit sub-group run; if it says local memory, the --exchange local run is a repeat and that is fine.

Flags: --list, --device D, --dataset-mib N (power of two, default 1024), --sweep (4, 64, 256, 512, 1024 MiB), --batch-log2 B (default 24), --batches N (default 5), --group-warps W (32-lane units per work-group, 1..8, default 1), --exchange auto|local|subgroup, --kernel path, --build-opts "..." (appended to clBuildProgram), --time event|wall.

What PASS looks like

In order, the run prints:

  1. Every platform and device, with vendor, driver, OpenCL C version, compute units, memory, local memory, the sub-group extension it lists, and (AMD) the wavefront width or (NVIDIA) the warp size. The chosen device is starred.
  2. build options: and exchange:. The exchange line is the one that matters on AMD; it says which path was taken and why, for example local-memory exchange with barrier (the sub-group size for a 32-item work-group is not 32; queried sub-group size 64) on a wave64 card, or sub_group_shuffle_xor (cl_khr_subgroup_shuffle), sub-group size 32 for a 32-item work-group on a wave32 one. Both are correct by construction (WAVEFRONT.md).
  3. kernel: (work-group limit and local memory of igneum_hash), program:, seed words:, day.
  4. Memory-hard packs: cache fill time on the device (twice) and on one host thread, then cache check: PASS (device == host all 67108864 words PASS, host FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head 16 vs Mac PASS, last line vs Mac PASS). The FNV is the same on every machine that has ever run this construction for day 2026-10-03.
  5. dataset build (or dataset fill) times, then dataset self-test: PASS (head 16 vs Mac PASS, element [MASK] vs Mac PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS).
  6. Six verify warp ... PASS lines: three standalone (bases 0, 4096, 1000000) and the same three read out of the warm-up batch.
  7. batch fingerprint: FNV-1a 64 of every output of the warm-up batch. At --batch-log2 13 every implementation so far prints f99fb375b3abeaf5 for igneum-genesis-mh (Apple OpenCL, pocl on both exchange paths, the emulator in seven configurations); an AMD run at --batch-log2 13 must print the same value. At the default 2^24 the value is a new reference to compare AMD against the next machine.
  8. timed: with device and wall time, then the rate line: Mhash/s and GB/s useful.
  9. The summary table and OVERALL: PASS. Exit code 0 on PASS, 1 on FAIL, 2 on an OpenCL error or a build failure (the build log is printed in full).

A FAIL with a few lanes differing points at the exchange; a FAIL in every lane at an arithmetic op or the dataset; a cache FAIL at the ChaCha fill. Send the whole printout either way, including the --list output and, on Linux, clinfo and the ROCm or Mesa version, or on Windows the Adrenalin driver version from the AMD Software panel.

What to paste back

All seven text files above, unedited. From them the bench log gets: device name and driver string, the exchange: line, the cache check line, the 96/96 vector result per pack, the rate at 1 GiB, and the sweep table. If a run looks odd, run it again and keep both.

Results so far (no AMD)

Machine, runtime Exchange igneum-genesis-mh closed-form packs Rate at 1 GiB
Apple M5 Max, Apple OpenCL 1.2 (deprecated runtime) local memory (runtime lists no sub-group extension) cache FNV = Mac, 96/96; also 96/96 at --group-warps 2 and 4 96/96 and 96/96 45.03 Mhash/s, 18.73 GB/s useful (Metal on the same chip: 45.2). Apple OpenCL on the M5 Max, not an AMD number
Apple M5 Max, pocl 7.2 CPU device, OpenCL 3.0, LLVM 23 sub_group_shuffle_xor (auto and --exchange subgroup; pocl's clGetKernelSubGroupInfoKHR fails with CL_INVALID_OPERATION, the probe kernel reports 32), and --exchange local cache FNV = Mac, 96/96 on both paths, same batch fingerprint as Apple and the emulator at 2^13 not run CPU, not meaningful
CPU emulator (emu/), 7 configurations incl. sub-group 64 both 96/96 in every configuration, identical batch fingerprint f99fb375b3abeaf5 not run none

Sweep on Apple OpenCL (3 batches, wall time): 4 MiB 573.7, 64 MiB 178.9, 256 MiB 94.3, 512 MiB 68.8, 1024 MiB 45.0 Mhash/s. Full tables in docs/bench-log.md.

Checking a pack without a GPU

emu/emu.sh igneum-genesis-mh 0 32 --sg 32     # local-memory exchange, work-group 32, 32-wide sub-group
emu/emu.sh igneum-genesis-mh 0 64 --sg 64     # local-memory exchange on a wave64 model, two units per work-group
emu/emu.sh igneum-genesis-mh 1 32 --sg 32     # sub_group_shuffle_xor, 32-wide sub-group
emu/emu.sh igneum-genesis-mh 1 64 --sg 64     # sub_group_shuffle_xor over a 64-wide sub-group (two units in one wave)

The emulator compiles kernel.cl unchanged as C++ (emu/emu_opencl.h supplies the OpenCL built-ins; the kernel's own prelude includes it when no OpenCL compiler is present) and runs work-groups on host threads with a barrier inside every exchange. It prints the cache check, dataset self-test, vectors and a fingerprint of the whole 2^13-hash batch, so configurations can be compared for every nonce, not only the vector warps. It says nothing about any GPU.

Regenerating a pack

cd proto-metal
swiftc -O -o igneum-bench main.swift -framework Metal
./igneum-bench --seed igneum-genesis --export-pack ../proto-cuda/packs/igneum-genesis-mh
./igneum-bench --closed-form --seed igneum-genesis --export-pack ../proto-cuda/packs/igneum-genesis
./igneum-bench --closed-form --seed igneum-hourly --export-pack ../proto-cuda/packs/igneum-hourly

The exporter writes kernel.cl next to kernel.cu from the same instruction list, refuses to write unless the Metal GPU matches the CPU interpreter on all 96 vector outputs, and for memory-hard packs unless the GPU cache equals the CPU cache word for word. The memory-hard core is emitted once (emitMemhardCore) in three dialects (Metal, CUDA C++, OpenCL C), so the constants in kernel.cl are the literals of memhard.h.