Exporter writes kernel.cl next to kernel.cu (same instruction list; memory-hard core emitted in a third, OpenCL C dialect with the same literals as memhard.h). Pack headers are now C99-safe so a plain C host can include them. proto-opencl/host.c: C99 + OpenCL 1.2 API, device list, runtime build, cache fill and FNV check, dataset build and self-test, 3 vector warps standalone and in batch, bench and sweep as host.cu, whole-batch fingerprint. The 32-lane exchange is sub_group_shuffle_xor only when the queried sub-group size for a 32-item work-group is exactly 32; otherwise a local-memory exchange with one barrier per exchange, so wave64 hardware cannot change the hash (WAVEFRONT.md). build.sh (macOS, Linux), build.bat (MSVC), README with the exact AMD-rig commands. Proven without AMD silicon: Apple OpenCL 1.2 on the M5 Max 96/96 on all three packs (45.0 Mhash/s at 1 GiB, Apple number, not AMD); pocl 7.2 CPU device 96/96 on both exchange paths including the real sub_group_shuffle_xor text; CPU emulator 7 configurations incl. 64-wide sub-groups, identical fingerprint f99fb375b3abeaf5 everywhere. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
8.4 KiB
Wavefront width and the 32-lane verification unit
Written 3 October 2026 for the OpenCL path (proto-opencl). It records the one place where AMD hardware differs
from Apple and NVIDIA in a way the lottery hash can feel, and how kernel.cl is built so that it cannot.
The unit is 32 lanes, by definition
The program pack defines a hash over 32 lanes. Six of its 64 instructions are exchanges (shfl): lane l reads a
register from lane l ^ m with m in {1, 2, 4, 8, 16}. The CPU verifier (proto-metal/main.swift, cpuWarp)
models exactly 32 lanes. Metal runs it on a 32-wide SIMD group (simd_shuffle_xor), CUDA on a 32-wide warp
(__shfl_xor_sync). The vectors in every pack are 3 warps x 32 lanes x 64 bits. Nothing in the definition mentions
hardware; 32 is a parameter of the hash.
What AMD hardware does
| Architecture | Native wave width | Notes |
|---|---|---|
| GCN (Radeon HD 7000 to Vega, approximate) | 64 | always wave64 |
| CDNA (Instinct MI100 to MI300, approximate) | 64 | always wave64 |
| RDNA 1 to 4 (RX 5000 to RX 9000, approximate) | 32 or 64 | the shader compiler picks per kernel; compute kernels are often wave64 unless the compiler decides otherwise, and nothing in OpenCL lets the program demand wave32 |
Figures from memory, labelled approximate. NVIDIA is 32 everywhere; Apple is 32 on every chip the Mac tool has
seen (it warns if threadExecutionWidth is not 32). In OpenCL terms the hardware wave is the sub-group:
get_sub_group_size() and clGetKernelSubGroupInfoKHR(CL_KERNEL_MAX_SUB_GROUP_SIZE_FOR_NDRANGE_KHR) report it.
So on an AMD card a sub-group may be 64 lanes, and a sub-group shuffle (sub_group_shuffle_xor) spans 64 lanes.
Two things could go wrong if the kernel simply used sub-group shuffles:
- A wave64 sub-group holds two logical 32-lane units.
lane ^ mwithm < 32never leaves the unit's own aligned half, so the xor exchange itself would still be correct. The emulator below demonstrates this. But asub_group_broadcast(x, 0)(used only by the optional wide-load leverwload, off by default) would broadcast lane 0 of the wave, which is wrong for the upper unit. - The mapping from sub-group lane to work-item is not something the OpenCL specification promises to be linear and aligned. In practice it is (consecutive work-items in linear local id order), on AMD, NVIDIA, Intel and pocl alike, but "in practice" is not a guarantee a consensus rule should rest on.
The rule kernel.cl follows
The exchange is selected at compile time by IGNEUM_EXCHANGE, and proto-opencl/host.c chooses the value:
IGNEUM_EXCHANGE |
Exchange | When host.c selects it |
|---|---|---|
| 0 (default) | local memory: each lane writes its register to __local memory, barrier(CLK_LOCAL_MEM_FENCE), reads slot lid ^ m |
always available; chosen whenever any condition below fails |
| 1 | sub_group_shuffle_xor (extension cl_khr_subgroup_shuffle, OpenCL C 2.0 or 3.0) |
the device lists the extension, --group-warps 1 (work-group of exactly 32, so one work-group is one unit), the variant compiles, and clGetKernelSubGroupInfoKHR on igneum_hash reports a sub-group size of exactly 32 for a 32-item work-group |
| 2 | intel_sub_group_shuffle_xor (cl_intel_subgroups) |
as 1, on Intel devices without the khr shuffle extension |
In words: the 32-lane group is defined by the work-group of 32 (one work-group = one unit) whenever sub-group
shuffles are used, and the local-memory exchange is used whenever the sub-group size is not exactly 32. On a
wave64 device the queried size is 64, so AMD GCN, CDNA and RDNA-in-wave64 all take the local-memory path, and the
hash does not depend on the wave width at all: a work-group barrier and __local memory mean the same thing at any
wave width. If the sub-group size cannot be queried (clGetKernelSubGroupInfoKHR missing, as on pocl), host.c
runs a probe kernel from the same build that reports get_sub_group_size() for a 32-item work-group; this is
weaker than the per-kernel query because a compiler may choose the wave width per kernel, and the printout says so.
--exchange local forces path 0 on any device; --exchange subgroup demands path 1 or 2 and fails otherwise.
The run prints which path it took and why, on the exchange: line.
The local-memory exchange uses two buffers of IGNEUM_GROUP words that alternate (counter xk), so one barrier per
exchange suffices: a lane can only overwrite buffer b at exchange k + 2 after passing barrier k + 1, and every
lane reaches barrier k + 1 only after its read of buffer b at exchange k. Control flow is uniform (the program
has no branches), so every work-item reaches every barrier. --group-warps W packs W units into one work-group of
32 W items; lid ^ m stays inside the unit because m < 32, and the Apple runs below show W = 1, 2 and 4 bit-exact.
What was proven without AMD silicon (3 October 2026)
| Check | Path | Result |
|---|---|---|
| Apple M5 Max, Apple OpenCL 1.2 runtime, pack igneum-genesis-mh | local memory (the runtime lists no sub-group extension), work-group 32 | cache FNV 48c4f5bf24166b2e = Mac, 96/96 vectors standalone and in batch, 45.03 Mhash/s at 1 GiB |
Same, --exchange local --group-warps 2 and --group-warps 4 |
local memory, work-groups of 64 and 128 | 96/96 and 96/96 |
| Same, closed-form packs igneum-genesis and igneum-hourly | local memory | 96/96 and 96/96 |
pocl 7.2 CPU device (OpenCL 3.0, LLVM 23), Khronos ICD loader, --exchange auto and --exchange subgroup |
sub_group_shuffle_xor, OpenCL C 3.0 (pocl's clGetKernelSubGroupInfoKHR fails with CL_INVALID_OPERATION, the probe kernel reports a sub-group of 32 for a 32-item work-group) |
cache FNV = Mac, 96/96 |
Same, --exchange local |
local memory | cache FNV = Mac, 96/96 |
CPU emulator, IGNEUM_EXCHANGE 0, work-group 32, sub-group 32 |
local memory | PASS, batch fingerprint f99fb375b3abeaf5 |
CPU emulator, IGNEUM_EXCHANGE 0, work-group 64, sub-group 64 (wave64, two units per wave) |
local memory | PASS, same fingerprint |
CPU emulator, IGNEUM_EXCHANGE 0, work-group 32, sub-group 64 (wave64 half empty) |
local memory | PASS, same fingerprint |
CPU emulator, IGNEUM_EXCHANGE 1, work-group 32, sub-group 32 |
sub_group_shuffle_xor |
PASS, same fingerprint |
CPU emulator, IGNEUM_EXCHANGE 1, work-group 32, sub-group 64 |
sub_group_shuffle_xor over a half-empty wave64 |
PASS, same fingerprint |
CPU emulator, IGNEUM_EXCHANGE 1, work-group 64, sub-group 64 (wave64, two units in one shuffle domain) |
sub_group_shuffle_xor |
PASS, same fingerprint |
CPU emulator, IGNEUM_EXCHANGE 1, work-group 64, sub-group 32 (two sub-groups per work-group) |
sub_group_shuffle_xor |
PASS, same fingerprint |
The fingerprint is FNV-1a 64 over the 2^13 outputs of the batch at base nonce 0; seven emulator configurations
produced the same one, and host.c prints the same fingerprint for its warm-up batch (--batch-log2 13): Apple's
OpenCL (local memory) and pocl (sub-group shuffles, and local memory) all print f99fb375b3abeaf5. So both exchange
implementations, both wave widths, two real OpenCL compilers and the emulator give identical hashes for every nonce
in the batch, not only for the three vector warps. An AMD run at --batch-log2 13 must print the same value. The emulator rows with sub-group 64 and IGNEUM_EXCHANGE 1 show that the xor
exchange alone would survive a wave64 sub-group (point 1 above); host.c still refuses that configuration on a real
device because of point 2 and because of sub_group_broadcast. The conservative rule costs nothing in correctness
and, on a wave64 card, the local-memory path is what runs.
What is not proven here
- No AMD compiler has compiled
kernel.cl, and no AMD device has run it. The emulator is clang; pocl is LLVM on a CPU; Apple's OpenCL is Apple's compiler. Syntax every one of them accepts can still trip AMD's front end, which would be a build log (printed in full by host.c), not a silent difference. - No AMD hash rate exists. The cost of the local-memory exchange relative to sub-group shuffles on AMD is unknown;
on Apple the local-memory path runs at the same 45 Mhash/s as Metal's
simd_shuffle_xorkernel because the hash is bound by the 104 random dataset loads, and the same is expected elsewhere, but expected is not measured. - Whether RDNA compiles
igneum_hashas wave32 (sub-group 32, path 1) or wave64 (path 0) is a driver decision. The run will say which on theexchange:line. Both are correct by construction; only the speed may differ.