igneum/proto-opencl/WAVEFRONT.md
igneum-labs dc84789e34 read-width experiment (gate 1): load classes W=4/16/64, per-load width mix, scratch RMW variant behind a generator flag; 20 packs; CPU verifier and acceptance mirror; emulator shims
Nothing changes for the default class: the pinned packs are byte-identical (tests/packs.rs), the v2 draw stream is untouched.
LoadClass {mix, load_slots, scratch}: fixed widths w16, w64, w64x4 (4 loads of 64 B), era mixes 50/35/15 and 25/50/25 drawn per load with one extra below(100) roll, and the scratch variant scr0/2/4/8 (persistent warps, 1 MiB per warp, tagged lazy fill, measurement only). A wide load reads the W-aligned address and folds every word: x = dst ^ w0; x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]. Program ids carry the class. proto-opencl/host.c taken from opencl-rdna4 23810df (--memprobe, select read-back).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 19:46:54 +00:00

9.1 KiB

Wavefront width and the 32-lane verification unit

Written 3 October 2026 for the OpenCL path (proto-opencl). It records the one place where AMD hardware differs from Apple and NVIDIA in a way the lottery hash can feel, and how kernel.cl is built so that it cannot.

The unit is 32 lanes, by definition

The program pack defines a hash over 32 lanes. Six of its 64 instructions are exchanges (shfl): lane l reads a register from lane l ^ m with m in {1, 2, 4, 8, 16}. The CPU verifier (proto-metal/main.swift, cpuWarp) models exactly 32 lanes. Metal runs it on a 32-wide SIMD group (simd_shuffle_xor), CUDA on a 32-wide warp (__shfl_xor_sync). The vectors in every pack are 3 warps x 32 lanes x 64 bits. Nothing in the definition mentions hardware; 32 is a parameter of the hash.

What AMD hardware does

Architecture Native wave width Notes
GCN (Radeon HD 7000 to Vega, approximate) 64 always wave64
CDNA (Instinct MI100 to MI300, approximate) 64 always wave64
RDNA 1 to 4 (RX 5000 to RX 9000, approximate) 32 or 64 the shader compiler picks per kernel; compute kernels are often wave64 unless the compiler decides otherwise, and nothing in OpenCL lets the program demand wave32

Figures from memory, labelled approximate. NVIDIA is 32 everywhere; Apple is 32 on every chip the Mac tool has seen (it warns if threadExecutionWidth is not 32). In OpenCL terms the hardware wave is the sub-group: get_sub_group_size() and clGetKernelSubGroupInfoKHR(CL_KERNEL_MAX_SUB_GROUP_SIZE_FOR_NDRANGE_KHR) report it.

So on an AMD card a sub-group may be 64 lanes, and a sub-group shuffle (sub_group_shuffle_xor) spans 64 lanes. Two things could go wrong if the kernel simply used sub-group shuffles:

  1. A wave64 sub-group holds two logical 32-lane units. lane ^ m with m < 32 never leaves the unit's own aligned half, so the xor exchange itself would still be correct. The emulator below demonstrates this. But a sub_group_broadcast(x, 0) (used only by the optional wide-load lever wload, off by default) would broadcast lane 0 of the wave, which is wrong for the upper unit.
  2. The mapping from sub-group lane to work-item is not something the OpenCL specification promises to be linear and aligned. In practice it is (consecutive work-items in linear local id order), on AMD, NVIDIA, Intel and pocl alike, but "in practice" is not a guarantee a consensus rule should rest on.

The rule kernel.cl follows

The exchange is selected at compile time by IGNEUM_EXCHANGE, and proto-opencl/host.c chooses the value:

IGNEUM_EXCHANGE Exchange When host.c selects it
0 (default) local memory: each lane writes its register to __local memory, barrier(CLK_LOCAL_MEM_FENCE), reads slot lid ^ m always available; chosen whenever any condition below fails
1 sub_group_shuffle_xor (extension cl_khr_subgroup_shuffle, OpenCL C 2.0 or 3.0) the device lists the extension, --group-warps 1 (work-group of exactly 32, so one work-group is one unit), the variant compiles, and clGetKernelSubGroupInfoKHR on igneum_hash reports a sub-group size of exactly 32 for a 32-item work-group
2 intel_sub_group_shuffle_xor (cl_intel_subgroups) as 1, on Intel devices without the khr shuffle extension

In words: the 32-lane group is defined by the work-group of 32 (one work-group = one unit) whenever sub-group shuffles are used, and the local-memory exchange is used whenever the sub-group size is not exactly 32. On a wave64 device the queried size is 64, so AMD GCN, CDNA and RDNA-in-wave64 all take the local-memory path, and the hash does not depend on the wave width at all: a work-group barrier and __local memory mean the same thing at any wave width. If the sub-group size cannot be queried (clGetKernelSubGroupInfoKHR missing, as on pocl), host.c runs a probe kernel from the same build that reports get_sub_group_size() for a 32-item work-group; this is weaker than the per-kernel query because a compiler may choose the wave width per kernel, and the printout says so. --exchange local forces path 0 on any device; --exchange subgroup demands path 1 or 2 and fails otherwise. The run prints which path it took and why, on the exchange: line.

The local-memory exchange uses two buffers of IGNEUM_GROUP words that alternate (counter xk), so one barrier per exchange suffices: a lane can only overwrite buffer b at exchange k + 2 after passing barrier k + 1, and every lane reaches barrier k + 1 only after its read of buffer b at exchange k. Control flow is uniform (the program has no branches), so every work-item reaches every barrier. --group-warps W packs W units into one work-group of 32 W items; lid ^ m stays inside the unit because m < 32, and the Apple runs below show W = 1, 2 and 4 bit-exact.

What was proven without AMD silicon (3 October 2026)

Check Path Result
Apple M5 Max, Apple OpenCL 1.2 runtime, pack igneum-genesis-mh local memory (the runtime lists no sub-group extension), work-group 32 cache FNV 48c4f5bf24166b2e = Mac, 96/96 vectors standalone and in batch, 45.03 Mhash/s at 1 GiB
Same, --exchange local --group-warps 2 and --group-warps 4 local memory, work-groups of 64 and 128 96/96 and 96/96
Same, closed-form packs igneum-genesis and igneum-hourly local memory 96/96 and 96/96
pocl 7.2 CPU device (OpenCL 3.0, LLVM 23), Khronos ICD loader, --exchange auto and --exchange subgroup sub_group_shuffle_xor, OpenCL C 3.0 (pocl's clGetKernelSubGroupInfoKHR fails with CL_INVALID_OPERATION, the probe kernel reports a sub-group of 32 for a 32-item work-group) cache FNV = Mac, 96/96
Same, --exchange local local memory cache FNV = Mac, 96/96
CPU emulator, IGNEUM_EXCHANGE 0, work-group 32, sub-group 32 local memory PASS, batch fingerprint f99fb375b3abeaf5
CPU emulator, IGNEUM_EXCHANGE 0, work-group 64, sub-group 64 (wave64, two units per wave) local memory PASS, same fingerprint
CPU emulator, IGNEUM_EXCHANGE 0, work-group 32, sub-group 64 (wave64 half empty) local memory PASS, same fingerprint
CPU emulator, IGNEUM_EXCHANGE 1, work-group 32, sub-group 32 sub_group_shuffle_xor PASS, same fingerprint
CPU emulator, IGNEUM_EXCHANGE 1, work-group 32, sub-group 64 sub_group_shuffle_xor over a half-empty wave64 PASS, same fingerprint
CPU emulator, IGNEUM_EXCHANGE 1, work-group 64, sub-group 64 (wave64, two units in one shuffle domain) sub_group_shuffle_xor PASS, same fingerprint
CPU emulator, IGNEUM_EXCHANGE 1, work-group 64, sub-group 32 (two sub-groups per work-group) sub_group_shuffle_xor PASS, same fingerprint

The fingerprint is FNV-1a 64 over the 2^13 outputs of the batch at base nonce 0; seven emulator configurations produced the same one, and host.c prints the same fingerprint for its warm-up batch (--batch-log2 13): Apple's OpenCL (local memory) and pocl (sub-group shuffles, and local memory) all print f99fb375b3abeaf5. So both exchange implementations, both wave widths, two real OpenCL compilers and the emulator give identical hashes for every nonce in the batch, not only for the three vector warps. An AMD run at --batch-log2 13 must print the same value. The emulator rows with sub-group 64 and IGNEUM_EXCHANGE 1 show that the xor exchange alone would survive a wave64 sub-group (point 1 above); host.c still refuses that configuration on a real device because of point 2 and because of sub_group_broadcast. The conservative rule costs nothing in correctness and, on a wave64 card, the local-memory path is what runs.

RDNA 4, measured (5 October 2026)

An RX 9070 XT (gfx1201, Adrenalin 26.9.2, OpenCL driver 3683.0 PAL,LC) on PC 1 reports AMD wavefront width 32, lists cl_khr_subgroups without a shuffle extension, and clGetKernelSubGroupInfoKHR on igneum_hash answers a sub-group of 32 for a 32-item work-group: wave32, so a work-group of 32 is one full wave and path 0 (local memory) runs with its barriers elided by the compiler. --group-warps 1, 2, 4, 8 give 18.02, 18.06, 18.07 and 18.04 MH/s, the same number: the exchange path and the work-group shape cost nothing there. The rate is set by the card's random-read throughput at the dataset size (docs/bench-log.md, "the 9070 XT on the eGPU").

What is not proven here

  • No AMD compiler has compiled kernel.cl, and no AMD device has run it. The emulator is clang; pocl is LLVM on a CPU; Apple's OpenCL is Apple's compiler. Syntax every one of them accepts can still trip AMD's front end, which would be a build log (printed in full by host.c), not a silent difference.
  • No AMD hash rate exists. The cost of the local-memory exchange relative to sub-group shuffles on AMD is unknown; on Apple the local-memory path runs at the same 45 Mhash/s as Metal's simd_shuffle_xor kernel because the hash is bound by the 104 random dataset loads, and the same is expected elsewhere, but expected is not measured.
  • Whether RDNA compiles igneum_hash as wave32 (sub-group 32, path 1) or wave64 (path 0) is a driver decision. The run will say which on the exchange: line. Both are correct by construction; only the speed may differ.