268 lines
18 KiB
Markdown
268 lines
18 KiB
Markdown
# igneum-bench-cl (proto-opencl)
|
|
|
|
The portable third path for Igneum's random-program proof-of-work kernels, after Apple Metal (`proto-metal`) and
|
|
NVIDIA CUDA (`proto-cuda`). It runs the same program packs through OpenCL 1.2, which is what an AMD card exposes on
|
|
both Windows (Adrenalin driver) and Linux (ROCm, or Mesa), and checks every output bit for bit against the Mac.
|
|
|
|
This is a test harness, not a miner. No pool, no network, no wallet, no mining protocol. It fills a dataset, checks the
|
|
device against known answers, and times the kernel. Nothing here earns anything.
|
|
|
|
Status on 4 October 2026: every pack is now generator version 2 (`igneum-pow export`, spec 01 sections 1.4.2 to
|
|
1.4.6; 128 loads per hash, program id in `program.h`). Apple OpenCL on the M5 Max gives 96/96 on all four packs, batch
|
|
fingerprint `f2a95d5bb84d961e` at 2^13 for `igneum-genesis-mh` (`8e22ad069cb2a8c3` for `igneum-devnet-v4-epoch0`),
|
|
identical to the emulator in the sub-group 32 and wave64 configurations; 27.9 Mhash/s at 1 GiB, down from 45.0 under
|
|
the retired 80-distinct-load genesis program, exactly the 128-distinct-load projection of the census. The AMD gfx1036
|
|
and RTX 5090 OpenCL runs of 3 October were on version 1 packs and are owed again.
|
|
|
|
Status on 3 October 2026: no AMD device has run this yet. Everything that could be proven without one has been
|
|
(`WAVEFRONT.md`, "What was proven"): Apple's deprecated OpenCL 1.2 runtime on the M5 Max, pocl 7.2 on the Mac's CPU,
|
|
and a CPU emulator with 32- and 64-wide sub-groups all give the Mac's 96/96 vectors and cache FNV for the memory-hard
|
|
pack. The AMD run itself is the next step and the commands for it are below.
|
|
|
|
## The one-click worker mode (--pack, 4 October 2026)
|
|
|
|
`host.c --serve --pack <dir>` serves a pack read at run time (`../proto-cuda/nvrtc/packfile.h`: program.h, seeds.txt,
|
|
vectors.h, kernel_bound.cl), whatever pack the exe was built against, so one prebuilt exe serves every hourly program.
|
|
The first pair is built and self-tested exactly like a prepared one (cache head, last line and FNV-1a 64, dataset head,
|
|
last word and 64 samples, the three vector warps through `igneum_hash_bound` with the pack's seed words), and every
|
|
prepared pair is self-tested too; the `ready` line ends in `path prebuilt-generic`. The Windows package ships this as
|
|
`igneum-worker-opencl.exe`, built by `../proto-cuda/nvrtc/build-windows.sh` with `IGNEUM_CL_DYNAMIC` (`cl_dynamic.h`:
|
|
OpenCL.dll opened with LoadLibrary, every entry point a function pointer, no import library; nothing to install but
|
|
the driver) and the Khronos headers fetched by `fetch-redist.sh`. `test-generic.sh` checks the mode here through Apple
|
|
OpenCL with the two packs `proto-cuda/nvrtc/emu/test.sh` writes (PASS on 4 October 2026: 192 found lines, prepare and
|
|
swap, 15 sampled hashes equal to `igneum-pow hash-bound`). The bench and the compiled-in serve mode are unchanged.
|
|
|
|
## 5 October 2026: the RX 9070 XT (gfx1201, RDNA 4) on PC 1
|
|
|
|
Measured in `docs/bench-log.md` ("the 9070 XT on the eGPU"). What changed in `host.c`:
|
|
|
|
- `--list` folds the same card listed by two platforms of one vendor (an old driver's OpenCL registration left behind
|
|
after an update: PC 1 had 3652.0 and 3683.0) into one entry; the older one prints as ` dup [N] ...` and the app's
|
|
parser (`detect.rs`) never makes a card of it. Before, the app ran two workers on one 9070 XT at half rate each.
|
|
Indices stay flat, so `--device N` still reaches the hidden entry for a comparison.
|
|
- Every path prints the kernel as the driver compiled it: `kernel: ... preferred multiple, local and private memory,
|
|
sub-group size`. Private memory above 0 means spilled registers. The sub-group size is queried on the local-memory
|
|
path too (the gfx1201 answers 32: wave32).
|
|
- `--serve` reads back the hits and 34 sentinel words through a GPU-side select pass instead of 8 bytes per nonce
|
|
(`--readback full` or `IGNEUM_READBACK=full` keeps the old path). The stats line every 200 jobs carries the bytes up
|
|
and down per chunk and the mean device time of the hash kernel, the select pass, the read-back and the host scan.
|
|
- `--memprobe [--probe-mib N]`: no pack. Dependent random 4-byte loads against lanes in flight (latency and the
|
|
random-read ceiling), eight independent loads per lane, random 64-byte lines, a coalesced stream and an integer
|
|
chain, at 4, 64 and 1024 MiB. The hash is 128 dependent random 4-byte loads, so the 1024 MiB chase ceiling divided
|
|
by 128 is the card's hash-rate ceiling for this program class.
|
|
|
|
## Layout
|
|
|
|
```
|
|
proto-opencl/
|
|
host.c C99 host: device list, runtime kernel build, cache + dataset fill, self-tests, vectors, bench, sweep, --serve, --pack
|
|
cl_dynamic.h Windows one-click build: OpenCL.dll loaded at run time (IGNEUM_CL_DYNAMIC)
|
|
test-generic.sh the --pack mode checked here through Apple OpenCL (needs proto-cuda/nvrtc/emu/test.sh's packs)
|
|
test_host.c device-free unit tests of host.c's rules (the duplicate-platform fold); run with test-host.sh
|
|
gpu-telemetry.c igneum-gpu-telemetry: AMD power, temperature, fan, clocks and busy per card (ADLX on Windows, amdgpu sysfs on Linux),
|
|
one line per card per sample; the app's AMD card row reads it (engine.rs amd_telemetry_line)
|
|
build.sh macOS (-framework OpenCL, or the Khronos ICD loader) and Linux (-lOpenCL)
|
|
build.bat Windows (MSVC cl.exe + OpenCL.lib)
|
|
WAVEFRONT.md wave32 vs wave64 on AMD, and why the kernel cannot tell the difference
|
|
emu/ CPU emulator: compiles kernel.cl as C++ and runs it with a 32- or 64-wide sub-group
|
|
../proto-cuda/packs/<seed>/kernel.cl the OpenCL C kernel of each pack, written by igneum-pow export (generator v2)
|
|
```
|
|
|
|
The packs stay in `proto-cuda/packs/` because the three implementations share one `program.h`, `vectors.h` and
|
|
`memhard.h` per seed; `kernel.cl` sits next to `kernel.cu` and `program.metal`. Four packs are checked in, all
|
|
generator version 2: `igneum-genesis-mh` and `igneum-devnet-v4-epoch0` (memory-hard dataset, the current
|
|
construction; the second is the devnet's own epoch 0 derivation), `igneum-genesis` and `igneum-hourly` (closed-form
|
|
dataset, kept for comparison).
|
|
|
|
## What the OpenCL path does differently
|
|
|
|
| Item | CUDA (`host.cu`) | OpenCL (`host.c`) |
|
|
|---|---|---|
|
|
| Kernel compile | ahead of time by nvcc | at runtime by the driver from `kernel.cl`, with `-D IGNEUM_GROUP=32W -D IGNEUM_EXCHANGE=M` |
|
|
| 32-lane exchange | `__shfl_xor_sync` | `sub_group_shuffle_xor` only when the device's sub-group size is exactly 32 for a 32-item work-group; otherwise `__local` memory and a barrier (`WAVEFRONT.md`) |
|
|
| Timing | cudaEvent | event profiling (`CL_PROFILING_COMMAND_START/END`); wall time on Apple, whose OpenCL timestamps are unusable |
|
|
| Device choice | `--device D` | `--list`, then `--device D` by flat index over all platforms; default is the first GPU |
|
|
| Language | C++17 | C99, so MSVC builds it without nvcc or any C++ toolchain |
|
|
|
|
Everything else (dataset fill or build, cache check against the Mac's FNV, dataset self-test, 3 vector warps standalone
|
|
and inside the warm-up batch, 5 timed batches of 2^24, the sweep, the summary table, exit codes) mirrors `host.cu`
|
|
line for line. Rates have the same definition: Mhash/s, and GB/s useful = loads per hash x 4 bytes x hashes/s.
|
|
|
|
## Build
|
|
|
|
Linux (AMD ROCm, AMD Adrenalin for Linux, or Mesa rusticl; any of them registers with the system ICD loader):
|
|
|
|
```
|
|
sudo apt install ocl-icd-opencl-dev opencl-headers clinfo # Debian or Ubuntu; Fedora: ocl-icd-devel opencl-headers clinfo
|
|
clinfo | head -40 # the AMD device must appear here before anything else matters
|
|
cd proto-opencl
|
|
./build.sh igneum-genesis-mh # -> ./igneum-bench-cl-igneum-genesis-mh
|
|
./build.sh igneum-genesis # closed-form packs
|
|
./build.sh igneum-hourly
|
|
```
|
|
|
|
The one command behind the script, for a pack `P`:
|
|
|
|
```
|
|
cc -std=c99 -O2 -Wall -Wextra -I ../proto-cuda/packs/P -DIGNEUM_KERNEL_PATH='"../proto-cuda/packs/P/kernel.cl"' -o igneum-bench-cl-P host.c -lOpenCL -ldl
|
|
```
|
|
|
|
Windows (AMD Adrenalin driver installed; it provides `OpenCL.dll` and the AMD ICD, nothing else is needed to run):
|
|
|
|
1. Install Visual Studio 2022 or 2026 Build Tools with "Desktop development with C++". No CUDA, no nvcc.
|
|
2. Get the OpenCL headers and the import library `OpenCL.lib`, any one of these:
|
|
`vcpkg install opencl:x64-windows` then `set OPENCL_SDK=C:\vcpkg\installed\x64-windows`;
|
|
or the Khronos OpenCL-SDK release zip (GitHub KhronosGroup/OpenCL-SDK) unpacked anywhere, `set OPENCL_SDK=<that dir>`;
|
|
or an installed CUDA Toolkit, which ships both under `%CUDA_PATH%` (the script finds it on its own).
|
|
3. From an "x64 Native Tools Command Prompt":
|
|
|
|
```
|
|
cd proto-opencl
|
|
build.bat igneum-genesis-mh
|
|
build.bat igneum-genesis
|
|
build.bat igneum-hourly
|
|
```
|
|
|
|
The one command behind the script:
|
|
|
|
```
|
|
cl /nologo /O2 /W3 /std:c11 /I "%OPENCL_SDK%\include" /I "..\proto-cuda\packs\P" /DIGNEUM_KERNEL_PATH="\"../proto-cuda/packs/P/kernel.cl\"" /Fe:igneum-bench-cl-P.exe host.c /link "%OPENCL_SDK%\lib\OpenCL.lib"
|
|
```
|
|
|
|
macOS (correctness check only; Apple deprecated OpenCL in macOS 10.14 and the runtime is 1.2):
|
|
|
|
```
|
|
./build.sh igneum-genesis-mh # Apple OpenCL.framework
|
|
./build.sh igneum-genesis-mh khr # Khronos ICD loader from Homebrew, for pocl: brew install opencl-headers opencl-icd-loader pocl
|
|
OCL_ICD_VENDORS=/opt/homebrew/etc/OpenCL/vendors SDKROOT=$(xcrun --show-sdk-path) ./igneum-bench-cl-igneum-genesis-mh-khr --list
|
|
```
|
|
|
|
(`SDKROOT` is needed because pocl links its kernels with the system linker and otherwise cannot find `-lSystem`.)
|
|
|
|
## Run on the AMD rig
|
|
|
|
Run from `proto-opencl/` so the default kernel path resolves, or pass `--kernel`. Exactly this, in this order, and paste
|
|
the whole stdout back:
|
|
|
|
Windows:
|
|
|
|
```
|
|
cd proto-opencl
|
|
igneum-bench-cl-igneum-genesis-mh.exe --list
|
|
igneum-bench-cl-igneum-genesis-mh.exe > amd-mh-auto.txt
|
|
igneum-bench-cl-igneum-genesis-mh.exe --exchange local > amd-mh-local.txt
|
|
igneum-bench-cl-igneum-genesis-mh.exe --group-warps 2 > amd-mh-gw2.txt
|
|
igneum-bench-cl-igneum-genesis-mh.exe --sweep > amd-mh-sweep.txt
|
|
igneum-bench-cl-igneum-genesis.exe > amd-genesis.txt
|
|
igneum-bench-cl-igneum-hourly.exe > amd-hourly.txt
|
|
```
|
|
|
|
Linux:
|
|
|
|
```
|
|
cd proto-opencl
|
|
./igneum-bench-cl-igneum-genesis-mh --list
|
|
./igneum-bench-cl-igneum-genesis-mh | tee amd-mh-auto.txt
|
|
./igneum-bench-cl-igneum-genesis-mh --exchange local | tee amd-mh-local.txt
|
|
./igneum-bench-cl-igneum-genesis-mh --group-warps 2 | tee amd-mh-gw2.txt
|
|
./igneum-bench-cl-igneum-genesis-mh --sweep | tee amd-mh-sweep.txt
|
|
./igneum-bench-cl-igneum-genesis | tee amd-genesis.txt
|
|
./igneum-bench-cl-igneum-hourly | tee amd-hourly.txt
|
|
```
|
|
|
|
If `--list` shows more than one device (an iGPU, a CPU runtime, an NVIDIA card in the same box), add `--device D`
|
|
with the AMD card's index to every line. If the auto run says `exchange: sub_group_shuffle_xor`, also run
|
|
`--exchange subgroup` once so the log has an explicit sub-group run; if it says local memory, the `--exchange local`
|
|
run is a repeat and that is fine.
|
|
|
|
Flags: `--list`, `--device D`, `--dataset-mib N` (power of two, default 1024), `--sweep` (4, 64, 256, 512, 1024 MiB),
|
|
`--batch-log2 B` (default 24), `--batches N` (default 5), `--group-warps W` (32-lane units per work-group, 1..8,
|
|
default 1), `--exchange auto|local|subgroup`, `--kernel path`, `--build-opts "..."` (appended to clBuildProgram),
|
|
`--time event|wall`.
|
|
|
|
## What PASS looks like
|
|
|
|
In order, the run prints:
|
|
|
|
1. Every platform and device, with vendor, driver, OpenCL C version, compute units, memory, local memory, the
|
|
sub-group extension it lists, and (AMD) the wavefront width or (NVIDIA) the warp size. The chosen device is starred.
|
|
2. `build options:` and `exchange:`. The exchange line is the one that matters on AMD; it says which path was taken and
|
|
why, for example `local-memory exchange with barrier (the sub-group size for a 32-item work-group is not 32; queried
|
|
sub-group size 64)` on a wave64 card, or `sub_group_shuffle_xor (cl_khr_subgroup_shuffle), sub-group size 32 for a
|
|
32-item work-group` on a wave32 one. Both are correct by construction (`WAVEFRONT.md`).
|
|
3. `kernel:` (work-group limit and local memory of `igneum_hash`), `program:`, `seed words:`, `day`.
|
|
4. Memory-hard packs: cache fill time on the device (twice) and on one host thread, then `cache check: PASS (device ==
|
|
host all 67108864 words PASS, host FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head 16 vs Mac PASS,
|
|
last line vs Mac PASS)`. The FNV is the same on every machine that has ever run this construction for day
|
|
2026-10-03.
|
|
5. `dataset build` (or `dataset fill`) times, then `dataset self-test: PASS (head 16 vs Mac PASS, element [MASK] vs Mac
|
|
PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS)`.
|
|
6. Six `verify warp ... PASS` lines: three standalone (bases 0, 4096, 1000000) and the same three read out of the
|
|
warm-up batch.
|
|
7. `batch fingerprint`: FNV-1a 64 of every output of the warm-up batch. At `--batch-log2 13` every implementation so far
|
|
prints `f2a95d5bb84d961e` for the version 2 igneum-genesis-mh (Apple OpenCL, the emulator in the sub-group 32 and
|
|
wave64 configurations; the version 1 value was `f99fb375b3abeaf5`); an AMD run at `--batch-log2 13` must print the
|
|
same value. At the default 2^24 Apple OpenCL prints `25f96e7dce90bd4e` (`3cc4fbf90fa6366c` for the devnet pack).
|
|
8. `timed:` with device and wall time, then the `rate` line: Mhash/s and GB/s useful.
|
|
9. The summary table and `OVERALL: PASS`. Exit code 0 on PASS, 1 on FAIL, 2 on an OpenCL error or a build failure (the
|
|
build log is printed in full).
|
|
|
|
A FAIL with a few lanes differing points at the exchange; a FAIL in every lane at an arithmetic op or the dataset; a
|
|
cache FAIL at the ChaCha fill. Send the whole printout either way, including the `--list` output and, on Linux,
|
|
`clinfo` and the ROCm or Mesa version, or on Windows the Adrenalin driver version from the AMD Software panel.
|
|
|
|
## What to paste back
|
|
|
|
All seven text files above, unedited. From them the bench log gets: device name and driver string, the `exchange:`
|
|
line, the cache check line, the 96/96 vector result per pack, the rate at 1 GiB, and the sweep table. If a run looks
|
|
odd, run it again and keep both.
|
|
|
|
## Results so far (no AMD)
|
|
|
|
Generator version 2 packs (4 October 2026):
|
|
|
|
| Machine, runtime | Exchange | igneum-genesis-mh | igneum-devnet-v4-epoch0 | closed-form packs | Rate at 1 GiB |
|
|
|---|---|---|---|---|---|
|
|
| Apple M5 Max, Apple OpenCL 1.2 | local memory | cache FNV = Rust, 96/96, fingerprint `f2a95d5bb84d961e` at 2^13 (also at `--group-warps 2`) | cache FNV `448274a57f508cbc`, 96/96, `8e22ad069cb2a8c3` | 96/96 and 96/96 | 27.9 Mhash/s, 14.3 GB/s useful, 128 loads per hash (every pack within 1 percent of this rate) |
|
|
| CPU emulator (`emu/`), sub-group 32 and wave64 | both | 96/96, `f2a95d5bb84d961e` | 96/96, `8e22ad069cb2a8c3` | not run | none |
|
|
|
|
Version 1 packs (3 October 2026, retired vectors; the kernel text is unchanged, so the exchange and arithmetic
|
|
findings stand):
|
|
|
|
| Machine, runtime | Exchange | igneum-genesis-mh | closed-form packs | Rate at 1 GiB |
|
|
|---|---|---|---|---|
|
|
| Apple M5 Max, Apple OpenCL 1.2 (deprecated runtime) | local memory (runtime lists no sub-group extension) | cache FNV = Mac, 96/96; also 96/96 at `--group-warps 2` and `4` | 96/96 and 96/96 | 45.03 Mhash/s, 18.73 GB/s useful (Metal on the same chip: 45.2). Apple OpenCL on the M5 Max, not an AMD number |
|
|
| Apple M5 Max, pocl 7.2 CPU device, OpenCL 3.0, LLVM 23 | `sub_group_shuffle_xor` (auto and `--exchange subgroup`; pocl's `clGetKernelSubGroupInfoKHR` fails with CL_INVALID_OPERATION, the probe kernel reports 32), and `--exchange local` | cache FNV = Mac, 96/96 on both paths, same batch fingerprint as Apple and the emulator at 2^13 | not run | CPU, not meaningful |
|
|
| CPU emulator (`emu/`), 7 configurations incl. sub-group 64 | both | 96/96 in every configuration, identical batch fingerprint `f99fb375b3abeaf5` | not run | none |
|
|
|
|
Sweep on Apple OpenCL (3 batches, wall time): 4 MiB 573.7, 64 MiB 178.9, 256 MiB 94.3, 512 MiB 68.8, 1024 MiB 45.0 Mhash/s.
|
|
Full tables in `docs/bench-log.md`.
|
|
|
|
## Checking a pack without a GPU
|
|
|
|
```
|
|
emu/emu.sh igneum-genesis-mh 0 32 --sg 32 # local-memory exchange, work-group 32, 32-wide sub-group
|
|
emu/emu.sh igneum-genesis-mh 0 64 --sg 64 # local-memory exchange on a wave64 model, two units per work-group
|
|
emu/emu.sh igneum-genesis-mh 1 32 --sg 32 # sub_group_shuffle_xor, 32-wide sub-group
|
|
emu/emu.sh igneum-genesis-mh 1 64 --sg 64 # sub_group_shuffle_xor over a 64-wide sub-group (two units in one wave)
|
|
```
|
|
|
|
The emulator compiles `kernel.cl` unchanged as C++ (`emu/emu_opencl.h` supplies the OpenCL built-ins; the kernel's own
|
|
prelude includes it when no OpenCL compiler is present) and runs work-groups on host threads with a barrier inside
|
|
every exchange. It prints the cache check, dataset self-test, vectors and a fingerprint of the whole 2^13-hash batch,
|
|
so configurations can be compared for every nonce, not only the vector warps. It says nothing about any GPU.
|
|
|
|
## Regenerating a pack
|
|
|
|
```
|
|
cd igneum-pow && cargo build --release
|
|
./target/release/igneum-pow export --seed igneum-genesis --out ../proto-cuda/packs/igneum-genesis-mh
|
|
./target/release/igneum-pow export --closed-form --seed igneum-genesis --out ../proto-cuda/packs/igneum-genesis
|
|
./target/release/igneum-pow export --closed-form --seed igneum-hourly --out ../proto-cuda/packs/igneum-hourly
|
|
```
|
|
|
|
The Rust exporter (since 4 October 2026 the pack source) writes `kernel.cl` and `kernel_bound.cl` next to `kernel.cu`
|
|
from the same instruction list. The memory-hard core is emitted once (`emit_memhard_core`) in three dialects (Metal,
|
|
CUDA C++, OpenCL C), so the constants in `kernel.cl` are the literals of `memhard.h`. The Metal cross-check of a pack is
|
|
`proto-metal/igneum-bench --seed <seed> --export-pack <scratch dir>` and a diff of its `vectors.json` (see
|
|
`../proto-cuda/README.md`, "Regenerating a pack").
|