igneum/proto-opencl
igneum-labs 7eed16a29a Pre-public scrub, the text pass (7 October 2026, 19:5x UK): no founder name, personal login, earlier business or personal address in any tracked text file, and a gate check that keeps it so
The sweep (main's item 1): 199 tracked text files, 783 lines. The founder's full name, first name and possessive become "the founder" (sentence starts capitalised); the lowercase operating-system user name in WSL paths and commands becomes <user>; the second owner login becomes "the second owner login"; the three earlier businesses and the two other brands become "the other business", "the earlier entity", "the earlier business" and "another brand"; the Chrome profile rule names the igneum.network profile, not the profile's label. The standing commit login igneum-labs is not a founder term here: the fresh-repository step renames it in the history (docs/plans/history-rewrite.md, tools/repo/fresh-repo.sh).

The patterns never appear in plain text in the tree (a plaintext list would be the hit): tools/ci/founder-strings.b64 (perl regex, tab, a sample per row) is read by tools/ci/founder-strings-check.sh (every tracked text file, perl, known-failed first: the self-test plants each row's sample in a fixture and the hit must name the file), by tools/community/discord-hooks.mjs (the guard's founder and business rows; the test takes its fixtures from the samples) and by tools/repo/fresh-repo.sh (the business names of the rewrite rules). site/forbidden-strings.txt carries the same patterns as b64: lines, decoded case-insensitive by site/scrub.mjs and tools/ci/launch-gates-check.mjs (whose fixture now plants an encoded made-up name). The check runs in the gate's tree checks on every merge.

Not in this commit, by main's word: the 105 commit messages and 40 personal-identity commits that need the history rewrite (listed, not run), and the secrets found by gitleaks over the history (reported with owners).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-07 18:39:50 +00:00
..
emu read-width: scratch per warp is a class parameter (32 or 128 KiB, under the 6 GB working-set cap), distinct-address rule bounds dataset loads only; Metal pack harness; OpenCL --bench-pack, scratch args and 16-byte probe; packfile class fields; OpenCL emulator persistent launch 2026-10-05 19:58:20 +00:00
.gitignore proto-opencl: OpenCL path for AMD, proven on Apple OpenCL, pocl and a wave64 CPU emulator 2026-10-03 16:52:24 +00:00
build.bat OpenCL build.bat: library detection without parentheses in blocks (Program Files (x86) broke the parser) 2026-10-03 17:04:39 +00:00
build.sh proto-opencl: OpenCL path for AMD, proven on Apple OpenCL, pocl and a wave64 CPU emulator 2026-10-03 16:52:24 +00:00
cl_dynamic.h Miner dev fee: app switch and UI line, HiveOS package, Linux worker cross-build, docs and public text 2026-10-04 21:46:41 +00:00
dot4-probe.c Counter ASIC 2.0 layer 6: SRAM mirror analysis (cited bit cells N7 to N2 and 18A, area and cost per node, no cache growth rule in the spec, options A to E for the project lead); layer 7 dp4a probes for Metal, CUDA and OpenCL (standalone, no lottery kernel) 2026-10-05 20:09:10 +00:00
family-probe.c Counter ASIC 3.0 PC 1 AMD: the family probe picks the card by name on its newest platform (never ordinal 0; device name, CUs and platform on every row; refuses without a match), the family job run b with its own kit, the ordinal-0 fault and the AMD OpenCL C builtin finding recorded 2026-10-06 17:04:07 +00:00
gpu-telemetry.c Pre-public scrub, the text pass (7 October 2026, 19:5x UK): no founder name, personal login, earlier business or personal address in any tracked text file, and a gate check that keeps it so 2026-10-07 18:39:50 +00:00
host.c Counter ASIC 3.0 PC 1 AMD: OpenCL family probe (every 1.13.2 family through AMD's builtins, khr/intel sub-groups, media ops, C forms and local emulation; bit-exact on Apple OpenCL) and the clBuildProgram time line in the OpenCL worker 2026-10-06 15:39:51 +00:00
README.md Merge release-0.3.12 (fda4684) into ember-tune: 0.3.11's six-section View and card order kept, Ember Tune's line and switches re-added on it; the tune fields move into hotplug::apply_pref; the power-cap plan keeps present(); both CI test lists; 132 app tests, 26 UI tests, every gate green 2026-10-06 08:20:46 +00:00
test-generic.sh OpenCL worker: generic --pack serve mode and OpenCL.dll loaded at run time (one-click AMD worker) 2026-10-04 09:55:05 +00:00
test-host.sh read-width experiment (gate 1): load classes W=4/16/64, per-load width mix, scratch RMW variant behind a generator flag; 20 packs; CPU verifier and acceptance mirror; emulator shims 2026-10-05 19:46:54 +00:00
test_host.c read-width experiment (gate 1): load classes W=4/16/64, per-load width mix, scratch RMW variant behind a generator flag; 20 packs; CPU verifier and acceptance mirror; emulator shims 2026-10-05 19:46:54 +00:00
WAVEFRONT.md read-width experiment (gate 1): load classes W=4/16/64, per-load width mix, scratch RMW variant behind a generator flag; 20 packs; CPU verifier and acceptance mirror; emulator shims 2026-10-05 19:46:54 +00:00

igneum-bench-cl (proto-opencl)

The portable third path for Igneum's random-program proof-of-work kernels, after Apple Metal (proto-metal) and NVIDIA CUDA (proto-cuda). It runs the same program packs through OpenCL 1.2, which is what an AMD card exposes on both Windows (Adrenalin driver) and Linux (ROCm, or Mesa), and checks every output bit for bit against the Mac.

This is a test harness, not a miner. No pool, no network, no wallet, no mining protocol. It fills a dataset, checks the device against known answers, and times the kernel. Nothing here earns anything.

Status on 4 October 2026: every pack is now generator version 2 (igneum-pow export, spec 01 sections 1.4.2 to 1.4.6; 128 loads per hash, program id in program.h). Apple OpenCL on the M5 Max gives 96/96 on all four packs, batch fingerprint f2a95d5bb84d961e at 2^13 for igneum-genesis-mh (8e22ad069cb2a8c3 for igneum-devnet-v4-epoch0), identical to the emulator in the sub-group 32 and wave64 configurations; 27.9 Mhash/s at 1 GiB, down from 45.0 under the retired 80-distinct-load genesis program, exactly the 128-distinct-load projection of the census. The AMD gfx1036 and RTX 5090 OpenCL runs of 3 October were on version 1 packs and are owed again.

Status on 3 October 2026: no AMD device has run this yet. Everything that could be proven without one has been (WAVEFRONT.md, "What was proven"): Apple's deprecated OpenCL 1.2 runtime on the M5 Max, pocl 7.2 on the Mac's CPU, and a CPU emulator with 32- and 64-wide sub-groups all give the Mac's 96/96 vectors and cache FNV for the memory-hard pack. The AMD run itself is the next step and the commands for it are below.

The one-click worker mode (--pack, 4 October 2026)

host.c --serve --pack <dir> serves a pack read at run time (../proto-cuda/nvrtc/packfile.h: program.h, seeds.txt, vectors.h, kernel_bound.cl), whatever pack the exe was built against, so one prebuilt exe serves every hourly program. The first pair is built and self-tested exactly like a prepared one (cache head, last line and FNV-1a 64, dataset head, last word and 64 samples, the three vector warps through igneum_hash_bound with the pack's seed words), and every prepared pair is self-tested too; the ready line ends in path prebuilt-generic. The Windows package ships this as igneum-worker-opencl.exe, built by ../proto-cuda/nvrtc/build-windows.sh with IGNEUM_CL_DYNAMIC (cl_dynamic.h: OpenCL.dll opened with LoadLibrary, every entry point a function pointer, no import library; nothing to install but the driver) and the Khronos headers fetched by fetch-redist.sh. test-generic.sh checks the mode here through Apple OpenCL with the two packs proto-cuda/nvrtc/emu/test.sh writes (PASS on 4 October 2026: 192 found lines, prepare and swap, 15 sampled hashes equal to igneum-pow hash-bound). The bench and the compiled-in serve mode are unchanged.

5 October 2026: the RX 9070 XT (gfx1201, RDNA 4) on PC 1

Measured in docs/bench-log.md ("the 9070 XT on the eGPU"). What changed in host.c:

  • --list folds the same card listed by two platforms of one vendor (an old driver's OpenCL registration left behind after an update: PC 1 had 3652.0 and 3683.0) into one entry; the older one prints as dup [N] ... and the app's parser (detect.rs) never makes a card of it. Before, the app ran two workers on one 9070 XT at half rate each. Indices stay flat, so --device N still reaches the hidden entry for a comparison.
  • Every path prints the kernel as the driver compiled it: kernel: ... preferred multiple, local and private memory, sub-group size. Private memory above 0 means spilled registers. The sub-group size is queried on the local-memory path too (the gfx1201 answers 32: wave32).
  • --serve reads back the hits and 34 sentinel words through a GPU-side select pass instead of 8 bytes per nonce (--readback full or IGNEUM_READBACK=full keeps the old path). The stats line every 200 jobs carries the bytes up and down per chunk and the mean device time of the hash kernel, the select pass, the read-back and the host scan.
  • --memprobe [--probe-mib N]: no pack. Dependent random 4-byte loads against lanes in flight (latency and the random-read ceiling), eight independent loads per lane, random 64-byte lines, a coalesced stream and an integer chain, at 4, 64 and 1024 MiB. The hash is 128 dependent random 4-byte loads, so the 1024 MiB chase ceiling divided by 128 is the card's hash-rate ceiling for this program class.

Layout

proto-opencl/
  host.c          C99 host: device list, runtime kernel build, cache + dataset fill, self-tests, vectors, bench, sweep, --serve, --pack
  cl_dynamic.h    Windows one-click build: OpenCL.dll loaded at run time (IGNEUM_CL_DYNAMIC)
  test-generic.sh the --pack mode checked here through Apple OpenCL (needs proto-cuda/nvrtc/emu/test.sh's packs)
  test_host.c     device-free unit tests of host.c's rules (the duplicate-platform fold); run with test-host.sh
  gpu-telemetry.c igneum-gpu-telemetry: AMD power, temperature, fan, clocks and busy per card (ADLX on Windows, amdgpu sysfs on Linux),
                  one line per card per sample; the app's AMD card row reads it (engine.rs amd_telemetry_line)
  build.sh        macOS (-framework OpenCL, or the Khronos ICD loader) and Linux (-lOpenCL)
  build.bat       Windows (MSVC cl.exe + OpenCL.lib)
  WAVEFRONT.md    wave32 vs wave64 on AMD, and why the kernel cannot tell the difference
  emu/            CPU emulator: compiles kernel.cl as C++ and runs it with a 32- or 64-wide sub-group
../proto-cuda/packs/<seed>/kernel.cl   the OpenCL C kernel of each pack, written by igneum-pow export (generator v2)

The packs stay in proto-cuda/packs/ because the three implementations share one program.h, vectors.h and memhard.h per seed; kernel.cl sits next to kernel.cu and program.metal. Four packs are checked in, all generator version 2: igneum-genesis-mh and igneum-devnet-v4-epoch0 (memory-hard dataset, the current construction; the second is the devnet's own epoch 0 derivation), igneum-genesis and igneum-hourly (closed-form dataset, kept for comparison).

What the OpenCL path does differently

Item CUDA (host.cu) OpenCL (host.c)
Kernel compile ahead of time by nvcc at runtime by the driver from kernel.cl, with -D IGNEUM_GROUP=32W -D IGNEUM_EXCHANGE=M
32-lane exchange __shfl_xor_sync sub_group_shuffle_xor only when the device's sub-group size is exactly 32 for a 32-item work-group; otherwise __local memory and a barrier (WAVEFRONT.md)
Timing cudaEvent event profiling (CL_PROFILING_COMMAND_START/END); wall time on Apple, whose OpenCL timestamps are unusable
Device choice --device D --list, then --device D by flat index over all platforms; default is the first GPU
Language C++17 C99, so MSVC builds it without nvcc or any C++ toolchain

Everything else (dataset fill or build, cache check against the Mac's FNV, dataset self-test, 3 vector warps standalone and inside the warm-up batch, 5 timed batches of 2^24, the sweep, the summary table, exit codes) mirrors host.cu line for line. Rates have the same definition: Mhash/s, and GB/s useful = loads per hash x 4 bytes x hashes/s.

Build

Linux (AMD ROCm, AMD Adrenalin for Linux, or Mesa rusticl; any of them registers with the system ICD loader):

sudo apt install ocl-icd-opencl-dev opencl-headers clinfo     # Debian or Ubuntu; Fedora: ocl-icd-devel opencl-headers clinfo
clinfo | head -40                                            # the AMD device must appear here before anything else matters
cd proto-opencl
./build.sh igneum-genesis-mh                                  # -> ./igneum-bench-cl-igneum-genesis-mh
./build.sh igneum-genesis                                     # closed-form packs
./build.sh igneum-hourly

The one command behind the script, for a pack P:

cc -std=c99 -O2 -Wall -Wextra -I ../proto-cuda/packs/P -DIGNEUM_KERNEL_PATH='"../proto-cuda/packs/P/kernel.cl"' -o igneum-bench-cl-P host.c -lOpenCL -ldl

Windows (AMD Adrenalin driver installed; it provides OpenCL.dll and the AMD ICD, nothing else is needed to run):

  1. Install Visual Studio 2022 or 2026 Build Tools with "Desktop development with C++". No CUDA, no nvcc.
  2. Get the OpenCL headers and the import library OpenCL.lib, any one of these: vcpkg install opencl:x64-windows then set OPENCL_SDK=C:\vcpkg\installed\x64-windows; or the Khronos OpenCL-SDK release zip (GitHub KhronosGroup/OpenCL-SDK) unpacked anywhere, set OPENCL_SDK=<that dir>; or an installed CUDA Toolkit, which ships both under %CUDA_PATH% (the script finds it on its own).
  3. From an "x64 Native Tools Command Prompt":
cd proto-opencl
build.bat igneum-genesis-mh
build.bat igneum-genesis
build.bat igneum-hourly

The one command behind the script:

cl /nologo /O2 /W3 /std:c11 /I "%OPENCL_SDK%\include" /I "..\proto-cuda\packs\P" /DIGNEUM_KERNEL_PATH="\"../proto-cuda/packs/P/kernel.cl\"" /Fe:igneum-bench-cl-P.exe host.c /link "%OPENCL_SDK%\lib\OpenCL.lib"

macOS (correctness check only; Apple deprecated OpenCL in macOS 10.14 and the runtime is 1.2):

./build.sh igneum-genesis-mh            # Apple OpenCL.framework
./build.sh igneum-genesis-mh khr        # Khronos ICD loader from Homebrew, for pocl: brew install opencl-headers opencl-icd-loader pocl
OCL_ICD_VENDORS=/opt/homebrew/etc/OpenCL/vendors SDKROOT=$(xcrun --show-sdk-path) ./igneum-bench-cl-igneum-genesis-mh-khr --list

(SDKROOT is needed because pocl links its kernels with the system linker and otherwise cannot find -lSystem.)

Run on the AMD rig

Run from proto-opencl/ so the default kernel path resolves, or pass --kernel. Exactly this, in this order, and paste the whole stdout back:

Windows:

cd proto-opencl
igneum-bench-cl-igneum-genesis-mh.exe --list
igneum-bench-cl-igneum-genesis-mh.exe                          > amd-mh-auto.txt
igneum-bench-cl-igneum-genesis-mh.exe --exchange local         > amd-mh-local.txt
igneum-bench-cl-igneum-genesis-mh.exe --group-warps 2          > amd-mh-gw2.txt
igneum-bench-cl-igneum-genesis-mh.exe --sweep                  > amd-mh-sweep.txt
igneum-bench-cl-igneum-genesis.exe                             > amd-genesis.txt
igneum-bench-cl-igneum-hourly.exe                              > amd-hourly.txt

Linux:

cd proto-opencl
./igneum-bench-cl-igneum-genesis-mh --list
./igneum-bench-cl-igneum-genesis-mh                            | tee amd-mh-auto.txt
./igneum-bench-cl-igneum-genesis-mh --exchange local           | tee amd-mh-local.txt
./igneum-bench-cl-igneum-genesis-mh --group-warps 2            | tee amd-mh-gw2.txt
./igneum-bench-cl-igneum-genesis-mh --sweep                    | tee amd-mh-sweep.txt
./igneum-bench-cl-igneum-genesis                               | tee amd-genesis.txt
./igneum-bench-cl-igneum-hourly                                | tee amd-hourly.txt

If --list shows more than one device (an iGPU, a CPU runtime, an NVIDIA card in the same box), add --device D with the AMD card's index to every line. If the auto run says exchange: sub_group_shuffle_xor, also run --exchange subgroup once so the log has an explicit sub-group run; if it says local memory, the --exchange local run is a repeat and that is fine.

Flags: --list, --device D, --dataset-mib N (power of two, default 1024), --sweep (4, 64, 256, 512, 1024 MiB), --batch-log2 B (default 24), --batches N (default 5), --group-warps W (32-lane units per work-group, 1..8, default 1), --exchange auto|local|subgroup, --kernel path, --build-opts "..." (appended to clBuildProgram), --time event|wall.

What PASS looks like

In order, the run prints:

  1. Every platform and device, with vendor, driver, OpenCL C version, compute units, memory, local memory, the sub-group extension it lists, and (AMD) the wavefront width or (NVIDIA) the warp size. The chosen device is starred.
  2. build options: and exchange:. The exchange line is the one that matters on AMD; it says which path was taken and why, for example local-memory exchange with barrier (the sub-group size for a 32-item work-group is not 32; queried sub-group size 64) on a wave64 card, or sub_group_shuffle_xor (cl_khr_subgroup_shuffle), sub-group size 32 for a 32-item work-group on a wave32 one. Both are correct by construction (WAVEFRONT.md).
  3. kernel: (work-group limit and local memory of igneum_hash), program:, seed words:, day.
  4. Memory-hard packs: cache fill time on the device (twice) and on one host thread, then cache check: PASS (device == host all 67108864 words PASS, host FNV-1a 64 48c4f5bf24166b2e vs Mac 48c4f5bf24166b2e PASS, head 16 vs Mac PASS, last line vs Mac PASS). The FNV is the same on every machine that has ever run this construction for day 2026-10-03.
  5. dataset build (or dataset fill) times, then dataset self-test: PASS (head 16 vs Mac PASS, element [MASK] vs Mac PASS, 64 random points vs host derivation PASS, 64 Mac samples PASS).
  6. Six verify warp ... PASS lines: three standalone (bases 0, 4096, 1000000) and the same three read out of the warm-up batch.
  7. batch fingerprint: FNV-1a 64 of every output of the warm-up batch. At --batch-log2 13 every implementation so far prints f2a95d5bb84d961e for the version 2 igneum-genesis-mh (Apple OpenCL, the emulator in the sub-group 32 and wave64 configurations; the version 1 value was f99fb375b3abeaf5); an AMD run at --batch-log2 13 must print the same value. At the default 2^24 Apple OpenCL prints 25f96e7dce90bd4e (3cc4fbf90fa6366c for the devnet pack).
  8. timed: with device and wall time, then the rate line: Mhash/s and GB/s useful.
  9. The summary table and OVERALL: PASS. Exit code 0 on PASS, 1 on FAIL, 2 on an OpenCL error or a build failure (the build log is printed in full).

A FAIL with a few lanes differing points at the exchange; a FAIL in every lane at an arithmetic op or the dataset; a cache FAIL at the ChaCha fill. Send the whole printout either way, including the --list output and, on Linux, clinfo and the ROCm or Mesa version, or on Windows the Adrenalin driver version from the AMD Software panel.

What to paste back

All seven text files above, unedited. From them the bench log gets: device name and driver string, the exchange: line, the cache check line, the 96/96 vector result per pack, the rate at 1 GiB, and the sweep table. If a run looks odd, run it again and keep both.

Results so far (no AMD)

Generator version 2 packs (4 October 2026):

Machine, runtime Exchange igneum-genesis-mh igneum-devnet-v4-epoch0 closed-form packs Rate at 1 GiB
Apple M5 Max, Apple OpenCL 1.2 local memory cache FNV = Rust, 96/96, fingerprint f2a95d5bb84d961e at 2^13 (also at --group-warps 2) cache FNV 448274a57f508cbc, 96/96, 8e22ad069cb2a8c3 96/96 and 96/96 27.9 Mhash/s, 14.3 GB/s useful, 128 loads per hash (every pack within 1 percent of this rate)
CPU emulator (emu/), sub-group 32 and wave64 both 96/96, f2a95d5bb84d961e 96/96, 8e22ad069cb2a8c3 not run none

Version 1 packs (3 October 2026, retired vectors; the kernel text is unchanged, so the exchange and arithmetic findings stand):

Machine, runtime Exchange igneum-genesis-mh closed-form packs Rate at 1 GiB
Apple M5 Max, Apple OpenCL 1.2 (deprecated runtime) local memory (runtime lists no sub-group extension) cache FNV = Mac, 96/96; also 96/96 at --group-warps 2 and 4 96/96 and 96/96 45.03 Mhash/s, 18.73 GB/s useful (Metal on the same chip: 45.2). Apple OpenCL on the M5 Max, not an AMD number
Apple M5 Max, pocl 7.2 CPU device, OpenCL 3.0, LLVM 23 sub_group_shuffle_xor (auto and --exchange subgroup; pocl's clGetKernelSubGroupInfoKHR fails with CL_INVALID_OPERATION, the probe kernel reports 32), and --exchange local cache FNV = Mac, 96/96 on both paths, same batch fingerprint as Apple and the emulator at 2^13 not run CPU, not meaningful
CPU emulator (emu/), 7 configurations incl. sub-group 64 both 96/96 in every configuration, identical batch fingerprint f99fb375b3abeaf5 not run none

Sweep on Apple OpenCL (3 batches, wall time): 4 MiB 573.7, 64 MiB 178.9, 256 MiB 94.3, 512 MiB 68.8, 1024 MiB 45.0 Mhash/s. Full tables in docs/bench-log.md.

Checking a pack without a GPU

emu/emu.sh igneum-genesis-mh 0 32 --sg 32     # local-memory exchange, work-group 32, 32-wide sub-group
emu/emu.sh igneum-genesis-mh 0 64 --sg 64     # local-memory exchange on a wave64 model, two units per work-group
emu/emu.sh igneum-genesis-mh 1 32 --sg 32     # sub_group_shuffle_xor, 32-wide sub-group
emu/emu.sh igneum-genesis-mh 1 64 --sg 64     # sub_group_shuffle_xor over a 64-wide sub-group (two units in one wave)

The emulator compiles kernel.cl unchanged as C++ (emu/emu_opencl.h supplies the OpenCL built-ins; the kernel's own prelude includes it when no OpenCL compiler is present) and runs work-groups on host threads with a barrier inside every exchange. It prints the cache check, dataset self-test, vectors and a fingerprint of the whole 2^13-hash batch, so configurations can be compared for every nonce, not only the vector warps. It says nothing about any GPU.

Regenerating a pack

cd igneum-pow && cargo build --release
./target/release/igneum-pow export --seed igneum-genesis --out ../proto-cuda/packs/igneum-genesis-mh
./target/release/igneum-pow export --closed-form --seed igneum-genesis --out ../proto-cuda/packs/igneum-genesis
./target/release/igneum-pow export --closed-form --seed igneum-hourly --out ../proto-cuda/packs/igneum-hourly

The Rust exporter (since 4 October 2026 the pack source) writes kernel.cl and kernel_bound.cl next to kernel.cu from the same instruction list. The memory-hard core is emitted once (emit_memhard_core) in three dialects (Metal, CUDA C++, OpenCL C), so the constants in kernel.cl are the literals of memhard.h. The Metal cross-check of a pack is proto-metal/igneum-bench --seed <seed> --export-pack <scratch dir> and a diff of its vectors.json (see ../proto-cuda/README.md, "Regenerating a pack").