the project lead watched the 9070 XT at 90% usage with its fans barely turning and the app could not say what it drew: the
draw, temperature and MH per watt line came from nvidia-smi only, and the earlier per-watt figure used the board
rating. proto-opencl/gpu-telemetry.c prints one line per AMD card per sample (bus from SetupAPI by the display
device's name, kind, name, watts, temp_c, fan_rpm, fan_pct, mclk_mhz, gclk_mhz, util_pct, source), built by
build-windows.sh against vendor/adlx (the SDK clone), shipped by make-payload.sh and push-inputs.sh. The engine
runs it with -l 5 beside nvidia-smi (Source::AmdTelemetry, tick_amd_telemetry), parse_amd_telemetry fills
power_w, temp_gpu, fan_pct, fan_rpm, mclk_mhz, util_pct and telemetry_at on the AMD card matched by kind and
ordinal, so eff_mhw and the dashboard's existing line show it; app.js shows fan and memory clock when present.
Tests: three on the parser with lines captured on PC 1 and the Mac fixture; the sysfs path ran on a fixture tree.
Measured over 20:27:45 to 20:29:41 UTC with both cards mining (docs/bench-log.md, under the 9070 XT ceiling table):
9070 XT 198.9 W (193 to 212), 64 C, 657 rpm, 2,505 MHz memory, 3,290 MHz shader, 100% busy, 17.73 MH/s =
0.089 MH/W; RTX 5090 307.6 W, 69 C, 44% fan, 122.30 MH/s = 0.398 MH/W.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Nothing changes for the default class: the pinned packs are byte-identical (tests/packs.rs), the v2 draw stream is untouched.
LoadClass {mix, load_slots, scratch}: fixed widths w16, w64, w64x4 (4 loads of 64 B), era mixes 50/35/15 and 25/50/25 drawn per load with one extra below(100) roll, and the scratch variant scr0/2/4/8 (persistent warps, 1 MiB per warp, tagged lazy fill, measurement only). A wide load reads the W-aligned address and folds every word: x = dst ^ w0; x = (rotl(x, 11) * 0x9e3779b1) ^ w[j]. Program ids carry the class. proto-opencl/host.c taken from opencl-rdna4 23810df (--memprobe, select read-back).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
host.c --serve --pack <dir> serves a pack read at run time (packfile.h), whatever pack the exe was built against: the
first pair is built and self-tested like a prepared one (pairBuffers, pairSelfTest), prepared pairs are self-tested
too, the ready line says path prebuilt-generic. cl_dynamic.h (IGNEUM_CL_DYNAMIC) opens OpenCL.dll with LoadLibrary and
fills a function pointer per entry point, so the mingw build links no import library. test-generic.sh checks the mode
through Apple OpenCL with the packs of the CUDA emulation test: PASS (192 found, prepare and swap, 15 sampled hashes
equal igneum-pow hash-bound), built against a different placeholder pack than it served.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
igneum-pow 0.2.0: generator v2 draws exactly 16 load slots from instructions 1..63, a
load's source from the registers written earlier and not read by a load since, the other
48 ops from the ten non-load weights; accept.rs is spec 01 section 1.4.6 (static: no
stale load source, every register injected; dynamic: 64 units on the seed-keyed
closed-form dataset, no constant bit, no lane-constant site, under 164 saturated, bias
within 136 of 1024, distinct addresses above 245,760); a rejected candidate is replaced
by the next attempt of the seed (seed || k_le32), 32 a consensus fault. Packs carry the
generator version, attempt and program id. Version 1 kept as generate_v1 for the census.
Packs: igneum-genesis, igneum-hourly, igneum-genesis-mh regenerated by igneum-pow export;
new igneum-devnet-v4-epoch0 (devnet genesis hash, day bytes 20730). Checks: Rust 39 of
39 tests; Metal natively via the Swift port (export cross-check 3 of 3 warps, identical
programs and vectors on five seeds incl. three with attempt 1, fuzz 2,000 of 2,000);
CUDA emu 4 of 4 packs; OpenCL emu 2 packs x 2 configurations; Apple OpenCL 4 of 4 packs
at 27.9 Mhash/s. Census 20,000: 5.225 percent rejected, accepted distinct mean 127.887.
Spec 01 0.2 (1.4.2, 1.4.3, 1.4.6, 1.11, 1.15, 1.16, 1.17), igneum-pow README, the CUDA,
OpenCL and Metal test notes, bench-log entry, ledger M5 and M6 Fixed.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Exporter writes kernel.cl next to kernel.cu (same instruction list; memory-hard core emitted in a third, OpenCL C
dialect with the same literals as memhard.h). Pack headers are now C99-safe so a plain C host can include them.
proto-opencl/host.c: C99 + OpenCL 1.2 API, device list, runtime build, cache fill and FNV check, dataset build and
self-test, 3 vector warps standalone and in batch, bench and sweep as host.cu, whole-batch fingerprint. The 32-lane
exchange is sub_group_shuffle_xor only when the queried sub-group size for a 32-item work-group is exactly 32;
otherwise a local-memory exchange with one barrier per exchange, so wave64 hardware cannot change the hash
(WAVEFRONT.md). build.sh (macOS, Linux), build.bat (MSVC), README with the exact AMD-rig commands.
Proven without AMD silicon: Apple OpenCL 1.2 on the M5 Max 96/96 on all three packs (45.0 Mhash/s at 1 GiB, Apple
number, not AMD); pocl 7.2 CPU device 96/96 on both exchange paths including the real sub_group_shuffle_xor text;
CPU emulator 7 configurations incl. 64-wide sub-groups, identical fingerprint f99fb375b3abeaf5 everywhere.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>