From 5dc4414cc1620e86caf6218bd19df9afded9f81e Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Sat, 3 Oct 2026 17:07:58 +0000 Subject: [PATCH] AMD gfx1036 (9800X3D iGPU): memory-hard pack 96/96 PASS, third vendor bit-exact Co-Authored-By: Claude Fable 5.1 --- docs/bench-log.md | 14 ++++++++++++++ 1 file changed, 14 insertions(+) diff --git a/docs/bench-log.md b/docs/bench-log.md index 37d3e2b69..fbeacae84 100644 --- a/docs/bench-log.md +++ b/docs/bench-log.md @@ -183,3 +183,17 @@ pocl 7.2 CPU device (OpenCL 3.0, LLVM 23, Khronos ICD loader, `brew install pocl CPU emulator (`proto-opencl/emu`, kernel.cl compiled as C++, 1 GiB dataset built on 256 host threads): 7 configurations all PASS with the identical batch fingerprint f99fb375b3abeaf5 over 2^13 outputs: exchange 0 with work-group 32/sub-group 32, 64/64, 32/64; exchange 1 (sub-group shuffles) with 32/32, 32/64, 64/64 (wave64 carrying two 32-lane units in one shuffle domain), 64/32. Cross-implementation fingerprint at `--batch-log2 13`, base nonce 0, igneum-genesis-mh: Apple OpenCL f99fb375b3abeaf5, pocl sub-group f99fb375b3abeaf5, pocl local f99fb375b3abeaf5, emulator f99fb375b3abeaf5 (all 7). At 2^24 Apple OpenCL prints 98af644e993239e2 (reference for the AMD run). proto-cuda emulator re-run after the header changes (program.h, vectors.h, memhard.h now C99-safe): PASS. Not demonstrated: any AMD compile or run, any AMD hash rate, the cost of the local-memory exchange on AMD, whether RDNA compiles igneum_hash as wave32 or wave64. Next: run the seven commands in `proto-opencl/README.md` on the AMD rig and paste the logs. + +## 3 October 2026, AMD gfx1036 (Ryzen 7 9800X3D integrated RDNA 2 graphics, 1 compute unit), AMD OpenCL 2.1 driver 3652.0 + +Pack igneum-genesis-mh, memory-hard dataset, 1024 MiB, exchange via local memory (the driver lists no sub-group shuffle extension), wavefront 32. + +| Check | Result | +|---|---| +| 256 MiB cache, device vs host vs Mac | PASS, 17.5 ms device fill | +| Dataset build from the cache, 1 GiB | 392 ms, 42.8 M items/s | +| Dataset self-test, 4 checks | PASS | +| Vectors, 3 warps, standalone and in batch | 96/96 PASS | +| Hash rate | 4.38 Mhash/s on one compute unit, 1.82 GB/s useful | + +Reading: the third GPU vendor. The same memory-hard program now produces identical hashes on Apple Metal, NVIDIA CUDA, Apple OpenCL and AMD OpenCL, cache and dataset included. The AMD number is from a two-CU integrated chip sharing system memory and is a correctness result only; the discrete AMD card is still to come. The local-memory exchange path, which wave64 cards will also use, is now proven on AMD silicon.