igneum/proto-opencl/emu/emu_opencl.h
igneum-labs 660f0eb16c proto-opencl: OpenCL path for AMD, proven on Apple OpenCL, pocl and a wave64 CPU emulator
Exporter writes kernel.cl next to kernel.cu (same instruction list; memory-hard core emitted in a third, OpenCL C
dialect with the same literals as memhard.h). Pack headers are now C99-safe so a plain C host can include them.

proto-opencl/host.c: C99 + OpenCL 1.2 API, device list, runtime build, cache fill and FNV check, dataset build and
self-test, 3 vector warps standalone and in batch, bench and sweep as host.cu, whole-batch fingerprint. The 32-lane
exchange is sub_group_shuffle_xor only when the queried sub-group size for a 32-item work-group is exactly 32;
otherwise a local-memory exchange with one barrier per exchange, so wave64 hardware cannot change the hash
(WAVEFRONT.md). build.sh (macOS, Linux), build.bat (MSVC), README with the exact AMD-rig commands.

Proven without AMD silicon: Apple OpenCL 1.2 on the M5 Max 96/96 on all three packs (45.0 Mhash/s at 1 GiB, Apple
number, not AMD); pocl 7.2 CPU device 96/96 on both exchange paths including the real sub_group_shuffle_xor text;
CPU emulator 7 configurations incl. 64-wide sub-groups, identical fingerprint f99fb375b3abeaf5 everywhere.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-03 16:52:24 +00:00

49 lines
2.4 KiB
C++

// CPU emulation shim: the slice of OpenCL C that the generated kernel.cl uses, as C++.
// kernel.cl includes this file when __OPENCL_VERSION__ is not defined (its own prelude does that), so the exact
// generated text compiles with clang++ or g++ and runs on host threads. emu_main.cpp supplies the runtime below.
// The sub-group width is a runtime setting (32 or 64) so the kernel can be checked as a 64-wide hardware wave
// (AMD GCN/CDNA, RDNA in wave64) would run it: two logical 32-lane units in one wave. Not part of any deliverable
// that runs on a GPU, and it says nothing about AMD hardware; only PASS/FAIL matters, rates are noise.
#pragma once
#include <cstdint>
#include <cstddef>
typedef unsigned int uint;
typedef unsigned long ulong;
static_assert(sizeof(uint) == 4, "uint must be 32 bits");
static_assert(sizeof(ulong) == 8, "ulong must be 64 bits (LP64 host: macOS or Linux)");
// Address-space qualifiers and the kernel attribute mean nothing on the host.
#define __kernel
#define __global
#define __constant
#define __local
#define __private
#define CLK_LOCAL_MEM_FENCE 1
#define CLK_GLOBAL_MEM_FENCE 2
#define IGNEUM_KERNEL_HASH
// Work-group local memory: one arena per work-group instance, shared by its work-items (emu_local_words).
#define IGNEUM_LOCAL_WORDS(name, n) uint* name = emu_local_words((uint)(n))
// Work-item functions (thread-local state set by emu_launch).
size_t get_global_id(uint dim);
size_t get_local_id(uint dim);
size_t get_group_id(uint dim);
size_t get_local_size(uint dim);
size_t get_global_size(uint dim);
uint get_sub_group_size(void);
uint get_sub_group_local_id(void);
uint get_sub_group_id(void);
uint get_num_sub_groups(void);
// Synchronisation and exchange.
void barrier(int flags); // work-group barrier
uint sub_group_shuffle_xor(uint v, uint mask); // cl_khr_subgroup_shuffle semantics over the emulated sub-group
uint intel_sub_group_shuffle_xor(uint v, uint mask); // same semantics
uint sub_group_broadcast(uint v, uint lane);
uint* emu_local_words(uint n);
// Integer built-ins with OpenCL semantics.
static inline uint mul_hi(uint a, uint b) { return (uint)(((uint64_t)a * (uint64_t)b) >> 32); }
// rotate(v, i): bits shifted left by i modulo the bit width (OpenCL C spec 6.3 and 6.12.3).
static inline uint rotate(uint x, uint n) { n &= 31u; return n == 0u ? x : ((x << n) | (x >> (32u - n))); }